Testing
How we check that features work, tools stay fast, private data stays private and people can complete their tasks.
We test all our packages, especially the client command-line tools (CLIs) people and AI agents use. Here you can see what we check before an update reaches you: whether features work, tasks stay fast and private data is handled safely. Each package needs checks that match its job, whether it is a shared library, a client tool or the documentation site.
What we test
| Area | What we check |
|---|---|
| Functionality | Features do what they promise, return the right results and explain failures clearly. |
| Compatibility | Packages work together, updates preserve existing data, and client tools install and run on supported systems. |
| Search quality | Search finds useful sources and puts the best results near the top. |
| Performance | Everyday tasks finish promptly, including with large archives, without excessive memory use or downloads. |
| Security | Credentials stay out of logs and repositories, data shared with agents stays within the request, and actions respect permissions. |
| User experience (UX) | People can find a feature, understand what happened and complete their task. |
We combine automatic checks with manual reviews. Passing tests gives evidence for the cases checked; it does not guarantee every possible situation will work.
Functionality and released packages
We check individual parts of a package, then check how they work together. For client CLIs, we also install and run the package a user would receive. This catches problems that can disappear in a developer's setup, such as a missing file or an option that never reaches the shared library.
For example, a search restricted to one chat must not return messages from another chat. An answer found in a reply must still obey the requested author and date. Exact matching and searches for related evidence are checked separately; the search architecture explains how they work.
Tests for a package stay alongside its code. Shared testing tools and experiments live in cli-testing. Changes to a shared library also need checks in the client tools that use it. The published-package search check, for example, runs the actual Telegram and MAX tools against fictional messages with network access blocked.
Routine checks run automatically during development. Longer performance comparisons and checks using real accounts run separately. Everyday tests use fictional data and temporary storage; real-account checks use approved test accounts, and private results are not published.
Search quality and ranking
Finding a message about a topic is different from finding the answer. We test questions with known answers, questions with no answer in the archive, and replies that contain the answer without repeating the question's words.
We also check the order. A useful answer in first place saves more work than the same answer in tenth place.
| Measure | What it tells us |
|---|---|
| Success@10 | How often at least one useful result appears in the first ten. |
| Recall@10 | How much of the known relevant material appears in the first ten. |
| Precision@10 | What share of the first ten results is relevant. |
| MRR | How close to the top the first relevant result appears. |
| nDCG@10 | Whether the most useful results appear first when some results are better than others. |
For example, 43.8% Success@10 means that 43.8% of the tested questions had at least one relevant result in the first ten. It does not mean every relevant message was found or that the answer was first.
We check whether useful messages were found before sorting them, too: changing their order cannot recover a missing answer. Once a question helps us improve search, it becomes a development example. Fresh questions are needed to judge whether the improvement works elsewhere. Each report says which measures it includes. The search evaluation guide explains the calculations and experiments.
Performance and resource use
We measure how long a whole task takes, how much memory it uses and how that changes as the data grows. For search, that includes finding messages, sorting results, adding useful replies and preparing the output.
We look at typical waiting time and slower runs. Reports may call these p50 (half the measured runs finish within this time) and p95 (95% finish within it). We also compare the first run with later runs, and measure preparing an archive separately from searching it.
If a feature uses a downloadable model, we measure download size and time, loading time and memory use. A small download does not necessarily mean low memory use. Comparisons record the computer and workload so a result from one machine is not presented as a promise for every device. The search experiments contain the detailed measurements.
Security and privacy
We check what could expose your data or let a tool do something you did not request:
- Credentials in logs. Passwords, access tokens and login codes should not appear in logs, error reports or results sent to an AI agent.
- Secrets in repositories. We check files and change history for accidentally committed credentials or private data before publishing code.
- Data sent to agents. Results should contain the data requested for the task and its selected context. Asking about one chat should not expose unrelated chats, other accounts or login details.
- Actions and permissions. Reading a message should not send one. We check that an agent cannot make a forbidden change, and that instructions hidden in a message do not grant extra permissions.
- Unexpected input and files. We review how tools handle invalid requests, damaged or very large files, and file names that could save a download outside the chosen folder. These cases should fail safely and respect limits.
We use fictional data and test credentials, inspect outputs and logs, scan repositories, and review dependencies for known security problems. Some issues need a person to trace what the code does; an automatic scan alone cannot prove that nothing can leak. The security method has the detailed procedure.
For the controls you can set yourself, read security and permissions. An agent's access through other tools needs its own controls.
User experience and documentation
We check real tasks: can someone find the right command, understand an error and tell whether the task succeeded? For the documentation site, we try opening guides, following links, using search and copying instructions.
We check keyboard use, mobile screens, accessibility and every supported language. Automatic checks find broken links, missing translations and some accessibility problems. Manual reviews check that instructions are readable, examples make sense and the page gives a useful next step.
During development we check code for mistakes, run feature tests, and check documentation and a short set of browser tasks. Before publication, the automated build runs broader checks on the complete site, its links and browser behavior. Exact contributor commands live in the repositories' developer guides, including cli-testing's testing guide.
To explore a particular check, follow the linked testing tools or experiments. Their reports should show what was tested, what passed and what still needs attention.