Attachments: sending, downloading and extracting content
Send, download and extract text from MAX attachments; compare local readers, agent tools and explicitly enabled OCR APIs.
Choose how to send or download an attachment, learn which files the CLI reads itself, which dependencies you need and when to ask an agent or external model to read a file. After extraction, you can save the text for search: content: finds the original message.
Sending delivers a file to a chat, downloading saves its bytes, and extraction obtains its content for reading and search. These are separate capabilities: you can send and download an XLSX file, although the built-in extractor cannot yet read its cells.
What you can send
The table shows how the CLI chooses an attachment type. Whether a particular file is accepted and how it is processed also depends on MAX. Sending requires an explicit command; extracting text sends nothing to the chat.
--voice is a separate path for Ogg Opus voice messages, not arbitrary audio. A personal account sends a voice message separately, without text or other attachments. Images are identified by extension, including with --file; here --as-file changes how a video is sent but does not turn such an image into a document. See sending and bots.
What you can download
This section applies to personal accounts. Downloading saves bytes; it does not extract text or convert a document to another format.
max messages download "Учебная группа" 204 --output-dir ./files --json
max attachments list --chat "Учебная группа" --needs-text --jsonThe saved localPath is available on the machine running the CLI. A path alone does not transfer a file to a remote agent: the agent needs access to that file or a separate transfer of its bytes.
How content is read
By default, the agent reads scans and photos with its own tools. attachments extract --ocr explicitly enables an API for bulk recognition. Without this flag, extraction does not call a model. Check whether your installed version supports the flag with max attachments extract --help.
| Format | Programmatically, locally | Explicit API: extract --ocr | When an agent is needed |
|---|---|---|---|
| TXT, MD, MARKDOWN, CSV, TSV, JSON, LOG | Reads UTF-8 text without extra packages | Still reads locally | The agent detects the encoding and converts to UTF-8; automatic detection is not built in yet |
Other files with MIME type text/* or application/json | Reads UTF-8 text | Still reads locally | If the MIME type is absent and the extension is unsupported, an external reader is needed |
| DOCX | Extracts text with mammoth | Still reads locally | This extraction does not recognize embedded images or preserve the exact layout |
| PDF with a text layer | Extracts text with unpdf | Reads text pages locally | Check column order, tables and extracted text accuracy |
| PDF with scanned or mixed pages | Without text: needs-agent; ordinary extraction reads the existing text layer in a mixed PDF | Reads page text; converts pages without text to images using unpdf and @napi-rs/canvas, then uses a model to recognize them | The usual route for scans; also when an engine is unavailable, limits apply or an API fails |
| JPG, JPEG, PNG, WEBP | needs-agent | Sends a supported image to a vision model | By default, the agent reads it itself |
| GIF, HEIC, TIF, TIFF, BMP | needs-agent, without built-in conversion | Automatic OCR of these formats is unsupported | The agent needs a suitable viewer or conversion to PNG/JPEG/WEBP |
| DOC, PPT, XLS | No built-in reader for older binary formats | Does not add a reader for these formats | Convert with an installed office application, then read the text or pages |
| ODT, ODS, XLSX, PPTX | No dedicated built-in extractor | Does not add a reader for these formats | Read with a dedicated library or export the document, sheets or slides using an available tool |
| RTF | No dedicated built-in extractor | Does not add an RTF reader | Convert with a program that understands RTF commands and encoding |
| EPUB | No dedicated built-in extractor | Does not add an EPUB reader | Read chapters in book order using an EPUB tool |
| ZIP | Does not traverse the archive contents | Does not recognize archive contents | List the files, unpack those needed and process each according to its format |
| Voice message | Separate local speech model: messages transcribe | This OCR does not recognize speech | Configure the model and language; see voice transcription |
| Other audio, video, animation and sticker files | Not read by the attachment text extractor | Not recognized by this OCR | Arbitrary audio needs an available speech tool and a suitable format; video needs audio or individual frames |
CSV and JSON become searchable text here, not structured database tables. HTML/XML with a text MIME type is read as source text, not as a browser page. PDF/DOCX extraction saves text, not the original layout. OCR can make mistakes in numbers, reading order and formatting; verify important information against the original.
Voice messages are processed separately from documents: the speech model is downloaded once and then runs locally. Commands, language selection and limits are covered in voice transcription.
Required dependencies
| Task | Package |
|---|---|
| Read a PDF text layer | unpdf |
| Read DOCX text | mammoth |
| Convert PDF pages to images for API OCR | unpdf with rendering support and @napi-rs/canvas |
| Read a supported image through an API | PDF/Word packages are unnecessary; a configured vision-compatible API is required |
| An agent reads a file with its own tools and saves the text | These CLI packages are optional; the agent needs its own way to open the file |
Packages are optional and are not installed automatically with the CLI. engine-missing means the required package is absent or cannot load. It is not an AI model refusal. unpdf reads PDFs and their text layers but does not itself perform OCR on a scan.
Install the package in an environment where the CLI can load it. For a global npm installation using the same prefix:
npm install -g unpdf mammoth @napi-rs/canvasFor a local installation, add the required packages to the same project. With another package manager, a global installation in a separate environment does not guarantee availability: repeat extraction after installation and check that engine-missing disappears. Rendering was verified with unpdf 1.8.1 and @napi-rs/canvas 1.0.10; an older unpdf may read text but lack the necessary rendering functions.
Agent: read and save
max attachments list --chat "Учебная группа" --needs-text --json
# Агент открывает localPath, читает все страницы и сохраняет буквальный текст в scan.txt.
max attachments text set "Учебная группа" 204 --attachment 1 --text-file ./scan.txt --json
max messages search 'content:умножение' --chat "Учебная группа" --offline --json--attachment numbering starts at 1. Preserve the original language and page order; do not replace a transcription with a summary. Do not mark an entire PDF as read after processing just one page. Agent-written text is protected from being overwritten by automatic extraction.
What the agent actually does
An agent cannot automatically open every file. It needs access to localPath, a program to read or convert the format and, for images, a model with vision. It chooses an available method, checks that the result is complete and saves the text through text set.
| Source file | How an agent can obtain text |
|---|---|
| Text in another encoding | Check the encoding marker or detect the encoding; convert with an available tool, then check readability |
| PDF with text | Use an available PDF reader, such as pdftotext; verify column and table order |
| Scanned PDF | Find the page count, convert each page to an image, for example with pdftoppm, and read every image with a vision model |
| DOCX/ODT and presentations | Read with a library or installed office application; export pages for visual verification if needed |
| Spreadsheet | Read each sheet with a library or export sheets to CSV; preserve sheet names and rows, and verify numbers and formulas |
| EPUB | Read the chapter list and each chapter’s HTML in book order; unpacking alone does not guarantee the correct order |
| ZIP | Inspect the contents, select the needed files and use the appropriate reading method for each |
These are examples of possible tools, not programs the CLI installs for an agent. If a required tool is unavailable, the agent must report incomplete processing. A remote agent needs the file itself transferred; a path string does not provide that access.
What affects quality
| Method | What affects the result |
|---|---|
| Programmatic text reading | Correct encoding, complete text layer, format support and paragraph/column/cell order; this does not read text inside an image |
| Agent with available tools | All of the above, plus its vision model quality, image resolution, access to every page, context limits and careful verification |
| API OCR | Vision model selected, scan resolution and quality, language, small print, rotation, tables and handwriting; consistent processing of many files helps repeatability but does not guarantee accuracy |
An API is not necessarily more accurate than an agent: they may use similar models. Its advantages here are a managed queue, parallel processing and reuse of results. An agent can combine precise programmatic text reading with visual checks of difficult sections. For a digital document, obtain its original text first instead of recognizing an image of it. With any OCR, check numbers, names and important tables against the original.
API: explicitly chosen bulk processing
An external model can read text in supported images and scanned PDF pages. The CLI sends it an image of each required page, receives the literal text and saves it in the same content: index. PDFs with a text layer and DOCX files still use programmatic reading; selecting an API does not add support for older Office formats or ZIP.
You need a vision model, its endpoint and an API key. Step-by-step setup for OpenAI, Anthropic and compatible servers, key storage and the models.ocr task configuration are covered in the external models guide. After setup:
max attachments extract --chat "Учебная группа" --ocr --concurrency 4 --limit 20 --jsonRepeated extraction uses the file hash and model target; good saved text is preserved on errors, cancellation or incomplete responses. Agent-written text is not overwritten. Images and scanned pages go to the provider only with explicit --ocr; calls are billed under its terms. --offline --ocr cannot be combined; there is no automatic switch from an agent to an API.
The extraction limit is 50 MiB per file; local text is limited to 2 million characters. API OCR accepts PDFs of up to 20 pages and images of up to 4 MiB and 20 million pixels, with neither side exceeding 8000 pixels. File concurrency is 1–8, default 4; pages within a file run sequentially. By default the API processes up to 100 files; --limit accepts 1–500. Continue with the returned cursor. A provider response of 429 stops further API calls in that run, without retries. The command returns statuses and message links, not full recognized text.
Downloading, extraction and search are covered in more detail in search. This description matches the CLI source code; the existence of a command does not mean every possible file of that format has been tested against live MAX.