GLM-OCR: RAM, speed and settings on a Mac
GLM-OCR is Z.ai's small vision model that reads a page into markdown, with tables and LaTeX formulas. It is a 1.4 GB download in two files. Below is what we measured running it in PocketWebTools for Mac on a 16 GB M1 Pro, the one setting that matters, and where it lost a 14-page comparison.
Which Mac runs it
- Mac memory
- 8 GB and up
- Model file
- 0.95 GB
The text model, 8‑bit.
- Companion file
- 0.48 GB
The projector that turns the page image into tokens the model can read.
- Memory while reading
- 2.4 GB
Peak for the whole process.
- Alongside a chat model
- Works on 16 GB
It ran in 3.6 seconds with Gemma 4 12B still loaded.
GLM-OCR speed on an M1 Pro
Measured in the app's own engine on a 16 GB M1 Pro.
- A one-page report
- 7.0 to 7.9 s
At a 1,600-token image budget. 3.4 s of that is reading the image.
- A screenshot
- 5.5 s
- A 3-page PDF
- 28.0 s
One page with text already in it, one image page and one landscape page.
- Pages that are mostly table
- 14 to 19 s
OvisOCR2 took 20 to 25 s on the same pages.
- Writing speed
- 150 to 175 tokens/s
GLM-OCR next to OvisOCR2 on 14 pages
Nine pages rendered from HTML so the exact text is known (an invoice, the same invoice rotated, blurred and noisy, prose, small prose, a rough scan, a code screenshot, a merged-cell table, two columns, formulas), one page in nine languages and four pages from arXiv papers.
- Clean English text
- A tie
Both missed 0.00 to 0.22% of the characters on every English page, got every number, kept reading order and wrote formulas as LaTeX.
- A 13-column table
- Columns slipped
GLM-OCR wrote 14 or 15 cells in 19 of 22 rows, so numbers sat under the wrong header. OvisOCR2 wrote 13 cells in every row.
- A page with a second table
- Dropped the table
On one arXiv page it left a whole table out with no sign that anything was missing. OvisOCR2 returned it.
- Headings and captions
- Sometimes dropped
A table title and a running header went missing.
- Arabic
- 44% of characters missed
It invented the sentence. OvisOCR2 missed 2%.
- Hindi
- 10% missed
The one language where it beat OvisOCR2, which missed 39%.
Settings the app runs GLM-OCR with
- Files
- GLM-OCR-Q8_0.gguf + mmproj-GLM-OCR-Q8_0.gguf
The ggml-org GGUF pair, pinned to one revision and checked against its SHA-256 after download.
- Image token budget
- 1,600
The setting that matters. At 1,600 the text was identical to the full budget on our report page; at 1,024 it already dropped characters.
- Context window
- 8,192 tokens
- Engine
- llama.cpp multimodal, on Metal
- Page rendering
- 300 or 200 dpi
PDF pages that already contain text are not sent to the model at all.
What the app does with GLM-OCR
Extract
Drop PDFs, scans or screenshots, or a whole folder, and get plain text, markdown, or a searchable PDF with the text placed on the printed words.
The quicker pick for tables
OvisOCR2 is the app's default document model. GLM-OCR stays as the option for pages that are mostly table, where it finished sooner.
Extract, then ask
Send to Chat hands the extracted text to the chat screen, so you can ask a chat model about the document.
Where it falls short
- Check wide tables. On a 13-column table most rows came out with one or two cells too many, which moves numbers under the wrong header without any warning.
- It can leave out a table, a heading or a caption. If a page had two tables, count them in the result.
- Do not use it for Arabic: it missed 44% of the characters on our page. For other scripts the app has dots.ocr.
- Every number here is from a small set of pages on a 16 GB M1 Pro, not a benchmark suite.
Run GLM-OCR in PocketWebTools for Mac
One app for chat, transcription, documents, voices, photos and video, with every model on your Mac. $99 once, free lifetime updates.
GLM-OCR questions
- How much RAM does GLM-OCR need?
- About 2.4 GB while reading a page, measured as the peak for the whole process on an M1 Pro. The two files are 1.4 GB on disk together. PocketWebTools for Mac files it under 8 GB Macs.
- How fast is GLM-OCR on an M1 Pro?
- About 7 seconds for a page of text at a 1,600-token image budget, of which 3.4 seconds is reading the image. A screenshot took 5.5 seconds and pages that are mostly table took 14 to 19.
- Does GLM-OCR run on llama.cpp?
- Yes, through llama.cpp's multimodal path with two GGUF files: the text model and its projector. Set the image token budget to 1,600; at 1,024 it started dropping characters on our test page.
- GLM-OCR or OvisOCR2?
- OvisOCR2 for most documents. The two tied on clean English text, but on our 14 pages only OvisOCR2 kept a 13-column table in its columns and returned every table. GLM-OCR was quicker on pages that are mostly table, 14 to 19 seconds against 20 to 25, and better on Hindi.
- Which languages does GLM-OCR read?
- The app offers it for Chinese, English, French, Spanish, Russian, German, Japanese and Korean. On our nine-language page it did badly on Arabic (44% of characters missed) and missed 10% on Hindi.
- Can I use GLM-OCR commercially?
- Yes. It is released under the MIT license, which allows commercial use. The app ships the license text and adds no restrictions of its own.