Skip to content
PocketWebToolsPocketWebTools
PocketWebTools for Mac

GLM-OCR: RAM, speed and settings on a Mac

GLM-OCR is Z.ai's small vision model that reads a page into markdown, with tables and LaTeX formulas. It is a 1.4 GB download in two files. Below is what we measured running it in PocketWebTools for Mac on a 16 GB M1 Pro, the one setting that matters, and where it lost a 14-page comparison.

MITMeasured on a 16 GB M1 ProUpdated
1.4GB
Download, two files
2.4GB
Memory while reading
About 7s a page
A page of text
150 to 175tokens/s
Writing speed

Which Mac runs it

Mac memory
8 GB and up
Model file
0.95 GB

The text model, 8‑bit.

Companion file
0.48 GB

The projector that turns the page image into tokens the model can read.

Memory while reading
2.4 GB

Peak for the whole process.

Alongside a chat model
Works on 16 GB

It ran in 3.6 seconds with Gemma 4 12B still loaded.

GLM-OCR speed on an M1 Pro

Measured in the app's own engine on a 16 GB M1 Pro.

A one-page report
7.0 to 7.9 s

At a 1,600-token image budget. 3.4 s of that is reading the image.

A screenshot
5.5 s
A 3-page PDF
28.0 s

One page with text already in it, one image page and one landscape page.

Pages that are mostly table
14 to 19 s

OvisOCR2 took 20 to 25 s on the same pages.

Writing speed
150 to 175 tokens/s

GLM-OCR next to OvisOCR2 on 14 pages

Nine pages rendered from HTML so the exact text is known (an invoice, the same invoice rotated, blurred and noisy, prose, small prose, a rough scan, a code screenshot, a merged-cell table, two columns, formulas), one page in nine languages and four pages from arXiv papers.

Clean English text
A tie

Both missed 0.00 to 0.22% of the characters on every English page, got every number, kept reading order and wrote formulas as LaTeX.

A 13-column table
Columns slipped

GLM-OCR wrote 14 or 15 cells in 19 of 22 rows, so numbers sat under the wrong header. OvisOCR2 wrote 13 cells in every row.

A page with a second table
Dropped the table

On one arXiv page it left a whole table out with no sign that anything was missing. OvisOCR2 returned it.

Headings and captions
Sometimes dropped

A table title and a running header went missing.

Arabic
44% of characters missed

It invented the sentence. OvisOCR2 missed 2%.

Hindi
10% missed

The one language where it beat OvisOCR2, which missed 39%.

Settings the app runs GLM-OCR with

Files
GLM-OCR-Q8_0.gguf + mmproj-GLM-OCR-Q8_0.gguf

The ggml-org GGUF pair, pinned to one revision and checked against its SHA-256 after download.

Image token budget
1,600

The setting that matters. At 1,600 the text was identical to the full budget on our report page; at 1,024 it already dropped characters.

Context window
8,192 tokens
Engine
llama.cpp multimodal, on Metal
Page rendering
300 or 200 dpi

PDF pages that already contain text are not sent to the model at all.

What the app does with GLM-OCR

Extract

Drop PDFs, scans or screenshots, or a whole folder, and get plain text, markdown, or a searchable PDF with the text placed on the printed words.

The quicker pick for tables

OvisOCR2 is the app's default document model. GLM-OCR stays as the option for pages that are mostly table, where it finished sooner.

Extract, then ask

Send to Chat hands the extracted text to the chat screen, so you can ask a chat model about the document.

Where it falls short

  • Check wide tables. On a 13-column table most rows came out with one or two cells too many, which moves numbers under the wrong header without any warning.
  • It can leave out a table, a heading or a caption. If a page had two tables, count them in the result.
  • Do not use it for Arabic: it missed 44% of the characters on our page. For other scripts the app has dots.ocr.
  • Every number here is from a small set of pages on a 16 GB M1 Pro, not a benchmark suite.

Run GLM-OCR in PocketWebTools for Mac

One app for chat, transcription, documents, voices, photos and video, with every model on your Mac. $99 once, free lifetime updates.

See everything in the Mac app

GLM-OCR questions

How much RAM does GLM-OCR need?
About 2.4 GB while reading a page, measured as the peak for the whole process on an M1 Pro. The two files are 1.4 GB on disk together. PocketWebTools for Mac files it under 8 GB Macs.
How fast is GLM-OCR on an M1 Pro?
About 7 seconds for a page of text at a 1,600-token image budget, of which 3.4 seconds is reading the image. A screenshot took 5.5 seconds and pages that are mostly table took 14 to 19.
Does GLM-OCR run on llama.cpp?
Yes, through llama.cpp's multimodal path with two GGUF files: the text model and its projector. Set the image token budget to 1,600; at 1,024 it started dropping characters on our test page.
GLM-OCR or OvisOCR2?
OvisOCR2 for most documents. The two tied on clean English text, but on our 14 pages only OvisOCR2 kept a 13-column table in its columns and returned every table. GLM-OCR was quicker on pages that are mostly table, 14 to 19 seconds against 20 to 25, and better on Hindi.
Which languages does GLM-OCR read?
The app offers it for Chinese, English, French, Spanish, Russian, German, Japanese and Korean. On our nine-language page it did badly on Arabic (44% of characters missed) and missed 10% on Hindi.
Can I use GLM-OCR commercially?
Yes. It is released under the MIT license, which allows commercial use. The app ships the license text and adds no restrictions of its own.