dsh-ab-ocr
Verifieddsh-ab-ocr Β· v0.1.4 Β· MIT
--- description: "The out-of-tree ocr tool: recognizes a scanned PDF or image page by page, merges the pages in order only after the last one, and saves one Markdown document with the heading levels the model decided." kind: "package-reference" ---
Install
dsh plugin add dsh-ab-ocr Confirm the layer applied with dsh --profile default --dump-config β see the install guide.
Source
Published to npm without a public repository. Inspect the package contents before installing.
Tags
Readme
description: "The out-of-tree ocr tool: recognizes a scanned PDF or image page by page, merges the pages in order only after the last one, and saves one Markdown document with the heading levels the model decided." kind: "package-reference"
@deepseek-ai/dsh-ab-ocr
Summary
dsh-ab-ocr gives this machine's dsh profiles one way to read a document whose text cannot be selected. One ocr call renders, recognizes, and merges every page of a PDF or a still image, and saves the result as one Markdown document: page numbers and repeated edge lines removed, the whitespace a line box leaves behind removed, and paragraphs rejoined across page breaks.
The outline is a two-step shape. The merge infers a heading level for every line it can, and writes the lines a reader might call headings to a small outline file beside the document. The conversation's model reads that file, decides the levels, and calls ocr again with a levels value; the second call rewrites the document from the geometry it stored, without recognizing anything again. With no second call the document is still complete β it simply carries the inferred outline.
The OCR engine is RapidOCR, whose PP-OCRv6 detection, classification, and recognition weights ship inside the wheel, so installing the dependencies is installing the model. The environment is created from inside the package by pnpm run setup.
Table of Contents
- Use this package
- Install the engine
- Configuration
- What one call records
- Where the files go
- Understand the implementation
- Model Experience
- Evidence
./invariant- Known Limitations and Deferred Work
- Dev Note
Use this package
Mount the row in a profile when the agent should read a scan, a photographed page, a screenshot, or a PDF whose text cannot be selected. The row injects ctx.tools and ctx.systemPrompt: the tool reaches the model catalog through the ordinary schema assembly, and the prompt through the ordinary section assembly. Mounting it adds no provider β every file it writes goes through the filesystem capability the composition already has.
Mount it in a profile
# <DSH_HOME>/profiles/<profile>/cordis.patch.yml
# The package is out-of-tree: its own bundle patch (./cordis.patch.yml) already
# inserts the `ocr` row, so the profile only supplies the config. Do NOT re-insert
# the row here β a second `- insert: id: ocr` would register it twice and the
# harness would warn "tool ocr is already registered" at startup.
- id: ocr
config:
dpi: 200
maxPixels: 12000000
maxPages: 2000
maxDocuments: 20
timeoutMs: 1800000
startupTimeoutMs: 60000
maxOutlineCandidates: 400
engineLifetime: perDocument
writePageFiles: true
outlineCandidateRatio: 1.05
// <DSH_HOME>/profiles/<profile>/package.json
"dsh-ab-ocr": "link:<plugins root>/plugins/dsh-ab-ocr"
Every bound that block states is required: the row is where this deployment says how much work, memory, time, and disk it is willing to carry, and a row that omits one fails to load naming the field. The recognition heuristics have documented defaults and appear here only when this deployment wants a different value.
Run pnpm install in the profile directory after editing either file.
Install the engine
The Python environment is separate from pnpm and is created by a script inside this package:
pnpm --dir <plugins root>/plugins/dsh-ab-ocr run setup
It finds a base interpreter, creates python/.venv, installs python/requirements.txt, runs the worker's own --self-test, and reports the interpreter, the package versions, and the model files the wheel carries. It is idempotent: a run against a working environment verifies and exits without touching the network.
| Flag | Effect |
|---|---|
--python <path> |
Base interpreter to build the environment from (or DSH_OCR_PYTHON) |
--prefetch |
Also load the engine once, so the first recognition is not the one that waits for the models |
--force |
Remove the environment and build it again |
--json |
Print the summary as one JSON object |
Installing into the environment needs permission to create directories with a private mode, so run it outside a restricted shell. The tool reports what is missing β and names this script β when a call arrives without an environment.
Configuration
Every field is overridable from the row. A field the deployment owns is required β it has no value inside the package, because a hidden default would state an agreement the deployment never made β and a field whose default is a documented resolution rule or a recognition heuristic is defaulted.
| Field | Required | Default | Meaning |
|---|---|---|---|
pythonPath |
(auto) | Interpreter carrying the OCR dependencies | |
workerScript |
(auto) | Worker script to run | |
outputDir |
(beside source) | Where the Markdown files go | |
pageDirName |
.ocr-pages |
Per-document directory, relative to each output file | |
writePageFiles |
yes | β | Write each page's text to its own file as it finishes |
dpi |
yes | β | PDF render resolution |
maxPixels |
yes | β | Ceiling on one rendered page; 0 disables it |
textScore |
0.5 | Recognition confidence floor | |
detectHeadings |
true |
Render the inferred outline into the document | |
headingMinRatio |
1.18 | Glyph-height ratio that makes a line a heading | |
indentRatio |
1.0 | Left-edge offset, in body glyph heights, that marks an indented first line | |
paragraphGapRatio |
0.85 | Vertical gap, in body glyph heights, that separates two paragraphs | |
outlineCandidateRatio |
1.05 | Glyph-height ratio above which a line becomes an outline candidate | |
maxOutlineCandidates |
yes | β | Most candidates one document's outline file may list |
detectColumns |
true |
Split a page at a vertical gutter | |
removePageNumbers |
true |
Drop folios | |
removeRunningHeads |
true |
Drop lines repeated in the edge bands | |
runningHeadRatio |
0.6 | Fraction of pages such a line must appear on | |
runningHeadMinPages |
3 | Absolute page minimum for that rule | |
engineLifetime |
yes | β | Rebuild the engine per document, or share it across the call |
timeoutMs |
yes | β | Ceiling on one recognition; 0 disables it |
startupTimeoutMs |
yes | β | Ceiling on the worker's startup check before the call is refused |
maxPages |
yes | β | Pages a document may carry without an explicit selection; 0 disables it |
maxDocuments |
yes | β | Files one call may name |
The required group is what the tool may cost this machine: the render resolution, the per-page pixel ceiling, the page and file ceilings, the two time ceilings, the outline candidate ceiling, the engine lifetime, and whether per-page text is kept. The defaulted group is what a document reader wants: where the interpreter, the worker, and the artifacts are found, and how a page is read and merged.
startupTimeoutMs exists because the startup check is a separate process from the batch. An environment whose engine never loads has to be refused inside the deployment's startup window rather than held for the whole recognition ceiling.
What one call records
// ocr({ path: "D:/docs/ι¨η½²θ§θ.pdf" })
{
"documents": [{
"input": "D:/docs/ι¨η½²θ§θ.pdf",
"output": "D:/docs/ι¨η½²θ§θ.md",
"pages": 3, "totalPages": 3, "selection": [1, 2, 3],
"lines": 16, "headings": 3, "chars": 289,
"droppedPageNumbers": 1, "droppedRunningHeads": 0, "joinedAcrossPages": 0,
"pageDir": "D:/docs/.ocr-pages/ι¨η½²θ§θ",
"digest": "9c41acae7ce7",
"outlinePath": "D:/docs/.ocr-pages/ι¨η½²θ§θ/9c41acae7ce7/outline.json",
"outlineTruncated": false,
"corrections": 0,
"seconds": 31.56
}],
"failures": []
}
A call carrying levels reports corrections instead of a recognition, and rewrites the Markdown the earlier recognition produced. A document that fails does not discard the others; each failure carries its own path and reason.
Where the files go
For a source named <stem>.<ext> under a base directory <base> β the source's own directory, or outputDir:
| Artifact | Path |
|---|---|
| Markdown | <base>/<stem>.md |
| A second recognition of the same source | <base>/<stem>.<digest12>.md |
| Per-page text | <base>/.ocr-pages/<stem>/<digest12>/page-00001.txt |
| Recognition record | <base>/.ocr-pages/<stem>/<digest12>/record.json |
| Outline sheet | <base>/.ocr-pages/<stem>/<digest12>/outline.json |
| Most recent recognition | <base>/.ocr-pages/<stem>/latest.json |
The source's extension is replaced by .md and every space is removed, so the name is space-free on disk and in a link. Tabs, non-breaking spaces, and the full-width space are removed with it; characters a file system rejects are replaced, a reserved Windows device name gains a leading underscore, and a name that compacts to nothing becomes document.
digest12 is the first twelve hex characters of a hash over the source path and content, the page selection, and the merge settings. The worker reports the source's own content digest as the document starts, so the name follows what the document is rather than how large it happens to be: a source edited without changing size is a different recognition, and an unchanged one is always the same recognition. That is how a levels call finds what to correct without scanning.
<stem>.md is never overwritten with different content: when it already holds another document, the new one takes the digest name. Repeating a recognition whose result is unchanged reuses the file it wrote. writePageFiles: false suppresses only the per-page text; the record and the outline are always written, because the correction pass reads them.
Understand the implementation
Design commitments
- One page in memory. A page is rendered, recognized, reported, and dropped before the next page is rendered.
- A page is saved as it finishes. The text is written when the page event arrives, so an interrupted run keeps every page it completed.
- The document is merged only after the last page. The merge is a separate pass in the worker, over records it kept; nothing is written to a document while a page is still being recognized.
- The engine is released per document.
engineLifetime: perDocumentrebuilds it for each document, so a batch never holds more than one. - Nothing reads a PDF's text layer. Every page is rendered and recognized, which is what makes a scan work and what costs time on a text PDF.
- Every write goes through ctx.fs. See Why every write goes through ctx.fs.
Why every write goes through ctx.fs
This deployment enforces its file policy in two different places, and only one of them reaches a plugin that opens files itself.
| Layer | What it confines |
|---|---|
| The process sandbox | A command started through the sandbox provider, and everything that command spawns |
ctx.fs |
Every mutation a plugin performs through the capability, at the path level |
The base bundle mounts the sandbox-enforcing filesystem provider, so ctx.fs.writeText fences each write against the session mode and the writable roots. A plugin that calls the host filesystem directly is fenced by neither. The worker therefore opens only the document it was given and writes to stdout; the page text, the geometry, the merged Markdown, and the outline all travel back as events, and this module writes them through ctx.fs.
The capability is read with ctx.get('fs') when a call runs and is not listed in the row's inject. Cordis satisfies a declared injection by waiting for the service, so declaring fs would make the row a provider requirement and leave it unmounted in a composition that has none.
Why every write carries a fence
A sandboxing backend fences a mutation by a per-call policy, and that policy β not the backend's own default β is the only thing naming the workspace the calling session runs in. Calling writeText without one makes the backend fall back to the deployment's root, which is the directory the server was launched from; a session working anywhere else then has every artifact refused, one page at a time. src/sandbox.ts owns that seam:
callFence(ctx, request)resolves the calling session's policy once per call and refuses a composition that mounts a confining backend with noctx.sandboxPolicy.saveText(fs, target, content, fence, signal)is the one path every persisted text goes through, so the fence is applied in one place. Every other write in the plugin (writeJson,writeDocument, the page files, the correction pass) delegates to it.- The resolved fence's
workspaceRootis also the base a relativeoutputDirresolves against, so a call that names no directory writes inside the only root the fence will accept.process.cwd()is not consulted at all when a policy is available.
The guard has to live at the call rather than at load: dsh-fs-sandbox declares sandboxPolicy as a declared injection, so in a composition that mounts the backend without the policy service the backend is never constructed β ctx.get('fs') is undefined while this row applies, and the row mounts cleanly. The first call is the earliest point at which the composition's real shape is knowable.
A refused write is restated rather than passed through: the denial marker, the path, the mode, and the workspace root the artifact has to move under, with the backend's refusal kept as the cause and the FS_SANDBOX_DENIED code re-stamped, so a caller that keys a retry off the code still sees it.
Source map
| Path | Role |
|---|---|
src/index.ts |
plugin entry: identity, routing section, tool registration, pending-call presentation |
src/config.ts |
the deployment's bounds and the recognition heuristics, and their schema |
src/plan.ts |
request planning: interpreter, worker, and artifact resolution; the worker spec |
src/recognize.ts |
one call: startup check, the batch, the correction pass, the artifact writes |
src/events.ts |
the worker's event stream and the per-call state it accumulates |
src/documents.ts |
the artifact files of one recognition, and the report it produces |
src/artifacts.ts |
artifact paths and the recognition digest |
src/records.ts |
the stored record and outline formats, their readers, and the JSON write |
src/sandbox.ts |
the per-call fence, the guard for a composition that cannot resolve one, and the one policy-stamped write path |
src/levels.ts |
the levels argument |
src/filename.ts |
output naming |
src/worker.ts |
process transport: spec handoff, line framing, timeout, cancellation |
src/render.ts |
model-facing result text |
scripts/setup.mjs |
the environment installer |
python/ocr_worker.py |
worker entry: engine lifetime, page loop, merge and assemble modes |
python/source.py |
PDF and image page sources |
python/layout.py |
boxes to lines, column detection, outline levels |
python/clean.py |
folios, running heads, line joining |
python/assemble.py |
the merge pass, the candidate list, and the level overrides |
The per-page pass
For each page the worker renders the page, recognizes it, and turns the engine's boxes into reading-order lines. It reports the page's text and its geometry on a page event; this module writes the text to page-NNNNN.txt and accumulates the geometry. The page's pixels are dropped before the next render. A page split by a vertical gutter is read column by column; every other page is read top to bottom, with boxes that share a row joined left to right.
The page files carry five digits, so their order on disk is their order in the document past the 9999th page.
Page order and identity
The page selection is resolved to the explicit list of page numbers it names β "1-5,8" is 1, 2, 3, 4, 5, 8, and nothing else β and the worker emits pages in that ascending order with the source's own page number as the index. The merge looks each page's lines up by that number rather than by position, so a selection that does not start at page 1 merges like any other.
This module validates what the worker reports: a page number must be a positive integer, strictly greater than the one before it, and within the document's own page count. A violation aborts the call with a message naming the page, rather than producing a document with pages out of order or missing.
The merge pass
When the last page is done the worker merges its records in one pass.
- Outline. Heading levels come from the trailing outline number (
1.1,第 3 η«,δΈγ), from a named section (ζθ¦,References), or from glyph height measured against the document's own body size. - Furniture. Folios and repeated edge lines are removed.
- Whitespace. Blank runs collapse, spaces that exist only because of a line box are removed between CJK characters, and a Latin word split by a hyphen at a line break is rejoined.
- Page breaks. A paragraph cut in half by a page break is rejoined when the page's first kept line continues it.
The outline pass
The merge also collects the lines a reader might call headings, and writes them to the outline file as candidates:
| Field | Meaning |
|---|---|
id |
h1, h2, β¦ in document order |
page |
The page the line came from |
text |
The line as the document renders it |
level |
The level the document used: the correction when one was applied, else the inferred level, or null for body text |
inferred |
The level the merge inferred, before any correction |
ratio |
The line's glyph height over the document's body height |
numbered |
The section number the line opens with, or null |
A line becomes a candidate when the merge would call it a heading, when it opens with a section number, when it is taller than outlineCandidateRatio of the body height, or when it names a known section. Those last two are what put body lines in the list as well: a line the merge left alone is exactly the kind of line a reader can see is not a heading, and the model needs to be able to say so.
The model answers with levels, a compact list of h<id>=<level>. 1 to 6 makes that line a heading at that depth; 0 leaves it as body text, which is how a line the recognition promoted is demoted again. A second ocr call carrying levels reads the stored geometry, re-runs the same merge with those levels applied, and rewrites the Markdown. It does not render or recognize anything, so it costs the merge rather than the document β three pages answered in about a second where the recognition took thirty.
The corrections become part of the record, and the outline keeps both the applied level and the inferred one. Recognizing the same source again therefore reproduces the corrected document at the path it already owns rather than writing an uncorrected one beside it, and a caller re-reading the outline can see that its correction took.
Export shape
The entry module named-exports name, inject, Config, and apply and carries no default export. Config applies the documented defaults and requires every bound the deployment owns. The entry re-exports the pure core its suites drive β the request planning, the artifact digest and naming, the correction parser, the finished-document reader, the result renderer, and the line framing β and the orchestration and artifact-writing modules behind recognize, so the model-facing surface and the testable core are both reachable from the one . export.
Model Experience
System-prompt section
What the model sees
tool:ocr at order 2975, after the built-in tool band. It says when to reach for the tool, that the call saves a document and an outline file, that the heading levels can be corrected by calling again with levels, and that the document is what to pass to present. The text is empty wherever ocr is not visible in that scope.
Token effect
Three sentences, constant β 438 characters measured.
KV Cache effect
None; the section is a function of tool visibility, not of the conversation.
Tool schema
What the model sees
Five parameters: path, paths, pages, outputDir, and levels. path and paths are alternatives.
Token effect
Roughly constant β 1275 characters of JSON Schema measured on a booted row, of which the parameters are 953 and the description 277. The plugin's whole fixed per-request cost is 1713 characters against 1561 before the outline pass existed: the correction step and the present hand-off cost about ten percent more per request.
KV Cache effect
None. The schema does not vary with the conversation.
Tool-call history and result
What the model sees
The call as made, then one block per document (source path, saved path, page counts, line and heading counts, character count, what was removed, and the outline file), followed by one block per failure. A correction reports what it corrected instead.
Token effect
Proportional to the number of documents and failures, not to document size: the file path is reported instead of its contents, and the outline is a file rather than a list in the result. A three-page document costs 341 characters of result. The recognized text enters the conversation only if the model subsequently reads the Markdown it was told about.
KV Cache effect
None on its own. The result enters the transcript on the following turn like any other tool result.
Evidence
| File | Proves |
|---|---|
tests/index.test.mjs |
Request planning, page validation, the artifact layout, the correction pass, every rejection, the deployment bounds that have no default, and the heuristics that keep one |
tests/transport.test.mjs |
A real child process: page events written to the right files, an out-of-order page aborting the call, a silent worker, the time ceiling, and one failed document leaving the rest of the batch intact |
tests/filename.test.mjs |
The stem rules, the digest name, and the five-digit page file |
tests/render.test.mjs |
The model-facing result of a recognition, a correction, a failure, and a truncated outline |
tests/presentation.test.mjs |
The presentation metadata, and that no result text carries deployment vocabulary |
tests/load-path.test.mjs |
The module form through the real Loader: no default export, inject kept |
tests/loader-composition.test.mjs |
A real cordis.yml boots, the row's bounds are live, an omitted bound fails the load, the section reaches the assembled prompt, and a call dispatched through the registered tool drives the worker and writes every artifact through the mounted capability |
tests/hmr-safety.test.mjs |
The registrations leave with their contributing fiber |
tests/sandbox-write.test.mjs |
The artifact writes under a real workspace-write backend: the fence is required, lands the document inside the session workspace, refuses a path outside it, restates the refusal, and leaves a non-denial alone |
tests/conventions.test.mjs |
The source rules the skill scaffolds a package with |
tests/deployment.mjs |
The bounds a suite configures the tool with, so no suite inherits one from the package |
python/tests/ |
The pure layout, cleaning, merge, page-numbering, candidate, and override rules β 89 cases, no third-party import |
./invariant
The package publishes none. An invariant companion is warranted only when independent observations of one owned relation can diverge; here the merge is a pure function of the page records one process produced, and the artifact layout is derived from a digest of its own inputs. The worker's and the plugin's suites cover those relations instead.
Known Limitations and Deferred Work
- The row must state every bound the deployment owns. There is no fallback: a profile that mounts this tool without
dpi, the pixel, page, file, outline, and time ceilings, the engine lifetime, and the per-page policy fails to load. That is deliberate β a hidden default would decide how much memory and time this machine spends without anyone having agreed to it β but it does mean a copied patch row from an earlier release needs those values added. - A read-only session saves nothing. Every write goes through the sandboxed filesystem capability and is stamped with the calling session's per-call policy, so the call fails with the standard denial marker when the session forbids writing, and a write whose path falls outside the session workspace is refused by name.
- A composition that mounts a confining filesystem without
ctx.sandboxPolicyis refused at the first call. The tool borrows both capabilities rather than declaring them, so the row itself still mounts; the misconfiguration surfaces as one readable error on the first call instead of as artifacts refused one page at a time. - The outline pass is one model turn. A document whose structure the model gets wrong needs a second
levelscall. A correction replaces the levels it names and leaves the rest as the recognition inferred them, and the worker reports the inferred level alongside the applied one, so a correction can be revised rather than only repeated. - The candidate list is capped. Past
maxOutlineCandidatesthe remaining candidates are dropped and the result says the outline is incomplete. - Nothing here reads a PDF's embedded text layer. Every page is rendered and recognized, which is what makes a scan work and what costs time on a text PDF.
- Recognition quality is the engine's. Small type, dense tables, handwriting, and mathematical notation are recognized poorly or not at all.
- Tables are not reconstructed. Cells on one row become one line, so a table is emitted as prose.
- Column detection finds one gutter. A page with three columns, or with full-width headings between columns, is read as two.
- A running head needs repetition to be recognized. A document shorter than
runningHeadMinPages, or a page selection covering fewer pages, keeps its running head as body text. - The tool needs a local filesystem. Its renderer opens the source by path, so a remote filesystem backend cannot serve it.
- The engine's first use loads its models, so expect a few seconds before the first page is recognized;
--prefetchmoves that cost to installation. - The environment is large β about 260 MB, of which 30 MB is the model weights β and lives inside the package directory at
python/.venv.
Dev Note
Built with the plugin-development skill's scaffolder, node <checkout>/.dsh/skills/dsh-plugin-development/scripts/src/index.mjs ocr --tool --plugins-root ., which wrote the package skeleton, the profile patch row, and the profile link dependency in one run.
pnpm build # tsc -p tsconfig.json && tsdown
pnpm test # node --test over the built output
pnpm run setup # create python/.venv and verify it
pnpm test:python # the pure layout, cleaning, and merge rules
pnpm test runs against lib/, so build first. The Python suite has no third-party dependency and runs under any Python 3; it does not need the environment pnpm run setup creates.
The two halves are specified by one wire protocol, documented in python/README.md. The host validates what the worker reports rather than trusting it, and every field the host reads is emitted by a test fixture as well as by the real worker.