Skip to content

dsh-ab-ocr

Verified

dsh-ab-ocr Β· v0.1.4 Β· MIT

--- description: "The out-of-tree ocr tool: recognizes a scanned PDF or image page by page, merges the pages in order only after the last one, and saves one Markdown document with the heading levels the model decided." kind: "package-reference" ---

Install

dsh plugin add dsh-ab-ocr

Confirm the layer applied with dsh --profile default --dump-config β€” see the install guide.

Source

Published to npm without a public repository. Inspect the package contents before installing.

Tags

Readme


description: "The out-of-tree ocr tool: recognizes a scanned PDF or image page by page, merges the pages in order only after the last one, and saves one Markdown document with the heading levels the model decided." kind: "package-reference"

@deepseek-ai/dsh-ab-ocr

Summary

dsh-ab-ocr gives this machine's dsh profiles one way to read a document whose text cannot be selected. One ocr call renders, recognizes, and merges every page of a PDF or a still image, and saves the result as one Markdown document: page numbers and repeated edge lines removed, the whitespace a line box leaves behind removed, and paragraphs rejoined across page breaks.

The outline is a two-step shape. The merge infers a heading level for every line it can, and writes the lines a reader might call headings to a small outline file beside the document. The conversation's model reads that file, decides the levels, and calls ocr again with a levels value; the second call rewrites the document from the geometry it stored, without recognizing anything again. With no second call the document is still complete β€” it simply carries the inferred outline.

The OCR engine is RapidOCR, whose PP-OCRv6 detection, classification, and recognition weights ship inside the wheel, so installing the dependencies is installing the model. The environment is created from inside the package by pnpm run setup.

Table of Contents


Use this package

Mount the row in a profile when the agent should read a scan, a photographed page, a screenshot, or a PDF whose text cannot be selected. The row injects ctx.tools and ctx.systemPrompt: the tool reaches the model catalog through the ordinary schema assembly, and the prompt through the ordinary section assembly. Mounting it adds no provider β€” every file it writes goes through the filesystem capability the composition already has.

Mount it in a profile

# <DSH_HOME>/profiles/<profile>/cordis.patch.yml
# The package is out-of-tree: its own bundle patch (./cordis.patch.yml) already
# inserts the `ocr` row, so the profile only supplies the config. Do NOT re-insert
# the row here β€” a second `- insert: id: ocr` would register it twice and the
# harness would warn "tool ocr is already registered" at startup.
- id: ocr
  config:
    dpi: 200
    maxPixels: 12000000
    maxPages: 2000
    maxDocuments: 20
    timeoutMs: 1800000
    startupTimeoutMs: 60000
    maxOutlineCandidates: 400
    engineLifetime: perDocument
    writePageFiles: true
    outlineCandidateRatio: 1.05
// <DSH_HOME>/profiles/<profile>/package.json
"dsh-ab-ocr": "link:<plugins root>/plugins/dsh-ab-ocr"

Every bound that block states is required: the row is where this deployment says how much work, memory, time, and disk it is willing to carry, and a row that omits one fails to load naming the field. The recognition heuristics have documented defaults and appear here only when this deployment wants a different value.

Run pnpm install in the profile directory after editing either file.

Install the engine

The Python environment is separate from pnpm and is created by a script inside this package:

pnpm --dir <plugins root>/plugins/dsh-ab-ocr run setup

It finds a base interpreter, creates python/.venv, installs python/requirements.txt, runs the worker's own --self-test, and reports the interpreter, the package versions, and the model files the wheel carries. It is idempotent: a run against a working environment verifies and exits without touching the network.

Flag Effect
--python <path> Base interpreter to build the environment from (or DSH_OCR_PYTHON)
--prefetch Also load the engine once, so the first recognition is not the one that waits for the models
--force Remove the environment and build it again
--json Print the summary as one JSON object

Installing into the environment needs permission to create directories with a private mode, so run it outside a restricted shell. The tool reports what is missing β€” and names this script β€” when a call arrives without an environment.

Configuration

Every field is overridable from the row. A field the deployment owns is required β€” it has no value inside the package, because a hidden default would state an agreement the deployment never made β€” and a field whose default is a documented resolution rule or a recognition heuristic is defaulted.

Field Required Default Meaning
pythonPath (auto) Interpreter carrying the OCR dependencies
workerScript (auto) Worker script to run
outputDir (beside source) Where the Markdown files go
pageDirName .ocr-pages Per-document directory, relative to each output file
writePageFiles yes β€” Write each page's text to its own file as it finishes
dpi yes β€” PDF render resolution
maxPixels yes β€” Ceiling on one rendered page; 0 disables it
textScore 0.5 Recognition confidence floor
detectHeadings true Render the inferred outline into the document
headingMinRatio 1.18 Glyph-height ratio that makes a line a heading
indentRatio 1.0 Left-edge offset, in body glyph heights, that marks an indented first line
paragraphGapRatio 0.85 Vertical gap, in body glyph heights, that separates two paragraphs
outlineCandidateRatio 1.05 Glyph-height ratio above which a line becomes an outline candidate
maxOutlineCandidates yes β€” Most candidates one document's outline file may list
detectColumns true Split a page at a vertical gutter
removePageNumbers true Drop folios
removeRunningHeads true Drop lines repeated in the edge bands
runningHeadRatio 0.6 Fraction of pages such a line must appear on
runningHeadMinPages 3 Absolute page minimum for that rule
engineLifetime yes β€” Rebuild the engine per document, or share it across the call
timeoutMs yes β€” Ceiling on one recognition; 0 disables it
startupTimeoutMs yes β€” Ceiling on the worker's startup check before the call is refused
maxPages yes β€” Pages a document may carry without an explicit selection; 0 disables it
maxDocuments yes β€” Files one call may name

The required group is what the tool may cost this machine: the render resolution, the per-page pixel ceiling, the page and file ceilings, the two time ceilings, the outline candidate ceiling, the engine lifetime, and whether per-page text is kept. The defaulted group is what a document reader wants: where the interpreter, the worker, and the artifacts are found, and how a page is read and merged.

startupTimeoutMs exists because the startup check is a separate process from the batch. An environment whose engine never loads has to be refused inside the deployment's startup window rather than held for the whole recognition ceiling.

What one call records

// ocr({ path: "D:/docs/ιƒ¨η½²θ§„θŒƒ.pdf" })
{
  "documents": [{
    "input": "D:/docs/ιƒ¨η½²θ§„θŒƒ.pdf",
    "output": "D:/docs/ιƒ¨η½²θ§„θŒƒ.md",
    "pages": 3, "totalPages": 3, "selection": [1, 2, 3],
    "lines": 16, "headings": 3, "chars": 289,
    "droppedPageNumbers": 1, "droppedRunningHeads": 0, "joinedAcrossPages": 0,
    "pageDir": "D:/docs/.ocr-pages/ιƒ¨η½²θ§„θŒƒ",
    "digest": "9c41acae7ce7",
    "outlinePath": "D:/docs/.ocr-pages/ιƒ¨η½²θ§„θŒƒ/9c41acae7ce7/outline.json",
    "outlineTruncated": false,
    "corrections": 0,
    "seconds": 31.56
  }],
  "failures": []
}

A call carrying levels reports corrections instead of a recognition, and rewrites the Markdown the earlier recognition produced. A document that fails does not discard the others; each failure carries its own path and reason.

Where the files go

For a source named <stem>.<ext> under a base directory <base> β€” the source's own directory, or outputDir:

Artifact Path
Markdown <base>/<stem>.md
A second recognition of the same source <base>/<stem>.<digest12>.md
Per-page text <base>/.ocr-pages/<stem>/<digest12>/page-00001.txt
Recognition record <base>/.ocr-pages/<stem>/<digest12>/record.json
Outline sheet <base>/.ocr-pages/<stem>/<digest12>/outline.json
Most recent recognition <base>/.ocr-pages/<stem>/latest.json

The source's extension is replaced by .md and every space is removed, so the name is space-free on disk and in a link. Tabs, non-breaking spaces, and the full-width space are removed with it; characters a file system rejects are replaced, a reserved Windows device name gains a leading underscore, and a name that compacts to nothing becomes document.

digest12 is the first twelve hex characters of a hash over the source path and content, the page selection, and the merge settings. The worker reports the source's own content digest as the document starts, so the name follows what the document is rather than how large it happens to be: a source edited without changing size is a different recognition, and an unchanged one is always the same recognition. That is how a levels call finds what to correct without scanning.

<stem>.md is never overwritten with different content: when it already holds another document, the new one takes the digest name. Repeating a recognition whose result is unchanged reuses the file it wrote. writePageFiles: false suppresses only the per-page text; the record and the outline are always written, because the correction pass reads them.


Understand the implementation

Design commitments

  • One page in memory. A page is rendered, recognized, reported, and dropped before the next page is rendered.
  • A page is saved as it finishes. The text is written when the page event arrives, so an interrupted run keeps every page it completed.
  • The document is merged only after the last page. The merge is a separate pass in the worker, over records it kept; nothing is written to a document while a page is still being recognized.
  • The engine is released per document. engineLifetime: perDocument rebuilds it for each document, so a batch never holds more than one.
  • Nothing reads a PDF's text layer. Every page is rendered and recognized, which is what makes a scan work and what costs time on a text PDF.
  • Every write goes through ctx.fs. See Why every write goes through ctx.fs.

Why every write goes through ctx.fs

This deployment enforces its file policy in two different places, and only one of them reaches a plugin that opens files itself.

Layer What it confines
The process sandbox A command started through the sandbox provider, and everything that command spawns
ctx.fs Every mutation a plugin performs through the capability, at the path level

The base bundle mounts the sandbox-enforcing filesystem provider, so ctx.fs.writeText fences each write against the session mode and the writable roots. A plugin that calls the host filesystem directly is fenced by neither. The worker therefore opens only the document it was given and writes to stdout; the page text, the geometry, the merged Markdown, and the outline all travel back as events, and this module writes them through ctx.fs.

The capability is read with ctx.get('fs') when a call runs and is not listed in the row's inject. Cordis satisfies a declared injection by waiting for the service, so declaring fs would make the row a provider requirement and leave it unmounted in a composition that has none.

Why every write carries a fence

A sandboxing backend fences a mutation by a per-call policy, and that policy β€” not the backend's own default β€” is the only thing naming the workspace the calling session runs in. Calling writeText without one makes the backend fall back to the deployment's root, which is the directory the server was launched from; a session working anywhere else then has every artifact refused, one page at a time. src/sandbox.ts owns that seam:

  • callFence(ctx, request) resolves the calling session's policy once per call and refuses a composition that mounts a confining backend with no ctx.sandboxPolicy.
  • saveText(fs, target, content, fence, signal) is the one path every persisted text goes through, so the fence is applied in one place. Every other write in the plugin (writeJson, writeDocument, the page files, the correction pass) delegates to it.
  • The resolved fence's workspaceRoot is also the base a relative outputDir resolves against, so a call that names no directory writes inside the only root the fence will accept. process.cwd() is not consulted at all when a policy is available.

The guard has to live at the call rather than at load: dsh-fs-sandbox declares sandboxPolicy as a declared injection, so in a composition that mounts the backend without the policy service the backend is never constructed β€” ctx.get('fs') is undefined while this row applies, and the row mounts cleanly. The first call is the earliest point at which the composition's real shape is knowable.

A refused write is restated rather than passed through: the denial marker, the path, the mode, and the workspace root the artifact has to move under, with the backend's refusal kept as the cause and the FS_SANDBOX_DENIED code re-stamped, so a caller that keys a retry off the code still sees it.

Source map

Path Role
src/index.ts plugin entry: identity, routing section, tool registration, pending-call presentation
src/config.ts the deployment's bounds and the recognition heuristics, and their schema
src/plan.ts request planning: interpreter, worker, and artifact resolution; the worker spec
src/recognize.ts one call: startup check, the batch, the correction pass, the artifact writes
src/events.ts the worker's event stream and the per-call state it accumulates
src/documents.ts the artifact files of one recognition, and the report it produces
src/artifacts.ts artifact paths and the recognition digest
src/records.ts the stored record and outline formats, their readers, and the JSON write
src/sandbox.ts the per-call fence, the guard for a composition that cannot resolve one, and the one policy-stamped write path
src/levels.ts the levels argument
src/filename.ts output naming
src/worker.ts process transport: spec handoff, line framing, timeout, cancellation
src/render.ts model-facing result text
scripts/setup.mjs the environment installer
python/ocr_worker.py worker entry: engine lifetime, page loop, merge and assemble modes
python/source.py PDF and image page sources
python/layout.py boxes to lines, column detection, outline levels
python/clean.py folios, running heads, line joining
python/assemble.py the merge pass, the candidate list, and the level overrides

The per-page pass

For each page the worker renders the page, recognizes it, and turns the engine's boxes into reading-order lines. It reports the page's text and its geometry on a page event; this module writes the text to page-NNNNN.txt and accumulates the geometry. The page's pixels are dropped before the next render. A page split by a vertical gutter is read column by column; every other page is read top to bottom, with boxes that share a row joined left to right.

The page files carry five digits, so their order on disk is their order in the document past the 9999th page.

Page order and identity

The page selection is resolved to the explicit list of page numbers it names β€” "1-5,8" is 1, 2, 3, 4, 5, 8, and nothing else β€” and the worker emits pages in that ascending order with the source's own page number as the index. The merge looks each page's lines up by that number rather than by position, so a selection that does not start at page 1 merges like any other.

This module validates what the worker reports: a page number must be a positive integer, strictly greater than the one before it, and within the document's own page count. A violation aborts the call with a message naming the page, rather than producing a document with pages out of order or missing.

The merge pass

When the last page is done the worker merges its records in one pass.

  • Outline. Heading levels come from the trailing outline number (1.1, 第 3 η« , 三、), from a named section (ζ‘˜θ¦, References), or from glyph height measured against the document's own body size.
  • Furniture. Folios and repeated edge lines are removed.
  • Whitespace. Blank runs collapse, spaces that exist only because of a line box are removed between CJK characters, and a Latin word split by a hyphen at a line break is rejoined.
  • Page breaks. A paragraph cut in half by a page break is rejoined when the page's first kept line continues it.

The outline pass

The merge also collects the lines a reader might call headings, and writes them to the outline file as candidates:

Field Meaning
id h1, h2, … in document order
page The page the line came from
text The line as the document renders it
level The level the document used: the correction when one was applied, else the inferred level, or null for body text
inferred The level the merge inferred, before any correction
ratio The line's glyph height over the document's body height
numbered The section number the line opens with, or null

A line becomes a candidate when the merge would call it a heading, when it opens with a section number, when it is taller than outlineCandidateRatio of the body height, or when it names a known section. Those last two are what put body lines in the list as well: a line the merge left alone is exactly the kind of line a reader can see is not a heading, and the model needs to be able to say so.

The model answers with levels, a compact list of h<id>=<level>. 1 to 6 makes that line a heading at that depth; 0 leaves it as body text, which is how a line the recognition promoted is demoted again. A second ocr call carrying levels reads the stored geometry, re-runs the same merge with those levels applied, and rewrites the Markdown. It does not render or recognize anything, so it costs the merge rather than the document β€” three pages answered in about a second where the recognition took thirty.

The corrections become part of the record, and the outline keeps both the applied level and the inferred one. Recognizing the same source again therefore reproduces the corrected document at the path it already owns rather than writing an uncorrected one beside it, and a caller re-reading the outline can see that its correction took.

Export shape

The entry module named-exports name, inject, Config, and apply and carries no default export. Config applies the documented defaults and requires every bound the deployment owns. The entry re-exports the pure core its suites drive β€” the request planning, the artifact digest and naming, the correction parser, the finished-document reader, the result renderer, and the line framing β€” and the orchestration and artifact-writing modules behind recognize, so the model-facing surface and the testable core are both reachable from the one . export.


Model Experience

System-prompt section

What the model sees

tool:ocr at order 2975, after the built-in tool band. It says when to reach for the tool, that the call saves a document and an outline file, that the heading levels can be corrected by calling again with levels, and that the document is what to pass to present. The text is empty wherever ocr is not visible in that scope.

Token effect

Three sentences, constant β€” 438 characters measured.

KV Cache effect

None; the section is a function of tool visibility, not of the conversation.

Tool schema

What the model sees

Five parameters: path, paths, pages, outputDir, and levels. path and paths are alternatives.

Token effect

Roughly constant β€” 1275 characters of JSON Schema measured on a booted row, of which the parameters are 953 and the description 277. The plugin's whole fixed per-request cost is 1713 characters against 1561 before the outline pass existed: the correction step and the present hand-off cost about ten percent more per request.

KV Cache effect

None. The schema does not vary with the conversation.

Tool-call history and result

What the model sees

The call as made, then one block per document (source path, saved path, page counts, line and heading counts, character count, what was removed, and the outline file), followed by one block per failure. A correction reports what it corrected instead.

Token effect

Proportional to the number of documents and failures, not to document size: the file path is reported instead of its contents, and the outline is a file rather than a list in the result. A three-page document costs 341 characters of result. The recognized text enters the conversation only if the model subsequently reads the Markdown it was told about.

KV Cache effect

None on its own. The result enters the transcript on the following turn like any other tool result.


Evidence

File Proves
tests/index.test.mjs Request planning, page validation, the artifact layout, the correction pass, every rejection, the deployment bounds that have no default, and the heuristics that keep one
tests/transport.test.mjs A real child process: page events written to the right files, an out-of-order page aborting the call, a silent worker, the time ceiling, and one failed document leaving the rest of the batch intact
tests/filename.test.mjs The stem rules, the digest name, and the five-digit page file
tests/render.test.mjs The model-facing result of a recognition, a correction, a failure, and a truncated outline
tests/presentation.test.mjs The presentation metadata, and that no result text carries deployment vocabulary
tests/load-path.test.mjs The module form through the real Loader: no default export, inject kept
tests/loader-composition.test.mjs A real cordis.yml boots, the row's bounds are live, an omitted bound fails the load, the section reaches the assembled prompt, and a call dispatched through the registered tool drives the worker and writes every artifact through the mounted capability
tests/hmr-safety.test.mjs The registrations leave with their contributing fiber
tests/sandbox-write.test.mjs The artifact writes under a real workspace-write backend: the fence is required, lands the document inside the session workspace, refuses a path outside it, restates the refusal, and leaves a non-denial alone
tests/conventions.test.mjs The source rules the skill scaffolds a package with
tests/deployment.mjs The bounds a suite configures the tool with, so no suite inherits one from the package
python/tests/ The pure layout, cleaning, merge, page-numbering, candidate, and override rules β€” 89 cases, no third-party import

./invariant

The package publishes none. An invariant companion is warranted only when independent observations of one owned relation can diverge; here the merge is a pure function of the page records one process produced, and the artifact layout is derived from a digest of its own inputs. The worker's and the plugin's suites cover those relations instead.

Known Limitations and Deferred Work

  • The row must state every bound the deployment owns. There is no fallback: a profile that mounts this tool without dpi, the pixel, page, file, outline, and time ceilings, the engine lifetime, and the per-page policy fails to load. That is deliberate β€” a hidden default would decide how much memory and time this machine spends without anyone having agreed to it β€” but it does mean a copied patch row from an earlier release needs those values added.
  • A read-only session saves nothing. Every write goes through the sandboxed filesystem capability and is stamped with the calling session's per-call policy, so the call fails with the standard denial marker when the session forbids writing, and a write whose path falls outside the session workspace is refused by name.
  • A composition that mounts a confining filesystem without ctx.sandboxPolicy is refused at the first call. The tool borrows both capabilities rather than declaring them, so the row itself still mounts; the misconfiguration surfaces as one readable error on the first call instead of as artifacts refused one page at a time.
  • The outline pass is one model turn. A document whose structure the model gets wrong needs a second levels call. A correction replaces the levels it names and leaves the rest as the recognition inferred them, and the worker reports the inferred level alongside the applied one, so a correction can be revised rather than only repeated.
  • The candidate list is capped. Past maxOutlineCandidates the remaining candidates are dropped and the result says the outline is incomplete.
  • Nothing here reads a PDF's embedded text layer. Every page is rendered and recognized, which is what makes a scan work and what costs time on a text PDF.
  • Recognition quality is the engine's. Small type, dense tables, handwriting, and mathematical notation are recognized poorly or not at all.
  • Tables are not reconstructed. Cells on one row become one line, so a table is emitted as prose.
  • Column detection finds one gutter. A page with three columns, or with full-width headings between columns, is read as two.
  • A running head needs repetition to be recognized. A document shorter than runningHeadMinPages, or a page selection covering fewer pages, keeps its running head as body text.
  • The tool needs a local filesystem. Its renderer opens the source by path, so a remote filesystem backend cannot serve it.
  • The engine's first use loads its models, so expect a few seconds before the first page is recognized; --prefetch moves that cost to installation.
  • The environment is large β€” about 260 MB, of which 30 MB is the model weights β€” and lives inside the package directory at python/.venv.

Dev Note

Built with the plugin-development skill's scaffolder, node <checkout>/.dsh/skills/dsh-plugin-development/scripts/src/index.mjs ocr --tool --plugins-root ., which wrote the package skeleton, the profile patch row, and the profile link dependency in one run.

pnpm build        # tsc -p tsconfig.json && tsdown
pnpm test         # node --test over the built output
pnpm run setup    # create python/.venv and verify it
pnpm test:python  # the pure layout, cleaning, and merge rules

pnpm test runs against lib/, so build first. The Python suite has no third-party dependency and runs under any Python 3; it does not need the environment pnpm run setup creates.

The two halves are specified by one wire protocol, documented in python/README.md. The host validates what the worker reports rather than trusting it, and every field the host reads is emitted by a test fixture as well as by the real worker.