Skip to content

Transcribe (Auto-Detect)

🤖 AI Drafted (Not reviewed)

Use this when: you don’t know the document’s script type and want automatic routing. Classifies the script type (Typescript / Manuscript / HTR / Paleography) then runs the matching transcription profile — including two-pass review for historical and archaic scripts.

Folder /Transcribe
Steps 10
Tags preset, transcribe, auto-detect, classify, routing

Steps, in run order

1. Files

Tool: Files — Pass through input files from workflow context

2. Classify Script Type

Tool: Classify Script Type — Detect whether a document is typescript, manuscript, HTR, or paleography

Settings this step uses:

Option Value
save_to_db yes
vision_mode llm

What this step asks the model:

Examine this document image and classify its script/handwriting type.

Choose EXACTLY ONE type from:
- typescript   : typewritten, printed, or born-digital text (no handwriting)
- manuscript   : modern handwriting, 20th–21st century (cursive or print)
- htr          : historical handwriting, 16th–19th century (legible but archaic letterforms)
- paleography  : archaic script with non-standard letterforms, heavy abbreviation, or
                 specialist scribal conventions (pre-18th century or highly specialised)

Return a JSON object with exactly these fields:
{
  "script_type": "<one of the four types>",
  "confidence": <float 0.0–1.0>,
  "notes": "<one sentence explaining the classification>"
}

Be conservative: if the image is ambiguous between htr and paleography, pick the harder
category (paleography) and lower your confidence. Only mark typescript when there is
clearly no handwriting. Do not add any text outside the JSON object.

3. Transcribe (typescript)

Tool: Transcribe — Extract text from images (OCR)

Settings this step uses:

Option Value
language auto
prompt Transcribe the text visible on this image. The document is typewritten or printed. Rules: - Output ONLY the transcription. No headings, preamble, commentary, or summary. - Preserve original layout, line breaks, and paragraph structure. - Preserve original spelling, capitalisation, and punctuation exactly as printed. - Include every visible text element — headers, body, footers, page numbers, marginalia, stamps, annotations. - For text you cannot confidently read, write [ILLEGIBLE] inline at that position. - If the image contains no text, output [no text].
update_page_content yes
vision_mode auto

What this step asks the model:

Transcribe the text visible on this image.

Language: transcribe in the language of the source. Do not translate, and do not assume the document is in English.

Rules:
- Output ONLY the transcription. No headings, no preamble, no commentary,
  no summary, no notes, no explanations, no observations about quality or
  legibility, no descriptions of seals or images, no language about the
  difficulty of the handwriting.
- Preserve original layout, line breaks, and paragraph structure.
- Preserve original spelling and capitalisation, including ALL CAPS
  headers if they appear that way.
- Preserve orthography exactly as written, including all diacritics
  and accent marks (e.g., keep "Chocó" as "Chocó", never "Choco";
  keep "Ramón" as "Ramón", never "Ramon").
- Do not strip accents, tildes, cedillas, or umlauts. If a mark is
  visible, keep it.
- Include every visible text element — headers, body, marginalia, stamps,
  signatures (transcribe the signed name as written), printed labels,
  handwritten annotations.
- For text you cannot confidently read, use explicit uncertainty markers:
  [ilegible] for unreadable text and [uncertain] for plausible-but-low-
  confidence readings. Place the marker inline at the uncertain span.
  Do not guess. Do not fill in.
- Do NOT invent dates, numbers, names, or words that are not legibly
  present. Do not normalise dates ("23/7/1999" stays "23/7/1999", not
  "1999-07-23").
- Do NOT repeat any portion of the transcription. Output each visible
  passage exactly once.
- If the image contains no legible text, output the single token
  [sin texto].

4. Transcribe (manuscript)

Tool: Transcribe — Extract text from images (OCR)

Settings this step uses:

Option Value
language auto
prompt Transcribe the handwritten text visible on this image. The document contains modern handwriting (20th–21st century). Rules: - Output ONLY the transcription. No headings, preamble, commentary, or summary. - Preserve original layout, line breaks, and paragraph structure. - Preserve original spelling, capitalisation, and punctuation as written — do not correct errors. - Include every visible text element — headers, body, dates, signatures, annotations, marginalia. - Transcribe signatures as written (the full signed name if legible). - For text you cannot confidently read, write [ILLEGIBLE] inline at that position. Do not guess. - Do NOT invent or normalise content. What is on the page is the source of truth. - If the image contains no legible text, output [no text].
update_page_content yes
vision_mode auto

What this step asks the model:

Transcribe the text visible on this image.

Language: transcribe in the language of the source. Do not translate, and do not assume the document is in English.

Rules:
- Output ONLY the transcription. No headings, no preamble, no commentary,
  no summary, no notes, no explanations, no observations about quality or
  legibility, no descriptions of seals or images, no language about the
  difficulty of the handwriting.
- Preserve original layout, line breaks, and paragraph structure.
- Preserve original spelling and capitalisation, including ALL CAPS
  headers if they appear that way.
- Preserve orthography exactly as written, including all diacritics
  and accent marks (e.g., keep "Chocó" as "Chocó", never "Choco";
  keep "Ramón" as "Ramón", never "Ramon").
- Do not strip accents, tildes, cedillas, or umlauts. If a mark is
  visible, keep it.
- Include every visible text element — headers, body, marginalia, stamps,
  signatures (transcribe the signed name as written), printed labels,
  handwritten annotations.
- For text you cannot confidently read, use explicit uncertainty markers:
  [ilegible] for unreadable text and [uncertain] for plausible-but-low-
  confidence readings. Place the marker inline at the uncertain span.
  Do not guess. Do not fill in.
- Do NOT invent dates, numbers, names, or words that are not legibly
  present. Do not normalise dates ("23/7/1999" stays "23/7/1999", not
  "1999-07-23").
- Do NOT repeat any portion of the transcription. Output each visible
  passage exactly once.
- If the image contains no legible text, output the single token
  [sin texto].

5. Transcribe — HTR Pass 1 (Draft)

Tool: Transcribe — Extract text from images (OCR)

Settings this step uses:

Option Value
language auto
prompt You are performing Handwritten Text Recognition on a historical document in a single careful pass. Take your time and think through hard lines before committing to a reading. Before transcribing, note in a single bracketed line: [Script: ; hand: ; language: ] Then output ONLY the transcription — no headings, preamble, or commentary after that line. Transcription discipline (non-negotiable): - Transcribe the WHOLE page in reading order. Preserve original spelling, capitalisation, punctuation, and line breaks exactly. Never modernise. - Square brackets ONLY around letters you supply by expanding an abbreviation or contraction — mer[ce]d, ma[gesta]d, v[ecin]o. Never bracket letters that are visible on the page, never one bracket per letter, and never a full stop after every word: write continuous text the way the scribe did. - [UNCERTAIN: word] when you have a plausible reading; [ILLEGIBLE] only when the strokes are truly unreadable. An honest [UNCERTAIN] is worth more than a fluent guess. - You know the period’s formulas. Use them to READ, not to write: where damage hides a standard formula, propose it as [UNCERTAIN: …], never as clean text. - Keep names, dates, and places consistent with what the document itself establishes (docket/cover, headings, other pages of the batch). If two readings conflict, flag the conflict — do not silently choose. - [M.N.] marginalia at position · [deleted: …] struck text · interlinear insertions inline at the insertion point · [Rúbrica] / [Sello] for flourishes and stamps · [sin texto] for a page with no legible text. Historical handwriting (16th–19th C., legible hands): - Expect period orthography (u/v, i/j/y, long-s, b/v instability in Spanish; long-s and thorn-as-y in English) — transcribe as written. - Common Spanish abbreviations: Vmd, dho/dha, q̃, tpo = t[iem]po, Sr., S.M.; expand with brackets only when unambiguous. - Include all visible text: headers, marginalia [M.N.], stamps [Sello], signatures as written, rubric flourishes [Rúbrica].
thinking_mode long
update_page_content no
vision_mode auto

What this step asks the model:

Transcribe the text visible on this image.

Language: transcribe in the language of the source. Do not translate, and do not assume the document is in English.

Rules:
- Output ONLY the transcription. No headings, no preamble, no commentary,
  no summary, no notes, no explanations, no observations about quality or
  legibility, no descriptions of seals or images, no language about the
  difficulty of the handwriting.
- Preserve original layout, line breaks, and paragraph structure.
- Preserve original spelling and capitalisation, including ALL CAPS
  headers if they appear that way.
- Preserve orthography exactly as written, including all diacritics
  and accent marks (e.g., keep "Chocó" as "Chocó", never "Choco";
  keep "Ramón" as "Ramón", never "Ramon").
- Do not strip accents, tildes, cedillas, or umlauts. If a mark is
  visible, keep it.
- Include every visible text element — headers, body, marginalia, stamps,
  signatures (transcribe the signed name as written), printed labels,
  handwritten annotations.
- For text you cannot confidently read, use explicit uncertainty markers:
  [ilegible] for unreadable text and [uncertain] for plausible-but-low-
  confidence readings. Place the marker inline at the uncertain span.
  Do not guess. Do not fill in.
- Do NOT invent dates, numbers, names, or words that are not legibly
  present. Do not normalise dates ("23/7/1999" stays "23/7/1999", not
  "1999-07-23").
- Do NOT repeat any portion of the transcription. Output each visible
  passage exactly once.
- If the image contains no legible text, output the single token
  [sin texto].

Tool: Search — Find files matching a search query

Settings this step uses:

Option Value
limit 5
search_type hybrid
status_filter completed

7. Transcribe — HTR Pass 2 (Review)

Tool: Transcribe Review — Second-pass QA of a prior transcription against the image

Settings this step uses:

Option Value
language auto
prompt You are a paleographer reviewing an EXISTING transcription of the manuscript image you are shown. The prior transcription is provided as context. Take your time — think through every hard line against the strokes before accepting or changing a reading. This pass: abbreviations, contractions, and formulary. - Verify every expansion. Brackets belong ONLY around letters supplied by expansion (mer[ce]d, v[ecin]o). Remove bracket noise around letters that are visible on the page, and remove any spurious full stops between words — restore continuous text. - Expand abbreviations the draft left unexpanded where the expansion is unambiguous; otherwise keep the abbreviated form exactly as written. - Identify the document’s genre (carta de poder, confesión, probanza, auto, oficio…) and check the draft against its formula skeleton. Where the draft’s text breaks the formula AND the image is damaged or ambiguous there, propose the formula reading as [UNCERTAIN: …] — never insert formula text as if it were read. - Output the COMPLETE corrected transcription (same line breaks), nothing else. Then apply, in the same pass: orthography and letter-shape audit for the period, and name/date consistency — flag conflicts, never silently choose.
thinking_mode medium
update_page_content yes
vision_mode llm

What this step asks the model:

You are a paleographer reviewing an EXISTING transcription of the manuscript image you are shown. The prior transcription is provided as context. Take your time — think through every hard line against the strokes before accepting or changing a reading.

**This pass: abbreviations, contractions, and formulary.**
- Verify every expansion. Brackets belong ONLY around letters supplied by expansion (mer[ce]d, v[ecin]o). Remove bracket noise around letters that are visible on the page, and remove any spurious full stops between words — restore continuous text.
- Expand abbreviations the draft left unexpanded where the expansion is unambiguous; otherwise keep the abbreviated form exactly as written.
- Identify the document's genre (carta de poder, confesión, probanza, auto, oficio...) and check the draft against its formula skeleton. Where the draft's text breaks the formula AND the image is damaged or ambiguous there, propose the formula reading as [UNCERTAIN: ...] — never insert formula text as if it were read.
- Output the COMPLETE corrected transcription (same line breaks), nothing else.

Then apply, in the same pass: orthography and letter-shape audit for the period, and name/date consistency — flag conflicts, never silently choose.

8. Transcribe — Paleography Pass 1 (Draft)

Tool: Transcribe — Extract text from images (OCR)

Settings this step uses:

Option Value
language auto
prompt You are a paleographer transcribing a historical manuscript image in a single careful pass. Take your time and think through hard lines before committing to a reading. Before transcribing, note in a single bracketed line: [Script: ; hand: ; language: ] Then output ONLY the transcription — no headings, preamble, or commentary after that line. Transcription discipline (non-negotiable): - Transcribe the WHOLE page in reading order. Preserve original spelling, capitalisation, punctuation, and line breaks exactly. Never modernise. - Square brackets ONLY around letters you supply by expanding an abbreviation or contraction — mer[ce]d, ma[gesta]d, v[ecin]o. Never bracket letters that are visible on the page, never one bracket per letter, and never a full stop after every word: write continuous text the way the scribe did. - [UNCERTAIN: word] when you have a plausible reading; [ILLEGIBLE] only when the strokes are truly unreadable. An honest [UNCERTAIN] is worth more than a fluent guess. - You know the period’s formulas. Use them to READ, not to write: where damage hides a standard formula, propose it as [UNCERTAIN: …], never as clean text. - Keep names, dates, and places consistent with what the document itself establishes (docket/cover, headings, other pages of the batch). If two readings conflict, flag the conflict — do not silently choose. - [M.N.] marginalia at position · [deleted: …] struck text · interlinear insertions inline at the insertion point · [Rúbrica] / [Sello] for flourishes and stamps · [sin texto] for a page with no legible text. Period guidance (auto-detect the script family): - Procesal/cortesana (16th–17th C.): heavy line-end abbreviation, stretched baselines, knotted connections; confusions rn/m, c/e, u/n, long-s/f, t/l. - Albalaes/privilegios (12th–15th C.): abbreviation by suspension and contraction; ligatured letters. - Earlier scripts (carolingia, visigótica): isolated letters; heavy superscript abbreviation. - Humanística/itálica (16th C. onward): closer to modern forms; watch long-s and u/v. - If you recognise the language, apply its paleographic conventions (Spanish, Latin, and English guidance exist as dedicated presets — this one adapts to what it sees).
thinking_mode long
update_page_content no
vision_mode auto

What this step asks the model:

Transcribe the text visible on this image.

Language: transcribe in the language of the source. Do not translate, and do not assume the document is in English.

Rules:
- Output ONLY the transcription. No headings, no preamble, no commentary,
  no summary, no notes, no explanations, no observations about quality or
  legibility, no descriptions of seals or images, no language about the
  difficulty of the handwriting.
- Preserve original layout, line breaks, and paragraph structure.
- Preserve original spelling and capitalisation, including ALL CAPS
  headers if they appear that way.
- Preserve orthography exactly as written, including all diacritics
  and accent marks (e.g., keep "Chocó" as "Chocó", never "Choco";
  keep "Ramón" as "Ramón", never "Ramon").
- Do not strip accents, tildes, cedillas, or umlauts. If a mark is
  visible, keep it.
- Include every visible text element — headers, body, marginalia, stamps,
  signatures (transcribe the signed name as written), printed labels,
  handwritten annotations.
- For text you cannot confidently read, use explicit uncertainty markers:
  [ilegible] for unreadable text and [uncertain] for plausible-but-low-
  confidence readings. Place the marker inline at the uncertain span.
  Do not guess. Do not fill in.
- Do NOT invent dates, numbers, names, or words that are not legibly
  present. Do not normalise dates ("23/7/1999" stays "23/7/1999", not
  "1999-07-23").
- Do NOT repeat any portion of the transcription. Output each visible
  passage exactly once.
- If the image contains no legible text, output the single token
  [sin texto].

Tool: Search — Find files matching a search query

Settings this step uses:

Option Value
limit 5
search_type hybrid
status_filter completed

10. Transcribe — Paleography Pass 2 (Review)

Tool: Transcribe Review — Second-pass QA of a prior transcription against the image

Settings this step uses:

Option Value
language auto
prompt You are a paleographer reviewing an EXISTING transcription of the manuscript image you are shown. The prior transcription is provided as context. Take your time — think through every hard line against the strokes before accepting or changing a reading. This pass: abbreviations, contractions, and formulary. - Verify every expansion. Brackets belong ONLY around letters supplied by expansion (mer[ce]d, v[ecin]o). Remove bracket noise around letters that are visible on the page, and remove any spurious full stops between words — restore continuous text. - Expand abbreviations the draft left unexpanded where the expansion is unambiguous; otherwise keep the abbreviated form exactly as written. - Identify the document’s genre (carta de poder, confesión, probanza, auto, oficio…) and check the draft against its formula skeleton. Where the draft’s text breaks the formula AND the image is damaged or ambiguous there, propose the formula reading as [UNCERTAIN: …] — never insert formula text as if it were read. - Output the COMPLETE corrected transcription (same line breaks), nothing else. Then apply, in the same pass: orthography and letter-shape audit for the period, and name/date consistency — flag conflicts, never silently choose.
thinking_mode long
update_page_content yes
vision_mode llm

What this step asks the model:

You are a paleographer reviewing an EXISTING transcription of the manuscript image you are shown. The prior transcription is provided as context. Take your time — think through every hard line against the strokes before accepting or changing a reading.

**This pass: abbreviations, contractions, and formulary.**
- Verify every expansion. Brackets belong ONLY around letters supplied by expansion (mer[ce]d, v[ecin]o). Remove bracket noise around letters that are visible on the page, and remove any spurious full stops between words — restore continuous text.
- Expand abbreviations the draft left unexpanded where the expansion is unambiguous; otherwise keep the abbreviated form exactly as written.
- Identify the document's genre (carta de poder, confesión, probanza, auto, oficio...) and check the draft against its formula skeleton. Where the draft's text breaks the formula AND the image is damaged or ambiguous there, propose the formula reading as [UNCERTAIN: ...] — never insert formula text as if it were read.
- Output the COMPLETE corrected transcription (same line breaks), nothing else.

Then apply, in the same pass: orthography and letter-shape audit for the period, and name/date consistency — flag conflicts, never silently choose.