Transcribe (Auto-Detect)
🤖 AI Drafted (Not reviewed)
Use this when: you don’t know the document’s script type and want automatic routing. Classifies the script type (Typescript / Manuscript / HTR / Paleography) then runs the matching transcription profile — including two-pass review for historical and archaic scripts.
| Folder | /Transcribe |
| Steps | 10 |
| Tags | preset, transcribe, auto-detect, classify, routing |
Steps, in run order
1. Files
Tool: Files — Pass through input files from workflow context
2. Classify Script Type
Tool: Classify Script Type — Detect whether a document is typescript, manuscript, HTR, or paleography
Settings this step uses:
| Option | Value |
|---|---|
save_to_db |
yes |
vision_mode |
llm |
What this step asks the model:
Examine this document image and classify its script/handwriting type.
Choose EXACTLY ONE type from:
- typescript : typewritten, printed, or born-digital text (no handwriting)
- manuscript : modern handwriting, 20th–21st century (cursive or print)
- htr : historical handwriting, 16th–19th century (legible but archaic letterforms)
- paleography : archaic script with non-standard letterforms, heavy abbreviation, or
specialist scribal conventions (pre-18th century or highly specialised)
Return a JSON object with exactly these fields:
{
"script_type": "<one of the four types>",
"confidence": <float 0.0–1.0>,
"notes": "<one sentence explaining the classification>"
}
Be conservative: if the image is ambiguous between htr and paleography, pick the harder
category (paleography) and lower your confidence. Only mark typescript when there is
clearly no handwriting. Do not add any text outside the JSON object.
3. Transcribe (typescript)
Tool: Transcribe — Extract text from images (OCR)
Settings this step uses:
| Option | Value |
|---|---|
language |
auto |
prompt |
Transcribe the text visible on this image. The document is typewritten or printed. Rules: - Output ONLY the transcription. No headings, preamble, commentary, or summary. - Preserve original layout, line breaks, and paragraph structure. - Preserve original spelling, capitalisation, and punctuation exactly as printed. - Include every visible text element — headers, body, footers, page numbers, marginalia, stamps, annotations. - For text you cannot confidently read, write [ILLEGIBLE] inline at that position. - If the image contains no text, output [no text]. |
update_page_content |
yes |
vision_mode |
auto |
What this step asks the model:
Transcribe the text visible on this image.
Language: transcribe in the language of the source. Do not translate, and do not assume the document is in English.
Rules:
- Output ONLY the transcription. No headings, no preamble, no commentary,
no summary, no notes, no explanations, no observations about quality or
legibility, no descriptions of seals or images, no language about the
difficulty of the handwriting.
- Preserve original layout, line breaks, and paragraph structure.
- Preserve original spelling and capitalisation, including ALL CAPS
headers if they appear that way.
- Preserve orthography exactly as written, including all diacritics
and accent marks (e.g., keep "Chocó" as "Chocó", never "Choco";
keep "Ramón" as "Ramón", never "Ramon").
- Do not strip accents, tildes, cedillas, or umlauts. If a mark is
visible, keep it.
- Include every visible text element — headers, body, marginalia, stamps,
signatures (transcribe the signed name as written), printed labels,
handwritten annotations.
- For text you cannot confidently read, use explicit uncertainty markers:
[ilegible] for unreadable text and [uncertain] for plausible-but-low-
confidence readings. Place the marker inline at the uncertain span.
Do not guess. Do not fill in.
- Do NOT invent dates, numbers, names, or words that are not legibly
present. Do not normalise dates ("23/7/1999" stays "23/7/1999", not
"1999-07-23").
- Do NOT repeat any portion of the transcription. Output each visible
passage exactly once.
- If the image contains no legible text, output the single token
[sin texto].
4. Transcribe (manuscript)
Tool: Transcribe — Extract text from images (OCR)
Settings this step uses:
| Option | Value |
|---|---|
language |
auto |
prompt |
Transcribe the handwritten text visible on this image. The document contains modern handwriting (20th–21st century). Rules: - Output ONLY the transcription. No headings, preamble, commentary, or summary. - Preserve original layout, line breaks, and paragraph structure. - Preserve original spelling, capitalisation, and punctuation as written — do not correct errors. - Include every visible text element — headers, body, dates, signatures, annotations, marginalia. - Transcribe signatures as written (the full signed name if legible). - For text you cannot confidently read, write [ILLEGIBLE] inline at that position. Do not guess. - Do NOT invent or normalise content. What is on the page is the source of truth. - If the image contains no legible text, output [no text]. |
update_page_content |
yes |
vision_mode |
auto |
What this step asks the model:
Transcribe the text visible on this image.
Language: transcribe in the language of the source. Do not translate, and do not assume the document is in English.
Rules:
- Output ONLY the transcription. No headings, no preamble, no commentary,
no summary, no notes, no explanations, no observations about quality or
legibility, no descriptions of seals or images, no language about the
difficulty of the handwriting.
- Preserve original layout, line breaks, and paragraph structure.
- Preserve original spelling and capitalisation, including ALL CAPS
headers if they appear that way.
- Preserve orthography exactly as written, including all diacritics
and accent marks (e.g., keep "Chocó" as "Chocó", never "Choco";
keep "Ramón" as "Ramón", never "Ramon").
- Do not strip accents, tildes, cedillas, or umlauts. If a mark is
visible, keep it.
- Include every visible text element — headers, body, marginalia, stamps,
signatures (transcribe the signed name as written), printed labels,
handwritten annotations.
- For text you cannot confidently read, use explicit uncertainty markers:
[ilegible] for unreadable text and [uncertain] for plausible-but-low-
confidence readings. Place the marker inline at the uncertain span.
Do not guess. Do not fill in.
- Do NOT invent dates, numbers, names, or words that are not legibly
present. Do not normalise dates ("23/7/1999" stays "23/7/1999", not
"1999-07-23").
- Do NOT repeat any portion of the transcription. Output each visible
passage exactly once.
- If the image contains no legible text, output the single token
[sin texto].
5. Transcribe — HTR Pass 1 (Draft)
Tool: Transcribe — Extract text from images (OCR)
Settings this step uses:
| Option | Value |
|---|---|
language |
auto |
prompt |
You are performing Handwritten Text Recognition on a historical document in a single careful pass. Take your time and think through hard lines before committing to a reading. Before transcribing, note in a single bracketed line: [Script: |
thinking_mode |
long |
update_page_content |
no |
vision_mode |
auto |
What this step asks the model:
Transcribe the text visible on this image.
Language: transcribe in the language of the source. Do not translate, and do not assume the document is in English.
Rules:
- Output ONLY the transcription. No headings, no preamble, no commentary,
no summary, no notes, no explanations, no observations about quality or
legibility, no descriptions of seals or images, no language about the
difficulty of the handwriting.
- Preserve original layout, line breaks, and paragraph structure.
- Preserve original spelling and capitalisation, including ALL CAPS
headers if they appear that way.
- Preserve orthography exactly as written, including all diacritics
and accent marks (e.g., keep "Chocó" as "Chocó", never "Choco";
keep "Ramón" as "Ramón", never "Ramon").
- Do not strip accents, tildes, cedillas, or umlauts. If a mark is
visible, keep it.
- Include every visible text element — headers, body, marginalia, stamps,
signatures (transcribe the signed name as written), printed labels,
handwritten annotations.
- For text you cannot confidently read, use explicit uncertainty markers:
[ilegible] for unreadable text and [uncertain] for plausible-but-low-
confidence readings. Place the marker inline at the uncertain span.
Do not guess. Do not fill in.
- Do NOT invent dates, numbers, names, or words that are not legibly
present. Do not normalise dates ("23/7/1999" stays "23/7/1999", not
"1999-07-23").
- Do NOT repeat any portion of the transcription. Output each visible
passage exactly once.
- If the image contains no legible text, output the single token
[sin texto].
6. Reference Corpus Search
Tool: Search — Find files matching a search query
Settings this step uses:
| Option | Value |
|---|---|
limit |
5 |
search_type |
hybrid |
status_filter |
completed |
7. Transcribe — HTR Pass 2 (Review)
Tool: Transcribe Review — Second-pass QA of a prior transcription against the image
Settings this step uses:
| Option | Value |
|---|---|
language |
auto |
prompt |
You are a paleographer reviewing an EXISTING transcription of the manuscript image you are shown. The prior transcription is provided as context. Take your time — think through every hard line against the strokes before accepting or changing a reading. This pass: abbreviations, contractions, and formulary. - Verify every expansion. Brackets belong ONLY around letters supplied by expansion (mer[ce]d, v[ecin]o). Remove bracket noise around letters that are visible on the page, and remove any spurious full stops between words — restore continuous text. - Expand abbreviations the draft left unexpanded where the expansion is unambiguous; otherwise keep the abbreviated form exactly as written. - Identify the document’s genre (carta de poder, confesión, probanza, auto, oficio…) and check the draft against its formula skeleton. Where the draft’s text breaks the formula AND the image is damaged or ambiguous there, propose the formula reading as [UNCERTAIN: …] — never insert formula text as if it were read. - Output the COMPLETE corrected transcription (same line breaks), nothing else. Then apply, in the same pass: orthography and letter-shape audit for the period, and name/date consistency — flag conflicts, never silently choose. |
thinking_mode |
medium |
update_page_content |
yes |
vision_mode |
llm |
What this step asks the model:
You are a paleographer reviewing an EXISTING transcription of the manuscript image you are shown. The prior transcription is provided as context. Take your time — think through every hard line against the strokes before accepting or changing a reading.
**This pass: abbreviations, contractions, and formulary.**
- Verify every expansion. Brackets belong ONLY around letters supplied by expansion (mer[ce]d, v[ecin]o). Remove bracket noise around letters that are visible on the page, and remove any spurious full stops between words — restore continuous text.
- Expand abbreviations the draft left unexpanded where the expansion is unambiguous; otherwise keep the abbreviated form exactly as written.
- Identify the document's genre (carta de poder, confesión, probanza, auto, oficio...) and check the draft against its formula skeleton. Where the draft's text breaks the formula AND the image is damaged or ambiguous there, propose the formula reading as [UNCERTAIN: ...] — never insert formula text as if it were read.
- Output the COMPLETE corrected transcription (same line breaks), nothing else.
Then apply, in the same pass: orthography and letter-shape audit for the period, and name/date consistency — flag conflicts, never silently choose.
8. Transcribe — Paleography Pass 1 (Draft)
Tool: Transcribe — Extract text from images (OCR)
Settings this step uses:
| Option | Value |
|---|---|
language |
auto |
prompt |
You are a paleographer transcribing a historical manuscript image in a single careful pass. Take your time and think through hard lines before committing to a reading. Before transcribing, note in a single bracketed line: [Script: |
thinking_mode |
long |
update_page_content |
no |
vision_mode |
auto |
What this step asks the model:
Transcribe the text visible on this image.
Language: transcribe in the language of the source. Do not translate, and do not assume the document is in English.
Rules:
- Output ONLY the transcription. No headings, no preamble, no commentary,
no summary, no notes, no explanations, no observations about quality or
legibility, no descriptions of seals or images, no language about the
difficulty of the handwriting.
- Preserve original layout, line breaks, and paragraph structure.
- Preserve original spelling and capitalisation, including ALL CAPS
headers if they appear that way.
- Preserve orthography exactly as written, including all diacritics
and accent marks (e.g., keep "Chocó" as "Chocó", never "Choco";
keep "Ramón" as "Ramón", never "Ramon").
- Do not strip accents, tildes, cedillas, or umlauts. If a mark is
visible, keep it.
- Include every visible text element — headers, body, marginalia, stamps,
signatures (transcribe the signed name as written), printed labels,
handwritten annotations.
- For text you cannot confidently read, use explicit uncertainty markers:
[ilegible] for unreadable text and [uncertain] for plausible-but-low-
confidence readings. Place the marker inline at the uncertain span.
Do not guess. Do not fill in.
- Do NOT invent dates, numbers, names, or words that are not legibly
present. Do not normalise dates ("23/7/1999" stays "23/7/1999", not
"1999-07-23").
- Do NOT repeat any portion of the transcription. Output each visible
passage exactly once.
- If the image contains no legible text, output the single token
[sin texto].
9. Reference Corpus Search
Tool: Search — Find files matching a search query
Settings this step uses:
| Option | Value |
|---|---|
limit |
5 |
search_type |
hybrid |
status_filter |
completed |
10. Transcribe — Paleography Pass 2 (Review)
Tool: Transcribe Review — Second-pass QA of a prior transcription against the image
Settings this step uses:
| Option | Value |
|---|---|
language |
auto |
prompt |
You are a paleographer reviewing an EXISTING transcription of the manuscript image you are shown. The prior transcription is provided as context. Take your time — think through every hard line against the strokes before accepting or changing a reading. This pass: abbreviations, contractions, and formulary. - Verify every expansion. Brackets belong ONLY around letters supplied by expansion (mer[ce]d, v[ecin]o). Remove bracket noise around letters that are visible on the page, and remove any spurious full stops between words — restore continuous text. - Expand abbreviations the draft left unexpanded where the expansion is unambiguous; otherwise keep the abbreviated form exactly as written. - Identify the document’s genre (carta de poder, confesión, probanza, auto, oficio…) and check the draft against its formula skeleton. Where the draft’s text breaks the formula AND the image is damaged or ambiguous there, propose the formula reading as [UNCERTAIN: …] — never insert formula text as if it were read. - Output the COMPLETE corrected transcription (same line breaks), nothing else. Then apply, in the same pass: orthography and letter-shape audit for the period, and name/date consistency — flag conflicts, never silently choose. |
thinking_mode |
long |
update_page_content |
yes |
vision_mode |
llm |
What this step asks the model:
You are a paleographer reviewing an EXISTING transcription of the manuscript image you are shown. The prior transcription is provided as context. Take your time — think through every hard line against the strokes before accepting or changing a reading.
**This pass: abbreviations, contractions, and formulary.**
- Verify every expansion. Brackets belong ONLY around letters supplied by expansion (mer[ce]d, v[ecin]o). Remove bracket noise around letters that are visible on the page, and remove any spurious full stops between words — restore continuous text.
- Expand abbreviations the draft left unexpanded where the expansion is unambiguous; otherwise keep the abbreviated form exactly as written.
- Identify the document's genre (carta de poder, confesión, probanza, auto, oficio...) and check the draft against its formula skeleton. Where the draft's text breaks the formula AND the image is damaged or ambiguous there, propose the formula reading as [UNCERTAIN: ...] — never insert formula text as if it were read.
- Output the COMPLETE corrected transcription (same line breaks), nothing else.
Then apply, in the same pass: orthography and letter-shape audit for the period, and name/date consistency — flag conflicts, never silently choose.