Skip to content

Transcribe Typescript

🤖 AI Drafted (Not reviewed)

Use this when: you have typewritten or printed documents — typescripts, books, pamphlets, official records with clear print. Uses Apple Vision (on-device, fast) when configured; falls back to the LLM vision model for scanned pages with poor contrast.

Folder /Transcribe
Steps 2
Tags preset, transcribe, typescript, ocr, printed

Steps, in run order

1. Files

Tool: Files — Pass through input files from workflow context

2. Transcribe (typescript)

Tool: Transcribe — Extract text from images (OCR)

Settings this step uses:

Option Value
language auto
prompt Transcribe the text visible on this image. The document is typewritten or printed. Rules: - Output ONLY the transcription. No headings, preamble, commentary, or summary. - Preserve original layout, line breaks, and paragraph structure. - Preserve original spelling, capitalisation, and punctuation exactly as printed. - Include every visible text element — headers, body, footers, page numbers, marginalia, stamps, annotations. - For text you cannot confidently read, write [ILLEGIBLE] inline at that position. - If the image contains no text, output [no text].
update_page_content yes
vision_mode auto

What this step asks the model:

Transcribe the text visible on this image.

Language: transcribe in the language of the source. Do not translate, and do not assume the document is in English.

Rules:
- Output ONLY the transcription. No headings, no preamble, no commentary,
  no summary, no notes, no explanations, no observations about quality or
  legibility, no descriptions of seals or images, no language about the
  difficulty of the handwriting.
- Preserve original layout, line breaks, and paragraph structure.
- Preserve original spelling and capitalisation, including ALL CAPS
  headers if they appear that way.
- Preserve orthography exactly as written, including all diacritics
  and accent marks (e.g., keep "Chocó" as "Chocó", never "Choco";
  keep "Ramón" as "Ramón", never "Ramon").
- Do not strip accents, tildes, cedillas, or umlauts. If a mark is
  visible, keep it.
- Include every visible text element — headers, body, marginalia, stamps,
  signatures (transcribe the signed name as written), printed labels,
  handwritten annotations.
- For text you cannot confidently read, use explicit uncertainty markers:
  [ilegible] for unreadable text and [uncertain] for plausible-but-low-
  confidence readings. Place the marker inline at the uncertain span.
  Do not guess. Do not fill in.
- Do NOT invent dates, numbers, names, or words that are not legibly
  present. Do not normalise dates ("23/7/1999" stays "23/7/1999", not
  "1999-07-23").
- Do NOT repeat any portion of the transcription. Output each visible
  passage exactly once.
- If the image contains no legible text, output the single token
  [sin texto].