Skip to content

Clean Dates (page)

🤖 AI Drafted (Not reviewed)

Group near-duplicate dates within each page using a focused LLM call. Re-points claims at the merged canonical entity.

Tool id dates_page_cleanup
Category llm
Uses a language model yes
Needs a generative model no
Runs over many items no
Item handling batch
Structured output no
Human-verified not yet

What it reads

Port Type Required What it is
Text (text) text no Passthrough text from the upstream extractor.
Records (records) array yes Per-page records [{doc_id, text}, …] from the upstream Aggregate node. Page cleanup uses doc_id to scope its DB read to one page at a time.

What it emits

Port Type Required What it is
Text (text) text Raw text response
Value (value) any Parsed value
Texts (texts) array Per-item texts
Values (values) array Per-item values
Results (results) json Full results
Records (records) array Per-document text records [{doc_id, text}, …].
Artifacts (artifacts) json Artifact IDs

Options

Option Type Default What it does
choices array Valid choices. (Not shown in the editor.)
chunk_size_chars integer 0 Chunk large input text above this character budget (0=auto).
match_mode string prefer Match mode. One of: prefer, strict, inform.
max_items integer 10 List max items.
max_tokens integer 8192 Max response.
max_words integer 50 Word limit.
metadata_field string Save to field.
model_name string Model name.
output_format string text Response format. One of: text, boolean, choice, number, words, list, json.
prompt string Custom prompt.
provider_name string LLM provider. One of: openai, anthropic, google, ollama, lmstudio, groq, together, deepseek, mistral, openrouter, dashscope, xai, perplexity, fireworks, deepl.
quality_gate boolean yes Stop the run if output is unreadable.
reference_values object Known values to match. (Not shown in the editor.)
save_to_db boolean yes Save to library.
save_to_file boolean no Export to file.
temperature number 0.7 Creativity.
thinking_mode string off Chain-of-thought reasoning depth. One of: off, short, medium, long.

The prompt it sends

This is what the tool asks a model, with every option left at its default. Changing the options above changes this text.

You are an expert archivist deduplicating dates extracted from a document. Different entries may refer to the same date via spelling variants.

Duplicate rule: Two entries refer to the same date if their normalised YYYY-MM-DD strings are identical.

You are deduplicating, not curating. Every numbered input MUST appear in your output as a canonical or as an alias — total across all groups must equal 3. Do NOT invent new entries.

Title Case the canonical (re-case ALL-CAPS entries). Keep accents (María, José, Chocó). Entries with no duplicates become their own group with empty aliases.

Return ONLY valid JSON, no prose, no fences:
{"groups": [{"canonical": "...", "aliases": ["...", "..."]}, ...]}

---
Items to deduplicate:
1. Don Mateo Restrepo
2. Don Mateo
3. D. Mateo