Clean Events (page)
🤖 AI Drafted (Not reviewed)
Group near-duplicate events within each page using a focused LLM call. Re-points claims at the merged canonical entity.
|
|
| Tool id |
events_page_cleanup |
| Category |
llm |
| Uses a language model |
yes |
| Needs a generative model |
no |
| Runs over many items |
no |
| Item handling |
batch |
| Structured output |
no |
| Human-verified |
not yet |
What it reads
| Port |
Type |
Required |
What it is |
Text (text) |
text |
no |
Passthrough text from the upstream extractor. |
Records (records) |
array |
yes |
Per-page records [{doc_id, text}, …] from the upstream Aggregate node. Page cleanup uses doc_id to scope its DB read to one page at a time. |
What it emits
| Port |
Type |
Required |
What it is |
Text (text) |
text |
— |
Raw text response |
Value (value) |
any |
— |
Parsed value |
Texts (texts) |
array |
— |
Per-item texts |
Values (values) |
array |
— |
Per-item values |
Results (results) |
json |
— |
Full results |
Records (records) |
array |
— |
Per-document text records [{doc_id, text}, …]. |
Artifacts (artifacts) |
json |
— |
Artifact IDs |
Options
| Option |
Type |
Default |
What it does |
choices |
array |
— |
Valid choices. (Not shown in the editor.) |
chunk_size_chars |
integer |
0 |
Chunk large input text above this character budget (0=auto). |
match_mode |
string |
prefer |
Match mode. One of: prefer, strict, inform. |
max_items |
integer |
10 |
List max items. |
max_tokens |
integer |
8192 |
Max response. |
max_words |
integer |
50 |
Word limit. |
metadata_field |
string |
— |
Save to field. |
model_name |
string |
— |
Model name. |
output_format |
string |
text |
Response format. One of: text, boolean, choice, number, words, list, json. |
prompt |
string |
— |
Custom prompt. |
provider_name |
string |
— |
LLM provider. One of: openai, anthropic, google, ollama, lmstudio, groq, together, deepseek, mistral, openrouter, dashscope, xai, perplexity, fireworks, deepl. |
quality_gate |
boolean |
yes |
Stop the run if output is unreadable. |
reference_values |
object |
— |
Known values to match. (Not shown in the editor.) |
save_to_db |
boolean |
yes |
Save to library. |
save_to_file |
boolean |
no |
Export to file. |
temperature |
number |
0.7 |
Creativity. |
thinking_mode |
string |
off |
Chain-of-thought reasoning depth. One of: off, short, medium, long. |
The prompt it sends
This is what the tool asks a model, with every option left at its default. Changing the options above changes this text.
You are an expert archivist deduplicating events extracted from a document. Different entries may refer to the same event via spelling variants.
Duplicate rule: Two entries refer to the same event if they describe the same incident, transaction, signing, meeting, voyage, ruling, death, or transfer — even when worded differently or seen from different angles. Pick the most precise and concise description as canonical, in evidentiary phrasing ('the file records that X', 'Y is reported to have...'), with the alternative wordings as aliases.
You are deduplicating, not curating. Every numbered input MUST appear in your output as a canonical or as an alias — total across all groups must equal 3. Do NOT invent new entries.
Title Case the canonical (re-case ALL-CAPS entries). Keep accents (María, José, Chocó). Entries with no duplicates become their own group with empty aliases.
Return ONLY valid JSON, no prose, no fences:
{"groups": [{"canonical": "...", "aliases": ["...", "..."]}, ...]}
---
Items to deduplicate:
1. Don Mateo Restrepo
2. Don Mateo
3. D. Mateo