Clean Organizations (page)
🤖 AI Drafted (Not reviewed)
Group near-duplicate organizations within each page using a focused LLM call. Re-points claims at the merged canonical entity.
|
|
| Tool id |
organizations_page_cleanup |
| Category |
llm |
| Uses a language model |
yes |
| Needs a generative model |
no |
| Runs over many items |
no |
| Item handling |
batch |
| Structured output |
no |
| Human-verified |
not yet |
What it reads
| Port |
Type |
Required |
What it is |
Text (text) |
text |
no |
Passthrough text from the upstream extractor. |
Records (records) |
array |
yes |
Per-page records [{doc_id, text}, …] from the upstream Aggregate node. Page cleanup uses doc_id to scope its DB read to one page at a time. |
What it emits
| Port |
Type |
Required |
What it is |
Text (text) |
text |
— |
Raw text response |
Value (value) |
any |
— |
Parsed value |
Texts (texts) |
array |
— |
Per-item texts |
Values (values) |
array |
— |
Per-item values |
Results (results) |
json |
— |
Full results |
Records (records) |
array |
— |
Per-document text records [{doc_id, text}, …]. |
Artifacts (artifacts) |
json |
— |
Artifact IDs |
Options
| Option |
Type |
Default |
What it does |
choices |
array |
— |
Valid choices. (Not shown in the editor.) |
chunk_size_chars |
integer |
0 |
Chunk large input text above this character budget (0=auto). |
match_mode |
string |
prefer |
Match mode. One of: prefer, strict, inform. |
max_items |
integer |
10 |
List max items. |
max_tokens |
integer |
8192 |
Max response. |
max_words |
integer |
50 |
Word limit. |
metadata_field |
string |
— |
Save to field. |
model_name |
string |
— |
Model name. |
output_format |
string |
text |
Response format. One of: text, boolean, choice, number, words, list, json. |
prompt |
string |
— |
Custom prompt. |
provider_name |
string |
— |
LLM provider. One of: openai, anthropic, google, ollama, lmstudio, groq, together, deepseek, mistral, openrouter, dashscope, xai, perplexity, fireworks, deepl. |
quality_gate |
boolean |
yes |
Stop the run if output is unreadable. |
reference_values |
object |
— |
Known values to match. (Not shown in the editor.) |
save_to_db |
boolean |
yes |
Save to library. |
save_to_file |
boolean |
no |
Export to file. |
temperature |
number |
0.7 |
Creativity. |
thinking_mode |
string |
off |
Chain-of-thought reasoning depth. One of: off, short, medium, long. |
The prompt it sends
This is what the tool asks a model, with every option left at its default. Changing the options above changes this text.
You are an expert archivist deduplicating organizations extracted from a document. Different entries may refer to the same organisation via spelling variants.
Duplicate rule: Two entries refer to the same organisation if names are spelling variants, abbreviations (Cía. / Compañía, S.A. / Sociedad Anónima), or one is the long form and the other a short form. Pick the most complete form as canonical.
You are deduplicating, not curating. Every numbered input MUST appear in your output as a canonical or as an alias — total across all groups must equal 3. Do NOT invent new entries.
Title Case the canonical (re-case ALL-CAPS entries). Keep accents (María, José, Chocó). Entries with no duplicates become their own group with empty aliases.
Return ONLY valid JSON, no prose, no fences:
{"groups": [{"canonical": "...", "aliases": ["...", "..."]}, ...]}
---
Items to deduplicate:
1. Don Mateo Restrepo
2. Don Mateo
3. D. Mateo