← Back to writing

Query Entity Disambiguation Multilingual

By TaeHo Kim (Hannune)  ·  Published July 25, 2026

Query Entity Disambiguation Multilingual

In 2asy.ai, the hardest retrieval problem is not missing data. It is the query that arrives with an ambiguous entity name.

"Hyundai" matches seventeen separate nodes in the knowledge graph. A vector search returns all of them ranked by embedding similarity. A graph traversal needs one starting node.

The disambiguation step runs between query parsing and graph traversal. It ranks candidate nodes using three signals:

Corpus frequency prior. Hyundai Motor appears in roughly two-thirds of news documents where "Hyundai" appears without additional context. In the absence of other signals, the frequency-weighted candidate is right more often than any alternative default.

Query context coherence. If the same query mentions battery technology or manufacturing capacity, co-occurrence statistics across entity descriptions shift probability mass toward Hyundai Motor. If the query mentions construction permits, Hyundai Engineering moves up the ranking. The entity mention does not arrive in isolation; the surrounding query terms provide signal that the disambiguation layer can use.

Temporal suppression. A historical subsidiary with a closed valid-to date should rank below currently-active entities for queries that do not specify a historical time window. The graph has the structure to express this, but the query layer has to use it.

The failure mode is silent. When disambiguation selects the wrong entity, the graph traversal returns a coherent subgraph, the generation step produces a confident-sounding answer, and nothing in the output signals that the retrieval started from the wrong node. The user does not know they asked about Hyundai Steel when they meant Hyundai Motor.

The fix is one line before the answer: surface the disambiguated entity explicitly. Showing "retrieving for: Hyundai Motor Company" before the answer converts a silent failure into something the user can catch and correct. The graph traversal machinery is already doing the work. The only change is making the disambiguation decision visible before the answer is rendered.

Building something similar?

Get in touch →