← Back to writing

Benchmarking CJK Entity-Name Resolution

By TaeHo Kim (Hannune)  ·  Published October 9, 2026

A Korean name can change under romanization; a company can gain or lose a corporate suffix; Chinese names can switch scripts; and Japanese names can cross from kanji or kana into Latin letters. We built and ran a small CJK entity-name benchmark to see how our ER API handled those pairs. The numbers below are measurements from our own benchmark, not independent validation.

What we tested

The task was pair classification: given two names, decide whether they refer to the same entity. This set includes company names, personal names, and an organization name. It is not a company-only benchmark. Positive examples came from Wikidata-derived labels and, for some Korean categories, rule-generated variants. Negative examples pair different entities.

We evaluated a held-out set of 100 labeled pairs. Its Wikidata QIDs did not overlap with those used in the earlier 149-pair development set or the separate training set. The table shows the final held-out composition.

CategoryMatchNo match
kr_romanization1312
kr_corp_suffix1312
zh_variants1312
ja_variants1212
ja_rename10
Total5248

What the pair matcher scored

On the held-out 100 pairs, the /v1/splink-pairs path scored F1 0.9903: 51 true positives, 0 false positives, 1 false negative, and 48 true negatives. The earlier 149-pair development set scored F1 0.9686 on the same pair-comparison path. These are self-run benchmark measurements, not estimates of accuracy on customer data.

The missed pair was “ZERO1” and “Pro Wrestling Zero1,” two names for the same wrestling promotion. The short name was ambiguous, and the LLM fallback returned no match with low confidence. That single miss matters more than a rounded score: a name that needs outside context can still defeat this matcher.

What sits behind ER API

ER API builds on Splink, the open-source record-linkage engine. The tested path adds CJK normalization and a selective LLM fallback for pairs that the earlier matching steps leave unresolved. Splink deserves credit for the underlying pair-matching work; the CJK and LLM steps are our additions.

Limits of this result

  • We made, labeled, and ran this benchmark ourselves. No independent evaluator validated the score.
  • The holdout contains 100 pairs derived from Wikidata data. Some Korean variants were generated by rules, so the sample does not capture every kind of real-world name noise.
  • The held-out Korean negatives did not include the especially similar romanized names that caused false positives in the earlier development set. Their absence here does not show that issue was fixed.
  • LLM judgments can vary between runs. The result above is the recorded benchmark run.
  • The F1 scores describe pair comparison through /v1/splink-pairs. Registry-backed /v1/match depends on which entities are in the registry and had much lower recall in this evaluation. These scores are not a claim about that path or a production-wide guarantee.

Try ER API on your own names

Open the API console or email contact@hannune.ai.