Your shopping cart is empty!
WordLift's new analyzer leaves an entity blank rather than guess

- WordLift released Content Analysis v3, which turns text into stable Wikidata entity identities across eight languages.
- The system abstains when the evidence is thin instead of guessing.
- On its own evaluation sets it already scores higher than Google Cloud Natural Language in German, English and Spanish, and the company writes up where it still fails.
The company, which started fifteen years ago with the IKS and Apache Stanbol projects, argues that entity resolution is becoming core infrastructure for AI. So WordLift published the third version of Content Analysis and explained how it works, what it measured and where it still fails. I could not open the original (the site returned a 403), so what follows relies on the post's announcement.
From a word to a permanent identifier
The idea is easy to explain and hard to build. The word "Apple" in a text can be a fruit, a company or a surname. Entity resolution means picking the right meaning and tying it to a stable record, here Wikidata, where every object has a permanent ID. Content Analysis v3 does this for eight languages and returns a set of identities that other systems can reference instead of a flat list of words.
The developers treat this task as the same kind of foundation for AI that semantic markup once was for the web. Search engines and generative models lean on entities when they decide what a page is about and whom to trust. That makes it useful for a site owner to know which entities from their text a machine sees and which it loses. A practical way to check is in our checklist for auditing your AI entity footprint.
A system that knows how to decline
The most interesting detail in the announcement is abstention. When the text gives thin evidence, v3 leaves the entity unresolved and slots nothing in at random. For an analyst that is the difference between a false link that quietly corrupts data later and an honest gap you can see. Where output loads into a knowledge graph, that behavior is worth more than a couple of extra points of recall.
| Claim | Detail |
|---|---|
| Identifiers | Stable Wikidata records |
| Languages | Eight |
| Comparison | Higher than Google Cloud Natural Language in German, English and Spanish on WordLift's own sets |
| Weak spots | The company separately describes where the system fails |
Why to treat these numbers with care
The comparison runs on the company's own evaluation sets, with no independent benchmark in sight. Beating Google Cloud Natural Language in three languages out of eight means there is no such claim for the other five. For smaller markets the question is sharper: whether Ukrainian is among those eight languages is not visible from the announcement, and local companies and shops are usually thinly represented in Wikidata. If the record does not exist, even the best tool has nothing to match against.
A simple practical point about data follows. Registering your brand and key products in Wikidata and keeping a consistent Organization schema on the site helps any resolver, whether it is WordLift or Google. There is a piece on the basics of this approach, a single source of truth for AI.
What interested me more than the percentages is that the company writes openly about its own failures. That tone is rare among SEO tools, and I am curious whether any competitor will answer in kind.


