NO_CANDIDATE and AMBIGUOUS: Reading MCP for Wikidata Outcomes
Anyone who has tried to link messy real-world records to Wikidata knows the hardest part is not finding a result. It is deciding whether the result is trustworthy enough to use.
That is why the outcome language in the open-source “Wikidata + Google Knowledge Graph MCP” matters. The project is built to let AI agents search Wikidata, inspect Knowledge Graph MCP lookup selected facts, and connect local records to Wikidata QIDs with visible evidence and explicit uncertainty. It does not pretend that every lookup has a neat answer. Instead, it uses deterministic resolution outcomes such as AUTO_MATCH, HOLD, AMBIGUOUS, and NO_CANDIDATE.
Of those, NO_CANDIDATE and AMBIGUOUS are often the most revealing. They tell you less about a single record than about the quality of your evidence, the shape of your data, and the discipline of the resolver. In practice, these are the outcomes that keep teams from over-linking, which is one of the easiest ways to pollute a catalog, a content pipeline, or a knowledge layer.
This is where MCP for Wikidata becomes more than a convenience. It becomes a restraint system.
What the server is actually doing
The project published as revanalex/wikidata-google-knowledge-mcp is an MCP server and CLI. It is designed for use in MCP clients such as Claude Code, Cursor, and Codex. On the Wikidata side, it can search without an account or API key. On the Google side, the Knowledge Graph Search API is optional rather than required.
That optional Google layer matters because the server does not frame Google as an oracle. The documented cross-check relies on exact identifier joins: /m/ against Wikidata property P646 and /g/ against P2671. Even when those align, the project treats agreement between Google and Wikidata as provider concordance, not proof of identity. That is a careful distinction, and a professional one. Two systems agreeing can strengthen confidence. It does not erase the need for judgment.
The server also stays intentionally bounded. By default, it returns three candidates, and at most five, instead of dumping a long raw search result set into the agent’s context. That design says a lot. It assumes a useful resolver should narrow, not flood. When you are working through entity resolution, more options do not always improve accuracy. Past a certain point, they mainly create surface area for rationalization.
The documented toolset reflects this practical orientation. There are tools to search, inspect entities, see related information, resolve, and check status. The CLI adds batch work and evidence export. Selected-fact retrieval can include ranks, qualifiers, and references on request, which is important because a bare label match is rarely enough for difficult records.
Why NO_CANDIDATE is often a healthy result
Teams tend to react badly to NO_CANDIDATE at first. It feels like failure. A search was run, a record was presented, and the system came back empty-handed. If your mental model is “every decent record should map somewhere,” then NO_CANDIDATE feels like the system is underperforming.
In many cases, it is doing the opposite.
A disciplined resolver should be willing to say, “I do not see a defensible candidate in the bounded set.” That is exactly what NO_CANDIDATE communicates. Not “there is definitely no corresponding entity anywhere.” Not “this thing does not exist in Wikidata.” Only that, given the search and resolution logic being used, no candidate rose to the threshold for consideration or action.
That difference is not semantic hair-splitting. It changes how you respond operationally.
In a loose workflow, an ambiguous name often gets forced onto the nearest plausible QID. The record looks resolved, the dashboard looks cleaner, and six months later someone discovers a musician was linked to a politician or a town was linked to a football club. The cleanup cost is usually higher than the original lookup cost. NO_CANDIDATE protects against that kind of false certainty.
I have seen this pattern in entity matching work outside Wikidata as well. The most expensive mistakes rarely come from blank fields. They come from confident wrong links that spread downstream into search indexes, analytics, recommendation layers, and editorial tools. A blank can be revisited. A bad identity claim tends to harden over time.
With this server, the bounded search behavior reinforces that discipline. If the top three, or even top five, candidates do not produce a credible match, the system does not manufacture one. That is exactly what many data teams need from MCP for Wikidata, especially when the downstream environment includes automation.
What AMBIGUOUS tells you that NO_CANDIDATE does not
If NO_CANDIDATE means “nothing credible surfaced,” AMBIGUOUS means something more interesting: credible possibilities did surface, but the available evidence did not separate them cleanly enough.
That distinction is easy to miss until you work with real records. A weak query and a crowded identity space can produce very different failure modes. NO_CANDIDATE usually points to insufficient or poorly aligned candidates. AMBIGUOUS points to a choice problem.
A classic case is a local record with a common label and little context. You might search a person’s name and receive several plausible Wikidata entities. Perhaps each has some overlap with the sparse local metadata. Maybe one shares a profession while another shares a geography. Maybe dates are absent or too broad to eliminate either. In that situation, a resolver that returns AMBIGUOUS is not hesitating out of weakness. It is refusing to fake precision.
That behavior becomes more valuable when evidence is inspectable. Since the project supports selected-fact retrieval with ranks, qualifiers, and references on request, you can inspect the competing candidates in a structured way. Often the answer is not hidden in the label at all. It is in a date qualifier, a role, a jurisdiction, a language, or a referenced statement that clarifies which entity your local record was really describing.
The important point is that AMBIGUOUS is not a dead end. It is a prompt to gather more disambiguating detail.
Reading these outcomes as workflow signals
A lot of teams make the mistake of treating resolver outputs as final business outcomes. In practice, they are often workflow signals.
AUTO_MATCH and HOLD are straightforward enough. The former suggests the deterministic logic found a strong enough match to link automatically. The latter suggests caution or review. But NO_CANDIDATE and AMBIGUOUS deserve slightly different handling because they point to different next actions.
Here is the simplest way to think about them:
- NO_CANDIDATE usually calls for broader or better input evidence.
- AMBIGUOUS usually calls for narrower or more discriminating input evidence.
That sounds subtle, but it is operationally useful. If you are seeing many NO_CANDIDATE outcomes, the issue may be shallow source records, poor query text, or records that are simply not represented clearly enough for bounded retrieval. If you are seeing many AMBIGUOUS outcomes, the issue may be that your records sit in a crowded name space and need stronger differentiators such as dates, category, place, or another known identifier.
The system’s deterministic resolution logic helps here because it removes some of the guesswork from interpretation. When a resolver behaves the same way under the same conditions, teams can learn from its outcomes rather than treating them as mood swings.
Why bounded search changes the meaning of failure
One of the most interesting design choices in this project is the bounded candidate set. By default, three candidates come back. Five is the maximum. That means both NO_CANDIDATE and AMBIGUOUS occur in a constrained decision environment.
This is not a trivial implementation detail. It changes what the outcomes mean.
In an unbounded search system, NO_CANDIDATE might simply mean the right result was buried on page four and never surfaced meaningfully. In a bounded system, it means the resolver is prioritizing precision and inspectability over exhaustive recall in a single pass. That makes the outcome more actionable. You know the system did not scan and dump dozens of names for a human to sort manually. It looked within a small, intentional window and still did not find a defensible path.
Likewise, AMBIGUOUS in a bounded system means the ambiguity exists among the best few surfaced candidates, not across an endless field of noise. That is a much better problem to hand to a person or a higher-level workflow. You are comparing a handful of serious contenders rather than wading through search sprawl.
This design also pairs well with MCP clients. Context windows are finite, and too much search output can blur the very evidence you need to inspect. A concise, bounded candidate set is not just cleaner. It is more compatible with how agent workflows actually operate.
The role of Google Knowledge Graph, and its limits
Because the project explicitly supports an optional Google cross-check, it is worth being precise about what that adds and what it does not.
If your workflow uses MCP for google knowledge graph and wikidata together, the appeal is obvious. You have two external knowledge providers that can sometimes reinforce one another. The project documents exact ID joins through /m/ and /g/ mappings into Wikidata properties P646 and P2671. That is much better than vague label similarity. Exact identifier concordance is a concrete signal.
But the project is equally clear that this is concordance, not identity proof. That language matters a great deal in production settings. When two providers line up, you can treat the alignment as evidence that the same real-world subject is being referenced. You still should not skip due diligence, especially when your local source record is thin or internally inconsistent.
This is where many people overread MCP for google knowledge graph. They assume a cross-provider match settles the matter. It does not. It strengthens a case. It does not absolve you from reading the surrounding facts.
That becomes even more important with AMBIGUOUS. Sometimes Google and Wikidata may agree on one candidate’s external identifiers, which can break a tie. Sometimes no such exact join exists, and the ambiguity remains. Sometimes a local record is so sparse that even cross-provider support does not give enough confidence to act automatically. The value of the project is that it does not hide these distinctions.
When to inspect ranks, qualifiers, and references
There is a tendency in entity resolution work to stop at labels, aliases, and maybe a top-level description. For easy records, that is often enough. For the outcomes discussed here, it usually is not.
The ability to retrieve selected facts, and to include ranks, qualifiers, and references on request, is one of the most practical parts of this server. It lets you move from “the names look similar” to “the statements align in a way that is explainable.”
Suppose you are deciding between two plausible entities for a local record. The label alone may be useless. A qualifier attached to a role statement could separate them. A rank could tell you which statement is preferred versus deprecated or less favored. A reference can show whether a claim is documented in a way that supports your trust in using it for disambiguation.
You do not need every possible fact for every record. In fact, that would defeat the project’s bounded spirit. But for AMBIGUOUS, especially, targeted fact inspection is often the difference between safe delay and safe resolution.
For NO_CANDIDATE, the same detailed retrieval can help in a different way. It can confirm that near-miss candidates are genuinely near misses, not overlooked matches. That matters if you are deciding whether to retry with richer local metadata or escalate to manual review.
A practical reading pattern for teams
If you are introducing this resolver into a linking workflow, it helps to standardize how people read the outcomes. One useful pattern is this:
- Treat NO_CANDIDATE as a signal to improve the query or source record before trying to force a link.
- Treat AMBIGUOUS as a signal to compare candidate evidence and look for one missing discriminator.
- Treat Google agreement, when used, as corroboration rather than final authority.
- Export evidence when decisions need auditability or later review.
That may sound conservative, but conservatism is exactly what keeps identity systems usable over time. Once a bad link enters a batch process, it rarely stays isolated.
The CLI’s batch and evidence-export capabilities make this especially relevant. Batch linking can accelerate good workflows, but it also magnifies sloppy assumptions. Evidence export gives teams something much better than a gut feeling or an opaque score. It gives them a record of why a link was accepted, held, or refused.
What these outcomes do for governance
There is also a governance angle here that is easy to overlook.
This project is read-only. It does not edit Wikidata, Google, or user data. It is not official Wikimedia or Google software, and it is not an export of the Google Knowledge Graph. Those boundaries matter because they frame the server as an interpretation and linking layer, not a source of truth that can silently rewrite upstream systems.
In that context, NO_CANDIDATE and AMBIGUOUS become governance-friendly outcomes. They are explicit. They can be logged. They can be audited. They expose uncertainty instead of laundering it into a false-positive link.
For organizations using MCP for wikidata in production, that explicitness is often more valuable than a superficially higher match rate. A deterministic resolver that clearly says “not enough here” is easier to trust, easier to monitor, and easier to improve than one that returns a match for almost everything but cannot explain why.
Wikidata itself already has a broader MCP context through standardized tools for exploring and querying via the Wikidata API and Query Service. This project sits in a more specific niche. It is focused on practical search, selected fact reading, and record-to-QID resolution with inspectable evidence. That focus makes the meaning of its outcomes especially important.
Edge cases worth respecting
A common mistake with deterministic systems is to expect determinism to eliminate judgment. It does not. It just makes the boundary lines clearer.
There will always be edge cases. Some local records are too sparse to resolve. Some entity spaces are too crowded without another identifier. Some candidates will Wikidata MCP look compelling until one qualifier breaks the tie. Some will remain unresolved even after careful inspection.
In that sense, NO_CANDIDATE and AMBIGUOUS are not signs that the resolver failed to do its job. They are signs that it refused to do someone else’s job. A machine can bound, compare, and surface evidence. It cannot invent missing context from nowhere.
That restraint is what makes the project useful in serious workflows. It is also what makes MCP for google knowledge graph and wikidata a better phrase than a magical one. The value is in combined evidence, bounded retrieval, and explicit uncertainty. Not in pretending that every record can be neatly snapped into place.
The real test of a resolver
The best test of any resolution system is not how often it matches. It is how often you still trust the data six months later.
By that measure, NO_CANDIDATE and AMBIGUOUS are not secondary statuses. They are core features. They preserve the integrity of your links by making uncertainty visible and actionable. They encourage teams to gather better evidence, inspect facts carefully, and use cross-provider agreement responsibly. They also fit the reality of MCP client workflows, where too much noisy output is often worse than a clear refusal.
If you are working with MCP for Wikidata, or combining MCP for google knowledge graph with Wikidata in a practical linking pipeline, these two outcomes deserve close attention. They are the part of the system that says, with discipline, “not yet” or “not clearly enough.” In data work, that is often the most trustworthy answer you can get.