Independent directory.  Not affiliated with or endorsed by Anthropic. “Claude” and “Claude Certified Architect” are trademarks of Anthropic.
← All articles

Hiring an AI specialist

How to vet an AI consultant

Four checks, and together they take about an hour. Verify the credential at its source rather than accepting a screenshot. Ask what broke on the last project and listen to how specific the answer is. Read the proposal for its assumptions rather than its total. And give the same brief to two people, because where they disagree is where the real information is.

The short version

A credential proves someone passed an exam on a date; it is a floor, not a recommendation. What separates people who have delivered from people who have studied is that the first group can describe a failure in detail. Everything else — the deck, the confidence, the terminology — is available to anyone who has read enough.

How do I verify a Claude certification?

Ask for the verification link, not a screenshot or a PDF. Anthropic issues four Claude certifications through proctored, identity-verified exams and badges them on Credly, and every badge has a public page showing the holder’s name, the exact credential, and the dates it runs between. Anyone who genuinely holds one can send that link in under a minute.

Then check three things on the page itself: that the name matches the person in front of you, that the credential is the one they claimed rather than a different tier, and that it has not expired. Certifications run twelve months. A badge awarded eighteen months ago is a historical fact, not a current qualification, and someone presenting it as current has told you something.

One thing that is not a red flag: the email on the badge differing from their business address. People routinely sit exams under a personal account. It warrants a question, not a rejection.

What should I ask about their previous work?

One question does most of the work: tell me about a project that went wrong, and what you changed afterwards.

Anyone who has actually deployed something has a ready answer, and it is usually specific and slightly boring — the integration that took three times longer than quoted, the accuracy that held in testing and collapsed on real data, the client who would not release the sample they needed. The specificity is the signal. People who have only studied answer this abstractly, in terms of general lessons rather than particular Tuesdays.

Follow it with: what did you do differently on the next one? A practitioner who has iterated will describe a process change — they now insist on a data sample before quoting, or they run a parallel period as standard. That is what experience actually looks like from the outside.

What does a good proposal tell me?

Read it for assumptions rather than the number at the bottom. Every proposal contains guesses about your integration difficulty, your exception rate and your data quality, and those guesses are where the price comes from. A proposal that states them explicitly is written by someone who knows they are guesses.

Specifically, look for:

A confident fixed price arriving without any questions is the single most common warning sign in this market. It means either heavy padding to cover unknown risk, or a conversation in month two about why the scope has changed.

Should I ask for references?

Yes, and ask the referee a better question than whether they were happy. Nobody supplies an unhappy reference, so satisfaction tells you nothing. Ask what surprised them, what took longer than expected, and whether the system is still in use.

That last one is the most revealing question available and almost nobody asks it. A deployment still running a year later has survived model changes, staff turnover and the initial novelty wearing off. One quietly abandoned after four months looked exactly the same at handover.

How do I judge technical competence if I'm not technical?

You do not need to assess the architecture. You need to assess whether they can explain it, and whether the explanation stays consistent when you push.

Ask them to describe the approach as they would to your finance director. Someone with a real command of it will simplify without becoming vague, and will happily say “I don’t know yet, that depends on what we find”. Someone without it reaches for jargon under pressure, because specificity is where the gap would show.

Two more that work regardless of your own expertise: ask what they would do if the approach did not work, and ask what they would need from your team. A practitioner who has delivered knows exactly how much of your people’s time they will need and will tell you unprompted, because underestimating it has burned them before.

Is a small firm safer than an individual?

Not inherently. A firm gives you cover if someone is ill and more than one pair of hands at delivery. An individual gives you the person who does the work, usually moves faster, and cannot quietly substitute a junior after the pitch — which is a real risk with larger suppliers and worth asking about explicitly: who exactly will do this work, and will they be the person I have been speaking to?

What makes either safe is the same: a verifiable credential, evidence of something delivered and still running, and a first engagement small enough that being wrong is survivable.

What is the single best thing I can do?

Give the same brief to two or three people and compare how they read it, not what they charge.

One will propose a rules engine with a model handling exceptions. Another will propose an agent with a human approval step. A third will say the upstream data problem should be fixed first and the rest becomes easy. That disagreement maps the actual solution space, which no single proposal can do — and it tells you which of them understood your problem rather than recognising a pattern.

It also costs you nothing but the time to send the same document twice.

What if they have no Claude certification?

It is not disqualifying, and treating it as such would rule out people who were building production systems before the certification existed. What it means is you have lost one objective signal and need more of the others: work you can look at, references you can call, a public track record.

The question worth asking is why not, and the answer is usually revealing. “I have been doing this since before it launched and my clients judge me on delivery” is a fine answer from someone with a portfolio. “Certifications don’t mean anything” from someone with no demonstrable work is a different answer entirely.

How much should the first engagement be worth?

Small enough that being wrong is annoying rather than damaging. A scoped discovery, or a single narrow deliverable, is a far better test than any interview — you learn how they communicate when something is harder than expected, whether they raise problems early or late, and whether the written work is any good.

Treat it explicitly as a trial for both sides. Good practitioners are also assessing whether you are a client worth having, and one who says so directly is usually one who has been burned by a chaotic engagement and has learned to check.

Does it matter where they are based?

Less than timezone overlap and more than most people assume for regulated work. For a straightforward automation, a practitioner four hours away who is genuinely available for the meetings that matter is fine.

Where location does matter is data residency and applicable law. If your data cannot leave a jurisdiction, that constrains where processing happens and sometimes where the practitioner can work from. Raise it in the brief rather than discovering it during onboarding.