The short answer

Ask what they have running in production, how they evaluate whether it works, what happens to your data, and who owns the code. The strongest single signal is whether they can describe a project that went badly and what they changed afterwards. Firms that have only ever succeeded have not shipped much.

We are an AI development company, so read this with the appropriate scepticism. It is written the way we would want to be evaluated, which includes the questions we find uncomfortable.

Production experience

1. What have you built that is running in production right now, and what does it do all day?

Good: specifics. What it does, how much traffic, what broke, what it costs to run. Red flag: demos, pilots and proofs of concept presented as production systems. The gap between a demo and a system that has survived a year of real users is the entire skill.

2. What percentage of your projects reached production?

Good: an honest number below 100%, with the reasons. Red flag: "all of them". Industry pilot-to-production rates are not close to 100% for anyone, and claiming otherwise means they are counting differently or not counting.

3. Tell me about a project that went badly.

Good: a real story with a diagnosis and a specific change to how they work. Red flag: a humblebrag about caring too much, or an inability to name one. This is the highest-information question on the list.

Engineering practice

4. How do you evaluate whether an AI system is working?

Good: fixed evaluation sets built from real cases, run on every change, with retrieval measured separately from generation. Red flag: "we test it thoroughly" or a demo as evidence. Without evals nobody knows whether a change improved anything, including them.

5. What happens when the model gives a wrong answer?

Good: a designed answer covering grounding, citations, confidence thresholds, escalation and monitoring. Red flag: "modern models rarely hallucinate". They do, and a partner who will not say so has not run one in front of real users.

6. Which model do you use, and what happens when it is deprecated?

Good: an abstraction layer, per-task routing, and a migration process that starts with re-running evals. Red flag: a single vendor hard-wired throughout, or surprise that deprecation is a thing.

7. How do you handle our data?

Good: clear statements on what leaves your infrastructure, whether anything is used for training (it should not be), where it is processed, and how residency requirements are met. Red flag: vagueness, or having to check.

Commercials

8. How do you price, and what does the quote exclude?

Good: a scoping phase producing a fixed quote, with running costs stated separately and honestly. Red flag: a firm number for a complex build before anyone has looked at your data, or a build quote with no mention of ongoing cost.

9. Who owns the code and the models?

Good: you do, unambiguously, in writing. Red flag: a licence back to their platform, or ownership contingent on a support contract. Ask specifically about prompts, evaluation sets and fine-tuned weights, which are the parts people forget to name.

10. What happens after launch?

Good: a support model with named response times, a monitoring plan, and a stated cost. Red flag: "we are always here" with nothing written down. Most AI systems need attention within the first month; find out now what that costs.

Fit

11. Who exactly will work on this?

Good: named people, their experience, and how much of their time you get. Red flag: senior people in the pitch who vanish at kickoff. Ask directly whether the person answering will be on the build.

12. What would make you tell us not to do this?

Good: a clear answer. Data not ready, the process not stable enough to automate, the ROI not there, a product that already does it for a fraction of the price. Red flag: enthusiasm for everything. A partner who has never talked a client out of a project is selling capacity, not judgement.

Agency, company or consultancy?

Worth saying plainly: the labels mean nothing. "AI development company", "AI development agency", "AI consultancy" and "machine learning development company" are used interchangeably, and the word chosen tells you about their marketing, not their engineering. What distinguishes firms is whether they build and run systems or only advise, whether the people who pitch are the people who build, and whether they have production scar tissue. Ask about those; ignore the noun.

How to run the evaluation

  1. Shortlist three. More than that and you are comparing sales processes, not capability.
  2. Give all three the same brief, including your constraints and your data problems.
  3. Ask the twelve questions above, in a live conversation rather than by email.
  4. Buy a small paid scoping engagement from your top choice before committing to a build. It is the cheapest possible test of how someone thinks, and if their scoping concludes the project is not worth doing, you have saved far more than it cost.

Our own answers are on the AI development page, and our shortlist criteria are set out in top AI development companies in 2026. If you want to put us through the twelve questions, book a call.

Common questions

What should I look for in an AI development company?

Production systems rather than demos, a real evaluation practice with fixed test sets, clear data handling commitments, unambiguous code ownership, and a stated post-launch support model with costs. The strongest single signal is whether they can describe a project that went badly and what they changed as a result.

What are the red flags when hiring an AI development company?

A claimed 100% success rate, a firm price for a complex build before anyone has examined your data, enthusiasm for every idea you raise, a single model provider hard-wired through the architecture, vagueness about whether your data is used for training, and senior people in the pitch who do not appear at kickoff.

Is there a difference between an AI development company, agency and consultancy?

Not in practice. The terms are used interchangeably and describe marketing preference rather than capability. What actually differs is whether the firm builds and operates systems or only advises, whether the people who pitch are the people who build, and how much production experience they have.

Should I pay for a scoping phase before committing to a build?

Yes. A short paid scoping engagement that examines your data and systems is the cheapest way to test how a partner thinks before committing five or six figures. Good firms credit the fee toward the build, and a scoping phase that concludes the project is not worth doing has saved you far more than it cost.

Related reading