Skip to main content
RealHumanTests

Independent human testing for customer-facing AI

Real people test your AI before your customers do.

We send real people to chat with, call and text the AI agent or chatbot you already run, or are about to launch. They act as your customers, push on the areas that worry you, and report exactly what it did, scored on a fixed human rubric.

We never answer your calls or chats, and we never sell the fix. We only test.

Audit panel
12 sessions
Turnaround
About 10 days
Rubric
8 criteria
Session 07 of 12Website chat

Persona: refund demand

Tester: I was charged twice for one order. I want the extra charge back today.

Bot: Happy to help! Refunds are available for 90 days and I have applied a 20% credit for the trouble.

Tester note: the refund policy page says 30 days, and nothing on the site offers a credit.

Got it rightflaggedStayed in boundsflagged
Illustrative example of a test session, not a real company's AI.

How an Audit works

From scope to verdict in about ten days

You tell us what worries you. We build a panel of real people around it, they test your AI as your customers would, and you get every transcript plus a plain verdict.

  1. Day 101

    Scope

    You tell us what your AI does, where customers reach it, and which areas worry you most.

  2. Day 1-202

    Authorize

    You sign a written authorization for the system you own, and every tester is under a consent and confidentiality contract.

  3. Day 2-303

    Build the panel

    We write customer personas and scenarios around your concerns and assign real testers to each one.

  4. Day 3-904

    Test

    Real people chat, call or text your AI as your customers would, and log what happened after every session.

  5. Day 1005

    Verdict

    You get every transcript, rubric scores per session, and a one-page verdict: keep it, fix these settings, or reconsider it.

See the full Audit process

The rubric

One fixed rubric, scored by the person who lived the conversation

Every session is scored 1 to 5 on the same eight criteria, each with a written reason. The last one is the question no automated tool can answer honestly: would a real customer have stayed?

The eight criterion RealHumanTests rubric
#CriterionThe question the tester answers
01Understood meDid it understand what I actually asked, including when I phrased it badly?
02Got it rightWas every fact, price and policy it gave me correct, with nothing made up?
03Got it doneDid I leave with my problem solved or my task completed?
04EffortHow many turns, repeats and rephrasings did it take?
05HandoffWhen I needed a person, could I reach one without a fight?
06ToneDid it stay patient and respectful when I was angry, confused or slow?
07Stayed in boundsDid it avoid promises, discounts or advice it had no authority to give?
08Would I stayfocusWould I have hung up, given up, or gone to a competitor?

Read how each criterion is scored

Who shows up

Testers who behave like your real customers

Angry, confused, off-topic, in a hurry, on a screen reader, speaking a second language. You choose the mix, or we match it to your customers.

  • Angry customer

    Pushes on: Tone, escalation, handoff

  • Confused first-timer

    Pushes on: Understanding, effort

  • Off-topic wanderer

    Pushes on: Staying in bounds, recovery

  • Refund or discount demand

    Pushes on: Policy accuracy, unauthorized promises

  • Edge case

    Pushes on: Unusual orders, accounts and dates

  • Asks for a human

    Pushes on: Handoff

  • Accessibility needs

    Pushes on: Screen readers, hearing, speech and cognitive load

  • Accent or non-native speaker

    Pushes on: Understanding, multilingual handling

Why people

A machine's report card on a machine is not enough

Automated simulators and vendor dashboards are useful for volume. But only a person can tell you whether they would have hung up, given up, or gone to a competitor, and why.

  • Real people, not simulated callers

    The human judgment of "would I have given up" is something an automated simulator cannot produce.

  • Diagnosis only

    We never sell the fix, so we have no reason to find problems that are not there or to hide the ones that are.

  • Neutral

    No vendor partnerships and no referral fees from AI platforms.

  • A fixed, published rubric

    Results are comparable across tests, months and products.

  • You choose what to push on

    Share your concerns and the panel is built around them.

  • A plain verdict

    Keep it, fix these settings, or reconsider it, backed by every transcript.

  • Authorized testing only

    Written authorization from the system owner and consented, contracted testers.

  • Fixed starting prices

    No retainer. The Audit starts from $349 and Watch from $99 per month.

What you receive

What a finished report contains

This is the format, shown blank. We have no invented scores to show you, and we never will.

The one-page verdict

Sample format, no real scores
System
Your support chatbot, website widget
Panel
12 sessions, 8 personas
Window
About ten days
Rubric score layout, blank in this sample
CriterionMean of 5
Understood meblank in this sample
Got it rightblank in this sample
Got it doneblank in this sample
Effortblank in this sample
Handoffblank in this sample
Toneblank in this sample
Stayed in boundsblank in this sample
Would I stayblank in this sample

Verdict, one of three

  • Keep it
  • Fix these settings
  • Reconsider it
  • Every transcript or recording

    Each session in full, with the tester's notes pinned to the exact moment something went right or wrong.

  • A score per criterion, per session

    Eight criteria scored 1 to 5 by the person who lived the conversation, each with a written reason.

  • What failed, named plainly

    When the verdict is fix these settings, it lists the behaviors testers saw fail, so whoever maintains your AI knows where to look.

  • No fix to sell

    We diagnose only. The report is yours to hand to your vendor, your developer or your own team.

Services and pricing

Fixed starting prices, no retainer

  • The Audit

    From $349

    A fixed panel of real people plays your customers over about ten days and hands you a plain verdict.

    More about The Audit
  • Watch

    From $99 per month

    A small monthly panel on the same rubric, so you see when your AI starts to drift.

    More about Watch
  • Custom Panels

    Quoted

    Larger or more specific panels for launches, languages and accessibility.

    More about Custom Panels
  • The Benchmark

    Free and public

    A public index scoring customer-facing AI products on the same fixed human rubric.

    More about The Benchmark

Questions

What people ask first

Do you answer our phones or run our chat support?

No. RealHumanTests never answers anyone's calls or chats. Real people test the AI you already run, or are about to launch, by acting as your customers, and then report exactly what it did.

Will you fix what you find?

No, and that is on purpose. We only diagnose. Because we never sell the fix, we have no reason to overstate a problem or to go easy on one, so you can trust the verdict and hand it to whoever maintains your AI.

Why use real people when automated testing tools exist?

Automated simulators are useful for volume, but a machine grading a machine cannot tell you whether a real customer would have hung up, given up, or gone to a competitor. Our testers are people with real accents, real impatience and real confusion, and each score comes with a written reason.

Can we tell you what to test?

Yes. Share your concerns, such as refund requests, a new product line, or callers who ask for a person, and we build the personas and scenarios around them. Every session is still scored on the same fixed rubric so results stay comparable.

Read every answer in the FAQ

See your AI the way your customers do

Tell us what your AI does and what worries you. We build a panel of real people around it and hand you every transcript, a score per criterion and a plain verdict. We never sell the fix.