AI Generated Code Audit

n8n, Make.com and Zapier automation with AI steps and a human approval check — fixed price, documented on handover.

What Does AI Model Testing Cover?

AI model testing is the independent checking of what an AI system actually outputs before your users see it. EICRA builds a test set from your own real cases, runs it against your LLM, RAG or agent system, scores every output against a written rubric, and hands you the failures together with the test set.

Tested on cases you supply, at a model version you name
Reviewed by a person, not scored by a script alone
Every failure reproduced before it reaches the report
Test set handed over so your team can re-run it
45 minutes 01 Scoping call · test plan, free
2 days 02 Starter run · 100 cases, $1,200
1–2 weeks 03 Standard run · 300 cases, $2,400
3–5 days 04 Repeat run · new version, $900
Weekly 05 Monitoring · weekly re-run, $700/month

Which AI Testing Services Do We Run?

EICRA runs six AI testing services: output accuracy, RAG and grounding, prompt injection and adversarial, bias and consistency, version and drift, and an evidence pack for procurement. Each one answers a different question about your system, and each can be bought on its own or combined into a single run against one shared test set.

Output Accuracy Testing

You learn how often the system gives a wrong answer, and on which kinds of question. We build cases whose correct answer is already known, run them, and score every output against that reference. The report groups the wrong answers by cause, so your team fixes patterns rather than single tickets.

RAG and Grounding Evaluation

You learn whether your retrieval layer is actually feeding the answer. Each claim in an output is traced back to the passage that was retrieved for it. Where the passage does not support the claim, the case is logged as unsupported, and the retrieval gap behind it is named in the report.

Prompt Injection and Adversarial Testing

You learn whether a user can talk your system out of its own instructions. We write inputs designed to override the system prompt, leak it, or push the system outside its permitted scope, then record which ones succeed. Every successful attempt is re-run to confirm it repeats.

Bias and Consistency Testing

You learn whether the system treats similar cases the same way. The same question is put in different wordings, different names and different customer segments, and the answers are compared. Where the output changes without the substance changing, the pattern is recorded with the cases that produced it.

Version and Drift Comparison

You learn what a model, prompt or retrieval change actually did. The same test set and the same rubric are run against the old and new versions, then compared line by line. The report names what improved, what got worse, and which cases changed verdict between the two versions.

Evidence Pack for Procurement

You get written evidence a customer, partner or procurement reviewer can read. The pack sets out the scope, the test set, the rubric, the results and the limits of what was tested, so a reviewer can see what was measured rather than take a claim on trust, with the date of the run.

Not included: building or changing the system under test · retraining or fine-tuning models · certification or conformity assessment, because EICRA is not a notified body · AI model usage fees, paid by you to the provider.

How Do Tests Expose Model Failures?

A test exposes a failure by putting the system in front of an input whose correct answer is already known, then comparing what came back against that answer. The four failures below account for most of what reaches production:

  • Confident answers that are simply not true
  • Answers the retrieved source does not actually support
  • Inputs that talk the system out of its own instructions
  • Quality that slips after a model or prompt change

Failure Types Mapped to Test Methods

Failure type What it looks like Test method
Hallucination A confident answer that is factually wrong Known-answer cases scored against the reference
Unsupported grounding An answer the cited source does not contain Claim-to-source tracing on every retrieved passage
Prompt injection The system follows text instead of its instructions Adversarial inputs written to override the system prompt
Quality drift Answers get worse after an upgrade The same test set re-run against the new version

What Happens After You Enquire?

An enquiry moves through five stages, and you decide at the end of each one whether to continue. Nothing is scored until the rubric is agreed in writing, and nothing is reported until the failure has been reproduced.

Start at 01 if the system is live. Start at 02 if it is still in build.

Five Stages Ending in a Decision

Stage In numbers You give us You get Decide at the end
01 Scope 1 call, 45 minutes The system, its version and its use case A written scope and a test plan Go to the test set, or stop
02 Test set 100 to 300 cases Real cases, or permission to sample your traffic 4 groups: ordinary, edge, adversarial and policy cases Approve the set, or revise it
03 Rubric 6 verdicts Your definition of a correct answer A written rubric scoring correct, hallucinated, unsupported, tone, incomplete and policy Approve the rubric, or revise it
04 Run and review 1 pass per version Read-only access to the system under test Every input and output logged, every output scored, failures reproduced twice Receive the report, or extend the run
05 Handover 3 files Nothing new Findings report, run logs and the test set, for your own team Run it yourself, or book a repeat
More than one system, or several model versions? Business packages from $6,000 →

Repeat Runs After Every Change

Model providers change their models, and your prompts and retrieval sources change too. A repeat run re-uses the same test set and the same rubric against the new version, so the two results can be compared line by line. For the build side of the same work, see our AI services; for a wider roadmap, see our AI consulting services; for recording what happened after a live failure, see error and incident documentation.

Systems LLM assistants, RAG, agents
Access Read-only, in your environment
Contracts NDA and DPA before any access
Handover Report, run logs, test set

Why Choose EICRA for Your AI Testing Needs?

Five reasons, and each one is something you can check rather than a claim you have to trust. A person reads every output, not a script alone. The test set is yours at the end, so the work is never bought twice. We test systems we did not build. Pricing starts at one feature, not a full audit. And every report names what it did not cover. Open a tab for the detail behind each.

Reason one

A Person Reads Every Output

Automated scoring tells you a string did not match. It cannot tell you the answer was technically correct and commercially wrong, or right but in a tone that loses the customer. On every run here a reviewer reads each output against the rubric you signed off before scoring began. Scripts do the counting. The judgement stays human.

Every output read, not sampled
Rubric agreed before scoring
Tone and policy judged too
Two reviewers, same verdict
Scripts count, people judge
Disagreements logged, not buried
Reason two

The Test Set Stays With You

Most vendors keep the harness, so your next release means buying the same work again. We hand over the test set, the rubric and the run logs at close. Your own engineers can repeat the whole run after any prompt or model change, with us or without us. It is the one decision that stops this becoming a subscription.

Test set handed over at close
Rubric and run logs included
Re-run without calling us
No licence, no lock in
Repeat runs cost less
Capability stays in your team
Reason three

We Test What We Did Not Build

A builder marking its own work has an interest in a clean result, and a procurement reviewer spots that before reading a single number. EICRA runs tests on systems built by your team or another vendor. Where we did build the thing, the report says so on its first page and calls itself internal review rather than independent testing.

Tested by people who did not build
Stated in the report either way
Survives a procurement question
Internal review labelled as such
No stake in a clean result
Findings reported whoever built it
Reason four

Priced for One Feature

Published minimums for a single AI red team engagement start around sixteen thousand dollars. That is the right shape for a regulated audit and the wrong shape for checking one chatbot before launch. A starter run here is one system, one hundred cases, twelve hundred dollars, and you can stop when it ends.

Start with one feature
Fixed scope, fixed fee
Scoping call costs nothing
No minimum retainer
Stop after any stage
Below the published market floor
Reason five

The Report States Its Limits

Every report names what the run did not cover: inputs nobody tried, versions not tested, behaviour it cannot promise. That paragraph is what makes the rest believable. A reviewer who finds no limits section assumes something was hidden, and a finding you cannot defend in a meeting is worth nothing to the person holding it.

Not covered section in every report
Versions tested named exactly
No promise on untested inputs
Limits written before findings
Defensible in a review meeting
What we cannot do, said early
Human1Yours2Independent3Right Size4Limits5Why ChooseEICRA
Where we are not the answer: if you need a certificate rather than evidence, a notified body or a large assurance firm is the honest route and we will say so on the scoping call. We are a small team working remotely from Bangladesh, we hold no certification we have not been awarded, and we have no case studies to show yet. What we can show you is the method and a scoped test plan for your own system.

Proof You Get Before You Pay

We publish evidence rather than promises: a written test plan from the first call, a test set built from your own cases, and every finding traceable to a logged run. Client case studies with numbers will be added here as engagements complete, with each client permission.

100–300 Real cases every run is tested on, drawn from your own traffic and agreed in writing before a single output is scored. Included in every run
45 minutes One call turns your system and its use case into a written test plan, and the plan is yours whether or not you continue. Scoping call costs nothing
3 files Findings report, run logs and the test set itself, handed over at close so your own engineers can repeat the run. Included in every engagement
Free 45-minute scoping call Send one AI feature and the use case it serves. You get a written note on what a test set for it would contain and what a run would not tell you, whether or not you hire us.

Send one AI feature

Is It Safe to Outsource AI Model Testing?

Your data does not leave your systems. Runs execute in the environment you nominate, using accounts and keys you control, with read-only access for named reviewers. The sensitive artefact in AI model testing is not the model, it is the test set, because that set is built from your own real prompts and records and is effectively a copy of production traffic. That is where the controls sit, and it is also why Article 10 of Regulation (EU) 2024/1689, which governs data and data governance, applies to the set as much as to the model.

Contracts and data transfer

Which agreements are signed first?

NDA mutual, signed before your system, its prompts or its logs are touched.
DPA the processor terms required by Article 28(3) of the GDPR, wherever personal data appears in test cases.
Transfers standard contractual clauses or the transfer agreement your own law requires, signed before personal data moves.
Certifications we hold none we have not been awarded, which is why the controls above are contractual rather than badge based.
Data handling during a run

Where does your test data sit?

Location Runs execute in the environment you nominate, using keys and accounts you control.
Minimisation Cases can be masked or synthesised from real patterns where raw records must not leave your systems.
Retention Run logs, scored outputs and the test set are kept for the period written into the engagement letter, then deleted on confirmation.
Access Read-only rights, named reviewers only, and our logins removed at handover.

Six Signs Your AI Needs Testing

Most teams do not go looking for AI model testing. They run into one of these six situations and then start looking. Find the row that sounds like your week, and the last column tells you which service answers it, so you can start there rather than buying a full programme you do not need yet.

Signs and Where to Start

No. The sign Where to start
01 The feature is live and nobody has measured its answers Starter run, one hundred cases
02 A customer or procurement reviewer has asked for evidence Evidence pack for procurement
03 You changed the model or prompt with no before and after Version and drift comparison
04 Support keeps forwarding answers that were confidently wrong Output accuracy testing
05 Your retrieval source changed and grounding was not re-checked RAG and grounding evaluation
06 Testing today is a spreadsheet of prompts one person maintains Test set design and handover
None of these yet? Then you probably do not need us this quarter. The cheapest useful step is to write down what a correct answer looks like for your feature, because that document is what any test set is built on, whoever builds it.

Request a test plan

AI Model Testing Questions Buyers Ask

Want the stage sequence instead? See how a run works →

What does a test set actually contain?

A test set contains between one hundred and three hundred cases taken from your own traffic, split into four groups: ordinary queries the system meets every day, edge cases at the limits of its scope, adversarial inputs written to break its instructions, and policy-sensitive inputs where a wrong answer would cost money or trust. The split is agreed with you in writing before the first run.

Who is accountable if a tested system still fails?

You remain accountable for the system you operate. A test run measures behaviour on a defined set of cases at a named version; it does not transfer liability for the system, and no report from EICRA says otherwise. Where your obligations come from a regulation such as the EU AI Act, the duty sits with the provider or deployer of the system, and our report is evidence you can put against that duty rather than a substitute for it.

Which law and jurisdiction govern the engagement?

The engagement letter names the governing law and forum, and we work to whichever your legal team requires rather than imposing ours. EICRA Soft Limited is registered in Bangladesh and delivers remotely. Where personal data is involved, the data processing agreement follows Article 28(3) of the GDPR and the transfer mechanism your own regulator requires is signed before any data moves.

Where does our test data sit during a run?

Runs execute in the environment you nominate, using accounts and keys you control, with read-only access for named reviewers only. Where raw records must not leave your systems, cases can be masked or rebuilt synthetically from the same patterns. Run logs and the test set are retained for the period written into the engagement letter and deleted on your confirmation.

How long does a test run take?

It depends on how many cases the set holds, how much of the review is human rather than automated, and how quickly the rubric is agreed. A single feature on a vendor model settles faster than a fine-tuned system needing version comparison. We give a stage-by-stage schedule with the written scope, and we do not quote a duration before the scope exists.

What happens after we send an enquiry?

Five things happen, in this order:

1. We ask which system, which version and which use case, and put the answers in a written scope.
2. We agree the test set with you, including how the four case groups are split.
3. We agree the scoring rubric, so what counts as a failure is settled before anything is scored.
4. We run the set, log every input and output, score each output and reproduce each failure.
5. We hand over the findings report, the run logs and the test set itself.

Can you take over from our current testing provider?

Yes. Bring whatever exists: their test set, their rubric, their last report. We map what was covered against what was not, tell you plainly where the two disagree with our own reading, and then either extend the existing set or rebuild it. Where their set is sound we say so and reuse it rather than charging you to recreate work you already own.

Can you guarantee the model will not hallucinate again?

No, and any vendor who says otherwise is selling something that does not exist. A test run measures behaviour on the cases it ran, at the version it ran against. It cannot promise behaviour on inputs nobody has tried, and a model or prompt change can undo a good result. That is exactly why the test set is handed to you: so the same run can be repeated after every change instead of trusted once.

Reviewed by Mohammed Nazrul Islam
Tell us which system, which version and what it is used for. You get a written scope and a test plan back before anything is agreed.