PRAPI Research · 2026-10-08

How Operators Stop LLMs From Inventing Numbers in Production

Large language models still invent numbers in customer-facing output, and most teams find out only after a wrong figure has shipped. We asked operators who run LLM features in production for the one concrete mechanism that actually stopped it. Their answers converge on a single rule: a model may phrase an answer, but it may not be the source of a fact.

12 contributors cited


Every operator we heard from described the same failure in almost the same words. A fluent model states a number that was never real, it reads well, and nobody catches it until a customer does. The fixes they shipped are not clever prompts. They are structural rules about where a number is allowed to come from.

Pattern 1: No number without a source it can point to

The most common mechanism was refusing to let any figure reach a customer unless it traces to a verifiable source. Marat Shigapov of Cloud Cube put it plainly:

no claim without a page. Every statement in the report must have page number and the quote from client's document. If the model can not show the quote, the only allowed answer is "Not found", and it is a normal status, same as "Stated".

Rahul Agrawal of QuickIntell, building for healthcare revenue-cycle work, drew the same line:

a number isn't allowed into customer-facing output unless it is tied to a source the system can point at.

And Nick Sawinyh of Gunsnation enforces it at write time:

no number ships unless it traces to a source file. The model never gets to define its own test.

Pattern 2: Take number-generation away from the model entirely

Several operators stopped the problem upstream by never letting the model produce a figure at all. Felix Römer of SmartKeys described the cheapest, dumbest gate that caught the most:

The mechanism that worked was not letting the model produce numbers at all. In the AI support workflow at the online retailer where I run e-commerce and IT, every number a customer sees, order status, tracking link, invoice amount, delivery window, comes from a lookup against the ERP, inserted into the reply as a field. The model phrases the answer. It never computes or recalls one. If the lookup returns nothing, the field is empty and the reply cannot be sent

Jimi Patel of eStore Factory does it with keyed facts:

every number in our customer-facing output has to carry a provenance key, or it does not render.

Scott Stouffer of Market Brew draws the boundary in application code:

I draw a hard boundary around measured numbers. Our audit-email renderer formats numerical fields from the audit evidence in application code and returns before the email-writing LLM is called. When an audit-specific draft is required without usable evidence, generation fails.

And Deian Isac of Takibi Base keeps generation away from retrieval completely:

the knowledge layer never generates the answer. Agents ask through the Takibi CLI or API and get exact cited passages from your documents, plus an evidence support score. No paraphrase. No "helpful" summary that quietly invents a figure. If the source does not say it, Takibi does not return it.

Pattern 3: A deterministic gate that audits every figure

Where the model does draft freely, operators bolt a verification step behind it. Sudhanshu Dubey of Errna framed the shape of it:

decoupling the creative generation from a deterministic verification gate that acts as a rigid auditor.

Aleksa Baburska of Devox Software runs the same idea as a second pass:

we never treat the LLM output as the source of truth. It does create numbers and it can't be totally fixed (for ever) unless it becomes a RAG. For us, the most effective mechanism was a source-bound verification gate.

And Ritwick Dey of Panto AI built it out end to end:

a post-generation numeric claim-checker that (1) extracts all numeric, date, and named-entity claims from the LLM output using deterministic parsing (regex + AST for structured responses), (2) looks up each claim against authoritative sources

Pattern 4: When the system is unsure, a person decides

The last pattern is about what happens at the edge. Dinesh Goel of Robylon blocks rather than guesses:

The agent may not state a number unless that number came from a tool call in the same turn. Every figure in a reply is checked against the order system, policy doc, or ticket data it came from. If a number has no source, the reply is blocked.

Shane Larrabee of FatLab Web Support keeps a human on every outbound message:

no AI-written message reaches a client until a person has approved it, and the name on the message says who did the work.

Where it still leaks

The operators were candid about the gaps. The hardest failure is not an invented number but a correct citation of bad data. Rahul Agrawal of QuickIntell warned:

A stale coverage date from a payer portal will pass with high confidence and a clean citation.

A source-binding gate proves a figure came from somewhere; it cannot prove the somewhere was right, or that the model read it correctly. Stale lookups, misread documents, and unit mismatches all survive a clean provenance check.

The through-line

Across every answer, one rule held: the model may phrase, but it may not be the source of a fact. Refuse, route to a person, or leave the field blank before you let it guess.

Contributors

  1. We made a service which reads software development proposals and shows to client what is inside and what is missing. One invented price or term, and the whole report is garbage. Our mechanism: no claim without a page. Every statement in the report must have page number and the quote from client's document. If the model can not show the quote, the only allowed answer is "Not found", and it is a normal status, same as "Stated". The model invents when you force it to answer. What it misses: the page proves that the quote exists, not that the model understood it right. One line "testing is included" can become "Stated" when nothing concrete is promised. It is not invented number, but for the client it is the same wrong report. Marat Shigapov / Founder & CEO / CloudCube IT
  2. The mechanism that works for us is field-level provenance with a confidence gate: a number isn't allowed into customer-facing output unless it is tied to a source the system can point at. QuickIntell builds AI agents for healthcare revenue-cycle work, where an invented number is a wrong patient balance, a wrong coverage date or a wrong authorization reference. Every value an agent uses carries where it came from (the payer document, portal screen, call transcript or system record), a timestamp and a confidence score. Output is assembled from those fields, not generated freely by the model. If a field has no source, the output says it's missing rather than filling the gap. Below the confidence threshold, the case goes to a person with the evidence attached, and high-impact updates to the practice's billing or clinical system wait for staff approval. What it caught: in call-handling, the model would happily 'confirm' details a patient hadn't actually given, such as an appointment time, because the transcript was ambiguous. The provenance check flags those as unsourced, and they route to review instead of being written back. What it missed: it doesn't protect against a source that is itself wrong. A stale coverage date from a payer portal will pass with high confidence and a clean citation. We've learned that a correct citation of bad data is the harder failure, and the only fix we have is treating conflicting sources as an exception rather than letting the newest one win. The cost of fabrication in our world is rework and a patient who was told the wrong thing, which is why we prefer 'unknown, routed to a person' over a confident guess. Rahul Agrawal Founder & CEO, QuickIntell
  3. The mechanism that worked was not letting the model produce numbers at all. In the AI support workflow at the online retailer where I run e-commerce and IT, every number a customer sees, order status, tracking link, invoice amount, delivery window, comes from a lookup against the ERP, inserted into the reply as a field. The model phrases the answer. It never computes or recalls one. If the lookup returns nothing, the field is empty and the reply cannot be sent, which is a cheap, dumb gate that has caught more than any clever one. Second gate: a confidence score on every draft. Below the threshold, or whenever the customer's reply contains something the system cannot place, the ticket goes to a person. What got through anyway: the model paraphrased a lookup result. The ERP said "4 to 7 days", the reply said "by Thursday", which sounded helpful and was wrong. No number was invented, a meaning was. It cost one annoyed customer and an afternoon of adding a rule that date fields are quoted verbatim, never rephrased. The gap I still have: a wrong answer that is confidently phrased and fully grounded in a stale source. The score cannot see that the source is the problem. Humans catch those, which is why they still read a sample every week.
  4. One mechanism: every number in our customer-facing output has to carry a provenance key, or it does not render. In SellerQI the model never writes figures. The pipeline resolves values first from SP-API and the Advertising API: ad spend, ACOS, sessions, unit session percentage, FBA fee per unit, aged inventory counts. Each is stored as a keyed fact. The model gets template slots, not a free text field, and writes prose around placeholders. A post-generation gate then scans the finished string for any numeral, percentage or currency symbol that did not come from a resolved key. If it finds one, the output is rejected, regenerated, and the rejection is logged. What it caught: invented ACOS values, fee figures that looked plausible but sat a tier off the real one, and date ranges the model quietly widened from 30 days to 90. What it missed: everything wrong without being numeric. The model would report the correct ad spend and then attach the wrong cause, saying a listing lost sales to suppression when it was actually out of stock. Right number, wrong story. The fix there was narrower: for causal claims the model picks from a fixed list of diagnoses, each tied to a condition the pipeline can actually test. One got through early on. A summary called a reimbursement recoverable when the FBA claim window for that discrepancy had already closed. The number was real, just stale. The cost was credibility with that seller, which is the expensive kind of error. We now stamp every fact with the timestamp of the report it came from and refuse to render anything past its freshness window.
  5. I built Takibi Base because I got tired of agents inventing numbers. The mechanism is simple: the knowledge layer never generates the answer. Agents ask through the Takibi CLI or API and get exact cited passages from your documents, plus an evidence support score. No paraphrase. No "helpful" summary that quietly invents a figure. If the source does not say it, Takibi does not return it. That is the grounding step. Claim-check is the citation itself — every returned passage points at the source. When the score is weak, the agent should refuse or escalate instead of guessing. [THEIR STORY: one production miss that still got through, what it cost, and the gate you added after.] The lesson I keep: do not ask an LLM to be honest about facts it does not have. Retrieve the evidence. Cite it. Leave generation for prose that does not need a number.
  6. At Market Brew, I draw a hard boundary around measured numbers. Our audit-email renderer formats numerical fields from the audit evidence in application code and returns before the email-writing LLM is called. When an audit-specific draft is required without usable evidence, generation fails. The output is marked for human review. A regression test makes every LLM request fail and checks that the percentages match the supplied source fields. This verifies the composition boundary; it does not prove the upstream data is correct. The gap showed up in a separate customer-facing ASK widget. A customer reviewing stored conversations found a question about one chapter answered with results from the previous chapter, plus a claim that no matches existed for the requested chapter. The customer's site did contain relevant commentary. Restricting output to source references or a no-match response had worked in customer testing, but it still allowed a false statement about available content. Our support team investigated the mismatch; the records do not establish its root cause or a financial loss. The documented consequence was an incorrect answer and follow-up investigation. Another complaint in the same review concerned empty "coming soon" references; we subsequently added a source-text fallback. That change addressed the empty-reference case, not proof that the wrong-chapter case was solved. My takeaway is to validate measured numbers in code, then separately test citation relevance and abstentions. "No matches" is a factual claim that also needs verification. Scott Stouffer Co-Founder + CTO, Market Brew
  7. The mechanism that's worked for us is boring: no AI-written message reaches a client until a person has approved it, and the name on the message says who did the work. We use AI on our support tickets. It reads the request, works out which site it's about, investigates read-only, and often drafts the reply. Apart from one short automatic acknowledgment, everything a client receives is approved by a person first. If the AI did most of the work, the message goes out under the AI's name, not mine. That's deliberate. It tells the client who handled it, and it gives us a record, so if something goes wrong we can look back at any ticket and see exactly who took charge. Our miss happened on the inside, not with a client. AI had changed a checkout form on a client site and left our team a note to test it, complete with standard test card numbers. The instructions were specific and confident. The payment gateway was in live mode, so those numbers could never work, and a team member who doesn't usually handle payments lost a significant amount of time trying to figure out why. The AI didn't invent anything outright. It just didn't know the state of the system, and nobody questioned the assumption. What we changed: AI output inside our ticket system now comes from its own dedicated, limited user, so anyone reading a note knows it came from AI and should ask what the AI couldn't see. The approval gate protects what clients read. Labeling protects what our own team reads, and that turned out to need the same care.
  8. At Robylon we build AI customer-support agents, so a made-up number is our worst failure: a refund amount, an order date, a delivery window. The mechanism that helped most is simple. The agent may not state a number unless that number came from a tool call in the same turn. Every figure in a reply is checked against the order system, policy doc, or ticket data it came from. If a number has no source, the reply is blocked. It is either rewritten without the number or handed to a human agent. What it caught: confident guesses on refund timelines and plan prices, the kind of detail a model half remembers from an old help article. What it missed: numbers that were real but stale. If a brand had not updated its policy page, the check passed, because the source itself was wrong. The fix there was not more AI. It was one owner for the knowledge base and a review date on every policy article. My takeaway: grounding stops invention, but it cannot fix a bad source. You need a claim check and a person who owns the truth. Dinesh Goel, Founder, Robylon (robylon.ai), AI customer-support agents.
  9. I'm Aleksa Baburska, Director of Solution Acceleration at Devox Software, where I lead dev teams building customer-facing platforms and AI-enabled features. Concerning your request I can say first of all we never treat the LLM output as the source of truth. It does create numbers and it can't be totally fixed (for ever) unless it becomes a RAG. For us, the most effective mechanism was a source-bound verification gate. The model drafts an answer, but any specifics like numbers, names, prices, and so on is additionally checked against retrieved source data before reaching the customer. It actually add a second step to the verification process. A second step extracts every factual claim and compares it against approved sources. They could be CRM records or product databases. If a claim can't be matched, it is removed or escalated for human review. Despite it's the most obvious methods, there are others. Some hallucinations are plausible, only slightly off but still present the significant reputational risk. Additionally, this is the same reason they are harder to detect because they sound operationally normal. In this case, a verification layer can't cope. Instead, we add the rule so if a number cannot be traced to a specific record, it's flagged. Happy to provide a follow-up quote or expand on the verification workflow if useful. Best, Aleksa Baburska Director of Solution Acceleration, Devox Software
  10. Stop LLM numerical hallucinations by decoupling the creative generation from a deterministic verification gate that acts as a rigid auditor. In high-stakes production environments, prompting and basic RAG are insufficient; you need a post-processing filter that extracts every numerical value and cross-references it against the raw structured source data. If a model outputs "179.10," the gate must verify that this specific sequence exists in the source or is the product of a pre-defined calculation script. This specifically targets "hallucination by interpolation," where models attempt to calculate figures like discounted prices on the fly-often producing numbers that look plausible but do not exist in the database. The critical failure mode for these gates is unit mismatch. We encountered a deployment where the gate allowed a response because the number "500" appeared in both the context and the output. The source referred to 500 units of inventory, but the LLM confidently presented it to the customer as a 500-dollar price point. This led to a surge in support tickets from customers trying to hold the company to a volume count as a price. The takeaway is that grounding must be entity-aware. You must validate not just the digits, but the associated metadata and units. Moving from probabilistic generation to deterministic validation is the only way to build genuine enterprise trust.
  11. A rule, enforced at write time: no number ships unless it traces to a source file. The model never gets to define its own test. Context. I build data infrastructure and developer tools, and LLMs draft a lot of the surrounding text in that work: docs, summaries, release notes. Anything that leaves the building is customer-facing in practice, and a wrong figure in a public doc is worse than a missing one. So the pipeline has a check that runs on every write to a document. Any specific figure, percentage, dollar amount, rate, or audience size that isn't in the verified corpus gets flagged. The corpus is a file I maintain by hand. If a claim needs a number that isn't in the file, the claim is cut rather than softened. What it catches: the plausible rounding-up a model does when a sentence wants a statistic. A round percentage of teams, an audience size, a growth rate, none of them sourced. Those get flagged, and the rule has changed how the drafts get requested, since there's no point asking for a number the check will cut. What it missed, and what it cost: a fact, not a number. In a code comment, a model attributed a feed-validation tool to a third party that hadn't built it. The comment read as authoritative, a later documentation pass copied it, and the documentation went live. Nobody checked a sentence that had the shape of source code. The fix was a correction on a live site and a new rule: no prose comments in generated code at all, because a confident comment is a source nobody verifies. The general lesson is that a numeric guardrail catches numeric fabrication and nothing else. Names, attributions, and citations need the same treatment: a claim is allowed only if it points at something checkable. The single highest-yield habit I've found is to verify that a cited thing exists before evaluating what it says. It takes seconds, and the inversion catches a class of error that reading carefully never will, because the fabricated version is written to read well.
  12. Hi Tom, My name is Ritwick Dey, Co-Founder and CTO of Panto AI Inc. I’ve been featured on Tech Crunch, Stackoverflow and CTO Club. I am an active Open Source Contributor with my plugin having 70 Million Installs. Here’s my Input on your recent help a Journalist query: Preventing LLMs from inventing numbers requires an enforceable canonicalization + verification pipeline tied to authoritative data. We implemented one concrete mechanism: a post-generation numeric claim-checker that (1) extracts all numeric, date, and named-entity claims from the LLM output using deterministic parsing (regex + AST for structured responses), (2) looks up each claim against authoritative sources (live billing API, user profile DB, policy docs) with provenance, and (3) applies hard rules to accept/reject (exact-match for balances/IDs, tolerance thresholds for amounts/percentages, and a required source link for any policy citation). Claim-check / grounding: every number must map to a canonical record (customer*id → ledger*balance) returned with timestamped provenance. If no canonical match exists, the response is marked “unsupported.” We also surface model token-level confidence; low log-prob for a numeric token increases scrutiny. Verification gate & human-review: any mismatch beyond configured thresholds (>$5 or >1% for monetary values, or any unlinked policy claim) blocks publishing and routes to a human-review queue with a 2-hour SLA for critical flows. What it caught and missed: the gate stopped an LLM-generated $20 rounding error from going to customers. It missed a fabricated policy phrase (“30‑day prorated fee”) when our document cache was stale; that slipped into a support reply and cost an enterprise renewal (~$150k revenue impact) plus three days of engineering/legal triage. Fixes: canonicalize live reads (avoid stale caches for policy), tighten thresholds, and expand provenance requirements. The practical trade-off is latency vs. safety; for customer-facing financial claims, favor safety. Cheers! Ritwick Dey Co-Founder and CTO at Panto AI www.getpanto.ai LinkedIn: https://www.linkedin.com/in/ritwickdey/ Headshot: https://pantopublic.blob.core.windows.net/pubicdata/photo/ritwick_dey.jpg Opensource: https://marketplace.visualstudio.com/items?itemName=ritwickdey.LiveServer

This report is drawn from a PRAPI reverse-pitch research brief distributed to operators building customer-facing LLM features. Contributors submitted through MentionMatch, Connectively, and PRAPI's own intake form in response to a single question: what concrete mechanism stopped your model from inventing numbers in production, what it caught, and what it missed. Submissions were screened with an automated AI-authorship filter and scored for fit and editorial quality; quotes are published verbatim, with each contributor's consent to named citation.

Research conducted by PRAPI. PR system for founders running portfolios. Try PRAPI →