Docs / KaizenFlow AI

Answers you can check.

Most AI assistants will answer anything. KaizenFlow AI is built so that when it cannot back a number with your data, it says so instead of guessing. Here is every check an answer passes, where that rule is not enforced yet, and how to audit an answer yourself.

Updated Sep 26, 2026Reading time 17 minAudience Skeptics and finance reviewers

Why an AI answer needs a paper trail

An assistant that answers anything, whether or not data backs the answer, is worse than no assistant on a plant floor. KaizenFlow AI is built around a different rule: when it cannot back a number, it should say so instead of showing one. This page explains, check by check, how that rule is enforced, and where it is not enforced yet.

On a plant floor, a confident wrong number costs hours. A supervisor sends a crew to the wrong press. A shift lead chases a "drift" that is really the normal difference between two machines. Each miss teaches the team to stop asking.

In a finance review it costs more. One inflated projection, found late, discredits every figure before it, and the improvement program loses its budget case.

This is not hypothetical. While building KaizenFlow we caught projections that multiplied a running production total instead of an hourly rate, inflating output roughly twelvefold, and a card priced off an averaged machine state code. Both are fixed, and both became rules described below.

That is why we say the unit of trust is a verified dollar. A saving is a claim until it is measured. An answer is a claim until you can see what it was built from. Every check here keeps claims and proof apart.

Two paths, one standard

KaizenFlow AI shows up in two places. Ask Anything and AI Chat answer your questions: one assistant that queries your plant's records directly. AI Suggestions proposes improvements, drafted by up to nine specialist analysts (downtime, quality, capacity, energy, safety, supply chain, workforce, sustainability and vision). The specialists work only on suggestions, not on your questions. The checks differ between the two paths, so this page labels which path each one guards.

The checks every answer passes

An answer is shaped by scope, data and math rules before you see it. A suggestion passes a write side (saved only if it passes) and a read side (checked again when the board shows it). A saving should not count until an actual result is on record, and the audit steps below show how to check that it is.

CheckPathWhat it checksWhat you see if it trips
Sign-in scopeAnswersCompany comes from your login; queries of plant records are limited to your company, starting from the selected plantYour company's records. The assistant can also read your other plants, for example to compare them.
Empty-data ruleAnswersTested before release: a tool with no records must return nothing or say it has no data (known exceptions are listed below)A message such as "No data found for this metric and period", or what to connect
Honest mathAnswersCounters become hourly rates, state codes become shares of readings, spread is described per machineA statement that a rate cannot be computed yet, instead of a raw running total
Named basisAnswersDollar figures use your plant's unit economics or a disclosed defaultA note that the figure is a planning estimate, not a measurement
Chart honestyAnswers (Ask Anything)The assistant is instructed to plot the full window, match titles to the data and invent no pointsA chart with no data points is not drawn
Schema and quality gateSuggestionsComplete card, evidence with numbers, actionable steps, enough depthCard not saved; if none pass, "Analysis finished, but no suggestions passed the quality gate"
Numeric sanitySuggestionsFigures agree with their own stated math and with physical limitsCorrected figure, or the impossible sentence removed
Plant cross-checkSuggestionsStated throughput and reject levels against your measured tilesA sentence off by an order of magnitude is removed
Feed filtersSuggestionsLow known confidence, cards priced at exactly $0Card hidden from the board
No double countingSuggestionsDuplicate opportunities, rejected cards in totalsOne card per matched opportunity; board totals count each saving once
People and measurementSuggestionsChanges made with the status control are recorded; a card can be marked verified without an actual resultThe savings tool keeps marked, measured and verified apart; check that a verified card carries an actual result
ANSWERS: ASK ANYTHING AND AI CHAT A TOOL WITH NO DATA IS BUILT TO SAY SO, NOT TO FILL IN A NUMBER SUGGESTIONS: AI SUGGESTIONS BOARD WRITE SIDE: SAVED ONLY IF IT PASSES READ SIDE: CHECKED AGAIN WHEN THE BOARD SHOWS IT COUNT A SAVING ONLY WHEN AN ACTUAL RESULT IS ON RECORD Sign-in scope Plant's own data Honest math Named basis Answer Specialists draft Merge and dedupe Schema check Quality gate Numeric sanity Plant cross-check Feed filters No double count People decide Verified company and plant from your login tools query your records rates, state mix, per machine your economics or a disclosed default text, tools used, charts in Ask up to nine domains, run on your data one card per opportunity complete cards only evidence, depth and steps fix own math, drop impossible claims against measured plant tiles confidence floor, no $0 cards each saving summed once review, accept, implement check actual result
The guard stack in order. Answers are shaped before you see them; suggestions are checked before they are saved and again when they are shown. A saving should count only once an actual result is on record.

Each check, in plain English

ANSWERS: ASK ANYTHING AND AI CHAT

The empty-data rule

Checks: an automated test calls the assistant's data tools against a company account with no records. Each must return nothing it cannot back (zeros, blanks, empty lists) or say plainly it has no data. Ask Anything is also instructed to use its tools for real data before drawing a chart, and never to make data up.

Prevents: a new plant seeing a confident OEE, or a tidy list of machines that do not exist.

You see: the assistant relays messages such as "No data found for this metric and period", "No downtime events found" or "No alerts in this period". The test runs in our automated build checks: a new tool that returns confident figures on an empty account fails the build, and the known-exceptions list is meant only to shrink. We name those exceptions in the limits section.

Per machine, not pooled

Checks: the facility data profile the AI reads describes multi-machine metrics one machine at a time, then rolls them up explicitly: facility mean, each machine's mean, variation within each machine. No pooled variation is computed, and recent readings carry the name of their machine.

Prevents: pricing the normal gap between a fast station and a slow one as "volatility", or narrating five machines' readings as one sensor's sequence.

You see: spread between machines described as the line's design. Where spread cannot be measured yet, the profile says so instead of reporting zero.

Rates, not running totals; time shares, not averages

Production and reject counters are cumulative, so they are converted to per-hour rates per machine, the same way your KPI screens do. Machine states and fault codes are categories, so they are reported as the share of readings in each state, not averaged. Without enough history for a rate, the answer says it cannot compute one, and output or revenue projections are withheld.

Dollar answers name their basis

Checks: dollar answers (cost of quality, benchmark gaps, what-if scenarios, scheduling) use your facility's cost per unit, revenue per unit and cost per downtime hour. Where those are not set, a default is used ($10, $25 and $5,000 respectively), the result is marked as an assumed default, and the assistant is instructed to say so.

Prevents: a generic number dressed up as yours.

You see: a note that the figure is a planning estimate, not a measurement, plus what to set to make it real. A cost line with no data should be reported as unavailable, not $0.

Illustrative example, sample plant, not a customer
You

What is poor quality costing us over the last 30 days?

KaizenFlow AI

Measured scrap over the last 30 days is about 1,240 units. Your facility has no cost per unit set, so I used the default of $10 per unit: roughly $12,400 in scrapped material. Treat this as a planning estimate, not a measurement. Set your facility's unit economics and this figure will use your real cost. Rework and appraisal costs are unavailable because no data for them is connected.

Basis: production and reject counters for this facility, rolling 30 days; unit cost is an assumed default

No invented roster

Skills and training questions need workforce records. When none are connected, the skill-matrix answer says the data is not available and what would enable it. An earlier version could return a sample roster; a fictional operator with a fictional skill gap is exactly the answer that breaks trust on the floor, so it was removed. Crew optimization has not been fixed the same way yet, and is listed in the limits section.

No pretend capabilities

  • Scheduling. No schedule optimizer is connected today. Asked to optimize, the assistant reports measured throughput, changeover, utilization and OEE and states that optimization is not available. No invented sequence, percentage or dollar figure.
  • Energy. The energy tool reports which monitored power series and how many measured hours a figure covers, and the assistant is instructed never to call it metered or a facility total. Areas with no data are left unscored and are not to be ranked as weaknesses.
  • What we track for you. The supported-metrics list is labelled a product-wide catalogue, not your plant's inventory.
  • Plan adherence. Days with no plan show "no plan on record" and are excluded, not counted as 100%.
  • TEEP. With no shift schedule, utilization and TEEP are reported as unknown, not 0%. See the OEE and TEEP guide.

Charts show only what was measured

Charts appear in Ask Anything. The assistant is instructed to plot the full window the tool returned, to match chart titles to the plotted window, and never to invent, interpolate or backfill points. A chart with no data points is not drawn.

Honest when the AI is unavailable

If the AI service fails, Ask Anything says why in an amber notice, such as "AI service timed out" or "Daily AI budget reached". It never substitutes a demo answer. AI Chat shows a general unavailable message rather than the cause. A failed suggestion run saves nothing.

Scope and access

Your company comes from your sign-in, not from anything typed, and the assistant's queries of plant records filter by it: naming another company or plant in a question does not reach its records. Questions are length-limited. Only engineers, managers and admins can ask; see the security overview.

SUGGESTIONS: AI SUGGESTIONS BOARD

The quality gate

Checks: by default, every suggestion, on-demand or scheduled, passes one shared chain before it is saved. A schema check requires a title and description of real length, a known category, a priority and evidence, rejects negative savings or out-of-range confidence, and requires steps for high-priority cards. A scored gate then looks for depth, stated confidence, numbers in the evidence and actionable steps. It checks that numbers are present, not that they are right; the next checks and your team handle that.

You see: failed cards are not saved. If every card fails, you see "Analysis finished, but no suggestions passed the quality gate". An empty board beats a board of weak cards.

Numeric sanity: math that checks itself

Checks: without asking the model anything, a figure stated next to its own formula is recomputed and corrected if wrong. Physically impossible sentences are removed: more than 168 hours in a week, more hours recovered than the card's own stated gap, OEE or availability above 100%. Legitimate ratios, such as a return above 100%, are untouched.

You see: a corrected number, or a shorter card. The original is kept in server logs.

Cross-checked against your own tiles

On the suggestions board and detail view, a sentence stating a throughput or rejected-parts level an order of magnitude away from your plant's measured rate is removed. The measured rate uses the same window and method as the dashboard KPI tile. Older cards are repaired the same way on display, without changing the stored record.

Feed filters

The board hides cards whose recorded confidence is below the confidence floor, shown read-only in Settings under AI Engine as "Min Confidence Threshold". Cards priced at exactly $0 are dropped. Cards with no dollar figure stay: an unquantified idea is honest, a zero is not.

One opportunity, one card

Cards that a matching rule judges to be the same improvement in the same category are merged, within a run and across everything saved, keeping the one a person already acted on, then the most confident. Board totals count a matched saving once, because a saving can be realized once. Board totals exclude rejected cards, and the dashboard widget shows only cards someone can still act on.

When the honest answer is "I don't know yet"

A plant in its first days has little history, which is exactly when an assistant is tempted to fill gaps. Here is what you will see instead:

  • Metric questions that return "No data found for this metric and period" until records arrive.
  • TEEP and utilization reported as unknown until a shift schedule is configured.
  • Energy figures declared unavailable, with a note on what to connect.
  • What-if scenarios that say there is not enough throughput history for projections.
  • Dollar answers labelled as planning estimates until your unit economics are set.
  • A suggestions board that may stay empty: scheduled analysis needs a minimum amount of recent data, and a run whose cards fail the quality gate saves nothing.
  • Projected-impact ranges on suggestions labelled "±25% planning band" until you have enough verified wins in a category to use your own history.

This is a feature. A "not yet" is designed to name the data to connect next, and it protects the credibility of the first real numbers.

Illustrative example, sample plant, not a customer
You

Where are we losing the most time this week: availability, performance or quality?

KaizenFlow AI

I can only partly answer this yet. OEE over the last 7 days averages 61.4%, from 3 days of machine data. No shift schedule is configured, so utilization and TEEP are unknown, not 0%, and I cannot size schedule losses. No downtime events with reasons have been recorded, so I cannot break the availability loss into causes. To answer fully: configure the shift schedule and connect downtime reason codes from your MES or operator entry.

Basis: OEE records for this facility, rolling 7 days (3 days present); shift schedule and downtime events not available

Start with the inventory question

On day one, ask "Give me an overview of what data you have for this plant". It is designed to answer from your facility's actual records, not the product catalogue, so you can see which questions your data can support yet. The prompt library marks what data each question needs, and Integrations lists the ways to connect it.

People decide, measurement verifies

The checks filter what the AI says. What your plant does stays with people, with two exceptions named below.

Status changes are recorded. Suggestions move through new, reviewed, accepted, in progress, implemented and verified, or rejected. Only engineers, managers and admins change status or record results. A change made with the card's status control writes a permanent event and an audit entry with the user, the old status, the new status and an optional reason. Some other moves, such as recording results, a work order created by the assistant or a verification from a measurement, are logged less completely or not at all. Outcomes feed the next analysis run, so the AI sees what your team accepted and rejected.

Flagged cards wait for review. High-impact, safety and quality suggestions are flagged, and a flagged card that is still new cannot become a work order until someone moves it to reviewed. The flag is not yet shown on the card itself.

The assistant can create a work order. If you ask Ask Anything or AI Chat to act on a suggestion, it can create a work order and move that suggestion to in progress, with no separate confirmation step. Treat that request as the decision itself, and review the card first. The downtime analyst that drafts suggestions has the same tool, so check your work orders after a suggestion run.

A forecast is not a saving. A card can reach verified in several ways: a person sets the status, a person records results, or a before-and-after measurement is run on an implemented suggestion. Not every path requires an actual savings figure, so a verified status is not proof on its own. Asked what implemented suggestions saved, the assistant's savings tool returns three bases: marked implemented, measured (a recorded implementation date), and verified with an actual savings figure, and only that last group is summed as verified savings. Some portfolio rollups fall back to the AI estimate for a verified card with no actual figure, so check the basis before you report a number.

Your track record frames each forecast. Each card's confidence panel shows the share of predicted savings your plant actually realized in that category. With enough verified wins, the projected range comes from your history, not a generic planning band.

Finance reconciles. Good practice before accepting any measured result: compare it with a baseline normalized for volume, product mix and scheduled hours, so a busy month is not mistaken for an improvement. See KaizenFlow for finance.

Can the assistant change a machine setting?

No. None of the assistant's tools write to equipment. It reads the records your connectors bring in.

Can it change anything at all?

Two things. It can create a work order from a suggestion, as described above. And asking for an SPC analysis can add an SPC alert to your notifications when it finds recent out-of-control points.

Who can see my AI Chat conversation?

AI Chat keeps a short recent history for each plant, and it is shared by everyone in your company who uses chat for that plant. Clearing it clears it for everyone. Ask Anything keeps its history only in your browser tab and loses it on reload.

What these checks cannot do

A page about honesty has to be honest about its own limits. These are the ones we know about.

Garbage in, faithful garbage out

Every check assumes the records are real. A mis-scaled sensor, a counter that resets mid-shift, careless reason codes or a connector mapped to the wrong machine will be reported faithfully. Ask about data quality and connector health regularly, and see monitoring older machines for legacy signals.

Suggestion dollars are estimates

The AI model writes suggestion dollar figures itself. It is given per-machine baselines, but no formula recomputes its figure. The checks catch self-contradicting, impossible and order-of-magnitude errors, not a figure that is plausible but optimistic. That is why a saving should not count until it is measured.

  • The plant cross-check is narrow. It covers throughput and rejected parts on the suggestions board and detail view. Shift handovers, briefings and other surfaces show the text as saved.
  • Removed sentences are not marked. A card edited by numeric sanity does not currently show that it was edited.
  • Confidence is self-assessed by the model, not a calibrated probability. Your track record is the better guide.
  • The workflow does not force an order. A user with rights can move a card to any status, including verified. Before counting a verified card, confirm it carries an actual result.
  • The empty-data rule is a release test, not a live filter. Known exceptions remain: predictive-maintenance equipment counts, value-stream bottlenecks and crew optimization. With no downtime records, reliability can read 100% available, and a TEEP breakdown can show quality at 100%. Treat these with caution on thin data.
  • Windows are rolling and capped. Ranges count back from now in UTC, not your shift calendar, and each question type has a maximum lookback. The assistant's instructions do not include today's date, so "yesterday" or "last Tuesday" is approximated as hours or days back. Read the window in the answer.
  • Most answers are plant-level. Most tools cannot filter to one machine or line.
  • Ask Anything has no memory between questions. Restate the metric and window in follow-ups.
  • No input filter is complete. Do not paste secrets into any AI box. See the IT review page.

How to audit an answer yourself

You should not have to take our word for any of this. A skeptical supervisor or finance reviewer can follow these steps.

  1. Pin the window and the plant. Find the window in the answer or chart title (rolling, UTC) and confirm the plant selected in the app.
  2. Read the "Tools used" row. Chips under each answer name the data tools queried. If no listed tool could have returned downtime data, ask where a downtime figure came from.
  3. Match it to a dashboard tile. "Current OEE" is the same windowed mean the KPI dashboards use, and throughput rates are derived the same way. Compare over the same window.
  4. Find the basis of every dollar. "Planning estimate, not a measurement" means a default, not your economics.
  5. Ask what the AI can see. Use the inventory, data-quality and connector-health questions below.
  6. Open a suggestion's confidence panel. It shows what the card is grounded in, why now, the key assumption with a "Check:" line, the reasoning, the impact range and its method, and your track record.
  7. Test the assumption on the floor. If it fails, reject the card; rejections feed the next run.
  8. Separate the savings bases. Confirm savings answers split marked, measured and verified, and that every verified card carries an actual result.
  9. Trace who moved it. Changes made with the status control are recorded with the user, the time, and the old and new status. Moves made by recording results, by the assistant or by a measurement may be missing or incomplete there.
  10. Reconcile. Compare verified savings with your normalized baseline before a finance review.
  • Give me an overview of what data you have for this plantWhich metrics and records actually exist for this facilityNeeds: any metric data
  • How good is our data? Are any sensors stale or missing?Data-quality scorecard for connected sources and streamsNeeds: connected data sources
  • Are all our data connectors healthy?Connector status and heartbeatsNeeds: active connectors
  • How much have the AI suggestions we implemented actually saved us?Savings split into marked, measured and verifiedNeeds: recorded actual savings
  • Which suggestions were rejected, and why?Rejected cards, with any reasons your team recordedNeeds: rejected suggestions

For the full question set, see the prompt library. To see these checks on your own data, start a pilot.

Audit it on your own data

The pilot ends in a before and after savings report measured against a normalized baseline and reconciled with your finance team. Nothing counts until it settles.