Docs / KaizenFlow AI

KaizenFlow AI user guide

Ask your plant a question in plain English and get an answer you can check. This guide covers what KaizenFlow AI does, what it deliberately will not do, and how to get answers that hold up in a finance review.

Updated Sep 26, 2026Reading time 39 minAudience Every plant role

What KaizenFlow AI is

THE PROBLEM

A modern plant produces more data than anyone has time to read: machine states, part counters, downtime reasons, scrap, energy, plans, work orders. The answer to "why did we miss plan this week" is usually spread across five screens and two systems, and someone has to stitch it together by hand.

Plants are data-rich and answer-poor. KaizenFlow AI does the stitching. You ask in plain English. It queries your plant's own data in KaizenFlow, does the arithmetic over the window each data tool covers, and tells you what it found, including when it found nothing.

THE SURFACES

KaizenFlow AI is not one feature. It is a set of surfaces that read the same data under the same rules.

SurfaceWhat it doesWhere to find it
Ask AnythingAnswers one plain-English question about a facility from its data, with charts where they help.Sidebar: Analytics & AI, then Ask Anything
AI ChatThe same data tools in a conversation. It remembers recent turns for a facility, so follow-ups work. Answers are text.Sidebar: Communication, then AI Chat
AI SuggestionsUp to nine specialists review your data on a schedule and on demand, and propose improvement cards with evidence and an estimated impact.Sidebar: Analytics & AI, then AI Suggestions
Confidence panelOn a suggestion card, each part shown when the card has it: evidence, why now, the key assumption and how to check it, reasoning, an impact range and your track record.Tap or hover a card's confidence
Shift handoverAn AI-written narrative of the last shift, covering a window of 1 to 24 hours (8 by default).Sidebar: Shift Intelligence
AI Project PlannerBuilds DMAIC, TPM, Kaizen Blitz and similar plans from methodology steps plus your live metric gaps and suggestions. It does not call the AI model to build a plan.Sidebar: Strategy & Finance, then AI Project Planner

THE RULE UNDER ALL OF IT

Proof, not promises. The assistant is instructed to take every figure from its data tools and never make data up. The tools are built to say "no data" instead of returning a comfortable default, and to label dollar figures that rest on assumed costs as planning estimates. Where that is not yet fully true, this guide says so in place, not in fine print. The full account of the checks lives in How KaizenFlow AI earns trust.

Getting to it

From the sidebar

  • Ask Anything sits in the Analytics & AI group, next to AI Suggestions.
  • AI Chat sits in the Communication group.

From the keyboard

  • Ctrl+K (Cmd+K on Mac) opens the command palette. Type ask or chat and press Enter.
  • / focuses the global search bar. ? opens the full shortcuts panel.
  • There is no dedicated shortcut for Ask or Chat. On the Ask page, the question box is focused on load.

Which plant it answers about

Ask Anything answers about the facility selected in the app-wide facility selector. The Ask page has no picker of its own, so switch facility first, then ask. If your facility has not finished loading, you get a short message asking you to try again in a moment.

AI Chat has its own facility selector. Both can reach your company's other facilities for comparisons.

Who can use it

RoleAsk Anything and AI ChatAI Suggestions
Admin, ManagerAsk and chatGenerate, review, change status, record results. Assign owners.
EngineerAsk and chatGenerate, review, change status, record results
ViewerSees the links, but questions are refusedRead only

Known gap: viewers see a generic error

A Viewer who opens Ask Anything and asks a question gets "Sorry, I could not process that question. Please try again." rather than a message about permissions. If that happens on every question, check your role with your KaizenFlow admin before anything else.

Language

The interface ships in five languages, but the Ask and Chat pages are in English and your interface language is not passed to the AI. For the most predictable answers, ask in English.

What it can and cannot do

Ask Anything and AI Chat draw on the same 44 data tools. Here is what they cover. Where the data a topic needs is missing, the tool says so.

TopicAsk it like thisWhat comes back, and what it needs
OEEWhat's our current OEE?The windowed average the dashboards use, reported as the current value instead of a spot reading. Needs OEE readings.
Loss breakdownWhere are we losing the most time: availability, performance or quality?TEEP, OEE factors and a loss waterfall in hours. TEEP and the waterfall need a shift schedule. Quality comes from scrap; availability and performance are flagged as estimates unless you stream them separately.
ThroughputWhat's our throughput rate over the last 24 hours?Units per hour, computed per machine. Needs several hours of counter data.
Scrap and SPCIs our scrap rate in statistical control? What's the Cpk?Trend, control analysis, Cp and Cpk. Spec limits are optional.
Cost of qualityWhat is our cost of quality over the last 30 days?Prevention, appraisal and failure costs, with their economics basis.
DowntimeWhat were the top downtime reasons over the last 7 days?Event count, hours, categories, top 10 reasons, recent examples with equipment.
Reliability and riskWhat's our MTBF and MTTR, and which equipment is most at risk in the next 24 hours?MTBF, MTTR, worst equipment, and a risk forecast learned from downtime history.
Anomalies and alertsWere there any anomalies in the last 24 hours?Unusual readings against the recent baseline. Alerts by severity; in a busy window the alert counts top out at 50.
EnergyWhat's our energy per unit and carbon footprint over the last 30 days?Intensity and carbon, naming the power series covered. Needs power readings.
Plan vs actualHow did we do against plan over the last 7 days?Adherence on days with a plan. Days without one are excluded, not scored 100%. Needs a production plan, for example from your ERP.
Plants and patternsWhich plant is best on OEE? Does OEE vary by weekday?Your own facilities ranked, which needs more than one facility with data. Weekday and hour patterns from several weeks of data.
What-ifWhat if we cut changeover time by 20%?Output change from the measured hourly rate, with dollars and their basis. With too little throughput history, projections are withheld.
Root cause, handoverWhy is OEE low this week? Summarize the last 8 hours for handover.A structured root-cause analysis, which uses its own window. A shift summary.
Safety, supply chainAny near misses in the last 7 days? Any supply chain risks?Events and risks, or an empty result or plain statement when none are on record.
Improvement recordHow much have the suggestions we implemented actually saved?Three labelled figures: marked implemented, measured, and verified.
Documents, past casesWhat does our SOP say about changeover? Have we tried this before?Passages from documents your team uploaded. Similar past suggestions.
Data healthAre all our data connectors healthy?Connector heartbeats and a data-quality scorecard.

Two to start with. The prompt library has many more, by role.

  • Give me an overview of what data you have for this plantThe fastest way to learn what the AI can answer for youNeeds: any data
  • Show me OEE trends for the last 30 daysUsually a line chart over the full window, with a dashed target line where a target is setNeeds: OEE readings

Not a help desk, and three topics to avoid for now

Ask and Chat have data tools for your plant and a search over documents your team uploaded. None of their tools covers KaizenFlow's own help content, so for how-to questions about the screens, use the Academy. Do not rely yet on predictive-maintenance equipment lists, value-stream bottleneck maps or crew sizing: their tools can still return placeholder content, and the value-stream map can do so even when your plant has real data. Asked to optimize a schedule, the scheduling tool reports measured changeover, utilization and OEE, and states that no optimization is available.

What it never does

It never writes to machines, PLCs or setpoints: no AI tool controls equipment. It never writes to your ERP, MES, CMMS or other connected systems; it reads KaizenFlow's copy of the data. It never edits production schedules or job sequences. And Ask Anything never remembers your previous question.

The two things it can change

Work orders. Asked to act on a specific suggestion, Ask or Chat can create a work order from it and move the suggestion to In progress. It refuses suggestions already underway or closed, and some unreviewed high-impact ones. There is no separate confirmation step, so ask only when you mean it.

SPC alerts. An SPC analysis that finds recent out-of-control points records an SPC violation alert alongside your other alerts.

Three things that shape every answer

  • Time is rolling, not calendar. Windows count back from now in UTC. The AI is not given today's date, so "yesterday" or "last Tuesday" becomes an approximate number of hours back, not a calendar day in your time zone. Say "the last 24 hours" or "the last 7 days" and the window is unambiguous.
  • Machine-level questions are limited. Most tools report at facility level and cannot filter to one machine or line. Equipment appears in ranked lists and recent downtime examples. Only vision inspection filters by line.
  • Each tool has a maximum lookback. A longer request is shortened without warning, so check the window the answer states. The limits are in Data it reads.

How a question gets answered

The real path of an Ask Anything question. AI Chat follows the same shape, with recent conversation added, up to four tool rounds instead of five, and text-only answers.

  1. You ask. Type a question of 3 to 2,000 characters and press Enter or Ask, or, before your first question, click one of eight suggested questions. One question runs at a time.
  2. It is scoped before anything runs. Your company comes from your sign-in, and the facility from the app-wide selector. Your role is checked: Engineer, Manager or Admin.
  3. The model reads your question. Its standing instructions: use the data tools for real figures and never make data up. There is no fixed menu of question types; the model decides which tools fit, and turns "this week" into a number of hours or days. Plant questions cross domains, so letting it chain tools lets one question follow the evidence.
  4. Tools gather data, in rounds. In each of up to five rounds the model can call several tools and use the results to choose the next call: find the biggest loss, then pull the downtime reasons behind it. Each round has a time limit.
  5. Tools do the arithmetic where the data lives. For most metrics and for downtime, averages and totals are computed in the database over the full window, not a sample. OEE uses the dashboards' windowed average. Counters become per-hour rates, machine by machine, from the most recent readings, so on a long window check that a throughput rate covers the period you asked about. State codes become the share of readings in each state, not an average.
  6. Tools say when they cannot back a number. Tools are built to answer "no data" rather than a zero when they find no rows; the known exceptions are listed under Anatomy of an answer. TEEP without a shift schedule comes back unknown. Dollars based on assumed unit costs are labelled as such.
  7. The model writes the answer and any charts. The model is instructed to plot the full-window series in trend charts, to match each chart title's window to the plotted data, and never to invent, interpolate or backfill points. These are instructions, not a check on the output, so compare a chart's title with the window stated in the text. If the rounds run out mid-gathering, a final pass writes the best answer from what was collected; if even that fails, you are asked to try again.
  8. You get the whole answer at once. Answers are not streamed. You see "Analyzing your question with AI tools..." until it is ready, then the text, charts and a Tools used: row. Multi-round questions take longer. If the AI service cannot answer, you get a labelled notice, not a guess.
01 · INPUT One question, 3 to 2,000 characters 02 · SCOPE Your company and the selected facility 03 · MODEL Chooses its own tools, up to 5 rounds 04 · TOOLS 44 data tools, they do the math 05 · DATABASE Your plant data, filled by connectors RESULTS QUERY 06 · COMPOSE Text and chart specs, told to match window WHEN DONE 07 · RENDER Empty charts dropped, sent whole, no stream 08 · ANSWER Text, charts and Tools used chips ACCENT: THE DATA-GATHERING LOOP. THE MODEL PICKS TOOLS; THE TOOLS DO THE MATH.
An Ask Anything question from input to answer. The model never touches your machines or connected systems: it calls tools, and the tools query KaizenFlow's database.

Anatomy of an answer

Every Ask Anything answer has up to three parts:

  1. The written answer, formatted with headings, bullets and tables where useful.
  2. Charts, when one helps: line, area, bar, scatter, pie (drawn as a donut), gauge or table. A gauge shows the value against its target and turns green when the target is met. Line, area, bar and scatter charts draw a target, when there is one, as a dashed line. Any chart can carry a one-line insight underneath. A chart with no data points is not drawn, though an empty table can still show its title.
  3. Tools used: a row of chips naming the data tools that ran, for example query_metrics or get_downtime_events.

Reading the basis

There is no separate "sources" box. You read the basis of an answer from three places: the Tools used chips tell you what data was queried, the window is stated in the answer's text or chart title (the model is instructed to match chart titles to the plotted data), and the scope is the facility selected when you asked. If an answer does not state its window, ask again and name one. In the examples below, the Basis line is our annotation assembled from those three places; the product does not print it as one line.

When data is missing or a figure cannot be backed

Each Ask tool is tested in our build against an empty plant: it must return nothing it cannot back, or say plainly that the data is absent. Three tools are still on a known-exceptions list (the predictive-maintenance, value-stream and crew-sizing tools named above), and the build fails if any other tool returns figures for an empty plant without declaring that the data is absent. The test checks that the absence is declared, not that every figure beside it is neutral, which is how the two soft spots below still pass. You will see answers like these:

  • "No data found for this metric and period." or "No downtime events found."
  • TEEP and utilization reported as unknown, not 0%, when no shift schedule is configured.
  • Projections withheld, and cost line items reported as unavailable rather than $0, when they cannot be computed.
  • Dollar figures marked as planning estimates when your facility's unit economics are not set, using $10 cost per unit, $25 revenue per unit and $5,000 per downtime hour.

Two known soft spots

With no downtime records, the reliability tool can report 100% availability; with no quality readings, the loss breakdown can report 100% quality. On a new plant, read a perfect score as "no data yet" until you confirm the data exists. If availability and performance come back as the same number, they are an estimated split, not two measurements. And "which metrics do you track?" returns the product-wide catalogue, labelled as such; ask for an overview of this plant's data instead.

Illustrative example, sample plant, not a customer. Ask Anything, a clean answer with a chart.
You

What were the top downtime reasons over the last 7 days?

KaizenFlow AI

Over the last 7 days the plant logged 46 downtime events totalling 38.2 hours. 41 events have a recorded duration; the other 5 are counted but not included in the hours.

By time lost, the top three reasons are material shortage (12 events, 11.4 h), tool change (9 events, 7.9 h) and hydraulic fault (3 events, 6.1 h). Unplanned stops are 58% of events but 71% of minutes, so they are fewer and longer than planned ones.

Chart: bar, "Downtime hours by reason, last 7 days".

Basis: Tools used get_downtime_events. Window: rolling 7 days back from now, UTC. Scope: selected facility, all equipment. Method: totals over every event in the window.

Illustrative example, sample plant, not a customer. Ask Anything, when a figure cannot be backed.
You

What's our TEEP this week, and where is the biggest loss?

KaizenFlow AI

TEEP and utilization are unknown for this facility, not 0%. No shift schedule is configured, so scheduled time cannot be separated from unscheduled time.

What can be measured over the last 7 days: OEE is 63.8% and quality is 94.3%, taken from the scrap rate. This facility does not stream availability and performance separately, so they are an estimated split of what remains (about 82.3% each), not measurements. What is solid: availability and performance together cost more than quality does.

The loss waterfall in hours also needs a shift schedule. To get TEEP and the full waterfall, configure this facility's shift schedule.

Basis: Tools used get_teep_breakdown. Window: rolling 7 days. Scope: selected facility. Missing input named: shift schedule. Availability and performance flagged as estimated.

Follow-up questions

AI Chat keeps the recent conversation for each facility, so follow-ups such as "that one" carry context. Ask Anything answers every question on its own.

Illustrative example, sample plant, not a customer. AI Chat, where earlier turns carry over.
You

Which equipment had the worst reliability over the last 30 days?

KaizenFlow AI

Plant-wide there were 22 unplanned failures: MTBF 30.3 hours, MTTR 38 minutes. The machine with the most failures was CNC-07: 9 unplanned failures, MTBF 79.1 hours, MTTR 52 minutes.

Basis: Tools used get_reliability_metrics. Window: rolling 30 days. Scope: selected facility. Method: unplanned stops count as failures; equipment ranked by failure count.

You

What have the stops on that one been about?

KaizenFlow AI

Taking "that one" as CNC-07. Its 9 unplanned stops over the last 30 days were 5 tool breakages, 3 low-coolant stops and 1 spindle alarm. This comes from the per-machine breakdown in the reliability data, which covers unplanned stops only; planned stops such as changeovers on CNC-07 are not separated out, because downtime reports cannot be filtered to one machine.

Basis: Tools used get_reliability_metrics. Window: rolling 30 days. Scope: selected facility, CNC-07's top reasons by count, unplanned stops only.

You

Have we tried anything on tool breakage before?

KaizenFlow AI

Two earlier suggestions match. Changing inserts on a tool-life count instead of on failure is marked implemented. A switch to a premium insert grade was rejected, and the recorded reason is cost.

Basis: Tools used find_similar_past_cases, get_suggestion_history. Scope: your company's past suggestions.

In Ask Anything, the second question above would arrive with no context. Restate it in full instead: "What were CNC-07's top unplanned stop reasons over the last 30 days?"

Chat history is shared

AI Chat memory belongs to the facility, not to you. Everyone in your company chatting about that facility shares one thread, and Clear Chat clears it for all of them. It keeps the last 20 messages, and the thread expires 7 days after its last message. Ask Anything's history stays in your browser tab and is gone on reload.

When there is no answer

In Ask Anything, an amber box names why the AI service could not answer, such as "AI service timed out" or "AI provider rate limit". A red "Sorry, I could not process that question." means the request itself failed, most often because it was under 3 or over 2,000 characters, came from a Viewer role, or hit the request limit. The message does not say which, so check those three in that order. "Could not reach the AI service" means your connection dropped. None of these is ever replaced by a made-up answer.

Data it reads, and how fresh it is

How data reaches the AI

The AI never talks to your machines or business systems. Connectors bring data into KaizenFlow's database, and the AI's tools query that database. That separation is why it can read what it needs and change nothing on the floor.

Connector categoryExamplesWhat it lets the AI answer
Shop floor and IIoTOPC UA, MQTT, Modbus TCP, MTConnect, RESTOEE, throughput, machine states, anomalies, energy
Historians and SCADAIgnition, Kepware KEPServerEX, AVEVA, PIThe same machine-level metrics, from the tags your historian already collects
MESPlex, DELMIAworks, Tulip, TrakSYS, L2LDowntime reasons, production, quality
ERPSAP S/4HANA, Business Central, Epicor Kinetic, Infor CloudSuitePlan vs actual
CMMS, quality, PLMFiix, UpKeep, Maximo, MasterControl, ETQ Reliance, ArenaContext for cross-system questions
WorkforceADP, BrainierWorkforce and skill questions
Data platforms and filesSnowflake, Databricks, SQL, OneLake, CSV or Excel upload, Google Sheets, webhooks, emailed reportsWhatever the source holds, once it lands as KaizenFlow metrics or events

Energy data arrives over OPC UA or MQTT; there is no dedicated utility-meter connector. A few catalog entries are roadmap placeholders that move no data yet, so check Integrations or ask us before planning around one.

Which surface sees what

SurfaceReadsWindow
Ask AnythingAll 44 data tools, for the selected facility and, for comparisons, your other facilitiesPer-tool limits, below
AI ChatThe same tools, plus the facility's recent conversationPer-tool limits, below
AI SuggestionsEach specialist's own tools, plus the data profile, past suggestions and uploaded documentsScheduled runs analyze the last 24 hours
Shift handoverMetrics, downtime, alerts and tasks for the shift1 to 24 hours, 8 by default

How fresh it is

ProcessCadence
Polled connectors syncEvery 5 minutes
MQTT dataContinuously, as messages arrive
Hourly and daily rollups recomputedEvery 30 minutes
Alert rules checkedEvery 5 minutes
Anomaly scan (last 2 hours against a 7-day baseline)Every 30 minutes, and it can trigger a focused AI analysis
Scheduled AI suggestion runEvery 4 hours on active facilities, skipped when too little new data has arrived

All schedules run in UTC. Ask's tools aggregate stored readings when you ask, so for a polled connector an answer can be a few minutes behind the machine.

How far back it can look

Question typeMaximum lookback
Anomalies, vision inspection7 days
Metrics (OEE, throughput, scrap and others), SPC30 days
Downtime, plan vs actual, cross-system correlation90 days
Cost of quality180 days
Energy, safety365 days

A longer request is shortened without warning: "OEE for the last 6 months" is answered over 30 days. Retention also caps history: metric data older than the retention period your organization sets in Settings under Data Retention is removed by a daily cleanup, and the AI cannot see removed data.

What "no data yet" means on a new plant

  • A tool says no data. The connector is not live yet, or does not carry that metric.
  • TEEP is unknown, or dollars are planning estimates. Configure the facility's shift schedule and unit economics.
  • No rate yet. Counters need several hours of history before a per-hour rate can be computed.
  • No scheduled suggestions. A scheduled run needs a minimum amount of new data in the last 24 hours before it starts.
  • A second plant looks empty while the first looks busy. In a company with several facilities, polled connector data can land on the first facility rather than the one you expect. Ask us to confirm where each connector writes before comparing plants.

First question on a new plant

Ask "Give me an overview of what data you have for this plant." It answers from what this facility actually has on record, so you know which questions will get a real answer today.

The specialist ensemble

KaizenFlow has nine AI specialists. They work in AI Suggestions, the background analysis that proposes improvements, not in Ask Anything or AI Chat. Each brings the methods of its discipline and a narrower set of tools.

SpecialistLaneLooks at
Downtime and AvailabilityTPM, changeovers (SMED), root cause (5 Why, fishbone, fault tree), MTBF and MTTR, breakdown ParetoDowntime events, downtime risk forecast, reliability, patterns
Quality and DefectsSPC and Cp/Cpk, DMAIC, poka-yoke, first-pass yield, scrap, cost of qualitySPC, cost of quality, root cause, anomalies, benchmarks
Capacity and UtilizationTEEP and the OEE-to-TEEP gap, shift scheduling, constraints, cycle time. Always on.Loss breakdown, plan vs actual, plant comparison, what-if
Energy and SustainabilityISO 50001, compressed air, HVAC, motors and drives, peak demandEnergy analysis, anomalies, patterns
Safety and ComplianceIncident and near-miss trends, lockout/tagout, guarding, PPE. Instructed never to trade safety for savings.Safety events, anomalies, active playbooks
Supply Chain and MaterialsSupplier on-time delivery, lead-time variability, safety stock, Kanban, single-source riskSupply chain data, anomalies, multi-variable scenarios
Workforce and LaborSkill matrix, training, overtime and fatigue, cross-trainingWorkforce data, skill matrix, patterns, shift handover
Sustainability and ESGCarbon per unit, waste, water, ISO 14001Sustainability data, energy, anomalies
Computer Vision and Visual QualityVisual defect detection, camera coverage, precision and recall, false-reject tuning, label checksVision data, anomalies

How a run is routed

  1. A run starts. Every 4 hours when there is new data, when the anomaly scan asks for a focused look, or on demand.
  2. Triage picks specialists from your data, not from words. A specialist joins when the facility has the metrics its lane depends on. Capacity always joins, as do categories where your team has accepted valuable suggestions before.
  3. Your settings apply last. Specialists can be switched off per company in Settings under AI Engine, so your company may run fewer than nine. Capacity cannot be switched off.
  4. Specialists work in parallel. Each gathers data with its own tools and drafts findings.
  5. One synthesis merges the drafts. Cards matched as the same opportunity collapse to the strongest one, so that saving is counted once. The match works on category and wording, so two cards that describe one fix differently can both survive; check look-alike cards before adding up their savings. A quality check can send the run back for up to two more passes.
  6. Cards pass the write-side checks. By default, every card is validated and scored before it is saved, and a card that falls short is not saved. Where a card shows its working, figures that contradict their own arithmetic are corrected, and some provably impossible statements are removed.

Why an ensemble here, and one assistant in Ask

Finding improvements is a survey problem. A downtime engineer and an energy engineer see different money in the same plant. A focused lens with its own methods avoids generic advice, and the merge step stops two lenses claiming the same saving.

Answering a question is a direct problem: you want one answer quickly, from whatever data it touches. So Ask and Chat use one assistant with every tool, not a relay between specialists.

What a suggestion's dollar figure is

The savings estimate on a suggestion card is written by the AI, which is given per-machine baselines from your data. No fixed formula recalculates it. The automated checks catch contradictions and impossible statements in the card's text; they do not prove the dollar estimate right. Treat it as an estimate until your team records a measured actual saving against it. A card's Verified status on its own is not that proof, so check that a verified card carries a recorded actual saving. The trust page explains every check.

From answer to action

Ask Anything answers questions. Improvement suggestions come from a separate process, the AI Suggestions analysis, which is where KaizenFlow's nine specialists work. Here is how one suggestion travels from first draft to a verified, measured result. The rule at every step: the AI ranks, people decide.

HOW A SUGGESTION IS PRODUCED

  1. An analysis starts. An Engineer or above runs it from the AI Suggestions page, or it runs on schedule every 4 hours over the last 24 hours, skipping when nothing new arrived. An anomaly scan every 30 minutes can add a focused run.
  2. Your data picks the specialists. Specialists are chosen by which metrics your facility has and by which categories paid off before, not by wording. Capacity always runs; the other eight can be switched off per organization in Settings, under AI Engine, so your organization may run fewer than nine. Selected specialists work in parallel.
  3. Drafts are merged, not voted on. One synthesis step merges the drafts under explicit instructions: never invent data, one card per opportunity, and state why now, the key assumptions and how to test them. Near-duplicates are collapsed by matching the improvement and its title, which catches most repeats of one saving but not every rewording.
  4. By default, every card is validated before it is saved. It must pass a structure check and a deterministic quality score, then goes through a numeric sanity pass that edits text but never drops a card. Figures whose stated arithmetic can be checked are recomputed when they contradict it; physically impossible sentences, such as more than 168 hours in a week, are removed. A removed sentence leaves no marker on the card today. If no card passes, nothing is saved and you see Analysis finished, but no suggestions passed the quality gate.
  5. The board filters again. It hides cards below the confidence floor (set in server configuration and shown read-only as Min Confidence Threshold in Settings), drops cards priced at $0, collapses duplicates across runs, and removes throughput or rejected-parts statements that are off from your plant's measured rate by an order of magnitude. That last check runs on the suggestions board and card detail, and catches gross errors, not small ones.

HOW A SUGGESTION IS PRICED

Plainly: the dollar estimate on a suggestion card is written by the AI model, and no formula recomputes it. What we control are the inputs and the checks. The model gets per-machine baselines instead of pooled plant statistics, so normal spread between stations is not mistaken for volatility, and the data tools report known state codes, such as run state, as counts rather than averages.

Treat it as a ranked hypothesis. The confidence popover on each card shows what it is grounded in, why now, the key assumption and how to check it, and an impact range drawn from your own verified history, or a plus or minus 25% planning band until you have enough. The confidence score starts as the model's self-assessment and can later be recalibrated from observed outcomes. See the trust and verification guide.

PEOPLE DECIDE

A suggestion moves through New, Reviewed, Accepted, In progress, Implemented and Verified, or is Rejected. Changing status and recording results need the Engineer role or above; Managers and Admins assign owners. Status changes made on the suggestions board write an immutable event and an audit entry, and outcomes feed the next analysis run. A work order created from Ask Anything or AI Chat moves the card to In progress without writing either.

Two things to know today

Safety, quality, high-value and very-high-confidence suggestions are flagged for review and cannot become a work order while still New. The flag is not yet shown as a badge, and moving the card to Reviewed clears the block, so review is a team discipline, not a hard lock. Also, Ask Anything and AI Chat can turn an existing suggestion into a real work order and set it In progress if a question asks for it, with no separate confirmation step. Only ask for a work order when you mean it.

FROM SAVING TO VERIFIED RESULT

  1. Implement and mark it. Marking a card Implemented timestamps it and, where it can, creates a proof card.
  2. Record the measured result. An Engineer or above records actual savings and throughput gain, and the card becomes Verified with a lifecycle event. The results form opens filled with the AI estimate and saves as Verified even with its fields empty, so replace the estimate with your measured figure before you save. A card can also reach Verified through a direct status change, so before you count a saving, check that a measured figure, not the estimate, is on record.
  3. Reconcile with finance. Compare each card's predicted dollars with its recorded result, line by line, and report as verified savings only Verified cards whose recorded result you have checked against a measurement. The savings ledger page describes the predicted-versus-verified view this step is built on.

The ledger at design-partner stage

The predicted-versus-verified ledger exists as a report in the platform, not yet as a screen in the app. It reads the raw suggestion records, while Ask reads the deduplicated board, so the two totals can differ; reconcile line by line and count only Verified cards whose recorded result you have checked. A nightly calibration job can also replace the recorded actual savings on older cards with a figure computed from plant-wide metric changes, so keep the measured figure in your own records too. The ROI Tracker page is not the ledger: it fills gaps with example figures.

Illustrative example, sample plant, not a customer
You

How much have the AI suggestions we implemented actually saved us?

KaizenFlow AI

Three numbers, and they differ. 6 suggestions are marked Implemented or Verified. 4 have a recorded implementation date. 3 are Verified with actual savings on record, totalling $41,800. Check those 3 against your measurements before you report the $41,800 as verified savings.

Basis: deduplicated suggestion board; status, implementation date and recorded actual savings on Verified cards

THE AI RANKS 01 ANALYZE Specialists read your plant's data 02 MERGE Drafts merged, one card per opportunity 03 GATE Structure, quality and numeric checks 04 BOARD Low-confidence and $0 cards hidden PEOPLE DECIDE 05 DECIDE Review, accept or reject 06 EXECUTE Work order, owner and due date 07 MEASURE Record the result; card is Verified 08 RECONCILE Predicted vs verified, line by line
The AI produces, checks and ranks suggestions. From step 05 on, the decisions belong to people: status changes made on the suggestions board are written to the audit trail, and a saving should count only once a measured result is on record.

Use it by role

Access follows your KaizenFlow role, not your job title: Engineer, Manager and Admin accounts can ask; Viewer accounts cannot. Each prompt below maps to a question the data tools can answer.

Operator

Less time reconstructing the last shift. Asking needs the Engineer role or above, so where operators hold Viewer accounts, a supervisor can run these at handover and share the answer. Training: operator track.

  • Summarize the last shift for handoverMetrics, downtime, alerts and tasks from the last 8 hoursNeeds: recent metrics, downtime, alerts
  • What alerts fired in the last 24 hours, and are any still unacknowledged?Alerts by severity and type, with how many are unacknowledgedNeeds: alert events
  • Any safety incidents or near misses in the last 7 days?Recorded incidents and near missesNeeds: safety event records

Supervisor

The downtime picture and overnight exceptions in one question each, before the start-of-shift meeting. Training: supervisor track.

  • What were the top downtime reasons over the last 7 days, by events and by minutes?Top reasons, categories and recent examples with equipmentNeeds: downtime events with reason codes
  • Were there any anomalies in the last 24 hours?Readings that broke from their recent baselineNeeds: recent metric history
  • Which open actions are overdue against SLA?Escalation status for open action itemsNeeds: action items with SLAs

Quality and continuous improvement

Ask whether a process is in control, what poor quality costs, and what the team already tried. Training: quality and CI track.

  • Is our scrap rate in statistical control over the last 30 days? What is the Cpk?Control status and Cp/Cpk; out-of-control points can raise an SPC alertNeeds: scrap rate history
  • What is our cost of quality over the last 30 days, and which figures use assumed unit costs?Prevention, appraisal and failure costs, assumptions labelledNeeds: throughput counter, scrap, unit costs
  • Which suggestions were rejected, and why?Rejected suggestions, with a reason where one was recordedNeeds: suggestion history

Engineer

Rank where reliability work pays back, and find what the plant's own documents say. Training: engineer track.

  • What is our MTBF and MTTR over the last 30 days, and which equipment is worst?Reliability metrics and a worst-equipment listNeeds: downtime with equipment and durations
  • Does our OEE vary by day of the week or time of day?Recurring patterns across several weeksNeeds: several weeks of OEE history
  • What does our SOP say about changeover on the press line?Relevant passages from manuals your team uploadedNeeds: uploaded SOPs and manuals

Plant manager

The loss structure, plan adherence and a quick sizing of an idea before committing people to it. Training: plant manager track.

  • Where are we losing the most time: availability, performance or quality?Loss waterfall in hours and the biggest lossNeeds: metrics and a shift schedule
  • How did we do against plan over the last 7 days?Plan vs actual; days with no plan on record are excludedNeeds: ERP plan and actuals
  • What if we cut changeover time by 20%?Projection from the measured throughput rate, or withheldNeeds: 7 days of throughput, unit economics

Executive and finance

Separate claimed savings from verified savings, and compare plants on the same definitions. Training: executive track; see also KaizenFlow for finance.

  • How much have the AI suggestions we implemented actually saved us?Marked, dated and measured counts reported separatelyNeeds: suggestions with recorded results
  • Which plant is performing best on OEE over the last 30 days?Comparison across your own facilitiesNeeds: OEE for more than one facility
  • How does our OEE compare to the industry?Benchmark position; assumed economics labelledNeeds: facility OEE

IT administrator

A health check on the data feeding the AI; bad answers usually start with a stale connector. Training: IT administration track.

  • Are all our data connectors healthy?Connector status and heartbeatsNeeds: active connectors
  • How good is our data? Are any sensors stale or missing?Data-quality scorecardNeeds: connected metric streams
  • Give me an overview of what data you have for this plantWhat this facility holds, not the product catalogueNeeds: any facility data

Best practices

First week: verify against answers you already know

  1. Start with the data profile. Ask what data the AI has for your plant. If a data type is missing there, questions that depend on it will come back thin or unavailable, so fix the source before judging the answers.
  2. Match a dashboard tile. Ask for current OEE. Ask reports the same windowed average the KPI dashboards use, so they should agree for the same window.
  3. Match your downtime log. Ask for unplanned downtime events and hours over the last 30 days and compare with your records.
  4. Read the Tools used row. If a downtime question shows no downtime tool, rephrase.

Phrasing that works

Ask resolves time as rolling hours or days back from now, in UTC, is not told today's date, and cannot filter most data by machine. Phrase for that.

WeakStrongWhy
How's OEE?What was our average OEE over the last 7 days, and is it trending up or down?Names a window and asks for direction.
What happened last Tuesday?What were the top downtime reasons over the last 7 days?Calendar days are approximated; rolling windows are exact.
Show OEE for the last 6 monthsShow me OEE trends for the last 30 daysMetric history caps at 30 days; longer requests shorten silently.
Why is Press 4 slow?Which equipment had the most downtime over the last 30 days, and why?Ranked lists name equipment; most data has no machine filter.
And last week?What was our scrap rate over the last 14 days, and is it getting better or worse?Ask does not remember your previous question.
How much money are we losing?What is our cost of quality over the last 30 days, and which figures use assumed unit costs?Surfaces planning estimates.
What do you track for us?Give me an overview of what data you have for this plantYour plant's data, not the product catalogue.
Optimize our scheduleWhat was our measured changeover time and utilization over the last 7 days?No schedule optimizer is connected.
Is quality OK?Is our scrap rate in statistical control? What is the Cpk?Names a method with a testable answer.

Follow-up patterns

Every Ask question is answered on its own, so restate the metric, window and scope in each follow-up. AI Chat keeps recent turns, but draws no charts and its thread is shared by everyone on that facility.

Illustrative example, sample plant, not a customer
You

What was our average scrap rate over the last 7 days, and is it getting better or worse?

KaizenFlow AI

Scrap rate averaged 3.1% over the last 7 days against a 2.5% target, and the trend is rising, so it is getting worse. The chart plots the full 7-day trend with the target as a dashed line.

Basis: scrap rate readings, rolling 7 days back from now, full-window average

You

Is our scrap rate over the last 30 days in statistical control, and what is the Cpk?

KaizenFlow AI

No. 4 points fall outside the control limits, all in the last 5 days. Cpk is 0.82. No spec limits were given, so limits are based on the target. Because points are out of control, check your notifications for an SPC alert on scrap rate.

Basis: SPC analysis of scrap rate, rolling 30 days, target-based limits

A team prompt library and shift routines

Keep a shared list of prompts checked against known answers, each with a fixed window. Start from the prompt library. A simple routine: at shift start, the handover for the last 8 hours; mid-shift, anomalies in the last 4 hours; weekly, top downtime reasons and scrap trend over 7 days.

When not to use the AI

  • Safety-critical steps. Follow the controlled SOP itself, not a summary of it.
  • Real-time control. No AI tool writes to equipment, and polled data can be minutes old.
  • Financial reporting. Dollar figures are estimates until a measured result is recorded, and a recorded figure is only as good as the measurement behind it.
  • Judging individuals. Plant data answers questions about processes, not people.
  • Predictive-maintenance lists, value-stream maps and crew optimization. These answers can currently contain content not backed by your data. Use the downtime-risk forecast and reliability questions instead.

Troubleshooting

Run the checks in order; most problems are role, window or missing data.

"Sorry, I could not process that question. Please try again."

  1. Check your role. This is what a Viewer sees on every question.
  2. Check the length. Questions must be 3 to 2,000 characters. Pasted logs often exceed this.
  3. Check the pace. A burst of requests from you or your organization can hit the rate limit.

Fix: get the Engineer role or above, shorten the question, or retry after a minute. If the message instead reads Could not reach the AI service. Please check your connection and try again., the request never arrived: check your network and reload.

"Still loading your facility, please try again in a moment."

  1. Wait for the app to load. Ask uses the facility selected across the app and has no picker of its own.

Fix: reload, confirm the facility selector shows your plant, and ask again.

An amber box instead of an answer

  1. Read the heading. It names the cause: AI billing limit reached, AI provider rate limit, AI configuration problem, Daily AI budget reached, AI service timed out, or AI service unavailable.
  2. Match the owner. Billing, configuration and budget need an administrator. Rate limits and timeouts usually clear on their own.

Fix: retry shortly for rate limits and timeouts. A daily budget resets at the next UTC day; its limit is set in server configuration, not in Settings. The page never substitutes demo answers.

"I couldn't finish gathering the data for that just now."

  1. Count the topics. A question spanning downtime, quality, energy and plan needs many data rounds. The AI gets a fixed number of rounds, and a data step that runs past its time limit ends in this message.

Fix: split it into one question per topic, each with a named window.

The spinner runs for a long time

  1. Know what is happening. Analyzing your question with AI tools... stays until the whole answer is ready; nothing is streamed.
  2. Check the scope. Broad questions trigger several tool rounds, each with its own time limit.

Fix: narrow the question. If it still hangs, reload and ask a smaller version.

"No data found for this metric and period." or a missing figure

  1. Ask what data exists. Give me an overview of what data you have for this plant.
  2. Check connectors. Are all our data connectors healthy? A stale connector explains most gaps.
  3. Check the window against retention. Metric readings are kept only for your organization's retention period, and some questions cap shorter.

Fix: connect or repair the source (see integrations). By design, the data tools are tested against an empty account and must return nothing they cannot back, or say the data is absent, so missing energy data reads unavailable, not 0. The known exceptions are listed under Best practices and Other known limits.

Illustrative example, sample plant, not a customer
You

What's our energy per unit and carbon footprint this month?

KaizenFlow AI

That figure is unavailable. No power or energy readings are connected for this facility, so I can't compute energy per unit or carbon. Connecting a power series through OPC-UA or MQTT would enable it.

Basis: energy analysis returned no data for this facility; no figure reported

The number does not match my dashboard or report

  1. Compare windows. Ask counts rolling hours back from now in UTC, not calendar days in plant time.
  2. Check for a cap. A 6-month OEE request is answered over 30 days without a warning. Read the window stated in the answer.
  3. Check the facility. In a multi-plant organization, confirm with your admin which facility each connector's data is assigned to.
  4. Check the unit. Throughput and part counts are per-hour rates, not raw counter totals.

Fix: restate the question with a window that matches the report.

TEEP or utilization shows UNKNOWN

  1. Check the shift schedule. Without one, scheduled time cannot be computed, so the AI reports UNKNOWN rather than 0%.

Fix: configure the shift schedule. The OEE and TEEP guide explains why TEEP needs it.

A dollar figure looks too round

  1. Look for the planning-estimate note. Without your unit economics, answers use $10 cost per unit, $25 revenue per unit and $5,000 per downtime hour.
  2. Check for a rate. With too little throughput history, projections are withheld rather than guessed.

Fix: set the facility's cost per unit, revenue per unit and cost per downtime hour.

My follow-up ignored the previous answer

  1. Check the page. Ask answers every question independently.

Fix: restate the metric, window and scope, or use AI Chat for multi-turn drill-downs.

Other known limits

  • With zero downtime records, reliability reports 100% availability. Read it as "no downtime data".
  • Alert totals stop at 50 in a busy window. Ask for a shorter window.
  • Root cause analysis ignores a requested window today.
  • A chart with no data points is not drawn, and AI Chat draws no charts.
  • Your interface language is not passed to the AI and its instructions are in English, so expect answers in English.

Privacy, security and data handling

SCOPE

Your organization comes from your sign-in, never from the question. The AI's queries of your plant data are filtered to it, and by default to your selected facility, and facility comparisons cover only your own facilities.

PERMISSIONS

KaizenFlow has four roles: Admin, Manager, Engineer and Viewer. Asking, chatting and running analysis need Engineer or above; Viewer accounts cannot use them. Managers and Admins see LLM Usage and Cost: feature, model, tokens, cost and latency per call, without prompt text.

WHAT REACHES THE MODEL

The model receives system instructions, your question and the data tool results; in AI Chat, also recent turns. The AI never connects to your machines or connectors: it reads the KaizenFlow database they fill, and no AI tool writes to equipment.

Personal data

There is no automatic redaction of personal data before a model call. If you connect workforce, skill-matrix or safety records that contain names, those names can appear in what the AI reads. Connect them only if that fits your policy.

The AI provider is set in server configuration. Our security page states that AI analysis runs through OpenAI and Anthropic and that we do not use your data to train third-party models. That commitment rests on the providers' terms, not on a setting inside the app.

RETENTION

  • Metric readings: kept for the retention period your organization sets in Settings, with options from 90 days to Unlimited.
  • Ask history: on the page in your browser tab only, lost on reload.
  • AI Chat: the last 20 messages per facility, for up to 7 days after the latest message, shared across your organization.
  • AI call records: model calls are logged, including the prompt and response text, which can be shortened.

An optional daily AI spending limit, set in server configuration and counted per organization, is off unless configured; once reached, it pauses AI analysis until the next UTC day; Ask then shows Daily AI budget reached. For the full control set, see the IT review guide.

FAQ

Which roles can use Ask Anything?

Engineer, Manager and Admin. Viewers see Ask Anything in the menu, but their questions return a generic error.

Does Ask Anything remember my previous question?

No. Each question is answered on its own, and on-screen history is lost on reload. Restate the metric, window and scope every time. AI Chat keeps recent turns.

Is my AI Chat conversation private?

No. Chat history is kept per facility, not per user. Everyone in your organization working on that facility shares one thread, and Clear Chat clears it for all of them.

Does Ask Anything use the nine specialists?

No. The nine specialists power the AI Suggestions analysis, and an organization can switch all but one of them off. Ask Anything and AI Chat use a single assistant that calls the data tools directly.

Can the AI change machine settings or write to my ERP?

No. No AI tool writes to equipment or to connected systems such as your ERP. Inside KaizenFlow, the AI can turn an existing suggestion into a work order if a question asks for it, with no separate confirmation step, and an SPC question that finds out-of-control points can create an SPC alert.

How current is the data?

Polled connectors sync every 5 minutes, MQTT data arrives continuously, and rollups are recomputed every 30 minutes. The AI reads whatever has reached the database when you ask.

How far back can it look?

It depends on the question: metric trends cap at 30 days, downtime at 90 days, energy and safety at a year. Metric history is also limited to your organization's data retention setting.

Are suggestion savings calculated from my data?

The estimate is written by the AI model from per-machine baselines in your data, then checked for arithmetic and plausibility. No formula recomputes it. Count a saving only once a measured result is on record.

Which languages does it support?

The interface is available in five languages: English, Spanish, German, Chinese and Japanese. Your interface language is not passed to the AI and its instructions are in English, so expect answers in English.

Can I export or share an answer?

Not yet. There is no export, copy or feedback control, and Ask history is lost on reload. Copy the text from the page to keep it.

Does KaizenFlow train AI models on my data?

Our security page states that we do not use your data to train third-party models. That commitment rests on the providers' terms, not on a setting inside the app.

What is the verified savings ledger?

A report of predicted and verified dollars for each suggestion, with realization and portfolio totals, built for finance to reconcile. At design-partner stage it is a report in the platform, not yet a screen in the app. Its verified dollars are the actual savings recorded on each card, which may be a typed-in figure, the prefilled AI estimate or a calibration figure, so check each line against a measurement before you report it.

Ask it about your own lines

The eight-week pilot connects one facility and the systems it already runs on, then ends in a verified savings report reconciled with your finance team.