Soft control
Prompt
Expresses desired behaviour.
Evidence-gated analysis
Reliability in GenAI analysis does not come from bounding the model to the right material and giving it good instructions. It comes from a system that governs which claims may be formed, on what evidence they may be accepted, and when the analysis must be blocked.
Soft control
Expresses desired behaviour.
Soft control
Orchestrates model activity.
Assurance
Binds claims to approved sources.
Hard control
Accepts, blocks, or escalates the output.
I have noticed that a surprisingly strong belief persists around artificial intelligence, and generative AI in particular: that following a familiar set of principles will automatically make a GenAI analysis accurate and almost free from the risk of serious error.
Even full adherence to these principles does not guarantee analytical accuracy. It does not necessarily make the result more reliable than a single, well-prompted rapid analysis of the same material.
The methods above may suffice for clearly bounded, low-risk tasks that mainly extract information from source material. They do not, by themselves, constitute an adequate quality-assurance architecture for analysis that requires abstract reasoning, is complex, or contains multiple dependencies among observations, interpretations, and conclusions.
Connecting several agents does not automatically improve the result either. In some cases a multi-agent system can even underperform a single agent. Any benefit depends on the division of labour, the agents’ independence, the available evidence, the verification mechanisms, and the nature of the task—not on the number of agents. If every agent is built on the same principles and uses similar material, similar reasoning, and similar checking logic, errors are not necessarily corrected. They may be repeated, amplified, or given the appearance of a formally valid but false consensus.
If the aim is to produce, repeatedly, highly accurate analysis grounded in given background information and defined research questions, the target level of accuracy must first be defined in a measurable way. A level above 95 percent could mean, for example, that more than 95 percent of the analysis’s factual claims can be traced to approved evidence that genuinely supports them, and to the criteria set. Even that does not, by itself, guarantee coverage, the logical soundness of conclusions, or that the research questions have been answered. Those properties require their own acceptance criteria and measures.
When the analysis is abstract in character—for example strategic, concerned with defining a company’s needs, or concerned with assessing the effects of climate change—more is required of both the analytical architecture and the cognitive model behind it. GenAI may be an important part of that system. It should not, on its own, control the selection of evidence, the progress of the analysis, the acceptance of quality, or the publication decision.
These stages do not guarantee an error-free analysis. They create the conditions for analysis that is traceable, repeatedly reliable, and stoppable when evidence is insufficient.
To avoid leaving the argument as rhetoric alone, I tested how the accuracy of an AI-assisted analytical apparatus might be measured. The comparison held the same materials, the same knowledge base, and the same output format. The difference is whether the rule lives only in a prompt, or is enforced by the system.
I prepared nine background-material packages from poor to good. Three represented clear cases (BLOCK, WARN, and GO). The remaining six were grey-area packages at the boundaries. For each package, a correct answer was created together with experts.
Ten analyses of each designed package with the apparatus, and ten of each of three randomly selected authentic customer materials. Experts produced reference answers against the same predefined criteria, without seeing apparatus output. A prompted model then received a sample report, the materials, the knowledge base, and an instruction to use only those materials.
Three anchors—BLOCK, WARN, GO—and six grey-area packs at the boundaries.
Ten runs for each designed package and ten for each of three customer materials.
Ten matched analyses for each designed package, same materials and knowledge base.
Relative reduction versus prompted AI on the designed packages.
Repeated runs were used to observe both accuracy and variance. Only the enforcement of the rule changed.
Experts establish the reference against the same criteria, without seeing apparatus output.
Each designed and customer package is run ten times through the analysis apparatus.
The model receives a sample report, the materials, the knowledge base, and a source restriction.
Items are classified as correct or incorrect; mean, range, and median are calculated.
Designed packages with the apparatus: mean 94%, range 85–98%, median 96%. Customer materials: mean 93%. Prompted AI on the same designed packages: mean 71%.
Mean accuracy · median 96% · observed range 85–98%
Mean accuracy · median 75% · observed range 59–86%
(29 − 6) ÷ 29 = 79%
Accuracy improved by 23 percentage points. Expressed as error, the mean fell from 29% to 6%—a 79% relative reduction in errors. The 79% figure is not a percentage-point change.
Band is the observed range. Disc is the mean. Tick is the median. Scale starts at 50%.
The BLOCK material had the widest range in both conditions: 85–95% for the apparatus (mean 91%) and 59–76% for prompted AI (mean 68%). WARN and GO remained tight for the apparatus (95% and 97%, each ±2 percentage points). In the prompted condition, WARN and GO each obtained a mean of 76%. None of the customer materials was BLOCK-level material.
Designed packages only. Whiskers show the full observed range. Customer packs are omitted because none were BLOCK-level.
Mean on the first line. Range and median (M) on the second.
| Condition | Overall | BLOCK | WARN | GO |
|---|---|---|---|---|
| Apparatus · designed | 94%85–98 · M 96 | 91%85–95 | 95%93–97 | 97%95–99 |
| Apparatus · customer | 93%88–95 · M 93 | — | — | — |
| Prompted AI · designed | 71%59–86 · M 75 | 68%59–76 | 76%73–86 · M 76 | 76%73–86 · M 76 |
A translator from company reality into structured signals, testable need hypotheses, and safe discovery paths.
The system is designed to keep evidence, interpretation, hypotheses, and actions distinguishable. Weak or unknown evidence remains gated; it is not converted into confirmed buying intent.
Three related presentations describe the cognitive model, evidence-governed flow, foresight methodology, and an example of the knowledge-base material used by the system.
The answer is only as trustworthy as its evidence path—not as impressive as its model.