Ask the question in English. Get the number, and the query that produced it.
Every question went through the one person who could write the query.
Operational data arrived from several systems on different schedules and in different shapes. It landed in a warehouse, and from there any real question — what moved, where, compared with when — required somebody who knew both the business and the schema. That person becomes a queue. The questions that get asked are the ones worth waiting for, which is a quiet tax on how a company understands itself.
The brief was not “add a chatbot”. It was to make the warehouse answer for itself, without putting a confident wrong number in front of anyone.
Pipelines in, an assistant on top, and guards around the assistant.
Ingestion and the warehouse
Scheduled pipelines pull from upstream systems and file uploads, normalise into a modelled schema, and record what arrived and when. A freshness check answers “is this number current?” before anybody acts on it.
Questions in plain English
A question becomes SQL against the real schema and returns both the figure and a written summary of what it means. Schema and worked examples are assembled into the prompt and cached, rather than pasted in by hand.
It holds a conversation
“Now break that down by region” works, because recent turns carry into the next query. Follow-ups are how people actually interrogate data; one-shot questions are a demo.
Guarded before it runs
Generated SQL is validated before execution, execution is read-only and row-bounded, and a runtime error triggers an automatic retry rather than an error message. The model is never the last line of defence.
It learns the vocabulary
When somebody corrects an answer, the question and the right query are kept as an example and the prompt is rebuilt. The assistant gets better at how this company speaks rather than staying where it launched.
It defers to agreed numbers
Existing reports are searchable by keyword and can be run directly. Where finance has already agreed a definition, the assistant uses it instead of inventing a second version of the truth.
A tool server AI clients log into
An MCP server with a full OAuth flow — client registration, authorisation, tokens, JWT verification and sessions — so an AI assistant can use the warehouse as a set of tools without anyone pasting a credential into a chat window.
Dashboards, same identity
Embedded analytics behind single sign-on, so the platform and the reporting tool are one login and one set of permissions rather than two.
Access control
Managed identity with multi-factor authentication, user administration, roles, and scoping so a user sees the slice of data their role allows — enforced on the query, not in the interface.
A wrong answer that looks right is worse than no answer.
A model that writes plausible SQL against the wrong column returns a number nobody can tell is false. It does not look like a failure. It looks like a result, and it travels into a meeting.
The query is checked rather than trusted, and runs against a read-only path with a row ceiling. A question cannot become a write, and it cannot become a table scan that takes the warehouse down.
The query that produced the figure is available with it. Anyone who knows the schema can check the answer in seconds, which is what makes the tool trustworthy to the people who were previously the queue.
Every correction becomes an example the next answer is built from. The failure mode of a system like this is staying exactly as good as it was on launch day; feeding corrections back is what prevents that.
The client is under NDA and the system holds their commercial data, so there are no screenshots here and the sector is not named. We demonstrate it live on a screen-share, including the assistant answering against a real database.