Start here

How it works

Every message passes through the same eight stages. Language models are used where judgement is needed, such as reading a question or choosing a dataset. Every fact, count and query comes from deterministic code that checks the model’s work.

Eight stages from route to summary; chat and catalogue questions leave after the first stage; checks can stop the pipeline with a clarifying question
The pipeline. Chat and catalogue questions leave after routing. Spatial questions go through all eight stages. Wherever a model is used, deterministic rules check its output, and the plan can end in a clarifying question instead of an answer.

Stage by stage

StageWhat happensWhat you see
1 RouteThe message is sorted into chat, a question about the catalogue, or a spatial question. Its language is noted. A question asking for two different kinds of thing is split into parts, each written as a complete question with its own places and conditions.Nothing for chat; otherwise the first steps in the thinking panel.
2 UnderstandThe question is broken into concepts, named places, grouping (per district, neighbourhood or municipality), time conditions and required conditions. Places are matched to official codes by fixed rules.“Recognised places: …”
3 Find dataWhen a place is named, only datasets with data in that municipality are searched. Each concept is then searched separately by meaning, in Dutch and English, and the results are merged.Candidate datasets with scores
4 SelectDatasets get roles: the target that is counted, constraints that restrict it, and a join partner it is linked to. District boundaries are attached when needed.“Selected: …” with a reason
5 DatasetRelated registers of the same kind in the asked area, including ones the search did not rank, are combined or one is chosen, based on measured overlap. Each condition is bound to a real column and a real value.Shown on the Stage
6 PlanA structured plan is written, never SQL. A per-area question that comes back grouped on a missing column is asked again once, to group by the official boundaries. More than thirty checks then verify every name, value, threshold and place, repair what can be repaired (for example, holding a per-area plan to the named place), and stop what cannot.The plan, or a clarifying question
7 ExecuteCode turns the plan into a safe database query and runs it. A suspicious zero is checked before it is reported.The query and “returned N of M”
8 ExplainThe answer is written from the actual rows and the executed query, with a scope note stating what was counted and where.The streamed answer and downloads

Two design rules that shape everything

Judgement in the model, facts in code

What a question means and which dataset fits are judgements. Column names, values, coverage, counts and place boundaries are facts, and facts are always checked by code against the real data.

Repeatable decisions

Decision steps use fixed settings with no random sampling, which keeps plans reproducible and lets every change be tested against a suite of real municipal questions before it ships.

Further reading on the GeoAI blog: One question, eight stops and Why our AI isn’t allowed to write SQL.