DataBotch turns fragmented public signals into a clear, traceable estimation of tourism activity.
The project is explicitly OSINT-first: it relies on open/public sources (OpenStreetMap, Data Tourisme, Wikipedia pageviews, transport and weather data, etc.) and keeps a full audit trail of collection and reasoning.
A non-deterministic ETL backbone, augmented by specialized agents, to transform raw open data into explainable knowledge for decision-making.
flowchart TB
Q[<b>Input question</b>\n <i>How many tourists in a specific geographic space?</i>]
Q --> S[<b>Identify raw data sources</b><i>\nOSM, Data Tourisme, SNCF, Meteo, Wikipedia,\nmunicipal open data portals, regional/national portals<i>]
S --> R[<b>Collect raw data</b>\nHeterogeneous formats and granularities]
R --> T[<b>Standardize</b>\nCleaning, schema alignment, entity resolution]
T --> C[<b>Cross sources and Computation</b>\nFusion, inference, nowcasting model, consistency checks]
C --> A[<b>Answer</b>\nBest possible precision, explicit confidence/caveats]
classDef d fill:#dbeafe,stroke:#1d4ed8,color:#0f172a;
classDef i fill:#d1fae5,stroke:#0f766e,color:#052e16;
classDef k fill:#fef3c7,stroke:#b45309,color:#451a03;
classDef a fill:#fee2e2,stroke:#b91c1c,color:#450a0a;
class R d;
class T i;
class C k;
class A a;
Adapted from Grundstein (2003) about Knowledge Theory.
A classical ETL pipeline is designed for stable, pre-defined source lists. It cannot do three things that are essential here:
- Adapt source selection per query: which sources are relevant depends on the city, its size, and data availability. Agents decide dynamically which sources to query rather than hitting all of them blindly.
- Standardize heterogeneous sources on the fly: schemas, granularities, and semantics differ across open data providers. Agents normalize and reconcile them rather than failing or requiring manual mapping.
- Improve cross-source inference: combining signals from SNCF traffic, Wikipedia views, accommodation capacity, and weather is not a deterministic join. Agents weigh, arbitrate, and flag inconsistencies.
Live demo runs against the Kimi gateway with KIMI_FUSION_MODEL=kimi-k2.6 and FUSION_PROMPT_VARIANT=legacy — the values shipped in .env.example. Rationale and conditions for re-challenging the default: docs/dev/fusion-model-choice.md.
The deterministic backbone is operational: city resolution, multi-source collection, structured logbook, and disk caching are all in place and tested. Every external call is logged and every result is reproducible.
The pipeline is built around three non-negotiable principles: full traceability (no black-box output), graceful degradation (one failing source never stops the run), and OSINT ethics (all sources are public, auditable, and inspectable).