Skip to content
FAVIERPaulPublic

About

Data Botch is a project which aims to automate data identification, collection, and processing to estimate collective human behaviors as tourism flows. Using agentic orchestration, it transforms raw datasets into actionable insights ideal for research, urban planning, or marketing.

Resources

Stars

1 star

Watchers

0 watching

Forks

Latest commit

 

History

79 Commits

Folders and files

Repository files navigation

DataBotch - Open Data to Actionable Tourism Knowledge

Python Mode Data Output

Why this project exists

DataBotch turns fragmented public signals into a clear, traceable estimation of tourism activity.

The project is explicitly OSINT-first: it relies on open/public sources (OpenStreetMap, Data Tourisme, Wikipedia pageviews, transport and weather data, etc.) and keeps a full audit trail of collection and reasoning.

In one sentence

A non-deterministic ETL backbone, augmented by specialized agents, to transform raw open data into explainable knowledge for decision-making.


From raw data to decision-ready knowledge

flowchart TB
    Q[<b>Input question</b>\n <i>How many tourists in a specific geographic space?</i>]
    Q --> S[<b>Identify raw data sources</b><i>\nOSM, Data Tourisme, SNCF, Meteo, Wikipedia,\nmunicipal open data portals, regional/national portals<i>]
    S --> R[<b>Collect raw data</b>\nHeterogeneous formats and granularities]
    R --> T[<b>Standardize</b>\nCleaning, schema alignment, entity resolution]
    T --> C[<b>Cross sources and Computation</b>\nFusion, inference, nowcasting model, consistency checks]
    C --> A[<b>Answer</b>\nBest possible precision, explicit confidence/caveats]

    classDef d fill:#dbeafe,stroke:#1d4ed8,color:#0f172a;
    classDef i fill:#d1fae5,stroke:#0f766e,color:#052e16;
    classDef k fill:#fef3c7,stroke:#b45309,color:#451a03;
    classDef a fill:#fee2e2,stroke:#b91c1c,color:#450a0a;

    class R d;
    class T i;
    class C k;
    class A a;
Loading

Adapted from Grundstein (2003) about Knowledge Theory.


Why this would not scale without agents

A classical ETL pipeline is designed for stable, pre-defined source lists. It cannot do three things that are essential here:

  1. Adapt source selection per query: which sources are relevant depends on the city, its size, and data availability. Agents decide dynamically which sources to query rather than hitting all of them blindly.
  2. Standardize heterogeneous sources on the fly: schemas, granularities, and semantics differ across open data providers. Agents normalize and reconcile them rather than failing or requiring manual mapping.
  3. Improve cross-source inference: combining signals from SNCF traffic, Wikipedia views, accommodation capacity, and weather is not a deterministic join. Agents weigh, arbitrate, and flag inconsistencies.

Demo

Live demo runs against the Kimi gateway with KIMI_FUSION_MODEL=kimi-k2.6 and FUSION_PROMPT_VARIANT=legacy — the values shipped in .env.example. Rationale and conditions for re-challenging the default: docs/dev/fusion-model-choice.md.


Where we stand

The deterministic backbone is operational: city resolution, multi-source collection, structured logbook, and disk caching are all in place and tested. Every external call is logged and every result is reproducible.

The pipeline is built around three non-negotiable principles: full traceability (no black-box output), graceful degradation (one failing source never stops the run), and OSINT ethics (all sources are public, auditable, and inspectable).

About

Data Botch is a project which aims to automate data identification, collection, and processing to estimate collective human behaviors as tourism flows. Using agentic orchestration, it transforms raw datasets into actionable insights ideal for research, urban planning, or marketing.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages