Cover image for 'From Open Data to AI-Ready Data'

From Open Data to AI-Ready Data

This report is jointly published by OCF and the Taiwan Office/Global Innovation Hub of the Friedrich Naumann Foundation for Freedom. Drawing on case studies from Spain, Germany, Cambodia, and France, it explores how government open data can become machine-readable, legally reusable, and useful for citizen-government dialogue in the age of AI.

About the Report

As artificial intelligence develops rapidly, traditional government open data must evolve toward AI-ready open data. This report examines formats, metadata, licensing, privacy protection, cross-agency coordination, and public-private collaboration to show how data stewards can make open data easier for AI tools to understand and responsibly reuse.

The report is also intended as a communication resource between government and civic communities. It helps public agencies identify practical paths for data governance, while giving advocates, civic technologists, and open-source communities concrete examples for dialogue with agencies at different stages of readiness.

Research Outputs / Document Versions

This page will continue collecting research outputs, document versions, and collaboration entry points for the AI-Ready Open Data study. Different versions serve different purposes:

Online Preview

Case Studies

The research presents five international cases, showing how AI-ready open data can begin under different institutional and technical conditions:

Spain
ImpulsaDATA
Interviewee: Carlos Alonso Peña, Director of the Spanish Data Office Division
A companion-style support program that helps public institutions navigate overlapping regulations. Its diagnose-gap-act method helps agencies identify which high-value datasets can be safely released.
Germany
The German Commons
Interviewees: Lukas Gienapp and Martin Potthast, Co-Founders
The project consolidates 41 clearly licensed German-language data sources and releases an automated code repository covering data cleaning, personally identifiable information removal, and other workflow steps for fully open German LLM development.
Cambodia
Open Development Cambodia
Interviewees: Wen-Ling (Amy) Lai, Vimoil Ourn, and Try Thy
In a context where transparency and data protection laws are still developing, ODC provides neutral technical assistance, builds trust with government partners, and strengthens data management, integration, and sharing capacity.
Germany
Parla
Interviewee: Ingo Hinterding, Lead Prototyping at CityLAB Berlin
Jointly developed by the Senate Chancellery of Berlin and CityLAB Berlin, Parla extracts more than 11,000 public parliamentary documents and uses a generative AI chatbot to improve administrative search and responses to citizen inquiries.
France
CroissantLLM
Interviewee: Manuel Faysse, Co-Founder
Built on cross-ministry datasets from France's government open data initiative platform, CroissantLLM trained a 1.3-billion-parameter, truly bilingual French-English open language model with an emphasis on transparency and performance.

Key Takeaways

Adapted from the report's Executive Summary, these points offer a quick-read view of what AI-ready open data requires:

01
Design AI usability at the point of release
Government open data must move beyond publication alone, with formats, structures, and reuse contexts designed for AI training and applications.
02
Reduce conversion costs
Open data is valuable because it is accessible and often carries lower privacy and copyright risk, but poor formats still force developers to spend substantial effort cleaning and converting it.
03
Make machine-readable formats the default
CSV works well for numerical and statistical data, while Markdown is better suited to text-based materials that need to preserve structure, context, and logic.
04
Align law, licensing, and governance
Open data policy must be coordinated with data protection, intellectual property rules, and international open licensing standards to reduce institutional risk and developer transaction costs.
05
Keep data persistent and reliable
Public agencies should shift from reactive disclosure to proactive release through centralized platforms and resilient, multi-tier storage strategies.
06
Build trust before asking for more openness
Advocacy can begin by understanding government needs and choosing lower-sensitivity, high-impact topics as practical entry points for collaboration.
07
Start with civil servants' pain points
Non-adversarial advocacy and tailored technical support can show how data governance improves internal efficiency instead of adding new burdens.
08
Treat AI readiness as democratic infrastructure
Structured and interoperable data helps public servants improve administration while enabling citizens to analyze government records and strengthen oversight with AI tools.
09
Address privacy before data becomes corpus
Because data is difficult to remove once it enters a model, agencies should test for personal data before release or training and use context-preserving substitutes for sensitive information.
10
Use local corpora to sustain cultural diversity
High-quality local data can reduce linguistic marginalization and temporal bias, preserve Taiwan's context, and support a more resilient open data ecosystem.

Activities and Presentations

The following records collect activities, presentations, discussions, and public references related to this research. Event notes, slides, shared notes, and follow-up articles can be added here over time.

2026

g0v Summit Forum: Open Data in the Age of AI - Taiwan's Next Step in Co-Creation from International Public-Private Collaboration

2026/05/24
Time: 10:00-11:00
Venue: g0v Summit 2026, R3

This forum focused on the next-stage challenges of open data in the age of AI, drawing from international public-private collaboration experiences to discuss how Taiwan can promote more usable, governable, and collaborative AI-ready open data.

Consensus-Building Workshop for the Next Stage of Open Data Advocacy

2026/04/14
Format: Workshop / discussion case

This workshop served as an important case for next-stage open data advocacy and AI-ready open data discussions, helping organize observations and consensus from communities, advocates, and stakeholders on future data governance directions.

2025

From Open Data to AI-Ready Data Webinar: Online Presentation

2025/12/29
Time: 17:00-19:00
Format: Online event

This online presentation shared the research findings, international case observations, and follow-up discussion directions from the From Open Data to AI-Ready Data report.

Related Links

Open Technology and Advocacy

Publication and License

This report is jointly published by the Open Culture Foundation (OCF) and the Taiwan Office/Global Innovation Hub of the Friedrich Naumann Foundation for Freedom (FNF). It is released under the CC BY 4.0 license.

財團法人開放文化基金會 (OCF)
電話:+886-2-2395-1505
Email:[email protected]
地址:105057 臺北市松山區八德路四段 636 之 1 號