This report is jointly published by OCF and the Taiwan Office/Global Innovation Hub of the Friedrich Naumann Foundation for Freedom. Drawing on case studies from Spain, Germany, Cambodia, and France, it explores how government open data can become machine-readable, legally reusable, and useful for citizen-government dialogue in the age of AI.
As artificial intelligence develops rapidly, traditional government open data must evolve toward AI-ready open data. This report examines formats, metadata, licensing, privacy protection, cross-agency coordination, and public-private collaboration to show how data stewards can make open data easier for AI tools to understand and responsibly reuse.
The report is also intended as a communication resource between government and civic communities. It helps public agencies identify practical paths for data governance, while giving advocates, civic technologists, and open-source communities concrete examples for dialogue with agencies at different stages of readiness.
This page will continue collecting research outputs, document versions, and collaboration entry points for the AI-Ready Open Data study. Different versions serve different purposes:
Official reading and citation
The fully designed version for reading, downloading, citation, and external sharing.
Formal organization and long-term maintenance
This version supports formal organization, version control, long-term maintenance, issue discussion, and document evolution.
Community reading, discussion, and feedback
This version is suitable for community reading, shared notes, suggested edits, and additional case contributions.
The research presents five international cases, showing how AI-ready open data can begin under different institutional and technical conditions:
Spain
Interviewee: Carlos Alonso Peña, Director of the Spanish Data Office Division
A companion-style support program that helps public institutions navigate overlapping regulations. Its diagnose-gap-act method helps agencies identify which high-value datasets can be safely released.
Germany
Interviewees: Lukas Gienapp and Martin Potthast, Co-Founders
The project consolidates 41 clearly licensed German-language data sources and releases an automated code repository covering data cleaning, personally identifiable information removal, and other workflow steps for fully open German LLM development.
Cambodia
Interviewees: Wen-Ling (Amy) Lai, Vimoil Ourn, and Try Thy
In a context where transparency and data protection laws are still developing, ODC provides neutral technical assistance, builds trust with government partners, and strengthens data management, integration, and sharing capacity.
Germany
Interviewee: Ingo Hinterding, Lead Prototyping at CityLAB Berlin
Jointly developed by the Senate Chancellery of Berlin and CityLAB Berlin, Parla extracts more than 11,000 public parliamentary documents and uses a generative AI chatbot to improve administrative search and responses to citizen inquiries.
France
Interviewee: Manuel Faysse, Co-Founder
Built on cross-ministry datasets from France's government open data initiative platform, CroissantLLM trained a 1.3-billion-parameter, truly bilingual French-English open language model with an emphasis on transparency and performance.
Adapted from the report's Executive Summary, these points offer a quick-read view of what AI-ready open data requires:
01
Government open data must move beyond publication alone, with formats, structures, and reuse contexts designed for AI training and applications.
02
Open data is valuable because it is accessible and often carries lower privacy and copyright risk, but poor formats still force developers to spend substantial effort cleaning and converting it.
03
CSV works well for numerical and statistical data, while Markdown is better suited to text-based materials that need to preserve structure, context, and logic.
04
Open data policy must be coordinated with data protection, intellectual property rules, and international open licensing standards to reduce institutional risk and developer transaction costs.
05
Public agencies should shift from reactive disclosure to proactive release through centralized platforms and resilient, multi-tier storage strategies.
06
Advocacy can begin by understanding government needs and choosing lower-sensitivity, high-impact topics as practical entry points for collaboration.
07
Non-adversarial advocacy and tailored technical support can show how data governance improves internal efficiency instead of adding new burdens.
08
Structured and interoperable data helps public servants improve administration while enabling citizens to analyze government records and strengthen oversight with AI tools.
09
Because data is difficult to remove once it enters a model, agencies should test for personal data before release or training and use context-preserving substitutes for sensitive information.
10
High-quality local data can reduce linguistic marginalization and temporal bias, preserve Taiwan's context, and support a more resilient open data ecosystem.
The following records collect activities, presentations, discussions, and public references related to this research. Event notes, slides, shared notes, and follow-up articles can be added here over time.
2026/05/24
Time: 10:00-11:00
Venue: g0v Summit 2026, R3
This forum focused on the next-stage challenges of open data in the age of AI, drawing from international public-private collaboration experiences to discuss how Taiwan can promote more usable, governable, and collaborative AI-ready open data.
2026/04/14
Format: Workshop / discussion case
This workshop served as an important case for next-stage open data advocacy and AI-ready open data discussions, helping organize observations and consensus from communities, advocates, and stakeholders on future data governance directions.
2025/12/29
Time: 17:00-19:00
Format: Online event
This online presentation shared the research findings, international case observations, and follow-up discussion directions from the From Open Data to AI-Ready Data report.
Open Technology and Advocacy
This report is jointly published by the Open Culture Foundation (OCF) and the Taiwan Office/Global Innovation Hub of the Friedrich Naumann Foundation for Freedom (FNF). It is released under the CC BY 4.0 license.