About the Workshop

As foundation models scale, available training data sources have rapidly depleted. However, several forms of valuable data artifacts such as medical records, legal, and financial documents are restricted from use in model training due to their sensitive nature. In addition, the strong reasoning capabilities in current generative models have opened the possibility for highly personalizable AI applications but these remain bottlenecked by limited access to high quality user data. Hence, it is of immense value to responsibly unlock these data sources (for example: using data transformation or constrained training paradigms) or to generate synthetic alternatives. In this workshop, we aim to bring together domain experts in data, privacy, model training, and legal policy, to advance the frontier of responsibly leveraging such sensitive data with foundation models.

Topics of interest include (but are not limited to):

  1. Data Transformation: De-identification, Anonymization, Pseudonymization.
  2. Synthetic Data Generation: Controlled Regeneration, Data Diversity.
  3. Novel Training Paradigms: DP, Federated Learning, Architectural Solutions.
  4. Evaluation & Auditing: Privacy attack benchmarks, Utility-Privacy tradeoffs.
  5. Policy: Compliance, New regulations on data sharing.

Call for Papers

We invite long papers with novel research contributions (up to 8 pages long) as well as short papers (up to 4 pages) reflecting preliminary studies or negative results.

Submissions are managed via OpenReview. Accepted papers are non-archival, and concurrent submissions are allowed. Please follow the COLM 2026 template.

Key Dates

All deadlines are 23:59 AoE (anywhere on earth)

  • Submissions open: May 27, 2026
  • Submission deadline: June 23, 2026 June 28, 2026 AoE
  • Acceptance Notifications: July 24, 2026
  • Workshops day at COLM: October 9, 2026

Workshop Schedule

08:45 - 09:00 Welcome and Opening Remarks
09:00 - 09:35 Invited Talk 1
"Data-Efficient Foundation Model Pre-Training" by Sewon Min
09:35 - 10:10 Invited Talk 2
"Private Text Generation with LLMs: Mechanisms, Applications, and Audits" by Krishna Pillutla
10:10 - 10:45 Invited Talk 3
"Overclaiming in Frontier Coding Agents" by Nouha Dziri
10:45 - 11:00 Coffee Break
11:00 - 12:00 Oral Presentations
12:00 - 13:30 Lunch
13:30 - 14:05 Invited Talk 4
"The Legal Life of a Dataset: What Makes Training Data Usable?" by Janel Thamkul
14:05 - 14:40 Invited Talk 5
"Emergent Misalignment through the lens of Memorization" by Niloofar Mireshghallah
14:40 - 15:15 Invited Talk 6
"Creating RL Environments for Responsible Frontier Model Post-Training" by Alex Dimakis
15:15 - 15:30 Coffee Break
15:30 - 17:00 Poster Session
17:00 - 17:05 Closing Remarks

Speakers

Committee

Program Committee

If you wish to join the program committee, please signup here.

Contact

For questions or inquiries, please reach out to us at: redatacolm2026@gmail.com.