Home Content News Codeberg Bans AI-Generated Repositories To Protect Open-Source Commons

Codeberg Bans AI-Generated Repositories To Protect Open-Source Commons

0
1
Codeberg
Codeberg

Codeberg e.V. updates its terms to ban unreviewed LLM projects and prohibit code harvesting for AI training, citing maintainer strain and resource limits.

In a decisive move to safeguard developer infrastructure and open-source integrity, on 23 July, 2026, Codeberg e.V. passed two major community-driven motions during its annual assembly, updating its Terms of Use to explicitly ban hosted repositories that mostly consist of generative AI or LLM code lacking human oversight.

Approved by a 71% majority vote, with 358 in favour, 144 against, and 14 abstentions, the policy directly addresses operational strain, copyright ambiguity, and security risks introduced by unreviewed ‘vibe-coded’ projects. Alongside the ban, Codeberg formally ratified an Institutional Data Policy that codifies a zero-tolerance stance against data harvesting, guaranteeing that no user-hosted source code or developer metadata on the platform will ever be used to train AI models.

Automated, unreviewed ‘ghost projects’ rapidly drain Codeberg’s shared CI/CD computational infrastructure, server storage, and bandwidth while offering minimal long-term value to human developer communities. Simultaneously, maintainers face acute exhaustion from managing an influx of low-quality pull requests and automated ‘Looks Good To Me’ rubber-stamping, which elevates the risk of introducing security vulnerabilities into active projects.

“The widespread use of LLMs in FLOSS is instead becoming a multidimensional attack on the trust between contributors and the very idea of convivial collaboration itself,” Codeberg, in its blog, stated. “As we want to center on human collaboration, we will not actively support or engage in the creation of LLMs and will not put our limited resources to use for storing single-use software that would pollute our FLOSS commons.”

Beyond infrastructure bottlenecks, Codeberg cited broader legal and environmental hazards associated with generative AI deployment. The ambiguous copyright status of machine-generated code presents ongoing risks to standard open-source licensing models, leaving projects vulnerable to ownership disputes. Additionally, Codeberg highlighted the environmental burden of model training alongside the operational disruption caused by aggressive, unauthorised web scraping bots crawling its repositories. Codeberg stated enforcement will be handled on a manual, case-by-case basis driven by moderation and community flagging.

LEAVE A REPLY

Please enter your comment!
Please enter your name here