In a significant policy shift, GitHub has announced that it will start using customer interaction data to train its AI models, effective April 24. This decision affects users of Copilot Free, Pro, and Pro+ plans, while Copilot Business and Enterprise users, along with students and teachers, are exempt due to existing contractual terms.
The data GitHub intends to collect includes “inputs, outputs, code snippets, and associated context.” Users will have the ability to opt out of this data collection, aligning with established industry practices prevalent in the U.S., as opposed to the opt-in requirements commonly found in Europe. To opt out, users can navigate to their settings and disable the option that allows GitHub to use their data for AI model training.
Mario Rodriguez, GitHub’s chief product officer, encourages participation, stating that user data will help improve the AI’s understanding of development workflows, enhance code suggestions, and assist in identifying potential bugs before they reach production. GitHub justifies this approach by noting that similar policies are in place at companies like Anthropic and JetBrains.
Rodriguez claims that incorporating interaction data has led to notable improvements in AI model performance, particularly citing increased acceptance rates for AI-generated suggestions. The specific types of data GitHub seeks include accepted or modified model outputs, code snippets, context surrounding cursor positions, comments, documentation, file names, repository structures, interactions with Copilot features, and user feedback.
This policy change raises questions about the privacy of GitHub’s private repositories, which are traditionally defined as accessible only to the user and those they explicitly share access with. However, if a user enables model training on their interaction data, snippets from private repositories may be collected during their use of Copilot.
Community response to this announcement appears largely negative, with users expressing their discontent through emoji reactions—59 thumbs down compared to just three rocket ships, which signify enthusiasm. Only GitHub’s VP of developer relations, Martin Woodward, has publicly endorsed the initiative among the 39 comments at the time of this article’s publication.
Despite the backlash, it’s important to note that GitHub’s Copilot utilizes OpenAI’s Codex, a model fine-tuned on publicly available code from GitHub, indicating that the collection of data for AI training is already a common practice in the industry.
This article was produced by NeonPulse.today using human and AI-assisted editorial processes, based on publicly available information. Content may be edited for clarity and style.








