For thirty years, the open-source software movement was the closest thing humanity had to a digital miracle.
Millions of individual human beings—students in university dorms, late-night hobbyists, retired engineers, underpaid maintainers—wrote millions of lines of code. They designed operating system kernels, math libraries, audio codecs, encryption algorithms, and web frameworks.
They didn't do it to build proprietary monopolies. They slapped GPL, MIT, or Apache licenses on their repositories and gave their work away to the world for free.
The social contract was simple and sacred:
- You are free to read this code.
- You are free to run this code.
- You are free to modify this code.
- BUT if you build on it, you must respect the author's copyright, preserve the original license attribution notice, and (if it’s copyleft like the GPL) contribute your modifications back to the global digital commons so everyone benefits.
That contract powered the modern world. Linux runs every supercomputer, every Android phone, and 95% of the internet. Git, PostgreSQL, Python, and GCC exist because humanity decided to share its collective intellectual heritage.
And then, in 2021, Microsoft, GitHub, and OpenAI looked at that thirty-year mountain of human generosity, smiled, and executed the largest intellectual property heist in human history.
How to Launder Copyright with Linear Algebra
Microsoft didn't hire an army of engineers to write a code-completion engine from scratch. They didn't need to. They already owned GitHub—the central digital repository where humanity’s open-source code lived.
They pointed web scrapers at every public repository on GitHub. They hoovered up billions of lines of code:
- Code written by hobbyists on weekends.
- Code written by university researchers on public science grants.
- Strictly copyleft GPLv2 and GPLv3 code whose legal terms explicitly forbid inclusion in closed-source proprietary commercial software without releasing source code.
They dumped all of it into the training pipeline of OpenAI’s Codex and GPT models.
And then Microsoft’s high-paid legal team came up with the ultimate intellectual property defense: "The model isn't copying code; it’s learning statistical probabilities! Neural network weights aren't derivative works, they are fair use!"
Think about how brilliant that legal laundering maneuver is:
- If an individual developer copies 50 lines of GPL-licensed Linux kernel code into a commercial closed-source SaaS product, Microsoft's lawyers will sue them into oblivion for copyright infringement.
- But if Microsoft passes those exact same 50 lines of GPL code through a transformer neural network, serializes the mathematical weights into a floating-point matrix, and uses those weights to emit those exact same 50 lines of code on an end-user's screen inside VS Code? Suddenly, all legal obligations vanish into thin air!
No attribution. No author name. No GPL license preservation. No link to the original repository.
They laundered human labor through linear algebra, erased every trace of the people who built it, and then turned around and slapped an ₹800/month or ₹1,600/month per-seat subscription paywall on the autocomplete box.
The Receipts: It Literally Emits Verbatim Code
When researchers and programmers first called Microsoft out on this, executives claimed: "The model synthesizes novel code! It doesn't emit verbatim copies from its training data."
Then developers started testing it, and the receipts were devastating.
Tim Davis, a computer science professor at Texas A&M, tweeted an example where GitHub Copilot verbatim regurgitated entire copyrighted functions from his sparse matrix mathematics library—complete with his original function names, variable names, and mathematical edge-case comments—while claiming the code was "newly generated."
Armin Ronacher (creator of Flask) showed Copilot verbatim generating the famous fast inverse square root algorithm from Quake III—complete with John Carmack’s original 1999 profane code comments verbatim.
The model was not magically "understanding computer science." It was an over-parameterized statistical compression engine that had memorized vast swaths of open-source repositories and was happily disgorging verbatim proprietary code into users' commercial projects with all copyright metadata stripped away.
The Moral Bankruptcy of Selling Open Source to Broke Kids
As a student, what makes me genuinely sick is the predatory power dynamic here.
GitHub built its entire multi-billion-dollar valuation on the backs of broke kids, indie hackers, and open-source volunteers who hosted their code on GitHub for free. We gave them our code, our bug reports, our documentation, and our network effects because GitHub promised to be the home of open source.
Now, a high schooler or university student trying to learn programming sits down with VS Code. They are greeted by aggressive banners urging them to pay ₹800 a month for Copilot autocomplete.
Whose code is helping that student write their Python loops? It is code written by open-source volunteers who never received a dime from Microsoft. Microsoft acts as the middleman landlord, charging a subscription toll on the collective commons of the programming community.
They privatized the profits, socialized the labor, and left open-source maintainers holding the bag.
The Death of Copyleft and the Rise of Code Vultures
The long-term danger of this heist is existential for open source.
Why would any independent engineer spend five hundred hours writing an open-source library under the GPL today, knowing that Microsoft, Google, and Anthropic will immediately scrape it, feed it to a proprietary frontier model, and sell access to the highest bidder without a word of credit or compensation?
It destroys the psychological foundation of open source. It encourages developers to close-source their repositories, switch to source-available "anti-AI" licenses, or retreat behind private Git servers.
The corporate tech giants have acted like industrial fishing trawlers that scraped the ocean floor clean of all fish, dumped the catch into luxury restaurant kitchens, and now wonder why the ocean is barren.
How to Resist the Enclosure of Code
If you care about software freedom, you cannot be passive about this:
- Migrate Off GitHub: You don't need to give Microsoft your code. Support independent, non-corporate Git hosting. Self-host Forgejo or Gitea on a ₹400/month VPS. Use Codeberg—a non-profit, community-driven Git host run in Germany that explicitly rejects corporate AI scraping and respects software freedom.
- Audit Your Tools: Stop paying for GitHub Copilot. If you want local code completion that respects your privacy and doesn't feed corporate landlords, run open-source local models like DeepSeek-Coder or Qwen-Coder via Ollama or llama.cpp directly on your own GPU. It runs 100% offline, costs zero dollars a month, and doesn't phone home to Redmond.
- Demand Legal Accountability: Support the ongoing class-action lawsuits against GitHub Copilot for DMCA Section 1202 violations (the intentional removal of copyright management information).
Open source was built by people who believed that knowledge belongs to humanity, not to quarterly balance sheets. Don't let them rewrite history into a paid autocomplete subscription.