
The Art of Changing or: How I stopped worrying and learned to love the model
MARCH 21ST, 2026

I've been battling with the argument around ethical concerns that many in the software space have against AI having been trained on open source software. Admittedly, this concept is more nebulous for me, and I'm not sure I've decided where I land on it yet. But Cory Doctorow mentioned something recently that I think hits the nail square on the head, and I haven't found a position to refute his claim. His statement can be summed up as: it's not bad that virtually all open-source code has been hoovered up into AI models, that's exactly what open source is meant to do and to be. In other circles, we'd call this a search engine or a software directory. (Sorry, can't find the source–it was in a recent daily email he sends, but if someone does, I'll update the link here!)
And I think he's right. To assume there's some moral code that's been broken because models were trained on this code, I think is indefensible in a way, if that original code was also created based on learning/reading/using still other prior open source code. Virtually all code uses libraries and code written by others. It's how we've experienced such a massive growth in the technology industry over the last 40 years. I suppose there may be purists who have exactly zero NPM modules in their web app, and they wrote 100% of it in vanilla Javascript, but even then, the language they're using to write their software is, itself, a specification they're using, that they didn't directly write themselves. Unless one were to pull a TempleOS and build the entire OS from the ground-up or something I suppose.
Point is, I'm failing to come up with a valid argument against Cory's point: LLM learning on software code that's publicly available is just the next iteration of what we've been doing–and many have been advocating for–for decades, and it falls squarely into Stewart Brand's claim that "Information wants to be free". Though it's important to remember that the full phrase was something like Information Wants To Be Free. Information also wants to be expensive. That tension will not go away. So this tension continues. I just didn't think it'd be the open source side that'd advocate for keeping it expensive.
And in trying to be as forthright about my own biases here as I can, this position also happens to strengthen my long-held discomfort with the concept of intellectual property itself. Despite having filed and been granted several software patents in my own career, I suspected, and now I think I fully believe that intellectual property is simply too artificial to exist and probably ought to be abolished. Property, I believe, needs some fundamental properties (ha) to be defined as Property–among those are scarcity. Information, like fire, is not scarce. Jefferson's statement, paraphrased: He who lights his candle at mine receives light without darkening me is prescient here.
I think it's possible that I'm finally coming around to understanding–and agreeing with–Stallman's Libre vs Gratis argument about Free Software, where he states: I believe that all generally useful information should be free. By "free" I am not referring to price, but rather to the freedom to copy the information and to adapt it to one's own uses ... When information is generally useful, redistributing it makes humanity wealthier no matter who is distributing and no matter who is receiving.
Interestingly, it seems to be the most vehement antagonists against companies training models on open source software also happen to be the biggest proponents of GPL-level software licensing. This is a pretty distinct contradiction, and leads me to believe there are other forces at play within the antagonists' minds pushing their revulsion towards LLM training. Probably in large part their (rightful) revulsion at the greed and scale of many of these companies.
The only argument I can come up with in favor of the antagonistic view is that the receiving side (OpenAI/Anthropic/etc) should also ensure their data is freely available–which today, most of them are not. A valid argument certainly, but nothing we haven't experienced before, and continue to experience today. Google, though it's search index is free (gratis) to use, it's not free (libre) to download and use and repurpose. We know this in our bones. Yet anyone can create a new [search engine](https://kagi.com). But now that capital and investment has shown that LLMs can indeed be powerful, that's all the more reason why We (capital We, as in society, humanity), should be driving towards open-weight models that can run locally on standard workstations and laptops. That is the natural continuation of information wanting to (and continuing to) be free.
I wonder if the negative sentiment would subside if a truly open model emerged that wasn't profit-driven, and was able to run locally on most standard-priced computers–not massive GPU clusters. We might be a year or two out from that, but the latest Apple M5 silicon may put that realization much close to our future reality, as they can run exceptionally good models locally today.
Would there still be pushback if you could run a local Claude-Code-competitive model on your own machine, fully and privately? I think the narrative will change then, and I think that'll reveal the real reasons for the antagonism.
Importantly, perhaps for those who cannot afford local-compute, we can look to places like libraries to run these models on behalf of their community. Libraries will have a very important role to play here for access to these models, regardless of income level.

clear.txt
Builders, researchers, and the genuinely curious. Infrastructure informed by independent culture. all killer, no filler
