Jake Makes AI
Open Washing

Open Weights Is Not Open Source

They handed you a locked box and called it freedom. You can't see inside and you can't rebuild it.

An engineer squinting through the keyhole of a wooden crate that is propped open but wrapped in chains and padlocks

Meta calls Llama "open source." Mistral says it. So does half of Hugging Face on any given Tuesday. It is the warmest word in the industry, the one that makes a trillion-dollar company sound like a scrappy hacker collective sharing code in a garage. And most of the time it is a lie, told on purpose, because the lie is worth more than the truth.

Here is what you actually get when a big lab "opens" a model. You get the weights. A giant blob of numbers, gigabytes of them, that you can download and run. That is genuinely useful, and I am not knocking it. But open source has a definition that predates the AI boom by three decades, and that definition is not "you can run the binary." It is "you can see how this was built, change it, and rebuild it yourself." The weights give you none of that.

You do not get the training data. You do not get the exact recipe, the filtering steps, the data mix, the reinforcement pipeline that turned raw numbers into something that answers questions. Without those, you cannot reproduce the model. You cannot audit it. You cannot even fully explain why it does what it does. You got the cake. You did not get the ingredients or the oven, and you certainly cannot bake another one.

A binary you can run is not the same as a thing you can understand.

The Open Source Initiative, the group that has literally maintained the definition of the term since 1998, looked at this and drew a line. Their open-source AI definition says a real open model has to release enough about its training data for someone else to recreate it. By that standard, Llama does not qualify. And it is not just the data. Llama ships with a license that says if your product hits 700 million monthly users, you have to go ask Meta for permission. Real open-source licenses do not have a "unless you get too successful" clause. That is not the GPL. That is a rental agreement with extra steps.

So why do it? Why does a company spend hundreds of millions training a model and then give away the expensive part? Because "open" is the cheapest marketing they will ever buy. It floods the ecosystem with developers who build on your model for free, tune it, evangelize it, and lock their own products to your architecture. It buys you goodwill with regulators who love the word. It lets you look like the good guy standing against the closed labs, while you keep the two things that actually matter, the data and the process, locked in the same vault as everyone else. You get the halo of the commons without giving up the commons.

The tell is what they never release. Nobody hands you the training set, because the training set is where the lawsuits live. It is scraped books, pirated archives, the entire open web whether it wanted to be there or not. "Open source" would mean showing you the receipts, and the receipts are radioactive. So they open the one part that carries no liability and keep the one part that would end them in court. Then they call the whole thing open, and a room full of journalists writes it down.

I want to be fair, because the alternative is worse. Open weights beat closed weights. Being able to run a capable model on your own hardware, offline, with no company metering every token, is a real gift and it matters for privacy, for cost, for anyone who does not want to be a tenant on someone else's platform forever. I use these models. I am glad they exist. That is exactly why the language should be honest.

Call it open weights. That is what it is, and it is a good thing. But when a company borrows the credibility of a thirty-year-old movement built by people who gave away everything, including the source, to describe a locked box with a usage cap, that is not generosity. It is a costume with the tag still on. The word "open" did real work for real decades. Watching it get hollowed out into a press-release adjective is its own small tragedy, and we are all just quietly agreeing to let it happen.

Download the weights. Run them. Enjoy them. Just do not let anyone tell you that you were handed the keys, when all you got was permission to sit in the car.

§
Post-ready for LinkedIn
Half the industry is calling their models "open source." Most of them are lying, and it's on purpose, because the word is worth more than the truth. Here's what you actually get when a big lab "opens" a model. You get the weights. A giant blob of numbers you can download and run. Useful, sure. You don't get the training data. You don't get the recipe. So you can't reproduce it, can't audit it, can't rebuild it. You got the cake, not the ingredients or the oven. Open source has meant one thing since 1998... you can see how it was built and remake it yourself. Weights fail that test. The Open Source Initiative said so out loud. And then there's the fine print. Llama's license says if you hit 700 million users, go ask Meta for permission. Real open-source licenses don't have an "unless you get too successful" clause. Why do it? Because "open" is the cheapest marketing a trillion-dollar company will ever buy. Free developer army, regulator goodwill, the good-guy halo. All while the two things that matter, the data and the process, stay locked in the same vault as everyone else. Call it open weights. It's a good thing. It just isn't what they're calling it. What's the last "open source" model you actually could have retrained from scratch?
← All essays