Deepseek released its model weights publicly, but "open source" is more complicated than that
Deepseek, a Chinese AI company, has released the weights (the trained parameters) of its large language models to the public. This is a real step toward openness — you can download the model files and run them on your own computer. However, the company did not release the full source code for the training process, the data used to train the model, or the infrastructure that built it. So Deepseek is partially open: the finished product is available, but the recipe and ingredients are not.
This matters because "open source" means different things to different people. In traditional software, open source means you can see and modify the code that makes the program run. With AI models, the situation is murkier. You can use Deepseek's model weights for research, commercial projects, or personal experiments without paying the company. But you cannot fully understand how it was built or reproduce it from scratch without their proprietary training code and data.
Key Takeaways
- Deepseek released model weights publicly under a license that allows research and commercial use, which is more open than most competing AI models from other companies.
- The company did not release the training code, training data, or the full technical details of how the model was built, so you cannot fully reproduce it yourself.
- You can download and run Deepseek models locally on your own hardware without paying Deepseek or asking permission, which is a genuine form of openness.
- The license terms allow commercial use, but you should read the specific license document before building a business around the model.
What Deepseek actually released
Deepseek released the model weights for its Deepseek-67B and Deepseek-7B models under a license called the Deepseek License Agreement. The weights are the numerical parameters that make the model work — think of them as the "brain" of the AI. You can download these files from Hugging Face, a platform where researchers share AI models, and run them on your own computer using software like Ollama or LLaMA.cpp.
The company also released some technical documentation describing how the model works at a high level. This includes information about the model's architecture, its training approach, and its performance on various benchmarks. However, this documentation does not include the actual code used to train the model, the exact dataset it was trained on, or the specific hardware setup required to train it from scratch.
What Deepseek did not release
Deepseek kept several things private. The training code — the software that took raw data and turned it into the finished model — remains proprietary. The training dataset, or at least the full details of what data was used and how it was processed, is not public. The infrastructure and computational resources used to train the model are also not disclosed in detail.
This is common among AI companies, even those that release model weights. OpenAI, Meta, and others have done similar things: release the finished model but keep the training process secret. The reasoning is usually that the training process is where the real competitive advantage lies, and releasing it would allow competitors to replicate the work too easily.
How Deepseek's openness compares to other AI models
Deepseek is more open than proprietary models like ChatGPT or Claude, which you can only access through a company's website or API. You cannot download those models or run them on your own hardware without permission. Deepseek is roughly on par with Meta's Llama models, which also released weights publicly under a license that allows research and commercial use.
Deepseek is less open than truly open-source projects like Mistral or some community-driven models, where researchers have released not just the weights but also detailed training code and information about the datasets. However, those models are typically smaller and less powerful than Deepseek.
What you can actually do with Deepseek
Under the Deepseek License Agreement, you can download the model weights and run them locally. You can use the model for research, experimentation, and commercial projects. You do not need to ask Deepseek for permission or pay them a fee. You can modify the weights or fine-tune the model on your own data to specialize it for a particular task.
What you cannot do is claim you built the model yourself or remove Deepseek's attribution. You also cannot use the model to train another model and then claim that new model is open source if you have not disclosed your changes. The exact restrictions depend on the specific license version, so read the license document before building something you plan to sell or distribute widely.
Why this matters for your decision
If you want to run an AI model on your own computer without relying on a company's servers, Deepseek is a real option. You can download it, install it, and use it for free. If you are building a commercial product and want to avoid paying API fees to a third party, Deepseek's weights give you that option.
However, if you need to understand exactly how a model was trained, audit its training data for bias, or reproduce the training process yourself, Deepseek will not give you that level of transparency. For most users — people who just want to run a capable AI model locally — this does not matter. For researchers studying AI training methods or companies with strict requirements about model provenance, it does.
The practical difference between "open weights" and "open source"
The AI industry has started using the term open weights to describe models like Deepseek, where the trained parameters are public but the training process is not. This is distinct from open source, which traditionally means the source code is available. The distinction matters because it clarifies what you actually get: a working model, but not the full recipe.
In practice, open weights is enough for most use cases. You can run the model, modify it, and build products with it. You just cannot see inside the training process or fully understand why the model makes the decisions it does. If that limitation is acceptable for what you are trying to do, Deepseek is genuinely useful and genuinely free to use.
Frequently Asked Questions
Can I use Deepseek commercially without paying Deepseek?
Yes. The Deepseek License Agreement allows commercial use. You can build a product using Deepseek's model weights, sell it, and keep the revenue. You do not owe Deepseek a percentage or a fee. You do need to include attribution and follow the license terms, so read the full license before you launch.
Can I modify Deepseek's model weights?
Yes. You can fine-tune the model on your own data, adjust it for a specific task, or merge it with other models. The license permits this. However, if you distribute your modified version, you should disclose that you modified it and follow the license requirements for attribution.
Is Deepseek safer or more transparent than other AI models?
Deepseek is not inherently safer or more transparent than other open-weights models like Llama. Both release weights but keep training data and methods private. Safety and transparency depend on the specific model version and how it was trained. You should test any model you plan to use for potential biases or harmful outputs.
Do I need special hardware to run Deepseek locally?
Deepseek's smaller models (7B parameters) can run on a laptop with 8GB of RAM, though it will be slow. The larger models (67B) need a GPU with at least 24GB of memory, like an RTX 4090 or an A100. You can also run it on a cloud provider's GPU if you do not want to buy hardware.
What is the difference between Deepseek's free API and downloading the model weights?
Deepseek offers a paid API where you send requests to their servers and pay per token. Downloading the weights lets you run the model on your own hardware for free, but you handle all the setup and computing costs yourself. The API is faster and easier if you do not have a powerful GPU; local use is cheaper if you run the model frequently.