An open-weight model is an AI model, often a large language model, whose trained parameter files are available for other people to download or use. The weights are the numerical values learned during training. Having access to them can make it possible to run inference on your own infrastructure or adapt it for a particular task with techniques such as fine-tuning.
The phrase is precise: it describes the availability of the weights, not every part of the project.
What may be open
An open-weight release may include some combination of:
- Model weight files.
- Architecture or configuration files.
- Tokenizer files.
- Code for inference or fine-tuning.
- Documentation and evaluation results.
- A license describing allowed uses.
These pieces are separate. A project can publish weights while keeping its training data, training code, or data-cleaning process private.
Open-weight is also not automatically the same as open source. Open-source software under the Open Source Initiative definition requires more than downloadable weights, including broad freedoms to use, modify, and redistribute source code. For models, the term can be contested because training data, training code, usage rights, and redistribution rights may not be open. Always read the model license, not just the release headline.
Why use an open-weight model?
Running an available model yourself can provide:
- More control over where data is processed.
- Lower per-request cost at sustained high volume, if the hardware stays busy enough to justify hardware, operations, and staffing costs.
- The ability to work offline or in a private network.
- Custom fine-tuning or parameter-efficient adaptation, such as adapters and LoRA.
- More control over latency, hardware, and model versions.
These benefits come with operational work. A team must provide suitable hardware, manage updates, secure the service, monitor quality, and handle failures.
Running an open-weight model in practice
The first practical question is memory. A rough estimate is:
weight memory ~= parameters x bytes per parameter
For example, a 7B-parameter model needs about 14 GB for 16-bit weights, about 7 GB for 8-bit weights, and about 3.5 GB for 4-bit weights before extra serving memory. The parameters-and-weights article covers the math and quantization details. Real serving also needs room for activations, runtime overhead, and the KV cache used during generation.
Teams then choose a runtime. For local experiments, that may be a desktop or single-machine runtime. For production, it is usually a GPU serving stack with batching, concurrency limits, request queues, metrics, and rollback. Batching can improve throughput by running several requests together, but it can also increase latency for an individual request if the queue waits too long. Quantized variants can reduce memory and sometimes improve speed, but they still need quality checks on your task.
A release is not a guarantee
Before using an open-weight model, use a short adoption checklist:
- License review: check commercial use, redistribution, hosted-service rights, attribution, and restricted use cases. Access may be gated even when the weights are downloadable after approval.
- Provenance: download from official sources, record the exact version, and verify checksums or signatures when provided.
- Task benchmark: test the model on your own prompts, documents, languages, and failure cases.
- Latency and cost targets: include hardware, utilization, energy, operations, and staffing, not only raw token cost.
- Safety evaluation: check refusal behavior, data leakage risks, bias, and misuse paths.
- Compatibility: confirm the tokenizer, architecture config, runtime, adapters, and model file format work together.
- Update and rollback plan: know how you will patch, replace, or revert the model.
The downloaded weights may also be quantized or converted. That is a deployment-time compression and compatibility step, not a form of post-training adaptation. It can reduce memory use, but it may affect quality and should be evaluated.
Hosted versus self-hosted
With a hosted API, a provider runs the model and handles much of the infrastructure. With an open-weight model, your team may run inference directly or use a third party that hosts the model.
Neither option is always better. Self-hosting can improve control but increases responsibility. A hosted service can simplify operations but may provide less control over data location, model updates, and runtime behavior.
The key idea
Open-weight means that trained model parameters are available, not that every dataset, training process, or use is unrestricted. The release can enable private deployment and customization, but licensing, provenance, security, infrastructure, and evaluation responsibilities remain with the user.