← Back to concepts
6 min read

Open-weight AI models

An open-weight model is an AI model, often a large language model, whose trained parameter files are available for other people to download or use. The weights are the numerical values learned during training. Having access to them can make it possible to run inference on your own infrastructure or adapt it for a particular task with techniques such as fine-tuning.

The phrase is precise: it describes the availability of the weights, not every part of the project.

What may be open

An open-weight release may include some combination of:

  • Model weight files.
  • Architecture or configuration files.
  • Tokenizer files.
  • Code for inference or fine-tuning.
  • Documentation and evaluation results.
  • A license describing allowed uses.

These pieces are separate. A project can publish weights while keeping its training data, training code, or data-cleaning process private.

Open-weight is also not automatically the same as open source. Open-source software under the Open Source Initiative definition requires more than downloadable weights, including broad freedoms to use, modify, and redistribute source code. For models, the term can be contested because training data, training code, usage rights, and redistribution rights may not be open. Always read the model license, not just the release headline.

What an open-weight release may include The release package has trained weights as the central open item. Around it are optional architecture configuration, tokenizer files, inference or fine-tuning code, documentation, evaluations, and a license. Training data, training code, and data-cleaning process are shown outside the package as items that may remain private. Trained weights available to download the defining open item Architecture config Tokenizer files Inference code Documentation and evals License terms Fine-tuning code Training data Training code Data cleaning Often not included, even when the weights are available.
Open-weight describes access to the trained parameters; the surrounding code, data, and permissions are separate release choices.

Why use an open-weight model?

Running an available model yourself can provide:

  • More control over where data is processed.
  • Lower per-request cost at sustained high volume, if the hardware stays busy enough to justify hardware, operations, and staffing costs.
  • The ability to work offline or in a private network.
  • Custom fine-tuning or parameter-efficient adaptation, such as adapters and LoRA.
  • More control over latency, hardware, and model versions.

These benefits come with operational work. A team must provide suitable hardware, manage updates, secure the service, monitor quality, and handle failures.

Running an open-weight model in practice

The first practical question is memory. A rough estimate is:

weight memory ~= parameters x bytes per parameter

For example, a 7B-parameter model needs about 14 GB for 16-bit weights, about 7 GB for 8-bit weights, and about 3.5 GB for 4-bit weights before extra serving memory. The parameters-and-weights article covers the math and quantization details. Real serving also needs room for activations, runtime overhead, and the KV cache used during generation.

Teams then choose a runtime. For local experiments, that may be a desktop or single-machine runtime. For production, it is usually a GPU serving stack with batching, concurrency limits, request queues, metrics, and rollback. Batching can improve throughput by running several requests together, but it can also increase latency for an individual request if the queue waits too long. Quantized variants can reduce memory and sometimes improve speed, but they still need quality checks on your task.

A release is not a guarantee

Before using an open-weight model, use a short adoption checklist:

  • License review: check commercial use, redistribution, hosted-service rights, attribution, and restricted use cases. Access may be gated even when the weights are downloadable after approval.
  • Provenance: download from official sources, record the exact version, and verify checksums or signatures when provided.
  • Task benchmark: test the model on your own prompts, documents, languages, and failure cases.
  • Latency and cost targets: include hardware, utilization, energy, operations, and staffing, not only raw token cost.
  • Safety evaluation: check refusal behavior, data leakage risks, bias, and misuse paths.
  • Compatibility: confirm the tokenizer, architecture config, runtime, adapters, and model file format work together.
  • Update and rollback plan: know how you will patch, replace, or revert the model.

The downloaded weights may also be quantized or converted. That is a deployment-time compression and compatibility step, not a form of post-training adaptation. It can reduce memory use, but it may affect quality and should be evaluated.

Checks before adopting an open-weight model A team moves through several checks before deployment: license and use restrictions, compatible tokenizer and runtime, hardware and memory needs, task evaluation, provenance and safety, and security and monitoring. A quantized or converted model loops back through evaluation because deployment compression can change quality. License allowed use Runtime tokenizer, code Hardware memory, latency Evaluate your task Run monitor Quantized copy re-evaluate compression Availability is only the first step; deployment still needs licensing, provenance, capacity, evaluation, and monitoring.
Adopting an open-weight model is a checklist, not just a download: legal, runtime, hardware, quality, and operations all matter.

Hosted versus self-hosted

With a hosted API, a provider runs the model and handles much of the infrastructure. With an open-weight model, your team may run inference directly or use a third party that hosts the model.

Neither option is always better. Self-hosting can improve control but increases responsibility. A hosted service can simplify operations but may provide less control over data location, model updates, and runtime behavior.

Hosted API compared with self-hosting an open-weight model Two paths are compared from an application request to a model response. In the hosted API path, the provider owns serving infrastructure, scaling, model updates, and much of operations, while the application gives up some control over data location and runtime behavior. In the self-hosted path, the team controls the weights, runtime, hardware, data location, and versioning, but must operate and secure the service itself. Hosted API App request provider endpoint Provider operates serving, scaling, updates runtime details Response less ops less control Self-hosted open-weight model App request your endpoint Team operates hardware, runtime, security monitoring, failures Response more control more work
Hosted services shift operational work to a provider; self-hosting an open-weight model shifts more control and responsibility to your team.

The key idea

Open-weight means that trained model parameters are available, not that every dataset, training process, or use is unrestricted. The release can enable private deployment and customization, but licensing, provenance, security, infrastructure, and evaluation responsibilities remain with the user.