Open-weight adoption is now large enough to matter economically
The shift from proprietary model APIs toward open-weight models is no longer only a developer preference. It is visible in production traffic.
Vercel's September AI Gateway index reported that open-weight models processed 56% of gateway tokens in August, up from 7% in December 2025. Those models accounted for only 14% of gateway spending. The same report said the average token cost had fallen by more than half over five months and declined another 23.2% in August.
That creates a different market question from the usual debate over which model ranks highest on a benchmark. If more production workloads can move to cheaper, customizable weights, economic value may migrate away from proprietary model access and toward the infrastructure that serves those models.
For listed options markets, the direct exposure is not OpenAI or Anthropic, which remain private. The observable public-market layer is further down the stack: semiconductors, networking, cloud infrastructure and the platforms that host inference.
The central research tension is whether lower model costs reduce the amount of compute required for a given application or instead make enough new inference economical that aggregate infrastructure demand continues to expand.
Startups are gaining portability rather than escaping infrastructure
The user-facing narrative around open weights often emphasizes independence. That part is real, but it can be misunderstood.
Hugging Face reported nearly three million public model repositories by August 2026 and described open models as increasingly embedded in company workflows. TechCrunch reported that developers and startups are using open models to avoid depending on a single closed provider, customize models for narrow workloads and gain more control over deployment.
OpenRouter data supports the same pattern from another angle. DeepSeek increased its share of token traffic substantially during the first half of 2026, including among AI-native companies and large organizations. OpenRouter described the emerging environment as one in which users mix models rather than standardize permanently on one provider.
That does not mean companies stop consuming external infrastructure.
An open-weight model can be run on company-controlled hardware, but it can also be served through a managed cloud endpoint. OpenAI's own open-weight documentation makes that distinction explicit: its open models can run on infrastructure controlled by the user or through hosting providers, and the economics depend on workload and operational approach.
The strategic change is therefore portability. A company can move the model, fine-tune it, change the serving provider or operate it privately. The underlying inference still has to run somewhere.
Cloud platforms are absorbing the open-weight transition
The major cloud vendors increasingly treat open-weight models as infrastructure workloads rather than competitors to their platforms.
AWS added six fully managed open-weight models to Amazon Bedrock in February, including DeepSeek, MiniMax, GLM, Kimi and Qwen families. The service runs those models through a distributed inference system designed to provide serverless capacity and OpenAI-compatible endpoints.
Microsoft has taken a similar route through Foundry. Its Fireworks AI integration provides managed inference for open models and custom-weight deployments. Microsoft said Fireworks was already processing more than 13 trillion tokens per day when the integration was announced.
Google's Vertex AI Model Garden also supports managed and self-deployed open models, including DeepSeek, Llama, Gemma and other families.
This matters because the migration away from a closed model provider does not necessarily mean migration away from hyperscale cloud spending. A startup can replace a proprietary model API while continuing to rent the accelerator, networking, storage and orchestration layer from a large cloud platform.
That shifts the economic question from model ownership to where inference is executed.
Cheaper models can pressure software margins without lowering chip usage
Open-weight models create an obvious pricing challenge for closed model providers. Vercel's August data is unusually clear: a majority of tokens ran through open-weight models while most spending still went to closed models.
The difference between volume share and spending share shows how wide the unit-cost gap can be.
But lower spending per token does not map mechanically to lower semiconductor demand. Hardware demand depends on total computation, not simply the invoice attached to one model endpoint.
NVIDIA has highlighted this distinction directly. In February it said inference providers running open models on Blackwell systems were reducing cost per token by as much as tenfold through a combination of hardware and optimized inference software. Lower cost can make additional workloads economically viable even when the same amount of intelligence becomes cheaper.
The unresolved variable is usage elasticity. If a company cuts the cost of one million tokens but keeps total token consumption unchanged, the infrastructure burden can fall. If it responds by deploying more agents, longer contexts, more frequent inference and additional automated tasks, total compute can stay high or increase.
Open-weight adoption therefore creates a margin question at the model layer and a volume question at the hardware layer.
The semiconductor options market is not pricing one common open-weight outcome
The latest verifiable options snapshot used here is September 18, before the September 21 session.
NVDA showed 30-day at-the-money implied volatility of 28.4% and an IV rank of zero within its prior-year range. SMH, the broad semiconductor ETF, showed 29.0% 30-day implied volatility and the same zero IV rank.
Broadcom sat higher at 35.1%, with an IV rank of 44. AMD was much higher at 49.3%, with an IV rank of 58.
The spread is too wide to interpret as one unified market view on open-weight AI. AMD's implied volatility was roughly twenty volatility points above SMH, while NVDA was slightly below the ETF. Broadcom occupied the middle.
That cross-section is more consistent with company-specific catalysts and different realized-volatility histories than with a sector-wide event premium tied to model commoditization.
It also prevents a simplistic conclusion. If open weights were already being treated as a direct threat to semiconductor economics, a broad volatility repricing across the hardware basket would be one observable possibility. That pattern is not visible in the September 18 snapshot.
The absence of a common premium does not prove that the theme is irrelevant. It means the theme has not yet become a clean options-market factor that can be separated from company events.
NVDA and AVGO represent different parts of the same open-model stack
The open-weight transition is not equally relevant to every semiconductor company.
NVIDIA is exposed through accelerator demand, networking and inference software. It also supports open-model ecosystems directly, including its own Nemotron models and optimized deployment software. If open models expand total inference volume, the company can participate even when the model creator is not a proprietary U.S. lab.
Broadcom is exposed through a different channel. Its AI business includes networking and custom accelerators used in large-scale data-center infrastructure. A world with more model portability can increase the importance of efficient serving architectures and specialized infrastructure without requiring every workload to use the same accelerator.
AMD is another distinct exposure because it is competing for accelerator and data-center CPU workloads while carrying much higher current implied volatility than the broad semiconductor ETF.
The correct comparison is therefore not simply open models versus closed models. It is how workload portability changes the mix of accelerators, networking, custom silicon, CPUs and cloud serving.
Options can help measure whether one layer begins to reprice relative to another.
Model portability may reduce concentration risk while increasing infrastructure competition
A second-order effect is vendor concentration.
When a startup depends on one closed model API, model pricing, capacity limits, policy changes and outages can all sit behind one external dependency. Open weights make it easier to move the workload between providers or host it under different operational constraints.
OpenRouter's model-routing business is evidence that this multi-model behavior is becoming a product category of its own. Its infrastructure lets users route requests across many models and providers, including price and fallback controls.
That portability can weaken pricing power at the model layer while intensifying competition at the infrastructure layer.
Cloud providers can respond by making open models easier to deploy. Inference providers can compete on throughput and token cost. Semiconductor vendors can compete on the cost of serving the same workload. The result is not less competition for compute. It is a different form of competition around compute.
For options research, that suggests relative volatility may become more informative than outright semiconductor volatility.
The next test is whether lower model prices create more total inference
The strongest evidence will come from usage rather than benchmark scores.
Vercel already shows open-weight token share rising from 7% to 56% in eight months while the average cost per token falls. The next question is whether total gateway token volume continues to accelerate as the mix becomes cheaper.
Cloud disclosures can provide another test. If managed open-model services expand while AI infrastructure spending remains elevated, that would support the interpretation that model commoditization is redirecting spending rather than eliminating it.
Semiconductor earnings can provide a third test through accelerator shipments, networking revenue, custom-silicon programs and inference-related commentary.
The options test is then straightforward. Researchers can compare NVDA, AMD and AVGO implied volatility with SMH around major model-cost, cloud-capacity and infrastructure-spending updates. A persistent company-specific divergence would suggest the economics are shifting between hardware layers. A coordinated move across the basket would suggest a broader change in expected AI infrastructure demand.
Open-weight adoption is already large enough to challenge the idea that production AI must remain tied to a small number of proprietary model vendors. What it has not yet done is establish a clear semiconductor outcome. The model layer is becoming cheaper and more portable. The unresolved options question is whether that portability compresses total compute requirements or simply moves a larger volume of inference onto a more competitive infrastructure stack.
Primary sources & disclosures
- Vercel, Open-weight models take 56% of token volume, September 17, 2026
- Hugging Face, State of Open Models: Summer 2026 Observations, August 14, 2026
- TechCrunch, The real AI race may no longer be at the frontier, July 14, 2026
- OpenRouter, DeepSeek V4 Is Earning Agentic Token Share, June 30, 2026
- OpenAI, Open-weight models overview, updated August 2026
- AWS, Amazon Bedrock adds six fully managed open-weight models, February 10, 2026
- Microsoft Azure, Fireworks AI open-model inference on Foundry, 2026
- NVIDIA, Open-source models on Blackwell reduce inference cost per token, February 12, 2026
- OptiView, NVDA options statistics, September 18, 2026
- OptiView, AMD options statistics, September 18, 2026
- OptiView, AVGO options statistics, September 18, 2026
- OptiView, SMH options statistics, September 18, 2026