Three years ago, Etched was three Harvard dropouts with a Thiel Fellowship and a bet that Nvidia's chips were the wrong tool for running deployed AI models. On August 18, 2026, the San Jose startup closed a $700 million Series D at a $21 billion valuation, doubling the $10.3 billion mark it set only 26 days earlier with its Sequoia-led Series C. The round was led by Jane Street, the quantitative trading firm that had become Etched's first paying customer months before the deal closed. Total funding now stands at roughly $1.9 billion, the company claims more than $1 billion in signed customer contracts, and Etched employs about 400 people drawn from Nvidia, Broadcom, Google's TPU team, and SK Hynix. The narrow bet is the point: Etched's chip, Sohu, is not trying to be an all-purpose GPU. It runs transformer inference, and only transformer inference, hardwired into silicon.

The valuation trajectory is the headline. $5 billion in December 2025, $10.3 billion in late July 2026, $21 billion on August 18, 2026. That is a 4x in eight months and a 2x in 26 days, with no revenue to anchor the number and only one publicly named customer. The valuation is the test, not the order book. The right question for an operator looking at this is not "is Etched overvalued," it is "what does the existence of a $21 billion pure-inference ASIC tell us about the AI compute cycle of 2026-2028, and where is the next dollar of margin going to land?" This guide works through the deal, the chip, the customer-led round structure, and what an operator can reasonably conclude about the inference-versus-training split, the ASIC-versus-GPU competition, and the venture capital cycle underwriting it.

The deal at a glance

FieldDetail
RoundSeries D
Amount raised$700 million
Post-money valuation$21 billion
Round leadJane Street (also first customer)
Other investorsKleiner Perkins, Sequoia, Andreessen Horowitz, Tiger Global, Bain Capital Ventures, Blackstone, Peter Thiel, Neo, Stripes, Primary, Positive Sum, Diffusion, Argo
Total funding to date~$1.9 billion
Headcount~400 (drawn from Nvidia, Broadcom, Google TPU, SK Hynix)
Announced customer contracts>$1 billion (signed, not recognized revenue)
Named customerJane Street (production deployment in own data center)
Valuation trajectory$5B Dec 2025 → $10.3B Jul 23, 2026 → $21B Aug 18, 2026
Process nodeTSMC N4P
Memory144 GB HBM3E per chip
ArchitectureTransformer-only ASIC, hardwired inference, no CUDA layer
Time to working silicon from tape-out44 days (vs typical 6+ months)
Etched AI valuation trajectory, Dec 2025 to Aug 2026 Post-money valuation in $B across 3 rounds. Customer-led Series D doubled the C in 26 days. $25B $15B $0 Dec 2025 Jul 23, 2026 Aug 18, 2026 +26 days $5B Series B $5B $10.3B $21B Series C (Jul 23) Series D (Aug 18, lead: Jane Street) Source: TechCrunch, WSJ, Reuters 2026 funding round coverage
Etched AI valuation trajectory across three rounds in eight months: $5B Series B (Dec 2025), $10.3B Series C (Jul 23, 2026), $21B Series D (Aug 18, 2026). The Series D doubled the Series C in 26 days, led by Jane Street, the company's first paying customer.
Etched's Sohu chip, the first transformer-specific ASIC, designed to run large language models at higher throughput than Nvidia GPUs at lower cost per token. (Etched)
Etched's Sohu chip, the first transformer-specific ASIC, designed to run large language models at higher throughput than Nvidia GPUs at lower cost per token. (Etched)

What the chip actually is

Sohu is an application-specific integrated circuit designed for one thing: running transformer inference at the lowest possible cost per token. It is not a GPU. It does not run training. It does not run any non-transformer workload, including older recurrent neural networks, classical machine learning models, or the rising tier of mixture-of-experts (MoE) transformer derivatives that some frontier labs are now shipping.

That narrowness is the bet. General-purpose GPUs handle inference by routing it through a software abstraction layer called CUDA, which dispatches AI calculations across thousands of parallel computing cores. The abstraction is powerful and flexible, which is why Nvidia dominates the AI compute market. The cost is overhead, scheduling instructions, launching compute kernels, managing thread allocation across cores, that idle large fractions of the GPU's compute capacity at any given moment during inference. Sohu eliminates that overhead by hardwiring the transformer computation directly into silicon. The result is a chip that, on transformer inference workloads, can deliver more tokens per dollar and per watt than a general-purpose GPU at the same process node.

The technical specifics matter. The chip is manufactured on TSMC's N4P, a 4nm-class process that is not the leading edge (TSMC's N3 and the upcoming N2 are smaller) but is mature and high-yield. Each chip carries 144 GB of HBM3E, the latest generation of high-bandwidth memory, which is the right choice for inference at scale where the model weights need to fit in memory to avoid the latency penalty of off-chip storage. Etched's approach to power management, which the company calls Low Voltage Inference, runs the chip's math blocks at below half the operating voltage of conventional AI chips. That packs more transistors into the same thermal envelope, which in turn allows the chip to process more tokens before thermal throttling kicks in.

The customer-led round is the unusual part

AI hardware mega-rounds are common in 2026. The Etched round is unusual because Jane Street was both the lead investor and the first paying customer, and had been a shareholder for months before the deal closed. This is a structure that is becoming more common in late-cycle AI funding. The Wall Street Journal reported that Jane Street tested Etched's hardware before committing capital, ran production trading workloads through it, and decided the chips were worth backing. The firm published a statement saying, in its words,

We tested the chip and are pleased with the early results

and now has a production rack running in its own data center.

For Jane Street, the structure is rational. The firm converts microseconds of latency directly into trading profit or loss, which makes inference throughput a core input to revenue. If Sohu delivers a meaningful latency or cost-per-token improvement over the Nvidia H100 or B200 at the workloads Jane Street runs, the equity position is a small price to pay for the option to capture that margin. For Etched, having Jane Street as both customer and lead investor is a credential that no other inference-chip startup can match in 2026. The round also includes Hudson River Trading, Jump Trading, and Two Sigma, a cluster of quant trading firms whose business models convert speed directly into money. The right read of the round is not "VCs are chasing AI" but "the customers who most need inference throughput are now backing the chip they want to deploy."

What Sohu cannot do

The narrowness of the chip is also its limit. Three things Sohu cannot do that buyers should know about before signing a $1 billion contract. First, mixture-of-experts architectures. MoE models are technically transformer-derived, but their dynamic routing structure, where different inputs activate different subsets of the model's parameters, undermines Sohu's fixed-function advantage. MoE is gaining ground precisely because it offers better efficiency at scale. If MoE becomes the dominant inference architecture, Sohu's value proposition narrows further. Second, training. Sohu is inference-only. Organizations running their own training pipelines still need Nvidia or AMD hardware for that phase. Sohu enters the stack only at serving time. Third, multi-workload serving. Cloud providers and AI labs that run a heterogeneous mix of inference workloads, including classical machine learning, older RNN architectures, and other non-transformer models, cannot run Sohu on those workloads. The addressable market for Sohu is the subset of AI inference spend that runs transformer-only models, which is large but not all of it.

Etched's Sohu chip
Etched's Sohu chip, the transformer-only inference ASIC. (Etched)

The AI inference market in 2026

The AI inference market is under pressure, and the pressure is structural. The AI API price war that broke out in August 2026, with Anthropic, Google, and OpenAI all cutting inference costs in quick succession, drives higher volume, which creates demand for exactly the throughput Sohu delivers. Lower prices drive higher usage. Higher usage demands more throughput. Specialized silicon is the logical endpoint of that compression. Etched's timing is not accidental. The market size estimates for AI inference in 2026 cluster around $20-25 billion, with projections reaching $35-40 billion by 2030. Dedicated ASICs are estimated to capture up to 45 percent of inference workloads by the end of the decade. The total addressable market for a transformer-only ASIC like Sohu is the intersection of those two projections, which is what the $21 billion valuation is implicitly underwriting.

Nvidia H100
Nvidia H100, the incumbent GPU against which Sohu competes. Inference is 5-10x more power-efficient than training at scale. (Nvidia)

Who the competitors are

Etched is not the only inference-chip startup chasing this opportunity. Three competitors deserve specific attention.

Cerebras Systems builds the Wafer-Scale Engine, a chip that is roughly the size of an entire silicon wafer and designed for both training and inference. The WSE-3 is the current generation, and the company announced new server systems in August 2026 specifically targeting AI chatbot inference. Cerebras's pitch is throughput at scale, with the wafer-scale design eliminating the latency of inter-chip communication. Cerebras has been around since 2016, has raised more than $700 million, and has a real customer book. The valuation is rumored to be in the $4-6 billion range, materially below Etched's $21 billion, but with broader applicability across training and inference.

Groq builds the Language Processing Unit, an inference-focused ASIC with a deterministic architecture that delivers very low latency for batch-1 inference (the single-query, real-time case). Groq has been around since 2016, raised $640 million, and was last valued at $2.8 billion. Groq's customer book includes large enterprise deployments and a recent partnership with Aramco for Saudi Arabia-scale deployments. The pitch is real-time inference where every millisecond matters.

The hyperscalers, Google, Amazon, Microsoft, Meta, are building their own inference silicon. Google's TPU is the longest-running custom inference chip, now in its sixth generation with the TPU v6 (Trillium). Amazon's Trainium and Inferentia are the second-longest-running, with Inferentia 3 announced in 2025. Microsoft's Maia, announced in late 2024, is the most recent entrant. Meta's MTIA, the Meta Training and Inference Accelerator, is the fourth. These chips are not sold externally. They are built for the hyperscalers' own workloads and run at the hyperscalers' own operating cost. They do not compete for the addressable market Etched is chasing, the AI labs and mid-tier cloud providers who need transformer inference throughput but do not have hyperscaler-scale custom silicon programs.

What an operator should conclude

The Etched valuation is the canonical example of a late-cycle, customer-led, proof-of-concept-via-deployment AI hardware round. The right way to read it is not as a prediction that Sohu will replace Nvidia, which it will not, but as a signal that inference has become a separable market from training, with its own economics, its own customer base, and its own capital cycle. Inference economics in 2026 are about cost per token at scale, and the chip that delivers the most tokens per dollar and per watt wins the order, even with Nvidia still dominating the market.

Three concrete takeaways. First, if you are buying inference capacity, the API price war is real, and the price compression is structural. The inference compute layer is commoditizing faster than the training layer. Second, if you are investing in AI hardware, the customer-led round structure is the model that survives the late-cycle squeeze. A quant trading firm testing the chip and then leading the round is a fundamentally different risk profile from a VC leading the round on the basis of a pitch deck. Third, if you are building AI infrastructure, the inference-versus-training split matters for capacity planning. Inference workloads have different latency requirements, different memory access patterns, and different cost-per-token economics than training. The right architecture for 2027-2028 is a mix, with training on Nvidia or AMD and inference on a mix of Nvidia, custom silicon at hyperscalers, and purpose-built ASICs like Sohu at the AI labs and mid-tier clouds.

The risk in the Etched valuation is the standard late-cycle venture risk. A $21 billion valuation against $1 billion in signed contracts is a 21x revenue multiple, assuming all $1 billion converts to recognized revenue, which it may not. The chip has not shipped at scale. The customer book is one named customer plus an undisclosed number of contracts. The chip does not handle MoE architectures, which are gaining share. Nvidia's software stack (CUDA) is a durable moat, and switching costs for AI labs are real. If Etched fails to convert the $1 billion in contracts to recognized revenue at scale, the valuation resets fast. If it succeeds, the inference ASIC market is the next AI infrastructure sub-category, and Etched is the first mover with a hyperscaler-quality customer reference.

Frequently asked questions

What is Etched Sohu

Etched Sohu is a transformer-only inference ASIC manufactured on TSMC's N4P process with 144 GB of HBM3E memory per chip. Unlike Nvidia GPUs, Sohu hardwires the transformer computation directly into silicon, eliminating the CUDA software abstraction overhead that idles large fractions of GPU compute during inference.

What is the Sohu chip used for

Transformer inference. Sohu runs large language models and other transformer-based AI models at serving time, where the model weights fit in memory and the workload is generating tokens for an end user or API call. Sohu does not run training, does not run MoE architectures efficiently, and does not run non-transformer workloads.

Why is Etched valued at $21 billion

The valuation reflects (1) the implied share of the $20-25 billion 2026 AI inference market that dedicated ASICs could capture, (2) the $1 billion in signed customer contracts, (3) the speed of inference commoditization, which makes dedicated silicon more valuable, and (4) the scarcity premium for a credible Nvidia challenger with a hyperscaler-quality customer reference.

Who are Etched's competitors

Cerebras Systems (wafer-scale AI accelerator), Groq (Language Processing Unit for real-time inference), and the hyperscalers' own custom silicon (Google TPU, Amazon Trainium and Inferentia, Microsoft Maia, Meta MTIA). The hyperscalers' silicon does not compete for Etched's addressable market because it is internal-only.

Why did Jane Street lead the round

Jane Street is Etched's first paying customer and had been a shareholder for months. The quant trading firm tested Etched's hardware in production, ran real trading workloads through it, and decided the chips were worth backing. Leading the round secured Etched's roadmap alignment with Jane Street's workload needs.

Is the Etched valuation too high

At 21x the $1 billion in signed contracts, the multiple is high relative to public semiconductor comparables. The valuation depends on (a) converting signed contracts to recognized revenue at scale, (b) maintaining the customer book against Nvidia's response, (c) handling the MoE architecture shift, and (d) shipping at scale. If any of those break, the valuation resets fast.

Sources