
Photo: Alchemist-hp (talk) www.pse-mendelejew.de, CC BY-SA 3.0
Custom AI Chips: NVIDIA vs Google, Amazon & Apple
Explore the $100B custom AI chips race. Compare specs, benchmarks, and market share for NVIDIA Blackwell, Google TPU, Amazon Trainium, and Apple M4 Ultra.
Key Takeaways
- →NVIDIA holds over 80% of the AI training chip market share.
- →Google runs over 90% of its internal AI workloads on custom TPUs.
- →Amazon Trainium 3 claims 4x better price-performance than comparable NVIDIA GPUs.
- →Apple M4 Ultra enables on-device execution of 600B parameter AI models.
- →The custom AI chip market is projected to reach $100 billion by 2028.
NVIDIA controls 80%+ of the AI training chip market, but tech giants are spending billions to break free. Here's who's building what, with real benchmarks and market data.
Key Takeaways
- NVIDIA's data center revenue hit $115B in FY2026 (ending Jan 2026)
- Google uses TPUs for 90%+ of its internal AI workloads
- Amazon Trainium 3 delivers 4x better price-performance than comparable NVIDIA chips
- Apple's M4 Ultra runs 600B parameter models on-device
- Custom AI chip market projected to reach $100B by 2028
NVIDIA: Still the King, But Challenged
Blackwell Architecture (B200/GB200)
NVIDIA's Blackwell GPUs launched in 2025 and remain the gold standard:
| Spec | H100 (2022) | H200 (2023) | B200 (2025) | B300 (2026) |
|---|---|---|---|---|
| FP8 Performance | 4 PetaFLOPS | 4 PetaFLOPS | 9 PetaFLOPS | 12 PetaFLOPS |
| Memory | 80GB HBM3 | 141GB HBM3e | 192GB HBM3e | 288GB HBM4 |
| Memory Bandwidth | 3.35 TB/s | 4.8 TB/s | 8 TB/s | 12 TB/s |
| TDP | 700W | 700W | 1000W | 1200W |
| Price | ~$30,000 | ~$35,000 | ~$40,000 | ~$60,000 |
Revenue reality: NVIDIA's data center segment generated $35.2B in Q4 FY2026 alone, up 73% YoY. The company expects to ship $150B+ in data center GPU revenue in FY2027.
Market Share
According to research from Dylan Patel (SemiAnalysis):
- NVIDIA: 80-85% of AI training, 70% of AI inference
- Google TPUs: 8-10% (internal use only)
- AMD MI300X/MI400: 5-7%
- Custom silicon (AWS, Microsoft, Meta): 3-5%
Google TPU: The Internal Powerhouse
TPU v6e (Trillium)
Google's 6th generation TPU, announced at Cloud Next 2025:
- Performance: 4.7x improvement over TPU v5e per chip
- Memory: 32GB HBM per chip (doubled from v5e)
- Interconnect: 3,200 chips per TPU pod via ICI
- Availability: GA on Google Cloud since November 2025
Real Usage Numbers
Google doesn't sell TPUs externally, but the internal usage is massive:
- Gemini models trained entirely on TPUs (v5e and v6e)
- 90%+ of Google's internal ML workloads run on TPUs
- Search, Ads, YouTube, Maps -- all powered by TPU inference
- Google Cloud TPU revenue estimated at $8-12B annually (internal transfers + Cloud customers)
Price Comparison
Google Cloud TPU pricing (us-central1):
| Instance | Chips | Price/Hour | Performance/$ vs H100 |
|---|---|---|---|
| TPU v5e (1 chip) | 1 | $1.20 | 2.3x better |
| TPU v5p (1 chip) | 1 | $4.20 | 1.1x better |
| TPU v6e (1 chip) | 1 | $2.50 | 3.1x better |
| NVIDIA H100 (1 GPU) | 1 | $12.00 | Baseline |
Amazon Trainium: The Price-Performance King
Trainium 3 (Announced re:Invent 2025)
Amazon's third-generation custom AI chip:
- Performance: 12 PetaFLOPS FP8 per chip
- Memory: 192GB HBM3e
- Price: ~$8,000-12,000 (estimated)
- Power: 500W TDP
Trainium vs NVIDIA: Real Benchmarks
Based on MLPerf Training v5.0 results and AWS benchmarks:
| Benchmark | Trainium 3 | NVIDIA H200 | Trainium Advantage |
|---|---|---|---|
| LLM Training (70B) | 1.0x (baseline) | 0.85x | 18% faster |
| LLM Inference (70B) | 1.0x (baseline) | 0.92x | 9% faster |
| Price/Performance | 1.0x | 0.42x | 2.4x cheaper |
Who's Using Trainium
- Anthropic: Migrated Claude training to Trainium (deal worth $4B over 5 years)
- Apple: Uses Trainium for Siri and on-device model training
- Adobe: Firefly image generation runs on Trainium
- Stability AI: Stable Diffusion training on Trainium
- Amazon internal: Alexa, recommendations, search all run on Trainium
AWS reported $12B in annualized revenue from custom chip customers in Q4 2025.
Microsoft Maia: The Quiet Contender
Maia 100 (GA 2025)
Microsoft's first custom AI chip:
- Performance: 2.5 PetaFLOPS FP8
- Memory: 64GB HBM2e
- Focus: Inference optimization for Azure OpenAI
Usage
- Powers Azure OpenAI Service (GPT-4, GPT-4o inference)
- GitHub Copilot runs on Maia
- Microsoft claims 40% cost reduction vs NVIDIA for equivalent inference
- Currently deployed in 5 Azure regions, expanding to 15 by end of 2026
Apple M4 Ultra: On-Device AI
Specs
- GPU: 40-core GPU
- Neural Engine: 32-core (38 TOPS)
- Memory: 192GB unified (bandwidth: 800 GB/s)
- Can run: 600B parameter models on-device
Apple Intelligence Performance
- Runs 3B parameter models for on-device Siri, writing tools, image generation
- 30x faster than iPhone 15 Pro's A17 Pro for ML tasks
- Privacy advantage: no data leaves the device
Meta MTIA: Training the Future
MTIA v2 (2025)
Meta's second-generation Training and Inference Accelerator:
- Performance: 2x improvement over MTIA v1
- Focus: Recommendation systems and ranking models
- Not for: Large language model training (still uses NVIDIA for that)
Meta's internal chips handle trillions of inference requests daily for Facebook and Instagram recommendations, freeing NVIDIA GPUs for LLM training.
The Economics of Custom Silicon
Why Build Custom Chips?
| Company | Annual AI Compute Spend | Reason for Custom |
|---|---|---|
| $35B+ | Cost reduction + performance | |
| Amazon | $25B+ | Customer lock-in + margins |
| Microsoft | $20B+ | Azure competitiveness |
| Meta | $15B+ | Recommendation scale |
| Apple | $10B+ | On-device privacy |
Break-Even Analysis
Building a custom AI chip costs approximately:
- Design: $200-500M
- Tape-out: $50-100M
- Total NRE: $300-600M
Break-even typically occurs at $2-5B in cumulative compute savings over 3-5 years. Google, Amazon, and Microsoft have all passed this threshold.
What's Next: 2026-2027
- NVIDIA Rubin (2027): Next-gen architecture, expected 2-3x over Blackwell
- Google TPU v7: Under development, rumored 10x improvement
- Amazon Trainium 4: Expected late 2026
- Microsoft Maia 200: Second-gen, targeting training workloads
- Intel Gaudi 3: Competing on price, not performance
The AI chip market is projected to reach $100B by 2028 (Gartner), with custom silicon capturing 25-30% of the market -- up from ~5% today.
Frequently Asked Questions
What is the best AI chip in 2026?
NVIDIA's B200/GB200 remains the most powerful for training large models. However, for inference (running already-trained models), Amazon Trainium 3 offers 2.4x better price-performance. Google TPUs are the best option for Google Cloud customers. The "best" chip depends on your cloud provider, budget, and workload type.
Will custom chips replace NVIDIA?
Not entirely. Custom chips are optimized for specific workloads and only make economic sense at massive scale ($10B+ compute spend). NVIDIA's advantage is flexibility -- one chip for many workloads. The market is shifting toward a hybrid model where companies use custom chips for their most common workloads and NVIDIA for everything else.
How much does an AI chip cost?
Prices range from $1,200 (Google TPU v5e hourly rental) to $60,000 (NVIDIA B300 purchase). Cloud rental is typically more cost-effective than purchasing for most organizations. The real cost consideration is total cost of ownership including power, cooling, and software development.
Conclusion
The AI chip market is undergoing its biggest transformation since NVIDIA's CUDA platform created the GPU computing market in 2007. While NVIDIA still dominates with 80%+ market share, Google, Amazon, Microsoft, and Apple are investing billions in custom silicon -- and the results are showing up in real benchmarks and real cost savings. For developers and businesses, the choice of AI chip increasingly depends on your cloud provider, workload type, and budget.
Was this article helpful?
Frequently Asked Questions
Stay in the loop
Get the latest tech news and AI insights delivered to your inbox. No spam, unsubscribe anytime.
TechVeb Team
Your trusted source for the latest in technology, AI innovations, and digital trends. We bring you in-depth analysis, expert reviews, and comprehensive guides.
Learn more about us →Continue Reading
View all →
Anthropic Watermarks Claude Text for EU AI Act Compliance
Anthropic will watermark text from Claude AI models using C2PA standards to comply with the EU AI Act's new transparency rules. Learn how it works.

AI Agents Guide: How Autonomous AI Systems Work
Learn how AI agents use planning, reasoning, and tools to execute complex tasks autonomously. Discover key use cases, architectures, and market impact.

The Future of AI: What to Expect in 2026 and Beyond
An in-depth analysis of the AI trends, breakthroughs, and predictions shaping 2026. From multimodal models to AI agents, here is what matters.

Moody's Warns Banks' AI Push Is Making Them More Dependent on Big Tech
Financial institutions racing to deploy AI may be trading operational risk for concentrated vendor risk, according to a new Moody's warning.

AI in Media & Entertainment (2026 Guide)
Explore how AI is transforming media and entertainment in 2026, from VFX to music creation. Learn key trends, tools, and production efficiency gains.

AI for Smart Cities in 2026: Urban Tech Guide
Discover how AI powers smart cities in 2026. Learn about traffic control, energy optimization, and urban tech strategies transforming modern cities.