CATEGORY
Model platforms
Foundation-model platforms, open releases and infrastructure for deploying AI at scale.
EDITOR'S SELECTION
Latest analysis
NVIDIA puts Groq 3 LPX into production for agent inference
The Groq 3 LPX extension to Vera Rubin targets the token-generation stage of long-context agents, while NVIDIA reports early adoption by Nebius, CoreWeave and SpaceXAI.
· Ravi Kapoor · 8 min readAn ONNX portability checklist for moving models between runtimes
ONNX makes model exchange possible, but it cannot guarantee identical behavior in every runtime. Use this validation sequence before treating a converted graph as deployable.
· Elena Morris · 10 min readLiquid AI releases draft models for faster local inference
Liquid AI's LFM2.5-DSpark checkpoints use speculative decoding to accelerate several LFM2.5 models. The vendor reports large gains, but teams should reproduce them on their own hardware and workloads.
· Elena Morris · 8 min readWhat speculative decoding changes for local AI agents
Draft-and-verify inference can make an agent feel faster without replacing its target model. The benefit is workload-specific, and poor evaluation can hide memory and reliability costs.
· Ethan Brooks · 9 min readContext windows are capacity limits, not memory guarantees
A large context window tells you how much a model can accept, not whether it will use every detail accurately. Test retrieval, instruction stability, latency and cost together.
· Elena Morris · 9 min readHow to evaluate an AI model API before you commit
A leaderboard cannot tell you which model belongs in your product. Measure quality, latency, cost, safety and operational fit together.
· Elena Morris · 10 min readThe model race is becoming local, multilingual and sovereign
Sarvam’s open 30B and 105B models ignited a broader conversation about who builds a country’s AI layer—and whether global benchmarks capture the languages and workflows that matter locally.
· Ravi Kapoor · 7 min readOpenAI launches GPT-6 Astra for ChatGPT Work, Codex and the API
OpenAI's newest flagship model targets computer use and professional workflows, with API pricing beginning at $10 per million input tokens and $50 per million output tokens.
· Ravi Kapoor · 10 min readHugging Face Candle fixes an ONNX Reshape compatibility gap
A focused Candle commit now honors ONNX's allowzero behavior and corrects -1 dimension inference. The patch shows why model portability depends on operator edge cases, not file conversion alone.
· Ethan Brooks · 8 min readNVIDIA Cosmos 3 turns the world model into a robotics stack
NVIDIA's Cosmos 3 announcement connects perception, simulation and action generation. Here is what the full-stack approach means for robotics teams—and what still needs independent testing.
· Ethan Brooks · 8 min readGPT-5.6 arrives across ChatGPT, Codex and the OpenAI API
OpenAI's July flagship release brings one model family to consumer chat, coding agents and the API, followed by substantial price cuts for its Luna and Terra variants.
· Ravi Kapoor · 7 min readClaude Sonnet 5 brings a more agentic default model to Claude Code
Anthropic's new Sonnet model targets coding, tool use and autonomous work while keeping a lower price point than its Opus-class frontier models.
· Ethan Brooks · 8 min readFLUX 3 expands Black Forest Labs from images into video and audio
The early-access model jointly learns from images, video and sound, signaling a move from specialist image generation toward a multimodal foundation for visual intelligence.
· Noah Chen · 7 min read