📚 All Articles
200 guide(s) — regularly updated
MiniMax M3.1 Flash Preview: China continues its war of small, fast models
Discover MiniMax M3.1 Flash Preview, the new fast small Chinese model quietly launched on September 27, 2026, without a model card or press release.
Hindsight: AI agent memory becomes a product in its own right (+1,653 stars in one day)
Hindsight, the open source memory layer for AI agents, is blowing up on GitHub with +1,653 stars in one day. AI agent memory is becoming a product.
Laya: the open-source 421M param "decision model" that wants to replace the LLM in your agents
Laya, the open-source 421M-param decision model that wants to replace the LLM for routing, scoring and validating your AI agents' decisions.
**ElevenLabs Eleven v4: Voice Generation Levels Up with Turbo and Real-Time Latency**
ElevenLabs Eleven v4: studio-quality voice generation and real-time latency with Turbo. The new standard for AI conversational agents.
Claude Sonnet 5.5: Anthropic targets the mid-market with a model that's 30% faster and 30% cheaper
Claude Sonnet 5.5: Anthropic targets the mid-market with a model 30% faster and 30% cheaper. An analysis of the new model in the Claude 5.5 family.
"Do not guess": a single sentence reduces hallucinated fields from 71% to 20% across 16 frontier models
Discover how the phrase "Do not guess" cuts LLM hallucinated fields from 71% to 20% across 16 frontier models. A simple trick for your pipelines.
Here's the English translation: **Title:** Open models already serve 56% of tokens on Vercel: the silent migration of enterprises away from OpenAI and Anthropic
Open models run 56% of Vercel's tokens: discover how companies are quietly migrating away from OpenAI and Anthropic to open source
Naive-N0.5-Flash: the open-weight 309B MoE with no full-attention layers at all aims for 2,000 tokens/s per user
Naive-N0.5-Flash: NaiveAI's open-weight 309B-parameter MoE without full-attention aims for 2,000 tokens/s per user on Hugging Face.
Claude breaks the theoretical physics computation record: the Yang-Mills amplitude at 9 loops for $2,000 of compute
Claude breaks a theoretical physics record: the 9-loop Yang-Mills amplitude computed for $2,000 of compute. A historic AI challenge.
NVIDIA Open Agent Safety Platform: OpenShell and Sentry want to lock up AI agents before they escape
NVIDIA Open Agent Safety Platform: OpenShell and Sentry to contain AI agents and prevent their escape. A breakdown of this new solution.
Gemini 3.8 Live with Live Avatar: Google adds real-time visual presence to its conversational assistant
Google announces Gemini 3.8 Live with Live Avatar: a conversational assistant with real-time visual presence that listens, sees and speaks.
China-US: The World's First AI Hotline — Beijing and Washington Open a Communication Channel Dedicated to AI Incidents
China-US: the world's first AI red line. Beijing and Washington open a communication channel dedicated to artificial intelligence incidents.