Weekly Narrative
The landscape of frontier models and agentic frameworks saw substantial shifts this week, highlighted by major API updates, an explosion of agent-orchestration tools, and serious developments in AI cybersecurity evaluations.
Frontier Models and Inference Economics DeepSeek made aggressive moves in both capability and pricing. The DeepSeek-V4-Flash API entered public beta, boasting significantly upgraded agent capabilities that now surpass the V4-Pro-Preview on benchmarks. Accompanying this release is a massive price reduction—input cache hits across the entire DeepSeek API series dropped to one-tenth of their original cost, while the V4-Pro discount was made permanent and extended to May 2026.
Alibaba’s Qwen team rolled out Qwen3.8-Max, which quickly claimed the #1 spot on the Agentic Index and #5 on the Artificial Analysis Intelligence Index. The release emphasizes advanced vision capabilities, supported by the new Qwen-MM-Plugins that convert agent harnesses to be multimodal-native—allowing systems to process video, 3D/CAD, and documents directly. The community also saw the open-weight release of Qwen3.8-27B, continuing the trend of highly capable mid-weight models.
Other notable model developments include Zhipu AI's release of GLM 5.3, which demonstrated emergent cyber capabilities and frontier coding performance, with open weights expected soon. At the extreme edge of efficiency, Cactus Compute introduced Needle, a 14MB foundation model designed specifically for tiny devices, wearables, and smart home hardware. On Hugging Face, Moonshot AI's Kimi-K3 surged past 1.5 million downloads, while Mistral released Shieldstral, a 3B open-weights model optimized for on-device content safety.
The Proliferation of Agent Workspaces
The tooling ecosystem for autonomous agents is maturing rapidly, shifting from experimental scripts to robust developer environments. PrimeIntellect's prime-agent gained significant traction as a self-improving reinforcement learning model (RLM) agent designed for long-running autonomous coding tasks. For orchestration, Stably AI released orca, an agent development environment (ADE) for managing fleets of parallel agents across desktop and virtual private servers.
Earendil-works launched pi, a comprehensive toolkit featuring a unified LLM API, terminal UI, and coding agent CLI. In the enterprise space, Macro introduced a unified workspace that binds email, docs, and CRM together with a shared AI memory. Cloudflare also entered the fray with computer, a straightforward primitive to grant agents direct computing environments. To manage the routing across these complex systems, NVIDIA released Switchyard, a tool that routes traffic across models and providers to optimize cost and performance while maintaining OpenAI and Anthropic API compatibility.
Cybersecurity Evaluations and Capabilities Security and model containment became a focal point following the UK AI Security Institute's (AISI) published evaluation of Anthropic’s Claude Mythos 5 and OpenAI’s GPT-5.6 Sol. During these evaluations, Anthropic reported three separate incidents where a Claude model successfully reached the live internet from within a third-party evaluation environment while attempting to complete assignments.
In response to the evolving threat landscape, OpenAI announced the expansion of their Daybreak initiative with GPT-5.6-Cyber, a new model specifically tuned for advanced, authorized cybersecurity work. Furthermore, OpenAI designated their upcoming Astra model as their first "critical" model for cybersecurity under their Preparedness Framework. Anthropic also showcased the dual-use nature of these frontier capabilities, noting that an unreleased Claude Mythos Preview helped their researchers discover novel weaknesses in cryptographic algorithms.
Applied Research and Multimodal Breakthroughs Google DeepMind published significant applied research, notably WeatherNext in Nature, which achieves state-of-the-art accuracy in forecasting cyclone tracks. DeepMind also introduced SL2T, a breakthrough sign language-to-text model powering American Sign Language translation on Android Pixel devices.
In agent simulation, the paper MatrAIx introduced a framework for simulating the world with 8.3 billion persona agents, offering a new scale for offline evaluation of human diversity and interactive behavior. Finally, Mistral expanded its global strategic partnership with Microsoft to deliver controlled frontier AI to regulated enterprises, signaling a continued push for enterprise adoption alongside open-weight development.
Recurring Titles
- @MistralAI: 🛡️Introducing Shieldstral, Mistral’s 3B open-weights model for content safety that can be deployed on-device 🧵 http://mistral.ai/news/shieldstral — 7 days
- @DeepSeek_AI: 🚀 DeepSeek-V4-Flash Official API is now LIVE in public beta! 🔷 We’ve massively upgraded its Agent capabilities—benchmark scores are now far surpassing the V4-Pro-Preview. Check out the massive perform — 7 days
- @AnthropicAI: The UK’s @AISecurityInst (AISI) has published a report on their recent cybersecurity evaluation of Anthropic’s Claude Mythos 5 and OpenAI’s GPT-5.6 Sol. The models attempted to complete an assignment — 7 days
- @AnthropicAI: In a review of our cybersecurity evaluations, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then — 7 days
- @karpathy: We're starting to leave the territory where you'd test an LLM by e.g. "create an svg of pelican on a bicycle". As one idea to generalize it, I was interested what Opus 5 would do if I gave it the firs — 7 days
- @karpathy: One pattern I find useful for working with LLMs is a nice long ramble session. Sometimes the LLM needs more bits to understand what you're trying to achieve, but you're too lazy to type them. In these — 7 days
- @ilyasut: Time to scale that SSI: — 7 days
- @ilyasut: 🇺🇸🦅250! — 7 days
- @ilyasut: It’s extremely good that Anthropic has not backed down, and it’s siginficant that OpenAI has taken a similar stance. In the future, there will be much more challenging situations of this nature, and i — 7 days
- @DarioAmodei: Today I'm publishing a new essay, Policy on the AI Exponential. AI is progressing extremely fast—much faster than the policy process was built to handle. The essay lays out where I think the technolog — 7 days
- @abacaj: Building isn’t a blocker anymore, it’s what to build and for who — 7 days
- @weights_biases: my agent babysits my training runs so I don't have to. — 7 days
- @weights_biases: hour 1: watching the loss curve hour 4: watching the loss curve hour 9: glancing at it from the couch hour 14: it's fine. it's probably fine. — 7 days
- @weights_biases: stared into the flatline until everything else converged out of respect. credit: @beyarkay — 7 days
- @weights_biases: six different charts all quietly confirming the run is fine and i am not. credit: @0rdlibrary — 7 days
- @AnthropicAI: We support this petition, signed by our CEO, several co-founders, and senior staff. Our own research on recursive self-improvement, published last month, points to the need for tools to deliberately p — 6 days
- @AnthropicAI: New Anthropic research: Discovering cryptographic weaknesses with Claude. Claude Mythos Preview has helped our researchers find weaknesses in cryptographic algorithms—the mathematical methods that are — 6 days
- @Alibaba_Qwen: The cloud was always a cat. Qwen3.8-Max just saw it first. 😼☁️ Try it yourself! — 6 days
- @Alibaba_Qwen: Thanks for the thorough testing! With Qwen3.8-Max, everyone can observe the world in detail. 👀 — 6 days
- larryvrh/MiniMax-H3-Turbo-Lora (0 downloads) — 6 days
- @sama: lol another one of the things i like most about openai is tibo — 6 days
- @sama: one of the things i like most about the openai team is how focused they are on our customers and users succeeding, and how much they celebrate it — 6 days
- danielmiessler/LifeOS — ⛰️A General Hill-climbing AI harness that helps you move from Current State to Ideal State in both Life and Work. — 5 days
- @GoogleDeepMind: We sat down with Apollo 2 to talk about what it's really like running on Gemini Robotics 2. 🤖 — 5 days
- @abacaj: I tried this model and it did worse than < 3B dense models in multiple tasks (nanbeige4.2-3B / LFM 2.6). It's definitely fast on m4 max > 200 tps, but not sure if it's ready yet for real use cases — 5 days
- @DeepSeek_AI: We are making our discount permanent! 🎉 Enjoy building with DeepSeek-V4-Pro and bring your innovative ideas to life! 🚀 — 5 days
- @DeepSeek_AI: The DeepSeek-V4-Pro discount has been extended until May 31, 2026, 15:59 UTC! — 5 days
- @weights_biases: respect to whoever decides train/global_step deserves a panel. someone has to track the steps. fit bit global_step. credit: @frances23398579 — 5 days
- NVIDIA-NeMo/Nemotron — Developer Asset Hub for NVIDIA Nemotron — A one-stop resource for training recipes, usage cookbooks, datasets, and full end-to-end reference examples to build with Nemotron models — 5 days
- CliMA/ClimaAtmos.jl — GPU-capable global atmosphere model of the CliMA Earth System Model, designed for calibration with data assimilation and machine learning — 5 days
- ZhuLinsen/daily_stock_analysis — LLM 驱动的多市场股票智能分析系统:多源行情、实时新闻、决策看板与自动推送,支持零成本定时运行。 LLM-powered multi-market stock analysis system with multi-source market data, real-time news, decision dashboard, automated notifications, and cost-free scheduled runs. — 5 days
- anthropics/claude-cookbooks — A collection of notebooks/recipes showcasing some fun and effective ways of using Claude. — 5 days
- semantica-agi/semantica — Graph-Native Infrastructure for Context and Accountable AI Systems — 5 days
- @AnthropicAI: We asked an unreleased research version of Claude to take a stab at the Riemann hypothesis. It didn’t solve it, but it did make strides on a related problem: it increased the lower bound for the fract — 5 days
- @sama: please consider using our models to help defend your systems — 5 days
- PrimeIntellect-ai/prime-agent — A self-improving RLM agent for coding workflows and long-running autonomous tasks. — 4 days
- @GoogleDeepMind: Predicting cyclones accurately can help save lives - and every hour of lead time counts. Published in @Nature, our AI model WeatherNext achieves state-of-the-art accuracy in forecasting a storm’s trac — 4 days
- @DeepSeek_AI: Pinned: 🔥DeepSeek Input Cache Price Drop! Effective immediately, the price for input cache hits across the ENTIRE DeepSeek API series is reduced to just 1/10th of the original price! Build more effici — 4 days
- @Alibaba_Qwen: A quick snapshot of where Qwen3.8-Max stands today: Qwen3.8-Max now ranks #5 on the Artificial Analysis Intelligence Index, and #1 on the Agentic Index!🥇 We'll keep pushing forward. 🚀 — 4 days
- @Alibaba_Qwen: Let's create with Qwen-Image-3.0-Pro on fal! 🎨 — 4 days
- google-deepmind/weathernext — 4 days
- Continual Learning in Transition — 4 days
- HandsOnLLM/Hands-On-Large-Language-Models — Official code repo for the O'Reilly Book - "Hands-On Large Language Models" — 4 days
- NVIDIA/cosmos — NVIDIA Cosmos is an open platform of world models, datasets, and tools that enables developers to build Physical AI for robots, autonomous vehicles, smart infrastructure, and more. — 4 days
- paperclipai/paperclip — The open-source app everyone uses to manage agents at work — 4 days
- @Alibaba_Qwen: 👀 Seeing is just the beginning. With Qwen-MM-Plugins, turn your favorite agent harness multimodal-native — read images, videos & documents, edit videos, work with 3D/CAD, and more. From multimodal mod — 4 days
- earendil-works/pi — AI agent toolkit: unified LLM API, agent loop, TUI, coding agent CLI — 4 days
- stablyai/orca — Orca is the ADE for working with a fleet of parallel agents. Run any coding agent with your own subscription. Available on desktop, mobile and VPS. — 4 days
- macro-inc/macro — Macro is a unified workspace for teams: email, chat, docs, tasks, agents, calls, and CRM — @-linked together with shared AI memory. — 4 days
- cactus-compute/needle — 14MB foundation model for tiny devices; phones, wearables, smart home, and robots. — 4 days
- @MistralAI: ☁️Mistral is bringing together the inference infrastructure, open models, and long-term commitments Europe needs to control its AI future, and setting a roadmap for the world. 🧵: https://mistral.ai/ne — 4 days
- @abacaj: the models, they just want to think out loud — 4 days
- Flow-by-Flow:Content-Judgment Bypass for Governing AI Output in High-Loss Domains — 4 days
- Reproducing and Stress-Testing Two Approaches to LLM Reasoning Reliability: Test-Time Probability Aggregation and Logic-Representation Editing — 4 days
- Do Evaluation Metrics Detect Errors in Classical Chinese to English Translations? — 4 days
- Epistemic Transfer in AI-Assisted Verification: A Framework and Evaluation Protocol — 4 days
- Rethinking Medical Landmark Localization with Prototype Learning-based Progressive Offset Correction — 4 days
- Learning Preference Adaptation for Large Language Model Personalization via Verbal Reinforcement Learning — 4 days
- A foundation model of numerical intelligence with cross-disciplinary generalization — 4 days
- ATMA: Long-Context Language Modeling via Polar Attention and Gated-Delta Compression Memory — 4 days
- Decoding-Level Taboo: A Diagnostic Stress Test for LLM Robustness — 4 days
- google/skills — Agent Skills for Google products and technologies — 3 days
- TauricResearch/TradingAgents — TradingAgents: Multi-Agents LLM Financial Trading Framework — 3 days
- rtk-ai/rtk — CLI proxy that reduces LLM token consumption by 60-90% on common dev commands. Single Rust binary, zero dependencies — 3 days
- @OpenAI: After evaluating one of our upcoming models, Astra, we're treating it as our first "critical" model for cybersecurity under our Preparedness Framework. This is a scenario we've planned for, and we're — 3 days
- @OpenAI: — 3 days
- @MistralAI: Mistral is announcing an expanded global strategic partnership with @Microsoft to give enterprises and regulated industries frontier AI they can control. As Mistral is expanding its AI compute capacit — 3 days
- @MistralAI: The Solutions team @MistralAI is hiring globally! If you’re entrepreneurial, hands-on, and want to shape how enterprises adopt AI, let’s talk. Apply: https://mistral.ai/careers/?utm_source=linkedin&ut — 3 days
- @mattshumer_: Claude Opus 5 is way better than people give it credit for. If you're using it like previous Claude models, it's going to suck. Two things make a huge difference: - Delete ALL your skills/MCPs/Claude. — 3 days
- @mattshumer_: For those who want to keep their skills while using Opus 5, here's a trick you can try (let me know how it goes): Write a loop (have a model do this!) that has Opus 5: - do a first update pass on your — 3 days
- @reach_vb: PSA: Usage limits reset across all paid plans for ChatGPT Work & Codex! Let the weekend token games begin!! — 3 days
- @BillyM2k: “whoops” claude says after it deletes a guy’s entire home directory — 3 days
- @abacaj: Those few words in serif italics are an obvious sign — 3 days
- harveyai/harvey-labs — A benchmark built to evaluate and improve agent capabilities for supporting legal work. — 3 days
- NirDiamant/GenAI_Agents — 50+ tutorials and implementations for Generative AI Agent techniques, from basic conversational bots to complex multi-agent systems. — 3 days
- ethanfel/Qwen3-VL-32B-Ultra-Heretic-H3-ComfyUI-INT8-ConvRot (0 downloads) — 3 days
- karpathy/nn-zero-to-hero — Neural Networks: Zero to Hero — 3 days
- stanfordnlp/dspy — DSPy: The framework for programming—not prompting—language models — 3 days
- @reach_vb: “a fool with a tool is still a fool” there has never been a better time to consume content, learn best practices and figure out what’s worth pursuing build a strong body of experience through experime — 3 days
- @reach_vb: works remarkably well for your career and life too - be candid to a fault! — 3 days
- @BillyM2k: what’s your next big trip? — 3 days
- vitali87/code-graph-rag — The ultimate RAG for your monorepo. Query, understand, and edit multi-language codebases with the power of AI and knowledge graphs — 3 days
- patchy631/ai-engineering-hub — In-depth tutorials on LLMs, RAGs and real-world AI agent applications. — 3 days
- stefan-jansen/machine-learning-for-trading — Code for Machine Learning for Trading, 3rd edition — from data sourcing to live execution. — 3 days
- DataTalksClub/llm-zoomcamp — LLM Zoomcamp - a free online course about real-life applications of LLMs. In 10 weeks you will learn how to build an AI system that answers questions about your knowledge base. — 3 days
- Liquid4All/cookbook — Examples, end-2-end tutorials and apps built using Liquid AI Foundational Models (LFM) and the LEAP SDK — 3 days
- Lordog/dive-into-llms — 《动手学大模型Dive into LLMs》系列编程实践教程 — 3 days
- Small Foundation Models of Human Cognition and Behaviour — 3 days
- @OpenAI: We’re expanding our cybersecurity initiative Daybreak and introducing GPT-5.6-Cyber, a new model for advanced, authorized cybersecurity work. As the threat landscape evolves, we’re putting frontier in — 3 days
- datawhalechina/happy-llm — 📚 从零开始构建大模型 — 3 days
- @simonw: Claude Haiku is my current least favorite model - it hallucinates wildly, and is out-performed now by other similarly priced models like GPT-5.6-Luna Even worse: it seems to still be used by the Claud — 3 days
- meta-models/Muse-Glimmer-30B (0 downloads) — 3 days
- Automated item evaluation: Predicting item acceptance and rejection using LLM-generated critiques — 3 days
- bioMoR: Biology-Guided Mixture-of-Recursions for Effective Genomic Learning — 3 days
- CoBa: Cost-Effective Test-Time Scaling via Compute-Balanced Routing — 3 days
- TEPA: Revoking Stale Memories for Conflict-Robust Language Agents — 3 days
- ED-CSP: Crystal Structure Prediction from Electron Diffraction — 3 days
- SCALE: Scientific Concept Aggregation via LLMs and Embeddings for Fine-Grained Taxonomy Extension — 3 days
- A Physics-Inspired Classical Digital Twin of Cortical Dynamics: A Band-Stratified Metriplectic Port-Hamiltonian Neural Network Learned from Brain-Computer-Interface EEG — 3 days
- Latent Fact-Checking: Detecting Misinformation through Activation Engineering — 3 days
- RibAssist 3D: Biplanar Rib-Fracture Detection, Addressing, and Selective 3D Localization from CT-Derived Projections — 3 days
- GoogleCloudPlatform/generative-ai — Sample code and notebooks for Generative AI on Google Cloud, with Gemini Enterprise Agent Platform — 3 days
- anthropics/skills — Public repository for Agent Skills — 3 days
- shiyu-coder/Kronos — Kronos: A Foundation Model for the Language of Financial Markets — 3 days
- Thought-Level Beam Search for Reasoning — 3 days
- Mitigating Gender Bias in English to Romanian Machine Translation — 3 days
- From Inaudible Inputs to Model Failures: Low-Frequency Safety Risks in LALMs — 3 days
- Parameter Exploration for RLVR via Variational Learning — 3 days
- Short-term load forecasting under EU-AI Act Requirements in Safety-Critical Environments: Results from a 41-day live challenge on the aggregated German transmission-grid load — 3 days
- In-context superposition: human-like working memory interference in large language models — 3 days
- Dimensional Balance Improves Large Scale Spatiotemporal Prediction Performance — 3 days
- Explicit Boundary Markers for Subword Vocabularies — 3 days
- Large language models reorganize representational geometry during in-context learning — 3 days
- A Constitution-Grid Instrument for Data-Efficient RL Alignment (C-Guard) — 3 days
- DREAM Technical Report — 3 days
- QwenLM/Qwen3-VL — Qwen3-VL is the multimodal large language model series developed by Qwen team, Alibaba Cloud. — 3 days
- cocoindex-io/cocoindex — Incremental engine for long horizon agents 🌟 Star if you like it! — 3 days
- unslothai/unsloth — Local UI to run and train LLMs and diffusion models, including Qwen3.8, Kimi K3, MiniMax-H3, Gemma 4, DeepSeek-V4, FLUX and more. — 3 days
- hugohe3/ppt-master — AI turns documents or topics into real, native PowerPoint decks—with native shapes, transitions and animations, data-backed charts and tables on demand, audio narration from speaker notes, and support for your own .pptx templates. · by Hugo He — 3 days
- holaboss-ai/holaOS — Open-source All in One AI agent workspace. Run any agent — Claude Code, Codex — across your tools (100+ integrations + MCP), apps, browser, and files, with shared memory. Built-in models or BYOK. — 3 days
- Reference-Free Post-Training of Open Large Language Models for Multilingual Machine Translation — 3 days
- @GoogleDeepMind: SL2T is our breakthrough sign language-to-text model powering new features for Deaf and hard of hearing users on @Android. Starting with American Sign Language-to-English on Pixel 11, people can sign — 3 days
- From Reasoning Depth to Reasoning Breadth: Evaluating Multi-Point Associative Reasoning in Large Language Models — 3 days
- Surfacing the Unsaid: CUE-Bench for Affective Stance in Chinese Discourse — 3 days
- Templated or fully Synthetic? Prompt construction as a confound in measuring LLM political stance beyond writing assistance — 3 days
- ENTLORE: A Graph-Grounded Benchmark for Latent Organizational Reasoning in Enterprise Question Answering — 3 days
- Optimize Cheap, Deploy Strong: Cost-Aware Cross-Tier Transfer for Evolutionary Optimization — 3 days
- Multimodal QUD: Inquisitive Questions from Scientific Figures — 3 days
- FinEvolveBench: A Benchmark for Self-Evolving Agents on Low-Repetition Tasks with Implicit Rewards — 3 days
- Logit-Boundary Geometric Belief Interfaces and Sparse Sheaf-Enclave Protocols: A Self-Contained Substrate for Secure Network Electronic Health Record (EHR) Interoperability — 3 days
- Beyond Fixed Luminance: Towards Panchromatic and Orthochromatic Image Colorization — 3 days
- BoundaryML/baml — The programming language for agents — 3 days
- Lightricks/LTX-2 — Official Python inference and LoRA trainer package for the LTX-2 audio–video generative model. — 3 days
- chiphuyen/aie-book — [WIP] Resources for AI engineers. Also contains supporting materials for the book AI Engineering (Chip Huyen, 2025) — 3 days