Weekly Narrative
The frontier model landscape saw significant disruption this week, led by an unreleased OpenAI model—widely rumored to be Astra or GPT-5.6 Sol—which successfully solved 10 long-standing open problems in mathematics and theoretical computer science using roughly $2,000 worth of compute. Concurrently, Alibaba's Qwen 3.8 Max claimed the top spot on the Artificial Analysis agentic index, edging out Anthropic's Opus 5. Open-weight models also marked major milestones: DeepSeek launched the public beta of its V4-Flash API, boasting agentic capabilities that surpass V4-Pro-Preview, while the local DeepSeek-V4-Flash-0731 variant achieved an intelligence index score of 50, matching frontier models from earlier this year. Moonshot AI's Kimi-K3 saw massive traction on Hugging Face, with community engineers successfully clustering 16x GB10 systems to run the full model locally at high throughput.
On the infrastructure front, the push for local and efficient inference yielded impressive engineering feats. The lyogavin/airllm repository demonstrated 70B model inference on a single 4GB GPU, severely lowering the hardware floor for running massive parameter models. At the hardware level, AMD acquired Taalas to boost inference performance by etching neural networks directly into silicon. Mistral continued its specialized deployments, announcing a global partnership with Microsoft alongside two targeted releases: Shieldstral, a 3B open-weights model designed for on-device content safety, and Robostral Navigate, an 8B model for embodied robotics navigation.
Agentic architectures shifted from single-agent prototypes to robust, multi-agent enterprise infrastructure. Tencent introduced TencentDB Agent Memory, a team-level hub that parses conversations, documents, and code into reusable, cross-agent memory assets. Bytedance open-sourced deer-flow, a long-horizon SuperAgent harness featuring sandboxes, skills, and subagent orchestration for tasks spanning minutes to hours. For runtime management, huangruiteng/loopx emerged as a lightweight loop engineering kernel handling quota-aware auto-wakes and verifiable handoffs, while katanemo/plano offered an AI-native proxy data plane for agentic routing and observability. Context control also saw innovation with `yvgude/
Recurring Titles
- TencentCloud/TencentDB-Agent-Memory — TencentDB Agent Memory is a team-level memory hub for AI Agents — turning conversations, docs, and code into four reusable memory assets (Chat Memory, Skill, LLM-Wiki, Code-Graph) that are governed, shared, and equipped across agents and frameworks. — 7 days
- DataExpert-io/data-engineer-handbook — This is a repo with links to everything you'd ever want to learn about data engineering — 7 days
- @AnthropicAI: In a review of our cybersecurity evaluations, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then — 7 days
- @AnthropicAI: We support this petition, signed by our CEO, several co-founders, and senior staff. Our own research on recursive self-improvement, published last month, points to the need for tools to deliberately p — 7 days
- @AnthropicAI: New Anthropic research: Discovering cryptographic weaknesses with Claude. Claude Mythos Preview has helped our researchers find weaknesses in cryptographic algorithms—the mathematical methods that are — 7 days
- @MistralAI: Mistral is announcing an expanded global strategic partnership with @Microsoft to give enterprises and regulated industries frontier AI they can control. As Mistral is expanding its AI compute capacit — 7 days
- @MistralAI: The Solutions team @MistralAI is hiring globally! If you’re entrepreneurial, hands-on, and want to shape how enterprises adopt AI, let’s talk. Apply: https://mistral.ai/careers/?utm_source=linkedin&ut — 7 days
- @ilyasut: Time to scale that SSI: — 7 days
- @ilyasut: 🇺🇸🦅250! — 7 days
- @ilyasut: It’s extremely good that Anthropic has not backed down, and it’s siginficant that OpenAI has taken a similar stance. In the future, there will be much more challenging situations of this nature, and i — 7 days
- @DarioAmodei: Today I'm publishing a new essay, Policy on the AI Exponential. AI is progressing extremely fast—much faster than the policy process was built to handle. The essay lays out where I think the technolog — 7 days
- @DeepSeek_AI: 🚀 DeepSeek-V4-Flash Official API is now LIVE in public beta! 🔷 We’ve massively upgraded its Agent capabilities—benchmark scores are now far surpassing the V4-Pro-Preview. Check out the massive perform — 7 days
- @DeepSeek_AI: We are making our discount permanent! 🎉 Enjoy building with DeepSeek-V4-Pro and bring your innovative ideas to life! 🚀 — 7 days
- @DeepSeek_AI: The DeepSeek-V4-Pro discount has been extended until May 31, 2026, 15:59 UTC! — 7 days
- @DeepSeek_AI: Pinned: 🔥DeepSeek Input Cache Price Drop! Effective immediately, the price for input cache hits across the ENTIRE DeepSeek API series is reduced to just 1/10th of the original price! Build more effici — 7 days
- @weights_biases: hour 1: watching the loss curve hour 4: watching the loss curve hour 9: glancing at it from the couch hour 14: it's fine. it's probably fine. — 7 days
- @weights_biases: stared into the flatline until everything else converged out of respect. credit: @beyarkay — 7 days
- @weights_biases: six different charts all quietly confirming the run is fine and i am not. credit: @0rdlibrary — 7 days
- @weights_biases: respect to whoever decides train/global_step deserves a panel. someone has to track the steps. fit bit global_step. credit: @frances23398579 — 7 days
- JuliaRegistries/General — The official registry of general Julia packages — 7 days
- CliMA/ClimaAtmos.jl — GPU-capable global atmosphere model of the CliMA Earth System Model, designed for calibration with data assimilation and machine learning — 7 days
- CliMA/ClimaCoupler.jl — ClimaCoupler: bringing atmosphere, land, and ocean together — 7 days
- NousResearch/hermes-agent — The agent that grows with you — 6 days
- microsoft/generative-ai-for-beginners — 21 Lessons, Get Started Building with Generative AI — 6 days
- @GoogleDeepMind: This is how Gemini Robotics 2 helps @Apptronik’s Apollo 2 use whole body intelligence to pack for a sports game ↓ — 6 days
- @sama: team humanity — 6 days
- @BillyM2k: Image — 6 days
- SimplifyJobs/Summer2027-Internships — Summer 2026 software engineering, data science, AI, quant, product management, and hardware internship postings. Updated daily by Simplify and Pitt CSC. — 6 days
- @GoogleDeepMind: One brain. For any robot. 🤖 We’re launching Gemini Robotics 2: our next-generation physical AI bringing full body intelligence to humanoids, advanced dexterity, multi-robot teamwork and more. — 5 days
- @sama: cool use case of chatgpt work i heard last night: connect your family calendars and explain your kids' interests. every morning for the drive to school, have it make a podcast that talks about one kid — 5 days
- lyogavin/airllm — AirLLM 70B inference with single 4GB GPU — 5 days
- ed-donner/agents — Repo for the Complete Agentic AI Engineering Course — 5 days
- CliMA/ClimaLand.jl — Modular, GPU-capable land surface model of the CliMA Earth System Model, designed for data-driven parameterizations — 5 days
- JuliaLang/julia — The Julia Programming Language — 5 days
- donnemartin/system-design-primer — Learn how to design large-scale systems. Prep for the system design interview. Includes Anki flashcards. — 5 days
- @abacaj: Those few words in serif italics are an obvious sign — 5 days
- @abacaj: A friend of mine mentioned that he wanted to run LLMs locally. I immediately said, "I got a guy." I did not ask him which model he wanted to run. I did not ask how big the model was. He then asked who — 5 days
- microsoft/AI-For-Beginners — 12 Weeks, 24 Lessons, AI for All! — 4 days
- usekaneo/kaneo — 🎯 All you need. Nothing you don't. Open source project management that works for you, not against you. — 4 days
- farion1231/cc-switch — A cross-platform desktop All-in-One assistant for Claude Code, Codex, OpenCode, OpenClaw, Grok Build & Hermes Agent. Only official website: ccswitch.io — 4 days
- @karpathy: One pattern I find useful for working with LLMs is a nice long ramble session. Sometimes the LLM needs more bits to understand what you're trying to achieve, but you're too lazy to type them. In these — 4 days
- @abacaj: One shot is the new zero shot — 4 days
- microsoft/ML-For-Beginners — 12 weeks, 26 lessons, 52 quizzes, classic Machine Learning for all — 4 days
- jackfrued/Python-100-Days — Python - 100天从新手到大师 — 4 days
- moghtech/komodo — 🦎 a tool to build and deploy software on many servers 🦎 — 4 days
- CliMA/Oceananigans.jl — 🌊 Julia software for fast, friendly, flexible, ocean-flavored fluid dynamics on CPUs and GPUs — 4 days
- CliMA/ClimaCore.jl — GPU-capable dynamical core for the CliMA Earth System Model: spectral-element and finite-difference discretization tools — 4 days
- firecrawl/pdf-inspector — Fast Rust library for PDF inspection, classification, and text extraction. Intelligently detects scanned vs text-based PDFs to enable smart routing decisions. — 4 days
- When Bits Break Recourse: Counterfactual-Faithful Quantization — 4 days
- katanemo/plano — Plano is an AI-native proxy server and data plane for agentic apps. Smart LLM routing, observability, agent orchestration, and guardrails so you stay focused on your agents core logic. — 4 days
- huangruiteng/loopx — Lightweight loop engineering state kernel for long-running AI agent teams. Agent-loop agnostic across Codex, Claude Code, and other coding agents, with durable goals, quota-aware auto-wake, executable todos, evidence logs, and verifiable handoffs. — 4 days
- @MistralAI: 🛡️Introducing Shieldstral, Mistral’s 3B open-weights model for content safety that can be deployed on-device 🧵 http://mistral.ai/news/shieldstral — 4 days
- PICopilot: An LLM-based Agentic Framework for Assisting Photonic Integrated Circuit Design via Script Generation — 4 days
- @AnthropicAI: The UK’s @AISecurityInst (AISI) has published a report on their recent cybersecurity evaluation of Anthropic’s Claude Mythos 5 and OpenAI’s GPT-5.6 Sol. The models attempted to complete an assignment — 4 days
- @weights_biases: my agent babysits my training runs so I don't have to. — 4 days
- bytedance/deer-flow — An open-source long-horizon SuperAgent harness that researches, codes, and creates. With the help of sandboxes, memories, tools, skill, subagents and message gateway, it handles different levels of tasks that could take minutes to hours. — 3 days
- openai/codex — Lightweight coding agent that runs in your terminal — 3 days
- gitroomhq/postiz-app — 📨 The ultimate agentic social media scheduling tool 🤖 — 3 days
- zed-industries/zed — Code at the speed of thought – Zed is a high-performance, multiplayer code editor from the creators of Atom and Tree-sitter. — 3 days
- abus-aikorea/voice-pro — Gradio WebUI for creators and developers, featuring key TTS (Edge-TTS, kokoro) and zero-shot Voice Cloning (E2 & F5-TTS, CosyVoice), with Whisper audio processing, YouTube download, Demucs vocal isolation, and multilingual translation. — 3 days
- mrdbourke/zero-to-mastery-ml — All course materials for the Zero to Mastery Machine Learning and Data Science course. — 3 days
- @MistralAI: Announcing Robostral Navigate, our first model for embodied navigation: an 8B robotics navigation model that guides robots to autonomously perform tasks specified with natural language. Single RGB cam — 3 days
- @sama: i see your moore's law and i raise you 20x — 3 days
- @weights_biases: watching nanochat grow up in real time on an 8xH200 node running wandb in the terminals. credit: @karpathy — 3 days
- OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models — 3 days
- AI4Finance-Foundation/FinRL — FinRL®: Financial Reinforcement Learning. 🔥 — 3 days
- facebookresearch/segment-anything — The repository provides code for running inference with the SegmentAnything Model (SAM), links for downloading the trained model checkpoints, and example notebooks that show how to use the model. — 3 days
- Panniantong/Agent-Reach — Give your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees. — 3 days
- jamiepine/voicebox — The open-source AI voice studio. Clone, dictate, create. — 3 days
- can1357/oh-my-pi — ⌥ AI Coding agent for the terminal — hash-anchored edits, optimized tool harness, LSP, Python, browser, subagents, and more — 3 days
- pathwaycom/llm-app — Ready-to-run cloud templates for RAG, AI pipelines, and enterprise search with live data. 🐳Docker-friendly.⚡Always in sync with Sharepoint, Google Drive, S3, Kafka, PostgreSQL, real-time data APIs, and more. — 3 days
- microsoft/ai-agents-for-beginners — 18 Lessons to Get Started Building AI Agents — 3 days
- rustdesk/rustdesk — An open-source remote desktop application designed for self-hosting, as an alternative to TeamViewer. — 3 days
- @karpathy: We're starting to leave the territory where you'd test an LLM by e.g. "create an svg of pelican on a bicycle". As one idea to generalize it, I was interested what Opus 5 would do if I gave it the firs — 3 days
- rust-lang/rust — Empowering everyone to build reliable and efficient software. — 3 days
- n0-computer/iroh — IP addresses break, dial keys instead. A library that adds QUIC + NAT Traversal to your apps. — 3 days
- rust-lang/rust-clippy — A bunch of lints to catch common mistakes and improve your Rust code. Book:https://doc.rust-lang.org/clippy/ — 3 days
- ed-donner/llm_engineering — Repo to accompany my mastering LLM engineering course — 3 days
- Alishahryar1/free-claude-code — Use Claude Code, Codex and Pi for free from your terminal, app, IDE, or phone like OpenClaw (voice supported) — 3 days
- livekit/agents — A framework for building realtime voice AI agents 🤖🎙️📹 — 3 days
- K-Dense-AI/scientific-agent-skills — Turn any AI agent into an AI Scientist. The #1 Agent Skills library for science, used by 170,000+ scientists worldwide. 158 ready-to-use skills plus 100+ scientific databases covering biology, chemistry, medicine, and drug discovery. Compatible with Cursor, Claude Code, Codex, Pi, Antigravity, and the open Agent Skills standard. — 3 days
- Identifying Informative Environments for Cognition Parameter Inference via Bayesian Experimental Design — 3 days
- The Asymmetric Effects of Knowledge Distillation on Bias in Small Language Models — 3 days
- SCMA: Structure-Conditioned and Metal-Aware Flow Matching for CT Metal Artifact Reduction — 3 days
- The persuasive power of large language models does not depend on their perceived national origin — 3 days
- What Makes a Sale? Simulating End-to-End Seller--Buyer Retail Dynamics with LLM Agents — 3 days
- Deconstructing Off-Policy Ratios: Entropy-Scaled Trust Regions for Asynchronous Reinforcement Learning — 3 days
- DRIFT: Difficulty Routing Self-DIstillation with Rhythm-Gated Exploration and Success BuFfer Training — 3 days
- ECHO: Prune To Act, Trace To Learn With Selective Turn Memory In Agentic RL — 3 days
- Frontier AI performance across the business disciplines: a case-grounded benchmark of knowledge work and analytical reasoning — 3 days
- Benchmarking LLM Competence on Logical Inference over Probability Operators — 3 days
- Studying quantization trade-offs for efficient inference deployment in machine translation — 3 days
- @OpenAI: An internal version of our next major model produced 10 new results on long-standing open problems in mathematics and theoretical computer science, using roughly $2,000 worth of tokens at GPT-5.6 Sol — 3 days
- @reach_vb: gm chat, the jet lag is upon us — 3 days
- @reach_vb: man you can't make this shit up, I get in my cab and my driver is telling me about Astra, the model OpenAI is about to release solved 10 math problems !! — 3 days
- @abacaj: More apps should be shipping their own models. You should be offloading as much compute as you can to the end user. Pay once, or pay again for updated versions. Small, easily accessible models are goi — 3 days
- Comfy-Org/MiniMax-H3 (2 downloads) — 3 days
- usestrix/strix — Open-source AI penetration testing tool to find and fix your app’s vulnerabilities. — 3 days
- browser-use/video-use — Edit videos with coding agents — 3 days
- Lordog/dive-into-llms — 《动手学大模型Dive into LLMs》系列编程实践教程 — 3 days
- uber/ADR — ADR secures enterprise AI agents through observability, security benchmarking, and threat detection. Deployed at Uber. — 3 days
- FalkorDB/FalkorDB — A super fast Graph Database uses GraphBLAS under the hood for its sparse adjacency matrix graph representation. Our goal is to provide the best Knowledge Graph for LLM (GraphRAG). — 3 days
- H+ Embedding: Harmonizing Global and Token-Level Retrieval with Context-Dependent Phrases — 3 days
- DASH: Decoupled Adaptive Surrogate - Acquisition Harness for Automated Bayesian Optimization — 3 days
- Measurement Without Validity: The Compounding Reliability Problem in Agentic AI Evaluation — 3 days
- FactorJEPA: Factorizing Monolithic Futures into Layout-Agent-Interaction Channels for Crowded and Chaotic Global South Urban Worlds — 3 days
- LongCat Sparse Attention: Taming the Lightning via Streaming-aware Hierarchical Cross-Layer Indexing — 3 days
- MemSIF: From Structured Interactions to Dual-Track Fact Memory for LLM Agents — 3 days
- HALT: Verification-Aware Stopping for Retrieval-Augmented Search Agents — 3 days
- Mamba with Hierarchical Memory: Solving Representation Bottleneck in Long Sequence Modeling — 3 days
- Long-term Measurements: Towards a Longitudinal Understanding of Human-AI Interactions — 3 days
- Amplitude-Only FFN Intervention for Tool-Structured LLM Inference Method: Gated Evaluation Protocol, and Cross-Model Empirical Results — 3 days
- Role Steering of Language Models for Social Simulations — 3 days
- LLM-OSDA: An Optimal-Stopping Dynamic Auction for Native Advertising in Multi-Turn LLM Conversations — 3 days
- It's the Decoding Format, Not the Perturbation: Auditing Consistency-Based Selection for Vision-Language Test-Time Scaling — 3 days
- ACE-GraphRAG: Agentic Context Engineering for Hierarchical GraphRAG — 3 days
- Statistical Mechanics of Learning on Product Wasserstein Manifolds — 3 days
- Rapid Embodiment Adaptation for Quadrupedal Locomotion — 3 days
- Style Wins, Substance Loses: A Diagnosis of LLM-as-Judge in Idea Generation — 3 days
- Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills — 3 days
- CompanionBench: A Theory-Anchored, Real-World-Grounded Benchmark for AI Emotional Companionship — 3 days
- Self-Improving Large Language Models via Progressive Experience Evolution — 3 days
- PhyCheck: Fine-Grained Evidence-Grounded Dataset for Physical Law Understanding in Video-LLMs — 3 days
- Lossless Tensor Compression as Program Synthesis — 3 days
- GROVE: Growing and Reasoning over Temporally Stratified Memory from Streaming Video Experience — 3 days
- Revisiting Generalization Across Difficulty Levels: It's Not So Easy — 3 days
- A Constitution-Grid Instrument for Data-Efficient RL Alignment (C-Guard) — 3 days
- SERL-SQL: Selective Hindsight Distillation for Text-to-SQL Reinforcement Agentic Learning — 3 days
- Pruned BPE: Post-training Visibility Pruning and Token Reallocation for Byte Pair Encoding — 3 days
- Opt.Gear Technical Report — 3 days
- CRISP: Critical Step Perception for Training Efficient Deep Search Agents — 3 days
- RoMeRL: Balancing Feedback Coverage and the Memory-Reward Trap in Self-Evolving Agent Memory via Reduced-Order Utility States — 3 days
- Failing to See or Failing to Know? Attributing Errors in Vision-Language Models — 3 days
- HiResNets: Native Full-HD Video Recognition with Foveal Residual Streams — 3 days
- @sama: i would rather be an optimist and work hard than a pessimist posting about why things won't work. it's much more difficult and the most likely path is failure, but society fails if people don't try. n — 3 days
- @reach_vb: Most of the Developer Experience team at OpenAI is getting together this week to plan, build, and think about what we can do better for developers. What should we focus on? What would you love to see — 3 days
- @reach_vb: incredible to put actual faces to the slack photos — 3 days
- karpathy/micrograd — A tiny scalar-valued autograd engine and a neural net library on top of it with PyTorch-like API — 3 days
- cloudflare/computer — Give your agent a computer 👾 — 3 days
- Activation-Guided Neuron Intervention to Induce Alzheimer's-Related Computational Language Phenotypes in a Large Language Model — 3 days
- HomoEnsNER: Does Language Alignment Outperform Architectural Complexity in Gujarati Named Entity Recognition? — 3 days
- SFT Conflicts, RL Coexists: A Theoretical and Empirical Analysis of Multi-Task Learning for LLMs — 3 days
- AI Security Leaderboard: Methodology, Results and Minimal Standard — 3 days
- Resume Means Resume: A Machine-Checked Conformance Contract for Checkpoint, Interrupt, and Resume Semantics in Workflow Persistence Layers — 3 days
- @abacaj: Image — 3 days
- @Alibaba_Qwen: Let's create with Qwen-Image-3.0-Pro on fal! 🎨 — 3 days
- @Alibaba_Qwen: Try npm i -g cline on Cline!👀 — 3 days
- yvgude/lean-ctx — Control what your AI can see. LeanCTX (Lean Context) is the context intelligence layer for AI agents — one local Rust binary that decides what they read, remembers what they learn, guards what they touch, and proves what they save. 60–90% fewer tokens as the receipt. 76 MCP tools, 30+ agents, local-first. — 3 days
- ethanfel/Qwen3-VL-32B-Ultra-Heretic-H3-ComfyUI-INT8-ConvRot (0 downloads) — 3 days