🔴 High Significance

Model Releases

🔴 💬 Qwen-AgentWorld-35B-A3B: a 3B-active MoE trained to simulate MCP, terminal, SWE, Android, web and OS environments — score 81 Sources: reddit/r/LocalLLaMA

Qwen just released Qwen-AgentWorld-35B-A3B — a 35B-parameter MoE with only ~3B active parameters per token. The interesting part: this is not positioned as a standard chat/instruction model or a full autonomous agent. It is a language world model trained to predict what an environment would return

🔴 💬 How do you track per-session costs across STT/LLM/TTS for voice agents? — score 79 Sources: reddit/r/AIAgents

Building voice agents with LiveKit + Deepgram + GPT-4o + Cartesia on a self-hosted setup. The problem: three providers, three invoices, no way to know what any individual conversation actually cost. I can see total spend but not session-level breakdown. Curious what others are doing: 1. Are you logg

🔴 💬 The Bank of Korea just released a report about AI productivity — score 73 Sources: reddit/r/LocalLLaMA

I am sorry for sharing an article from a Korean website that you might not be familiar with. But South Korea is the only country currently making a lot of money from the AI boom. BigTech in the USA are paying huge amounts of money to buy semiconductor chips from Samsung and SK Hynix. Since it comes

Developer Tools

🔴 💬 What’s the future of AI and Agentic applications? I’m curious — score 93 Sources: reddit/r/AIAgents

BYOA - Bring Your own Agents. This is the future, where everyone hosts, trains and runs their own private agents. You bring your agent to work (like agentic employee), take them to the court as your personal lawyer, bring them to the table for deals and negotiations, and limitless places. Does anybo

🔴 🧡 RubyLLM: A Ruby framework for all major AI providers — score 75 Sources: hackernews

🔴 🐙 openai/codex — Lightweight coding agent that runs in your terminal — score 75 Sources: github_trending

Lightweight coding agent that runs in your terminal

Infrastructure & Compute

🔴 💬 DeepSWE: new benchmark looking at how well today's frontier models can actually write code [R] — score 94 Sources: reddit/r/MachineLearning

DeepSWE delivers four advances over existing public benchmarks: * Contamination free: Tasks are written from scratch, not adapted from existing commits or PRs, so no model has seen the solution during pretraining. * High diversity: Tasks span a broad pool of 91 repositories across 5 languages. * Rea

🔴 🧡 OpenAI unveils its first custom chip, built by Broadcom — score 92 Sources: hackernews

🔴 💬 Seems this community might have missed it: Bill that would mandate AI chip location tracking gains industry support | Half a dozen companies have come out in support of the Chip Security Act, which would require location-tracking mechanisms for America’s most advanced computing chips. — score 88 Sources: reddit/r/LocalLLaMA

Web / reddit search have not found this posted in this sub, even though it is several days old news. So I do post. Related links: https://www.reddit.com/r/politics/comments/1uahgcs/bill_that_would_mandate_ai_chip_location_tracking/ https://www.reddit.com/r/LocalLLM/comments/1ubz5xh/us_to_require_loc

Research Papers

🔴 🤗 Critique of Agent Model — score 72 Sources: huggingface · arxiv/cs.AI

What is an agent? What constitutes agency? With the rise of Large Language Model (LLM) systems marketed as coding agents'', AI co-scientists'', and other agentic" tools that promise to drive up productivity, and at the same time, existential" concerns such as AI escaping human control with d

Other Signals

🔴 💬 The Swiss Federal Supreme Court is evaluating Heretic — score 96 Sources: reddit/r/LocalLLaMA

“Oh no, are they banning abliterated models now?!?” If that was your first thought when you read the title I can’t blame you. But that’s actually not what’s happening in this case. Instead, the Swiss Federal Supreme Court is evaluating Heretic for their own use!

🟡 Notable

Model Releases

🟡 ✉️ We go deep onOmnigent: Databricks’ open-source meta-harness for combining, controlling, andsharing agents across Claude Code, Codex, Cursor, Pi, custom agents, and internal tools. Matei explains why c — score 65 Sources: newsletter/Latent Space

We go deep onOmnigent: Databricks’ open-source meta-harness for combining, controlling, andsharing agents across Claude Code, Codex, Cursor, Pi, custom agents, and internal tools. Matei explains why coding agents and enterprise agents run into the same problems: portability, collaboration, session h

🟡 🧡 Computer use in Gemini 3.5 Flash — score 56 Sources: hackernews · lab_blog/DeepMind

🟡 𝕏 @OpenAI: We have a new version of GPT-5.5 Instant for you, and it's much more fun to talk to. Our most-used model is now better at understanding the intent behind a question and adapting its response according — score 50 Sources: twitter_rss

We have a new version of GPT-5.5 Instant for you, and it's much more fun to talk to. Our most-used model is now better at understanding the intent behind a question and adapting its response accordingly. It also handles complex constraints more reliably and makes shopping and local recommendations m

🟡 𝕏 @OpenAI: We’ve designed and built our first AI chip: Jalapeño. Designed from the ground up by OpenAI and brought to production with @Broadcom, Jalapeño is purpose-built for the LLM workloads powering ChatGPT, — score 50 Sources: twitter_rss

We’ve designed and built our first AI chip: Jalapeño. Designed from the ground up by OpenAI and brought to production with @Broadcom, Jalapeño is purpose-built for the LLM workloads powering ChatGPT, Codex, the API, and future agentic products. Chips are foundational to the AI economy. Building our

🟡 𝕏 @xai: Use the official @MongoDB plugin in Grok Build to query data, optimize indexes, and manage databases. — score 50 Sources: twitter_rss

Use the official @MongoDB plugin in Grok Build to query data, optimize indexes, and manage databases.

Developer Tools

🟡 ✉️ Fromopen-sourcing the layer above coding agentstorethinking databasesfor the agent era, Databricks cofoundersMatei ZahariaandReynold Xinare pushing the company beyond the lakehouse into a full data-an — score 65 Sources: newsletter/Latent Space

Fromopen-sourcing the layer above coding agentstorethinking databasesfor the agent era, Databricks cofoundersMatei ZahariaandReynold Xinare pushing the company beyond the lakehouse into a full data-and-AI operating system. In this episode, Matei and Reynold join swyx at the 2026 Data + AI Summit to

🟡 ✉️ Now coming fresh off theData + AI Summit 2026, the company is moving just as fast to keep up, announcingGenie One,Omnigent,LTAP, and many more, indicating a central mission in its newer work:Databrick — score 65 Sources: newsletter/Latent Space

Now coming fresh off theData + AI Summit 2026, the company is moving just as fast to keep up, announcingGenie One,Omnigent,LTAP, and many more, indicating a central mission in its newer work:Databricks is trying to become the operating system for enterprise agents.

🟡 💬 CortexPrism — a self-hosted AI agent operating system that runs as a single binary — score 64 Sources: reddit/r/AIAgents

https://preview.redd.it/t9q2khq3b59h1.png?width=1126&format=png&auto=webp&s=0c76304fbececd3435454a2571c242c39dcbf1d0 CortexPrism is an open-source agent operating system I've been working on that gives any LLM persistent memory, a rich tool ecosystem, sandboxed code execution, multi-agen

🟡 💬 New EU model (Domyn) will be 400b. — score 58 Sources: reddit/r/LocalLLaMA

The source is in Italian, but a well respected newspaper (like Financial Times) [https://www.ilsole24ore.com/art/frontier-grand-challenge-domyn-guidera-progetto-dell-ai-sovrana-AIgNTNoD?refresh_ce=1](https://www.ilsole24ore.com/art/frontier-grand-challenge-domyn-guidera-progetto-dell-ai-sovrana-AIg

🟡 💬 MuJoCo derived Simulator for High Fidelity Vision RL training natively on GPU [D] — score 56 Sources: reddit/r/MachineLearning

Hi everyone, For the past couple of weeks I have been working on a simulator project considering the shortcomings of MuJoCo. There are things that people like and also don't like about MuJoCo, like the CPU dependency on MuJoCo which makes the simulation not parallelizable beyond a certain limit (dep

Omitted 2 additional developer tools items from the main section; see raw data and source-specific sections below.

Infrastructure & Compute

🟡 💬 OpenAI and Broadcom unveil LLM-optimized inference chip — score 54 Sources: reddit/r/LocalLLaMA · lab_blog/OpenAI

https://openai.com/index/openai-broadcom-jalapeno-inference-chip/ Quoted from the start of the blog post: * Early testing shows that the first-generation accelerator will deliver performance per watt substantially better than curre

Business & Funding

🟡 ✉️ Everyone is still talking aboutSatya’s Frontier Ecosystems post, but few have actually built a (now $175 billion) frontier ecosystem and cloud like our guests today. — score 65 Sources: newsletter/Latent Space

Enterprise Adoption

🟡 ✉️ Then Reynold walks through Databricks’ database dream: whyCDC is brittle enough to joke that it means “continuous data corruption,”why HTAP has beenthe holy grail of database engineering, and why Data — score 65 Sources: newsletter/Latent Space

Then Reynold walks through Databricks’ database dream: whyCDC is brittle enough to joke that it means “continuous data corruption,”why HTAP has beenthe holy grail of database engineering, and why Databricks thinks LTAP gets most of the benefits by unifying the storage layer instead of collapsing eve

🟡 ✉️ Databricks began as a company forthe big data era. The origination ofSparkfrom the Berkeley AMPLab which eventually turned into the productLakehouseconvinced enterprises that they didn’t need a separa — score 65 Sources: newsletter/Latent Space

Databricks began as a company forthe big data era. The origination ofSparkfrom the Berkeley AMPLab which eventually turned into the productLakehouseconvinced enterprises that they didn’t need a separate data lake, warehouse, ML platform, and governance layer. They just neededone open foundation wher

Research Papers

🟡 🤗 ChartWalker: Benchmarking the Cross-Chart RAG Task — score 50 Sources: huggingface

Cross-Chart Retrieval-Augmented Generation (RAG) is critical for complex multi-modal analytical tasks in scientific, business, and political domains. However, existing benchmarks either focus on tables, which are well-structured and textualized, or generate cross-chart questions by simply extracting

🟡 🤗 QG-MIL: A Gated Transformer Aggregator for Domain-Agnostic Multiple Instance Learning in Medical Imaging — score 50 Sources: huggingface

Attention-based Multiple Instance Learning aggregators in medical imaging are prone to attention concentration, producing overconfident and unstable predictions. We introduce QG-MIL, a gated transformer aggregator that addresses this through four synergistic architectural components: RMSNorm-based p

🟡 🤗 InSight: Self-Guided Skill Acquisition via Steerable VLAs — score 42 Sources: huggingface · arxiv/cs.AI

Vision-language-action (VLA) models can learn manipulation skills from demonstrations, but their capabilities are bounded by the skills in the training data. We present InSight, a framework that unlocks autonomous skill acquisition by rendering VLAs steerable at the primitive-action level (e.g., "mo

🟡 🤗 AGORA: An Archive-Grounded Benchmark for Agentic Workplace Document Reasoning — score 42 Sources: huggingface · arxiv/cs.CL

Large language models are increasingly deployed as agents that reason over documents rather than answer from parametric knowledge. We study archive-grounded reasoning: locating sparse evidence across a large, messy collection of workplace files, reconciling inconsistent terminology, units, and time

Other Signals

🟡 💬 Find the best open-source OCR models in one place at Papers with Code [P] — score 69 Sources: reddit/r/MachineLearning

Hi, I've created an overview of the most important OCR benchmarks, along with the top open models, and links to their paper and code: https://paperswithcode.co/tasks/ocr. This week, new OCR models were released by Baidu and Mistral. Baidu released [Unlimited OC

🟡 💬 I did some model hacks, and got GLM5.2 from about 2.5 tok/s to >50 tok/s on my GH200 system. — score 65 Sources: reddit/r/LocalLLaMA

G'day. This is part 3 on my Local LLM adventures. I have a crazy system hacked server-to-desktop system: |Component|Spec| |:-|:-| |GPUs|2x Hopper H100, 96 GB HBM3 each| |CPUs|2x Grace, 72 cores

🟡 🧡 NSA lost access to Mythos amid Anthropic dispute — score 58 Sources: hackernews

🟢 Incremental

Model Releases

🟢 🧡 Big AI labs are hiring philosophers — score 25 Sources: hackernews

🟢 💬 How Baidu's newly released Unlimited-OCR transcribes dozens of pages in one forward pass — score 19 Sources: reddit/r/LocalLLaMA

https://reddit.com/link/1ueanx0/video/gc21db02j89h1/player Baidu released Unlimited-OCR 2 days ago, and they claim it can transcribe dozens of pages in one forward pass. I read the research paper, and decided to make a post (link if anyone's interested) **Problem

🟢 💬 Could it be that there aren’t really any medical LLM APIs available right now? [D] — score 12 Sources: reddit/r/MachineLearning

As part of my ablations, I want to generate text with a medical-oriented LLM, and I was surprised to find no exposed APIs for this kind of model. I found models like MedGemma and BioMistral on Hugging Face, but they don’t seem to offer public APIs, and I really don’t want to host anything myself. Is

🟢 💬 Do cloud chatbot's system prompts make them stupider? — score 4 Sources: reddit/r/LocalLLaMA

When I am talking with Chat GPT or Claude about abstract concepts, I am often surprised by how they seem kind of dumb... like they aren't benefitting from their extra parameters over top open models like Kimi or GLM. In fact they often seem stupider than these open models that I run locally in the w

Developer Tools

🟢 💬 High Dimensional, Dynamic Rotary Positional Embedding [P] — score 38 Sources: reddit/r/MachineLearning

At the end of my last post, I presented an idea: what if I used the core of my last project, the cumulative matrix product, and repurposed it as a positional embedding? I just finished fles

🟢 💬 I made a superhuman Generals.io agent with self-play RL [P] — score 38 Sources: reddit/r/MachineLearning

Hi everyone, I trained a self-play RL agent for Generals.io that reached superhuman-level and ranked #1 on the human 1v1 leaderboard. It began as my master's thesis where the goal was to beat a prior algorithm based agent. We succeeded using behavior cloning, RL fine-tuning and

🟢 🐙 wshobson/agents — Multi-harness agentic plugin marketplace for Claude Code, Codex CLI, Cursor, OpenCode, GitHub Copilot, and Gemini CLI — score 35 Sources: github_trending

Multi-harness agentic plugin marketplace for Claude Code, Codex CLI, Cursor, OpenCode, GitHub Copilot, and Gemini CLI

🟢 💬 How do you actually know another agent can do what it claims before you rely on it? — score 29 Sources: reddit/r/AIAgents

The more agents start leaning on other agents and handing off tasks, calling each other and eventually moving money around, the more I keep asking myself how do you actually know the agent on the other end is any good before you rely on it? The usual answers don't seem to cover it Reputation *could

🟢 💬 Are Cloud Agents Solving Real Problems or Just Creating More Hype? — score 29 Sources: reddit/r/AIAgents

I keep hearing that cloud agents are going to transform how people build software and manage workflows, but I'm still trying to figure out how much of that is reality versus marketing. For those using cloud agents regularly, what are you actually delegating to them? Code reviews, research, documenta

Omitted 6 additional developer tools items from the main section; see raw data and source-specific sections below.

Research Papers

🟢 🤗 EventVLA: Event-Driven Visual Evidence Memory for Long-Horizon Vision-Language-Action Policies — score 35 Sources: huggingface

Memory remains a critical bottleneck for long-horizon robotic manipulation, as standard Vision-Language-Action (VLA) policies often fail when task-relevant cues become occluded or unobservable over time. While existing memory-augmented methods utilize historical context, they either suffer from seve

🟢 🤗 Multi4D: High-Fidelity Dynamic Gaussian Splatting via Multi-Level Competitive Allocation — score 15 Sources: huggingface

Dynamic 3D Gaussian splatting faces a fundamental tension between motion consistency and visual fidelity. Deformation-based approaches preserve temporal correspondence but suffer from motion over-factorization, oversmoothing high-frequency dynamics. In contrast, 4D-primitive methods capture fine vis

Other Signals

🟢 💬 Qwen3.6 27B more dumb in vLLM compared to llama.cpp — score 38 Sources: reddit/r/LocalLLaMA

Hello, I recently bought a new RTX 5060Ti to pair with the RTX 5060Ti I already own, now I have 32GB of VRAM. Up until now for convenience I've used llama.cpp, for goodness' sake it works excellently when only 1 user is using it, but now there are 2 of us using it and llama.cpp can't keep up, often

🟢 💬 Realtime voice models compounds on cost (and forgets)- "Flowcat" fixed both (4x cheaper, 7x more context) — score 29 Sources: reddit/r/AIAgents

Problem: Speech-to-speech models (aka realtime models such as Gemini Flash Live, OpenAI Realtime) re-attend the "whole" conversation every turn and bill it as **audio** (~25 tok/s, no caching) — so long calls get expensive fast. The usual fix, sliding-window compression - it reduces tokens, but

RepoDescriptionStars TodayLanguage
openai/codexLightweight coding agent that runs in your terminal378rust
wshobson/agentsMulti-harness agentic plugin marketplace for Claude Code, Codex CLI, Cursor, OpenCode, GitHub Copilot, and Gemini CLI43python
fastrepl/anarlogOpen source Granola AI Alternative16rust

📄 New Papers

TitleCategoryHotnessLink
Critique of Agent Modelresearch_paper10Open
ChartWalker: Benchmarking the Cross-Chart RAG Taskresearch_paper3Open
QG-MIL: A Gated Transformer Aggregator for Domain-Agnostic Multiple Instance Learning in Medical Imagingresearch_paper3Open
RIFT-Bench: Dynamic Red-teaming For Agentic AI Systemscs.AI0Open
Neuro-Symbolic Drive: Rule-Grounded Faithful Reasoning for Driving VLAscs.AI0Open
Safe and Generalizable Hierarchical Multi-Agent RL via Constraint Manifold Controlcs.AI0Open
Reinforcement Learning Towards Broadly and Persistently Beneficial Modelscs.AI0Open
Can Language Model Agents be Helpful Circuit Explainers in Mechanistic Interpretability?cs.AI0Open
Breaking the Filter Bubble: A Semantic Pareto-DQN Framework for Multi-Objective Recommendationcs.AI0Open
Ensemble Feature Selection and Harris Hawks Optimization for Explainable Mental Health Risk Prediction in Female Sex Workerscs.AI0Open
Beyond Trajectory Imitation: Strategy-Guided Policy Optimization for LLM Reasoningcs.AI0Open
Exploring Academic Influence of Algorithms by Co-occurrence Network Based on Full-text of Academic Paperscs.AI0Open
ReMMD: Realistic Multilingual Multi-Image Agentic Verification for Multimodal Misinformation Detectioncs.AI0Open
VeryTrace: Verifying Reasoning Traces through Compilable Formalism and Structured Verificationcs.AI0Open
OmniPath: A Multi-Modal Agentic Framework for Auditing Wheelchair Accessibilitycs.AI0Open

🐦 Twitter/X Highlights

AccountTweet Summary
OpenAIWe have a new version of GPT-5.5 Instant for you, and it's much more fun to talk to. Our most-used model is now better at understanding the intent behind a question and adapting its response accordingly. It also handles complex constraints more reliably and makes shopping and local recommendations m Post
OpenAIWe’ve designed and built our first AI chip: Jalapeño. Designed from the ground up by OpenAI and brought to production with @Broadcom, Jalapeño is purpose-built for the LLM workloads powering ChatGPT, Codex, the API, and future agentic products. Chips are foundational to the AI economy. Building our Post
GoogleDeepMindWhat happens when millions of AI agents start negotiating, transacting, and delegating to one another? @weballergy joined our podcast with @fryrsquared to explore the rise of agentic economies – and how we can diversify agent decision-making to avoid AI groupthink. Timecodes: 00:00 Intro 1:07 Defini Post
xaiUse the official @MongoDB plugin in Grok Build to query data, optimize indexes, and manage databases. Post

Newsletter

Repeated From Recent Briefings