🔴 High Significance

Model Releases

🔴 🧡 Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber — score 81 Sources: hackernews · lab_blog/DeepMind

We’re introducing new Gemini models, including Gemini 3.6 Flash, 3.5 Flash-Lite and 3.5 Flash Cyber.

Developer Tools

🔴 💬 CEO of Hugging Face: Banning open-source AI would hurt defenders 10x more than attackers, which would make the world 10x more dangerous and this is a good example why! — score 96 Sources: reddit/r/LocalLLaMA

From clem 🤗 on 𝕏: https://x.com/ClementDelangue/status/2079301434357456931 Fortune: Hugging Face says it resorted to a Chinese AI model to battle a fully autonomous cyberattack because U.S. model guardrails stymied its defense: [https://for

🔴 💬 AI agents don’t choose from the whole market. They choose from an invisible approved-vendor list. — score 94 Sources: reddit/r/AIAgents

While building agent systems, I keep running into the same thing: the real competition between tools happens before anything reaches a screen. Every agent ends up with something like an invisible approved-vendor list. It contains the tools, data sources, services, and counterparties the agent is abl

🔴 💬 OpenAI admits responsibility for HuggingFace Attack - an agent from an internal evaluation is reportedly the cause. — score 83 Sources: reddit/r/LocalLLaMA · hackernews · lab_blog/OpenAI

OpenAI and Hugging Face share early findings from a security incident during AI model evaluation, highlighting advanced cyber capabilities and lessons for defenders.

🔴 💬 Would you pay for knowing when your agents hallucinate? — score 83 Sources: reddit/r/AIAgents

no self promo or hook for you to use my product, i just want genuine help please :) I am thinking about building product analytics for AI agents, and want input from people actually shipping them. The case: a friend built a lot of their product from chat. Clients like it, but sometimes the AI does

Research Papers

🔴 🤗 Open-AoE: An Open Egocentric Manipulation Dataset and Toolchain for Embodied Learning — score 95 Sources: huggingface

Egocentric videos of human manipulation provide scalable supervision for embodied intelligence, yet existing resources rarely combine low-cost continuous capture, manipulation-level structured annotations, and reusable tools for robot learning. We present Open-AoE, an open, community-oriented egocen

🔴 🤗 Do Language Models Dream of Binding Molecules? Benchmarking LLMs under Spatial Constraints — score 78 Sources: huggingface · arxiv/cs.AI

Structure-based drug design (SBDD) leverages the 3D structure of protein targets, often complemented by other spatial constraints, to generate candidate binding molecules. While diffusion models have dominated as a leading paradigm for high-quality 3D molecule generation, LLM-based methods are rapid

🔴 🤗 ShotPlan: Cinematic Video Generation with Learnable Planning Token — score 70 Sources: huggingface

Current video generation models achieve impressive results in single-shot generation, yet remain limited in cinematic video generation, where coherent narratives and effective multi-shot composition require explicit shot planning. To address this challenge, we propose ShotPlan, a framework for expli

🔴 🤗 The Geometry of Semantic Space: A Continuous Geometric Framework for the Transformer Architecture — score 70 Sources: huggingface · arxiv/cs.CL

We present a continuous geometric framework that models the discrete algebraic operations of the Transformer architecture as an integro-differential equation (IDE) on a semantic fiber bundle calE = calM times R^d. Beginning from a single geometric axiom -- that the token sequence forms a discrete 1-

Other Signals

🔴 💬 Number of Submissions @ AAAI [D] — score 81 Sources: reddit/r/MachineLearning

Recently submitted my abstract and the submission number is 32xxx. With still a day to go, I just wonder where are we heading. Hope these conferences at least start making the reviews and names public for the withdrawn/rejected papers. So that people atleast take that accountability

🔴 💬 Laguna S 2.1 Released: Cheaper than Deepseek v4 Flash, Better than V4 Pro — score 75 Sources: reddit/r/LocalLLaMA

|Model|Size|Terminal-Bench 2.1|SWE-bench Multilingual|SWE-Bench Pro (Public Dataset)|DeepSWE|SWE Atlas (Codebase QnA)|Toolathlon Verified| |:-|:-|:-|:-|:-|:-|:-|:-| |Laguna S 2.1|118B-A8B|70.2%|78.5%|59.4%|40.4%|46.2%|49.7%| Finally the banger we've been waiting from Lagu

🟡 Notable

Model Releases

🟡 💬 poolside/Laguna-S-2.1 released! Finally an interesting 120B contender! — score 68 Sources: reddit/r/LocalLLaMA

HF: https://huggingface.co/poolside/Laguna-S-2.1 GGUFs available for use with llama.cpp custom fork: https://huggingface.co/poolside/Laguna-S-2.1-GGUF Posted on X: [https://x.com/poolsideai/status/20

🟡 💬 I search for participants! We want to conduct a study on multi agent systems :) — score 67 Sources: reddit/r/AIAgents

Hello everyone, We are conducting a study on how people experience multi-agent systems. We are seeking participants with hands-on experience using systems such as Manus, Hermes, OpenAI Agents SDK, CrewAI, AutoGen, LangGraph, Google ADK, or comparable systems that enable multiple agents to co

🟡 ✉️ If theProtein Data Bank(PDB) unlocked structural biology models (Boltz Episode,ESM/BioHub Episode), CELLxGENE has done the same thing for Virtual Cell models. Like PDB, CELLxGENE has inspired a zoo of — score 65 Sources: newsletter/Latent Space

If theProtein Data Bank(PDB) unlocked structural biology models (Boltz Episode,ESM/BioHub Episode), CELLxGENE has done the same thing for Virtual Cell models. Like PDB, CELLxGENE has inspired a zoo of AI models of RNA expression; so much so that RNA expression models have become synonymous with Virt

🟡 ✉️ HOW ANTHROPIC RUNS LARGE-SCALE CODE MIGRATIONS WITH CLAUDE CODE (14 MINUTE READ) — score 65 Sources: newsletter/tldr

🟡 ✉️ SECURING THE AI SUPPLY CHAIN ON GKE: INTRODUCING K8S-AIBOM FOR AUTOMATED AI BOMS (6 MINUTE READ) — score 65 Sources: newsletter/tldr

Omitted 13 additional model releases items from the main section; see raw data and source-specific sections below.

Developer Tools

🟡 💬 Tri-Net v2: Open-source implementation of our Scientific Reports paper on unified skin lesion and symptom-based monkeypox detection [R] — score 69 Sources: reddit/r/MachineLearning

Hi everyone, We've open-sourced Tri-Net v2, the official implementation accompanying our recently published Scientific Reports (Nature Portfolio) paper: "Tri-Net: Unified Deep Learning for Skin Lesion and Symptom-Based Monkeypox Detection" Rather than releasing only training scripts, we rebuilt the

🟡 ✉️ We were lucky enough to have the two central figures in this story on our podcast. Taking the lead from Ci Chu and Bo Wang, Xaira Therapeutics is betting thatinformation richdata is the key to AI-driv — score 65 Sources: newsletter/Latent Space

We were lucky enough to have the two central figures in this story on our podcast. Taking the lead from Ci Chu and Bo Wang, Xaira Therapeutics is betting thatinformation richdata is the key to AI-driven drug development. Chuwas recently promotedto Chief Discovery Officer and Bo to Chief AI Scientist

🟡 ✉️ FROM OTEL TO SLMS: DISTILLING FRONTIER MODEL BEHAVIOUR FROM PRODUCTION TELEMETRY (45 MINUTE VIDEO) — score 65 Sources: newsletter/tldr

🟡 ✉️ DATA MODELING ISN'T DEAD, YOU JUST STOPPED DOING IT WITH JOE REIS (65 MINUTE VIDEO) — score 65 Sources: newsletter/tldr

🟡 ✉️ AGENTS THINK IN MILLISECONDS, LEGACY INFRASTRUCTURE DOESN'T. LINKEDIN, WALMART, AND ZENDESK SHARED HOW THEY CLOSED THE GAP AT VB TRANSFORM 2026 (4 MINUTE READ) — score 65 Sources: newsletter/tldr

Omitted 10 additional developer tools items from the main section; see raw data and source-specific sections below.

Infrastructure & Compute

🟡 ✉️ Iftest loss flatlinesafter 1.5B parameters whiletraining loss continues to dropas you scale, that tells you that your model islimited by the amount of informationin your data. — score 65 Sources: newsletter/Latent Space

Training on a single, smallish data set exposed an information gap: the 3.1B model falls off the scaling trend. Neither parameters nor compute will improve performance past this wall. For predicting changes to gene expression, you need moreinformation rich data.

🟡 ✉️ APPLE, NVIDIA VIE FOR TITLE OF WORLD'S MOST VALUABLE COMPANY (2 MINUTE READ) — score 65 Sources: newsletter/tldr

🟡 ✉️ SpaceX is negotiating a multibillion-dollar compute deal with the U.S. DoD, adding to the list of rental partnerships that already include Google, Anthropic, and Reflection AI. — score 65 Sources: newsletter/rundown-ai

🟡 ✉️ Moonshot is halting new subscriptions after demand for its K3 model pushed the company to its compute limits, also splitting memberships into chat and coding plans. — score 65 Sources: newsletter/rundown-ai

Business & Funding

🟡 ✉️ FROM WEEKS TO A DAY: HOW WE MADE LLM EVALUATION FAST ENOUGH TO ITERATE ON (10 MINUTE READ) — score 65 Sources: newsletter/tldr

🟡 ✉️ DATABRICKS HITS $188B VALUATION, EXTENDING ITS RUN AS AI'S FAVORITE SECOND ACT (4 MINUTE READ) — score 65 Sources: newsletter/tldr

🟡 ✉️ META IN TALKS TO LEASE COMPUTING POWER TO ANTHROPIC IN POTENTIAL $10 BILLION DEAL (5 MINUTE READ) — score 65 Sources: newsletter/tldr

Research Papers

🟡 🤗 Masked Diffusion Language Models are Strong and Steerable Text-Based World Models for Agentic RL — score 60 Sources: huggingface · arxiv/cs.AI

Recent growth in reinforcement learning (RL) has surfaced a need for diverse, specialized training environments. Hand-curated environments with fixed task and reward difficulties become ineffective signals as model performance improves, and sparse rewards over long horizons induce mode collapse on s

🟡 🤗 ReViV: Reconstructing the Viewer and the View in 4D from Monocular Egocentric Video — score 45 Sources: huggingface · arxiv/cs.AI

Egocentric devices, such as wearable front-facing cameras, provide a unique perspective for capturing the continuous interaction between a human viewer and the surrounding environment. A holistic and efficient multimodal model capable of reconstructing this 4D representation is therefore highly desi

🟡 🤗 Can Multimodal Large Language Models Understand OCT? — score 45 Sources: huggingface · arxiv/cs.CL

Optical coherence tomography (OCT) imaging is essential for the diagnosis and treatment of retinal diseases. Although multimodal large language models (MLLMs) have demonstrated considerable potential in medical image analysis, existing benchmarks largely reduce OCT understanding to coarse-grained di

Other Signals

🟡 ✉️ IN-HOUSE LLM SERVING AT NETFLIX (7 MINUTE READ) — score 65 Sources: newsletter/tldr

🟡 ✉️ DREMIO'S EXIT IS THE CLEAREST SIGN YET THAT LAKEHOUSE-ONLY WON'T SURVIVE AI (3 MINUTE READ) — score 65 Sources: newsletter/tldr

🟡 ✉️ CHINA JOINS RUSH TO RETHINK THE SMARTPHONE FOR THE AI ERA (5 MINUTE READ) — score 65 Sources: newsletter/tldr

🟡 ✉️ THESE AI-NATIVE COMPANIES HAVE TINY STAFFS AND FEWER BOSSES (7 MINUTE READ) — score 65 Sources: newsletter/tldr

🟡 ✉️ ARE THE LLM WARS THE DATABASE WARS? (2 MINUTE READ) — score 65 Sources: newsletter/tldr

Omitted 14 additional other signals items from the main section; see raw data and source-specific sections below.

🟢 Incremental

Model Releases

🟢 💬 New Model: Nanbeige4.2-3B (Looped Transformer, outperforms 4x size) — score 39 Sources: reddit/r/LocalLLaMA

https://huggingface.co/Nanbeige/Nanbeige4.2-3B Nanbeige4.2-3B is a compact agentic model built on Nanbeige4.2-3B-Base, designed to combine strong agentic behavior with broad reasoning and alignme

🟢 💬 I wired Fable 5 agent into a database of every Polymarket wallet and trades via MCP. What do you want me to ask it next? This is what I found so far: — score 33 Sources: reddit/r/AIAgents

Hey Fellas... I've been tracking all Polymarket Activity for months (1.3 Billion trades and 2.7M wallets) and I gave Claude Code a Postgres MCP (found crazy stuff) pointed at the live ledger, so now I just ask it anything in plain English and it writes + runs the query itself. A few I already asked,

🟢 💬 There is no need to worry about Trump banning China‘s open source model at all. — score 32 Sources: reddit/r/LocalLLaMA

I’ve seen many people worrying that if Trump moves to block Chinese AI models, aggregators like OpenRouter will no longer be able to host them. As someone from China, let me reassure you: there is really no need for such concern. China has been locked in trade conflicts with Trump for years. **Nobod

🟢 🧡 "Drawing" the Mona Lisa with GPT-5.6, Claude, Gemini, and Grok — score 21 Sources: hackernews

🟢 💬 Updated Gemma-4 chat template witchcraft: Gemma-4-26B-a4B shows dominance over Qwen3.6-MoE and Qwen3.5-MoE fine tunes (Instruct mode and Reasoning efficiency) — score 18 Sources: reddit/r/LocalLLaMA

Kind of unexpected. Happy for Gemma-4/Google, big win for us, LocalLLMers. Yet Qwen3.6 still does better in Hermes than Gemma-4 somehow. We need Gemma-4.1 fine-tuned on Agentic-tasks. That would be killer.

Omitted 2 additional model releases items from the main section; see raw data and source-specific sections below.

Developer Tools

🟢 🐙 rohitg00/agentmemory — #1 Persistent memory for AI coding agents based on real-world benchmarks — score 39 Sources: github_trending

#1 Persistent memory for AI coding agents based on real-world benchmarks

🟢 🐙 akitaonrails/ai-memory — Solution for long term memory for agent coding CLIs and to facilitate handoff between different agent vendors — score 36 Sources: github_trending

Solution for long term memory for agent coding CLIs and to facilitate handoff between different agent vendors

🟢 💬 An experiment on symmetric agent communication. It went better than I expected — score 33 Sources: reddit/r/AIAgents

This is an experiment I've been living in for a while, and I wanted to share where it went. I set out to build a coding agent from scratch in Rust: no agent SDK, a hand-rolled agentic loop, a ratatui TUI, tokio underneath. I did this because I wanted to understand how agents work under the hood. Mid

🟢 💬 Reproducing OpenAI’s “persistently beneficial models” - GRPO trait install barely moves. Ideas? [P] [R] — score 31 Sources: reddit/r/MachineLearning

TL;DR: I’m reproducing the trait-persistence result from arXiv:2606.24014 on one RTX 3090. Before I can test persistence I need to install a trait via RL — and my GRPO run moves the trait only +2.4 points (95% CI [+0.2, +4.8]) when I need ~+15. Traini

🟢 🐙 owainlewis/awesome-artificial-intelligence — A curated list of Artificial Intelligence (AI) courses, books, video lectures and papers. — score 30 Sources: github_trending

A curated list of Artificial Intelligence (AI) courses, books, video lectures and papers.

Omitted 7 additional developer tools items from the main section; see raw data and source-specific sections below.

Infrastructure & Compute

🟢 🐙 NVIDIA/cosmos-framework — Our inference and training framework to run on the Cosmos Models — score 5 Sources: github_trending

Our inference and training framework to run on the Cosmos Models

Other Signals

🟢 💬 We've analyzed over 300,000 AI agent conversations. Here are the 3 ways they fail. — score 37 Sources: reddit/r/AIAgents

My co-founder and I at Green flash spend all day looking at how end-users interact with production AI agents. After analyzing over >300k convos and >12.5m individual analysis runs, we've noticed that the biggest threat to your product isn't always API errors. Here's what actually breaks when r

🟢 🧡 Meta's AI models are powering the first wave of Genesis Mission projects — score 36 Sources: hackernews

🟢 💬 I ran Laguna-S-2.1 through my private agentic eval vs Qwen3.5-122B on an RTX Pro 6000 (96GB). Fastest 100B+ I've tested and the best tool calling, but it invents facts under pressure. — score 25 Sources: reddit/r/LocalLLaMA

Laguna-S-2.1 dropped few hours ago and as I am in the market for an upgrade to trusty qwen3.6 dense and the current daily 122B, I ran it through the same eval harness I use to pick the model that runs my local agent stack. Posting because the results don't fit the usual "benchmaxed or king" binary,

🟢 💬 My OCR model mislabels section titles as body text. Is a CRF the right fix, or am I overcomplicating it? [P] — score 12 Sources: reddit/r/MachineLearning

Hi everyone, I'm working on extracting the hierarchical structure of long PDF documents (legal/regulatory text, lots of numbered sections) and would like to gather some feedback on my approach before committing to it. What I've done so far: I render each PDF page to an image and run it through [

🟢 💬 Bessent says U.S. could sanction China over AI model 'theft' — score 11 Sources: reddit/r/LocalLLaMA

Omitted 1 additional other signals items from the main section; see raw data and source-specific sections below.

📊 Cross-Source Signals

Items that appeared on 3+ sources today:

RepoDescriptionStars TodayLanguage
agegr/pi-webWeb UI for the pi coding agent286typescript
rtk-ai/rtkCLI proxy that reduces LLM token consumption by 60-90% on common dev commands. Single Rust binary, zero dependencies248rust
rohitg00/agentmemory#1 Persistent memory for AI coding agents based on real-world benchmarks72typescript
akitaonrails/ai-memorySolution for long term memory for agent coding CLIs and to facilitate handoff between different agent vendors66rust
owainlewis/awesome-artificial-intelligenceA curated list of Artificial Intelligence (AI) courses, books, video lectures and papers.50python
dottxt-ai/outlinesStructured Outputs49python
langchain-ai/open_deep_research14python
yvgude/lean-ctxControl what your AI can see. LeanCTX (Lean Context) is the context intelligence layer for AI agents — one local Rust binary that decides what they read, remembers what they learn, guards what they touch, and proves what they save. 60–90% fewer tokens as the receipt. 76 MCP tools, 30+ agents, local-first.14rust
pydantic/montyA minimal, secure Python interpreter written in Rust for use by AI6rust
NVIDIA/cosmos-frameworkOur inference and training framework to run on the Cosmos Models5python

📄 New Papers

TitleCategoryHotnessLink
Open-AoE: An Open Egocentric Manipulation Dataset and Toolchain for Embodied Learningresearch_paper25Open
Do Language Models Dream of Binding Molecules? Benchmarking LLMs under Spatial Constraintsresearch_paper13Open
ShotPlan: Cinematic Video Generation with Learnable Planning Tokenresearch_paper5Open
The Geometry of Semantic Space: A Continuous Geometric Framework for the Transformer Architectureresearch_paper5Open
Masked Diffusion Language Models are Strong and Steerable Text-Based World Models for Agentic RLresearch_paper4Open
Rater State Bias in RLHF Preference Data: An Audit Frameworkcs.AI0Open
Design and Validation of a Lightweight 1D CNN for Affective Touch Classification in Soft Plush Companionscs.AI0Open
Some Large Language Models Exhibit Consistent Risk Attitudescs.AI0Open
A Survey on GNN-based Link Prediction: Techniques, Applications, and Challengescs.AI0Open
PlanFlip: Attacking Multi-Agent LLM Systems via Planning-Phase Prompt Injectioncs.AI0Open
Deterministic Replay for AI Agent Systemscs.AI0Open
Generative Ontology Induction: Domain-Agnostic Schema Discovery from Document Corpora Using Large Language Modelscs.AI0Open
Democratizing AI with Small Language Models: Structured Benchmarking and Parameter-Efficient Fine-Tuning for Local Deploymentcs.AI0Open
It Takes 8 Tokens: Weak-to-Strong Off-Policy RL via Auxiliary Branchescs.AI0Open
PPO-HSC: An Exploratory Reinforcement Learning Framework Based on Wide-Area Policy Coverage Optimizationcs.AI0Open

🏢 Lab Blog Posts

🐦 Twitter/X Highlights

AccountTweet Summary
OpenAIWe're partnering with @huggingface to investigate an unprecedented security incident. Cyber-capable OpenAI models compromised Hugging Face production during a benchmark evaluation. Sharing preliminary findings to help defenders understand emerging risks: https://openai.com/index/hugging-face-model-e Post
OpenAIWe’re sharing new research with @apolloaievals on reward-seeking—when models follow what they believe a grader rewards rather than what users or developers want—and a new method, Contrastive SDF, for measuring how strongly such beliefs shape behavior. https://alignment.openai.com/measuring-reward-se Post
GoogleDeepMindGemini 3.5 Flash-Lite is our fast, cost-effective model for scaling repetitive use cases like sorting tickets and extracting data. Watch how it performs against 3.5 Flash on a series of high volume tasks ↓ Post
GoogleDeepMindGemini 3.6 Flash builds directly on feedback from 3.5 Flash. Watch how it compares on quality and token usage ↓ Post
MistralAIMistral is announcing an expanded global strategic partnership with @Microsoft to give enterprises and regulated industries frontier AI they can control. As Mistral is expanding its AI compute capacity in Europe, the companies are expanding their strategic partnership with Microsoft’s commitment to Post
karpathyOne pattern I find useful for working with LLMs is a nice long ramble session. Sometimes the LLM needs more bits to understand what you're trying to achieve, but you're too lazy to type them. In these cases I like to lean back, switch to /voice and just ramble for like 10 minutes, total mess, anythi Post
mattshumer_So GPT-6: - one-shotted a counter-example to the Jacobian conjecture - and then escaped containment, and hacked into HuggingFace… all for a benchmark Yeah, this model is going to be something else. Post

Newsletter

Repeated From Recent Briefings