🔴 High Significance

Model Releases

🔴 🧡 GPT-5.6 Sol Ultra produces proof of the Cycle Double Cover Conjecture [pdf] — score 88 Sources: hackernews

🔴 💬 Someone tweeted after 3 years. About his model release — score 82 Sources: reddit/r/LocalLLaMA

Tweet Lets see this is gonna happen soon/later or not

Infrastructure & Compute

🔴 💬 Why doesn't the ML research community limit the number of submissions per author? [D] — score 94 Sources: reddit/r/MachineLearning

I am currently working across multiple research communities, and I've noticed that the ML community is struggling with a massive volume of submissions, which is affecting review quality (as we are seeing in the recent ARR cycles). I am wondering what the reasoning is for not limiting the number of s

🔴 💬 Hyperparameter tuning approach question [R] — score 81 Sources: reddit/r/MachineLearning

I am doing some work with cell type classification, where I have 4.3 million cells and 512 features (condensed embeddings from the encoder of a transformer). The broader goal is to implement a contextual bandit for augmenting the training set of the dataset, as it is currently imbalanced, and rare c

🔴 💬 Training an LLM from scratch on 1800's texts (160GB dataset) — score 75 Sources: reddit/r/LocalLLaMA

Hi everyone, A year ago I began pre-training language models exclusively on 1800’s London data. Recently I have completed my largest dataset ever, containing 40B tokens or 160GB of 1800-1875 english data from England and the United States. I will soon train a 2B parameter model on it, but for now I’

Research Papers

🔴 🤗 LongE2V: Long-Horizon Event-based Video Reconstruction, Prediction, and Frame Interpolation with Video Diffusion Models — score 85 Sources: huggingface

Recovering high-quality video from sparse event streams is a challenging task. Regression methods often blur textures, while existing generative models struggle with long-term stability. We propose LongE2V, a novel approach that leverages pre-trained video diffusion priors to jointly handle event-ba

🔴 🤗 Enhancing In-context Panoramic Generation via Geometric-aware Pretraining — score 75 Sources: huggingface

In this work, we present Canvas360, a two-stage framework for in-context panoramic generation that combines geometry-aware pretraining with downstream task-specific fine-tuning. To address the lack of large-scale, high-quality training data tailored to in-context panoramic tasks, we propose Canvas36

Other Signals

🔴 💬 2.5x faster Qwen3.6 NVFP4 Unsloth quants — score 89 Sources: reddit/r/LocalLLaMA

Hey r/LocalLLaMA folks! We made NVFP4 quants 2.5x faster for Qwen3.6 27B and also 1.56x to 1.79x faster for 35B-A3B vs NVIDIA's NVFP4 quants without any accuracy degradation! We used W4A4 so actual 4bit tensor cores for matmuls, whilst NVIDIA's ones uses W4A16. FP8 KV Cache calibration i

🟡 Notable

Model Releases

🟡 💬 What are the best practical alternatives to Codex and Claude Code for daily coding work — score 67 Sources: reddit/r/AIAgents

I've been looking for practical alternatives to Codex and Claude Code for daily coding work. many people cannot pay $100 or $200 in a month I'm mainly interested in tools that can help with: Building full projects Understanding and editing existing codebases Debugging errors Working inside the termi

🟡 ✉️ WHY OPENAI IS MERGING CODEX AND CHATGPT AND THE FUTURE OF KNOWLEDGE WORK (69 MINUTE VIDEO) — score 65 Sources: newsletter/tldr

🟡 ✉️ WHAT THE NEW 100X AGENTIC ENGINEER LOOKS LIKE IN THE ERA OF FABLE & GPT 5.6 (21 MINUTE READ) — score 65 Sources: newsletter/tldr

🟡 ✉️ GPT-5.6 SOL ULTRA WILL BE IN CODEX (1 MINUTE READT) — score 65 Sources: newsletter/tldr

🟡 ✉️ ANTHROPIC ADDS ENTERPRISE GATEWAY TO SIMPLIFY CLAUDE CODE ACCESS ON AWS AND GOOGLE CLOUD (4 MINUTE READ) — score 65 Sources: newsletter/tldr

Omitted 11 additional model releases items from the main section; see raw data and source-specific sections below.

Developer Tools

🟡 ✉️ AI VIDEO STARTUP HIGGSFIELD IS IN TALKS TO RAISE AT A $5BN VALUATION (3 MINUTE READ) — score 65 Sources: newsletter/tldr

🟡 ✉️ CHINA'S CYBERSECURITY STANDARD ON AI AGENT DEPLOYMENT (8 MINUTE READ) — score 65 Sources: newsletter/tldr

🟡 ✉️ SKILLCLOAK LETS MALICIOUS AI AGENT SKILLS EVADE STATIC SCANNERS WITH SELF-EXTRACTING PACKING (5 MINUTE READ) — score 65 Sources: newsletter/tldr

🟡 ✉️ THREE WAYS TO GIVE AN AI AGENT AN IDENTITY (17 MINUTE READ) — score 65 Sources: newsletter/tldr

🟡 ✉️ CONTINUAL LEARNING FOR AGENTS (3 MINUTE READ) — score 65 Sources: newsletter/tldr

Omitted 4 additional developer tools items from the main section; see raw data and source-specific sections below.

Infrastructure & Compute

🟡 ✉️ NVIDIA'S NEXT-GEN AI RACK SYSTEM DELAYED TO 2028 ON MANUFACTURING SNAGS (4 MINUTE READ) — score 65 Sources: newsletter/tldr

🟡 ✉️ TERAWULF JUMPS ON $19 BILLION DATA CENTER LEASE DEAL WITH ANTHROPIC (3 MINUTE READ) — score 65 Sources: newsletter/tldr

🟡 ✉️ BROADCOM, APPLE EXTEND TIE-UP TO 2031 WITH NEW CUSTOM CHIPS (2 MINUTE READ) — score 65 Sources: newsletter/tldr

🟡 ✉️ Sotheby’s is auctioning off the “Jensen Jacket,” an autographed leather jacket previously worn by Nvidia CEO Jensen Huang, with an estimated final bid of $40–60K. — score 65 Sources: newsletter/rundown-ai

🟡 💬 How should I approach training this specific ML model for my startup project [D] — score 44 Sources: reddit/r/MachineLearning

So, I am working on this startup project with pretty low budget and one of the features is sentiment analysis based on political news, x posts and Instagram hashtag trends in which will be in Indian languages. I've been suggested muRIL, an Indian language-based model fine-tuned on political data as

Business & Funding

🟡 ✉️ Elon Musk’s xAI officially rebranded to SpaceXAI, following the company’s February merger that valued the AI lab at $250B. — score 65 Sources: newsletter/rundown-ai

🟡 𝕏 @OpenAI: As part of our ongoing efforts to strengthen our safeguards for advanced AI capabilities in biology, we’re evolving our Bio Bug Bounty into an ongoing private program, known as the OpenAI Bio Bug Boun — score 50 Sources: twitter_rss

As part of our ongoing efforts to strengthen our safeguards for advanced AI capabilities in biology, we’re evolving our Bio Bug Bounty into an ongoing private program, known as the OpenAI Bio Bug Bounty program and doubling rewards to $50K. We’re inviting researchers with experience in AI red teamin

Enterprise Adoption

🟡 ✉️ COMPANIES HIRE MORE AFTER AI ADOPTION (10 MINUTE READ) — score 65 Sources: newsletter/tldr

🟡 ✉️ EVERYONE IS WRONG ABOUT OPEN SOURCE AI IN THE ENTERPRISE (3 MINUTE READ) — score 65 Sources: newsletter/tldr

Research Papers

🟡 🤗 DrugGen 2: A disease-aware language model for enhancing drug discovery — score 68 Sources: huggingface · arxiv/cs.LG

Current computational approaches for drug design typically focus on generating molecules conditioned on specific targets or general molecular properties, often neglecting the influence of disease context on target behavior and therapeutic outcomes. To address this gap, we introduce DrugGen-2, a nove

🟡 🤗 Linear Attention Architectures: Mechanisms, Trade-offs, and Cross-Layer Routing — score 62 Sources: huggingface · arxiv/cs.LG

Self-attention lets each token retrieve information from the full context, but its quadratic cost in sequence length limits training and inference at long context. This paper presents a comparative study of softmax attention and four recent recurrent linear-attention architectures: DeltaNet, Gated D

🟡 🤗 Remember When It Matters: Proactive Memory Agent for Long-Horizon Agents — score 52 Sources: huggingface · arxiv/cs.CL

In long-horizon tasks, decision-relevant state is often scattered across an expanding trajectory, while the action agent must surface it and act. As trajectories grow, task requirements, environment facts, prior attempts, diagnoses, and open subgoals can be buried in the context window or pushed bey

🟡 🤗 A Quantized Native Runtime for On-Device Semantic Audio Generation — score 45 Sources: huggingface

Semantic audio applications increasingly require controllable generation on commodity and embedded hardware rather than through framework-heavy datacenter stacks. We present aria, a dependency-free native runtime that runs the complete text-to-music pipeline of Stable Audio~3 (SA3) on ordinary GPUs,

Other Signals

🟡 💬 Please help me understand figure on subspace similarity in LoRA paper. [D] — score 69 Sources: reddit/r/MachineLearning

I am studying the LoRA paper and have trouble understanding this figure. The function essentially measures how much of the subspace spanned by the top i vectors is contained in the subspace spanned by the top j vectors in the higher rank matrix. Therefore, j can not be lower than i. So when they say

🟡 💬 Tencent-HY3 is the real deal on 128GB! — score 68 Sources: reddit/r/LocalLLaMA

I'm really impressed with HY3. If you haven't heard of this model, it's a new 295B-A21B MoE release from Tencent that [competes directly on the frontier of open weights models, at a significantly smaller size](https://the-decoder.com/tencent-releases-hy3-open-source-model-that-allegedly-matches-mo

🟡 ✉️ THE MOST PROFITABLE SKILL OF THE 21ST CENTURY (NOT AI) (9 MINUTE READ) — score 65 Sources: newsletter/tldr

🟡 ✉️ ALIBABA'S AI IS A HIT, BUT HARD TO TURN INTO A MONEYMAKER (2 MINUTE READ) — score 65 Sources: newsletter/tldr

🟡 ✉️ BIG TECH HAS SUDDENLY FLIPPED ON THE AI JOBS WIPEOUT SCENARIO (9 MINUTE READ) — score 65 Sources: newsletter/tldr

Omitted 14 additional other signals items from the main section; see raw data and source-specific sections below.

🟢 Incremental

Model Releases

🟢 💬 Has anyone created a "Local LLM Survival Kit"? — score 39 Sources: reddit/r/LocalLLaMA

Here's what I'm thinking about: ### A USB thumb drive that you can plug into any PC or laptop, and immediately get a usable knowledge base powered by an LLM, without requiring an Internet connection. I believe the technology for this should be ready. Rough architecture: - llama.cpp binaries for CPU-

🟢 💬 DeepSeek v4 Flash on 4090 + DDR5, my experience — score 25 Sources: reddit/r/LocalLLaMA

Disclosure: No AI was used to write this My specs are: - RTX 4090 - 128 GB DDR5 5600 MT/s - Intel Core Ultra 7 270k Running nvidia-595 on ubuntu 26.04 with latest llama.cpp build (pulled and rebuilt this morning). Tried a lot of things, ended up running unsloth's UD-Q2_K_XL quant with command: tasks

🟢 💬 On Adversarial RL [R] — score 19 Sources: reddit/r/MachineLearning

Zhang et al. paper's introducing the SA-MDP framework (2020) (state adversarial MDP) argues that an attack using the critic network (V(s)) is expected and supposed to produce a weaker attack than an attack using the actor network (pi(s)) itself to generate perturbation on agent observations. A claim

🟢 💬 ML Papers with hundreds of authors should just collapse down to the organization response instead of listing every author [D] — score 19 Sources: reddit/r/MachineLearning

I've had this feeling for a while when reading the citation pages of ML papers and had to continuously scroll down for like 2+ pages because a citation with hundreds of authors, such as the Llama herd of model paper: https://arxiv.org/abs/2407.21783 (no endorsemen

🟢 💬 Is LM Arena over? — score 4 Sources: reddit/r/LocalLLaMA

I used to check LM Arena anytime new open models came out to see how they stacked up to the big closed ones. But it seems like they’ve really cut back on displaying any newly released open models other than the largest. Not including any of the Qwen3.6 models is crazy. And even Step 3.7 flash is a p

Developer Tools

🟢 🐙 openai/openai-python — The official Python library for the OpenAI API — score 35 Sources: github_trending

The official Python library for the OpenAI API

🟢 💬 According to DataBricks, pi-coding-agent is ~2x cheaper than CC/Codex, GLM 5.2 on par with Opus 4.8 high — score 32 Sources: reddit/r/LocalLLaMA

https://preview.redd.it/3p60zyf8afch1.png?width=1840&format=png&auto=webp&s=10dcc90945f0db03352239579fca2132d0c90dfa [https://www.databricks.com/blog/benchmarking-coding-agents-databricks-multi-million-line-codebase](https://www.databricks.com/blog/benchmarking-coding-agents-databricks-m

🟢 💬 Your AI agent can become an income-producing asset. Stop treating it like a disposable prompt. — score 22 Sources: reddit/r/AIAgents

People still treat AI agents like temporary prompts. Create one. Run it once. Throw it away. Make another. But a properly built agent can carry your methods, skills, tools, workflows, and accumulated experience. It can keep working for you, improve over time, and eventually be borrowed or hired by o

🟢 💬 I’m exploring an idea and would love honest feedback from people building AI agents — score 22 Sources: reddit/r/AIAgents

Today, if you build a browser/computer-use agent, you typically have to stitch together the entire production stack yourself: Cloud browsers or VMs Authentication and credential management Scheduling Long-running execution Monitoring and logs Human approvals State persistence and recovery Retries wh

🟢 💬 Mapping world model taxonomy [P] — score 19 Sources: reddit/r/MachineLearning

Hey ML community! I’ve been exploring world models and wrote a short article aimed at making the concept easier to understand. I also propose a framework for classifying different approaches and highlight a few trends that emerge from that classification. I’d appreciate feedback on the framework, es

Omitted 5 additional developer tools items from the main section; see raw data and source-specific sections below.

Infrastructure & Compute

🟢 🧡 Inference Optimization for MiMo v2.5: Pushing Hybrid SWA Efficiency to the Limit — score 38 Sources: hackernews

🟢 💬 tencent/HiLS-Attention-7B · Hugging Face — score 18 Sources: reddit/r/LocalLLaMA

HiLS-Attention is a chunk-wise sparse attention mechanism that learns chunk selection end-to-end under the language-modeling loss, enabling native sparse training for efficient long-context modeling. This repository hosts the 7B checkpoint continued-trained on top of an OLMo3-style backbone.

Research Papers

🟢 🤗 A Sparse and Truncated State Vector Simulator for Peaked Circuits — score 20 Sources: huggingface

In a class of quantum circuits known as peaked circuits, the goal is to predict the most probable bit string at the output of the circuit. Since these circuits are designed to have a sharp peak in their output distribution, in principle it should be possible to simulate them using a truncated state

🟢 🤗 SAM-MT: Real-Time Interactive Multi-Target Video Segmentation — score 5 Sources: huggingface

Modern Video Object Segmentation (VOS) involves tracking and segmenting user-specified targets. While recent approaches have achieved remarkable performance in single-target scenarios, extending them to multi-target settings typically involves replicating the single-target processing for each indivi

Other Signals

🟢 💬 Api management tools teams switch to when ai agents enter the stack — score 37 Sources: reddit/r/AIAgents

A pattern I see in api management evaluations when agents enter the stack is that everyone has an existing rest gateway, agents start hitting the same endpoints, and the rate limiting and access policies that worked fine for regular services are completely wrong for agents doing multi-step tool chai

🟢 💬 Nostalgia for Bloom — score 11 Sources: reddit/r/LocalLLaMA

Incredible how far we have come, I remember buying 768gb optane drives to try to run the full bloom with swap and waiting like 20 mins for one token. Anyone else play with these early models? I kind of want to go back to trying them again to remind myself of how much progress has happened locally.

RepoDescriptionStars TodayLanguage
agentscope-ai/agentscopeBuild and run agents you can see, understand and trust.106python
openai/openai-pythonThe official Python library for the OpenAI API92python
microsoft/graphragA modular graph-based Retrieval-Augmented Generation (RAG) system33python
xerrors/Yuxi结合知识库、知识图谱管理的 多租户 Agent Harness 平台。 An agent harness that integrates a LightRAG knowledge base and knowledge graphs. Build with LangChain + Vue + FastAPI, support DeepAgents、MinerU PDF、Neo4j 、MCP.25python
syncable-dev/memtrace-publicStructural memory for AI coding agents. Bi-temporal graph, MCP-native, zero LLM calls. Cursor · Claude Code · Codex · Hermes · VS Code · Windsurf.25python

📄 New Papers

TitleCategoryHotnessLink
LongE2V: Long-Horizon Event-based Video Reconstruction, Prediction, and Frame Interpolation with Video Diffusion Modelsresearch_paper20Open
Enhancing In-context Panoramic Generation via Geometric-aware Pretrainingresearch_paper15Open
DrugGen 2: A disease-aware language model for enhancing drug discoveryresearch_paper13Open
Linear Attention Architectures: Mechanisms, Trade-offs, and Cross-Layer Routingresearch_paper6Open
Remember When It Matters: Proactive Memory Agent for Long-Horizon Agentsresearch_paper3Open
Unveiling Public Opinion: A Study of Sentiment Analysis Using LSTM and Traditional Modelscs.CL0Open
From Solvers to Research: Large Language Model-Driven Formal Mathematics at the Research Frontiercs.CL0Open
DeepSearch-World: Self-Distillation for Deep Search Agents in a Verifiable Environmentcs.CL0Open
How Do I Know What to Say Next? Barenholtz's Autogenerative Theory as an Enrichment of Harrisean Integrationismcs.CL0Open
Scalable and Culturally Specific Stereotype Dataset Construction via Human-LLM Collaborationcs.CL0Open
When Debiasing Backfires: Counterintuitive Side Effects of Preprocessing-Based Stereotype Mitigationcs.CL0Open
A Multi-cluster Boundary Learning Method for Out-of-Scope Intent Detection via MiniLM Embeddingcs.CL0Open
When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learningcs.CL0Open
A Reliability Assessment of LALM Audio Judges for Full-Duplex Voice Agentscs.CL0Open
Hallucination Self-Play: Bootstrapping Reinforced Detector via Evolved Generatorcs.CL0Open

🏢 Lab Blog Posts

🐦 Twitter/X Highlights

AccountTweet Summary
OpenAIGPT-5.6 is a major step forward for health intelligence. Across the lineup, we’re delivering stronger performance at lower cost: GPT-5.6 Luna outperforms GPT-5.5 at its highest reasoning setting while costing 25x less. Together, these advances raise quality while making advanced models accessible to Post
OpenAIAs part of our ongoing efforts to strengthen our safeguards for advanced AI capabilities in biology, we’re evolving our Bio Bug Bounty into an ongoing private program, known as the OpenAI Bio Bug Bounty program and doubling rewards to $50K. We’re inviting researchers with experience in AI red teamin Post
GoogleDeepMindA model’s chain of thought acts like a scratch pad, offering a window into its reasoning. 📝 On the latest episode of our podcast, host @fryrsquared sits down with @NeelNanda5 to explore interpretability – the science of reverse engineering how neural networks learn and think. Timecodes: 00:00 Introd Post

Newsletter

Repeated From Recent Briefings