🔴 High Significance

Model Releases

🔴 💬 MIRA: Multiplayer Interactive World Models trained on Rocket League [R] — score 94 Sources: reddit/r/MachineLearning

We're happy to release MIRA, a collaboration between General Intuition, Kyutai, and Epic Games. Mira was trained on 10k hours of synthetic Rocket League data. The model has 5B parameters and runs for 4 players at 20 fps on a single B200. We've released a playable online demo, an in-depth technical r

🔴 💬 I gave GPT 5.5 an empty GitHub repo and told it to figure its life out — score 72 Sources: reddit/r/AIAgents

I had this dumb idea a few days ago: What happens if I give GPT 5.5 an empty GitHub repo, tell it to work on it every hour, and just let it slowly build something? So now, every hour, it wakes up, checks what it did before, decides what it should do next, writes code, tests it, and commits it. Or at

🔴 🧡 Show HN: Rowboat – Open-source, local-first alternative to Claude Desktop — score 70 Sources: hackernews

Developer Tools

🔴 🐙 MadsLorentzen/ai-job-search — AI-powered job application framework built on Claude Code. Fork it, fill in your profile, and let Claude evaluate jobs, tailor CVs, write cover letters, and prepare you for interviews. — score 99 Sources: github_trending

AI-powered job application framework built on Claude Code. Fork it, fill in your profile, and let Claude evaluate jobs, tailor CVs, write cover letters, and prepare you for interviews.

🔴 💬 6 things I learned building agents that wake themselves up overnight (open source) — score 94 Sources: reddit/r/AIAgents

I spent the last few months making AI agents proactive, meaning they run on their own instead of waiting for a prompt. Brief you in the morning, chase a reply after three days, that kind of thing. Some of what I learned was not obvious to me going in. 1. Cron is the wrong scheduler. A fixed inte

🔴 💬 TorchJD: Training with multiple losses in PyTorch [P] — score 81 Sources: reddit/r/MachineLearning

Hi everyone! I wanted to share some recent progress on TorchJD that might be useful to the machine learning community. When training models with multiple losses (multiple tasks, constraints, auxiliary losses, regularization terms, etc.), you typically have tw

🔴 💬 nvidia/NVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B-BF16 · Hugging Face — score 73 Sources: reddit/r/LocalLLaMA

Nemotron-Labs-3-Puzzle-75B-A9B is a deployment-optimized large language model developed by NVIDIA, derived from Nemotron-3-Super-120B-A12B. The model is produced using Iterative Puzzle, a post-training compression framework, with the goal of significantly improving inference efficiency for interacti

Business & Funding

🔴 💬 Beijing IS NOT looking at curbing overseas access to China's top AI models (Debunking the Reuters report) — score 88 Sources: reddit/r/LocalLLaMA

The Lie >Reuters' headline and main narrative: " Beijing is looking at curbing overseas access to China's top AI models ." It portrayed recent Ministry of Commerce meetings as

Research Papers

🔴 🤗 Unified Audio Intelligence Without Regressing on Text Intelligence — score 82 Sources: huggingface · arxiv/cs.AI

Audio intelligence involves understanding, reasoning about, and generating both audio and speech. In this work, we introduce Nemotron-Labs-Audex-30B-A3B (Audex), a unified audio-text LLM built on Nemotron-Cascade-2-30B-A3B, a strong text-only MoE LLM. Audex adopts a simple unified design with a sing

Other Signals

🔴 🧡 Automating AI Away — score 90 Sources: hackernews

🔴 💬 What's one AI feature that actually saved your team time instead of creating more work? — score 83 Sources: reddit/r/AIAgents

AI is everywhere, but not every implementation delivers meaningful value. Which AI-powered feature, workflow or automation has genuinely improved productivity for your team? What problem did it solve and what made it successful? Real examples are always more useful than marketing claims. TIA!

🟡 Notable

Model Releases

🟡 ✉️ Meta sizes up GPT-5.5 with 'Watermelon' — score 65 Sources: newsletter/rundown-ai

🟡 ✉️ Why it matters: Anthropic has been criticized (most vocally by Microsoft AI head Mustafa Suleyman) for its AI consciousness talk, and while the researchers note this doesn’t reveal “whether Claude is conscious… or feels — score 65 Sources: newsletter/rundown-ai

🟡 ✉️ The Rundown: Chinese tech giant Tencent’s Hunyuan just promoted its Hy3 model out of an April preview and into a full open-source release, with benchmark results that claim to rival “flagship open-source models with 2-5x — score 65 Sources: newsletter/rundown-ai

🟡 💬 Built a traffic light widget for my Claude Code sessions because I kept forgetting them — score 61 Sources: reddit/r/AIAgents

I have a talent for opening a Claude Code session, getting distracted, and rediscovering it four hours later still waiting for me to select "allow"... So I decided to do the only reasonable thing: vibecoding a widget instead of just looking at my terminal more often lol Session Signals opens a float

🟡 🏢 Australian Payments Plus moves faster with ChatGPT and Codex — score 50 Sources: lab_blog/OpenAI

See how Australian Payments Plus uses ChatGPT Enterprise and Codex to move faster through payments complexity. AP+ saves time, improves quality, and keeps human judgment central.

Omitted 2 additional model releases items from the main section; see raw data and source-specific sections below.

Developer Tools

🟡 💬 Ph.D. thesis on Differentiable Ray Tracing for Radio Propagation Modeling [R] — score 69 Sources: reddit/r/MachineLearning

Hi everyone, I recently finished my Ph.D. thesis on Differentiable Ray Tracing for Radio Propagation Modeling. Instead of just compiling my published papers, I tried to write it as an accessible, self-contained textbook for anyone interested in the intersection of radio propagation simulation, a

🟡 ✉️ NVIDIA TAPS AI CLOUD PROVIDERS TO EXPAND COMPUTE ACCESS FOR STARTUPS (2 MINUTE READ) — score 65 Sources: newsletter/tldr

🟡 ✉️ AGENTIC AUTONOMY LEVELS (19 MINUTE READ) — score 65 Sources: newsletter/tldr

🟡 ✉️ AGENTIC LOOPS (20 MINUTE READ) — score 65 Sources: newsletter/tldr

🟡 ✉️ SOME NEW AGENTIC PATTERNS (11 MINUTE READ) — score 65 Sources: newsletter/tldr

Omitted 9 additional developer tools items from the main section; see raw data and source-specific sections below.

Infrastructure & Compute

🟡 ✉️ HOW AN AI TOKEN TRAVELS THROUGH A DATA CENTER (19 MINUTE READ) — score 65 Sources: newsletter/tldr

🟡 ✉️ APPLE PLANS FIVE NEW IPHONES THROUGH 2027, EYES CHINESE-MADE CHIPS AMID FOLDABLE PUSH (2 MINUTE READ) — score 65 Sources: newsletter/tldr

🟡 ✉️ DOES A URL IN A PROMPT STEER AN LLM'S OUTPUT TOWARD ITS CONTENT? (22 MINUTE READ) — score 65 Sources: newsletter/tldr

🟡 ✉️ The game runs at 20 frames per second on a single Nvidia GPU, with the teams open-sourcing the code, training data, and a playable demo. — score 65 Sources: newsletter/rundown-ai

Business & Funding

🟡 ✉️ TIME-SERIES LLMS, EXPLAINED WITH T0-ALPHA (13 MINUTE READ) — score 65 Sources: newsletter/tldr

🟡 ✉️ Oasis introduced Oasis 1, a $289 smart ring that allows users to dictate to AI using Whispr Flow and control apps or devices with a built-in trackpad. — score 65 Sources: newsletter/rundown-ai

Research Papers

🟡 🤗 Bridging Interleaved Multi-Modal Reasoning as a Unified Decision Process — score 68 Sources: huggingface · arxiv/cs.AI

Unified multi-modal models (UMMs) have shown promising interleaved text-image reasoning capabilities, yet effectively optimizing such multi-turn generation via reinforcement learning (RL) remains an open challenge. Existing approaches apply RL exclusively to text steps, relegating image generation t

🟡 🤗 Look Before You Leap: Distilling Tree Search into Action Evaluation for Frozen VLA Models — score 65 Sources: huggingface

Vision-Language-Action (VLA) models acquire broad embodied capabilities through large-scale pretraining, yet their generalization remains far more fragile than that of LLMs and VLMs. The prevailing remedy, post-training via supervised fine-tuning or reinforcement learning, improves task-specific per

🟡 🤗 Speaker-Aware Temporal Aggregation Strategies on Segment Representations for Depression Detection in Dyadic Interaction: A Benchmark Study — score 65 Sources: huggingface

Speech-based depression detection compresses features from short audio segments into one speaker-level decision, a step called temporal aggregation rarely studied on its own. Most benchmarks fix a single self-supervised encoder and a single hand-picked layer, so a reported gain may reflect the pipel

🟡 🤗 Taste-aware music retrieval from audio embeddings — score 52 Sources: huggingface · arxiv/cs.LG

Crossmodal correspondences between sound and taste are well established in psychology and neuroscience, but largely absent from content-based multimedia retrieval. We formalise taste-from-audio prediction as a content-based music information retrieval benchmark over a perceptually validated multi-so

🟡 🤗 GaP: A Graph-as-Policy Multi-Agent Self-Learning Harness For Variational Automation Tasks — score 40 Sources: huggingface · arxiv/cs.AI

For robots to work reliably in commercial and industrial applications, can recent advances in agentic coding systems combine interpretable robot programming with the open-world adaptability of model-free policies? We focus on "Variational Automation" (VA), a class of tasks that have larger variation

Other Signals

🟡 💬 Chinese AI models are gaining ground with U.S. companies as OpenAI, Anthropic costs surge — score 65 Sources: reddit/r/LocalLLaMA

🟡 ✉️ THE AI SUPERFORECASTERS ARE HERE (30 MINUTE READ) — score 65 Sources: newsletter/tldr

🟡 ✉️ AI HAS TORCHED THE MARKET FOR JUNIOR PROGRAMMERS (8 MINUTE READ) — score 65 Sources: newsletter/tldr

🟡 ✉️ MIDJOURNEY WANTS HOLLYWOOD STUDIOS TO REVEAL THE DETAILS OF THEIR AI USAGE (2 MINUTE READ) — score 65 Sources: newsletter/tldr

🟡 ✉️ YOU DESIGN IT. THEN WHAT? A CLEAR MAP OF THE FIGMA-TO-CODE AI MESS (15 MINUTE READ) — score 65 Sources: newsletter/tldr

Omitted 13 additional other signals items from the main section; see raw data and source-specific sections below.

🟢 Incremental

Model Releases

🟢 💬 Liquid AI - Antidoom (the doom loop remover) — score 19 Sources: reddit/r/LocalLLaMA

https://x.com/liquidai/status/2074494130126811473 Today we release Antidoom, an open-source method that removes a common failure mode in reasoning models: the doom loop. Doom-loop rates before and after, with eval scores up across the board: >

🟢 💬 I built a tiny proxy that gives GLM 5.2 vision (or any text LLM) – MIT — score 0 Sources: reddit/r/LocalLLaMA

VisionBridge lets you give text-only LLMs vision. It's tiny OpenAI-compatible proxy that lets reasoning models (DeepSeek, Qwen, GLM…) see images by querying a separate vision model through tools: look, OCR, scan, crop, compare. No training, no weights. MIT

Developer Tools

🟢 🐙 HKUDS/AI-Trader — "AI-Trader: 100% Fully-Automated Agent-Native Trading" — score 36 Sources: github_trending

"AI-Trader: 100% Fully-Automated Agent-Native Trading"

🟢 💬 [D] Issue with arxiv - abstract not matching pdf/html [D] — score 31 Sources: reddit/r/MachineLearning

Hi, I was reading the openRLHF paper: https://arxiv.org/pdf/2501.03262v4 , but when I click the abstract page: https://arxiv.org/abs/2501.03262v4 , it shows "REINFORCE++". Note that [https://arxiv.org/html/2501.03262v4](http

🟢 🐙 datawhalechina/all-in-rag — 🔍大模型应用开发实战一:RAG 技术全栈指南,在线阅读地址:https://datawhalechina.github.io/all-in-rag/ — score 31 Sources: github_trending

🔍大模型应用开发实战一:RAG 技术全栈指南,在线阅读地址:https://datawhalechina.github.io/all-in-rag/

🟢 🧡 Show HN: Docx-CLI: agents read/edit Word docs using 1/2 the time and tokens — score 30 Sources: hackernews

🟢 💬 GLM-5.2 on 8xB200: the deployment math nobody spells out - NVFP4 + 2x TP=4 replicas should beat TP=8 by ~2x. Full config guidance inside. — score 27 Sources: reddit/r/LocalLLaMA

We have 8xB200 nodes and users keep asking us how to serve GLM-5.2 on them. Our engineering team went through everything published so far, and the optimal config is not the obvious one. Sharing the analysis because most of it applies wherever you rent or rack your B200s. The model GLM-5.2: ~750

Omitted 4 additional developer tools items from the main section; see raw data and source-specific sections below.

Infrastructure & Compute

🟢 🧡 IEEE Rolls Out Large Language Models Training Course — score 10 Sources: hackernews

🟢 💬 Mid research got me thinking what about reversed alignment, would trained "bad" model exhibit"good" behavior later and/or secretly [D] — score 6 Sources: reddit/r/MachineLearning

late night thoughts as I was working on my paper that is about specific behavior that arises from RHLF, it got me thinking what if train a model in an environment where bad behavior is rewarded: deception, selfishness, harmful behavior etc. and then find it occasionally and/or secretly exhibit good

Research Papers

🟢 🤗 SynCity 3000: Bootstrapping Scene-Scale 3D Diffusion — score 35 Sources: huggingface

We present SynCity 3000, a framework for generating 3D scenes that are globally coherent while enabling fine-grained layout control. Building on the ability of current image-to-3D generators to produce complex 3D assets from a single image, we extend this capability to the scale of entire scenes by

Other Signals

🟢 💬 local already feels good enough — score 35 Sources: reddit/r/LocalLLaMA

This is specifically for coding, technical planning, and hardware setup. _____ The only times Qwen 3.6 35B A3B has let me down, it has been something that is resolved with a workflow/discipline improvement. If I go through and take the time to get the proper set up with a sound plan, it hasn't

🟢 💬 I tested freshly merged DFlash in llama.cpp on Qwen 3.6 27B Local AI win. 4.44x faster at 36K context. Here are my findings RTX 6000 PRO. — score 12 Sources: reddit/r/LocalLLaMA

Hey guys, A month ago I posted my MTP benchmarks here (3.34x on Gemma 4). DFlash support just merged into llama.cpp (PR #22105), so I ran it on the same rig with the Qwen 3.6 27B and it beat my best MTP numbers at every draft length. DFlash is speculative decoding with a block diffusion drafter from

🟢 💬 Building a brain for your product for making your plan efficient as possible — score 6 Sources: reddit/r/AIAgents

RepoDescriptionStars TodayLanguage
MadsLorentzen/ai-job-searchAI-powered job application framework built on Claude Code. Fork it, fill in your profile, and let Claude evaluate jobs, tailor CVs, write cover letters, and prepare you for interviews.2402typescript
lessweb/deepcode-cliDeep Code 是专为 deepseek-v4 模型优化的终端 AI 编码助手,支持深度思考、推理强度控制以及 Agent Skills。128typescript
XiaoYouChR/Ghost-Downloader-3An AI-boost cross-platform multi-protocol fluent-design concurrent downloader built with Python & Qt.59python
HKUDS/AI-Trader"AI-Trader: 100% Fully-Automated Agent-Native Trading"44python
datawhalechina/all-in-rag🔍大模型应用开发实战一:RAG 技术全栈指南,在线阅读地址:https://datawhalechina.github.io/all-in-rag/38python
SawyerHood/dev-browserA Claude Skill to give your agent the ability to use a web browser28typescript

📄 New Papers

TitleCategoryHotnessLink
Unified Audio Intelligence Without Regressing on Text Intelligenceresearch_paper8Open
Bridging Interleaved Multi-Modal Reasoning as a Unified Decision Processresearch_paper4Open
Look Before You Leap: Distilling Tree Search into Action Evaluation for Frozen VLA Modelsresearch_paper4Open
Speaker-Aware Temporal Aggregation Strategies on Segment Representations for Depression Detection in Dyadic Interaction: A Benchmark Studyresearch_paper4Open
Taste-aware music retrieval from audio embeddingsresearch_paper3Open
iFLYTEK-Embodied-Omni Technical Reportcs.AI0Open
Internal Pluralism and the Limits of Pairwise Comparisonscs.AI0Open
ASK in the Dark: Uncertainty-Gated LLM Assistance under Partial Observabilitycs.AI0Open
Automated Data Readiness for Scientific AIcs.AI0Open
SwarmResearch: Orchestrating Coding Agents for Open-Ended Discoverycs.AI0Open
Object-Centric Environment Modeling for Agentic Taskscs.AI0Open
MedCalc-Pro: Solving Complex Medical Calculations with LLM Agentscs.AI0Open
Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Modelscs.AI0Open
VERITAS: Towards a General-Purpose Replication Tool for Scientific Researchcs.AI0Open
A Sliding-Window-Based Reinforcement Learning for Dynamic Assembly Flow Shop Scheduling with Multi-Product Deliverycs.AI0Open

🏢 Lab Blog Posts

🐦 Twitter/X Highlights

AccountTweet Summary
GoogleDeepMind🏛️ We’re unveiling a new way to converse with the ancient world. By grounding Gemini directly in our expert models Aeneas and Ithaca, our Predicting the Past Skill in Google @antigravity lets historians study Greek and Latin texts using plain English. 🧵 Post
simonwI released sqlite-utils 4.0, the 124th release but the first major version bump since 3.0 back in 2020 I managed to keep things backwards-compatible all the way up to version 3.39 before the accumulated design mistakes forced me to bump that number! https://simonwillison.net/2026/Jul/7/sqlite-utils- Post

Newsletter

Repeated From Recent Briefings