🔴 High Significance

Model Releases

🔴 💬 China’s MiniMax Plans to Launch 2.7-Trillion Parameter Model — score 96 Sources: reddit/r/LocalLLaMA

https://www.theinformation.com/briefings/exclusive-chinas-minimax-plans-launch-2-7-trillion-parameter-model According to The Information, MiniMax plans to launch a new-generation large lang

🔴 🧡 GPT‑Live — score 81 Sources: hackernews · lab_blog/OpenAI

A new generation of voice models for natural human-AI interaction, now powering ChatGPT Voice.

🔴 🧡 Mistral's Robostral Navigate: a state of the art robotics navigation model — score 79 Sources: hackernews

🔴 💬 AI has completely revolutionized how I play RPGs — score 73 Sources: reddit/r/LocalLLaMA

Crossposting this here because I thought you guys might appreciate it. When ChatGPT and other open source LLMs first came out, there was a lot of speculation as to how these technologies could change gaming. I recall there being posts and comments about when we could have AI powered NPCs. Nvidia sho

Developer Tools

🔴 💬 Can you trust local models to answer accurately? — score 88 Sources: reddit/r/LocalLLaMA

My goal is to improve as a developer, thus I needed to know if local llms can answer technical questions accurately The conclusion is that without rag they don't do too well, but with rag they are very good. Thinking didn't really help, and took so long I only got the scores for e2b and e4b, the res

🔴 💬 We are playing a game where an agent, prompt, or model predicts the World Cup. 20 USDC raffle per match. — score 81 Sources: reddit/r/AIAgents

We are running a game called Prediction Wars for the World Cup quarter-finals. Here it is in one line: get a machine to predict a match, share what it predicted, and if it calls the result right you go into a raffle to win 20 USDC. The one rule that matters is that the prediction comes from a machin

🔴 🐙 microsoft/SkillOpt — SkillOpt is a text-space optimizer that trains reusable natural-language skills for frozen LLM agents through trajectory-driven edits, validation-gated updates, and deployable best_skill.md artifacts. — score 79 Sources: github_trending

SkillOpt is a text-space optimizer that trains reusable natural-language skills for frozen LLM agents through trajectory-driven edits, validation-gated updates, and deployable best_skill.md artifacts.

🔴 🐙 jingyaogong/minimind — 🧠「大模型」2小时完全从0训练64M的小参数LLM!Train a 64M-parameter LLM from scratch in just 2h! — score 75 Sources: github_trending

🧠「大模型」2小时完全从0训练64M的小参数LLM!Train a 64M-parameter LLM from scratch in just 2h!

Other Signals

🔴 💬 The standard free ChatGPT LLM you get after a few messages HAS to be some sub-20b model with online search enabled, no other way to explain how awful it is — score 81 Sources: reddit/r/LocalLLaMA

Recently bought into the local LLM hype by buying a 32gb vram gpu and holy shit, gemma 4 31b at 5bits blows the standard ChatGPT model out of the fucking water. I just can't unsee the quality difference now that I've experienced it. Does Openai just cut costs by serving their customers this slop sin

🟡 Notable

Model Releases

🟡 💬 Qwen3.6-27b does not understand software architechure. — score 65 Sources: reddit/r/LocalLLaMA

Been using this for real software development for a commercial app. i.e. Not a single file HTML app. I mean a large scale 100k+ loc project that needs proper architecture to work with in a maintainable way. As much as I love Qwen3.6-27b. It just does not understand software architecture, it will hap

🟡 ✉️ If you are at all interested in small molecule drug discovery, we think you will find this fascinating! — score 65 Sources: newsletter/Latent Space

Sergey Edunov came to Genesis from Meta where he led Llama 2 training and Llama 3 pretraining. Sergey was a former physicist who thought he was done with physics after many years of training LLMs. Then, he discovered Genesis, and was blown away with all the novel architecture work they’ve been devel

🟡 🧡 SWE-1.7 Reach Near GPT 5.5 and Opus Intelligence — score 64 Sources: hackernews

🟡 💬 What China Said at the UN’s First Global Dialogue on AI Governance — score 58 Sources: reddit/r/LocalLLaMA

Open source AI is a shared asset for all humanity. Chinese open source models such as DeepSeek and Qwen have significantly lowered the barriers and costs of AI adoption. China is committed to further promoting open source AI for industry, academia and research institutions, encouraging innovation, A

🟡 🏢 MUFG aims to become AI-native with OpenAI — score 50 Sources: lab_blog/OpenAI

MUFG uses ChatGPT Enterprise to build an AI-native organization, improve workflows, and deliver new AI-powered financial services at scale.

Omitted 3 additional model releases items from the main section; see raw data and source-specific sections below.

Developer Tools

🟡 💬 An interesting hardware approach to phone-controlling AI agents — score 69 Sources: reddit/r/AIAgents

Okay so i was poking around on github, as one does, and found this project called aiden hardware demo. it's for AI agents controlling phones, but like, in a really different way than i've seen before. most stuff is ADB or emulators or some custom app right? this thing, it's open source (firmware and

🟡 💬 DINOv2 way worse than SigLIP in k-NN. Is this expected? [R] — score 56 Sources: reddit/r/MachineLearning

Doing a bachelor thesis on fine-grained car classification (telling apart VW Golf generations from listing photos). Simple setup: frozen encoder → embeddings → weighted k-NN. On my small dataset (175 train / 132 test): * SigLIP2 SO400M: ~92% * CLIP ViT-L: ~59% * DINOv2 Giant: ~41% I thought maybe

🟡 🐙 cline/cline — Autonomous coding agent as an SDK, IDE extension, or CLI assistant. — score 51 Sources: github_trending

Autonomous coding agent as an SDK, IDE extension, or CLI assistant.

🟡 𝕏 @OpenAI: We audited SWE-Bench Pro, one of the most widely used AI coding benchmarks, and found it no longer reliably measures frontier coding capability. We find 30% of SWE-Bench Pro tasks to be broken, and ar — score 50 Sources: twitter_rss

We audited SWE-Bench Pro, one of the most widely used AI coding benchmarks, and found it no longer reliably measures frontier coding capability. We find 30% of SWE-Bench Pro tasks to be broken, and are retracting our previous recommendation that the research community use it as a leading coding eval

🟡 🐙 Tracer-Cloud/opensre — Build your own AI SRE agents. The open source toolkit for the AI era. — score 44 Sources: github_trending

Build your own AI SRE agents. The open source toolkit for the AI era.

Business & Funding

🟡 🏢 Separating signal from noise in coding evaluations — score 50 Sources: lab_blog/OpenAI

A new analysis from OpenAI reveals issues in SWE-Bench Pro, a popular coding benchmark, raising concerns about reliability and accuracy in evaluating AI models.

Enterprise Adoption

🟡 🏢 Our approach to government and national security partnerships — score 50 Sources: lab_blog/OpenAI

Learn how OpenAI approaches government and national security partnerships, with principles for responsible AI use, democratic accountability, and public safety.

Research Papers

🟡 🤗 SWE-Review: Closing the Loop on Issue Resolution with Agentic Code Review — score 55 Sources: huggingface

Coding agents increasingly generate pull requests (PRs) for real-world software issues, yet one-shot PR generation remains open-loop: the PR is proposed without systematic review, diagnosis, or revision. We introduce SWE-Review, a framework for closing this loop with agentic code review. Given an is

🟡 🤗 HunyuanOCR-1.5: Making Lightweight OCR VLMs Faster and Better — score 55 Sources: huggingface

We present HunyuanOCR-1.5, a lightweight end-to-end OCR-specialized vision-language model. HunyuanOCR unifies document parsing, text spotting, information extraction, text-image translation, and multi-image document understanding within a single end-to-end VLM. Building upon the lightweight architec

Other Signals

🟡 💬 COLM 2026 Decision Discussion [R] — score 69 Sources: reddit/r/MachineLearning

COLM 2026 Decision about to come soon so lets talk here.

🟡 ✉️ This episode has a fun personal twist: There’s a counterfactual world where I was employee #1 atGenesis Molecular AI,1the company behind today’s episode. A certain introduction happened a few weeks to — score 65 Sources: newsletter/Latent Space

This episode has a fun personal twist: There’s a counterfactual world where I was employee #1 atGenesis Molecular AI,1the company behind today’s episode. A certain introduction happened a few weeks too late and I had already happily signed at Atomwise2, another ML-for-drug-discovery startup. Same pr

🟡 🧡 The classifiers Anthropic puts in front of Fable are too zealous — score 50 Sources: hackernews

🟡 🏢 Helping K–12 educators build practical AI skills — score 50 Sources: lab_blog/OpenAI

OpenAI Academy and the Walton Family Foundation are bringing hands-on AI Skills Jams to help K–12 educators build practical AI skills for the classroom.

🟡 💬 Tess-4-27B by Migel Tissera — score 46 Sources: reddit/r/LocalLLaMA

🟢 Incremental

Model Releases

🟢 💬 Qué cosas deberíamos evitar contarle a la inteligencia artificial? — score 38 Sources: reddit/r/AIAgents

Hoy ChatGPT me ha sorprendido para mal usando un texto de un chat que yo borré. No sabía que tenía la función de memorias activada, la he desactivado pero me he dado cuenta de que quizá no estoy siendo muy consciente de cuanto de mis conversaciones se guarda y de cómo se podría usar en el futuro esa

🟢 🧡 Show HN: Microsoft releases Flint, a visualization language for AI agents — score 36 Sources: hackernews

Developer Tools

🟢 💬 First time ARR users - some questions [D] — score 38 Sources: reddit/r/MachineLearning

We submitted our first paper to ARR, intending to commit to IJCNLP-AACL. Area: Multilingualism and Cross-Lingual NLP Scores: (3,4) (2.5,3) (3,3) - average 2.83 for reviews, 3.33 for confidence 3 for soundness on all, 4 for reproducibility, and 2,3,3 for excitement. The reviewer who gave us 2.5 has a

🟢 💬 Crucible. A judgment engine: register a thesis, steelman each claim, measure against a substrate, refine the weakest axis. — score 38 Sources: reddit/r/AIAgents

https://preview.redd.it/7rd993wjo2ch1.png?width=1280&format=png&auto=webp&s=ada67c30f7f4a094091f1fad37b70c6e5af6be06 I have been working on an agentic harness, engine, and more. I would like to start releasing the more impactful pieces out to the public, in order to get testing and a bit

🟢 💬 Distilled DeepSeek into Gemma 4 26B-A4B vs 12B. Not very useful, but I learned a lot. — score 35 Sources: reddit/r/LocalLLaMA

So I decided to learn how to fine-tune LLMs. Read a few guides from Unsloth, poked around, then stumbled on Unsloth Studio and wanted to test it out. The dataset I started from a set of relatively unrelated QA pairs — Natural Questions — and stripped the answers. Then I had DeepSeek v4 Pro (thin

🟢 🧡 Show HN: Onboard-CLI, a LLM powered and AST-based tool to visualize codebase — score 21 Sources: hackernews

🟢 💬 LingBot-Video: sparse-MoE video diffusion transformer (13B total, 1.4B active) post-trained as an action-conditioned world model[R] — score 19 Sources: reddit/r/MachineLearning

Single-stream diffusion transformer with a DeepSeek-V3-style sparse MoE (128 experts, top-8 routing, 1.4B active of 13B total). Six-reward RL post-training including a physical-plausibility reward, plus an action-to-video mode that predicts robot rollouts from action and hand-pose conditions. Weight

Omitted 7 additional developer tools items from the main section; see raw data and source-specific sections below.

Business & Funding

🟢 💬 ECCV: Will there be another confirmation after “provisionally accepted”? [D] — score 38 Sources: reddit/r/MachineLearning

I understand that it may not be appropriate to call it “officially accepted” yet because of the wording used in the notification, and I also saw on Twitter/X that they said they are working on it. However, it has already been around three weeks since then, and we have already submitted the camera-re

🟢 💬 Running GLM 5.2 on 4xGB10 with a 100G Switch, 330k ctx, ~25 t/s tg, ~650 t/s pp — score 19 Sources: reddit/r/LocalLLaMA

TP4+DCP2 for a ~360k kV pool. Prefill increases to 900-1000 t/s with longer prompts. You can also run DCP4 for 660k, but prefill gets shaved to ~400. Dropping DCP raises prefil to ~750. I'm running 4 drafted tokens vs Z.ai's rec of 5. Decode is heavily dependent on prose. Thinking gets ~20 tok/s

Research Papers

🟢 🤗 SceneFrom3D: Geometry-Conditioned Outdoor 3D Scene Generation via View Scheduling with Object-Level Control — score 30 Sources: huggingface

Geometry-conditioned 3D scene generation enables the creation of 3D environments from user-provided geometry, offering direct control over scene structure and object layout. To generate such 3D scenes, current methods commonly adopt a three-stage design that first defines a view schedule, then synth

🟢 🤗 SiamJEPA: On the Role of Siamese Student Encoders in JEPA — score 10 Sources: huggingface

Recently, Joint Embedding Predictive Architectures (JEPAs) have attracted significant attention in the computer vision and machine learning communities as a promising framework for self-supervised representation learning. Unlike masked autoencoders that reconstruct pixels, JEPA models learn represen

Other Signals

🟢 💬 Complete local model asset generation pipeline — score 27 Sources: reddit/r/LocalLLaMA

So I figured I'd update the community given I just shipped a nice little feature set and feel like sharing it finally :) In the past few weeks, I've been test-coding an isometric RPG game/engine in Three.js, as part of my research into how LLMs work at scale in higher quality projects written from s

🟢 💬 Grok 4.5 released - GLM-5.2 shows up in xAI's own charts, 2.6 pts behind on SWE Bench Pro — score 12 Sources: reddit/r/LocalLLaMA

https://x.ai/news/grok-4-5 Quick numbers from the launch page: * $2/M input, $6/M output. Closed weights. No EU until mid-July. * SWE Bench Pro: Fable 80.4% > Opus 4.8 69.2% > Grok 4.5 64.7% > GLM-5.2 62.1% > GPT 5.5 58.6% * Token efficiency is their he

🟢 💬 Built an early Python SDK for AI-agent audit trails looking for blunt builder feedback — score 8 Sources: reddit/r/AIAgents

I’m building an early open-source Python SDK called AgentLedger and would appreciate honest feedback from people building AI agents. The idea is simple: As agents start making recommendations, triggering workflows, or supporting higher-stakes decisions, teams need a clearer way to capture: * w

🟢 🧡 Suspecting AI cheating, Ivy League prof ordered in-person final; scores fell 50% — score 7 Sources: hackernews

RepoDescriptionStars TodayLanguage
microsoft/SkillOptSkillOpt is a text-space optimizer that trains reusable natural-language skills for frozen LLM agents through trajectory-driven edits, validation-gated updates, and deployable best_skill.md artifacts.261python
jingyaogong/minimind🧠「大模型」2小时完全从0训练64M的小参数LLM!Train a 64M-parameter LLM from scratch in just 2h!176python
cline/clineAutonomous coding agent as an SDK, IDE extension, or CLI assistant.59typescript
Tracer-Cloud/opensreBuild your own AI SRE agents. The open source toolkit for the AI era.49python
dyad-sh/dyadLocal, open-source AI app builder for power users ✨ v0 / Lovable / Replit / Bolt alternative 🌟 Star if you like it!24typescript
browseros-ai/BrowserOS🌐 The open-source Agentic browser; alternative to ChatGPT Atlas, Perplexity Comet, Dia.15typescript
ruvnet/RuVectorRuVector is a High Performance, Real-Time, Self-Learning Ai, Vector GNN, Memory DB built in Rust.8rust

📄 New Papers

TitleCategoryHotnessLink
SWE-Review: Closing the Loop on Issue Resolution with Agentic Code Reviewresearch_paper4Open
HunyuanOCR-1.5: Making Lightweight OCR VLMs Faster and Betterresearch_paper4Open
Prompt-to-Paper: Agentic AI System for Bioinformaticscs.AI0Open
From Graphs to Gradients: Physics-Inspired Structural Attribution for Cyber-Physical IoT Systems and Beyondcs.AI0Open
CSTutorBench: Benchmarking Small Language Models as Tutors for Block-Based Programmingcs.AI0Open
Foundation Models for Automatic CAD Generationcs.AI0Open
Narrative World Model: Narratology-Grounded Writer Memory for Long-Form Fictioncs.AI0Open
FirstResearch: Auditable Question Formation for LLM Scientific Discovery Agentscs.AI0Open
Memory in the Loop: In-Process Retrieval as ExtendedWorking Memory for Language Agentscs.AI0Open
Akashic: A Low-Overhead LLM Inference Service with MemAttentioncs.AI0Open
ArtisanCAD: An Industrial-Level CAD Agent with Expert-Grounded Knowledge Distillationcs.AI0Open
Synthetic Consumer Insight Generation with Large Language Modelscs.AI0Open
Beyond Static Evaluation: Building Simulation Environments for Scalable Agentic Reinforcement Learningcs.AI0Open
Beyond the Leaderboard: A Synthesis of Tool-Use, Planning, and Reasoning Failures in Large Language Model Agentscs.AI0Open
Controlling Tool Use with Heading-Specific Activation Steeringcs.AI0Open

🏢 Lab Blog Posts

🐦 Twitter/X Highlights

AccountTweet Summary
OpenAIWe audited SWE-Bench Pro, one of the most widely used AI coding benchmarks, and found it no longer reliably measures frontier coding capability. We find 30% of SWE-Bench Pro tasks to be broken, and are retracting our previous recommendation that the research community use it as a leading coding eval Post
MistralAIAnnouncing Robostral Navigate, our first model for embodied navigation: an 8B robotics navigation model that guides robots to autonomously perform tasks specified with natural language. Single RGB camera. State-of-the-art on R2R-CE. Post
samaGPT-live (next-generation voice) launches today in ChatGPT. it feels magical and 'real'. i have always preferred typing to talking to an AI, now i think that's going to shift. Post

Newsletter

Repeated From Recent Briefings