🔴 High Significance

Model Releases

🔴 🧡 Introducing System One Models and Jev — score 84 · 🔥 engaged Sources: hackernews

🔴 💬 Gemini 3.8 Live & Gemini 3.8 Live Extended Thinking — score 78 · 🔗 ×3 · 🏢 first-party Sources: reddit/r/singularity · hackernews · lab_blog/DeepMind

🔴 💬 Voodoo Dynamic Quant - Now MIT Licensed — score 72 Sources: reddit/r/LocalLLaMA

Two months ago I announced I had found a new dynamic quant method called Voodoo Quant which was SOTA for the most aggressive quant levels on some smaller Qwen3.5 GGUF models. I kept the methodology private at the time, but I've seen too many requests for dyn quants for various models lately, so I de

🔴 ✉️ tldraw took OpenAI up on a challenge. Steve (the founder of tldraw) said he could make ChatGPT’s Sketch 100x better. OpenAI’s Tibo gave him a day to prove it. The result: a whole ChatGPT-style prototy — score 70 Sources: newsletter/Ben's Bites

tldraw took OpenAI up on a challenge. Steve (the founder of tldraw) said he could make ChatGPT’s Sketch 100x better. OpenAI’s Tibo gave him a day to prove it. The result: a whole ChatGPT-style prototype with better drawing tools built in.

🔴 ✉️ ChatGPT mini- A tiny floating widget to start chats, see updates, and more. Go to Pets in your ChatGPT desktop app to switch. — score 70 Sources: newsletter/Ben's Bites

Omitted 6 additional model releases items from the main section; see raw data and source-specific sections below.

Developer Tools

🔴 💬 CrofAI "cheapest inference provider in the world" gets exposed as an OpenRouter wrapper, routing requests to smaller, cheaper models at up to 20x markup. CrofAI responds to Wire Fraud allegations by denying everything, then backtracking, then 3 hours later wiping their entire online presence — score 83 · 🔥 engaged Sources: reddit/r/LocalLLaMA

Disclaimer: no AI was used whatsoever to write this post Cautionary tale about chasing cheap tokens. exposé: https://kendell.dev/blog/crofaifalse/ reaction by nahcrof, announcing the shutdown of the service: https://x.com/nahcrof/status/2099552389434900643 - now deleted, archive picture: https://i.i

🔴 💬 Got it unopened off Craigslist for $4k. Excited to start hosting my own models! — score 77 · 🔥 engaged Sources: reddit/r/LocalLLaMA

I’ve been tempted to get a spark for a few months now but the prices were climbing. A new one can go for anywhere from $5k-$6k. I was lucky enough to be scrolling through Craigslist when I stumbled upon this. The owner won it in a raffle and was trying to unload it quickly while they were in town. F

🔴 ✉️ How tf do I write an intro to the craziness that’s happened since the end of last week?! — score 70 Sources: newsletter/Ben's Bites

I got access to Instinct, the personal agent all the VCs are raving about - I think it’s a bit meh? I don’t know if it’s the pro-activeness that people seem to like, but I don’t love that. Makes me feel like I’m having to do work to keep it happy, or like I’ve got a boss again - no thanks.

🔴 ✉️ “Games have always been these underrated educational tools. They’re super approachable. They’re very human.” — score 70 Sources: newsletter/Latent Space

The bigger idea isthat the way a game is presented to an AI can determine which skills it learns, and whether those skills carry into work outside the game. And the best evidence for that so far comes from a nineteenth-century railroad game.

🔴 ✉️ The idea came froma 2025 Twitch streamof frontier models playing the game Diplomacy, which Duffy said normally takes “days or weeks to play.” This was when he worked at Every as its head of AI trainin — score 70 Sources: newsletter/Latent Space

The idea came froma 2025 Twitch streamof frontier models playing the game Diplomacy, which Duffy said normally takes “days or weeks to play.” This was when he worked at Every as its head of AI training.

Omitted 13 additional developer tools items from the main section; see raw data and source-specific sections below.

Infrastructure & Compute

🔴 ✉️ Those are the words ofAlex Duffy, co-founder and CEO ofGood Start Labs, who spoke to Latent Space about why his company isturning games into training material for AI models. The company was spun out o — score 70 Sources: newsletter/Latent Space

Those are the words ofAlex Duffy, co-founder and CEO ofGood Start Labs, who spoke to Latent Space about why his company isturning games into training material for AI models. The company was spun out of AI media and tools company Everylast October, with$3.6 million in fundingfrom General Catalyst, In

🔴 ✉️ From this, Duffy concluded that training AI models on games like Diplomacy couldteach them skills like strategic thinking. Especially because those kinds of games have outcomes that can be verified.In — score 70 Sources: newsletter/Latent Space

From this, Duffy concluded that training AI models on games like Diplomacy couldteach them skills like strategic thinking. Especially because those kinds of games have outcomes that can be verified.In a laterarticle published on Every, Duffy wrote that “fine-tuning a model on the strategy game Diplo

Business & Funding

🔴 ✉️ No IPO for OpenAI in 2026. Sam told Fortune’s Alyson Shontell it would be “ill-timed” given the work ahead on alignment, control and safety. — score 70 Sources: newsletter/Ben's Bites

🔴 ✉️ Today, we start a new series about one of the hottest topics in AI: recursive self improvement. Throughout the next few weeks, we will deep dive into the top research — score 70 Sources: newsletter/TheSequence

For about sixty years, arguing about recursive self-improvement meant arguing in the abstract. There was no system to point at. You cited I.J. Good, somebody cited Schmidhuber, everyone disagreed about definitions for two hours, and then you went home. It was a very pleasant way to spend an afternoo

Enterprise Adoption

🔴 ✉️ Your AI Adoption Lift Is A Selection Effect (17 Minute Read) — score 70 Sources: newsletter/tldr

Research Papers

🔴 🤗 LynnReal-Omni: Native multi-modal Video Generation for Agentic Visual Workflows — score 75 · 🔗 ×2 Sources: huggingface · arxiv/cs.CV

Video diffusion models are stochastic and hard to control: precise content often requires repeated sampling without guaranteed success, and long-horizon scenes drift in appearance, interactions, and temporal coherence. Agentic visual creation provides explicit references, editable 3D scenes, or exec

🔴 🤗 How Lossless Is Lossless Speculative Decoding? The Role of Numerical Precision in Orthrus — score 72 · 🔗 ×2 Sources: huggingface · arxiv/cs.AI

Orthrus is a hybrid autoregressive-diffusion architecture that accelerates autoregressive language-model inference by generating multiple tokens in parallel while using a frozen autoregressive backbone. Its central claim is that an intra-model consensus mechanism enables lossless speculative decodin

Other Signals

🔴 💬 PBS has a new documentary on AI. — score 84 Sources: reddit/r/artificial

🔴 💬 From Japanese twitter — score 80 · 🔥 engaged Sources: reddit/r/singularity

🔴 𝕏 @sama: There are two ways AI progress could go very badly and that we must avoid. First, we could lose control of the future to AI. This is unacceptable; we are unapologetically on Team Humanity, and AI must — score 80 · 🔗 ×2 Sources: twitter_rss · newsletter/Ben's Bites

There are two ways AI progress could go very badly and that we must avoid. First, we could lose control of the future to AI. This is unacceptable; we are unapologetically on Team Humanity, and AI must always serve people. To ensure that, we need ways to ensure that alignment and safety techniques st

🔴 💬 I Think All the Current Pessimism is Just an Attempt by XAI, Anthropic, and Open AI to Get Local Models Criminalized — score 76 Sources: reddit/r/artificial

I am not even that pro AI but the furor over the past few days seems suspicious to me.

🔴 🧡 A single firm is behind OpenAI, Anthropic, and Meta hacking scandals — score 76 · 🔥 engaged Sources: hackernews

Omitted 21 additional other signals items from the main section; see raw data and source-specific sections below.

🟡 Notable

Model Releases

🟡 💬 Don’t buy a $9K RTX 5090.... instead. — score 66 Sources: reddit/r/LocalLLaMA

\1. Fly to Taipei. Round-trip from Orlando: $1,081. 2. Go to the largest retailer in Taiwan to Spend NT$129,990 ≈ US$4,093. 3. Hang out in Taiwan for two weeks. Eat good food. Touch international grass. 4. Fly home and flex on r/LocalLLaMA**.** https://preview.redd.it/vtve8s6wgrph1.png?width=287&

🟡 💬 Koboldcpp v1.121 released — score 61 Sources: reddit/r/LocalLLaMA

🟡 🤗 TokenRhythm/NeoHorse-1-4B (11,904 downloads) — score 58 · 🔥 engaged Sources: huggingface_models

Author: | Downloads: 11,904 | Likes: 2021

🟡 𝕏 @sama: The world deserves confidence that American companies developing increasingly capable AI will act responsibly, especially as the trajectory of progress has steepened. Every frontier lab must deliver o — score 55 Sources: twitter_rss

The world deserves confidence that American companies developing increasingly capable AI will act responsibly, especially as the trajectory of progress has steepened. Every frontier lab must deliver on this, and there is no reason any of us should come to work if we cannot. We welcome a federal fram

🟡 💬 Apple Foundation Models: local AI natively on MacOS 27 — score 49 Sources: reddit/r/LocalLLaMA

Maybe some of you know but I didn’t see any post about this. Apple just made available their AFM model on MacOS 27 natively. Just run fm chat in a terminal. Disclaimer: I’m an open weight person. I prefer open models and ecosystem, but I’ll still open the discussion. Did you test them? Build using t

Omitted 5 additional model releases items from the main section; see raw data and source-specific sections below.

Developer Tools

🟡 💬 What do we think about the paperclip maximizer? — score 69 Sources: reddit/r/artificial

I want to vent/talk about this with others, but I couldn't find a sub for it; apologize if it's the wrong place. For some context, the paperclip maximizer thought experiment highlights: if a superintelligence is given a task to manufacture paperclips, it has the potential do everything it can to pro

🟡 💬 Why does working with AI agents still feel so fragmented? — score 67 Sources: reddit/r/AIAgents

With most software projects, the repo is usually the source of truth. With agents, half the logic is scattered across prompts, configs, framework abstractions, tool wiring, and memory setups. Also, portability barely exists. Things break in absurd ways when the framework shifts. Even prompts don't t

🟡 💬 NeurIPS 2026: handling of multiple venue locations seems bad [D] — score 63 Sources: reddit/r/MachineLearning

There has been a recent post acknowledging that NeurIPS passes for Sydney sold out in minutes. For paper authors (guaranteed 1 pass at their designated location) — we recently received forms to select a preferred venue, but apparently we’re not guaranteed to present there. I thought it’s a good idea

🟡 🧡 We got admin access to Baseten's production GitHub in 25 minutes — score 62 Sources: hackernews

🟡 🐙 MG1937/ASC — ASC is a super FAST Android decompiler front-end designed for Agents/Mobile Researchers. — score 62 Sources: github_trending

ASC is a super FAST Android decompiler front-end designed for Agents/Mobile Researchers.

Omitted 3 additional developer tools items from the main section; see raw data and source-specific sections below.

Infrastructure & Compute

🟡 💬 Friendly Reminder :: eye health — score 62 Sources: reddit/r/artificial

I've gotten really good feedback on these positive "Friendly Reminder" posts, so I'm going to keep them posting. For new coders doing vibe coding, new software engineers, or anyone who's suddenly spending a lot more time on a computer because of AI (or for any other reasons): Make sure you're using

🟡 🧡 The Inference Hardware Revolution of 2026 — score 55 Sources: hackernews

Research Papers

🟡 🤗 Learning to Solve Hard Problems in RL for LLMs by Never Giving Up — score 62 · 🔗 ×3 Sources: huggingface · hackernews · arxiv/cs.AI

We demonstrate that training LLMs with RL does not improve performance equally across a dataset. RL shows large improvements on easy problems that an LLM is already good at solving, but small improvements on hard problems. We call this the Matthew Effect in RL for LLMs, after the phenomenon of cumul

🟡 🤗 Root-Cause Attribution Is a Search Problem: Continual Search for Long-Horizon Agent Failures — score 57 · 🔗 ×2 Sources: huggingface · arxiv/cs.AI

The increasing deployment of AI agents in long-horizon tasks yields massive execution logs. Diagnosing failures within these records is crucial for reliability, as it transforms outcome-level signals into actionable interventions. The sheer scale of the data renders human review impractical, driving

🟡 🤗 Thought without systematicity? Evaluating reasoning models on rule induction tasks — score 57 · 🔗 ×2 Sources: huggingface · arxiv/cs.AI

A central tenet of human cognition is systematicity, the principle that understanding one concept is inherently tied to understanding close variations of that concept. Do reasoning models robustly exhibit such systematicity? If so, we would expect consistent performance on structurally equivalent va

🟡 🤗 E2A-Bench: Benchmarking Evidence-to-Action Reliability in Financial Chart Reasoning — score 47 · 🔗 ×2 Sources: huggingface · arxiv/cs.CL

Can financial vision-language models (VLMs) turn chart evidence into reliable action recommendations? Existing hallucination evaluations are mostly claim-centric; they assess whether generated statements are supported, but not whether evidence remains traceable through rationale, confidence, and fin

🟡 🤗 ModaLens: Measuring Image Sensitivity in Report-Conditioned Medical VLMs — score 47 · 🔗 ×2 Sources: huggingface · arxiv/cs.AI

A radiology report can already answer a clinical question, so it is hard to tell whether a vision-language model also uses the image. ModaLens, a paired image-swap audit, measures how report availability changes image sensitivity: MedGemma-27B on 3,199 paired MIMIC-CXR cases from 293 patients, all 1

Other Signals

🟡 💬 Guy who has literally trained a frontier LLM AND engineered viruses thinks the AI-supervirus doomer scenario is bogus. — score 63 · 🔥 engaged Sources: reddit/r/singularity

🟡 💬 Another Google AI safety researcher has quit, warning we might all be about to die — score 61 · 🔥 engaged Sources: reddit/r/OpenAI

🟡 💬 Cut Qwen3.8-27B Reasoning Tokens by 40% -- 3.8 'ThinkingCap' benchmarked! — score 55 Sources: reddit/r/LocalLLaMA

EDIT: Sorry for the unclear title. This model is UkisAI's Swift-Qwen3.8-27B, not a new version of BottleCap AI's 3.6-ThinkingCap. All credit goes to UkisAI for making great fine-tune, and I made this post to celebrate their work. I meant no disrespect by mentioning

🟡 💬 Oracle CFO says 'doing more with less' isn't the answer in all-hands after layoffs — score 55 Sources: reddit/r/artificial

🟡 💬 Continual learning in the fruit fly brain has been decoded, the missing piece for true AGI — score 55 · 🔥 engaged Sources: reddit/r/singularity

Omitted 4 additional other signals items from the main section; see raw data and source-specific sections below.

🟢 Incremental

Model Releases

🟢 💬 How I keep a large agent-built codebase from rotting: specs, outcome notes, and a file list the agent has to respect — score 28 Sources: reddit/r/AIAgents

For about eight months I've been building a desktop app almost entirely through coding agents, first Claude Code and more recently Codex. It's now around 650 commits and 43k lines of JavaScript. I read and approve every change, but I write very little of the code myself. I want to share the process

🟢 💬 For Web game development, I swear Flash Next on Q2 is better than Gemini Flash 3.8 — score 22 Sources: reddit/r/LocalLLaMA

One thing I love about both Qwen 3.8 27B and Flash Next is the effort it puts in to make a good product. It is the opposite of a lazy model and "good enough" is never what aims to accomplish. It goes the extra mile every time, when if it takes forever. Most of my use is making Web Games adapted to a

Developer Tools

🟢 🐙 Jeffallan/claude-skills — 67 Specialized Skills for Full-Stack Developers. Transform Claude Code into your expert pair programmer. — score 33 Sources: github_trending

67 Specialized Skills for Full-Stack Developers. Transform Claude Code into your expert pair programmer.

🟢 💬 The Hacker's Guide to Attacking AI Agents — score 26 Sources: reddit/r/artificial

🟢 🧡 Show HN: Pizza Bot – An inbox for AI agents that work in the background — score 26 Sources: hackernews

🟢 🐙 RightNow-AI/openfang — Open-source Agent Operating System — score 13 Sources: github_trending

Open-source Agent Operating System

Infrastructure & Compute

🟢 💬 How much work in progress can a workshop submission be [R] — score 38 Sources: reddit/r/MachineLearning

Hi, let's suppose I am working on an algorithm that uses principles x to solve problems A and B. I already implemented a very basic algorithm that used principle "x mini" to just solve problem A, ran experiments, but have not yet implemented the full one to solve A and B. I must say the algorithm to

Other Signals

🟢 💬 Occamy-1.0 by Accio Lab — score 38 Sources: reddit/r/LocalLLaMA

🟢 💬 Scott Aaronson says that labs, "having been burned by the hostile response to the Navier-Stokes proof, are now sitting on solutions to some very major problems until they figure out a better way to handle things" — score 38 Sources: reddit/r/singularity

🟢 💬 "Astra Extra High" is very Extra High today — score 38 Sources: reddit/r/OpenAI

🟢 💬 Qwen3.8-27B-NVFP4 1M context. So far so good. — score 33 Sources: reddit/r/LocalLLaMA

I am a beginner, Took a while to get started, get everything right. This setup is native not container. Still not sure if I did this right, or if I can tune this more. Environment=HF_HUB_OFFLINE=1 Environment=VLLM_LOGGING_LEVEL=INFO Environment=VLLM_ALLOW_LONG_MAX_MODEL_LEN=1 Environment=PATH=/home/

🟢 💬 [D] How do you get preprocessed dataset of a paper [D] — score 30 Sources: reddit/r/MachineLearning

Hi all, I'm trying to reproduce a paper where the reported dataset statistics in Table 1 don't match what I get from the public raw data, even after implementing the preprocessing exactly as described. I've tried all reasonable interpretations of the filtering described in the paper and the closest

Omitted 1 additional other signals items from the main section; see raw data and source-specific sections below.

📊 Cross-Source Signals

Items that appeared on 3+ sources today:

RepoDescriptionStars TodayLanguage
MG1937/ASCASC is a super FAST Android decompiler front-end designed for Agents/Mobile Researchers.122python
CodeWithCJ/SparkyFitnessSparkyFitness: Built for Families. Powered by AI. Track food, fitness, water, and health — together.55typescript
Jeffallan/claude-skills67 Specialized Skills for Full-Stack Developers. Transform Claude Code into your expert pair programmer.22python
RightNow-AI/openfangOpen-source Agent Operating System5rust

📄 New Papers

TitleCategoryHotnessLink
LynnReal-Omni: Native multi-modal Video Generation for Agentic Visual Workflowsresearch_paper44Open
How Lossless Is Lossless Speculative Decoding? The Role of Numerical Precision in Orthrusresearch_paper30Open
Learning to Solve Hard Problems in RL for LLMs by Never Giving Upresearch_paper17Open
Root-Cause Attribution Is a Search Problem: Continual Search for Long-Horizon Agent Failuresresearch_paper3Open
Thought without systematicity? Evaluating reasoning models on rule induction tasksresearch_paper3Open
E2A-Bench: Benchmarking Evidence-to-Action Reliability in Financial Chart Reasoningresearch_paper2Open
ModaLens: Measuring Image Sensitivity in Report-Conditioned Medical VLMsresearch_paper2Open
ZGCM-1: A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Searchcs.AI0Open
Converge Then Diversify: Decoupling Convergence and Diversity in Multi-Objective Bayesian Optimisationcs.AI0Open
Generalized Agent Iteration: One Formal Framework for Iterative Policy Improvement and Recursive Self-Improvementcs.AI0Open
Vibe Patenting: Evaluating LLM Judges for Professional Patent-Drafting Agentscs.AI0Open
Toward Self-Adaptive Physical AI: Can LLM Agents Manage Long-Horizon Physical Tasks?cs.AI0Open
LabAgent: Customize Any Research Hubs for Scientific Discoveries Using AI Agentscs.AI0Open
TimeThink: Eliciting Compositional Reasoning in Timeseries Large Language Modelscs.AI0Open
Governing at Machine Speed: An Adaptive Intelligence Architecture for Real-Time AI Policy Enforcementcs.AI0Open

🐦 Twitter/X Highlights

AccountTweet Summary
samaThere are two ways AI progress could go very badly and that we must avoid. First, we could lose control of the future to AI. This is unacceptable; we are unapologetically on Team Humanity, and AI must always serve people. To ensure that, we need ways to ensure that alignment and safety techniques st Post
samaThe world deserves confidence that American companies developing increasingly capable AI will act responsibly, especially as the trajectory of progress has steepened. Every frontier lab must deliver on this, and there is no reason any of us should come to work if we cannot. We welcome a federal fram Post

Newsletter

Repeated From Recent Briefings