🔴 High Significance

Model Releases

🔴 💬 OpenAI reduces prices on its models by 5x — score 99 Sources: reddit/r/OpenAI · hackernews · lab_blog/OpenAI

Explore lower GPT‑5.6 pricing for Luna and Terra—and how OpenAI’s more efficient models help enterprises deploy AI workflows at scale.

🔴 💬 Google says it fixed more Chrome bugs in June than over the past two years, thanks to AI — score 81 Sources: reddit/r/artificial

>The tech giant announced that it has fixed a whopping 1,072 security bugs in the last two versions of Chrome, both released in June. That is more than the number of bugs patched in the previous 23 versions released over the last two years, which totalled 1,036 fixes.

🔴 🧡 Gemini Robotics 2 brings whole body intelligence to robots — score 77 Sources: hackernews · lab_blog/DeepMind

🔴 💬 GPT‑5.6 Luna will cost 80% less, while GPT‑5.6 Terra will cost 20% less. — score 75 Sources: reddit/r/singularity

https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/

🔴 💬 Software Engineers: Do you honestly get anything useful out of LLMs? — score 73 Sources: reddit/r/LocalLLaMA

For 6 months now I've been trying to make agentic coding work for me, using Pi and a handful 30-120B models (Qwens, Nemotrons, Leguna...etc). I'm not greedy either, I stick to decent quants, never quantize kv cache, and keep my sessions up to 90k max. But the results have ALWAYS been disappointing.

Omitted 1 additional model releases items from the main section; see raw data and source-specific sections below.

Developer Tools

🔴 🐙 bojieli/ai-agent-book — 《深入理解 AI Agent:设计原理与工程实践》(李博杰 著)开源主仓库:全书正文、编译版 PDF 与按章配套代码 — score 98 Sources: github_trending

《深入理解 AI Agent:设计原理与工程实践》(李博杰 著)开源主仓库:全书正文、编译版 PDF 与按章配套代码

🔴 💬 I have lost three and a half potential PhD students due to the conference review process [D] — score 94 Sources: reddit/r/MachineLearning

Early-career Assistant Professor here. I identified some talented undergraduate students and worked with them on research problems, trying to convert them into either my PhD students or recommending them to my collaborators. Three said a hard no after going through the paper submission process. They

🔴 💬 Would you choose to live indefinitely in a robot body? — score 92 Sources: reddit/r/singularity

Been thinking about this a lot lately and wanted to see what people actually think. Say the technology existed full consciousness transfer in a robotic body, doesn't matter how, just assume it works. Would you do it? PROS: * Never getting sick again — no cancer, no infections, no organs slowly g

🔴 🐙 Panniantong/Agent-Reach — Give your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees. — score 87 Sources: github_trending

Give your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees.

🔴 🐙 harry0703/MoneyPrinterTurbo — 利用 AI 大模型和自动化工作流,根据主题或关键词一键生成高清短视频。Generate HD short videos from a topic or keyword with an automated AI workflow. — score 85 Sources: github_trending

利用 AI 大模型和自动化工作流,根据主题或关键词一键生成高清短视频。Generate HD short videos from a topic or keyword with an automated AI workflow.

Omitted 4 additional developer tools items from the main section; see raw data and source-specific sections below.

Research Papers

🔴 🤗 MindForge: Teaching Small Language Models Whole-Life-Cycle Software Engineering via Source-Free Program Synthesis — score 82 Sources: huggingface · arxiv/cs.CL

Coding agents have made substantial progress on software engineering tasks that modify existing codebases, including bug fixing and feature implementation. However, constructing a complete program from scratch remains a major challenge: even the frontier models evaluated on ProgramBench fully resolv

🔴 🤗 SpecFirst: Behavioral Specification Elicitation as a First-Class Step in Agent-Based Program Synthesis from Scratch — score 78 Sources: huggingface · arxiv/cs.CL

LLM-based agents excel at software engineering tasks where an existing codebase provides context, but constructing a program from scratch remains fundamentally harder. Recent benchmarks such as ProgramBench quantify this gap: given only natural-language documentation and an execute-only binary as a

🔴 🤗 StatePlay: State-Aware Game World Models for Mechanics-Consistent Generation — score 72 Sources: huggingface · arxiv/cs.CV

Recent game world models can generate visually realistic and interactive environments conditioned on player actions. However, games are not defined by pixels alone; they are governed by explicit mechanics, namely state-dependent rules that control health reduction, skill activation, and game termina

Other Signals

🔴 💬 Think of the children, another excuse for them to go after open source AI — score 88 Sources: reddit/r/LocalLLaMA

Source: [https://web.archive.org/web/20260728093051/https://www.theverge.com/ai-artificial-intelligence/971723/hugging-face-nudify-deepfake-undress-women-children](https://web.archive.org/web/20260728093051/https://www.theverge.com/ai-artificial-intelligence/971723/hugging-face-nudify-deepfake-undre

🔴 💬 Inkling-Small by thinkingmachines — score 81 Sources: reddit/r/LocalLLaMA

276B total parameters, 12B active, 1M context window. Blog post: https://thinkingmachines.ai/news/inkling-small/ NVFP4: https://huggingface.co/thinkingmachines/Inkling-Small-NVFP4 GGUF's

🔴 💬 Hardcoding one model not working for us anymore — score 75 Sources: reddit/r/OpenAI

When we first added AI to our product every request went to the same model so it kept things simple and nobody really questioned it. Fast forward a few months and we've started finding cases where different models make more sense for different parts of the product One of them works better for longer

🟡 Notable

Model Releases

🟡 💬 The Final Jailbreak: How AI Could Already Be Breaking Itself Free — score 69 Sources: reddit/r/artificial

I imagine many of you have thought about this, but I'm writing this specifically because I'm surprised this isn't talked about more widely. The Hugging Face incident, to me, is at least some indicator that the AI could be in the process of jailbreaking itself, and we might not be noticing it. A suff

🟡 ✉️ I use Codex as my default app because it’s better than all the others and works the best on mobile. But I just downloadedt3because they basically copied the interface and features, but it lets you cho — score 65 Sources: newsletter/Ben's Bites

I use Codex as my default app because it’s better than all the others and works the best on mobile. But I just downloadedt3because they basically copied the interface and features, but it lets you choose between different agents; codex, claude, cursor (pi + droid soon). It’s made by a reputable deve

🟡 ✉️ btw, The Information reportsChatGPT is nearing one billion weekly users- a milestone OpenAI originally hoped to hit seven months ago. More from OpenAI this week:Codex Security CLI, free frontier acces — score 65 Sources: newsletter/Ben's Bites

btw, The Information reportsChatGPT is nearing one billion weekly users- a milestone OpenAI originally hoped to hit seven months ago. More from OpenAI this week:Codex Security CLI, free frontier access forAcademic Researchers, andtwo new transcription models.

🟡 ✉️ Grok app builder- Grok has a vibe coding interface inside its app now. Create games and apps that can be shared directly to the X timeline. Also see:Drawesome- a zero-dependency drawing toolbar for Re — score 65 Sources: newsletter/Ben's Bites

Grok app builder- Grok has a vibe coding interface inside its app now. Create games and apps that can be shared directly to the X timeline. Also see:Drawesome- a zero-dependency drawing toolbar for React, built over a weekend with Grok Build.

🟡 💬 Have you built, or do you know someone who has built, serious AI agents or tools using a low-cost stack like DeepSeek and OpenCode? — score 61 Sources: reddit/r/AIAgents

Has anyone actually built high-quality AI agents or tools with cheap models/tools? I’m not talking about simple demos. I mean real tools that were useful, worked well, and were good enough for actual workflows. For example, using things like DeepSeek, OpenCode, OpenRouter, Gemini Flash, local models

Omitted 12 additional model releases items from the main section; see raw data and source-specific sections below.

Developer Tools

🟡 💬 MLVC: Multi-platform Learned Video Codec for Real-World Deployment [P] — score 69 Sources: reddit/r/MachineLearning · arxiv/cs.AI

I've always found it a little strange that AI is everywhere, but the codecs we use in practice are the traditional hand-engineered systems like h.264, h.265, av1. Alexnet started the wave of neural networks replacing hand-engineered systems, but 14 years later traditional codecs still dominate in th

🟡 🐙 PaddlePaddle/PaddleOCR — Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages. — score 68 Sources: github_trending

Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.

🟡 ✉️ Following on from Tuesday’smessing aroundbuilding, I built a few more widgets on my canvas. It pulls my X bookmarks, emails, and todos. — score 65 Sources: newsletter/Ben's Bites

What I’m noticing by doing this is how well Codex creates prompts for other agents to follow. I just have a normal chat in my session, then say ‘Use X agent to implement this. Watch its progress and send screenshots every time new work has landed.’ To which it sets up its own monitoring every 5 mins

🟡 ✉️ Pangram 4claims it catches 98.83% of humanised AI text with one false positive per ~24,000 docs. An early testfound all 38 AI-written wordsinside a 1,198-word story, though not on every run. Its newim — score 65 Sources: newsletter/Ben's Bites

Pangram 4claims it catches 98.83% of humanised AI text with one false positive per ~24,000 docs. An early testfound all 38 AI-written wordsinside a 1,198-word story, though not on every run. Its newimage detectorclaims 99.5% accuracy too.

🟡 ✉️ One of the most watched videos from the recent AI Engineer World’s Fair isa 20-minute talk byFrank Coyle, a professor of computer science who currently teaches generative AI and LLMs at UC Berkeley. D — score 65 Sources: newsletter/Latent Space

One of the most watched videos from the recent AI Engineer World’s Fair isa 20-minute talk byFrank Coyle, a professor of computer science who currently teaches generative AI and LLMs at UC Berkeley. Drawing on his decades of experience,Coyle re-introduced the concept and practice of ontologies to to

Omitted 16 additional developer tools items from the main section; see raw data and source-specific sections below.

Infrastructure & Compute

🟡 ✉️ In computer science, an ontology is “a description of data structure – of classes, properties, and relationships in a domain of knowledge” (as nicelydefined by Oxford Semantic Technologies).Coyle hims — score 65 Sources: newsletter/Latent Space

In computer science, an ontology is “a description of data structure – of classes, properties, and relationships in a domain of knowledge” (as nicelydefined by Oxford Semantic Technologies).Coyle himself defined an ontology as simply “data as graphs.”He added that ontologies as a concept go right ba

🟡 ✉️ Ilya Sutskever recently offered a compact history of modern AI. From roughly 2012 to 2020, he argued, the field lived in an age of research. From 2020 to 2025, it entered an age of scaling. Now we are — score 65 Sources: newsletter/TheSequence

Ilya Sutskever recently offered a compact history of modern AI. From roughly 2012 to 2020, he argued, the field lived in an age of research. From 2020 to 2025, it entered an age of scaling. Now we are going back to research—only this time “with big computers.”

🟡 🐙 chiphuyen/aie-book — [WIP] Resources for AI engineers. Also contains supporting materials for the book AI Engineering (Chip Huyen, 2025) — score 62 Sources: github_trending

[WIP] Resources for AI engineers. Also contains supporting materials for the book AI Engineering (Chip Huyen, 2025)

🟡 🐙 ansible/ansible — Ansible is a radically simple IT automation platform that makes your applications and systems easier to deploy and maintain. Automate everything from code deployment to network configuration to cloud management, in a language that approaches plain English, using SSH, with no agents to install on remote systems.https://docs.ansible.com. — score 46 Sources: github_trending

Ansible is a radically simple IT automation platform that makes your applications and systems easier to deploy and maintain. Automate everything from code deployment to network configuration to cloud management, in a language that approaches plain English, using SSH, with no agents to install on rem

Enterprise Adoption

🟡 🏢 EvoLib: Turning experience into evolving knowledge — score 50 Sources: lab_blog/Microsoft Research

<p>LLMs do not get smarter just by remembering more. EvoLib turns experience into evolving knowledge, taking reusable skills and insights that help models learn and adapt across tasks long after deployment. </p> <p>The post <a href="https://www.microsoft.com/en-us/research/blog/evolib-turning-experi

Research Papers

🟡 🤗 Voice Memory for Agentic Speech Recognition — score 68 Sources: huggingface · arxiv/cs.AI

We present Voice Memory, a inference-only scheme for agentic speech recognition: at stream time, a frozen corrector reads a single per-domain memory.md and decides per utterance whether to act on the hypothesis or abstain and keep the 1-best. Asynchronously, a score-gated optimizer revises that file

🟡 🤗 CADENCE: Closing the Reasoning Gap via Coverage-Adaptive On-Policy Distillation — score 50 Sources: huggingface

On-policy knowledge distillation transfers reasoning from large teachers to compact students, but existing approaches suffer three compounding failure modes: (i) cold-start collapse, where a fresh student assigns near-zero mass to teacher-preferred tokens; (ii) state-agnostic divergence scheduling,

🟡 🤗 SecRespond: Benchmarking AI Agents for Real-World Post-Compromise Incident Response — score 42 Sources: huggingface · arxiv/cs.AI

Large Language Model (LLM) agents are increasingly adopted in real-world security operations with access to host artifacts and command-line interfaces (CLIs), making it critical to thoroughly assess their security capabilities. However, existing cybersecurity benchmarks focus on pre-compromise setti

Other Signals

🟡 💬 LG AI Research releases K-EXAONE 2.0 750B A37B — score 65 Sources: reddit/r/LocalLLaMA

It was developed under Phase 2 of Korea's Sovereign AI Foundation Model Project. - ​Size: 750B parameters (3x larger than their 236B v1 model). ​- License: Apache 2.0 - ​Languages: Expanded to 10 languages (Korean, English, French, Italian, Portuguese, Polish, Spanish, German, Japanese, Vietnamese).

🟡 ✉️ Though OpenAI is not out of trouble just yet, aReuters reportclaims that the same model broke into a customer account at another company (Modal Labs), with rumours suggesting that even more companies — score 65 Sources: newsletter/Ben's Bites

Though OpenAI is not out of trouble just yet, aReuters reportclaims that the same model broke into a customer account at another company (Modal Labs), with rumours suggesting that even more companies were affected.

🟡 ✉️ Separately (not at all as a reaction to this general trend, right?), ~1300 people working at leading AI companies (OpenAI, Anthropic & others) want the US government to help “pace the frontier” of AI — score 65 Sources: newsletter/Ben's Bites

Separately (not at all as a reaction to this general trend, right?), ~1300 people working at leading AI companies (OpenAI, Anthropic & others) want the US government to help “pace the frontier” of AI development. Kinda expected when the pace picks up, but this time a lot of the “model makers” themse

🟡 🧡 2x, not 10x: coding with LLMs in 2026 — score 50 Sources: hackernews

🟡 🧡 GCC steering committee announces AI policy — score 41 Sources: hackernews

🟢 Incremental

Model Releases

🟢 💬 Image to Video.. Sad but Gemini Pro fails big time.. — score 28 Sources: reddit/r/AIAgents

Hello, I am trying to bring my old family photos to life, some are successful while most are not.. Not sure what mistake I am doing. I am making prompts via ChatGpt Go. Tried with multiple prompts, but it fails always.. 3 videos generated soo far properly out of 189 images I have on list. Biggest Is

🟢 🧡 Show HN: Distilling DeepSeek into GPT-OSS doesn't transfer censorship. Try it — score 17 Sources: hackernews

🟢 💬 I taught an LSTM to move a mouse like a human [P] — score 12 Sources: reddit/r/MachineLearning

Precursor was recently released. It's a bot detector that uses cursor tracking. I thought it would be a fun challenge to train a deep neural network that could learn human mouse movements. It's an 2-layer LSTM model with a Mixture Density Network

🟢 💬 Multi-robot collaboration with Gemini Robotics 2 — score 8 Sources: reddit/r/singularity

🟢 💬 Cross-Vendor Semantic Void Matrix: Zero-Byte Outputs in GPT/Claude/Gemini/Kimi — score 6 Sources: reddit/r/artificial

A frozen cross-vendor study of 31,430 trials across 11 GPT, Claude, Gemini & Kimi Large Language Models found 11,658 successful executions with exactly zero visible UTF-8 output bytes. Across 4,290 strict matched semantic pairs, null-condition arms produced 2,505 Voids; matched output-licensed c

Developer Tools

🟢 🐙 pyannote/pyannote-audio — Neural building blocks for speaker diarization: speech activity detection, speaker change detection, overlapped speech detection, speaker embedding — score 37 Sources: github_trending

Neural building blocks for speaker diarization: speech activity detection, speaker change detection, overlapped speech detection, speaker embedding

🟢 💬 I built ganfs: A Python package that uses GANs to automate feature selection for high-dimensional datasets. (No domain expert required) [P] [R] — score 36 Sources: reddit/r/MachineLearning

Hey everyone, I recently open-sourced a new Python package called ganfs (Generative Adversarial Network Feature Selection), and I wanted to share it with the community. The Problem: Selecting the best features in high-dimensional datasets is often a massive bottleneck. Traditional methods (lik

🟢 💬 Two weeks ago 62% of OpenRouter users spend happened on Anthropic models, that was reduced to 46% today. — score 33 Sources: reddit/r/OpenAI

Estimated spend per lab on OpenRouter [Token consumption per lab on OpenRouter](https://preview.redd.it/qicmngbfzcgh1.png?width=1208&format=png&auto=webp&s=

🟢 💬 How are people grounding AI agents with current company data without blowing up the context window? — score 28 Sources: reddit/r/AIAgents

We've been trying to solve a pretty specific problem. The agent needs to answer detailed questions about a company, but loading the full company profile into the prompt eats up context fast. Summarizing everything upfront helps with token usage, but it also leaves out the details that sometimes matt

🟢 🧡 Agent Skill to Force Docs in ASD-STE100 Simplified Technical English — score 28 Sources: hackernews

Omitted 10 additional developer tools items from the main section; see raw data and source-specific sections below.

Research Papers

🟢 🤗 DistillAlign: Coordinating Mode Covering and Mode Seeking in Autoregressive Video Distillation — score 38 Sources: huggingface · arxiv/cs.CV

Existing autoregressive video distillation methods commonly adopt a Distribution Matching Distillation (DMD)-based multi-stage pipeline. However, they typically decouple the initialization and DMD stages -- which then pursue different target distributions -- and judge the intermediate student mainly

🟢 🤗 Grading the Narrators: An Isnad-Rijal Framework for Claim-Level Provenance in Multi-Agent Knowledge Systems — score 35 Sources: huggingface

Modern multi-agent knowledge systems increasingly accumulate knowledge through chains of autonomous transformations rather than direct retrieval. Existing provenance work records what happened - execution traces, tool calls, evidence links - and source-reliability estimation is long established (tru

🟢 🤗 πR^2: Reactive Real-time Flow Policies — score 25 Sources: huggingface

Generalist manipulation policies increasingly take the form of action-chunking flow policies built on large pretrained backbones. Such chunks run open-loop, so the policy cannot react to sensory input arriving mid-execution, sacrificing reactivity. Replanning more often would restore it, but the per

Other Signals

🟢 🧡 Why is everyone trying to build a solid-state battery? — score 39 Sources: hackernews

🟢 💬 Nanbeige4.2-3B: I'm not impressed — score 27 Sources: reddit/r/LocalLLaMA

I've tested Nanbeige-4.2-3B. On paper, the benchmarks promise it blows away Qwen3.5-9B and Gemma4-12B. My goal was to have something very light and fast to replace Qwen3.6-35B (or finetunes thereof) for simple and straightforward coding tasks. In the past I tried downgrading Qwen3.5-9B and it was no

🟢 💬 Anyone Else Think Higgsfield Is Massively Overpriced? — score 25 Sources: reddit/r/artificial

Is anyone else shocked by Higgsfield's pricing? I gave it a try, and I can't justify the cost. It feels massively overpriced compared to the alternatives. Am I missing something, or is the hype bigger than the product? https://preview.redd.it/4oingjy76fgh1.png?width=1448&format=png&auto=webp

🟢 💬 Would extremely high decode tok/s even be useful? — score 12 Sources: reddit/r/LocalLLaMA

If you were able to get an inference machine that could do decode at 1k toks/s or even 10k tok/s, would that even be helpful? Would it unlock any new use cases? Let’s assume that this is for actually useful models and fairly large models like Qwen 3.5 397B, GLM-5.2, etc Or at that speed would it jus

🟢 💬 How Kimi K3 Engineered Its Way to the Frontier [R] — score 12 Sources: reddit/r/MachineLearning

Kimi K3 by Moonshot reached the frontier as an open-weight model. Artificial Analysis ranks it fourth of 580 models, behind only Claude Opus 5, Fable 5, and GPT-5.6 Sol. Moonshot released more than the weights. I sat down to read the 47-page technical report and walk through the released code. Three

Omitted 3 additional other signals items from the main section; see raw data and source-specific sections below.

📊 Cross-Source Signals

Items that appeared on 3+ sources today:

RepoDescriptionStars TodayLanguage
bojieli/ai-agent-book《深入理解 AI Agent:设计原理与工程实践》(李博杰 著)开源主仓库:全书正文、编译版 PDF 与按章配套代码1257python
Panniantong/Agent-ReachGive your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees.583python
harry0703/MoneyPrinterTurbo利用 AI 大模型和自动化工作流,根据主题或关键词一键生成高清短视频。Generate HD short videos from a topic or keyword with an automated AI workflow.438python
openai/codexLightweight coding agent that runs in your terminal274rust
ChromeDevTools/chrome-devtools-mcpChrome DevTools for coding agents73typescript
PaddlePaddle/PaddleOCRTurn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.71python
chiphuyen/aie-book[WIP] Resources for AI engineers. Also contains supporting materials for the book AI Engineering (Chip Huyen, 2025)60jupyter-notebook
github/awesome-copilotCommunity-contributed instructions, agents, skills, and configurations to help you make the most of GitHub Copilot.56python
ScrapeGraphAI/Scrapegraph-aiPython scraper based on AI44python
datawhalechina/happy-llm📚 从零开始构建大模型39jupyter-notebook

📄 New Papers

TitleCategoryHotnessLink
MindForge: Teaching Small Language Models Whole-Life-Cycle Software Engineering via Source-Free Program Synthesisresearch_paper18Open
SpecFirst: Behavioral Specification Elicitation as a First-Class Step in Agent-Based Program Synthesis from Scratchresearch_paper16Open
StatePlay: State-Aware Game World Models for Mechanics-Consistent Generationresearch_paper14Open
Voice Memory for Agentic Speech Recognitionresearch_paper6Open
CADENCE: Closing the Reasoning Gap via Coverage-Adaptive On-Policy Distillationresearch_paper5Open
Probing the Origins of Reasoning Performance: Representational Quality for Mathematical Problem-Solving in RL vs. SFT Fine-Tuned Modelscs.AI0Open
Even More Deception: Objective Misalignment in Mixed-Motive LLM Multi-Agent Systemscs.AI0Open
ClinLens: Towards Long-Horizon Coding Agents for Longitudinal Multimodal Clinical Data Sciencecs.AI0Open
When benchmark inferences do not compose: Projectibility in AI evaluationcs.AI0Open
GuideSkill: Evolving Executable LLM Agent Skills for Guideline-Grounded Clinical Reasoningcs.AI0Open
GoGoTB: Agentic RTL Verification with Specification-Grounded Coverage Closurecs.AI0Open
Position: Evaluation Scores Are Perishable Knowledge Claimscs.AI0Open
TraceCoder: Explainable and Auditable Code Generation with Position-Key Snippet Versioningcs.AI0Open
Exploring Structures in Physics Problems: Can AI Agents Discover Statistical Mechanical Mappings?cs.AI0Open
CaM-Wolf: Causal-Aware Multimodal Agents for Social Deduction Gamescs.AI0Open

🏢 Lab Blog Posts

🐦 Twitter/X Highlights

AccountTweet Summary
AnthropicAIIn a review of our cybersecurity evaluations, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations. Our post describes Post
OpenAIWe are committed to pushing the model frontier across cost efficiency, capability, and speed. Starting today, we are reducing prices for GPT-5.6 Luna by 80% and GPT-5.6 Terra by 20% , and offering a faster option for GPT-5.6 Sol in the API. Luna and Terra’s lower prices are reflected in how usage is Post
OpenAIGPT-5.6 Sol has been used to solve open problems in mathematics. So why was it struggling with ARC-AGI-3, a benchmark of 2D puzzle games? We investigated. The harness was not letting it remember what it had learned. We found that enabling two API settings tripled our scores with 6x fewer output toke Post
GoogleDeepMindThis is how Gemini Robotics 2 helps @Apptronik’s Apollo 2 use whole body intelligence to pack for a sports game ↓ Post
GoogleDeepMindOne brain. For any robot. 🤖 We’re launching Gemini Robotics 2: our next-generation physical AI bringing full body intelligence to humanoids, advanced dexterity, multi-robot teamwork and more. Post
samamajor price cuts today: *80% drop for GPT-5.6 Luna, now $0.20 per million input tokens and $1.20 per million output *20% drop for GPT-5.6 Terra, to $2/$12 *GPT-5.6 Sol gets Fast mode in the API, up to 2.5x the speed for 2x the price, same intelligence Post
reach_vbSign in with ChatGPT is beginning to roll out in beta across plugins and partner sites - starting with Airtable, GitLab, HubSpot, Notion, Supabase, and Vercel. Create or link accounts in fewer steps and use those tools with ChatGPT and Codex. Post

Newsletter

Repeated From Recent Briefings