๐ด High Significance
Model Releases
๐ด ๐ฌ I vibe code apps for a living. Here are my three tips. โ score 94
Sources: reddit/r/AIAgents
I am by no means the best in the world at vibe coding, definitely not. But I run an agency that builds ready-made internal software for large companies (mainly Fortune 500 companies). So I have spent thousands of hours working with agents, making mistakes and (hopefully) learning from my mistakes. A
๐ด ๐งก Identity verification on Claude โ score 88
Sources: hackernews
๐ด ๐ฌ Qwen is never going to open source Qwen 3.7, aren't they? โ score 82
Sources: reddit/r/LocalLLaMA
Well, this was predictable. After Qwen fired Junyang Lin, the next models are no longer open source. Ignoring the small models for a minute, theyโve fully locked down all the big models. No Deepseek/GLM competitor. And all the rumors on chinese weibo now say that the small model Qwen team is gone, a
๐ด ๐ข Predicting model behavior before release by simulating deployment โ score 75
Sources: lab_blog/OpenAI ยท newsletter/tldr
OpenAI introduces Deployment Simulation, a method to predict AI model behavior before deployment using real conversation data to improve safety and evaluation accuracy.
Developer Tools
๐ด ๐ chopratejas/headroom โ Compress tool outputs, logs, files, and RAG chunks before they reach the LLM. 60-95% fewer tokens, same answers. Library, proxy, MCP server. โ score 99
Sources: github_trending
Compress tool outputs, logs, files, and RAG chunks before they reach the LLM. 60-95% fewer tokens, same answers. Library, proxy, MCP server.
๐ด ๐ calesthio/OpenMontage โ World's first open-source, agentic video production system. 12 pipelines, 52 tools, 500+ agent skills. Turn your AI coding assistant into a full video production studio. โ score 94
Sources: github_trending
World's first open-source, agentic video production system. 12 pipelines, 52 tools, 500+ agent skills. Turn your AI coding assistant into a full video production studio.
๐ด ๐ NousResearch/hermes-agent โ The agent that grows with you โ score 92
Sources: github_trending
The agent that grows with you
๐ด ๐ jamiepine/voicebox โ The open-source AI voice studio. Clone, dictate, create. โ score 89
Sources: github_trending
The open-source AI voice studio. Clone, dictate, create.
๐ด ๐ ZhuLinsen/daily_stock_analysis โ LLM ้ฉฑๅจ็ๅคๅธๅบ่ก็ฅจๆบ่ฝๅๆ็ณป็ป๏ผๅคๆบ่กๆ
ใๅฎๆถๆฐ้ปใๅณ็ญ็ๆฟไธ่ชๅจๆจ้๏ผๆฏๆ้ถๆๆฌๅฎๆถ่ฟ่กใ LLM-powered multi-market stock analysis system with multi-source market data, real-time news, decision dashboard, automated notifications, and cost-free scheduled runs. โ score 87
Sources: github_trending
LLM ้ฉฑๅจ็ๅคๅธๅบ่ก็ฅจๆบ่ฝๅๆ็ณป็ป๏ผๅคๆบ่กๆ ใๅฎๆถๆฐ้ปใๅณ็ญ็ๆฟไธ่ชๅจๆจ้๏ผๆฏๆ้ถๆๆฌๅฎๆถ่ฟ่กใ LLM-powered multi-market stock analysis system with multi-source market data, real-time news, decision dashboard, automated notifications, and cost-free scheduled runs.
Omitted 5 additional developer tools items from the main section; see raw data and source-specific sections below.
Infrastructure & Compute
๐ด ๐ฌ Hi Reddit, I posted my Build Your Own LLM workshop to Youtube teaching ML, LLM and math intuition [P] โ score 94
Sources: reddit/r/MachineLearning
Hi internet friends, I recorded a workshop about building your own LLM without any math / ML prerequisites. It covers everything from machine learning fundamentals, deep neural networks, transformer architecture, and pre/post-training. The only prerequisite is being comfortable with learning through
Research Papers
๐ด ๐ค Context-Aware RL for Agentic and Multimodal LLMs โ score 95
Sources: huggingface
Large language models (LLMs) often fail when answering requires identifying a small but decisive piece of evidence within a long or complex context, such as a single line in a tool trace or a subtle detail in an image. We propose ContextRL, a context-aware reinforcement learning (RL) method that imp
๐ด ๐ค The Data Manifold under the Microscope โ score 85
Sources: huggingface
A significant gap exists between theory and practice in deep learning. Generalization and approximation error bounds are often derived for simplified models or are too loose to be informative. Many rely on the manifold hypothesis and on geometric regularity such as intrinsic dimension, curvature, an
๐ด ๐ค Freeing the Law with LOCUS: A Local Ordinance Corpus for the United States โ score 70
Sources: huggingface
Progress in legal AI increasingly depends on access to authoritative legal text at scale. Yet one of the most consequential layers of American law remains largely absent from existing machine-readable corpora: local ordinances. Local codes govern zoning, housing, business licensing, public health, n
๐ด ๐ค Rethinking Shrinkage Bias in LLM FP4 Pretraining: Geometric Origin, Systemic Impact, and UFP4 Recipe โ score 70
Sources: huggingface
FP4 training promises substantial reductions in memory and computation cost for LLM pretraining, yet current FP4 hardware paths and recipes, including NVIDIA Blackwell/Rubin-class systems and AMD MI350-series GPUs, remain centered on E2M1 data elements. In this study, we identify a fundamental limit
Other Signals
๐ด ๐ฌ Tokenomics โ score 96
Sources: reddit/r/LocalLLaMA
๐ด ๐ฌ Vercel CEO: "Almost shocked" by how good GLM-5.2 is at coding โ score 89
Sources: reddit/r/LocalLLaMA
Guillermo Rauch (Vercel CEO) says he is "genuinely impressed, almost shocked" by GLM-5.2's coding performance. What has your experience with GLM-5.2 been so far? Source: X post
๐ด ๐ฌ Would you let an ML PhD student graduate without a top-tier paper? [D] โ score 81
Sources: reddit/r/MachineLearning
Suppose youโre a PhD advisor in machine learning. Your student has been in the program for 4 years, has done solid work, and has a coherent thesis direction but they havenโt published in an A*ML venue or top journal. No NeurIPS/ICML/ICLR/CVPR/etc., and no equivalent top venue in their subfield eith
๐ด ๐ฌ Gemma 4 QAT seems to respond significantly better to KV cache quantization โ score 75
Sources: reddit/r/LocalLLaMA
Results from KL Divergence on wikitext with 16k context I know some users, including myself, were disappointed with Gemma 4's sensitivity to KV cache quantization. Seems like Q8_0 on QAT models might be back on the menu. KLD measures divergence from the base (in this case, full 16-bit KV cache). 99
๐ก Notable
Model Releases
๐ก ๐ฌ 8-16 MI50s Minimax M3 @19 tps TG (peak) โ score 68
Sources: reddit/r/LocalLLaMA
TL;DR Speeds are not too ugly for this old 2018 hardware but imo, not very usable for agentic coding (if you compare with qwen3.6 27B on 8 MI50 @ 50 tps TG 800 tps PP). More concerning is that the reasoning output is very very long and still didnโt check about the quality of code outputโฆ As said
๐ก โ๏ธ The AI scaling debate always focuses on the question of โhow do we get more GPUs?โ but the better question may be:how do we make the most of ones we already have. โ score 65
Sources: newsletter/Latent Space
For context, older frontier-scale training runs were already much higher than 10%. GPT-3 was around21% MFU. Gopher was around32%. Megatron-Turing NLG was around30%. PaLM reached around46%. And our guest Anjney says best-in-class MFU today is closer to60โ70%.
๐ก โ๏ธ ANTHROPIC SHIPS MAJOR CLAUDE DESIGN OVERHAUL WITH DESIGN SYSTEM IMPORTS, CODE ROUND-TRIPS, AND A FIX FOR ITS TOKEN-BURNING PROBLEM (13 MINUTE READ) โ score 65
Sources: newsletter/tldr
๐ก โ๏ธ VERCEL LAUNCHES ENTERPRISE CONTROLS FOR AGENTIC AI INFRASTRUCTURE (3 MINUTE READ) โ score 65
Sources: newsletter/tldr
๐ก โ๏ธ INSIDE CLAUDE MANAGED AGENTS: REVERSE ENGINEERING THE SECURITY BOUNDARIES OF ANTHROPIC'S HOSTED AGENT RUNTIME (12 MINUTE READ) โ score 65
Sources: newsletter/tldr
Omitted 25 additional model releases items from the main section; see raw data and source-specific sections below.
Developer Tools
๐ก ๐ฌ 30 Core Agentic Engineering Concepts, Explained Simply โ score 69
Sources: reddit/r/AIAgents
๐ก ๐ Alishahryar1/free-claude-code โ Use claude code and codex for free in the terminal, VSCode extension, and discord like OpenClaw (voice supported) โ score 68
Sources: github_trending
Use claude code and codex for free in the terminal, VSCode extension, and discord like OpenClaw (voice supported)
๐ก ๐ withastro/flue โ The sandbox agent framework. โ score 65
Sources: github_trending
The sandbox agent framework.
๐ก โ๏ธ BUILDING AN AI DATABASE FOR AGENTIC GTM OPERATIONS (11 MINUTE READ) โ score 65
Sources: newsletter/tldr
๐ก โ๏ธ FINE-TUNING A CLINICAL AI MODEL TO FRONTIER PARITY (8 MINUTE READ) โ score 65
Sources: newsletter/tldr
Omitted 32 additional developer tools items from the main section; see raw data and source-specific sections below.
Infrastructure & Compute
๐ก โ๏ธ Itโs not necessarily that xAI is uniquely incompetent(itโs clear they have talented folks)but rather the priorities may be flipped in the GPU arms race. โ score 65
Sources: newsletter/Latent Space
๐ก โ๏ธ Dario Amodei and Demis Hassabis reportedly proposed a โU.S.-led AI coalitionโ at G7, with international cooperation on model access, chip exports, and safety risks. โ score 65
Sources: newsletter/rundown-ai
Business & Funding
๐ก โ๏ธ OpenAI reported $3.7B in Q1 cash burn against $5.7B in revenue in 2026, with both tripling from a year earlier, according to documents seen by The Information. โ score 65
Sources: newsletter/rundown-ai
Research Papers
๐ก ๐ค LedgerAgent: Structured State for Policy-Adherent Tool-Calling Agents โ score 50
Sources: huggingface
Policy-adherent tool-calling agents in customer-service domains must maintain task states across turns while calling tools and obeying domain policies. Task states consist of relevant facts, identifiers, constraints, and conditions observed through user interaction and tool calls. In standard agents
๐ก ๐ค LegalHalluLens: Typed Hallucination Auditing and Calibrated Multi-Agent Debate for Trustworthy Legal AI โ score 50
Sources: huggingface
AI systems deployed in legal workflows hallucinate at rates that aggregate metrics report at ~52%, but this average conceals where errors concentrate and in which direction they run, leaving compliance officers without an actionable signal for trustworthy deployment. We present LegalHalluLens, an au
Other Signals
๐ก ๐ฌ A slightly improved DVD-JEPA demo [P] โ score 69
Sources: reddit/r/MachineLearning
Hey! I came across this post, which I found quite neat as a minimal demonstration of JEPA. However, as the comments pointed out, there was some room for improvement. So I added a few things such as environment noise and a fair* comparison to
๐ก โ๏ธ ANTHROPIC EMPLOYEES ACCUSE TRUMP ADMINISTRATION OF TARGETING THEM (11 MINUTE READ) โ score 65
Sources: newsletter/tldr
๐ก โ๏ธ BUILDING TREX: CODE EXECUTION AND ARTIFACT GENERATION FOR AI CODE REVIEW (9 MINUTE READ) โ score 65
Sources: newsletter/tldr
๐ก โ๏ธ THE FOUNDER'S PLAYBOOK: BUILDING AN AI-NATIVE STARTUP (5 MINUTE READ) โ score 65
Sources: newsletter/tldr
๐ก โ๏ธ CAN GZIP BE A LANGUAGE MODEL? (5 MINUTE READ) โ score 65
Sources: newsletter/tldr
Omitted 22 additional other signals items from the main section; see raw data and source-specific sections below.
๐ข Incremental
Model Releases
๐ข ๐งก Show HN: Recall โ fully-local project memory for Claude Code โ score 38
Sources: hackernews
๐ข ๐ฌ Qwen 3.6 27b Abliterated (apostate) โ score 25
Sources: reddit/r/LocalLLaMA
I've been working on a project called Apostate and have finally released my first large model with it on Hugging Face. Qwen 3.6 27B with safety alignment removed down from 92% to 7.6% refusal rate with minimal impact on the model's capabilities (0.120 KL).
๐ข ๐งก Claude: Elevated Error Rates for Opus 4.8, Opus 4.7, Opus 4.6, and Sonnet 4.6 โ score 12
Sources: hackernews
๐ข ๐ฌ I forked ik_llama.cpp and added a "--numa mirror" mode to maximize performance on multi-socket CPU systems. Just sharing and looking for testers! โ score 0
Sources: reddit/r/LocalLLaMA
GitHub: https://github.com/mikechambers84/ik_llama.cpp/tree/numa-mirror Be sure to checkout the
numa-mirrorbranch. Sharing this for anyone else who's trying to use their multi-socket CPU systems for inference. I've been wanti
๐ข ๐ฌ If youโre using AI agents (Claude / Cursor / Copilot)โฆ Youโre probably missing one critical layer: ๐ a safety + cost firewall โ score 0
Sources: reddit/r/AIAgents
โ โ Right now, most setups look like this: AI Agent โ Full access โ Your system โ No guardrails. No validation. No limits. โ Which means your agent can: โ โข run destructive commands โข leak ".env" secrets โข get stuck in API loops โ $$$ โข execute
Developer Tools
๐ข ๐ vercel-labs/agent-browser โ Browser automation CLI for AI agents โ score 39
Sources: github_trending
Browser automation CLI for AI agents
๐ข ๐ฌ Question for people who own profitable agents โ score 38
Sources: reddit/r/AIAgents
Would you guys take an investment directly into your agent if it meant giving up a % of your revenues to the person that invested in it? Or would you bootstrap? I am wondering if this is an effective way for people who have standalone revenue-producing AI agents to get actual funding for compute cos
๐ข ๐ฌ What are y'all using for observability in your agent systems? โ score 38
Sources: reddit/r/AIAgents
a lil bit about me since this is my first post here i'm ajay. i've had a couple exits before so i'm not completely new to startups, but the space me and the team are building in right now is relatively new to us. since everyone's building multi-agent systems these days, i've been curious about the i
๐ข ๐ BuilderIO/agent-native โ A framework for building agent-native applications. โ score 35
Sources: github_trending
A framework for building agent-native applications.
๐ข ๐ FlowiseAI/Flowise โ Build AI Agents, Visually โ score 30
Sources: github_trending
Build AI Agents, Visually
Omitted 4 additional developer tools items from the main section; see raw data and source-specific sections below.
Infrastructure & Compute
๐ข ๐ THUDM/slime โ slime is an LLM post-training framework for RL Scaling. โ score 37
Sources: github_trending
slime is an LLM post-training framework for RL Scaling.
๐ข ๐ EricLBuehler/mistral.rs โ Fast, flexible LLM inference โ score 6
Sources: github_trending
Fast, flexible LLM inference
Research Papers
๐ข ๐ค ReSyn: A Generalized Recursive Regular Expression Synthesis Framework โ score 25
Sources: huggingface
Existing Programming-By-Example (PBE) systems often rely on simplified benchmarks that fail to capture the high structural complexity of real-world regexes, such as deeper nesting and frequent use of union operations. To overcome the resulting performance drop, we propose ReSyn, a synthesizer-agnost
๐ข ๐ค Configurable Clinical Information Extraction with Agentic RAG: What Works, What Breaks, and Why โ score 25
Sources: huggingface
Patient contexts span hundreds of heterogeneous documents and thousands of structured data points, yet the document-level metadata that AI systems need for retrieval and triage is absent or incomplete. Standard retrieval-augmented generation fails on this data, mishandling temporal reasoning, cross-
๐ข ๐ค Duration Aware Scheduling for ASR Serving Under Workload Drift โ score 25
Sources: huggingface
Scheduling policies in large-scale Automatic Speech Recognition (ASR) serving pipelines play a key role in determining end-to-end (E2E) latency. Yet, widely used serving engines rely on first-come-first-served (FCFS) scheduling, which ignores variability in request duration and leads to head-of-line
๐ข ๐ค The FID Lottery: Quantifying Hidden Randomness in Generative-Model Evaluation โ score 5
Sources: huggingface
The Frechet Inception Distance (FID) is the de facto arbiter of image generation, yet most papers report just a single number from a single trained model using a single sampling seed. How reproducible is that number if we retrain the model, or merely resample from it? In this paper, we treat FID as
Other Signals
๐ข ๐ฌ Local LLM Inference Optimization: The Complete Guide โ score 39
Sources: reddit/r/LocalLLaMA
I compiled a year of local LLM experiments into a practical llama.cpp optimization guide, covering VRAM fitting, KV cache, MoE placement, MTP, CPU tuning, and common OOM traps. Pass this to an LLM of your choice and get on the local model train. [https://carteakey.dev/blog/local-inference/local-llm-
๐ข ๐ฌ I pretrained and post trained a 500M parameter LLM and 330M parameter Image generator from scratch โ score 34
Sources: reddit/r/LocalLLaMA
Hey folks Hope you are doing well I started HobbyLM as an side project last month Initially I wrote an Agent harness using Claude SDK which takes notes on various LLM architecture does ablation studies to find optimised or well fit architecture for this model training then I pretrained HobbyLM archi
๐ข ๐ฌ Best local model for vision - 2nd benchmark update - 21 Jun 2026 โ score 32
Sources: reddit/r/LocalLLaMA
I previously posted the first results of my VLM benchmark. There were a few useful comments and observations I took into account, to revise and expand my benchmark: * I initially did not take into
๐ข ๐ฌ [ECCV 2026] Paper Decision Appeals Discussion [D] โ score 26
Sources: reddit/r/MachineLearning
With the release of meta-reviews, ECCV sent out a google form for dissatisfied authors to submit an appeal for the following reasons: 1. Policy errors, e.g., reviewers or Area Chairs applied a policy that does not exist, or reviewers or Area Chairs applied policies that are not applicable for the co
๐ข ๐ฌ Best current methods for finetuning whisper on domain specific vocabulary? [P] โ score 25
Sources: reddit/r/MachineLearning
Hey everyone, Iโm wondering whether there are any newer or more effective methods for fine tuning whisper on domain specific speech. Iโm working on a project where the model needs to reliably detect certain specific words and technical terms. The vocabulary and context are mostly in spanish. Does an
Omitted 1 additional other signals items from the main section; see raw data and source-specific sections below.
๐ Trending Repos
| Repo | Description | Stars Today | Language |
|---|---|---|---|
| chopratejas/headroom | Compress tool outputs, logs, files, and RAG chunks before they reach the LLM. 60-95% fewer tokens, same answers. Library, proxy, MCP server. | 2624 | python |
| calesthio/OpenMontage | World's first open-source, agentic video production system. 12 pipelines, 52 tools, 500+ agent skills. Turn your AI coding assistant into a full video production studio. | 987 | python |
| NousResearch/hermes-agent | The agent that grows with you | 700 | python |
| jamiepine/voicebox | The open-source AI voice studio. Clone, dictate, create. | 614 | typescript |
| ZhuLinsen/daily_stock_analysis | LLM ้ฉฑๅจ็ๅคๅธๅบ่ก็ฅจๆบ่ฝๅๆ็ณป็ป๏ผๅคๆบ่กๆ ใๅฎๆถๆฐ้ปใๅณ็ญ็ๆฟไธ่ชๅจๆจ้๏ผๆฏๆ้ถๆๆฌๅฎๆถ่ฟ่กใ LLM-powered multi-market stock analysis system with multi-source market data, real-time news, decision dashboard, automated notifications, and cost-free scheduled runs. | 568 | python |
| garrytan/gstack | Use Garry Tan's exact Claude Code setup: 23 opinionated tools that serve as CEO, Designer, Eng Manager, Release Manager, Doc Engineer, and QA | 454 | typescript |
| bytedance/deer-flow | An open-source long-horizon SuperAgent harness that researches, codes, and creates. With the help of sandboxes, memories, tools, skill, subagents and message gateway, it handles different levels of tasks that could take minutes to hours. | 442 | python |
| mukul975/Anthropic-Cybersecurity-Skills | 754 structured cybersecurity skills for AI agents ยท Mapped to 5 frameworks: MITRE ATT&CK, NIST CSF 2.0, MITRE ATLAS, D3FEND & NIST AI RMF ยท agentskills.io standard ยท Works with Claude Code, GitHub Copilot, Codex CLI, Cursor, Gemini CLI & 20+ platforms ยท 26 security domains ยท Apache 2.0 | 361 | python |
| topoteretes/cognee | Cognee is the open-source AI memory platform for agents. Give your AI agents persistent long-term memory across sessions with a self-hosted knowledge graph engine. | 347 | python |
| Alishahryar1/free-claude-code | Use claude code and codex for free in the terminal, VSCode extension, and discord like OpenClaw (voice supported) | 258 | python |
๐ New Papers
| Title | Category | Hotness | Link |
|---|---|---|---|
| Context-Aware RL for Agentic and Multimodal LLMs | research_paper | 13 | Open |
| The Data Manifold under the Microscope | research_paper | 10 | Open |
| Freeing the Law with LOCUS: A Local Ordinance Corpus for the United States | research_paper | 8 | Open |
| Rethinking Shrinkage Bias in LLM FP4 Pretraining: Geometric Origin, Systemic Impact, and UFP4 Recipe | research_paper | 8 | Open |
| LedgerAgent: Structured State for Policy-Adherent Tool-Calling Agents | research_paper | 7 | Open |
| LegalHalluLens: Typed Hallucination Auditing and Calibrated Multi-Agent Debate for Trustworthy Legal AI | research_paper | 7 | Open |
| ReSyn: A Generalized Recursive Regular Expression Synthesis Framework | research_paper | 6 | Open |
| Configurable Clinical Information Extraction with Agentic RAG: What Works, What Breaks, and Why | research_paper | 6 | Open |
| Duration Aware Scheduling for ASR Serving Under Workload Drift | research_paper | 6 | Open |
| The FID Lottery: Quantifying Hidden Randomness in Generative-Model Evaluation | research_paper | 5 | Open |
๐ข Lab Blog Posts
- OpenAI: Predicting model behavior before release by simulating deployment
- Anthropic: Jun 17, 2026 Announcements Anthropic opens Seoul office and announces new partnerships across the Korean AI ecosystem
- Anthropic: Jun 12, 2026 Announcements Statement on the US government directive to suspend access to Fable 5 and Mythos 5
- Anthropic: Jun 12, 2026 Announcements Results from the first Anthropic Public Record
- Anthropic: Jun 12, 2026 Announcements TCS and Anthropic partner to bring Claude to regulated industries
- Anthropic: Jun 11, 2026 Announcements DXC will integrate Claude into the systems banks, airlines, and other regulated industries rely on
- Anthropic: Jun 11, 2026 Announcements Introducing Claude Corps
- Anthropic: Jun 9, 2026 Announcements Claude Fable 5 and Claude Mythos 5
- Anthropic: Jun 3, 2026 Announcements Introducing the Services Track and Partner Hub of the Claude Partner Network
- Anthropic: Jun 3, 2026 Policy What we learned mapping a yearโs worth of AI-enabled cyber threats
- Anthropic: Jun 2, 2026 Announcements Expanding Project Glasswing
- OpenAI: Samsung Electronics brings ChatGPT and Codex to employees
- OpenAI: New usage analytics and updated spend controls for enterprises
- OpenAI: Improving health intelligence in ChatGPT
- OpenAI: Using AI to help physicians diagnose rare genetic diseases affecting children
- OpenAI: A near-autonomous AI chemist improves a challenging reaction in medicinal chemistry
- OpenAI: Introducing LifeSciBench
- DeepMind: Unlocking UK house-building with AI-accelerated planning
- DeepMind: Securing the future of AI agents
๐ฆ Twitter/X Highlights
| Account | Tweet Summary |
|---|---|
| xai | Grok models are now available on Databricks Agent Bricks. Bring SpaceXAI's latest models to your enterprise data to power capable AI agents. https://x.ai/news/grok-databricks Post |
| AnthropicAI | New Frontier Red Team blog: Phase 2 of Project Fetch, where we test how well Claude can program a robodog. Opus 4.7, on its own, was ~20x faster than last year's best human team aided by Opus 4.1. (The robodog, alas, still failed to fetch a beach ball.) https://www.anthropic.com/research/project-fet Post |
| GoogleDeepMind | Pinned: Instead of assuming AI will always do what we intend, we ask: what if it doesn't? Thatโs why weโve developed our AI Control Roadmap: a framework for building and managing the advanced AI we deploy within Google. ๐งต Post |
| GoogleDeepMind | Weโre working with @SciTechgovuk, >@mhclg and @i_dot_ai on a new AI housing application planning prototype. ๐ก By cutting down the time spent on repetitive tasks, it could help planning officers focus their attention on complex projects and reduce processing times by up to 50%. โ https://goo.gle/4xzq Post |
| MistralAI | We're taking on the hardest problems in the real world ๐๏ธ๐ ๐ซโ๏ธ Today at The AI Now Summit, held at the Louvre, we announced AI solutions for aerospace, automotive, energy, and physics. Deployed in production at @Airbus , @BMW, @EDFofficiel , and more. More below: Post |
| MistralAI | Mistral AI made the TIME100 Most Influential Companies list for 2026 โ and the top 10 for AI. Why we're proud: customers run frontier models in production on their own terms, on their own infrastructure. Thank you to our customers for their trust and for joining us on the journey. Grateful to our in Post |
| karpathy | This is a super exciting release - Claude Fable 5 is the same underlying model as Mythos but with added safeguards. The benchmarks are great and it's SOTA on everything by a margin but I'll add that qualitatively also, this is a major-version-bump-deserving step change forward (imo of the same ord Post |
| simonw | I just released the first release candidate for sqlite-utils v4, adding a migrations system (previously released independently as sqlite-migrate) and support for nested transactions: https://simonwillison.net/2026/Jun/21/sqlite-utils-40rc1/ Post |
Newsletter
- tldr, rundown-ai: AGENTIC CODING AND PERSISTENT RETURNS TO EXPERTISE (25 MINUTE READ)
- Latent Space: The AI scaling debate always focuses on the question of โhow do we get more GPUs?โ but the better question may be:how do we make the most of ones we already have.
- Latent Space: Itโs not necessarily that xAI is uniquely incompetent(itโs clear they have talented folks)but rather the priorities may be flipped in the GPU arms race.
- tldr: BUILDING AN AI DATABASE FOR AGENTIC GTM OPERATIONS (11 MINUTE READ)
- tldr: FINE-TUNING A CLINICAL AI MODEL TO FRONTIER PARITY (8 MINUTE READ)
- tldr: ANTHROPIC EMPLOYEES ACCUSE TRUMP ADMINISTRATION OF TARGETING THEM (11 MINUTE READ)
- tldr: ANTHROPIC SHIPS MAJOR CLAUDE DESIGN OVERHAUL WITH DESIGN SYSTEM IMPORTS, CODE ROUND-TRIPS, AND A FIX FOR ITS TOKEN-BURNING PROBLEM (13 MINUTE READ)
- tldr: BUILDING A DESIGN SYSTEM SPECCED FOR ENGINEERS AND AGENTS (12 MINUTE READ)
- tldr: AI CODING AGENTS TAUGHT ROBOTS HOW TO INSTALL GPUS AND CUT ZIP TIES (4 MINUTE READ)
- tldr: BUILDING TREX: CODE EXECUTION AND ARTIFACT GENERATION FOR AI CODE REVIEW (9 MINUTE READ)
- tldr: THE FOUNDER'S PLAYBOOK: BUILDING AN AI-NATIVE STARTUP (5 MINUTE READ)
- tldr: CAN GZIP BE A LANGUAGE MODEL? (5 MINUTE READ)
- tldr: EXA AGENT (5 MINUTE READ)
- tldr: DESIGNING WITH UNCERTAINTY: HOW AI SUPERCHARGES PROBABILISTIC THINKING (19 MINUTE READ)
- tldr: AI IS NOT A NEW MARKETING PROBLEM. IT IS A NEW BRAND INTERFACE (5 MINUTE READ)
- tldr: WORLD LEADERS WANT US AI, BUT NOT US CONTROL OVER ACCESS (3 MINUTE READ)
- tldr: DATA PRIMACY COULD BUILD MORE RELIABLE AI INFRASTRUCTURE (5 MINUTE READ)
- tldr: OKTA AND GOOGLE CLOUD LINK IDENTITY TO AI AGENTS AND BROWSERS (3 MINUTE READ)
- tldr: VERCEL LAUNCHES ENTERPRISE CONTROLS FOR AGENTIC AI INFRASTRUCTURE (3 MINUTE READ)
- tldr: A NEEDLE IN A STACK OF NEEDLES: HUNTING INFOSTEALERS WITH AI (10 MINUTE READ)
- tldr: SENTINELONE EXPANDS PURPLE AI TO INVESTIGATE THREATS AUTONOMOUSLY (3 MINUTE READ)
- tldr: LIGHTHOUSE FOR AI AGENTS (GITHUB REPO)
- tldr: INSIDE CLAUDE MANAGED AGENTS: REVERSE ENGINEERING THE SECURITY BOUNDARIES OF ANTHROPIC'S HOSTED AGENT RUNTIME (12 MINUTE READ)
- tldr: CHAINGUARD, JPMORGAN, BNY TEAM UP TO SECURE OPEN SOURCE FROM AI THREATS (2 MINUTE READ)
- tldr: JETBRAINS MARKETPLACE ECOSYSTEM SECURITY UPDATE: ADDRESSING MALICIOUS THIRD-PARTY AI PLUGINS (4 MINUTE READ)
- tldr: ADYEN LEANS INTO AGENTIC COMMERCE (2 MINUTE READ)
- tldr: STRIPE AND AWS ENABLE AI AGENT PAYMENTS FOR CONTENT OWNERS AND PUBLISHERS (2 MINUTE READ)
- rundown-ai: How Iโm preparing for an AI doctor
- rundown-ai: Save hours on meeting prep with Google Gemini
- rundown-ai: OpenAI poaches a transformer pioneer from Google
- rundown-ai: Codex - OpenAIโs agentic coding tool, with new Record & Replay for creating reusable skills
- rundown-ai: Brain - Perplexityโs self-improving memory for its Computer agent
- rundown-ai: Firefly Studio - Adobeโs upgraded all-in-one platform to generate and edit with AI
- rundown-ai: Crosby - Agentic law firm for sales teams to speed up time to signature
- rundown-ai: Anthropicโs Chris Ciauri told reporters during an event in Korea that the company is โvery confidentโ that its Mythos and Fable models will become available again in the โcoming days.โ
- rundown-ai: Adobe rolled out new agentic skills for Firefly AI Assistant, also extending its creative agent into public beta across Photoshop, Premiere, Illustrator, InDesign, and Frame.
- rundown-ai: Databricks launched a slate of new agentic tools at its Data + AI Summit, including LTAP for running AI apps and analytics, and an AI-run customer data platform called CustomerLake.
- rundown-ai: Former White House AI advisor Dean W. Ball is joining OpenAI to lead Strategic Futures, a new team that will help shape frontier AI policy.
- rundown-ai: Community AI workflows
- rundown-ai: Read our last AI newsletter: Inside the deadlock keeping Mythos offline
- rundown-ai: RSVP to next workshop on June 25: Get consultant-grade strategy from AI
- rundown-ai: AI, world leaders meet at G7 as Mythos standoff continues
- rundown-ai: Bloomberg released a letter from U.S. Commerce Sec. Howard Lutnick, who warned Anthropic against distributing Mythos/Fable to โforeign persons.โ
- rundown-ai: Why it matters: Anthropic employees are coming to the same conclusion we initially did โ that this is a relationship issue more than a safety one. But details like WaPoโs expanded Mythos list point to a clear situation t
- rundown-ai: Pew: Americans using AI more but trust it less
- rundown-ai: Design-driven governance for agentic AI
- rundown-ai: Grok Imagine 1.5 - xAIโs newly upgraded image-to-video model
- rundown-ai: Exa Agent - Exaโs cost-effective, frontier-level web research API
- rundown-ai: Eve - Vercelโs open-source framework for turning a file directory into an agent
- rundown-ai: GLM 5.2 - Z AIโs powerful new open-weights model
- ... plus 3 more newsletter-only items in raw data