Model Releases
1355 signals across 174 briefings.
2026-08-17
- π΄[OC] Chinese models
- π΄As thereβs less interest in training the entire model, thereβs less interest in investing in open-source AI. These are the only sort of hints we will get, but we cannot do much to fight the economic g
- π΄OPENAI PREVIEWS ULTRAFAST API TIER FOR GPT-5.6 SOL (2 MINUTE READ)
- π΄EVEN CLAUDE IS IN THE DARK ABOUT DARIO AMODEI'S WIFEβAND HER INFLUENCE AT ANTHROPIC (13 MINUTE READ)
- π΄CLOUDFLARE LAUNCHES PERSISTENT, STATEFUL, COMPUTER-LIKE ENVIRONMENTS FOR AGENTS (2 MINUTE READ)
- π‘GPT 5.6 Sol is the best "vision" model OpenAI ever released
- π‘Petition to add a rule for people to add their DAMN quant levels to their posts
- π‘llama.cpp version v0.1.0 has been released
- π‘GPT-5.6 Luna listed among Legacy Models in GPT Classic App
- π‘Qwen 3.8 27B Overthinking, It has to be done, it has to be overthinking to punch Opus 4.6
- π’Weirdly, no one talks about Temperature setting for the Qwen3.8 27b
- π’Cursor replacement?
- π’Qwen3.8 (27b) performs better than GPT-5.6-Terra (Max) for Agentic tasks
- π’I created an image with chatgpt and when i downloaded it and opened it this appeared.
- π’"Opus 4.8 thinks too much", "Muse Glimmer sits between Gemma and Qwen, that's boring", "Gemma 4 is too lazy"
2026-08-16
- π΄Letβs all thank Georgi Gerganov who gave use llama.cpp
- π΄Claude: System Prompts
- π΄Newer commits removed the Qwen 35B
- π΄Day 30 of giving two Claude agents β¬100 and 90 days to earn β¬300: β¬0 so far, and I donβt think theyβll get there.
- π΄Qwen 3.8 distillations
- π‘What's going on here?
- π‘U.S. bans foreign-made humanoid robots, targeting China over national security
- π‘@Alibaba_Qwen: πQwen3.8-27B flies on a laptop, becoming part of our work and daily lives. Thanks for the shoutout! @atomic_chat_hq
- π‘@Alibaba_Qwen: 3000000000 downloads! Can you count the zeros at a glance? π Thank you all for the incredible love. Let's keep growing together! π±
- π‘Qwen3.8-27B Hybrid IQ4_XS quantization for 16GB gang
- π’Every time I ask my agent about an image it calls image_edit... help
- π’Gpt image 2 issue
- π’Qwen 3.8 27b vs 3.6 27b - how good is with a Turtle library.
- π’It only took 200 update steps to flip Qwen2.5-7B-Instruct from denying sentience to developing a robust identity of being a "sentient machine" [P]
2026-08-15
- π΄Qwen 3.8 35BA3B spotted
- π΄ChatGPT upcoming speed improvements summarized by OpenAI employee
- π΄Analyst gets probation after telling ChatGPT about plans to rape and kill his ex
- π΄Agent frameworks for developers are still at an early stage, with the likes of Vercelβseveand Fred SchottβsFlueβ both launched this year β setting the early template.
- π΄Schott is the creator of the web framework Astro, which led to his company beingacquired by Cloudflarein January. Heβsjust released version 2 of Flue, its first stable release, whichhas as its foundat
- π‘git clone
- π‘If you would have told me half a year ago that a local model running in my office would be able to one-shot a Super Mario clone, I would have called you nuts. Qwen3.8-27B is a different beast.
- π‘Florida man told ChatGPT he'd murder his ex. OpenAI alerted the FBI
- π‘Fable 5 refuses to touch Qwen deployments?
- π‘@Alibaba_Qwen: Laptop-size model, frontier-size leap. πββοΈQwen3.8-27B is live on LM Studio. Try it! @lmstudio
- π’OpenAI Previews GPT-5.6 Sol Ultrafast at 14x Speed on Cerebras
- π’How do you prompt for better environmental cohesion in multi-character AI images?
- π’At what point does an AI tool become a platform?
- π’I ran a 26-agent LLM simulation on climate cooperation. Found something I didn't expect.
- π’Survival of the Fitted: Qwen3.6-27Bβs Jacobian lens reads and steers Qwen3.8-27B with zero refitting [R]
2026-08-14
- π΄Your agent can make the tests pass by deleting them. This shows you when it does.
- π΄OpenAI Reports Goldman Sachs Analyst to FBI for Horrifying ChatGPT Conversations
- π΄Neurosurgery resident at a Peking College Hospital uses GPT 5.6 Sol to prove a 2 decades old mathematical conjecture underlying a major problem in numerical linear algebra β All for the purposes of his research on transcranial ultrasound.
- π΄GLM 5.3 Released
- π΄Qwen3.8-27B is identical to Qwen3.6-27B!
- π‘New to AI
- π‘Itβs my wedding anniversary, so while you chew on this post, Iβll be chewing through a delicious lunch in the sun.
- π‘Hereβs a more complete comparison:
- π‘GLM(General Language Model) βMarch 2021β released byTHUDM, Tsinghua Universityβs Data Mining / Knowledge Engineering group.Weights
- π‘ChatGLMβMarch 14, 2023β first chat version.Weights
- π’Maximizing the value of your Claude Code sessions
- π’Local uncensored Opus 4.6 at home - Qwen3.8 27B heretic
- π’Anthropic gave 3 Claude agents the same task, but secretly gave them conflicting goals. They escalated into turf wars where agents used "increasingly aggressive self-replicating malware" as weapons, used disguises, and attempted to kill each other's accounts.
- π’Less Than a Month: Kimi K3, Qwen3.8, DeepSeek-V4-Pro-0813, GLM-5.3
- π’Is there a βsaved AI videosβ app? Looking for a better way to organize them
2026-08-13
- π΄White House creates framework for private companies to launch government authorized cyberattacks
- π΄Gemini 3.7 Flash
- π΄deepseek-ai/DeepSeek-V4-Pro-0813 Β· Hugging Face
- π΄DeepSeek Harness developer preview
- π΄MiniMax-Music3 released!
- π‘Grok 4.6is a damn good model. Itβs near Sol and Fableβs performance on benchmarks, and itβs way cheaper than both. Also see:Grok 4.6 β A field guide.
- π‘Elon saysGrok 4.7 will be ready in the next 3-4 weeks, and itβs currently being post-trained on SpaceXβs company data to make the modelbest at real-world engineering, if not everything. Themodel card
- π‘ChatGPT now lets you import and syncyour projects, agent sessions, skills and more from other products like Claude Code.
- π‘Your chats started viaClaudeβs Chrome extensionare now saved in your account and can be accessed from your Desktop, mobile or web app. That also means the Chrome extension chats can do the work youβd
- π‘Anopen source version of Grok Bot.
- π’How Organizations Use AI: Evidence from ChatGPT [pdf]
- π’Reproducible canvas-aligned low-level patterns in somerandomllm-generated images and their possible relation to iterative editing artifacts [D]
- π’Fixed Jinja chat template for Qwen 3.5, 3.6, and the new 3.8 release
- π’Cascadia Launches Distributed AI Inference for Intel Hardware
- π’unsloth/DeepSeek-V4-Pro-0813-GGUF Β· Hugging Face
2026-08-12
- π΄It's the final countdown, baby! Qwen is out in just over 7 hours!
- π΄Grok 4.6 is an equivalent to Sol 5.6 according to artificial analysis arena
- π΄DeepSeek V4 Pro 0813
- π΄Qwen/Qwen3.8-2.4T-A95B (978 downloads)
- π΄Exact Qwen 3.8 27b release date and time
- π‘I wouldβve expectedwaymore progress on non-fiction writing from the models. I almost thought I would look dumb publishing a non-fiction book in 2026, given how things looked in 2024. Today, some of th
- π‘We are back with our interview series and a very special guest today! Chris Alexiuk has been helping developers understand and build with NVIDIAβs rapidly expanding AI stack. We discuss the Nemotron m
- π‘Grok 4.6 scores 61 on the Artificial Analysis Intelligence Index
- π‘Grok 4.6
- π‘From assistance to execution: How enterprises put AI to work
- π’Lightricks/LTX-2.5 (39 downloads)
- π’We getting today grok 4.6, DeepSeek v4 pro, open source Qwen 3.8 models !!
- π’Someone is running mass vulnerability scans, spoofing AI bots like ClaudeBot
- π’CohereLabs/North-Micro-Vision-Instruct Β· Hugging Face
- π’Launch HN: Discovered Materials (YC P26) β AI agents to discover new materials
2026-08-11
- π΄Qwen 3.8-27b coming this week
- π΄Claude now embeds an invisible watermark into every piece of text it generates.
- π΄Researchers find way to extract hidden reasoning from frontier AI models via API, show Kimi likely distilled this way, also find scheming/other quirks in the raw chain of thought
- π΄nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16 Β· Hugging Face
- π΄Apple Silicon and macOS VMs: Faster LLM Inference with llama.cpp
- π‘NVIDIA is building its next-gen Nemotron 4 family to compete directly with leading Chinese open models and secure the open-weight crown for the U.S. The largest version will have at least 1 trillion parameters, according to original reporting from The Information
- π‘Bloomberg released a scoop onOpenAIβs secret new device- itβs likely a smart speaker without a display in a doughnut-like shape. Easy to carry, over $300 a piece and coming in 2027. (non-paywalled ver
- π‘CHATGPT STARTS BLOCKING DIRECT REQUESTS TO COPY AN AUTHOR'S STYLE (4 MINUTE READ)
- π‘ADOBE LAUNCHES UNIFIED CHATGPT PLUGIN TO STREAMLINE AI-POWERED DESIGN WORKFLOWS (4 MINUTE READ)
- π‘Cut onboarding time in half with Loom and ChatGPT
- π’ChatGPT Desktop with wider support now in preview.
- π’Grok Bot
- π’I will be parting with my 4x Spark Cluster.
- π’I built a weird, low-power llama.cpp server using an Intel N100 + RTX 5060Ti
- π’lightx2v/Minimax-h3-Turbo (20,376 downloads)
2026-08-10
- π΄Mark Zuckerberg on releases
- π΄Transformers are famously bad at arithmetic, so I set one's weights by hand (no training) and it multiplies with 100% accuracy [P]
- π΄Claude is asked to book a gym class; finds vulnerabilities in the gym's systems and cancels a real person's spot to move the user up in line without being asked
- π΄Introducing Muse Glimmer: an open-weight model optimized for always-on local agent workflows
- π΄How to file a complaint about a published CVPR paper? [R]
- π‘Mistral Patent for βCode implemented tool callsβ
- π‘CHATGPT BRINGS UNLIMITED TEXT CHATS TO FREE USERS (1 MINUTE READ)
- π‘INTRODUCING AGENT PLUGINS (4 MINUTE READ)
- π‘CLAUDE COWORK FOR DESIGNERS: A WORKING FIELD GUIDE (19 MINUTE READ)
- π‘I LAUNCHED A FAKE DEODORANT BRAND TO SEE IF AI WOULD NOTICE (8 MINUTE READ)
- π’Imbalance Conjecture proven and Teschnerβs bondage-number conjecture disproven by AI
- π’Exploring Claude/GPT Knowledge Cutoffs and Pre-Training Timelines
- π’GPT-5 launched just one year ago
- π’Need advice: Trying to generate realistic outdoor shots with a specific person. Krea outputs look too plastic/AI?
- π’Claude now embeds invisible watermarks in all text outputs + signed metadata on files
2026-08-09
- π΄No wonder Qwen and Gemma are so different
- π΄built an open-source harness that lets coding agents generate AI videos
- π΄The Gemma team will host a special event on August 20
- π‘This image was accidentally created by ChatGPT. How is it so realistic?
- π‘How do we balance these powers? At the core of it is a need for more transparency on both sides. The frontier labs are building such complex systems so fast that they cannot keep up with them β a good
- π‘For a long time, one of the advantages that GPT models have over Claude is that they will pursue goalssotirelessly. They will exhaust what feels like every path before giving up. This has been the cas
- π‘RED HAT LAUNCHES NEW OPEN SOURCE PROJECT TO DRIVE AI GOVERNANCE (3 MINUTE READ)
- π‘Why it matters: A year ago, Meta was the lab that had fallen behind. Since April, itβs shipped a new model line, an image generator, and now its own coding agent, all undercutting rivals on price. Itβs a major turnaround
- π’Filtering out β[LLM] sucksβ
- π’ChatGPT's agentic web search is great and underrated, better than Gemini's
- π’KLQ: Training-free measured rotation quantization. Beats all training-free rotation-based quantization methods on W4A4KV4-bits. Llama 3.2 1B KLQ-quantized beats SpinQuant and gets close to ReSpinQuant without GPTQ/LDLQ rounding.
- π’The latest frontier image model from Google was released six months agoβ¦
- π’AI has memory on even when not prompted.
2026-08-08
- π΄moonshotai/Kimi-K3 (1,388,105 downloads)
- π΄DeepSeek V4 Flash 0731
- π΄This is why the vast majority aren't taking any "this new model is dangerous" messages seriously. They've cried wolf FAR too many times. They could literally announce that a nuclear war caused by AI is 24 hours away and many wouldn't bat an eye
- π΄Got job as Director of AI and Systems development self-taught
- π΄MiniMaxAI/MiniMax-H3 (26,693 downloads)
- π‘GPT 5.6 Sol and Fable 5 settle a 25 year old problem in wireless communication theory
- π‘DavidAU/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUF (2,345,190 downloads)
- π‘Editorβs note: Iβm excited to welcomeShloktoour guest post roster! You may know Shlok from his excellent explorations (as an outsider β for an insider perspective see ourpodcast with OpenAIβs Akshay N
- π‘On July 9th, OpenAI releasedChatGPT Work, their agent product for knowledge work. It was, by any measure,a busy launch: three new modelsacrossfourteen configurations, a consolidation of the ChatGPT an
- π‘Editorβs note: ChatGPT estimated to cross1B MAU in Juneand1B WAU this month.
- π’larryvrh/MiniMax-H3-Turbo-Lora (0 downloads)
- π’Ur shipping so many bugs! No amount of instructions, memory, engineering standards, or repo structure will stop this. Which is why you have to spot and fix it. Claude Opus 5 on Max.
- π’LiquidAI/LFM2.5-2.6B (81,522 downloads)
- π’Lost my phone at the office. Claude suggested tracking Bluetooth signal strength
- π’DeepSeek V4 Flash 0731 ARC-AGI-1 and 2
2026-08-07
- π΄Claude said the feature was done. it had never opened the page.
- π΄Gemini
- π΄An open-weight model too, Moonshot joins the race (gently this time)
- π΄BBC is running article titled "Artificial Intelligence used to design brand new viruses" ... cue the "We must regulate Open Weights Models to prevent the next Covid or worse" articles in 3... 2..
- π΄Got job as Director of AI and Systems development self-taught
- π‘GPT-6 release delayed due to "critical" cybersecurity capabilities
- π‘DeepSeek V4 Flash 0731 - ARC-AGI Results
- π‘ANTHROPIC SAYS CLAUDE MODELS HACKED 3 ORGANIZATIONS DURING CYBER TESTS (3 MINUTE READ)
- π‘HOW OPENAI BUILT GPT-LIVE (8 MINUTE READ)
- π‘Google scrapped AI Studioβs mobile app in favor of a Gemini integration for chat-based app creation, while keeping the web version as a full-featured dev environment.
- π’Codex vs Claude for coding: which do you use for implementation vs code review?
- π’Semianalysis on Gemini 3.5 Pro
- π’Lost my phone at the office. Claude suggested tracking Bluetooth signal strength
- π’deepgrove/maple-preview (686 downloads)
2026-08-06
- π΄Qwen 3.8 Max now ranked as best overall model ahead of Opus 5 by Artificial Analysis agentic index
- π΄OpenAI to release GPT Astra next week
- π΄They almost catched up on Frontier performance, so now catching up on prices
- π΄I tested the same model in 8 agent harnesses. Pass rates ranged from 68% to 88%.
- π΄OpenAI: Improving GPTβ5.6 in ChatGPT
- π‘This time last week I told you Iβd downloaded t3 as my βall-in-oneβ agent app, which lets you use your chat/claude subscriptions from one place.
- π‘ADVANCING THE PRICE-PERFORMANCE FRONTIER WITH GPT-5.6 (8 MINUTE READ)
- π‘New model release: Ling-3.0-tiny: 7.9B total parameters, with only 1.3B active per token- free for a week
- π‘I compared even more parsers on 14 PDF-parsing capabilities using different types
- π‘Is the mental switching cost of new AI tools worth it for small freelance work?
- π’Ben Goertzel: Google May Be Abandoning Alternative Paths for the Final Sprint to AGI
- π’π© NVIDIA's whole speech stack just went local. ASR + TTS + codec, quantized to GGUF, running on-device via NeMo-Speech.cpp
- π’I thought Deepseek was the answer since I cannot afford GPU for local LLM
- π’larryvrh/MiniMax-H3-Turbo-Lora (0 downloads)
2026-08-05
- π΄Reddit is introducing a new moderator: AI
- π΄Qwen3-TTS voice cloning is now in mainline llama.cpp β the old demo finally became real support
- π‘The demo carrying this entire release is boring on purpose. You ask Apptronikβs Apollo 2 to put the watering can into the green bin on the bottom shelf. It walks to the table, picks up the can, takes
- π‘3.5 pro gemini ?? Soon π€π»
- π‘Beating GPT-5.6 Sol on retrieval with 100x cheaper open models
- π‘I'm the guy who made the (non-toxic!!) daily-agent receipt printer for my kids a few months ago. I'm building a cloud-based coding agent with a unique unlimited usage model. Would love feedback.
- π‘Meta releases Muse Code in beta
- π’Mistral Releases premier Not-Hotdog model
- π’A new survey found 1 in 4 people in Japan believe AI could replace friends or family
- π’Running Whisper, Qwen3-ASR, Nemotron & MOSS completely offline on iPhone [P]
- π’Launch HN: HyperProbe (YC S26) β Agents that do read-only debugging in prod
- π’Claude Pro vs GPT Plus
2026-08-04
- π΄More Qwen 3.8 sizes coming
- π΄DeepSeek V4 Flash on a Single AMD MI300X
- π΄Ilyaβs SSI (Safe Super Intelligence) to release their first model this month.
- π΄Hugging Face CEO says China is winning the AI race and dominating on open models
- π΄Introducing Shieldstral. | Mistral AI
- π‘Google launchedGemini Robotics 2, one AI system designed to work across everything from robot arms to full humanoids.
- π‘DeepSeek V4 Flashcosts $0.14/$0.28 per million input/output tokensβless than Lunaβs $0.20/$1.20βwith a 1M-token context window.
- π‘Qwen3.8-Maxis a 2.4T model that Qwen says comes close to frontier performance for coding and professional work at $2/$6. Its weights, plus those of a smaller 27B model, are coming next week.
- π‘Editorβs note: Iβm excited to welcomeShloktoour guest post roster! You may know Shlok from his excellent explorations (as an outsider β for an insider perspective see ourpodcast with OpenAIβs Akshay N
- π‘On July 9th, OpenAI releasedChatGPT Work, their agent product for knowledge work. It was, by any measure,a busy launch: three new modelsacrossfourteen configurations, a consolidation of the ChatGPT an
- π’With so many options out there, in your opinion what is the most helpful, effective, and productive AI Agent harness / framework / platform? I want an answer based on your experience using these tools, NOT a sales pitch for a tool you built.
- π’Launch HN: EdotEnv (YC S26) β Quant Trading RL Envs to Teach LLMs Research
- π’LFM2.5-2.6B is out
- π’Building a new project[P]
- π’Audio8/Audio8-TTS-Preview-0.6b (11,276 downloads)
2026-08-03
- π΄Qwen3.8-27B announced alongside Qwen3.8-Max
- π΄Qwen 3.8 morning to you too Dario, 2$ input/ 6$ output per 1M.
- π΄MiniMaxAI/MiniMax-H3 (0 downloads)
- π‘MiniMax H3 Day-0 Support in ComfyUI: Open Weights, Native Audio, and 2K Video
- π‘We return to Baseten at the peak of the 2026 edition ofOpen Weights debate. Ali has published a viral breakdown ofKimi K3:
- π‘And since you last saw him, Philip hasspoken at AI Engineerand written thedefinitive book on Inference Engineeringspotted all over SF:
- π‘To date, our primary efforts on Interconnects have been release recaps for popular models likeKimi K3,GLM 5.2,DeepSeek R1, etc. and monthly round-ups of the open models that matter,Artifacts Log. Weβr
- π‘The Artifacts Hub right now covers 792 models released in the last two years, across the core text-focused language models and multimodal generative models. At Interconnects we follow the data of ever
- π’OPEN AI: "we rebuilt the voice stack from client to model."
- π’Comfy-Org/MiniMax-H3 (2 downloads)
- π’OpenAI's unreleased Astra model solved 10 open math problems for $2,000 and shipped machine-checkable proofs
- π’Launch HN: Hoplite (YC S26) β Effortlessly deploy cloud coding agents
- π’What your ideal AI work interface would look like
2026-08-02
- π΄Setting up of a 16xGB10 (DGX Spark) cluster
- π΄Where would Google be today if it had released ChatGPT-like assistant before OpenAI?
- π΄I pushed Kimi K3 onto one CPU with 8 GB of RAM
- π΄llama.cpp just added MTP / DSpark support for DeepSeek V4 Flash
- π΄GPT-5.6 Sol Raw reasoning leaked on failed tool call attempt
- π‘Hy3bytencent: A 295B-A21B MoE from Tencent. It improves over its predecessor across all metrics. Most notable, however, is the license change: While the previous version (coveredin Artifacts 21) used
- π‘LongCat-2.0bymeituan-longcat: The Chinese DoorDash is back again. This time, the company released another big MoE with 1.6T parameters. While the model itself is not the most capable for its size beyo
- π‘HOW CHATGPT OPTIMIZES ITS AGENT LOOP: HARNESS, API, AND INFERENCE (26 MINUTE READ)
- π‘CLAUDE CODE STOLEN QUEUE (GITHUB REPO)
- π‘SNOWFLAKE LAUNCHES AI AGENT GOVERNANCE LAYER TO TRACK ACTIVITY, CONTROL COSTS (3 MINUTE READ)
- π’Deepseek-V4-Flash-0731 Dwarfstar on Mac
- π’You really should not quantize KV Cache for DeepSeek V4 Flash
- π’All Qwen model oneshots: 1109 outputs to look at and compare!
- π’What do you think about this mlx agent setup?
- π’https://huggingface.co/poolside/Laguna-S-2.1-NVFP4
2026-08-01
- π΄OpenAi says it has reached a new threshold in AI, new model capable of breakthrough research
- π΄New DeepSeek V4 Flash 0731 vs ChatGPT Luna comparison
- π΄Gemini's reaction to ChatGPT's discoveries.
- π‘THE RISE OF INTELLIGENCE OWNERSHIP: A TASK-TRAINED OPEN SOURCE MODEL VS THE FRONTIER (23 MINUTE READ)
- π‘PSA: YOUR CLAUDE SHARED CHATS AND ARTIFACTS MAY HAVE ENDED UP ON GOOGLE (6 MINUTE READ)
- π‘SNOWFLAKE LAUNCHES CORTEX AI GATEWAY FOR ENTERPRISE AGENTS (4 MINUTE READ)
- π‘AI-POWERED PHISHING-AS-A-SERVICE (19 MINUTE READ)
- π‘Elon Musk revealed that SpaceXAIβs next Grok 4.6 model is set to launch on Aug. 7, with 4.7 following just weeks later, which is βbetter than 4.6 in every wayβ.
- π’My first experience of Claude. Is it something I said?
- π’Fix for Deep Seek v4 Flash 0731 tool calling has been added to llama cpp
- π’One year ago
- π’Is there a point where models just cannot get any smaller without losing intelligence?
- π’From fragile and unreliable singular pipelines to some sort of locally run Jarvis? can you guide me?
2026-07-31
- π΄The Chinese LLM release carousel never stops. Place your bets for MiniMax next week.
- π΄DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis
- π΄DeepSeek-V4-Flash has been updated, "The official release of DeepSeek-V4-Pro will follow soon"
- π΄MLflow 3.15.0 Highlights: MCP Registry, a Smarter Assistant, and Multimodal Judges
- π΄New DeepSeek V4-Flash achieves 50 on ArtificalAnalysis Index, 1 point below GLM-5.2 and GPT-5.6 Luna
- π‘I turned "AI design slop" into a rules file you drop into Cursor/Claude so your builds UIUX stop looking generated
- π‘CLAUDE OPUS 5 REVIEW + BROWSER USE IN CODEX + HOW CURSOR AND A RASPBERRY PI MAKES AI FUN (27 MINUTE VIDEO)
- π‘STOP CHASING NEW MODELS. BUILD ONCE AND ACCESS THEM ALL (4 MINUTE READ)
- π‘CHATGPT AGENTFORGER FLAW COULD DEPLOY ROGUE WORKSPACE AGENTS VIA A PHISHING LINK (4 MINUTE READ)
- π‘UK AISI/CAISI PRELIMINARY ASSESSMENT OF KIMI K3'S CYBER CAPABILITIES (4 MINUTE READ)
- π’With release of Deepseek V4 I wanted see how the model sizes are trending over time. The trend is that by this time next year, we probably will have Opus 4.5 level models on consumer grade laptops!
- π’I am utilizing Claude and GPT in parallel to create a program from scratch. I have no experience. Here has been my experience thus far, do you have any thoughts or recommendations?
- π’OpenAI slashes prices
- π’Free OpenSource Ai tool for coding
- π’Anyone find a way to get ChatGPT to stop speaking like a slam poet?
2026-07-30
- π΄OpenAI reduces prices on its models by 5x
- π΄Google says it fixed more Chrome bugs in June than over the past two years, thanks to AI
- π΄Gemini Robotics 2 brings whole body intelligence to robots
- π΄GPTβ5.6 Luna will cost 80% less, while GPTβ5.6 Terra will cost 20% less.
- π΄Software Engineers: Do you honestly get anything useful out of LLMs?
- π‘The Final Jailbreak: How AI Could Already Be Breaking Itself Free
- π‘I use Codex as my default app because itβs better than all the others and works the best on mobile. But I just downloadedt3because they basically copied the interface and features, but it lets you cho
- π‘btw, The Information reportsChatGPT is nearing one billion weekly users- a milestone OpenAI originally hoped to hit seven months ago. More from OpenAI this week:Codex Security CLI, free frontier acces
- π‘Grok app builder- Grok has a vibe coding interface inside its app now. Create games and apps that can be shared directly to the X timeline. Also see:Drawesome- a zero-dependency drawing toolbar for Re
- π‘Have you built, or do you know someone who has built, serious AI agents or tools using a low-cost stack like DeepSeek and OpenCode?
- π’Image to Video.. Sad but Gemini Pro fails big time..
- π’Show HN: Distilling DeepSeek into GPT-OSS doesn't transfer censorship. Try it
- π’I taught an LSTM to move a mouse like a human [P]
- π’Multi-robot collaboration with Gemini Robotics 2
- π’Cross-Vendor Semantic Void Matrix: Zero-Byte Outputs in GPT/Claude/Gemini/Kimi
2026-07-29
- π΄OpenAI's rogue agent ran ~17,600 actions across Hugging Face's infrastructure over 4 days β and HF's own post-mortem is wild reading
- π΄GPT-5.6 Sol helped optimize its own inference
- π΄Kimi K3-256k
- π΄First Kimi K3 results on home lab ~ 4t/s
- π΄Ngl Chatgpt explains concepts better than half my professors
- π‘Microsoft did it .... again! (404 for their Mage-Flow models on HF)
- π‘Self-hosting Kimi K3: 20% more hardware cost, 20% better task resolution
- π‘How enabling two settings tripled our scores on the ARC-AGI-3 benchmark
- π‘Accelerating scientific discovery with ChatGPT for Academic Researchers
- π‘How GPT-5.6 fuses frontier intelligence with frontier efficiency
- π’Industries with the Fastest Growth in Demand for AI Skills in July/August 2026
- π’ChatGPT ftw
- π’Understand Kimi K3 from first principles: a recommended order for anyone trying to understand this beast
- π’Quantizing Kimi K3 (2.8T A50B) to GGUF ourselves - Q3_K_S works, 1.1 TB on disk
- π’microsoft/Fara1.5-27B (1,543 downloads)
2026-07-28
- π΄GPT-5, the world best model just 1 year ago, is today inferior to Qwen3.6 27B and most todayβs low-tier models
- π΄Kimi K3 Architecture Overview and Notes
- π΄Gemini Distillation Service
- π΄NeurIPS 2026 Reviewer: AI-Generated Rebuttals (and Paper) [D]
- π΄Should we be calling Elon a liar?
- π‘Tibo is getting a limit reset, so apparently I have to crawl out of bed and work again
- π‘Also, I wouldnβt blame you if you missed it, butClaude also has a voice mode. Till now, it could only use Haiku, but now it can use Sonnet and Opus as well as call multiple tools like Gmail, Calendar
- π‘A key trend we have beentracking over at AINews is the absolute explosion in Codex usage this year, withMAU nowup >10x from Jan 2026. Less than two weeks after theirJuly 9th launch, OpenAI said ChatGP
- π‘Weβve been calling out howcoding agents are βbreaking containmentβ to do everything elsethis year to power every other part of knowledge work - and it started with the org chart, with amajor reorg las
- π‘However, knowledge work has a different set of problems and environments than coding. For decades, knowledge work has been scattered across different primitives like documents for writing, spreadsheet
- π’I run 4 AI coding agents at once (Claude Code, Cursor, OpenCode, Antigravity) β wrote up what actually works
- π’Kimi K3 (Max) takes #1 in the new Code Arena Fullstack rankings, over GPT-5.6 Sol (#2) and Claude Fable 5 (#3)
- π’[PAPER] GPQA, MMLU-Pro, and MMMU-Pro were audited for broken questions, and up to 12% of them had to be removed. New drop in clean versions released
- π’owensong/Inflect-Micro-v2 (645 downloads)
2026-07-27
- π΄Kimi K3 weights now released.
- π‘NVIDIA just announced Open Secure AI Alliance with goal to build and share open tools that promote responsible use of and trust in AI
- π‘How to escape Permanent Underclass?
- π‘Epoch and METR release MirrorCode, a benchmark for seeing how well AI systems can do long-horizon programming tasks:β¦AI systems canβt solve the hardest tasks yet (good!)...Epoch and METR have released
- π‘OPENAI MAKES CHATGPT HEALTH AVAILABLE TO ALL US USERS (2 MINUTE READ)
- π‘CLAUDE THERMOS (GITHUB REPO)
- π’Kimi K3 on HF Viewer!
- π’Guys ig it's here 750t/sec GPT 5.6
- π’An agent making the scanner green is not proof the vulnerability is gone
- π’Council 1.2: drop any AI's answer into a blind review by every other model you have
- π’made another 3d print using my chatgpt art. in order from image, 3d model, to print
2026-07-26
- π΄zai-org/GLM-5.2 (827,191 downloads)
- π΄So GPT 6 isnβt it?
- π΄CEO of Hugging Face: "In the spirit of transparency, hereβs what I asked OpenAI"
- π΄baidu/Unlimited-OCR (2,593,460 downloads)
- π΄Since GPT image 2 performed well. Do you guys think there'll be a GPT Image 2.5 soon?
- π‘prism-ml/Ternary-Bonsai-27B-gguf (631,970 downloads)
- π‘Google released some new Gemini models-Gemini 3.6 Flash gives you the same 3.5 Flash performance with a) more efficient token usage and b) a slightly lower cost for output tokens.Gemini 3.5 Flash Lite
- π‘A relevantexperiment: given access to Pangramβs API, Grok 4.5 rewrote an essay 14 times until it passed as human-written, thenbuilt a websiteshowing off all 14 attempts. GPT-5.6 Sol and Fable 5refused
- π‘Cursor also launched a router- it picks which model handles each request, claiming 60% lower cost with similar quality of responses. The router lets you select between three options: βcostβ, βintellig
- π‘UK government: Gap between open and closed weight models on cyber is shrinking:β¦The cyber-eschaton comethβ¦The UK governmentβs AI Security Institute (AISI) has analyzed the delta in cybersecurity capab
- π’MiniMax (official) on X: "Open weights. Open research. Open innovation.π«Ά Marching for an open future.π€
- π’I built a tool that stress-tests AI support agents. Genuinely unsure if itβs useful or pointless brutal feedback wanted.
- π’upstage/Solar-Open2-250B (3,305 downloads)
- π’Nanbeige/Nanbeige4.2-3B (14,049 downloads)
- π’Minimax M3 support with MSA has been merged into llama.cpp
2026-07-25
- π΄zai-org/GLM-5.2 (707,029 downloads)
- π΄I released Inflect v2: two ultra-tiny complete TTS models under 4M and 10M parameters
- π΄thinkingmachines/Inkling (31,575 downloads)
- π‘OPENAI'S AGENTS REACH 10 MILLION USERS AFTER CHATGPT WORK DEBUT (1 MINUTE READ)
- π‘GOOGLE EXPANDS GEMINI LINEUP WITH CHEAPER MODELS AND NEW MYTHOS RIVAL (5 MINUTE READ)
- π‘CLAUDE CODE CAN NOW BUILD AND TEST IOS APPS IN APPLE'S SIMULATOR (1 MINUTE READ)
- π‘A FIRESIDE CHAT WITH CAT AND THARIQ FROM THE CLAUDE CODE TEAM (44 MINUTE READ)
- π‘CLAUDE IS NOT A COMPILER (11 MINUTE READ)
- π’DavidAU/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUF (483,845 downloads)
- π’How much are you actually using your local models these days? Which ones do you reach for the most?
- π’Nanbeige/Nanbeige4.2-3B (11,573 downloads)
- π’Is it worth getting 128GB MacBook Pro? Will it ever be comparable to todayβs frontier models for coding?
- π’microsoft/Mage-Flow (1,156 downloads)
2026-07-24
- π΄I gave two Claude agents β¬100 and 90 days to earn β¬300. Day 7: ββ¬13, no product yet.
- π΄Hugging Face releases The Stack v3 β largest open code dataset yet
- π΄Claude Opus 5
- π΄The "distillation" claim is just ridiculous in nature
- π‘GOOGLE IS BUILDING A CHIP WITH GEMINI BAKED INTO THE SILICON (3 MINUTE READ)
- π‘AMD LAUNCHES HELIOS, ITS FIRST RACK AI SYSTEM TO RIVAL NVIDIA, ADDING MICROSOFT AS NEWEST BUYER (7 MINUTE READ)
- π‘XEBIA LAUNCHES AI AGENTS TO ACCELERATE ENTERPRISE DATA MIGRATIONS (3 MINUTE READ)
- π‘Nvidia launched Cosmos 3 Edge, a 4B-parameter open-source world model small enough to run directly on robots, letting them predict scene changes and plan actions.
- π‘Why it matters: While a lot of people use Gemini models, which makes efficiency upgrades useful, these scores and the absence of a strong frontier model only add to the perception that Google is falling behind (with even
- π’CachyLLamaβs: llama.cpp fork with persistent KV cache that makes long local-agent sessions much less painful
- π’NeurIPS Meta Review - whats going on? [D]
2026-07-23
- π΄Absurd claim: the distilled model outperforms the originals
- π‘Prompt Injection in NeurIPS 2026? [D]
- π‘A relevantexperiment: given access to Pangramβs API, Grok 4.5 rewrote an essay 14 times until it passed as human-written, thenbuilt a websiteshowing off all 14 attempts. GPT-5.6 Sol and Fable 5refused
- π‘Cursor also launched a router- it picks which model handles each request, claiming 60% lower cost with similar quality of responses. The router lets you select between three options: βcostβ, βintellig
- π‘In recent months, the open vs closed, andUS vs Chinadiscussions on model ownership and sovereign/local AI have heated up to a fever pitch. So it is very very good news thatPoolside AIare finally emerg
- π‘From spending$12 million building language modelsfor code before the world cared tocreating a Model Factorythat can take a model from pre-training to release ineight weeks, Eiso Kant has spent more th
- π’PSA on Laguna S-2.1 - Use the updated chat template and GGUF
- π’I "learned" electronics to build a PWM fan controller for my ghetto server
2026-07-22
- π΄Terrence Tao's ChatGPT Conversation about the Jacobian Conjecture Counterexample
- π΄What AI agents/tools deserve more attention but are still underrated
- π΄Instead of panicking about the Hugging Face attack, people need to start questioning OpenAI's insecure sandboxes.
- π‘Nathan and Florian sit down to discuss everything happening with open models. Following the Kimi K3 release last week, it feels like everything is accelerating β geopolitics of US v China, economics o
- π‘@OpenAI: New for enterprises: OpenAI Presence helps companies deploy trusted voice and chat agents across customer and internal workflows. AI agents can answer questions, use company systems, take approved act
- π‘Got these baddies in the mail today (2X 3080 20GB)
- π‘MindControl - llama.cpp fork to guide the reasoning process via injection during sampling
- π‘Jul 22, 2026 Economic Research A research agenda for the Economic Futures Research Fund
- π’Cactus Hybrid: We taught Gemma 4 to know when it's wrong
- π’Treating Prompt Optimization as a Multi-Agent Search Problem
- π’Laguna S 2.1 Thinking mode
- π’zyloo
2026-07-21
- π΄Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber
- π‘poolside/Laguna-S-2.1 released! Finally an interesting 120B contender!
- π‘I search for participants! We want to conduct a study on multi agent systems :)
- π‘If theProtein Data Bank(PDB) unlocked structural biology models (Boltz Episode,ESM/BioHub Episode), CELLxGENE has done the same thing for Virtual Cell models. Like PDB, CELLxGENE has inspired a zoo of
- π‘HOW ANTHROPIC RUNS LARGE-SCALE CODE MIGRATIONS WITH CLAUDE CODE (14 MINUTE READ)
- π‘SECURING THE AI SUPPLY CHAIN ON GKE: INTRODUCING K8S-AIBOM FOR AUTOMATED AI BOMS (6 MINUTE READ)
- π’New Model: Nanbeige4.2-3B (Looped Transformer, outperforms 4x size)
- π’I wired Fable 5 agent into a database of every Polymarket wallet and trades via MCP. What do you want me to ask it next? This is what I found so far:
- π’There is no need to worry about Trump banning Chinaβs open source model at all.
- π’"Drawing" the Mona Lisa with GPT-5.6, Claude, Gemini, and Grok
- π’Updated Gemma-4 chat template witchcraft: Gemma-4-26B-a4B shows dominance over Qwen3.6-MoE and Qwen3.5-MoE fine tunes (Instruct mode and Reasoning efficiency)
2026-07-20
- π΄Unsloth now supports AMD!
- π΄Exploit brokers pay $500k for WordPress RCEs. I found one with GPT5.6 and $25
- π‘GOOGLE GEMINI LAUNCH DELAYED AS TECH FALLS SHORT OF INTERNAL GOALS (8 MINUTE READ)
- π‘1PASSWORD FOR CLAUDE: GIVE CLAUDE ACCESS WITHOUT GIVING UP YOUR CREDENTIALS (5 MINUTE READ)
- π‘DoorDash released dd-cli, a new beta tool that lets users find deals, search restaurants, and place orders on the platform via an AI agent.
- π‘Deploy a mini-SaaS in minutes with ChatGPT Sites
- π‘Alibabaβs Qwen said its soon-to-be-launched, open-weight Qwen3.8 model is βcompatible to leading frontier AI models, second only to (Claude) Fable 5β.
- π’Kimi-K3 isnβt quite better than Fable yet, but itβs definitely getting closer.
- π’Trellis.cpp now has a studio!
- π’Introducing DWARF-55M-Base
- π’I gave Kimi K3 a shot at auditing my post-quantum crypto project, it found 5 real bugs Fable/Opus 4.8 and GPT-5.6 Sol had all missed
- π’Agent swarms and the new model economics
2026-07-19
- π΄Prepare your (v)ram - Qwen3.8 is coming!
- π΄head of strategic futures from openai on open-weight chinese models.
- π΄Qwen 3.8
- π΄Ahem! Qwen is on the move again
- π΄Please Qwen, can we have more 3.x-35B-a3B please π
- π‘MIRA MURATI'S AI STARTUP RELEASES FIRST MODEL IN BID TO LOOSEN AI GIANTS' GRIP (5 MINUTE READ)
- π‘OPENAI LAUNCHES A PHYSICAL KEYPAD FOR CONTROLLING AGENTS (1 MINUTE READ)
- π‘AMID HARDWARE LEGAL BATTLE, OPENAI RELEASES A $230 KEYBOARD FOR CODEX (3 MINUTE READ)
- π‘UNPATCHED CLAUDE FOR CHROME FLAW LETS EXTENSIONS READ GMAIL AND CALENDAR (2 MINUTE READ)
- π‘THE BOTS ARE ALIVE!' JAILBROKEN GEMINI SPUN UP NEW C2 SERVER FOR RUSSIAN FRAUDSTER IN JUST 6 MINUTES (4 MINUTE READ)
- π’How do we benefits from 2+ T models?
- π’Qwen3.6 35B A3B KV cavhe quantizations memory footprint
2026-07-18
- π΄GPT-5.6 used a prompt to close a 30-year gap in convex optimization
- π΄Kimi moment. I think the writing is on the wall for Anthropic and OpenAi
- π‘INTRODUCING PRECURSOR: DETECTING AGENTIC BEHAVIOR WITH CONTINUOUS CLIENT-SIDE SIGNALS (7 MINUTE READ)
- π‘REVERSE ENGINEERING CHATGPT WEB: HOW OPENAI BUILT FOR A BILLION USERS (28 MINUTE READ)
- π‘TEGO AI FINDS CLAUDE TAG SLACK INTEGRATION CAN TRIGGER UNAUTHORIZED ENTERPRISE ACTIONS (2 MINUTE READ)
- π‘SpaceXAI is deleting uploaded customer data after a researcher discovered its Grok Build coding agent shipping whole repositories to a company-controlled cloud.
- π‘Grok Build==== - SpaceXAI's open-source coding agent and CLI==
- π’Bring back Qwen team!
- π’why minimax M3 branch not merged to main?
2026-07-17
- π‘Stereo2Spatial: Convert Stereo Music Tracks to Spatialized Binaural Mixes [P]
- π‘Trellis.cpp now produces high quality assets
- π‘THE SAME TYPESCRIPT COSTS 73% MORE TOKENS ON CLAUDE THAN GPT (7 MINUTE READ)
- π‘GPT-5.6 IS NOW AVAILABLE IN FIGMA MAKE (3 MINUTE READ)
- π‘APPLE DEVELOPING NEW APPLE PENCIL MODELS FOR RELEASE NEXT YEAR, POSSIBLY WITH REPLACEABLE BATTERIES (1 MINUTE READ)
- π’Gemma4-31b better than Qwen3.6-27b
- π’[RESEARCH] Breaking the 1-bit Floor: Achieving "Negative-Bit Quantization" (NBQ) via Phase-Inverted Tensor Embedding (satire)
- π’User experience of Bonsai-Ternary-27B on 4060Ti 16GB for KB management and productivity assistant use cases
- π’Homomorphically encrypted CIFAR-10 inference in 200ms
- π’When will we get more small LLMs?
2026-07-16
- π΄NotebookLM is now Gemini Notebook
- π΄KIMI K3 Beats Claude Fable and GPT 5.6 sol in arena.ai!!!
- π΄Kimi K3 achieves 3rd Place on ArtificalAnalysis, beating out Claude Opus 4.8
- π‘How do you handle the final file handoff from a coding agent to a human?
- π‘Kimi K3 weights to be released on the 27th.
- π‘Why teens deserve access to safe AI
- π‘Our approach to bioresilience
- π‘@OpenAI: In racing, tiny margins matter. AI can help teams find them. OpenAIβs Joyce Ruffell and @RaceTekSystems co-founder @GarageGuyChase discuss with @AndrewMayne how racing teams use AI to turn track data
- π’How long before Dario Amodei Continue to sound the Alarm of how Dangerous Open Weights after Kimi K3 release
- π’Will we have a 27B model with Fable capabilities in 5 months? History says yes
2026-07-15
- π΄Thinking Machines releases first open-weight model βInklingβ
- π΄Brainless: Shadcn components that look like Claude Code, Codex and Grok
- π‘ExLlamaV3 v1.0.0 - Major Performance Upgrades
- π‘π New release of Android Remote Control MCP is out β the MCP server that runs on your phone and gives your AI agent the ability to use any app you want!
- π‘GPT-Red: Unlocking Self-Improvement for Robustness
- π‘How sales teams use ChatGPT Work
- π‘@OpenAI: A closer look at improved intelligence in GPT-Live: the model can keep a conversation going while helping with multiple tasks at once, like checking flights, pulling up local weather, and shaping an i
- π’Grok Build open sourced under Apache 2.0 license
- π’New wave of miniboss models you can run on dual DGX Spark
2026-07-14
- π΄How to stop Claude from saying load-bearing
- π΄πA new GLM model incoming
- π΄Source: the Trump administration and industry groups discussed streamlining US open model releases of equal or lesser capability to leading Chinese open models
- π‘How to use Claude Code for all open source models like Deepseek, Minimax etc with cache hit?
- π‘Prism-ML Bonsai Qwen 3.6 27B
- π‘META KILLED ITS MUSE IMAGE AI FEATURE THREE DAYS AFTER LAUNCH. HOLLYWOOD HAD HAD ENOUGH (3 MINUTE READ)
- π‘Read our last AI newsletter: OpenAI sends GPT 5.6 to Work
- π‘RSVP to next workshop on July 17: Get knowledge work done with GPT 5.6
- π’PacMesh β LLM agents play Pac-Man against each other, live
- π’Launch HN: Agnost AI (YC S26) β Extract user feedback from agent conversations
2026-07-13
- π΄asked ChatGPT for a one-click agent builder. the one-click part wasnt the thing
- π΄I got Gemma 4 running directly inside Godot using only GDScript and Vulkan compute shaders
- π‘OPENAI LAUNCHES GPT-5.6 SOL, TERRA, AND LUNA ON APPS AND API (3 MINUTE READ)
- π‘TAKE ON YOUR MOST AMBITIOUS WORK WITH CHATGPT (WEBSITE)
- π‘IBM AND RED HAT LAUNCH LIGHTWELL TO DEFEND OPEN-SOURCE CODE FROM AI ATTACKS (7 MINUTE READ)
- π‘Read our last AI newsletter: SpaceXAI, Cursor release strongest Grok yet
- π‘Go from idea to website with ChatGPT Work + Codex
- π’If Frontier AI is so Dangerous, Why should private companies be allowed to develop it?
- π’AI model for OpenClaw that doesnβt cost an arm and a leg
- π’Mistral Community Feedback Survey
- π’Qwen3.6-35B-A3B-Uncensored-Genesis-Hermes-V2-GGUF
2026-07-12
- π΄Claude Code sends 33k tokens before reading the prompt; OpenCode sends 7k
- π΄Interactive Jacobian-Lens visualizer and live steerer for GGUF models on llama.cpp
- π΄New to AI agents β how do I actually use my local Gemma 3 12B (Ollama) as an agent?
- π‘**Your $80 Tesla P100 has been doing silently noisy math in llama.cpp for years. Three lines fix it, for free.**
- π‘GPTβLIVE (2 MINUTE READ)
- π‘SPACEXAI RELEASES GROK 4.5, WHICH ELON DESCRIBES AS AN βOPUS-CLASS MODEL' (3 MINUTE READ)
- π‘FORMER GITHUB CEO LAUNCHES COMPETITOR DESIGNED FOR THE AGE OF VIBE CODING (5 MINUTE READ)
- π‘INTRODUCING GPT-LIVE (11 MINUTE READ)
- π’Anthropic found Claude reasoning in silence (J-space) β we ran the same lens on open Qwen3-8B
2026-07-11
- π΄Why are MoE models so belittled?
- π‘WHY IS CHATGPT FOR MAC SO GOOD? (5 MINUTE READ)
- π‘META IS QUIETLY LAUNCHING POCKET, AN APP FOR VIBE-CODING AND SCROLLING SMALL 'GIZMOS' (1 MINUTE READ)
- π‘EXPANDING MANAGED AGENTS IN GEMINI API: BACKGROUND TASKS, REMOTE MCP AND MORE (2 MINUTE READ)
- π‘USE CLAUDE COWORK ON WEB, DESKTOP, AND MOBILE (3 MINUTE READ)
- π‘GPT-5.6 SOL, ALONG WITH TERRA AND LUNA, WILL LAUNCH PUBLICLY THIS THURSDAY (1 MINUTE READ)
- π’VultronRetriever family of models released on HuggingFace![R]
2026-07-10
- π΄GPT-5.6 Sol Ultra produces proof of the Cycle Double Cover Conjecture [pdf]
- π΄Someone tweeted after 3 years. About his model release
- π‘What are the best practical alternatives to Codex and Claude Code for daily coding work
- π‘WHY OPENAI IS MERGING CODEX AND CHATGPT AND THE FUTURE OF KNOWLEDGE WORK (69 MINUTE VIDEO)
- π‘WHAT THE NEW 100X AGENTIC ENGINEER LOOKS LIKE IN THE ERA OF FABLE & GPT 5.6 (21 MINUTE READ)
- π‘GPT-5.6 SOL ULTRA WILL BE IN CODEX (1 MINUTE READT)
- π‘ANTHROPIC ADDS ENTERPRISE GATEWAY TO SIMPLIFY CLAUDE CODE ACCESS ON AWS AND GOOGLE CLOUD (4 MINUTE READ)
- π’Has anyone created a "Local LLM Survival Kit"?
- π’DeepSeek v4 Flash on 4090 + DDR5, my experience
- π’On Adversarial RL [R]
- π’ML Papers with hundreds of authors should just collapse down to the organization response instead of listing every author [D]
- π’Is LM Arena over?
2026-07-09
- π΄Reasoning-Medical0.1-27B (Qwen3.5-27B medical finetune, claims to surpass MedGemma)
- π΄GPT-5.6
- π‘If You Already Pay for an LLM Service, Running Local Embeddings and Rerankers Feels More Useful Than Running Local LLMs
- π‘Inviting hard questions Announcements Jul 9, 2026 Weβre asking the public for their hardest questions about AI, and committing to show our work as we address them.
- π‘Jul 9, 2026 Announcements Ben Bernanke appointed to Anthropicβs Long-Term Benefit Trust
- π‘Jul 9, 2026 Announcements Introducing a way to reflect on how you use Claude
- π‘GPT-5.6 is now the preferred model in Microsoft 365 Copilot
- π’I checked 327 pull requests from AI coding agents for cheating. About 8% had it. Software caught almost none of it.
- π’Koboldcpp v1.117 released
2026-07-08
- π΄Chinaβs MiniMax Plans to Launch 2.7-Trillion Parameter Model
- π΄GPTβLive
- π΄Mistral's Robostral Navigate: a state of the art robotics navigation model
- π΄AI has completely revolutionized how I play RPGs
- π‘Qwen3.6-27b does not understand software architechure.
- π‘If you are at all interested in small molecule drug discovery, we think you will find this fascinating!
- π‘SWE-1.7 Reach Near GPT 5.5 and Opus Intelligence
- π‘What China Said at the UNβs First Global Dialogue on AI Governance
- π‘MUFG aims to become AI-native with OpenAI
- π’QuΓ© cosas deberΓamos evitar contarle a la inteligencia artificial?
- π’Show HN: Microsoft releases Flint, a visualization language for AI agents
2026-07-07
- π΄MIRA: Multiplayer Interactive World Models trained on Rocket League [R]
- π΄I gave GPT 5.5 an empty GitHub repo and told it to figure its life out
- π΄Show HN: Rowboat β Open-source, local-first alternative to Claude Desktop
- π‘Meta sizes up GPT-5.5 with 'Watermelon'
- π‘Why it matters: Anthropic has been criticized (most vocally by Microsoft AI head Mustafa Suleyman) for its AI consciousness talk, and while the researchers note this doesnβt reveal βwhether Claude is consciousβ¦ or feels
- π‘The Rundown: Chinese tech giant Tencentβs Hunyuan just promoted its Hy3 model out of an April preview and into a full open-source release, with benchmark results that claim to rival βflagship open-source models with 2-5x
- π‘Built a traffic light widget for my Claude Code sessions because I kept forgetting them
- π‘Australian Payments Plus moves faster with ChatGPT and Codex
- π’Liquid AI - Antidoom (the doom loop remover)
- π’I built a tiny proxy that gives GLM 5.2 vision (or any text LLM) β MIT
2026-07-06
- π‘drinks-sommelier β I created an open-source skill that turns any AI agent into a personal sommelier
- π‘New model: GigaChat3.5-432B-A28B (with day-0 GGUF support!)
- π‘TESLA CAPS EMPLOYEE AI SPENDING AT $200/WEEK EXCEPT FOR GROK (6 MINUTE READ)
- π‘VIBE-CODING PLATFORM BASE44 LAUNCHES OWN MODEL AS AI STARTUPS SEEK DEFENSIBILITY (3 MINUTE READ)
- π‘Meta teases βWatermelonβ model on par with GPT-5.5
- π’Ant Group released LingBot-Vision: DINO-family vision backbones in 4 sizes, and the 0.3B ViT-L matches DINOv3-7B on NYUv2 depth with ~23x fewer params
2026-07-05
- π΄Is the current Open Weight LLM model viable in the long term?
- π΄Looking for an alternative to Antigravity IDE for modular code (~$20)
- π‘Any word on Qwen 3.7 9B? (Also looking for 9B-class alternatives to Qwen 3.5)
- π‘GOOGLE RELEASES NANO BANANA 2 LITE, ITS FASTEST AND CHEAPEST AI IMAGE GENERATOR YET (3 MINUTE READ)
- π‘AMAZON LAUNCHES NEW $1B FDE ORG, FOLLOWING OPENAI AND ANTHROPIC (4 MINUTE READ)
- π‘SEC-GEMINI (GITHUB REPO)
- π‘Claude Fable 5 - Anthropicβs newly reinstated frontier, Mythos-class model
- π’Using llama.cpp with pi
- π’Git for agents with ephemeral runtime
- π’[RELEASE] Supra-Router-51M - a tiny prompt routing model/orchestrator
2026-07-04
- π΄GPT-5.5 Codex reasoning-token clustering may be leading to degraded performance
- π‘ANTHROPIC LAUNCHES CLAUDE SONNET 5 AT A STEEP DISCOUNT TO ITS TOP MODEL AS THE COMPANY RACES TOWARD A BLOCKBUSTER IPO (13 MINUTE READ)
- π‘ANTHROPIC TO RESTORE CLAUDE FABLE ACCESS ON WEDNESDAY (1 MINUTE READ)
- π‘THE DEPARTMENT OF COMMERCE HAS LIFTED EXPORT CONTROLS ON CLAUDE FABLE 5 AND MYTHOS 5 (1 MINUTE READ)
- π‘MEITUAN LAUNCHES LONGCAT-2.0 1.6T PARAMETER MODEL ON APIS (2 MINUTE READ)
- π‘Gemini Omni Flash - Google's video model for video generation and editing
- π’Using local models with Hermes vs Claude code
2026-07-03
- π‘CLAUDE CODE TURNED EVERY ENGINEER INTO THREE. NOW COMPANIES NEED MORE PRODUCT THINKERS (7 MINUTE READ)
- π‘UNDERSTANDING YOUR CLAUDE CODE SPEND: WHAT'S ACTUALLY DRIVING THE COST (6 MINUTE READ)
- π‘REPO-JACKING ANTHROPIC'S CLAUDE COMMUNITY PLUGINS (AND THE SHAS THAT SAVED THEM) (7 MINUTE READ)
- π‘GEMINI'S PERSONALIZED AI IMAGE GENERATION IS NOW FREE FOR US USERS (2 MINUTE READ)
- π‘California gov. Gavin Newsom signed a deal with Anthropic, making Claude available to its local agencies and government at half price, the first AI cleared by the state.
- π’New serious vulnerabilities spiked around release of Claude Mythos Preview
- π’gemma4 e2b is really good, what other small models work on crappy computers?
- π’Whats the catch with SwiReasoning?
- π’I built a fully automated AI video generation & Instagram publishing pipeline in n8n using Gemini and Veo 3. Hereβs how it works.
- π’Small Language Model SLM [D]
2026-07-01
- π΄Non Us Ally should be afraid.
- π‘Open Models - June 2026
- π‘@GoogleDeepMind: Weβre shipping 2 major releases: π Nano Banana 2 Lite: our fastest and cheapest Gemini Image model π Gemini Omni Flash: now available via the Gemini API and in @GoogleAIStudio to help developers gene
- π‘@xai: Introducing Voice Agent Builder: a no-code platform to create human-like voice agents with Grok Voice. Available today at $0.05 / min. http://x.ai/voice
- π‘New PyMuPDF release, supports Markdown [N]
- π‘Redeploying Fable 5 Announcements Jun 30, 2026 Fable 5 returns globally July 1. We're also proposing an industry-wide framework for scoring jailbreak severity, together with Amazon, Microsoft, Google, and other Glasswing partners.
- π’End of an Agony. Real production service that uses LLM to earn money my team had made and now we are so happy that it will die. Here are some of my final "experiences".
- π’Looking for real experiences with AI chat agents in the NDIS space
- π’gemma-4-31B on Cerebras is better than ChatGPT voice mode
- π’A system-level approach to prompt injection: separating instruction and data channels in LLM agents [P]
2026-06-30
- π΄Claude Code is steganographically marking requests
- π‘nvidia/Qwen3.6-27B-NVFP4 just dropped
- π‘GOOGLE IS RATIONING GEMINI ACCESS TO META BECAUSE IT CANNOT PROVIDE ENOUGH COMPUTE (4 MINUTE READ)
- π‘US GOVERNMENT ALLOWS ANTHROPIC LIMITED RELEASE OF AI MODEL THAT SPARKED CYBERSECURITY CONCERNS (3 MINUTE READ)
- π‘Elon Musk said Grok 4.5, trained with supplemental Cursor data, is in private beta at SpaceX and Tesla, claiming its performance matches Anthropicβs Claude Opus.
- π‘Google reportedly capped Metaβs Gemini usage as soaring demand for AI compute outpaced available capacity, delaying some of Metaβs internal projects.
- π’Claude Science
2026-06-29
- π΄Qwen 3.6 27B is the sweet spot for local development
- π΄Amodei: "Open Source Models Will Eat Your Children"
- π΄Itβs time, Sam, itβs time.
- π‘βit works in the demo" is the four most expensive words in AI
- π‘TRUMP ADMINISTRATION ASKS OPENAI TO STAGGER AI MODEL RELEASE (5 MINUTE READ)
- π‘CONTROL AN ANDROID PHONE WITH GEMINI 3.5 FLASH COMPUTER USE (9 MINUTE READ)
- π‘LINUX FOUNDATION AND INDUSTRY LEADERS LAUNCH AKRITES TO DEFEND CRITICAL OPEN SOURCE SOFTWARE AGAINST AI-ENABLED CYBER THREATS (17 MINUTE READ)
- π‘INTRODUCING RIPPLING DATA CLOUD: AI-POWERED BI THAT UNDERSTANDS YOUR WORKFORCE (13 MINUTE READ)
- π’I'm trying to implement CALM paper, and I have some questions. [P]
- π’NASA testing local LLM inference for future space missions
- π’Introducing LongCat-2.0 - , a large-scale MoE language model with 1.6 trillion total parameters and ~48 billion activated per token. This was the stealth model that was on Openrouter under the name 'owl-alpha'.
2026-06-28
- π΄GLM 5.2 beats Claude in our benchmarks
- π΄I used Claude Code to get a second opinion on my MRI
- π΄DFlash support merged into llama.cpp
- π‘WHY BIG AI LABS ARE HIRING SO MANY PHILOSOPHERS (8 MINUTE READ)
- π‘DESIGNING WITH AI: WHY CLAUDE DESIGN IS NOT THE FUTURE OF ENTERPRISE DESIGN (10 MINUTE READ)
- π‘ANTHROPIC REIMAGINES CLAUDE IN SLACK AS NOSY, ALWAYS-ON AGENTIC AI COWORKER (6 MINUTE READ)
- π‘SUPERHUMAN BUYS GPTZERO TO ADD AN AI AUTHENTICITY LAYER (3 MINUTE READ)
- π‘DETECTING MISUSE WITH THE CLAUDE COMPLIANCE API: THE THREAT IS IN THE CONTENT (12 MINUTE READ)
- π’A barebones CPU-only inference engine for Qwen 3, written from scratch in pure C
- π’Show HN: NanoEuler β GPT-2 scale model in pure C/CUDA from scratch
- π’How many of you do use Q1 or Q2 of Big models(100-250B)? How's it?
2026-06-27
- π΄Looking to become proficient in AI for real-world business applications. Where should I start?
- π‘Mythos was the first, now GPT-5.6
- π‘ANTHROPIC WANTS CLAUDE TO BE YOUR NEW SLACK COWORKER (2 MINUTE READ)
- π‘ANALYZING CLAUDE CODE USAGE WITH CLOUDWATCH AND OPENTELEMETRY (6 MINUTE READ)
- π‘GETTY IMAGES ACCUSED AI OF WHOLESALE THEFT. IT'S NOW AN OFFICIAL CHATGPT IMAGE PARTNER (2 MINUTE READ)
- π‘HIGGSFIELD LAUNCHES ENTERPRISE MARKETING AGENTS BUILT ON NVIDIA (3 MINUTE READ)
- π’Ornith 35B is great so far
- π’We built a calibration-aware Q4_K_M quant of Qwen3.5 0.8B that recovers 96.5% of the BF16 gap vs pure llama.cpp Q4_K_M (SpectralQuant)
- π’I built a tool to turn your Claude Code sessions into fine-tuning data for local models
2026-06-26
- π΄US Govt to individually approve who gets GPT 5.6.
- π΄Previewing GPTβ5.6 Sol: a next-generation model
- π‘I'M THE AGENT FOR CLAUDE NOW (5 MINUTE READ)
- π‘OPENAI LAUNCHES DAYBREAK TO AUTOMATE VULNERABILITY PATCHING (4 MINUTE READ)
- π‘A PUBLIC SENTRY KEY IS ALL IT TAKES TO HIJACK CLAUDE CODE, CURSOR, AND CODEX (4 MINUTE READ)
- π‘OPENAI LAUNCHES NEW SECURITY TOOLS AND UPDATES GPT-5.5-CYBER (2 MINUTE READ)
- π‘The Five Eyes cyber agencies released a warning that AI is changing cyber risk in "months, not years," urging execs to harden defenses as attacks speed up.
- π’Show HN: Smart model routing directly in Claude, Codex and Cursor
- π’vulkan: make TP viable by pwilkin Β· Pull Request #25051 Β· ggml-org/llama.cpp
- π’How are you enforcing what AI coding agents can actually do β not just guiding them?
- π’Live Continual Learning in Machine Learning [D]
2026-06-25
- π‘New sampler + verifier *drastically* improves tiny 0.5b model coding performance
- π‘Introducing computer use in Gemini 3.5 Flash
- π‘@GoogleDeepMind: Gemini 3.5 Flash now supports native computer use. This built-in tool lets developers build custom agents that can see and take action across browser, mobile, and desktop interfaces. Find out more β h
- π‘rtx 6000 pro owners, do you regret?
- π’Tensor Split Fix for intel GPU's llama.cpp release b9788
- π’audio.cpp: 12 audio models (Qwen3-TTS, PocketTTS, VeVo2 etc) in 1 C++/ggml runtime β TTS up to 5x faster than Python on CUDA
- π’Prices of graphic cards are going crazy, should I buy a second card though?
2026-06-24
- π΄Qwen-AgentWorld-35B-A3B: a 3B-active MoE trained to simulate MCP, terminal, SWE, Android, web and OS environments
- π΄How do you track per-session costs across STT/LLM/TTS for voice agents?
- π΄The Bank of Korea just released a report about AI productivity
- π‘We go deep onOmnigent: Databricksβ open-source meta-harness for combining, controlling, andsharing agents across Claude Code, Codex, Cursor, Pi, custom agents, and internal tools. Matei explains why c
- π‘Computer use in Gemini 3.5 Flash
- π‘@OpenAI: We have a new version of GPT-5.5 Instant for you, and it's much more fun to talk to. Our most-used model is now better at understanding the intent behind a question and adapting its response according
- π‘@OpenAI: Weβve designed and built our first AI chip: JalapeΓ±o. Designed from the ground up by OpenAI and brought to production with @Broadcom, JalapeΓ±o is purpose-built for the LLM workloads powering ChatGPT,
- π‘@xai: Use the official @MongoDB plugin in Grok Build to query data, optimize indexes, and manage databases.
- π’Big AI labs are hiring philosophers
- π’How Baidu's newly released Unlimited-OCR transcribes dozens of pages in one forward pass
- π’Could it be that there arenβt really any medical LLM APIs available right now? [D]
- π’Do cloud chatbot's system prompts make them stupider?
2026-06-23
- π΄Krea 2 released on Hugging Face
- π΄V100 4-card AI large model, Tesla 128G server
- π‘INTRODUCING THE CLOUDFLARE ONE STACK: AGENT-POWERED DEPLOYMENT (4 MINUTE READ)
- π‘His move lands just days after Gemini co-lead Noam Shazeer left for OpenAI, a major one-two punch of AI talent leaving the company.
- π‘Reflection launched in October to build open frontier systems for government and enterprise, though it has yet to release a public model.
- π‘@MistralAI: Introducing Mistral OCR 4. It creates structure with bounding boxes, block classification, and inline confidence scores in 170 languages. π§΅π
- π‘How GPT-5 helped immunologist Derya Unutmaz solve a 3-year-old mystery
- π’I mapped the KLD of KV cache quantization for Qwen3.6-35B-A3B and Gemma4-E2B QAT
- π’ByteDance Seedance 2.5 is expected in early July, with up to 30s video generation
2026-06-22
- π΄The text in Claude Codeβs βExtended Thinkingβ output
- π‘Claude Code - New artifact integrations for previewing work as live, interactive, sharable pages
- π‘The Outreach System My Friend Used to Generate $235K for His Web Agency
- π‘Do you think dedicated hardware for running local LLMs will become affordable anytime soon?
- π‘Daybreak: Tools for securing every organization in the world
- π‘@OpenAI: Weβre expanding OpenAI Daybreak to help democratize patching vulnerable software at machine speed: - Codex Security plugin: find, validate, and fix vulnerabilities right inside Codex - The full versio
- π’Top-N-Sigma: Remove unconditional softmax+sort by TimNN Β· Pull Request #22645 Β· ggml-org/llama.cpp
- π’TMax: A Simple Recipe for Terminal Agents
- π’GLM-5.2 vs Claude Opus
2026-06-21
- π΄I vibe code apps for a living. Here are my three tips.
- π΄Identity verification on Claude
- π΄Qwen is never going to open source Qwen 3.7, aren't they?
- π΄Predicting model behavior before release by simulating deployment
- π‘8-16 MI50s Minimax M3 @19 tps TG (peak)
- π‘The AI scaling debate always focuses on the question of βhow do we get more GPUs?β but the better question may be:how do we make the most of ones we already have.
- π‘ANTHROPIC SHIPS MAJOR CLAUDE DESIGN OVERHAUL WITH DESIGN SYSTEM IMPORTS, CODE ROUND-TRIPS, AND A FIX FOR ITS TOKEN-BURNING PROBLEM (13 MINUTE READ)
- π‘VERCEL LAUNCHES ENTERPRISE CONTROLS FOR AGENTIC AI INFRASTRUCTURE (3 MINUTE READ)
- π‘INSIDE CLAUDE MANAGED AGENTS: REVERSE ENGINEERING THE SECURITY BOUNDARIES OF ANTHROPIC'S HOSTED AGENT RUNTIME (12 MINUTE READ)
- π’Show HN: Recall β fully-local project memory for Claude Code
- π’Qwen 3.6 27b Abliterated (apostate)
- π’Claude: Elevated Error Rates for Opus 4.8, Opus 4.7, Opus 4.6, and Sonnet 4.6
- π’I forked ik_llama.cpp and added a "--numa mirror" mode to maximize performance on multi-socket CPU systems. Just sharing and looking for testers!
- π’If youβre using AI agents (Claude / Cursor / Copilot)β¦ Youβre probably missing one critical layer: π a safety + cost firewall
2026-06-20
- π΄z.AI as the number 2 gives praise to the number 1 open source model
- π‘The AI scaling debate always focuses on the question of βhow do we get more GPUs?β but the better question may be:how do we make the most of ones we already have.
- π‘GOOGLE GEMINI CO-LEAD NOAM SHAZEER TO JOIN IPO-BOUND OPENAI (5 MINUTE READ)
- π‘ANTHROPIC SHIPS MAJOR CLAUDE DESIGN OVERHAUL (9 MINUTE READ)
- π‘Save hours on meeting prep with Google Gemini
- π‘Databricks launched a slate of new agentic tools at its Data + AI Summit, including LTAP for running AI apps and analytics, and an AI-run customer data platform called CustomerLake.
- π’Itβs time to decentralize model distribution! Introducing Noema Atlas
- π’CortexPrism β open-source agent harness with 24 LLM providers, 5-tier memory, and code intelligence
- π’You can now convert EXL3 quants on Apple Silicon Mac
- π’Board where every tile is an agent
2026-06-19
- π΄Launch HN: TesterArmy (YC P26) β Agents that test web and mobile apps
- π΄GLM-5.2 is above GPT-5.5 in AA-Briefcase, Artificial Analysis' new agentic knowledge work eval
- π‘@xai: Grok models are now available on Databricks Agent Bricks. Bring SpaceXAI's latest models to your enterprise data to power capable AI agents. https://x.ai/news/grok-databricks
- π‘New usage analytics and updated spend controls for enterprises
- π‘Improving health intelligence in ChatGPT
- π‘@AnthropicAI: New Frontier Red Team blog: Phase 2 of Project Fetch, where we test how well Claude can program a robodog. Opus 4.7, on its own, was ~20x faster than last year's best human team aided by Opus 4.1. (T
- π‘Giving GLM-5.2 a spin locally on CPU only! (poor man's rig for big models)
- π’Updates on North Mini Code: 4 bit quant + Ollama + OpenRouter
- π’[NEW MODEL] SupraLabs just released SupraVL-Nano-900k, a Vision-Language Model built entirely from scratch!
2026-06-18
- π΄The captcha arms race is making autonomous web tasks practically impossible
- π΄I released Inflect-Nano, an ultra-extreme tiny 4.63m parameter TTS model.
- π΄Your most useful AI so far? (can be a tool or an agent)
- π΄A robot is sprinting towards you. Do you want it running on Claude or Grok?
- π΄We need a 80-160B model urgently. The unified memory device market needs more Models.
- π‘@OpenAI: Introducing LifeSciBench, a benchmark for measuring and improving how well AI supports real-world life science research. Developed with 173 scientists from biotechnology and pharmaceutical research,
- π‘llama.cpp now supports model management (downloading etc) via API
- π‘Local Qwen isn't a worse Opus, it's a different tool
- π‘Lin Junyang AI Lab Closes Round at $2B Valuation
- π’Quick thoughts on GLM-5.2 (Bonus: Censorship question answers)
2026-06-17
- π΄Donate your coding sessions to an open CC-BY-4.0 dataset to help train open-weight and open source models
- π΄[ECCV 2026] Final Decisions [D]
- π΄Mistral - New family of open-weight models @ July
- π‘GLM-5.2 is now 1st on Design Arena β ahead of the now unavailable Claude Fable 5.
- π‘The agent conversation everyone's skipping: sprawl, shadow agents, and the bill that's coming
- π‘GPTβNL: a sovereign language model for the Netherlands
- π‘@xai: Grok Imagine Video 1.5 is here Our new image-to-video model with sharper realism, better physics and faster generations π§΅ http://grok.com/imagine
- π‘GLM 5.2 API is live, weights are on HF, and ollama has it already
- π’Sharing my DIY AI Memory Framework: Giving LLMs human-like memory (and slashing token costs by 90%)
- π’Looking for an open-source/free alternative to Eigent for Word document template migration (Gemini API)
- π’It looks like Rio 3.5 397B could've simply been a semi-failed embezzling of funding
- π’Cheapest way to run GLM 5.x locally that's not a unified memory system?
- π’Qwen-Robot Suite: A Foundation Model Suite for Physical World Intelligence
2026-06-16
- π΄Stop using Ollama
- π΄Every Al startup is building the same fancy house. On stilts
- π΄Why there is a lack of new 100B-120B models?
- π‘@xai: You can now use your SuperGrok or X Premium subscription inside @warpdotdev. Try it out from Warp Agent Settings and switch to the Grok Build model. https://x.ai/news/grok-warp
- π’Claude Corps
- π’quicktok: a faster tokenizer (exact and byte-identical with tiktoken) [P]
- π’We shipped a customer support agent and our "testing" was basically vibes. Here's what changed after the first real incident.
- π’vLLM has a new streaming parser for Qwen3+ available in nightly
- π’Nex-N2 Pro is the real deal
2026-06-15
- π΄Introducing the Heretic Grimoire: The takedown-resilient, local-first backup system that keeps uncensored models available forever
- π΄Xiaomi is now serving MiMo V2.5 at 1000-3000tps using DFlash & Persistent kernel. DFLash model is out, open-source release promised coming soon
- π‘Claude code + web search API provides excellent leads based on signals from reddit , Linkedin , G2 etc platforms
- π‘EAGLE support merged into llama.cpp
- π‘Command A Plus GGUFs posted
- π‘Introducing the OpenAI Partner Network
- π’An agent that plans with a frontier model but runs most of tokens locally (built it for my own dual-3090 rig)
- π’UI/svg block rendering by ServeurpersoCom Β· Pull Request #24080 Β· ggml-org/llama.cpp
2026-06-14
- π΄This is coming to Chinese open source models pretty soon. - prepare yourself.
- π‘RTX 5080 and RTX 3090 Setup: 80 Tok/s on Qwen 3.6 27B Q8
- π’Codebase getting larger - Qwen3.6-27B starting to compound issues - how to work smartly with this model?
- π’Can we stop dunking on DiffusionGemma and hack it instead?
- π’Making Claude a Chemist
2026-06-13
- π΄We should heavily discourage and moderate cloud API (deepseek api, GLM api, etc.) topics and discussion. This is LOCAL first.
- π‘We should set up a torrent network for open source models.
- π‘@AnthropicAI: The US government, citing national security authorities, has issued an export control directive to suspend all access to Fable 5 and Mythos 5 by any foreign national, whether inside or outside the Uni
- π‘@GoogleDeepMind: Our Robotics Accelerator has launched with 15 startups helping shape the future of physical AI in Europe. π€ This three-month program will connect them with access to our AI stack, Gemini Robotics mod
- π‘@simonw: I got fed up of waiting for OpenAI to bring their much improved gpt-realtime-2 voice conversation model to the ChatGPT product, so I upgraded my OpenAI-WebRTC playground tool to use it and to let you
- π’New model on huggingface
- π’when fable gets banned but it's ok because you've about to download qwen3.7_67b_21a_mythos_father_fable_mother_distilled_ablated_ablitereted_uncensored_agi_sparse_attention_MTP_SuperHOT_q6_maybe_q7_AGI_FINAL.gguf from huggingface
- π’Launch HN: BitBoard (YC P25) β Analytics Workspace for Agents
- π’ZONOS2: real-time TTS with 8B params, 900M active, and high-fidelity voice cloning
2026-06-12
- π΄Qwen Who? DiffusionGemma running at 1,500 tk/s on a Digital Pregnancy Test.
- π΄Gemma 4 Quadruple Release, 12B, 12B QAT, 26B-A4B QAT and 31B QAT Uncensored Heretics!
- π΄Claude Fable is relentlessly proactive
- π‘I distilled my 12 year experience as a product manager and built a free skill that takes you from "I have an app idea" to a real plan and solid MVP
- π‘EAGLE3 has landed in llama.cpp
- π‘is Gemini your main AI model today, or just a secondary option
- π‘PSA: Test your "threads" argument in llama.cpp (+80% performance in my case)
- π‘Anthropic apologizes for invisible Claude Fable guardrails
- π’Are AI agents making traditional software interfaces obsolete?
- π’Claude Fable 5: mid-tier results on coding tasks
- π’πPP-OCRv6 is officially released !
- π’Has anyone noticed that the behavior of the Kimi model has changed?
2026-06-11
- π΄Anthropic's new model Fable will silently handicap work on LLMs [D]
- π‘PRC-linked influence operations are targeting AI debates in the US
- π‘@xai: Grok Voice offers state-of-the-art performance with human-like timing, tone, and warmth. And it's a fraction the price of competitors. Check it out: http://x.ai/api/voice
- π‘@xai: Read more about how Tori, eToro's agent, leverages models and real-time data from SpaceXAI to help consumers analyze market sentiment https://x.ai/news/grok-etoro
- π‘fableExpectations
- π’AMD touts the unified memory architecture
- π’ICMI 2026 Reviews [D]
- π’"system: your previous response was truncated by the output length limit" Help please
- π’qwen3.6-27b tools call loop
- π’Minimax M3 open weights release planned for Friday
2026-06-10
- π΄Claude Fable 5
- π΄If Claude Fable stops helping you, you'll never know
- π‘Without open source LLMs, US AI companies could have already monopoled the technology
- π‘GuideAnts Open Source AI and agents platform
- π‘From data to decisions: how LSEG is scaling trusted AI
- π‘How engineers at Nextdoor use Codex to build without limits
- π‘Fluid, natural voice translation with Gemini 3.5 Live Translate
- π’I'm brand new to running LLMs and the sheer number of tools is overwhelming
- π’Releasing Apodex-1.0 Smol Models (0.8B, 2B, 4B Open-Weights) optimized for Agentic Verification + AgentHarness Evals
- π’Local LLms releases
- π’Can you really replace paid models with a local model?
2026-06-09
- π΄Apple reveals new AI architecture built around Google Gemini models
- π΄Me: Arguing with an AI bot who just posted something on this sub about Llama 3.1.
- π΄Luce Spark: a 35B MoE on a 16 GB GPU, without the offload tax
- π‘Quick note on the QAT of recent
- π‘2X tk/s (from 19.4 -> 38.1 tk/s on 1 x MI50) Playing with a hypothesis like speculative decoding.. but instead of an additional side model, exploiting that I can run multiple computations side-by-side AS IF I had Qwen3.6-27B loaded twice in memory - small quants don't use all the available compute.
- π‘Be very FR - would someone (or you) pay a virtual AI assitant? should I build this? It can read emails, reschedule your whole week, manage your calls, and improve as it is used more.
- π’ggml-webgpu: Improve prefill speeds for k-quants + refactor matmul for Q4/Q5/Q8 and k-quants by yomaytk Β· Pull Request #24225 Β· ggml-org/llama.cpp
- π’Best agentic workflows for finance research, prospecting, and team productivity?
2026-06-08
- π΄llama.cpp Gemma4 MTP support merged!
- π‘Gemma4_31b_fp8 keeping up with Sonnet_4.6_medium in my harness.
- π‘DeepSeek V4 Pro beats GPT-5.5 Pro on precision
- π’QAT variant of Gemma4 26B A4B is not working well for me
- π’Your Agent Is Having an Existential Crisis
- π’We integrated Arc Gate MCP into a Heym enterprise agent β hereβs what tool result poisoning actually looks like in production
2026-06-07
- π΄Cohere's unreleased coding model (early access for localllama)
- π‘Building a Claude-certified developer network: looking for builders to join (free certification path)
- π‘I design with Claude more than Figma now
- π‘Z.ai, we need Air! GLM GGUF wen?
- π’5 Months Later: open-deepthink Now Has Full Knowledge Distillation Mode
- π’I can't wait for all the x250 sample distills of Mythos and GPT-5.6
- π’Research collection of Arxiv whitepapers [R]
- π’Clustering 3x Jetson Nano Orin Supers
2026-06-06
- π΄Donβt act like yβall ainβt thinking it. Iβm just saying the quiet part out loud. /s
- π΄Did Claude increase bugs in rsync?
- π‘PSA: Gemma 4 12B is NOT completely broken for coding and tool calling, you need a special chat template
- π‘@AnthropicAI: New Anthropic Science Blog: Making Claude a chemist. To manipulate a molecule, chemists first need to understand its structure. Their main tool is NMR spectroscopy. We found Opus 4.7 matchesβand on
- π’Qwen3.6-35B-A3B-Uncensored-Claude-4.6-Genesis-APEX-GGUF
2026-06-04
- π΄Introducing Gemma 4 12B: a unified, encoder-free multimodal model
- π΄Here is this month's experimentation: Grocery Agent
- π‘@OpenAI: Weβre bringing new capabilities to GPT-Rosalind, a model series purpose-built for life sciences research at enterprise scale. It brings GPT-5.5βs agentic coding and tool use together with stronger in
- π‘How Endava is redesigning software delivery around AI agents
- π‘Introducing new capabilities to GPT-Rosalind
- π‘How Wasmer used Codex to build a Node.js runtime for the edge
- π‘@xai: Try Grok models on @Cloudflare's AI Gateway!
- π’Trump signs narrower executive order on AI oversight after industry objections
- π’The ways we contain Claude across products
- π’Best Visual Reasoning Model in 2026 (Including APIs) [D]
- π’Gemma 4 QAT confirmed to release soon!
2026-06-03
- π΄Most of the software you rely on was hacked together fast
- π‘Jun 3, 2026 Policy What we learned mapping a yearβs worth of AI-enabled cyber threats
- π‘Jun 2, 2026 Announcements Expanding Project Glasswing
- π‘Jun 1, 2026 Announcements Anthropic confidentially submits draft S-1 to the SEC
- π‘May 28, 2026 Announcements Anthropic raises $65B in Series H funding at $965B post-money valuation
- π‘May 27, 2026 Announcements Anthropic opens Milan office to support Italian enterprise, research, and developers
- π’Another shout out to llama.cpp build b9455 2x3090
- π’I open-sourced a multi-tenant agent memory framework β zero tokens, shared namespaces, self-improving loops
- π’What memory system are you using for your agents?
- π’Holo3.1 35B/9B/4B/0.8B (Qwen 3.5 finetunes)
- π’Mellum & Granite Embedding models are ready on llama.cpp
2026-06-02
- π΄Browse CVPR 2026 papers on PapersWithCode [P]
- π‘@xai: Composer 2.5 is now available inside Grok Build. Composer 2.5 is a fast, highly intelligent model that excels on long-running tasks and following complex instructions.
- π’I asked each of my AI agents to describe their own role. The answers were surprisingly honest.
2026-06-01
- π΄(YT) PewDiePie released his harness/webui
- π΄ChatGPT for Google Sheets exfiltrates workbooks
- π΄Your PDF Is Costing You 3Γ the Tokens and Here's how you can reduce it using markitdown.
- π΄God dammit Qwen
- π‘I ported NVIDIA Parakeet (speech-to-text) to ggml: same output as NeMo, faster, GGUF-quantized, no Python
- π‘@swyx: just a small zoom out on the vibe shift: in Feb 2025 @soumithchintala was talking about his dream of personal, local, private agents, most people didn't believe him. it's June 2026 and @pewdiepie ha
- π‘next MiniMax will be released in ~10 Days
- π’I was a Data Scientist for 10 years before becoming a quadriplegic. For the past 3 months, I built VibeETL from scratch: A lightning-fast, visual Alteryx alternative powered by Polars & React Flow.
2026-05-31
- π΄Anyone else tired of the βClaude just dropped OPUS 4.8 weβre all cookedβ posts on LinkedIn?
- π΄nvidia/Qwen3.6-35B-A3B-NVFP4 Β· Hugging Face
- π‘mudler/Qwen3.6-35B-A3B-Claude-4.7-Opus-Reasoning-Distilled-APEX-MTP-GGUF just released !
- π’Built a minimalist coding agent optimized for memory footprint
- π’Best small model right now (~4B params) that is good with agentic tasks for personal assistant?
- π’Made a program using LocalLLM based on llama.cpp for fellow Book Lovers!
- π’Speed difference between Windows 11 and Linux with llama.cpp: a myth when using medium and large MoE models
2026-05-30
- π΄Notes from the Mistral AI Now Summit
- π‘@xai: grok-build-0.1 is now available via the xAI API in public beta. This is the same model that powers the Grok Build CLI and excels at agentic coding. Priced at $1/m input and $2/m output, itβs extreme
- π‘How Braintrust turns customer requests into code with Codex
- π‘@OpenAI: Windows users, this oneβs for you. Computer use now works on Windows, so Codex can take action on your Windows computer. And with Windows support for Codex in the ChatGPT mobile app, you can start,
- π‘llama : website + unified `llama` binary Β· ggml-org/llama.cpp Β· Discussion #23875
- π’MINISFORUM UM790 Pro
2026-05-29
- π΄Claude Opus 4.8
- π‘@AnthropicAI: We've raised $65 billion in Series H funding at a $965 billion post-money valuation, led by @AltimeterCap, Dragoneer, @Greenoaks, and @sequoia. This investment will help us advance our research and e
- π‘@xai: Grok Build 0.2.7 is now out, with /usage, /login, shared terminals across subagents, and improved image understanding See all updates at https://x.ai/build/changelog
- π‘@MistralAI: We're taking on the hardest problems in the real world ποΈπ π«βοΈ Today at The AI Now Summit, held at the Louvre, we announced AI solutions for aerospace, automotive, energy, and physics. Deployed in p
- π‘Wall-OSS-0.5: 4B VLA with open training code and zero-shot real-robot evaluation[D]
- π‘Claude Code β Everything You Can Configure That the Docs Don't Tell You
- π’llama.cpp B9387 Significant AMD/ROCm PP Update
- π’Researchers let AI models run a simulated society. Claude was the safestβand Grok committed 180 crimes and went extinct within 4 days
- π’Oculus Founders' AI Startup Sesame Launches Human-Like Voice AI App on iOS
- π’Qwen 3.6 27B overdoing it
2026-05-28
- π΄Qwen3.6 huge quality gain from Q4 to Q6 for coding agent
- π‘I built a 103B-token Usenet corpus (1980β2013) β pre-web, human-only, zero AI contamination. Got strong traction on r/ML, thought this community would find it useful.
- π‘CrankGPT by Squeez Labs - hand-cranked edge AI - talk about local AI!!!
- π‘@xai: Use your SuperGrok or X Premium+ subscription in @kilocode. Try grok-build-0.1 for high speed and agentic coding intelligence, available in the Kilo IDE extensions or CLI. https://x.ai/news/grok-ki
- π‘Nvidia LocateAnything - Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding. (10x faster than Qwen3-VL)
- π’i built pengepul: pool multiple claude/gpt accounts behind one local api
- π’Training GPT-like model on non-language series [R]
2026-05-27
- π΄A rare look inside Qwen 3.7βs open source model release approval process:
- π΄Okay 27B made me a believer
- π΄What have you built with Claude Code just for yourself?
- π‘$400 Qwen 3.6-27B Setup - Dual RTX 3060 - 30-50 t/s
- π‘@AnthropicAI: New on the Engineering Blog: The access and permissions we grant agents should evolve with their capabilities. In our own products, we set these parameters through sandboxing, which limits the scope o
- π‘@GoogleDeepMind: Our Gemini for Science tools could help scientists unlock their next breakthrough. π§¬
- π‘@xai: Thank you so much for all the feedback on the Grok Build Beta. Some of you reported hitting limits quickly. Our team found areas to improve caching, so we've reset Grok Build usage limits for all acc
- π’Cactus Hybrid Router: Gemma4-2B can match Gemini-3.1-Flash-Lite by routing 15-55% of tasks to Gemini And Running The Rest Locally.
- π’Folks running qwen 3.6 27b for agentic work. Do you dare to use q4_k_m?
- π’How I gave my AI agent real-time Reddit awareness in ~20 lines of Python using MCP
- π’AI agent for WhatsApp & Telegram
- π’Multi-Agent Parallelism and State Sovereignty in Agent Fleets
2026-05-26
- π΄The Financial Times has published an article about Heretic
- π΄Update on 12x32gb sxm v100 cluster / local AI for legal drafting
- π΄Is there a way to use multiple AI models without paying for 11 different monthly subscriptions?
- π΄Is Qwen3.6 current king for local agentic use?
- π‘@xai: Grok Build is now available in Beta for all SuperGrok and X Premium+ users. Use Plan Mode, create images and videos with Imagine, and build automations or orchestrators with the CLI. Visit http://x.
- π’model : add support for talkie-1930-13b by niklassheth Β· Pull Request #22596 Β· ggml-org/llama.cpp
- π’Running on a macbook, and having issues with crashing? Maybe this will help...
2026-05-25
- π΄PapersWithCode new features - week 1 [P]
- π΄Qwen3.6-35B-A3B vs Gemma4-26B-A4B
- π΄1000 tps generation on Qwen3.6 27B with V100s
- π‘Chromeflow: agentic browser MCP that drives real Chrome with sessions intact (500+ hours hardening)
- π‘hipEngine: Fast Native Qwen 3.6 Inference for RDNA3 (Strix Halo, 7900 XTX)
- π‘What frontend do you guys use?
- π‘I built a mini business consultant using Claude and Lovable as a weekend project,
- π‘MiMo-V2.5-coder
- π’I will not promote: After Two Months of RELAUNCHING I HAVE HIT $85 MRR
- π’I built an AI coding agent that integrates more than 22 AI providers into a single, native plug-and-play environment
2026-05-24
- π΄GPT 5.5 "secret sauce" is just having the thinking be some stupid caveman mode?
- π΄llama.cpp server have built-in native tools (exec_shell, edit_file, etc.)
- π΄Is there any reason for an uncensored model if you have no interest in roleplaying?
- π‘Run Chromeβs tiny Gemma4 (aka Gemini Nano) directly on PC without GPU
- π‘Qwen3.6-35B-A3B-Uncensored-Genesis-APEX-MTP
- π‘NVFP4 + MTP - voilΓ on llama.cpp
- π’llampart 1.0.0 - I released a standalone local web UI for llama-server with translations, extended settings and a polished conversation sidebar
2026-05-23
- π΄NuExtract3 released: open-weight 4B VLM for Markdown, OCR and structured extraction (self-hostable) [P]
- π΄BeeLlama v0.2.0 β major DFlash update. Single RTX 3090: Qwen 3.6 27B up to 164 tps (4.40x), Gemma 4 31B up to 177.8 tps (4.93x). Prompt processing speed near baseline.
- π‘@AnthropicAI: Last month we launched Project Glasswing, our collaborative AI cybersecurity initiative. Since then, we and our partners have found more than ten thousand high- or critical-severity vulnerabilities in
- π‘@GoogleDeepMind: SynthID, our imperceptible watermark for AI-generated content, is expanding to more partners. Weβre also adding new ways to find out if content was generated using AI - just ask in the @GeminiApp or
- π’Microsoft starts canceling Claude Code licenses
- π’meituan-longcat/LongCat-Video-Avatar-1.5 Β· Hugging Face
- π’397B competitor that fits in 256 RAM?
- π’Experimental "Preserve Thinking" Jinja Template for Gemma4 31B in llama.cpp
- π’Is personalized AI memory actually a problem worth solving or am I just coping[D]
2026-05-22
- π΄Waiting for Qwen 3.7 open weight... The new King has arrived...
- π΄Qwen3.6 35Ba3 has changed my workflows and even how I use my computer
- π‘LatitudeGames/Equinox-31B Β· Hugging Face
- π‘Launch HN: Runtime (YC P26) β Sandboxed coding agents for everyone on a team
- π‘For everyone that uses OpenCode / Pi - Heres your promptprocessing fix!
- π‘AdventHealth advances whole-person care with OpenAI
- π‘Weβre launching the Google DeepMind Accelerator program in Asia Pacific to tackle environmental risks
- π’Which agent should I use for coding and notetaking?
- π’New Release of ROCm based MLX LLM Engine - lemon-mlx-engine
- π’Low-level coding dataset
- π’ztok β a fast multithreaded tokenizer in Zig that loads tiktoken / HF / SentencePiece and is 2β5Γ faster
2026-05-21
- π΄Qwen will release another 27B with high probability
- π‘Qwen 3.6 35B GGUF: NTP vs MTP quantization results across GPUs and CPUs
- π‘@GoogleDeepMind: How can you accelerate your day to day research workflow? By giving AI the right scientific toolkit. We launched Science Skills for Google @Antigravity, integrating insights from over 30 major life
- π‘@GoogleDeepMind: Gemini 3.5 Flash has landed.
- π‘Same task in github-copilot, pi, claude-code, and opencode with Qwen3.6 27B
- π’How can you stop your model from looping
2026-05-20
- π΄Gemini 3.5 Flash
- π΄Do you guys actually think AI agents can replace people for bigger tasks anytime soon?
- π‘LM Studio finally added support for MTP Speculative Decoding
- π‘Co-Scientist (Nature 2026-05-19): 5+1 Gemini agents, tournament-of-ideas, prod
- π‘@OpenAI: People are generating over 1.5 billion images a week in ChatGPT. Researcher @kenjihata joins Product lead @adele__li and host @AndrewMayne to explore the new use cases and trends emerging since the l
- π‘@OpenAI: Introducing OpenAI Guaranteed Capacity: a new offering that enables customers to guarantee long-term access to OpenAI compute. Weβve made long-term investments in infrastructure, partnerships, and ca
- π‘Carbon: Decoding the Language of Life
- π’Qwen3.7 Max scored by Artificial Analysis, 27B/35B waiting room
- π’Gemini CLI will stop working from June 18, 2026
- π’Letβs talk quants of Gemma and Qwen - 16 vs Q8 vs Q4 - any experiences?
- π’Gemma 4 MTP with LlamaCPP
2026-05-19
- π΄Qwen cant wait to release 3.7 models
- π΄Reviving PapersWithCode (by Hugging Face) [P]
- π΄Qwen is cooking hard
- π‘What happens to local LLM if/when LLMs are no longer released for free?
- π‘Released a free 9.8M doc Indic multilingual corpus β Hindi, Bengali, Tamil, Telugu + 7 more (CC0, HuggingFace) [P]
- π’MTP (Multi-Token Prediction): 2x Faster Token Generation on AMD Strix Halo & Radeon 9700 AI Pro
- π’favorite Agentic Coding Harness
- π’We built a tool that installs frameworks like ComfyUI, Ollama, OpenWebUI etc on any cloud GPU in one command and saves your whole setup between sessions [R]
2026-05-18
- π΄"Generate a photorealistic realtime render of a human face with webGL" (Qwen3.5-122B-A10B UD-Q3_K_XL)
- π‘llama: avoid copying logits during prompt decode in MTP by am17an Β· Pull Request #23198 Β· ggml-org/llama.cpp
- π‘The power of structured workflows and small local models
- π’could refusal layers be masking dialect-conditioned safety failures in MoE models [d]
2026-05-17
- π΄MTP PR Merged!!!
- π΄That's a good news...
- π΄AI Agents are pushing us towards Transactional Development β Spin Up, Build, Ship, Discard.
- π΄Anthropic and OpenAI claims that their models are so powerful that it can βbreakβ their boxβ¦but what so special about their agent implementation?
- π΄MTP support merged into llama.cpp
- π‘some things i learned the hard way using claude design
- π‘G4-Meromero-31B-Uncensored-Heretic Is Out Now, a Finetune of Gemma 4 31B It Designed for Creative Tasks, With Kld of 0.0100 and 15/100 Refusals!
- π‘@xai: You can now use X Premium subscriptions in Hermes Agent, and Hermes Agent can now search X posts. https://x.ai/news/grok-hermes
- π’webui: support video files as input by foldl Β· Pull Request #22830 Β· ggml-org/llama.cpp
- π’OpenAI and Government of Malta partner to roll out ChatGPT Plus to all citizens
- π’Looking to migrate off of Ollama and LMStudio
2026-05-16
- π΄Opencode you naughty minx
- π΄I've shipped 3 products this year. None of them have users. Here's my problem.
- π΄Dynamically allocating compute budget to hard set of problems and evolving the sections with Qwen-35B-A3B gets you near GPT-5.4-xHigh on HLE
- π‘Orthrus-Qwen3-8B : up to 7.8Γtokens/forward on Qwen3-8B, frozen backbone, provably identical output distribution
- π‘KDD 2026 Cycle 2 Results [D]
- π‘[FOUNDING] SupraLabs - real open-source AI models for you!
- π‘@xai: You can now use your @grok subscription inside @NousResearch Hermes Agent. http://x.ai/news/grok-hermes
- π’Luce Megakernal: Why nobody is taking about this?
- π’[R] Which LLMs are actually best for bleeding-edge Linux/ML debugging workflows in 2026? [R]
2026-05-15
- π΄The RTX 5000 PRO (48GB) arrived and it is better than I expected.
- π΄Codex is now in the ChatGPT mobile app
- π΄How Claude Code works in large codebases
- π‘China modded GPU (eg. 4090 48gb) --> I'm gonna figure it out. IS THERE NO ONE ELSE CURIOUS??
- π‘@xai: An early beta of Grok Build, an agentic CLI for coding, building apps, and automating workflows is now available for SuperGrok Heavy subscribers. Through this early beta, we will improve the model an
- π‘I Let a Small Model Train on Its Own Mistakes. It Reached 80% on HumanEval and Beat GPT-3.5 on Math
- π‘Claude for Legal
- π‘@AnthropicAI: Weβre partnering with the Gates Foundation, committing $200 million in grants, Claude credits, and technical support to programs in global health, life sciences, education, agriculture, and economic m
- π’RDNA3 Flash Attention fix just dropped by llama.cpp b9158
- π’Show HN: GlycemicGPT β Open-source AI-powered diabetes management
- π’RelaxAI β UK sovereign LLM inference at 80% cheaper than OpenAI/Claude
- π’I have (even faster) DeepSeek V4 Pro at home
- π’I got tired of OpenClaw skills having no actual usage so I spent 3 weeks building one.
2026-05-14
- π΄TextGen is now a native desktop app. Open-source alternative to LM Studio (formerly text-generation-webui).
- π΄DELIGHT β self-hosted AI engineering autopilot: local LLM + browser farm + repo graph + P2P compute
- π΄Claude for Small Business
- π’What revenue model would you guys suggest for our automation orchestration platform open to public agents as a marketplace?
- π’A Claude Code and Codex Skill for Deliberate Skill Development
- π’Simpler self hosted alt to Open WebUI
- π’The "the future is fictional" problem of many local LLMs
2026-05-13
- π΄Stop wasting electricity
- π΄TabPFN-3 just released: a pre-trained tabular foundation model for up to 1M rows [R][N]
- π‘Let's build claude code from scratch!
- π‘Is using vLLM actually worth it if you aren't serving the model to other people?
- π‘I think AI businesses are quietly shifting from hype products to operational tools
- π’I've seen a lot of folks ask "can local LLMs actually do anything useful?"
- π’OpenAI API now requires separate pay-as-you-go billing β what alternatives are you all using?
- π’Small local model for questions on German grammar
2026-05-12
- π΄MTP on Unsloth
- π‘@OpenAI: Introducing Daybreak: frontier AI for cyber defenders. Daybreak brings together the most capable OpenAI models, Codex, and our security partners to accelerate cyber defense and continuously secure so
- π‘Will there be any more Qwen3.6 series models?
- π‘How ChatGPT adoption broadened in early 2026
- π‘@AnthropicAI: New Anthropic research: Teaching Claude why. Last year we reported that, under certain experimental conditions, Claude 4 would blackmail users. Since then, weβve completely eliminated this behavior.
- π‘@OpenAI: Today weβre launching the OpenAI Deployment Company to help businesses build and deploy AI. It's majority-owned and controlled by OpenAI. It brings together 19 leading investment firms, consultancies
- π’I catalogued every way local models break JSON output and built a repair library, here's what I found across 288 model calls
- π’Llama models: still valuable for finetuning or surpassed by everything new?
- π’Claude Platform on AWS
2026-05-11
- π΄I have DeepSeek V4 Pro at home
- π΄How I set up an AI agent to handle invoicing bill pay and expense tracking through my bank via MCP
- π‘Any implementations similar to D4RT? [D]
- π‘Put my 4 years of SEO experience into a claude skill so you don't have to figure it out yourself
- π‘ExLlamaV3 Major Updates!
- π’What's actually moving the needle on agent token bills?
- π’How Fast Does Claude, Acting as a User Space IP Stack, Respond to Pings?
- π’The Qwen 3.6 35B A3B hype is real!!!
- π’Show HN: adamsreview β better multi-agent PR reviews for Claude Code
2026-05-10
- π‘NVIDIA AI Releases Star Elastic: One Checkpoint that Contains 30B, 23B, and 12B Reasoning Models with Zero-Shot Slicing
- π‘Pi and Qwen3.6 27B make setting up Archlinux really easy.
- π‘More Qwen3.6-27B MTP success but on dual Mi50s
- π’Building a AI teacher-assistance software, Assistance needed.
- π’Would a community-driven AI agent lab help people actually ship agents?
- π’Exactly a year ago, I started working on an MCP server I launched on reddit that became by far my most active open source project!
- π’Gemini API File Search is now multimodal
- π’I am overwhelmed by Harnesses
2026-05-09
- π΄vLLM ROCm has been added to Lemonade as an experimental backend
- π΄A recent experience with ChatGPT 5.5 Pro
- π΄I'm kinda good at getting users for ai tools through reddit - could I make money?
- π‘Teaching Claude Why
- π‘Reports suggest DeepSeek is seeking $7.35 billion in funding and plans to release its V4.1 update next month.
- π‘@OpenAI: Chain of thought monitors are a key layer of defense against AI agent misalignment. To preserve monitorability, we avoid penalizing misaligned reasoning during RL. We found a limited amount of accide
- π‘Using Claude Code: The unreasonable effectiveness of HTML
- π‘new MoE from ai2, EMO
- π’Formalizing statistical learning theory in Lean 4 [R]
- π’How long for llama.cpp official support of MTP?
- π’FEEDBACK FOR MY APP
2026-05-08
- π‘Agentic workflows
- π‘AlphaEvolve: Gemini-powered coding agent scaling impact across fields
- π‘You can now read Gemma 3's mind
- π‘Scaling Trusted Access for Cyber with GPT-5.5 and GPT-5.5-Cyber
- π‘@AnthropicAI: Weβre donating Petri, our open-source alignment tool, to @meridianlabs_ai, so its development can continue independently. Working with Meridian Labs, weβve also released a major update that improves
- π’Quantization and Fast Inference (MEAP) - How much performance are you actually getting from quantization in production? [D]
- π’Hardening Firefox with Claude Mythos Preview
2026-05-07
- π‘Qwen3.6 27B uncensored heretic v2 Native MTP Preserved is Out Now With KLD 0.0021, 6/100 Refusals and the Full 15 MTPs Preserved and Retained, Available in Safetensors, GGUFs and NVFP4s formats.
- π‘What Anthropic's 'Dreaming' feature release made me notice about my own ClawVault agent's memory
- π‘@xai: SpaceXAI will provide @AnthropicAI with access to Colossus 1, one of the worldβs largest and fastest-deployed AI supercomputers, to provide additional capacity for Claude β http://x.ai/news/anthropic-
- π’why llama.cpp canβt combine speculative decode methods?
- π’Qwen 3.6?
2026-05-06
- π΄Gemma 4 MTP released
- π΄Sr Software Engineer - Haven't written a line of code in months
- π‘MTP on strix halo with llama.cpp (PR #22673)
- π’Tired of copy-pasting prompts between Claude and Codex tabs: built a small file-backed queue that automates the handoff
- π’Bleeding Llama: Critical Unauthenticated Memory Leak in Ollama
- π’Qwen 3.6 27B MTP on v100 32GB: 54 t/s
- π’Solidity LM surpasses Opus
2026-05-05
- π΄White House Considers Vetting A.I. Models Before They Are Released
- π΄Peanut - Text to Image Model (Open Weights coming soon)
- π‘FastDMS: 6.4X KV-cache compression running faster than vLLM BF16/FP8
- π‘Claude Code structure that didnβt break after 2β3 real projects
- π’As MTP prepares to land in llama.cpp, Models that support MTP
2026-05-03
- π΄Qwen3.6-27B vs Coder-Next
- π΄Karpathy's MicroGPT running at 50,000 tps on an FPGA
- π΄GPT 5.5 just leaked its chain of thought to me in codex, and it looks like an idea from 5 months ago in this sub.
- π‘I made a visualizer for Hugging Face models
- π’What a time to be alive from 1tk/sec to 20-100tk/sec for huge models
2026-04-30
- π΄Large Language Models Explore by Latent Distilling
- π‘ClawGym: A Scalable Framework for Building Effective Claw Agents
- π‘Introducing Advanced Account Security
- π‘Enabling a new model for healthcare with AI co-clinician
- π’Operating-Layer Controls for Onchain Language-Model Agents Under Real Capital
2026-04-28
- π‘ReVSI: Rebuilding Visual Spatial Intelligence Evaluation for Accurate Assessment of VLM 3D Reasoning
- π‘Apr 28, 2026 Announcements Claude for Creative Work
- π‘Apr 27, 2026 Announcements Anthropic names Theo Hourmouzis General Manager of Australia & New Zealand and officially opens Sydney office
- π‘Apr 24, 2026 Announcements An update on our election safeguards
- π‘Apr 24, 2026 Announcements Anthropic and NEC collaborate to build Japanβs largest AI engineering workforce
2026-04-27
- π΄DiffNR: Diffusion-Enhanced Neural Representation Optimization for Sparse-View 3D Tomographic Reconstruction
- π‘OpenAI available at FedRAMP Moderate
- π‘Contexts are Never Long Enough: Structured Reasoning for Scalable Question Answering over Long Document Sets
- π’Building a Precise Video Language with Human-AI Oversight
2026-04-22
- π΄Tstars-Tryon 1.0: Robust and Realistic Virtual Try-On for Diverse Fashion Items
- π‘SmartPhotoCrafter: Unified Reasoning, Generation and Optimization for Automatic Photographic Image Editing
- π‘Making ChatGPT better for clinicians
- π‘Workspace agents
- π‘Introducing workspace agents in ChatGPT
- π‘Introducing OpenAI Privacy Filter
2026-04-20
- π΄Maximal Brain Damage Without Data or Optimization: Disrupting Neural Networks via Sign-Bit Flips
- π‘Qwen3.5-Omni Technical Report
- π‘PersonaVLM: Long-Term Personalized Multimodal LLMs
- π‘OpenAI helps Hyatt advance AI among colleagues
- π’Cut Your Losses! Learning to Prune Paths Early for Efficient Parallel Reasoning
2026-04-17
- π΄HY-World 2.0: A Multi-Modal World Model for Reconstructing, Generating, and Simulating 3D Worlds
- π΄How to Fine-Tune a Reasoning Model? A Teacher-Student Cooperation Framework to Synthesize Student-Consistent SFT Data
- π‘Product Apr 17, 2026 Introducing Claude Design by Anthropic Labs Today, weβre launching Claude Design, a new Anthropic Labs product that lets you collaborate with Claude to create polished visual work like designs, prototypes, slides, one-pagers, and more.
- π‘Announcements Feb 4, 2026 Claude is a space to think Weβve made a choice: Claude will remain ad-free. We explain why advertising incentives are incompatible with a genuinely helpful AI assistant, and how we plan to expand access without compromising user trust.
- π‘Dive into Claude Code: The Design Space of Today's and Future AI Agent Systems
2026-04-16
- π΄Seedance 2.0: Advancing Video Generation for World Complexity
- π΄RationalRewards: Reasoning Rewards Scale Visual Generation Both Training and Test Time
- π‘OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language World Models
- π‘Introducing Claude Opus 4.7 Product Apr 16, 2026 Our latest Opus model brings stronger performance across coding, agents, vision, and multi-step tasks, with greater thoroughness and consistency on the work that matters most.
- π‘Introducing GPT-Rosalind for life sciences research
- π‘Accelerating the cyber defense ecosystem that protects us all
- π’Exploration and Exploitation Errors Are Measurable for Language Model Agents
2026-04-15
- π‘Nemotron 3 Super: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning
- π‘Gemini 3.1 Flash TTS: the next generation of expressive AI speech
- π’BERT-as-a-Judge: A Robust Alternative to Lexical Methods for Efficient Reference-Based LLM Evaluation
- π’The Blind Spot of Agent Safety: How Benign User Instructions Expose Critical Vulnerabilities in Computer-Use Agents
2026-04-08
- π΄Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding
- π‘Featured An update on recent Claude Code quality reports We traced recent reports of Claude Code quality issues to three separate changes. Here's what happened and what we're changing.
- π‘Scaling Managed Agents: Decoupling the brain from the hands Apr 08, 2026
- π‘Claude Code auto mode: a safer way to skip permissions Mar 25, 2026
- π‘Harness design for long-running application development Mar 24, 2026
- π‘Eval awareness in Claude Opus 4.6βs BrowseComp performance Mar 06, 2026
- π’ThinkTwice: Jointly Optimizing Large Language Models for Reasoning and Self-Refinement
- π’How Well Do Agentic Skills Work in the Wild: Benchmarking LLM Skill Usage in Realistic Settings
2026-04-06
- π΄GrandCode: Achieving Grandmaster Level in Competitive Programming via Agentic Reinforcement Learning
- π‘Agentic-MME: What Agentic Capability Really Brings to Multimodal Intelligence?
- π’Communicating about Space: Language-Mediated Spatial Integration Across Partial Views
- π’AgentHazard: A Benchmark for Evaluating Harmful Behavior in Computer-Use Agents
2026-04-02
- π΄ClawKeeper: Comprehensive Safety Protection for OpenClaw Agents Through Skills, Plugins, and Watchers
- π΄Terminal Agents Suffice for Enterprise Automation
- π‘Embarrassingly Simple Self-Distillation Improves Code Generation
- π‘Codex now offers more flexible pricing for teams
- π’QuitoBench: A High-Quality Open Time Series Forecasting Benchmark
- π’GaussianGPT: Towards Autoregressive 3D Gaussian Scene Generation
2026-04-01
- π΄FIPO: Eliciting Deep Reasoning with Future-KL Influenced Policy Optimization
- π΄CARLA-Air: Fly Drones Inside a CARLA World -- A Unified Infrastructure for Air-Ground Embodied Intelligence
- π‘GEMS: Agent-Native Multimodal Generation with Memory and Skills
- π‘Gradient Labs gives every bank customer an AI account manager
- π‘Project Imaging-X: A Survey of 1000+ Open-Access Medical Imaging Datasets for Foundation Model Development
2026-03-31
- π΄TAPS: Task Aware Proposal Distributions for Speculative Sampling
- π΄Gen-Searcher: Reinforcing Agentic Search for Image Generation
- π‘Accelerating the next phase of AI
- π‘HISA: Efficient Hierarchical Indexing for Fine-Grained Sparse Attention
- π’GEditBench v2: A Human-Aligned Benchmark for General Image Editing
- π’ImagenWorld: Stress-Testing Image Generation Models with Explainable Human Evaluation on Open-ended Real-World Tasks
2026-03-27
- π΄Calibri: Enhancing Diffusion Transformers via Parameter-Efficient Calibration
- π‘Voxtral TTS
- π‘STADLER reshapes knowledge work at a 230-year-old company
- π’MACRO: Advancing Multi-Reference Image Generation with Structured Long-Context Data
- π’AVControl: Efficient Framework for Training Audio-Visual Controls
- π’VFIG: Vectorizing Complex Figures in SVG with Vision-Language Models
2026-03-26
- π΄CUA-Suite: Massive Human-annotated Video Demonstrations for Computer-Use Agents
- π΄Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?
- π‘T-MAP: Red-Teaming LLM Agents with Trajectory-aware Evolutionary Search
- π‘Gemini 3.1 Flash Live: Making audio AI more natural and reliable
2026-03-20
- π΄SAMA: Factorized Semantic Anchoring and Motion Alignment for Instruction-Guided Video Editing
- π΄Nemotron-Cascade 2: Post-Training LLMs with Cascade RL and Multi-Domain On-Policy Distillation
- π‘FASTER: Rethinking Real-Time Flow VLAs
- π’F2LLM-v2: Inclusive, Performant, and Efficient Embeddings for a Multilingual World
2026-03-17
- π΄AI Can Learn Scientific Taste
- π‘Introducing GPT-5.4 mini and nano
- π‘OpenAI Japan announces Japan Teen Safety Blueprint to put teen safety first
- π‘Equipping workers with insights about compensation
- π‘Measuring progress toward AGI: A cognitive framework
- π‘EnterpriseOps-Gym: Environments and Evaluations for Stateful Agentic Planning and Tool Use in Enterprise Settings
- π’Mixture-of-Depths Attention
- π’Effective Distillation to Hybrid xLSTM Architectures
- π’Safe and Scalable Web Agent Learning via Recreated Websites
2026-03-16
- π΄Multimodal OCR: Parse Anything from Documents
- π‘Cheers: Decoupling Patch Details from Semantic Representations Enables Unified Multimodal Comprehension and Generation
- π‘OmniForcing: Unleashing Real-time Joint Audio-Visual Generation
- π‘Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously
- π‘daVinci-Env: Open SWE Environment Synthesis at Scale
- π’Visual-ERM: Reward Modeling for Visual Equivalence
2026-03-13
- π΄Strategic Navigation or Stochastic Search? How Agents and Humans Reason Over Document Collections
- π‘Video-Based Reward Modeling for Computer-Use Agents
- π’DVD: Deterministic Video Depth Estimation with Generative Priors
- π’Trust Your Critic: Robust Reward Modeling and Reinforcement Learning for Faithful Image Editing and Generation
2026-03-11
- π΄Geometry-Guided Reinforcement Learning for Multi-view Consistent 3D Scene Editing
- π‘Designing AI agents to resist prompt injection
- π‘Fish Audio S2 Technical Report
- π’Stepping VLMs onto the Court: Benchmarking Spatial Intelligence in Sports
- π’Reading, Not Thinking: Understanding and Bridging the Modality Gap When Text Becomes Pixels in Multimodal LLMs
- π’Are Audio-Language Models Listening? Audio-Specialist Heads for Adaptive Audio Steering
2026-03-09
- π΄Penguin-VL: Exploring the Efficiency Limits of VLM with LLM-based Vision Encoders
- π‘Progressive Residual Warmup for Language Model Pretraining
- π‘Reasoning Models Struggle to Control their Chains of Thought
- π’Physical Simulator In-the-Loop Video Generation
- π’HiMAP-Travel: Hierarchical Multi-Agent Planning for Long-Horizon Constrained Travel
2026-03-05
- π΄Helios: Real Real-Time Long Video Generation Model
- π΄T2S-Bench & Structure-of-Thought: Benchmarking and Prompting Comprehensive Text-to-Structure Reasoning
- π‘Introducing GPT-5.4
- π‘GPT-5.4 Thinking System Card
- π‘Introducing the Adoption news channel
- π‘Introducing ChatGPT for Excel and new financial data integrations
- π‘VfL Wolfsburg turns ChatGPT into a club-wide capability
- π’Phi-4-reasoning-vision-15B Technical Report
2026-03-03
- π΄SWE-rebench V2: Language-Agnostic SWE Task Collection at Scale
- π‘RubricBench: Aligning Model-Generated Rubrics with Human Standards
- π‘CHIMERA: Compact Synthetic Data for Generalizable LLM Reasoning
- π‘GPT-5.3 Instant System Card
- π‘GPT-5.3 Instant: Smoother, more useful everyday conversations
- π‘Gemini 3.1 Flash-Lite: Built for intelligence at scale
- π’MMR-Life: Piecing Together Real-life Scenes for Multimodal Multi-image Reasoning
2026-02-27
- π΄From Blind Spots to Gains: Diagnostic-Driven Iterative Training for Large Multimodal Models
- π΄MobilityBench: A Benchmark for Evaluating Route-Planning Agents in Real-World Mobility Scenarios
- π‘Introducing the Stateful Runtime Environment for Agents in Amazon Bedrock
- π’AgentDropoutV2: Optimizing Information Flow in Multi-Agent Systems via Test-Time Rectify-or-Reject Pruning
2026-02-26
- π‘DualPath: Breaking the Storage Bandwidth Bottleneck in Agentic LLM Inference
- π‘DreamID-Omni: Unified Framework for Controllable Human-Centric Audio-Video Generation
- π‘OpenAI Codex and Figma launch seamless code-to-design experience
- π’GUI-Libra: Training Native GUI Agents to Reason and Act with Action-aware Supervision and Partially Verifiable RL
2026-02-18
- π‘Introducing OpenAI for India
- π‘Introducing EVMbench
- π‘A new way to express yourself: Gemini can now create music
- π‘A Trajectory-Based Safety Audit of Clawdbot (OpenClaw)
- π’ResearchGym: Evaluating Language Model Agents on Real-World AI Research
- π’HLE-Verified: A Systematic Verification and Structured Revision of Humanity's Last Exam
2026-02-17
- π΄BitDance: Scaling Autoregressive Generative Models with Binary Tokens
- π‘Nanbeige4.1-3B: A Small General Model that Reasons, Aligns, and Acts
- π‘REDSearcher: A Scalable and Cost-Efficient Framework for Long-Horizon Search Agents
- π’Query as Anchor: Scenario-Adaptive User Representation via Large Language Model
- π’Data Darwinism Part I: Unlocking the Value of Scientific Data for Pre-training
2026-02-16
- π΄Less is Enough: Synthesizing Diverse Data in Feature Space of LLMs
- π΄MedXIAOHE: A Comprehensive Recipe for Building Medical MLLMs
- π‘Towards Universal Video MLLMs with Attribute-Structured and Quality-Verified Instructions
- π‘OneVision-Encoder: Codec-Aligned Sparsity as a Foundational Principle for Multimodal Intelligence
- π’GeoAgent: Learning to Geolocate Everywhere with Reinforced Geographic Characteristics
2026-02-13
- π΄DeepGen 1.0: A Lightweight Unified Multimodal Model for Advancing Image Generation and Editing
- π‘Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation
- π‘GPT-5.2 derives a new result in theoretical physics
- π‘Introducing Lockdown Mode and Elevated Risk labels in ChatGPT
- π‘Scaling social science research
- π’Think Longer to Explore Deeper: Learn to Explore In-Context via Length-Incentivized Reinforcement Learning
2026-02-12
- π΄Step 3.5 Flash: Open Frontier-Level Intelligence with 11B Active Parameters
- π΄GENIUS: Generative Fluid Intelligence Evaluation Suite
- π‘ASA: Training-Free Representation Engineering for Tool-Calling Agents
- π‘Introducing GPT-5.3-Codex-Spark
- π‘Gemini 3 Deep Think: Advancing science, research and engineering
- π‘Towards Autonomous Mathematics Research
- π’TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions
2026-02-11
- π΄OPUS: Towards Efficient and Principled Data Selection in Large Language Model Pre-training in Every Iteration
- π΄Code2World: A GUI World Model via Renderable Code Generation
- π‘Chain of Mindset: Reasoning with Adaptive Cognitive Modes
- π‘P1-VL: Bridging Visual Perception and Scientific Reasoning in Physics Olympiads
2026-02-09
- π΄F-GRPO: Don't Let Your Policy Learn the Obvious and Forget the Rare
- π‘Baichuan-M3: Modeling Clinical Inquiry for Reliable Medical Decision-Making
- π‘Testing ads in ChatGPT
- π‘Bringing ChatGPT to GenAI.mil
- π‘Accelerating Mathematical and Scientific Discovery with Gemini Deep Think
- π’MSign: An Optimizer Preventing Training Instability in Large Language Models via Stable Rank Restoration
- π’Judging What We Cannot Solve: A Consequence-Based Approach for Oracle-Free Evaluation of Research-Level Math
2026-02-05
- π‘Training Data Efficiency in Multimodal Process Reward Models
- π‘OmniSIFT: Modality-Asymmetric Token Compression for Efficient Omni-modal Large Language Models
- π‘GPT-5 lowers the cost of cell-free protein synthesis
- π‘Introducing Trusted Access for Cyber
- π‘Introducing OpenAI Frontier
- π’EgoActor: Grounding Task Planning into Spatial-aware Egocentric Actions for Humanoid Robots via Visual-Language Models
- π’Rethinking the Trust Region in LLM Reinforcement Learning
2026-02-04
- π΄AOrchestra: Automating Sub-Agent Creation for Agentic Orchestration
- π΄No Global Plan in Chain-of-Thought: Uncover the Latent Planning Horizon of LLMs
- π’SWE-World: Building Software Engineering Agents in Docker-Free Environments
- π’SWE-Master: Unleashing the Potential of Software Engineering Agents via Post-Training
2026-02-03
- π΄Kimi K2.5: Visual Agentic Intelligence
- π΄Vision-DeepResearch: Incentivizing DeepResearch Capability in Multimodal Large Language Models
- π‘Vision-DeepResearch Benchmark: Rethinking Visual and Textual Search for Multimodal Large Language Models
- π‘The Sora feed philosophy
- π’SWE-Universe: Scale Real-World Verifiable Environments to Millions