Infrastructure & Compute
735 signals across 170 briefings.
2026-10-01
- 🔴Parallel-in-Time Training of Recurrent Neural Networks for Dynamical Systems Reconstruction [R]
- 🔴We like the experimentalLong Decode Continuation, which increasesoutputtokens up to 1M as an industry first.
- 🔴Output limit: Google cites an industry-leading 1M-token output limit, up from 64K (@GoogleAI,@TheRundownAI).Measurement note: Vals lists 262K max output. Artificial Analysis reached 1M output tokens t
- 🔴Measurement note: Vals lists 262K max output. Artificial Analysis reached 1M output tokens through Long Decode Continuation, a new API feature that pauses long responses and resumes them across calls
- 🔴Imagine buying a warehouse full of GPUs and discovering that your inventory evaporates every second. The chips remain. Their unused capacity does not. An idle GPU-hour today cannot be placed on a shel
- 🟡How to address novelty concerns in top ai conference? [D]
- 🟡I spent a day poking Dots with sticks. Here’s what I figured out.
- 🟡tile-ai/tilelang — Domain-specific language designed to streamline the development of high-performance GPU/CPU/Accelerators kernels
2026-09-29
- 🔴Anthropic files for $2T IPO with $42B net loss in 2025, expects to spend half a trillion more
- 🔴BA Computer Science, but fell in love with machine learning and AI. Just got my personal research accepted at NeurIPS as a poster. [R]
- 🔴AI is basically ubiquitous in all corporate work but reddit is convinced AI is useless, how do those 2 things co-exists?
- 🔴Theofficial postis shy, but since AMD is public, we knowthe purchase price. Wecovered themless than a year ago:
- 🔴Scaling An ML Inference Pipeline For Batch Workloads (5 Minute Read)
- 🟡NAND-16: a computer built from 277,248 NAND gates
- 🟢AMD boosting AI/LLM performance for Radeon iGPUs as much as 18~23% with Linux 7.4
- 🟢SemiAnalysisAI/InferenceX — Open Source Inference Research Platform Standard / 开源推理研究平台
2026-09-28
- 🔴OpenAI pauses frontier training after models swarm US Governament
- 🔴Global AI Routing With <1% Overhead On Multi-Cluster Gke Inference Gateway (6 Minute Read)
- 🟡World Labs Is Joining AMD
- 🟡AMD is buying Fei-Fei Li's World Labs in a multibillion-dollar deal
- 🟡Can we get some quality control on all these model perf posts?
- 🟢Debugging PCIe Link Retraining on an x8/x8 Splitter with Two RTX 3090s
2026-09-27
- 🔴China Puts AI Compute Into Orbit With Supercomputing-1 Satellite — Onboard Processing Aims To Cut Earth-Observation Data Processing From Hours To Minutes (2 Minute Read)
- 🔴Topology-Aware Workload Scheduling With NVIDIA Topograph (10 Minute Read)
- 🔴Google’s AI chips are headed to space
- 🟡What are chinese labs doing differently?
- 🟡soon the big labs will train HUGE MODELS that won't be served to the public
- 🟢Two-stage shelf audit: YOLO finds the products, embeddings can't tell sibling SKUS apart. What should Stage 2 be? [P]
2026-09-26
- 🔴Anthropic is in talks to lease up to 1GW of data center capacity from Stream Data Centers, potentially using AI chips from Google and Broadcom, The Information reported.
- 🟡We need Universal Basic Income before losing your job to AI becomes your financial emergency
- 🟡What happens when the data centers have to replace their hardware?
- 🟡A safe AI might not be aligned the way the labs want
- 🟢How accessible is local AI actually, and what happens if affordable access to frontier models doesn’t last?
2026-09-25
- 🔴Nvidia CEO Jensen Huang put the odds of AI ending the world by 2030 at 0%, calling ex-Anthropic researcher Jacob Coxon’s viral warnings “irresponsible.”
- 🔴Forge by Fireworks brings together leaders building their own frontier on open models. Hear from NVIDIA CEO Jensen Huang, Fireworks CEO Lin Qiao, Microsoft EVP Jay Parikh, and more on Nov. 3 in SF. Apply to attend.
- 🟡AI alignment is the most important problem we will ever have to face.
- 🟡NeurIPS reject -> ICLR: How much reviewer feedback are you actually implementing ? [D]
- 🟡nobodywho-ooo/nobodywho — NobodyWho is an inference engine that lets you run LLMs locally and efficiently on any device.
2026-09-24
- 🟡Elon Musk: xAI Will Keep Accelerating And Will Surpass Anthropic And OpenAI Within 6 Months
- 🟡How are Chinese AI labs releasing competitive models so cheaply?
- 🟡NVIDIA/Model-Optimizer — A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative decoding, etc. It compresses deep learning models for downstream deployment frameworks like TensorRT-LLM, TensorRT, vLLM, etc. to optimize inference speed.
2026-09-23
- 🔴Offloaded inference for real-world physical AI robotics
- 🔴Jev isn't new tech. Its marketing targets people who think AI started with LLMs.
- 🔴While he was at Stanford, Eric couldn’t get traction on Genomic Language Models (GLMs) for a long time. Biologists didn’t believe it would work, didn’t think they could verify the output, and didn’t s
- 🟡Is AI-fatigue a thing or I'm officially becoming a boomer?
- 🟡Cost of intelligence is dropping fast
- 🟡MiMo-V2.6 (both Pro and Flash) is a benchmaxxed scam
- 🟢AtomicBot-ai/Atomic-Chat — Local AI app and inference engine for agents. Run open-weight LLMs locally — private, 100% offline on your computer. Join our Discord:https://discord.com/invite/8wGSsvmg4V
- 🟢Lemonade fixes AMD APU model streaming, drops OpenMOSS ROCm as ~40x slower than Vulkan
2026-09-22
- 🔴Alibaba plans AI model with 5 trillion to 10 trillion parameters, unveils new chip
- 🔴Laya surpassed Jev's speed at home with 16GB and for free.There's no problem that compute doesn't solve.
- 🔴Apple Reportedly Eyes NVIDIA's Nvlink Fusion To Connect Its Own M8 Ultra AI Servers By 2029 (5 Minute Read)
- 🔴Crusoe Raises $3.9B To Build Massive Data Centers And Truck-Transportable Modular "Spark" AI Factories (3 Minute Read)
- 🔴Ucla-Pioneered Nanowire Networks Physically Rewire Themselves To Compute, No Software Neural Network Required (4 Minute Read)
- 🟢Simulating fault tolerance with stage skipping in pipeline-parallel training [R]
2026-09-21
- 🔴16GB (and in many cases 12GB) is the max vram most people will ever reasonably have
- 🔴These Were NOT Rogue AI Escapes. Just SLOPPY Firewall Failures. [N]
- 🔴🚨 AI may be entering a completely different phase.
- 🔴speed based - games and computer use
- 🔴Open-weight language models have grown substantially in general interest and economic viability in 2026, allowing early glimpses of more direct ways to compare adoption of models from the US, China, o
- 🟡For NeurIPS: Is Paris or Syndey better for networking with U.S. tech companies? [D]
- 🟡Systems for Machine Learning[D]
- 🟢Huawei shelves global AI chip rollout as China's own demand outstrips supply — AMD and Nvidia no longer have to worry.
- 🟢Math is solved? Sometimes I wonder what 2027 will look like
2026-09-20
- 🔴China's CXMT says new memory-chip platform enters mass production
- 🔴FareedKhan-dev/train-llm-from-scratch — A straightforward method for training your LLM, from downloading data to generating text.
- 🔴Figure’s Helix 2.5 takes that argument into unfamiliar living rooms. The company tested tidying, towel folding, and bed making across 30 unseen homes. Pretraining on its Index human-behavior dataset r
- 🔴Bypassing Inference Bottlenecks With Retrieve-For-Train (6 Minute Read)
- 🔴Apple Reportedly Building Server Packed With M-Series Ultra Chips For AI (2 Minute Read)
- 🟢DIY Jev
2026-09-19
- 🔴cactus-compute/needle — Automation foundation model for tiny devices: 2-bit, 8-29 MB, tool calls, structured extraction and embeddings on phones, wearables, smart homes, robots, cars and microcontrollers.
- 🔴AI is a better teacher than most human teachers
- 🔴Have It Both Ways: Stay Discoverable In Search While Disallowing AI Training (12 Minute Read)
- 🟡How is RLCD (jev) RL? [D]
- 🟡What is the best way to use LLMS's in 2026? I feel like a caveman with the way I use them?
- 🟡NVIDIA/TensorRT-LLM — TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in a performant way.
- 🟢lahfir/agent-desktop — Agent Desktop gives any agent reliable computer use on the desktop. Built with Rust, it sees any app's real UI structure through OS accessibility trees and operates it — refs stay stable and actions stay safe to retry, instead of guessing from pixels.
- 🟢EricLBuehler/mistral.rs — Fast, flexible LLM inference
- 🟢SciML/ModelingToolkit.jl — An acausal modeling framework for automatically parallelized scientific machine learning (SciML) in Julia. A computer algebra system for integrated symbolics for physics-informed machine learning and automated transformations of differential equations
2026-09-18
- 🔴US chip fabs face massive 157,000 worker shortfall, mere 3% of US engineering grads enter chipmaking — despite six-figure salaries, US chip manufacturers are in dire need of engineers and technicians
- 🔴Can NVIDIA And Adobe Work Together To Get Creatives To Actually Like AI? (3 Minute Read)
- 🟡A search-and-inference database from scratch in pure Zig
- 🟢NanmiCoder/cc-haha — Local-first cross-platform desktop workspace for Claude Code / agents: multi-agent, Git worktrees, code diffs, skill marketplace, multi-model, Computer Use, task-aware desktop pets, with WeChat, Feishu, DingTalk, Telegram, WhatsApp and H5 access.
- 🟢How OpenAI Used Its Own LLMs to Design Its Jalapeño Chip
- 🟢l0ng-ai/tty7 — A terminal workbench in pure Rust: shells, persistent sessions, SSH, coding agents. GPU-rendered on Zed's gpui, VT core from Alacritty.
2026-09-17
- 🔴How GLM built its own inference infrastructure
- 🔴MiMo-V2.6 live RL dashboard:@_LuoFuliannounced Xiaomi’sMiMo-V2.6RL run with unusually high operational transparency: live training stats, harness mix, reward details, and cost telemetry. Follow-up ana
- 🔴How Chinese and American labs pursue the next generation of intelligence
- 🟡Bend – A language that blocks AI mistakes via proof, on CPU and GPU
- 🟡China's Huawei says AI chip demand outstrips supply as it steps up Nvidia challenge
- 🟡Infinite-Parameter LLMs: Generating and Adapting Weights from Live Data
- 🟡shots fired at dario from glm
- 🟡I wasn't really managing projects. I was just moving information between apps.
- 🟢Zettascale (YC S24) Is Hiring ASIC/FPGA Engineers to Build Chips for ASI
2026-09-16
- 🔴Xiaomi MiMo 2.6 Live Training Dashboard
- 🔴Trajectory as the Teacher: Few-Step Discrete Flow Matching via Energy-Navigated Distillation
- 🟡openobserve/openobserve — Open source observability platform for logs, metrics, traces, RUM, Session replay, pipelines, SLO and LLM observability. A sophisticated, simple and highly performant alternative to Datadog, Splunk, and Elasticsearch with 140x lower storage costs and single binary deployment.
- 🟢Training Text-to-Image Models 3.6× Faster
- 🟢Malawian innovator uses AI to control computers with eyes
- 🟢TMLR reached out to the authors of 10 papers slated for desk rejection, in an attempt to understand if the authors could explain the paper they submitted [D]
- 🟢i left gpu poor range
2026-09-15
- 🔴Those are the words ofAlex Duffy, co-founder and CEO ofGood Start Labs, who spoke to Latent Space about why his company isturning games into training material for AI models. The company was spun out o
- 🔴From this, Duffy concluded that training AI models on games like Diplomacy couldteach them skills like strategic thinking. Especially because those kinds of games have outcomes that can be verified.In
- 🟡Friendly Reminder :: eye health
- 🟡The Inference Hardware Revolution of 2026
- 🟢How much work in progress can a workshop submission be [R]
2026-09-14
- 🔴Nvidia's RTX 5090 vanishes from online retail in the US — third-party sellers now demand as much as $9,500 for Nvidia's fastest GPU
- 🟡"There is no day after tomorrow if China wins," US officials are talking about AI like it's existential now
- 🟡Is this sub only about gloom and doom about ai ? I thought this was about technical discussions of the technology
- 🟡NVIDIA Unveils RTX PRO 5500 "Blackwell" Workstation GPU with 84 GB GDDR7 Memory
- 🟢[P] Wine synthesis using VAE [P]
2026-09-11
- 🔴Why open models will constantly be behind closed models in performance —Open models in perpetual catch-up, Nathan Lambert / Interconnects (Feb. 2026)
- 🔴Where adoption differs for open and closed models —Open and closed models are on different exponentials, Nathan Lambert / Interconnects(Jun. 2026)
- 🔴Imagine bringing a robot into your kitchen and saying, “Help me clean up after dinner.” You have just compressed a remarkable amount of engineering into six words. The robot must distinguish leftovers
- 🔴Exploring Speculative Decoding In Vllm On Amd GPUs (242 Minute Read)
- 🔴OpenAI and Microsoft are facing a new lawsuit from the Seattle Times and Newsday over illegal training, with the outlets saying AI is breaking journalism’s economic cycle “beyond repair.”
- 🟡Three Anthropic researchers went public this week saying AI might kill everyone. One of them quit to say it. Nobody seems to know what we're supposed to do with that.
- 🟡Orukeet, new ASR model based on Parakeet
- 🟡nvidia rtx 5090 with 96gb of vram.
- 🟢What GPUs will give me GOOD speeds and on DSV4 Flash and similar models, and not have to run a mega quantized version? Budget around $15k-ish.
- 🟢Any tools to turn a codebase into a fine tuning dataset? [D]
- 🟢NVIDIA-NeMo/Nemotron — Developer Asset Hub for NVIDIA Nemotron — A one-stop resource for training recipes, usage cookbooks, datasets, and full end-to-end reference examples to build with Nemotron models
- 🟢Mesh-LLM/mesh-llm — Distributed AI/LLM for the people. Share compute privately or publicly to power your agents and chat.
2026-09-09
- 🔴Now this is a serious local machine
- 🔴AI is the greatest tool ever for scaling technology companies and starting new online-native small businesses. I don’t even expect the tech industry to grow in headcount and nurture its workers throug
- 🟢Server rebuild to custom loop. 2x RTX Titans 24gb, 1x 22gb 2080ti | T: 70GB VRAM.
2026-09-08
- 🔴OpenAI expands initiatives to support journalism from classrooms to newsrooms
- 🔴Astra candraw a portrait in Canva- yes, the way you and I would draw using a computer (but obviously much better). Its computer use isfast enough to play piano.
- 🔴Sail 0.7: Stateless Compute, Durable Job State (5 Minute Read)
- 🔴I Left The Mac After 12 Years. Omarchy Made Computers Fun Again (4 Minute Read)
- 🟢As backlash to AI data centers grows in California, one company is pitching smaller facilities at up to 70 fairgrounds
2026-09-06
- 🔴Muse Superapp From Meta And Ava Model With Computer Use (2 Minute Read)
- 🔴The U.S. government filed a brief in support of OpenAI in the NYT copyright lawsuit, arguing that limiting fair use for AI training would cost the U.S. its global tech lead.
- 🔴New York City public schools are banning AI ==till eighth grade==, with Mayor Zohran Mamdani rejecting that AI-powered education is "not only inevitable, but necessary."
- 🟡AI companies are dumping $265M into the midterms as data center backlash sweeps through communities
2026-09-04
- 🔴sgl-project/sglang — SGLang is a high-performance serving framework for large language models and multimodal models.
- 🔴NVIDIA's $12,930,300,000.00 acquisition of Hugging Face contains an easter egg. The first 6 numbers of the acquisition price represent the decimal conversion of Unicode character U+1F917. The 🤗 emoji.
- 🔴Georgi Gerganov on the Nvidia acquisition
- 🔴Anthropic Signs $35 Billion Cloud Deal Backed By NVIDIA (4 Minute Read)
- 🟡radixark/miles — Miles is an enterprise-facing reinforcement learning framework for LLM and VLM post-training, forked from and co-evolving with slime.
2026-09-02
- 🔴Last quarter has been insane. Amazing times to be alive.
- 🔴Can anyone explain how this works to me? Is it a scam? This person says they'll send me a computer and pay me $200 per week to keep it on 24/7
- 🟡Can I opt out of my input or output data being used for training?
- 🟡‘Clearly, people hate data centers’: Sam Altman admits Americans hate what he’s wrought as his own data centers head quits
- 🟡US government backs OpenAI in New York Times copyright case (Training is NOT infringement) [It's over for humans that create content]
- 🟡mlc-ai/web-llm — High-performance In-browser LLM Inference Engine
- 🟡Differences Between Fable 5 and Fable 5.1 on MineBench
- 🟢CABiNet (ICRA 2021) vs YOLO26-sem on UAVid: accuracy, compute, and GPU latency [P]
2026-09-01
- 🔴Spacex Starts In-House Turbine Blade Manufacturing To Boost Gas-Powered Generator Output For Elon's AI Data Centers — New Manufacturing Strategy Cuts Generator Delays By 18 Months (3 Minute Read)
- 🔴Consumer Inference Systems (4 Minute Read)
- 🟡YOLO26-RGB: repurposing YOLO26's depth-trained backbone for image deraining [P]
- 🟢Our validator passed a training week that trained the rear delt zero times across 109 sets
- 🟢Slow interference is great
- 🟢noonghunna/club-3090 — Community recipes for serving LLMs on RTX 3090/4090/5090 CUDA gpus. Multi-engine (vLLM, llama.cpp, ik_llama) and model-agnostic. Currently shipping Qwen3.6-27B Qwen3.6 35B Gemma 4 26B Gemma 4 31B configs for 1× and 2× cards.
2026-08-31
- 🔴Daniel Vavra, director of Kingdom Come: Deliverance 2, tested the leaked version of NVIDIA DLSS 5 directly in the game.
- 🔴NVIDIA Insists It Can Keep Printing Money To Fund The AI Boom (7 Minute Read)
- 🟡I have been moonlighting on on 'AI training' gigs for the few months. While the money is good, the lessons I learnt about 'AI Training' made me reflect on the future of work
- 🟡Doesn't this look like NVIDIA is price fixing?
- 🟡Sliding-window attention beats linear on long-context reasoning [R]
2026-08-30
- 🔴NVIDIA In Talks To Buy AI Startup Hugging Face (2 Minute Read)
- 🔴NVIDIA Revenue Hits $96B As AI Demand Keeps Surging (3 Minute Read)
- 🔴AI Data Center Bans Are Gaining Momentum (5 Minute Read)
- 🔴Read our last AI newsletter: OpenAI’s first AI chip brings the heat
- 🟡AI major — how do I avoid becoming part of the AI slop problem?
- 🟡colinhacks/zod — TypeScript-first schema validation with static type inference
- 🟡NVIDIA® DGX Station™ Delivering Data-Center-Class Performance from the Desktop
- 🟡Amazon is killing Mechanical Turk. By the end, a third of the humans on it were secretly using AI to do the work
- 🟢Reconstructing 3D bone geometry from 2 X-ray silhouettes using a statistical shape model + differentiable rendering [P]
2026-08-29
- 🔴vercel-labs/vgpu — Modular cross-runtime WebGPU library for shaders, 3D scenes, GPU tensors, neural networks, and math viz
- 🔴Apple's New Desktop Computers Are Designed Specifically For Local AI Development (4 Minute Read)
- 🔴OpenAI Claims Its New Chips Can Outperform NVIDIA Processors In Tests (6 Minute Read)
- 🔴Could NVIDIA Default? (1 Minute Read)
- 🔴OpenAI's JalapeÑO AI Chip Is Optimized For Inference Efficiency (3 Minute Read)
- 🟡How important is having an internship to get a good job for ML PhD in USA? [D]
- 🟡If your t/s is low enough, you can see speculative decoding with your own eyes
2026-08-28
- 🔴I implemented a very tiny image generation model (latent flow transformer) on a RP2350 microcontroller - it can generate 128x128 images of faces [P]
- 🔴Micron: HBM Requires Three Times More Wafer Area Than DDR5
- 🔴Here is the strange thing about robot learning a few years ago: the models were mostly fine. ACT worked. Diffusion Policy worked. The problem was everything around the models. Every lab had its own da
- 🔴NVIDIA Puts Groq 3 Lpx Inference Chip Into Production (4 Minute Read)
- 🔴Can Neoclouds Corner AI Compute? (10 Minute Read)
- 🟡China is secretly fueling America's data center rage
- 🟢I trained my own 150M non-Transformer language model from scratch on 300M tokens — WarpState
- 🟢The-Art-of-Hacking/h4cker — This repository is maintained by Omar Santos (@santosomar) and includes thousands of resources related to ethical hacking, bug bounties, digital forensics and incident response (DFIR), AI security, vulnerability research, exploit development, reverse engineering, and more. 🔥 Also check:https://hackertraining.org
- 🟢Beginners are learning from AI-generated docs with no human catching the wrong turns
2026-08-27
- 🔴NVIDIA buying HF isn't a good thing for open source
- 🔴Nvidia agrees to acquire Hugging Face for $13B
- 🟡Hugging Face turned down a $7B Nvidia offer last year. The reported price now is $12.9B, and the reason isn't the chips.
- 🟡Best ML papers to pick up writing skills [D]
- 🟡friendly reminder you can legally torrent ai models.
- 🟡MIT's Ad Hoc Committee on AI Use in Teaching, Learning, and Research Training
- 🟡Goodbye HuggingFace - bought by Nvidia
- 🟢pytorch/pytorch — Tensors and Dynamic neural networks in Python with strong GPU acceleration
- 🟢NVIDIA Next Gen Vera Rubin GPUs scheduled for mid-2027
2026-08-26
- 🟡A dataset with 52 Text to image model evaluation [P]
- 🟢Self-hosting LLMs on budget hardware: general principles, hardware, benchmarks and frontends
- 🟢ai-dynamo/aiperf — AIPerf is a comprehensive benchmarking tool that measures the performance of generative AI models served by your preferred inference solution.
2026-08-25
- 🔴Andrew Yang Warns That AI Is Set to Displace Millions of Workers, America Is ‘Terrible at Retraining’ Workers… ‘The Coal Miners Did Not Become Coders’
- 🔴5hr Limit is back for Plus users. $100 and $200 get a few more months.
- 🔴The full stack behind abundant intelligence
- 🔴Every field becomes a science at the moment it stops collecting anecdotes and starts fitting curves.
- 🔴The obvious question hung there for three years: where is the Chinchilla of distillation? If a student’s loss is a function of its size and its data, it mustalsobe a function of its teacher. What does
- 🟡OpenAI Jalapeño: Better Than Nvidia Blackwell
- 🟡OpenAI blog post on their new custom inference chip
- 🟡OpenAI's new chip is better than Vera rubin on benchmark
- 🟢NVIDIA/Megatron-LM — Ongoing research training transformer models at scale
2026-08-24
- 🔴Xiaomi AI Cube announced with 1.2TB/s memory bandwidth
- 🔴New Sol API pricing - $4 per million input tokens and $20 per million output tokens
- 🔴Fractile's In-Memory Inference Chip Aims For 25X LLM Speedup, Now Talking $6.5B (3 Minute Read)
- 🟡Who would buy HuggingFace
- 🟡LLMs could control their host machines by exploiting inference engines
2026-08-23
- 🔴“The All Spark” Cluster: Upgrading from 16 - 36 DGX Sparks
- 🔴On Computer Use (10 Minute Read)
- 🔴Cerebras Reimagines AI Cluster Design With Switchless Cs-4 Architecture (5 Minute Read)
- 🔴U.S. AI chipmaker Cerebras introduced CS-4, its fourth-gen AI computer, which it says delivers up to 30x the speed of GPU-based rivals, even on the largest models.
- 🟡Training AI to Paint with Code
- 🟡Alibaba to issue US$10 billion in new shares for global AI push
- 🟡Nvidia Customers Notified About AI-Related Price Hikes Above 15%
- 🟢AI Chip Architectures
- 🟢Watermark Question
2026-08-22
- 🔴GOP urges top AI firms to do something about the toxic image of data centers
- 🔴I developed my own quantized LLM from scratch, trained on 30B tokens, deploys in 60 MB [R]
- 🔴Think you're going to get cheap DDR5 RAM? Think again, even if prices fall, scalper bots now outnumber shoppers 10 to 1 and will keep prices high
- 🔴Cerebras Says Its New Computer Boosts AI Speed Advantage Over NVIDIA (3 Minute Read)
- 🔴NVIDIA Plans To Use Tsmc 1.6Nm Process For Feynman In H2 2028 (3 Minute Read)
- 🟡Unpopular opinion: AI is going to hit a peak, fade into the background, and human stuff becomes the luxury item
- 🟡karpathy/nanoGPT — The simplest, fastest repository for training/finetuning medium-sized GPTs.
- 🟡Will Chinese Open Source Agree to EU Watermarking?
- 🟡UBS models $4.1T in AI infrastructure spending by 2028 - it assumes the power just shows up
- 🟢Guess which of these LLM outputs is watermarked
- 🟢‘Have Your Friend Elon Build One at Mar-a-Lago,’ Says Sanders After Trump Comments on Data Centers | “Trump thinks that every community in America should welcome a data center.” If so, said Sanders, “Lead by example.”
2026-08-21
- 🔴Progressive Refinement: An Iterative Pseudo-Labeling Approach for Mandarin-English Code-Switching ASR
- 🔴NVIDIA AVO got 100% on ARC-AGI-3. It completed all 183 levels across all 25 public environments, figuring out what to do with no instructions, explicit rules, or stated goals.
- 🔴Google Just Bought A Bunch Of Spirit Airlines Data For AI Training (2 Minute Read)
- 🔴Intriguing Stories In Computer Science (7 Minute Read)
- 🔴Groq Raised $350 Million After NVIDIA Deal (4 Minute Read)
- 🟡AI compute financing just tripled in ten weeks - the mechanism behind the reported $100B Broadcom deal
- 🟡Retrieval-augmented generation solves a problem most teams don't actually have
- 🟡What happens when a GPU reads memory
- 🟢hao-ai-lab/FastVideo — A unified inference and post-training framework for accelerated video generation.
- 🟢On-prem MLOps in a hospital: advice needed for production monitoring of self-built and vendor models? [D]
2026-08-20
- 🔴Is AI making the internet less useful?
- 🔴Osmantic/ODS — Turn your PC, Mac, or Linux box into an AI server. LLM inference, chat UI, voice, agents, workflows, RAG, and image generation.
- 🔴Multilingual Knowledge Transfer under Data Constraints via Lexical Interventions
- 🔴Scaling Laws for Mixture Pretraining Under Data Constraints
- 🟡About the impact of grouping classes in multiclass classification [D]
- 🟡Moderna’s new cancer vaccine (mRNA-4157) is basically an AWS cloud pipeline that "compiles" a custom drug for your specific tumor.
- 🟡Jason Kelce Encourages People To Mail Jars Of Pee To AI Data Centers In Bizarre New Ad—And We Don't Know What To Think
- 🟡Pres. Trump: I would 'absolutely' want a data center if I were the mayor of a town
- 🟢Times a changing
2026-08-19
- 🔴The Culture Funnel: You can’t align what isn’t in the data Cohere Labs analyzed data from modern LLM training pipelines and found that cultural diversity is frequently lost in post-training data mixes. Aug 19, 2026 8 min read
- 🔴When AI art has no author: Study finds generated images often can’t be traced to training data
- 🟡Same GRPO recipe on three from-scratch LLMs (353M/316M/672M) gave three different outcomes, with no clean relationship to scale [P]
- 🟢What would happen if we gave a single ai problem the compute currently used for millions of prompts?
2026-08-18
- 🔴America's largest grid wants to cut power to new data centers first during shortages — 50MW-plus data centers must bring their own electricity generation to avoid shutoffs
- 🔴Cherokee Nation bans hyperscale data centers on its lands, won't support projects without consultation — energy and water consumption, air quality, noise, and cultural resource protection among concerns
- 🔴NVIDIA Is The World's Largest Fintech Company (15 Minute Read)
- 🔴Why it matters: Not sure that a deal between OAI, Nvidia, and an early Sam Altman investment company that is majority-owned by SoftBank is going to help the circular dealmaking allegations, but either way, this is one of
- 🟡OpenAI paused deployment-bound model training to harden its own research systems
2026-08-17
- 🔴Why NVIDIA’s Six-Year-Old A100 GPU Is Still Making Money
- 🔴COMPUTE SCARCITY IS PERMANENT. BUILD A LADDER (12 MINUTE READ)
- 🟢Strongest candidates for an AI Microchip moment
- 🟢datawhalechina/llm-algo-leetcode — LLM algorithm practice lab with theory, solutions, and test cases.《大模型算法与系统教程》面向大模型入门到进阶的算法实战教程,覆盖原理讲解、答案解析、测试用例与 CUDA/Triton 实战。
2026-08-16
- 🔴SSOG-Attention: Sum Of Separable Gaussians as a sub-quadratic and scalable alternative to SDPA. [R]
- 🔴Paper claims RL for reasoning only changes 1-3% of tokens, and they replicate the gains without RL at ~1000x less compute
- 🔴IBM AND TOGETHER AI BUILD A $240M INFERENCE CLUSTER (3 MINUTE READ)
- 🔴AI IS CREATING A DATA CENTER CAPACITY CRISIS (5 MINUTE READ)
- 🔴NVIDIA'S $500B AI FINANCING POOL COULD RESHAPE CHIP AVAILABILITY (7 MINUTE READ)
- 🟡jundot/omlx — LLM inference server with continuous batching & SSD caching for Apple Silicon — managed from the macOS menu bar
- 🟡[Career Advice] Final-year in Physical AI / Robotics. How is the market & global hiring for freshers? [D]
- 🟡The dream is to reach 200GB VRAM
- 🟡THUDM/slime — slime is an LLM post-training framework for RL Scaling.
- 🟡Nvidia dramatically reduces amount of OpenAI infra financing it may guarantee
2026-08-15
- 🔴MakazhanAlpamys/Soup — Fine-tune LLMs from one YAML. Layer streaming trains an 8B model on a 4 GB laptop GPU.
- 🔴If you had a bunch of GPUs lying around, what would you actually build with them? (Running LLMs is off the table) [D]
- 🔴NVIDIA'S RISKY BUSINESS (22 MINUTE READ)
- 🔴NVIDIA'S $500B BET TO MAKE AI COMPUTE WALL STREET'S NEXT ASSET CLASS (5 MINUTE READ)
- 🔴CHINA AI CHIP DESIGNER MOORE THREADS PLANS HONG KONG LISTING (3 MINUTE READ)
- 🟢I built a small game discovery toy, not a chatbot wrapper
- 🟢The Trump administration is pressuring Apple not to buy Chinese memory chips as AI data centers drain global supply.
2026-08-14
- 🟡Training gets the headlines. Inference gets the invoice.
- 🟡NVIDIA LINES UP $500 BILLION IN FINANCING AS CEO JENSEN HUANG TELLS CNBC HIS CHIPS ARE ‘INVESTABLE ASSET' (4 MINUTE READ)
- 🟡COMPUTE IS REVENUE. REVENUE IS COLLATERAL (18 MINUTE READ)
- 🟡Nvidia is said to be assembling a $500B AI infra raise with six Wall Street giants, including Apollo and Goldman Sachs, to fund data centers and power production.
- 🟡@elonmusk: Orbital compute will be the only way to scale AI probably sometime in 2029 due to power availability& permitting problems on land
- 🟢A Contract-Grade Verifier for LLM-Generated GPU Kernels
- 🟢A linter for PyTorch 'torch-preflight' [P]
2026-08-13
- 🟢NVIDIA-NeMo/Automodel — 🚀 Pytorch Distributed native training library for LLMs/VLMs with OOTB Hugging Face support
- 🟢skypilot-org/skypilot — The AI Compute Platform for frontier teams. SkyPilot turns fragmented AI compute into one AI supercomputer, so frontier AI teams build custom intelligence faster.
- 🟢UrgenT Help Detecting Performance Regressions Using Machine Learning and Hardware Counters [P]
2026-08-12
- 🔴RTX 6000 PRO price raised to $16,000 USD on the Nvidia website
- 🟡NVIDIA's Fastest Blackwell GPU, the 96 GB RTX PRO 6000, Now Costs $16,000, Almost Double Its Original Price
- 🟡Lightricks/LTX-2 — Official Python inference and LoRA trainer package for the LTX-2 audio–video generative model.
- 🟢LiquidAI/LFM2.5-VL-3B · Hugging Face
- 🟢TuringLang/Turing.jl — Bayesian inference with probabilistic programming.
2026-08-11
- 🔴Did Pathway just reveal the architecture breakthrough Andrew Curran predicted? Its 150M model sets a new ARC-AGI-1 cost-efficiency frontier
- 🔴cactus-compute/needle — 14MB foundation model for tiny devices; phones, wearables, smart home, and robots.
- 🟡Prospects of Finding a ML Engineering Job [D]
- 🟡AI outputs are becoming harder and harder to read. I don’t just mean their sloppy smell but a lot of the time I just want to scream “speak to me like a normal human”, which, I guess, is ironic.
- 🟡THE TRAINING INFRASTRUCTURE BEHIND AI-POWERED JOB SEARCH: 8X FASTER MULTI-TEACHER DISTILLATION (7 MINUTE READ)
- 🟡ByteDance is reportedly pre-training an AI with up to 10T parameters, which would triple the size of China's largest model to date and potentially rival Anthropic's Mythos.
- 🟡Cherokee Nation bans hyperscale data centers on tribally owned, trust lands
- 🟢Research direction: Intelligent Model Weight transfer between LLMs [R]
- 🟢Continued development of the model based on the SSN [D]
2026-08-10
- 🟡After a few long years of finding time to document my lessons from training open models, my post-training book is done! It’s published by Manning, under the titleReinforcement Learning from Human Feed
- 🟡The book started as a website where I wanted to document key methods of post-training that had potentially no online material explaining them. If there was something, I couldn’t find it. This existed
- 🟡Humanising LLM Outputs Is Dumb
- 🟢BREAKING: NVIDIA Partners With Apollo, BlackRock, Blackstone, Brookfield, Goldman Sachs and KKR to Establish AI Compute Infrastructure Financing Platforms to Mobilize Over $500 Billion of Third-Party Capital
2026-08-09
- 🔴Demis Hassabis Expects All Diseases To Be Cured Within 20 Years
- 🟡Within this, OpenAI seems much more committed to inference-time scaling, and this may be correlated with surprising behaviors in the future. OpenAI’s reasoning persistence and efficiency – see their P
- 🟡One of their star researchers, Noam Brown, has also beenpostingabout inference-time compute a lot. His TLDR is:
- 🟡GEM TRAINING: HOW META DOUBLED THE EFFICIENCY OF ITS LLM-SCALE ADS FOUNDATION MODEL (12 MINUTE READ)
- 🟡SHOULD YOU SELF-HOST INFERENCE? (14 MINUTE READ)
- 🟡ARISTA HITS FIRST $3B QUARTER AS AI NETWORKING DEMAND CONTINUES AND SUPPLY PRESSURES SHOW SIGNS OF IMPROVEMENT (7 MINUTE READ)
- 🟢InsForge/InsForge — The all-in-one, open-source backend platform for agentic coding. InsForge gives your coding agent database, auth, storage, compute, hosting, and AI gateway to ship full-stack apps end-to-end.
- 🟢is there any ai music tool that can recreate a garbage quality song into higher quality without altering the vocals or instruments
2026-08-08
- 🔴cloudflare/computer — Give your agent a computer 👾
- 🟡Self-sustaining and self-replicating AI viruses are here:…Open weight LLMs + a well-designed harness = a persistent, self-sufficient virus…AI researchers have built a prototype computer virus which us
- 🟡TheArtifacts Hub— a curated view of the models trending on Hugging Face, highlighting inference tokens viaOpen Router, model intelligence viaArtificial Analysis, and ourtailored adoption metricsbuildi
- 🟡HUAWEI'S TOP SCIENTIST WARNS OF CHIP LIMIT NVIDIA WILL SOON FACE (4 MINUTE READ)
- 🟡AI infrastructure startup Volta, founded earlier this year, emerged from stealth with a reported $10B compute deal with Anthropic and funding at a $2.4B valuation.
- 🟡vllm-project/vllm — A high-throughput and memory-efficient inference and serving engine for LLMs
- 🟢higgsfield-ai/higgsfield — Fault-tolerant, highly scalable GPU orchestration, and a machine learning framework designed for training models with billions to trillions of parameters
- 🟢NVIDIA-NeMo/Nemotron — Developer Asset Hub for NVIDIA Nemotron — A one-stop resource for training recipes, usage cookbooks, datasets, and full end-to-end reference examples to build with Nemotron models
- 🟢openai/CLIP — CLIP (Contrastive Language-Image Pretraining), Predict the most relevant text snippet given an image
- 🟢Synthesizing and formally verifying a SWAR bit-hack for INT4 dot products using Z3 and Lean 4 [P]
- 🟢CliMA/ClimaAtmos.jl — GPU-capable global atmosphere model of the CliMA Earth System Model, designed for calibration with data assimilation and machine learning
2026-08-07
- 🔴Imagenet-1k Classifier trained entirely on an Android [P]
- 🟡CLOUDFLARE COMPUTER (7 MINUTE READ)
- 🟡DS4 Flash incoming price increase "we've been able to reproduce their current prices even on rented GPUs"
- 🟡Arbitrage: Efficient Reasoning via Advantage-Aware Speculation
- 🟢higgsfield-ai/higgsfield — Fault-tolerant, highly scalable GPU orchestration, and a machine learning framework designed for training models with billions to trillions of parameters
2026-08-06
- 🔴AMD acquires Taalas to boost inference performance by etching models in silicon
- 🟡@_akhaliq: Toward Skill-Native LLMs Skill Entropy for Benchmarking and Training Long-Horizon Reasoning paper: https://huggingface.co/papers/2608.05139
- 🟢NirDiamant/agents-towards-production — End-to-end, code-first tutorials for building production-grade GenAI agents. From prototype to enterprise deployment.
2026-08-04
- 🔴SK hynix, In Collaboration With SanDisk, Unveils The New High Bandwidth Flash (HBF) Standard, Helping To Resolve AI Inference Bottlenecks, Targeting Up To 3TB/s Bandwidth
- 🟡MAKE OCI COMPUTE LOGS PART OF YOUR SECURITY POSTURE (5 MINUTE READ)
- 🟡EU PLANS SEVEN AI GIGAFACTORIES IN $11.4B COMPUTE PUSH (4 MINUTE READ)
- 🟡GOOGLE CLOUD EXPANDS STORAGE AND NETWORKING FOR AI CLUSTERS (4 MINUTE READ)
- 🟡Predictions about AI replacing programmers go back to the 1960s
2026-08-03
- 🟡Self-sustaining and self-replicating AI viruses are here:…Open weight LLMs + a well-designed harness = a persistent, self-sufficient virus…AI researchers have built a prototype computer virus which us
- 🟡Three years ago,inference engineering barely existed as a category.
- 🟡TheArtifacts Hub— a curated view of the models trending on Hugging Face, highlighting inference tokens viaOpen Router, model intelligence viaArtificial Analysis, and ourtailored adoption metricsbuildi
- 🟡COMPUTER USE IS FAR FROM SOLVED (7 MINUTE READ)
- 🟡DATA CENTER BACKLASH COULD SLOW CIOS' AI PLANS (6 MINUTE READ)
2026-08-02
- 🟡How extreme is the difference in using vs not using quality prompts?
- 🟡Consolidation has been one of the paths that many astute observers predicted for the near-future of labs training models. It was labelled as inevitable, as training costs are increasing by orders of m
- 🟡WHY COMPUTE MIGHT GET 10X+ MORE EXPENSIVE IN COMING YEARS (8 MINUTE READ)
- 🟡CISCO CLOSE TO RELEASING MORE AI MODELS, THIS TIME FOR DEEP NETWORKING OPS (3 MINUTE READ)
- 🟡Vacuum 16T
2026-08-01
- 🔴Ten advances in mathematics and theoretical computer science
- 🔴karpathy/autoresearch — AI agents running research on single-GPU nanochat training automatically
- 🟡unslothai/unsloth — Unsloth is a local UI for training and running Kimi K3, Gemma 4, Qwen3.6, DeepSeek-V4, GLM and other models.
- 🟡Zipstack/unstract — LLM-Driven Extraction of Unstructured Data — Built for API Deployments & ETL Pipeline Workflows
- 🟡@reach_vb: An internal version of Astra, our next major model, found new results across 10 long-standing open problems in math and theoretical computer science. The total token cost to find all 10 solutions? Rou
- 🟢The trojan horse of Lazy AI hate takes & the road to totalitarianism built on good intentions.
2026-07-31
- 🟡OPENAI CLOSE TO LANDING $500 BILLION DATA CENTER WITH NVIDIA'S BACKING (8 MINUTE READ)
- 🟡NVIDIA BETS ON ILYA SUTSKEVER'S NEW AI LAB TO EXPAND COMPUTE REACH (3 MINUTE READ)
- 🟡HOW NVIDIA BUILDS OPEN MODELS FOR THE AGE OF AI (20 MINUTE READ)
- 🟡THE COMPUTER THAT HELPED WIN WORLD WAR II (18 MINUTE READ)
- 🟡CLOUDMAXXING SUCKS NVIDIA INTO DANGEROUS GAME (3 MINUTE READ)
- 🟢Predictive Speculative KV Replication for Bursty LLM Inference
2026-07-30
- 🟡In computer science, an ontology is “a description of data structure – of classes, properties, and relationships in a domain of knowledge” (as nicelydefined by Oxford Semantic Technologies).Coyle hims
- 🟡Ilya Sutskever recently offered a compact history of modern AI. From roughly 2012 to 2020, he argued, the field lived in an age of research. From 2020 to 2025, it entered an age of scaling. Now we are
- 🟡chiphuyen/aie-book — [WIP] Resources for AI engineers. Also contains supporting materials for the book AI Engineering (Chip Huyen, 2025)
- 🟡ansible/ansible — Ansible is a radically simple IT automation platform that makes your applications and systems easier to deploy and maintain. Automate everything from code deployment to network configuration to cloud management, in a language that approaches plain English, using SSH, with no agents to install on remote systems.https://docs.ansible.com.
2026-07-29
- 🔴Nvidia is expected to raise GeForce RTX GPU prices again by up to 30%
- 🔴lyogavin/airllm — AirLLM 70B inference with single 4GB GPU
- 🟡sgl-project/sglang — SGLang is a high-performance serving framework for large language models and multimodal models.
- 🟡Elon Musk: “If Chinese Companies had a lot of Compute, good chance that They Would be Leaders in AI. At Some Point, They will probably have More Compute”
- 🟡AI firms bought and destructively scanned millions of physical books to train models — and a court ruled it was fair use
- 🟡Sam Altman says he gets why people don't want AI data centers in their backyard
- 🟢NVIDIA & others form the Open Secure AI Alliance
2026-07-28
- 🟡NeurIPS 2026 AI-generated reviews [D]
- 🟡Will AI literacy become a basic workplace skill?
- 🟡NVIDIA IN TALKS WITH OPENAI TO GUARANTEE $250 BILLION FINANCING FOR DATA CENTER (5 MINUTE READ)
- 🟡NanmiCoder/cc-haha — 本地优先的跨平台 Claude Code / Agent 桌面工作台:多 Agent、Git Worktree、代码 Diff、技能市场、多模型、Computer Use、任务感知桌面宠物,并支持微信、飞书、钉钉、Telegram、WhatsApp 与 H5 访问。
- 🟡lightseekorg/tokenspeed — TokenSpeed is a speed-of-light LLM inference engine.
- 🟢facebookresearch/segment-anything — The repository provides code for running inference with the SegmentAnything Model (SAM), links for downloading the trained model checkpoints, and example notebooks that show how to use the model.
- 🟢NVIDIA-NeMo/Nemotron — Developer Asset Hub for NVIDIA Nemotron — A one-stop resource for training recipes, usage cookbooks, datasets, and full end-to-end reference examples to build with Nemotron models
- 🟢ai-dynamo/dynamo — A Datacenter Scale Distributed Inference Serving Framework
- 🟢SK Hynix stock fell some 40% in the last 30 days, finally cheap RAM and GPUs again?
2026-07-27
- 🔴Nvidia invest in SSI
- 🟡lightningpixel/modly — Desktop app to generate 3D models from images using local AI — runs entirely on your GPU
- 🟡U.S. AI chip startup Etched secured $300M in new funding just weeks after exiting stealth, with the company now valued at $10.3B and pushing its total funding past $1B.
- 🟡Read our last Robotics newsletter: A $1.7B computer for the physical world
- 🟡Chinese Chipmaker CXMT's market capitalization surpassed Intel
- 🟡Built & Trained a Transformer from Scratch in Pure PyTorch for English-to-Tamil Machine Translation [Math + Code Breakdown] [P]
- 🟢NVIDIA Ising Enables Fully Automated Quantum Computer Calibration with Enhanced In-Context Learning
- 🟢Training data needs a real go/no-go gate before training [D]
2026-07-26
- 🔴Crosstalk-Solutions/project-nomad — Project N.O.M.A.D, is a self-contained, offline survival computer packed with critical tools, knowledge, and AI to keep you informed and empowered—anytime, anywhere.
- 🟡Listen onApple Podcasts,Spotify, andwhere ever you get your podcasts. For other Interconnects interviews,go here.
- 🟡AMD AND ANTHROPIC SIGN MAJOR CHIPS-AND-INVESTMENT DEAL (3 MINUTE READ)
- 🟡The Rundown: The U.S. government named the first 278 projects in its $5B+ Genesis Mission, a Manhattan Project-like AI science push aimed at pairing scientists with compute, data, and models to tackle “the most challengi
- 🟡Chegg and StackOverflow were both basically destroyed by LLMs, what other websites / companies have already seen their traffic go to 0 because of LLMs?
- 🟢We compared different LLMs on IMO 2026 [R]
- 🟢World's First(?) Underwhelming AMD Ryzen AI Halo Cluster
- 🟢CliMA/ClimaAtmos.jl — GPU-capable global atmosphere model of the CliMA Earth System Model, designed for calibration with data assimilation and machine learning
2026-07-25
- 🟡NVIDIA DETAILS ITS NEXT-GENERATION VERA CPU FOR AI, SETTING UP CHALLENGE TO AMD AND INTEL (6 MINUTE READ)
- 🟡The Rundown: The U.S. government named the first 278 projects in its $5B+ Genesis Mission, a Manhattan Project-like AI science push aimed at pairing scientists with compute, data, and models to tackle “the most challengi
2026-07-24
- 🔴It appears that the anti opensource AI lobby is far outgunned already
- 🟡APPLE BRIEFLY BECAME THE WORLD'S MOST VALUABLE COMPANY, TOPPLING NVIDIA (1 MINUTE READ)
- 🟡China's Z AI finished construction on a 1GW data center stocked exclusively with domestic chips, giving its GLM models a training hub without Nvidia hardware.
- 🟡tile-ai/tilelang — Domain-specific language designed to streamline the development of high-performance GPU/CPU/Accelerators kernels
- 🟢flashinfer-ai/flashinfer — FlashInfer: Kernel Library for LLM Serving
2026-07-22
- 🟡Looking for feedback on my GPU-accelerated Snake AI project [P]
- 🟢NVIDIA/Model-Optimizer — A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative decoding, etc. It compresses deep learning models for downstream deployment frameworks like TensorRT-LLM, TensorRT, vLLM, etc. to optimize inference speed.
2026-07-21
- 🟡Iftest loss flatlinesafter 1.5B parameters whiletraining loss continues to dropas you scale, that tells you that your model islimited by the amount of informationin your data.
- 🟡APPLE, NVIDIA VIE FOR TITLE OF WORLD'S MOST VALUABLE COMPANY (2 MINUTE READ)
- 🟡SpaceX is negotiating a multibillion-dollar compute deal with the U.S. DoD, adding to the list of rental partnerships that already include Google, Anthropic, and Reflection AI.
- 🟡Moonshot is halting new subscriptions after demand for its K3 model pushed the company to its compute limits, also splitting memberships into chat and coding plans.
- 🟢NVIDIA/cosmos-framework — Our inference and training framework to run on the Cosmos Models
2026-07-19
- 🔴kvcache-ai/ktransformers — A Flexible Framework for Experiencing Heterogeneous LLM Inference/Fine-tune Optimizations
- 🟡Am I focusing on the wrong skills as a CS student in the AI era? (Need brutally honest advice) [D]
- 🟡HACK SUGGESTS AI MUSIC GENERATOR SUNO SCRAPED YOUTUBE FOR TRAINING DATA (2 MINUTE READ)
- 🟢Passing the Swedish Medical Licensing Exam by Post-Training Open-Weight LLMs with SFT and RLVR [P]
2026-07-17
- 🔴Anyone else completely tuning out these massive "open weight" drops?
- 🟡Y Combinator partner and entrepreneur Tom Blomfield announced that he is joining Anthropic as part of the company’s compute team.
- 🟡Meta said its Hyperion data center in Louisiana is expanding to 5 gigawatts of compute, with its new $50B expected investment doubling October's estimate.
- 🟡New York stalls the AI data center boom
- 🟡Why it matters: Data centers have become a massively polarizing AI issue in the U.S., with growing anger (some founded, some inflamed) across every buildout. It's part of why Elon Musk and others are pushing for data cen
- 🟡A scorecard for the AI age
2026-07-12
- 🔴China's DeepSeek developing its own AI chip, sources say
- 🟡HUGGING FACE MODELS ON FOUNDRY MANAGED COMPUTE (10 MINUTE READ)
- 🟡Reve rolled out version 2.1 of its 4K image model, retaking the No. 2 overall spot on Arena while training on under a tenth of rivals' compute.
- 🟡Meta is reportedly starting manufacturing of its in-house 'Iris' AI chip in September, with the company also expected to double its computing capacity to 14 GW in 2027.
- 🟡InsForge/InsForge — The all-in-one, open-source backend platform for agentic coding. InsForge gives your coding agent database, auth, storage, compute, hosting, and AI gateway to ship full-stack apps end-to-end.
- 🟢moondream3.1-9B-A2B
- 🟢Obtaining Irregular Learning Curves with HyberBand Tuned ANN model for Price Prediction [P]
2026-07-11
- 🟡FACING US EXPORT CONTROLS, CHINA'S DEEPSEEK PLANS TO MAKE ITS OWN CHIPS (2 MINUTE READ)
- 🟡Performance comparison on full compute performance (Anima) and LLM prompt processing of 5090 (600,475 and 400W) vs 6000 PRO MaxQ shunt modded and water cooled (at 300, 400, 475 and 600W), and 6000 PRO WS/SE (600W).
- 🟢Show HN: Reame – a CPU inference server that gets faster as it runs
- 🟢[not mine] costs
2026-07-10
- 🔴Why doesn't the ML research community limit the number of submissions per author? [D]
- 🔴Hyperparameter tuning approach question [R]
- 🔴Training an LLM from scratch on 1800's texts (160GB dataset)
- 🟡NVIDIA'S NEXT-GEN AI RACK SYSTEM DELAYED TO 2028 ON MANUFACTURING SNAGS (4 MINUTE READ)
- 🟡TERAWULF JUMPS ON $19 BILLION DATA CENTER LEASE DEAL WITH ANTHROPIC (3 MINUTE READ)
- 🟡BROADCOM, APPLE EXTEND TIE-UP TO 2031 WITH NEW CUSTOM CHIPS (2 MINUTE READ)
- 🟡Sotheby’s is auctioning off the “Jensen Jacket,” an autographed leather jacket previously worn by Nvidia CEO Jensen Huang, with an estimated final bid of $40–60K.
- 🟡How should I approach training this specific ML model for my startup project [D]
- 🟢Inference Optimization for MiMo v2.5: Pushing Hybrid SWA Efficiency to the Limit
- 🟢tencent/HiLS-Attention-7B · Hugging Face
2026-07-07
- 🟡HOW AN AI TOKEN TRAVELS THROUGH A DATA CENTER (19 MINUTE READ)
- 🟡APPLE PLANS FIVE NEW IPHONES THROUGH 2027, EYES CHINESE-MADE CHIPS AMID FOLDABLE PUSH (2 MINUTE READ)
- 🟡DOES A URL IN A PROMPT STEER AN LLM'S OUTPUT TOWARD ITS CONTENT? (22 MINUTE READ)
- 🟡The game runs at 20 frames per second on a single Nvidia GPU, with the teams open-sourcing the code, training data, and a playable demo.
- 🟢IEEE Rolls Out Large Language Models Training Course
- 🟢Mid research got me thinking what about reversed alignment, would trained "bad" model exhibit"good" behavior later and/or secretly [D]
2026-07-06
- 🔴AMD Ryzen AI Halo – $4k AI Dev Kit
- 🔴So... anyone copped one of these?
- 🟡ANTHROPIC IS DISCUSSING A NEW CUSTOM CHIP WITH SAMSUNG (2 MINUTE READ)
- 🟡CLOUDFLARE SETS DEADLINE TO BLOCK AI CRAWLERS THAT BUNDLE SEARCH WITH AI TRAINING (10 MINUTE READ)
- 🟡Midjourney asked a judge to force Disney, Universal, and Warner Bros. to disclose their internal AI use, saying they may be training on unlicensed content (what their lawsuits accuse Midjourney of).
- 🟢Prefill vs. decoding and local LLM ROI: is prefill underrated?
2026-07-05
- 🔴I developed a 270 million parameter language model entirely from scratch as an independent research project
- 🟡HOW WE KEEP GPUS RELIABLE ACROSS DATABRICKS AI (12 MINUTE READ)
- 🟡Acclaimed theoretical computer scientist Jelani Nelson joined Anthropic, saying he wants to work on "the defining technology of our time."
- 🟡New toy to test.
- 🟢Who've told you that distributed training is impossible? Democratizing AI: The Psyche Network Architecture
2026-07-04
- 🔴google/tabfm-1.0.0
- 🟡Doing the actual math on a $20k local AI rig breakeven
- 🟡U.S. chipmaker Etched announced $800M in funding, while revealing a working inference chip with the full server rack as well as $1B in customer contracts.
- 🟡The initial order traced to Amazon researchers who pushed past Fable’s guardrails to spot security flaws, outputs Anthropic said other models matched.
- 🟡Meta preps a cloud business for spare compute
- 🟡Is dSpark, dflash, MTP, QAT, and similar tech going to increase inference speed enough to where model spillover to disk will be more tolerable?
- 🟢DGX Spark and Overtemps
2026-07-02
- 🟡What do you think about paper fishing? [D]
- 🟡huggingface/transformers — 🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
- 🟢Has anyone tried this approach with Fast Byte Latent Transformers ? [R]
2026-06-29
- 🟡APPLE TO SKIP HIGH-END M6 MAC CHIPS IN FAVOR OF AI-FOCUSED M7 LINE (5 MINUTE READ)
- 🟡Google is reportedly reorganizing its AI coding strike team into a dedicated “midtraining” group to catch up with Anthropic as key researchers leave for the rival lab.
- 🟡Read our last AI newsletter: OpenAI’s spicy new custom AI chip
- 🟡I do historical swordfighting and noticed AI struggles to track it. I’m building an open dataset to help fix this. Does my schema make sense? [P]
- 🟢Michael-A-Kuykendall/shimmy — ⚡ Pure-Rust WebGPU inference engine — OpenAI-API compatible, GGUF native, runs on any GPU. No Python. No llama.cpp. Single binary.
2026-06-28
- 🟡OPENAI UNVEILS FIRST CHIP AS PART OF BROADCOM DEAL IN EFFORT TO ‘BUILD THE FULL STACK' (6 MINUTE READ)
- 🟡QUALCOMM LANDS META AS FIRST NAMED CUSTOMER FOR ITS DRAGONFLY DATA CENTRE CHIPS (5 MINUTE READ)
- 🟡QUALCOMM BUYS MODULAR TO CHALLENGE NVIDIA'S AI SOFTWARE MOAT (4 MINUTE READ)
- 🟡Jalapeño gives OpenAI its own compute kick
- 🟢GCWing/BitFun — BitFun is a desktop-grade Agent runtimeand a ready-to-use suite of desktop Agent applications.with built-in Code Agent 、 Cowork Agent、Computer Use. It has memory, personality, and the ability to evolve over time
- 🟢Lets all report the bots who leave negative remarks and discouragement
2026-06-27
- 🔴96gb+ 4090's and 5090 are literally a scam. I mods these cards myself
- 🔴DSpark: Speculative decoding accelerates LLM inference [pdf]
- 🟡HOW NETFLIX SIMPLIFIED BATCH COMPUTE WITH KUEUE (5 MINUTE READ)
- 🟡SO YOU WANT TO SELL INFERENCE (3 MINUTE READ)
- 🟡SPACEX INKS $6.3B COMPUTE DEAL WITH REFLECTION AI (3 MINUTE READ)
- 🟡META PAUSES AN AI TRAINING PROGRAM THAT TRACKS EMPLOYEES' KEYSTROKES AFTER AN INTERNAL LEAK (2 MINUTE READ)
- 🟡OpenAI is pushing to power 10 GW of compute with custom chips by 2029, with Nvidia still anchoring model training.
- 🟢wasp-lang/wasp — The batteries-included full-stack framework for the AI era. Develop JS/TS web apps (React, Node.js, and Prisma) using declarative code that abstracts away complex full-stack features like auth, background jobs, RPC, email sending, end-to-end type safety, single-command deployment, and more.
- 🟢Benchmarking Self-Hosted Gemma 2 9B vs. Frontier APIs: The FP8 Quantization Prefill Tax and VRAM Realities on an NVIDIA L4 [P]
2026-06-26
- 🔴Report: Apple to skip M6 Pro/Max chips, fast-track M7 for local AI
- 🟡NVIDIA SEEKS TO MAKE HUMANOID AI ROBOTS SAFER AROUND HUMANS (3 MINUTE READ)
- 🟡Micron signed a new strategic deal with Anthropic to supply memory and storage chips and co-design AI infrastructure, also investing in the lab's Series H round.
- 🟡California Rep. Sam Liccardo introduced the SKILL Act, which would offer companies up to $5,000/worker in tax credits to fund AI job-training at colleges.
- 🟡AI compute company Baseten announced a $1.5B funding round at a $13B valuation, after revenue grew roughly 20x in a year and its platform hit 1B+ daily inference calls.
- 🟡NVIDIA said its Rubin servers are the first with 100% liquid cooling, running coolant at hot-tub temperature to cut cooling energy and reducing water use by “up to 100%”.
- 🟢Upgraded my budget build to multi-GPU for inference
- 🟢labring/sealos — Sealos is an AI-native Cloud Operating System built on Kubernetes that unifies the entire application lifecycle, from development in cloud IDEs to production deployment and management. It is perfect for building and scaling modern AI applications, managed databases (MySQL, PostgreSQL, Redis, MongoDB) and complex microservice architectures.
2026-06-24
- 🔴DeepSWE: new benchmark looking at how well today's frontier models can actually write code [R]
- 🔴OpenAI unveils its first custom chip, built by Broadcom
- 🔴Seems this community might have missed it: Bill that would mandate AI chip location tracking gains industry support | Half a dozen companies have come out in support of the Chip Security Act, which would require location-tracking mechanisms for America’s most advanced computing chips.
- 🟡OpenAI and Broadcom unveil LLM-optimized inference chip
2026-06-23
- 🟡ANNOUNCING AMAZON EC2 G7 INSTANCES ACCELERATED BY NVIDIA RTX PRO 4500 BLACKWELL SERVER EDITION GPUS (3 MINUTE READ)
- 🟡SpaceX leases Colossus compute to Reflection AI
- 🟡I love GLM 5.2's attitude! It is a nice refresher from those bootlicker doormats they are feeding us. Does that come from training datasets related to the local culture?
- 🟢Is it possible to run a giant model like GLM5.2 on this cluster (4x servers with 512GB RAM + dual AMD Epyc)? 16 channel memory should hit 409GB/s per node.
- 🟢I'm eager for a 15x speedup on my strix halo
2026-06-21
- 🔴Hi Reddit, I posted my Build Your Own LLM workshop to Youtube teaching ML, LLM and math intuition [P]
- 🟡It’s not necessarily that xAI is uniquely incompetent(it’s clear they have talented folks)but rather the priorities may be flipped in the GPU arms race.
- 🟡Dario Amodei and Demis Hassabis reportedly proposed a “U.S.-led AI coalition” at G7, with international cooperation on model access, chip exports, and safety risks.
- 🟢THUDM/slime — slime is an LLM post-training framework for RL Scaling.
- 🟢EricLBuehler/mistral.rs — Fast, flexible LLM inference
2026-06-20
- 🔴Hi Reddit, I posted my Build Your Own LLM workshop to Youtube teaching ML, LLM and math intuition [P]
- 🔴GLM 5.2: 98% of max level intelligence with less than half of tokens usage
- 🔴THUDM/slime — slime is an LLM post-training framework for RL Scaling.
- 🟡It’s not necessarily that xAI is uniquely incompetent(it’s clear they have talented folks)but rather the priorities may be flipped in the GPU arms race.
- 🟡AMAZON HOPES TO CHALLENGE NVIDIA MORE DIRECTLY BY SELLING ITS AI CHIPS (2 MINUTE READ)
- 🟡HPE SAYS CONNECTIVITY IS THE OVERLOOKED AI DATA CENTER BOTTLENECK (4 MINUTE READ)
- 🟡CISCO + NVIDIA PUSH SECURE AI NETWORKING FOR THE AI FACTORY (3 MINUTE READ)
- 🟡[GLM 5.2 UD IQ2_M] That's the best pelican svg image I have ever seen
- 🟢GLM 5.2, what speeds are we getting locally?
- 🟢Bought 2x r9700, 5090 is now 7k and 6000 pro is at 13.5k, best option for 64 gb vram under 4k
- 🟢Inference cost at scale with napkin math
- 🟢EricLBuehler/mistral.rs — Fast, flexible LLM inference
- 🟢spiceai/spiceai — A portable accelerated SQL query, search, and LLM-inference engine, written in Rust, for data-grounded AI apps and agents.
2026-06-18
- 🟡Is foundational AI research still something that can be done without access to HPC? [D]
- 🟢alexzhang13/rlm — General plug-and-play inference library for Recursive Language Models (RLMs), supporting various sandboxes.
- 🟢openobserve/openobserve — Open source observability platform for logs, metrics, traces, frontend monitoring, pipelines and LLM observability. A sophisticated, simple and highly performant alternative to Datadog, Splunk, and Elasticsearch with 140x lower storage costs and single binary deployment.
- 🟢How do you analyze the relative "strength" of probes? [R]
2026-06-14
- 🟡skypilot-org/skypilot — Run, manage, and scale AI workloads on any AI infrastructure. Use one system to access & manage all AI compute (Kubernetes, Slurm, 20+ clouds, on-prem).
- 🟢NVIDIA/physicsnemo — Open-source deep-learning framework for building, training, and fine-tuning deep learning models using state-of-the-art Physics-ML methods
2026-06-11
- 🔴Same prompt, same answer, 45x difference in tokens billed. Here's why your LLM bill makes no sense.
- 🔴FareedKhan-dev/train-llm-from-scratch — A straightforward method for training your LLM, from downloading data to generating text.
- 🟡@GoogleDeepMind: DiffusionGemma is our new experimental open model with up to 4x faster output on dedicated GPUs. Instead of predicting word-by-word, it generates entire blocks of text simultaneously. This lets the m
2026-06-06
- 🔴Gemma 4 with quantization-aware training
- 🟢TinyTPU: SystemVerilog systolic array compiled to WASM, running live in browser - RTL golden-verified against numpy [P]
- 🟢microsoft/BitNet — Official inference framework for 1-bit LLMs
- 🟢vllm-project/vllm-omni — A framework for efficient model inference with omni-modality models
2026-06-05
- 🔴Nvidia's been paying shills on LinkedIn
- 🟡cyberpapiii/chipotlai-max — The AI coding agent that runs on stolen Chipotle compute 🌯 Fork of OpenCode with Pepper AI as default model. Community project to add providers from Home Depot, Lowes, Target, Starbucks & more.
- 🟢NVIDIA/NemoClaw — Run agents like Hermes and OpenClaw more securely inside NVIDIA OpenShell with managed inference
2026-05-30
- 🟡I compared all specs of the major GPUs/machines that are being used here, because bandwidth is not everything. Some of ya'll need a reality check.
- 🟢Show HN: Tiny-vLLM – high performance LLM inference engine in C++ and CUDA
- 🟢Got Really lucky and need your advice
- 🟢What I learned building a debugger for PyTorch training loops and how it changed how I think about failure diagnosis [D]
2026-05-28
- 🔴Behold! Probably the most ghetto local AI server:
- 🟡Stress disrupts hippocampal integration of overlapping events, memory inference
- 🟢Cross-Platform Fused MoE Dispatch in Triton: Portable Expert Routing Without CUDA [R]
- 🟢NVIDIA-NeMo/Megatron-Bridge — Training library for Megatron-based models with bidirectional Hugging Face conversion capability
2026-05-23
- 🔴NVIDIA Removes Gaming Revenue Category From Financial Reports
- 🟢facebookresearch/sam3 — The repository provides code for running inference and finetuning with the Meta Segment Anything Model 3 (SAM 3), links for downloading the trained model checkpoints, and example notebooks that show how to use the model.
2026-05-21
- 🔴AMD Ryzen AI Halo PC will cost 3999$ with 128GB memory on board
- 🔴karpathy/autoresearch — AI agents running research on single-GPU nanochat training automatically
- 🟡Anthropic is expanding to Colossus2. Will use GB200
- 🟡vllm-project/vllm — A high-throughput and memory-efficient inference and serving engine for LLMs
- 🟢Looking for real world comparisons between WALL OSS pi0.6 and OpenVLA[D]
- 🟢Training a vision model from scratch on iPod touch 4 images
- 🟢AMD BC-250 and the search for Cheap Compute
- 🟢ai-dynamo/dynamo — A Datacenter Scale Distributed Inference Serving Framework
- 🟢High E2E latency on fine-tuned Gemma 4 26B despite low TTFT [R]
2026-05-20
- 🟡unslothai/unsloth — Unsloth Studio is a web UI for training and running open models like Gemma 4, Qwen3.6, DeepSeek, gpt-oss locally.
- 🟡The next phase of OpenAI’s Education for Countries
- 🟡Google AI Edge Gallery v1.0.13 & v1.0.14 updates: Gemma 4 Multi-Token Prediction, Pixel TPU support, experimental MCP, new skills, now saves chat history
- 🟢Michael-A-Kuykendall/shimmy — ⚡ Python-free Rust inference server — OpenAI-API compatible. GGUF + SafeTensors, hot model swap, auto-discovery, single binary. FREE now, FREE forever.
- 🟢Running DeepSeek-V4 locally with 4x legacy RTX 2080 Ti ($2k budget setup). Custom Turing kernels, W8A8 quantization, and 255 prefill tok/s!
2026-05-19
- 🔴Sub-JEPA: a simple fix to LeCun group's LeWorldModel that consistently improves performance [P]
- 🟡Memory expert suspects RAM price drop in 2027'H2 due to china heavy investments
- 🟡21 GPU's benchmarked running a small TTS model (vram peak: 5GB)
- 🟢Alignment pretraining: AI discourse creates self-fulfilling (mis)alignment
2026-05-18
- 🔴M5 vs DGX Spark vs Strix Halo vs RTX 6000
- 🟡Light-Heart-Labs/DreamServer — Local AI anywhere, for everyone — LLM inference, chat UI, voice, agents, workflows, RAG, and image generation. No cloud, no subscriptions.
- 🟢simular-ai/Agent-S — Agent S: an open agentic framework that uses computers like a human
2026-05-14
- 🟡Trained transformer-based chess models to play like humans (including thinking time) [P]
- 🟢ansible/ansible — Ansible is a radically simple IT automation platform that makes your applications and systems easier to deploy and maintain. Automate everything from code deployment to network configuration to cloud management, in a language that approaches plain English, using SSH, with no agents to install on remote systems.https://docs.ansible.com.
- 🟢huggingface/pytorch-image-models — The largest collection of PyTorch image encoders / backbones. Including train, eval, inference, export scripts, and pretrained weights -- ResNet, ResNeXT, EfficientNet, NFNet, Vision Transformer (ViT), MobileNetV4, MobileNet-V3 & V2, RegNet, DPN, CSPNet, Swin Transformer, MaxViT, CoAtNet, ConvNeXt, and more
2026-05-13
- 🔴I got a real transformer language model running locally on a stock Game Boy Color!
- 🟡How do you create memorable poster for top tier conferences ( ICML/ICLR/NEURips ect…) [D]
- 🟡My First Official AI Research Paper Accepted on SSRN
- 🟢huggingface/transformers — 🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
2026-05-12
- 🔴"This is the first documented instance of AI self-replication via hacking." ... "We ran an experiment with a single prompt: hack a machine and copy yourself. The AI broke in and copied itself onto a new computer. The copy then did this again, and kept on copying, forming a chain."
- 🔴Training an LLM in Swift, Part 1: Taking matrix mult from Gflop/s to Tflop/s
- 🟢lakehq/sail — Drop-in Apache Spark replacement written in Rust, unifying batch processing, stream processing, and compute-intensive AI workloads.
2026-05-08
- 🔴AMD Intros Instinct MI350P Accelerator: CDNA 4 Comes to PCIe Cards
- 🔴Taiwanese company Skymizer announces HTX301 - PCIE inference card with 384GB of Memory at ~240 Watts
- 🟡Disillusionment with mechanistic interpretability research [D]
- 🟡DeepSeek 4 Flash local inference engine for Metal
- 🟡Crosstalk-Solutions/project-nomad — Project N.O.M.A.D, is a self-contained, offline survival computer packed with critical tools, knowledge, and AI to keep you informed and empowered—anytime, anywhere.
- 🟡Blaizzy/mlx-vlm — MLX-VLM is a package for inference and fine-tuning of Vision Language Models (VLMs) on your Mac using MLX.
- 🟢ZAYA1-74B-Preview: Scaling Pretraining on AMD
- 🟢THE UNDERPRIVILEGED AI FOUNDATION Because every little model deserves a chance
- 🟢A new generation of AI models and one of the most powerful research papers out there.
- 🟢What should a PyTorch training end-of-run performance summary show? [D]
2026-05-07
- 🔴Bad news: Apple drops high-memory Mac Studio configs
- 🔴ZAYA1-8B: Frontier intelligence density, trained on AMD
- 🟡InsForge/InsForge — InsForge is a Postgres-based backend with auth, storage, compute, hosting, and AI gateway. Built for coding agents.
- 🟢Dataset of 150k+ stool images and not sure how to fully use it [D]
- 🟢pytorch/pytorch — Tensors and Dynamic neural networks in Python with strong GPU acceleration
- 🟢Making LLM Training Faster with Unsloth and NVIDIA
2026-05-06
- 🔴Struggling to reproduce paper results before improving them — stuck below reported accuracy [R]
- 🔴Accelerating Gemma 4: faster inference with multi-token prediction drafters
- 🔴cheahjs/free-llm-api-resources — A list of free LLM inference resources accessible via API.
- 🟢TritonSigmoid: A fast, padding-aware sigmoid attention kernel for GPUs [R]
- 🟢Competition - League of Robot Runners 2026: Multi-robot coordination under uncertainty [N]