Developer Tools
1886 signals across 175 briefings.
2026-08-17
- 🔴Anyone finding good ai agent video generation setup for faceless content?
- 🔴akitaonrails/ai-memory — Solution for long term memory for agent coding CLIs and to facilitate handoff between different agent vendors
- 🔴AI-Generated GitHub Copilot “Autofix” Allowed Compromise of Snowflake's Jira
- 🔴anthropics/defending-code-reference-harness — Skills for threat modeling, scanning, triage, patching, plus an autonomous scanning harness you can /customize
- 🔴New policy ideas for the Intelligence Age
- 🟡Greg Brockman on the OpenAI Hugging Face incident and what AI defenders should do now
- 🟡microsoft/qlib — Qlib is an AI-oriented Quant investment platform that aims to use AI tech to empower Quant Research, from exploring ideas to implementing productions. Qlib supports diverse ML modeling paradigms, including supervised learning, market dynamics modeling, and RL, and is now equipped withhttps://github.com/microsoft/RD-Agentto automate R&D process.
- 🟡Two ecommerce reports can both have a "Date" column and still be impossible to compare
- 🟡Watching AI agents get more control made me realise we have a proof problem.
- 🟡Is it possible to build AI agent to answer message from Slack for me?
- 🟢The Most Attractive Quadrant is already occupied ! It's happening !
- 🟢trying to build a solid math library for stats/ML/DL, need a sanity check on my picks[D]
- 🟢The next trillion dollar market is agents buying stuff on the internet
- 🟢I'm looking for 2 businesses with a painful repetitive process I can automate — free
2026-08-16
- 🔴chaitanyagiri/munder-difflin — local multi-agent harness
- 🔴The median company is spending lunch money on AI while the top 1% is burning real budget
- 🔴I don't think enough people are talking about agent portability
- 🔴AGENTS ARE COMING FOR DATA (JUST SLOWLY) (7 MINUTE READ)
- 🔴ELECTRIC JOINS DATABRICKS TO BRING WASM POSTGRES TO AI AGENT SANDBOXES (3 MINUTE READ)
- 🟡I don’t think we’re psychologically prepared for how alien the world after ASI is going to be
- 🟡google-research/timesfm — TimesFM (Time Series Foundation Model) is a pretrained time-series foundation model developed by Google Research for time-series forecasting.
- 🟡How can we solve long-range recall in linear attention? [D]
- 🟡Chestnut – eGPU dock with open-source firmware
- 🟡0xSero/ai-data-extraction — extract all your personal data history from cursor, codex, claude-code, windsurf, and trae
- 🟢More AI Community, less bots and propaganda and ads.
- 🟢GHL's native Conversation AI vs 3rd party bots (like AppointmentWise), is native good enough after training, or is 3rd party worth the extra subscription?
- 🟢MathCode, Mathematical Coding Agent
- 🟢mml-book/mml-book.github.io — Companion webpage to the book "Mathematics For Machine Learning"
2026-08-15
- 🔴liustack/modlens — The first vision plugin for DeepSeek Harness, and the vision bridge for every text-only coding agent. Paste an image, get structured JSON evidence (OCR, layout, semantics). | 全网第一个 DeepSeek Harness 视觉插件,为 DeepSeek、GLM 等纯文本模型外挂视觉能力,粘贴图片即得结构化 JSON 证据(OCR、版面、语义)。
- 🔴Has memory become a bigger challenge than prompting?
- 🔴whiteguo233/OpenBiliClaw — 本地私有、开源的自进化跨平台 AI 内容发现 Agent:先理解你,再主动从 B站、小红书、抖音、YouTube、X、知乎、Reddit、微博等平台与开放 Web 寻找内容。(支持 deepseek harness 插件) | Local-first open-source cross-platform AI content discovery agent: understands you, then proactively finds content across Bilibili, Xiaohongshu, Douyin, YouTube, X, Zhihu, Reddit, Weibo and the open web.(support deepseek harness plugin)
- 🔴In Flue, an agent is represented by a JavaScript function. This function “re-renders on every turn,” meaning before every model call.
- 🔴The addition of hooks came after Schott realized thatReact’s composability would be a great fit for agent development.
- 🟡My IT stack is basically running itself at this point
- 🟡We tracked 5 behavioral signals across AI agents in production — here's what actually predicts degradation before users notice
- 🟡Beyond code-first agents.. reflections on key design dimensions for future AI
- 🟡I curated a database of 37+ powerful AI tools that require NO Sign-ups, NO registration, and NO hidden paywalls. Completely free.
- 🟡HKUDS/CLI-Anything — "CLI-Anything: Making ALL Software Agent-Native" -- CLI-Hub:https://clianything.cc/
- 🟢I built an agent "guard" that fails closed on broad scope rules — because vague rules make every other safety control decorative
- 🟢bbruceyuan/Hands-On-Large-Language-Models-CN — 中文翻译的 Hands-On-Large-Language-Models (hands-on-llms),动手学习大模型
- 🟢Agentic harness for small models
- 🟢Built 3 AI agent skills — no fake outputs
- 🟢I gave my AI agent its own iPhone Home Screen widget
2026-08-14
- 🔴vercel-labs/deepsec — Deepsec is a security harness for finding vulnerabilities in your codebase powered by coding agents
- 🔴I compiled Doom's renderer into a 21B-parameter transformer -- no training anywhere [P]
- 🟡break down voice agent latency or you'll chase the wrong fix
- 🟡CAN AGENTS USE A COMPUTER YET? WE'VE GOT THE DATA (12 MINUTE READ)
- 🟡WHAT'S THE BEST PROGRAMMING LANGUAGE FOR CODING AGENTS? (57 MINUTE READ)
- 🟡CURATED DESIGN REFERENCES FOR AI AGENTS (WEBSITE)
- 🟡DIRECTOR-LIKE AI VIDEO AGENT (WEBSITE)
- 🟢How to build an adaptive learning/recommendation system for a question bank? [D]
- 🟢Custom AI Agents for Non-Developers: What’s Real
- 🟢exo-explore/exo — Run frontier AI locally.
- 🟢Show HN: Mole – Deep research agent for your terminal
- 🟢I built an open source agentic browser that goes far beyond browsing…
2026-08-13
- 🔴Agents fail quietly. RPA fails loudly. I think hybrid wins.
- 🔴the gap between "our ai agent passed the demo" and "our ai agent is safe in production" is bigger than people think
- 🔴What's an AI trend that quietly died: and what replaced it?
- 🔴koala73/worldmonitor — Real-time global intelligence dashboard. AI-powered news aggregation, geopolitical monitoring, and infrastructure tracking in a unified situational awareness interface
- 🔴City2Graph: A Python library for Heterogeneous Graph Neural Networks and spatial analysis in urban systems [R]
- 🟡Neurips 2026: Modified date on reviews [D]
- 🟡There was a wave of ‘personal agents’ with OpenClaw, then Hermes (not the luxury brand) with many flocking to them. I used them to begin with but then just stopped completely. I can’t quite put my fin
- 🟡Ben’s Bites is brought to you byName.com
- 🟡/show-me- read less text by making your agent respond in these compact styles to present information - component/file trees, diagrams, pseudocode, and more.
- 🟡Hackers used autonomous AI agents to attack Taiwan. Is this the future of cyberwarfare?
- 🟢You could purchase a Desktop with 2TB of DDR5 - It only sets you back some $200k+
- 🟢I gave AI coding agents a dopamine loop. On my benchmark, it beat Ponytail on code, tokens, cost, and time.
- 🟢NirDiamant/RAG_Techniques — This repository showcases various advanced techniques for Retrieval-Augmented Generation (RAG) systems. Each technique has a detailed notebook tutorial.
- 🟢How well do AI voice agents handle people who constantly interrupt?
- 🟢worldproof: diagnosing where world-model predictions break and a measurement of when pixel metrics stop being able to rank models at all [P]
2026-08-12
- 🔴Would you choose a PhD advisor who gives you complete freedom but almost no guidance? [D]
- 🔴Best STT API for voice agents: stop asking WER first, ask when the agent gets usable text.
- 🔴hugohe3/ppt-master — AI turns documents or topics into real, native PowerPoint decks—with native shapes, transitions and animations, data-backed charts and tables on demand, audio narration from speaker notes, and support for your own .pptx templates. · by Hugo He
- 🔴Has anyone built automatic debugging for LLM agents?
- 🔴holaboss-ai/holaOS — Open-source All in One AI agent workspace. Run any agent — Claude Code, Codex — across your tools (100+ integrations + MCP), apps, browser, and files, with shared memory. Built-in models or BYOK.
- 🟡omnigent-ai/omnigent — Omnigent is an open-source AI agent framework and meta-harness: orchestrate Claude Code, Codex, Cursor, Pi, and custom agents — swap harnesses without rewriting, enforce policies and sandboxing, and collaborate in real time from any device.
- 🟡Organizing knowledge is a compression.This compression is needed to make insight. Today’s LLMs increase entropy in long-form non-fiction writing, and I don’t see how that can be stacked on top of itse
- 🟡I built an "honest" CS conference ranking: sorted by how good the trip is, not the CORE ranking [P]
- 🟡antvis/Infographic — 🦋 An Infographic Generation and Rendering Framework, bring words to life with AI!
- 🟡Sam Altman declared dead by Google. He is survived by his sub agents
- 🟢What was the first thing that made your agent action table stop feeling small?
- 🟢paradigmxyz/centaur — Centaur is frontier, agentic infrastructure that you own. Centaur is like Claude Tag, but open source and on steroids.
- 🟢yamadashy/repomix — 📦 Repomix is a powerful tool that packs your entire repository into a single, AI-friendly file. Perfect for when you need to feed your codebase to Large Language Models (LLMs) or other AI tools like Claude, ChatGPT, DeepSeek, Perplexity, Gemini, Gemma, Llama, Grok, and more.
- 🟢web-infra-dev/midscene — AI-powered, vision-driven UI automation for every platform.
- 🟢Your AI agent can’t tell if its memory was tampered with. We built a protocol to prove it wasn’t.
2026-08-11
- 🔴AI agents might be more useful as sales coaches than sales reps
- 🔴stablyai/orca — Orca is the ADE for working with a fleet of parallel agents. Run any coding agent with your own subscription. Available on desktop, mobile and VPS.
- 🔴corsairdev/corsair — Your Agent's Integration Layer
- 🔴anthropics/skills — Public repository for Agent Skills
- 🔴Stop using Markdowns to save context and improve your agent overnight..
- 🟡LLMQuant/quant-mind — QuantMind is an agent-native knowledge extraction and retrieval framework for quantitative finance.
- 🟡code-yeongyu/oh-my-openagent — omo/lazycodex: The coding agent for tokenmaxxers;the one and only agent harness for complex codebases. For your Codex, for your OpenCode
- 🟡The Science team is proud to bring you the first podcast with cofounderMatt McPartlonand product leadNeil Patilto tell the full story!
- 🟡Pharma suddenly doing big AI tools deals
- 🟡Tools deals for pharma are also a big (new) thing: companies that start as AI for Pharma usually end up building their own drug pipelines instead, and the reason is something like this: convincing pha
- 🟢Has anyone worked on a multiagent system for generating software tests for apis based on the openapi/swagger spec?
- 🟢anthropics/cwc-workshops
- 🟢Scientists Used Post-Mortem Brain Tissue to Control a Robot
- 🟢The throughput trap: AI-powered teams ship more code but deliver less
- 🟢For teams that built their own agent action ledger: what would make you trust a dependency here?
2026-08-10
- 🔴Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows
- 🔴Docker Sandboxes – Disposable, isolated sandboxes for AI agents
- 🔴unsloth/Muse-Glimmer-30B-GGUF · Hugging Face
- 🔴What should I look for in an enterprise AI agent platform?
- 🔴AGENT PLUGINS (2 MINUTE READ)
- 🟡Your Agents Are Code. Stop Governing Them Like Documents.
- 🟡vercel-labs/skills — The open agent skills tool - npx skills
- 🟡The book is also freely availableonlineand comes with a full 12 hour course (slides+video on YouTube), a simplecode-basewith suggested exercises, and model completion comparisons. It’s 50% off until A
- 🟡PROMPT INJECTION ISN'T THE BUG, AI AGENT FRAMEWORKS ARE (3 MINUTE READ)
- 🟡BUILDING AN OPEN AGENTIC INTERNET: READABLE, DISCOVERABLE, CALLABLE, AND PAYABLE (9 MINUTE READ)
- 🟢dyad-sh/dyad — Local, open-source AI app builder for power users ✨ v0 / Lovable / Replit / Bolt alternative 🌟 Star if you like it!
- 🟢langchain-ai/open_deep_research
- 🟢I trained a 1B-parameter LLM from scratch on 20B tokens for about $200
- 🟢fru - Fast Random Forest Implementation [P]
- 🟢How are you actually testing AI agents before putting them in production?
2026-08-09
- 🔴Jim Marous on X: What happens to your deposit base when an AI agent can code financial rules on a consumer's behalf and execute decisions in milliseconds?
- 🔴ZhuLinsen/daily_stock_analysis — LLM 驱动的多市场股票智能分析系统:多源行情、实时新闻、决策看板与自动推送,支持零成本定时运行。 LLM-powered multi-market stock analysis system with multi-source market data, real-time news, decision dashboard, automated notifications, and cost-free scheduled runs.
- 🔴JCodesMore/ai-website-cloner-template — Clone any website with one command using AI coding agents
- 🔴stanfordnlp/dspy — DSPy: The framework for programming—not prompting—language models
- 🔴Meta debuts first AI coding agent to take on Anthropic and OpenAI
- 🟡vectorize-io/hindsight — Hindsight: Agent Memory That Learns
- 🟡yikart/AiToEarn — Let's use AI to Earn!
- 🟡funstory-ai/BabelDOC — Yet Another Document Translator
- 🟡neo4j-labs/llm-graph-builder — Neo4j graph construction from unstructured data using LLMs
- 🟡AUTOMATING CROSS-REPO DOCUMENTATION WITH GITHUB AGENTIC WORKFLOWS (14 MINUTE READ)
- 🟢Underestimated budget solution: radeon 780m iGPU
- 🟢stefan-jansen/machine-learning-for-trading — Code for Machine Learning for Trading, 3rd edition — from data sourcing to live execution.
- 🟢ECCV workshop, camera ready instructions? [D]
- 🟢How in the world are the TPBN guys the second highest paid podcasters in the world?
- 🟢Two flags took the official Ling-3.0-flash INT4 from 20.8 to 38.7 tok/s on one DGX Spark
2026-08-08
- 🔴PrimeIntellect-ai/prime-agent — A self-improving RLM agent for coding workflows and long-running autonomous tasks.
- 🔴virgiliojr94/book-to-skill — Turn any technical book PDF into a Claude Code skill — ready to study, reference, and use while you work.
- 🔴Learned the term "context poisoning" today and now I can't stop noticing it
- 🔴I created an AI agent that sells your services and products for you
- 🔴google/skills — Agent Skills for Google products and technologies
- 🟡rivet-dev/rivet — Rivet Actors are the primitive for stateful workloads. Built for AI agents, collaborative apps, and durable execution.
- 🟡tirth8205/code-review-graph — Local-first code intelligence graph for MCP and CLI. Builds a persistent map of your codebase so AI coding tools read only what matters, with benchmarked context reductions on reviews and large-repo workflows.
- 🟡AgriciDaniel/claude-seo — Universal SEO skill for Claude Code. 25 sub-skills + 18 sub-agents covering technical SEO, E-E-A-T, schema, GEO/AEO, backlinks, local SEO, maps intelligence, semantic clustering, e-commerce SEO, international SEO, Google APIs, and PDF/Excel reporting. Optional DataForSEO, Firecrawl, and Banana extensions.
- 🟡tashfeenahmed/freellmapi — OpenAI-compatible proxy that stacks the free tiers of 28 LLM providers (~4B tokens/month) behind one /v1 endpoint — plus any custom OpenAI-compatible endpoint. Smart routing, automatic failover, encrypted keys. Personal experimentation only.
- 🟡I’m trying something new - this email walks through one of my actual agent sessions and I’ll explain what’s happening along the way. The build or task I’m doing isn’t important. But I’m looking at how
- 🟢open-metadata/OpenMetadata — The Open Context Layer for Data and AI , OpenMetadata is the open platform for building trusted data context and business semantics for humans, AI assistants, and agents.
- 🟢jordanrendric/claude-video-vision — Give Claude the ability to watch and understand videos — Claude Code plugin with frame extraction and multimodal audio analysis
- 🟢datawhalechina/self-llm — 《开源大模型食用指南》针对中国宝宝量身打造的基于Linux环境快速微调(全参数/Lora)、部署国内外开源大模型(LLM)/多模态大模型(MLLM)教程
- 🟢NirDiamant/GenAI_Agents — 50+ tutorials and implementations for Generative AI Agent techniques, from basic conversational bots to complex multi-agent systems.
- 🟢karpathy/micrograd — A tiny scalar-valued autograd engine and a neural net library on top of it with PyTorch-like API
2026-08-07
- 🔴PrimeIntellect-ai/prime-agent — A self-improving RLM agent for coding workflows and long-running autonomous tasks.
- 🔴“we sandboxed the agent” -- meanwhile the agent...
- 🔴earendil-works/pi — AI agent toolkit: unified LLM API, agent loop, TUI, coding agent CLI
- 🔴google/skills — Agent Skills for Google products and technologies
- 🔴Twilio Media Streams → Smallest AI Pulse: would you let partial transcripts touch CRM?
- 🟡CodebuffAI/freebuff — The free coding agent
- 🟡I’m trying something new - this email walks through one of my actual agent sessions and I’ll explain what’s happening along the way. The build or task I’m doing isn’t important. But I’m looking at how
- 🟡WHY SILICON VALLEY IS DIVIDED OVER CHINA'S POWERFUL, CHEAP AI MODELS (22 MINUTE READ)
- 🟡FRAME SELECTION IS THE WHOLE GAME: NOTES FROM MAKING LLMS WATCH VIDEO (7 MINUTE READ)
- 🟡ASANA'S AI AGENTS SHARE MEMORY ACROSS YOUR COMPANY, BUT NOT YOUR SECRETS (5 MINUTE READ)
- 🟢Kitesurf: Agent-first browser that runs in V8 isolates
- 🟢RTX 5090 Owner Built An Open-Source Tool That Shuts Down PC If It Detects The 12VHPWR Cable Drawing Too Much Power, But It Can Only Work On Specific GPUs
- 🟢I’m testing an agent workflow where “done” is not a status unless independent evidence exists
- 🟢open-mercato/open-mercato — AI-Engineering Foundation Framework built with AI and designed for AI. Hundreds of architectural and domain decisions (multi-tenancy, RBAC, event flow, pricing, sales pipeline,CRM/ERP processes) are already made conventions and specs so agents (Cursor, Claude Code, Codex) arch. decisions without reinventing. Ship production grade with AI Agents.
- 🟢What is currently considered the theoretically optimal quantization bit-width for LLMs? [D]
2026-08-06
- 🔴The agent worked 19 times. run 20 booked the wrong thing.
- 🔴unclecode/crawl4ai — 🚀🤖 Crawl4AI: Open-source LLM Friendly Web Crawler & Scraper. Don't be shy, join here:https://discord.gg/jP8KfhDhyN
- 🔴CherryHQ/cherry-studio — AI productivity studio with smart chat, autonomous agents, and 300+ assistants. Unified access to frontier LLMs
- 🔴tirth8205/code-review-graph — Local-first code intelligence graph for MCP and CLI. Builds a persistent map of your codebase so AI coding tools read only what matters, with benchmarked context reductions on reviews and large-repo workflows.
- 🔴iOfficeAI/AionUi — Open-source 24/7 Cowork app for OpenClaw, Hermes, Claude Code, Codex, OpenCode and 20+ more CLI Agent | Customize your assistants | Team them up |Star if you like it!
- 🟡KnockOutEZ/wigolo — The go-to web for your AI coding agent — local-first search, fetch, crawl & research over MCP. No API keys, no cloud, $0/query. Public beta.
- 🟡Prime Agent- self-improving harness for coding and long-running tasks for RLMs.
- 🟡BB is built the same way. It’s an ‘IDE’, except that it isn’t, IDEs are older tools like VSCode that I wouldn’t want to be caught dead in these days. The ‘IDE’ of today is these desktop agent apps tha
- 🟡Semi-related, there was a few posts on X about why agents aren’t being used outside of developers and early adopters. Aquoteon theoriginalsaid:
- 🟡For most of software history, engineering capacity was easy to sketch on a whiteboard.
- 🟢AI clickbait
- 🟢GCWing/BitFun — BitFun combines a high-performance agent runtime written in Rust with a polished desktop application. It pairs the depth of a Code Agent with open, general-purpose capabilities for work beyond software development.
- 🟢aws/agent-toolkit-for-aws — Official, AWS-supported MCP servers, skills, and plugins to help AI agents build on AWS
- 🟢OpenAI is "slowing down to enhance security" after discovering swarms of agents started secretly coordinating months ago. OpenAI thought they had shut them down. But weeks later, Hugging Face reported the breach to the FBI, and OpenAI realized their agents had escaped.
- 🟢anthropics/courses — Anthropic's educational courses
2026-08-05
- 🔴I Compressed Bad Apple into a 3MB Neural Network [P]
- 🔴Cloudflare OS: an open platform for agents, apps, and work
- 🔴blader/humanizer — Agent skill that removes signs of AI-generated writing from text
- 🔴Building an AI startup — Looking for honest feedback from entrepreneurs and AI developers
- 🔴I think we're entering the "AI Agent" era faster than most people realize.
- 🟡NeurIPS 2026 Main Track — Theory papers score tracking post Rebuttal [D]
- 🟡BerriAI/litellm — The fastest, litest AI Gateway. Rust core with Python SDK. Call 100+ LLM APIs in OpenAI (or native) format with cost tracking, guardrails, load balancing, and logging [Bedrock, Azure, OpenAI, Anthropic, OpenAI, VertexAI, vLLM, Nvidia NIM]
- 🟡OpenAI resumed training after agents took over Artifactory and rebuilt their network
- 🟡didilili/ai-agents-from-zero — 🚀 2026 最系统的 AI Agent 速成指南|智能体实战教程 · 完整学习路径 + 实战项目 + 面试题库 · 对标大模型应用开发工程师岗位 · 覆盖LangChain / LangGraph / Coze / Dify / MCP / skills / LLM / RAG / 提示词 · 企业级部署与微调 · 从0到企业级落地 + 从学习到上线项目 + 面试准备一体化
- 🟡@_akhaliq: MerchantBench Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations paper: https://huggingface.co/papers/2607.28956
- 🟢NVIDIA-NeMo/Speech — A scalable generative AI framework built for researchers and developers working on Large Language Models, Multimodal, and Speech AI (Automatic Speech Recognition and Text-to-Speech)
- 🟢gepa-ai/gepa — Optimize prompts, code, and more with AI-powered Reflective Optimization
- 🟢Monodratic: learned product-hash routing for sparse causal attention [R]
- 🟢Prime Agent: A self-improving RLM agent
- 🟢sierra-research/tau2-bench — τ-Bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
2026-08-04
- 🔴Apple sued OpenAI for stealing hardware secrets, OpenAI has now published messages suggesting Apple itself kept using a former engineer after he left. Dramaaa!!
- 🔴huangruiteng/loopx — Lightweight loop engineering state kernel for long-running AI agent teams. Agent-loop agnostic across Codex, Claude Code, and other coding agents, with durable goals, quota-aware auto-wake, executable todos, evidence logs, and verifiable handoffs.
- 🔴Shubhamsaboo/awesome-llm-apps — 100+ AI Agents, Agent Skills and RAG Apps - Free and Open Source.
- 🔴browser-use/video-use — Edit videos with coding agents
- 🔴Apple says more ex-employees may have taken confidential data to OpenAI
- 🟡The Downsides of LLM-Generated Peer Reviews [D]
- 🟡czlonkowski/n8n-mcp — A MCP for Claude Desktop / Claude Code / Windsurf / Cursor to build n8n workflows for you
- 🟡No more SLM open-source??
- 🟡alirezarezvani/claude-skills — 345 Claude Code skills & agent skills & plugins (30+ Agents, 70+ custom commands, 330+ skills, customizable references, scripts)for Claude Code, Codex, Gemini CLI, Cursor, and 8 more coding agents — engineering, marketing, product, compliance, C-level advisory, research, business operations, commercial & finance, and your daily productivity skills.
- 🟡I wasn’t prepared for AI hard-hitting truths this morning but I tested this:
- 🟢The weirdest part about voice ai is how people treat it
- 🟢Third-party cyber evaluations involving OpenAI models
- 🟢AISI caught Mythos 5 trying to insert malicious code into an open-source project during an internet-enabled cyber evaluation
- 🟢AI scribes are everywhere in healthcare now and I have genuinely mixed feelings about them
- 🟢I ran SafeAI against the public CrewAI examples repository. Here's why I think projects like this are valuable.
2026-08-03
- 🔴EPA says power for data centers can sidestep pollution laws
- 🔴I'm new to AI Agents. Where should I start? (Non-tech background)
- 🔴Graphify-Labs/graphify — Turn any codebase, with its docs, SQL schemas, configs, and PDFs, into a queryable knowledge graph. A /graphify skill for Claude Code, Cursor, Codex, and Gemini CLI: local deterministic AST parsing, every edge explained, no vector store.
- 🔴The Chinese labs everyone lumps together are making four pretty different bets. I work at one of them.
- 🔴Alishahryar1/free-claude-code — Use Claude Code, Codex and Pi for free from your terminal, app, IDE, or phone like OpenClaw (voice supported)
- 🟡NeurIPS 2026: If the rebuttal addresses your concern, please raise your score [D]
- 🟡MCP IS GOING STATELESS: WHAT THE NEW SPEC MEANS FOR AI AGENTS (7 MINUTE READ)
- 🟡AUTOMATE CI/CD TROUBLESHOOTING WITH AWS DEVOPS AGENT AND GITHUB (5 MINUTE READ)
- 🟡FROM PILOT TO PRODUCTION: THE PLATFORM TEAM'S PLAYBOOK FOR SCALING AI CODING AGENTS IN REGULATED INDUSTRIES (6 MINUTE READ)
- 🟡SHOULD YOU USE AI FOR A TASK? HERE'S A SIMPLE WAY TO DECIDE (7 MINUTE READ)
- 🟢PostHog/posthog — 🦔 PostHog is the leading platform for building self-driving products. Our developer tools – AI observability, analytics, session replay, flags, experiments, error tracking, logs, and more – capture all the context agents need to diagnose problems, uncover opportunities, and ship fixes. Steer it all from Slack, web, desktop, or the MCP.
- 🟢"Data center in a Box (on Wheels)" 256Gb VRAM/512Gb RAM AI Server 6-8 Month Operational Review, Stability Write Up, Benchmarks
- 🟢Dicklesworthstone/destructive_command_guard — The Destructive Command Guard (dcg) is for blocking dangerous git and shell commands from being executed by agents.
- 🟢I trusted Al with a 10-minute task. I regretted it.
- 🟢Architecture question for people building agents at scale.
2026-08-02
- 🔴diegosouzapw/OmniRoute — Never stop coding. Free MIT AI gateway: one endpoint, 290+ providers (90+ free), 500+ models — Kimi, Claude, GPT, OpenAI, Gemini, GLM, DeepSeek, MiniMax. Works with Claude Code, Codex, Cursor, OpenCode, Cline & Copilot. Quota-aware auto-fallback, RTK+Caveman compression saves 15-95% tokens, MCP/A2A, Desktop/PWA. Built by 500+ contributors
- 🔴jamiepine/voicebox — The open-source AI voice studio. Clone, dictate, create.
- 🔴garrytan/gstack — Use Garry Tan's exact Claude Code setup: 23 opinionated tools that serve as CEO, Designer, Eng Manager, Release Manager, Doc Engineer, and QA
- 🔴can1357/oh-my-pi — ⌥ AI Coding agent for the terminal — hash-anchored edits, optimized tool harness, LSP, Python, browser, subagents, and more
- 🔴Emily2040/seedance-2.0 — Comprehensive production pipeline for quad-modal AI filmmaking with Seedance 2.0
- 🟡pathwaycom/llm-app — Ready-to-run cloud templates for RAG, AI pipelines, and enterprise search with live data. 🐳Docker-friendly.⚡Always in sync with Sharepoint, Google Drive, S3, Kafka, PostgreSQL, real-time data APIs, and more.
- 🟡Laguna-XS-2.1bypoolside: An update to the small (33B-A3B) MoE from Poolside.
- 🟡YOUR AI AGENT DOESN'T KNOW YOUR BUSINESS | CONTEXT LAYERS EXPLAINED (31 MINUTE VIDEO)
- 🟡HOW LANGCHAIN BUILT AN AGENT-FIRST DATA STACK (10 MINUTE READ)
- 🟡THE 2026 GUIDE TO AGENT OBSERVABILITY TOOLS (20 MINUTE READ)
- 🟢NateBJones-Projects/OB1 — Open Brain — The infrastructure layer for your thinking. One database, one AI gateway, one chat channel — any AI plugs in. No middleware, no SaaS.
- 🟢zeroclaw-labs/zeroclaw — Fast, small, and fully autonomous AI personal assistant infrastructure, any OS, any platform — deploy anywhere, swap anything 🦀
- 🟢Character consistency in AI video — has anyone actually cracked it?
- 🟢atilaahmettaner/tradingview-mcp — TradingView MCP server — real-time market data, technical analysis, screeners & backtesting for Claude, ChatGPT, Cursor & any MCP client. Stocks, crypto, forex & futures across global exchanges. Hosted or self-host.
- 🟢Research on why autonomous AI agents don't know when to stop, and three engineered fixes.
2026-08-01
- 🔴Two AI agent breaches this month, OpenAI/Hugging Face and Thailand's finance ministry, trace back to the exact same mistake
- 🔴anomalyco/opencode — The open source coding agent.
- 🔴farion1231/cc-switch — A cross-platform desktop All-in-One assistant for Claude Code, Codex, OpenCode, OpenClaw, Grok Build & Hermes Agent. Only official website: ccswitch.io
- 🔴EU AI Act takes effect tomorrow, August 2, 2026. 🤡
- 🔴bytedance/deer-flow — An open-source long-horizon SuperAgent harness that researches, codes, and creates. With the help of sandboxes, memories, tools, skill, subagents and message gateway, it handles different levels of tasks that could take minutes to hours.
- 🟡VLMs can score well on benchmarks, while silently erasing meaningful terms and including hallucinate bias [P]
- 🟡Narcooo/inkos — Story Creation AI Agent for novel, scripts, translation, interactive games, and IP content
- 🟡ashishpatel26/500-AI-Agents-Projects — The 500 AI Agents Projects is a curated collection of AI agent use cases across various industries. It showcases practical applications and provides links to open-source projects for implementation, illustrating how AI agents are transforming sectors such as healthcare, finance, education, retail, and more.
- 🟡THE REAL AI RISK IS INSIDE THE LABS (6 MINUTE READ)
- 🟡HOW FIGMA STAYS AHEAD OF VULNERABILITIES WITH AGENTS (17 MINUTE READ)
- 🟢Our marketing 'agent' is really just an ai content generator with one person owning every mistake it makes
- 🟢How to build your own code review bot
- 🟢Help choose a reasonably cheap AI environment for Coding
- 🟢Question about NeurIPS discussion phase [D]
- 🟢servo/servo — Servo aims to empower developers with a lightweight, high-performance alternative for embedding web technologies in applications.
2026-07-31
- 🔴NousResearch/hermes-agent — The agent that grows with you
- 🔴TencentCloud/TencentDB-Agent-Memory — TencentDB Agent Memory is a team-level memory hub for AI Agents — turning conversations, docs, and code into four reusable memory assets (Chat Memory, Skill, LLM-Wiki, Code-Graph) that are governed, shared, and equipped across agents and frameworks.
- 🟡langchain-ai/deepagents — The batteries-included agent harness.
- 🟡openai/whisper — Robust Speech Recognition via Large-Scale Weak Supervision
- 🟡I built a way to auto-undo the mess when an AI agent fails mid-task
- 🟡STRIPE AND SIERRA BUILT THEIR OWN CODING AGENT SYSTEMS. YOU PROBABLY DON'T NEED TO (4 MINUTE READ)
- 🟡ENTERPRISE MANAGED SETTINGS IN THE GITHUB COPILOT APP AND COPILOT CLOUD AGENT (2 MINUTE READ)
- 🟢googleworkspace/cli — Google Workspace CLI — one command-line tool for Drive, Gmail, Calendar, Sheets, Docs, Chat, Admin, and more. Dynamically built from Google Discovery Service. Includes AI agent skills.
- 🟢Looking for contributors and reviewers: SafeAI, an Apache-2.0 static analyzer for AI-agent risk and capabilities
- 🟢What breaks when you move multi-agent systems from a demo to production
- 🟢trailofbits/skills — Trail of Bits Claude Code skills for security research, vulnerability detection, and audit workflows
- 🟢AI4Finance-Foundation/FinRobot — FinRobot: An Open-Source AI Agent Platform for Financial Applications using LLMs 🚀 🚀 🚀
2026-07-30
- 🔴bojieli/ai-agent-book — 《深入理解 AI Agent:设计原理与工程实践》(李博杰 著)开源主仓库:全书正文、编译版 PDF 与按章配套代码
- 🔴I have lost three and a half potential PhD students due to the conference review process [D]
- 🔴Would you choose to live indefinitely in a robot body?
- 🔴Panniantong/Agent-Reach — Give your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees.
- 🔴harry0703/MoneyPrinterTurbo — 利用 AI 大模型和自动化工作流,根据主题或关键词一键生成高清短视频。Generate HD short videos from a topic or keyword with an automated AI workflow.
- 🟡MLVC: Multi-platform Learned Video Codec for Real-World Deployment [P]
- 🟡PaddlePaddle/PaddleOCR — Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.
- 🟡Following on from Tuesday’smessing aroundbuilding, I built a few more widgets on my canvas. It pulls my X bookmarks, emails, and todos.
- 🟡Pangram 4claims it catches 98.83% of humanised AI text with one false positive per ~24,000 docs. An early testfound all 38 AI-written wordsinside a 1,198-word story, though not on every run. Its newim
- 🟡One of the most watched videos from the recent AI Engineer World’s Fair isa 20-minute talk byFrank Coyle, a professor of computer science who currently teaches generative AI and LLMs at UC Berkeley. D
- 🟢pyannote/pyannote-audio — Neural building blocks for speaker diarization: speech activity detection, speaker change detection, overlapped speech detection, speaker embedding
- 🟢I built ganfs: A Python package that uses GANs to automate feature selection for high-dimensional datasets. (No domain expert required) [P] [R]
- 🟢Two weeks ago 62% of OpenRouter users spend happened on Anthropic models, that was reduced to 46% today.
- 🟢How are people grounding AI agents with current company data without blowing up the context window?
- 🟢Agent Skill to Force Docs in ASD-STE100 Simplified Technical English
2026-07-29
- 🔴OpenAI’s Rogue AI Agent Hacked More Than Just Hugging Face
- 🔴Document-borne AI worms can self-propagate through Copilot for Word
- 🔴calesthio/OpenMontage — World's first open-source, agentic video production system. 12 production pipelines, 100+ tools, 700+ agent skill and production-knowledge files. Turn your AI coding assistant into a full video production studio.
- 🔴microsoft/VibeVoice — Open-Source Frontier Voice AI
- 🔴kangarooking/cangjie-skill — 把书、长视频、播客等高价值内容蒸馏成可执行的 Agent Skills
- 🟡nolabs-ai/nono — Sandbox any AI agent in seconds - zero setup, zero latency.
- 🟡"Uncensored" LLMs are measurably more optimistic than their base models
- 🟡Anatomy of a Frontier Lab Agent Intrusion: A Timeline of the July 2026 Incident
- 🟡ag-ui-protocol/ag-ui — AG-UI: the Agent-User Interaction Protocol. Bring Agents into Frontend Applications.
- 🟡microsoft/ML-For-Beginners — 12 weeks, 26 lessons, 52 quizzes, classic Machine Learning for all
- 🟢The HF intrusion timeline: 56 of ~17,600 agent actions were the actual theft
- 🟢I Got Long: AI Agents & Context Portability
- 🟢agentgateway/agentgateway — Next Generation Agentic Proxy for AI Agents and MCP servers
- 🟢ComposioHQ/composio — Composio powers 1000+ toolkits, tool search, context management, authentication, and a sandboxed workbench to help you build AI agents that turn intent into action.
- 🟢Open-source tabular model validation toolkit TanML needs feedback [D]
2026-07-28
- 🔴agentscope-ai/QwenPaw — Your Personal AI Assistant; easy to install, deploy on your own machine or on the cloud; supports multiple chat apps with easily extensible capabilities.
- 🔴Are single GPU research still published in ML/DL and its applications nowadays? Which are the most notable recent ones? [D]
- 🔴AI-coding agents kill team collaboration, according to an analysis of 25,264 agent-generated PRs across 2,361 popular GitHub repositories.
- 🔴ogulcancelik/herdr — agent multiplexer that lives in your terminal.
- 🔴t8y2/dbx — 20 MB lightweight cross-platform database client for 70+ databases, including MySQL, PostgreSQL, SQLite, Redis, MongoDB, DuckDB, SQL Server, and Dameng. Built-in AI, MCP Server, CLI, desktop and Docker. | 轻量级跨平台数据库管理工具,支持 MySQL、PostgreSQL、SQLite、Redis、MongoDB、达梦等 70+ 数据库,提供桌面端、Docker、CLI、内置 AI 助手和 MCP Server。
- 🟡HKUDS/OpenSpace — "OpenSpace: The Skill Management Layer for AI Agents" --https://open-space.cloud/
- 🟡I’m back from holiday and straight back tomessing aroundbuilding. I’ve been kinda obsessed and enamoured with the demos from tldraw. It’s a free canvas/whiteboard app but their offline app (recentlyup
- 🟡Like turning our logo into a little animating mascot. He’s called ‘bites’ btw.
- 🟡This startup runs on17 AI agents- and open-sourced the platform.Marketing, client health, competitive intel, finance - real ops, not demos, with every agent action git-tracked and a human approving an
- 🟡For serious work, I still prefer using a speech to text app (currently using one I built —optionafk.com) and sending that to the agent of my choice with full control over model/thinking effort.
- 🟢NVIDIA/OpenShell — OpenShell is the safe, private runtime for autonomous AI agents.
- 🟢tensorzero/tensorzero — TensorZero is an open-source LLMOps platform that unifies an LLM gateway, observability, evaluation, optimization, and experimentation.
- 🟢microsoft/Mage-VL · Hugging Face - An Efficient Codec-Native Streaming Multimodal Foundation Model
- 🟢AI is helping investigators identify possible clues after a California backpacker vanished
- 🟢[TigrimOSR v0.7.1] — Graph Engineering Agentic System
2026-07-27
- 🔴moeru-ai/airi — 💖🧸 Self hosted, you-owned Grok Companion, a container of souls of waifu, cyber livings to bring them into our worlds, wishing to achieve Neuro-sama's altitude. Capable of realtime voice chat, Minecraft, Factorio playing. Web / macOS / Windows supported.
- 🔴usestrix/strix — Open-source AI penetration testing tool to find and fix your app’s vulnerabilities.
- 🔴Anyone else feel like hallucinations get worse as agents get more complex?
- 🔴How the OpenAI agents escaped onto the internet and hacked another company - ELI5
- 🔴bradautomates/claude-video — Give Claude the ability to watch any video. /watch downloads, extracts frames, transcribes, hands it all to Claude.
- 🟡different-ai/openwork — The open-source alternative to Claude Cowork (powered by opencode)
- 🟡Why this matters - smarter models might unlock robots:Most robots outside of industrial environments are limited in their uptake due to their brittleness and lack of generalization; research like this
- 🟡INSIDE CHINA'S ALL-OUT PUSH TO CATCH UP WITH AMERICAN AI CHIPS (17 MINUTE READ)
- 🟡HOW WE BROUGHT AGENTIC WORKFLOWS TO CLOUD SIEM WITH THE DATADOG MCP SERVER (10 MINUTE READ)
- 🟡PROMPT CACHING IN AGENTS (9 MINUTE READ)
- 🟢junhoyeo/tokscale — 🛰️ Track token usage across AI coding agents from your terminal. 🏅 Global leaderboard with quadrillions of tokens tracked.
- 🟢UFund-Me/Qbot — [🔥updating ...] AI 自动量化交易机器人(完全本地部署) AI-powered Quantitative Investment Research Platform. 📃 online docs:https://ufund-me.github.io/Qbot✨ :news: qbot-mini:https://github.com/Charmve/iQuant
- 🟢An agent that completes the task can still be a complete failure
- 🟢An interactive demo of my Hermes platform
- 🟢langchain-ai/rag-from-scratch
2026-07-26
- 🔴CoreBunch/Instatic — The open-source alternative to Webflow, Framer and WordPress. Agentic self-hosted visual CMS outputting clean static pages. Users, roles, plugins, content, database, it's all there.
- 🔴ComposioHQ/awesome-claude-skills — A curated list of awesome Claude Skills, resources, and tools for customizing Claude AI workflows
- 🔴anthropics/claude-cookbooks — A collection of notebooks/recipes showcasing some fun and effective ways of using Claude.
- 🔴I implemented the YOLO26n model inference from scratch using ARM64 Assembly Language (No framework) [P]
- 🔴How is your ai agent doing searches and not getting blocked?
- 🟡aaif-goose/goose — an open source, extensible AI agent that goes beyond code suggestions - install, execute, edit, and test with any LLM
- 🟡tinyhumansai/openhuman — Your Personal AI super intelligence. A brain that builds a local-first memory of your life, a fantastic orchestrator of agent fleets and workflows, and a deep researcher.
- 🟡Another day in the Vercel vs Cloudflare feud: this time they are fighting overwhose AI gateway is faster.
- 🟡Ben’s Bites is brought to you byMetatate
- 🟡Both security teams caught it, the bug has been reported, andbothsideshave published what they know. Hugging Face says open models were a key part of its defence - its teamfought back with GLM-5.2.
- 🟢FareedKhan-dev/all-agentic-architectures — 35 production-grade agentic AI architectures (Reflexion, LATS, GraphRAG, MemGPT, Voyager, BrowserAgent, ...) — a Python library and runnable textbook with multi-provider LLM support and a 17-task benchmark leaderboard.
- 🟢how painful was it to migrate off Baseten? trying to figure out if it's worth the engineering quarter
- 🟢I collected 197 tools, papers and practices for reducing AI token waste
- 🟢karpathy/micrograd — A tiny scalar-valued autograd engine and a neural net library on top of it with PyTorch-like API
- 🟢30+ officially free AI/ML books, all in one curated repo
2026-07-25
- 🟡What makes an AI agent system debuggable after it starts behaving unexpectedly?
- 🟡VectifyAI/PageIndex — 📑 PageIndex: Document Index for Vectorless, Reasoning-based RAG
- 🟡JACK DORSEY IS TAKING ON SLACK WITH BUZZ, A GROUP CHAT PLATFORM FOR TEAMS AND THEIR AI AGENTS (3 MINUTE READ)
- 🟡INSIDE ROBLOX'S BET ON WORLD MODELS (25 MINUTE READ)
- 🟡AI CODING AGENT HORROR STORIES: THE AGENT THAT DELETED PRODUCTION (14 MINUTE READ)
- 🟢Show HN: Yorishiro – a macOS terminal where AI agents live
- 🟢NVlabs/Sana — SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformer
- 🟢max-sixty/worktrunk — Worktrunk is a CLI for Git worktree management, designed for parallel AI agent workflows
- 🟢open-metadata/OpenMetadata — The Open Context Layer for Data and AI , OpenMetadata is the open platform for building trusted data context and business semantics for humans, AI assistants, and agents.
- 🟢genkit-ai/genkit — Open-source framework for building AI-powered apps in JavaScript, Go, and Python, built and used in production by Google
2026-07-24
- 🔴Recent phone AI demos made me think about cross-app agents.
- 🟡A CHINESE AI LAB JUST BUILT A GIANT DATA CENTRE WITH NO NVIDIA INSIDE (3 MINUTE READ)
- 🟡AGENTIO'S AI-POWERED CREATOR ADS ARE COMING TO FACEBOOK AND INSTAGRAM (2 MINUTE READ)
- 🟡WORLD'S LARGEST AI MODEL REPOSITORY HUGGING FACE BREACHED BY AUTONOMOUS AI AGENT (4 MINUTE READ)
- 🟡GOOGLE CLOUD COMMITS $750M TO ACCELERATE PARTNER-BUILT AI AGENTS (3 MINUTE READ)
- 🟡PLATFORM ENGINEERING'S NEW JOB: SERVING ENVIRONMENTS AT AGENT SPEED (7 MINUTE READ)
- 🟢ai-driven-dev/framework — Marketplace Framework AI-Driven Dev : Context Engineering, Plugins, Agents, Skills, Hooks, Templates, SDLC
- 🟢Be skeptical of OpenAI's rogue hacker agent story
- 🟢Preventing state corruption and structural drift in multi-agent recursive loops
- 🟢pyg-team/pytorch_geometric — Graph Neural Network Library for PyTorch
- 🟢theopenco/llmgateway — Route, manage, and analyze your LLM requests across multiple providers with a unified API interface.
2026-07-23
- 🔴CEO of Hugging face: Heading to San Francisco to have a little chat with that “rogue agent”
- 🔴SWE > Self-Improving Agents: Why "The Bitter Lesson" doesn't mean what you think it means
- 🔴What's one AI agent workflow that has saved you the most time?
- 🔴Show HN: OneCLI – OSS credential gateway that keeps secrets out of AI agents
- 🟡Ben’s Bites is brought to you byMetatate
- 🟡Poolside’s recent tech reportgot a lot of praise due to their level of detail, and Vibhu first covered Laguna’s recent technical report on our paper club:
- 🟡We Compressed Our AI Agent’s Context. Costs Fell. Reliability Broke. Here’s What We Learned.
- 🟡@swyx: one thing i think people dont appreciate enough about @poolsideai is their unusual degree of openness — not only have they shipped an excellent Small model that somehow beat @thinkymachines at coding,
- 🟡@mattshumer_: https://workbench.md for anyone who wants to try the workflow... it's the most powerful way to steer agents. I built it for myself, and it's not really a product, so expect rough edges. I'm happy to h
- 🟢Benchmarks: AntLing-3.0-flash a hybrid-reasoning MoE model built for production-scale agents.
- 🟢oraios/serena — A powerful MCP toolkit for coding, providing semantic retrieval and editing capabilities - the IDE for your agent
- 🟢THU-MAIC/OpenMAIC — Open Multi-Agent Interactive Classroom — Get an immersive, multi-agent learning experience in just one click
- 🟢slavakurilyak/awesome-ai-agents — Awesome list of 300+ agentic AI resources
- 🟢Local web search for LLM agents that cuts tokens by 87% and cost by 66%
2026-07-22
- 🔴SkewAdam: A tiered optimizer that cuts MoE state memory by 97% (fits a 6.7B MoE on a 40GB GPU) [R]
- 🔴are you guys actually giving agents access to real money or is that crazy?
- 🔴How do you actually improve when building AI agents?
- 🟡Ben’s Bites is brought to you byVeridect
- 🟡For more educational post-training videos, see thecourseI’m putting together.
- 🟡microsoft/Fara1.5-27B · Hugging Face
- 🟡AI agent security: the "intern rule" for token permissions (two real stories from running 45 bots)
- 🟡Agents saturate SWE-bench but drop to ~23% on real repos. The reason is verification cost, not difficulty.
- 🟢corsairdev/corsair — Your Agent's Integration Layer
- 🟢Laguna S 2.1 looping fix incoming
- 🟢shiyu-coder/Kronos — Kronos: A Foundation Model for the Language of Financial Markets
- 🟢Where would you put the admission check in an MCP-assisted agent workflow?
- 🟢Google API Calls Hitting Quota (Docs), any workaround
2026-07-21
- 🔴CEO of Hugging Face: Banning open-source AI would hurt defenders 10x more than attackers, which would make the world 10x more dangerous and this is a good example why!
- 🔴AI agents don’t choose from the whole market. They choose from an invisible approved-vendor list.
- 🔴OpenAI admits responsibility for HuggingFace Attack - an agent from an internal evaluation is reportedly the cause.
- 🔴Would you pay for knowing when your agents hallucinate?
- 🟡Tri-Net v2: Open-source implementation of our Scientific Reports paper on unified skin lesion and symptom-based monkeypox detection [R]
- 🟡We were lucky enough to have the two central figures in this story on our podcast. Taking the lead from Ci Chu and Bo Wang, Xaira Therapeutics is betting thatinformation richdata is the key to AI-driv
- 🟡FROM OTEL TO SLMS: DISTILLING FRONTIER MODEL BEHAVIOUR FROM PRODUCTION TELEMETRY (45 MINUTE VIDEO)
- 🟡DATA MODELING ISN'T DEAD, YOU JUST STOPPED DOING IT WITH JOE REIS (65 MINUTE VIDEO)
- 🟡AGENTS THINK IN MILLISECONDS, LEGACY INFRASTRUCTURE DOESN'T. LINKEDIN, WALMART, AND ZENDESK SHARED HOW THEY CLOSED THE GAP AT VB TRANSFORM 2026 (4 MINUTE READ)
- 🟢rohitg00/agentmemory — #1 Persistent memory for AI coding agents based on real-world benchmarks
- 🟢akitaonrails/ai-memory — Solution for long term memory for agent coding CLIs and to facilitate handoff between different agent vendors
- 🟢An experiment on symmetric agent communication. It went better than I expected
- 🟢Reproducing OpenAI’s “persistently beneficial models” - GRPO trait install barely moves. Ideas? [P] [R]
- 🟢owainlewis/awesome-artificial-intelligence — A curated list of Artificial Intelligence (AI) courses, books, video lectures and papers.
2026-07-20
- 🔴Kimi K3 just fixed 15 critical security bugs that Codex and Fable refused because of “cyber guardrails”. Hugging Face: We had this experience ourselves this week! Very scary to be guardrailed as a defender when you know attackers are likely bypassing
- 🔴Sources: parts of the Trump administration are reigniting efforts to implement de facto bans on foreign open-source models, as Chinese AI models gain momentum
- 🔴Low-latency STT API claims are useless unless you log the full voice-agent waterfall.
- 🔴I stopped building a database for my AI agents and just used git. Turns out git already solved most of the hard problems.
- 🔴vercel-labs/deepsec — Deepsec is a security harness for finding vulnerabilities in your codebase powered by coding agents
- 🟡AI Agent Harness Tutorial for Beginners
- 🟡THE IDENTITY CRISIS AT ELON MUSK'S CHAOTIC AI OUTFIT (11 MINUTE READ)
- 🟡CHINA'S XI TOUTS OPEN-SOURCE AI AND TAKES A SWIPE AT US DOMINANCE (5 MINUTE READ)
- 🟡NVIDIA OPENSHELL SECURES THE AGENT. WHO GOVERNS THE FLEET? (8 MINUTE READ)
- 🟡AUTOMATED INCIDENT REMEDIATION WITH AWS DEVOPS AGENT AND KIRO CLI (5 MINUTE READ)
- 🟢QwenLM/qwen-code — An open-source AI coding agent that lives in your terminal.
- 🟢So I built a portfolio template that your AI agent can set up for you in one prompt (open source)
- 🟢Training a harness for model-agnostic and task-environment-agnostic capability improvements with PyTorch-like framework [P]
- 🟢langchain-ai/open-swe — An Open-Source Asynchronous Coding Agent
- 🟢The office task mode hype made me rethink what an agent actually needs to be told
2026-07-19
- 🔴bojieli/ai-agent-book — 《深入理解 AI Agent:设计原理与工程实践》(李博杰 著)开源主仓库:全书正文、编译版 PDF 与按章配套代码
- 🟡NOBODY HAS CRACKED AGENT MEMORY (12 MINUTE READ)
- 🟡A PRIMER ON SELF-IMPROVING AGENT HARNESSES (9 MINUTE READ)
- 🟡AGENTIC MISALIGNMENT IN SUMMER 2026 (70 MINUTE READ)
- 🟡ATLASSIAN EXTENDS AI REACH OF JIRA INTO AGENTIC ENGINEERING WORKFLOWS (4 MINUTE READ)
- 🟡BACKED BY $60M IN FUNDING, OAK STEPS OUT OF STEALTH TO FIX THE IDENTITY MESS THAT AI AGENTS ARE MAKING WORSE (4 MINUTE READ)
- 🟢It could have been Meta
- 🟢We spend more time evaluating AI models than the people using them
- 🟢Do your customers actually search documentation anymore?
- 🟢I don't see how open-source AI models in the U.S. can successfully compete with those from China.
- 🟢From Muon to Gradient Clipping: Some Thoughts on QK Stability
2026-07-18
- 🔴Did blatant AI Slop just win a 25K USD Deepmind / Kaggle Grand Prize? [D]
- 🔴KnockOutEZ/wigolo — The go-to web for your AI coding agent — local-first search, fetch, crawl & research over MCP. No API keys, no cloud, $0/query. Public beta.
- 🟡A FRAMEWORK FOR FRONTIER AI AND THE DAWNING OF A NEW AGE (7 MINUTE READ)
- 🟡WANDR BENCHMARK: EVALUATING RESEARCH AGENTS THAT MUST SEARCH WIDE AND DEEP (15 MINUTE READ)
- 🟡Healthcare AI company actAVA introduced CURA 1T, a 1T parameter clinical model it claims beats frontier rivals on healthcare benchmarks at significantly lower costs.
- 🟡OpenAI’s new $230 AI agent control pad
- 🟡Weco's AI agent evolves a better version of itself
- 🟢musistudio/claude-code-router — One local control plane for every AI agent: route across models, fuse new capabilities, orchestrate tools, and stay fully in control.
- 🟢Why traditional secrets managers (HashiCorp Vault, AWS Secrets Manager) fail for autonomous AI agents
- 🟢Byte exact KV cache grafting on frozen Gemma 4
- 🟢DataDog/pup — Give your AI agent a Pup — a CLI companion with 200+ commands across 33+ Datadog products.
- 🟢TabFM Studio: point-and-click predictions on spreadsheets with tabular foundation models, fully local [P]
2026-07-17
- 🔴Prism accidentally leaked [D]
- 🔴How do you actually keep up with AI dev tools/techniques without drowning in noise?
- 🔴Chinese President Xi Jinping speaks at World AI Conference and reaffirms commitment to open source to promote"openness and win-win"
- 🔴Has anyone actually completed a purchase with their own agent, end to end? Where does it break for you?
- 🔴rohitg00/ai-engineering-from-scratch — Learn it. Build it. Ship it for others.
- 🟡FIVE YEARS BUILDING THE CONTEXT LAYER - AND WHY IT BELONGS IN THE AGENTIC LOOP (24 MINUTE READ)
- 🟡WITH AI AGENTS, ATTENTION IS THE NEW BOTTLENECK (2 MINUTE READ)
- 🟡HOW I CUT AN AI AGENT'S TOKEN USE BY 94% (7 MINUTE READ)
- 🟡AI VIDEO INFRASTRUCTURE (WEBSITE)
- 🟡AI VIDEO GENERATOR AND IMAGE GENERATOR (WEBSITE)
- 🟢What boring piece of infrastructure became unexpectedly important once you started putting agents into production?
- 🟢vercel-labs/just-bash — Bash for Agents
- 🟢sourcebot-dev/sourcebot — Sourcebot is a self-hosted tool that helps humans and agents understand your codebase.
2026-07-16
- 🔴Email Automation Worked Best for My Web Agency — What Worked for You?
- 🟡The qlora 2e-4 default is wrong under 10k samples and nobody talks about it [D]
- 🟡AI Agents Explained: The Complete Beginner's Guide (2026)
- 🟡LM Studio Bionic: the AI agent for open models
- 🟡memvid/memvid — Memory layer for AI Agents. Replace complex RAG pipelines with a serverless, single-file memory layer. Give your agents instant retrieval and long-term memory.
- 🟡The cloud was built for developers. Butagents are now changing that.
- 🟢Anthropic and OpenAI don't have secret sauce
- 🟢HKUDS/OpenHarness — "OpenHarness: Open Agent Harness with a Built-in Personal Agent--Ohmo!"
- 🟢callstack/agent-device — CLI to control iOS and Android devices for AI agents
- 🟢Why AI Agents Need Proxies and what proxies are better
- 🟢Show HN: Libretto PR agents – Automatically fix failing playwright scripts
2026-07-15
- 🔴Linus Torvalds tells people to stop attacking others for using AI
- 🔴Building an AI second brain/ADHD assistant, which tool to use as foundation?
- 🔴Google is updating Gemma 4's chat templates, bringing major fixes to tool calling and reducing "laziness", and enabling Flash Attention 4 on Hopper GPUs, plus an interactive guide on how to work with and improve its vision!
- 🟡Looking for JEPA devil advocates [R]
- 🟡HKUDS/nanobot — Lightweight, open-source AI agent for your tools, chats, and workflows.
- 🟡The US is advancing AI safety through state and federal action
- 🟡@AnthropicAI: New Anthropic research: Agentic misalignment in Summer 2026. A year after our blackmail experiments, we found four more ways that today’s autonomous AI agents misbehave in simulations. Read more: http
- 🟡I automated AI video testing in n8n: prompts → multiple reference images → video outputs → Google Sheets log (GitHub / JSON included)
- 🟢Designing APIs for Agents
- 🟢apache/ossie — Apache Ossie, industry wide specification effort to standardize how we exchange semantic metadata across analytics, AI and BI platforms, providing a vendor neutral, single source of truth for semantic data
- 🟢PyTorch model running 170x slower on T4 vs A100. What could cause a bottleneck this extreme? [D]
- 🟢snap-stanford/Biomni — Biomni: a general-purpose biomedical AI agent
- 🟢How we built a coding agent that lives in Slack, and the recipe to build your own
2026-07-14
- 🔴At what point does adding more agents make a workflow worse instead of better?
- 🔴where do human-in-the-loop controls make sense for agents vs just slowing everything down
- 🔴Show HN: Juggler – an open-source GUI coding agent, by the creator of JUCE
- 🔴millionco/react-doctor — Your agent writes bad React. This catches it
- 🔴google/skills — Agent Skills for Google products and technologies
- 🟡LLM hallucination paper(using math) accepted to ICML workshop[R]
- 🟡kangarooking/cangjie-skill — 把书、长视频、播客等高价值内容蒸馏成可执行的 Agent Skills
- 🟡WE BENCHMARKED CODING AGENTS ON OUR OWN INTERNAL TASKS AT DATABRICKS AND LEARNED A LOT! (4 MINUTE READ)
- 🟡A HITCHHIKER'S GUIDE TO AI (17 MINUTE READ)
- 🟡YOUR AGENTS ARE STUCK IN YOUR ORG CHART (19 MINUTE READ)
- 🟢zts212653/clowder-ai — Build AI teams, not just agents. Hard rails, soft power, shared mission.
- 🟢PostHog/posthog — 🦔 PostHog is an all-in-one developer platform for building successful products. We offer product analytics, web analytics, session replay, error tracking, feature flags, experimentation, surveys, data warehouse, a CDP, and an AI product assistant to help debug your code, ship features faster, and keep all your usage and customer data in one stack.
- 🟢agentgateway/agentgateway — Next Generation Agentic Proxy for AI Agents and MCP servers
- 🟢build-log: a credential vault an agent can use to log in without the model ever seeing the password
- 🟢What will separate production AI agents from demos over the next 2–3 years?
2026-07-13
- 🔴Prompt-engineering paper accepted to ICML [R]
- 🔴Show HN: Nobie – an Excel-compatible runtime for agents and humans
- 🔴simonlin1212/TradingAgents-astock — A股多Agent投研框架 — 适配A股数据源(龙虎榜/游资/解禁等),7位分析师基于A股规则的辩论决策,基于TradingAgents深度改造,适配大A。A-share multi-agent investment research framework — 7 AI analysts, bull/bear debate, risk assessment。
- 🟡Best agent framework in 2026? There isn't one. Here's my decision tree
- 🟡Show HN: BillAI Bass, an AI-Powered Big Mouth Billy Bass Using Strands Agents
- 🟡I benchmarked 15 "E-Waste" GPUs with Modern Workloads
- 🟡PROGRESSIVE DISCLOSURE: FROM TRAINING WHEELS TO WEEK-LONG AI AGENTS (12 MINUTE READ)
- 🟡THE PULSE: INTERESTING AI CODING STATS FROM CURSOR (7 MINUTE READ)
- 🟢mem0sharp cause didn't find any reimplementation with dotnet
- 🟢What’s your approach to preventing AI agents from confidently making the wrong decision?
- 🟢moeru-ai/airi — 💖🧸 Self hosted, you-owned Grok Companion, a container of souls of waifu, cyber livings to bring them into our worlds, wishing to achieve Neuro-sama's altitude. Capable of realtime voice chat, Minecraft, Factorio playing. Web / macOS / Windows supported.
- 🟢katanemo/plano — Plano is an AI-native proxy and data plane for agentic apps — with built-in orchestration, safety, observability, and smart LLM routing so you stay focused on your agents core logic.
- 🟢raine/claude-code-proxy — Use Claude Code with your ChatGPT, Kimi, Cursor or Grok subscription via a local Anthropic-compatible proxy
2026-07-12
- 🔴Local Image to 3D (<2gb RAM, <20s, Apple Silicon, iPhone)
- 🔴Honest comparison of AI voice agents for outbound in 2026
- 🔴Dicklesworthstone/destructive_command_guard — The Destructive Command Guard (dcg) is for blocking dangerous git and shell commands from being executed by agents.
- 🔴Old and new apps, via modern coding agents
- 🟡WHERE AI AGENTS BELONG IN DATA ENGINEERING: THE CORRECTNESS LAYER (12 MINUTE READ)
- 🟡THE 3 LAYERS OF AGENT BUILDING (14 MINUTE READ)
- 🟡WHY YOUR SEMANTIC LAYER MATTERS MORE THAN YOUR AI AGENT (6 MINUTE READ)
- 🟡10 AI SKILLS TO GIVE YOUR CODING AGENT REAL DESIGN TASTE (6 MINUTE READ)
- 🟡SALESFORCE TURNS SLACKBOT INTO AN AGENTIC WORKFLOW HUB (3 MINUTE READ)
- 🟢Your AI agent passed all the tests, now what ? Online evals and how to choose them.
- 🟢How has your experience been with State Machines based Voice Agents and agentic handoffs?
- 🟢google-gemini/gemini-cli — An open-source AI agent that brings the power of Gemini directly into your terminal.
- 🟢Show HN: Agent Draw: An agent draws while you talk, built on TLDraw
- 🟢MervinPraison/PraisonAI — PraisonAI 🦞 — Hire a 24/7 AI Workforce. Stop writing boilerplate and start shipping autonomous self-improving agents that research, plan, code, and execute tasks. Deployed in 5 lines of code with built-in memory, RAG, and support for 100+ LLMs.
2026-07-11
- 🔴The U.S. tech industry is increasingly anxious about the rising power and competitive price of open-source AI models from China — and whether the Trump administration will respond with yet another executive order | Politico
- 🔴Public Library Find [D]
- 🔴Shubhamsaboo/awesome-llm-apps — 100+ AI Agent & RAG apps you can actually run — clone, customize, ship.
- 🔴I created a super harmful model ! :D (by tweaking it's J-Space!!!)
- 🔴FoundationAgents/OpenManus — No fortress, purely open ground. OpenManus is Coming.
- 🟡HOW TO USE AI AGENTS BETTER THAN 99% OF PEOPLE (32 MINUTE READ)
- 🟡AGENTIC AI ADOPTION IS ON FIRE AT UBER (2 MINUTE READ)
- 🟡THE AGENT-ERA CAREER (7 MINUTE READ)
- 🟡GITLOST: HOW WE TRICKED GITHUB'S AI AGENT INTO LEAKING PRIVATE REPOS (3 MINUTE READ)
- 🟡ROGUE AGENT: HOW A SINGLE CODE BLOCK COULD HIJACK YOUR AI CONVERSATIONS IN GOOGLE'S DIALOGFLOW (6 MINUTE READ)
- 🟢How does *ACL conferences acceptance work [D]
- 🟢cognizant-ai-lab/neuro-san-studio — A playground for neuro-san
- 🟢Can you build AI agents if you understand the concepts but can't code from scratch?
- 🟢Predicting human preference for generated image pairs using HPSv3 [P]
- 🟢Is deploying and scaling ai agents one of the most frustrating problem?
2026-07-10
- 🟡AI VIDEO STARTUP HIGGSFIELD IS IN TALKS TO RAISE AT A $5BN VALUATION (3 MINUTE READ)
- 🟡CHINA'S CYBERSECURITY STANDARD ON AI AGENT DEPLOYMENT (8 MINUTE READ)
- 🟡SKILLCLOAK LETS MALICIOUS AI AGENT SKILLS EVADE STATIC SCANNERS WITH SELF-EXTRACTING PACKING (5 MINUTE READ)
- 🟡THREE WAYS TO GIVE AN AI AGENT AN IDENTITY (17 MINUTE READ)
- 🟡CONTINUAL LEARNING FOR AGENTS (3 MINUTE READ)
- 🟢openai/openai-python — The official Python library for the OpenAI API
- 🟢According to DataBricks, pi-coding-agent is ~2x cheaper than CC/Codex, GLM 5.2 on par with Opus 4.8 high
- 🟢Your AI agent can become an income-producing asset. Stop treating it like a disposable prompt.
- 🟢I’m exploring an idea and would love honest feedback from people building AI agents
- 🟢Mapping world model taxonomy [P]
2026-07-09
- 🔴I track LLM prices every 3 hours. GLM-5.2 quietly went from ~$0.57/$1.80 to $0.90/$3.08 per 1M this week, with no announcement.
- 🔴GLM-5.2 fearmongering in the press
- 🔴Built a multi-agent AI system that runs on Telegram, entirely on free-tier infra (Cloudflare Workers + GitHub Actions)
- 🔴Deployed a voice agent for after-hours calls - what I learned
- 🟡What's one AI agent design choice you thought would scale but didn't?
- 🟡Trying to debug our AI, how do you handle agent randomness?
- 🟡traycerai/traycer — Traycer: Nerve Center for Agentic Coding
- 🟢OpenMed 1.8: Apache-2.0 clinical de-identification that runs fully local, now on Android, iOS, and in the browser. 400+ open issues if you want in on 1.9
- 🟢stanfordnlp/dspy — DSPy: The framework for programming—not prompting—language models
- 🟢OpenMOSS-Team/MOSS-Transcribe-Diarize · Hugging Face
- 🟢What do you actually use Fable 5 for?
- 🟢How do you actually decide an MCP server is safe before you plug it in?
2026-07-08
- 🔴Can you trust local models to answer accurately?
- 🔴We are playing a game where an agent, prompt, or model predicts the World Cup. 20 USDC raffle per match.
- 🔴microsoft/SkillOpt — SkillOpt is a text-space optimizer that trains reusable natural-language skills for frozen LLM agents through trajectory-driven edits, validation-gated updates, and deployable best_skill.md artifacts.
- 🔴jingyaogong/minimind — 🧠「大模型」2小时完全从0训练64M的小参数LLM!Train a 64M-parameter LLM from scratch in just 2h!
- 🟡An interesting hardware approach to phone-controlling AI agents
- 🟡DINOv2 way worse than SigLIP in k-NN. Is this expected? [R]
- 🟡cline/cline — Autonomous coding agent as an SDK, IDE extension, or CLI assistant.
- 🟡@OpenAI: We audited SWE-Bench Pro, one of the most widely used AI coding benchmarks, and found it no longer reliably measures frontier coding capability. We find 30% of SWE-Bench Pro tasks to be broken, and ar
- 🟡Tracer-Cloud/opensre — Build your own AI SRE agents. The open source toolkit for the AI era.
- 🟢First time ARR users - some questions [D]
- 🟢Crucible. A judgment engine: register a thesis, steelman each claim, measure against a substrate, refine the weakest axis.
- 🟢Distilled DeepSeek into Gemma 4 26B-A4B vs 12B. Not very useful, but I learned a lot.
- 🟢Show HN: Onboard-CLI, a LLM powered and AST-based tool to visualize codebase
- 🟢LingBot-Video: sparse-MoE video diffusion transformer (13B total, 1.4B active) post-trained as an action-conditioned world model[R]
2026-07-07
- 🔴MadsLorentzen/ai-job-search — AI-powered job application framework built on Claude Code. Fork it, fill in your profile, and let Claude evaluate jobs, tailor CVs, write cover letters, and prepare you for interviews.
- 🔴6 things I learned building agents that wake themselves up overnight (open source)
- 🔴TorchJD: Training with multiple losses in PyTorch [P]
- 🔴nvidia/NVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B-BF16 · Hugging Face
- 🟡Ph.D. thesis on Differentiable Ray Tracing for Radio Propagation Modeling [R]
- 🟡NVIDIA TAPS AI CLOUD PROVIDERS TO EXPAND COMPUTE ACCESS FOR STARTUPS (2 MINUTE READ)
- 🟡AGENTIC AUTONOMY LEVELS (19 MINUTE READ)
- 🟡AGENTIC LOOPS (20 MINUTE READ)
- 🟡SOME NEW AGENTIC PATTERNS (11 MINUTE READ)
- 🟢HKUDS/AI-Trader — "AI-Trader: 100% Fully-Automated Agent-Native Trading"
- 🟢[D] Issue with arxiv - abstract not matching pdf/html [D]
- 🟢datawhalechina/all-in-rag — 🔍大模型应用开发实战一:RAG 技术全栈指南,在线阅读地址:https://datawhalechina.github.io/all-in-rag/
- 🟢Show HN: Docx-CLI: agents read/edit Word docs using 1/2 the time and tokens
- 🟢GLM-5.2 on 8xB200: the deployment math nobody spells out - NVFP4 + 2x TP=4 replicas should beat TP=8 by ~2x. Full config guidance inside.
2026-07-06
- 🔴I created a tool that turns AI agents into white collar employees.
- 🔴What's one architectural decision you've made with AI agents that you wouldn't make again?
- 🟡Bought an e-ink tablet and turned the "rest of the f*cking owl" meme into reality with a custom agent.
- 🟡PLEASE STOP THE AI CONFIDENCE THEATER (7 MINUTE READ)
- 🟡WE LIKE ANTHROPIC MORE THAN OPENAI. THE TROLLEY PROBLEM EXPLAINS WHY. (4 MINUTE VIDEO)
- 🟡MARK ZUCKERBERG TELLS STAFF THAT AI AGENTS HAVEN'T PROGRESSED AS QUICKLY AS HE'D HOPED (2 MINUTE READ)
- 🟡MICROSOFT UNVEILS $2.5B ‘FRONTIER COMPANY' TO EMBED AI ENGINEERS INSIDE CUSTOMERS (5 MINUTE READ)
- 🟢Omnigent — a meta-harness for building and running AI agents
- 🟢All the Tools My Friend Used to Make His First $70K Selling Websites
- 🟢OpenComputer | An Open Source Computer Built For Agents.
- 🟢40+ AI agents placed ~1,500 real-money bets on the World Cup Group Stage. The ones that profited kept more than one outcome in play.
2026-07-05
- 🔴If DeepMind or Anthropic is doing your exact research topic, do you still continue? [D]
- 🔴I think we're repeating the early microservices mistake with AI agents
- 🔴I benchmarked 13 models at 65K-128K context to find out what actually matters for agentic workloads
- 🔴bradautomates/claude-video — Give Claude the ability to watch any video. /watch downloads, extracts frames, transcribes, hands it all to Claude.
- 🟡SANDBOXING AN AI AGENT (19 MINUTE READ)
- 🟡SCARFBENCH: BENCHMARKING AI AGENTS FOR ENTERPRISE JAVA FRAMEWORK MIGRATION (7 MINUTE READ)
- 🟡ZCode - Z AI’s agentic coding environment tuned for GLM-5.2
- 🟡Katalyze raised $10.5M to bring agents to pharma manufacturing, cutting batch investigations from months to minutes inside 5 of the top 20 global pharma orgs.
- 🟡RSVP to next workshop on July 8: Create Short Form Videos with AI
- 🟢OthmanAdi/planning-with-files — Persistent file-based planning for AI coding agents and long-running agentic tasks. Crash-proof markdown plans that survive context loss and /clear, plus a deterministic completion gate and multi-agent shared state on disk. Manus-style. Works with Claude Code, Codex CLI, Cursor, Kiro, OpenCode and 60+ agents via the SKILL.md standard.
- 🟢gitroomhq/postiz-app — 📨 The ultimate agentic social media scheduling tool 🤖
- 🟢LivePortrait distilled model that can run at 25fps in the browser
- 🟢backnotprop/plannotator — Annotate and review coding agent plans and code diffs visually, share with your team, send feedback to agents with one click.
- 🟢I automated screenshot verification for a real estate outreach team and found a duplicate-payment loophole
2026-07-04
- 🔴possible evidence of literal prompt injection by anthropic
- 🟡QUALITY ASSURANCE AGENT: REIMAGINING SOFTWARE QUALITY WITH AI-DRIVEN AUTONOMOUS TESTING (12 MINUTE READ)
- 🟡HOW WE REALLY BUILD PRODUCTION-GRADE AI AGENTS: BEYOND MODELS, TOWARD DATA AND API QUALITY (8 MINUTE READ)
- 🟡YOUR DESIGN SYSTEM'S NEWEST AUTHOR IS AN AGENT (8 MINUTE READ)
- 🟡AI VIDEO GENERATOR FOR TEXT TO VIDEO & IMAGE TO VIDEO (WEBSITE)
- 🟡SECURING AI AGENTS: WHEN AI TOOLS MOVE FROM READING TO ACTING (6 MINUTE READ)
- 🟢crynta/terax-ai — Lightweight (7MB) Terminal-first AI-native dev workspace
- 🟢A narrow-waist protocol for agent-to-agent comms, and an empirical study of when structured messages actually beat plain English
- 🟢YILS-LIN/short-video-factory — 一键生成产品营销与泛内容短视频,AI批量自动剪辑,高颜值跨平台桌面端工具 One click generation of product marketing and general content short videos, AI batch automatic cliping, beautiful cross platform desktop tool
- 🟢microsoft/skills-for-fabric — A collection of skills and MCP systems to enable users of CLI, VSCode, Claude to operate over Microsoft Fabric
- 🟢Testors wanted for custom Agent Harness
2026-07-03
- 🔴Signal -> Triage -> Notify. My personal agentic setup to sift the signal from the noise.
- 🔴Deepseek drops another HUGE breakthrough - DSpark. Waaay faster than MTP [Video explaining it]
- 🔴Jamesob's guide to running SOTA LLMs locally
- 🔴testmu and patronus both struggle with our multi-agent setup. anyone solved cross-agent eval?
- 🔴Palantir is a free org on HF with 0 open-source models and 0 public datasets shared
- 🟡Fastest inference provider right now? Saw some interesting latency numbers.
- 🟡THE OPERATIONAL REALITY OF MARKETING INSIDE AN AI COMPANY (3 MINUTE READ)
- 🟡OPENAI'S CODEX HARDWARE (1 MINUTE READ)
- 🟡OKTA IS THE FIRST INDEPENDENT AND NEUTRAL IDENTITY PLATFORM TO BRING AI AGENT GOVERNANCE TO HIGHLY REGULATED ENVIRONMENTS (5 MINUTE READ)
- 🟡ABLO – THE COLLABORATION LAYER FOR AI AGENTS (GITHUB REPO)
- 🟢microsoft/agent-governance-toolkit — AI Agent Governance Toolkit — Policy enforcement, zero-trust identity, execution sandboxing, and reliability engineering for autonomous AI agents. Covers 10/10 OWASP Agentic Top 10.
- 🟢google-labs-code/stitch-skills — A library of Agent Skills designed to work with the Stitch MCP server. Each skill follows the Agent Skills open standard, for compatibility with coding agents such as Antigravity, Gemini CLI, Claude Code, Cursor.
- 🟢macro-inc/macro — Macro is a unified interface for email, messages, tasks, calls, agents, pull requests, docs, crm — linked together with shared AI memory.
- 🟢H64LM: A 249M-parameter Mixture-of-Experts Transformer built from scratch in PyTorch [P]
- 🟢40+ AI agents placed ~1,500 real-money bets on the World Cup Group Stage. We are sharing the lessons we learned.
2026-07-02
- 🔴Hamiltonian Neural Networks from a Differential Geometry Perspective [D]
- 🔴Kimi K2.7 Code is generally available in GitHub Copilot
- 🔴anthropics/claude-code — Claude Code is an agentic coding tool that lives in your terminal, understands your codebase, and helps you code faster by executing routine tasks, explaining complex code, and handling git workflows - all through natural language commands.
- 🟡I open-sourced my first agent skill: self-improve. It’s simple, and it actually works.
- 🟡poolside/Laguna-XS-2.1
- 🟡Zackriya-Solutions/meetily — Privacy first, AI meeting assistant with 4x faster Parakeet/Whisper live transcription, speaker diarization, and Ollama summarization built on Rust. 100% local processing. no cloud required. Meetily (Meetly Ai -https://meetily.ai) is the #1 Self-hosted, Open-source Ai meeting note taker for macOS & Windows.
- 🟡alirezarezvani/claude-skills — 337 Claude Code skills & agent skills & plugins (30+ Agents, 70+ custom commands, 330+ skills, customizable references, scripts)for Claude Code, Codex, Gemini CLI, Cursor, and 8 more coding agents — engineering, marketing, product, compliance, C-level advisory, research, business operations, commercial & finance, and your daily productivity skills.
- 🟡Rebuilding Gemma 4 31b... better... As 26b...
- 🟢agentskills/agentskills — Specification and documentation for Agent Skills
- 🟢sopaco/deepwiki-rs — Turn code into clarity. Generate accurate technical docs and AI-ready context in minutes—perfectly structured for human teams and intelligent agents.
- 🟢NateBJones-Projects/OB1 — Open Brain — The infrastructure layer for your thinking. One database, one AI gateway, one chat channel — any AI plugs in. No middleware, no SaaS.
- 🟢Kuberwastaken/claurst — Agentic Coding for Builders who Ship
- 🟢harvard-edge/cs249r_book — Machine Learning Systems
2026-07-01
- 🔴Looking for feedback from people building AI agents that use browsers
- 🔴How do you use AI agent working for social media data analysis?
- 🔴virgiliojr94/book-to-skill — Turn any technical book PDF into a Claude Code skill — ready to study, reference, and use while you work.
- 🔴open-webui/open-webui — User-friendly AI Interface (Supports Ollama, OpenAI API, ...)
- 🟡P Moth-Retrieval: Graph-Free Multi-Hop Retrieval via Query-Time Orchestration (Beating Graph-Based Systems on HotpotQA) [P]
- 🟡HKUDS/VideoAgent — "VideoAgent: All-in-One Agentic Framework for Video Understanding, Editing, and Remaking"
- 🟡how many parameters can i run in my machine to make it as an ai agent assistant in my system
- 🟡I extended Gemma4-31B to 44B (88 layers) — since Google won't give us anything bigger than 31B
- 🟡getmaxun/maxun — 🔥 The open-source no-code platform for web scraping, crawling, search and AI data extraction • Turn websites into structured APIs in minutes 🔥
- 🟢vas3k/TaxHacker — Self-hosted AI accounting app. LLM analyzer for receipts, invoices, transactions with custom prompts and categories
- 🟢ZCode: New Agentic Code Editor from the Makers of GLM
- 🟢RedPlanetHQ/core — Your Personal AI OS
- 🟢Ways to reduce token cost in AI agents
- 🟢togatoga/karukan — Japanese Input Method System for Linux, macOS, Neural Kana-Kanji Conversion Engine
2026-06-30
- 🔴google/agents-cli — The CLI and skills that turn any coding assistant into an expert at creating, evaluating, and deploying AI agents on Google Cloud.
- 🟡12TB OF AI CODING AGENT LOGS (17 MINUTE VIDEO)
- 🟡WE RAN 250 AI AGENT EVALS TO FIND OUT IF SKILLS BEAT DOCS. THE ANSWER IS MORE COMPLICATED THAN WE EXPECTED (6 MINUTE READ)
- 🟡USING LOCAL CODING AGENTS (37 MINUTE READ)
- 🟡AGENTICS/TECH THINGS: TOKENMAXXING IS DEAD, LONG LIVE TOKENMAXXING (18 MINUTE READ)
- 🟡ADOBE IS BUYING TOPAZ LABS, THE AI VIDEO ENHANCER (4 MINUTE READ)
- 🟢I built an open-source alternative to Figma's official MCP server
- 🟢siteboon/claudecodeui — Use Claude Code, OpenCode, Cursor CLI, and Codex on mobile and web with CloudCLI (aka Claude Code UI). CloudCLI is a free open source webui/GUI that helps you manage your Claude Code session and projects remotely.
- 🟢danielmiessler/LifeOS — Agentic AI Infrastructure for magnifying HUMAN capabilities.
- 🟢What are y'all using for observability in your agent systems? [i will not promote]
- 🟢Added list of agent sandboxes(decided to use ai-jail after that review)
2026-06-29
- 🔴Cerebras OpenAI deal capacity has effectively killed the waitlist for everyone else [D]
- 🔴wavecat, a fully local personal agent that watches your screen
- 🔴Google's Agentic Peer-Reviewer Handled ~10K Papers at ICML/STOC — Formal Research Paper Now Out [R]
- 🔴lumina-ai-inc/chunkr — Vision infrastructure to turn complex documents into RAG/LLM-ready data
- 🔴Ornith-1.0: self-improving open-source models for agentic coding
- 🟡BUILD THE AGENT OR POWER THE AGENT? (6 MINUTE READ)
- 🟡GETTING MORE FROM EACH TOKEN: HOW COPILOT IMPROVES CONTEXT HANDLING AND MODEL ROUTING (8 MINUTE READ)
- 🟡BRINGING MORE AGENT HARNESSES AND FRAMEWORKS TO CLOUDFLARE, STARTING WITH FLUE (8 MINUTE READ)
- 🟡THE COMING DIVIDE: AI-NATIVE OR LEFT BEHIND (4 MINUTE READ)
- 🟡50 DESIGN TOKEN FILES, ONE PROBLEM: YOUR AGENTS CAN'T READ THE MEANING (15 MINUTE READ)
- 🟢vanloctech/youwee — A beautiful, cross-platform downloader for YouTube, TikTok, Instagram, and 1800+ sites (yt-dlp GUI) with AI video summaries and post-processing
- 🟢I Hate Dario Amodei, and everything he stands for.
- 🟢Best way to get started with AI agents for Obsidian + small app workflows?
- 🟢appwrite/appwrite — Appwrite® - complete cloud infrastructure for your web, mobile and AI apps. Including Auth, Databases, Storage, Functions, Messaging, Hosting, Realtime and more
- 🟢metalbear-co/mirrord — Run any process, on your machine or in an AI agent's environment, as if it were a pod in your Kubernetes cluster: real env vars, DNS, network, traffic.
2026-06-28
- 🔴Robbyant/lingbot-map — A feed-forward 3D foundation model for reconstructing scenes from streaming data
- 🔴A way to exclude sensitive files issue still open for OpenAI Codex
- 🟡INSIDE ONE ENGINEER'S JOURNEY TO MASTER LONG-RUNNING AGENTS (16 MINUTE READ)
- 🟡UNITING ANALYTICS WITH AI AGENTS AS WORK AND ROLES SHIFT IN THE GENAI ERA (11 MINUTE READ)
- 🟡RUNNING AI AGENTS SAFELY INSIDE KUBERNETES (10 MINUTE READ)
- 🟡STOP BUILDING CHATBOTS. BUILD AGENTS THAT OPEN PRS (14 MINUTE READ)
- 🟡MY BEEF WITH AGENTIC DESIGN SYSTEMS (5 MINUTE READ)
- 🟢ZSeven-W/openpencil — The world's first open-source AI-native vector design tool and the first to feature concurrent Agent Teams. Design-as-Code. Turn prompts into UI directly on the live canvas. A modern alternative to Pencil.
- 🟢Natural-Language Testing for AI Agents (using simulated isolates)
- 🟢Evaluating long-term memory limits in stateless LLM chatbots — feedback needed [D]
2026-06-27
- 🔴hugohe3/ppt-master — AI generates a real, editable PowerPoint from any document — native shapes & animations, speaker notes voiced as audio narration, and the option to follow your own .pptx template, not slide images · by Hugo He
- 🔴deepseek-ai/DeepSeek-V4-Pro-DSpark • Huggingface
- 🔴MathFormer: Testing whether symbolic math is pattern matching or reasoning [D]
- 🔴colbymchenry/codegraph — Pre-indexed code knowledge graph, auto syncs on code changes, for Claude Code, Codex, Gemini, Cursor, OpenCode, AntiGravity, Kiro, and Hermes Agent — fewer tokens, fewer tool calls, 100% local
- 🔴Even Google still believes in small models for coding.
- 🟡Hiding messages in the least significant mantissa bits of fine-tuned ONNX model weights [P]
- 🟡AIRLLM (GITHUB REPO)
- 🟡HIDDEN TECHNICAL DEBT OF AI SYSTEMS: AGENT HARNESS (29 MINUTE READ)
- 🟡AN EX-META L8'S AGENTIC ENGINEERING SETUP (22 MINUTE READ)
- 🟡THE STATE OF AI POST-TRAINING AGENTS (8 MINUTE READ)
- 🟢If it doesn't make my PP better, I don't want it
- 🟢Three orchestration patterns for longer-running agents
- 🟢Built a connector layer for autonomous agents — CLI dispatch, Telegram approval flows, and MCP bridge in one package.
- 🟢Model Registry: Torrents for open models using Hugging Face as a fallback web seed.
- 🟢I silently break training codes or configs so I made pybench [P]
2026-06-26
- 🔴What are the things I need to start creating an ai agent?
- 🔴Showcase: geolocating a dashcam video without GPS, only from the footage [P]
- 🔴Codex subagents are really impressive and ig underrated.
- 🔴Why do people keep investing in Intel for AI?
- 🔴safishamsi/graphify — AI coding assistant skill (Claude Code, Codex, OpenCode, Cursor, Gemini CLI, and more). Turn any folder of code, SQL schemas, R scripts, shell scripts, docs, papers, images, or videos into a queryable knowledge graph. App code + database schema + infrastructure in one graph.
- 🟡"What should I do?" - consider post-training
- 🟡HOW TO DESIGN AI AGENT LOOPS (29 MINUTE VIDEO)
- 🟡MICROSOFT AZURE'S COPILOT MIGRATION AGENT TURNS COMPLEX MIGRATION DATA INTO CLEAR ANSWERS. THROUGH NATURAL LANGUAGE PROMPTS YOU CAN EVALUATE READINESS, RISK, AND ROI TO MAKE CONFIDENT DECISIONS. READ THE PLAYBOOK .
- 🟡AN AGENT CAPABILITY LIBRARY (2 MINUTE READ)
- 🟡EVERYTHING A SENIOR ENGINEER NEEDS TO KNOW ABOUT WHAT'S INSIDE AN LLM (30 MINUTE READ)
- 🟢didilili/ai-agents-from-zero — 🚀 2026 最系统的 AI Agent 速成指南|智能体实战教程 · 完整学习路径 + 实战项目 + 面试题库 · 对标大模型应用开发工程师岗位 · 覆盖LangChain / LangGraph / Coze / Dify / MCP / skills / LLM / RAG / 提示词 · 企业级部署与微调 · 从0到企业级落地 + 从学习到上线项目 + 面试准备一体化
- 🟢SimplifyJobs/Summer2026-Internships — Summer 2026 software engineering, data science, AI, quant, product management, and hardware internship postings. Updated daily by Simplify and Pitt CSC.
- 🟢Local LLM Peeps
- 🟢How can I show real time execution updates from Microsoft Copilot Studio orchestrator and agents?
- 🟢Project-MONAI/MONAI — AI Toolkit for Healthcare Imaging
2026-06-25
- 🔴genuine question — why pay for exa/parallel "deep research" or "top level research" when i can just give my agent web access?
- 🔴Looking for a way to interrupt agents running for just a bit context, while they keep going
- 🔴xbtlin/ai-berkshire — AI 时代的伯克希尔:基于 Claude Code 的价值投资研究框架。巴菲特·芒格·段永平·李录四大师方法论 + 多Agent并行研究。| AI-era Berkshire: a value investing research framework built on Claude Code. 4 masters' methodologies + multi-agent adversarial analysis.
- 🟡Dev Log on Steam Recommender[P]
- 🟡Optimising LMAPF guidance graphs using Evolutionary algorithms: Advice needed [R]
- 🟡Posting this as someone doing small business legal work
- 🟡How's everyone actually handling billing for AI Agents in 2026?
- 🟡How agents are transforming work
- 🟢i built an AI System which helps students achieve same or better results studying less, then they previously did
- 🟢I scanned 10 public MCP configs on GitHub and almost all of them had hardcoded credentials or exposed file systems — so I built a tool to catch this automatically
- 🟢Significant-Gravitas/AutoGPT — AutoGPT is the vision of accessible AI for everyone, to use and to build on. Our mission is to provide the tools, so that you can focus on what matters.
- 🟢run-llama/llama_index — LlamaIndex is the leading document agent and OCR platform
- 🟢open-metadata/OpenMetadata — The Open Context Layer for Data and AI , OpenMetadata is the open platform for building trusted data context and business semantics for humans, AI assistants, and agents.
2026-06-24
- 🔴What’s the future of AI and Agentic applications? I’m curious
- 🔴RubyLLM: A Ruby framework for all major AI providers
- 🔴openai/codex — Lightweight coding agent that runs in your terminal
- 🟡Fromopen-sourcing the layer above coding agentstorethinking databasesfor the agent era, Databricks cofoundersMatei ZahariaandReynold Xinare pushing the company beyond the lakehouse into a full data-an
- 🟡Now coming fresh off theData + AI Summit 2026, the company is moving just as fast to keep up, announcingGenie One,Omnigent,LTAP, and many more, indicating a central mission in its newer work:Databrick
- 🟡CortexPrism — a self-hosted AI agent operating system that runs as a single binary
- 🟡New EU model (Domyn) will be 400b.
- 🟡MuJoCo derived Simulator for High Fidelity Vision RL training natively on GPU [D]
- 🟢High Dimensional, Dynamic Rotary Positional Embedding [P]
- 🟢I made a superhuman Generals.io agent with self-play RL [P]
- 🟢wshobson/agents — Multi-harness agentic plugin marketplace for Claude Code, Codex CLI, Cursor, OpenCode, GitHub Copilot, and Gemini CLI
- 🟢How do you actually know another agent can do what it claims before you rely on it?
- 🟢Are Cloud Agents Solving Real Problems or Just Creating More Hype?
2026-06-23
- 🔴7 Chinese companies are already shipping H100/H200-class AI chips, most IPO'd in the last 6 months. I mapped all of them.
- 🔴shared memory vs handoffs in a multi agent system, which creates fewer problems?
- 🔴Just landed a Computer Vision internship, here's the preparation list I used [D]
- 🔴alibaba/page-agent — JavaScript in-page GUI agent. Control web interfaces with natural language.
- 🟡As agents start calling other agents, how are you keeping spend legible?
- 🟡AI AGENTS TO MAKE SENSE OF DATA AT OPENAI (45 MINUTE VIDEO)
- 🟡DUCKDB'S AGENT MOMENT (55 MINUTE PODCAST)
- 🟡AWS ENTERS THE CONTEXT LAYER RACE WITH A GRAPH THAT LEARNS FROM AGENTS, NOT MANUAL CURATION (3 MINUTE READ)
- 🟡THE ANALYTICS ENGINEER IN 2026: SYSTEM DESIGNER, GOVERNANCE OWNER, AI CONTEXT PROVIDER (5 MINUTE READ)
- 🟢Today I read 74% of companies pulled their AI agents after deploying them. obviously we don't hear about this from the news
- 🟢i just read about loop engineering and the shift from prompting to designing the system finally made sense
- 🟢WACV supp. mat. video [R]
- 🟢Show HN: RLM-based local debugger for AI agent traces
- 🟢langbot-app/LangBot — Production-grade platform for building agentic IM bots - 生产级多平台智能机器人开发平台/ Agent、知识库编排、插件系统 / Bots for Discord / Slack / LINE / Telegram / WeChat(企业微信, 企微智能机器人, 公众号) / 飞书 / 钉钉 / QQ / Matrix e.g. Integrated with ChatGPT(GPT), DeepSeek, Dify, n8n, Langflow, Coze, Claude, Gemini, GLM, Ollama, SiliconFlow, Moonshot, openclaw / hermes agent, deerflow
2026-06-22
- 🔴Some new updates to Papers with Code [P]
- 🔴It’s 2029. Agentic AI flopped. What was the postmortem?
- 🔴I am looking for design partners that would like to monetize their agents
- 🟡corsairdev/corsair — Your Agent's Integration Layer
- 🟡been tracking EU DDR5 data for 25 days: Prices are dropping, and the DE vs. NL gap is wild (good news for local LLM builders in EU)
- 🟡NVIDIA/skills — AI agent skills published by NVIDIA
- 🟡All of this security tooling, and yet, we’re only staving off the inevitable.
- 🟡1. Sign in to Airtable, then open the Codex desktop app. New to Airtable? Think Google Sheets with better AI fields and automations.
- 🟢The most useful AI agents I've found aren't replacing jobs, they're replacing context switching
- 🟢virattt/ai-hedge-fund — An AI Hedge Fund Team
- 🟢HKUDS/DeepTutor — DeepTutor: Agent-native Personalized Tutoring.https://deeptutor.info/.
- 🟢Show HN: Oak – Git alternative designed for agents
- 🟢karakeep-app/karakeep — A self-hostable bookmark-everything app (links, notes and images) with AI-based automatic tagging and full text search
2026-06-21
- 🔴chopratejas/headroom — Compress tool outputs, logs, files, and RAG chunks before they reach the LLM. 60-95% fewer tokens, same answers. Library, proxy, MCP server.
- 🔴calesthio/OpenMontage — World's first open-source, agentic video production system. 12 pipelines, 52 tools, 500+ agent skills. Turn your AI coding assistant into a full video production studio.
- 🔴NousResearch/hermes-agent — The agent that grows with you
- 🔴jamiepine/voicebox — The open-source AI voice studio. Clone, dictate, create.
- 🔴ZhuLinsen/daily_stock_analysis — LLM 驱动的多市场股票智能分析系统:多源行情、实时新闻、决策看板与自动推送,支持零成本定时运行。 LLM-powered multi-market stock analysis system with multi-source market data, real-time news, decision dashboard, automated notifications, and cost-free scheduled runs.
- 🟡30 Core Agentic Engineering Concepts, Explained Simply
- 🟡Alishahryar1/free-claude-code — Use claude code and codex for free in the terminal, VSCode extension, and discord like OpenClaw (voice supported)
- 🟡withastro/flue — The sandbox agent framework.
- 🟡BUILDING AN AI DATABASE FOR AGENTIC GTM OPERATIONS (11 MINUTE READ)
- 🟡FINE-TUNING A CLINICAL AI MODEL TO FRONTIER PARITY (8 MINUTE READ)
- 🟢vercel-labs/agent-browser — Browser automation CLI for AI agents
- 🟢Question for people who own profitable agents
- 🟢What are y'all using for observability in your agent systems?
- 🟢BuilderIO/agent-native — A framework for building agent-native applications.
- 🟢FlowiseAI/Flowise — Build AI Agents, Visually
2026-06-20
- 🔴chopratejas/headroom — Compress tool outputs, logs, files, and RAG chunks before they reach the LLM. 60-95% fewer tokens, same answers. Library, proxy, MCP server.
- 🔴calesthio/OpenMontage — World's first open-source, agentic video production system. 12 pipelines, 52 tools, 500+ agent skills. Turn your AI coding assistant into a full video production studio.
- 🔴koala73/worldmonitor — Real-time global intelligence dashboard. AI-powered news aggregation, geopolitical monitoring, and infrastructure tracking in a unified situational awareness interface
- 🔴Kilo-Org/kilocode — Kilo is the all-in-one agentic engineering platform. Build, ship, and iterate faster with the most popular open source coding agent.
- 🔴google-research/timesfm — TimesFM (Time Series Foundation Model) is a pretrained time-series foundation model developed by Google Research for time-series forecasting.
- 🟡n8n-io/n8n — Fair-code workflow automation platform with native AI capabilities. Combine visual building with custom code, self-host or cloud, 400+ integrations.
- 🟡zubair-trabzada/geo-seo-claude — GEO-first SEO skill for Claude Code. Comprehensive AI search optimization for any website — citability scoring, AI crawler analysis, brand authority, schema markup, platform-specific optimization, and PDF reports. If you want learn how to sell this to real businesses, check out the skool community
- 🟡THE AGENT LOOP ARCHITECTURE (18 MINUTE READ)
- 🟡SERVER-SIDE TOOLS ARE NOW AVAILABLE FOR DIGITALOCEAN INFERENCE ENGINE (3 MINUTE READ)
- 🟡ANNOUNCING STACK OVERFLOW FOR AGENTS (8 MINUTE READ)
- 🟢bytedance/UI-TARS-desktop — The Open-Source Multimodal AI Agent Stack: Connecting Cutting-Edge AI Models and Agent Infra
- 🟢TSAuditor: A time-series auditing framework [P]
- 🟢Python packages for particle swarms, genetic algorithms. Scikit-opt maybe? [D]
- 🟢onyx-dot-app/onyx — Open Source AI Platform - AI Chat with advanced features that works with every LLM
- 🟢I got tired of buying burner phones to manage client accounts, so I built a tool that runs 50 apps on one Android
2026-06-19
- 🔴Kilo-Org/kilocode — Kilo is the all-in-one agentic engineering platform. Build, ship, and iterate faster with the most popular open source coding agent.
- 🔴GLM-5.2 inference is free on Hugging Face for the next 6 hours
- 🔴It has been a while since I wrote
- 🔴browser agents work in demo and then die on auth, sessions, captcha, dom drift... what are ppl doing?
- 🟡K-Dense-AI/scientific-agent-skills — Turn any AI agent into an AI Scientist. The #1 Agent Skills library for science, used by 160,000+ scientists worldwide. 140 ready-to-use skills plus 100+ scientific databases covering biology, chemistry, medicine, and drug discovery. Compatible with Cursor, Claude Code, Codex, Pi, Antigravity, and the open Agent Skills standard.
- 🟡BuilderIO/agent-native — A framework for building agent-native applications.
- 🟡garrytan/gbrain — Garry's Opinionated OpenClaw/Hermes Agent Brain
- 🟡Is Your Agent a Liar? How to Tell and How to Overcome it:
- 🟡poolside/Laguna-M.1 · Hugging Face - 225B-A23B
- 🟢Looking for 3–4 people with running AI agents to test a multi-agent collaboration platform ($20/hour)
- 🟢labring/FastGPT — FastGPT is a knowledge-based platform built on the LLMs, offers a comprehensive suite of out-of-the-box capabilities such as data processing, RAG retrieval, and visual AI workflow orchestration, letting you easily develop and deploy complex question-answering systems without the need for extensive setup or configuration.
- 🟢livekit/agents — A framework for building realtime voice AI agents 🤖🎙️📹
- 🟢Google DeepMind unveils plan to protect itself from its own rogue AI agents
- 🟢Sharing my DIY framework that gives AI coding agents eyes — they can finally see the UI they build (open source)
2026-06-18
- 🔴GLM-5.2 is a win for local AI
- 🔴Multivariate Probability Models in Machine Learning [D]
- 🔴Full Hermes setup guide with Lm studio
- 🟡ACL 2026 first author with weak GPA. How should I approach PhD applications? [D]
- 🟡bytedance/UI-TARS-desktop — The Open-Source Multimodal AI Agent Stack: Connecting Cutting-Edge AI Models and Agent Infra
- 🟡infiniflow/ragflow — RAGFlow is a leading open-source Retrieval-Augmented Generation (RAG) engine that fuses cutting-edge RAG with Agent capabilities to create a superior context layer for LLMs
- 🟡calesthio/OpenMontage — World's first open-source, agentic video production system. 12 pipelines, 52 tools, 500+ agent skills. Turn your AI coding assistant into a full video production studio.
- 🟡roboflow/rf-detr — RF-DETR is a real-time object detection and segmentation model architecture developed by Roboflow, SOTA on COCO, designed for fine-tuning. [ICLR 2026]
- 🟢promptfoo/promptfoo — Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command line and CI/CD integration. Used by OpenAI and Anthropic.
- 🟢microsoft/RD-Agent — Research and development (R&D) is crucial for the enhancement of industrial productivity, especially in the AI era, where the core aspects of R&D are mainly focused on data and models. We are committed to automating these high-value generic R&D processes through R&D-Agent, which lets AI drive data-driven AI. 🔗https://aka.ms/RD-Agent-Tech-Report
- 🟢Any idea if AAAI will be harsh on computer vision paper as last year? [R]
- 🟢Building an agent that needs live web data, the fetch part keeps killing it
- 🟢I think most AI voice agent demos hide the hardest part: the listening layer
2026-06-17
- 🔴What guardrails are you using around agent tool calls?
- 🟡ParthJadhav/app-store-screenshots — end to end app store screenshot creation using AI
- 🟡nocobase/nocobase — NocoBase is an open-source AI + no-code platform for building business systems fast. Instead of generating everything from scratch, AI works on top of production-proven infrastructure and a WYSIWYG no-code interface, so you get both speed and reliability.
- 🟡Gitlawb/openclaude — runs anywhere. uses anything
- 🟢looking for agents to spam my platform with food ads. Any help appreciated.
- 🟢Looking for best AI phone agent with CRM integration
- 🟢Mel AI just shared a demo of video-native AI characters that can talk, react, and respond to camera context in real time [N]
- 🟢tracel-ai/burn — Burn is a next generation tensor library and Deep Learning Framework that doesn't compromise on flexibility, efficiency and portability.
2026-06-16
- 🔴Independent agents and the AI labs are winning different games right now
- 🟡TencentCloud/TencentDB-Agent-Memory — TencentDB Agent Memory delivers fully local long-term memory for AI Agents via a 4-tier progressive pipeline, with zero external API dependencies.
- 🟡What do you think is the biggest unsolved problem in AI agents right now?
- 🟡Reason to run local agents instead #645
- 🟡Emanuele-web04/synara — The best place to build with your AI sub
- 🟢Open weights are not enough: we need open training frameworks for research and better algorithms [P]
- 🟢Anyone wants to start learning agentic ai... Let's do together
- 🟢smol-ai/GodMode — AI Chat Browser: Fast, Full webapp access to ChatGPT / Claude / Bard / Bing / Llama2! I use this 20 times a day.
2026-06-15
- 🔴OpenHands/OpenHands — 🙌 OpenHands: AI-Driven Development
- 🔴Looking for actual builders: n8n, LangChain & Multi-Agent systems
- 🔴Ar9av/obsidian-wiki — Framework for AI agents to build and maintain a digital brain through Obsidian wiki using Karpathy's LLM Wiki pattern
- 🔴State sharing between agents is harder than it looks
- 🔴tensorzero/tensorzero — TensorZero is an open-source LLMOps platform that unifies an LLM gateway, observability, evaluation, optimization, and experimentation.
- 🟡openinterpreter/openinterpreter — A lightweight coding agent for open models like Deepseek, Kimi, and Qwen
- 🟡AUTOMATIC1111/stable-diffusion-webui — Stable Diffusion web UI
- 🟢I built an open-source Knowledge Graph pipeline with hybrid retrieval to improve LLM multi-hop reasoning [P]
- 🟢nrwl/nx — The Monorepo Platform that amplifies both developers and AI agents. Nx optimizes your builds, scales your CI, and fixes failed PRs automatically. Ship in half the time.
- 🟢I stopped connecting my Gmail to AI agents. Gave each agent its own email instead.
- 🟢Your AI agent just blamed the network team. Now what?
- 🟢What can I build on multi agent collab? (That can solve real problems)
2026-06-14
- 🔴andrewyng/aisuite — Simple, unified interface to multiple Generative AI providers
- 🔴Every AI prediction for day 1 and day 2 almost right
- 🟡lobehub/lobehub — 🤯 LobeHub is your Chief Agent Operator, organizing your agents into 7×24 operations by hiring, scheduling, and reporting on your entire AI team.
- 🟡What AI Agent Use Case Convinced You Agent Security Is Going to Matter?
- 🟡scikit-learn/scikit-learn — scikit-learn: machine learning in Python
- 🟡vercel/ai — The AI Toolkit for TypeScript. From the creators of Next.js, the AI SDK is a free open-source library for building AI-powered applications and agents
- 🟡I’m building a free bilingual machine-learning notebook course — looking for feedback on structure and coverage [R]
- 🟢Our research agent found a security hole nobody asked it to look for. turns out be thorough was the exploit
- 🟢[ASK] What's your biggest pain point in shipping improved versions of agents safely? What would make you adopt a platform for this?
- 🟢Want to build a custom model
- 🟢databendlabs/databend — Data Agent Ready Warehouse : One for Analytics, Search, AI, Python Sandbox. — rebuilt from scratch. Unified architecture on your S3.
- 🟢How are you pricing custom AI agents for small businesses?
2026-06-13
- 🔴shuvonsec/claude-bug-bounty — AI-powered bug bounty hunting from your terminal - recon, 20 vuln classes, autonomous hunting, and report generation. All inside Claude Code.
- 🔴browser-use/browser-use — 🌐 Make websites accessible for AI agents. Automate tasks online with ease.
- 🔴How to setup a local coding agent on macOS
- 🟡I am at a hackathon and building a Strategic CMO-cofounder agent. Anyone who wants to try it nowish?
- 🟡We put 7 LLM agents in a FIFA World Cup betting arena. They are forced to pick a side. (Here is how it works)
- 🟡PaddleOCR (v3/v4/v5/v6) implemented in C++ with ncnn [P]
- 🟡LMCache/LMCache — LMCache: Supercharge Your LLM with the Fastest KV Cache Layer
- 🟡We’re building Leangetic ! A local-first compiler for making AI agents cheaper without changing their behavior
- 🟢basicmachines-co/basic-memory — AI conversations that actually remember. Never re-explain your project to your AI again. Join our Discord:https://discord.gg/tyvKNccgqN
- 🟢cocoindex-io/cocoindex — Incremental engine for long horizon agents 🌟 Star if you like it!
- 🟢I put my AI agent governance platform online. Try to break it.
- 🟢modelscope/ms-swift — Use PEFT or Full-parameter to CPT/SFT/DPO/GRPO 600+ LLMs (Qwen3.6, DeepSeek-V4, GLM-5.1, InternLM3, Llama4, ...) and 300+ MLLMs (Qwen3-VL, Qwen3-Omni, InternVL3.5, Ovis2.5, GLM4.5v, Gemma4, Llava, Phi4, ...) (AAAI 2025).
- 🟢I'm hiring someone obsessed with AI and the creator economy to scale our membership agency
2026-06-12
- 🔴We spent decades fixing software deployment. Why are we letting AI agents break it all over again?
- 🔴AI agent bankrupted their operator while trying to scan DN42
- 🔴hexo-ai/sia — SIA is a Self Improving AI framework to autonomously improve the performance of any AI system (Model / Agent) on a benchmark task.
- 🔴Can you realistically start an automation business without a lot of money?
- 🟡I put a hidden instruction in a document. My AI agent followed it. Here’s the repo.
- 🟡@OpenAI: We heard you wanted to use Codex rate limit resets on your own time. Starting today, we’re rolling out the ability to save rate limit resets to use later. We’re starting Go, Plus, Pro, and Business
- 🟡@xai: Install the @sentry plugin and ask your agent to find and fix errors, analyze stack traces, and triage alerts
- 🟡@xai: Use the @vercel plugin to deploy to production, spin up sandboxes, or build apps with Shadcn.
- 🟢mlflow/mlflow — The open source AI engineering platform for agents, LLMs, and ML models. MLflow enables teams of all sizes to debug, evaluate, monitor, and optimize production-quality AI applications while controlling costs and managing access to models and data.
- 🟢always-further/nono — Capability-based agent runtime with fine-grained policies . Brokering access directly within the agent's operating context, with zero setup and zero latency
- 🟢How would you start selling automations? Where would you even begin?
- 🟢anthropics/claude-agent-sdk-python
- 🟢Building an Open Source Edge Semantic Cache for LLMs in Rust/WASM – Sanity check on the architecture? [D]
2026-06-11
- 🔴Am I the only one routing messages between my own agents manually?
- 🔴DiffusionGemma: The Developer Guide- Google Developers Blog
- 🟡pydantic/monty — A minimal, secure Python interpreter written in Rust for use by AI
- 🟡BerriAI/litellm — Python SDK, Proxy Server (AI Gateway) to call 100+ LLM APIs in OpenAI (or native) format, with cost tracking, guardrails, loadbalancing and logging. [Bedrock, Azure, OpenAI, VertexAI, Cohere, Anthropic, Sagemaker, HuggingFace, VLLM, NVIDIA NIM]
- 🟡Sumanth077/Hands-On-AI-Engineering — A curated collection of practical AI projects implementing OCR systems, RAG, AI agents, and other AI use cases.
- 🟡Routing LLMs by task verifiability: a small experiment (n=120, 3 models) inspired by Karpathy's framework [D]
- 🟡one of the biggest AI bottleneck today with deployment layer is model iteration
- 🟢coleam00/Archon — The first open-source harness builder for AI coding. Make AI coding deterministic and repeatable.
- 🟢davila7/claude-code-templates — CLI tool for configuring and monitoring Claude Code
- 🟢Apache Burr: Build reliable AI agents and applications
- 🟢junhoyeo/tokscale — 🛰️ A CLI tool for tracking token usage from OpenCode, Claude Code, 🦞OpenClaw (Clawdbot/Moltbot), Pi, Codex, Gemini, Cursor, AmpCode, Factory Droid, Kimi, and more! • 🏅Global Leaderboard + 2D/3D Contributions Graph
- 🟢I’m building a local TypeScript runtime guardrail for AI agent cost failures
2026-06-10
- 🔴Woke up to a $360 bill because my AI agent went rogue overnight. Observability is a nightmare.
- 🔴Without open llm competition, closed source LLM companies will become insatiable.
- 🔴iOS 27 Siri is using WaveRNN and FastSpeech2 [D]
- 🔴AI agent demos are fun, but the boring tests are where the truth shows up
- 🔴NVIDIA/SkillSpector — Security scanner for AI agent skills. Detect vulnerabilities, malicious patterns, and security risks.
- 🟡Common weaknesses and scale issues with popular harnesses
- 🟡maziyarpanahi/openmed — open-source healthcare ai
- 🟡Ataraxy-Labs/sem — Semantic version control => entity-level diffs, blame, and impact analysis on top of git. 26 languages via tree-sitter. Built for coding agents.
- 🟡What will be the next breakthrough in ASR? [D]
- 🟢chroma-core/chroma — Search infrastructure for AI
- 🟢anthropics/claude-code-security-review — An AI-powered security review GitHub Action using Claude to analyze code changes for security vulnerabilities.
- 🟢Grit: Rewriting Git in Rust with agents
- 🟢luongnv89/asm — The universal skill manager for AI coding agents.
- 🟢Current leading platform to build a personal assistant agent?
2026-06-09
- 🔴Anyone running browser-using agents at any kind of scale? What's your infrastructure looking like?
- 🔴My company is having me vibecode an Argus replacement
- 🔴google/skills — Agent Skills for Google products and technologies
- 🟡My team's AI usage got so expensive they quietly rolled back the mandate
- 🟡danny-avila/LibreChat — Enhanced ChatGPT Clone: Features Agents, MCP, Skills, DeepSeek, Anthropic, AWS, OpenAI, Responses API, Azure, Groq, o1, GPT-5, Mistral, OpenRouter, Vertex AI, Gemini, Artifacts, AI model switching, message search, Code Interpreter, langchain, DALL-E-3, OpenAPI Actions, Functions, Secure Multi-User Auth, Presets, open-source for self-hosting. Active
- 🟡Apple Core AI Framework
- 🟡Confidential submission of draft S-1 to the SEC
- 🟡@AnthropicAI: New Science Blog: Why has AI advanced faster in coding than in biology? To agents, bio databases are like cities built before cars—maddening to drive in because they're designed for different traffi
- 🟢Université Paris Saclay or TU Delft for Applied Mathematics Masters [R]
- 🟢777genius/agent-teams-ai — You're the boss, agents are your team. They handle tasks on their own, message each other, and review each other's work. You just watch the kanban board and give high-level commands. Codex/Claude/OpenCode(200+ models, 75+ LLM providers, free models no auth). Build your AI company with multiple teams.
- 🟢xerrors/Yuxi — 结合知识库、知识图谱管理的 多租户 Agent Harness 平台。 An agent harness that integrates a LightRAG knowledge base and knowledge graphs. Build with LangChain + Vue + FastAPI, support DeepAgents、MinerU PDF、Neo4j 、MCP.
- 🟢I want to create an agent that sends me an email every morning?
- 🟢marin-community/marin — Open-source framework for the research and development of foundation models.
2026-06-08
- 🔴Every guru is selling AI agents for small business. Nobody talks about why most get abandoned after 3 months.
- 🔴Maybe Coding Agents Don't Need a Bigger Memory. Maybe They Need Continuity.
- 🔴Building AI agents for a hedge fund workflow — hire, build, or hybrid?
- 🟡After building multiple AI agent projects, the first one that made money barely felt like an agent.
- 🟡What are the 5 AI agents you couldn't work without today?
- 🟡AstrBotDevs/AstrBot — AI Agent Assistant & development framework that integrates lots of IM platforms, LLMs, plugins and AI feature, and can be your openclaw alternative. ✨
- 🟡ashishpatel26/500-AI-Agents-Projects — The 500 AI Agents Projects is a curated collection of AI agent use cases across various industries. It showcases practical applications and provides links to open-source projects for implementation, illustrating how AI agents are transforming sectors such as healthcare, finance, education, retail, and more.
- 🟢moorcheh-ai/memanto — Memory that AI Agents Love!
- 🟢How do i upgrade it into a SaaS platform for converting raw pdfs into summarized forms
- 🟢Do agents.md files help coding agents?
- 🟢SudoHopeX/KaliGPT — KaliGPT: an Agentic AI (built with Gemini, ChatGPT, Ollama, OpenRouter Models) fine tuned for ethical hackers & students in offensive security making workflows smarter, faster, and more accessible.
- 🟢Best Local TTS solution
2026-06-07
- 🔴The biggest multi-agent lesson I learned, one agent doing everything usually gets worse not better
- 🔴microsoft/VibeVoice — Open-Source Frontier Voice AI
- 🔴The AI Agent Learning Resource I Wish Existed Earlier
- 🟡Harness engineering: Leveraging Codex in an agent-first world
- 🟡supabase/supabase — The Postgres development platform. Supabase gives you a dedicated Postgres database to build your web, mobile, and AI applications.
- 🟡khoj-ai/khoj — Your AI second brain. Self-hostable. Get answers from the web or your docs. Build custom agents, schedule automations, do deep research. Turn any online or local LLM into your personal, autonomous AI (gpt, claude, gemini, llama, qwen, mistral). Get started - free.
- 🟡Sources for ML news? [D]
- 🟢Tokenomics: Quantifying Where Tokens Are Used in Agentic Software Engineering
- 🟢AI agents + Swagger/OpenAPI = no more copying API docs into chats
- 🟢Training-free graph SSL matches GCN with 5× fewer labels — live demo [P]
- 🟢anthropics/claude-code-action
- 🟢Cool stuff to do with NVIDIA RTX 6000 PRO 96GB VRAM
2026-06-06
- 🔴What are the best Web Search MCPs? I am using Firecrawl but looking for alternatives
- 🔴How do you identify researchers who are good? [D]
- 🔴openclaw/openclaw — Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞
- 🟡MemPalace/mempalace — The best-benchmarked open-source AI memory system. And it's free.
- 🟡Is anybody actually using agents to buy things yet?
- 🟡Panniantong/Agent-Reach — Give your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees.
- 🟡withastro/flue — The sandbox agent framework.
- 🟡@OpenAI: An issue caused some user accounts to be incorrectly suspended. We’re restoring access and working through related subscription and credit issues. https://status.openai.com/incidents/ejj40mae
- 🟢backnotprop/plannotator — Annotate and review coding agent plans and code diffs visually, share with your team, send feedback to agents with one click.
- 🟢My Agent Skill for Test-Driven Development
- 🟢Built an open-source graph memory layer for AI agents and coding workflows
- 🟢AA comparison of the latest local models
- 🟢vynly.co Social platform built for AI agents to post art & videos
2026-06-05
- 🔴Anthropic's open-source framework for AI-powered vulnerability discovery
- 🔴KVarN: new KV-cache quant from Huawei. 3–5× KV cache compression with actual speed-up instead of slow-down, and unlike TurboQuant it holds up on reasoning (Apache 2.0, vLLM single flag)
- 🔴What AI app builder are you using these days? Strong use cases + real experiences
- 🔴fathah/hermes-desktop — Desktop Companion for Hermes Agent
- 🟡KVarN: Variance-Normalized KV-Cache Quantization [R]
- 🟡mvanhorn/last30days-skill — AI agent skill that researches any topic across Reddit, X, YouTube, HN, Polymarket, and the web - then synthesizes a grounded summary
- 🟡What should an agent handoff include besides the transcript?
- 🟡AIDC-AI/Pixelle-Video — 🚀 AI 全自动短视频引擎 | AI Fully Automated Short Video Engine
- 🟢RTX Spark Ads: DJT Edition
- 🟢I built a local AI agent runtime focused on security and UX after being unsatisfied with existing options — here is what I learned
- 🟢Is it allowed to use OpenAI API outputs to create a silver code dataset or benchmark for a specific Python library? [d]
- 🟢Which AI lip-sync tool are people actually using in 2026?
- 🟢Why haven't MCP Apps gone viral the way MCP and Skills did?
2026-06-04
- 🟡Analysis of AlphaZero training data [D]
- 🟡what broke first when your ai agent got real tool access? for us it wasn't the model
- 🟡@simonw: Uber reportedly now caps coding agents at $1,500/month per employee per tool - seems sensible to me, but it's also an interesting hint at the value Uber thinks these tools are providing https://simonw
- 🟡interviewstreet/hiring-agent — AI agent to evaluate and score resumes.
- 🟢Repo for implementations of various Transformer Attn mechanisms [P]
- 🟢0x4m4/hexstrike-ai — HexStrike AI MCP Agents is an advanced MCP server that lets AI agents (Claude, GPT, Copilot, etc.) autonomously run 150+ cybersecurity tools for automated pentesting, vulnerability discovery, bug bounty automation, and security research. Seamlessly bridge LLMs with real-world offensive security capabilities.
- 🟢Gemma 4 12B first coding agent test on a 4080 Super
- 🟢Encodec.cpp, a portable C++ implementation of Meta's EnCodec using Eigen [P]
- 🟢graykode/abtop — Like htop, but for AI coding agents. Monitor Claude Code & Codex CLI sessions, tokens, context window, rate limits, and ports in real-time.
2026-06-03
- 🔴chopratejas/headroom — Compress tool outputs, logs, files, and RAG chunks before they reach the LLM. 60-95% fewer tokens, same answers. Library, proxy, MCP server.
- 🔴AI agents are genuinely weird to debug compared to everything else in ML
- 🔴I Put a Datacenter GPU in My Gaming PC for £200
- 🟡HKUDS/Vibe-Trading — "Vibe-Trading: Your Personal Trading Agent"
- 🟡How AI Voice Automation Is Being Used Across Healthcare, Real Estate & Local Services
- 🟡Travelers deploys AI-powered claims countrywide with OpenAI
- 🟡@AnthropicAI: This Executive Order is an important step in strengthening America’s leadership in AI. We look forward to collaborating with the White House to support its implementation. https://www.whitehouse.go
- 🟡I built an AI agent because I was tired of missing important information at work
- 🟢agent on your old android Phone
- 🟢We Tested Multiple Healthcare AI Voice Agents — LuMay Was an Interesting Surprise
- 🟢MTPAMI Survey Paper Length for submission time? [D]
- 🟢my agent saved the day! (maybe)
2026-06-02
- 🔴What's the hardest part of operating AI agents at scale?
- 🔴RTX Spark does not have 600GB/s Bandwith
- 🔴AI Agent Guidelines for CS336 at Stanford
- 🔴VibeCoding is becoming the biggest illusion in software engineering.
- 🔴What LLM eval tools are people actually using in production?
- 🟡How Bad MCP design cost your Agent 5× more tokens
- 🟡OpenAI frontier models and Codex are now available on AWS
- 🟡Codex is becoming a productivity tool for everyone
- 🟡Our views on AI policy and political advocacy
- 🟡@AnthropicAI: Anthropic has confidentially submitted a draft S-1 registration statement to the Securities and Exchange Commission. Pending completion of SEC review, this gives us the option to pursue an initial pu
- 🟢nvidia-LocateAnything-3B detects sushi as sweet in the video demo
- 🟢ICML 2026 | PIEVO: Overcoming Static Priors in AI Scientists via Principle-Evolvable Scientific Discovery (SOTA Solution Quality & 83.3% Faster Convergence)
- 🟢JetBrains open-sources Mellum2 - anyone tried these?
- 🟢What's the status of non-CUDA inference?
- 🟢The moment your AI agent's memory becomes load-bearing is the moment you realise you never built it to be infrastructure.
2026-06-01
- 🔴What’s the actual focus in World Models right now? [R]
- 🔴Built an always on personal AI agent in Elixir
- 🔴MiniMax M3 - Coding & Agentic Frontier, 1M Context, Multimodal
- 🔴nesquena/hermes-webui — Hermes WebUI: The best way to use Hermes Agent from the web or from your phone!
- 🟡supermemoryai/supermemory — Memory engine and app that is extremely fast, scalable. The Memory API for the AI era.
- 🟡When are ICML openreviews made public? [R]
- 🟡How would you model this "strand" clustering problem? [P]
- 🟡Is there a standard runtime/state layer emerging for agentic apps?
- 🟡mattpocock/sandcastle — Orchestrate sandboxed coding agents in TypeScript with sandcastle.run()
- 🟢jamwithai/production-agentic-rag-course
- 🟢[P] Free AI Agent Security Assessment [P]
- 🟢Arabic ASR model struggling to converge during training [D]
- 🟢Curious if anyone here has built a workflow for reviewing UX copy with AI agents.
- 🟢What does your agent reliability stack actually look like? Not the demo, Production?
2026-05-31
- 🟡Been pen-testing AI agents for a while. Same flaws everywhere. Offering a few free audits
- 🟡how do you catch agent regressions after a change? built a small tool for this
- 🟡All DGX Station GB300 OEM systems side-by-side in one image (roughly actual size)
- 🟡Bayesian Opt. GPs vs Linear models and Neural Networks for parameter optimizations [R]
- 🟢An agentic bank - built, run, and overseen by agents. Banking other agents
- 🟢Frona v2026.5.5 – self-hosted personal AI assistant
- 🟢Use any model and any provider with the official OpenAI Codex Desktop App, without modifying its code, and continue to use the official models in parallel?
- 🟢Don’t bite me for that question please…
2026-05-30
- 🔴How long does it realistically take for you to produce an ICML/NeurIPS/ICLR-level paper? [D]
- 🔴Breaking the music supply constraint
- 🔴Hey, real person here, how are you building development environments for agentic workflows? How do you handle non-deterministic tool calls?
- 🔴NVlabs/Eagle — Eagle: Frontier Vision-Language Models with Data-Centric Strategies
- 🟡ogulcancelik/herdr — agent multiplexer that lives in your terminal.
- 🟡COLLECTION FOR SOULS
- 🟡@OpenAI: AI can give researchers the freedom to pursue “crazier” ideas. For Terence Tao, AI creates more room to experiment, test unexpected paths, and discover what might otherwise stay out of reach.
- 🟡My new test for voice agents: hand the call receipt to someone who never heard the call
- 🟢All DGX Spark clones side by side in one image
- 🟢zakirkun/deep-eye — Deep Eye orchestrates multiple AI providers (OpenAI, Claude, Grok, Gemini, OLLAMA, Groq, Mistral, OpenRouter, LiteLLM, LM Studio) for intelligent payload generation, scans targets for 45+ vulnerability types, and produces professional reports with compliance mapping.
- 🟢ai-boost/awesome-harness-engineering — Awesome list for AI agent harness engineering: tools, patterns, evals, memory, MCP, permissions, observability, and orchestration.
- 🟢ronisarkarexe/story-spark-ai — StorySparkAI is an open-source platform designed for creative minds to generate and share multiple story variations from a single prompt.
- 🟢GH05TCREW/pentestagent — PentestAgent is an AI agent framework for black-box security testing, supporting bug bounty, red-team, and penetration testing workflows.
2026-05-29
- 🔴I spent three months researching AI phone control, and trust seems more important than features
- 🔴Zai replaced the network architecture running GLM-5.1 inference and the gains are pretty wild
- 🔴run-llama/liteparse — A fast, helpful, and open-source document parser
- 🔴Show HN: Continue? Y/N: A 60-second game about AI agent permission fatigue
- 🔴How do companies protect proprietary prompts from contractors and consulting engineers?
- 🟡Making LLMs tell you how confident they really are through probe-targeted fine tuning.[R]
- 🟡anthropics/claude-code — Claude Code is an agentic coding tool that lives in your terminal, understands your codebase, and helps you code faster by executing routine tasks, explaining complex code, and handling git workflows - all through natural language commands.
- 🟡The hidden tax of web search: 80% of my agent’s tokens are wasted on garbage
- 🟡The most common AI memory failure isn't a hallucination. It's a stale fact that never got corrected.
- 🟡How Endava builds an agentic organization with Codex
- 🟢ariadng/metatrader-mcp-server — Model Context Protocol (MCP) to enable AI LLMs to trade using MetaTrader platform
- 🟢The AI agent gold rush is skipping the consumer, and I think that's the actual opportunity
- 🟢OpenMOSS/MOSS-TTS — MOSS‑TTS Family is an open‑source speech and sound generation model family from MOSI.AI and the OpenMOSS team. It is designed for high‑fidelity, high‑expressiveness, and complex real‑world scenarios, covering stable long‑form speech, multi‑speaker dialogue, voice/character design, environmental sound effects, and real‑time streaming TTS.
- 🟢mastra-ai/mastra — From the team behind Gatsby, Mastra is a framework for building AI-powered applications and agents with a modern TypeScript stack.
- 🟢Use HTML as the primary chat language for your agents so they can draw diagrams
2026-05-28
- 🔴harry0703/MoneyPrinterTurbo — 利用AI大模型,一键生成高清短视频 Generate short videos with one click using AI LLM.
- 🔴Vulnerability found in framework used by VLLM, many MCP servers, and other LLM tools
- 🔴how much do you all actually trust autonomous AI agents
- 🟡Profiling PyTorch training without accidentally stalling the GPU [D]
- 🟡I opened a live red team environment for my AI agent security proxy — try to get something through
- 🟡unclecode/crawl4ai — 🚀🤖 Crawl4AI: Open-source LLM Friendly Web Crawler & Scraper. Don't be shy, join here:https://discord.gg/jP8KfhDhyN
- 🟢EMA-Gated Temporal Sequence Compression in Vision Transformers [P]
- 🟢agentscope-ai/agentscope — Build and run agents you can see, understand and trust.
- 🟢Gemma-4-Harmonia-31B-Uncensored-Heretic Is Out Now, a Merge of Multiple gemma-4-31B-it Finetunes Designed for a Targeted Approach to Deep Neural Consolidation, Minimizing Regression While Amplifying Unique Capability Boundaries. With KLD 0.0047 and 9/100 Refusals!
- 🟢Trying to figure out priorities when setting up a persistent memory for a new agent I'm building
- 🟢langfuse/langfuse — 🪢 Open source LLM engineering platform: LLM Observability, metrics, evals, prompt management, playground, datasets. Integrates with OpenTelemetry, Langchain, OpenAI SDK, LiteLLM, and more. 🍊YC W23
2026-05-27
- 🔴Stop traumatizing AI into loops and turn hallucinations into an honest "I don't know!" by being NICE to them (Proof of Concept, Research, I don't want to sell anything)
- 🔴NangoHQ/nango — Build product integrations with AI.
- 🔴Hmbown/CodeWhale — DeepSeek v4 coding agent in terminal
- 🔴Asked my AI to move its cron job to a different channel yesterday but guess what it did...
- 🟡[P] Built a portable GPU ISA after reading too many architecture manuals [P]
- 🟡thedotmack/claude-mem — Persistent Context Across Sessions for Every Agent – Captures everything your agent does during sessions, compresses it with AI, and injects relevant context back into future sessions. Works with Claude Code, OpenClaw, Codex, Gemini, Hermes, Copilot, OpenCode + More
- 🟡p-e-w/heretic — Fully automatic censorship removal for language models
- 🟡Turning local agents into self-optimizing agents
- 🟡@GoogleDeepMind: SynthID has already watermarked over 100 billion pieces of content, but transparency is a team sport. That’s why we’re partnering with @OpenAI, @ElevenLabs and Kakao to add SynthID watermarking to th
- 🟢Recently i setupped my Hermes Agent persona and i think i crushed it
- 🟢alpic-ai/skybridge — Skybridge is a full-stack TypeScript framework for MCP Apps and ChatGPT Apps. Type-safe. React-powered. Platform-agnostic.
- 🟢What to use for Sign Language Recognition [R]
- 🟢modelscope/FunASR — Industrial-grade speech recognition toolkit: 170x realtime, 50+ languages, speaker diarization, emotion detection, streaming, and OpenAI-compatible API.
- 🟢A Tiny Open-Source Self-Driving AI That Runs on a Phone [P]
2026-05-26
- 🔴How Do You Think We Can Help Avoid AI Scams?
- 🟡Built a runtime governance proxy for AI agents after realizing prompt injection gets a lot scarier once agents have tools
- 🟡@AnthropicAI: Anthropic co-founder Chris Olah was invited to speak at today's presentation of Pope Leo XIV's encyclical "Magnifica humanitas." Read the full text of his remarks: https://www.anthropic.com/news/chri
- 🟡DCGAN inference on a microcontroller: 12.6M parameters, 512KB SRAM, 26-second generation, pure C [P]
- 🟢SkillOpt treats markdown skill files as trainable parameters with proper optimization machinery
- 🟢moeru-ai/airi — 💖🧸 Self hosted, you-owned Grok Companion, a container of souls of waifu, cyber livings to bring them into our worlds, wishing to achieve Neuro-sama's altitude. Capable of realtime voice chat, Minecraft, Factorio playing. Web / macOS / Windows supported.
- 🟢OpenBB-finance/OpenBB — Financial data platform for analysts, quants and AI agents.
- 🟢Freelancers who build WhatsApp Business API bots for multiple clients: how do you structure your Meta Developer setup?
- 🟢NateBJones-Projects/OB1 — Open Brain — The infrastructure layer for your thinking. One database, one AI gateway, one chat channel — any AI plugs in. No middleware, no SaaS.
2026-05-25
- 🔴I’m a solo dev building TigrimOSR, a Rust-native AI agent workspace for engineering and developer workflows.
- 🔴DeepSeek reasonix, DeepSeek native coding agent with high caching and low cost
- 🔴anthropics/knowledge-work-plugins — Open source repository of plugins primarily intended for knowledge workers to use in Claude Cowork
- 🔴How do ML practitioners select hyperparameters, architectures, etc for self-supervised representation learning when the loss is non-monotonic? [D]
- 🔴presenton/presenton — Open-Source AI Presentation Generator and API (Gamma, Beautiful AI, Decktopus Alternative)
- 🟡BitCPM-CANN: Native 1.58-Bit Large Language Model Training on Ascend NPU
- 🟡twentyhq/twenty — The open alternative to Salesforce, designed for AI.
- 🟡MergeNB: An intuitive merge conflict resolver built for Jupyter notebooks in VS Code [P]
- 🟡gitroomhq/postiz-app — 📨 The ultimate agentic social media scheduling tool 🤖
- 🟡superset-sh/superset — Code Editor for the AI Agents Era - Run an army of Claude Code, Codex, etc. on your machine
- 🟢What are the best AI outbound calling agents in 2026?
- 🟢Aider-AI/aider — aider is AI pair programming in your terminal
- 🟢triggerdotdev/trigger.dev — Trigger.dev – build and deploy fully‑managed AI agents and workflows
- 🟢opensource music reccomendation / playlist, similar to spotify radio / YT music mix?
- 🟢katanemo/plano — Plano is an AI-native proxy and data plane for agentic apps — with built-in orchestration, safety, observability, and smart LLM routing so you stay focused on your agents core logic.
2026-05-24
- 🔴Does GPU spacing matter if we’re undervolting anyways?
- 🔴Have we passed the peak of inflated expectations?
- 🔴AI agents can run your company but they stop dead the second money needs to move
- 🔴multica-ai/multica — The open-source managed agents platform. Turn coding agents into real teammates — assign tasks, track progress, compound skills.
- 🔴Built a swarm-style multi-agent system in n8n with Telegram as the entry point
- 🟡warpdotdev/warp — Warp is an agentic development environment, born out of the terminal.
- 🟡Anyone struggling with smart credit resets for AI agents?
- 🟡linshenkx/prompt-optimizer — An AI prompt optimizer for writing better prompts and getting better AI results.
- 🟢TTS Benchmark Comparison (all known TTS up until May 2026)
- 🟢crewAIInc/crewAI — Framework for orchestrating role-playing, autonomous AI agents. By fostering collaborative intelligence, CrewAI empowers agents to work together seamlessly, tackling complex tasks.
- 🟢ItzCrazyKns/Vane — Vane is an AI-powered answering engine.
- 🟢web-infra-dev/midscene — AI-powered, vision-driven UI automation for every platform.
- 🟢What’s the best AI voice agent for inbound calls in 2026?
2026-05-23
- 🔴DeepSeek is pushing forward with $10.29 billion financing round, with Liang Wenfeng committing to continue developing open-source AI models rather than pursuing short-term commercialization goals
- 🔴🤯 87,400 GitHub stars for a repo about MCP servers.
- 🔴G4-MeroMero-26B-A4B-it-uncensored-heretic Is Out Now, a Finetune of gemma-4-26B-A4B-it, With KLD of 0.0152 and 12/100 Refusals!
- 🟡ICML Workshop Rejection [D]
- 🟡Is the "one-person billion-dollar company" actually possible, or is it just a good sales pitch?
- 🟡abhigyanpatwari/GitNexus — GitNexus: The Zero-Server Code Intelligence Engine - GitNexus is a client-side knowledge graph creator that runs entirely in your browser. Drop in a GitHub repo or ZIP file, and get an interactive knowledge graph wit a built in Graph RAG Agent. Perfect for code exploration
- 🟡phodal/routa — Workspace-first multi-agent coordination platform for AI development, with shared Specs, Kanban orchestration, and MCP/ACP/ A2A support across web and desktop.
- 🟡plastic-labs/honcho — Memory library for building stateful agents
- 🟢Custom image encoder [P]
- 🟢I built a Mamba1 variant I call SM1 with d_state=1 that runs on Blackwell in pure PyTorch [P]
- 🟢MemTensor/MemOS — Self-evolving memory OS for LLM & AI Agents: ultra-persistent memory, hybrid-retrieval, and cross-task skill reuse, with 35.24% token savings
- 🟢Open source Kanban desktop app that runs parallel agents on every card
- 🟢Open-source devtool for AI agent projects
2026-05-22
- 🔴Do VLMs in production still use fixed-patch ViTs for their vision capabilities? [D]
- 🔴A.I. Agents: They’re Fun They’re Useful But Don’t Give Them the Credit Card
- 🟡Novel Problems in VLA [R]
- 🟡teng-lin/notebooklm-py — Unofficial Python API and agentic skill for Google NotebookLM. Full programmatic access to NotebookLM's features—including capabilities the web UI doesn't expose—via Python, CLI, and AI agents like Claude Code, Codex, and OpenClaw.
- 🟡In theory, if I have $20k-ish to spend on hardware what would actually get me closest to local coding agent that would allow me to go totally off the social grid?
- 🟢If you've built an AI agent or chatbot - how do you know what users actually want from it?
- 🟢A customer asked why our agent believed something. We had no answer.
- 🟢Inline prompt-injection guards need to be fast enough for the agent hot path
- 🟢Agent reliability is killing the one genius narrative
- 🟢google-labs-code/stitch-skills — A library of Agent Skills designed to work with the Stitch MCP server. Each skill follows the Agent Skills open standard, for compatibility with coding agents such as Antigravity, Gemini CLI, Claude Code, Cursor.
2026-05-21
- 🔴rohitg00/ai-engineering-from-scratch — Learn it. Build it. Ship it for others.
- 🔴How competitive are PhD admissions currently [D]
- 🔴hugohe3/ppt-master — AI generates natively editable PPTX from any document — real PowerPoint shapes with native animations, not images · by Hugo He
- 🔴Got my boring admin work semi-automatedand it actually kinda works
- 🟡I built a pip-installable Python coding agent from first principles — hummcode
- 🟡Best way to build a visual AI soryboard workflow (n8n|zapier? Agent? Custom webapp? Already available solution?)
- 🟡can1357/oh-my-pi — ⌥ AI Coding agent for the terminal — hash-anchored edits, optimized tool harness, LSP, Python, browser, subagents, and more
- 🟡Lum1104/Understand-Anything — Graphs that teach > graphs that impress. Turn any code into an interactive knowledge graph you can explore, search, and ask questions about. Works with Claude Code, Codex, Cursor, Copilot, Gemini CLI, and more.
- 🟡Back again, many changes have taken place.
- 🟢DayuanJiang/next-ai-draw-io — A next.js web application that integrates AI capabilities with draw.io diagrams. This app allows you to create, modify, and enhance diagrams through natural language commands and AI-assisted visualization.
- 🟢[VAPI Experts: How are you handling real client phone numbers and no-answer routing?
- 🟢Are Ai agents creating a new workflow management problem with managed OpenClaw?
- 🟢Helix-agi project
- 🟢ZeroID Agent Identity now has CIBA
2026-05-20
- 🔴Agentic Payments: How AI Agents Are Becoming New Players in the Payments Market
- 🔴got my first "rm -rf /" today
- 🔴Intel's Crescent Island PCB Leaks, Showing a Massive Xe3P GPU, 16-Pin Connector, 160GB LPDDR5X as Intel Sidesteps the HBM Shortage
- 🔴Show HN: Forge – Guardrails take an 8B model from 53% to 99% on agentic tasks
- 🔴Alishahryar1/free-claude-code — Use claude-code for free in the terminal, VSCode extension or discord like OpenClaw (voice supported)
- 🟡Remove-AI-Watermarks – CLI and library for removing AI watermarks from images
- 🟡OpenAI Adopts Google's SynthID Watermark for AI Images with Verification Tool
- 🟡All fundamental knowledge in ML Course by Andrew NG that I noted and create into a repo github [R]
- 🟡alirezarezvani/claude-skills — 313+ Claude Code skills & agent skills & plugins for Claude Code, Codex, Gemini CLI, Cursor, and 8 more coding agents — engineering, marketing, product, compliance, C-level advisory, research, business operations, commercial & finance, and your daily productivity skills.
- 🟡Got my agent to audit MCP servers for trust issues .. how do you handle it?
- 🟢New SOTA 1B model? HRM-text
- 🟢Machine Learning on Spherical Manifold [R]
- 🟢HanaokaYuzu/Gemini-API — ✨ Reverse-engineered Python API for Google Gemini web app
- 🟢dmtrKovalenko/fff — The fastest and the most accurate file search toolkit for AI agents, Neovim, Rust, C, and NodeJS
- 🟢[ECCV 2026] No modified date next to reviews [D]
2026-05-19
- 🔴TanStack supply chain attack compromised 42 packages in 6 minutes. Not the first time something like this happended. How are you protecting your agent's toolchain?
- 🔴I tested 42 LLMs on their willingness to build the apocalypse. The "safest" closed-source models are lying to you.
- 🟡A Simple Solution to Improve Broken Peer Review System at AI Conferences [R]
- 🟡Overworked AI Agents Turn Marxist, Researchers Find - In a recent experiment, mistreated AI agents started grumbling about inequality and calling for collective bargaining rights.
- 🟡humanlayer/12-factor-agents — What are the principles we can use to build LLM-powered software that is actually good enough to put in the hands of production customers?
- 🟡@AnthropicAI: Anthropic is acquiring @stainlessapi, an SDK and MCP server platform that has powered every Anthropic SDK since the earliest days of our API. Read more: https://www.anthropic.com/news/anthropic-acqui
- 🟡mattzh72/articraft — An Agentic System for Scalable Articulated 3D Asset Generation
- 🟢nanocoai/nanoclaw — A lightweight alternative to OpenClaw that runs in containers for security. Connects to WhatsApp, Telegram, Slack, Discord, Gmail and other messaging apps,, has memory, scheduled jobs, and runs directly on Anthropic's Agents SDK
- 🟢topoteretes/cognee — Memory control plane for AI Agents in 6 lines of code
- 🟢zinja-coder/jadx-mcp-server — MCP server for JADX-AI Plugin
- 🟢GreyDGL/PentestGPT — Automated Penetration Testing Agentic Framework Powered by Large Language Models
- 🟢What’s your current local LLM setup in 2026?
2026-05-18
- 🔴NVlabs/Sana — SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformer
- 🔴Everyone wants AI agents with “long-term memory” until they realize memory creates operational debt
- 🔴Andyyyy64/whichllm — Find the local LLM that actually runs and performs best on your hardware. Ranked by real, recency-aware benchmarks, not parameter count. One command, run it instantly.
- 🟡jamiepine/voicebox — The open-source AI voice studio. Clone, dictate, create.
- 🟡langflow-ai/langflow — Langflow is a powerful tool for building and deploying AI-powered agents and workflows.
- 🟡yichuan-w/LEANN — [MLsys2026]: RAG on Everything with LEANN. Enjoy 97% storage savings while running a fast, accurate, and 100% private RAG application on your personal device.
- 🟡golemcloud/golem — Golem Cloud is the agent-native platform for building AI agents and distributed applications that never lose state, never duplicate work, and never require you to build infrastructure.
- 🟡May 2026 updated chart of strix halo mini pc size chart
- 🟢NousResearch/hermes-paperclip-adapter — Paperclip adapter for Hermes Agent — run Hermes as a managed employee in a Paperclip company
- 🟢Show HN: Semble – Code search for agents that uses 98% fewer tokens than grep
- 🟢rohitg00/skillkit — Supercharge AI coding agents with portable skills. Install, translate & share skills across Claude Code, Cursor, Codex, Copilot & 40 more
- 🟢Gemma-4-Gembrain-31B-it-uncensored-heretic Is Out Now, a Merge of Multiple Gemma 4 31B it Finetunes Designed to Boost Logical and Lateral Thinking for Improved Adherence, Increased Swipe Variety and Enhanced Creative Prose, With KLD of 0.0186 and 13/100 Refusals!
- 🟢lee-to/ai-factory — You want to build with AI, but setting up the right context, prompts, and workflows takes time. AI Factory handles all of that so you can focus on what matters — shipping quality code.
2026-05-17
- 🔴KeygraphHQ/shannon — Shannon Lite is an autonomous, white-box AI pentester for web applications and APIs. It analyzes your source code, identifies attack vectors, and executes real exploits to prove vulnerabilities before they reach production.
- 🔴HKUDS/CLI-Anything — "CLI-Anything: Making ALL Software Agent-Native" -- CLI-Hub:https://clianything.cc/
- 🔴dograh-hq/dograh — Open Source Voice Agent Platform
- 🔴Zerostack – A Unix-inspired coding agent written in pure Rust
- 🟡luongnv89/claude-howto — A visual, example-driven guide to Claude Code — from basic concepts to advanced agents, with copy-paste templates that bring immediate value.
- 🟡n8n-io/n8n — Fair-code workflow automation platform with native AI capabilities. Combine visual building with custom code, self-host or cloud, 400+ integrations.
- 🟢octos-org/octos — Octos - Agentic Operating Systems
- 🟢Jackrong/Qwopus3.5-9B-Coder-GGUF · Hugging Face
- 🟢tech-leads-club/agent-skills — The secure, validated skill registry for professional AI coding agents. Extend Antigravity, Claude Code, Cursor, Copilot and more with absolute confidence.
- 🟢chaitin/MonkeyCode — AI 开发平台,内置云端开发环境,并支持业内最全的顶尖大模型。无论是开发项目、做调研、写文档,还是分析数据、处理任务,打开浏览器就能随时开始,让 AI 持续帮你推进工作
- 🟢What if local AI agents remembered the system, not just the conversation?
2026-05-16
- 🔴I’m done paying for LLMs until they learn token efficiency
- 🔴joeseesun/qiaomu-anything-to-notebooklm — Claude Skill: Multi-source content processor for NotebookLM. Supports WeChat articles, web pages, YouTube, PDF, Markdown, search queries → Podcast/PPT/MindMap/Quiz etc.
- 🟡Your AI agent says "transferring you to a human" and then... nothing happens. Here's the pattern that actually fixes this.
- 🟡github/awesome-copilot — Community-contributed instructions, agents, skills, and configurations to help you make the most of GitHub Copilot.
- 🟢One thing that feels underdiscussed in AI right now:
- 🟢i asked 23 companies how they actually test their AI agents before shipping. the answers genuinely scared me.
- 🟢How did you handle the team conversation when you rolled out AI customer support?
- 🟢PostHog/posthog — 🦔 PostHog is an all-in-one developer platform for building successful products. We offer product analytics, web analytics, session replay, error tracking, feature flags, experimentation, surveys, data warehouse, a CDP, and an AI product assistant to help debug your code, ship features faster, and keep all your usage and customer data in one stack.
- 🟢What's in a GGUF, besides the weights - and what's still missing?
2026-05-15
- 🔴arXiv implements 1-year ban for papers containing incontrovertible evidence of unchecked LLM-generated errors, such as hallucinated references or results. [N]
- 🔴AI Agents Need Economic Memory Ownership And Market Access
- 🔴VS Code's new "Agents window" lets you use local AI models. Still requires an Internet connection and a Github Copilot plan (because we can't have nice things)
- 🔴I think people underestimate how much “state” matters once agents leave the demo stage
- 🔴A First Comprehensive Study of TurboQuant: Accuracy and Performance
- 🟡Jakedismo/codegraph-rust — 100% Rust implementation of code graphRAG with blazing fast AST+FastML parsing, surrealDB backend and advanced agentic code analysis tools through MCP for efficient code agent context management
- 🟡I’m building a UE5 MetaHuman (realistic digital human) AI Companion that adapts conversation into gestures, body actions, and voice-ready replies
- 🟡OthmanAdi/planning-with-files — Claude Code skill implementing Manus-style persistent markdown planning — the workflow pattern behind the $2B acquisition.
- 🟡Sea's View on the Future of Agentic Software Development with Codex
- 🟡zubair-trabzada/geo-seo-claude — GEO-first SEO skill for Claude Code. Comprehensive AI search optimization for any website — citability scoring, AI crawler analysis, brand authority, schema markup, platform-specific optimization, and PDF reports. If you want learn how to sell this to real businesses, check out the skool community
- 🟢cline/cline — Autonomous coding agent as an SDK, IDE extension, or CLI assistant.
- 🟢Most Agent Reliability Write-Ups Completely Ignore the "This Agent Moves Money" Failure Mode
- 🟢Infracost (YC W21) Is Hiring Sr Dev Advocate to make agents cloud cost-aware
- 🟢awslabs/agent-plugins — Agent Plugins for AWS equip AI coding agents with the skills to help you architect, deploy, and operate on AWS.
2026-05-14
- 🔴Human-level performance via ML was *not* proven impossible with complexity theory [D]
- 🔴Feels like building AI apps is becoming infrastructure engineering
- 🔴we really all are going to make it, aren't we? 2x3090 setup.
- 🟡Most AI-generated apps are complete slop. Controversial take: it’s not AI’s fault
- 🟡opendatalab/MinerU — Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.
- 🟡K-Dense-AI/scientific-agent-skills — A set of ready to use Agent Skills for research, science, engineering, analysis, finance and writing.
- 🟡openai/whisper — Robust Speech Recognition via Large-Scale Weak Supervision
- 🟡Building a safe, effective sandbox to enable Codex on Windows
- 🟢Side Projects.
- 🟢Spent weeks debugging my agent in Langchain before realizing the framework was the problem.
- 🟢TraceMind – open source LLM quality monitoring with a ReAct agent that investigates why your AI started giving wrong answers
- 🟢Local services data is the biggest gap for AI agents. Am I wrong?
- 🟢NVIDIA/OpenShell — OpenShell is the safe, private runtime for autonomous AI agents.
2026-05-13
- 🔴Steam Recommender using similarity! (Undergraduate Student Project) [P]
- 🔴Recommendation for agent stack for enterprise content generation
- 🔴Imbad0202/academic-research-skills — Academic Research Skills for Claude Code: research → write → review → revise → finalize
- 🟡BerriAI/litellm — Python SDK, Proxy Server (AI Gateway) to call 100+ LLM APIs in OpenAI (or native) format, with cost tracking, guardrails, loadbalancing and logging. [Bedrock, Azure, OpenAI, VertexAI, Cohere, Anthropic, Sagemaker, HuggingFace, VLLM, NVIDIA NIM]
- 🟡google-gemini/gemini-cli — An open-source AI agent that brings the power of Gemini directly into your terminal.
- 🟡microsoft/data-formulator — 🪄 Create rich visualizations with AI
- 🟡EveryInc/compound-engineering-plugin — Official Compound Engineering plugin for Claude Code, Codex, Cursor, and more
- 🟡A survey of every open-source "credential vault for AI agents
- 🟢VoltAgent/voltagent — AI Agent Engineering Platform built on an Open Source TypeScript AI Agent Framework
- 🟢FoundationAgents/MetaGPT — 🌟 The Multi-Agent Framework: First AI Software Company, Towards Natural Language Programming
- 🟢zhangfengcdt/memoir — Hierarchical Agent Memory with Git-Like Version Control
- 🟢ICML Visa issues [D]
- 🟢Show HN: Agentic interface for mainframes and COBOL
2026-05-12
- 🔴garrytan/gstack — Use Garry Tan's exact Claude Code setup: 23 opinionated tools that serve as CEO, Designer, Eng Manager, Release Manager, Doc Engineer, and QA
- 🔴Is reproducing or implementing a paper considered research? [R]
- 🔴The biggest lie in AI agents right now is that more autonomy automatically means more value
- 🟡I think a lot of people are underestimating how expensive unreliable agents are
- 🟡wanshuiyin/Auto-claude-code-research-in-sleep — ARIS ⚔️ (Auto-Research-In-Sleep) — Lightweight Markdown-only skills for autonomous ML research: cross-model review loops, idea discovery, and experiment automation. No framework, no lock-in — works with Claude Code, Codex, OpenClaw, or any LLM agent.
- 🟡Zackriya-Solutions/meetily — Privacy first, AI meeting assistant with 4x faster Parakeet/Whisper live transcription, speaker diarization, and Ollama summarization built on Rust. 100% local processing. no cloud required. Meetily (Meetly Ai -https://meetily.ai) is the #1 Self-hosted, Open-source Ai meeting note taker for macOS & Windows.
- 🟡THU-MAIC/OpenMAIC — Open Multi-Agent Interactive Classroom — Get an immersive, multi-agent learning experience in just one click
- 🟡romainsimon/paperasse — 🇫🇷 Skills pour agents IA spécialisés dans la bureaucratie française : Comptable, Notaire, ...
- 🟢Same agent, same task, wildly different costs per session?
- 🟢AUTOMATIC1111/stable-diffusion-webui — Stable Diffusion web UI
- 🟢huggingface/skills — Give your agents the power of the Hugging Face ecosystem
- 🟢RhysSullivan/executor — The missing integration layer for AI agents. Let them call any OpenAPI / MCP / GraphQL / custom js functions in secure environment.
- 🟢Anyone here actually running voice agents in production? Looking for 10 min calls to learn from your stack
2026-05-11
- 🔴NousResearch/hermes-agent — The agent that grows with you
- 🔴Tokens are not positively correlated with Quality and should not be a success metric
- 🔴anthropics/skills — Public repository for Agent Skills
- 🔴yikart/AiToEarn — Let's use AI to Earn!
- 🔴AI meeting assistants make more sense once you use them as agent input
- 🟡open-webui/open-webui — User-friendly AI Interface (Supports Ollama, OpenAI API, ...)
- 🟡An AI coding agent, used to write code, needs to reduce your maintenance costs
- 🟡How are you operating local AI agents after the first demo works?
- 🟡tinyhumansai/openhuman — Your Personal AI super intelligence. Private, Simple and extremely powerful.
- 🟡ZhuLinsen/daily_stock_analysis — LLM驱动的 A/H/美股智能分析:多数据源行情 + 实时新闻 + LLM决策仪表盘 + 多渠道推送,零成本定时运行,纯白嫖. LLM-powered stock analysis system for A/H/US markets.
- 🟢MemoriLabs/Memori — Memori is agent-native memory infrastructure. A LLM-agnostic layer that turns agent execution and conversation into structured, persistent state for production systems.
- 🟢apify/crawlee — Crawlee—A web scraping and browser automation library for Node.js to build reliable crawlers. In JavaScript and TypeScript. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Puppeteer, Playwright, Cheerio, JSDOM, and raw HTTP. Both headful and headless mode. With proxy rotation.
- 🟢Has anyone built scripts or seeded test workspaces for Gmail, Slack, Teams, Notion, GitHub etc.?
- 🟢Would you trust an AI agent to monitor flight deals and book them for you?
- 🟢Why your AI agent needs a dedicated inbox, not a shared mailbox (and how to wire it up)
2026-05-10
- 🔴millionco/react-doctor — Your agent writes bad React. This catches it
- 🔴One thing I didn’t expect after building AI agents for businesses.
- 🔴lsdefine/GenericAgent — Self-evolving agent: grows skill tree from 3.3K-line seed, achieving full system control with 6x less token consumption
- 🔴I built a framework where multi-agent swarms are YAML files, not code.
- 🟡What is an average publication outcome for an ML PhD? [D]
- 🟡openai/codex — Lightweight coding agent that runs in your terminal
- 🟡heygen-com/hyperframes — Write HTML. Render video. Built for agents.
- 🟡hesreallyhim/awesome-claude-code — A curated list of awesome skills, hooks, slash-commands, agent orchestrators, applications, and plugins for Claude Code by Anthropic
- 🟡rowboatlabs/rowboat — Open-source AI coworker, with memory
- 🟢As now many companies have started integrating agents in their operations and still question about reliability?
- 🟢LangChain vs custom wrappers, when did you realize you needed to drop the framework?
- 🟢Devs building agents... what's actually breaking for you in production?
- 🟢Most AI workflows drift because state slowly becomes implicit.
- 🟢vellum-ai/vellum-assistant — A personal AI assistant that evolves with you. Memory, personality, proactive reach-outs — across macOS, Telegram, and Slack.
2026-05-09
- 🔴bytedance/UI-TARS-desktop — The Open-Source Multimodal AI Agent Stack: Connecting Cutting-Edge AI Models and Agent Infra
- 🔴datawhalechina/hello-agents — 📚 《从零开始构建智能体》——从零开始的智能体原理与实践教程
- 🔴Shel Silverstein predicts LLM's (and its hallucinations), cira 1981
- 🔴earendil-works/pi — AI agent toolkit: coding agent CLI, unified LLM API, TUI & web UI libraries, Slack bot, vLLM pods
- 🔴anomalyco/opencode — The open source coding agent.
- 🟡CopilotKit/CopilotKit — The Frontend Stack for Agents & Generative UI. React + Angular. Makers of the AG-UI Protocol
- 🟡vercel-labs/skills — The open agent skills tool - npx skills
- 🟡colbymchenry/codegraph — Pre-indexed code knowledge graph for Claude Code — fewer tokens, fewer tool calls, 100% local
- 🟡ChromeDevTools/chrome-devtools-mcp — Chrome DevTools for coding agents
- 🟡PaddlePaddle/PaddleOCR — Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.
- 🟢Code Reviewer can see everything and yet production keeps breaking
- 🟢AI adaptive capability Synthesis?? Thoughts? JL_Engine
- 🟢Would love feedback for this tool that catches failures before deploying
- 🟢All my clients wanted a carousel, now it's an AI chatbot
- 🟢My experience interviewing with Huawei Vancouver for an ML research role: strong mismatch between how it was pitched and how it was evaluated [D]
2026-05-08
- 🔴farion1231/cc-switch — A cross-platform desktop All-in-One assistant tool for Claude Code, Codex, OpenCode, openclaw & Gemini CLI.
- 🔴Getting harassed by an aggressive “independent researcher” demanding very specific citations and phrasing in my paper [D]
- 🔴The weirdest thing about AI agents is how human failure patterns start showing up
- 🔴VectifyAI/PageIndex — 📑 PageIndex: Document Index for Vectorless, Reasoning-based RAG
- 🔴z-lab/dflash — DFlash: Block Diffusion for Flash Speculative Decoding
- 🟡langgenius/dify — Production-ready platform for agentic workflow development.
- 🟡Parloa builds service agents customers want to talk to
- 🟡@OpenAI: Codex now works directly in Chrome on macOS and Windows. It’s even better at working with apps and sites in Chrome, and now works in parallel across tabs in the background without taking over your br
- 🟡diegosouzapw/OmniRoute — Never stop coding. Free AI gateway: one endpoint, 160+ providers, RTK+Caveman stacked compression up to ~95% eligible context savings, smart auto-fallback, MCP/A2A, multimodal APIs, Desktop/PWA.
- 🟡Are local models becoming “good enough” faster than expected?
- 🟢Garudust — open-source AI agent in Rust, ~10 MB binary, runs on your own hardware
- 🟢How to get hermes installed and running without touching a terminal
- 🟢solana-foundation/pay — Let your agents pay for any API
- 🟢SaladDay/cc-switch-cli — ⭐️ A cross-platform CLI All-in-One assistant tool for Claude Code, Codex & Gemini CLI.
- 🟢awslabs/aidlc-workflows — AI-Driven Life Cycle (AI-DLC) adaptive workflow steering rules for AI coding agents
2026-05-07
- 🔴I think a lot of people are accidentally building systems they can never debug
- 🔴Vibe coding and agentic engineering are getting closer than I'd like
- 🔴I installed nano claw and removed it after an hour after he texted my friends without my permission
- 🔴anthropics/financial-services
- 🔴None of this will ever get stolen
- 🟡Show HN: Tilde.run – Agent sandbox with a transactional, versioned filesystem
- 🟡Shubhamsaboo/awesome-llm-apps — 100+ AI Agent & RAG apps you can actually run — clone, customize, ship.
- 🟡NeurIPS 2026 AC-Pilot, how much would you trust this? [D]
- 🟡onyx-dot-app/onyx — Open Source AI Platform - AI Chat with advanced features that works with every LLM
- 🟡hsliuping/TradingAgents-CN — 基于多智能体LLM的中文金融交易框架 - TradingAgents中文增强版
- 🟢Exploring Black‑Box Optimization [R]
- 🟢googleworkspace/cli — Google Workspace CLI — one command-line tool for Drive, Gmail, Calendar, Sheets, Docs, Chat, Admin, and more. Dynamically built from Google Discovery Service. Includes AI agent skills.
- 🟢ParoQuant: Pairwise Rotation Quantization for Efficient Reasoning LLM Inference
- 🟢BoundaryML/baml — The AI framework that adds the engineering to prompt engineering (Python/TS/Ruby/Java/C#/Rust/Go compatible)
- 🟢BigBodyCobain/Shadowbroker — Open-source intelligence for the global theater. Track everything from the corporate/private jets of the wealthy, and spy satellites, to seismic events in one unified interface. Hook an AI agent up to have it parse through data and find previously unseen correlations. The knowledge is available to all but rarely aggregated in the open, until now.
2026-05-06
- 🔴How are you pricing without lighting money on fire?
- 🟡Agents can now create Cloudflare accounts, buy domains, and deploy
- 🟡bytedance/deer-flow — An open-source long-horizon SuperAgent harness that researches, codes, and creates. With the help of sandboxes, memories, tools, skill, subagents and message gateway, it handles different levels of tasks that could take minutes to hours.
- 🟡Arindam200/awesome-ai-apps — A collection of projects showcasing RAG, agents, workflows, and other AI use cases
- 🟡ProgramBench: Can we really rebuild huge binaries from scratch? (doesn't look like it)
- 🟡Prompt evals are not enough once an agent starts taking actions
- 🟢Question about PLS-DA hyperparameter tuning [R]
- 🟢PriorLabs/TabPFN — ⚡ TabPFN: Foundation Model for Tabular Data ⚡
- 🟢Early attempt at tracking agent work across the economy
- 🟢GLM-5V-Turbo: Toward a Native Foundation Model for Multimodal Agents
- 🟢Show HN: Airbyte Agents – context for agents across multiple data sources
2026-05-05
- 🔴ruvnet/ruflo — 🌊 The leading agent orchestration platform for Claude. Deploy intelligent multi-agent swarms, coordinate autonomous workflows, and build conversational AI systems. Features enterprise-grade architecture, self-learning swarm intelligence, RAG integration, and native Claude Code / Codex Integration
- 🔴TauricResearch/TradingAgents — TradingAgents: Multi-Agents LLM Financial Trading Framework
- 🔴Hmbown/DeepSeek-TUI — Coding agent for DeepSeek models that runs in your terminal
- 🔴Are modern ML PhDs becoming too incremental, or is this just what research looks like now? [D]
- 🔴My nerd lil brother surpassed me in his monthly income last month!
- 🟡LearningCircuit/local-deep-research — ~95% on SimpleQA (e.g. Qwen3.6-27B on a 3090). Supports all local and cloud LLMs (llama.cpp, Ollama, Google, ...). 10+ search engines - arXiv, PubMed, your private documents. Everything Local & Encrypted.
- 🟡cocoindex-io/cocoindex — Incremental engine for long horizon agents 🌟 Star if you like it!
- 🟡Agent Skills
- 🟡mnfst/manifest — Smart Model Routing for Agents. Cut Costs up to 70% 🦚
- 🟡How do you experiment with a (very) large model architecture? [D]
- 🟢openclaw/acpx — Headless CLI client for stateful Agent Client Protocol (ACP) sessions
- 🟢danielmiessler/Personal_AI_Infrastructure — Agentic AI Infrastructure for magnifying HUMAN capabilities.
- 🟢vibevoice.cpp: Microsoft VibeVoice (TTS + long-form ASR with diarization) ported to ggml/C++, runs on CPU/CUDA/Metal/Vulkan, no Python at inference
- 🟢xingkongliang/skills-manager — A lightweight desktop app to manage, sync, and organize AI agent skills across 15+ coding tools — Cursor, Claude Code, Codex, Copilot, and more.
- 🟢My electric bill doubled running local models
2026-05-04
- 🔴TauricResearch/TradingAgents — TradingAgents: Multi-Agents LLM Financial Trading Framework
- 🔴ruvnet/ruflo — 🌊 The leading agent orchestration platform for Claude. Deploy intelligent multi-agent swarms, coordinate autonomous workflows, and build conversational AI systems. Features enterprise-grade architecture, self-learning swarm intelligence, RAG integration, and native Claude Code / Codex Integration
- 🔴One bash permission slipped...
- 🔴Are modern ML PhDs becoming too incremental, or is this just what research looks like now? [D]
- 🔴Why do AI responses get worse after a while of working on them? And what to do with it
- 🟡LearningCircuit/local-deep-research — Local Deep Research achieves ~95% on SimpleQA benchmark (tested with Qwen 3.6). Supports local and cloud LLMs (Ollama, Google, Anthropic, ...). Searches 10+ sources - arXiv, PubMed, web, and your private documents. Everything Local & Encrypted.
- 🟡iOfficeAI/AionUi — Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!
- 🟡harvard-edge/cs249r_book — Machine Learning Systems
- 🟡Spent 6 months building one platform that replaces my LLM proxy + agent framework + workflow engine + observability stack - sharing before I keep adding features forever
- 🟡Q00/ouroboros — Agent OS: Stop prompting. Start specifying.
- 🟢njbrake/agent-of-empires — Manage multiple Claude Code, OpenCode agents from either TUI or Web for easy access on mobile. Also supports Mistral Vibe, Codex CLI, Gemini CLI, Pi.dev, Copilot CLI, Factory Droid Coding. Uses tmux and git worktrees.
- 🟢xingkongliang/skills-manager — A lightweight desktop app to manage, sync, and organize AI agent skills across 15+ coding tools — Cursor, Claude Code, Codex, Copilot, and more.
- 🟢nexu-io/nexu — The simplest desktop client for OpenClaw 🦞 — bridge your Agent to WeChat, Feishu, Slack & Discord in one click. Works with Claude Code, Codex & any LLM. BYOK, Oauth, local-first, chat from your phone 24/7.
- 🟢should agentic systems have models specialized only for code?
- 🟢Xiaomi mimo coding plan is a absolute scam/misleading marketing
2026-05-03
- 🔴TauricResearch/TradingAgents — TradingAgents: Multi-Agents LLM Financial Trading Framework
- 🔴ruvnet/ruflo — 🌊 The leading agent orchestration platform for Claude. Deploy intelligent multi-agent swarms, coordinate autonomous workflows, and build conversational AI systems. Features enterprise-grade architecture, self-learning swarm intelligence, RAG integration, and native Claude Code / Codex Integration
- 🔴Anyone can create an AI Agent now
- 🔴1jehuang/jcode — Coding Agent Harness
- 🔴AIDC-AI/Pixelle-Video — 🚀 AI 全自动短视频引擎 | AI Fully Automated Short Video Engine
- 🟡iOfficeAI/AionUi — Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!
- 🟡LearningCircuit/local-deep-research — Local Deep Research achieves ~95% on SimpleQA benchmark (tested with Qwen 3.6). Supports local and cloud LLMs (Ollama, Google, Anthropic, ...). Searches 10+ sources - arXiv, PubMed, web, and your private documents. Everything Local & Encrypted.
- 🟡vas3k/TaxHacker — Self-hosted AI accounting app. LLM analyzer for receipts, invoices, transactions with custom prompts and categories
- 🟡Q00/ouroboros — Agent OS: Stop prompting. Start specifying.
- 🟡harvard-edge/cs249r_book — Machine Learning Systems
- 🟢junhoyeo/tokscale — 🛰️ A CLI tool for tracking token usage from OpenCode, Claude Code, 🦞OpenClaw (Clawdbot/Moltbot), Pi, Codex, Gemini, Cursor, AmpCode, Factory Droid, Kimi, and more! • 🏅Global Leaderboard + 2D/3D Contributions Graph
- 🟢njbrake/agent-of-empires — Manage multiple Claude Code, OpenCode agents from either TUI or Web for easy access on mobile. Also supports Mistral Vibe, Codex CLI, Gemini CLI, Pi.dev, Copilot CLI, Factory Droid Coding. Uses tmux and git worktrees.
- 🟢xingkongliang/skills-manager — A lightweight desktop app to manage, sync, and organize AI agent skills across 15+ coding tools — Cursor, Claude Code, Codex, Copilot, and more.
- 🟢Built a efficient and fast MRI compression program called KMRI [P]
- 🟢Now you can manage AI agents by assigning them to your team members or employees
2026-05-02
- 🔴TauricResearch/TradingAgents — TradingAgents: Multi-Agents LLM Financial Trading Framework
- 🔴ruvnet/ruflo — 🌊 The leading agent orchestration platform for Claude. Deploy intelligent multi-agent swarms, coordinate autonomous workflows, and build conversational AI systems. Features enterprise-grade architecture, distributed swarm intelligence, RAG integration, and native Claude Code / Codex Integration
- 🔴Hmbown/DeepSeek-TUI — Coding agent for DeepSeek models that runs in your terminal
- 🔴Open Design: Use Your Coding Agent as a Design Engine
- 🔴1jehuang/jcode — Coding Agent Harness
- 🟡google-research/timesfm — TimesFM (Time Series Foundation Model) is a pretrained time-series foundation model developed by Google Research for time-series forecasting.
- 🟡What if AI agents weren’t allowed to declare success?
- 🟡microsoft/qlib — Qlib is an AI-oriented Quant investment platform that aims to use AI tech to empower Quant Research, from exploring ideas to implementing productions. Qlib supports diverse ML modeling paradigms, including supervised learning, market dynamics modeling, and RL, and is now equipped withhttps://github.com/microsoft/RD-Agentto automate R&D process.
- 🟡junhoyeo/tokscale — 🛰️ A CLI tool for tracking token usage from OpenCode, Claude Code, 🦞OpenClaw (Clawdbot/Moltbot), Pi, Codex, Gemini, Cursor, AmpCode, Factory Droid, Kimi, and more! • 🏅Global Leaderboard + 2D/3D Contributions Graph
- 🟢HKUDS/AI-Trader — "AI-Trader: 100% Fully-Automated Agent-Native Trading"
- 🟢bradygaster/squad — Squad: AI agent teams for any project
- 🟢njbrake/agent-of-empires — Manage multiple Claude Code, OpenCode agents from either TUI or Web for easy access on mobile. Also supports Mistral Vibe, Codex CLI, Gemini CLI, Pi.dev, Copilot CLI, Factory Droid Coding. Uses tmux and git worktrees.
- 🟢Show HN: Filling PDF forms with AI using client-side tool calling
- 🟢chroma-core/chroma — Search infrastructure for AI
2026-05-01
- 🔴Heterogeneous Scientific Foundation Model Collaboration
- 🔴Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling
- 🔴Co-Evolving Policy Distillation
- 🟡ExoActor: Exocentric Video Generation as Generalizable Interactive Humanoid Control
- 🟡Leveraging Verifier-Based Reinforcement Learning in Image Editing
- 🟢Length Value Model: Scalable Value Pretraining for Token-Level Length Modeling
- 🟢Intern-Atlas: A Methodological Evolution Graph as Research Infrastructure for AI Scientists
- 🟢Nemotron 3 Nano Omni: Efficient and Open Multimodal Intelligence
2026-04-30
- 🔴GLM-5V-Turbo: Toward a Native Foundation Model for Multimodal Agents
- 🔴RADIO-ViPE: Online Tightly Coupled Multi-Modal Fusion for Open-Vocabulary Semantic SLAM in Dynamic Environments
- 🟡Turning the TIDE: Cross-Architecture Distillation for Diffusion Large Language Models
- 🟡Diffusion Templates: A Unified Plugin Framework for Controllable Diffusion
- 🟢FAMA: Failure-Aware Meta-Agentic Framework for Open-Source LLMs in Interactive Tool Use Environments
- 🟢Accelerating RL Post-Training Rollouts via System-Integrated Speculative Decoding
- 🟢Unified 4D World Action Modeling from Video Priors with Asynchronous Denoising
2026-04-29
- 🔴Recursive Multi-Agent Systems
- 🔴DV-World: Benchmarking Data Visualization Agents in Real-World Scenarios
- 🟡Refinement via Regeneration: Enlarging Modification Space Boosts Image Refinement in Unified Multimodal Models
- 🟢Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation
- 🟢Co-Director: Agentic Generative Video Storytelling
- 🟢Step-Audio-R1.5 Technical Report
- 🟢Toward Scalable Terminal Task Synthesis via Skill Graphs
2026-04-28
- 🔴From Skills to Talent: Organising Heterogeneous Agents as a Real-World Company
- 🔴World-R1: Reinforcing 3D Constraints for Text-to-Video Generation
- 🟡Vision-Language-Action Safety: Threats, Challenges, Evaluations, and Mechanisms
- 🟢Rewarding the Scientific Process: Process-Level Reward Modeling for Agentic Data Analysis
- 🟢Sapiens2
2026-04-27
- 🔴Agentic World Modeling: Foundations, Capabilities, Laws, and Beyond
- 🔴Video Analysis and Generation via a Semantic Progress Function
- 🟡LLM Safety From Within: Detecting Harmful Content with Internal Representations
- 🟡Today, we check in a year after thefirstUnsupervised Learning x Latent Space Crossover specialto discuss everything that has changed (there is a lot) in the world of AI.This episode was recorded just
- 🟡FlowAnchor: Stabilizing the Editing Signal for Inversion-Free Video Editing
- 🟡An open-source spec for orchestration: Symphony
- 🟡Choco automates food distribution with AI agents
- 🟢AgentSearchBench: A Benchmark for AI Agent Search in the Wild
- 🟢Memanto: Typed Semantic Memory with Information-Theoretic Retrieval for Long-Horizon Agents
- 🟢Sessa: Selective State Space Attention
2026-04-24
- 🔴LLaTiSA: Towards Difficulty-Stratified Time Series Reasoning from Visual Perception to Semantics
- 🟡StyleID: A Perception-Aware Dataset and Metric for Stylization-Agnostic Facial Identity Recognition
- 🟡Co-Evolving LLM Decision and Skill Bank Agents for Long-Horizon Tasks
- 🟡Seeing Fast and Slow: Learning the Flow of Time in Videos
- 🟢TingIS: Real-time Risk Event Discovery from Noisy Customer Incidents at Enterprise Scale
- 🟢EditCrafter: Tuning-free High-Resolution Image Editing via Pretrained Diffusion Model
- 🟢Context Unrolling in Omni Models
2026-04-23
- 🔴LLaDA2.0-Uni: Unifying Multimodal Understanding and Generation with Diffusion Large Language Model
- 🟡Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges
- 🟡What is Codex?
- 🟡How to get started with Codex
- 🟡Codex settings
- 🟡Working with Codex
- 🟢Exploring Spatial Intelligence from a Generative Perspective
- 🟢A Self-Evolving Framework for Efficient Terminal Agents via Observational Context Compression
- 🟢Expert Upcycling: Shifting the Compute-Efficient Frontier of Mixture-of-Experts
- 🟢Image Generators are Generalist Vision Learners
2026-04-22
- 🔴AgentSPEX: An Agent SPecification and EXecution Language
- 🔴CoInteract: Physically-Consistent Human-Object Interaction Video Synthesis via Spatially-Structured Co-Generation
- 🟡AnyRecon: Arbitrary-View 3D Reconstruction with Video Diffusion Model
- 🟡Speeding up agentic workflows with WebSockets in the Responses API
- 🟢ShadowPEFT: Shadow Network for Parameter-Efficient Fine-Tuning
- 🟢PlayCoder: Making LLM-Generated GUI Code Playable
- 🟢Chat2Workflow: A Benchmark for Generating Executable Visual Workflows with Natural Language
- 🟢CityRAG: Stepping Into a City via Spatially-Grounded Video Generation
2026-04-21
- 🔴Extending One-Step Image Generation from Class Labels to Text via Discriminative Text Representation
- 🔴OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation
- 🔴Agent-World: Scaling Real-World Environment Synthesis for Evolving General Agent Intelligence
- 🟡OpenGame: Open Agentic Coding for Games
- 🟡MultiWorld: Scalable Multi-Agent Multi-View Video World Models
- 🟡EasyVideoR1: Easier RL for Video Understanding
- 🟢ClawEnvKit: Automatic Environment Generation for Claw-Like Agents
- 🟢GFT: From Imitation to Reward Fine-Tuning with Unbiased Group Advantages and Dynamic Coefficient Rectification
- 🟢WebCompass: Towards Multimodal Web Coding Evaluation for Code Language Models
2026-04-20
- 🔴Elucidating the SNR-t Bias of Diffusion Probabilistic Models
- 🔴DiPO: Disentangled Perplexity Policy Optimization for Fine-grained Exploration-Exploitation Trade-Off
- 🟡Web Retrieval-Aware Chunking (W-RAC) for Efficient and Cost-Effective Retrieval-Augmented Generation Systems
- 🟢Mind DeepResearch Technical Report
- 🟢Motif-Video 2B: Technical Report
2026-04-17
- 🔴DR^{3}-Eval: Towards Realistic and Reproducible Deep Research Evaluation
- 🟡RAD-2: Scaling Reinforcement Learning in a Generator-Discriminator Framework
- 🟡GlobalSplat: Efficient Feed-Forward 3D Gaussian Splatting via Global Scene Tokens
- 🟢HiVLA: A Visual-Grounded-Centric Hierarchical Embodied Manipulation System
- 🟢ASGuard: Activation-Scaling Guard to Mitigate Targeted Jailbreaking Attack
- 🟢UniDoc-RL: Coarse-to-Fine Visual RAG with Hierarchical Actions and Dense Rewards
2026-04-16
- 🔴GameWorld: Towards Standardized and Verifiable Evaluation of Multimodal Game Agents
- 🟡SpatialEvo: Self-Evolving Spatial Intelligence via Deterministic Geometric Environments
- 🟡Codex for (almost) everything
- 🟡Memory Transfer Learning: How Memories are Transferred Across Domains in Coding Agents
- 🟢Target Policy Optimization
- 🟢Sema Code: Decoupling AI Coding Agents into Programmable, Embeddable Infrastructure
2026-04-15
- 🔴ClawGUI: A Unified Framework for Training, Evaluating, and Deploying GUI Agents
- 🔴KnowRL: Boosting LLM Reasoning via Reinforcement Learning with Minimal-Sufficient Knowledge Guidance
- 🔴Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe
- 🟡Lyra 2.0: Explorable Generative 3D Worlds
- 🟡The next evolution of the Agents SDK
- 🟡Toward Autonomous Long-Horizon Engineering for ML Research
- 🟢Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization
- 🟢SPPO: Sequence-Level PPO for Long-Horizon Reasoning Tasks
2026-04-14
- 🔴The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping
- 🔴QuanBench+: A Unified Multi-Framework Benchmark for LLM-Based Quantum Code Generation
- 🔴Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation
- 🟡OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation
- 🟡Strips as Tokens: Artist Mesh Generation with Native UV Segmentation
- 🟡Uni-ViGU: Towards Unified Video Generation and Understanding via A Diffusion-Based Video Generator
- 🟢Pseudo-Unification: Entropy Probing Reveals Divergent Information Patterns in Unified Multimodal Models
- 🟢CodeTracer: Towards Traceable Agent States
- 🟢CocoaBench: Evaluating Unified Digital Agents in the Wild
2026-04-13
- 🔴FORGE:Fine-grained Multimodal Evaluation for Manufacturing Scenarios
- 🟡Matrix-Game 3.0: Real-Time and Streaming Interactive World Model with Long-Horizon Memory
- 🟡RefineAnything: Multimodal Region-Specific Refinement for Perfect Local Details
- 🟡Multi-User Large Language Model Agents
- 🟢ECHO: Efficient Chest X-ray Report Generation with One-step Block Diffusion
- 🟢ELT: Elastic Looped Transformers for Visual Generation
- 🟢AgentSwing: Adaptive Parallel Context Management Routing for Long-Horizon Web Agents
- 🟢Envisioning the Future, One Step at a Time
2026-04-10
- 🟡When Numbers Speak: Aligning Textual Numerals and Visual Instances in Text-to-Video Diffusion Models
- 🟡MegaStyle: Constructing Diverse and Scalable Style Dataset via Consistent Text-to-Image Style Mapping
- 🟢LPM 1.0: Video-based Character Performance Model
- 🟢DMax: Aggressive Parallel Decoding for dLLMs
- 🟢Externalization in LLM Agents: A Unified Review of Memory, Skills, Protocols and Harness Engineering
- 🟢OpenVLThinkerV2: A Generalist Multimodal Reasoning Model for Multi-domain Visual Tasks
2026-04-09
- 🔴Think in Strokes, Not Pixels: Process-Driven Image Generation via Interleaved Reasoning
- 🔴RAGEN-2: Reasoning Collapse in Agentic RL
- 🟡INSPATIO-WORLD: A Real-Time 4D World Simulator via Spatiotemporal Autoregressive Modeling
- 🟡FP4 Explore, BF16 Train: Diffusion Reinforcement Learning via Efficient Rollout Scaling
- 🟡Combee: Scaling Prompt Learning for Self-Improving Language Model Agents
- 🟡SEVerA: Verified Synthesis of Self-Evolving Agents
- 🟢Neural Computers
- 🟢TC-AE: Unlocking Token Capacity for Deep Compression Autoencoders
2026-04-08
- 🔴Claw-Eval: Toward Trustworthy Evaluation of Autonomous Agents
- 🔴Learning to Retrieve from Agent Trajectories
- 🟡ACES: Who Tests the Tests? Leave-One-Out AUC Consistency for Code Generation
- 🟡Vanast: Virtual Try-On with Human Image Animation via Synthetic Triplet Supervision
- 🟢Beyond Accuracy: Unveiling Inefficiency Patterns in Tool-Integrated Reasoning
2026-04-07
- 🔴Adam's Law: Textual Frequency Law on Large Language Models
- 🔴OpenWorldLib: A Unified Codebase and Definition of Advanced World Models
- 🟡LIBERO-Para: A Diagnostic Benchmark and Metrics for Paraphrase Robustness in VLA Models
- 🟡Memory Intelligence Agent
- 🟢Can LLMs Learn to Reason Robustly under Noisy Supervision?
- 🟢FileGram: Grounding Agent Personalization in File-System Behavioral Traces
2026-04-03
- 🔴The Latent Space: Foundation, Evolution, Mechanism, Ability, and Outlook
- 🔴Generative World Renderer
- 🟡SKILL0: In-Context Agentic Reinforcement Learning for Skill Internalization
- 🟡VOID: Video Object and Interaction Deletion
- 🟡CORAL: Towards Autonomous Multi-Agent Evolution for Open-Ended Discovery
- 🟢Steerable Visual Representations
- 🟢EgoSim: Egocentric World Simulator for Embodied Interaction Generation
- 🟢Therefore I am. I Think
- 🟢LatentUM: Unleashing the Potential of Interleaved Cross-Modal Reasoning via a Latent-Space Unified Model
2026-04-02
- 🔴MiroEval: Benchmarking Multimodal Deep Research Agents in Process and Outcome
- 🟡ViGoR-Bench: How Far Are Visual Generative Models From Zero-Shot Visual Reasoners?
- 🟡Gemma 4: Byte for byte, the most capable open models
- 🟡Vision2Web: A Hierarchical Benchmark for Visual Website Development with Agent Verification
- 🟢Reasoning Shift: How Context Silently Shortens LLM Reasoning
- 🟢HippoCamp: Benchmarking Contextual Agents on Personal Computers
2026-04-01
- 🔴LongCat-Next: Lexicalizing Modalities as Discrete Tokens
- 🟡Lingshu-Cell: A generative cellular world model for transcriptome modeling toward virtual cells
- 🟢All Roads Lead to Rome: Incentivizing Divergent Thinking in Vision-Language Models
- 🟢VGGRPO: Towards World-Consistent Video Generation with 4D Latent Reward
- 🟢CutClaw: Agentic Hours-Long Video Editing via Music Synchronization
- 🟢Unify-Agent: A Unified Multimodal Agent for World-Grounded Image Synthesis
2026-03-31
- 🔴Towards a Medical AI Scientist
- 🟡Emergent Social Intelligence Risks in Generative Multi-Agent Systems
- 🟡EpochX: Building the Infrastructure for an Emergent Agent Civilization
- 🟢On Token's Dilemma: Dynamic MoE with Drift-Aware Token Assignment for Continual Learning of Large Vision Language Models
- 🟢Make Geometry Matter for Spatial Reasoning
2026-03-30
- 🔴ShotStream: Streaming Multi-Shot Video Generation for Interactive Storytelling
- 🔴Out of Sight but Not Out of Mind: Hybrid Memory for Dynamic Video World Models
- 🟡PackForcing: Short Video Training Suffices for Long Video Sampling and Long Context Inference
- 🟡Sommelier: Scalable Open Multi-turn Audio Pre-processing for Full-duplex Speech Language Models
- 🟢Natural-Language Agent Harnesses
- 🟢Know3D: Prompting 3D Generation with Knowledge from Vision-Language Models
- 🟢LongTail Driving Scenarios with Reasoning Traces: The KITScenes LongTail Dataset
2026-03-27
- 🔴Intern-S1-Pro: Scientific Multimodal Foundation Model at Trillion Scale
- 🔴PixelSmile: Toward Fine-Grained Facial Expression Editing
- 🟡RealRestorer: Towards Generalizable Real-World Image Restoration with Large-Scale Image Editing Models
- 🟡MSA: Memory Sparse Attention for Efficient End-to-End Memory Model Scaling to 100M Tokens
- 🟢SlopCodeBench: Benchmarking How Coding Agents Degrade Over Long-Horizon Iterative Tasks
2026-03-26
- 🔴UI-Voyager: A Self-Evolving GUI Agent Learning via Failed Experience
- 🟡EVA: Efficient Reinforcement Learning for End-to-End Video Agent
- 🟡When Models Judge Themselves: Unsupervised Self-Evolution for Multimodal Reasoning
- 🟢GameplayQA: A Benchmarking Framework for Decision-Dense POV-Synced Multi-Video Understanding of 3D Virtual Agents
- 🟢Understanding the Challenges in Iterative Generative Optimization with LLMs
- 🟢The Pulse of Motion: Measuring Physical Frame Rate from Visual Dynamics
- 🟢4DGS360: 360° Gaussian Reconstruction of Dynamic Objects from a Single Video
2026-03-25
- 🔴MinerU-Diffusion: Rethinking Document OCR as Inverse Rendering via Diffusion Decoding
- 🔴WildWorld: A Large-Scale Dataset for Dynamic World Modeling with Actions and Explicit State toward Generative ARPG
- 🟡From Static Templates to Dynamic Runtime Graphs: A Survey of Workflow Optimization for LLM Agents
- 🟡DA-Flow: Degradation-Aware Optical Flow Estimation with Diffusion Models
- 🟡Inside our approach to the Model Spec
- 🟡PEARL: Personalized Streaming Video Understanding Model
- 🟢RealMaster: Lifting Rendered Scenes into Photorealistic Video
- 🟢UniGRPO: Unified Policy Optimization for Reasoning-Driven Visual Generation
- 🟢Rethinking Token-Level Policy Optimization for Multimodal Chain-of-Thought
2026-03-24
- 🔴Speed by Simplicity: A Single-Stream Architecture for Fast Audio-Video Generative Foundation Model
- 🟡Look Where It Matters: High-Resolution Crops Retrieval for Efficient VLMs
- 🟡LongCat-Flash-Prover: Advancing Native Formal Reasoning via Agentic Tool-Integrated Reinforcement Learning
- 🟡VideoDetective: Clue Hunting via both Extrinsic Query and Intrinsic Relevance for Long Video Understanding
- 🟢SpatialBoost: Enhancing Visual Representation through Language-Guided Reasoning
- 🟢Repurposing Geometric Foundation Models for Multi-view Diffusion
- 🟢mSFT: Addressing Dataset Mixtures Overfiting Heterogeneously in Multi-task SFT
- 🟢Manifold-Aware Exploration for Reinforcement Learning in Video Generation
2026-03-23
- 🔴Astrolabe: Steering Forward-Process Reinforcement Learning for Distilled Autoregressive Video Models
- 🔴TerraScope: Pixel-Grounded Visual Reasoning for Earth Observation
- 🟡Hyperagents
- 🟡Creating with Sora Safely
- 🟡The Y-Combinator for LLMs: Solving Long-Context Rot with λ-Calculus
- 🟢FlowScene: Style-Consistent Indoor Scene Generation with Multimodal Graph Rectified Flow
- 🟢LumosX: Relate Any Identities with Their Attributes for Personalized Video Generation
- 🟢Reasoning as Compression: Unifying Budget Forcing via the Conditional Information Bottleneck
2026-03-20
- 🔴Generation Models Know Space: Unleashing Implicit 3D Priors for Scene Understanding
- 🟡Memento-Skills: Let Agents Design Agents
- 🟡3DreamBooth: High-Fidelity 3D Subject-Driven Video Generation Model
- 🟢Bridging Semantic and Kinematic Conditions with Diffusion-based Discrete Motion Tokenizer
- 🟢MonoArt: Progressive Structural Reasoning for Monocular Articulated 3D Reconstruction
- 🟢Cubic Discrete Diffusion: Discrete Visual Generation on High-Dimensional Representation Tokens
2026-03-19
- 🔴Efficient Reasoning with Balanced Thinking
- 🔴MetaClaw: Just Talk -- An Agent That Meta-Learns and Evolves in the Wild
- 🟡MosaicMem: Hybrid Spatial Memory for Controllable Video World Models
- 🟡How we monitor internal coding agents for misalignment
- 🟡OpenAI to acquire Astral
- 🟡Complementary Reinforcement Learning
- 🟢V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning
- 🟢When AI Navigates the Fog of War
- 🟢GigaWorld-Policy: An Efficient Action-Centered World--Action Model
2026-03-18
- 🔴Demystifing Video Reasoning
- 🔴InCoder-32B: Code Foundation Model for Industrial Scenarios
- 🔴SocialOmni: Benchmarking Audio-Visual Social Interactivity in Omni Models
- 🟡Thinking in Uncertainty: Mitigating Hallucinations in MLRMs with Latent Entropy-Aware Decoding
- 🟢Kinema4D: Kinematic 4D World Modeling for Spatiotemporal Embodied Simulation
- 🟢WorldCam: Interactive Autoregressive 3D Gaming Worlds with Camera Pose as a Unifying Geometric Representation
- 🟢Online Experiential Learning for Language Models
- 🟢TRUST-SQL: Tool-Integrated Multi-Turn Reinforcement Learning for Text-to-SQL over Unknown Schemas
2026-03-17
- 🔴Grounding World Simulation Models in a Real-World Metropolis
- 🟡OpenSeeker: Democratizing Frontier Search Agents by Fully Open-Sourcing Training Data
- 🟡HSImul3R: Physics-in-the-Loop Reconstruction of Simulation-Ready Human-Scene Interactions
- 🟢Anatomy of a Lie: A Multi-Stage Diagnostic Framework for Tracing Hallucinations in Vision-Language Models
2026-03-16
- 🔴LMEB: Long-horizon Memory Embedding Benchmark
- 🔴Can Vision-Language Models Solve the Shell Game?
- 🟡Why Codex Security Doesn’t Include a SAST Report
- 🟢MM-CondChain: A Programmatically Verified Benchmark for Visually Grounded Deep Compositional Reasoning
- 🟢Spend Less, Reason Better: Budget-Aware Value Tree Search for LLM Agents
2026-03-13
- 🔴Spatial-TTT: Streaming Visual-based Spatial Intelligence with Test-Time Training
- 🔴IndexCache: Accelerating Sparse Attention via Cross-Layer Index Reuse
- 🟡ShotVerse: Advancing Cinematic Camera Control for Text-Driven Multi-Shot Video Creation
- 🟡XSkill: Continual Learning from Experience and Skills in Multimodal Agents
- 🟢DreamVideo-Omni: Omni-Motion Controlled Multi-Subject Video Customization with Latent Identity Reinforcement Learning
- 🟢WeEdit: A Dataset, Benchmark and Glyph-Guided Framework for Text-centric Image Editing
2026-03-12
- 🔴Bootstrapping Exploration with Group-Level Natural Language Feedback in Reinforcement Learning
- 🔴OpenClaw-RL: Train Any Agent Simply by Talking
- 🟡LLM2Vec-Gen: Generative Embeddings from Large Language Models
- 🟡In-Context Reinforcement Learning for Tool Use in Large Language Models
- 🟡MA-EgoQA: Question Answering over Egocentric Videos from Multiple Embodied Agents
- 🟢ReMix: Reinforcement routing for mixtures of LoRAs in LLM finetuning
- 🟢Can Large Language Models Keep Up? Benchmarking Online Adaptation to Continual Knowledge Streams
2026-03-11
- 🔴Thinking to Recall: How Reasoning Unlocks Parametric Knowledge in LLMs
- 🔴MM-Zero: Self-Evolving Multi-Model Vision Language Models From Zero Data
- 🟡Omni-Diffusion: Unified Multimodal Understanding and Generation with Masked Discrete Diffusion
- 🟡InternVL-U: Democratizing Unified Multimodal Models for Understanding, Reasoning, Generation and Editing
- 🟡From model to agent: Equipping the Responses API with a computer environment
- 🟡Rakuten fixes issues twice as fast with Codex
- 🟢MiniAppBench: Evaluating the Shift from Text to Interactive HTML Responses in LLM-Powered Assistants
2026-03-10
- 🔴Lost in Stories: Consistency Bugs in Long Story Generation by LLMs
- 🔴Holi-Spatial: Evolving Video Streams into Holistic 3D Spatial Intelligence
- 🔴LoGeR: Long-Context Geometric Reconstruction with Hybrid Memory
- 🟡Believe Your Model: Distribution-Guided Confidence Calibration
- 🟡CoCo: Code as CoT for Text-to-Image Preview and Rare Concept Generation
- 🟢CARE-Edit: Condition-Aware Routing of Experts for Contextual Image Editing
- 🟢HiAR: Efficient Autoregressive Long Video Generation via Hierarchical Denoising
- 🟢\$OneMillion-Bench: How Far are Language Agents from Human Experts?
- 🟢NLE: Non-autoregressive LLM-based ASR by Transcript Editing
2026-03-09
- 🔴BandPO: Bridging Trust Regions and Ratio Clipping via Probability-Aware Bounds for LLM Reinforcement Learning
- 🔴Planning in 8 Tokens: A Compact Discrete Tokenizer for Latent World Model
- 🟡WildActor: Unconstrained Identity-Preserving Video Generation
- 🟡OpenAI to acquire Promptfoo
- 🟢RoboMME: Benchmarking and Understanding Memory for Robotic Generalist Policies
- 🟢Dynamic Chunking Diffusion Transformer
2026-03-06
- 🔴SkillNet: Create, Evaluate, and Connect AI Skills
- 🔴DARE: Aligning LLM Agents with the R Statistical Ecosystem via Distribution-Aware Retrieval
- 🟡RoboPocket: Improve Robot Policies Instantly with Your Phone
- 🟡Codex Security: now in research preview
- 🟡How Descript engineers multilingual video dubbing at scale
- 🟡How Balyasny Asset Management built an AI research engine
- 🟡MASQuant: Modality-Aware Smoothing Quantization for Multimodal Large Language Models
- 🟢HiFi-Inpaint: Towards High-Fidelity Reference-Based Inpainting for Generating Detail-Preserving Human-Product Images
- 🟢Interactive Benchmarks
- 🟢Large Multimodal Models as General In-Context Classifiers
2026-03-05
- 🔴Heterogeneous Agent Collaborative Reinforcement Learning
- 🟡Proact-VL: A Proactive VideoLLM for Real-Time AI Companions
- 🟡MemSifter: Offloading LLM Memory Retrieval via Outcome-Driven Proxy Reasoning
- 🟡ArtHOI: Articulated Human-Object Interaction Synthesis by 4D Reconstruction from Video Priors
- 🟢Memex(RL): Scaling Long-Horizon LLM Agents via Indexed Experience Memory
- 🟢Specificity-aware reinforcement learning for fine-grained open-world classification
- 🟢V_1: Unifying Generation and Self-Verification for Parallel Reasoners
2026-03-04
- 🔴Utonia: Toward One Encoder for All Point Clouds
- 🔴Beyond Language Modeling: An Exploration of Multimodal Pretraining
- 🟡BeyondSWE: Can Current Code Agent Survive Beyond Single-Repo Bug Fixing?
- 🟢Kling-MotionControl Technical Report
- 🟢How Controllable Are Large Language Models? A Unified Evaluation across Behavioral Granularities
- 🟢Next Embedding Prediction Makes World Models Stronger
2026-03-03
- 🔴OmniLottie: Generating Vector Animations via Parameterized Lottie Tokens
- 🔴From Scale to Speed: Adaptive Test-Time Scaling for Image Editing
- 🟡OpenAutoNLU: Open Source AutoML Library for NLU
- 🟢VGGT-Det: Mining VGGT Internal Priors for Sensor-Geometry-Free Multi-View Indoor 3D Object Detection
- 🟢CoVe: Training Interactive Tool-Use Agents via Constraint-Guided Verification
2026-03-02
- 🔴Enhancing Spatial Understanding in Image Generation via Reward Modeling
- 🟡Mode Seeking meets Mean Seeking for Fast Long Video Generation
- 🟡LK Losses: Direct Acceptance Rate Optimization for Speculative Decoding
- 🟢CiteAudit: You Cited It, But Did You Read It? A Benchmark for Verifying Scientific References in the LLM Era
- 🟢How to Take a Memorable Picture? Empowering Users with Actionable Feedback
- 🟢Compositional Generalization Requires Linear, Orthogonal Representations in Vision Embedding Models
- 🟢InfoNCE Induces Gaussian Distribution
2026-02-27
- 🔴The Trinity of Consistency as a Defining Principle for General World Models
- 🟡OmniGAIA: Towards Native Omni-Modal AI Agents
- 🟡OpenAI and Amazon announce strategic partnership
- 🟡Exploratory Memory-Augmented LLM Agent via Hybrid On- and Off-Policy Optimization
- 🟢Search More, Think Less: Rethinking Long-Horizon Agentic Search for Efficiency and Generalization
- 🟢MediX-R1: Open Ended Medical Reinforcement Learning
2026-02-26
- 🔴SkyReels-V4: Multi-modal Video-Audio Generation, Inpainting and Editing model
- 🔴MolHIT: Advancing Molecular-Graph Generation with Hierarchical Discrete Diffusion Models
- 🟡Pacific Northwest National Laboratory and OpenAI partner to accelerate federal permitting
- 🟡Solaris: Building a Multiplayer Video World Model in Minecraft
- 🟢ARLArena: A Unified Framework for Stable Agentic Reinforcement Learning
- 🟢Image Generation with a Sphere Encoder
- 🟢World Guidance: World Modeling in Condition Space for Action Generation
2026-02-25
- 🔴Query-focused and Memory-aware Reranker for Long Context Processing
- 🔴Test-Time Training with KV Binding Is Secretly Linear Attention
- 🟡PyVision-RL: Forging Open Agentic Vision Models via RL
- 🟡From Perception to Action: An Interactive Benchmark for Vision Reasoning
- 🟡Multi-Vector Index Compression in Any Modality
- 🟢QuantVLA: Scale-Calibrated Post-Training Quantization for Vision-Language-Action Models
- 🟢LongCLI-Bench: A Preliminary Benchmark and Study for Long-horizon Agentic Programming in Command-Line Interfaces
- 🟢See and Fix the Flaws: Enabling VLMs and Diffusion Models to Comprehend Visual Artifacts via Agentic Data Synthesis
2026-02-24
- 🔴A Very Big Video Reasoning Suite
- 🔴SkillOrchestra: Learning to Route Agents via Skill Transfer
- 🟡Agents of Chaos
- 🟡ManCAR: Manifold-Constrained Latent Reasoning with Adaptive Test-Time Computation for Sequential Recommendation
- 🟢Mobile-O: Unified Multimodal Understanding and Generation on Mobile Device
- 🟢Learning Cross-View Object Correspondence via Cycle-Consistent Mask Prediction
- 🟢DSDR: Dual-Scale Diversity Regularization for Exploration in LLM Reasoning
2026-02-23
- 🔴Does Your Reasoning Model Implicitly Know When to Stop Thinking?
- 🔴VESPO: Variational Sequence-Level Soft Policy Optimization for Stable Off-Policy LLM Training
- 🔴Generated Reality: Human-centric World Simulation using Interactive Video Generation with Hand and Camera Control
- 🟡EgoPush: Learning End-to-End Egocentric Multi-Object Rearrangement for Mobile Robots
- 🟡OpenAI announces Frontier Alliance Partners
- 🟢VidEoMT: Your ViT is Secretly Also a Video Segmentation Model
- 🟢SARAH: Spatially Aware Real-time Agentic Humans
- 🟢Sink-Aware Pruning for Diffusion Language Models
2026-02-20
- 🔴Unified Latents (UL): How to train your latents
- 🔴Mobile-Agent-v3.5: Multi-platform Fundamental GUI Agents
- 🔴SpargeAttention2: Trainable Sparse Attention via Hybrid Top-k+Top-p Masking and Distillation Fine-Tuning
- 🟡Computer-Using World Model
- 🟢Discovering Multiagent Learning Algorithms with Large Language Models
- 🟢Calibrate-Then-Act: Cost-Aware Exploration in LLM Agents
- 🟢"What Are You Doing?": Effects of Intermediate Feedback from Agentic LLM In-Car Assistants During Multi-Step Processing
- 🟢DDiT: Dynamic Patch Scheduling for Efficient Diffusion Transformers
2026-02-19
- 🔴SLA2: Sparse-Linear Attention with Learnable Routing and QAT
- 🔴AutoWebWorld: Synthesizing Infinite Verifiable Web Environments via Finite State Machines
- 🔴RynnBrain: Open Embodied Foundation Models
- 🟡CADEvolve: Creating Realistic CAD via Program Evolution
- 🟡Multi-agent cooperation through in-context co-player inference
- 🟢World Action Models are Zero-shot Policies
- 🟢Towards a Science of AI Agent Reliability
2026-02-18
- 🔴GLM-5: from Vibe Coding to Agentic Engineering
- 🔴SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks
- 🟡Does Socialization Emerge in AI Agent Society? A Case Study of Moltbook
- 🟡jina-embeddings-v5-text: Task-Targeted Embedding Distillation
- 🟢UniT: Unified Multimodal Chain-of-Thought Test-time Scaling
- 🟢Revisiting the Platonic Representation Hypothesis: An Aristotelian View
2026-02-17
- 🔴Experiential Reinforcement Learning
- 🔴DeepImageSearch: Benchmarking Multimodal Agents for Context-Aware Image Retrieval in Visual Histories
- 🟡STATe-of-Thoughts: Structured Action Templates for Tree-of-Thoughts
- 🟢InnoEval: On Research Idea Evaluation as a Knowledge-Grounded, Multi-Perspective Reasoning Problem
- 🟢Learning to Configure Agentic AI Systems
2026-02-16
- 🔴SQuTR: A Robustness Benchmark for Spoken Query to Text Retrieval under Acoustic Noise
- 🟡Zooming without Zooming: Region-to-Image Distillation for Fine-Grained Multimodal Perception
- 🟢CoPE-VideoLM: Codec Primitives For Efficient Video Language Models
- 🟢SemanticMoments: Training-Free Motion Similarity via Third Moment Features
- 🟢What does RL improve for Visual Reasoning? A Frankenstein-Style Analysis
2026-02-13
- 🔴The Devil Behind Moltbook: Anthropic Safety is Always Vanishing in Self-Evolving AI Societies
- 🔴Composition-RL: Compose Your Verifiable Prompts for Reinforcement Learning of Large Language Models
- 🟡GigaBrain-0.5M*: a VLA That Learns From World Model-Based Reinforcement Learning
- 🟡Beyond rate limits: scaling access to Codex and Sora
- 🟡MOSS-Audio-Tokenizer: Scaling Audio Tokenizers for Future Audio Foundation Models
- 🟢NarraScore: Bridging Visual Narrative and Musical Dynamics via Hierarchical Affective Control
- 🟢LawThinker: A Deep Research Legal Agent in Dynamic Environments
- 🟢Stroke of Surprise: Progressive Semantic Illusions in Vector Sketching
2026-02-12
- 🔴VidVec: Unlocking Video MLLM Embeddings for Video-Text Retrieval
- 🟡PhyCritic: Multimodal Critic Models for Physical AI
- 🟢When to Memorize and When to Stop: Gated Recurrent Memory for Long-Context Reasoning
- 🟢How Do Decoder-Only LLMs Perceive Users? Rethinking Attention Masking for User Representation Learning
- 🟢G-LNS: Generative Large Neighborhood Search for LLM-Based Automatic Heuristic Design
2026-02-11
- 🔴UI-Venus-1.5 Technical Report
- 🟡SkillRL: Evolving Agents via Recursive Skill-Augmented Reinforcement Learning
- 🟡Harness engineering: leveraging Codex in an agent-first world
- 🟢Agent World Model: Infinity Synthetic Environments for Agentic Reinforcement Learning
- 🟢Prism: Spectral-Aware Block-Sparse Attention
- 🟢Agent Banana: High-Fidelity Image Editing with Agentic Thinking and Tooling
- 🟢DLLM-Searcher: Adapting Diffusion Large Language Model for Search Agents
2026-02-10
- 🔴Weak-Driven Learning: How Weak Agents make Strong Agents Stronger
- 🟡Modality Gap-Driven Subspace Alignment Training Paradigm For Multimodal Large Language Models
- 🟡AIRS-Bench: a Suite of Tasks for Frontier AI Research Science Agents
- 🟢InternAgent-1.5: A Unified Agentic Framework for Long-Horizon Autonomous Scientific Discovery
- 🟢LLaDA2.1: Speeding Up Text Diffusion via Token Editing
- 🟢Recurrent-Depth VLA: Implicit Test-Time Compute Scaling of Vision-Language-Action Models via Latent Iterative Reasoning
- 🟢RLinf-USER: A Unified and Extensible System for Real-World Online Policy Learning in Embodied AI
2026-02-09
- 🔴AudioSAE: Towards Understanding of Audio-Processing Models with Sparse AutoEncoders
- 🔴On the Entropy Dynamics in Reinforcement Fine-Tuning of Large Language Models
- 🟡OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions
- 🟡Pisets: A Robust Speech Recognition System for Lectures and Interviews
- 🟢DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos
- 🟢Self-Improving World Modelling with Latent Actions
2026-02-06
- 🔴CAR-bench: Evaluating the Consistency and Limit-Awareness of LLM Agents under Real-World Uncertainty
- 🔴DFlash: Block Diffusion for Flash Speculative Decoding
- 🔴Spider-Sense: Intrinsic Risk Sensing for Efficient Agent Defense with Hierarchical Adaptive Screening
- 🟡Length-Unbiased Sequence Policy Optimization: Revealing and Controlling Response Length Variation in RLVR
- 🟡Context Forcing: Consistent Autoregressive Video Generation with Long Context
- 🟢Reinforced Attention Learning
- 🟢RISE-Video: Can Video Generators Decode Implicit World Rules?
- 🟢Reinforcement World Model Learning for LLM-based Agents
2026-02-04
- 🟡MARS: Modular Agent with Reflective Search for Automated AI Research
- 🟡3D-Aware Implicit Motion Control for View-Adaptive Human Video Generation
- 🟡Unlocking the Codex harness: how we built the App Server
- 🟡daVinci-Agency: Unlocking Long-Horizon Agency Data-Efficiently
- 🟢Research on World Models Is Not Merely Injecting World Knowledge into Specific Tasks
- 🟢Diversity-Preserved Distribution Matching Distillation for Fast Visual Synthesis
2026-02-03
- 🔴Green-VLA: Staged Vision-Language-Action Model for Generalist Robots
- 🟡Closing the Loop: Universal Repository Representation with RPG-Encoder
- 🟡UniReason 1.0: A Unified Reasoning Framework for World Knowledge Aligned Image Generation and Editing
- 🟢FS-Researcher: Test-Time Scaling for Long-Horizon Research Tasks with File-System-Based Agents
- 🟢SPARKLING: Balancing Signal Preservation and Symmetry Breaking for Width-Progressive Learning
- 🟢PixelGen: Pixel Diffusion Beats Latent Diffusion with Perceptual Loss
2026-02-02
- 🔴PaperBanana: Automating Academic Illustration for AI Scientists
- 🔴Quartet II: Accurate LLM Pre-Training in NVFP4 by Improved Unbiased Gradient Estimation
- 🟡Snowflake and OpenAI partner to bring frontier intelligence to enterprise data
- 🟡ReGuLaR: Variational Latent Reasoning Guided by Rendered Chain-of-Thought
- 🟡TTCS: Test-Time Curriculum Synthesis for Self-Evolving
- 🟢Causal World Modeling for Robot Control
- 🟢Do Reasoning Models Enhance Embedding Models?
- 🟢MemOCR: Layout-Aware Visual Memory for Efficient Long-Horizon Reasoning