Research Papers
567 signals across 91 briefings.
2026-08-17
- π‘Modular Cognitive Architecture Emerges in Large Language Models
- π‘Agents Catching Agents: Shortcut Cascades and Benchmark Gaming in Clinical Multi-Agent Systems
- π‘Nanbeige4.2-3B on Apple Silicon: Fixing Deployment Bugs and Decreasing Looped Transformer Memory Overhead
- π‘Amplified Does Not Mean Predictive: Reasoning Behaviors in Thinking Models
- π‘A Pathway to General-Purpose Scientific AI: Multimodal Comprehension of Scientific Images
- π’Is this Citation on Point?
2026-08-14
- π΄Context-Matched Distillation: Teacher Causality for Autoregressive Video Distillation
- π΄Specification-first convergence with an AI coding agent: a case study of dismantling a core architectural invariant across 189 files in a 717k-line codebase with no test oracle and no human code review
- π‘PixSDS: Why Latent SDS Makes Noisy Pixels
2026-08-13
- π΄Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence
- π΄Self-Evolving Embodied Agents via Skill-Harness Evolution
- π‘Gaze Target Estimation Anywhere with Concepts
- π‘Ready Cohorts: Bounding GPU Opportunity and Avoiding Host Round Trips in LLM-Agent Control
- π‘Poor Man's Agentic Modeling: Simulating Large LLM-Agent Societies on a Laptop
2026-08-12
- π΄AdvFD: Boosting Visual Generation via Adversarial Fr'echet Distance Loss
- π‘DistilVDR: A Compact End-to-End Visual Document Retriever via Dual-Student Distillation
- π‘InSight-doc: Agentic Visual Perception for Long-Document Understanding
- π‘Reference-Free Post-Training of Open Large Language Models for Multilingual Machine Translation
- π’Power law graph attention: exact generalization of scaled dot-product attention, empirical collapse at inference
2026-08-11
- π΄Business Arena: Benchmarking LLM Agents in a Realistic Marketplace
- π‘Cultivar: A Contrastive and Locale-Oriented Translation Benchmark for Investigating Contamination and Localisation Robustness
- π‘SymDiag: Explainable Diagnosis for LLM Reasoning via Neuro-Symbolic Verification
- π‘Don't Scroll Back: Missing-Evidence Memory for Streaming Dialogue Summarization
- π‘Gaming Without an Attacker: Benchmark Fingerprinting in LLM-Driven Search Under Selection Pressure
- π‘A Hybrid Nested Harness for Decoupling Structure and Parameters in LLM-Driven Optimization
2026-08-08
- π΄Interpretable MEG Decoding of Perceived Speech: Cortical Sources and the Stimulus Features That Drive Retrieval
- π΄DataSpace: Benchmarking Data Agents for Verifiable Analytics over Heterogeneous Workspaces
- π΄Activity Frames: Deterministic Screen-Activity Compilation for Agent Memory and Replay
- π‘KVAE: Family of Tokenizers for Multimodal Generative Models
- π‘MameLoshnLM: Yiddish Language Model and Evaluation Benchmark
- π‘Continual Learning in Transition
- π’Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills
- π’FactorJEPA: Factorizing Monolithic Futures into Layout-Agent-Interaction Channels for Crowded and Chaotic Global South Urban Worlds
- π’Task-Conditional Flow Matching for Balanced Multilingual Text Embedding Adaptation
- π’GaussianSelector: Lightweight Human-Guided Object Selection in 3D Gaussian Splatting with Graph Optimization
2026-08-07
- π΄DataSpace: Benchmarking Data Agents for Verifiable Analytics over Heterogeneous Workspaces
- π‘KVAE: Family of Tokenizers for Multimodal Generative Models
- π‘Activity Frames: Deterministic Screen-Activity Compilation for Agent Memory and Replay
- π‘MameLoshnLM: Yiddish Language Model and Evaluation Benchmark
- π‘Continual Learning in Transition
- π‘Task-Conditional Flow Matching for Balanced Multilingual Text Embedding Adaptation
2026-08-06
- π΄GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks
- π΄ABSeeker: Training Long-Horizon Search Agents via Answer-Backtracked Credit Assignment
- π΄FocusMem: Factorizing Content, Readout, and Trust in Latent GUI Memory
- π‘SIGNPOST-Bench: Benchmarking Text-Vision Conflict Resolution in Multimodal Large Language Models
- π‘Self-Evolving Coding Agents
2026-08-05
- π΄Know When to Stop: Segment-Level Credit Assignment for Reducing Overthinking
- π΄ARCHead: Activation-Metric Residual Correction for Large Language Model Output Heads
- π΄When Attention Goes Blind: Numerical Failure in ALiBi Positional Encodings
- π‘LegalPincite: Multi-level Legal Information Retrieval Dataset
- π‘Better, Stronger, Faster, and Broader: Structured All-Mask Prediction for MLLM-Based Segmentation
- π‘When Many Answers Are Valid, Voting Fails: Symbolic Verification for Best-of-K Causal Reasoning in LLMs
- π‘ChronoLens: Measuring Language Change Across Time, Languages, and Linguistic Levels
- π’CURV: Enhancing Chart Understanding Through Curriculum Visual Grounded Reasoning
- π’ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning
2026-08-04
- π‘Wnuan: Staged Post-Training for Question Answering over Proprietary Enterprise Knowledge
- π‘Loud or Silent? A Reusable Framework for Per-Modality Failure Analysis in Multimodal Clinical AI
- π‘Compute Globally, Materialize Locally: The Memory Contract of Sparse Event-KV
- π‘SG-WAM: Self-Guided World Modeling in Geometry-Aware Policy Space
- π‘GEOID-Flood: A Large-Scale Multi-Modal Benchmark Dataset for Flood Segmentation
2026-08-03
- π΄SAF-OPD: Stable Advantage Fusion for On-Policy Distillation
- π΄RL^2-VLA: Adaptive RL Latent Compositional Steering with Test-Time Scaling for Vision-Language-Action Models
- π‘Toward Robust and 3D-Aware RGB-NIR Imaging in the Dark
- π’Not All Tokens Deserve Equal Credit: Counterfactual Sensitivity Credit Reallocation for Long-CoT Reasoning
- π’In the Driver's Seat: A Multi-Company Study on the Reality of Autonomous Driving System Testing
- π’SGTP: Sampling-based Game-Theoretic Planning for Real-Time Multi-Vehicle Autonomous Racing
2026-07-31
- π΄ShadowDancer: Teaching Video World Models Any Action by Learning Unified Dynamics Representations from a Video and Its Shadow
- π΄Ξ²-OPSD: Deriving with Policy Optimization, Training with Self-Distillation
- π‘Ξ£-Mem: An Online Reliability Memory for LLM-based Multi-Agent Systems
- π‘Echoverse: Deep, Evolving Environments for Training Computer-Use Agents at Scale
- π‘Fairness Pruning: Locating Demographic Bias in GLU-MLP Layers via Differential Activations
- π‘Beyond Geometric Complementarity: Coherent Overlap in Sparse Mixture-of-Experts Routing
- π‘Is Deep Research Reliable? Misleading Knowledge Induces False Conclusions
- π’Pedestrian Archetypes Extension -- More Pedestrian Models for Autonomous Vehicle Safety Testing
2026-07-30
- π΄MindForge: Teaching Small Language Models Whole-Life-Cycle Software Engineering via Source-Free Program Synthesis
- π΄SpecFirst: Behavioral Specification Elicitation as a First-Class Step in Agent-Based Program Synthesis from Scratch
- π΄StatePlay: State-Aware Game World Models for Mechanics-Consistent Generation
- π‘Voice Memory for Agentic Speech Recognition
- π‘CADENCE: Closing the Reasoning Gap via Coverage-Adaptive On-Policy Distillation
- π‘SecRespond: Benchmarking AI Agents for Real-World Post-Compromise Incident Response
- π’DistillAlign: Coordinating Mode Covering and Mode Seeking in Autoregressive Video Distillation
- π’Grading the Narrators: An Isnad-Rijal Framework for Claim-Level Provenance in Multi-Agent Knowledge Systems
- π’ΟR^2: Reactive Real-time Flow Policies
2026-07-29
- π΄CodeNib: A Multi-View Data System for Serving Repository Context to Coding Agents
- π΄Reinforcement Learning for Code Optimization
- π΄How Fast Can Reward Models Score? A Systems Study of C++ and PyTorch Inference Runtimes for RLHF
- π΄Agent Retrieval Bench: Evaluating Repository Context Retrieval for Coding Agents
- π‘Human-in-the-Loop Signature Bootstrapping for UAV Hyperspectral PFM-1 Mine Detection
- π‘Uncovering Latent Reasoning Strategies in Language Models
- π‘OPERA: Offline Policy-guided Expert Routing and Adaptation for Universal Biomedical Image Analysis
- π’Edge-Aware Thermal Infrared UAV Swarm Tracking
2026-07-28
- π΄Sol-Attn: Accelerating Video Generation Inference via On-the-Fly Attention Sparsification
- π΄Chamaileon: Cross-Context Binder Design with Contextualized Modeling and Mixed Sampling
- π‘UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models
- π‘A Vocabulary for Multi-Agent Automated Research Systems
- π‘Memory Efficient Audio Synthesis with Decoupled Temporal Depth Diffusion Transformers
- π‘TRACE: Business Rule-Grounded Reasoning Curriculum for Knowledge-Preserving Parametric Tool Retrieval in Enterprise LLMs
- π‘Bitcoin Price Direction Prediction via Regime-Aware Multi-Modal Fusion of Social Sentiment and Technical Features
- π’Characterizing Warp Divergence from Pascal to Blackwell
2026-07-27
- π΄DataPrep-Bench: Benchmarking LLMs as Training Data Preparators
- π΄Agentic Context Management: Solving Agent Memory and Cost by Treating Them as Lifecycle and Architecture Problems
- π΄Interactive Training 2: Auditable Control Plane for Live Model Training
- π‘Three-Body Scattering for Generative Modeling
- π‘O-VAD: Industrial Video Anomaly Detection through Object-Centric Tracking and Reasoning
- π‘IDEAgent: Agentic Quality-Diversity Search for Research Idea Generation
- π‘Spectral Prior for Reducing Exposure Bias in Diffusion Models
- π‘Multi-Head Latent Control: A Unified Interface for LLM Agent Decision Making
- π’VisCo: Leveraging Large Language Models as Intrinsic Encoders for Visual Token Compression
- π’Multimodal Speaker Verification as a Threat to Speaker Anonymization
2026-07-26
- π΄K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs
- π΄SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation
- π΄LLMs Get Lost in Evolving User Intent
- π‘Self-Supervised Learning of Structured Dynamics from Videos
- π‘Color Pass-Through via Camera-Display Coupling
- π‘Sample-Efficient Learning from Agent Experience
- π’Multi-Turn On-Policy Distillation with Prefix Replay
- π’OpenForgeRL: Train Harness-native Agents in Any Environment
- π’FinanceComplexQA: Benchmarking Agentic Reasoning on Industrial-grade Financial Documents
- π’Dataset Distillation by Influence Matching
2026-07-24
- π΄Color Pass-Through via Camera-Display Coupling
- π΄LLMs Get Lost in Evolving User Intent
- π‘SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation
- π‘Self-Supervised Learning of Structured Dynamics from Videos
- π‘Sample-Efficient Learning from Agent Experience
- π‘OpenForgeRL: Train Harness-native Agents in Any Environment
- π’FinanceComplexQA: Benchmarking Agentic Reasoning on Industrial-grade Financial Documents
- π’Dataset Distillation by Influence Matching
2026-07-23
- π΄Scaling Laws for Hypernetwork-Based Knowledge Injection in Large Language Models
- π΄DocOps: A Verifiable Benchmark for Autonomous Agents in Complex Document Operations
- π΄FVAttn: Adaptive Sparse Attention with Runtime Load Balancing for Video Generation
- π‘ENTRAP-VL: A Taxonomic Probe for Dual Contextual Entrainment in Vision-Language Models
- π‘Train the Model, Not the Reader: Decodability Supervision for Verifiable Activation Explanations
- π‘Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning
- π’Moving Alphabet: A Controlled Study of Training Data for Text-to-Video Generation
- π’SLAM in Low-Light Environments: Project Report
2026-07-22
- π΄Text Template Tokens Are Implicit Semantic Registers in Diffusion Transformers
- π΄Generative World Renderer at the Speed of Play
- π‘AgentDebugX: An Open-Source Toolkit for Failure Observability, Attribution, and Recovery in LLM Agents
- π‘AutoIndex: Learning Representation Programs for Retrieval
- π‘Transcription Policy as a Latent Variable: Activating Controllable Verbatim ASR with Word-Level Timing
- π‘Where Should Optimizer State Live? Tiered State Allocation for Memory-Efficient Mixture-of-Experts Training
- π‘EduPanel: A Three-Agent LLM Judge for Teaching Videos -- Reliability, Complementarity, and Human Trust Calibration
- π’Computational Humor with Multimodal LLMs: Methods, Datasets, Evaluation, and Challenges
- π’Delineate Anything v2: A Global Foundation Model for Field Delineation
2026-07-21
- π΄Open-AoE: An Open Egocentric Manipulation Dataset and Toolchain for Embodied Learning
- π΄Do Language Models Dream of Binding Molecules? Benchmarking LLMs under Spatial Constraints
- π΄ShotPlan: Cinematic Video Generation with Learnable Planning Token
- π΄The Geometry of Semantic Space: A Continuous Geometric Framework for the Transformer Architecture
- π‘Masked Diffusion Language Models are Strong and Steerable Text-Based World Models for Agentic RL
- π‘ReViV: Reconstructing the Viewer and the View in 4D from Monocular Egocentric Video
- π‘Can Multimodal Large Language Models Understand OCT?
2026-07-20
- π΄RESOURCE2SKILL: Distilling Executable Agent Skills from Human-Created Multimodal Resources
- π΄When Does Muon Help Agentic Reinforcement Learning?
- π‘REBASE: Reference-Background Subspace Elimination for Training-Free In-Context Segmentation
- π‘Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents
- π‘Behavioral Privacy Leakage in Agentic Negotiation: Formalizing and Mitigating Inference Attacks via Randomized Policies
2026-07-17
- π΄LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget
- π΄VIABench: A Comprehensive Video Benchmark Collected from Blind Individuals for Visual Impairment Assistance
- π‘AsySplat: Efficient Asymmetric 3D Gaussian Splatting for Long-Sequence Scene Modeling
- π‘Token Time Continuous Diffusion for Language Modeling
- π‘RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination
- π‘SUFLECA: Scaling Up Feature Learning for CAD-to-image Alignment
- π’Chat2Scenic: An Iterative RAG-Based Framework for Scenario Generation in Autonomous Driving
- π’Hierarchical Denoising For Multi-Step Visual Reasoning
2026-07-16
- π΄Registers Matter for Pixel-Space Diffusion Transformers
- π΄Self-Improvements in Modern Agentic Systems: A Survey
- π΄Discrete Diffusion Models: A Unified Framework from Tokenization to Generation
- π‘Self in Space: Benchmarking Self-Awareness and Spatial Cognition in UAV Embodied Intelligence
- π‘From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World
- π‘Generative Compilation: On-the-Fly Compiler Feedback as AI Generates Code
- π’AffectFlow-DINO: Uncertainty-Aware Multi-Task Affect Estimation via Conditional Rectified Flow
2026-07-14
- π΄EgoSteer: A Full-Stack System Towards Steerable Dexterous Manipulation from Egocentric Videos
- π΄Metacognition in LLMs: Foundations, Progress, and Opportunities
- π΄Proxy Exploration and Reusable Guidance: A Modular LLM Post-Training Paradigm via Proxy-Guided Update Signals
- π‘Multi-Agent LLMs Fail to Explore Each Other
- π‘MET: Theory-Grounded and Culture-Aware Multilingual Moral Reasoning
- π‘Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model
- π‘Evidence-Backed Video Question Answering
- π’Latent-Identity Tuning in Text-to-Image Personalization Models
- π’A Theory of Contrastive Learning with Natural Images
2026-07-13
- π΄Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading
- π‘From RGB Generation to Dense Field Readout: Pixel-Space Dense Prediction with Text-to-Image Models
- π‘PanoWorld: Real-World Panoramic Generation
- π‘Phone Segmentation and Recognition through Phonological Activation Mapping
- π’VaseMuseum: Digital Intelligent Museum for Ancient Greek Pottery
2026-07-10
- π΄LongE2V: Long-Horizon Event-based Video Reconstruction, Prediction, and Frame Interpolation with Video Diffusion Models
- π΄Enhancing In-context Panoramic Generation via Geometric-aware Pretraining
- π‘DrugGen 2: A disease-aware language model for enhancing drug discovery
- π‘Linear Attention Architectures: Mechanisms, Trade-offs, and Cross-Layer Routing
- π‘Remember When It Matters: Proactive Memory Agent for Long-Horizon Agents
- π‘A Quantized Native Runtime for On-Device Semantic Audio Generation
- π’A Sparse and Truncated State Vector Simulator for Peaked Circuits
- π’SAM-MT: Real-Time Interactive Multi-Target Video Segmentation
2026-07-09
- π΄Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning
- π΄Sparse Delta Memory: Scaling the State of Linear RNNs through Sparsity
- π‘AgentLens: Production-Assessed Trajectory Reviews for Coding Agent Evaluation
- π‘RoboTALES: Learning Reasoning-Guided Robot Policies via Task-Aligned Simulated Futures
- π’Wake up for Touch! Mask-isolated Tactile Alignment Learning in MLLMs
- π’Imagined Rollouts are Kinematic, Not Dynamic: A Diagnosis of Long-Horizon World-Model Failure
2026-07-08
- π‘SWE-Review: Closing the Loop on Issue Resolution with Agentic Code Review
- π‘HunyuanOCR-1.5: Making Lightweight OCR VLMs Faster and Better
- π’SceneFrom3D: Geometry-Conditioned Outdoor 3D Scene Generation via View Scheduling with Object-Level Control
- π’SiamJEPA: On the Role of Siamese Student Encoders in JEPA
2026-07-07
- π΄Unified Audio Intelligence Without Regressing on Text Intelligence
- π‘Bridging Interleaved Multi-Modal Reasoning as a Unified Decision Process
- π‘Look Before You Leap: Distilling Tree Search into Action Evaluation for Frozen VLA Models
- π‘Speaker-Aware Temporal Aggregation Strategies on Segment Representations for Depression Detection in Dyadic Interaction: A Benchmark Study
- π‘Taste-aware music retrieval from audio embeddings
- π‘GaP: A Graph-as-Policy Multi-Agent Self-Learning Harness For Variational Automation Tasks
- π’SynCity 3000: Bootstrapping Scene-Scale 3D Diffusion
2026-07-06
- π΄VLA-Corrector: Lightweight Detect-and-Correct Inference for Adaptive Action Horizon
- π‘Securing the AI Agent: A Unified Framework for Multi-Layer Agent Red Teaming
- π’Interpretation-Oriented Cloud Removal via Observation-Anchored Residual Flow with Geo-Contextual Alignment
- π’AnyBokeh: Physics-Guided Any-to-Any Bokeh Editing with Optical Fingerprint Transfer
2026-07-03
- π΄AGVBench: A Reliability-Oriented Benchmark of Data Augmentation for Vein Recognition
- π΄InstanceControl: Controllable Complex Image Generation without Instance Labeling
- π‘WARP: Weight-Space Analysis for Recovering Training Data Portfolios
- π’Scaling Laws for Grid-Based Approximate Nearest Neighbor Search in High Dimensions
2026-07-02
- π΄Are Performance-Optimization Benchmarks Reliably Measuring Coding Agents?
- π΄Graph-Native Reinforcement Learning Enables Traceable Scientific Hypothesis Generation through Conceptual Recombination
- π‘Rank-Aware Hyperbolic Alignment for Vision-Language Dataset Distillation
- π‘SciIR: A Large-scale Training Dataset and Benchmark for Scientific Image Reasoning Generation
- π‘PixelEyes: Decoupling Perception and Reasoning for Pinpoint Visual Evidence Seeking
- π‘GRPO, Dr. GRPO, and DAPO Are Three Operations on One Number: The Group-Standard-Deviation Identity
- π‘Building to the Test: Coding Agents Deliver What You Check, Not What You Requested
- π’CogSENet: Blind Image Deblurring with Blur-Conditioned Semantic Routing and Explicit Frequency Fusion
2026-06-30
- π΄One Model, Many Latencies: Universal Speech Enhancement for Diverse Real-Time Applications
- π΄SWE-Together: Evaluating Coding Agents in Interactive User Sessions
- π΄RocketSmith: Agentic Additive Manufacturing of High-Powered Rockets
- π΄SAM2Matting: Generalized Image and Video Matting
- π‘A Gravitational Interpretation of Fine-Tuning Reversion
- π’MirrorPPR: Exemplar-Based Portrait Photo Retouching
2026-06-29
- π΄Qwen-RobotNav Technical Report: A Scalable Navigation Model Designed for an Agentic Navigation System
- π‘Simplified Sparse Attention via Gist Tokens
- π’CogniRoute: Learning to Route Social Evidence in Omni-Modal Models
- π’The Galaxy's Guide to the Tokenizer: A Benchmark for Scientific Foundation Models
- π’To Run or Not to Run: Analyzing the Cost-Effectiveness of Code Execution in LLM-Based Program Repair
- π’MemoBench: Benchmarking World Modeling in Dynamically Changing Environments
- π’How Much Static Structure Do Code Agents Need? A Study of Deterministic Anchoring
2026-06-26
- π΄LISA: Likelihood Score Alignment for Visual-condition Controllable Generation
- π‘Information-Aware KV Cache Compression for Long Reasoning
- π‘EO-WM: A Physically Informed World Model for Probabilistic Earth Observation Forecasting
- π‘When Does Combining Language Models Help? A Co-Failure Ceiling on Routing, Voting, and Mixture-of-Agents Across 67 Frontier Models
2026-06-25
- π΄Lite Any Stereo V2: Faster and Stronger Efficient Zero-Shot Stereo Matching
- π΄What Intermediate Layers Know: Detecting Jailbreaks from Entropy Dynamics
- π‘Physics Question Scene Graph: Fine-grained Evaluation of Physical Plausibility in Text-to-Video Generation
- π‘Do Thinking Tokens Help with Safety?
- π‘Distill Once, Adapt Life-Long: Exploring Dataset Distillation for Continual Test-Time Adaptation
2026-06-24
- π΄Critique of Agent Model
- π‘ChartWalker: Benchmarking the Cross-Chart RAG Task
- π‘QG-MIL: A Gated Transformer Aggregator for Domain-Agnostic Multiple Instance Learning in Medical Imaging
- π‘InSight: Self-Guided Skill Acquisition via Steerable VLAs
- π‘AGORA: An Archive-Grounded Benchmark for Agentic Workplace Document Reasoning
- π’EventVLA: Event-Driven Visual Evidence Memory for Long-Horizon Vision-Language-Action Policies
- π’Multi4D: High-Fidelity Dynamic Gaussian Splatting via Multi-Level Competitive Allocation
2026-06-23
- π΄Vera: A Layered Diffusion Model for Content-Preserving Video Editing
- π΄A Verifiable Search Is Not a Learnable Chain-of-Thought
- π‘When Agents Commit Too Soon: Diagnosing Premature Commitment in LLM Agents
- π‘TROPT: An Open Framework for Unifying and Advancing Discrete Text Optimization
- π‘Go-with-the-Track: Video Compositing and Motion Control with Point Tracking
- π‘Libretto: Giving LLM Agents a Sense of Musical Structure
- π‘Robusto-2: Benchmarking Humans & VLMs for Autonomous Driving in Lima & New York City
- π’Lift4D: Harmonizing Single-View 3D Estimation for 4D Reconstruction In-the-Wild
- π’ShotcreteDepth: A Bi-modal Dataset for Robust Robotic Depth Perception in Shotcrete Construction Environments
2026-06-21
- π΄Context-Aware RL for Agentic and Multimodal LLMs
- π΄The Data Manifold under the Microscope
- π΄Freeing the Law with LOCUS: A Local Ordinance Corpus for the United States
- π΄Rethinking Shrinkage Bias in LLM FP4 Pretraining: Geometric Origin, Systemic Impact, and UFP4 Recipe
- π‘LedgerAgent: Structured State for Policy-Adherent Tool-Calling Agents
- π‘LegalHalluLens: Typed Hallucination Auditing and Calibrated Multi-Agent Debate for Trustworthy Legal AI
- π’ReSyn: A Generalized Recursive Regular Expression Synthesis Framework
- π’Configurable Clinical Information Extraction with Agentic RAG: What Works, What Breaks, and Why
- π’Duration Aware Scheduling for ASR Serving Under Workload Drift
- π’The FID Lottery: Quantifying Hidden Randomness in Generative-Model Evaluation
2026-06-20
- π΄Context-Aware RL for Agentic and Multimodal LLMs
- π΄The Data Manifold under the Microscope
- π΄LedgerAgent: Structured State for Policy-Adherent Tool-Calling Agents
- π΄Freeing the Law with LOCUS: A Local Ordinance Corpus for the United States
- π‘Rethinking Shrinkage Bias in LLM FP4 Pretraining: Geometric Origin, Systemic Impact, and UFP4 Recipe
- π‘ReSyn: A Generalized Recursive Regular Expression Synthesis Framework
- π‘LegalHalluLens: Typed Hallucination Auditing and Calibrated Multi-Agent Debate for Trustworthy Legal AI
- π‘Configurable Clinical Information Extraction with Agentic RAG: What Works, What Breaks, and Why
- π‘Duration Aware Scheduling for ASR Serving Under Workload Drift
- π’The FID Lottery: Quantifying Hidden Randomness in Generative-Model Evaluation
2026-06-19
- π΄S-Agent: Spatial Tool-Use Elicits Reasoning for Spatial Intelligence
- π΄Playful Agentic Robot Learning
- π΄DragMesh-2: Physically Plausible Dexterous Hand-Object Interaction with Articulated Objects
- π‘JanusMesh: Fast and Zero-Shot 3D Visual Illusion Generation via Cross-Space Denoising
- π‘FlowBender: Feedback-Aware Training for Self-Correcting Conditional Flows
- π‘DF3DV-1K: A Large-Scale Dataset and Benchmark for Distractor-Free Novel View Synthesis
- π’HumanScale: Egocentric Human Video Can Outperform Real-Robot Data for Embodied Pretraining
- π’No Resource, No Benchmarks, No Problem? Evaluating and Improving LLMs for Code Generation in No-Resource Languages
2026-06-18
- π΄SAE Interventions are Unreliable: Post-Intervention Recovery of Suppressed Behavior
- π‘Native Active Perception as Reasoning for Omni-Modal Understanding
- π‘Sumi: Open Uniform Diffusion Language Model from Scratch
- π‘Learning User Simulators with Turing Rewards
- π’IndustryBench-MIPU: Benchmarking Multi-Image Attribute Value Extraction for Industrial Products
2026-06-17
- π΄ACE-Ego-0: Unifying Egocentric Human and Robotic Data for VLA Pretraining
- π΄LoopCoder-v2: Only Loop Once for Efficient Test-Time Computation Scaling
- π΄Learning from the Self-future: On-policy Self-distillation for dLLMs
- π‘Show the Signal, Hide the Noise: Spectral Forcing for Pixel-Space Diffusion
- π‘ProCUA-SFT Technical Report
- π‘ChLogic: Evaluating Robustness of Logical Reasoning in Chinese Expressions
- π’Variable-Width Transformers
- π’Text-Vision Co-Instructed Image Editing
- π’MotionVLA: Vision-Language-Action Model for Humanoid Motion
2026-06-16
- π΄DreamX-World 1.0: A General-Purpose Interactive World Model
- π΄Where Did It Go Wrong? Process-Level Evaluation of Web Agents with Semantic State Tracking
- π΄Memento: Reconstruct to Remember for Consistent Long Video Generation
- π΄GD^2PO: Mitigating Multi-Reward Conflicts via Group-Dynamic reward-Decoupled Policy Optimization
- π‘Hierarchical Advantage Weighting for Online RL Fine-Tuning of VLAs from Sparse Episode Outcomes
- π‘SP^3: Spherical Priors for Plug-and-Play Restoration
- π‘Who Flips? Self- and Cross-Model Counterarguments Reveal Answer Instability in LLMs
- π’MMDiff: Extending Diffusion Transformers for Multi-Modal Generation
- π’Selective Control under Noisy Perception: Governance Failures Hidden by Aggregate Metrics in Modular Networks
- π’PermaVid: Consistent Video Generation Across Edits via Disentangled Context Memory
2026-06-15
- π‘RhymeFlow: Training-Free Acceleration for Video Generation with Asynchronous Denoising Flow Scheduling
- π‘P3D-Bench: Benchmarking MLLMs for Parametric 3D Generation and Structural Reasoning
- π‘AdaSR: Adaptive Streaming Reasoning with Hierarchical Relative Policy Optimization
- π’WaveDiT: Distribution-Aware Wavelet Flow Matching for Efficient 3D Brain MRI Synthesis
2026-06-12
- π΄HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers
- π΄From 2D Grids to 1D Tokens: Reforming Shared Representations for Multimodal Image Fusion
- π‘Where, What, Why, and Importance: Structured Defect Grounding for Text-to-Image Feedback
- π‘ArogyaSutra: A Multi-Agent Framework for Multimodal Medical Reasoning in Indic Languages
- π’Revisiting Articulated Parts Perception in Robot Manipulation
- π’Leveraging Morphology for Historical Script Metrological Analysis
2026-06-11
- π΄Reason, Then Re-reason: Cross-view Revisiting Improves Spatial Reasoning
- π΄Grammar-Constrained Decoding Can Jailbreak LLMs into Generating Malicious Code
- π‘Time-Series Foundation Model Embeddings for Remaining Useful Life Estimation
- π’FlowLet: Conditional 3D Brain MRI Synthesis using Wavelet Flow Matching
2026-06-10
- π΄Kwai Keye-VL-2.0 Technical Report
- π΄Role-Agent: Bootstrapping LLM Agents via Dual-Role Evolution
- π‘Interpreting and Steering a Text-to-Speech Language Model with Sparse Autoencoders
- π‘UniPET: a universal network for high-quality PET image denoising across varied dose reduction factors
- π‘U-TTT: Towards Generalizable PET Image Denoising via Test-Time Training
2026-06-09
- π΄SwiftVR: Real-Time One-Step Generative Video Restoration
- π‘Reasoning Arena: Trace Tournaments When Verifiable Rewards Fall Short
- π‘Hardening Agent Benchmarks with Adversarial Hacker-Fixer Loops
- π‘Chiaroscuro Attention: Spending Compute in the Dark
- π‘Optical Reasoning: Rethinking Images as an Expressive Reasoning Medium Beyond Text
- π‘PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment
- π’Text-to-Image Models Need Less from Text Encoders Than You Think
2026-06-08
- π΄MMAE: A Massive Multitask Audio Editing Benchmark
- π΄AnchorWorld: Embodied Egocentric World Simulation with View-based Evolution Customization
- π‘Direct 3D-Aware Object Insertion via Decomposed Visual Proxies
- π‘Physics in 2-Steps: Locking Motion Priors Before Visual Refinement Erases Them
- π‘Entropy as a Structural Prior: How a Log-Barrier on DiT Belief Space Drives Musical Diversity and Development
- π‘LLM Explainability with Counterfactual Chains and Causal Graphs
- π’Streaming Video Generation with Streaming Force Control
2026-06-06
- π΄MAOAM: Unified Object and Material Selection with Vision-Language Models
- π‘AffordanceVLA: A Vision-Language-Action Model Empowering Action Generation through Affordance-Aware Understanding
- π’SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
- π’BRepCLIP: Contrastive Multimodal Pretraining on BRep Primitives for CAD Understanding
2026-06-05
- π΄AdaPlanBench: Evaluating Adaptive Planning in Large Language Model Agents under World and User Constraints
- π΄Complexity-Balanced Diffusion Splitting
- π‘LLMs Can Leak Training Data But Do They Want To? A Propensity-Aware Evaluation of Memorization in LLMs
- π‘Towards One-to-Many Temporal Grounding
- π‘Towards Truly Multilingual ASR: Generalizing Code-Switching ASR to Unseen Language Pairs
2026-06-04
- π΄OVO-S-Bench: A Hierarchical Benchmark for Streaming Spatial Intelligence in Multimodal LLMs
- π΄Reproducing, Analyzing, and Detecting Reward Hacking in Rubric-Based Reinforcement Learning
- π‘Unlocking Feature Learning in Gated Delta Networks at Scale
- π‘STRIDE: Training Data Attribution via Sparse Recovery from Subset Perturbations
- π‘Deep Embedded Multiplicative DMD for Algebra-Preserving Koopman Learning
2026-06-03
- π΄World Models Meet Language Models: On the Complementarity of Concrete and Abstract Reasoning
- π‘Decoupled Residual Denoising Diffusion Models for Unified and Data Efficient Image-to-Image Translation
- π‘Small RL Controller, Large Language Model: RL-Guided Adaptive Sampling for Test-Time Scaling
- π’PaddleOCR-VL-1.6: Expanding the Frontier of Document Parsing with Under-Optimized Region Refinement and Progressive Post-Training
- π’Ξ±Depth: Learning Single-Pass Soft Boundary Decomposition for Stereo Conversion
2026-06-02
- π΄LongLive-RAG: A General Retrieval-Augmented Framework for Long Video Generation
- π‘FineVerify: Scaling Test-Time Compute with Fine-Grained Self-Verification for Agentic Search
- π‘Can Predicted Dynamics Exist in the Physical World?
- π‘Adapting Multilingual Embedding Models to Turkish via Cross-Lingual Tokenizer Surgery and Offline Distillation
- π‘EVA01: Unified Native 3D Understanding and Generation via Mixture-of-Transformers
- π‘Confidence-Adaptive SwiGLU for Mixture-of-Experts
- π’ChartArena: Benchmarking Chart Parsing across Languages, Scenarios, and Formats
2026-06-01
- π΄Recovering Policy-Induced Errors: Benchmarking and Trajectory Synthesis for Robust GUI Agents
- π΄PEEK: Picking Essential frames via Efficient Knowledge distillation
- π‘Emergent Languages in Populations of Language Model Agents: From Token Efficiency to Oversight Evasion
- π‘Seeing Isn't Knowing: Do VLMs Know When Not to Answer Spatial Questions (and Why)?
- π‘DecMem: Towards Minute-Long Consistent World Generation with Decoupled Memory
- π‘iVGR: Internalizing Visually Grounded Reasoning for MLLMs with Reinforcement Learning
- π‘Benchmarking Composed Image Retrieval for Applied Earth Observation
- π’DRIFT: Decoupled Rollouts and Importance-Weighted Fine-Tuning for Efficient Multi-Turn Optimization
2026-05-29
- π΄UniSteer: Text-Guided Flow Matching in Activation Space for Versatile LLM Steering
- π΄When Cloud Agents Meet Device Agents: Lessons from Hybrid Multi-Agent Systems
- π΄RUBRIC-ARROW: Alternating Pointwise Rubric Reward Modeling for LLM Post-training in Non-verifiable Domains
- π‘PhyGenHOI: Physically-Aware 4D Generation of Dynamic Human-Object Interactions
- π‘Thinking Before Constraining: A Unified Decoding Framework for Large Language Models
- π’Geometry Matters: 3D Foundation Priors for Learning Semantic Correspondence
- π’Towards Consistent Video Geometry Estimation
2026-05-28
- π΄DenoiseRL: Bootstrapping Reasoning Models to Recover from Noisy Prefixes
- π‘GE-Sim 2.0: A Roadmap Towards Comprehensive Closed-loop Video World Simulators for Robotic Manipulation
- π‘SkillGrad: Optimizing Agent Skills Like Gradient Descent
- π‘ESC-Skills: Discovering and Self-Evolving Skills for Emotional Support Conversations
- π‘Revealing Algorithmic Deductive Circuits for Logical Reasoning
- π’Clark Hash: Stateless Sparse Johnson-Lindenstrauss Quantization for Neural Embeddings
2026-05-27
- π΄Share More, Search Less: Collaborative Parallel Thinking for Efficient Test-Time Scaling
- π‘Rethinking VLM Representation for VLA Initialization
- π‘VitaBench 2.0: Evaluating Personalized and Proactive Agents in Long-Term User Interactions
- π‘Learning to Act under Noise: Enhancing Agent Robustness via Noisy Environments
- π’Learning High-Frequency Continuous Action Chunks in Latent Space
2026-05-26
- π΄SkillEvolBench: Benchmarking the Evolution from Episodic Experience to Procedural Skills
- π΄Anticipate and Learn: Unleashing Idle-Time Compute in Proactive Agents
- π΄InstructSAM: Segment Any Instance with Any Instructions
- π‘CRONOS: Benchmarking Counterfactual Physical Consistency in Video Models
- π‘A Comprehensive Dataset for Human vs. AI Generated Image Detection
- π‘SimuWoB: Simulating Real-World Mobile Apps for Fast and Faithful GUI Agent Benchmarking
- π‘Reinforcing Few-step Generators via Reward-Tilted Distribution Matching
- π’HorizonStream: Long-Horizon Attention for Streaming 3D Reconstruction
- π’Pixel-Level Pavement Distress Assessment Using Instance Segmentation
2026-05-25
- π΄SciAtlas: A Large-Scale Knowledge Graph for Automated Scientific Research
- π΄See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding
- π‘PhotoFlow: Agentic 3D Virtual Photography Missions
- π‘RankE: End-to-End Post-Training for Discrete Text-to-Image Generation with Decoder Co-Evolution
- π‘SCOPE: Simulating Cross-game Operations in Playable Environments for FPS World Models
- π‘Good Token Hunting: A Hitchhiker's Guide to Token Selection for Visual Geometry Transformers
- π’Geo-Align: Video Generation Alignment via Metric Geometry Reward
2026-05-22
- π΄TransitLM: A Large-Scale Dataset and Benchmark for Map-Free Transit Route Generation
- π΄LatentOmni: Rethinking Omni-Modal Understanding via Unified Audio-Visual Latent Reasoning
- π‘Q-ARVD: Quantizing Autoregressive Video Diffusion Models
- π‘GenEvolve: Self-Evolving Image Generation Agents via Tool-Orchestrated Visual Experience Distillation
- π‘More Context, Larger Models, or Moral Knowledge? A Systematic Study of Schwartz Value Detection in Political Texts
- π’OmniPro: A Comprehensive Benchmark for Omni-Proactive Streaming Video Understanding
2026-05-21
- π΄You Only Need Minimal RLVR Training: Extrapolating LLMs via Rank-1 Trajectories
- π΄IndusAgent: Reinforcing Open-Vocabulary Industrial Anomaly Detection with Agentic Tools
- π‘Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMs
- π‘OCTOPUS: Optimized KV Cache for Transformers via Octahedral Parametrization Under optimal Squared error quantization
2026-05-20
- π΄When Vision Speaks for Sound
- π΄PixVerve: Advancing Native UHR Image Generation to 100MP with a Large-Scale High-Quality Dataset
- π‘TideGS: Scalable Training of Over One Billion 3D Gaussian Splatting Primitives via Out-of-Core Optimization
- π‘SAGA: A Sequence-Adaptive Generative Architecture for Multi-Horizon Probabilistic Forecasting with Adaptive Temporal Conformal Prediction
- π‘Omni-DuplexEval: Evaluating Real-time Duplex Omni-modal Interaction
- π‘Where Does Authorship Signal Emerge in Encoder-Based Language Models?
- π‘CopT: Contrastive On-Policy Thinking with Continuous Spaces for General and Agentic Reasoning
2026-05-19
- π΄Code-as-Room: Generating 3D Rooms from Top-Down View Images via Agentic Code Synthesis
- π΄SkillsVote: Lifecycle Governance of Agent Skills from Collection, Recommendation to Evolution
- π‘NGM: A Plug-and-Play Training-Free Memory Module for LLMs
- π‘A2RBench: An Automatic Paradigm for Formally Verifiable Abstract Reasoning Benchmark Generation
- π‘TOBench: A Task-Oriented Omni-Modal Benchmark for Real-World Tool-Using Agents
- π‘SafeDiffusion-R1: Online Reward Steering for Safe Diffusion Post-Training
- π‘Monitoring the Internal Monologue: Probe Trajectories Reveal Reasoning Dynamics
2026-05-18
- π΄PhysBrain 1.0 Technical Report
- π΄FashionChameleon: Towards Real-Time and Interactive Human-Garment Video Customization
- π’From Plans to Pixels: Learning to Plan and Orchestrate for Open-Ended Image Editing
- π’Unlocking Dense Metric Depth Estimation in VLMs
- π’OmniHumanoid: Streaming Cross-Embodiment Video Generation with Paired-Free Adaptation
2026-05-15
- π΄FrontierSmith: Synthesizing Open-Ended Coding Problems at Scale
- π΄Realiz3D: 3D Generation Made Photorealistic via Domain-Aware Learning
- π‘ViMU: Benchmarking Video Metaphorical Understanding
- π‘LiSA: Lifelong Safety Adaptation via Conservative Policy Induction
- π‘CurveBench: A Benchmark for Exact Topological Reasoning over Nested Jordan Curves
- π‘BEAM: Binary Expert Activation Masking for Dynamic Routing in MoE
2026-05-14
- π΄FrameSkip: Learning from Fewer but More Informative Frames in VLA Training
- π‘RealICU: Do LLM Agents Understand Long-Context ICU Data? A Benchmark Beyond Behavior Imitation
- π‘Offline Preference Optimization for Rectified Flow with Noise-Tracked Pairs
- π‘Vividh-ASR: A Complexity-Tiered Benchmark and Optimization Dynamics for Robust Indic Speech Recognition
- π‘PersonalAI 2.0: Enhancing knowledge graph traversal/retrieval with planning mechanism for Personalized LLM Agents
- π‘F-GRPO: Factorized Group-Relative Policy Optimization for Unified Candidate Generation and Ranking
- π’From Pixels to Concepts: Do Segmentation Models Understand What They Segment?
2026-05-13
- π΄Debiased Model-based Representations for Sample-efficient Continuous Control
- π΄A Causal Language Modeling Detour Improves Encoder Continued Pretraining
- π‘Multi-Stream LLMs: Unblocking Language Models with Parallel Streams of Thoughts, Inputs and Outputs
- π‘Pion: A Spectrum-Preserving Optimizer via Orthogonal Equivalence Transformation
- π‘WildRelight: A Real-World Benchmark and Physics-Guided Adaptation for Single-Image Relighting
2026-05-12
- π΄TMAS: Scaling Test-Time Compute via Multi-Agent Synergy
- π΄SlimSpec: Low-Rank Draft LM-Head for Accelerated Speculative Decoding
- π‘FORTIS: Benchmarking Over-Privilege in Agent Skills
- π‘LLiMba: Sardinian on a Single GPU -- Adapting a 3B Language Model to a Vanishing Romance Language
- π‘Path-Coupled Bellman Flows for Distributional Reinforcement Learning
- π‘SplatWeaver: Learning to Allocate Gaussian Primitives for Generalizable Novel View Synthesis
- π’Can Muon Fine-tune Adam-Pretrained Models?
- π’Training-Free Dense Hand Contact Estimation with Multi-Modal Large Language Models
- π’CapVector: Learning Transferable Capability Vectors in Parametric Space for Vision-Language-Action Models
- π’RoboMemArena: A Comprehensive and Challenging Robotic Memory Benchmark
2026-05-11
- π΄MACE-Dance: Motion-Appearance Cascaded Experts for Music-Driven Dance Video Generation
- π‘Rethinking State Tracking in Recurrent Models Through Error Control Dynamics
- π‘Gated QKAN-FWP: Scalable Quantum-inspired Sequence Learning
- π’Sparse Autoencoders as Plug-and-Play Firewalls for Adversarial Attack Detection in VLMs
- π’R^3-SQL: Ranking Reward and Resampling for Text-to-SQL
2026-05-08
- π΄When to Trust Imagination: Adaptive Action Execution for World Action Models
- π΄RemoteZero: Geospatial Reasoning with Zero Human Annotations
- π‘When No Benchmark Exists: Validating Comparative LLM Safety Scoring Without Ground-Truth Labels
- π‘Are We Making Progress in Multimodal Domain Generalization? A Comprehensive Benchmark Study
- π’TIDE: Every Layer Knows the Token Beneath the Context
- π’Sparkle: Realizing Lively Instruction-Guided Video Background Replacement via Decoupled Guidance
2026-05-07
- π΄Stream-R1: Reliability-Perplexity Aware Reward Distillation for Streaming Video Generation
- π΄Stream-T1: Test-Time Scaling for Streaming Video Generation
- π΄OpenSearch-VL: An Open Recipe for Frontier Multimodal Search Agents
- π‘Lightning Unified Video Editing via In-Context Sparse Attention
- π‘Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation
- π‘When to Think, When to Speak: Learning Disclosure Policies for LLM Reasoning
- π’Parameter-Efficient Multi-View Proficiency Estimation: From Discriminative Classification to Generative Feedback
- π’Diffusion Model as a Generalist Segmentation Learner
2026-05-06
- π΄PatRe: A Full-Stage Office Action and Rebuttal Generation Benchmark for Patent Examination
- π΄X2SAM: Any Segmentation in Images and Videos
- π‘StateSMix: Online Lossless Compression via Mamba State Space Models and Sparse N-gram Context Mixing
- π‘The TTS-STT Flywheel: Synthetic Entity-Dense Audio Closes the Indic ASR Gap Where Commercial and Open-Source Systems Fail
- π‘ESARBench: A Benchmark for Agentic UAV Embodied Search and Rescue
- π’Skills-Coach: A Self-Evolving Skill Optimizer via Training-Free GRPO
2026-05-05
- π΄MolmoAct2: Action Reasoning Models for Real-world Deployment
- π΄Persistent Visual Memory: Sustaining Perception for Deep Generation in LVLMs
- π΄Repetition over Diversity: High-Signal Data Filtering for Sample-Efficient German Language Modeling
- π‘OceanPile: A Large-Scale Multimodal Ocean Corpus for Foundation Models
- π‘Graph Rewiring in GNNs to Mitigate Over-Squashing and Over-Smoothing: A Survey
- π‘Representation in large language models
- π‘AcademiClaw: When Students Set Challenges for AI Agents
- π‘Hierarchical Abstract Tree for Cross-Document Retrieval-Augmented Generation
- π’Code World Model Preparedness Report
- π’Motion-Aware Caching for Efficient Autoregressive Video Generation
- π’Generative Modeling with Orbit-Space Particle Flow Matching
- π’Perceptual Flow Network for Visually Grounded Reasoning
2026-05-04
- π΄UniVidX: A Unified Multimodal Framework for Versatile Video Generation via Diffusion Priors
- π΄Learning while Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies
- π΄From Skill Text to Skill Structure: The Scheduling-Structural-Logical Representation for Agent Skills
- π‘Web2BigTable: A Bi-Level Multi-Agent LLM System for Internet-Scale Information Search and Extraction
- π‘Online Self-Calibration Against Hallucination in Vision-Language Models
- π‘LASE: Language-Adversarial Speaker Encoding for Indic Cross-Script Identity Preservation
- π‘Learning to Act and Cooperate for Distributed Black-Box Consensus Optimization
- π‘AnalogRetriever: Learning Cross-Modal Representations for Analog Circuit Retrieval
- π’Talker-T2AV: Joint Talking Audio-Video Generation with Autoregressive Diffusion Modeling
2026-05-03
- π΄Efficient Training on Multiple Consumer GPUs with RoundPipe
- π΄Claw-Eval-Live: A Live Agent Benchmark for Evolving Real-World Workflows
- π΄Nemotron 3 Nano Omni: Efficient and Open Multimodal Intelligence
- π‘Step-level Optimization for Efficient Computer-use Agents
- π‘Compliance versus Sensibility: On the Reasoning Controllability in Large Language Models
- π‘Learning from Noisy Preferences: A Semi-Supervised Learning Approach to Direct Preference Optimization
- π’Instruction-Guided Poetry Generation in Arabic and Its Dialects
- π’ViPO: Visual Preference Optimization at Scale
- π’FlashRT: Towards Computationally and Memory Efficient Red-Teaming for Prompt Injection and Knowledge Corruption
- π’Safety Drift After Fine-Tuning: Evidence from High-Stakes Domains
2026-05-02
- π΄Efficient Training on Multiple Consumer GPUs with RoundPipe
- π΄Claw-Eval-Live: A Live Agent Benchmark for Evolving Real-World Workflows
- π΄Nemotron 3 Nano Omni: Efficient and Open Multimodal Intelligence
- π‘Step-level Optimization for Efficient Computer-use Agents
- π‘Compliance versus Sensibility: On the Reasoning Controllability in Large Language Models
- π‘Instruction-Guided Poetry Generation in Arabic and Its Dialects
- π‘Learning from Noisy Preferences: A Semi-Supervised Learning Approach to Direct Preference Optimization
- π’ViPO: Visual Preference Optimization at Scale
- π’FlashRT: Towards Computationally and Memory Efficient Red-Teaming for Prompt Injection and Knowledge Corruption
- π’Safety Drift After Fine-Tuning: Evidence from High-Stakes Domains