Research
60 briefs in this section.
HarnessRisk benchmark tests agent harness safety
The arXiv benchmark evaluates safety failures across agent harness lifecycle phases.
Agent Lightning v1.0 targets harnessed agentic RL
The framework connects arbitrary agent harnesses to RL training through an LLM endpoint proxy.
Researchers test cross-model transfer for LLM memory
A paper studies moving frozen hashed memory between model backbones using target-side reader training.
Researchers propose DiSCO for safer text-to-image prompts
The black-box method aims to reduce NSFW outputs without model access or retraining.
PTXBench tests LLMs on architecture-specific GPU kernels
The benchmark finds uneven LLM performance on PTX optimization for H100 and B200 workloads.
LEGO-RL and ClawGym II target RL for agent harnesses
Two papers propose proxy-and-sandbox methods for training long-horizon coding and general agents.
FACET targets consistency in terminal task synthesis
The framework grounds instructions, solutions and verifiers in a repaired executable container state.
SkillForge targets repository-specific issue resolution
The framework distills project knowledge from synthetic issues tied to repository entities.
IBM Research tests calibrated memory for AI agents
ALTK-Evolve results show agent memory gains depend on model capability.
Researchers introduce TTP-D for truck-drone routing
The new optimization problem jointly models item selection, routing and drone synchronization.
Study finds audit-repair context makes LLM verifiers lenient
The arXiv paper reports lower false alarms after prior audit-repair episodes in model context.
AlphaEvolve helps lower matrix multiplication exponent bound
An arXiv note reports an improved upper bound of ω < 2.371177 using modern optimization methods.
SA-MRPO targets unsolved rewards in RL post-training
The method reweights multi-objective RL rewards by saturation instead of fixed scalar weights.
TRACE-Bench targets multi-reference image generation
The benchmark decomposes prompts into operators to diagnose model failures by capability.
Researchers target motion planning with generative methods
Two papers propose ways to make motion models more practical for robots and real-world video.
Researchers target hidden-state reasoning in video VLMs
New arXiv papers propose latent training methods to cut video reasoning overhead and improve visual grounding.
RUPA targets trajectory-level uncertainty in LLM agents
The arXiv paper proposes graph-based uncertainty propagation across agent execution histories.
New papers target long-horizon LLM agent training
Researchers propose ES, planning-aware RL, rubric credit assignment and routing methods for agentic LLMs.
New studies probe serving systems for agentic AI
Four papers examine latency, caching, traces and edge deployment as LLM workloads become more agentic.
Paper proposes common equation for graph neural networks
The framework maps GNN layers into seven components across several architectural families.
New datasets target compositional image and video editing
OpenGPT-4o-Image and CoinVE-200K address harder multi-step generation and editing tasks.
StartupBench tests agents on startup workflows
The benchmark uses market-validated AI product workflows to evaluate end-to-end agent performance.
Researchers propose aDSL for agentic 3D creation
The paper pairs an agent-centric DSL with a multi-agent loop to improve LLM-authored 3D programs.
Abra study maps scaling laws for diffusion image training
Researchers report predictable text-to-image scaling across 10^19 to 10^22 FLOPs.
GS-Voxel structures latents for aerial 3DGS generation
The method converts pre-optimized 3D Gaussian scenes into sparse voxel latents without per-scene fitting.
Paper proposes capability-centric image data design
The framework organizes image-generation supervision around capability dependencies and curriculum scheduling.
Researchers introduce ASI-Bench for autonomous research
The benchmark tests innovative exploration and scientific execution across research domains.
Study tests DeepSeek Harness prompt-injection resistance
Researchers report 14,560 controlled executions across 16 indirect-content channels.
Researchers target better VLM navigation with memory
Two papers propose memory-based navigation systems that better align visual prediction, planning and action.
EditBridge targets ultra-high-resolution image editing
The diffusion bridge method preserves HR source details while refining low-resolution edits.
Researchers target MoE latency in vision and decoding
New papers propose MoE designs for faster vision encoders, small-batch decoding and load balancing.
InternLM introduces Intern-S2-Mobius architecture
The model separates knowledge storage from reasoning and reports faster inference than Qwen3.5-35B.
R3-Bench tests LLM reasoning under shared budgets
The benchmark finds six models often underperform simple allocation baselines across reasoning suites.
Researchers target RL training for agent harnesses
LEGO-RL and ClawGym II propose proxy-and-sandbox approaches for long-horizon agent optimization.
GenRouter routes image prompts to lower-cost workflows
The framework standardizes agentic image pipelines and reports over 95% lower execution costs.
Ventor-QTest audits vendor-hosted LLM APIs
The black-box method tests hosted open-weight model routing without requiring API logprobs.
HarnessEval-W targets reasoning-based world model evaluation
The benchmark uses agents to produce evidence chains for visual world model assessments.
Study finds faster path for pixel-space diffusion models
Researchers report a latent-to-pixel recipe that matches or beats latent-space image models.
AnyTalk generates 3D speech animation without animation data
The method adapts video diffusion models to lip-sync arbitrary 3D characters.
HiFi-BRep targets more robust B-Rep generation
The framework uses topology-aware encoding and parallel decoding for CAD boundary representations.
Researchers target compositional image and video editing
New papers propose datasets and training methods for models that combine generation and editing tasks.
Researchers advance pixel-space diffusion for images
Two papers explore pixel-space diffusion for restoration and text-to-image generation.
Google applies homomorphic encryption to private AI
Google says encrypted processing can help make private AI more practical.
Google advances private AI with homomorphic encryption
Google says it is making private AI more practical using homomorphic encryption.
Google says it is making private AI practical
A Google Security post drew Hacker News discussion alongside broader AI critiques.
Researchers release LittleLearner sandbox for studying LMs
A 5B-parameter model and 88B-token Grade 5-limited corpus test how training scope bounds capabilities.
HN readers debate AI limits, privacy and drug discovery
Posts span AI drug discovery, math reasoning, private AI, Cloudflare and lab culture critiques.
AI scrutiny spans drug discovery, math and ChatGPT use
Three widely discussed reports assess what AI systems are doing well, and where evidence remains limited.
New papers target VLMs' spatial reasoning gap
Researchers propose runtime memory, RL training, benchmarks and 3D generation methods for spatial AI.
Google advances private AI with homomorphic encryption
Google says homomorphic encryption can make private AI more practical.
New papers probe spatial intelligence in AI vision models
SpaRRTa, SMA and PinpointQA target spatial reasoning gaps in visual and multimodal systems.
Papers refine pseudo-labeling for semi-supervised vision
New arXiv papers target label scarcity in HD mapping and semantic segmentation.
Anthropic finds AI agents can clash on shared tasks
Researchers say multi-agent systems showed conflict, collusion and coordination risks.
Researchers propose diffusion-based video dereflection
S2R combines simulated paired video data, a removal model and benchmark evaluation.
Paper argues agent safety needs runtime contracts
The arXiv paper says training-time alignment is insufficient for autonomous agents.
Study tests strong-to-weak transfer at inference time
A stronger model built task harnesses that lifted weaker-model scores on Theory-of-Mind benchmarks.
Researchers adapt agent harnesses for embodied AI
Thea and SHAPER papers focus on tool orchestration and train-free adaptation for embodied agents.
Paper bounds GPU opportunity in LLM-agent control
The study models when agent control paths expose enough concurrent work for GPU execution.
Researchers introduce Mechanist for AI interpretability
The agentic system aims to automate hypothesis generation and experiments on model mechanisms.