
Researchers Shrink Depth-Detection AI to Run on Any Device
A new compact model achieves near-foundation-model accuracy while running 50x faster on phones and edge devices.
Papers, breakthroughs, benchmarks, and the long-arc trends shaping artificial intelligence. arXiv highlights and lab announcements, distilled.

A new compact model achieves near-foundation-model accuracy while running 50x faster on phones and edge devices.

New benchmark reveals current music transcription models still struggle with 38% accuracy, highlighting a major challenge for AI-driven music analysis.

Researchers develop self-learning framework that adapts surface computer vision for ocean environments, addressing a critical gap in marine robotics.

A new AI framework dramatically cuts computational costs for inverse design by merging deep learning with evolutionary algorithms.

New research enables smartphone-based heart attack detection in remote clinics without high-speed internet or powerful computers.

New training method doubles reasoning performance by having competing models evaluate one another's problem-solving approaches.

New technique makes reinforcement learning from human feedback practical by identifying which training steps actually matter.

New technique uses AI-generated code to read database files directly, sidestep traditional query engines, and achieve massive speedups on analytical workloads.

New dataset of 216,000+ verified skills aims to make autonomous AI systems more reliable and traceable.

Researchers introduce a method to filter noisy execution data and identify root causes of agent failures, boosting optimization by 40% on verification tasks.

New research reveals that operational guardrails, not just model architecture, fundamentally determine whether AI systems cooperate or exploit each other.

New analysis reveals how to simplify transformer attention mechanisms without sacrificing accuracy on large language models.

Researchers demonstrate how externalized memory improves language model accuracy while enabling human oversight of factual outputs.

SciReasoner bridges scientific accuracy and interpretability by treating atomic arrangements as inspectable evidence for AI reasoning.

ProxyPose uses video translation to simplify 6-DoF pose tracking, eliminating dependencies that have hindered computer vision systems for decades.

Researchers demonstrate that unified multimodal generation can consolidate diverse computer vision capabilities into one foundation model.

Lift3D-VLA combines spatial reasoning with action prediction to dramatically improve robotic task performance in simulation and the real world.

Researchers demonstrate that machine learning models can evaluate the accuracy of primate vocalizations and gestures without human-created reference standards.