Research
Curated daily papers, lab research blogs and national AI research programs across the US, China, EU, UK, Japan and beyond.
Curated daily papers, lab research blogs and national AI research programs across the US, China, EU, UK, Japan and beyond.
Latest in Research
30 storiesAutoregressive Boltzmann Generators
Efficient sampling of molecular systems at thermodynamic equilibrium is a hallmark challenge in statistical physics. This challenge has driven the development o…

Understanding the brain with AI-driven explanations and experiments
Researchers introduce generative causal testing, which translates black box models into clear hypotheses and verifies them in the scanner, revealing what specif…
Multilingual Hematology Visual Question Answering Dataset
Vision Language Models (VLMs) have shown promising capabilities in medical image analysis by jointly understanding visual and textual information for tasks such…
FUTO Swipe: Layout-Agnostic Neural Swipe Decoding
Neural swipe decoders are typically tied to the keyboard they were trained on, requiring a new corpus and training run for each layout. In this report, we docum…
Automatic Generation of Highlights for Academic Paper Via Prompt-based Learning
Highlights provide a concise summary of the main contributions of an academic paper and help readers quickly understand its focus. However, many journals do not…
Pre-Warm: Input-Conditioned Weight Initialization for Convolutional Neural Networks
We introduce Pre-Warm, a simple yet effective zero-training-cost method for data-conditioned initialization of the first convolutional layer. Before the first f…
UC-Search: Risk-Aware Test-Time Search for Delayed Constrained Time-Series Control
Time-series models are usually scored as forecasters, yet deployed systems often require delayed decisions under uncertainty and hard feasibility constraints. U…
Physics Question Scene Graph: Fine-grained Evaluation of Physical Plausibility in Text-to-Video Generation
Video generation models are increasingly capable of producing realistic videos, but they still struggle to generate videos that follow basic physical laws. Comp…
REViT: Roto-reflection Equivariant Convolutional Vision Transformer
In this paper, we propose a discrete roto-reflection group equivariant vision transformer with convolutional attention. Roto-reflection equivariant networks pre…
Data-Driven Evolution of Library and Information Science Research Methods (1990-2022): A Perspective Based on Fine-grained Method Entities
Since the 1990s, advancements in big data and information technology have increasingly driven data-centric research in the field of Library and Information Scie…
Efficient Remote Sensing Instance Segmentation with Linear-Time State Space Distilled Visual Foundation Models
The computational complexity of Transformers scales quadratically with the number of tokens, which significantly constrains the efficiency of vision models, par…
Stagnant Neuron: Towards Understanding the Plasticity Loss in Multi-Agent Reinforcement Learning Value Factorization Methods
Multi-Agent Reinforcement Learning (MARL) value factorization methods can suffer from a loss of plasticity, gradually failing to adapt when transferring to new …
Lifelong In-Context Learning with Transformers Requires Parametric Forms of Attention
Lifelong continual learning remains an obstacle on the path to human-like intelligence. Modern transformers show sparks of intelligence with in-context learning…
Efficient and Trainable Language Model Test-Time Scaling via Local Branch Routing
Test-time scaling improves language-model reasoning, but existing approaches often face a difficult trade-off: long chain-of-thought sampling remains single-thr…
Three Buddhist Vocabularies: Computational Stylometry of the English Pali Canon across Sutta, Vinaya, and Abhidhamma
We present a computational stylometric analysis of the Tipitaka across all three Pitakas in English translation, extending earlier work on the Sutta Pitaka alon…
From Sounds to Scenes: A Benchmark for Evaluating Context-Aware Auditory Scene Understanding in Large Audio Language Models
Recent Large Audio Language Models (LALMs) have achieved remarkable progress in audio perceptual tasks across individual acoustic layers, including speech, soun…
Long-Term Simulation Exposes Cognitive-Developmental Risks in AI Companions
AI companions powered by large language models increasingly interact with cognition-developing users, including children and adolescents, creating risks that ma…
Beyond Next-Observation Prediction: Agent-Authored World Modeling for Sequential Decision Making
Recent studies on world modeling for Large Language Model (LLM) agents typically formulate the learning objective as next-observation prediction. However, this …
PRISM: Feed-Forward Single-Image 3D Reconstruction via Geometric Warp-Residual Modeling
Reconstructing 3D scenes from a single image is a fundamental challenge in computer vision, with broad applications in virtual reality, robotics, and content cr…
Brevity is the Soul of Inference Efficiency: Inducing Concision in VLMs via Data Curation
Inference efficiency is typically pursued by shrinking the model: distillation, pruning, quantization, and sparse routing each lower per-token cost while treati…
Reclaim Evaluation: A Lossy Memory Is Worse Than an Empty One
A language model's memory can be worse than having no memory at all. Give a model a memory that kept a wrong conclusion but dropped the work behind it, and it e…
How Reliable Is Your Jailbreak Judge? Calibration and Adversarial Robustness of Automated ASR Scoring
Almost every paper on LLM jailbreaks and prompt injection reports an attack-success rate (ASR), and that number is assigned not by people but by an automated ju…
Fault of Our Stars: Behavioral Drivers of Rating-Sentiment Incongruence
When people share experiences online, they often express thoughts in two ways: a star rating and a written review. In sentiment analysis, ratings are widely use…
Evaluating LLMs on Real-World Software Performance Optimization
Software performance optimization is a notoriously complex and manual task. Despite the growing use of Large Language Models (LLMs) for code refinement, we stil…
Disease-Centric Vision-Language Pretraining with Hybrid Visual Encoding for 3D Computed Tomography
Vision-language pre-training (VLP) holds great promise for general-purpose medical AI by leveraging radiology reports as rich textual supervision, yet existing …
Energy-Efficient CNN Acceleration with MSDF Digit-Serial Arithmetic on FPGA
This paper presents an energy-efficient hardware acceleration of the convolutional layers in the U-Net architecture for image segmentation, implemented on FPGA.…
VPA-Guard: Defending and Benchmarking Image-to-Video Generation Against Visual Prompt Attacks
Recent advancements in Image-to-Video (I2V) generation have transformed input images from simple appearance references into interactive control interfaces where…
Low-Complexity Policy Tessellations in Structured Markov Decision Processes
We study optimal-policy geometry in structured Markov decision processes. While approximate dynamic programming and reinforcement learning typically approximate…
Constraint Tax in Open-Weight LLMs: An Empirical Study of Tool Calling Suppression Under Structured Output Constraints
Tool Calling and Structured Output are two core capabilities of modern Agent systems, yet their interaction under joint deployment conditions remains insufficie…
Expresso-AI: Explainable Video-Based Deep Learning Models for Depression Diagnosis
Given the widespread prevalence of depression and its consequential impact on individuals and society, it is crucial to obtain objective measures for early diag…