Research
Curated daily papers, lab research blogs and national AI research programs across the US, China, EU, UK, Japan and beyond.
Curated daily papers, lab research blogs and national AI research programs across the US, China, EU, UK, Japan and beyond.
Latest in Research
30 storiesThemis: An explainable AI-enabled framework for Reinforcement Learning with Human Feedback
Training safe Reinforcement Learning (RL) systems is inherently challenging, with no guarantee of avoiding unwanted behaviors. The most effective defenses again…
The Warrant Gap: Claim-Conditioned Re-scoring for Fact-Checking
Fact-checking systems built on LLMs achieve high verdict accuracy on standard benchmarks, yet routinely output Supports labels whose cited evidence does not lic…
Extended pseudo-spectral physics-informed neural networks for phase-field models
Phase-field models play a central role in the continuum description of phase separation, in which the bulk free-energy density and the interfacial thickness par…
CN-NewsTTS Bench: a target-level automatic benchmark for raw-input Chinese news TTS pronunciation
Chinese news text contains dense written forms such as scores, hyphenated model names, ranges, unit symbols, percentages, English abbreviations, and mixed Chine…
UniDrive: A Unified Vision-Language and Grounding Framework for Interpretable Risk Understanding in Autonomous Driving
Recent multimodal large language models (MLLMs) have shown strong potential for autonomous driving scene understanding, yet existing methods still face a fundam…
Revealing Training Data Exposure in Vision Language Large Models via Parameter Gradients
Vision-Language Large Models (VLLMs) trained on massive crawled corpora raise pressing copyright and data-provenance concerns. These concerns are particularly a…
Grad Detect: Gradient-Based Hallucination Detection in LLMs
Large Language Models (LLMs) have demonstrated remarkable capabilities across diverse tasks, yet they remain prone to generating hallucinations. Detecting these…
Less is More: Quality-Aware Training Data Selection for Scientific Summarization
Scientific long-document summarization datasets commonly treat author-written abstracts as gold reference summaries, although their quality and alignment with t…
DiffusionBench: On Holistic Evaluation of Diffusion Transformers
Diffusion transformer (DiT) research on image generation has converged to a single evaluation setup: class-conditional generation on ImageNet. While methods imp…
GroundEval: A Deterministic Replacement for LLM-as-Judge in Stateful Agent Evaluation
Before letting an agent operate over real context, can you prove it used the right evidence? GroundEval turns that question into a deterministic test of what th…
Error Highways: Scaling Predictive Coding to Very Deep Networks
Predictive coding networks (PCNs) offer a biologically-plausible, local-learning alternative to back-propagation of errors (backprop). Nevertheless, they have r…
Learning Moral Diversity: Modelling Individual Perspectives in Moral Classification of Texts
Understanding moral values in social media text offers insight into moral judgement formation, and supervised NLP models trained on crowdsourced data have achie…
The Origins of Stochasticity: Comprehensive Investigations on Uncertainty Quantification for Large Language Models
Recent advancements in Large Language Models (LLMs) have enabled sophisticated reasoning and content generation, yet their inherent stochasticity poses signific…
Policy-as-Data: Learning Generalizable HOI Diffusion Models from Simulated Physics
Synthesizing realistic Human-Object Interactions (HOI) is critical for creating embodied avatars and functional virtual environments. However, current data-driv…
KaLM-Reranker-V1: Fast but Not Late Interaction for Compressed Document Reranking
As retrieval systems scale, high-quality reranking becomes increasingly important. However, most existing rerankers, whether encoder-based or decoder-based, joi…
RaMem: Contextual Reinstatement for Long-term Agentic Memory
Long-term memory has become increasingly important for LLM agents that operate across extended interactions and evolving task contexts. Recent memory systems ha…
AI Scientists as Engines of Discovery: A Case for Development within Reformed Institutions
Agentic artificial intelligence (AI) systems are beginning to assist, accelerate, and partially automate scientific discovery, performing tasks that span litera…
VideoLatent: Video-Language Learning via Latent Self-Forcing
Recent advancements in chain-of-thought (CoT) reasoning have shown promise in enhancing video understanding and reasoning capabilities of multimodal large langu…
SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning
Vision-language models (VLMs) are increasingly deployed in consumer, medical, financial, and enterprise applications. This broad deployment expands the safety s…
InteractiveAvatar: Real-Time Streaming Video Generation for Consistent and Intent-Aware Avatars
Recent diffusion-based models have enabled realistic audio-driven avatar generation in real-time streaming. However, existing approaches struggle to maintain vi…
Graph-Enhanced Large Language Models for Spatial Search
There have been many recent improvements in the ability of Large Language Models (LLMs) to perform complex tasks and answer domain-specific questions through te…
EEG Benchmarking Needs a Task Specification Layer: NeuroDoc for Rulebook-Guided, Executable Benchmark Construction
Electroencephalography (EEG) foundation models increasingly rely on multi-dataset training and evaluation, yet public EEG datasets still lack a shared task spec…
Hybrid Compression: Integrating Pruning and Quantization for Optimized Neural Networks
Deep neural networks have witnessed remarkable advancements in recent years and have become integral to various applications. However, alongside these developme…
Provable Benefits of RLVR over SFT for Reasoning Models: Learning to Backtrack Efficiently
Recent advances in large language models (LLMs) have demonstrated that reinforcement fine-tuning of pretrained base models can lead to significant gains in reas…
Understanding Knowledge Distillation in Post-Training: When It Helps and When It Fails
Large language models (LLMs) achieve strong performance across many tasks, but their high computational cost limits deployment in resource-constrained environme…
Neural Operator Processes for Probabilistic Operator Learning under Partial Observations
Neural operators learn mappings between function spaces, but are typically developed with dense input-output training fields and fully observed inputs at infere…
The Impact of VAE Design on Latent Pose Representations for Diffusion-based Sign Language Production
Latent diffusion approaches to sign language production (SLP) rely on an initial stage that learns an encoding of sign pose sequences, enabling generative model…
Subject-Level Unknown-Identity Identification from Leap Motion Controller 2 Hand Landmarks
This work studies subject recognition from Leap Motion Controller 2 (LMC2) hand landmark data under a subject-level unknown-identity identification protocol on …
Generalized nonparametric regression in reproducing kernel Hilbert spaces: Consistency and rates of convergence
We develop a comprehensive theory for regularized M-estimation in reproducing kernel Hilbert spaces. Under mild conditions on the loss we establish existence an…
From Point Estimates to Distributions: GMM Pooling for MIL in Preterm Birth Prediction
Preterm birth (PTB) prediction can enable targeted surveillance and timely intervention, yet most ultrasound-based models use a single selected transvaginal ult…