Research
Curated daily papers, lab research blogs and national AI research programs across the US, China, EU, UK, Japan and beyond.
Curated daily papers, lab research blogs and national AI research programs across the US, China, EU, UK, Japan and beyond.
Latest in Research
30 storiesLeveraging Similarities in Multi-Armed Bandits
In many online learning and bandit problems, the actions we consider possess inherent similarities--for instance because they share latent traits, tags, or hier…
Selective Time Series Forecasting via Metalearning
Deep learning methods have achieved state-of-the-art in time series forecasting, yet their accuracy varies considerably across samples, as some instances remain…
SPIRAL: Learning to Search and Aggregate
Language model reasoning can be substantially improved at test time via scaffolds that scale inference compute across different primitives -- sequential reasoni…
Sentence-Level Contextual Entrainment in Large Language Models
Contextual entrainment, which is a newly discovered phenomenon in large language models (LLMs), refers to the tendency of a model to assign higher probabilities…
Flood Mapping from RGB imagery using a Vision Foundation Model
Timely, high-resolution maps of flood extent around settlements are essential for emergency response and damage assessment. We consider airborne RGB imagery for…
A Dual Edge Spatial Jacobian Image Graph for Interpretable Diabetic Retinopathy Grading
Automated diabetic retinopathy (DR) grading from colour fundus photographs can achieve strong predictive performance, but clinical interpretation requires more …
EERLoss: A Novel Loss Function for Training Deep Biometric Models. A Case Study in Keystroke Dynamics
Deep learning approaches to biometric verification are commonly trained by optimizing indirect objectives, creating a misalignment between the optimization proc…
Cost-Optimal Decision Diagrams for Stochastic Boolean Function Evaluation
In many decision-making scenarios, acquiring information incurs different costs. We consider the problem of constructing a deterministic evaluation strategy tha…
OrbitForge: Text-to-3D Scene Generation via Reconstruction-Anchored Video Synthesis
Generic text-to-video models can be used as rich open-world scene priors. Despite the high quality of today's generated videos, they do not directly yield relia…
Real vs. Complex Spectral Bases for Neural Operators: The Role of Green's Function Alignment
Fourier Neural Operators (FNO) learn solution operators of partial differential equations by parameterizing global convolutions in the complex Fourier domain. F…
Cage-based Texture Transfer with Geometric Filtering
Real-time texture transfer expands the creative horizon for interactive applications, enabling seamless detail projection in scenarios that range from digital c…
MJEPA: A Simple and Scalable Joint-Embedding Predictive Architecture for Audio-Visual Learning
Self-supervised learning from large-scale video data has emerged as a dominant paradigm for visual representation learning. Since audio and visual streams natur…
Communicability-Inspired Positional Encoding (CIPE)
Positional encodings (PEs) are essential for Transformers. Yet designing effective PEs for non-Euclidean graphs remains challenging. Such encodings should ideal…
Sarashina2.2-TTS: Tackling Kanji Polyphony in Japanese Speech Generation via Data Scaling and Targeted Data Synthesis
While large language model (LLM)-based text-to-speech (TTS) systems have achieved high-quality speech synthesis, most existing systems focus on English and Chin…
Conformal Recovery-Deadline Certificates for Runtime Assurance of Adapting Controllers
Runtime assurance (RTA) protects a safety-critical system by switching from an advanced controller to a verified safe controller when a monitored condition is v…
DFMU: Data-Frugal Machine Unlearning
Machine unlearning is an emerging domain that ensures the safe removal of elements (includes concepts, attributes, entity and class) from the trained model alon…
PolicyAlign: Direct Policy-Based Safety Alignment for Large Language Models
Safety alignment of large language models (LLMs) typically depends on high-quality supervision data, such as safe demonstrations or preference pairs. However, i…
Does Translation-Enhanced Speech Encoder Pre-training Affect Speech LLMs?
Connecting a pre-trained speech encoder to a Large Language Model (LLM) is the standard architecture for building Speech LLMs. However, a structural misalignmen…
Overview of HIPE-2026: Person-Place Relation Extraction from Multilingual Historical Texts
Was this person ever at that place, and if so, when? Answering such questions from noisy, multilingual historical documents is the central challenge of HIPE-202…
Otter Weather: Skillful and Computationally Efficient Medium-Range Weather Forecasting
State-of-the-art medium-range AI weather models can outperform traditional Numerical Weather Prediction (NWP) but require massive training budgets. This restric…
From Hallucination to Grounding: Diagnosing Visual Spatial Intelligence via CRISP
Current VLM evaluations often conflate language priors with genuine spatial reasoning. To address this, we introduce CRISP, a novel structural-diagnostic evalua…
Extracting Neural Materials from Multi-view Images
Neural materials can represent complex specular reflections and scattering effects in a compact, universal basis. However, acquiring and authoring such material…
Native space based pipelines outperform template space based pipeline in subcortical segmentation
Accurate segmentation of subcortical regions is critical for neurosurgical planning and functional research. Most automated methods rely on template space coreg…
Adversarial Domain Prompt Tuning and Generation for Single Domain Generalization
Single domain generalization (SDG) aims to learn a robust model, which could perform well on many unseen domains while there is only one single domain available…
Training the Orchestrator: A Supervised Approach to End-to-End PDDL Planning with LLM Agents
Translating natural-language planning intent into verified plans is a longstanding challenge: people communicate goals in language, while classical planners req…
Reliability-Guided Adaptive Ensembling for Robust Test-Time Adaptation
Test-time adaptation (TTA) can mitigate domain shift without source data, but it is highly brittle under adversarially contaminated test streams, where corrupte…
Enhancing Road Safety: An IoT-Based Accident Detection and Prevention Mechanism
Road traffic accidents remain a critical global crisis, consistently serving as a primary driver of preventable mortality and severe injury. These incidents are…
Curvature-Adaptive Consistency Flow Matching: Autonomous Trajectory Optimization via Reinforcement Learning
Consistency distillation has significantly accelerated the inference of diffusion models. In this work, we reveal an intriguing asymmetry: while Logit-Normal sa…
Code Isn't Memory: A Structural Codebase Index Inside a Coding Agent
Coding agents now interleave LLMs with retrieval over the working repository, and retrieval implementations vary widely across deployed harnesses. Inside a fixe…
Curvature-aware 3D length estimation of greenhouse cucumbers using RGB-D imaging and cubic spline arc-length integration
Commercial greenhouse cucumber production is graded by fruit length, which drives harvest scheduling, labour allocation, and logistics. Manual measurement with …