Research
Curated daily papers, lab research blogs and national AI research programs across the US, China, EU, UK, Japan and beyond.
Curated daily papers, lab research blogs and national AI research programs across the US, China, EU, UK, Japan and beyond.
Latest in Research
30 storiesCORE-BREW: LLR-Based Soft Decoding for Robust Multi-Bit LLM Watermarking
Reliable provenance for LLM outputs requires multi-bit watermarks that remain robust under editing while maintaining strict false-positive control. Existing ECC…
Deep Learning Approaches for 3D Medical Scene Completion: From Geometric Modeling to Generative Paradigms
Three-dimensional scene completion has evolved as a major problem in computer vision and robotics, and its applications are diverse, including autonomous naviga…
Project Ariadne: Prompt-Conditioned Route Generation for Synthesis Planning
Retrosynthetic planning seeks to connect a target molecule to commercially available starting materials through a multistep route. Classical planners construct …
Exploring the relationship between human-centric AI and firm idiosyncratic risks
Despite the extensive discussions of human-centric AI (HCAI) in Industry 5.0, its effects on firms' idiosyncratic risks (IR) remains underexplored. This is an i…
Neural Network-Based Parametric Model Reduction for Predicting Turbulent Flow for Different Vehicle Geometries
Numerical simulations in industrial applications often require performing numerous high-precision computations parameterized by specific experimental conditions…
Deep numerical schemes for systems of Ergodic BSDEs with applications to regime-switching forward utilities
In this paper, we introduce two neural-network-based numerical schemes for solving systems of coupled ergodic Backward Stochastic Differential Equations (eBSDEs…
Tractable Reasoning and Conjunctive Query Answering for Defeasible DL-Lite under Rational Closure
In Description Logics (DLs), reasoning under Rational Closure (RC) is a well-known and widely accepted non-monotonic formalism to handle defeasible knowledge. I…
Training-free Cross-domain Few-shot Segmentation via Robust Semantic Representation and Matching
Cross-domain Few-shot Segmentation (CD-FSS) aims to transfer knowledge learned from source domain to distinct target domains, segmenting unseen target classes w…
Prob-BBDM: a Probabilistic Brownian Bridge Diffusion Model for MRI sequence image-to-image translation
AI-driven image-to-image synthesis is rapidly advancing, with growing applications in medical imaging. Multi-modal image analysis plays a crucial role in optimi…
ZONOS2 Technical Report
We present ZONOS2 8B, our latest TTS model, which achieves state-of-the-art naturalness, prosody, and voice cloning fidelity. We improve upon Zonos-v0.1 across …
TIGER: Taming Identity, Geometry, and Generative Priors for High-Quality Face Video Restoration
Face Video Restoration (FVR) aims to recover high-fidelity facial videos from degraded input while preserving identity and semantic consistency across frames. E…
Managing Task Execution for Unknown Workloads in Batteryless IoT: A Hardware-Agnostic Evaluation
In recent years, the Internet of Things (IoT) paradigm has been shifting toward batteryless, energy-harvesting architectures. Sustaining reliable operation in t…
Accelerating Disaggregated RL for Visual Generative LLMs with Diffusion-Based Parallelism and Trainer-Assisted Generation
Reinforcement learning (RL) has become a dominant post-training paradigm, driving the emergence of high-performance RL systems such as veRL for autoregressive l…
Structural Kolmogorov-Arnold Convolutions: Learnable Function on the Values or the Filter Shape as Parameter-Efficient Alternative to Per-Edge Convolutional KANs
Convolutional Kolmogorov--Arnold Networks (KANs) replace the fixed weights of a convolutional kernel with learnable univariate functions. The dominant formulati…
ComputeFHE: A Privacy-Preserving General-Purpose Computation Library
Fully Homomorphic Encryption (FHE) enables computations to be performed directly on encrypted data while preserving data confidentiality. However, its practical…
Entity Resolution via Batched Oracle Queries
We consider an oracle that processes a limited batch of records at a time and clusters those that refer to the same real-world entity. We study how to interroga…
Cycle-Consistent Neural Explanation of Formal Verification Certificates
Formal verification produces machine-checkable certificates that attest to the satisfaction or violation of temporal properties, yet these certificates remain o…
Escaping the Self-Confirmation Trap: An Execute-Distill-Verify Paradigm for Agentic Experience Learning
Experience-driven self-evolution is critical for large language model (LLM) agents to improve through open-world interaction. However, existing experience learn…
Detecting AI Coding Agents in Open Source: A Validated Multi-Method Census of 180 Million Repositories
Generative AI coding agents are entering the open-source supply chain, yet their diverse and often invisible traces leave their prevalence poorly understood. We…
MedPCFM: Improving Medical Point Cloud Completion by Integrating Point Transformers and Flow Matching
Medical point cloud completion is important for anatomical reconstruction and downstream clinical workflows, yet generative modeling in this setting remains ins…
ReM-MoA: Reasoning Memory Sustains Mixture-of-Agents Scaling
Mixture-of-Agents (MoA) architectures improve inference-time scaling by organizing multiple LLM agents into layered reasoning pipelines. However, existing MoA v…
S1-Omni-Image: A Unified Model for Scientific Image Understanding, Generation, and Editing
We present S1-Omni-Image, an open-weight unified multimodal model for scientific image understanding, generation, and editing. Unlike general-purpose image gene…
CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference
Long-context large language model (LLM) inference is increasingly constrained by the memory footprint and decoding cost of key-value (KV) caches, limiting susta…
Advancing WordArt-Oriented Scene Text Recognition: Datasets and Methods
WordArt (artistic text) features highly customized fonts, textures, and layouts, making WordArt-oriented scene TExt Recognition (WATER) substantially more chall…
Red-Teaming the Agentic Red-Team
The use of agentic systems to perform offensive security operations has moved from a theoretical possibility to a commoditized capability. However, while the co…
AGORA: An Archive-Grounded Benchmark for Agentic Workplace Document Reasoning
Large language models are increasingly deployed as agents that reason over documents rather than answer from parametric knowledge. We study archive-grounded rea…
PointVG-R: Internalizing Geometric Reasoning in MLLMs for Precise Pointing Localization via Visual Chain of Thought
Pointing-based visual grounding requires models to precisely locate target objects by deciphering complex spatial relationships between the visual scene and poi…
To Compare, or Not to Compare: On Methodological Practices in Evaluating Social Bias
As Large Language Models are increasingly deployed in critical applications, robustly evaluating their social biases is paramount. However, the current literatu…
Uncertainty-Aware Longitudinal Forecasting of Alzheimer's Disease Progression Using Deep Learning
Longitudinal modelling of Alzheimer's disease progression is clinically useful only if it can describe not just the most likely next diagnosis, but how a patien…
Infinitesimal Causality
This paper introduces a categorical account of infinitesimal causality in Frobenius Markov categories equipped with tangent-bundle semantics. IDC captures the i…