SIGNAL
Tracking the global AI frontier — labs · research · agents · policy
Frontier Signal

Research

Curated daily papers, lab research blogs and national AI research programs across the US, China, EU, UK, Japan and beyond.

Curated daily papers, lab research blogs and national AI research programs across the US, China, EU, UK, Japan and beyond.

Latest in Research

30 stories
Offline Multi-agent Continual Cooperation via Skill Partition and Reuse
Research

Offline Multi-agent Continual Cooperation via Skill Partition and Reuse

Extracting skills from multi-agent offline dataset improves learning efficiency via sharing task-invariant coordination skills among tasks. In settings where ta…

TopoCast: A Topological Fidelity Framework for Evaluating Transformer-Based Time Series Forecasting
Research

TopoCast: A Topological Fidelity Framework for Evaluating Transformer-Based Time Series Forecasting

Deep learning-based models have achieved state-of-the-art performance in Time Series Forecasting (TSF), yet their evaluation remains dominated by pointwise erro…

C3-Bench: A Context-Aware Change Captioning Benchmark
Research

C3-Bench: A Context-Aware Change Captioning Benchmark

While Change Captioning systems have garnered substantial attention to respond to our evolving world, their true performance on diverse real-world change contex…

Optimizing Abstractive Summarization With Fine-Tuned PEGASUS
Research

Optimizing Abstractive Summarization With Fine-Tuned PEGASUS

Abstractive text summarization is the technique of generating a short and concise summary comprising the salient ideas of a source text without making a subset …

Low Variance Trust Region Optimization with Independent Actors and Sequential Updates in Cooperative Multi-agent Reinforcement Learning
Research

Low Variance Trust Region Optimization with Independent Actors and Sequential Updates in Cooperative Multi-agent Reinforcement Learning

Cooperative multi-agent reinforcement learning assumes each agent shares the same reward function and can be trained effectively using the Trust Region framewor…

KG-TRACE: A Neuro-Symbolic Framework for Mechanistic Grounding in Antimicrobial Resistance Prediction
Research

KG-TRACE: A Neuro-Symbolic Framework for Mechanistic Grounding in Antimicrobial Resistance Prediction

While WGS-based AMR prediction has reached high accuracy, existing models lack a mechanism to ground neural attributions in established biological pathways. We …

Pulmonary Embolism Risk Stratification from CTPA and Medical Records: Vascular Graphs Are Not All You Need
Research

Pulmonary Embolism Risk Stratification from CTPA and Medical Records: Vascular Graphs Are Not All You Need

Risk stratification for pulmonary embolism (PE) is critical for clinical decision-making. Stratification guidelines are based on patient medical records, parame…

DualEval: Joint Model-Item Calibration for Unified LLM Evaluation
Research

DualEval: Joint Model-Item Calibration for Unified LLM Evaluation

Current LLM evaluation relies on two complementary but often disconnected signals: static benchmarks with objective correctness labels and arena-style preferenc…

ProvenAI: Provenance-Native Traces of Evidence in Generated Answers
Research

ProvenAI: Provenance-Native Traces of Evidence in Generated Answers

Retrieval-augmented systems routinely present citations alongside generated answers, yet a citation does not confirm that the corresponding source meaningfully …

scBench-Long: Verifiable Benchmarking of Long-Horizon Single-Cell Biology
Research

scBench-Long: Verifiable Benchmarking of Long-Horizon Single-Cell Biology

Single-cell studies require analysts to convert raw measurements into specific biological claims through multi-step workflows and integration of metadata, assay…

Capacity-Controlled Multi-View Stylization of 3D Gaussian Splatting
Research

Capacity-Controlled Multi-View Stylization of 3D Gaussian Splatting

While 3D Gaussian Splatting (3DGS) provides an efficient and explicit representation for novel view synthesis, enforcing stylistic coherence across viewpoints r…

Diagnosing Task Insensitivity in Language Agents
Research

Diagnosing Task Insensitivity in Language Agents

Large language models can serve as capable long-horizon agents, but their out-of-distribution (OOD) generalization remains weak. We identify a key source of thi…

Cross-Head Attention Uplift Network with Inverse Propensity Score under Unobserved Confounding
Research

Cross-Head Attention Uplift Network with Inverse Propensity Score under Unobserved Confounding

Uplift modeling, crucial for estimating individual treatment effects (ITE), faces dual challenges: flexibly leveraging inter-group similarity to enhance discrim…

RecallRisk-BERT: A Multi-Task Framework for Post-Report Medical Device Recall Triage
Research

RecallRisk-BERT: A Multi-Task Framework for Post-Report Medical Device Recall Triage

Medical device recalls are a critical regulatory mechanism for protecting patient safety. The growing volume of FDA recall records presents challenges in post-r…

HarmVideoBench: Benchmarking Harmful Video Understanding in Large Multimodal Models
Research

HarmVideoBench: Benchmarking Harmful Video Understanding in Large Multimodal Models

Large vision-language models (LVLMs) have recently shown immense potential in automated content moderation, sparking growing interest in developing harmful-vide…

CORTEX: A Structured Reasoning Benchmark for Trustworthy 3D Chest CT MLLMs
Research

CORTEX: A Structured Reasoning Benchmark for Trustworthy 3D Chest CT MLLMs

Reasoning in multimodal large language models (MLLMs) has shown strong promise in medical imaging. However, this reasoning is usually free-form text judged only…

Don't Settle at the Mode! Mitigating Diversity Collapse in Pretrained Flow Models via Feature Self-Guidance
Research

Don't Settle at the Mode! Mitigating Diversity Collapse in Pretrained Flow Models via Feature Self-Guidance

State-of-the-art flow models generate stunning images from text or image prompts. However, they suffer from diversity collapse when generating multiple samples …

Utilizing Cognitive Signals Generated during Human Reading to Enhance Keyphrase Extraction from Microblogs
Research

Utilizing Cognitive Signals Generated during Human Reading to Enhance Keyphrase Extraction from Microblogs

Microblogging platforms generate massive amounts of short, noisy, and dispersed user content, making automatic keyphrase extraction (AKE) an important but chall…

Nemotron-TwoTower: Diffusion Language Modeling with Pretrained Autoregressive Context
Research

Nemotron-TwoTower: Diffusion Language Modeling with Pretrained Autoregressive Context

Diffusion language models offer a promising alternative to autoregressive models due to their potential for parallel and iterative generation. However, existing…

Humans Disengage, Reasoning Models Persist: Separating Difficulty Registration from Deliberation Allocation
Research

Humans Disengage, Reasoning Models Persist: Separating Difficulty Registration from Deliberation Allocation

Large reasoning models (LRMs) take longer on harder problems, just as humans do. This surface similarity hides an opposite pattern within items. When an LRM get…

DanceDuo: Bridging Human Movement and AI Choreography
Research

DanceDuo: Bridging Human Movement and AI Choreography

In recent years, advancements in deep learning and generative models have revolutionized music-driven dance generation. This paper introduces a novel platform, …

NeuraDock Visual Cognitive Load Agent Tutorial: A Quality-Gated Open-Source EEG Workflow for Alpha Dynamics and Real-Time Applications
Research

NeuraDock Visual Cognitive Load Agent Tutorial: A Quality-Gated Open-Source EEG Workflow for Alpha Dynamics and Real-Time Applications

This tutorial paper provides a step-by-step, reproducible walkthrough of NeuraDock Agent, an open-source EEG agent focused on Alpha dynamics and visual cognitiv…

Assessing Post-Reform Changes in Risk Disclosure Quality with a Multidimensional Text Analysis Approach
Research

Assessing Post-Reform Changes in Risk Disclosure Quality with a Multidimensional Text Analysis Approach

While corporate narrative disclosures provide crucial information to capital markets, comprehensively evaluating their qualitative changes over time remains cha…

Coarse-to-Fine: A Hybrid Self-Supervised Method for Non-rigid 3D Shape Matching
Research

Coarse-to-Fine: A Hybrid Self-Supervised Method for Non-rigid 3D Shape Matching

Non-rigid 3D shape matching is a fundamental task in computer vision and graphics. In this paper, we propose a hybrid self-supervised method based on a coarse-t…

Explainable Ensemble-Based Machine Learning Models for Detecting the Presence of Cirrhosis in Hepatitis C Patients
Research

Explainable Ensemble-Based Machine Learning Models for Detecting the Presence of Cirrhosis in Hepatitis C Patients

Hepatitis C is a liver infection caused by a virus, which results in mild to severe inflammation of the liver. Over many years, hepatitis C gradually damages th…

Temporally Consistent Label Interpolation for Robust Surgical Multi-Task Learning under Challenging Conditions
Research

Temporally Consistent Label Interpolation for Robust Surgical Multi-Task Learning under Challenging Conditions

Effective multi-task learning for surgical scene understanding is fundamentally hindered by annotation granularity mismatch; temporal workflow tasks such as pha…

FracEvent: Event-Camera Simulation via Fractional-Relaxation Pixel Dynamics
Research

FracEvent: Event-Camera Simulation via Fractional-Relaxation Pixel Dynamics

Event cameras asynchronously report brightness changes with microsecond-level temporal resolution, but real event data remain difficult to collect at scale beca…

CAT-Q: Cost-efficient and Accurate Ternary Quantization for LLMs
Research

CAT-Q: Cost-efficient and Accurate Ternary Quantization for LLMs

In this paper, we present CAT-Q, Cost-efficient and Accurate Ternary Quantization, for compressing and accelerating LLMs. Unlike existing state-of-the-art terna…

PersistentKV: Page-Aware Decode Scheduling for Long-Context LLM Serving on Commodity GPUs
Research

PersistentKV: Page-Aware Decode Scheduling for Long-Context LLM Serving on Commodity GPUs

Autoregressive large language model (LLM) serving is increasingly limited by key-value (KV) cache movement rather than dense matrix multiplication. Modern paged…

SKILL-DISCO: Distilling and Compiling Agent Traces into Reusable Procedural Skills
Research

SKILL-DISCO: Distilling and Compiling Agent Traces into Reusable Procedural Skills

Agents often repeatedly solve similar task instances from scratch, leading to unnecessary reasoning cost and long execution traces. Prior work has explored work…