SIGNAL
Tracking the global AI frontier — labs · research · agents · policy
Frontier Signal

Briefing

The fast feed: breaking AI news from the global tech press, deduplicated and time-ordered.

The fast feed: breaking AI news from the global tech press, deduplicated and time-ordered.

Latest in Briefing

30 stories
Zero-Shot Test-Time Canonicalization using Out-of-Distribution Scoring
Research

Zero-Shot Test-Time Canonicalization using Out-of-Distribution Scoring

Pretrained vision models often misclassify inputs that are rotated, scaled, or sheared, even though these affine transformations leave the object class unchange…

Supervised Reinforcement Learning for the Coordination of Distributed Energy Resources
Research

Supervised Reinforcement Learning for the Coordination of Distributed Energy Resources

The increasing integration of distributed energy resources (DERs) is crucial for power system decarbonization, yet unlocking DERs' flexibility is challenged by …

MEMPROBE: Probing Long-Term Agent Memory via Hidden User-State Recovery
Research

MEMPROBE: Probing Long-Term Agent Memory via Hidden User-State Recovery

Long-term memory promises LLM agents that grow more capable across sessions, maintaining an accurate, evolving understanding of the user that interaction forms.…

SER: Learning to Ground Video Reasoning with Semantic Evidence Rewards
Research

SER: Learning to Ground Video Reasoning with Semantic Evidence Rewards

Video MLLMs often struggle with fine-grained spatio-temporal reasoning, sometimes generating correct answers based on irrelevant frames or objects. Although out…

Forget to Improve: On-Device LLM-Agent Continual Learning via Budget-Curated Memory
Research

Forget to Improve: On-Device LLM-Agent Continual Learning via Budget-Curated Memory

On-device language-model agents improve by accumulating experience in retrieved memory rather than by updating weights. This memory is hard-bounded and exposed:…

Variational Inference via Entropic Transport Descent
Research

Variational Inference via Entropic Transport Descent

Particle-based variational inference (ParVI) methods approximate an intractable target distribution by evolving an ensemble of interacting samples. Existing app…

KidRisk: Benchmark Dataset for Children Dangerous Action Recognition
Research

KidRisk: Benchmark Dataset for Children Dangerous Action Recognition

Children are naturally energetic, and during their spontaneous activities, they often encounter potentially dangerous situations, especially when lacking parent…

Supervised Post-training of Speech Foundation Models for Robust Adaptation in Speech Deepfake Detection
Research

Supervised Post-training of Speech Foundation Models for Robust Adaptation in Speech Deepfake Detection

Large speech foundation models have shown strong potential for speech deepfake detection, but direct fine-tuning is limited by a mismatch between self-supervise…

Offline Multi-agent Continual Cooperation via Skill Partition and Reuse
Research

Offline Multi-agent Continual Cooperation via Skill Partition and Reuse

Extracting skills from multi-agent offline dataset improves learning efficiency via sharing task-invariant coordination skills among tasks. In settings where ta…

TopoCast: A Topological Fidelity Framework for Evaluating Transformer-Based Time Series Forecasting
Research

TopoCast: A Topological Fidelity Framework for Evaluating Transformer-Based Time Series Forecasting

Deep learning-based models have achieved state-of-the-art performance in Time Series Forecasting (TSF), yet their evaluation remains dominated by pointwise erro…

C3-Bench: A Context-Aware Change Captioning Benchmark
Research

C3-Bench: A Context-Aware Change Captioning Benchmark

While Change Captioning systems have garnered substantial attention to respond to our evolving world, their true performance on diverse real-world change contex…

Optimizing Abstractive Summarization With Fine-Tuned PEGASUS
Research

Optimizing Abstractive Summarization With Fine-Tuned PEGASUS

Abstractive text summarization is the technique of generating a short and concise summary comprising the salient ideas of a source text without making a subset …

Low Variance Trust Region Optimization with Independent Actors and Sequential Updates in Cooperative Multi-agent Reinforcement Learning
Research

Low Variance Trust Region Optimization with Independent Actors and Sequential Updates in Cooperative Multi-agent Reinforcement Learning

Cooperative multi-agent reinforcement learning assumes each agent shares the same reward function and can be trained effectively using the Trust Region framewor…

KG-TRACE: A Neuro-Symbolic Framework for Mechanistic Grounding in Antimicrobial Resistance Prediction
Research

KG-TRACE: A Neuro-Symbolic Framework for Mechanistic Grounding in Antimicrobial Resistance Prediction

While WGS-based AMR prediction has reached high accuracy, existing models lack a mechanism to ground neural attributions in established biological pathways. We …

Pulmonary Embolism Risk Stratification from CTPA and Medical Records: Vascular Graphs Are Not All You Need
Research

Pulmonary Embolism Risk Stratification from CTPA and Medical Records: Vascular Graphs Are Not All You Need

Risk stratification for pulmonary embolism (PE) is critical for clinical decision-making. Stratification guidelines are based on patient medical records, parame…

DualEval: Joint Model-Item Calibration for Unified LLM Evaluation
Research

DualEval: Joint Model-Item Calibration for Unified LLM Evaluation

Current LLM evaluation relies on two complementary but often disconnected signals: static benchmarks with objective correctness labels and arena-style preferenc…

ProvenAI: Provenance-Native Traces of Evidence in Generated Answers
Research

ProvenAI: Provenance-Native Traces of Evidence in Generated Answers

Retrieval-augmented systems routinely present citations alongside generated answers, yet a citation does not confirm that the corresponding source meaningfully …

scBench-Long: Verifiable Benchmarking of Long-Horizon Single-Cell Biology
Research

scBench-Long: Verifiable Benchmarking of Long-Horizon Single-Cell Biology

Single-cell studies require analysts to convert raw measurements into specific biological claims through multi-step workflows and integration of metadata, assay…

Capacity-Controlled Multi-View Stylization of 3D Gaussian Splatting
Research

Capacity-Controlled Multi-View Stylization of 3D Gaussian Splatting

While 3D Gaussian Splatting (3DGS) provides an efficient and explicit representation for novel view synthesis, enforcing stylistic coherence across viewpoints r…

Diagnosing Task Insensitivity in Language Agents
Research

Diagnosing Task Insensitivity in Language Agents

Large language models can serve as capable long-horizon agents, but their out-of-distribution (OOD) generalization remains weak. We identify a key source of thi…

Cross-Head Attention Uplift Network with Inverse Propensity Score under Unobserved Confounding
Research

Cross-Head Attention Uplift Network with Inverse Propensity Score under Unobserved Confounding

Uplift modeling, crucial for estimating individual treatment effects (ITE), faces dual challenges: flexibly leveraging inter-group similarity to enhance discrim…

RecallRisk-BERT: A Multi-Task Framework for Post-Report Medical Device Recall Triage
Research

RecallRisk-BERT: A Multi-Task Framework for Post-Report Medical Device Recall Triage

Medical device recalls are a critical regulatory mechanism for protecting patient safety. The growing volume of FDA recall records presents challenges in post-r…

HarmVideoBench: Benchmarking Harmful Video Understanding in Large Multimodal Models
Research

HarmVideoBench: Benchmarking Harmful Video Understanding in Large Multimodal Models

Large vision-language models (LVLMs) have recently shown immense potential in automated content moderation, sparking growing interest in developing harmful-vide…

CORTEX: A Structured Reasoning Benchmark for Trustworthy 3D Chest CT MLLMs
Research

CORTEX: A Structured Reasoning Benchmark for Trustworthy 3D Chest CT MLLMs

Reasoning in multimodal large language models (MLLMs) has shown strong promise in medical imaging. However, this reasoning is usually free-form text judged only…

Don't Settle at the Mode! Mitigating Diversity Collapse in Pretrained Flow Models via Feature Self-Guidance
Research

Don't Settle at the Mode! Mitigating Diversity Collapse in Pretrained Flow Models via Feature Self-Guidance

State-of-the-art flow models generate stunning images from text or image prompts. However, they suffer from diversity collapse when generating multiple samples …

S. 1003 Signed into Law
Policy

S. 1003 Signed into Law

On Friday, June 26, 2026, the President signed into law: S. 1003, the “Lulu’s Law,” which requires the Federal Communications Commission to is…

Track total merges by adoption phase in enterprise and organization reports
Create

Track total merges by adoption phase in enterprise and organization reports

Building on the AI adoption phase cohorts added to the Copilot usage metrics API, organization and enterprise reports now report the total number of pull reques…

Claude Code v2.1.195
Agents

Claude Code v2.1.195

An AI model programmed nonstop for 19 days on a single MirrorCode task that cost $2,600 to run
Briefing

An AI model programmed nonstop for 19 days on a single MirrorCode task that cost $2,600 to run

Epoch AI's new MirrorCode benchmark tests whether AI models can recreate complete programs without access to the original code. Claude Opus 4.7 leads with a 56 …

OpenAI's Claude Mythos competitor GPT-5.6 Sol launches under government-controlled access it calls unsustainable
Briefing

OpenAI's Claude Mythos competitor GPT-5.6 Sol launches under government-controlled access it calls unsustainable

OpenAI's new flagship GPT-5.6 Sol beats Anthropic's Claude Mythos 5 in coding benchmarks, but the US government is forcing a restricted rollout. OpenAI isn't ha…