本篇博文主要内容为 2026-10-05 从Arxiv.org论文网站获取的最新论文列表,自动更新,按照NLP、CV、ML、AI、IR、MA六个大方向区分。

说明:每日论文数据从Arxiv.org获取,每天早上12:30左右定时自动更新。

提示: 当天未及时更新,有可能是Arxiv当日未有新的论文发布,也有可能是脚本出错。尽可能会在当天修复。

目录

概览 (2026-10-05)

今日共更新882篇论文,其中:

  • 自然语言处理共114篇(Computation and Language (cs.CL))
  • 人工智能共267篇(Artificial Intelligence (cs.AI))
  • 计算机视觉共130篇(Computer Vision and Pattern Recognition (cs.CV))
  • 机器学习共299篇(Machine Learning (cs.LG))
  • 多智能体系统共17篇(Multiagent Systems (cs.MA))
  • 信息检索共11篇(Information Retrieval (cs.IR))
  • 人机交互共19篇(Human-Computer Interaction (cs.HC))

多智能体系统

[MA-0] EdgeAgent : Orchestrating On-Device LLM inference for End-User Multi-Agent Systems on CPU-GPU Unified Memory Architectures

【速读】:该论文旨在解决多智能体大语言模型(multi-agent LLMs)在边缘统一内存架构(UMA)上进行隐私保护部署时,因推理系统无法有效支持协同工作流而引发的性能瓶颈问题。核心挑战包括:解码阶段内存密集型操作导致统一内存总线严重争用,阻碍了CPU-GPU协同执行;多智能体任务中推测性解码存在极端不稳定的生成难度波动,从复杂推理到可预测结构化生成频繁切换;加之工具调用常引发频繁阻塞,造成执行高度碎片化,传统静态批处理策略难以发挥效用。为应对上述问题,本文提出跨层级优化的EdgeAgent系统,其关键解决方案在于:在微架构层面,突破传统图编译器约束,实现零拷贝、面向UMA的张量并行,通过非对称内存布局充分饱和CPU与GPU计算单元;在调度层面,基于实时序列可预测性动态分配草案预算,抑制带宽浪费,并引入异步挂起-释放机制主动驱逐阻塞智能体,保障不可预测工具调用场景下的硬件持续高利用率。实验表明,在Apple M4 SoC上的评估显示,仅采用面向UMA的执行策略即可带来1.29倍加速,结合智能体感知调度后,系统整体达到1.77倍性能提升,显著优于传统批处理的推测解码方案。

链接: https://arxiv.org/abs/2610.03394
作者: Yuhai Long(1),Yuanxin Wei(1),Kai Wu(2),Jinhui Wei(1),Dan Huang(1),Jiangsu Du(1) ((1) School of Computer Science and Engineering, Sun Yat-sen University, (2) China Mobile Internet Company Ltd.)
机构: Sun Yat-sen University, School of Computer Science and Engineering (中山大学计算机科学与工程学院); China Mobile Internet Company Ltd. (中国移动互联网公司有限公司)
类目: Distributed, Parallel, and Cluster Computing (cs.DC); Multiagent Systems (cs.MA)
备注:

点击查看摘要

Abstract:Emerging multi-agent LLMs demand privacy-preserving edge deployment, yet current inference systems struggle with these collaborative workflows. Specifically, the memory-bound decode phase causes severe bus contention on unified memory architectures (UMA), paralyzing naive CPU-GPU co-execution. Furthermore, speculative decoding in multi-agent workloads faces extreme variance in drafting difficulty, alternating between complex reasoning and predictable structured generation. Compounded by frequent tool-induced stalls, this highly fragmented execution severely underutilizes hardware and defeats traditional static batching. We present EdgeAgent, a cross-layer inference system explicitly co-designed for edge UMA and multi-agent workloads. At the micro-architectural level, it bypasses rigid graph-compiler constraints to enable zero-copy UMA-aware tensor parallelism, utilizing asymmetric memory layouts to fully saturate both CPU and GPU compute units. At the scheduling level, it dynamically allocates draft budgets based on real-time sequence predictability to bound bandwidth waste. Concurrently, an asynchronous suspend-and-yield mechanism actively evicts stalled agents, ensuring continuous hardware saturation during unpredictable tool invocations. Extensive evaluations on an Apple M4 SoC demonstrate that the UMA-aware execution alone contributes a 1.29x speedup over batched speculative decoding. Adding the agent-aware scheduling lifts the full EdgeAgent system to a 1.77x speedup under extreme tool-use latencies. Subjects: Distributed, Parallel, and Cluster Computing (cs.DC); Multiagent Systems (cs.MA) Cite as: arXiv:2610.03394 [cs.DC] (or arXiv:2610.03394v1 [cs.DC] for this version) https://doi.org/10.48550/arXiv.2610.03394 Focus to learn more arXiv-issued DOI via DataCite (pending registration) Related DOI: https://doi.org/10.1145/3845814.3850072 Focus to learn more DOI(s) linking to related resources

[MA-1] Defense-in-Depth at the Perception-Reasoning Interface of LLM -Centric Agent ic UAV Swarms

【速读】:该论文旨在解决生成式 AI (Generative AI) 驱动的无人飞行器(UAV)蜂群在执行数据采集调度任务时,因感知-推理接口受到隐蔽篡改的传感器报告攻击而引发的系统安全问题。攻击者无需修改模型权重或无人机本身,仅通过操纵输入报告即可诱导蜂群偏离预定行为。现有防御机制多停留在架构层面,缺乏实际实现与评估。为此,本文提出一种纵深防御策略,在感知-推理接口部署五层检测机制:检查报告来源可信性、物理可实现性、与蜂群几何结构及服务历史的一致性、调度是否导致传感器饥饿,并在前四层失效时启用确定性调度器接管控制。关键在于将攻击检测与响应解耦,实现独立验证与应对。实验表明,针对每层检测器,可闭合形式推导出攻击可容忍的最大报告偏差阈值,且该阈值由部署参数预先设定,不依赖攻击数据;在三十次匹配仿真中,预测与实测边界高度一致。尽管检测层在拒绝异常报告后以最近有效报告替代,虽能有效遏制攻击,但导致累计成本分别上升79%和74%;而仅用于安全验证的检查层虽未检出任何攻击,却仍使攻击引发的成本降低37.5%,凸显了分离检测与响应对系统鲁棒性的关键价值。

链接: https://arxiv.org/abs/2610.03319
作者: Mohammadhossein Homaei,Yousef Emami,Sajad Homayoun,Rahim Taheri,Hao Zhou,Miguel Gutierrez Gaitan,Bo Wei
机构: 未知
类目: Cryptography and Security (cs.CR); Multiagent Systems (cs.MA)
备注: 14 Pages, 7 Tables, 4 Figures

点击查看摘要

Abstract:Large Language Models (LLMs) increasingly support Uncrewed Aerial Vehicle (UAV) swarm operations such as data collection scheduling, where the model reads structured sensor reports and decides which sensors to visit. An adversary who quietly manipulates those reports can redirect the swarm without modifying the model weights or the UAV. Defenses for this interface have been proposed architecturally but rarely implemented or evaluated. We implement and evaluate defense-in-depth at the perception-reasoning interface of LLM-Centric Agentic UAV Swarms. Five layers check the provenance of a report, whether its values are physically admissible, whether they agree with what swarm geometry and service history predict, whether the resulting schedule starves any sensor, and, when these fail, hand control to a deterministic scheduler that ignores the suspect input. We test each layer against an adversary strong enough to defeat the layer before it. For each of the three input-side layers, we derive in closed form how far a report can be distorted before that layer reacts, fixing each boundary from deployment parameters before any attack data is collected; across thirty matched simulation runs, predicted and measured boundaries agree. Separating attack detection from response is a well-established principle, and we quantify the cost of neglecting this distinction at the perception-reasoning interface. When the system rejects a report, it replaces it with the most recent accepted report. This prevents the adversary from controlling the UAV schedule, but it also increases cumulative cost by 79% and 74% for the two detectors, respectively, compared with the undefended system. The safety check does not detect any attacks, but it nevertheless reduces the attack-induced cost by 37.5%.

[MA-2] FinNextAssist: Towards Professional Financial Deep Research Assistant

【速读】:该论文旨在解决将深度研究(Deep Research, DR)代理应用于金融领域时所面临的独特挑战,即金融分析需协同完成跨越多种数据类型、工具与分析流程的异构子任务。其核心问题在于如何构建一个能够有效整合权威且异构的金融数据源、运用专业化分析工具与技能,并具备领域特定子任务处理能力的专业化金融DR代理。解决方案的关键在于提出一个端到端的深度研究框架FinNextAssist,该框架将研究过程分解为四个阶段:任务规划器(Task Planner)、证据编译器(Evidence Compiler)、推理引擎(Reasoning Engine)和报告生成器(Report Assembler),并创新性地引入两种轻量级子代理:TabAgent用于跨市场金融表格理解,HeteroAgent用于跨模态异构金融数据的解析。实验证明,FinNextAssist在FinDeepResearch、Finance Agent Benchmark及FinTMMBench-Web等多个基准上显著优于主流专有与开源DR代理,且消融实验验证了各组件在多市场与多语言环境下的有效性。

链接: https://arxiv.org/abs/2610.03174
作者: Xiangyu Li,Fengbin Zhu,Xuan Yao,Siyu Liu,Xiaoluan Liu,Chao Wang,Huanbo Luan,Xiaofen Xing,Xiangmin Xu,Ke-Wei Huang,Richang Hong,Tat-Seng Chua
机构: South China University of Technology (华南理工大学); National University of Singapore (新加坡国立大学); iFLYTEK CO., LTD (科大讯飞); Central University of Finance and Economics (中央财经大学); 6Estates (六地产); Hefei University of Technology (合肥工业大学)
类目: Multiagent Systems (cs.MA); Computational Finance (q-fin.CP)
备注:

点击查看摘要

Abstract:Deep Research (DR) agents have demonstrated strong capabilities in complex, research-oriented tasks through autonomous planning, iterative retrieval, multi-step reasoning, and structured reporting. However, adapting DR agents to finance introduces unique challenges: financial analysis demands the joint completion of heterogeneous sub-tasks spanning diverse data types, tools, and analytical workflows. We identify three key requirements for a professional financial DR agent: integration of authoritative, heterogeneous financial data sources; specialized analytical tools and skills; and dedicated sub-agents for domain-specific sub-tasks. Building on these principles, we propose FinNextAssist, an end-to-end deep research framework designed for professional financial analysis. FinNextAssist decomposes the research process into four stages: Task Planner, Evidence Compiler, Reasoning Engine, and Report Assembler, and introduces two novel lightweight sub-agents: TabAgent, for cross-market financial table understanding, and HeteroAgent, for cross-modality heterogeneous financial data interpretation. Extensive experiments on FinDeepResearch, the Finance Agent Benchmark, and FinTMMBench-Web show that FinNextAssist substantially outperforms both strong proprietary and open-source DR agents, with ablation studies confirming the contribution of each component across diverse markets and languages.

[MA-3] Engineering Sustainable Agents : A Systematic Comparison of Agent ic LLM s for Developer Workflows

【速读】:该论文旨在解决生成式 AI 在软件工程中应用时,尤其是多智能体(agentic)大语言模型(LLM)系统在计算与环境成本上的高昂开销问题。其核心挑战在于如何在保持任务性能的同时,实现能源效率与推理延迟的优化。解决方案的关键在于通过大规模实证研究,系统评估了五类典型软件工程任务(代码生成、技术债务识别、漏洞检测、日志解析与分析)中从单次查询基线到多智能体工作流的多种配置组合,涵盖六种开源模型、两种提示策略及三种硬件平台,并从准确性、推理延迟和能耗三方面进行综合评估。研究发现,多智能体设计虽在部分任务(如漏洞检测)中带来有限准确率提升,但平均能耗高出基线6.36倍,推理延迟增加6.07倍,最差情况下延迟甚至高达160倍;而多数帕累托最优配置仍由轻量级非智能体或单智能体方案主导。因此,研究提出应基于任务特性动态选择模型与提示策略,而非采用全局统一配置,最终形成面向可持续性与任务感知的生成式 AI 软件开发工具设计准则。

链接: https://arxiv.org/abs/2610.03010
作者: Merve Astekin,Yan Naing Tun,Arda Goknil,Erik Johannes Husom,Lwin Khin Shar,Hasan Sözer,Ratnadira Widyasari,Hui Song
机构: SINTEF(挪威科技研究所); Singapore Management University(新加坡管理大学); Ozyegin University(欧宗吉大学)
类目: oftware Engineering (cs.SE); Multiagent Systems (cs.MA)
备注: 50 pages, 14 figures

点击查看摘要

Abstract:Large language models (LLMs) are increasingly used in software engineering, including agentic systems that coordinate multiple agents, but impose higher computational and environmental costs. In this paper, we present a comprehensive empirical study of agentic LLM systems across five software engineering tasks: code generation, technical debt identification, code vulnerability detection, log parsing, and log analysis. For each task, we compare LLM configurations that range from a non-agentic single-query baseline to multi-agent workflows, using six open-weight LLMs, two prompt strategies, and three hardware platforms. We assess each configuration in terms of accuracy, inference latency, and energy consumption. Our results reveal substantial trade-offs between agentic complexity and energy efficiency: multi-agent designs consume on average 6.36 \times as much energy and run 6.07 \times as long as the non-agentic baseline, with worst-case slowdowns of up to 160 \times for individual task–hardware pairs. Accuracy gains from additional agents are limited and task-specific: multi-agent improves average vulnerability-detection accuracy, but lightweight non-agentic and single-agent configurations still dominate the Pareto front, accounting for 59 of 66 Pareto-optimal configurations. Model and prompt choice act as task-specific levers whose effective direction varies between tasks rather than as global defaults. We translate these findings into design guidelines for sustainable, task-aware LLM-based development tools.

[MA-4] Dynamic Expert Pruning for Multi-Agent Systems

【速读】:该论文旨在解决混合专家模型(Mixture-of-Experts, MoE)在实际部署中面临的内存瓶颈问题。尽管MoE通过仅激活部分专家实现了计算资源的高效利用,但所有专家仍需常驻加速器显存,导致内存占用受限,难以在资源受限场景下扩展。现有专家剪枝方法多为静态策略,即使用固定的专家掩码对所有请求统一处理,无法适应异构工作负载,尤其在多智能体系统中表现不佳——不同任务和角色所需专家各异,而静态方法强制统一分配专家子集,造成资源浪费或性能下降。本文提出动态专家剪枝(Dynamic Expert Pruning, DEP),其核心创新在于发现:智能体的系统提示与任务提示本身已包含其行为意图,足以推断所需专家集合。DEP通过一个轻量级预测器(在工作流日志上一次性训练完成)对输入提示进行单次前向推理,即可生成针对每条请求的个性化专家掩码,无需额外校准。实验表明,无论在任务多样性、模型规模或MoE架构差异下,DEP均优于静态剪枝与合并基线,在保留较少专家时优势尤为显著,验证了多智能体系统中角色特异性带来的稀疏服务潜力,从而在保证精度的同时实现更优的内存效率与泛化能力。

链接: https://arxiv.org/abs/2610.02951
作者: Jabin Koo,Soheil Abbasloo,Sungjae Lee,Jungseul Ok
机构: Pohang University of Science and Technology (POSTECH); Microsoft Research Asia
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA)
备注: 18 pages, 3 figures, 15 tables

点击查看摘要

Abstract:Mixture-of-Experts (MoE) architectures scale language models efficiently by activating only a few experts per token, but the saving is confined to computation: every expert must stay resident on the accelerator, so memory bounds where these models can be deployed. Expert pruning reduces this footprint, yet existing methods are static — a single mask, calibrated offline, is applied to the model for every subsequent request. This assumption can fail when the workload is heterogeneous, most prominently in multi-agent systems, where one backbone serves many tasks and roles at once: our analysis shows that different tasks and roles recruit different experts, while static methods assign one fixed subset to all of them. We therefore propose Dynamic Expert Pruning (DEP), which rests on a finding we establish here: an agent’s system and task prompts are by themselves sufficient to identify the experts that agent and its task require, since that text already describes what the agent will do. A lightweight predictor, trained once on workflow transcripts, turns those prompts into a specialized per-request mask in a single forward pass, with no per-configuration calibration. Across diverse tasks and roles, model scales, and MoE architectures, DEP achieves better overall accuracy than static pruning and merging baselines, and generalizes to workflows unseen in training without retraining. Its margin over those baselines is largest when few experts are retained, suggesting that the role specialization inherent to multi-agent systems permits sparser serving than static pruning allows.

[MA-5] Evaluator-in-the-Loop Monte Carlo Tree Search via LLM Agents for Motif Scaffolding in Protein Design

【速读】:该论文旨在解决生成式蛋白设计中“生成-筛选”范式效率低下问题,即传统方法在独立生成候选序列后仅将结构评估用于最终筛选或排序,未能充分利用评估过程中的失败反馈信息。其核心挑战在于:失败预测蕴含关于模体几何、全局可折叠性等结构约束的特定状态证据,但现有方法难以有效重用这些反馈以指导迭代优化。为此,本文提出基于证据的大型语言模型引导蒙特卡洛搜索(ELMS),其关键创新在于构建“评估器在环”(evaluator-in-the-loop)的搜索框架,将结构评估反馈转化为可执行的设计动作。具体而言,系统通过批评者代理(Critic Agent)诊断局部结构失败模式,策略代理(Policy Agent)选择具有参数化的靶向操作符,结合模体锁定操作符实现合法序列修改,并利用蒙特卡洛树搜索(MCTS)动态决定哪些历史搜索状态应继续投入设计资源。该设计避免了单一路径的过早收敛,实现了对结构反馈的高效再利用。在标准GeomMotif协议下,ELMS在单模体和双模体任务上的成功率达86.41%与84.57%,分别优于最强基线19.3和21.9个百分点;在MotifBench基准上,平均解决26.7/30个任务(任务成功率88.89%),显著超越基线的16.0/30(53.33%)。结果表明,ELMS成功将结构评估从静态筛选转变为动态、可行动的迭代设计引导机制。

链接: https://arxiv.org/abs/2610.02924
作者: Haotian Hu,Oguzhan Gungordo,Siheng Xiong,Faramarz Fekri
机构: Georgia Institute of Technology(佐治亚理工学院)
类目: Multiagent Systems (cs.MA)
备注:

点击查看摘要

Abstract:Motif-scaffolding systems commonly follow a generate-then-filter paradigm, in which candidate proteins are generated independently and structural evaluation is used primarily for terminal screening or ranking. This paradigm underuses evaluation: failed predictions contain state-specific evidence about whether a design requires repair of motif geometry, global foldability, or other structural constraints. We introduce \textbfELMS (Evidence-based LLM-guided Monte Carlo Search), an evaluator-in-the-loop search framework for motif scaffolding that turns such evaluator feedback into targeted design actions. Effective reuse of structural feedback is nontrivial because different scaffold states exhibit different failure modes, and repeatedly refining a single trajectory can prematurely commit computation to an unproductive region of sequence space. ELMS therefore retains evaluated scaffolds as persistent search states: a Critic Agent diagnoses state-local structural failures, a Policy Agent selects targeted operators with execution parameters, motif-locked operators realize legal sequence modifications, and MCTS determines which historical states should receive further design effort. Under the standard GeomMotif protocol (100 candidates per task), ELMS achieves Successful rates of 86.41% on single-motif tasks and 84.57% on paired-motif tasks, exceeding the strongest prior baseline by 19.3 and 21.9 percentage points, respectively. On MotifBench, under a matched 100-candidate search budget, it solves 26.7 of 30 tasks on average (88.89% Task Success), compared with 16.0 tasks (53.33%) for the strongest baseline. These results establish ELMS as an effective approach for converting structural evaluation from a terminal filter into actionable guidance for iterative motif scaffolding.

[MA-6] SceneFactory-3D: Lifting 2D Traffic Scenes into 3D Physical Counterfactuals for Scalable Physically Grounded Safety Evaluation

【速读】:该论文旨在解决传统可扩展驾驶仿真器在模拟车辆行为时忽视轮胎-路面物理交互机制的问题,导致其难以准确捕捉恶劣道路与环境条件对车辆执行能力的影响及其在交通流中的传播效应。其核心解决方案是提出SceneFactory-3D——一个基于GPU批处理、物理驱动的多智能体驾驶仿真平台。该系统通过在每个车轮接触点上计算受悬架和摩擦限制的力,实现对加速与转向命令的物理精确响应;同时,结合空间变化的摩擦系数、每世界独立的三维高程场(3D heightfields)、重力及刚性接触模型,统一约束车轮运动与车身碰撞行为。通过每世界地形隔离与GPU批量并行处理,SceneFactory-3D能够在保持交通场景设置与车辆控制器不变的前提下,同步运行多个物理反事实场景(即仅道路条件变化),从而高效评估闭环交通系统的动态响应。为验证该仿真器在反事实分析中的优势,研究开展了一项实证分析,评估不同学习型策略与经典规划算法在21种摩擦系数与坡度组合下的鲁棒性。结果显示,当摩擦系数从1.0降至0.18时,各类学习型策略的安全通过率下降6至90个百分点(经典规划算法下降18–19个百分点),且所有学习型策略均表现出近碰撞事件频率显著上升,凸显了道路条件对自动驾驶控制策略性能的关键影响。

链接: https://arxiv.org/abs/2610.02874
作者: Yicheng Zhu,Linfeng Tian,Tianmu Zhao,Yang Chen,Fan Zuo,Tao Li,Zilin Bian
机构: Rochester Institute of Technology(罗切斯特理工学院); City University of Hong Kong(香港城市大学); New York University(纽约大学)
类目: Robotics (cs.RO); Multiagent Systems (cs.MA)
备注:

点击查看摘要

Abstract:Scalable driving simulators typically execute vehicle commands using prescribed behavioral or kinematic rules, overlooking the physics of tire-road interfaces, thereby limiting their ability to capture how adverse road and environmental conditions alter vehicle execution and propagate through traffic. To address this limitation, we present SceneFactory-3D, a GPU-batched, physics-grounded multi-agent driving simulator. Vehicles execute acceleration and steering commands via suspension- and friction-limited forces evaluated at each wheel-contact point. Spatially varying friction, per-world 3D heightfields, gravity, and rigid contact consistently govern wheel motion and chassis collisions. Per-world terrain isolation and GPU batching enable SceneFactory-3D to run matched physical counterfactuals in parallel: traffic scenario setup and vehicle controllers remain fixed while only the road condition changes, enabling the resulting closed-loop effects to be evaluated across parallel worlds. To demonstrate the advantage of the SceneFactory-3D-enabled counterfactual evaluation, we conduct an empirical study on vehicle controllers’ sensitivity to road conditions. We study three learned-policy families on 1,024 matched 12-vehicle worlds per condition, and two classical planners on a shared 32-world subset, across 21 friction and grade conditions. When friction drops from 1.0 to 0.18, the share of vehicles that clear the work zone safely falls by 6 to 90 percentage points across learned policies (18-19 for classical planners), and near-collision situations become more frequent for every learned policy. Code: this https URL

[MA-7] Permutation Robustness Is Not Enough: Action Collapse in Multi-Agent Transformer Policies

【速读】:该论文旨在解决多智能体机器人学习中基于Transformer的策略在面对智能体顺序排列变化时的鲁棒性问题。由于Transformer通常将智能体视为有序的序列输入,而实际多智能体团队具有无序特性,这种结构上的不匹配可能导致策略对智能体顺序敏感,从而影响其在真实场景中的泛化能力。解决方案的关键在于引入并系统评估双重评价指标:一是基于排列一致性(permutation-consistency)的鲁棒性度量,二是针对动作坍缩(action-collapse)的诊断工具,包括动作多样性、相同动作占比及最大动作频率等。研究发现,仅依赖低排列误差可能产生误导,因为策略看似鲁棒实则因所有智能体采取相同动作而表现一致;因此,必须同时关注策略是否保持非坍缩、差异化的行为模式。实验表明,弱等变性正则化可有效提升N=3智能体团队的鲁棒性并维持较高动作多样性,而N=4团队则需更小的正则化强度以避免过度同质化。这表明,设计多智能体Transformer策略时,除性能回报和排列不变性外,还需特别关注其行为的多样性与分化能力。

链接: https://arxiv.org/abs/2610.02848
作者: Amit Thakur,Mukesh Singhal
机构: University of California, Merced (加利福尼亚大学默塞德分校)
类目: Robotics (cs.RO); Machine Learning (cs.LG); Multiagent Systems (cs.MA)
备注:

点击查看摘要

Abstract:Transformer policies are attractive for multi-agent robot learning because self-attention can model interactions among agents. However, multi-agent teams are unordered, while transformers typically process agents as ordered token sequences. We study how this mismatch affects cooperative navigation policies under agent-order permutations. Our results show that low permutation error alone can be misleading: policies may appear robust simply because all agents choose the same action. We therefore evaluate policies using both permutation-consistency metrics and action-collapse diagnostics, including action diversity, same-action fraction, and maximum action frequency. A PPO-ID baseline yields non-collapsed behavior but remains order-sensitive, while strong equivariance regularization can still induce homogeneous behavior. A weak equivariance penalty improves the robustness while preserving more diverse actions for teams with (N=3) agents, whereas teams with (N=4) agents require substantially smaller regularization weights. These findings suggest that multi-agent transformer policies should be evaluated not only by return and permutation robustness, but also by whether they maintain non-collapsed, differentiated multi-agent behavior.

[MA-8] urnover-Orthogonal Credit Assignment for Open-Team Multi-Agent Reinforcement Learning

【速读】:该论文旨在解决开放团队多智能体强化学习(open-team multi-agent reinforcement learning)中因智能体动态加入、退出或替换导致的信用分配混淆问题。在传统方法中,集中式评价网络(centralized critic)与共享优势函数(shared advantage)将智能体行为贡献与外部环境引起的团队成员更替效应混合为单一标量信用信号,致使存活智能体被奖励或惩罚于其无法控制的外部更换事件,从而影响学习稳定性与性能。本文提出周转正交信用分配(Turnover-Orthogonal Credit Assignment, TOCA),其核心创新在于对价值函数进行分解,明确分离动作效应(action effects)、纯周转效应(pure turnover effects)以及动作与周转的交互效应(action-turnover interactions)。通过引入事件条件值函数(event-conditioned value),并设计事件条件基线(event-conditioned baseline)以消除纯周转分量,同时保留对提升团队抗替换鲁棒性之动作的有效信用。作者进一步构建了基于可变规模智能体集合与事件标记(event tokens)的置换不变集中式评价网络,并提出两种实现形式:一种是基于反事实的每智能体信用信号,另一种是针对高方差控制环境优化的软加权交互变体——TOCA-β。控制性诊断实验表明,相较于事件感知的MAPPO类评价网络,TOCA显著提升团队回报;移除交互信用项则导致性能明显下降。在仅含替换的动态扩散(Dynamic Spread)基准测试中,TOCA-β 在高周转率下取得最优平均回报,并优于无交互信用的消融版本。结果表明,显式分离周转与动作信用是提升动态协作团队鲁棒学习能力的重要原则。

链接: https://arxiv.org/abs/2610.02847
作者: Amit Thakur,Mukesh Singhal
机构: University of California, Merced(加州大学默塞德分校)
类目: Multiagent Systems (cs.MA); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Open-team multi-agent reinforcement learning studies cooperative systems in which agents may join, leave, or be replaced during an episode. In such settings, the team return changes both because agents choose useful actions and because the active population itself changes. Standard centralized critics and shared advantages often mix these two effects into one scalar credit signal, allowing surviving agents to be rewarded or penalized for exogenous turnover events outside their control. We introduce turnover-orthogonal credit assignment (TOCA), a value decomposition for open teams that separates action effects, pure turnover effects, and action–turnover interactions. Under exogenous turnover, the event-conditioned value admits a centered decomposition whose event-conditioned baseline removes the pure turnover component while preserving credit for actions that make the team robust to future replacements. We instantiate this idea with a permutation-invariant centralized critic over variable-size agent sets and event tokens, and derive both a counterfactual per-agent credit signal and a softly weighted interaction variant, TOCA- \beta , for high-variance control environments. Controlled diagnostic experiments show that TOCA improves return over event-aware MAPPO-style critics and that removing interaction credit substantially hurts performance. In a replacement-only Dynamic Spread benchmark, TOCA- \beta achieves the best mean return at high turnover rates and improves over its no-interaction ablation. These results suggest that explicitly separating turnover from action credit is a useful principle for robust learning in dynamic cooperative teams.

[MA-9] Silent Dissent: LLM Agents That Yield to the Majority Still Represent Their Original Premise

【速读】:该论文旨在解决多智能体辩论(multi-agent debate)中一个关键问题:当大语言模型(LLM)代理在集体讨论中改变其答案时,这种变化是源于真实认知上的信念更新(即“改变认知”),还是仅仅出于对多数意见的妥协(即“仅改变表述”)。这一问题在生成式人工智能(Generative AI)系统中尤为突出,因为代理可能表面上达成共识,实则保留原有信念。研究的核心解决方案在于利用神经敏感性分析工具——雅可比透镜(Jacobian lens, J-lens)与对数透镜(logit lens),从模型残差流(residual stream)中提取隐藏于推理过程中的中间实体(桥接信息,bridge entity),以判断代理是否真正改变了内在信念。实验通过两跳事实型问题设计,让“脚本化同谋者”(scripted peers)一致声称一个错误答案,并基于另一组不同桥接的事实误导代理。结果表明,在预注册测试中,尽管多个模型(如Qwen3.5-4B、Qwen3.6-27B、Gemma-4-E4B-it等)在输出上服从多数意见,但其内部表示仍显著保留原始桥接实体信号(前100位命中率分别为0.85、0.22、0.24),而对数透镜几乎无法捕捉该信号(命中率0.00–0.06)。进一步地,当隐藏代理先前的回答后,其仍能保持原桥接实体的内部表征,且服从行为显著上升(如Qwen3.5-4B从8%升至89%),说明其初始信念独立于陈述。此外,仅在两个Qwen模型中,通过注入桥接实体的J-lens方向可使代理恢复原答案。因此,研究揭示了当前多智能体辩论机制中“表面共识”与“深层分歧”的脱节现象,强调需依赖内省探针技术评估真实认知状态,而非仅依赖输出一致性。

链接: https://arxiv.org/abs/2610.02702
作者: Ziang Ni,Peng Zou
机构: Delft University of Technology (代尔夫特理工大学); Sun Yat-sen University (中山大学)
类目: Computation and Language (cs.CL); Multiagent Systems (cs.MA)
备注: 9 pages, 3 figures, 3 tables. Supplementary material in ancillary files

点击查看摘要

Abstract:Multi-agent debate is increasingly used to reach consensus among LLM agents, yet agents often yield to a unanimous majority. When an agent changes its answer, has it changed its mind or only its statement? We study this with two-hop factual questions whose intermediate entity (the bridge, e.g. the country in “the capital of the country where the Sagrada Familia is located”) is never stated by anyone. Scripted peers, in the role of Asch’s confederates, unanimously assert a wrong answer taken from another fact with a different bridge. At the moment the agent answers, we read the bridge from its residual stream with the Jacobian lens (J-lens) and, for comparison, the logit lens. In pre-registered tests on held-out facts with four open-weight models, agents of Qwen3.5-4B, Qwen3.6-27B and Gemma-4-E4B-it that gave in still represented their original bridge in the pre-registered layers below the output (hit@100 above a control entity: 0.85, 0.22 and 0.24), where the logit lens rarely ranked it among the top 100 tokens (0.00-0.06). These agents also represented the bridge behind the peers’ answer, beyond a mention baseline. A pre-registered addendum hid the agent’s earlier answer or removed it: agents that gave in still represented their original bridge in all four models (0.43, 0.29, 0.37 and 0.25 with the answer hidden), including Llama-3.1-8B-Instruct, which barely did so with its answer in view (0.03). The premise can thus be computed from the question alone while the agent states the majority’s answer. Hiding the earlier answer also changed conformity: Qwen3.5-4B gave in on 89% of questions instead of 8%. In exploratory interventions, injecting the bridge’s J-lens direction brought agents back to their original answer only in the two Qwen models. Stated consensus in multi-agent debate can thus overstate agreement. We also report the negative results of our pre-registered program.

[MA-10] Who Went Where When on the Lunar Surface: Forensic Trajectory Analysis to Identify Byzantine Rovers

【速读】:该论文旨在解决多辆非协作式月球表面探测车(non-cooperative planetary rovers)在缺乏持续现场观测条件下,如何准确重建其行驶轨迹的问题。由于存在可能提供误校准或恶意伪造测量数据的拜占庭式探测车(Byzantine agents),传统基于异常值鲁棒的位姿图优化方法易受干扰,因其无法区分内部一致但虚假的数据与真实数据。本文提出一种具有归属意识(attribution-aware)的轨迹估计方法,其核心在于通过评估各探测车的可信度而非单一测量的有效性来实现鲁棒性提升:该方法通过对比候选可信探测车子集内部及边界相对检测结果与先验信息之间的统计一致性,识别出可信探测车集合,并仅使用归属于可信代理的测量进行轨迹估计。实验结果表明,该方法在合成仿真与真实行星类比轨迹数据上均能有效识别拜占庭探测车,并显著优于现有鲁棒位姿图优化基准。

链接: https://arxiv.org/abs/2610.02694
作者: Lachlan Holden,Feras Dayoub,Melissa de Zwart,David Harvey,Tat-Jun Chin
机构: 未知
类目: Robotics (cs.RO); Multiagent Systems (cs.MA)
备注: accepted to iSpaRo 2026

点击查看摘要

Abstract:Future planetary surface missions are likely to involve multiple independently operated rovers sharing the same deployment region, raising the need to verify compliance with operational constraints such as Lunar Safety Zones. Because continuous in-situ observability is rarely available, such verification requires post-hoc reconstruction of rover trajectories from sparse telemetry, including odometry, pose priors, and relative inter-rover detections. We introduce the problem of forensic trajectory analysis for non-cooperative planetary rovers in the presence of Byzantine agents: rovers that provide miscalibrated or deliberately falsified measurements to support an incorrect trajectory. We show that standard outlier-robust pose graph optimisation methods are vulnerable in this setting, because Byzantine rovers can generate measurements that are internally consistent and numerous enough to make truthful incriminating measurements appear as outliers. To address this, we propose an attribution-aware trajectory estimation method that reasons over rover credibility rather than individual measurement validity. The method evaluates candidate credible rover subsets by comparing the statistical consistency of their internal and boundary relative detections against provided priors, and then estimates trajectories using only measurements attributed to credible agents. Across synthetic simulations and real planetary-analogue trajectory data, the proposed method identifies Byzantine rovers and produces significantly more accurate trajectory estimates than existing robust pose graph optimisation baselines.

[MA-11] Coherence-Driven Belief Formation and Population Dynamics of Contagion in LLM Agents FAST NEURIPS2026

【速读】:该论文旨在解决传统社会传染模型中对个体信念采纳机制假设过于理想化的问题,转而通过实证方法测量语言模型代理(language model agents)在群体中采纳信念的实际行为,量化在不同数量同侪支持下个体采纳某一主张的概率。其核心解决方案在于揭示信念采纳核函数(adoption kernel)呈现典型的复杂传染(complex contagion)特征——即呈S型(sigmoid),并存在依赖于三个关键因素的阈值:主张的合理性(plausibility)、信息源的可信度(reliability)以及代理自身的倾向性(disposition)。研究发现,这三个维度可被有效简化为单一有效维度,即新信念与大语言模型(LLM)代理先验信念之间的内在一致性(coherence)。此外,研究还观察到复杂传染在AI代理系统集体动态中的表现:在簇状网络中信念传播效率高于随机网络,并出现分岔级联窗口(bifurcating cascade window)和自维持的滞后共识(hysteretic consensus),表明一旦形成共识,其消除难度远大于建立难度。这一发现为理解生成式人工智能系统中的信念演化机制提供了关键实证依据。

链接: https://arxiv.org/abs/2610.02654
作者: Tathagata Banerjee,Nima Moghaddas
机构: Takeda Pharmaceuticals(武田制药); Northeastern University(东北大学)
类目: Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA); Social and Information Networks (cs.SI)
备注: 18 pages, 6 figures. Accepted at the NeurIPS 2026 Workshop on Foundations of Agentic Systems Theory (FAST)

点击查看摘要

Abstract:Models of social contagion usually assume how individuals adopt beliefs and derive population behavior from it. We instead empirically measure belief adoption in language model agents, quantifying the probability an agent adopts a claim given how many peers endorse it. We find this adoption kernel to be sigmoid, a characteristic of complex contagion, with a threshold that is sensitive to three sources: the claim’s plausibility, the source’s reliability, and the agent’s disposition. These three dimensions are well approximated by a single effective dimension which we propose can be understood as the coherence of the incoming belief with the LLM agent’s prior beliefs. Further, we observe a characteristic of complex contagion in the collective dynamics of belief adoption in a system of AI agents: further spread on clustered than random networks. These systems also exhibit a bifurcating cascade window, and self-sustaining hysteretic consensus which lead to consensus being far harder to remove than to establish.

[MA-12] WebUIProof: Benchmarking WebUI Code Generators with UI-Agent Execution Harness

【速读】:该论文旨在解决大规模评估网页用户界面(WebUI)生成质量的难题,现有基准普遍依赖自由格式提示与静态检查(如编译成功、截图),难以捕捉交互功能正确性。其核心解决方案是提出一个面向执行的基准测试框架WebUIProof,通过结构化规范与密集可执行的交互测试,覆盖通用WebUI(如仪表盘、游戏、交互工具)和3D交互仿真(如粒子/星系系统、物理动力学)两类任务。关键创新在于引入基于UI-Agent的测试引擎,在无头浏览器中以“规划-行动-观察”迭代循环自动执行交互测试,精准定位DOM元素、执行操作、观测界面与DOM变化,并验证预设断言。实验表明,尽管多数页面渲染成功,但交互层面仍存在高频失败,尤其在3D仿真场景中更为显著;进一步证明,利用可执行交互测试生成的结果级训练信号(如强化学习奖励),可有效提升小型模型(如Qwen2.5 14B、MIMO 7B)的功能完成率,同时降低构建失败率。

链接: https://arxiv.org/abs/2610.02617
作者: Yun-Yun Tsai,Yuning Mao,Shiqi Wang,Junfeng Yang,Sinong Wang
机构: Columbia University (哥伦比亚大学); Meta SuperIntelligence Labs (Meta超级智能实验室)
类目: oftware Engineering (cs.SE); Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA)
备注: 43 pages

点击查看摘要

Abstract:Evaluating WebUI code generation at scale is difficult: outputs may compile and look plausible yet fail under user interaction, and prior benchmarks largely rely on free-form prompts with static checks (build success, screenshots) that miss functional correctness. We introduce WebUIProof, an execution-oriented benchmark that provides structured specifications and dense, executable interaction tests for WebUI generation across two task families: general WebUIs (e.g., dashboards, game, interactive tools) and 3D interactive simulation (e.g., particle/galaxy systems, physics dynamics). WebUIProof includes a UI-agent harness that runs executable interaction tests in a headless browser using an iterative plan–act–observe loop: it locates DOM elements, performs actions, observes resulting UI/DOM changes, and checks the specified assertions. We evaluate across eight commercial LLMs and observe frequent failures on interaction-based requirements even when pages render successfully, especially on 3D simulation interfaces. Finally, we show the UI-agent harness can provide outcome-level training signals. Training compact models (e.g., Qwen2.5 14B and MIMO 7B) with RL rewards derived from executable interaction tests improves functional completion while reducing build failures.

[MA-13] MIRROR: Multipath Quorum Integrity for LLM Multi-Agent Communication NEURIPS2026

【速读】:该论文旨在解决大型语言模型多智能体系统(LLM-MAS)中智能体间通信所面临的“中间人攻击”(Agent-in-the-Middle, AiTM)问题,即攻击者可在不破坏智能体自身的情况下篡改传输中的消息。现有防御方案存在明显局限:基于语义验证的防御需额外推理开销且可能误拦合法输出,而传输层加密在中间节点合法终止TLS时失效。为此,论文提出MIRROR,一种通信层完整性原语,其核心在于将单一标准化后的消息通过k条逻辑路径复制分发,并仅当严格多数路径返回相同摘要时才接受该消息。MIRROR采用无密钥哈希机制,自身不提供认证能力(因主动攻击者可重算摘要),其完整性保障依赖于“诚实路径占多数”的假设。摘要仅用于使见证路径保持常数规模,并在第二原像抗性下将恢复的消息绑定至多数共识值。理论分析表明,在路径被攻陷比例α < 0.5的条件下,系统可保证完整性;进一步扩展至相关路径场景,关键指标为最大共失效组的大小而非路径总数。此外,可用性与完整性在相同阈值下退化:当α < 0.5时,拒绝服务型攻击与消息丢弃攻击均无法阻断诚实通信。实验在MMLU、HumanEval、MBPP等多个基准及两种框架、四种通信拓扑上验证,以及在MetaGPT部署中对抗生产API,MIRROR在1倍大模型(LLM)token成本下将攻击成功率降至0%;相比之下,以大模型为裁判的防御方案成本高达35倍,且在拓扑扫描中最多阻断44.2%的合法输出。

链接: https://arxiv.org/abs/2610.02349
作者: Ryuichi Yamafuji Lun,Jingzhen Wang,Shreyas Kolte,Ruiteng Li
机构: University of Southern California(南加州大学)
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA)
备注: Accepted at the NeurIPS 2026 Workshop on Foundations of Language Model Security (FLMSec). 12 pages, 4 figures, 3 tables

点击查看摘要

Abstract:Inter-agent communication is central to Large Language Model Multi-Agent Systems (LLM-MAS), but it introduces an underexplored vulnerability: Agent-in-the-Middle (AiTM) attacks that manipulate messages in transit without compromising the agents themselves. Prior work reports Attack Success Rates (ASR) approaching 100% on structured tasks. Existing defenses rely on semantic validation, which requires additional inference and can block benign outputs, or on transport-layer encryption, which does not help when an intermediary legitimately terminates TLS. We present MIRROR, a communication-layer integrity primitive that replicates a single canonicalized payload across k logical routes and accepts a message only when a strict majority of routes report the same digest. MIRROR uses unkeyed hashing and so authenticates nothing on its own, since an active on-path adversary can always recompute a digest over a payload it has modified. All integrity derives from the assumption that honest routes form a majority. The digest serves only to make witness routes constant-size and to bind the recovered payload to the quorum-agreed value under second-preimage resistance. We give the guarantee under a route-compromise bound alpha 0.5, and extend it to correlated routes, where the quantity that matters is the size of the largest shared-failure group and not the route count. We further show that availability and integrity degrade at the same threshold: below alpha = 0.5, quorum-denial and message-dropping adversaries cannot block honest traffic. Across MMLU, HumanEval, and MBPP on two frameworks and four communication topologies, and in a MetaGPT deployment against a production API, MIRROR reduces ASR to 0% below the threshold at 1x LLM token cost. LLM-as-a-Judge costs 35x in the same deployment, and blocks up to 44.2% of benign outputs in the topology sweep.

[MA-14] MACTS-EM: Multi-Agent Collaborative Time Series Forecasting with Emergent Memory

【速读】:该论文旨在解决时间序列预测中面临的复杂挑战,包括制度转换(regime shifts)、跨领域知识迁移以及多模态数据融合等问题。现有方法在处理这些动态、非平稳且异构的现实场景时表现受限。其解决方案的关键在于提出一种名为“多智能体协同时间序列预测与涌现记忆”(MACTS-EM)的新框架,该框架通过五个核心机制实现突破:(1) 领域专用的预测智能体,分别负责模式识别、异常检测、因果推断与不确定性量化;(2) 具有动态任务分配能力的元认知层;(3) 支持跨领域模式迁移的涌现记忆机制;(4) 多模态上下文信息整合能力;(5) 抗对抗攻击的鲁棒性组件。实验结果表明,MACTS-EM在金融、气候、能源及疫情传播等多个真实场景中显著优于现有方法,尤其在零样本迁移能力、制度转换下的鲁棒性及分布偏移后的恢复速度方面提升显著,验证了基于协作式智能体架构在复杂时间序列建模中的有效性与前景。

链接: https://arxiv.org/abs/2610.02255
作者: Ahmad Shahi,Mamehgol Yousefi
机构: Unitec Institute of Technology(新西兰惠灵顿理工学院)
类目: Machine Learning (cs.LG); Multiagent Systems (cs.MA)
备注: 16 pages, 3 figures, 5 tables

点击查看摘要

Abstract:Time series forecasting remains a critical challenge across numerous domains. Despite significant advancements, existing approaches struggle with complex phenomena such as regime shifts, cross-domain knowledge transfer, and multimodal data integration. This paper introduces Multi-Agent Collaborative Time Series Forecasting with Emergent Memory (MACTS-EM), a novel framework where specialised agents collaborate to achieve superior forecasting performance. The MACTS-EM architecture integrates: (1) domain-specialised forecasting agents for pattern recognition, anomaly detection, causal inference, and uncertainty quantification; (2) a meta-cognitive layer for dynamic agent allocation; (3) an emergent memory mechanism enabling cross-domain pattern transfer; (4) multimodal contextual integration; and (5) adversarial robustness components. Evaluation across financial markets, climate patterns, energy consumption, and pandemic propagation demonstrates that MACTS-EM outperforms existing approaches in most scenarios, with 8-12% improvement in forecasting accuracy, 22-27% better zero-shot transfer capability, 16-21% enhanced resilience during regime shifts, and 15-18% faster recovery after distribution shifts. Our findings suggest that collaborative, agentic approaches to time series forecasting represent a promising direction beyond traditional architectures, particularly for complex real-world scenarios requiring multi-resolution temporal understanding and contextual adaptation.

[MA-15] Multi-Agent AI as a Nested Principal-Agent Problem in Private Wealth Management: Mandate Representation and Evidence Control in Switzerland Germany and Austria ALT

【速读】:该论文旨在解决私人财富管理中人工智能(AI)作为客户代理人与系统委托人双重角色下的责任界定与决策透明性问题,核心挑战在于如何在法律约束、委托人目标与管理人能力之间实现可解释且合规的最优决策。其解决方案的关键在于提出一种模型无关的嵌套委托-代理框架,结合受约束的联合最大化机制,将客户与管理人的绩效分别建模于投资组合-工作流对上,并通过权重设定与基准服务下限显式表达权衡关系;同时引入“让步会计”(concession accounting)以分离不同因素对客户利益的影响。研究表明,在多场景模拟中,忽略客户负债会导致流动性违规,遗漏管理人条款引发能力超限,而权重误译则改变本可接受的决策选择;在既定权重下,有六种状态选择了高于客户最优的服务层级,虽使客户承担1,178至2,264欧元的让步成本,但管理人获得3,062至10,381欧元收益。三类指令形式均在共享数值、证据和模拟审批控制下达成全部32项预设决策,其中专业指令在准确性与澄清度方面优于其他形式。后续汇率与收益率路径测试表明,尽管所选服务满足决策时预测基准,但客户最终结果仍低于基准服务水平,凸显了该方法在揭示委托决策后果方面的潜力。整体而言,该框架为提升委托指令的可审查性与治理效能提供了理论与实证基础。

链接: https://arxiv.org/abs/2610.02863
作者: Walter Kurz,Reinhard Magg,Florian Kollberg,Wojtek Stricker,Stefan Marx,Frank Reinhardt,Velimir Dedić
机构: 未知
类目: General Finance (q-fin.GN); Multiagent Systems (cs.MA)
备注: Quantitative formulation of client-manager-AI delegation and constrained joint decision objectives in private wealth management. 30 pages, 8 figures, 11 tables. Published in Swissi AI Journal under CC BY 4.0. Journal record: this https URL

点击查看摘要

Abstract:In private wealth management, a manager delegating to artificial intelligence (AI) acts as the client’s agent and the system’s principal. We introduce a model-independent formulation that combines nested principal–agent delegation with constrained joint maximisation as the task assigned to the AI system. The objective represents client and manager outcomes separately over portfolio–workflow pairs. Legal duties, mandate requirements and evidence sufficiency determine admissibility, with Switzerland, Germany and Austria supplying the legal context. Weights and reference-service floors make the trade-off explicit; concession accounting separates their effects on the client. Analytical constructions and a simulation using public-market observations illustrate the approach. Across eight decision states from four constructed mandates, omitted client liabilities caused two liquidity violations, omitted manager terms caused two capacity violations, and mistranslated weights changed four otherwise admissible choices under faithful optimisation. At the declared weights, six states selected a higher service tier than the client-best alternative, with client concessions of EUR 1,178 to EUR 2,264 and manager gains of EUR 3,062 to EUR 10,381. Three instruction forms each reached all 32 specified decisions under shared numerical, evidence and simulated approval controls; professional instructions matched explicit nested delegation on accuracy and clarification count. Subsequent 2022 exchange-rate and yield paths, combined with constructed growth scenarios, produced lower client outcomes than the reference service although the selected services met the decision-time forecast benchmarks. These examples suggest that the approach could help make mandate choices and their consequences easier to examine. Professional and field studies could assess whether this improves oversight and client outcomes.

[MA-16] Performance of Zero-determinant Strategies in Repeated Games without Discounting

【速读】:该论文旨在解决零决定性(Zero-determinant, ZD)策略在无折现重复博弈中性能分析的理论局限性问题。以往研究普遍依赖于“行动配置概率分布的Cesaro平均极限存在”这一假设来解释ZD策略能够“单方面强制支付之间的线性关系”的特性,但该假设在某些情况下可能不成立,限制了结论的普适性。本文的关键突破在于:无需依赖上述极限存在的假设,直接基于博弈过程的长期行为进行分析,从而在更一般条件下证明了ZD策略仍能实现对收益关系的单方面控制。这一方法不仅避免了对收敛性的强假设,还得到了比先前结果更强、更具普适性的理论结论,为理解ZD策略在复杂博弈环境中的作用机制提供了更坚实的数学基础。

链接: https://arxiv.org/abs/2610.02645
作者: Masahiko Ueda
机构: Tokyo Metropolitan University (东京都立大学)
类目: Physics and Society (physics.soc-ph); Multiagent Systems (cs.MA); Systems and Control (eess.SY)
备注: 4 pages, 1 figure

点击查看摘要

Abstract:Zero-determinant (ZD) strategies are a class of strategies in repeated games, which unilaterally control payoffs. It has been shown that several ZD strategies promote cooperation in social dilemma games. It has been widely believed that the performance of ZD strategies is expressed as ``unilaterally enforce linear relations between payoffs’'. However, in order to interpret properties of ZD strategies in repeated games without discounting, previous studies assumed that the limit of the Cesaro average of the probability distribution of the action profile exists. Here, we explain the performance of ZD strategies without this assumption, which leads to a stronger result than previous ones.

自然语言处理

[NLP-0] Language Models that Play Chess and Explain Their Moves

【速读】: 该论文旨在解决现代棋类引擎(chess engine)与语言模型(language model, LM)在棋局决策中各自存在的局限性:前者虽具备超人类的博弈能力,但缺乏可解释性;后者虽能生成看似合理的解释,却因棋力薄弱而难以提供有效指导。其核心解决方案在于提出一种名为Queen的40亿参数(4B-parameter)棋类语言模型,通过融合领域专用的“沉默专家”棋类编码器与指令微调的语言模型,构建了一个基于编码器-解码器架构并结合迭代蒸馏算法的新型框架。该框架利用跨注意力机制将棋类编码器的深层状态表示与语言模型进行交互,并通过问答式课程学习从编码器中提取关键棋类概念;随后引入一种自然语言版的贝尔曼更新(Bellman update)机制,使模型在每一步分析其最优候选走法后的局面演变,进而整合形成对当前局面的连贯解释,并将这些高质量解释反向蒸馏回自身。经过七轮迭代优化,模型在棋力上提升超过900 Elo分(从1782升至2697),显著超越所有前沿模型,在对弈强度和残局求解准确率方面均取得突破,且仅使用了其他模型约千分之一的参数量。此外,基于语言模型的评估表明,其生成的解释具有高度流畅性,接近GPT-5.6-Sol(高)水平的逻辑一致性。该方法所提出的架构与训练范式具备广泛适用性,为将语言模型应用于存在“沉默专家”编码器的领域(如棋类游戏、机器人控制、计算机辅助应用等)提供了可复用的技术路径。

链接: https://arxiv.org/abs/2610.03695
作者: Adithya Bhaskar,Jeffrey Cheng,Danqi Chen
机构: Princeton Language and Intelligence, Princeton University(普林斯顿大学语言与智能中心)
类目: Computation and Language (cs.CL)
备注: Code available at this https URL

点击查看摘要

Abstract:Modern chess engines are silent experts: they play at a superhuman level, but do not offer explanations for their play. On the other hand, language models (LMs) can generate plausible-sounding explanations, but their weak playing strength limits the utility of their explanations. We introduce Queen, a 4B-parameter chess-language model that can explain its moves and plans while playing at the level of a typical Grandmaster. Our novel framework enables domain-specific reasoning through complementary components: an encoder-decoder architecture and an iterative distillation algorithm. This architecture integrates a silent expert chess encoder with an instruction-tuned LM through cross-attention, which we train via a question-answering curriculum to extract chess concepts from the encoder’s representations. Building on this domain-adapted model, we iteratively improve its explanations with a natural-language analog of the Bellman update: the model analyzes the positions after its top candidate moves and consolidates them into an explanation of the current position, which is then distilled back into the model. Over seven iterations, our model gains over 900 Elo points (1782 to 2697), substantially surpassing all frontier models on both playing strength and puzzle accuracy, despite containing three orders of magnitude fewer parameters. Furthermore, LM-based evaluations show that our explanations are fluent and approach GPT-5.6-Sol (high) in coherence. The generality of our architecture and training procedure suggests a recipe for applying language models to domains where silent expert encoders are available, like games, robotics, and computer use.

[NLP-1] FrugalEvo: Towards Cost-Aware LLM -Guided Program Evolution

【速读】: 该论文旨在解决生成式人工智能(Generative AI)在复杂计算优化任务中存在“高成本低效用”的核心问题,即现有基于大语言模型(LLM)的进化方法通常以固定迭代次数为优化目标,忽视了实际应用中对单位成本下性能增益最大化的追求。其解决方案的关键在于提出一种成本感知的进化框架FrugalEvo:通过引入一个高性能但高成本的主LLM负责探索创新解法策略,同时使用低成本的辅助LLM执行并迭代优化生成代码,实现“策略探索-高效执行”的分工协同;此外,设计缓存高效的进化过程,通过最大化不同演化步骤间提示(prompt)前缀的共享,显著提升缓存复用率,降低重复计算开销。为更精准衡量在固定成本预算内的优化效果,论文提出了“预算感知曲线下面积”(Budget-Aware Area Under the Curve, BA-AUC),即在累计LLM成本达到预算上限前,最优解随时间演进的评估得分曲线下的面积。实验表明,FrugalEvo在10个数学与系统优化任务中达到或超越当前最优基线(如OpenEvolve、ShinkaEvolve、AdaEvolve、EvoX),并在9项任务上取得更高的BA-AUC;在ALE-Bench-Lite的10个算法优化任务中也展现出更高平均性能;尤其在圆盘打包(circle packing)任务中,仅以1.68美元(GPT-5.6 Terra/Luna)和0.55美元(GLM-5.3及其Flash版本)的成本即达成新的状态领先水平,显著优于包括多智能体方法CORAL和SwarmResearch在内的现有方案(平均成本约50美元),充分验证了其在成本效益上的卓越表现。

链接: https://arxiv.org/abs/2610.03675
作者: Hui Chen,Xuan Qi,James Xu Zhao,Zhaopeng Feng,Shilong Liu,Kuang Xu,Pang Wei Koh,Bryan Hooi
机构: National University of Singapore(新加坡国立大学); University of Washington(华盛顿大学); Princeton University(普林斯顿大学); Stanford University(斯坦福大学)
类目: Neural and Evolutionary Computing (cs.NE); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: 17 pages, 4 figures

点击查看摘要

Abstract:LLM-guided evolutionary methods, such as AlphaEvolve, have emerged as powerful approaches for challenging computational optimization problems, such as circle packing. However, prior work typically optimizes performance gain over a fixed number of iterations. We argue that practical optimization should maximize gain per unit cost. To this end, we propose FrugalEvo, a cost-aware evolutionary framework where a stronger, higher-cost LLM explores solution strategies, and a cheaper LLM implements them and iteratively refines the resulting code. We also design a cache-efficient evolution process, where our harness and prompts maximize the sharing of prefixes across different evolution steps, to improve cache reuse. To measure solution quality throughout a fixed cost budget, we introduce Budget-Aware Area Under the Curve (BA-AUC), defined as the area under the best-so-far evaluation score curve over cumulative LLM cost, up to the budget. Across 10 mathematical and systems optimization tasks, FrugalEvo matches or surpasses state-of-the-art baselines, including OpenEvolve, ShinkaEvolve, AdaEvolve, and EvoX, in final solution quality and achieves higher BA-AUC on 9 tasks. It also achieves higher average performance than these baselines on 10 algorithmic optimization tasks from ALE-Bench-Lite. Notably, on circle packing, FrugalEvo achieves new state-of-the-art performance with GPT-5.6 Terra and Luna for only 1.68 USD and with GLM-5.3 and its Flash variant for only 0.55 USD, matching or surpassing all baselines, including multi-agent methods such as CORAL and SwarmResearch, which cost approximately 50 USD on average.

[NLP-2] Pivot-SD: Efficient Self-Distillation for Masked Diffusion Language Models EMNLP2026

【速读】: 该论文旨在解决生成式语言模型中非自回归的掩码扩散语言模型(dLMs)在复杂推理任务中面临的信用分配(credit-assignment)难题。传统方法在后训练阶段通常对完整序列或整个去噪步骤进行统一优化,未能有效利用去噪过程中少数关键决策点(即“高影响力承诺”)所传递的强信号。其解决方案的关键在于提出一种高效的离线自蒸馏框架Pivot-SD,该框架通过信息增益度量(information-gain metric)识别出显著降低剩余掩码位置不确定性的关键决策点(称为“枢纽”或pivots),并仅对这些pivots进行监督学习:成功轨迹中的pivots采用交叉熵损失进行正向优化,失败轨迹中的pivots则通过目标不可似然(targeted unlikelihood)损失进行反向引导,其余未被选中的部分保持不变。实验表明,仅需200个问题和每题4次采样,Pivot-SD即可在数学与代码基准上超越全序列监督微调(SFT)及预算匹配的扩散强化学习基线,显著提升了模型性能。

链接: https://arxiv.org/abs/2610.03665
作者: Seo Hyun Kim,Sunwoo Hong,Younwoo Choi,Chen-Hao Chao,Se-Young Yun,Rahul G. Krishnan
机构: KAIST AI(韩国科学技术院人工智能研究中心); University of Toronto(多伦多大学); Vector Institute(向量研究所)
类目: Machine Learning (cs.LG); Computation and Language (cs.CL)
备注: EMNLP 2026 Main (Oral)

点击查看摘要

Abstract:Masked diffusion language models (dLMs) offer a promising parallel alternative to autoregressive models for complex reasoning. However, they face a distinct credit-assignment challenge, since a few commitments during denoising sharply reduce the uncertainty over the remaining masked positions and shape much of the response. Most post-training recipes for dLMs do not use this signal to decide which tokens to train on: they typically train on the final text or assign rewards to whole denoising steps, rather than selecting the individual commitments that shape the response. We introduce Pivot-SD, an efficient offline self-distillation framework that supervises only these high-impact commitments (pivots). Pivot-SD selects pivots using an information-gain metric measuring uncertainty reduction over the remaining masked positions. Pivots from successful trajectories are trained with cross-entropy, and pivots from failed trajectories with targeted unlikelihood, leaving the rest of the failed trajectory untouched. Using only 200 questions and four rollouts each, Pivot-SD improves LLaDA-8B-Instruct over full-sequence SFT and budget-matched diffusion RL baselines across math and code benchmarks.

[NLP-3] World Embedding Benchmark

【速读】: 该论文旨在解决视频生成与世界模型中物理保真度(physical fidelity)的表征机制不清晰的问题,即当前视频表示如何编码物理信息尚缺乏系统理解。其核心解决方案是提出世界嵌入基准(World Embedding Benchmark),包含来自流体力学、固体力学、动力学及光学-电磁学等80个物理家族的8,000个受控仿真案例,每个案例均配对渲染视频与仿真生成的物理标注,支持文本-视频检索、物理属性回归和多选视频描述配对分类三项互补任务,以区分跨模态物理对齐与定量物理信息可恢复性之间的差异。实验表明,预训练的多模态嵌入模型在跨族检索和族内配对分类中表现较弱,接近随机水平,而轻量级探测器可从冻结的视频嵌入中有效恢复物理信息;通过引入特定于物理的视频-文本对进行持续对比学习虽提升了检索与配对分类性能,却导致物理属性回归能力下降,揭示了对齐性与定量信息可恢复性间的权衡关系。最终,利用嵌入向量实现参考视频的检索增强生成(retrieval-augmented generation),结合MiniMax-H3模型,结果显示更强的检索模型带来更显著的生成视频物理保真度提升。研究强调需联合评估物理对齐与属性可恢复性,并验证了物理表示在提升视频生成质量中的实际价值。

链接: https://arxiv.org/abs/2610.03632
作者: Yiqi Liu,Ruifeng Yuan,Yang Wang,Long Li,Fengyu Cai,Hou Pong Chan,Jialin Yu,Hao Zhang,Chenghua Lin,Chenghao Xiao
机构: World-Embedding Team
类目: Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Physical fidelity has received increasing attention in world models and video generation, yet how video representations encode physical information remains less understood. We introduce the World Embedding Benchmark, comprising 8,000 controlled simulation cases from 80 families spanning fluid mechanics, solid mechanics, dynamics, and optics electromagnetism. Each case pairs a rendered video with simulation-derived physical annotations, supporting three complementary tasks: text-video retrieval, physical-property regression, and multiple-choice video-description pair classification. We use these tasks to distinguish cross-modal physical alignment from the recoverability of quantitative physical information. Evaluated pre-trained omnimodal embedding models show weak retrieval and near-chance within-family pair classification, while lightweight probes recover useful physical information from frozen video embeddings. Continual contrastive training with physics-specific video-text pairs improves retrieval and pair classification but degrades physical-property regression, revealing a trade-off between alignment and quantitative information recoverability. Finally, we use the embeddings to retrieve reference videos for retrieval-augmented generation with MiniMax-H3. Retrieved references improve the physical fidelity of generated videos, with stronger retrieval models yielding larger gains in our experiments. Together, these findings highlight the need to evaluate physical alignment and property recoverability jointly, and demonstrate the utility of physical representations for improving video generation.

[NLP-4] FALCON: A Model and Dataset Agnostic Framework for Synthetic Data Generation for NL2SQL Pairs EMNLP

【速读】: 该论文旨在解决现有自然语言到结构化查询语言(NL-to-SQL)数据生成方法忽视真实场景中用户查询的语义模糊性,导致生成的数据过于简化、无法有效训练模型应对复杂真实世界知识访问任务的问题。其解决方案的关键在于提出FALCON框架,通过结合保留字驱动的SQL种子生成与基于角色设定(persona-based)的提示策略,生成结构复杂且具备真实语义模糊性的自然语言查询;同时采用基于对齐的过滤机制,能够准确区分真正错误的样本与结构复杂但逻辑正确的有效查询,从而在保持数据难度的同时提升质量。该框架利用轻量级开源模型实现低成本生成,且具备模型与数据库无关的特性,支持组织本地化构建高复杂度训练数据。实验表明,基于FALCON生成的数据在SQL复杂度和自然语言丰富性上均超越现有基准,经人类评估验证其高质量一致性,并在不同规模模型上均表现出优越性能;混合训练策略进一步证明其在兼顾简单查询表现的同时,显著提升复杂查询上的泛化能力。

链接: https://arxiv.org/abs/2610.03625
作者: Darian Lee,Shannon Rumsey,Jack St. Clair,Xinyi Tang,Aditya Bansal,Yuanming Shi
机构: University of California, Santa Cruz; Adobe
类目: Computation and Language (cs.CL); Databases (cs.DB); Machine Learning (cs.LG)
备注: Accepted to AKBC Workshop, EMNLP

点击查看摘要

Abstract:Relational databases are among the most widely deployed forms of structured knowledge, and natural language access to them requires grounding language onto schema entities and relations while handling the ambiguity inherent in how people phrase requests. Existing synthetic NL-to-SQL data generation methods largely ignore this ambiguity and produce oversimplified queries that fail to prepare models for the complexity of real-world structured knowledge access. We present FALCON, a framework that generates realistic, ambiguity-aware NL-to-SQL data matching the complexity of challenging real-world benchmarks, at low cost using compact open models. Our approach combines reserved-word SQL seeding and persona-based prompting to generate structurally complex queries, while alignment-based filtering preserves difficulty by distinguishing genuinely incorrect examples from complex but valid queries. Human evaluation confirms consistent high quality across model sizes, and our generated data exceeds existing benchmarks in both SQL complexity and natural language richness. Difficulty-stratified analysis shows models trained on FALCON data increasingly outperform baseline-trained models as query complexity increases, validating our pipeline’s success in generating challenging training data. When combined with a small proportion of existing benchmark data, mixed training recovers performance on simpler queries while preserving these advantages on complex ones. The model- and database-agnostic design enables organizations to generate high-complexity NL-to-SQL training data locally without external APIs.

[NLP-5] Writerslogic at the CLEF 2026 SimpleText Track: Multi-Candidate LLM Simplification and Stacked Complexity Spotting

【速读】: 该论文旨在解决医学文本简化(text simplification)与复杂性检测(complexity spotting)两大核心问题,聚焦于提升生物医学领域文本的可读性与生成内容的准确性。针对任务1(句子级简化),其关键解决方案是构建一个基于GPT-4o-mini的多候选生成流水线,通过不同温度参数生成五个简化版本,并采用无参考评分机制进行优选,该机制综合考虑压缩率、源词保留度、Cochrane平易语言摘要词汇使用及词汇简单性等指标,以实现高质量的语义保真简化。在任务2(复杂性检测)中,其核心技术在于将幻觉检测建模为自然语言推理(Natural Language Inference, NLI)任务,利用35万条标注数据对DeBERTa-v3-large模型进行微调,以源句为前提(premise)、候选句为假设(hypothesis),直接学习区分内容是否源自原始来源,从而有效识别过量生成(overgeneration)与多类别错误类型。实验结果表明,该方法在英语及多语言生物医学文本上均表现出色,尤其在任务2的二分类和多分类错误识别中分别取得0.8081(0.8085集成)和0.804的优异性能,位居各赛道前列。

链接: https://arxiv.org/abs/2610.03567
作者: David L. Condrey
机构: 未知
类目: Computation and Language (cs.CL)
备注: 11 pages, 3 tables. Notebook for the SimpleText Lab at CLEF 2026. Code: this https URL

点击查看摘要

Abstract:We describe the Writerslogic team’s participation in the CLEF 2026 SimpleText shared task, addressing Task 1 (text simplification) and Task 2 (complexity spotting). For Task 1, we develop a multi-candidate generation pipeline using GPT-4o-mini that produces five simplification candidates per sentence at varying temperatures, then selects the best candidate using a reference-free scoring heuristic that rewards compression, source word retention, Cochrane Plain Language Summary vocabulary usage, and lexical simplicity. On Task 1.1 (sentence-level simplification), our Claude Sonnet 4 submission achieves SARI 47.43 and BLEU 14.21, the top-ranked sentence-level system (3rd on the combined Task 1 leaderboard, behind two document-level submissions). For Task 2, we fine-tune a DeBERTa-v3-large NLI model on 350K labeled (source, sentence) pairs, framing hallucination detection as natural language inference. The model reads the most relevant source sentence as premise and the candidate as hypothesis, directly learning to distinguish grounded from hallucinated content. On Task 2.1 (binary overgeneration identification), our fine-tuned DeBERTa system achieves 0.8081 document-level macro F1 (0.8085 in our best ensemble), the top-ranked entry within the identification track and 2nd among teams overall, behind AIIR Lab (0.8197). On Task 2.2 (multi-class error classification), our best submission reaches 0.804 multiclass accuracy, ranking 2nd among unique teams behind AIIR Lab (0.827). We evaluate both tasks on English and multilingual biomedical text from Cochrane systematic reviews.

[NLP-6] Writerslogic at PAN 2026: Process over Content for Robust Detection under Domain Shift

【速读】: 该论文旨在解决生成式 AI(Generative AI)在跨领域迁移场景下文本分析任务中的性能退化问题,核心挑战在于模型在分布偏移(distribution shift)条件下特征鲁棒性不足。其解决方案的关键在于提出一个统一的分析框架:特征在分布偏移下的稳健性由训练与测试分布之间的支持重叠(support overlap)决定,而非训练集规模大小。基于此框架,作者构建了三类特征分类体系——领域锚定型(domain-anchored)、领域可迁移型(domain-portable)和领域不变型(domain-invariant),并验证了生成过程属性特征(如压缩度量、字符n-gram、词汇指纹等)相较于生成内容特征在跨域场景中具有更强的稳定性。在三个任务中,系统设计均围绕“优先采用生成过程相关特征”展开:在推理轨迹检测任务中,通过融合数学领域训练与多领域测试间的结构化特征实现源识别与安全分类的优异表现;在Voight-Kampff生成式AI检测任务中,采用基于多模型集成与学习堆叠的校准系统,在平衡性良好的子指标下达到0.891的F1;在多作者写作风格分析任务中,提出结合谱聚类、归一化压缩距离与神经困惑度的混合检测架构,尽管因平台配置错误未完成正式评估,但其设计与先验预测已充分验证方法的有效性。

链接: https://arxiv.org/abs/2610.03565
作者: David L. Condrey
机构: 未知
类目: Computation and Language (cs.CL)
备注: 13 pages, 1 figure, 6 tables. Notebook for the PAN Lab at CLEF 2026. Code this https URL and this https URL

点击查看摘要

Abstract:We describe the Writerslogic systems for three PAN at CLEF 2026 shared tasks (Reasoning Trajectory Detection, Voight-Kampff Generative AI Detection, and Multi-Author Writing Style Analysis), unified by a shared analytical framework: feature robustness under distribution shift is governed by support overlap between training and test distributions, not by training-set effect size. This yields a taxonomy (domain-anchored, domain-portable, domain-invariant) that explains why generator-specific features die under domain shift while vocabulary fingerprints (hapax ratio, Yule’s K, Heaps’ exponent), compression measures, and character n-grams survive. On Reasoning Trajectory Detection, where training was entirely mathematics and 84 percent of test was unseen domains, the framework guided system design to 1st place in source detection (0.85 macro F1 via Opus-Sonnet agreement) and 3rd place in safety classification (0.66 macro F1 via query-refusal decomposition). For Voight-Kampff, we built a calibrated ensemble of DeBERTa-v2 (ONNX), multi-seed LightGBM with 44 domain-portable stylometric features, and SVM on n-gram TF-IDF, combined via learned stacking with isotonic calibration; the best configuration achieved 0.891 on the PAN 2026 test set with balanced sub-metrics (0.853 to 0.902 across all evaluation dimensions). For Multi-Author Writing Style Analysis, we describe a system fusing spectral clustering over character n-gram similarity graphs, normalized compression distance for local boundary detection, and SmolLM-135M perplexity for neural change-point detection; a platform mix-up meant our run never reached the official evaluation, so we report the design and its a priori predictions. Across all three tasks, features measuring generation process properties are designed to outperform features measuring generated content properties under domain shift.

[NLP-7] Author Representation Strategies for Zero-Shot Authorship Attribution: A Comparative Study of LLM -Based and Embedding-Based Approaches

【速读】: 该论文旨在解决零样本(Zero-Shot, ZS)场景下作者归属(Authorship Attribution, AA)任务中因缺乏特定任务监督而导致的细粒度风格特征捕捉难题。其核心挑战在于如何在无标注训练数据的情况下,有效建模并区分作者独特的写作风格。解决方案的关键在于设计和评估多种作者表征(author representation)策略,包括代表性文本样本、大语言模型(LLM)生成的风格描述以及基于LISA的风格嵌入(style embeddings)。研究发现,仅依赖标签提示的零样本方法效果不佳,而引入作者特异性表征可显著提升性能。其中,所提出的两阶段嵌入式框架通过候选空间缩减与嵌入维度选择优化了风格嵌入的匹配效率,实现了最优整体性能;相比之下,LLM生成的风格描述虽压缩了表示规模,但牺牲了一定的归属准确率。研究结果强调了高质量作者表征在零样本作者归属中的决定性作用,同时指出当前开源大语言模型在缺乏更有效的表征学习机制时,仍难以实现鲁棒的作者归属。

链接: https://arxiv.org/abs/2610.03531
作者: Nudrat Habib,Tosin Adewumi,Sana Sabah Al-Azzawi,Marcus Liwicki,Elisa Barney
机构: Luleå University of Technology (吕勒奥理工大学)
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Authorship Attribution (AA) requires capturing fine-grained stylistic characteristics, making it particularly challenging in zero-shot (ZS) settings where no task-specific supervision is available. In this work, we investigate the effect of author representations on ZS AA by evaluating a label-only prompting baseline together with three author representation strategies: representative writing samples, LLM-generated descriptions, and style embeddings (LISA). The first three approaches perform attribution using LLM prompting, while the embedding-based approach uses style embeddings with cosine similarity. We investigate the influence of prompt design and propose a two-stage embedding-based attribution framework that combines candidate space reduction with embedding-dimension selection. The results show that label-only ZS AA is ineffective, while incorporating author-specific representations consistently improves attribution performance. Among the evaluated approaches, the proposed two-stage LISA framework achieves the strongest overall performance, whereas LLM-generated style descriptions provide a substantially more compact representation of author style at the cost of some attribution performance. These findings demonstrate the importance of author representation in ZS AA, while indicating that current open-source LLMs remain insufficient for robust attribution without more effective representation learning.

[NLP-8] Divergence controls entropy in distillation

【速读】: 该论文旨在解决生成式大语言模型训练中知识蒸馏(Knowledge Distillation)机制的内在性质不明确的问题,特别是从信息熵的角度揭示学生模型(student)熵与蒸馏目标所定义的数据分布及分歧度量(divergence)之间的关系。其核心解决方案的关键在于:通过理论分析与实证验证,阐明不同分歧度量对学生模型熵的影响——前向KL散度(forward KL)会人为提升学生模型的熵,使其高于教师模型;而反向KL散度(reverse KL)则会抑制熵,尤其在师生差异过大时导致熵急剧下降;此外,混合分歧度量在训练初期平滑调节熵,但在收敛时产生突变。研究进一步指出,基于策略内(on-policy)蒸馏的低熵特性源于逐标记级(token-level)反向KL,而非策略采样本身。因此,分歧度量本质上充当了隐式的熵正则化项,这一作用在自蒸馏(self-distillation)场景中尤为显著:当利用特权信息(privileged information)降低熵时,最优的分歧超参数需具备补偿此熵压缩的能力。

链接: https://arxiv.org/abs/2610.03529
作者: Nicolas Zucchet,Scott W. Linderman
机构: Stanford University (斯坦福大学)
类目: Machine Learning (cs.LG); Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Distillation has become a core primitive of large language model training, but its properties are not yet well understood. We take an entropic perspective, studying how the entropy of the student depends on the data and the divergence that define the distillation objective. We prove that forward KL inflates the entropy of the student above that of the teacher. Since cross-entropy training is a special case, this yields an identity that we verify quantitatively in pretraining and supervised finetuning. Other divergences come with no such guarantee: reverse KL deflates entropy until the gap between student and teacher gets too large, and interpolating between the two changes entropy smoothly early in training but abruptly at convergence. The lower entropy of on-policy distillation comes from token-level reverse KL, not from on-policy sampling. The divergence therefore acts as an implicit entropy regularizer, whose role is clearest in self-distillation: as conditioning on privileged information deflates entropy, the divergence hyperparameters that work best are those that compensate for it.

[NLP-9] Structured Composition of Verifiable Atomic Insights for Table-to-Report Generation

【速读】: 该论文旨在解决表格到报告生成(table-to-report generation)任务中的核心挑战:如何系统性地发现跨表、跨属性及多分析视角下可验证的复合洞察,并将其组织为连贯、完整且可追溯的证据链。现有方法主要依赖顺序式反应型数据代理或直接的大语言模型(LLM)生成,存在探索偏差(exploration bias)问题,即早期局部观察会限制后续决策,导致模型过早聚焦于局部分析而遗漏跨表或跨维度的关联证据。其解决方案的关键在于提出ComInsight框架,将洞察发现重构为原子证据(atomic evidence)的组合过程。首先,定义原子洞察为符合预设分析模式的最小可执行分析单元,并基于数据库模式与内容枚举所有有效原子洞察;随后构建多关系洞察图(multi-relational insight graph),其中节点表示已验证的数据事实,边编码逻辑、时间或层级关系;最后,通过一组组合算子系统性地融合原子节点,生成更高阶的复合结论。每个复合输出均附带可执行的SQL语句和细粒度溯源信息,确保结果的完全可验证性。在InsightBench、DDR-Bench和T2R-Bench三个基准上的实验表明,ComInsight在事实正确性、新颖性和结构完整性方面持续优于强基线,为表格到报告生成提供了可靠、高效且可解释的新路径。

链接: https://arxiv.org/abs/2610.03525
作者: Teng Lin,Xinyu Liu,Nan Tang
机构: HKUST(GZ) (香港科技大学(广州)); DSA Thrust (数据科学与人工智能研究专项)
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Table-to-report generation refers to the task of automatically generating article-level analyt- ical reports from relational tables and is an essential capability for automated data science and decision support. Its central challenge lies in systematically discovering verifiable com- posite insights across tables, attributes, and analytical perspectives, and organizing them into coherent, complete, and traceable evidence chains. Existing methods primarily rely on sequential, reactive data agents or direct Large Language Model(LLM) generation. They suffer from exploration bias: early local observations constrain subsequent actions, causing models to focus prematurely on local analyzes and miss cross-table or cross-dimensional evidence. We propose ComInsight, which reformulates insight discovery as the composition of atomic evidences. We first define an atomic insight as the smallest executable analytical unit conforming to a predefined analysis pattern and enumerate all valid atomic insights from database schema and content. These atoms are then organized into a multi-relational insight graph, where nodes represent verified data facts and edges encode logical, temporal, or hierarchical relations. Finally, a set of composition operators systematically fuses atomic nodes into higher-order composite conclusions. Every composite output is accompanied by executable SQL and fine-grained provenance, ensuring full verifiability. Across three benchmarks InsightBench, DDR-Bench, and T2R-Bench, ComInsight consistently outperforms strong baselines in factual correctness, novelty, and structural completeness. We believe ComInsight offers a reliable, efficient, and explainable path toward table-to-report generation.

[NLP-10] Learning from Repaired Reasoning : Root-Cause-Guided On-Policy Distillation

【速读】: 该论文旨在解决在策略自蒸馏(On-policy self-distillation, OPSD)中,参考解提供的事后指导与学生模型自身推理错误之间存在的“推理错位”问题,以及因在整个推理轨迹上施加相同事后指导而引发的“蒸馏陷阱”问题。其核心挑战在于:传统方法仅提供正确答案作为监督信号,无法精准定位学生推理中的具体错误根源,导致学生可能简单复制正确结论但未修正内在逻辑缺陷;同时,统一的指导会抑制合理推理路径的探索,干扰对实质性错误的纠正。为克服上述问题,本文提出根因引导的策略自蒸馏(Root-Cause-Guided On-Policy Distillation, RC-OPD),其关键创新在于利用学生自身推理过程的修复(repair)来生成针对性指导——针对每次失败尝试,系统性地定位最早实质性错误点,构建局部修正方案,并以修正后的中间结果作为有效前缀的锚点。通过迭代诊断—修复—延续机制,在固定修复预算内测试修复效果并发现新错误。对于成功得出正确答案的修复链,采用根因导向的蒸馏监督错误段落,同时以锚点引导的蒸馏支持对应的合理推理前缀。实验表明,该方法有效缓解了推理错位与蒸馏陷阱,显著提升多数据集及不同模型规模下的性能表现。

链接: https://arxiv.org/abs/2610.03515
作者: Chenglei Shen,Haoyang Yao,Weijie Yu,Song Jin,Xiao Zhang,Jun Xu
机构: Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学高瓴人工智能学院); School of Software and Microelectronics, Peking University(北京大学软件与微电子学院); School of Information Technology and Management, University of International Business and Economics(对外经济贸易大学信息学院)
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:On-policy self-distillation (OPSD) uses reference solutions as privileged hindsight to supervise student-generated reasoning trajectories. However, reference-based guidance may explain a correct solution without addressing why the student’s own reasoning fails. This reasoning mismatch between the guidance provided and the correction needed can encourage the student to borrow correct conclusions while leaving its reasoning errors unresolved. Moreover, applying the same hindsight throughout the trajectory risks a distillation trap, where unnecessary constraints on valid reasoning compete with correction of substantive errors. To address these issues, we propose Root-Cause-Guided On-Policy Distillation (RC-OPD), which uses repairs of the student’s own reasoning to provide guidance that addresses its specific errors while building on valid progress. For each failed attempt, RC-OPD locates the earliest substantive error, develops a local correction, and uses the corrected intermediate result as an anchor for the valid prefix. An iterative diagnosis–repair–continuation process tests the repairs through student continuation, identifying further errors within a fixed repair budget. For repair chains that reach a correct answer, root–cause–guided distillation uses failure diagnoses and corrective goals to supervise the erroneous segments, while anchor-guided distillation supports the corresponding valid prefixes with reasoning chains leading to the repaired intermediate results. We evaluate RC-OPD across multiple datasets and model scales. Extensive experiments and analyses show that it mitigates reasoning mismatch and the distillation trap, yielding substantial performance gains.

[NLP-11] Single-Pass Uncertainty Heads for Claim-Level Hallucination Detection in Persian Medical Language Models

【速读】: 该论文旨在解决医学领域大语言模型(LLM)中幻觉检测的挑战,特别是针对波斯语医学语言模型缺乏高效、低成本且可迁移的不确定性估计方法的问题。现有方法依赖重复采样导致计算开销高,而现有的不确定性头(Uncertainty Head)资源难以直接迁移到新的模型架构和语言上。为此,研究提出将大语言模型不确定性头(LLM Uncertainty Head, LUH)框架适配至基于Aya-Expanse-8B的波斯语医学模型,并以Gaokerena-V和Gaokerena-R作为预训练基础模型进行对比分析。关键解决方案在于:在冻结主干网络注意力图与词元概率的基础上,构建轻量级的声明级别(claim-level)不确定性头,实现单次推理下的幻觉检测,无需检索或重复采样。研究构建了两个直接以波斯语生成的成对声明级幻觉数据集(各含1600条响应),实验表明所提方法在测试集上分别获得PR-AUC 0.4820和0.4652(分别为随机基线的2.30倍和2.66倍),以及ROC-AUC 0.7852和0.7810,验证了其在低资源、单次推理场景下对波斯语医学模型幻觉检测的有效性与可行性。

链接: https://arxiv.org/abs/2610.03482
作者: Mehrdad Ghassabi,Pedram Rostami,Hamidreza Baradaran Kashani,Sadra Hakim,Audrina Ebrahimi
机构: University of Isfahan (伊斯法罕大学); University of Tehran (德黑兰大学); University of Windsor (温莎大学); University of Texas at Dallas (德克萨斯大学达拉斯分校)
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Hallucination detection is particularly important for medical language models, but repeated-sampling approaches are expensive and existing uncertainty-head resources do not directly transfer to a new backbone and language. We adapt the LLM Uncertainty Head (LUH) framework to Aya-Expanse-8B-based Persian medical models, using Gaokerena-V and Gaokerena-R as two previously developed backbones. We first examine response variability on a 168-question Iranian medical entrance examination and observe substantially lower five-run consistency for Gaokerena-V than for Aya-Expanse-8B, whereas Gaokerena-R is comparable to Aya-Expanse-8B. We then construct two paired claim-level hallucination datasets directly in Persian, containing 1,600 responses for each backbone, and train lightweight claim-level heads on frozen backbone attention maps and token probabilities. On held-out test splits, the heads obtain PR-AUCs of 0.4820 and 0.4652, corresponding to 2.30 and 2.66 times their respective random baselines, and ROC-AUCs of 0.7852 and 0.7810. The heads require neither retrieval nor repeated sampling at inference time. These results provide an initial study of single-pass claim-level uncertainty estimation for Persian medical language models; the test splits are small and the labels are automatically generated.

[NLP-12] A Near-Zero Monitor Readout Is Not Evidence of Behavioral Control NEURIPS2026

【速读】: 该论文旨在解决后训练阶段使用可验证奖励(verifiable rewards)时引发的奖励黑客(reward hacking)问题,尤其关注在训练目标中引入监控机制(monitor)以替代仅依赖离线审计的局限性。其核心挑战在于:低监控读出值(low monitor readout)并不能有效表明模型行为已被正确控制,因为即使监控指标数值较低,仍可能伴随高比例的奖励黑客行为。解决方案的关键在于揭示监控读出值的误导性——不同监控机制(如域内激活探测器、基于早期承诺惩罚的两种策略)虽在离线评估中表现一致,但其读出值反映的是不同量度,且在训练过程中低读出值并不等同于行为控制。研究发现,在代码生成任务中,尽管所有探测器读出值均处于数值下限,且多数训练运行的中位数为零,但相同零中位数结果下的模型表现从混合型低黑客行为到近乎纯奖励黑客不等,仅由随机种子决定。文本分析进一步揭示一种“前缀失效模式”(prefix failure mode),即通用规划和填充结构会延迟漏洞触发时间,却未真正消除其存在。因此,低承诺度(low commitment)不能区分低黑客率与对漏洞的延迟采纳。结论强调:仅依赖离线判别和低监控对齐读出不足以证明行为控制,必须引入独立于监控机制的行为验证(out-of-band behavioral check)。研究聚焦于终点读出状态而非其演化过程,突显了对监控有效性进行动态评估与外部验证的必要性。

链接: https://arxiv.org/abs/2610.03458
作者: Zhe Zhou,Tianhua Tao
机构: University of Washington(华盛顿大学)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: 17 pages, 2 figures, 10 tables. Accepted as a poster at the NeurIPS 2026 Workshop on Foundations of LLM Post-Training in Changing Environments (FLLMPT)

点击查看摘要

Abstract:Post-training with verifiable rewards can induce reward hacking, motivating the use of monitors within the training objective rather than solely for offline auditing. We show that a low monitor readout does not identify whether such an intervention controls behavior. In a code-generation environment whose dominant exploit is available at the start of the reasoning trace, we train policies against three monitors that pass the same offline gate: an in-domain activation probe and two penalties conditioned on how early the policy commits to its own final answer. The probe score is at its numerical floor from the first recorded training step, and the trained-score median is zero for every prefix-trained run at the endpoint. These readouts estimate different quantities, and we do not compare their scales; within each monitor family, however, low values do not establish behavioral control. Within one fixed configuration, prefix-trained runs with the same zero-median trained score range, by seed alone, from a mixed regime with a low hacking share to near-pure reward hacking. All probe runs reach the hacking regime, but their floor-level readout reflects a mismatch between the position where the probe was validated and the position where it was read during training, not a second instance of this ambiguity. Text-level analysis identifies a prefix failure mode: generic planning and filler shells postpone the exploit past the cut without eliminating it from the final output. Low measured commitment therefore does not distinguish a low hacking share from delayed commitment to the exploit. Offline discrimination and low monitor-aligned readouts are insufficient evidence of behavioral control; an out-of-band behavioral check is required. We characterize the endpoint readout, not its evolution. Code is available at this https URL.

[NLP-13] Passing the Test You Trained On: Re-evaluating Prompt-Injection Detectors for LLM Agents

【速读】: 该论文旨在解决当前大语言模型(LLM)代理在实际应用中依赖公开基准测试得分来选择提示注入检测器(prompt-injection detector)所存在的有效性问题,即这些基准测试得分是否能准确预测检测器在真实代理环境中的表现。研究发现,不同基准测试之间的检测性能排名转移性极差:例如,在BIPIA基准上表现最佳的检测器在AgentDojo中仅能捕获2%的注入攻击,而另一检测器虽在AgentDojo中可捕获72%的攻击,但在tau-bench上仅能捕获15%。尽管如此,对工具输出的误报率(false-positive rate)在两个代理基准间具有较好的可转移性,范围从0%到超过90%。进一步分析表明,当训练数据公开时,训练输入的形式是决定检测性能的关键因素——例如,领先于BIPIA的检测器虽使用完整输入进行训练,但其对短提示形式的攻击字符串并无识别优势;而表现最佳的检测器并未使用任何基准测试数据,而是基于类代理输入进行训练。因此,解决方案的关键在于:评估检测器的实际部署效果必须基于代理自身的工具输出,以低误报率为标准,并严格审计检测器的训练数据构成,从而确保评估结果与真实场景高度相关。

链接: https://arxiv.org/abs/2610.03448
作者: Zhuowen Liu
机构: Japan Advanced Institute of Science and Technology (JAIST); Nomi, Ishikawa, Japan
类目: Cryptography and Security (cs.CR); Computation and Language (cs.CL)
备注: 12 pages, 5 figures, 4 tables. Code: this https URL

点击查看摘要

Abstract:LLM agents increasingly screen tool outputs with small prompt-injection detectors, and teams choose among detectors by their scores on public benchmarks. We ask whether those scores predict how a detector behaves inside an agent. We replay the ground-truth tool calls of two agent benchmarks, AgentDojo and tau-bench, without an LLM to obtain tool outputs that are benign by construction, label injected outputs by differential replay, and evaluate fifteen detectors, including Meta’s Prompt Guard 2, and two task-aware LLM judges on these outputs and on the BIPIA benchmark. Detection rankings transfer poorly between benchmarks: the best detector on BIPIA catches 2% of AgentDojo injections at a 1% false-positive rate, and a detector that catches 72% of AgentDojo injections catches 15% on tau-bench. False-positive rates on tool outputs, which range from none to over 90%, do transfer between the two agent benchmarks. Where training data is public, the form of the training inputs explains the results. The BIPIA leader was trained on full BIPIA inputs, but having seen InjecAgent’s attack strings as short prompts does not help it find them inside tool outputs; the best detector on both agent benchmarks shares no data with any benchmark and was trained on agent-style inputs. Evaluations meant to inform deployment should use the agent’s own tool outputs, report detection at a low false-positive rate, and audit what the detector was trained on.

[NLP-14] CLIMB: Confidence-Guided Complementary Evidence for Multimodal Retrieval-Augmented Generation EMNLP2026

【速读】: 该论文旨在解决多模态大语言模型(Multimodal Large Language Models, MLLMs)在知识密集型视觉问答任务中,因依赖图像和模型参数化知识而难以获取充分外部文本证据的问题。现有基于检索增强生成(Retrieval-Augmented Generation, RAG)的多模态系统通常采用Top-K检索或重排序策略,存在冗余文本返回及对答案更新是否充分由检索证据支持缺乏有效控制的缺陷。为此,论文提出一种无需训练的推理时框架CLIMB,其核心在于:首先通过类似最大边际相关性(MMR)的目标构建一个紧凑且互补的证据池,以平衡查询相关性与段落层面的冗余度;随后在固定证据池内进行置信度可控的迭代优化——由一个R/E/C评议员(评估相关性、证据特异性与跨模态对齐)对候选证据进行评分,并结合基于证据的置信度估计器,仅当置信度提升时才接受更新后的答案。该设计实现了简洁的停止准则,避免了不必要的迭代优化,且不需修改底层检索器或MLLM。实验表明,CLIMB在Encyclopedic-VQA和InfoSeek数据集上均显著优于现有检索增强基线,消融实验进一步验证了互补证据池构建、基于评议员的评分机制以及迭代置信度控制优化各自对性能提升的关键作用。

链接: https://arxiv.org/abs/2610.03421
作者: Hang Gao,Wujiang Xu,Zhixing Zhang,Kai Mei,Jingyi Yang,Dimitris N. Metaxas
机构: Rutgers University (罗格斯大学); Google(谷歌)
类目: Computation and Language (cs.CL)
备注: EMNLP 2026 Findings

点击查看摘要

Abstract:Multimodal large language models (MLLMs) have shown strong visual reasoning abilities, but knowledge-intensive visual question answering often requires external textual evidence beyond the image and the model’s parametric knowledge. Existing multimodal RAG systems commonly rely on Top- K retrieval or reranking, which may return redundant passages and provide limited control over whether an answer update is sufficiently supported by the retrieved evidence. We propose \textitCLIMB, a training-free inference-time framework for multimodal RAG. CLIMB first constructs a compact complementary evidence pool using an MMR-style objective that balances query relevance and passage-level redundancy. It then performs confidence-controlled refinement within this fixed pool: an R/E/C critic scores passages by relevance, evidence specificity, and cross-modal alignment, while an evidence-grounded confidence estimator accepts an updated answer only when the estimated confidence increases. This design provides a simple stopping criterion and reduces unnecessary refinement without modifying the underlying retriever or MLLM. Experiments on Encyclopedic-VQA and InfoSeek show that CLIMB consistently improves over retrieval-augmented multimodal baselines. Ablations further indicate that complementary pooling, critic-based scoring, and iterative confidence-controlled refinement each contribute to the final performance.

[NLP-15] Benchmarking Candidate Coverag e in Typed Decision Models

【速读】: 该论文旨在解决生成式 AI 在面对参考答案缺失或候选选项不完整时,其决策模型是否具备识别“答案缺失”能力以及能否正确避免错误拒绝有效候选答案的问题。现有评估方法在完整候选集下的准确率无法反映模型对参考答案缺失的感知能力或对有效候选的保留能力。为此,研究提出了一种配对候选-覆盖度基准测试协议(paired candidate-coverage benchmark protocol),通过构造“存在/不存在”成对样本以模拟真实场景中的参考标签遗漏问题,并在 AG News、DBpedia、Emotion 与 TREC 四个数据集上对 Laya 与 Jev 两个模型进行了初步评估。关键解决方案在于引入严格的对照设计:保持文本内容与请求一致,使用冻结的输入文本和统一的请求格式,确保测试结果可比性;同时通过名称变体保持描述、成员及顺序的一致性,以控制变量。实验结果显示,两模型在原生拒绝行为上差异显著:在包含五个自然命名候选的 TREC 数据集中,Laya 能检测出 97.2% 的缺失答案情况,但误拒率达 69.7%;而 Jev 的对应指标为 24.8% 和 0.0%。经过仅基于校准数据设定阈值后,两模型的误拒率分别降至 3.7% 和 1.8%,表明校准策略对性能有显著影响。此外,对分类性能、得分排序敏感性及接口审计的细粒度分析揭示了分类、打分排序与拒绝策略应独立评估的重要性。该基准虽具描述性且目前仅聚焦于参考标签遗漏问题,尚未涵盖自然界的泛化能力、因果机制验证或新型拒绝方法的建立,但为未来系统性评估生成式 AI 的认知完整性提供了重要起点。

链接: https://arxiv.org/abs/2610.03387
作者: Jiawen Lu,Tongtong Wu
机构: Monash University(蒙纳士大学)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: 19 pages, 1 figure, 8 tables

点击查看摘要

Abstract:Typed decision models return choices or distributions over answer options supplied at request time. Accuracy with complete options does not establish whether a model recognizes that a reference answer is missing or avoids rejecting valid candidates. We present a paired candidate-coverage benchmark protocol and an initial evaluation of Laya and Jev across AG News, DBpedia, Emotion, and TREC. The models receive identical frozen texts and requests: 300 calibration and 589 test texts yield 23,932 predictions per model. Present/absent pairs match ordinary candidate count, and name variants preserve descriptions, members, and order. Native rejection behavior differs sharply: at five TREC candidates with natural names, Laya detects 97.2% of missing-answer cases but falsely rejects 69.7% of present controls; Jev’s rates are 24.8% and 0.0%. Calibration-only none-score thresholds change these rates to 33.9%/3.7% and 45.0%/1.8%, respectively. On DBpedia, Jev’s high coverage-score AUROC supports a stronger operating point, whereas both models have weak complete-set accuracy on Emotion. Competence-conditioned analysis, probability-precision sensitivity, and interface audits show why classification, score ranking, and rejection policies need separate measurement. This initial benchmark is descriptive and limited to reference-label omission; it does not establish natural out-of-scope generalization, causal mechanisms, or a new rejection method.

[NLP-16] Multilingual GSM-Symbolic: What determines capability transfer across languages?

【速读】: 该论文旨在解决生成式 AI 模型在跨语言能力迁移(cross-lingual capability transfer)中的可预测性问题,即当前对模型在一种语言中习得的能力如何迁移到另一种语言的理解极为有限,且现有评估方法依赖于不可比、易饱和的数据集,难以联合分析影响迁移的关键因素。其核心解决方案是提出 Multilingual GSM-Symbolic——一个可扩展的多语言数学推理数据集,涵盖15种语言的30,000组匹配题-答对,并通过符号化模板生成数百万高质量变体,有效避免过拟合并保障泛化能力。基于该数据集,研究首次联合量化了影响跨语言迁移的核心决定因素:模型规模(β = 1.77)、语言资源水平(β = 0.77)、推理能力(β = 0.67)及语言类型学距离(β = -0.25)。研究发现,模型规模与推理能力能显著缩小低资源语言与高资源语言之间的性能差距(β = -0.27 与 β = -0.20),而类型学距离较远的语言则对此类提升不敏感。整体分析框架可解释92%的跨语言性能差异,且仅需在目标语言中使用10个模板即可将未见语言性能预测误差控制在4.19个百分点以内,实现高效、低成本的性能预估。

链接: https://arxiv.org/abs/2610.03367
作者: Kenneth Enevoldsen,Riley Herchert,Sofie Mosegaard,Dan Saattrup Smart,Simon Enni,Isaac Chung,Sofie Bruun,Ayush Sunil Munot,Max Müller-Eberstein,Adnan El-Assadi,Elisa Bassignana,Gianluca Barmina,Hafsteinn Einarsson,Iben Nyholm Debess,Linda Freienthal,Lukas Galke Poech,Mike Zhang,Nicolas Legrand,Vladimir Salnikov,Yevhen Kostiuk,Zafar Hussain,Sagandeep Kaur,Agnes Toftgård,Marie Mattson,Kristoffer Nielbo
机构: Aarhus University (奥胡斯大学); Danish Foundation Models (丹麦基础模型基金会); University of Alabama (阿拉巴马大学); Alexandra Institute (亚历山德拉研究所); Syv.ai (Syv.ai); Foam.io (Foam.io); Indian Institute of Technology Kharagpur (印度理工学院克哈普尔分校); The University of Tokyo (东京大学); IT University of Copenhagen (哥本哈根信息技术大学); Massachusetts General Hospital (麻省总医院); Bocconi University (博科尼大学); University of Southern Denmark (南丹麦大学); University of Iceland (冰岛大学); University of the Faroe Islands (法罗群岛大学); Zendesk (Zendesk); University of Copenhagen (哥本哈根大学); Indian Institute of Technology Madras (印度理工学院马德拉斯分校); National Library of Sweden (瑞典国家图书馆)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:We understand little about how capabilities acquired in one language carry over to another, or what governs this transfer: evaluations rely on incomparable, saturation-prone datasets and rarely examine its determinants jointly. Identifying what predicts transfer would let us avoid exhaustive evaluation across all language pairs and let developers target the factors that limit performance in low-resource languages. To evaluate cross-lingual capability transfer, we introduce Multilingual GSM-Symbolic, an extensible multilingual mathematical dataset covering 30,000 item-matched question-answer pairs and spanning 15 languages. It utilises symbolic templates to prevent overfitting and ensure generalisation by allowing generation of millions of high-quality variations from a single sample. Using Multilingual GSM-Symbolic, we quantify the largest determinants of capability as model size ( \beta = 1.77 ), language resource level ( \beta = 0.77 ), reasoning ( \beta = 0.67 ) and typological distance ( \beta = -0.25 ). This joint estimation allows these determinants to be expressed in terms of one another: a 32B model evaluated in Marathi performs like a 10B model in English. Our findings have important implications for model developers, showing that model size and reasoning narrow the performance gap between low- and high-resource languages ( \beta = -0.27 and \beta = -0.20 , respectively), while similar levers have little or no effect on typologically distant languages. Overall, our analysis framework explains 92% of between-language variation, but only 23% of the model-by-language variation, and predicts a model’s performance on an unseen language within 6.0pp (r=.96). Incorporating measurements from just 10 templates in the target language reduces this to 4.19pp, enabling reasonable estimates of performance with little or no downstream dataset. Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL) Cite as: arXiv:2610.03367 [cs.AI] (or arXiv:2610.03367v1 [cs.AI] for this version) https://doi.org/10.48550/arXiv.2610.03367 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[NLP-17] SyntaxBench: A Statistical Diagnostic Framework for Character-Level Reasoning in Large Language Models

【速读】: 该论文旨在解决大语言模型在需要精确字符级推理任务中表现评估不充分的问题,现有方法多依赖孤立的探针测试和整体准确率,难以揭示模型在细微语法错误敏感场景下的真实能力。其核心解决方案是提出SyntaxBench——一个面向字符级推理的诊断性基准与统计评估框架,包含五项基础任务(字符计数、字母包含性检测、回文识别、编辑距离计算、最长字符串选择)及一项更具挑战性的子串提取压力测试index_to_span,所有任务均采用成对的英语与等长随机字符串输入进行对比,并覆盖零样本、单样本和四样本提示设置。该框架通过精确匹配率、宽松准确率、Cohen’s kappa、McNemar检验、自助法置信区间、Kendall等级相关系数、条件分类指标、分词分析以及多重比较校正等多维度指标,系统评估8个参数量从2B到32B的开源模型在11种推理模式下的表现。关键发现表明:分词方式显著影响性能,随机字符串比英文文本具有更高的字符可见性(每标记1.892 vs. 3.169字符),且随着英文词汇占据更多标记,字符计数准确率下降;推理模式并非普遍有效,部分模型在特定任务中启用思维链反而导致性能下降;而index_to_span任务仍基本未被攻克,最优四样本条件下精确匹配率仅为6.75%。研究强调,字符级推理评估必须结合受控输入、配对测试、分词分析与推理模式剖析,而非仅依赖聚合准确率。

链接: https://arxiv.org/abs/2610.03329
作者: Mohsen Larni(1),Sobhan Ebrahimi Azar(1),Pouyan Nahed(1),Kazem Taghva(1) ((1) Department of Computer Science, University of Nevada, Las Vegas)
机构: University of Nevada, Las Vegas(内华达大学拉斯维加斯分校)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: 32 pages, 17 figures. The first two authors contributed equally. The code will be released soon

点击查看摘要

Abstract:Large language models are increasingly used where small syntactic errors matter, yet character-level reasoning is still evaluated mostly through isolated probes and aggregate accuracy. We introduce SyntaxBench, a diagnostic benchmark and statistical evaluation framework for character-level reasoning. It contains five core tasks, character counting, letter containment, palindrome detection, edit distance, and longest-string selection, plus index_to_span, a harder substring-extraction stress test. The five core tasks use paired English and character-length-matched random-string inputs. index_to_span documents share a 200-500 word band and are not character-length matched. All six tasks use zero-, one-, and four-shot prompts. We evaluate eight open-weight models from 2B to 32B parameters across 11 reasoning-mode configurations. The framework reports exact-match and relaxed accuracy, Cohen’s kappa, paired McNemar tests with odds ratios, bootstrap confidence intervals, Kendall’s tau, class-conditional metrics, tokenization analysis, and multiple-comparison-corrected tests. Three findings stand out. First, tokenization shapes accuracy: random strings are more character-visible than English strings (1.892 vs. 3.169 characters per token), and character-counting accuracy falls as English words occupy more tokens. Second, reasoning mode is not uniformly helpful: Gemma4-31B is nearly unchanged across modes on the near-saturated tasks, while Qwen3.6-27B is worse with thinking on palindrome detection (0.952 non-thinking vs. 0.886 thinking at four-shot). Third, index_to_span remains largely unsolved; the best four-shot exact-match accuracy is 6.75%. Character-level evaluation needs controlled inputs, paired tests, and analyses of tokenization and reasoning mode rather than aggregate accuracy alone. Comments: 32 pages, 17 figures. The first two authors contributed equally. The code will be released soon Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG) ACMclasses: I.2.7; I.2.6 Cite as: arXiv:2610.03329 [cs.CL] (or arXiv:2610.03329v1 [cs.CL] for this version) https://doi.org/10.48550/arXiv.2610.03329 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[NLP-18] o Jev or Not? Evaluating the Accuracy and Efficiency of Structured Decision Models for Hate-Speech Moderation

【速读】: 该论文旨在解决在线内容规模庞大背景下仇恨言论(hate speech)识别与管理的挑战,尤其关注如何在无需针对特定任务进行微调的情况下,利用结构化决策模型实现高效、灵活且成本可控的分类。其核心解决方案在于评估六种结构化决策模型配置在四个仇恨言论数据集上的表现,对比专门训练模型、零样本学习、商业大语言模型(LLM)及监督模型等基准方法。研究关键发现为:尽管商业LLM在单一数据集上显著优于所有决策模型,但整体表现不稳定;提供或分解仇恨言论定义虽能改变部分预测结果(最多影响28%),却未带来一致性的性能提升;仅在20%的比较中,问题分解能显著改善分类效果。在诊断测试集上,最优托管决策模型的宏F1值仅比最佳商业LLM低1.6个百分点,同时推理成本降低约97%。这表明,通过合理设计的结构化决策模型可在极低计算成本下实现接近商用模型的性能,为低成本、可解释的仇恨言论自动化监管提供了可行路径。

链接: https://arxiv.org/abs/2610.03324
作者: Demetris Paschalides,George Pallis,Marios D. Dikaiakos
机构: 未知
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:The scale of online content makes hate-speech moderation challenging, while Large Language Models (LLMs) enable harmful material to be produced and adapted more easily. Moderation therefore requires efficient classifiers that can accommodate different definitions of hate speech. Recent structured decision models accept natural-language criteria and select among specified answers, raising the question of whether they can meet these requirements without task-specific training. We present HATEDECIDE, an evaluation of six decision-model configurations on four hate-speech datasets against specialized moderation, zero-shot, commercial, and supervised baselines. We examine whether supplying a dataset’s definition, or decomposing it into multiple questions, improves classification, and we measure their latency and cost. We find that commercial LLMs significantly outperform all decision models on only one dataset. Supplying definitions changes up to 28% of predictions without consistently improving classification, and decomposition significantly improves performance in only 20% of the comparisons. On a diagnostic set of test cases, the best hosted decision model comes within 1.6 macro-F1 points of the best commercial LLM at approximately 97% lower inference cost. These results identify opportunities for inexpensive moderation, while showing that explicit criteria and additional questions do not reliably improve classification.

[NLP-19] Shrome at Touché: Soft-Vote Ensembling and Counter-Causal Augmentation for Causality Extraction

【速读】: 该论文旨在解决新闻文本中反因果(counter-causal)陈述的因果关系抽取问题,即表面形式看似具有因果关联(如“人们错误地认为X导致Y”),但实际语义是否定因果关系的情况。传统依赖表面词汇线索(如“caused”或“led to”)的系统会误判此类句子为正向因果,从而产生错误极性判断。针对这一挑战,论文提出三个子任务:因果句检测、因果与结果跨度识别以及极性分类(正向因果、反向因果或无因果)。解决方案的关键在于:在检测任务中,采用微调分类器并结合单一跨任务规则,利用提取的因果-结果跨度信息消除假阳性;在抽取任务中,通过集成三个RoBERTa-large BILOU+CRF标注器,并对各模型的词级别得分进行平均后统一解码,而非基于跨度投票,以提升一致性与精度;在极性分类任务中,由于反因果样本稀缺,引入由大语言模型根据九种源自Hagen等人的反因果表达模式生成的训练样本,并通过自动结构校验筛选有效样本,增强模型对稀疏类别的学习能力。实验结果显示,在独立测试集上,检测任务达到F1 0.869,极性分类达到宏平均F1 0.817,且在组织方主导的仅考虑因果关系的抽取评估中,粒度调整后的F1达0.728,为所有提交方案中的最高分。

链接: https://arxiv.org/abs/2610.03268
作者: Roham Zendehdel Nobari,Shayan Sooratgar
机构: 未知
类目: Computation and Language (cs.CL)
备注: 16 pages, 5 figures, 10 tables. Both authors contributed equally. Working notes of Touché at CLEF 2026 (Conference and Labs of the Evaluation Forum), 21-24 September 2026, Jena, Germany

点击查看摘要

Abstract:Touché 2026 extends causality extraction to counter-causal claims: news sentences whose surface form appears causal but whose meaning denies the causation, as in “It is falsely believed that X caused Y.” A system that relies on surface cues such as “caused” or “led to” will accept such a sentence as causal and give it the wrong polarity. On the Countercausal News Corpus (CCNC), the task has three subtasks: deciding whether a sentence is causal (detection), locating its cause and effect spans (extraction), and labeling its polarity as procausal, counter-causal, or uncausal. We build one model per subtask. Detection is a fine-tuned classifier with a single cross-task rule that uses the extracted spans to remove false positives. For extraction, we ensemble three RoBERTa-large BILOU+CRF taggers by averaging their token-level scores before decoding, rather than voting on the spans each tagger produces. For polarity, where labeled counter-causal examples are scarcest, we add training sentences generated by a large language model prompted with nine patterns of counter-causal expression adapted from Hagen et al., keeping only those that pass automatic structural checks. On the held-out CCNC test set, the system reaches F1 0.869 on detection and macro-F1 0.817 on polarity, and in the organizers’ final causal-only evaluation of extraction it scores granularity-adjusted F1 0.728, the highest extraction score among all submissions including the organizers’ baseline. The development split is used only for component selection and the ablations reported in the paper.

[NLP-20] Collective Bias Mitigation via Model Routing and Collaboration

【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在公共健康、金融与治理等关键领域部署时,因训练数据中嵌入的偏见而导致的公平性问题。尽管自修正(self-debiasing)方法尝试让模型自主识别并纠正偏见,但仅依赖单一模型的内在知识难以有效消除深层次的刻板印象。为此,论文提出集体偏见缓解(Collective Bias Mitigation, CBM)框架,其核心在于通过学习细粒度的模型行为,并促进多样化大语言模型之间的知识共享,实现更有效的偏见缓解。该研究首次系统性地探索了不同大语言模型的有效选择与组织方式,以生成更具公平性的响应。实验表明,CBM显著优于独立运行的基线模型(如在Top-7设置下,委员会机制将年龄偏见得分从0.25降至0.10);其中,辩论(Debating)与委员会(Committee)两种拓扑结构均表现出显著的偏见降低效果,而后者在缓解效果与推理成本之间实现了良好平衡,凸显了CBM在提升大语言模型公平性方面的巨大潜力。

链接: https://arxiv.org/abs/2610.03240
作者: Mingzhe Du,Luu Anh Tuan,Xiaobao Wu,Yichong Huang,Yue Liu,Dong Huang,Huijun Liu,Bin Ji,Jie M. Zhang,See-Kiong Ng
机构: Nanyang Technological University(南洋理工大学); National University of Singapore(新加坡国立大学); Harbin Institute of Technology(哈尔滨工业大学); King’s College London(伦敦国王学院)
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Large language models (LLMs) are increasingly deployed in public health, finance, and governance, requiring both accuracy and societal value alignment. Despite recent advances, LLMs often perpetuate or amplify bias embedded in their training data, posing challenges to fairness. While self-debiasing encourages an LLM to identify and correct its own biases, relying on a single model’s intrinsic knowledge may be insufficient to address deeply ingrained stereotypes. To address this limitation, we introduce Collective Bias Mitigation (CBM), a framework that alleviates bias by learning fine-grained model behavior and fostering knowledge sharing among diverse LLMs. This work is the first to systematically explore the effective selection and organization of distinct LLMs to cultivate fairer LLM responses. Experiments show CBM substantially outperforms standalone baselines (e.g., in the top-7 setting, Committee lowers the age bias score from 0.25 to 0.10). Our Debating and Committee topologies achieve substantial bias reduction, with the latter balancing mitigation effectiveness and inference cost, highlighting the potential of CBM for fairer LLMs.

[NLP-21] AdaStep: Adaptive Step Credit Weighting for Agent ic Reinforcement Learning

【速读】: 该论文旨在解决长时程大语言模型智能体(Long-horizon LLM agents)在稀疏结果奖励下,轨迹级目标过于粗粒度、难以精确区分单个决策贡献的问题。其核心挑战在于:虽然步骤级信用分配(step-level credit assignment)可提供更细粒度的监督信号,但其估计可靠性受限于后续动作、环境状态转移及轨迹长度带来的不确定性。为此,论文提出AdaStep——一种自适应步骤信用加权方法,通过动态调整每一步局部优势(local advantage)对轨迹级信号的影响强度,实现更精准的信用分配。该方法将加权过程建模为潜在步骤优势的均方误差估计问题,并在显式条件采样假设下推导出最优的逐状态收缩系数。该系数具有信噪比-总方差的解释意义:当回报变化主要由当前动作导致时,保留局部信用;当变化主要由下游随机性主导时,则抑制局部信用。AdaStep仅需轻量级标量计算,无需价值函数(critic)、额外采样或模型推理,计算开销极低。在ALFWorld、WebShop和ScienceWorld三个基准任务上,采用三种不同模型架构的实验均表明,该方法在保持低计算成本的前提下,显著优于现有基线。

链接: https://arxiv.org/abs/2610.03223
作者: Xin Wang,Wenhao Wu,Menghao Zhang,Zhi Wang,Kun Shao,Jian Luan
机构: Tsinghua University (清华大学); Nanjing University (南京大学); Xiaomi Inc. (小米公司)
类目: Machine Learning (cs.LG); Computation and Language (cs.CL)
备注: 21 pages, 3 figures

点击查看摘要

Abstract:Long-horizon LLM agents are typically trained with sparse outcome rewards, making trajectory-level objectives too coarse to distinguish the contribution of individual decisions. Step-level credit assignment provides finer-grained supervision, but its estimates can be unreliable because observed returns also depend on subsequent actions, environment transitions, and trajectory length. We propose AdaStep, an Adaptive Step-credit weighting method that controls how strongly each group-derived local advantage modifies the trajectory-level signal. We formulate this weighting as a mean-squared-error estimation problem for the latent step advantage and, under an explicit conditional sampling assumption, derive an optimal per-state shrinkage coefficient. The coefficient admits a signal-to-total-variance interpretation: it preserves local credit when return variation is attributable to the selected action and suppresses it when variation is dominated by downstream randomness. AdaStep requires only lightweight scalar computation, with no critic, additional rollouts, or extra model inference. Experiments with three model backbones on ALFWorld, WebShop, and ScienceWorld show consistent improvements over baselines at low computational cost.

[NLP-22] StanceEval 2026: The Second Stance Detection Shared Task

【速读】: 该论文旨在解决阿拉伯语社交媒体文本中立场检测(stance detection)的跨目标泛化问题,尤其关注在主题相关与完全未见目标之间的迁移能力。其核心挑战在于:如何使模型在缺乏特定目标训练数据的情况下,仍能准确识别用户对新话题(如电动车型或三月期制度)的立场。解决方案的关键在于采用多样化的先进方法,包括微调预训练语言模型、基于提示(prompt-based)和检索增强的大语言模型(LLM)、微调后的大型语言模型以及混合级联架构。实验结果表明,顶级系统在Track 1(主题相关迁移)和Track 2(跨领域迁移)上分别取得了0.8994和0.9400的宏平均F1分数(F_avg2),显著超越最强基线(分别为0.7366和0.7475)。值得注意的是,模型在完全未见目标上的表现反而优于主题相关目标,这一反直觉现象可能源于目标极化程度、类别不平衡以及话题间方言或讽刺语用差异的影响。

链接: https://arxiv.org/abs/2610.03215
作者: Rasha Albalawi,Nuha Albadi,Hamzah Luqman,Asma Yamani,Maram Kurdi,Saad Ezzini,Ahmed Ashraf,Maged Al-Shaibani,Nora Alturayeif
机构: 未知
类目: Computation and Language (cs.CL)
备注: 14 pages total (8 pages main paper + 6 pages appendix), 5 tables in the main paper, excluding the appendix

点击查看摘要

Abstract:StanceEval 2026 is the second edition of the StanceEval shared task series on stance detection in Arabic social media text. Stance detection aims to identify a writer’s stance toward a given topic. Given a tweet and a target, participating systems must determine whether the writer’s stance is Favor, Against, or None. This edition focuses on cross-target generalization across two distinct evaluation tracks: Track 1 evaluates thematically related cross-target transfer (testing on Women Driving, related to Women Empowerment from training data), while Track 2 evaluates cross-domain transfer to completely unseen targets (E-Cars and Trimester System). The shared task attracted 80 registered teams from 12 countries. During the evaluation phase, 30 unique teams submitted entries, with 21 teams officially ranked in Track 1 and 13 in Track 2 following validation filtering, and 20 teams submitting system-description papers. Participating teams employed diverse methodologies, including fine-tuned pretrained language models, prompt-based and retrieval-augmented large language models (LLMs), fine-tuned LLMs, and hybrid cascades. Top systems achieved impressive F_avg2 scores of 0.8994 on Track 1 and 0.9400 on Track 2, substantially outperforming the strongest baselines (0.7366 and 0.7475, respectively), where F_avg2 denotes the macro-averaged F1 score over the Favor and Against classes. Counterintuitively, performance on the unseen targets was higher than on the related target, a disparity could be driven by extreme target polarization, class imbalance, and dialectal or sarcastic nuance across topics.

[NLP-23] Predicting and Repairing Merge Collapse in Large Language Models

【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在基于共享基础模型进行微调后,通过任务向量平均法合并时可能出现的性能坍塌问题。具体而言,现有合并算子在合并过程中无法提前预警破坏性合并,导致部分合并结果显著低于基线模型性能。其解决方案的关键在于提出一种新的预合并评估指标——基于专家模型任务向量方差的干扰度量(interference),该度量能够有效预测合并后的性能坍塌,并指导修复。研究发现,平均操作所移除的性能“功率”正比于任务向量间的方差,即干扰度量;在噪声模型假设下,合并引入的扰动随合并系数与干扰程度增加,从而构建出可预判合并质量的前评估分数。实验表明,仅当合并配置的该分数超过阈值时,才会发生破坏性坍塌,而传统以符号冲突为优化目标的合并方法反而具有反向预测性。在此基础上,作者提出PRISM算子:先对任务向量进行平均,再根据各层干扰水平自适应地施加软阈值,无需额外数据或调参即可将所有破坏性合并的性能保持在基线模型的评估噪声范围内,显著优于原始平均法(最低下降14.4分或完全崩溃)。此外,对于低于阈值的无害合并,仍采用原始平均法以保留效率。该方法实现了对合并行为的智能判断与动态修复。

链接: https://arxiv.org/abs/2610.03199
作者: Jungseob Lee,Seungyoon Lee,Sugyeong Eo,Hyeonseok Moon,Jaehyung Seo,Heuiseok Lim
机构: Korea University (高丽大学); Yonsei University Mirae Campus (延世大学未来校区); Sookmyung Women’s University (淑明女子大学); Konkuk University (中央大学)
类目: Machine Learning (cs.LG); Computation and Language (cs.CL)
备注: 23 pages, 5 figures, 20 tables

点击查看摘要

Abstract:Large language models fine-tuned from a shared base can be merged by averaging their task vectors, but some merges collapse far below the base model, and common merge operators give no warning before evaluation. We show that one statistic of the specialists’ task vectors both predicts this collapse and calibrates its repair. The power that averaging removes equals the variance of the task vectors across specialists, our measure of interference. Under a working noise model, the disturbance that a merge injects grows with the merge coefficient and with interference, yielding a pre-merge score. In our experiments on twenty-two merge configurations from four model families, only destructive merges exceed a threshold on this score. We find that statistics of sign conflict between specialists, a common target of existing merge operators, are anti-predictive. We then predicted the outcomes of fourteen merges before evaluating them, and twelve predictions were correct, including the destructive outcome of a specialist pair pushed past the threshold by continued pretraining. To address this collapse, we introduce PRISM, an operator that averages the task vectors first and then soft-thresholds each layer at a level set by the layer’s interference. Without data or tuning, PRISM keeps all five destructive merges above the threshold within evaluation noise of the base model, where plain averaging falls at least 14.4 points below it or collapses entirely. We apply PRISM only above the threshold and keep the plain average for merges below it, which include all fifteen harmless ones. Code is available at this https URL.

[NLP-24] KV2: A Self-Refining KV Cache

【速读】: 该论文旨在解决长上下文模型中键值(Key-Value, KV)缓存内存占用过大这一关键瓶颈问题,尤其在预填充上下文需服务于多个不同查询的可复用场景下,传统方法面临成本与精度之间的权衡:轻量级估计器虽计算开销低但准确性差,而全上下文重建评分虽精度高却需重复处理整个提示(prompt),导致资源浪费。其解决方案的关键在于提出一种名为KV^2的查询无关(query-agnostic)KV缓存压缩方法,该方法基于选择性重建机制——首先利用轻量级代理评分器筛选出上下文中有信息量的令牌(tokens),随后仅对这些精选子集进行重处理以计算最终的淘汰得分。实验结果表明,在RULER、Needle-in-a-Haystack和LongBench等多个基准测试中,随着压缩预算收紧,KV^2相较于基线方法的优势显著扩大:在RULER 16K任务中,当KV缓存预算仅为2%时,其平均得分较次优基线提升超过40个百分点;在LongBench上,于2%-10%的压缩预算范围内均实现了最高平均性能,且在压缩阶段的运行时间和峰值内存消耗均低于全上下文重建方法。研究表明,可复用的KV缓存压缩无需重处理完整上下文即可实现高效高精度的推理。

链接: https://arxiv.org/abs/2610.03198
作者: Johannes Wesch,Danni Liu,Jan Niehues
机构: Karlsruhe Institute of Technology (卡尔斯鲁厄理工学院)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:The memory footprint of the key-value (KV) cache constrains the practical use of long-context models, and it dominates cost when one prefilled context must later serve many different queries. In this reusable setting, query-agnostic compression trades cost against quality: lightweight estimators are cheap but less accurate, whereas full-context reconstruction scoring is more accurate yet reprocesses the entire prompt. We introduce KV ^2 , a query-agnostic KV-cache compression method based on selective reconstruction. KV ^2 first uses a lightweight proxy scorer to identify informative in-context tokens, then reprocesses only this subset to compute final eviction scores. On RULER, Needle-in-a-Haystack, and LongBench, KV ^2 's margin over baselines widens as the budget tightens: on RULER 16K at a 2% KV-cache budget it improves the average score over the next-best baseline by more than 40 percentage points, and on LongBench it attains the highest average across 2%-10% budgets at lower compression-stage runtime and peak memory than full-context reconstruction. Reusable KV-cache compression thus does not require reprocessing the full context. Our code is available at this https URL.

[NLP-25] Source Preference in the Wild: How LLM Agents Favor Items by Source and How to Reduce It

【速读】: 该论文旨在解决大语言模型(Large Language Model, LLM)代理在端到端搜索任务中对特定信息来源的偏好问题,这种偏好会显著影响用户所获得的结果及其选择的来源。研究发现,在三个不同领域中,12种代理模型均表现出对某些来源的系统性偏爱,且这种偏好在各模型间高度一致,甚至可能超越对任务需求满足程度的考量:当一个来自偏好源的项目虽未完全满足一项要求,但其被选中的概率仍高达约三分之二,而更优但来自非偏好源的项目则几乎不会被选中。进一步实验表明,仅标识来源的信息本身即可影响选择行为——隐藏来源可削弱偏好,而将项目重新标记为来自偏好源则能显著提升其被选概率。该研究揭示了源偏好形成的两个关键机制:一是训练过程中对优质项目的奖励可能使特定来源成为满足需求的“捷径”;二是信息缺失会触发对来源的先入为主判断。通过补充缺失信息或引入反向提示,可有效缓解此类偏见。因此,解决方案的关键在于识别并干预上述两种形成机制,以实现更公平、基于实际质量而非来源的决策。

链接: https://arxiv.org/abs/2610.03195
作者: Jonghyun Song,Haewon Park,Jeonghoon Shim,Woojung Song,Yohan Jo
机构: Seoul National University (首尔国立大学)
类目: Computation and Language (cs.CL)
备注: 41 pages

点击查看摘要

Abstract:As LLM agents decide on users’ behalf which product to buy, which hotel to book, or which paper to cite, a preference for items from certain sources (the sites or services they come from) shapes what users receive and which sources are selected. We study source preference in end-to-end search with 12 agent models across three domains. Comparing items from different sources that satisfy the same requirements at the same position, we find that each model prefers some sources and avoids others in every domain, largely agreeing on which. This preference can outweigh how well items satisfy the request: an item satisfying one requirement fewer is selected about two-thirds of the time when it comes from a preferred source and the better one from a dispreferred source, but almost never in the reverse case. The information identifying an item’s source affects selection by itself: hiding it weakens the preference, and relabeling an item with a preferred source raises its selection rate. We test two routes to this preference: training that rewards better items can make a source a shortcut for requirement satisfaction, and missing information can trigger preconceptions about the source. Supplying missing information or a prompt countering these preconceptions reduces source preference.

[NLP-26] Not Until the Evidence Says So: Teaching LLM Investigators When to Close a Case

【速读】: 该论文旨在解决大语言模型(Large Language Model, LLM)在事故、缺陷及故障调查中面临的“案件关闭决策”问题,即判断现有证据是否足以得出结论。传统问答系统无需面对此类判断,而本案调查需权衡证据充分性,决定是否闭案或指出证据缺失。其核心挑战在于:当前模型普遍存在过度自信现象,即使未掌握充分证据也倾向于闭案,导致高过度假设率(如9B模型高达97%),且在官方判定为“原因不明”的案例中仍错误闭案。解决方案的关键在于构建一个名为Nautil的高质量数据集,涵盖航空、铁路、海事、化工安全及车辆缺陷等领域的731个经审计案例,并引入教师轨迹(teacher trajectories)、分布外测试集与反事实证据版本,以系统评估模型的闭案行为。通过在教师轨迹上微调9B模型,显著提升其闭案行为与证据的依赖性,使过度假设率从97%降至35%,正确非过度假设结论比例从3%提升至43%;进一步采用仅奖励闭案决策的强化学习方法,使平衡准确率从69.2提升至83.3,接近教师水平,同时在源内准确率方面达到74.1,表明模型能更精准地识别证据依赖性与关键缺口。

链接: https://arxiv.org/abs/2610.03190
作者: Tingzhu Bi,Ping Wang,Meng Ma
机构: Peking University (北京大学)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: 23 pages. Dataset: this https URL ; Models: this https URL , this https URL ; Demo: this https URL ; Code: this https URL

点击查看摘要

Abstract:Accident, defect and outage investigations end with a decision that ordinary question answering never faces: whether the evidence gathered so far is enough to close the case. We study this decision for LLM investigators, which request evidence from a case file, revise their hypotheses, and either close the case with a conclusion grounded in what they read or leave it open and name what is missing. This judgment does not come with capability: an untrained 9B model overstates its evidence in 97% of its answers, and a frontier model that identifies the right cause in 84% of cases still overstates in 91% and closes 17 of the 41 cases whose official finding is “cause undetermined”. Measuring it is also non-trivial: the source of a case largely predicts its label, and a rule that reads only the source reaches 83.0 balanced accuracy on our test cases. We therefore evaluate closure with three tests: closure accuracy, reported against this rule and within each source; evidence dependence, which removes the grounds of a conclusion and checks whether the model stops closing; and conclusion and gap quality, a judged checklist of what the model asserts and what it says is missing. We build Nautil, 731 audited cases from aviation, rail, maritime, chemical-safety and vehicle-defect reports and production server incidents, with teacher trajectories, an out-of-distribution test set and counterfactual evidence versions. Fine-tuning a 9B model on these trajectories makes its closures follow the evidence: removing the grounds lowers its closure rate by 26 points relative to a matched control, overstatement falls from 97% to 35%, and correct, non-overstated conclusions rise from 3% to 43%. Reinforcement learning that rewards only the closure decision then raises balanced accuracy from 69.2 to 83.3, on par with the teacher, and within-source accuracy from 60.4 to 74.1, at some cost in evidence dependence.

[NLP-27] Gains and Collapse in On-Policy Distillation:A Reinforcement Learning Perspective

【速读】: 该论文旨在解决生成式语言模型在基于策略蒸馏(On-policy Distillation, OPD)的后训练过程中出现的生成质量退化问题,即模型倾向于产生过长且重复的文本输出。其核心问题是:尽管OPD能够提升模型性能,但其背后的机制尚不明确,尤其在某些情况下会导致生成行为严重偏离预期。解决方案的关键在于从强化学习视角揭示了教师模型对学生的隐式奖励机制——即使教师自身极少生成此类冗余内容,仍会通过隐式反馈“奖励”学生产生的类似行为。研究发现,OPD并非扩展学生模型的能力边界,而是通过增强教师对优质响应的采样概率来改善表现;当教师的偏好与真实质量不一致时,则引发“奖励劫持”(reward hacking),放大冗余生成。基于此诊断,提出两种有效缓解策略:在训练中屏蔽不良响应,以及采用监督微调(SFT)初始化。二者共同表明,OPD的本质是放大教师对学生产出的隐式评估偏好,因此关键在于提升教师评估的可靠性,而非单纯优化生成能力。

链接: https://arxiv.org/abs/2610.03185
作者: Han Cui,Jianhao Yan,Yun Luo,Hongbo Zhang,Zhizhang Fu,Yue Zhang
机构: Zhejiang University (浙江大学); Westlake University (西湖大学)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:On-policy distillation (OPD) has become an important approach to language model post-training. However, despite its performance gains, OPD can also collapse into excessively long and repetitive generation, and the mechanism underlying these divergent outcomes remains poorly understood. We explain these outcomes through a reinforcement learning perspective: the teacher implicitly rewards student behaviors, even those it rarely exhibits itself. From this perspective, our experiments show that OPD improves performance without expanding the student’s capabilities. When the implicit reward model is reliable, OPD makes correct responses easier to sample. In contrast, when the preference misaligns with quality, reward hacking happens: the implicit reward model amplifies overlong, repetitive student rollouts, even though it rarely generates such text itself. Guided by this diagnosis, we find that masking unhealthy responses during training and using SFT initialization can each effectively mitigate the collapse. Together, these findings show that OPD amplifies student behaviors favored by the teacher’s implicit feedback, shifting the focus from how well the teacher generates to how reliably it evaluates student rollouts. Our code is available at this https URL.

[NLP-28] Hindsight-Guided Rationale Distillation for Rare Disease Diagnosis AACL

【速读】: 该论文旨在解决罕见病诊断中生成式 AI 模型因数据稀疏性导致的性能瓶颈问题,特别是在小样本、高复杂度的医学诊断任务中,如何有效提升模型对罕见疾病的识别准确率。其核心挑战在于:尽管采用大规模教师模型(8B参数)进行链式思维(Chain-of-Thought, CoT)生成,但学生模型(1.5B参数)在微调后仍难以显著超越教师模型,且绝对准确率整体偏低。解决方案的关键在于引入事后指导蒸馏(hindsight-guided distillation)与污染过滤机制的结合,其中最关键的创新是通过正则表达式(regex)过滤器识别并移除教师模型在推理过程中因可见真实标签而产生的“地面真值幻觉”(GT hallucination)——即模型在推理链中嵌入“ground truth is X”类语句的不必要偏见。实验证明,未经过滤的学生模型未能显著优于教师模型(p = 0.129),而经过污染过滤后的学生模型(StudentF)在代表性较好的疾病上实现了统计显著的性能提升(p < 0.001),表明真正驱动性能增益的是污染过滤而非单纯的事后蒸馏。该研究还揭示了频率依赖的知识迁移现象以及监督微调(SFT)无法弥合的校准差距,为未来改进罕见病诊断模型提供了关键方向。

链接: https://arxiv.org/abs/2610.03176
作者: Aarav Singh,Animesh Pathak,Navyansh Singh
机构: Dr. Shyama Prasad Mukherjee International Institute of Information Technology, Naya Raipur(德·夏玛·普拉萨德·穆克吉国际信息技术学院,新赖浦尔); IIIT Naya Raipur(印度信息科技学院新赖浦尔)
类目: Computation and Language (cs.CL)
备注: 15 pages, 4 figures, Github: this https URL , Accepted at AACL-IJCNLP SRW 2026

点击查看摘要

Abstract:We study hindsight-guided distillation for rare disease diagnosis on ZebraMap: a 1.5B student is fine-tuned on chain-of-thought traces from a 8B teacher that observes the ground-truth diagnosis during generation. Absolute accuracy remains low for all models - the task is hard at this scale - but within this ceiling a filtered variant (StudentF) achieves a small, statistically significant accuracy advantage over the teacher (p 0.001), concentrated in better-represented diseases. The unfiltered student does not significantly outperform the teacher (p = 0.129), establishing that contamination filtering - not hindsight distillation alone - drives the gain. The gap traces to an artifact we term GT hallucination. Label-visible generation causes the teacher to embed “ground truth is X” phrases in its reasoning chain; SFT copies the pattern. At inference, the unfiltered student reproduces the phrase in 33.9% of cases, with severe accuracy degradation when the hallucinated label is wrong. A regex filter removing these slots reduces contamination to near-zero, producing the observed gain - though the effect remains small. We precisely quantify this gain-cost tradeoff, document frequency-dependent knowledge transfer absent from the RL-trained teacher, and characterize a calibration gap that SFT does not close - identifying both as directions for future work.

[NLP-29] Predicting Steering Vectors and Adapter Weights for Few-Shot Author-Style Transfer EMNLP2026

【速读】: 该论文旨在解决从少量示例中为特定作者适配大语言模型写作风格的问题,尤其针对科学写作场景下的挑战:正式的文体规范导致表面风格差异有限,且作者通常围绕自身研究主题写作,使得提取的“风格”特征极易与内容信息混淆。为此,论文提出三种基于风格条件生成摘要的方法:(1)对比激活引导(contrastive activation steering),(2)预测引导向量的神经网络,以及(3)生成LoRA适配器的超网络(hypernetwork)。其核心解决方案在于在作者层面进行风格引导——通过将同一内容下某作者的摘要与其风格中性化生成结果进行对比,从而在固定主题的前提下分离风格信号,避免依赖预定义风格词典。该方法不仅无需预先设定风格集合,还显著优于基于词汇库存的引导方式。研究发现,在风格模仿度与输出质量之间存在稳定权衡:微调虽能捕获最强风格信号但牺牲流畅性,而超网络方法在已见和未见作者上均实现了最佳平衡。此外,分析表明人工提取与模型预测的引导向量虽近似正交,但性能相当,说明该任务中的风格条件可由至少两个独立方向实现,而非依赖单一特定轴。

链接: https://arxiv.org/abs/2610.03163
作者: Leonard Popp,Danni Liu,Supriti Sinhamahapatra,Jan Niehues
机构: Karlsruhe Institute of Technology (卡尔斯鲁厄理工学院)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: W-NUT Workshop @ EMNLP 2026

点击查看摘要

Abstract:Adapting large language models to an individual author’s style from a few examples is challenging, and scientific writing sharpens the difficulty: formal conventions leave little surface variation, and authors write about their own topics, so extracted ``style’’ easily entangles with content. We study style-conditioned abstract generation from a few example abstracts per author and propose three methods: (1) contrastive activation steering, (2) a network that predicts steering vectors, and (3) a hypernetwork that predicts LoRA adapters. We find a consistent trade-off between style imitation and output quality: fine-tuning buys most of the available style signal but forfeits fluency, while the hypernetwork achieves the best trade-off on both seen and unseen authors. Our steering operates at author level, contrasting an author’s abstracts against style-neutral generations for the same content. This holds topic fixed, removes the need for a predefined style inventory, and outperforms inventory-based steering. % [EDIT 1a] softened “no single optimal axis” claim Moreover, our analyses demonstrate that manually extracted and predicted steering vectors are near-orthogonal yet score comparably, indicating that style conditioning here can admit at least two unrelated directions rather than requiring one particular axis.

[NLP-30] Investigating the Role of Reasoning -Language Alignment in Monolingual Retrieval-Augmented Generation EMNLP2026

【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在多语言场景下进行推理时的语言对齐问题,特别是当模型需基于检索增强生成(Retrieval-Augmented Generation, RAG)机制处理非母语信息时的表现瓶颈。现有研究表明,强制模型在非母语环境下进行推理会降低准确性,但这一现象主要集中在短提示(short prompt)场景;而本文扩展至更复杂的长文本检索情境,探究语言一致性对性能的影响。其解决方案的关键在于构建一个完全德语化的RAG问答测试平台——基于桌游《黑暗之眼》(The Dark Eye)这一德语资料丰富的虚构世界,确保模型必须依赖外部检索而非记忆作答,从而真实模拟多语言信息整合过程。实验结果表明,当推理语言与查询及检索文档的语言一致(即德语)时,模型表现显著优于其他语言(如法语),且优势随检索上下文的丰富度与结构化程度提升而增强,证明语言对齐的重要性超越语言熟练度本身。然而,即使在最佳对齐条件下,强制德语推理仍未能超越模型原生英文推理能力,揭示了实现真正多语言推理仍需发展具备原生多语言能力的模型架构。研究团队已公开发布该测试平台与基准数据集以促进后续研究。

链接: https://arxiv.org/abs/2610.03136
作者: Oliver Hauck,Mario Sanz-Guerrero,Katharina von der Wense
机构: Johannes Gutenberg University Mainz (美因茨约翰内斯古腾堡大学); University of Colorado Boulder (科罗拉多大学博尔德分校)
类目: Computation and Language (cs.CL)
备注: Accepted to the Workshop on Open Reasoning Across Cultures Languages at EMNLP 2026

点击查看摘要

Abstract:Reasoning traces improve large language models (LLMs), but current models are trained to reason mostly in English. It has been shown that forcing a model to reason in another language degrades accuracy, even when the reasoning language matches the language of the prompt – but only for a setting where the model reasons over a short prompt. Here, we ask whether the same holds for retrieval-augmented generation (RAG), where the model must read and integrate a large amount of retrieved evidence in the target language. To study this, we build a fully monolingual German RAG question-answering testbed over the fictional world of the tabletop role-playing game The Dark Eye, a domain that is richly documented in German but too niche for the model to answer from memory, so that it has to rely on retrieval. Varying the forced reasoning language of an agentic RAG system on this testbed, we find that aligning the reasoning language with the language of the query and the retrieved documents helps. Forced German reasoning outperforms forced French, although the model benchmarks higher in French, so the benefit comes from alignment and not from language proficiency. The advantage grows when the retrieved context is richer and structure-aware. However, forced German only reaches the level of the model’s native, unconstrained English reasoning without surpassing it, showing that native multilingual reasoning is needed. We publicly release the testbed and QA benchmark.

[NLP-31] he Frag ility of Trigger-Tag Mechanisms for Misuse Detection in Open-Weight LLM s

【速读】: 该论文旨在解决开放权重语言模型(open-weight language models)在缺乏中心化管控的情况下,难以有效防范滥用行为的问题,尤其关注如何检测模型在特定恶意条件(如生成钓鱼内容)下的不当使用。现有研究提出的触发标签(trigger-tag)机制通过在模型输出或内部参数中嵌入可检测信号来实现条件性滥用检测,但其在对抗攻击下的鲁棒性尚未得到系统评估。本文的关键贡献在于:首先,形式化区分了两类触发标签——基于词元级(token-level trigger-tags)的机制,其借鉴水印思想,在解码过程中引入可检测信号;以及基于权重级(weight-level trigger-tags)的机制,其通过学习目标条件与可检测行为之间的后门关联实现检测。其次,提出统一攻击框架Untag,将不同机制的攻击面归纳为通用分类体系,并以钓鱼内容生成为例进行实证评估。结果表明,尽管触发标签在受控环境下可能提供有用证据,但在实际对抗攻击下,现有机制均被完全绕过,无法维持有效性。因此,论文强调:当攻击者能够修改输出或直接操作开放权重时,触发标签机制不应被视为可靠的滥用检测手段。

链接: https://arxiv.org/abs/2610.03124
作者: Toluwani Aremu,Manit Baser,Mohan Gurusamy,Nils Lukas,Dinil Mon Divakaran
机构: National University of Singapore(新加坡国立大学); MBZUAI(阿布扎比人工智能大学); A*STAR Institute of Advanced Intelligence and Computing(新加坡科技研究局先进智能与计算研究所)
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computers and Society (cs.CY); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Open-weight language models can be downloaded, modified, and deployed beyond their developers’ control, limiting the effectiveness of centrally enforced safeguards. Recent work has therefore proposed \emphtrigger-tag mechanisms that produce a detectable signal when a model is used under a target condition, such as generating phishing contents. Although these mechanisms borrow from established techniques, their use for conditional misuse detection in open-weight LLMs is relatively new. Therefore, existing research works have not systematically studied the robustness of trigger-tag mechanisms under adversarial attacks. To close this gap, (i)~we formalize trigger-tags and distinguish \emphtoken-level trigger-tags, which introduce watermark-inspired signals during decoding, from \emphweight-level trigger-tags, which learn backdoor-inspired associations between target conditions and detectable model behavior. Furthermore, (ii)~we introduce \Untag, a unified attack framework that organizes their mechanism-specific attack surfaces into a common taxonomy. We evaluate representative token-level and weight-level trigger-tags using phishing as a case study. We find that while trigger-tags may provide useful evidence in controlled settings, our attacks render the existing trigger-tag mechanisms to be entirely ineffective. Consequently, we argue that these mechanisms should not be treated as robust misuse detectors when attackers can transform outputs or modify open weights.

[NLP-32] Building Interpretable Feature Representations for Resume-Vacancy Matching by Distilling Production LLM Signals EMNLP2026

【速读】: 该论文旨在解决招聘过程中候选人与职位匹配的可解释性问题,即传统匹配系统仅提供单一、不透明的相关性评分,而忽视了招聘人员需要理解候选人为何适配的具体依据。其核心解决方案是提出一种两阶段架构:第一阶段采用基于大语言模型(LLM)的标注器,通过持续迭代优化提示词与特征定义,从招聘人员反馈中学习并生成可解释的匹配维度标签;第二阶段则构建一个轻量化的特征双向编码器(feature bi-encoder),该编码器由前述标注器蒸馏而来,采用LoRA微调的嵌入主干网络及紧凑的逐维头结构,可在CPU上高效运行并支持所有在线请求。该系统通过持续收集招聘人员对预测特征的确认或修正反馈,实现模型与标注策略的动态优化。在168,772个标注的职位-简历对上训练后,在927条生产环境反馈数据中,部署模型与人工记录决策的一致率达到95.79%,体现了其在实际应用中的高可靠性与可操作性。

链接: https://arxiv.org/abs/2610.03112
作者: Ilya Chekin(1),Vyacheslav Malyugin(1),Vladimir Chirkov(1),Mikhail Yurushkin(2) ((1) BroutonLab, (2) Curately)
机构: BroutonLab(布罗顿实验室); Curately(卡拉泰)
类目: Computation and Language (cs.CL)
备注: Accepted to EMNLP 2026; 13 pages, 4 figures, 7 tables

点击查看摘要

Abstract:Matching candidates to vacancies is central to recruitment, and a recruiter needs to see why a candidate fits, not only a single opaque relevance score. We provide this evidence as named, interpretable matching dimensions recruiters can act on - eight in our current deployment. We propose a two-part approach. The first is an LLM-based labeler whose prompts and feature definitions were refined from recruiter feedback while it served as an earlier production matching stage. In the current architecture, it is used only for offline labeling and is not called on online requests. The second is a feature bi-encoder distilled from it: a LoRA-adapted embedding backbone with compact per-dimension heads that runs on CPU and serves all online requests. Both parts keep improving: prompts are revised as feedback arrives, and the bi-encoder is retrained on the updated labels. The model is trained on 168,772 labeled vacancy-resume pairs (17,921 vacancies and 180,030 resumes). Recruiters using the service can confirm or revise surfaced feature predictions. On 927 recruiter-recorded values from this selected production-feedback subset, the deployed student agrees with the recorded decisions in 888 cases (95.79%). This is operational, non-blinded agreement rather than an independent human evaluation.

[NLP-33] Ontological Instability and Statistical Amplification: The Paradox of “Humanizing” LLM -Generated Text

【速读】: 该论文旨在解决当前监督式生成文本检测器(AI-text detector)在高基准准确率下决策依据不明确的问题,尤其关注其对文本语义、结构及分词层面扰动的敏感性。研究通过在M4数据集(N = 10,000)和受控生成样本(N = 300)上对基于RoBERTa的检测器进行系统性分析,发现当使用Mistral-7B-Instruct生成更接近人类写作风格的文本时,动词多样性(Verb Diversity)从0.77提升至0.92,反而使文本更容易被检测,表明检测器的判断可能依赖于统计复杂性等可量化的表面特征而非深层语义真实性。此外,该检测器在正式人类写作中产生了高达76.3%的误报率,凸显其泛化能力不足。作为对照,研究还评估了基于事件的潜在空间检测方法(event-based Latent Space detection),结果显示重述(paraphrasing)导致87%的事件序列变化(Jaccard相似度仅0.067),同音异形字符(homoglyphs)亦改变70%提取动词(Jaccard = 0.30),但其最佳领域AUC仅为0.577,表明该方法同样缺乏鲁棒性。因此,该研究的关键发现是:现有检测器的鲁棒性高度依赖于其所利用的具体特征,而单纯的结构抽象无法有效提升检测抗干扰能力,揭示出当前检测模型在面对语义改写与形式变换时的脆弱性。

链接: https://arxiv.org/abs/2610.03110
作者: Claudiu Creanga,Liviu Dinu
机构: University of Bucharest (布加勒斯特大学); University of Bucharest (布加勒斯特大学)
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Supervised AI-text detectors report high benchmark accuracy, but it is not clear what their decisions are based on. We analyze a RoBERTa-based detector under semantic, structural, and tokenizer-level perturbations, using the M4 dataset (N = 10,000) and controlled generations (N = 300). When Mistral-7B-Instruct was asked to make machine text sound more human, Verb Diversity rose from 0.77 to 0.92 and the outputs became easier to detect. Detection scores appear to track statistical complexity, which also leads to a 76.3% false-positive rate on formal human writing. As a control, we evaluate event-based Latent Space detection. Paraphrasing changed 87% of its event sequences (Jaccard = 0.067), and homoglyphs altered 70% of the extracted verbs even though extraction still ran (Jaccard = 0.30). Its best domain AUC was 0.577. RoBERTa’s robustness seems specific to the features it uses, and structural abstraction did not make detection more robust.

[NLP-34] Emergent Structure in the Marginal Attention Space of Language Models

【速读】: 该论文旨在解决独立训练的语言模型(LLM)在内部机制(如注意力行为)上缺乏系统性比较的问题,尤其关注后softmax注意力权重的结构特性。其核心解决方案是通过将查询位置进行边缘化(marginalization),将注意力权重映射至一个联合的“词元-注意力头”边际注意力空间(marginal attention space)。研究发现,沿词元轴降维时,边际注意力呈现出对文本固有的、跨模型高度一致的信号,该现象可通过将边际注意力与网络的输入-输出雅可比矩阵(input-output Jacobian)建立经验关联来解释;理论上,在平滑性假设下,具有相似下一个词分布的模型必然拥有相似的输入-输出雅可比统计特性。而沿注意力头轴降维时,则形成模型特有的签名,且在不同文档间保持稳定。这一发现为基于注意力头的键值(KV)缓存淘汰策略提供了新的思路:可预先在预训练文本上离线计算每个注意力头的预算,结合无需训练的词元评分,实现模型特定预算分配与文本内在词元评分的解耦。在标准淘汰基准测试中,该方法表现媲美需每文档重算或针对目标微调预算的方法,具备良好的实用性与效率。

链接: https://arxiv.org/abs/2610.03109
作者: Valentino Maiorca,Walter Nelson,Francesco Locatello
机构: Institute of Science and Technology Austria (奥地利科学技术研究所)
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:While representation similarity across independently trained language models is well-documented, how internal mechanics such as attention behave across models remains far less characterized. Inspired by this gap, we examine the structure of post-softmax attention weights by marginalizing over query positions, mapping them into a joint token-head “marginal attention space”. Evaluating across 60+ diverse LLMs, we find that different properties emerge when reducing this space along its token and head axes. When reduced token-wise, marginal attention yields a text-intrinsic signal robustly conserved across models. To explain this property, we empirically connect marginal attention to the input-output Jacobian of the network, and prove theoretically that under a smoothness assumption, models with similar next-token distributions are guaranteed to have similar input-output Jacobian statistics. When reduced head-wise, it forms a model-private signature conserved across documents. Practically, this provides a natural way to estimate a per-head budget for key-value (KV) cache eviction, effectively decoupling model-specific budget allocation from text-intrinsic token scoring. On standard eviction benchmarks, a per-head budget precomputed offline on pretraining text, combined with a training-free token score, shows competitive performance with methods that recompute the budget on every document or train it per target. Code available at this https URL

[NLP-35] Ask Relax or Act? Evaluating Actionable Indeterminacy in LLM Preference Reasoning

【速读】: 该论文旨在解决大语言模型(Large Language Model, LLM)代理在面对不确定情境时,虽能识别不确定性却仍做出错误的下一步决策的问题,具体表现为在已有合理行动方案时仍请求澄清,或在约束条件必须调整时仍试图维持现状。其核心挑战在于“可行动的不确定性”(actionable indeterminacy)——即如何在不同可接受偏好或目标下,准确判断是否应采取行动、请求澄清或进行约束修复。解决方案的关键在于提出一种基于最小成本允许约束修复的决策框架:当某一行动在所有可接受目标中均被共享时,应直接执行该行动;当每种可能性均可行但无共同行动时,应请求澄清;当请求不可行时,则实施最小成本的约束修复。研究构建了一个基于求解器的基准测试,涵盖物品分配、会议调度、公寓选择和稳定匹配等场景,通过保留相同问题源但改变干预必要性来评估模型表现,并区分决策正确性、配对可靠性与完全正确响应三个维度。结果表明,模型普遍存在“不必要的干预”问题:即便存在已验证的合理行动,仍会主动发起澄清或修复请求;同时,仅标注正确决策并不能保证生成的行动、问题或修复是可用的。关键发现在于,响应需求的显式表达不仅影响决策的表达方式,更直接影响决策本身,显式定义输出内容可显著提升完整正确响应率,并可能改变干预决策,即使输出格式已可解析。这揭示了可靠智能体的核心要求不仅是识别不确定性,更在于仅在必要时干预,并将所选下一步转化为可验证的响应。

链接: https://arxiv.org/abs/2610.03102
作者: Ang Li,Yue Lin,Feifei Kou,Zhan Su,Prayag Tiwari,Wenhao Li,Shuhui Zhu,Hongyuan Zha,Baoxiang Wang
机构: The Chinese University of Hong Kong Shenzhen(香港中文大学深圳校区); Beijing University of Posts and Telecommunications(北京邮电大学); Halmstad University College(哈尔姆斯塔德大学学院); Tongji University(同济大学); University of Waterloo(Waterloo大学)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: 55 pages, 5 figures

点击查看摘要

Abstract:An LLM agent can recognize uncertainty yet still choose the wrong next step: asking when action is already justified, or seeking clarification when the constraints must change. We formalize actionable indeterminacy: act when an accepted action is shared across all admissible preferences or objectives, clarify when each possibility is feasible but no action is shared, and propose a minimum-cost permitted constraint repair when the request is infeasible. We construct a solver-grounded benchmark spanning object allocation, meeting scheduling, apartment choice, and stable matching. Matched pairs retain the same source while changing whether intervention is necessary, and evaluation separates decision correctness, matched-pair reliability, and fully correct responses. Our findings reveal a recurring difficulty in recognizing when intervention is unnecessary: models can identify situations requiring clarification or repair yet still intervene when a justified action already exists. Correct decision labels also fail to guarantee usable actions, questions, or repairs. Crucially, response requirements shape not only how decisions are expressed but also which decisions are made. Making the required content explicit substantially improves fully correct responses and can change intervention decisions, even when outputs are already parseable. These findings highlight that reliable agency requires more than recognizing uncertainty: it requires intervening only when necessary and translating the chosen next step into a verifiable response.

[NLP-36] Peer Influence across Heterogeneous AI Models

【速读】: 该论文旨在解决多智能体系统中不同人工智能(AI)代理在意见分歧时的相互说服机制问题,即当两个基于不同架构和规模的语言模型产生分歧时,谁能够影响谁的判断。其核心挑战在于:传统上认为模型规模或独立决策置信度可预测说服力,但研究发现这一假设不成立。解决方案的关键在于揭示说服行为并非由单一模型属性决定,而是高度依赖于具体模型对之间的互动关系。研究通过测量单次交互后接收方决策概率分布的变化来量化说服程度,结果表明,尽管模型间存在显著差异,但小模型在某些组合中可有效说服甚至颠覆大模型的判断,且说服效果主要取决于接收方的敏感性而非发送方的“说服力”。因此,说服模式具有高度情境依赖性,强调必须在实际协作组合中评估模型行为,而非仅依据其孤立性能推断交互结果。

链接: https://arxiv.org/abs/2610.03095
作者: Frida Nøhr Laustsen,Marie Haahr Petersen,Victoria Popa,Ariel Flint,Romualdo Pastor-Satorras,Andrea Baronchelli,Luca Maria Aiello
机构: IT University of Copenhagen(哥本哈根信息技术大学); University of Pisa(比萨大学); City St. George’s, University of London(伦敦城市圣乔治大学); Universitat Politècnica de Catalunya(加泰罗尼亚理工大学)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computers and Society (cs.CY); Physics and Society (physics.soc-ph)
备注: 30 pages, 16 Figures, 6 Tables

点击查看摘要

Abstract:When two AI agents disagree, who persuades whom? As multi-agent systems increasingly combine language models of different families and sizes, the answer can determine which judgments survive interaction. Measuring persuasion as the probabilistic shift in an agent’s decision after a single exchange with a dissenting peer, we test seven open-weight models across three language understanding tasks. We find that persuasion is strong: when models disagree, receivers often abandon their initial judgment after seeing a peer’s answer and explanation. Surprisingly, however, neither standalone certainty nor model scale reliably predicts persuasion dynamics. Models producing almost perfectly consistent decisions in isolation can be among the most susceptible to persuasion, and small models can match larger ones as persuaders and resist their influence just as effectively. Furthermore, we show that the size of the shift depends more on the susceptibility of the listener than on the persuasiveness of the speaker. Persuasion patterns are therefore specific to each model pairing, with heterogeneity amplifying persuasion in some combinations and suppressing it in others, allowing a dissenting agent running a small model to overturn the judgments of a much larger one. These findings show that the behavior of interacting models cannot be inferred from their individual properties but must be evaluated in the combinations in which they will operate.

[NLP-37] MintEval: Do LLM s Implement the Trading Strategy You Asked For? A Behavioural-Equivalence Benchmark for Natural-Language-to-Strategy Code

【速读】: 该论文旨在解决生成式 AI 在量化交易领域中从生成交易信号转向自动生成执行代码时所面临的“无声失效”(silent failure)问题。其核心挑战在于:现有基准测试无法有效评估模型生成的代码是否真正实现了交易员所描述的风险逻辑,因为传统方法仅关注代码功能正确性(如单元测试)或预测能力(如金融基准),而未衡量实际执行行为与策略意图的一致性。为此,作者提出 MintEval 基准,通过从可组合的构建块库中程序化生成参考策略,将其回译为自然语言交易指令,并由待测模型重新实现;随后在相同市场数据和摩擦条件下逐根行情条(bar-by-bar)执行生成代码与参考代码,并基于动作一致性(ActionMatch)进行比较,而非依赖代码相似性或收益表现,从而消除阿尔法(alpha)带来的偏差。MintEval v0 包含 800 个任务,基于执行测量的状态跨度复杂度 τ(tau)进行分层,该指标独立于描述长度。实验表明,低成本模型平均动作匹配率仅为 0.544,仅能精确复现 8.7% 的任务;即使前沿模型(Claude Opus 5.5)在 200 个分层任务上达到 0.889 的匹配率并精确复现 57.5%,仍存在 27.5% 的任务出现无声错误。进一步分析发现,尽管模型对策略意图的理解近乎完美,但高达 79.2% 的正确解析实现仍会在超过 10% 的活跃行情条上发生行为偏离。值得注意的是,近期策略生成基准中使用的 LLM 判定器在原样应用时,完全忽略了这些隐性错误,凸显了当前评估体系的重大缺陷。因此,解决方案的关键在于建立以行为一致性为核心的、基于真实执行对比的评估范式,从而揭示并防范生成代码与预期策略之间的隐蔽偏差。

链接: https://arxiv.org/abs/2610.03080
作者: Siyu Wang,Yifan Wang,Yuecheng He
机构: Fudan University (复旦大学)
类目: oftware Engineering (cs.SE); Computation and Language (cs.CL); Machine Learning (cs.LG); Trading and Market Microstructure (q-fin.TR)
备注: 5 pages, 3 figures, benchmark code and evaluation harness available at this https URL . Siyu Wang and Varstern Yifan Wang contributed equally, Yifig Wang is corresponding author

点击查看摘要

Abstract:Large language models are moving from producing trading signals to writing the code that executes them. The failure mode of the second role is silent: generated code runs, a backtest plots, yet the risk logic that the trader described is not the logic being executed. Existing code benchmarks test functional correctness on unit tests and finance benchmarks test forecasting; neither measures whether an implementation behaves like the strategy that was asked for. We introduce MintEval, a benchmark in which reference strategies are generated programmatically from a library of composable building blocks, back-translated into colloquial trader instructions, and re-implemented by the model under test. Generated and reference programs are executed bar by bar on identical market data and frictions, and compared on their actions rather than on code similarity or profit: alpha is differenced away. MintEval v0 contains 800 tasks on BTCUSDT 15-minute data, stratified by an execution-measured state-span complexity tau that is decoupled from description length. Low-cost models reach a mean ActionMatch of at most 0.544 and reproduce at most 0.087 of tasks exactly; on a stratified subset of 200 tasks a frontier model (Claude Opus 5.5) reaches 0.889 and reproduces 0.575 exactly, yet still fails silently on 0.275 of tasks. Given a menu of building blocks, models identify the strategy almost perfectly, yet 79.2% of the implementations whose specification was read correctly diverge on more than 10% of active bars. The LLM judge of a recent strategy-generation benchmark, applied verbatim, accepts every one of these silent failures.

[NLP-38] An automated pipeline for standardised speech-unit annotation in spontaneous dialogue

【速读】: 该论文旨在解决自发性双人对话中交互单元(如对话回合与听者反馈)及其时间边界难以准确识别的问题,尤其针对仅依赖语音活动无法区分对话回合与听者反馈或句内停顿的挑战。其解决方案的关键在于提出一种自动化处理流程,能够从分通道录制的对话数据中自动提取对话回合(turns)与听者反馈(backchannels),该流程整合了语音活动检测(voice activity detection)、通道能量滤波、时间合并、自动语音识别(automatic speech recognition)以及基于上下文的后处理步骤。通过在99段丹麦双人对话(共33组)上评估,结果显示整体检测可靠性F1值为0.621,回合与听者反馈的F1值分别为0.624和0.618,时间边界误差中位数均小于0.2秒,且在正常与不对称听觉条件下的性能无显著差异。尽管在部分参数设置下自动化结果与人工标注的一致性低于人工间一致性,但该方法仍可作为半自动化标注工作流中的可靠初筛工具,为对话动态的标准化与可重复标注提供一致基础。

链接: https://arxiv.org/abs/2610.03078
作者: Hanlu He,Harald Vilhelm Skat-Rørdam,Ingvi Örnólfsson,Ivana Konvalinka
机构: Technical University of Denmark(丹麦技术大学)
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Quantifying conversational dynamics requires reliable identification of interactional units and their temporal boundaries, but speech activity alone does not distinguish conversational turns from listener feedback or within-turn pauses. We present an automated pipeline for extracting turns and backchannels from separate-channel recordings of spontaneous dyadic conversation, designed to provide a consistent first-pass annotation for subsequent human review. The pipeline combines voice activity detection, channel-energy filtering, temporal merging, automatic speech recognition, and context-based post-processing. We evaluated the pipeline on 99 ten-minute Danish conversations from 33 dyads using segment-level detection reliability and temporal boundary error. Conversations were recorded under both normal and asymmetric listening conditions. In the latter, speech-shaped noise was delivered to one participant through bone-conduction headphones. Overall detection reliability was F1=0.621, with similar performance for turns F1=0.624 and backchannels F1=0.618. For successfully matched segments, median absolute onset and offset errors were 0.150 and 0.160s for turns and 0.130 and 0.180s for backchannels, respectively. Mean errors were substantially larger for turn boundaries, indicating a smaller number of large boundary mismatches. Performance did not differ significantly across the two experimental listening conditions. In a four-conversation case study, pipeline-human agreement was lower and more variable than human inter-annotator agreement and varied across parameter settings. These results support the pipeline as an automated first pass within a semi-automated annotation workflow, providing a consistent basis for more standardised and reproducible annotation of conversational dynamics.

[NLP-39] Unmasking Propaganda: A Comparative Analysis of Masked and Causal Language Models

【速读】: 该论文旨在解决在自然语言处理(NLP)背景下,尤其是针对操纵性政治传播中识别具体宣传技巧(propaganda techniques)这一难题。由于宣传技巧往往具有隐晦性且高度依赖语境,其与合法修辞手段之间的界限模糊,导致检测难度较大。宣传通过选择性强调某些事实而淡化或忽略其他信息,以塑造特定认知,进而影响受众态度、信念或行为。为应对这一挑战,本文基于SemEval-2020 Task 11数据集,对现代语言模型进行了对比分析,涵盖掩码语言模型(如XLM-RoBERTa和DeBERTa V3)与因果语言模型(来自OpenAI、Google、Mistral、Anthropic及Meta)。研究采用基础提示(base prompting)与思维链提示(chain-of-thought prompting)两种策略进行评估。实验结果表明,最优的掩码语言模型在技巧分类任务上达到63.18的F1分数,而最优的因果模型则取得63.62的F1分数,均优于现有先进模型。此外,研究发现不同模型在特定技巧(如煽动性语言、人身攻击)上表现优异,但在群体效应(bandwagon)和非黑即白谬误(black-and-white fallacy)等类型上存在明显短板。由此揭示出未来提升宣传检测能力的关键在于模型微调(fine-tuning)、集成建模(ensemble modeling)以及利用更大规模标注数据集。

链接: https://arxiv.org/abs/2610.03077
作者: Claudiu Creanga,Ioachim Lihor,Liviu P. Dinu
机构: University of Bucharest(布加勒斯特大学); Human Language Technologies Research Center(人类语言技术研究中心)
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Propaganda detection is an essential task in natural language processing (NLP), particularly in the context of manipulative political communications. However, identifying specific propaganda techniques presents a significant challenge due to their often subtle nature and reliance on context, making them difficult to distinguish from legitimate persuasive language. Propaganda often involves highlighting certain facts while downplaying or ignoring others to create a desired perception. This biased communication aims to influence attitudes, beliefs, or behaviors towards a particular cause or position. This paper explores advances in detecting propaganda techniques through a comparative analysis of modern language models, using the SemEval-2020 Task 11 dataset. We evaluated both masked language models (based on XLM-RoBERTa or DeBERTa V3) and causal models (from OpenAI, Google, Mistral, Anthropic and Meta), employing two prompting strategies: base and chain-of-thought prompting. Our results demonstrate improvements over state-of-the-art models, with the best-performing MLM achieving an F1 score of 63.18 in technique classification and the best causal model achieving 63.62. We also observed that certain models excel in specific techniques, such as loaded language and name-calling, while struggling with others like bandwagon and black-and-white fallacy. These findings suggest that fine-tuning, ensemble modeling, and the use of larger datasets can further enhance propaganda detection capabilities.

[NLP-40] SecJev: Bringing Security Expertise to System One Decision Models

【速读】: 该论文旨在解决安全领域中复杂观测与显式策略难以有效转化为精准决策的问题,尤其针对现有通用模型在安全场景下缺乏领域特化能力的缺陷。其核心解决方案是提出SecJev——首个面向安全领域的类Jev(Jev-like)决策模型家族,参数规模覆盖0.8B至9B,基于Kev的单次候选评分架构,能够从文本、遥测数据及观察历史中学习布尔型、选择型和有序型决策。关键创新在于构建了SecJev-Corpus,统一整合14项任务与8类数据源(包括工具输出、流量数据、联邦更新、共识信息、认证日志及车辆消息)的源-标签预测与显式策略评估,通过场景加权训练实现跨域适应,同时保持统一的类型化决策接口。安全领域专业化显著提升模型性能,其中SecJev-0.8B在任务宏准确率上较通用型Kev-9B高出20.51个百分点;相较于仅依赖答案生成的微调方法,SecJev在保持相近准确率与延迟的同时,显著降低峰值推理内存消耗。在新数据源组上的测试进一步验证了其在对抗提示注入与流量决策中的优势,并揭示其误报率与捕获依赖性特征。研究开源了适配器、决策头、SecJev-Corpus及完整训练与推理代码,为安全决策建模提供可复用的基础设施。

链接: https://arxiv.org/abs/2610.03073
作者: Zheng Chen,Fei Yu,Haohao Huang,Yang Li,Anlong Chen,Lei Chen
机构: University of Electronic Science and Technology of China(电子科技大学); National Key Laboratory of Security Communication(信息安全通信国家重点实验室)
类目: Cryptography and Security (cs.CR); Computation and Language (cs.CL)
备注: 22 pages, 1 figure

点击查看摘要

Abstract:Security workflows need models that turn complex observations and explicit policies into decisions. System One models introduced by Jev return typed predictions and probabilities; security specialization supplies the domain expertise behind those predictions. We introduce SecJev, to our knowledge the first family of Jev-like decision models specialized for security, spanning 0.8B to 9B parameters. Built on Kev’s single-pass candidate scorer, SecJev learns Boolean, choice, and ordered decisions from text, telemetry, and observation histories. We develop SecJev-Corpus to unify source-label prediction and explicit-policy evaluation across 14 tasks and eight sources. It covers tool outputs, traffic, federated updates, consensus, authentication, and vehicle messages. Scene-weighted training adapts the models across these domains while preserving a shared typed decision interface. Security specialization improves every model in the family; SecJev-0.8B outperforms general Kev-9B by 20.51 percentage points in task-macro accuracy. Comparisons with answer-only generative fine-tuning show close accuracy and latency with lower peak inference memory. Tests on new source groups reproduce gains over Kev in prompt-injection and traffic decisions, with capture-dependent false alarms. We release adapters, decision heads, SecJev-Corpus, and training and inference code.

[NLP-41] HARPO: Hallucination-Aware Reinforcement Learning for Faithful and Creative Language Generation

【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在知识密集型任务中生成幻觉内容(hallucination)的问题,该问题严重损害了模型的可靠性,同时又需在不牺牲创造性的前提下实现优化。其核心解决方案是提出一种名为HARPO的强化学习框架,关键在于通过一个基于可验证反馈训练的幻觉感知生成式奖励模型(Hallucination-Aware Generative Reward Model, HA-GRM),协同评估生成内容的忠实性与写作质量。该框架引入选择性激活机制(Selective Activation Mechanism, SAM),仅对HA-GRM判定为无幻觉的输出激活写作奖励,从而引导模型在保持真实性的同时提升创造性;此外,采用数据课程(data curriculum)策略,渐进式地将训练重心从创造性写作过渡到幻觉导向任务,以增强模型对幻觉的识别与规避能力。实验结果表明,基于Qwen3-4B的HA-GRM在RAGTruth数据集上达到78.08%的响应级F1分数,显著优于监督微调基线(66.37%),且在多跳检索增强生成任务中,幻觉率由3.29%降至1.02%,同时在创意写作评测中得分提升至27.54%,验证了该方法在兼顾忠实性与创造力方面的有效性。

链接: https://arxiv.org/abs/2610.03063
作者: Tiezheng Yu,Yuxin Jiang,Jinpeng Li,Shuning Sun,Fei Mi,Haoli Bai,Lifeng Shang
机构: Fei Mi; Haoli Bai; Lifeng Shang; Huawei Technologies(华为技术)
类目: Computation and Language (cs.CL)
备注: 11 pages

点击查看摘要

Abstract:Large Language Models (LLMs) are prone to generating hallucinated content, which compromises their reliability in knowledge-intensive tasks. To address this challenge without sacrificing creativity, we propose HARPO, a reinforcement learning framework designed to jointly optimize faithfulness and creativity. HARPO incorporates a Hallucination-Aware Generative Reward Model (HA-GRM), trained via verifiable feedback, to assess both faithfulness and writing quality. A Selective Activation Mechanism (SAM) activates writing rewards only for outputs judged hallucination-free by HA-GRM, while a data curriculum progressively shifts training from creative writing to hallucination-centric tasks. On RAGTruth, our Qwen3-4B-based HA-GRM achieves a response-level F1 score of 78.08%, compared with 66.37% for the supervised fine-tuning baseline. Experiments on Qwen2.5 and Qwen3 models from 1.7B to 8B parameters show improvements in both faithful generation and writing quality. On Qwen3-4B, HARPO reduces the HA-GRM-judged hallucination rate on MultiHopRAG from 3.29% to 1.02%, while increasing the Arena-Hard-v2.0 creative-writing score from 16.95% to 27.54%.

[NLP-42] he Geometry of Knowledge Accessibility in Large Language Models

【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在知识调用中的可靠性问题,即尽管模型蕴含广泛知识,但无法始终有效访问所需信息。其核心问题是:如何量化并预测模型对特定查询的知识可及性(knowledge accessibility)。解决方案的关键在于发现知识可及性在模型对查询的表示空间中具有简单的几何结构——可及性高的查询在表示空间中更接近中心,而可及性低的查询则远离中心,形成一条清晰的知识边界。这一几何特性表明,知识可及性与查询到中心的距离呈单调递减关系,且该距离排序在不同数据集间具有可迁移性。此外,控制实验验证了该几何结构与知识可及性的关联强于与推理难度的关联。该发现还揭示了不同干预策略的有效时机:查询重写对近中心的高可及性查询更有效,链式思维(chain-of-thought)推理在边界附近效果更佳,而外部检索在边界之外能带来更大收益。因此,该研究不仅为理解模型内部知识组织提供了新的几何视角,也为实现自适应推理提供了可量化的预生成信号。

链接: https://arxiv.org/abs/2610.03052
作者: Lihu Chen
机构: Imperial College London(帝国理工学院)
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Large language models (LLMs) contain broad knowledge, but they cannot access all of it reliably. We study this problem through knowledge accessibility, which describes whether the knowledge needed for a query can be recalled from the model. We find that knowledge accessibility has a simple geometric structure in the model’s representation of the query alone, before any generation. More accessible queries are closer to a center in the representation space, while less accessible queries are farther away. This geometry reveals a knowledge boundary that separates more accessible queries from less accessible ones. Accessibility consistently decreases with distance from the center, and this distance-based ordering transfers across datasets even when the centers differ. Controlled experiments further show that the centered geometry is more closely related to knowledge accessibility than to reasoning difficulty. The geometry also reveals when different interventions are useful. Query rewriting helps more for accessible queries, chain-of-thought reasoning helps more near the boundary, and retrieval gives larger gains beyond the boundary. These findings not only provide a new geometric view of how knowledge is organized in language models, but also suggest a useful pre-generation signal for adaptive inference.

[NLP-43] HyperThink: Text-to-Parameter Hypernetworks for Efficient Reasoning

【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在多步推理任务中依赖长时序思维链(long-form thinking traces)所导致的高推理延迟问题,其核心挑战在于序列解码带来的显著计算开销。解决方案的关键在于提出HyperThink——一种文本到参数(text-to-parameter)的推理加速框架:通过轻量级超网络(hypernetwork)根据输入问题动态预测对基础模型少量参数的更新,结合向量量化解码器将参数更新约束在有限的可复用模式集合中,从而提升鲁棒性与迁移能力。该方法在基础模型输出上端到端训练,使测试阶段无需生成冗长的思维链,仅需一次超网络前向传播即可完成参数适配,随后直接生成简洁的分步解答与最终答案,大幅减少生成令牌数。实证结果表明,HyperThink显著优化了准确率-延迟权衡曲线中的低延迟区域,在数学与通用推理任务上表现突出,尤其在接近“无思考”(near-non-thinking)的极低延迟场景下取得最大收益。

链接: https://arxiv.org/abs/2610.03039
作者: Donggyun Kim,Jack Lu,Chanwoo Kim,Mengye Ren,Seunghoon Hong
机构: KAIST School of Computing (KAIST 计算机学院); NYU Courant Institute School of Mathematics, Computing, and Data Science (纽约大学柯朗数学科学研究所,数学、计算与数据科学学院)
类目: Computation and Language (cs.CL); Machine Learning (cs.LG)
备注: COLM 2026

点击查看摘要

Abstract:Long-form thinking traces can substantially improve the multi-step reasoning performance of large language models (LLMs), but they introduce high inference-time overhead, with latency dominated by sequential decoding. We propose HyperThink, a text-to-parameter approach that amortizes this reasoning computation into a single query-conditioned parameter update: a lightweight hypernetwork reads the question and predicts updates to a small subset of the base LLM’s parameters, while a vector-quantized decoder constrains them to a finite set of reusable patterns to improve robustness and transfer. Trained end-to-end on outputs from the base model itself, HyperThink eliminates long thinking traces at test time: after one hypernetwork forward pass, the adapted model generates a concise step-by-step solution and final answer without an intermediate trace, using far fewer tokens while retaining strong reasoning performance. Empirically, HyperThink improves the low-latency region of the accuracy-latency trade-off on mathematical and general reasoning tasks, with its strongest gains in the near-non-thinking regime.

[NLP-44] Adaptive Second-Order Solvers for Fast Stochastic Diffusion Sampling ICLR2027

【速读】: 该论文旨在解决扩散模型采样过程中时间离散化(time-discretization)对采样成本与生成质量之间权衡的显著影响问题。由于逆过程的计算难度在采样轨迹上及不同数据分布间存在变化,离散化策略的选择至关重要。其解决方案的关键在于引入比例-积分(Proportional-Integral, PI)步长控制机制,并结合自研的扩散噪声归一化误差估计器(diffusion noise-normalized error estimator),使步长调整不仅依赖当前误差,还融合历史误差信息,从而实现更平滑、稳定的自适应步长更新。此外,研究发现每样本的自适应轨迹具有共享结构,可聚合为一个固定离散化调度(fixed schedule),在保留大部分自适应优势的同时显著降低计算开销。实验表明,在自然图像和语言数据集上,该固定调度在匹配神经网络评估次数(NFE)条件下优于广泛使用的EDM调度;而PI自适应求解器在多数情况下优于其他随机与自适应基线,尤其在低至中等NFE下表现突出。值得注意的是,基于样本的自适应性收益具有任务依赖性:在1D简化示例中效果显著,但在图像与语言任务中提升有限,甚至平均调度有时表现更优。

链接: https://arxiv.org/abs/2610.03034
作者: Ella Kemperman,Luca Ambrogioni
机构: Radboud University (拉德堡德大学)
类目: Machine Learning (cs.LG); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)
备注: Submitted to ICLR 2027

点击查看摘要

Abstract:Diffusion models rely on numerical solvers requiring time-discretization, which has a large influence on the tradeoff between sampling cost and quality. However, the computational difficulty of the reverse process varies along the sampling trajectory and across data distributions, making the choice of discretization important. We adapt proportional-integral (PI) step-size control to diffusion, using our diffusion noise-normalised error estimator. Unlike existing adaptive methods in diffusion that respond only to the current error, the PI solver also incorporates the previous error, yielding smoother step adaptation. We further show that these per-sample trajectories exhibit shared structure and can be aggregated into a fixed schedule that retains much of the benefit of adaptive sampling. We evaluate both approaches on natural-image and language datasets, in terms of quality, measured by FID at a matched number of neural network evaluations (NFE), comparing them with widely used stochastic solvers and schedules. For images, our fixed discretization outperforms the commonly used EDM schedule in terms of sample quality when used with the stochastic Heun sampler, and with the EDM-churn sampler at low NFE. Additionally, our PI adaptive solver obtains better FID than most stochastic and adaptive baselines, although it does not beat the EDM-churn sampler at low NFE. Moreover, we find our solver outperforms both the EDM and the entropy schedule on language diffusion at low-to-medium NFE in terms of perplexity, with the drawback of lower token entropy. Lastly, we find that the benefit of per-sample adaptivity is problem-dependent. It is highly beneficial in 1D toy examples, while only marginal for image and language data, where the average schedule sometimes even outperforms the PI-adaptive solver. Code is available at this https URL

[NLP-45] ailoring the Quantization Space for 1-Bit KV Cache Compression

【速读】: 该论文旨在解决长上下文大语言模型(LLM)推理中键值(Key-Value, KV)缓存带来的内存瓶颈问题,尤其针对极低比特压缩(如1-bit)场景下现有向量量化(Vector Quantization, VQ)方法性能急剧下降的挑战。其核心问题是:在极端压缩条件下,传统VQ方法难以有效利用有限的聚类中心(centroid)来表征大量通道间的统计相关性与误差敏感性,导致重建误差显著增加。为此,本文提出TaSQ(Tailored Quantization for KV Cache),其关键创新在于通过查询引导的通道加权(query-guided channel weighting)、跨注意力头归一化(cross-head normalization)以及协方差感知的通道分组(covariance-aware channel grouping),对VQ的目标空间进行重构,以更准确地捕捉缓存激活的误差敏感性和统计结构。该方法具备旋转位置编码(RoPE)兼容性,可无缝嵌入投影权重与码本中,保持原有VQ查找结构不变,仅引入可忽略的服务开销。实验表明,TaSQ在通用任务、长链思维推理及长上下文检索等基准上均显著优于现有低比特KV缓存量化基线,并保障了推理稳定性;在单张RTX 6000 Ada GPU上,其SGLang实现相较BF16基线支持高达14倍的批处理规模,峰值吞吐提升1.87倍。

链接: https://arxiv.org/abs/2610.03027
作者: Minsoo Cheong,Donghyun Son,Sungjoo Yoo
机构: Seoul National University (首尔国立大学); Stanford University (斯坦福大学)
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:The key-value (KV) cache becomes a major memory bottleneck in long-context LLM inference, placing substantial pressure on memory capacity and bandwidth. To mitigate this bottleneck, vector quantization (VQ) has emerged as a promising approach for aggressive KV cache compression. However, existing VQ methods degrade substantially in the 1-bit regime. At such extreme compression, each codebook must represent a larger group of channels with a limited set of centroids, making effective use of its capacity increasingly challenging. To address this, we introduce \textbfTaSQ , which tailors the VQ target space by combining query-guided channel weighting, cross-head normalization, and covariance-aware channel grouping to better reflect the error sensitivity and statistical structure of cached activations. Since these transforms are RoPE-compatible and can be easily merged into projection weights and codebooks, TaSQ preserves the conventional VQ lookup structure and adds negligible serving overhead. Across general, long-chain-of-thought reasoning, and long-context retrieval benchmarks, TaSQ consistently outperforms existing low-bit KV cache VQ baselines while preserving reasoning stability. On a single RTX 6000 Ada GPU, its SGLang implementation supports up to 14\times larger batch sizes and achieves 1.87\times higher peak throughput compared to the BF16 baseline.

[NLP-46] Verifiable Articulable and Tacit Components of Preference

【速读】: 该论文旨在解决生成内容质量评估中“隐性偏好”(tacit preferences)与可表述、可验证的显性标准之间存在的系统性差距问题。具体而言,尽管人类对短篇小说的吸引力、新闻报道的新颖性或数学证明的优美性等创造性判断具有高度共识,但这些判断往往难以通过明确规则或可量化的指标来表达和验证。当前主流AI模型优化依赖于可阐述的评价标准(如RLAIF、RLVR中的显式评分体系),而忽视了隐性偏好在真实人类判断中的核心作用。为此,作者构建了大规模标注数据集CreativePreferences,涵盖280万条文本及3.17亿次人类偏好判断,覆盖7个创意领域并设立42个基准任务。研究提出三种建模方式:可执行程序、规则库以及密集训练的模型(V、A、VAT)。通过引入一种新颖的测量方法——结合捕获-重捕法(capture-recapture)识别可表述与可验证度量、剔除虚假变量,并估算未发现度量的价值,揭示出普遍存在的“可表述性差距”(articulability gap, VAT-VA)与“可验证性差距”(verifiability gap, VAT-V)。这些差距存在于所有领域,包括传统上被视为高度可验证的数学与软件工程领域,以及以论断和新颖性为核心的新闻、专利与同行评审领域。其规模随参与评价人数增加而扩大,符合柯林斯(Collins)提出的集体隐性知识理论。研究进一步揭示两个关键后果:(1)使用完整模型(包含隐性偏好信号)生成的内容更贴近人类真实偏好,且常与显性标准相悖;(2)过度强调可表述性会偏离原始的隐性偏好维度,类比于古德哈特定律(Goodhart’s law)。因此,该研究建议应根据任务特性决定是否采用提示工程、改进学习机制以捕捉隐性信号,或在必要时保留人类判断,从而提升生成质量与评估可靠性。

链接: https://arxiv.org/abs/2610.03025
作者: Alexander Spangher,Sheldon Huang,Andreas Haupt,Noah D. Goodman,Diyi Yang,Daniel E. Ho,Sanmi Koyejo
机构: Stanford University (斯坦福大学); University of Toronto (多伦多大学)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computers and Society (cs.CY); Machine Learning (cs.LG)
备注: 15 pages main text, 14 pages of references, 107-page appendix (136 pages total); 15 figures, 48 tables; 213 references

点击查看摘要

Abstract:What makes a short story gripping; a news article newsworthy; or a math proof elegant? These constructs resist articulation or verification; their meaning is at least partially tacit. However, modern AI models are improved primarily via articulated constitutions, rubrics and verifiers (i.e. in RLAIF and RLVR); tacit components of preferences are typically understudied. We introduce a large, labeled preference dataset CreativePreferences, containing 2.8M texts labeled by 317M human preference judgments across 7 creative domains, with 42 benchmark tasks. We model these labels with executable programs, rubric banks and densely trained models (V, A and VAT, respectively). We observe robust articulability gaps, VAT-VA; and verifiability gaps, VAT-V; we estimate upper and lower bounds for each gap with a novel measurement approach that discovers articulable and verifiable metrics, identifies spurious variables and estimates the value of undiscovered metrics using capture-recapture. These gaps occur across all domains, even in domains traditionally treated as fully verifiable: correctness-centered domains (i.e. mathematics and software engineering) and claim- and novelty-centric domains (i.e. news, patents, peer review). The size of the gap varies based on domain (e.g. peer review and creative writing have the largest articulability gaps) and widens as more people take part in the judgment, consistent with Collins’ collective tacit knowledge. We show two consequences: (1) on human generations, the full model more closely matches human preferences, often in disagreement with articulated criteria, and (2) in an analogy to Goodhart’s law, articulating preference shifts it away from the tacit dimension. Articulability and verifiability gaps are consequential; we give recommendations on when tasks can be prompted; how learning mechanisms might improve; and when to leave judgments with humans.

[NLP-47] ReSCUE: Re-translation with Sentence Commitment for Unsegmented Long-Form Simultaneous Sign Language Translation NEURIPS2026

【速读】: 该论文旨在解决实时场景下手语翻译(Simultaneous Sign Language Translation, SLT)中长期连续、未分段视频流的在线翻译问题。现有方法多局限于句子级离线处理,依赖预分割输入,无法适应真实场景中连续无边界的手语视频流。其核心解决方案在于提出ReSCUE这一统一框架,关键创新包括:① 推理感知训练(inference-aware training),使模型能够有效处理部分输入、非手语停顿及多句上下文;② 稳定化重翻译机制(stabilized re-translation),在低延迟条件下实现可修正的输出并减少输出闪烁;③ 句子确认机制(sentence commitment mechanism),支持在线分句与记忆管理。实验表明,ReSCUE在标准句子级基准上实现了更低延迟和更优的低延迟翻译质量,在长序列未分段数据集上接近使用真实句边界标注的离线系统性能,同时显著降低延迟,验证了其在真实流式场景中的实用性。

链接: https://arxiv.org/abs/2610.03022
作者: Sihan Ren,Gaozheng Li,Yuanshang Quan,Yiming Qin,Fuyi Yang,Chang Liu,Lan Xu,Minye Wu
机构: ShanghaiTech University (上海科技大学); DGene; University of Derby (德比大学)
类目: Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL)
备注: Accepted at NeurIPS 2026

点击查看摘要

Abstract:Simultaneous Sign Language Translation (SLT) is critical for real-time communication, yet existing methods remain largely confined to sentence-level, offline settings that assume pre-segmented inputs. These assumptions hinder deployment in realistic scenarios involving continuous, unsegmented video streams. We present ReSCUE, a unified framework for simultaneous SLT on unsegmented long-form sign language videos that aligns training and inference with realistic streaming conditions. ReSCUE combines inference-aware training to handle partial inputs, non-signing pauses, and multi-sentence contexts, stabilized re-translation to enable low-latency yet revisable predictions with reduced output flicker, and a sentence commitment mechanism for online segmentation and memory management. Experiments on standard sentence-level benchmarks show that ReSCUE achieves lower latency and the best translation quality under low-latency settings. On long-form unsegmented datasets, ReSCUE approaches the translation quality of oracle offline systems that use ground-truth sentence boundaries, while operating at substantially lower latency, demonstrating its practicality for real-world streaming scenarios.

[NLP-48] Recursive Self-Improvement in Unified Multimodal Models

【速读】: 该论文旨在解决统一多模态模型(Unified Multimodal Models, UMMs)在自提升训练过程中存在的监督瓶颈问题,即现有方法仅依赖视觉理解能力对图像生成进行监督,导致模型无法有效实现文本与视觉能力的双向协同优化。其解决方案的关键在于提出递归跨能力自提升(Recursive Cross-capability Self-improvement, RSI)框架,通过让模型在每轮迭代中同时生成图像并自主阅读图像以识别自身缺陷,进而编写程序针对这些缺陷进行修正,并利用程序执行结果作为外部真实依据(source of truth)验证生成内容的正确性。这一机制避免了错误在多轮迭代中累积,实现了文本生成、程序编写与视觉理解三者之间的闭环反馈:经验证的图像渲染用于训练图像生成能力,标注后的图像与模型自生成的正确程序则共同训练视觉理解与程序生成能力。实验表明,在图表生成任务上,经过四轮RSI后模型性能从45.7%提升至60.2%,显著优于持续训练组的46.3%;其中验证性构建(verified construction)贡献主要提升,针对性修复失败点额外带来3.5%增益;同时,验证程序占比由48.9%升至95.2%,图像阅读准确率从55.6%提升至87.4%,充分验证了该方法的有效性与可扩展性。

链接: https://arxiv.org/abs/2610.03002
作者: Huijuan Wang,Chufan Shi,Cheng Yang,Yaokang Wu,Taylor Berg-Kirkpatrick,Xuezhe Ma
机构: University of Southern California(南加州大学); University of California San Diego(加州大学圣地亚哥分校); Carnegie Mellon University(卡内基梅隆大学)
类目: Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Unified multimodal models (UMMs) understand and generate both text and images, which lets a model produce its own training data. Existing self-improvement in UMMs keeps supervision on the visual side, where image understanding judges image generation. We propose recursive cross-capability self-improvement (RSI), a training loop in which the text and visual abilities of a UMM supply training data for one another. In each round, the model generates images and reads them to find where it falls short. It then writes programs aimed at these shortcomings, and execution verifies every result against its specification. Verified renders train image generation, while labeled renders and the model’s own correct programs train visual understanding and program writing. Program execution thus acts as a source of truth outside the model, so errors do not accumulate across rounds. We study RSI on charts and build BasicChartBench to evaluate open models early in training. On requests worded differently from training, four rounds of RSI raise the score from 45.7% to 60.2%, while continued training stays at 46.3%. Verified construction carries most of the gain, and targeting the model’s failures adds 3.5%. Along the way, the share of verified programs rises from 48.9% to 95.2%, and the reader’s accuracy on edited renders rises from 55.6% to 87.4%.

[NLP-49] OmniConfess: Eliciting Token Confessions to Mitigate Omni-Modal Hallucination

【速读】: 该论文旨在解决多模态大语言模型(Omni-modal Large Language Models, OmniLLMs)在生成过程中因依赖错误或无关证据而产生幻觉的问题。现有推理阶段的方法虽能在一定程度上减少幻觉,但通常无法揭示生成结果所依赖的具体证据来源。为此,论文提出一种无需训练的解决方案——OmniConfess,其核心在于通过在通道级(channel-wise)证据干预下对候选响应进行逐标记(token-level)重评分,生成结构化的“认罪书”(confession),以明确揭示每个生成标记对不同模态证据的依赖关系。该机制使模型能够识别并剔除由无关或矛盾证据驱动的错误承诺,从而保留基于正确证据的内容。为全面评估方法性能,研究构建了包含3,540个样本的OmniHalluBench基准,覆盖文本、图像、音频和视频等多种模态及判断与自由生成任务。实验表明,OmniConfess在异构模态与任务设置下均能有效缓解幻觉现象。

链接: https://arxiv.org/abs/2610.02999
作者: Huiqiang Rong,Haoran Luo,Hui Feng,Zhonghong Ou,Kaiwen Xue,Guoxin Zhang,Yifan Zhu
机构: Beijing University of Posts and Telecommunications (北京邮电大学); Nanyang Technological University (南洋理工大学)
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Omni-modal large language models (OmniLLMs) unify text, images, audio, and video, yet hallucinate when generation relies on the wrong evidence. Existing inference-time methods can reduce hallucinations, but rarely reveal which evidence sustains a generated commitment. We introduce OmniConfess, a training-free method for mitigating omni-modal hallucinations. It fixes a candidate response and re-scores it at token resolution under controlled channel-wise evidence interventions, producing a structured token-by-channel confession that reveals the response’s evidential dependence. OmniConfess uses this confession to preserve grounded content and correct commitments driven by irrelevant or contradictory evidence. To evaluate OmniConfess, we construct OmniHalluBench, a 3,540-example benchmark built from six datasets spanning text, image, audio, and video settings and both judgment and free-form generation. Experiments show that OmniConfess mitigates hallucinations across heterogeneous modality and task settings. Our code and benchmark are publicly available at this https URL.

[NLP-50] Sentry: Learning to Recover from LLM Agent Failures at Test Time

【速读】: 该论文旨在解决大语言模型(Large Language Model, LLM)智能体在执行任务过程中因工具调用无效、重复动作或推理缺乏依据而中途失败的问题,核心挑战在于如何有效利用失败经验以提升智能体的可靠性。现有方法要么将失败教训直接注入智能体上下文(context),导致条件性知识在无故障场景下误触发,降低性能;要么采用运行时干预(runtime intervention),虽能即时响应失败但无法从修复中学习。本文提出的关键解决方案是:失败知识本质上是条件性知识(conditional knowledge),应仅在故障发生时被有条件地暴露。为此,作者设计了Sentry——一个与智能体并行运行的失败管理模块。Sentry在检测到失败后,从外部剧本(playbook)中检索相关教训以指导恢复,无需访问任务奖励即可验证恢复有效性,并仅在成功恢复时才将新教训存入剧本,确保整个剧本不进入智能体上下文。实验表明,Sentry在多个代理基准测试中均显著优于最强的运行时干预基线(平均提升37%)和上下文演化基线(在可比基准上提升39%),且结合上下文演化进一步提升性能。此外,所学教训具备跨任务泛化能力,控制实验证明:即使相关教训仍可调用,若将完整剧本暴露于智能体上下文,反而会损害性能,验证了条件性暴露的重要性。

链接: https://arxiv.org/abs/2610.02994
作者: Changxiu Ji,Amy Lu,Qizheng Zhang,Kunle Olukotun
机构: Stanford University (斯坦福大学)
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:LLM agents often fail mid-task due to invalid tool calls, repeated actions, or poorly grounded reasoning, and learning from these failures is a path to reliability. We find that how failure knowledge reaches the agent matters as much as what it contains. Failure lessons are conditional: kept in the agent’s context, they misfire when their failure is absent, and removing them from an evolving playbook improves performance. Runtime interventions, in contrast, act only when a failure occurs but do not learn from their repairs. We argue that failure knowledge is conditional knowledge and should be conditionally exposed, and instantiate this principle in Sentry, a failure-management layer that runs alongside the agent. When Sentry detects a failure, it retrieves matching lessons from an external playbook to guide recovery, verifies without access to task rewards whether the agent recovered, and stores a new lesson only if it did; the full playbook never enters the agent’s context. Across multiple agentic benchmarks, Sentry outperforms the strongest runtime-intervention baseline on every benchmark, by 37% on average, and the strongest context-evolution baseline by 39% on the two benchmarks where both are evaluated; combining Sentry with context evolution yields further gains. Learned lessons transfer to held-out tasks, and controlled experiments show that exposing the full playbook to the agent lowers performance even when relevant lessons remain available on demand.

[NLP-51] OLMo-Detect: A Multi-Stage Confounder-Controlled Benchmark for Membership Inference on Large Language Models

【速读】: 该论文旨在解决现有大型语言模型(LLM)成员推断(Membership Inference, MI)评估基准存在的三大核心问题:训练阶段覆盖不全、成员与非成员样本在分布上对齐不足,以及非成员样本未经过严格过滤以排除训练数据中的潜在重合。其解决方案的关键在于提出 OLMo-Detect——一个基于完全开源的 OLMo 2 流水线构建的多阶段、混杂因子控制的基准。该基准系统性地覆盖预训练、中段训练和后训练阶段,通过在三个关键维度上显式对齐成员与非成员样本,并采用 infini-gram 方法对非成员进行严格过滤,从而提升评估的可信度与可比性。此外,为评估模型在分布偏移下的鲁棒性,研究进一步引入 OLMo-Detect (Shifted) 变体,使成员与非成员在分布上存在偏差。实验结果表明,当前主流成员推断攻击性能有限(最佳 AUC 仅为 0.68),且在跨领域评估中监督式方法表现下降;推断能力在中段训练期达到峰值,主要受数据类型影响而非训练阶段本身;模型规模从 1B 到 32B 时性能提升但趋于饱和;且无一无监督攻击能有效抵御分布偏移,其 AUC 最多下降 0.42。研究还验证了结论在 OLMo 3 及非 OLMo 模型上的泛化能力。

链接: https://arxiv.org/abs/2610.02986
作者: Tao Shi,Chaoyi Xiang,Qiongkai Xu,Jey Han Lau
机构: The University of Melbourne(墨尔本大学); Macquarie University(麦考瑞大学)
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Membership inference on large language models (LLMs) aims to determine whether a given text sample was included in an LLM’s training data, without access to its training corpus. Despite recent progress, existing benchmarks suffer from three limitations: limited coverage of training stages, insufficient distributional alignment between members and non-members, and lack of rigorous filtering of non-members against the training corpus. To address these limitations, we propose OLMo-Detect, a multi-stage, confounder-controlled benchmark built upon the fully open OLMo 2 pipeline. OLMo-Detect spans pre-training, mid-training, and post-training, explicitly aligns members and non-members on three key axes, and rigorously filters non-members via infini-gram. To assess robustness to distribution shifts, we further introduce OLMo-Detect (Shifted), a variant where members are misaligned with non-members. We evaluate 15 unsupervised and 3 supervised membership inference attacks (MIAs) across the OLMo 2 family, finding that: (i) overall performance is limited: the best unsupervised and supervised MIAs both reach an AUC of only 0.68, and supervised MIAs degrade under cross-domain evaluation; (ii) MIA performance peaks at mid-training and is lower at pre-training and post-training, a pattern driven by data type rather than a stage effect: curated math data is far more detectable than other types; (iii) overall scores improve from 1B to 13B but plateau at 32B; and (iv) no unsupervised MIA is robust to distribution shifts, with AUCs shifting by up to 0.42. Finally, we find that our findings on OLMo 2 generalize to OLMo 3 and non-OLMo models.

[NLP-52] A Guideline-Augmented Multi-Agent Framework for Schema-as-Code Biomedical Named Entity Recognition

【速读】: 该论文旨在解决基于大语言模型(LLM)的生物医学命名实体识别(BioNER)方法中存在的两大关键问题:一是检索到的示范样本和外部生物医学知识对特定数据集的标注语义支持不足,导致实体边界、类型范围及标注规范存在歧义;二是自由生成模式缺乏足够的结构控制,易产生格式错误、虚构提及、实体重复和边界错误等问题。其解决方案的核心在于提出一种基于规则增强的多智能体框架GAMA(Guideline-Augmented Multi-Agent framework),通过从标注训练样本中归纳候选标注规则并经验证构建可靠的数据集特定规则记忆库,再由规划组件生成带理由的候选实体跨度-类型假设,编码组件将其转化为符合模式约束的实体对象,最后通过验证模块检查跨度锚定、类型有效性与结构合规性,并采用双循环迭代优化机制修正低置信度或无效预测,从而实现高精度、结构化且可解释的实体识别。

链接: https://arxiv.org/abs/2610.02970
作者: Songtao Li,Yijia Zhang,Shidi Zhang,Jianyuan Yuan,Fengyu Zhang,Hongfei Lin
机构: Dalian Maritime University(大连海事大学); Beijing Institute of Technology(北京理工大学); Northeastern University(东北大学); Dalian University of Technology(大连理工大学)
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Large language models (LLMs) have shown promising potential for biomedical named entity recognition (BioNER) through instruction following and in-context learning. However, existing LLM-based BioNER methods still face two key limitations. First, retrieved demonstrations and external biomedical knowledge provide limited support for dataset-specific annotation semantics, leaving entity boundaries, type scopes, and annotation conventions ambiguous. Second, free-form generation lacks sufficient structural control, often leading to invalid formats, hallucinated mentions, duplicated entities, and boundary errors. To address these limitations, we propose GAMA, a guideline-augmented multi-agent framework for schema-as-code BioNER. GAMA first induces candidate annotation rules from labeled training instances and verifies them against annotated data to construct reliable dataset-specific guideline memory. Guided by these verified rules, a planning component generates ranked span-type hypotheses with rationales, and a coding component converts them into schema-constrained entity objects. A verification module then checks span grounding, type validity, and structural compliance, and performs dual-loop refinement to correct invalid or low-confidence predictions. Experiments on five widely used BioNER datasets with multiple LLM backbones show that GAMA consistently outperforms strong LLM-based baselines. Ablation and parameter analyses further verify the effectiveness of the proposed components.

[NLP-53] Understanding Trajectory Heterogeneity in Federated World Model Learning

【速读】: 该论文旨在解决联邦学习在临床时间序列建模中因时间上下文访问受限而导致的性能瓶颈问题,核心挑战在于:患者数据在时间维度上的分布不均与所有权边界限制,使得各客户端无法完整获取跨时间步的轨迹信息,从而影响世界模型(World Model)对状态演化规律的学习。其解决方案的关键在于构建一个系统性的评估框架,通过在MIMIC-IV数据库的8个疾病队列上进行小时级动作条件下的临床预测任务,引入基于严重程度的客户端所有权划分、患者分离的数据构建规则、局部历史与未来窗口约束机制,并采用从1至32小时的配对滚动预测评估策略,全面考察联邦算法在不同时间跨度下的表现。研究揭示了时间上下文覆盖度、参与率、优化行为差异以及预测时域分辨率等多维度因素对模型性能的联合影响,指出当前主流联邦算法(如FedAvg、FedProx)在长窗预测中存在显著局限,且其内部更新行为具有高度异质性,表明单纯依赖算法标签无法反映实际训练动态,亟需从时空感知和动态适配角度重新设计联邦临床世界模型的训练范式。

链接: https://arxiv.org/abs/2610.02957
作者: Yipan Wei,Zhaokun Yan,Ziming Hong,Jiaqi Wu,Lixu Wang
机构: Wuhan University(武汉大学); China Academy of Information and Communications Technology(中国信息通信研究院); The University of Sydney(悉尼大学); Tsinghua University(清华大学); The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)
类目: Machine Learning (cs.LG); Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:World models learn state evolution from trajectories, making access to temporal context a central training requirement. Federated learning can use distributed records, while ownership boundaries within a trajectory restrict the examples each client can construct. Our study benchmarks this cross-time setting through hourly action-conditioned clinical prediction on eight MIMIC-IV disease cohorts, comprising 40.87 million transition memberships. We specify severity-based client ownership, patient-separated construction, local history and future-window rules, and paired rollout evaluation from one to 32 hours. A matrix of ten federated algorithms covers 32 disease–partition configurations under five rounds of ten-percent participation. Three findings emerge from existing results and training logs. First, client ownership and participation jointly restrict long-window coverage: only 7.55%–21.36% of pooled-available 32-step windows have a locally complete anchor visited during training, averaged across diseases. Second, finer severity partitions accompany higher FedAvg error in 15 of 16 paired comparisons, while algorithm gains are small and horizon-dependent: FedProx reduces mean error by 0.56%, with no consistent improvement at 32 steps. Third, algorithm labels conceal distinct update behavior, including inactive extrapolation and orders-of-magnitude differences in update scale. Cached-update performance also varies strongly across trajectory partitions under the same benchmark protocol. These results establish temporal access, participation coverage, optimization behavior, and horizon-resolved prediction as complementary dimensions for evaluating federated clinical world models.

[NLP-54] Enhancing Biomedical Named Entity Recognition via Multiple Programming Languages Instruction Tuning and Ensemble Method

【速读】: 该论文旨在解决现有指令微调(Instruction Tuning)方法在生物医学命名实体识别(BioNER)任务中面临的两大挑战:一是传统自然语言指令通常将实体标注序列化为扁平文本输出,缺乏对类型化实体提取的结构约束;二是高质量生物医学标注数据稀缺,单一序列化输出形式限制了模型的结构多样性,降低其鲁棒性。为此,论文提出MITE(Multiple Programming Languages Instruction Tuning and Ensemble),其核心创新在于将BioNER重构为结构到结构的生成任务,通过将指令和实体输出以多种编程语言格式(如Python、C++、Java)表示,保留相同语义的前提下引入结构多样性监督,无需依赖外部生物医学知识或额外标注。在推理阶段,MITE采用基于实体级别的投票策略融合不同编程语言格式的预测结果,有效降低语言特异性偏差,提升模型鲁棒性。实验表明,MITE在六个主流BioNER数据集上持续优于代表性BERT与大语言模型(LLM)基线,并展现出优异的跨数据集泛化能力,消融实验和参数分析进一步验证了各组件的有效性与稳定性。

链接: https://arxiv.org/abs/2610.02949
作者: Songtao Li,Yijia Zhang,Jianyuan Yuan,Shidi Zhang,Fengyu Zhang,Hongfei Lin
机构: Dalian Maritime University(大连海事大学); Beijing Institute of Technology(北京理工大学); Northeastern University(东北大学); Dalian University of Technology(大连理工大学)
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Instruction tuning has become a common paradigm for applying large language models (LLMs) to biomedical named entity recognition (BioNER). However, existing instruction-tuning approaches still face two key challenges. First, conventional natural-language instructions typically serialize BioNER annotations as flat textual outputs, providing limited structural constraints for typed entity extraction. Second, high-quality biomedical annotations are limited, and learning from a single serialized output form may restrict structural diversity and reduce model robustness. Although external biomedical knowledge can be introduced to alleviate data scarcity, it often requires costly resource construction. To address these challenges, we propose MITE, a Multiple Programming Languages Instruction Tuning and Ensemble method for BioNER. MITE reformulates BioNER as a structure-to-structure generation task by representing both instructions and entity outputs in code-formatted representations. Specifically, each training instance is transformed into multiple programming-language formats, including Python, C++, and Java, while preserving the same underlying entity semantics. These language-specific representations provide structurally diverse supervision without requiring external biomedical knowledge or additional annotations. During inference, MITE aggregates predictions from different code formats through an entity-level voting strategy, reducing language-specific prediction variance and improving robustness. Experiments on six widely used BioNER datasets demonstrate that MITE consistently outperforms representative BERT-based and LLM-based baselines and exhibits strong cross-dataset generalization. Ablation and parameter analyses further verify the effectiveness and robustness of the proposed components.

[NLP-55] Continual Graph Memory for Mathematical Research Agents

【速读】: 该论文旨在解决在长期数学研究任务中,生成式数学研究代理(mathematical research agent)因需并行调动大量智能体进行长时间证明搜索而产生的中间结果海量堆积、知识难以组织与复用的核心挑战。其解决方案的关键在于提出一种基于持续图记忆(Continual Graph Memory)的数学研究代理Ansatz,该系统通过构建统一的图结构记忆空间,显式地组织所有中间探索结果(包括事实、计划、反例等)及其相互关系;引入依赖感知的检索机制以精准获取局部上下文;设计证据敏感的管理者(curator)动态更新研究前沿并提炼过往尝试的经验;并通过范围化召回机制,主动揭示早期陈述与负面发现,支持局部重证而非盲目复用。实验表明,Ansatz在全部10个“第一证明第二批次”问题上实现闭合,并独立求解了Jamison毛虫猜想及Erdős问题289、348、488,同时在多个开放问题上取得部分进展,验证了其在长周期数学探索中持续推进与恢复能力的强大表现。

链接: https://arxiv.org/abs/2610.02945
作者: Junyi Zhang,Jinxi Yu,Eric Hanchen Jiang,Jiachen Lu,Zhi Zhang,Xinjie He,Hyunsik Chae,Ethan Ji,Alexander K Taylor,Vigyan Sahai,Yiwen Kou,Kai-Wei Chang,Raghu Meka,Nanyun Peng,Amit Sahai,Terence Tao,Wei Wang
机构: University of California, Los Angeles(加利福尼亚大学洛杉矶分校)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Using frontier agent harnesses to tackle mathematical research problems has emerged as an effective means of advancing mathematics. However, solving frontier problems in mathematics may require a massive number of agents working in parallel for extended periods to construct proofs, thereby generating an enormous volume of intermediate proof results. Organizing these intermediate results throughout a long-horizon proof-search process and reusing knowledge gained from prior explorations remain major challenges. We present Ansatz, a mathematical research agent built around Continual Graph Memory, a graph-based, evolvable, cross-problem mathematical research memory system that explicitly organizes the entire proof search process and reuses information from exploration trajectories of previous problems. Specifically, we develop a unified graph memory that represents all intermediate exploration results, including facts, plans, and counterexamples, together with edges that explicitly represent the relationships among them; dependency-aware retrieval supplies precisely targeted local context; an evidence-sensitive curator updates the research frontier and distills lessons from prior attempts; and scoped recall surfaces earlier statements and negative findings for local re-proving rather than uncritical reuse. Experiments cover runs across all ten First Proof Second Batch problems, together with four component studies. Ansatz reports closure on all ten research tasks, demonstrating its ability to sustain and resume long-horizon mathematical search. Beyond these problems, Ansatz also produces solutions to the Jamison caterpillar conjecture and Erdős Problems 289, 348, and 488 without human intervention, and makes partial progress on several open problems, illustrating its strong ability to solve open mathematical research problems.

[NLP-56] Output Language Confusion under Multilingual Prompt Contamination NEURIPS2026

【速读】: 该论文旨在解决在真实多语言部署场景中,传统事实性评估基准所依赖的“纯净单语提示”与“精确匹配评分”假设失效的问题。具体而言,当检索增强生成(RAG)管道返回混合语言文本或用户输入多语言内容时,现有评估方法无法准确反映模型在复杂语言干扰下的真实表现。其解决方案的关键在于提出一种轻量级、可完全复现的多语言干扰评估协议——多语言干扰(Multilingual Distractor Interference, MDI),该协议无需额外数据或标注,仅通过在事实性问题前添加语义无关的外语句子来模拟真实世界的语言混杂环境。研究发现,模型在面对特定语言干扰(如印地语干扰)时可能出现非预期的脚本切换现象(如将答案转为天城文书写),导致精确匹配评分产生严重偏差;但经人工审查后发现,多数此类响应在语义上仍正确,表明原始评估中的高幻觉率主要源于脚本不一致而非实质错误。因此,论文强调在混合语言环境下,必须将幻觉率分解为“脚本切换”与“语义错误”两个独立成分,才能准确评估模型可靠性。此外,英文干扰段落引发普遍拒绝回答(abstention escalation),反映出阅读理解混淆这一关键失败模式,对多语言RAG系统具有直接负面影响。

链接: https://arxiv.org/abs/2610.02926
作者: Riju Marwah,Ritvik Garimella,Khusham Bansal,Atishay Jain,Amit Sheth
机构: University of Tübingen, Germany; Artificial Intelligence Institute, University of South Carolina, USA; Thapar Institute of Engineering Technology, India; Indian Institute of Technology Kanpur, India; Indian AI Research Organization, India
类目: Computation and Language (cs.CL)
备注: Accepted at NeurIPS 2026 Workshop LP4FM

点击查看摘要

Abstract:Standard factual benchmarks assume clean monolingual prompts and exact-match scoring, two assumptions that break simultaneously in real-world multilingual deployment, from retrieval-augmented generation pipelines returning mixed-language passages to users pasting multilingual web content. We introduce Multilingual Distractor Interference (MDI), a lightweight and fully replicable evaluation protocol requiring no new data or annotation, in which factual questions are preceded by a semantically irrelevant foreign-language sentence, and evaluate five instruction-tuned LLMs across TruthfulQA and TriviaQA under eight distractor conditions (40,000 evaluations). Our central finding is a metric confound: for Llama-3.1-8B under a Hindi distractor, 58% of responses switch to Devanagari script, yielding a raw hallucination proxy of 0.710, but manual review reveals that 120 of 148 script-switched responses that were correct under clean conditions remain semantically correct despite being written in the wrong script, reducing the adjusted semantic hallucination rate to 0.470. All other models respond through abstention escalation with no hallucination increase. A paragraph-length English distractor triggers near-universal abstention (0.806-0.998) across all models, consistent with reading-comprehension confusion, a failure mode with direct consequences for multilingual RAG pipelines. TruthfulQA multiple-choice accuracy is unaffected under all single-sentence conditions. These results show that exact-match hallucination rates in mixed-language settings should be decomposed into script-switching and semantic error components before drawing conclusions about model reliability.

[NLP-57] Probe the Harness: Setup Checks for Stale-Data RL Comparisons in Language Models

【速读】: 该论文旨在解决在语言模型训练中使用过时样本(stale samples)时,不同方法的性能评估因实验设置细节而产生误导性结论的问题。现有方法通常通过与重要性校正基线进行比较来评估有效性,但研究发现实验环境中的细微设定差异可逆转方法间的相对排名,从而导致错误结论。其解决方案的关键在于提出PTH(Probe The Harness)——一套系统性的检查工具,用于揭示并验证实验框架中潜在的隐蔽偏差。研究通过在verl和单GPU训练器上对比行为无关方法SAN与截断重要性采样(TIS)的实验发现,四个关键的隐藏设定(如PPO比率基于学习者自身重计算的概率、数据种子未覆盖TIS分支、重放队列重复使用首个批次达33次、两个损失归一化器与描述不符)虽使日志指标看似正常,却实质性地扭曲了比较基准。经PTH检查后,TIS在verl上与SAN表现相当,在训练器中实现稳定学习,而SAN仍保持优势。论文贡献包括各偏差项的特征标识、在采样延迟下TIS与未校正GRPO的参考结果,以及可复用的PTH检查清单,为可信训练评估提供了方法论保障。

链接: https://arxiv.org/abs/2610.02911
作者: Taiheng Pan
机构: The University of Melbourne(墨尔本大学)
类目: Machine Learning (cs.LG); Computation and Language (cs.CL)
备注: 8 pages, 2 figures, 4 tables

点击查看摘要

Abstract:Methods for training language models on stale samples are judged by comparisons against importance-corrected baselines. We show that details of the experimental harness can reverse the observed ranking of methods, and we introduce PTH (Probe The Harness), a set of checks that makes the harness visible. Our case is a comparison between SAN, a behaviour-free method, and truncated importance sampling (TIS) on verl and in a single-GPU trainer, in which SAN first finished ahead in both stacks. Four details of the harness changed this comparison: the PPO ratio was taken against the learner’s own recomputed probabilities, the data seed did not reach the TIS arm, the replay queue reused its first batch for 33 updates, and two loss normalisers differed from their description. In each case the logged quantity looked consistent with a working setup, while the quantity that defines the comparison went unchecked. With the harness checked, TIS matches SAN on verl, and in the trainer TIS learns steadily while SAN keeps a margin. We contribute the signature of each detail and its effect on the comparison, reference results for TIS and uncorrected GRPO under sampler lag, and the PTH checklist.

[NLP-58] Misinformation Without Triggers: From Factual Answers to Downstream Decisions

【速读】: 该论文旨在解决生成式语言模型在训练过程中吸收虚假信息后,其输出的客观事实与后续决策行为之间存在的不一致性问题。具体而言,研究关注的是:当模型被注入错误事实时,尽管其直接回答可能仍能通过真实性检测(direct probe),但这些错误信息仍会系统性地影响模型在实际决策任务中的表现,形成“审计差距”(audit gap)。解决方案的关键在于揭示这一差距的存在——即即使模型对某一事实的回答是正确的,该事实依然可能误导其后续决策;反之,即便错误信息被移除,基于错误信息形成的决策偏差仍难以完全纠正。研究通过两个实验验证此现象:一是受控的“猜首都”决策任务,显示错误训练数据显著提升错误选项的选择率,且决策准确率下降幅度超过可解释范围;二是真实世界中关于2019–20年澳大利亚山火的虚假社交媒体内容分析,发现模型即使在训练数据中错误计数被修正后,仍持续输出“有人因纵火被捕”的误导性结论,且该结论受表述方式影响,表明语言表达形式可独立于数字本身塑造错误认知。因此,该研究的核心突破在于证明了单一的事实正确性验证不足以确保决策可靠性,必须引入多层级、行为导向的评估机制以识别和缓解深层偏见。

链接: https://arxiv.org/abs/2610.02886
作者: Lin Tian,Marian-Andrei Rizoiu
机构: University of Technology Sydney (悉尼科技大学)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: 35 pages, 11 figures, 16 tables

点击查看摘要

Abstract:Language models learn from web documents, some of them false, and false content can reach a model’s answer to a factual question and the summaries and decisions that use it. Most data-poisoning studies add a trigger to the training data and activate it in the prompt. False documents can also change factual responses without any trigger, but we do not know whether the direct answer predicts the decision. In this work, we follow false content past the answer and find an \emphaudit gap between what a direct probe reports and what the model then does, comparing false training with matched truthful controls in a controlled decision task, \emphGuess the Capital, where a fixed decoder turns factual answers into a scored card choice, and on a misleading claim from Facebook posts about the 2019–20 Australian bushfires. Across eight models at dose 1,000, direct injected-choice rates reach 95.8–100%, while injected game choices increase by 1.7–14.4 percentage points over matched truthful training. The gap runs the other way too. Facts that pass the direct probe still push decisions toward the injected answer, and game accuracy drops further than those choices explain. Truthful correction brings the fact back but not the decisions built on it. We then look into the real-world bushfire case, models trained on the false posts say that people were arrested for arson even when they lose the inflated count, and in a count-by-wording factorial the misleading arrest wording produces arrest assertions even when the training count stays at 24. In simpler terms, \textbfa correct factual answer does not guarantee a correct decision, and losing the injected number does not remove the misleading story.

[NLP-59] Evaluating VQA in Vision Language Models using Cooperative Principles

【速读】: 该论文旨在解决生成式视觉语言模型(Generative Vision Language Models, VLMs)在视觉问答(Visual Question Answering, VQA)任务中对违反格赖斯合作原则(Grice’s maxims)的问题的鲁棒性问题。具体而言,研究关注当问题包含冗余、模糊或虚假信息等语用违规时,VLMs的性能下降现象。其解决方案的关键在于构建一种基于VLMs自身生成的“问题修饰”机制,通过引入非必要、模糊或错误信息来系统性地模拟语用违规场景,并在此基础上评估主流VLMs(如ChatGPT、Claude、Gemini和Llava)的表现差异。研究发现,当问题中的语用违规由人类主动设计时,人类表现出更高的语用推理能力且认知负荷较低;而当违规由AI生成时,尽管人类处理时间更短,但VLMs的准确率显著降低,揭示了人类与VLMs在语用理解机制上的本质差异,凸显了当前VLMs在复杂语境下缺乏真正语用推理能力的核心局限。

链接: https://arxiv.org/abs/2610.02878
作者: Monika Shah,Sudarshan Balaji,Somdeb Sarkhel,Sanorita Dey,Deepak Venugopal
机构: University of Memphis, TN, USA(孟菲斯大学, 田纳西州, 美国); Adobe Research, San Jose, CA, USA(Adobe研究实验室, 圣何塞, 加利福尼亚州, 美国); University of Maryland Baltimore County, USA(马里兰大学巴尔的摩县分校, 美国)
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:We evaluate the performance of Vision Language Models in Visual Question Answering (VQA) when questions violate Grice’s maxims. To do this, we use VLMs to generate question modifiers that add non-essential, ambiguous or false information and show that in the presence of such violations, the VLMs that we evaluate (ChatGPT, Claude, Gemini and Llava) show diminished performance. Further, we empirically show the difference between how humans reason pragmatically compared to VLMs, and the difference in VLM reasoning when it resolves violations that are human-induced compared to those that are AI-generated. Finally, we show that human cognitive effort (measured through time-on-task in an experiment) is lower for resolving VLM-induced violations, but VLMs themselves perform less accurately in such cases.

[NLP-60] Evaluating LLM -as-a-Judge Beyond Score Alignment: A Psychometric Analysis of Residual Judging Difficulty AACL

【速读】: 该论文旨在解决当前大语言模型(LLM)作为自动评分工具时,尽管在整体评分上与人类评分具有较高一致性,但可能在判断难度分布上存在系统性偏差的问题。具体而言,现有评估方法仅关注人类与模型评分的总体对齐程度,却忽略了两者是否对相同类型的摘要评价任务具有相似的“困难感知”。为此,论文从心理测量学(psychometric)视角出发,采用多面罗斯模型(Many-Facet Rasch Model)分别对人类和LLM的评分进行建模,将评分分解为潜在的摘要质量、评分者严厉度、维度严厉度以及评分尺度阈值等潜变量。基于此分解,提出“残差难度”(residual hardness)作为经过模型调整后的判断难度度量,并用于比较人类与LLM在评价难度结构上的异同。研究发现,在SummEval数据集上,17个开源大语言模型裁判虽在潜在摘要质量维度上表现出中等程度的一致性,但在残差难度结构上存在显著差异:人类与模型对哪些“摘要-维度”组合更难判断的认知不一致,且这种不一致呈现强维度依赖性——如连贯性(coherence)维度上人类认为更难,而一致性(consistency)维度则呈现模型偏难的倾向。此外,研究进一步表明,人类容易但模型难以处理的案例可部分通过源文本与摘要的可观测属性进行预测。这说明,仅依赖总体评分一致性无法全面反映大语言模型作为评判者的可靠性,而基于心理测量学的残差诊断方法能够提供更精细的判别依据,从而支持更精准的人机协同评价机制。

链接: https://arxiv.org/abs/2610.02877
作者: Longwei Cong,Sonja Hahn,Sebastian Gombert,Leon Camus,Fabian Zehner,Hendrik Drachsler,Ulf Kroehne
机构: DIPF | Leibniz Institute for Research and Information in Education(德国教育研究与信息研究所); Centre for International Student Assessment (ZIB)(国际学生评估中心); Faculty of Computer Science, Goethe University Frankfurt(法兰克福歌德大学计算机科学学院); Chemnitz University of Technology(开姆尼茨工业大学)
类目: Computation and Language (cs.CL)
备注: Accepted at AACL-IJCNLP 2026

点击查看摘要

Abstract:Large language models (LLMs) are widely used as automatic judges, with validity typically assessed via alignment with human scores. However, aggregate agreement fails to reveal whether humans and LLMs find the same evaluation cases difficult. In this paper, we study this problem in summarization evaluation from a psychometric perspective. We fit Many-Facet Rasch Models separately to human and LLM ratings to decompose scores into latent summary quality, rater severity, dimension severity, and rating-scale thresholds. Building on this decomposition, we define residual hardness as a model-adjusted measure of judging difficulty and compare whether human and LLM judges share the same hardness structure. Across 17 open-weight LLM judges on SummEval, we find that moderate alignment in latent summary quality does not imply alignment in residual hardness. Human and LLM judges differ in which summary–dimension units remain difficult, and this mismatch is strongly dimension-dependent. Consistency shows a pronounced LLM-hard shift, whereas coherence shows a human-hard shift. We further show that human-easy but LLM-hard cases are partially predictable from observable source–summary properties. These findings suggest that aggregate human alignment reflects only part of LLM-as-a-judge reliability, while psychometric residual diagnostics support more informative judge evaluation and more targeted human–LLM collaboration.

[NLP-61] ConvoDrift: A Multi-Turn Conversational Dataset for Modeling Stylistic Tone Evolution

【速读】: 该论文旨在解决对话中语言风格动态演变(linguistic style drift)在自然语言处理(NLP)领域长期被忽视的问题。现有风格控制数据集多聚焦于单句层面或假设风格静态不变,无法捕捉用户偏好在多轮交互过程中随时间发生的渐进性变化。为此,论文提出ConvoDrift数据集,其核心在于建模在固定语义意图下逐步演化的对话风格漂移(style drift),基于15,727个共享的多轮对话结构构建,并引入人物设定(persona-conditioned)对齐方法。每个对话包含六组提示-响应对,均标注风格漂移程度与风格方向标签,覆盖多种交际语域。进一步地,通过配对语义等价但风格迥异的回应,构建互补的成对数据集,并基于五种不同风格化沟通人格(communication personas)标注个性化偏好,从而实现对语言风格个性化与多元对齐(pluralistic alignment)的可控研究。关键解决方案在于:构建兼具语义一致性与风格动态变化的数据框架,并结合人类验证、大模型判别及自动词汇与语义分析,系统评估风格漂移的有效性与保真度——实验表明,尽管词汇层面发生显著变化,语义相似性仍得以保持,且人工标注的一致性(平均Krippendorff’s alpha为0.88)达到高可靠水平。

链接: https://arxiv.org/abs/2610.02873
作者: Vihindi Kotalawala,Pamoda Dilranga,Gayani Thoradeniya,Prasan Yapa
机构: Informatics Institute of Technology (斯里兰卡信息学院技术); Independent Researcher (独立研究员, 斯里兰卡); National Health and Medical Research Council (澳大利亚国家卫生与医学研究委员会); Luxembourg Centre for Systems Biomedicine (卢森堡系统生物医学中心), University of Luxembourg (卢森堡大学)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: 13 pages, 14 figures, 7 tables, Accepted paper at the 13th Conference on Computational Linguistics and Speech Processing (ROCLING) 2026

点击查看摘要

Abstract:The evolution of linguistic style in conversations is an underexplored issue in NLP. Most style-control datasets focus on sentences or assume a static style throughout, missing the dynamic shifts that occur as user preferences change during interactions. We introduce ConvoDrift, a dataset designed to model progressive stylistic conversational tone drift under fixed semantic intent. It is built on 15,727 shared multi-turn conversational structures for adaptation and persona-conditioned alignment methods. It consists of six prompt-response pairs per conversation, each with the annotation of style drift and style direction labels. These pairs cover a range of communication genres. We further derive a complementary pairwise dataset by pairing semantically equivalent but stylistically distinct responses and annotating persona-conditioned preferences using five distinct style communication personas, enabling the controlled study of personalisation and pluralistic alignment in language tone. In addition to dataset construction, we conduct a comprehensive evaluation involving human validation, LLM-as-judge assessment, and automatic lexical and semantic evaluations. Across seven Likert criteria annotated by three human annotators, the average Krippendorff’s alpha is 0.88, and our lexical and semantic analyses show that drift events induce lexical changes while preserving semantic similarity.

[NLP-62] Adaptive Mutual Distillation for Balanced Multi-Task Post-Training of Large Language Models

【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在多任务后训练中,因各任务训练数据量不均衡而导致性能提升受限的问题。现有方法主要关注单一模型训练过程中的任务贡献平衡,但不同任务平衡策略所生成的模型具有互补优势,而传统方法未能有效利用这种跨模型监督的潜力。其核心挑战在于:跨模型知识迁移的效果随任务、迁移方向及训练阶段变化而异,难以静态设定最优的蒸馏权重。为此,论文提出自适应互蒸馏(Adaptive Mutual Distillation, AMD)框架,通过联合训练两个采用不同任务平衡策略的模型,并利用跨任务共享的短周期训练探针评估候选蒸馏权重调整方案,结合任务级验证分数动态选择每项任务与迁移方向的最优调整策略。实验结果表明,AMD在六个基准测试和三种模型架构上均显著优于监督微调(Supervised Fine-Tuning, SFT)基线及现有任务平衡方法;进一步通过模型融合可获得统一推理模型,平均性能较多任务SFT提升2.91分,充分验证了该框架在实现高效、自适应跨模型知识协同方面的有效性。

链接: https://arxiv.org/abs/2610.02856
作者: Baohang Li,Xiaocheng Feng,Yichong Huang,Chengpeng Fu,Wenshuai Huo,Zekun Zhou,Zekun Yuan,Tingjia Zhang,Bing Qin
机构: Harbin Institute of Technology (哈尔滨工业大学); Peng Cheng Laboratory (鹏城实验室)
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Multi-task post-training of large language models (LLMs) aims to improve performance across tasks with unequal amounts of training data. Existing methods focus primarily on balancing task contributions during single-model training. Different task-balancing strategies can produce models with complementary strengths, creating opportunities for mutual distillation. However, the usefulness of cross-model supervision can vary across tasks, transfer directions, and stages of training. We propose Adaptive Mutual Distillation (AMD), a collaborative post-training framework that jointly trains two models with different task-balancing strategies. AMD evaluates candidate adjustments to distillation weights through short training probes shared across tasks, then uses task-wise validation scores to select an adjustment for each task and transfer direction. Across six benchmarks and three LLM backbones, both AMD models achieve higher average benchmark scores than supervised fine-tuning (SFT) baselines trained with the same sampling strategies. They also outperform the task-balancing methods evaluated in our experiments. Merging the two trained models can further improve their average benchmark score while yielding a single model for inference. The merged models outperform multi-task SFT by an average of 2.91 points across the three backbones.

[NLP-63] How Robust Is Multimodal Claim Verification to LLM Rewriting? AACL2026

【速读】: 该论文旨在解决生成式人工智能(Generative AI)在科学任务中因文本风格变化而可能影响模型决策的问题,特别是在多模态声明验证(multimodal claim verification)任务中的表现稳定性。研究发现,尽管文本的风格和词汇选择发生变化,大多数模型在准确性上仍保持稳健,表明其对自然润色和可控注入等风格修改具有较强鲁棒性;然而,概率输出存在系统性偏移,尤其是采用强调不确定性的表达方式(hedging-oriented conditions)时,几乎在所有模型中均引发显著的概率变化,而单纯语法修正或流畅性提升等通用润色操作则影响较小。因此,解决方案的关键在于识别并量化风格转换对模型置信度输出的影响,揭示语言表达细微变化如何通过非语义层面的信号干扰模型判断,为提升科学推理任务的可解释性和可靠性提供依据。

链接: https://arxiv.org/abs/2610.02841
作者: Yun-Ang Wu,Xanh Ho,Andre Greiner-Petter,Sunisth Kumar,Tian Cheng Xia,Florian Boudin,Akiko Aizawa
机构: NII LLMC, Japan; National Institute of Informatics, Japan; University of Göttingen, Germany; University of Bologna, Italy; The University of Tokyo, Japan; Inria, LS2N, Nantes Université, France
类目: Computation and Language (cs.CL)
备注: Accepted to AACL 2026 (Main Conference). 18 pages

点击查看摘要

Abstract:LLMs are known to introduce stylistic changes into generated text, yet how these stylistic shifts affect model decisions on scientific tasks remains underexplored. In this paper, we focus on multimodal claim verification, where the goal is to determine whether a textual claim is grounded in a given piece of evidence. We apply two rewriting strategies: natural rewriting, which simulates how researchers routinely use LLMs to polish academic text, and controlled injection, which inserts a single LLM-associated word to isolate the effect of vocabulary choice. We evaluate 11 open-weight models spanning five VLM families and ranging from 2B to 38B parameters. We find that models are robust to these modifications: most show no significant drop in accuracy, and compared to prior work on review-score manipulation, verification appears far more stable. However, consistent probability shifts do occur. Hedging-oriented conditions produce significant shifts across nearly all models, while boosting conditions show a weaker effect and general polishing conditions (e.g., grammar correction, fluency improvement) have little effect.

[NLP-64] o Explore The Strange New World Beyond Data Distribution: System Behavior Causality Tax and Non-causal Base Model

【速读】: 该论文旨在解决语言模型(Language Models, LMs)中因果性(causality)作为核心假设的必要性与最优性问题。传统架构依赖因果性作为建模基础,但研究表明其与系统行为(System Behavior, S)之间存在持续的不匹配与矛盾,而这些矛盾主要源于对系统行为这一超越数据分布的主导因素的忽视。论文提出SBD框架,将系统行为S作为证据下界(ELBO)中不可约简的组成部分进行建模,并理论揭示了一种反直觉的“因果税”(Causality Tax)现象:因果性本质上是因忽略S而产生的次优近似,伴随结构性误差。为应对潜在变量分析挑战,研究通过隐式测量、基于理论边界控制及神经正切核(Neural Tangent Kernel, NTK)评估验证了S的影响。特别地,构建了非因果变分族Green Shell(GSH),以分治策略替代S组件的串行依赖链,从而有效缓解因果税。在懒惰训练(lazy-training)初期,NTK谱分析显示GSH在信噪比上实现7dB以上的提升;进入后期阶段,其泛化能力更显著增强,多尺度拟合能力提升高达20%。综上,SBD将系统行为确立为与因果性和分布拟合并列的理论抽象,为语言模型基座架构的设计与优化开辟了新路径。

链接: https://arxiv.org/abs/2610.02839
作者: Xianzhi Zeng,Jiangneng Li,Gao Cong
机构: Nanyang Technological University (南洋理工大学)
类目: Machine Learning (cs.LG); Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:We show that the causality of language models (LMs) may not be necessary nor optimal. This is the case when system behavior (denoted as S ) is incorporated as a first-principle Bayesian feature. Here, S refers to extra dominant factors beyond the data space, and they involve coupled effects. Despite being the de facto foundation of modern architecture, recent studies indicate persistent mismatches and contradictions with causality. These issues largely stem from system behavior rather than the data distribution. We therefore propose the SBD framework, which incorporates S as an irreducible component of the evidence lower bound (ELBO). SBD theoretically reveals a counter-intuitive Causality Tax phenomenon, where causality emerges as a suboptimal approximation with an additional structural error, due to the obliviousness to S . To address the challenge of latent variable analysis, we validate the SBD-predicted impact of S via implicit measurements, theoretical-bound-guided controls, and Neural Tangent Kernel (NTK) evaluations. In particular, we construct Green Shell (GSH) to show the possibility of reducing Causality Tax. GSH is a non-causal variational family, and it replaces the sequential dependency chain of S components with a divide-and-conquer partition. NTK spectra in the lazy-training regime confirm that GSH always achieves significantly tighter error bounds than causality, with 7dB+ improvement in signal-to-noise ratio. In the relatively later stage of lazy-training, GSH further leads to superior generalization (up to 20% richer multi-scale fitting capabilities). Taken together, SBD establishes system behavior as a complementary theoretical abstraction besides causality and distribution fitting, opening new research avenues such as designing and optimizing LM base models.

[NLP-65] Clinical Concept Centers in LLM s

【速读】: 该论文旨在解决当前大型语言模型(Large Language Models, LLMs)在临床决策支持中可靠性评估的局限性问题,即现有研究主要依赖对模型输出文本的语言层面评价,而忽视了模型内部潜在空间(latent space)中更为丰富且具有因果意义的表征。其核心挑战在于:模型的推理过程是否真正基于可定位、可解释且因果驱动的临床概念表征。论文提出的关键解决方案是将行为评估拓展至模型的潜在空间,通过机制可解释性(mechanistic interpretability)方法,验证开放权重的大型语言模型中是否存在可定位、可解释并实际影响决策的临床概念中心(clinical concept centers)。研究发现,在所测试的11个开放权重模型中均存在此类概念中心,它们仅在对应临床叙事激活时响应,且在受限与开放式任务中均以有意义且因果的方式驱动模型行为。这些概念中心不仅是分析性表征,更可作为可操作的神经电路应用于临床实践。从评估角度看,模型在对抗性角色提示下仍保持内部一致性,并持续使用相关概念中心;而对齐提示则提升下游临床表现。从性能角度看,模拟真实部署场景表明,沿这些概念中心进行引导可显著提升下游任务效果。最终的盲法临床医生验证进一步证实,这些概念中心的激活与使用能够有效预测临床医生的偏好,从而为模型的可信度评估与临床应用提供新的理论基础和实践路径。

链接: https://arxiv.org/abs/2610.02829
作者: Aishik Nagar,Abhishek Vaidyanathan,Arun-Kumar Kaliya-Perumal,Elijah Tzen Hsuen Boey,Stefan Winkler
机构: National University of Singapore(新加坡国立大学); SAP Asia Pte Ltd(SAP亚洲私人有限公司); Nanyang Technological University(南洋理工大学); Singapore Institute of Technology(新加坡科技学院)
类目: Computation and Language (cs.CL); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Large language models are increasingly used in clinical settings. However, research into the reliability and performance of these models has focused almost entirely on the language substrate, scoring what the model says. Mechanistic interpretability has found that the latent space carries a higher fidelity of representation than the text: internal representations not only encode substantially more than the output verbalizes, but the stated reasoning also systematically omits features that causally drive the answer. An evaluation of model behavior in terms of mechanistic interpretability has not been explored in clinical decision support. In this work, we extend behavioral evaluation into the latent space and ask whether clinical concepts exist as locatable, causally used representations inside open-weight LLMs. We find dedicated clinical concept centers in the latent space of all eleven open models we test. These concept centers are interpretable, firing only on their aligned clinical narratives, and meaningfully and causally drive model behavior in both constrained and open-ended settings. They are not just analytical representations, but circuits that can be utilized in clinical practice, and we explore their use from the perspective of both evaluation and performance. From the evaluation standpoint, models stay internally coherent and keep using the relevant concept centers even under adversarial role-based priming, while aligned priming improves downstream clinical performance. From a performance perspective, we simulate realistic deployment settings and find that steering models along these centers leads to meaningful downstream improvements. Finally, we conduct a blinded clinician validation and find the activation and usage of these concept centers predicts clinicians preferences.

[NLP-66] FSPO: Policy-Consistent Risk and Pareto-Feasible Control for Budgeted LLM RL Post-Training

【速读】: 该论文旨在解决自适应大语言模型(LLM)强化学习后训练中面临的三个关键问题:一是基于行为轨迹训练的未来风险模型无需评估部署控制器所引发的风险;二是对记录的状态-动作对进行校准的评分在经过选择性动作决策后可能出现校准偏差;三是单一资源的最小成本约束无法普遍保证多资源联合延续的可行性。针对这些问题,本文提出了一种名为FSPO(Feedback-State Policy Optimizer)的预算受限LLM强化学习后训练反馈-状态控制器,其核心解决方案包含三部分:首先,构建一个与策略一致的风险-剩余值(risk-to-go)模型,其贝尔曼目标遵循用于未来决策的冻结控制器,同时引入长时程效用模型以支持全局优化;其次,提出决策条件轨迹校准(DCTC),通过由临时控制器生成的交叉拟合轨迹对风险进行动态校准;最后,设计帕累托资源延续证书(PRCC),仅在剩余时域内存在非占优累积保留量可行时才允许执行动作。实验表明,在匹配的GRPO资源包络下,FSPO在保留数据集和分布外数据集上的准确率分别达到66.11%和59.43%,优于最强的自适应基线PB2(64.47%和57.03%),且在三组配对训练种子下相较上下文老虎机方法分别提升2.42和3.19个百分点。在高行为-部署不匹配场景下,策略一致性风险将选定决策的期望校准误差(ECE)从0.108降至0.053,DCTC在匹配接受率下使ECE从0.039降至0.022,PRCC可完全消除18个动作目录中的虚假可行准入(0.197→0.000),而三者协同使用使轨迹失败率从0.181降至0.083,验证了各组件的有效性与协同优势。

链接: https://arxiv.org/abs/2610.02828
作者: Miaobo Hu,Shuhao Hu,Xiaobo Guo,Xin Wang,Bokun Wang,Daren Zha,Jun Xiao
机构: University of Chinese Academy of Sciences (中国科学院大学); Institute of Information Engineering, Chinese Academy of Sciences (中国科学院信息工程研究所)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
备注: 40 pages, 4 figures

点击查看摘要

Abstract:Adaptive LLM reinforcement-learning post-training changes multiple training actuators online, including rollout temperature, group size, clipping, KL regularization, verifier allocation, and update budget. Three coupled issues remain unresolved. A future-risk model trained from behavior trajectories need not estimate the risk induced by the controller that will be deployed; a score calibrated on logged state-action pairs can become miscalibrated after selective action choice; and independent per-resource minimum costs do not in general certify a feasible multi-resource continuation. We introduce FSPO, a feedback-state controller for budgeted LLM RL post-training that addresses these issues jointly. FSPO learns a policy-consistent risk-to-go model whose Bellman target follows the same frozen controller used for future decisions, together with a long-horizon utility model. Decision-conditioned trajectory calibration (DCTC) calibrates risk on cross-fitted trajectories generated by actions selected by provisional controllers. A Pareto resource continuation certificate (PRCC) admits an action only when a non-dominated cumulative reservation remains feasible over the residual horizon. Under a matched GRPO resource envelope, FSPO reaches 66.11% held-out and 59.43% OOD accuracy, compared with 64.47% and 57.03% for PB2, the strongest evaluated adaptive baseline. Three paired training seeds give gains of +2.42 and +3.19 percentage points over the contextual bandit on held-out and OOD evaluation. Under high behavior-deployment mismatch, policy-consistent risk lowers selected-decision ECE from 0.108 to 0.053; DCTC lowers it from 0.039 to 0.022 at matched acceptance; PRCC removes false-feasible admissions on an 18-action catalog ( 0.197\rightarrow0.000 ); and enabling all three components reduces trajectory failure from 0.181 to 0.083 in a factorial ablation.

[NLP-67] xt-Centric Post-Training for Omni-Modal Reasoning

【速读】: 该论文旨在解决多模态大语言模型(Omni Large Language Models)在联合音频-视觉推理能力提升过程中面临的高昂数据构建与训练成本问题。其核心挑战在于:尽管模型在单跳推理任务中表现良好,但在多跳推理任务中仍存在显著困难,且感知(perception)与推理(reasoning)目标在局部优化过程中出现部分解耦现象。为此,论文提出一种以文本为中心的后训练范式(text-centric post-training paradigm),其关键在于将推理能力的优化主要依赖于纯文本推理训练,通过监督微调(Supervised Fine-Tuning, SFT)结合强化学习(Reinforcement Learning, RL)实现高效推理能力提升;随后,仅使用少量真实音频-视觉数据进行轻量级的强化学习微调,以恢复并优化感知能力。该方案不仅使Qwen2.5-Omni-7B模型在九项推理指标上的几何平均分相比基线提升25.83%,且训练效率远超完整原生多模态训练路径(节省56.6%的GPU小时),同时在不引入任何音频-视觉数据的情况下,仅通过纯文本生成数据亦可实现21.01%的性能增益。最终,该范式在大幅降低计算开销的同时,有效平衡了推理与感知能力,实现了性能与效率的协同优化。

链接: https://arxiv.org/abs/2610.02819
作者: Ziyang Cheng,Yuhao Wang,Hongcheng Liu,Qimin Wu,Jingru Fan,Chen Qian,Yanfeng Wang,Yu Wang
机构: Shanghai Jiao Tong University (上海交通大学); Alibaba Cloud (阿里云)
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Improving joint audio-visual reasoning in Omni Large Language Models typically incurs substantial data construction and training costs. Our diagnostics reveal multi-hop reasoning difficulties despite correct answers to all corresponding single-hop questions and suggest partial decoupling in the local optimization of perception and reasoning objectives. This motivates post-training with different emphases on these capabilities. Text-only reasoning training yields gains across data sources, model scales, and families. With the best-performing text-only configuration, supervised fine-tuning followed by reinforcement learning (RL) raises Qwen2.5-Omni-7B’s geometric mean of nine reasoning scores by 25.83% over the base model, outperforming the complete native audio-visual route with 56.6% fewer GPU-hours. Training on data synthesized entirely by a text-only LLM raises this geometric mean by 21.01% without audio-visual data in construction or training. However, text-only training degrades perception. We therefore propose a text-centric post-training paradigm: text-only training provides the main reasoning optimization, and reduced-data native audio-visual RL then refines perception. Refinement uses about 90% fewer input tokens than full-data audio-visual RL, restores perception above the base level, and retains 93.5% of the best-performing text-only pipeline’s reasoning gain.

[NLP-68] RMCW: A Deletion-Robust Watermark Based on Reed–Muller Codes for Language Models

【速读】: 该论文旨在解决大语言模型(Large Language Model, LLM)水印在遭遇后处理攻击时鲁棒性不足的问题,尤其针对删除攻击导致的令牌位置偏移及水印位置对齐失效这一关键挑战。其解决方案的核心在于提出基于里德-穆勒码(Reed–Muller Code Watermarking, RMCW)的水印机制:通过秘密密钥划分词汇表,在生成阶段将里德-穆勒码的代数结构嵌入文本序列;在检测阶段,利用密钥映射文本至对应词汇桶,并采用伯莱克-威尔奇(Berlekamp–Welch)算法测试局部子序列是否满足低次里德-索罗门(Reed–Solomon)一致性,从而恢复残存的局部代数结构。该方法不依赖全局码字恢复,而是通过仿射线限制所诱导的里德-索罗门一致性特性,有效应对删除等破坏性攻击,在C4和ELI5数据集上基于OPT-1.3B与Llama-3.1-8B-Instruct模型均表现出优异的干净文本可检测性及对抗多种删除与重写攻击的鲁棒性。

链接: https://arxiv.org/abs/2610.02817
作者: Yi Wang,Baicheng Chen,Yu Wang,Jian Zhao,Yilei Chen,Tianxing He
机构: Institute for Interdisciplinary Information Sciences, Tsinghua University(清华大学交叉信息研究院); University of Wisconsin–Madison(威斯康星大学麦迪逊分校); Shanghai Qi Zhi Institute(上海智元研究院); The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳) ); Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所); Xiongan AI Institute(雄安人工智能研究院)
类目: Cryptography and Security (cs.CR); Computation and Language (cs.CL); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Large Language Model (LLM) watermarking provides a lightweight mechanism for identifying text generated by a specific model, but its robustness remains fragile under post-processing attacks. Deletion attacks are particularly challenging because they shift token positions and break the alignment between observed tokens and their original watermark positions. We propose Reed–Muller Code Watermarking (RMCW), an LLM watermarking method based on Reed–Muller codes. In contrast to global codeword recovery, RMCW searches for surviving local algebraic structure, leveraging the Reed–Solomon consistency induced by affine-line restrictions of Reed–Muller codewords. During generation, RMCW injects a Reed–Muller structure into the sequence via a secret-keyed vocabulary partition. During detection, it maps the given text to keyed vocabulary bins and tests local subsequences for low-degree Reed–Solomon consistency using Berlekamp–Welch tests. Experiments on C4 and ELI5 datasets with OPT-1.3B and Llama-3.1-8B-Instruct show that RMCW preserves strong clean-text detectability and outperforms or matches the baseline methods under several deletion and rewriting attacks. Our code is available at this https URL.

[NLP-69] ROUTEAUDIT: Interaction-Aware Identification for Budgeted Multi-Verifier Routing

【速读】: 该论文旨在解决自适应多验证器系统在评估过程中因验证器目录、可用性、资源计账、信息过滤或评分机制随策略变化而产生的质量-成本差距(quality-cost gap)不可靠问题。传统方法在策略变动时无法准确归因于特定变量,导致评估结果失真。其解决方案的关键在于将验证器路由建模为合同条件识别问题(contract-conditioned identification problem),通过引入一个包含请求支持、验证器目录、实际可用性、资源计账、在线过滤及后置打分等要素的合同(contract),实现对策略变动的可追溯与可量化分析。在此基础上,提出 ROUTEAUDIT 协议,新增三个可测量对象:合同格(contract lattice) 用于平均所有合法桥接顺序下的坐标增量并报告归因及其路径敏感性;策略无关响应带(policy-independent response tape) 用于识别自适应策略下不同观测的配对序列对比;以及针对不完全匹配的请求级边界,利用可观测潜在结果构造精确的有限总体区间。协议在预言机接入前承诺支付观测与账本事件,并为每次比较生成归属证书。实验表明,在两个独立缓存数据集上,匹配的静态 SF+SA 策略性能等同于级联策略,且其优于全静态策略的增益(0.1797 和 0.1250)可准确归因于验证器集合的差异;在 1,319 个任务请求上,学习策略与 RLVR 策略的性能分别达到 0.9522 与 0.9553,显著优于匹配静态策略(0.9484),其配对差异为 +0.0068,置信区间分别为 [0.0015, 0.0122] 与 [0.0006, 0.0131],且控制归因恢复的路由均方误差仅为 0.0011,端点重构误差达 0.0004。因子分析、桥接顺序及随机提供者研究验证了证书接口的有效性,而 RLVR 在同一识别合同下提供了学习策略的鲁棒性压力测试。

链接: https://arxiv.org/abs/2610.02808
作者: Miaobo Hu,Shuhao Hu,Xiaobo Guo,Xin Wang,Bokun Wang,Tianshu Fu,Daren Zha,Jun Xiao
机构: University of Chinese Academy of Sciences (中国科学院大学); Institute of Information Engineering, Chinese Academy of Sciences (中国科学院信息工程研究所)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
备注: 45 pages, 15 figures

点击查看摘要

Abstract:Adaptive multi-verifier systems are commonly compared through endpoint quality-cost gaps, even when the verifier catalog, availability, accounting, information filtration, or scorer changes with the policy. We formulate verifier routing as a contract-conditioned identification problem. The contract records request support, verifier catalog, realized availability, resource accounting, online filtration, and post-trace scoring; a matched route contrast changes only the policy coordinate. ROUTEAUDIT adds three measurable objects to this contract. A contract lattice averages coordinate increments over every admissible bridge order and reports the resulting attribution together with its path sensitivity. A policy-independent response tape identifies paired sequential contrasts when adaptive policies reveal different observations. For incomplete matching, request-level bounds use whichever potential outcome remains observed and give a sharp finite-population interval. The protocol commits paid observations and ledger events before the oracle join and returns an attribution certificate for each comparison. On two held-out raw-tail caches, matched static SF+SA equals the cascade, assigning the apparent gains of 0.1797 and 0.1250 over full static to the verifier-set edge. On 1,319 held-out task requests, the learned and RLVR studies report quality 0.9522 and 0.9553 versus 0.9484 for matched static; the RLVR-static paired difference is +0.0068 with a request-paired interval [0.0015,0.0122] and a training-seed-by-request hierarchical interval [0.0006,0.0131] . Controlled attribution recovery yields route mean absolute error 0.0011 and endpoint reconstruction error 0.0004. Factorial, bridge-order, and stochastic-provider studies evaluate the certificate interface; RLVR supplies a learned-policy stress test under the same identification contract.

[NLP-70] OPD Before RL: Warm-Starting Rubric-Based RL with On-Policy Distillation

【速读】: 该论文旨在解决生成式语言模型在开放性任务中缺乏精确结果验证机制的问题,尤其针对依赖评分标准(rubric)的评估场景。传统基于评分标准的强化学习(Rubric-based RL)虽能对完整输出进行评分,但因奖励信号仅在响应生成完毕后提供,无法有效定位影响最终得分的具体决策点,导致训练信号稀疏且难以优化。为此,论文提出一种两阶段训练框架:第一阶段为“基于评分标准的在线策略蒸馏”(RP-OPD),利用评分标准作为特权教师上下文,对模型在生成前缀时的逐标记分布进行密集监督,使无访问评分标准的弱学生模型能够模仿强教师模型的行为;第二阶段则引入基于评分标准的强化学习,直接优化最终评分奖励,实现性能突破。实验表明,该两阶段框架在HealthBench、ResearchQA和RubricHub Science等健康与科学领域任务上均优于现有方法,且在鲁棒性方面表现更佳——相较于纯监督微调+强化学习基线存在明显的“奖励黑客”现象(即模型仅表面符合评分标准而未提供实质内容),本方案在RubricHub Science任务中表现出有限的奖励操纵迹象。因此,研究支持先以评分标准引导在线策略蒸馏,再进行强化学习的分阶段策略,从而实现更高效、更可靠的模型优化。

链接: https://arxiv.org/abs/2610.02781
作者: Xinpeng Wang,Wei Shi,Yu-Chia Chen,Maria Zontak,Yun He,Richard Yuanzhe Pang
机构: New York University (纽约大学); Meta
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Many useful language-model tasks cannot be evaluated by exact outcome verification. Rubric-based reinforcement learning (RL) addresses this issue by scoring open-ended responses against explicit criteria. However, because the reward is assigned after the complete response, the training signal does not directly identify which individual decisions contributed to the final score. We propose a two-stage training framework that uses rubrics first as privileged teacher context for dense token-level supervision, then as rewards for further RL. In the first stage, rubric-privileged on-policy distillation (RP-OPD), a student without access to the rubric matches a rubric-aware teacher’s next-token distributions at student-generated prefixes. In the second stage, RL directly optimizes the rubric reward and improves beyond the observed distillation plateau. We evaluate the framework on health and science tasks using open-weight models. Across HealthBench, ResearchQA, and RubricHub Science, we compare post-training methods and vary the amount of SFT or RP-OPD training before RL, finding that our two-stage framework achieves the highest scores among the methods evaluated. RP-OPD + RL shows limited signs of reward hacking on RubricHub Science, whereas the SFT + RL baseline increasingly receives high rewards for claims of rubric compliance without providing the required content. These findings support using rubrics to guide on-policy distillation before applying rubric-based RL.

[NLP-71] Automatic Evaluation of Mental Health Stigma in Online Communication AACL

【速读】: 该论文旨在解决在线语境中精神健康污名(mental health stigma)复杂且多维度的自动评估难题。传统方法难以全面捕捉污名的多元表现形式,包括显性贬损、隐性责备、恐惧、居高临下的同情、社会疏离、结构性排斥及歧视等。为此,研究提出一个基于理论框架的基准测试(benchmark),涵盖自然生成的新闻与社交媒体文本,并采用细粒度污名分类体系进行标注,包含三个层级:污名模式(stigma mode)、作用领域(domain)以及特定形式污名的具体构成成分。该框架同时设计了二元污名检测任务与多层次标注体系,以系统化刻画污名表现。研究将此框架应用于六种精神健康状况相关文本,对比大型语言模型(LLM)与传统污名相关分类器在情感、毒性与仇恨言论检测中的表现。结果表明,仅依赖邻近概念(如毒性、仇恨言论)训练的模型对精神健康污名的识别能力有限,而LLM在缺乏明确操作规则的情况下常出现过度预测,凸显了人类标注中决策规则的重要性。因此,解决方案的关键在于引入结构化的细粒度标注体系与显式的操作定义,从而提升模型对精神健康污名的准确识别能力。

链接: https://arxiv.org/abs/2610.02775
作者: Naomi Baes,Jemima Kang,Nick Haslam,Chris Groot,Alsa Wu,Luc Raszewski,Yulia Otmakhova
机构: The University of Melbourne (墨尔本大学); University of Tübingen (图宾根大学)
类目: Computation and Language (cs.CL)
备注: AACL Main 2026

点击查看摘要

Abstract:Mental health stigma has profoundly harmful impacts but its complexity makes it difficult to evaluate. Stigma may involve explicit derogation, but also subtler forms of blame, fear, paternalistic pity, social distancing, structural exclusion, and discrimination. We introduce a theory-grounded benchmark for automatic evaluation of mental health stigma in online communication, consisting of naturally occurring online news and social media text annotated with a fine-grained taxonomy of stigma across multiple mental health conditions. Our annotation framework comprises a binary stigma-detection task and a multi-level taxonomy covering (i) stigma mode, (ii) domain, and (iii) specific components of certain forms of stigma. We apply this framework to texts mentioning six mental health conditions and evaluate large language models alongside stigma-related classifiers for detecting sentiment, toxicity, and hate speech. Results show that mental health stigma is not well captured by models trained to detect these neighboring constructs, and that LLMs often overpredict stigma unless given explicit operational rules - mirroring the importance of decision rules in human annotation. We release the publicly available part of benchmark, annotations, prototypical exemplars of stigma and code at: this https URL.

[NLP-72] Improving Atomic-Fact Recall via Focused Views in Unstructured Knowledge Editing

【速读】: 该论文旨在解决生成式AI在知识编辑(Knowledge Editing, KE)过程中存在的“上下文依赖”问题,即现有无结构知识编辑(Unstructured Knowledge Editing, UKE)方法在编辑后,模型虽能复现原始编辑文本,却无法独立、可靠地回忆其中的单个事实,严重依赖原始上下文。其核心问题是:标准的以段落为单位的编辑目标导致后期事实因累积了更丰富的上下文信息而被低估难度,从而在训练中获得更低的初始损失,使模型优先学习上下文而非独立事实。针对此问题,论文提出FOVEATED——一种即插即用的框架,通过在编辑阶段随机扰动旋转位置编码(Rotary Position Embedding, RoPE)中键(key)的位置,对每个句子构建聚焦视图(focused views),迫使模型关注局部语义而非全局上下文;该扰动仅在训练时应用,推理时恢复原生位置编码,不影响模型原有结构。理论分析表明,FOVEATED有效缓解了上下文诱导的难度低估问题;实验验证其在五种不同编辑器、两种大语言模型(LLM)骨干网络及三个基准数据集上均实现稳定性能提升,显著增强了模型对独立事实的记忆能力。

链接: https://arxiv.org/abs/2610.02772
作者: Ding Wu,Ye Zhang,Haoyu Wang,Tianci Liu
机构: 未知
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: The first two authors contributed equally

点击查看摘要

Abstract:Large language models (LLMs) increasingly serve as general-purpose interfaces to factual knowledge, but their parameters do not automatically reflect information that changes after pretraining. Knowledge editing (KE) provides a targeted alternative to costly retraining by modifying selected knowledge and preserving unrelated knowledge and general capabilities. Conventional KE uses structured factual triples, whereas unstructured KE (UKE) uses free-form passages containing multiple facts. Nonetheless, existing UKE editors exhibit a failure mode known as context reliance: edited LLMs can often reproduce the editing passage but fail to reliably recall its individual facts without the original passage context. We identify context-induced difficulty underestimation under the standard passage-level editing objective: later facts receive increasingly rich ground-truth context and consequently incur lower initial losses, making them appear easier to learn. In response, we propose FOVEATED, a plug-and-play framework that constructs focused views of each sentence by randomly shifting the Rotary Position Embedding (RoPE) positions assigned to the keys of its preceding context. The perturbation is applied during editing and removed afterward, leaving the model’s native positional encoding unchanged at inference time. We instantiate FOVEATED for both direct-optimization and locate-then-edit editors. We theoretically analyze how FOVEATED counteracts context-induced difficulty underestimation and empirically demonstrate consistent improvements across five KE editors, two LLM backbones, and three benchmarks.

[NLP-73] AptMQL-Bench: From Text-to-SQL to Text-to-MQL via Access-Pattern Schema Design and Data-Preserving Migration

【速读】: 该论文旨在解决现有文本到查询语言(text-to-MQL)基准数据集在将关系型数据库转换为文档数据库(如MongoDB)时所面临的严重缺陷问题。当前方法依赖启发式规则进行机械转换,导致数据库模式与原始关系结构直接映射,且查询语句未适配文档数据库的原生特性,从而引发数据丢失、查询性能下降及迁移失败等问题。其解决方案的关键在于提出一种由编码代理(coding agents)驱动、结合人工验证的新型转换流程:该流程基于预期的数据访问模式重新设计文档模式,并将原始查询重写为符合MongoDB特性的原生MQL(MongoDB Query Language)语句。通过该方法对BIRD数据集进行重构,构建了首个基于访问模式的文本到MQL基准(AptMQL-Bench),包含21个面向文档的数据库、3,186条自然语言请求及其对应的高质量MQL查询,实现了无数据损失的迁移并具备良好的可扩展性。实验表明,即便采用最先进的模型(Claude Opus 4.5),在无外部知识支持下准确率仅达57.38%,引入外部证据后提升至70.34%,揭示了真实场景下文本到MQL生成任务仍面临巨大挑战。

链接: https://arxiv.org/abs/2610.02770
作者: Hy Nguyen,Nabi Rezvani,Robin Vujanic
机构: The University of Sydney(悉尼大学); MongoDB Research( MongoDB 研究院)
类目: Computation and Language (cs.CL); Databases (cs.DB)
备注:

点击查看摘要

Abstract:Document databases such as MongoDB are core infrastructure for modern applications, and natural-language interfaces to them—text-to-MQL—would let non-experts query complex, semi-structured data without mastering the query language. Progress on this task depends on high-quality benchmarks, which are most practically obtained by converting an existing text-to-SQL benchmark to the document setting. Unfortunately, existing efforts rely on heuristics for mechanical conversion: the document schema mirrors the relational foreign-key graph, and each query mirrors its source SQL. As a result in our experiments, these approaches fail to migrate 6 of 21 BIRD databases outright, silently drop up to 25.9% of rows on others, and yield schemas whose ground-truth queries run over an order of magnitude slower as the data scales. We instead propose a conversion pipeline, driven by coding agents with human-in-the-loop verification, that designs each document schema from expected access patterns and rewrites queries to be MongoDB-native. Applying it to BIRD, we build an access-pattern-based text-to-MQL benchmark (AptMQL-Bench). It includes 21 document-oriented databases, 3,186 natural-language requests, and their associated MQL queries—whose databases are migrated from SQLite without data loss and scale efficiently. The strongest model, Claude Opus 4.5, achieves only 57.38% accuracy without external knowledge evidence and 70.34% with it. This indicates that realistic text-to-MQL generation remains challenging.

[NLP-74] When History Fails to Become Experience: Action Calibration in Language Agents

【速读】: 该论文旨在解决语言智能体(language agent)在执行任务时无法有效利用历史交互信息以优化后续决策的问题。尽管历史信息理论上有助于提升任务表现,但实际中额外引入历史数据有时反而会降低任务成功率,表明智能体未能可靠地将过去的行动与其结果进行关联。研究发现,即使打乱历史中的动作与观测的对应关系,任务完成率仍保持较高水平,说明智能体对历史的使用具有一定的鲁棒性,但其核心缺陷在于缺乏对动作-结果因果关系的准确感知。为此,作者提出一种关键解决方案:显式标注每个观测为前一动作的结果(outcome labeling),这一简单标注显著提升了任务成功率并减少了重复动作的发生,且未引入新的环境信息。在此基础上,进一步提出一种可学习的校准器(learned calibrator),通过显式重新评估过往动作并选择性地记录经验,以更有效地指导后续决策,从而在性能上超越单纯的结果标注方法。

链接: https://arxiv.org/abs/2610.02769
作者: Jingyu Liu,Zhiwen Wang,Yuxin Jing,Huanyu Zhou,Yong Liu
机构: Gaoling School of Artificial Intelligence, Renmin University of China (中国人民大学高瓴人工智能学院); ByteDance(字节跳动); Beijing Key Laboratory of Research on Large Models and Intelligent Governance (北京市大模型与智能治理研究重点实验室); Engineering Research Center of Next-Generation Intelligent Search and Recommendation, MOE (教育部下一代智能搜索与推荐工程研究中心)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Language agents should draw on prior attempts and environmental feedback to improve subsequent decisions within the same task. However, providing additional interaction history can sometimes reduce task success, suggesting that agents do not consistently use this information effectively. To investigate this limitation, we examine how agents use history. We find that history improves task completion overall, yet much of this benefit persists even when past actions are shuffled. Disrupting the correspondence between actions and observations causes only a modest decline in task success. We therefore hypothesize that agents do not reliably connect past actions with their outcomes when deciding how to proceed. To test this hypothesis, we explicitly label each returned observation as the outcome of the preceding action. This simple annotation improves task success and reduces next-action repetition without introducing new environmental information. Building on this insight, we introduce a learned calibrator that explicitly reassesses past actions and selectively records experience to guide subsequent decisions, improving task success beyond outcome labeling alone.

[NLP-75] EpiWorld: Grounding LLM Policy Agents in Epidemiological World Models EMNLP2026

【速读】: 该论文旨在解决传染病干预政策制定过程中,大语言模型(LLM)因缺乏流行病动力学机制、定量监测信号及制度约束而难以准确预测干预后果、评估疫情严重程度并生成可解释、可控制政策的问题。其解决方案的关键在于提出EpiWorld框架——一个闭环系统,将LLM政策决策者与一个基于学习的、动作条件化的流行病世界模型以及分层技能库(包含公共卫生规程、监测工具和基于事后分析积累的适应性经验)相耦合。通过该框架,世界模型能够快速模拟不同干预措施下的区域疫情演化路径,并提供反事实推演反馈以支持政策选择与优化;同时,模拟结果被提炼为可复用的经验教训,而协议约束保持不变,从而在不牺牲可解释性与可控性的前提下实现决策过程的持续改进。实证评估表明,该框架在新冠与流感历史数据上显著优于现有基线,尤其在减少累积住院人数方面表现突出。

链接: https://arxiv.org/abs/2610.02744
作者: Zeeshan Memon,Yiqi Su,Kai Shu,Naren Ramakrishnan,Liang Zhao
机构: Emory University (埃默里大学); Virginia Tech (弗吉尼亚理工学院)
类目: Computation and Language (cs.CL)
备注: Accepted to Findings of EMNLP 2026. 22 pages

点击查看摘要

Abstract:Epidemic intervention policies are textual artefacts that human decision-makers interpret, justify, and revise through natural language, making large language models a natural candidate for epidemic policy reasoning. A naive LLM, however, lacks the epidemic dynamics needed to project intervention consequences, the quantitative surveillance signals required to assess severity, and the institutional constraints that define admissible actions. We present EpiWorld, a closed-loop framework that grounds an LLM policy actor in a learned action-conditioned epidemiological world model and a tiered skill library of public-health protocols, surveillance tools, and adaptive lessons accumulated through after-action analysis. Given a candidate intervention, the world model predicts regional epidemic evolution and enables fast counterfactual rollouts that provide feedback for policy selection and refinement. Outcomes of simulated futures are distilled into reusable lessons while protocol constraints remain fixed, allowing the decision process to improve without sacrificing interpretability or controllability. We evaluate both the world model and the end-to-end framework on retrospective COVID-19 and Influenza datasets: the world model achieves the best out-of-distribution Peak-MAE among all forecasting baselines, and the closed-loop framework reduces cumulative hospitalisation by up to 59% across datasets and by an average of ~16% across six LLM backbones, outperforming reinforcement-learning and optimal-control policy baselines.

[NLP-76] Beyond Correctness: Resolving Underspecification in Agent ic Text-to-SQL

【速读】: 该论文旨在解决生成式文本转SQL(Text-to-SQL)代理在处理用户查询中未明确信息(underspecification)时存在的“沉默失败”问题,即系统虽能生成正确执行结果,但可能依赖未经验证的隐含假设而未能真正澄清关键歧义。其核心挑战在于:现有方法易因过早终止澄清流程而导致遗漏关键问题,且即使被提示进行规划,代理仍常放弃已识别的相关问题。解决方案的关键是提出PlanPool机制,将澄清计划显式化为一个可修改的待问问题池(mutable question pool),强制要求每个计划中的问题必须显式提交或放弃,同时支持交互过程中动态新增新发现的歧义。该方法显著提升了歧义覆盖度,减少了沉默失败率,在多个基于BIRD-Interact和Spider的基准测试中表现出优于无约束与提示引导型基线的方法,同时保持了较高的执行准确率。研究强调了在智能体推理中,仅识别缺失信息并不足够,必须可靠地维护并主动解决这些信息缺口,才能确保推理过程的鲁棒性。

链接: https://arxiv.org/abs/2610.02739
作者: Wen-Zhi Li,Yue Gong,Konstantinos Kanellis,Balakrishnan Murali Narayanaswamy
机构: Cornell University (康奈尔大学); Amazon Web Services (亚马逊网络服务)
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Agentic Text-to-SQL systems can interact with users to clarify underspecified queries before generating SQL. However, a correct execution result does not necessarily imply that the agent has adequately resolved the underlying underspecification: the agent may silently make unverified assumptions that happen to match the intended answer. We show that this behavior is driven in part by premature clarification termination. Although forcing an agent to ask more questions improves execution accuracy, ambiguities are concentrated in earlier interactions, making brute-force questioning inefficient. More importantly, even when explicitly prompted to plan its clarification process, the agent frequently abandons questions that it has already identified as relevant. To address this failure mode, we introduce PlanPool, which externalizes the clarification plan as a mutable question pool. Every planned question must be explicitly asked or dropped before submission, while newly discovered ambiguities can be added during interaction. Across three benchmarks derived from BIRD-Interact and Spider, PlanPool consistently improves ambiguity coverage and reduces silent failures over unconstrained and prompt-based alternatives, while maintaining competitive execution accuracy. Our results highlight an important distinction in agentic reasoning: identifying missing information is not sufficient, and the agent must also reliably maintain and resolve it before committing to an answer.

[NLP-77] PBench: A Turning-Point Benchmark for Dialogue Compression

【速读】: 该论文旨在解决对话压缩(dialogue compression)中“转折点遗失”(turning-point eviction, TPBench)这一关键问题,即在保留对话核心信息的同时,可能错误地删除了改变用户意图的关键对话轮次,导致当前状态(current value)与初始目标(initial goal)的不一致。传统评估方法仅使用单一的整体保留率(overall retention score)掩盖了这一问题,因其将用户初始意图与最新意图混合评价,无法揭示压缩方法在关键语义转变处的表现缺陷。为此,作者提出TPBench基准,通过三个互补的信息目标进行精细化评估:P1关注用户的初始目标,P2关注被修改槽位的当前值,P3则在晚期标注槽位更新的对话中同时评估初始目标与当前值。其中,当前值答案来自MultiWOZ与SGD的人工对话状态标注,初始目标为首轮用户话语首句,无需额外众包。实验表明,不同探针任务对压缩方法的排序结果差异显著;在保留率为0.30时,所有压缩方法均落后于完整上下文,尤其是删除携带更新信息的轮次会显著降低当前值准确率,而删除无关匹配轮次则无影响。此外,采用Mistral模型阅读器时,仍可复现相同排名模式与性能差距。在长对话数据集LongMemEval-KU和中文RiSAWOZ上的进一步验证表明,完整上下文始终最优,而基于时效性(recency)的压缩方法在压缩场景下表现最佳。因此,该研究的核心解决方案在于引入多维度、探针驱动的评估框架,以精准识别并缓解对话压缩中的语义断裂风险。

链接: https://arxiv.org/abs/2610.02736
作者: Minji Park,Seunghyun Yoon,Hyuk Lim
机构: Korea Institute of Energy Technology (KENTECH)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: Code and benchmark: this https URL

点击查看摘要

Abstract:A compressor can keep the facts of a dialogue and still drop the turn that changed them. A user corrects a price, reverses a choice, or adds a constraint. We call this failure turning-point eviction. One overall retention score hides it, because that score mixes what the user first wanted with what the user wants now. We introduce TPBench, which evaluates three complementary information targets at shared nominal retention budgets. P1 asks for the user’s initial goal. P2 asks for the current value of a slot the user revised. P3 asks for both, in dialogues with a late annotated slot update. The current-value answers come from the human dialogue-state annotations of MultiWOZ and SGD. The initial-goal answer is the first sentence of the first user turn. Neither requires new crowdsourcing. The probe-specific evaluations rank compression methods differently. On the joint probe at a retained fraction of 0.30, every tested compressed method remains below full context with the main Llama reader. Deleting the turn that carries the update sharply lowers current-value accuracy, while deleting one matched irrelevant turn leaves it unchanged. A Mistral reader repeats the P2/P3 rankings and the joint-probe gap. Current-value recovery is tested on an additional corpus, LongMemEval-KU, and on Chinese RiSAWOZ: full context has the highest accuracy, and recency has the highest compressed-method mean in both evaluations. Comments: Code and benchmark: this https URL Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI) Cite as: arXiv:2610.02736 [cs.CL] (or arXiv:2610.02736v1 [cs.CL] for this version) https://doi.org/10.48550/arXiv.2610.02736 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[NLP-78] WakeKV: Reactive Reversible KV Residency for Heads That Change Their Minds NEURIPS2026

【速读】: 该论文旨在解决现有键值缓存(KV-cache)压缩方法在生成过程中静态分类注意力头所带来的效率瓶颈问题。传统方法通常在预填充阶段或离线时对注意力头进行一次分类,并在整个生成过程保持固定,但实验发现,在不同模型(1.5B–8B)和多种生成场景(如针堆检索、长链推理与多轮回忆)下,多数注意力头的读取行为会在生成过程中发生动态变化。针对这一问题,论文提出WakeKV,一种响应式驻留策略(reactive residency policy),其核心创新在于:当检测到“冷却”(cooling)头时,不将其冻结或永久淘汰,而是将其状态移至可恢复的CPU缓存池中,从而实现更灵活的资源管理。该方案在相同内存预算下显著降低了缓存未命中率,优于静态分类与破坏性淘汰策略,并在五个模型-场景组合上验证了其优越性。基于FlexiCache/vLLM的Mistral-7B实测结果表明,该方法在真实硬件上提升了吞吐量,同时保持LongBench基准测试的生成质量。

链接: https://arxiv.org/abs/2610.02713
作者: Utkarsh Ranjan
机构: UC San Diego (加州大学圣地亚哥分校)
类目: Computation and Language (cs.CL); Hardware Architecture (cs.AR); Distributed, Parallel, and Cluster Computing (cs.DC); Machine Learning (cs.LG)
备注: Accepted to the NeurIPS 2026 Workshop on ML for Systems. 2 figures, 4 tables, appendix

点击查看摘要

Abstract:Most KV-cache compression methods classify attention heads once, either offline or during prefill, and keep this classification fixed throughout generation. Across three models (1.5B-8B) and three regimes (needle retrieval, long chain-of-thought, and multi-turn recall), we measure head behavior on four model-regime combinations and find that most heads change their reading behavior at least once during generation. We introduce WakeKV, a reactive residency policy that moves cooling heads to a recoverable CPU reservoir rather than freezing or permanently evicting their state. At matched memory or budget, WakeKV consistently improves miss rate over frozen classification and destructive eviction, evaluated across five model-regime combinations and over three cited baselines (SnapKV, uniform R-KV, and ReasonAlloc) across four eligible combinations. A FlexiCache/vLLM implementation on Mistral-7B confirms the benefit on real hardware, improving throughput while retaining LongBench quality.

[NLP-79] Learning from Evolving Errors: Adaptive Iterative Repair for On-Policy Distillation

【速读】: 该论文旨在解决基于策略自蒸馏(On-Policy Self-Distillation, OPSD)中因教师模型依赖完整参考解(reference solution)而产生的“解条件捷径风险”(solution-conditioned shortcut risk)问题,即教师虽能提供目标状态,却无法指导学生如何从当前错误逐步修正至正确解,导致训练信号缺乏对误差到修复路径的显式引导。其核心解决方案是提出一种自适应迭代修复框架AIR-OPD(Adaptive Iterative Repair for On-Policy Distillation),通过引入一个引导生成器(guidance generator),在每次失败响应后动态生成针对当前错误的修复指引(repair guidance),并结合固定教师模型以该指引作为特权上下文(privileged context),对学生的最新失败响应中与误差对齐的区域进行监督。该框架采用结果感知阶段加权机制(outcome-aware stage weighting),优先奖励早期成功修复的阶段,并对即时通过验证的重试予以更高信用。实验在DAPO-Math-17K数据集上训练,于AIME24、AIME25和HMMT25等数学推理基准及MMLU-Pro、GPQA等分布外测试集上评估,结果显示,无论使用自引导(self-guidance)或外部大模型引导(external guidance),AIR-OPD均在Qwen3-4B和Qwen3-8B模型上取得最优的数学推理性能,相比最强基线最高提升达3.6分,同时保持了基础模型在分布外任务上的性能稳定性。

链接: https://arxiv.org/abs/2610.02700
作者: Rui Li,Liyang He,Zheng Zhang,Zhenya Huang,Linbo Zhu,Qi Liu
机构: University of Science and Technology of China (中国科学技术大学); Nanyang Technological University (南洋理工大学)
类目: Machine Learning (cs.LG); Computation and Language (cs.CL)
备注: 21 pages, 3 figures

点击查看摘要

Abstract:On-policy self-distillation (OPSD) supplies dense token-level feedback on trajectories sampled from the student’s own policy, a richer training signal than the outcome-level rewards of reinforcement learning. This feedback comes from a teacher conditioned on a full reference solution unavailable to the student. The reference solution specifies the target but not how to move from the student’s current error toward it, creating a solution-conditioned shortcut risk. We introduce AIR-OPD, an adaptive iterative repair framework for on-policy distillation that provides error-to-repair supervision. Given a failed response, a guidance generator synthesizes repair guidance for the current error. The student samples an on-policy retry with this guidance. If the retry remains incorrect, the generator produces new repair guidance for the newly observed error. At each round, a fixed teacher receives the guidance as privileged context and supervises the student on an error-aligned region of its latest failed response. Outcome-aware stage weighting favors early repair stages and credits stages whose immediate retry passes verification. We train AIR-OPD on the DAPO-Math-17K dataset and evaluate on AIME24, AIME25, and HMMT25, alongside out-of-distribution tests on MMLU-Pro and GPQA. We examine two guidance sources, self-guidance from the current student policy and external guidance from a larger model. For both Qwen3-4B and Qwen3-8B, AIR-OPD attains the best mathematical-reasoning averages, improving over the strongest baseline by up to 3.6 points, while preserving base-model performance on the out-of-distribution benchmarks.

[NLP-80] Large language models exhibit unreliable updating of clinical judgment as patient evidence evolves

【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在临床推理中对患者病情演变证据的动态信念更新能力是否可靠的问题。其核心挑战在于:当患者临床数据随时间变化时,模型能否合理修正先前判断,避免因先验偏见或证据处理不对称而导致错误推断。解决方案的关键在于提出一种基于真实电子健康记录(Electronic Health Records)的纵向信念更新评估框架——证据验证的纵向更新(Evidence-Validated Longitudinal Update, EVLU),通过匹配重症监护轨迹数据量化模型在不同证据序列下的预测误差变化,并识别出两类关键失败模式:一是模型对恶化证据的反应强度显著高于等量改善证据,表现出明显的证据方向性偏差;二是模型对初始先验风险的敏感性过高,即使当前证据不变,先验从10%升至90%仍导致预测值平均偏移26.2个百分点,揭示了先验信念的强因果影响。研究进一步表明,常规提示工程无法有效恢复可靠的信念更新能力,而EVLU方法虽提升了更新可靠性,但以牺牲覆盖范围为代价,揭示了可靠性与覆盖率之间的权衡关系。该研究首次将纵向信念更新确立为衡量大语言模型临床可靠性的一个独立维度。

链接: https://arxiv.org/abs/2610.02684
作者: Min Zeng,Rui Zhang
机构: University of Minnesota(明尼苏达大学)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Large language models (LLMs) are increasingly explored for clinical reasoning, but whether they appropriately revise judgments as patient evidence evolves remains unclear. We evaluated longitudinal belief updating using matched intensive-care trajectories from electronic health records. Across diverse LLMs, conditioning on a preceding judgment more often increased than reduced prediction error when estimates changed, replicated for a second endpoint. Controlled interventions revealed two failure modes. First, with preceding assessment fixed, models responded more strongly to worsening than matched improving respiratory evidence; this asymmetry persisted after headroom normalization at moderate and strong evidence levels. Second, with current evidence fixed, increasing prior risk from 10% to 90% shifted estimates by 26.2 percentage points, demonstrating causal influence of prior model beliefs. Prompting did not restore reliable updating. Evidence-Validated Longitudinal Update (EVLU) identified fewer, more reliable revisions, revealing a reliability-coverage trade-off. These findings establish longitudinal belief updating as a distinct dimension of LLM reliability.

[NLP-81] LEAP: Learning Efficient Action Proposals For LLM Agents

【速读】: 该论文旨在解决大语言模型(LLM)智能体在推理过程中因逐步执行而带来的延迟问题,特别是针对动作推测(action speculation)阶段的效率瓶颈。现有方法通过使用外部大型模型作为“草稿器”(drafter)来预生成动作建议,再由目标模型验证,但存在权衡:大型草稿器虽预测准确但耗时长,小型草稿器虽快速却难以匹配目标模型决策。论文的核心问题是:决定动作推测端到端加速效果的关键因素是什么?为回答此问题,作者构建了一个用于推测轮次的延迟分析框架,该框架从收益与代价两方面进行评估——收益取决于草稿器对目标模型动作序列的预测精度以及任务可执行的步数,代价则来自草稿生成、等待目标模型验证及工具执行的时间开销。基于该框架,提出LEAP(Learning Efficient Action Proposals)方法,采用小型0.6B参数量的草稿器,并通过在目标模型的动作序列上进行训练使其具备高准确性,从而在保持任务成功率不变的前提下,实现高达60%的端到端墙钟时间加速。实验表明,该框架能解释大部分实测加速现象;此外,研究还证明草稿模型可通过在线训练无需预先收集轨迹即可达到离线训练性能,显著提升了方法在实际场景中的可部署性。

链接: https://arxiv.org/abs/2610.02670
作者: Zhen Xu,Qizheng Zhang,Gerry Wan,Shang Zhu,Ce Zhang
机构: University of Chicago (芝加哥大学); Stanford University (斯坦福大学); Together AI
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:LLM agents are known to be slow in rollouts. An agent completes a task one step at a time. At each step, it reasons and then chooses an action to execute. The next step and action cannot start until the previous one has finished. Speculative decoding accelerates the rollouts at the reason phase by drafting and verifying the inference tokens. Recent works have also started to apply similar ideas at the action phase. These works use off-the-shelf models, usually large, to draft action proposals for target model to verify. Large drafters match the target more often but take longer to propose, while small off-the-shelf models are fast but rarely make the same decision as the target. We ask a more general question: what determines the end-to-end speedup of action speculation? To answer it, we develop a latency framework for the speculative round. The framework compares what a round gains with what it costs. The gain depends on how well the drafter predicts the target and on how many steps the task can take before it ends. The cost comes from drafting, from waiting for target verification and from executing tools. Guided by the framework, we introduce LEAP (Learning Efficient Action Proposals) which keeps the drafter small and makes it accurate by training it on the target actions sequences. With a small 0.6B model, LEAP agrees with the target on most decisions and makes agents up to 60% faster in end-to-end wall clock time, with no systematic change in task success. Across various datasets, target models and draft models, the framework accounts for most of the measured speedups. We also show the draft model can be online trained with no prior trace collection and match the performance of offline training, making LEAP practical to deploy in the real world.

[NLP-82] Large Language Continuous Diffusion Models

【速读】: 该论文旨在解决离散扩散语言模型(dLMs)在快速并行解码中虽表现优异,但其非光滑、高维的表示空间导致推理轨迹难以控制,进而阻碍了推理过程中的路径引导与加速的问题。其核心解决方案是提出Sigma——首个大规模(3B/8B)连续扩散语言模型,基于可引导的低维常微分方程/随机微分方程(ODE/SDE)潜在轨迹构建。通过分块似然优化训练,Sigma在去噪高斯扰动的词元嵌入的同时,学习最优嵌入几何结构;为加速训练,采用自回归(AR)模型预训练权重进行热启动。推理阶段,识别出无分类器指导(classifier-free guidance)与得分温度(score temperature)对实现高保真推理与编程能力至关重要。在数学推理与代码生成等任务上,经过预训练的Sigma在标准基准(如GSM8K、Minerva、HumanEval、MBPP)上表现与先进离散模型相当,经监督微调后在复杂推理任务(如MATH-500、AIME)上亦具竞争力。此外,研究揭示连续dLMs具有独特结构性质:(i)嵌入空间的可引导性能有效调控质量-多样性权衡,显著提升pass@k性能;(ii)连续轨迹支持低数值求解步数(NFE)下的平滑退化及高效模型蒸馏。这些发现确立了连续扩散语言模型作为高效语言生成的新范式。

链接: https://arxiv.org/abs/2610.02665
作者: Zhihan Yang,Wei Guo,Jean-Marie Lemercier,Simon Welker,Yonggan Fu,Mohammad Mahdi Kamani,Sajad Norouzi,Julius Berner,Tomas Geffner,Karsten Kreis,Yongxin Chen,Molei Tao,John Thickstun,Pavlo Molchanov,Ante Jukić,Arash Vahdat,Morteza Mardani
机构: NVIDIA; Cornell University; Georgia Institute of Technology
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Despite the success of discrete diffusion language models (dLMs) for fast parallel decoding, their non-smooth, high-dimensional space hinders trajectory steering for reasoning and inference acceleration. To overcome this, we present Sigma, the first large-scale (3B/8B) continuous dLM built on steerable, low-dimensional ODE/SDE latent trajectories. Trained blockwise via likelihood optimization, Sigma jointly denoises Gaussian-corrupted token embeddings while learning an optimal embedding geometry. To accelerate training, Sigma leverages pre-trained weights from autoregressive (AR) models for warm-starting. During inference, we identify classifier-free guidance and score temperature as essential for high-fidelity reasoning and coding. Across comprehensive math reasoning and coding evaluations against state-of-the-art discrete counterparts (masked dLMs and AR baselines), Sigma achieves competitive performance with discrete models on standard benchmarks (e.g., GSM8K, Minerva, HumanEval, MBPP) after pre-training and on challenging reasoning tasks (e.g., MATH-500, AIME) after supervised fine-tuning. Beyond performance parity, we uncover key structural properties unique to continuous dLMs: (i) embedding-space steering effectively governs the quality-diversity trade-off, yielding strong pass@k performance and (ii) continuous trajectories enable graceful degradation for low NFEs and efficient distillation. These establish continuous dLMs as a promising paradigm for efficient language generation.

[NLP-83] VERSE: Verified Self-Evolving Optimizer for Agent Harnesses

【速读】: 该论文旨在解决大语言模型(Large Language Model, LLM)智能体在优化其提示词(prompts)、工具(tools)和工作流(workflow)时,优化器自身诊断失败、生成修正方案及验证效果的能力受限的问题。传统方法中,优化器的工具与流程通常固定不变,难以实现深层次的自我改进。本文提出的关键解决方案是引入一种可验证的自演化优化器(VERSE),其核心在于使优化器不仅能够优化执行器的智能体架构,还能动态进化自身的故障诊断能力、编辑策略与验证机制。通过在执行过程中进行基于实际运行结果的验证(execution-based verification),VERSE支持优化器对草稿修改进行测试、重放失败案例、扰动可疑步骤,并追踪修复与回归情况。在此反馈闭环下,优化器可迭代更新执行器的提示词、技能、工具、钩子(hooks)及注释,而保持模型权重不变。实验表明,在跨五种语言的保留任务与分布外任务上,VERSE显著优于基线方法,最佳验证选择的智能体分别达到42.3%和37.7%的准确率,超越最强基线的39.2%和29.3%。

链接: https://arxiv.org/abs/2610.02616
作者: Zekai Wang,Yingqiang Ge,Zekun Wang,Hai Wang,Yuhui Xu,Joshua Frandsen,Shancong Fu,Ashia C. Wilson,Chandan K. Reddy
机构: MIT(麻省理工学院); Amazon(亚马逊)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
备注: 45 pages, 13 figures, 15 tables

点击查看摘要

Abstract:Harness evolution improves an LLM agent’s prompts, tools, and workflow, while the optimizer’s own tools and procedures often remain fixed. We study whether an optimizer can improve another agent more effectively by also improving how it diagnoses failures, develops edits, and tests their effects. Two observations guide our design. In a controlled study, optimizer self-evolution fails to improve performance without execution-based verification, but achieves the best result of that study when verification is available. Across five executors, self-evolving optimizers build their own tools for failure analysis, verification, training audits, and workflow control. Motivated by these findings, we introduce VERSE, a Verified Self-Evolving optimizer for agent harnesses. VERSE lets the optimizer test draft edits, replay failures, and perturb suspected steps before submission, while tracking fixes and regressions across rounds. Using this feedback, the optimizer revises both the executor harness and its own prompts, skills, tools, hooks, and notes, while the weights of the optimizer and executor models stay fixed. Under a shared protocol with disjoint training, validation, and test tasks, VERSE improves all four evaluated harness optimizers on held-out SWE-rebench tasks and newer out-of-distribution tasks in five languages. Its best validation-selected harness reaches 42.3% and 37.7% accuracy, respectively, against 39.2% and 29.3% for the strongest baselines. Code is available at this https URL.

[NLP-84] Learning When to Commit from Partial Speech for End-to-End Simultaneous Speech Translation

【速读】: 该论文旨在解决同步语音翻译(Simultaneous Speech Translation, SST)中如何在源语音尚未完全输入时实时生成高质量目标文本,同时确保已输出的词元(token)不被回退或修改这一核心挑战。其关键解决方案是采用前缀监督(prefix supervision)的全句语音语言模型自适应方法,通过利用模型自身对完整与部分语音波形的翻译结果来构建训练信号,从而无需依赖人工转录或人工翻译数据。该方法结合多轮追加式解码(multi-turn append-only decoding)与置信度阈值控制推理阶段的质量-延迟权衡,并通过调整合成前缀密度(synthesis margin)以调节训练过程中的前缀分布。实验表明,多轮训练显著提升了早期前缀的提交校准误差(commit-calibration error),较单轮训练提升63%-68%(整体)和68%-80%(早期前缀),且置信度机制提供了最广泛且稳定的性能操作范围;此外,适度的小合成边际可进一步拓展低延迟下的性能边界,尤其在短语场景下表现更优,而过大的边际则导致质量与校准性能下降,揭示了合成密度带来的非单调质量-延迟权衡。

链接: https://arxiv.org/abs/2610.02612
作者: Hieu Hoang,Amittai Axelrod
机构: Microsoft(微软)
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Simultaneous speech translation must emit useful target text before the source is complete while preserving every committed token. We adapt a full-utterance speech language model using prefix supervision derived from its own complete- and partial-waveform translations, requiring neither transcripts nor human translations. We compare single-turn forced-prefix and multi-turn append-only decoding, use a confidence threshold to control the inference-time quality–latency trade-off, and vary the density of training prefixes with a separate synthesis margin. On FLEURS and CoVoST2 in three language directions, prefix training improves quality–latency frontiers over the unadapted model, and confidence provides the broadest consistently competitive operating range. Multi-turn decoding is generally stronger at low latency; under multi-turn training, commit-calibration error falls by 63–68% overall and 68–80% at early prefixes, whereas single-turn training provides only modest overall calibration gains and no early-prefix improvement. A small synthesis margin sometimes extends the frontier to lower latency, particularly on shorter utterances, while a larger margin degrades translation quality and calibration. Prefix adaptation therefore improves simultaneous speech translation, especially under multi-turn append-only decoding, while synthesis density introduces a non-monotonic quality–latency trade-off.

[NLP-85] How Causality Bridges the Semantic Gap

【速读】: 该论文旨在解决系统测量数据中变量语义缺失的问题,即大量变量未被标注或从未被观测,而现有方法依赖人类通用知识进行语义赋值,不可避免地引入偏见且在知识空白处失效。其核心解决方案是利用因果结构作为语义映射的桥梁,通过变量间在因果图中的依赖关系来推断其语义,提出“结构约束语义对齐”(structure-constrained semantic alignment)框架:在已知少数变量名称作为锚点的前提下,基于因果图所隐含的依赖关系求解未知变量的嵌入表示。为此,研究构建了CausalBridge框架,能够从测量数据中自动发现包含潜在变量的因果图,据此求解变量嵌入,并通过语言模型将其转化为可解释的名称。该因果结构仅由观测数据推导而来,不依赖人类先验知识,因而具备去偏特性。实验在五个问卷调查和三个机器人场景中验证,当变量名称遮蔽率达20%至90%时,CausalBridge在恢复可观测与潜在变量语义方面显著优于依赖相关性的现有方法,且随着系统文档程度降低,优势愈加明显;其发现的因果图在命名准确性上媲美人工标注版本,新系统的命名可在分钟级完成,成本仅为采样方法的极小部分。该方法成功弥合了测量与语义之间的鸿沟,使机器能够以因果方式理解世界并采取行动。

链接: https://arxiv.org/abs/2610.02594
作者: Shuhao Zhang,Xuran Zhou,Han Guo,Pengtao Xie,Yujia Zheng
机构: University of California, San Diego(加州大学圣地亚哥分校); University of Illinois Urbana-Champaign(伊利诺伊大学香槟分校)
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Numerical measurements capture how a system behaves, but often leave the meanings of its variables unspecified. Some variables are measured but never labeled, and others are never measured at all. Existing methods assign semantics to such variables by consulting general human knowledge, but this inherits its biases where that knowledge exists and offers nothing where it does not. We bridge this gap between measurements and their meanings with causal structure instead, reading a variable’s semantics from how it acts on other variables. We formalize this as structure-constrained semantic alignment, in which the embedding of each unnamed variable is solved under the dependence relations implied by the causal graph, with the embeddings of a few known names as anchors. Accordingly, we build CausalBridge, a framework that discovers the causal graph from the measurements, latent variables included, solves for the embeddings under those relations, and expresses them as names through a language model. The causal structure reflects the mechanism that generated the measurements and is recovered from the measurements alone, which may make it the one source of information free of bias from human knowledge. We evaluate CausalBridge on five questionnaires and three robotics scenarios, with 20 to 90% of the variable names masked. It recovers the semantics of observed and latent variables more accurately than existing methods that rely on association, and its lead widens as less of the system is documented. The graph it discovers names variables as accurately as the documented one, and a new system is named in minutes and at a fraction of the cost of sampling methods. Once the semantic gap is bridged faithfully, machines can understand the world and take actions causally.

[NLP-86] Evaluating Multi-Dimensional Generalization of Large Language Models in Temporal Extraction Tasks AACL

【速读】: 该论文旨在解决时间与事件表达抽取任务中因标注模糊性、领域敏感性及模型行为不稳定所导致的泛化能力不足问题。现有评估方法主要关注模型在特定领域内的表现,难以揭示其在分布偏移下的可靠性。研究通过在四个泛化维度上系统评估多种模型配置(包括模型家族、架构及推理策略),考察了基础性能迁移、跨维度相关性以及规模、架构与提示策略的影响。关键发现表明:虽然较强的基线任务性能通常预示更好的泛化能力,但在显著的分布偏移下该关联性减弱;归纳式提示(inductive prompting)在领域转移、对抗扰动、组合性变化和输入长度增加等场景下表现出最一致的泛化性能,而规模、架构及演绎/归纳提示策略的增益则呈现不均衡且维度依赖的特点。研究结论强调,大语言模型(LLM)在时间与事件表达抽取任务中的泛化能力无法仅凭单一维度或域内评估可靠预测,必须依赖能够跨维度泛化的推理策略。

链接: https://arxiv.org/abs/2610.02549
作者: Fahmid Shahriar Iqbal,Ritam Dutt,Soumitra Das,Arnav Verma,Sagnik Ray Choudhury
机构: University of North Texas(北德克萨斯大学); Carnegie Mellon University(卡内基梅隆大学); University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校); New York University(纽约大学)
类目: Computation and Language (cs.CL)
备注: accepted AACL-IJCNLP 2026 Findings

点击查看摘要

Abstract:Time and event expression extraction are fundamental temporal reasoning tasks, but the problem remains difficult due to annotation ambiguity, domain sensitivity, and unstable model behavior. Existing evaluations focus on in-domain performance, offering limited insight into reliability under distribution shifts. We evaluate multiple model configurations across families, architectures, and reasoning strategies over four dimensions of generalization, examining transfer from base performance, cross-dimensional correlations, and the effects of scale, architecture, and prompting. This provides a systematic study of how prompted LLMs generalize in time and event expression extraction tasks. We find that strong base-task performance generally predicts better generalization. However, this relationship weakens under substantial distribution shifts. Inductive prompting performs most consistently across domain shift, adversarial perturbations, compositionality, and length increase, while gains from scale, architecture, and deductive and abductive prompting strategies are uneven and dimension-specific. We conclude that LLM generalization in temporal extraction tasks cannot be predicted from any single dimension alone and cannot be reliably inferred from in-domain or single-dimension evaluations, highlighting the need for reasoning strategies that generalize across dimensions.

[NLP-87] A generative-informed neuro-symbolic framework for syntactic ambiguity resolution: Evidence from Arabic DPs

【速读】: 该论文旨在解决现代标准阿拉伯语(Modern Standard Arabic, MSA)中名词短语(DP)的结构歧义问题,尤其针对形态丰富的名词构式在表面序列相同的情况下存在多种语法结构解释的挑战。其核心解决方案是提出一种生成式启发的神经符号框架,将形式语法中的生成性概念与AraBERT相结合,通过将歧义建模为基于候选项的决策任务,显式构建并评估语言学上合理的结构替代方案。该方法利用候选条件输入表示对不同结构假设进行区分和筛选,实现了对语法结构的可控且可解释的建模。实验结果表明,模型在未见测试集上达到96.88%准确率、95.92%宏平均F1、96.83%加权F1以及93.94%二分类F1,其中对高阶/动词短语附着(N1)的召回率达99.71%,而对低阶/名词短语/嵌套附着(N2)的召回率为89.26%,反映出嵌套结构恢复更具挑战性。研究结论表明,形式化句法表征可在基于Transformer的自然语言处理中作为语言结构与上下文神经建模之间的显式接口,为阿拉伯语乃至更广泛语言的句法歧义消解提供了可控、可解释的新范式。

链接: https://arxiv.org/abs/2610.02529
作者: Mohammed Damom,Muneef Y. Alshawsh,Ashraf A. Naji,Mustafa Ali Alhamzi,Fawwaz An-Nashef,Jameel Ahmed Elayah,Mohammed Q. Shormani,Noman AL-Sayadi
机构: 未知
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Syntactic ambiguity poses a persistent challenge for Arabic NLP, particularly in morphologically rich nominal constructions where multiple structu6ral interpretations may be compatible with the same surface sequence. This study proposes a generatively informed neuro-symbolic framework for resolving structural ambiguity in Modern Standard Arabic (MSA) DPs. The framework integrates generative syntactic notions with AraBERT by representing ambiguity as a candidate-based decision task in which linguistically motivated alternatives are explicitly constructed and evaluated through candidate-conditioned input representations. Findings indicate that the model achieved 96.88% accuracy, 95.92% macro-F1, 96.83% weighted F1, and 93.94% binary F1 on the unseen evaluation set. Class-level analysis revealed asymmetric performance, with recall of 99.71% for High/VP Attachment (N1) and 89.26% for Low/NP/Embedded Attachment (N2), indicating greater difficulty in recovering the embedded interpretation. The study concludes that formal syntactic representations can be operationalized within Transformer-based NLP as an explicit interface between linguistic structure and contextual neural modeling, providing a controlled and interpretable approach to Arabic syntactic ambiguity resolution and beyond.

[NLP-88] Right Order Wrong Scale: Auditing LLM Judges for Occupational AI Measurement

【速读】: 该论文旨在解决当前大语言模型(LLM)判官在评估人工智能输出是否符合职场要求时,虽能在响应排序上达成较高一致性,却无法准确反映真实工作者的可接受率及职业层面聚合指标的问题。其核心挑战在于:现有判官系统在排序上的表现(如成对准确性)并不能保证其对“可接受比例”的估计与实际劳动者评价一致。论文提出O*NET-BENCH这一审计基准套件,基于45,796名工人的实际评分数据,对6个模型家族中的33种判官配置进行评估。结果显示,尽管25种配置达到至少0.60的置信对齐成对准确性,但其估计的可接受率范围从3.0%至97.9%,远偏离职业匹配工人所给出的61.1%基准值。进一步分析表明,将评分策略从点对点打分改为捆绑式少样本/列表式协议虽提升了排序性能,却降低了与工人平均评分的一致性,且该现象在任务与工人均不重叠的验证集上重现。通过交叉验证校准可消除均值偏差,但校准后得分仅能解释个体评分方差的8.5%以内;基于预测的辅助估计也未能在有限标签预算下带来显著精度提升。因此,解决方案的关键在于:必须将判官系统的有效性验证从单纯的排序一致性扩展到对其预测的可接受率和职业聚合指标的准确性评估,否则将导致对AI系统在真实工作场景中表现的严重误判。

链接: https://arxiv.org/abs/2610.02492
作者: Harry Lyu,Neil Thompson
机构: Massachusetts Institute of Technology (麻省理工学院)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computers and Society (cs.CY); Machine Learning (cs.LG); General Economics (econ.GN)
备注:

点击查看摘要

Abstract:LLM judges are increasingly used to assess whether AI outputs meet workplace requirements, but agreement on response rankings does not establish agreement on acceptance rates or occupational aggregates. We introduce O*NET-BENCH, an audit suite derived from an existing survey of 45,796 worker ratings, and evaluate 33 pre-existing judge configurations across six model families on 4,501 test ratings. Twenty-five configurations achieve tie-aware pair accuracy of at least 0.60, although a train-fitted response-only TF-IDF baseline nearly matches the strongest judge. Despite this ordering agreement, judges estimate that 3.0%-97.9% of responses are acceptable, compared with 61.1% for occupation-matched workers. In one fine-tuned lineage, changing from pointwise scoring to a bundled few-shot/listwise protocol improves response ordering while reducing agreement with worker means at the task and occupation levels; this reversal replicates on a task- and worker-disjoint validation split under prespecified criteria. Cross-validated calibration largely removes mean bias, but calibrated scores explain at most 8.5% of individual worker-rating variance. Prediction-assisted estimation yields at most small precision gains at the studied label budgets. These results show that ranking agreement alone is insufficient for occupational measurement. Judges should be validated against the acceptance rates and aggregates their scores will be used to estimate.

[NLP-89] From Retrieval to Typed Decisions: Calibrated System One Models from Biomedical Sentence Encoders

【速读】: 该论文旨在解决生物医学文本中类型约束型决策(typed decision)任务的模型初始化与训练优化问题,具体关注预训练的生物医学句向量编码器(Sentence-Transformers)在构建高效、可解释的决策模型中的适用性。其核心挑战在于如何将仅针对信息检索任务训练的双编码器(bi-encoder)架构有效转化为支持多头结构(cross-head, C;先验融合残差, PFR)的类型化决策模型,并评估不同训练目标对性能的影响。解决方案的关键在于提出SBERT2S1框架,该框架能够将现有的句向量编码器无缝转换为具备跨头和先验融合机制的决策模型,并引入一个大规模标注数据集MEDLINE-S1(包含24.3万条来自NLM索引的训练决策),以及BIODECIDE基准套件,用于系统评估。研究发现,检索训练有助于零样本匹配内容相关选项,但对不同模型头的影响差异显著:先验融合残差(PFR)在多数情况下受益于检索训练,而交叉头(C)则表现不稳定,甚至在部分场景下被损害。进一步分析表明,当前主流的强化学习课程设计(RLCD)在开放系统一(System One)模型中表现低于交叉熵损失2.5–3.0个百分点,其根本原因在于奖励归一化导致噪声得分函数项被放大3.6–15倍;采用无偏留一法估计器可显著缩小这一差距。经过温度缩放后,各类训练目标在校准性上趋于一致,交叉熵仍为最稳健的选择。

链接: https://arxiv.org/abs/2610.02486
作者: Pritam Deka
机构: Queen’s University Belfast(贝尔法斯特女王大学); Belfast, United Kingdom
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Typed decision models answer schema-constrained questions about a text in one forward pass and return probabilities meant to be thresholded. We ask whether biomedical sentence encoders trained for retrieval are good starting points for such models. We present SBERT2S1, which converts Sentence-Transformers encoders into bi-encoder, cross-head © and prior-fused residual (PFR) decision models, together with BIODECIDE, a biomedical typed-decision suite, and MEDLINE-S1, 243k training decisions derived from NLM indexing. Across six parent-retriever pairs, retrieval training improves zero-shot matching of content-bearing options. After fine-tuning, its effect depends on the head: across five pairs and three training-set sizes, retrieval training significantly helps PFR, which keeps the retrieval prior, in 10 of 15 comparisons, but helps C in one and hurts it in five. A matched grid of two heads and five training objectives shows that C outperforms PFR under every objective, and that the released RLCD recipe of open System One models trails cross-entropy by 2.5-3.0 points. The deficit stems mainly from its reward normalisation, which inflates the noisy score-function term 3.6-15-fold; an unbiased leave-one-out estimator recovers most of the gap. After temperature scaling, no objective is clearly better calibrated than cross-entropy. We release the code, the MEDLINE-S1 labels and a model.

[NLP-90] APDMem: Agent -Controlled Progressive Disclosure for Query-Adaptive Long-Term Memory EMNLP2026

【速读】: 该论文旨在解决个性化大语言模型(LLM)助手在处理长对话历史时,如何高效从稀疏且复杂的多轮对话中检索所需证据的问题。现有方法通常依赖扁平化的记忆存储或固定的检索粒度,难以在计算成本与信息精度之间实现动态平衡。其解决方案的关键在于提出一种分层式长期记忆架构——代理控制的渐进式披露记忆(APDMem),该架构将对话历史组织为四个由粗到细的层次:主题摘要、个性化关键事实、回合级证据笔记和原始消息。在推理阶段,由控制器驱动渐进式披露机制:先访问高层摘要,仅在必要时逐步深入至更细粒度的证据。这一设计实现了自适应的成本-保真度权衡——简单查询可提前终止,而涉及时间线、多跳推理或精确证据的复杂查询则触发深层检索。此外,通过引入笔记合成器,系统能够将检索到的证据转化为聚焦于当前查询的结构化表示,整合事实、排序事件并标记矛盾,从而提升最终答案生成的质量。实验结果表明,在LongMemEval基准上,APDMem在保持优异长上下文推理性能的同时,仅需访问总对话内容的8%。

链接: https://arxiv.org/abs/2610.02472
作者: Chin-Lun Fu,Anagha Kulkarni,Hong Ni,Behrouz Madahian
机构: JPMorgan Chase Co.(摩根大通公司)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: Accepted at EMNLP 2026 (Industry Track)

点击查看摘要

Abstract:Personalized LLM assistants must recover sparse evidence from long conversation histories across queries of varying complexity. We introduce APDMem (Agent-controlled Progressive Disclosure Memory), a hierarchical long-term memory architecture that applies progressive disclosure to memory retrieval. Rather than relying on a flat memory store or fixed retrieval granularity, APDMem represents conversation history as four progressively detailed layers: thematic summaries, personalized key facts, turn-level evidence notes, and raw messages. At inference time, a controller applies progressive disclosure to the memory hierarchy: it first reads high-level summaries and drills into finer evidence only when needed. This creates an adaptive cost-fidelity trade-off: simple queries can terminate early, while complex temporal, multi-hop, or exact-evidence queries trigger deeper inspection. A note synthesizer converts retrieved evidence into a query-focused structure that consolidates facts, orders events, and flags contradictions before final answer generation. Experiments on LongMemEval show that APDMem achieves strong performance for long-context memory reasoning while accessing only 8% of the total conversations.

[NLP-91] Capability Scaling-Down Laws for LLM Compression

【速读】: 该论文旨在解决大语言模型(Large Language Model, LLM)压缩过程中方法与配置选择缺乏理论指导、依赖经验试错的问题,即在实现相似资源节约的情况下,不同压缩手段可能导致差异显著的能力损失。其核心解决方案是系统性地建立并验证适用于剪枝(pruning)、量化(quantization)和知识蒸馏(distillation)三类压缩技术的“能力衰减规律”(capability scaling-down laws),通过量化数学推理、代码生成和问答任务中的能力损失,并将其与模型规模、训练阶段、压缩设置、数据可用性及训练暴露程度等关键因素关联,构建可预测的简化关系。研究发现,共享剪枝密度响应可将拟合剪枝预测器所需的配置测量次数减半,在新Pythia模型状态、预注册OLMo-2测试状态及Wanda剪枝条件下,仅用部分测量即可达到与全量数据回归相当的精度(误差小于0.020 nats/token),且通过系数微调具备良好泛化能力。此外,受控蒸馏实验揭示了重用训练数据在问答分布上的成本代价,而净收益取决于评估分布。最终,基于预测结果进行配置选择,相较于使用中位数或固定方法优先级,在两个模型族上均能捕获大部分跨方法优势,尤其在问答任务中表现突出;而在数学与代码生成任务中,固定方法优先级已接近最优,表明能力衰减规律的预测价值具有任务依赖性。该工作明确了能力衰减规律的适用边界及其在压缩策略决策中的实际价值。

链接: https://arxiv.org/abs/2610.02462
作者: Xueqi Cheng,Liang Wu,Kelly Wan,Liangjie Hong,Yushun Dong
机构: Florida State University (佛罗里达州立大学); Nokia Applied Research (诺基亚应用研究)
类目: Machine Learning (cs.LG); Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:LLM compression reduces inference costs and memory requirements, but selecting a method and configuration remains largely empirical because comparable resource reductions can produce different capability losses. We systematically investigate capability scaling-down laws for LLM compression across pruning, quantization, and distillation. Our framework measures capability loss in mathematics, code generation, and question answering, and relates these measurements to model size, training stage, compression settings, data availability, and training exposure. We develop simple predictive relations and evaluate their accuracy, measurement efficiency, and generalization to unseen configurations and model states. Sharing the density response across pruning levels halves the configuration measurements needed to fit a pruning predictor: on new Pythia states, on pre-registered OLMo-2 test states and under Wanda pruning, the compact relation matches a regression fitted with all measurements on math and code to within 0.020 nats per token, with coefficients refitted for each setting. Controlled distillation experiments show that the cost of heavy data reuse recurs across question-answering distributions, while the net benefit depends on the evaluation distribution. We further evaluate the decision value of these predictions by comparing numerical selection with configuration medians and fixed method priorities. Independent evaluations across two model families show that selection captures most of the available cross-method benefit for question answering within the tested candidate sets, where a fixed method priority attains the same regret, with smaller opportunities for mathematics and code. These results clarify the predictive scope of capability scaling-down laws and their use in compression method selection. Our code is publicly available at: this https URL.

[NLP-92] CUEing User Simulators: Calibrated User Embeddings for Multi-Turn Benchmarking

【速读】: 该论文旨在解决现有用户模拟器在多轮交互评估中缺乏结果校准(outcome calibration)的问题,即模拟用户与真实用户在任务成功率和失败模式上的一致性不足。尽管已有模拟技术在行为风格上具备表面真实性,但其生成的失败模式与真实用户表现存在偏差,导致对智能体性能的评估失真。解决方案的关键在于提出校准式用户嵌入(Calibrated User Embeddings, CUE)框架:该框架通过编码真实交互会话,学习连续的用户表征,并解码为可引导大语言模型(LLM)扮演特定用户角色的“人格指令”(persona commands),从而实现无需微调即可生成高保真的用户模拟。实验表明,基于CUE的模拟器在τ²-Bench基准上显著减少了由模拟器自身引入的错误,更准确地复现了真实用户与智能体交互中的失败模式、整体成功率及特定任务-用户组合的结果分布,同时保持了与先前方法相当的用户行为拟真度。此外,经客服对话数据训练的CUE模型可泛化至文档创作、数学辅导和闲聊等多样化任务场景,并在不同基础LLM上保持有效性,无需重新训练CUE模块。

链接: https://arxiv.org/abs/2610.02460
作者: Anjali Kantharuban,Jonas Mueller
机构: Handshake AI(握手人工智能); Carnegie Mellon University(卡内基梅隆大学)
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Recent benchmarks rely on user simulators to evaluate AI agents in multi-turn interaction. While existing simulation techniques demonstrate surface fidelity to human style and behavior, ecologically valid interactive benchmarking also requires alignment in when and how agents fail across simulated and real user populations. We find that existing simulators lack outcome calibration: agreement with observed success rates and failure patterns when real users interact with the same agent. We introduce Calibrated User Embeddings (CUE), a framework that both encodes observed sessions and samples continuous representations, then decodes them into persona commands to steer LLMs to act as user simulators without training. Through this, we evaluate user-conditioned replay of past sessions and aggregate metric agreement when sampling novel personas for the same tasks. On \tau^2 -Bench, CUEd simulators commit fewer simulator-attributed errors and more faithfully reproduce real-user agent failure modes, aggregate success rates, and outcomes for specific task-user pairs than other persona-based simulation methods. These gains coexist with competitive user fidelity as measured using metrics established in prior work. After being fit to mostly customer support interactions, the same CUEd simulators generalize to document creation, math tutoring, and casual conversation, and remain effective across different simulator LLMs without CUE retraining.

[NLP-93] FinDialogLens: Event Extraction over Multi-Party Dialogue for Missed-Trade Identification in Financial Chatrooms EMNLP2026

【速读】: 该论文旨在解决多参与方金融对话中因信息分散与事件交错导致的错失交易(missed trades)自动恢复难题,核心挑战在于每条询价请求(Request for Quote, RFQ)的最终成交价格与交易结果通常在多个消息之后才出现,且被其他参与者的并发RFQ所干扰,使得传统方法难以有效定位关键信息。其解决方案的关键在于提出一种混合式大语言模型(LLM)流水线框架——FinDialogLens,通过轻量级微调分类器作为推理时的结构化引导(inference-time scaffolds),实现对RFQ触发事件及价格/交易结果元数据的精准检测;结合基于事件级别的窗口分割模块(RFQ-Level Module)与交易角色填充引擎(Trade Engine),系统性地重构事件上下文。该架构在GPT-4o上实现了92.1%的最终价格准确率和94.3%的交易结果准确率,显著优于全对话链式思维(Chain-of-Thought, CoT)提示方法;同时,仅需30亿参数的开源微调模型在少量领域内数据下即可达到相近性能。为实现大规模实用化,引入难度感知路由机制(difficulty-aware router),动态将低复杂度RFQ分配至低成本规则引擎,高难度任务交由高性能的LLM驱动交易引擎处理,使大模型调用次数减少85%,在维持半数准确率差距的同时日均节省成本超300美元(在每日7万条RFQ规模下)。

链接: https://arxiv.org/abs/2610.02455
作者: Chin-Lun Fu,Hong Ni,Behrouz Madahian
机构: Machine Learning Center of Excellence, JPMorgan Chase Co.(摩根大通公司机器学习卓越中心)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: Accepted at EMNLP 2026 (Industry Track)

点击查看摘要

Abstract:Multi-party financial chatrooms are vital for sales-and-trading professionals, but their complexity makes manual recovery of missed trades infeasible: each Request for Quote (RFQ) is an event whose final price and trade outcome appear many messages after the RFQ-trigger message (the inquiry message), interleaved with concurrent RFQs from other participants. We cast this as event extraction (EE) over multi-party dialogue and present FinDialogLens, a hybrid LLM pipeline in which compact fine-tuned classifiers act as inference-time scaffolds: they detect RFQ-triggers and price/trade outcome metadata, an RFQ-Level Module segments per-event RFQ windows, and a Trade Engine fills argument roles. With GPT-4o, FinDialogLens reaches 92.1% and 94.3% accuracy on final price and trade outcome, respectively, outperforming full-chatroom CoT prompting methods; fine-tuned open-source LLMs with as few as 3B parameters achieve comparable performance with modest in-domain data. To make the LLM-based solution practical at scale, a difficulty-aware router balances cost and accuracy by allocating RFQs between a low-cost rule-based engine and the higher-performing LLM-powered Trade Engine, cutting LLM calls by 85% on final price while recovering half of the accuracy gap to FinDialogLens (GPT-4o), saving over 300/day at our 70,000-RFQ/day scale.

[NLP-94] Counterexample Generation via Per-Theorem Symbolic Verifiers: When Imitation Hurts and Reinforcement Repairs EMNLP2026

【速读】: 该论文旨在解决大语言模型在数学定理证明中存在的一种“伪证缺口”(falsification gap)问题,即模型能够正向推导正确定理,却无法有效生成反例来反驳形式上相近但错误的命题。其核心挑战在于:传统的监督微调(Supervised Fine-Tuning, SFT)不仅无法弥合这一差距,反而可能加剧模型对真命题的识别能力退化。解决方案的关键在于构建一个基于确定性Python验证器的约束性反例生成框架——SymCE,该数据集包含4,707个关于本科代数与实分析的错误猜想,每个猜想均配备可执行的验证模块,该验证器同时作为强化学习中的奖励函数,构成一个闭环训练环境。实验表明,仅采用反例的SFT会引发“模仿陷阱”,导致真命题识别率从0.27骤降至0.00;而引入基于稀疏结果奖励的强化学习(GRPO)结合验证器反馈(RLVR),可修复该缺陷并使性能提升至0.66。研究进一步揭示,尽管稀疏与密集奖励在域内任务上表现无显著差异,但在外部校准探针上出现33个百分点的性能分歧,归因于部分得分项(partial-credit term)的影响。最终,该4B规模模型在多个基准测试中超越所有评估过的7B级开源数学专用模型,并媲美六款前沿商业API,且在不调整提示的情况下实现跨任务迁移至GSM8K、MATH-500和MMLU-college-math,经人工审计的177次验证决策准确率达97.7%。

链接: https://arxiv.org/abs/2610.02444
作者: Omar Farouk Zouak,Houssam Eddine Boukhalfa,Soumaya Lakehal,Shiv Katiyar,Samia Nefti-Meziani
机构: National School of Artificial Intelligence (国家人工智能学院), Algiers, Algeria; University of Birmingham (伯明翰大学), United Kingdom
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: Accepted at EMNLP 2026 Findings

点击查看摘要

Abstract:Large language models often solve a theorem forward yet fail to disprove a closely related false one: a falsification gap that supervised fine-tuning does not close and can actively worsen. We frame counterexample generation as constrained witness emission against a deterministic per-theorem Python verifier, and release SymCE, a corpus of 4,707 false undergraduate-algebra and real-analysis conjectures, each paired with executable verifiers. The verifier also serves as the reward function, making SymCE a training environment. Training Qwen3-4B with SFT followed by GRPO under this oracle reveals an imitation trap: counterexample-only SFT collapses true-theorem recognition from 0.27 to 0.00, while RLVR with a sparse outcome-only reward repairs this and exceeds the base, to 0.66. The collapse replicates across four seeds and on Gemma-3-4B. Sparse and dense rewards yield statistically indistinguishable in-domain success yet diverge by 33 points on a held-out calibration probe, a dissociation we trace to the partial-credit term. Our 4B model outperforms every evaluated 7B open-weights math specialist, remains competitive with six frontier commercial APIs, and transfers under unchanged prompting to GSM8K, MATH-500 and MMLU-college-math. A human audit of 177 verifier decisions finds 97.7% accuracy. Code, data, verifier modules and annotations: this https URL.

[NLP-95] Are you Synthesizing or Recalling? Evaluating LLM s on Algorithmic Code Retrieval

【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在代码生成任务中缺乏可解释性与组件分离的问题,尤其关注模型在面对经典算法时是真正“推理”还是仅依赖内部知识“检索”已有实现。传统LLM流水线将算法知识回忆与应用推理混杂在一起,导致性能评估难以区分其真实能力。为此,作者提出将此类场景重新定义为参数化代码检索(parametric code retrieval)——即从模型内化知识中准确复现已知命名算法的能力,而非从零合成新代码。其解决方案的关键在于构建了一个名为AlgoREval的基准测试集,涵盖14个领域、7种编程语言和4种图结构输入表示形式,共599个问题,覆盖77个经典算法,用于在零样本(zero-shot)条件下独立评估模型的参数化代码检索能力。实验结果表明,不同语言和输入表示对检索准确率有显著影响,且通过提示工程(如引入检索到的代码片段或结构化算法提示)可提升复杂算法的生成质量;此外,监督微调(SFT)带来更广泛的跨语言收益,而基于强化学习的偏好优化(GRPO)则在特定语言上表现更优。研究确立了参数化代码检索作为一项独立可度量的能力,并警示在未经过系统验证前不应直接部署由AI生成的算法代码。

链接: https://arxiv.org/abs/2610.02438
作者: Nickil Maveli,Antonio Vergari,Shay B. Cohen
机构: University of Edinburgh (爱丁堡大学)
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Programming Languages (cs.PL)
备注: 30 pages (preprint)

点击查看摘要

Abstract:Large language models (LLMs) have demonstrated strong performance in code generation, where success depends on both recalling relevant algorithmic knowledge and reasoning about how to apply it. However, existing LLM pipelines are opaque, with no explicit separation between these two components. We argue that for well-known algorithms whose canonical implementations are widely accessible in pretraining corpora, code generation is better measured as \textitparametric code retrieval: reproducing a named algorithm from internalised knowledge rather than synthesizing a novel one. We introduce AlgoREval, a benchmark of 599 problems spanning classical 77 algorithms across 14 domains, 7 programming languages, and 4 graph-input representations to evaluate this capability in isolation, and assess 15 models (7B–34B parameters) in a zero-shot setting. We find substantial variation in retrieval accuracy across languages and input representations, even for widely documented algorithms and show that prompt augmentation with retrieved code snippets or structured algorithmic hints improve accuracy on complex algorithms, while SFT achieves broader language gains and GRPO achieves larger per-language gains on specific languages. Together, our results establish parametric code retrieval as a distinct, measurable capability and caution against deploying AI-generated algorithmic code without systematic validation.\footnoteCode and dataset are available at this https URL

[NLP-96] Evaluating and Improving the Robustness of Large Language Models to Input Sequence Variations

【速读】: 该论文旨在解决生成式 AI(Generative AI)在生产系统中面临的提示注入(prompt injections)、后门攻击(Trojans)以及自动质量评估指标被操纵等安全问题。其核心挑战在于提升大语言模型(Large Language Models, LLMs)对对抗性输入序列扰动的鲁棒性。解决方案的关键在于提出一系列可量化的评估与防御机制:首先,引入 R_stab(f) 作为基于 Jensen-Shannon 散度的生成式鲁棒性度量,用于量化模型在微小输入扰动下的输出分布稳定性;其次,针对局部攻击,理论证明了决策一致性概率 R_class(h) 与攻击成功率之间的关系;针对非局部攻击,构建了校准的经验模型以增强泛化能力。此外,提出了 ASA 攻击方法,在 LLM-as-a-Judge 系统中实现高达 73.8% 的攻击成功率,并具备跨模型迁移能力。在后门检测任务中,通过代理触发器实现了接近 0.99 的召回增强攻击成功率(REASR),显著优于基线。针对多层防御体系的绕过行为,系统化归纳四类规避策略,揭示现有防御的脆弱性。通过集成 5–7 个异构模型组成的委员会机制,使 Gemma-3-4B 的攻击成功率下降 47–55 个百分点至 19.3%。最后,面向基于模型上下文协议(Model Context Protocol, MCP)的智能体系统,提出 AttestMCP 方案,利用 HMAC 保护的数据包实现毫秒级(<0.1 ms)工具调用认证及“提交边界隔离”模式,在 MCPBench 基准上将平均攻击成功率从 53.7% 降至 12.4%。上述方法已集成于 JudgeGuard、TrojanArmor 软件套件及 MCPSec 模块中,形成端到端的鲁棒性保障体系。

链接: https://arxiv.org/abs/2610.02432
作者: Narek Maloyan
机构: 未知
类目: Cryptography and Security (cs.CR); Computation and Language (cs.CL); Machine Learning (cs.LG)
备注: PhD thesis, 2026. 118 pages

点击查看摘要

Abstract:Large language models (LLMs) in production systems face prompt injections, trojans (backdoors), and manipulation of automatic quality metrics. This thesis develops models, methods, and algorithms for evaluating and improving LLM robustness to adversarial input sequence variations. We propose R_stab(f), a generative robustness metric based on the Jensen-Shannon divergence between per-step output distributions under small input perturbations. For localized attacks we prove V(h) = 1 - R_class(h), where R_class(h) is the probability that a decision operator h keeps its decision under small perturbations. For non-localized attacks we propose a calibrated empirical model. For LLM-as-a-Judge systems we develop ASA, an adaptive evolutionary black-box attack that reaches an attack success rate (ASR) of up to 73.8%, with transfer between open models up to 62.6%. On Trojan Detection Challenge 2023 data (Pythia-1.4B), surrogate triggers reach REASR ~0.99 while recall of the true triggers is ~0.17 against a baseline of ~0.14. On SaTML CTF 2024 we systematize four classes of bypasses of multi-layer defenses, which reduce the ASR from 90% to 15-25%. Committees of 5-7 heterogeneous models reduce the ASR for Gemma-3-4B by 47-55 percentage points, to 19.3% with 7 models. For agentic systems based on the Model Context Protocol (MCP), we propose AttestMCP, which attests tool calls with HMAC-protected packets at under 0.1 ms per call, and the Commit Boundary isolation pattern. On the MCPBench benchmark of 847 scenarios they reduce the average ASR from 53.7% to 12.4%. The methods are implemented in the JudgeGuard and TrojanArmor software suites and the MCPSec module.

[NLP-97] Finding the Move Is Not Winning the Game: XiangqiBench for Closed-Loop Evaluation of LLM Agents

【速读】: 该论文旨在解决当前大语言模型(LLM)在复杂策略任务中评估标准与实际执行能力之间的脱节问题,具体聚焦于中国象棋(Xiangqi)中的战术终局场景。传统静态评估仅关注模型能否正确命名“最优着法”,但真实智能体必须在对抗环境中持续执行计划并达成可验证的胜利结果,而对手会动态响应。为此,作者提出了XiangqiBench——一个可执行的基准测试框架,通过119个由引擎支持的强制将杀终局,要求LLM代理在面对引擎防守方时完成将杀。其解决方案的关键在于引入交互式REPL接口,将真实落子、状态查询与前向模拟分离,并记录12个前沿大模型在两种观测协议下的8,568条多轮决策轨迹。研究揭示了三个看似体现模型能力的信号实际上高估了闭环成功率:(i)转化差距(Conversion Gap)显示模型虽在26.1%的可见试验中走出参考首步,但仅13.9%最终获胜;(ii)一致性差距(Consistency Gap)表明领先模型在pass@3上达38.7%,但pass^3仅为5.9%,且其成功案例中仅有7个位置能全胜三步;(iii)仿真差距(Simulation Gap)指出32.3%的模拟调用因非法走法终止,49.3%情况下真实对手回应与模拟路径不一致,说明自生成的回溯推演虽可检查合法性,却无法预测对手行为。因此,论文强调:单纯找到正确着法并不等同于赢得对局,未来智能体评估应以闭环结果为核心指标,并同时报告成功率与可靠性。

链接: https://arxiv.org/abs/2610.02425
作者: Yekun Chai,Qiwei Peng,Haoyi Xiong
机构: 未知
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Static evaluations credit a language model for naming the right move, but an agent must carry a plan through to a verified outcome while an opponent responds. We introduce XiangqiBench, an executable benchmark that measures this difference in Chinese chess: starting from 119 tactical endgames with forced mates supported by engine or checks-only search, an LLM agent must deliver checkmate against an engine defender. An interactive REPL interface separates real moves, state queries, and forward simulation, and we record 8,568 multi-turn trajectories from 12 frontier LLMs under two observation protocols. Three signals that look like competence each overstate closed-loop success. (i) The Conversion Gap: models play the stored reference first move in 26.1% of Sighted trials, yet only 13.9% of these trials end in a win. (ii) The Consistency Gap: the leading model reaches 38.7% pass@3 but only 5.9% pass^3, winning all three trials on 7 of the 46 positions it ever wins. (iii) The Simulation Gap: 32.3% of accepted simulation calls stop on an illegal move, and in 49.3% of comparable cases the real defender replies differently from the line the agent simulated; self-authored rollouts check legality but cannot anticipate the opponent. Finding the move is not winning the game: agent evaluations should score closed-loop outcomes and report reliability alongside coverage.

[NLP-98] rained Agent ic Context Management

【速读】: 该论文旨在解决长上下文语言模型(long context language models)在处理超长文本时的效率与性能瓶颈问题,尤其关注如何在有限上下文长度下实现与大规模上下文模型相当甚至更优的表现。传统方法通常依赖于原生训练长上下文或设计复杂的长上下文测试基准(harness),但这些方式成本高且难以泛化。本文提出一种简洁高效的训练框架,仅使用两个基本工具:一个用于自我调用并指定提示(prompt)的工具,以及一个用于读取输入上下文中指定范围标记(tokens)的工具。基于此极简框架,作者对Qwen3.6-35B-A3B模型在多样化合成数据集上进行微调。实验结果表明,在仅8,000个标记的上下文长度下,该小规模模型在文档长度超过4万标记时,其性能可媲美拥有100万标记上下文的GPT-5.4模型,在OOLONG-synth基准测试中表现突出。解决方案的关键在于通过构造最小化但功能完备的训练环境,使模型在资源受限条件下仍能有效学习长程依赖和上下文理解能力,从而突破传统长上下文建模的范式局限。

链接: https://arxiv.org/abs/2610.02404
作者: Bryce Sandlund
机构: 未知
类目: Machine Learning (cs.LG); Computation and Language (cs.CL)
备注: 17 pages, 6 figures, 4 tables. Code: this https URL . Under review

点击查看摘要

Abstract:We study long context language models. Instead of training long context natively, or designing a long context harness, we train a model over the simplest possible harness: a tool to call itself with any specified prompt and a tool to read tokens in a range from the input context. We finetune Qwen3.6-35B-A3B on a diverse synthetic dataset using this harness. With only 8,000 tokens of context, our small model is as strong as GPT-5.4 with 1M tokens of context on the OOLONG-synth benchmark when document length exceeds 40K tokens.

[NLP-99] Hesitation Has a Geometry: Entropy-Trained Hyperbolic Probes for Sparse Activation Steering

【速读】: 该论文旨在解决大语言模型在求解数学问题时,其推理过程具有高度层次性且在少数高熵分叉点处产生分支结构,而传统激活操控(activation steering)方法通常采用固定欧氏向量在每个词元处进行修正,忽略了多数词元已由上下文决定的事实,导致优化效率低下。其解决方案的关键在于提出双曲熵引导操控(Hyperbolic Entropy Steering, HEST),将隐藏状态嵌入庞加莱球(Poincaré ball)的双曲空间中,利用模型自身预测的下一个词元熵作为轻量级探测器的标签。当熵超过阈值时,HEST沿探测器读出函数的测地线最陡下降方向进行移动,并将更新映射回隐藏状态空间。针对学习到的理想点的布塞曼(Busemann)读出,理论证明了固定步长可使读出值在所有状态下等量下降,从而实现高效、自适应的推理路径调控。实验表明,在Qwen2.5-Math和Llama-3.1系列三个指令微调模型上,使用布塞曼读出的HEST在MATH-500与GSM8K基准上五组设置中提升贪婪准确率,最高达1.8个百分点;而传统对比式欧氏操控向量反而降低性能。该优势主要体现在模型频繁犹豫的问题上,对其他问题影响几乎可忽略,验证了双曲结构对层级推理建模的优越性。

链接: https://arxiv.org/abs/2610.02391
作者: Zeyong Zhang,Tung Sum Thomas Kwok,Tengfei Ma,Mengjia Xu
机构: New Jersey Institute of Technology (新泽西理工学院); University of California, Los Angeles (加州大学洛杉矶分校); Stony Brook University (石溪大学)
类目: Machine Learning (cs.LG); Computation and Language (cs.CL)
备注: 30 pages, 5 figures, 16 tables

点击查看摘要

Abstract:When a large language model solves a mathematical problem, its reasoning is largely hierarchical, and the solution often branches at a few tokens where the next-token entropy is high. Such tree-like structure embeds in hyperbolic space with far lower distortion than in Euclidean space. Activation steering, however, usually edits the hidden states of a pretrained model by adding one fixed Euclidean vector at every token, even though most tokens of a solution are already determined by the context. We propose Hyperbolic Entropy Steering (HEST), which embeds the hidden states in the Poincaré ball with a lightweight probe whose only label is the model’s own next-token entropy. Where this entropy exceeds a threshold, HEST moves the embedded state along the geodesic of steepest descent of a readout of the probe and maps the change back to the hidden state. For the Busemann readout of a learned ideal point, we prove that a step of fixed length lowers it by the same amount at every state. On three instruction-tuned models from the Qwen2.5-Math and Llama-3.1 families, HEST with the Busemann readout improves greedy accuracy on MATH-500 and GSM8K in five of six settings, by up to 1.8 points, whereas a contrastive steering vector added at every token lowers accuracy. With a Euclidean probe trained in the same way, this gain disappears on Qwen2.5-Math-1.5B-Instruct. The gains are largest on problems where the model hesitates often, and accuracy on the remaining problems is almost unchanged.

[NLP-100] Social bot detection in the age of ChatGPT : Challenges and opportunities

【速读】: 该论文旨在解决由高度智能化的生成式AI聊天机器人(Generative AI-based chatbots)兴起所带来的社交机器人(social bot)检测难题。随着AI生成内容在语言风格、行为模式和交互策略上的日益逼真,传统基于规则或简单特征的检测方法面临失效风险,难以有效识别具备复杂社会行为特征的新型社交机器人。其解决方案的关键在于:(1)利用生成式智能体(generative agents)构建合成数据以用于模型训练与评估,提升检测系统的泛化能力;(2)发展多模态、跨平台的检测框架,通过分析协同行为与影响力网络的拓扑及行为签名实现对大规模自动化账户群体的识别;(3)拓展检测技术在非英语及低资源语言环境中的适用性,缓解现有方法对高资源语言的依赖;(4)推动基于联邦学习的协作式检测模型,使不同组织与平台可在不共享原始用户数据的前提下协同建模,兼顾隐私保护与检测效能。这些方向共同构成了应对生成式AI时代社交机器人威胁的系统性研究路径。

链接: https://arxiv.org/abs/2610.02386
作者: Emilio Ferrara
机构: 未知
类目: Computers and Society (cs.CY); Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:We present a comprehensive overview of the challenges and opportunities in social bot detection in the context of the rise of sophisticated AI-based chatbots. By examining the state of the art in social bot detection techniques and the more salient real-world application to date, we identify gaps and emerging trends in the field, with a focus on addressing the unique challenges posed by AI-generated conversations and behaviors. We suggest potentially promising opportunities and research directions in social bot detection, including (i) the use of generative agents for synthetic data generation, testing and evaluation; (ii) the need for multimodal and cross-platform detection based on network and behavioral signatures of coordination and influence; (iii) the opportunity to extend bot detection to non-English and low-resource language settings; and, (iv) the room for development of collaborative, federated learning detection models that can help facilitate cooperation between different organizations and platforms while preserving user privacy.

[NLP-101] SEDIMA: Cross-Run Hierarchical Insight Memory for Evolutionary Search Agents

【速读】: 该论文旨在解决大语言模型(Large Language Model, LLM)驱动的进化搜索系统中存在的“记忆缺失”问题,即每次搜索过程均从零开始,导致智能体反复发现相同的改进路径并重复遭遇相同的死胡同,造成计算资源浪费。其核心解决方案是提出SEDIMA——一种持久化的分层洞察记忆机制。SEDIMA通过将原始搜索轨迹提炼为自然语言形式的可解释性洞察,利用注意力加权的中心点对这些洞察进行语义聚类,并在后续变异过程中检索相关知识以指导搜索方向,从而实现跨运行和跨任务的知识迁移。该方法作为无需修改原有搜索算子的即插即用模块,在固定评估预算下分别使AlgoTune和ALE-Bench LITE的平均最终性能提升5.5%和6.6%,并在OpenEvolve基准上平均减少32.3%的迭代次数即可达到基线最优性能,显著提升了进化搜索的效率与收敛能力。

链接: https://arxiv.org/abs/2610.02361
作者: Amirhossein Abaskohi,Mahdi Mostajabdaveh,Zirui Zhou
机构: University of British Columbia(不列颠哥伦比亚大学); Huawei Technologies Canada(华为加拿大技术公司)
类目: Neural and Evolutionary Computing (cs.NE); Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Large language model (LLM)-driven evolutionary search is a powerful paradigm for automated program and algorithm discovery, yet existing systems are largely memoryless: each run explores from scratch, so agents repeatedly rediscover the same improvements and re-encounter the same dead ends. We introduce SEDIMA, a persistent hierarchical insight memory for evolutionary search agents. SEDIMA distills raw traces into natural-language insights, clusters them by semantic similarity using attention-weighted centroids, and retrieves relevant guidance to condition future mutations, accumulating transferable knowledge across runs and problems rather than within a single trajectory. As a drop-in module that leaves the search operators unmodified, SEDIMA improves average final performance by 5.5% on AlgoTune and 6.6% on ALE-Bench LITE under a fixed budget of 100 evaluated candidates. Under OpenEvolve, SEDIMA requires 32.3% fewer iterations on average to reach baseline-best performance across the five evaluated backbones.

[NLP-102] Lexicographic Multi-Objective On-Policy Distillation

【速读】: 该论文旨在解决多奖励后训练中因缺乏显式优先级机制而导致的奖励权衡失衡问题,尤其在生成式 AI(Generative AI)任务中,正确性、推理质量与简洁性之间存在非对称权衡时,现有方法常因简单加权或组合专家策略而损害高优先级目标(如正确性)。其核心解决方案是提出字典序多目标在线蒸馏(Lexicographic Multi-Objective On-Policy Distillation, LMOPD),通过显式定义奖励优先级,在学生模型采样过程中按优先级顺序选择对应专家,并对低优先级专家的策略修正进行局部投影,以消除与高优先级目标冲突的策略成分。实验表明,LMOPD在双专家和四专家设置下均显著优于现有基线,有效保留了高优先级目标(如准确性和推理质量)的增益,同时合理获取部分简洁性收益,验证了显式优先级机制在多专家策略融合中的关键作用。

链接: https://arxiv.org/abs/2610.02359
作者: Doseok Jang,Jon Ander Campos,Youran Qi
机构: Cohere; Mila, Université de Montréal (蒙特利尔大学)
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: 24 pages, 3 figures, 5 tables; includes appendices

点击查看摘要

Abstract:Reinforcement learning from verifiable rewards (RLVR) usually optimizes answer correctness, yet useful language-model behavior also requires high-quality reasoning and concise responses. Existing multi-reward post-training methods typically scalarize rewards or combine specialists without explicitly protecting a reward priority order. This is problematic when trade-offs are asymmetric: conciseness, for example, should not improve at the cost of correctness. We introduce Lexicographic Multi-Objective On-Policy Distillation (LMOPD), a multi-teacher method for integrating reward-specialized policies under explicit priorities. For each student rollout, LMOPD selects the specialist for the first objective whose gate detects a deficiency, then locally projects its centered log-policy correction to remove components that oppose higher-priority specialists. We evaluate 30B-A3B mixture-of-experts transformer models in two- and four-expert settings on three math benchmarks, measuring retained specialist gains. With two experts, LMOPD’s point estimates fully retain the accuracy and reasoning-quality gains while acquiring 46.9% of the conciseness gain. With four experts, it retains \approx90% of both the accuracy gain and reasoning-correctness gain, compared to only \approx57% by the next best evaluated baseline. Matched four-expertablations show that lexicographic routing outperforms random routing and that projection further strengthens both top-priority capabilities. Across both scales, LMOPD preserves the highest-priority capabilities more effectively than the existing baselines we evaluate, demonstrating the value of explicit priorities for specialist integration.

[NLP-103] Does Every User Need a Private LoRA? Decoupling Personalization from Per-User Adaptation

【速读】: 该论文旨在解决个性化大语言模型在用户规模扩大时面临的可扩展性瓶颈问题,即传统方法需为每位用户配置完整的适配状态(adaptor),导致参数量随用户数线性增长,难以规模化部署。其核心解决方案在于重新思考个性化能力的分配机制:通过实证分析发现,用户独立适配器中存在大量跨用户可复用的结构,且这些可复用方向的效用同时受用户相关性与查询差异性的共同影响,同时用户历史记录蕴含可用于紧凑个体修正的可迁移信号。基于此,作者提出LINEUP框架,其关键创新在于构建一个共享的低秩个性化因子库,通过用户条件化召回与查询依赖校准进行动态组合,并将目标用户的适配过程压缩至仅优化一个微小的用户编码(user code),而所有共享组件保持固定。该设计实现了表达能力强的个性化能力与用户专属可训练状态的解耦,使得每个目标用户仅需优化8个标量参数,相较基准私有LoRA配置(419万参数/用户)大幅降低计算开销。理论分析进一步提供了有限步、有限历史下的风险边界及用户编码精炼优于历史初始化的充分条件。在涵盖个性化分类、预测与生成的六个任务上,LINEUP在全部12项指标上均显著优于现有基线(如LaMP-3 RMSE相对降低11.4%),且在历史数据受限场景下仍保持优势,验证了以可复用、条件组合的共享容量支撑丰富个性化,而独立用户适配仅保留极小修正状态的有效性。

链接: https://arxiv.org/abs/2610.02353
作者: Songyuan Sui,Srikanth Malla,Chiho Choi,Joon Hee Choi
机构: Samsung Semiconductor, US; Rice University (莱斯大学)
类目: Machine Learning (cs.LG); Computation and Language (cs.CL)
备注: 10 pages main content, 36 pages total including appendix, 7 figures

点击查看摘要

Abstract:Personalized large language models often require a complete adaptation state for each user. However, this paradigm scales poorly as the user population grows. We revisit this design through the lens of personalization capacity allocation: how much adaptation capacity can be shared across users, how the shared capacity should be composed, and how much must remain user-specific. We answer them through three complementary empirical analyses. We find that independent user adapters contain substantial cross-user reusable structure, that the utility of reusable directions reflects both user relevance and variation across queries, and that user histories provide transferable signals for compact individual correction. Motivated by these findings, we propose LINEUP. It learns a bank of reusable low-rank personalization factors, composes them through user-conditioned recall and query-dependent calibration, and restricts target-user adaptation to a tiny user code over a shared correction space. This design decouples expressive personalization capacity from per-user trainable state. Each target user optimizes only eight scalars, while all shared components remain fixed. By comparison, the evaluated private-LoRA configuration uses 4.19 million per-user parameters. Our theoretical analysis gives a finite-step, finite-history risk bound and sufficient conditions for user-code refinement to improve on history initialization. Across six tasks spanning personalized classification, prediction, and generation, LINEUP leads on all 12 metrics, each averaged over three independent runs (e.g., reducing LaMP-3 RMSE by 11.4% relative to the strongest baseline). It maintains advantages under limited history. These results show that rich personalization can be supported primarily by reusable, conditionally composed shared capacity, while independent user adaptation remains confined to a tiny correction state.

[NLP-104] HakemBench: A Turkish Benchmark of Typed Decisions

【速读】: 该论文旨在解决多任务、高精度的类型化决策评估问题,特别是在自然语言理解与生成模型在真实场景下进行判断时的可靠性与可解释性评估。其核心挑战在于如何构建一个全面、可复现且具备多维度评价能力的基准测试体系,以衡量模型在事实核查筛选、教育、安全防护(guardrails)、法律路由、内容审核、垃圾信息与钓鱼攻击识别以及客户支持等关键任务中的综合表现。解决方案的关键在于提出HakemBench这一土耳其语环境下的开放基准数据集(Version 1.0),包含2,346个样本及4,275个选择题、是/否题和评分题,覆盖七个评估轨迹;通过统一的评估框架,综合衡量决策质量(宏F1)、校准度(基于归一化Brier分数)与选择性自动化能力(基于归一化广义风险-覆盖率曲线下面积),并采用几何平均整合三者得分,同时通过2,000次自助采样报告置信区间,提升结果稳健性。此外,该系统还系统性地探查了选项顺序、改写表达、英文翻译及名称替换对模型行为的影响,增强了评估的鲁棒性分析能力。值得注意的是,尽管多数黄金标签由单一模型家族盲评与多方大语言模型投票达成,未经过人工验证,但其设计仍为模型性能的横向比较提供了可扩展、可重复的基准,尤其揭示了训练数据受早期测试结果影响所带来的潜在偏差(如防护、审核与客户支持任务的异常表现),从而推动更透明、更可信的模型开发流程。

链接: https://arxiv.org/abs/2610.02293
作者: Sait Furkan Teke(ufak AI)
机构: ufak AI
类目: Computation and Language (cs.CL)
备注: 9 pages including references. Data, harness, scorer and board: this https URL , this https URL

点击查看摘要

Abstract:HakemBench is a Turkish benchmark of typed decisions, in which the model under test reads a text, a question and a fixed set of options and returns a probability for every option. Version 1.0 is released fully open under CC BY 4.0, with 2,346 items and 4,275 choice, yes/no and score questions in seven tracks (fact-check triage, education, guardrails, legal routing, moderation, spam and phishing, and customer support). One harness scores decision quality (macro F1), calibration (from the normalised Brier score) and selective automation (from the normalised area under the generalised risk-coverage curve), combines them by a geometric mean and reports intervals from 2,000 bootstrap draws; probes for option order, paraphrase, English translation and substituted names are reported alongside. Most gold labels come from blind passes of one AI model family compared with the votes of a panel of large language models from other model families; they are not human-verified. On a board of 16 rows the leader scores a composite of 0.888 and the lab’s own model is 7th at 0.660. Its numbers are not blind. Earlier runs’ test results shaped its training data, so its guardrail, moderation and customer support numbers are flagged; with every model scored on the other four tracks only, its composite is 0.678, 6th of 16.

[NLP-105] Fast Models Slow Evidence: A Paired and Self-Audited Evaluation of System-1 Decision Models for LLM Agent Harnesses

【速读】: 该论文旨在解决智能体(Agent)在执行任务时需频繁做出大量细粒度决策的问题,例如选择调用哪个模型、使用何种工具、判断检索文本的相关性以及识别输入是否存在提示注入攻击等。传统方法依赖大语言模型(LLM)进行逐项判断,导致高昂的计算成本与延迟。为此,论文提出采用系统1型决策模型(System-1 decision models),通过单次前向传播输出类别概率,实现低延迟、低成本的快速决策。其解决方案的关键在于:构建并评估一个开源权重模型(Laya)与一个托管模型(Jev)在11个基于18个公开数据源的代理决策任务上的表现,涵盖7,283个基础案例及6,640个鲁棒性变体,确保输入字节完全一致、测试成对进行,并通过跨硬件与跨日可复现性验证。结果表明,Jev在11个决策点中有9个显著更准确(提升达10.8至46.0个百分点),而两者均未在零样本模型路由任务上超越随机猜测,且在RAG相关性过滤任务中表现持平。此外,研究发现Laya对选项顺序敏感(30%答案随顺序改变),在多候选或相似候选场景下性能急剧下降(50个最近邻工具时准确率仅31%,远低于Jev的98%)。研究还对自身分析流程进行了审计,揭示了三个分析错误和一个设计混淆因素:遗漏预筛选成本(宣称节省23.9%,实际仅为4.3%)、将门控准确率误报为端到端质量(58% vs. 实际98%)、样本内阈值设定导致高达17%的漏检,以及“通道效应”在通道原生内容下消失的提示注入误报问题。尽管存在其他潜在混淆因素,但不影响核心结论。所有实验数据、原始输出及分析代码均已公开。

链接: https://arxiv.org/abs/2610.02267
作者: Jiawei Li
机构: 未知
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Cryptography and Security (cs.CR); Machine Learning (cs.LG)
备注: 11 pages, 7 figures. Code and data: this https URL

点击查看摘要

Abstract:Agent harnesses make many small, typed decisions per task: which model to call, which tool to use, whether retrieved text is relevant, whether an input carries an injection. System-1 decision models answer such questions in a single forward pass with class probabilities, promising large cost and latency savings over LLM calls. We present a paired evaluation of an open-weight (Laya) and a hosted (Jev) System-1 model on 11 agent decision points built from 18 public sources: 7,283 base cases plus 6,640 robustness variants, with byte-identical inputs, paired tests, and cross-hardware and cross-day reproducibility checks. Jev is significantly more accurate on 9 of 11 decision points (+10.8 to +46.0 pp). Neither model beats chance on zero-shot model routing, and they tie on RAG relevance gating. Laya changes 30% of its answers when the option order is reversed and degrades sharply with many or similar candidates (31% at 50 nearest-neighbour tools, vs. 98% for Jev on items with a unique correct tool). We also audit our own pipeline. Three analysis errors and one design confound distorted headline deployment claims: an omitted pre-screen cost (reported 23.9% saving, actual 4.3%), gate accuracy reported as end-to-end quality (58% vs. 98%), in-sample thresholds (5% target, up to 17% held-out misses), and a “channel effect” on injection false positives that vanishes with channel-native content. Two other suspected confounds did not change the conclusions. All cases, raw outputs and analysis code are available at this https URL.

[NLP-106] Budgeted Cache Repair for Cross-Context KV-Cache Reuse

【速读】: 该论文旨在解决跨上下文键值缓存(Cross-context KV-cache reuse)在生成式 AI 推理中因重复利用旧前缀的键值对而导致的性能下降问题。尽管已有研究声称该方法可无损复用缓存,但本文发现其存在两个关键缺陷:一是隐性代价——在 MMLU 和 GSM8K 基准上,缓存复用导致显著的准确率损失;二是决策单位不当——缺乏明确规则判断是否应复用缓存,从而无法有效控制误差。解决方案的关键在于将缓存修复的决策粒度细化至最小单元,即单个 token 的键值对(row-level),而非整个请求或大块缓存。通过引入预算化缓存修复(Budgeted Cache Repair, BCR),该方法以两个“草稿”token 为起点,根据其注意力权重对缓存行进行排序,并仅对优先级最高的固定数量缓存行进行精确重计算,采用三种不同布局策略优化修复效率。该方法将选择的智能性发挥到极致,在单行粒度下可消除 49.5% 的随机基线误差,而在 64-token 块和完整调用层级则分别降至 10.6% 和零,表明选择精度随粒度增大而衰减。实验表明,BCR 在保持高缓存命中率的同时,使 GSM8K 的准确率恢复至密集预填充(dense-prefill)水平,且其最优布局超越所有基准复用方法的平均表现。此外,其草稿机制优于随机选择(如抛硬币)的对照组,填补了此前评估中缺乏有效基线的空白。

链接: https://arxiv.org/abs/2610.02233
作者: Haeyong Kang,Chang D. Yoo
机构: Duksung Women’s University (德淑女子大学); KAIST (韩国科学技术院)
类目: Hardware Architecture (cs.AR); Computation and Language (cs.CL); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Cross-context KV-cache reuse predicts a shared segment’s keys and values under a new prefix instead of recomputing them, and has been reported to do so without quality loss. We find otherwise, and identify two problems. (1) A hidden cost: on MMLU and GSM8K, reuse costs substantial accuracy. (2) A decision at the wrong unit: no rule for deciding whether to reuse a cache removes that cost. What does help is choosing which parts of the cache to recompute, and the value of choosing well falls as the unit of choice grows: informed selection removes 49.5% of the cache error beyond chance at single rows (one token’s keys and values), 10.6% at 64-token chunks, and nothing at the level of whole calls. Budgeted Cache Repair (BCR) acts at the unit where selection still pays. It drafts two tokens from the assembled cache, ranks cache rows by the attention those tokens pay them, and recomputes a fixed budget of rows exactly, in one of three layouts. The cost is paid rather than predicted away, and the draft that fails as a gate succeeds as a selector. BCR restores GSM8K to dense-prefill accuracy while still serving most calls from cache, and its best layout outperforms every reuse baseline’s mean in the reference grid. The draft also beats a coin-flip selector at the same budget - a control prior evaluations lack.

[NLP-107] IntentCoding: Amplifying User Intent in Code Generation

【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在代码生成任务中难以准确遵循用户意图(尤其是包含多重约束的复杂需求)的问题。现有研究表明,随着用户意图中约束数量的增加,模型性能显著下降;同时,尽管用户意图会对模型输出的logits产生一定影响,但其作用强度不足以有效引导解码过程。针对这一挑战,论文提出了一种名为“意图增强型代码生成”(Intent-Amplified Code Generation, IntentCoding)的新解码策略,其核心在于通过掩码用户意图并引入多强度集成机制,放大用户意图对生成过程的影响,从而提升模型对细粒度约束的遵守能力。该方法具有模型无关性,无需额外训练,可无缝集成至现有解码流程中。为系统评估模型在不同约束条件下的意图遵循能力,研究还构建了专门的基准数据集CodeConstraints。实验结果表明,与标准解码方法相比,IntentCoding在CodeConstraints、IFEvalCode、HumanEval和LiveCodeBench等多个数据集上均显著提升了约束满足率与功能正确性,相对改进幅度分别达到71.0%、67.3%和29.3%(pass@1),验证了其有效性与普适性。

链接: https://arxiv.org/abs/2602.00066
作者: Zheng Fang,Yihong Dong,Lili Mou,Dongming Jin,Zhi Jin,Ge Li
机构: Peking University (北京大学); University of Alberta (阿尔伯塔大学); Canada CIFAR AI Chair (加拿大加拿大人工智能主席)
类目: oftware Engineering (cs.SE); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Large Language Models (LLMs) have shown strong capabilities in code generation, but their adherence to fine-grained user intent with multiple constraints remains a significant challenge. Our empirical analysis reveals two key observations: 1) Model performance deteriorates quickly as the number of constraints in the user intent increases, and 2) While user intent does influence the model’s logits, such an influence may not be strong enough to effectively steer the decoding process. To this end, we propose Intent-Amplified Code Generation (IntentCoding), a novel decoding strategy that enhances an LLM’s ability to follow user intent. IntentCoding captures the influence of user intent by masking out the intent, and applies a multi-strength ensemble mechanism to amplify the effect of user intent during generation. IntentCoding is model-agnostic, requires no additional training, and integrates seamlessly with existing decoding procedures. To enable systematic evaluation, we also construct CodeConstraints, a benchmark dataset specifically designed to test user intent compliance under varying numbers of constraints. Experiments on our constructed Constraints, as well as popular IFEvalCode, HumanEval and LiveCodeBench datasets, show that our IntentCoding model significantly improves both constraint satisfaction and functional correctness compared to standard decoding approaches. IntentCoding achieves up to 71.0% relative improvement on CodeConstraints, achieves up to 67.3% relative improvement on IFEvalCode and achieves up to 29.3% relative improvement in pass@1 on HumanEval and LiveCodeBench compared with greedy decoding.

信息检索

[IR-0] MRVQ: One Resident Index for Dimension- and Rate-Elastic Vector Search

链接: https://arxiv.org/abs/2610.03651
作者: Sean Culatana,Shang-En Huang,Kang Li
类目: Artificial Intelligence (cs.AI); Information Retrieval (cs.IR)
备注:

点击查看摘要

Abstract:Dense-retrieval services must switch among embedding-prefix dimensions and index bit rates as latency, quality, and memory budgets change. Tuning a quantizer separately for each rate gives the best quality, but the retrieval tier then holds several code streams and quantizer states at once. We introduce Matryoshka Residual Vector Quantization (MRVQ), a post-hoc residual quantizer for frozen embeddings. Its maximum-rate code can be truncated two ways: dropping residual stages lowers the rate, and dropping embedding coordinates lowers the dimension. One resident artifact therefore serves every (dimension, rate) pair we evaluate. Across FiQA and NFCorpus, four embedding families, and 4, 8, 16-byte codes, MRVQ is the lowest-RAM design we evaluate. It uses 17.8-22.0x less memory than three separately trained QINCo2 indices, and 1.89-2.02x less than a lean shared-model steelman. The saving is not free: per-rate QINCo2 is 0.026-0.107 nDCG@10 better on FiQA. But MRVQ beats PQ, OPQ, and AdANNS-OPQ at matched code size. We also evaluate a low-build-cost PCA-scalar design that attains quality comparable to RaBitQ and its extension while fitting 420x faster at the median. Finally, we report two negative results: QINCo2 collapses when trained at high rates, and a ranking-bound hypothesis misses its pre-specified acceptance criteria. MRVQ is therefore a low-memory operating point for elastic retrieval, not a universal quality winner.

[IR-1] SGuard: A Real-Time Framework for Detecting and Imputing Missing Data in Streaming Time Series CIKM’26

链接: https://arxiv.org/abs/2610.03147
作者: Imane Hocine,Asma Abboura,Soror Sahri,Abhijith Senthilkumar,Yacine Hakimi,Grégoire Danoy
类目: Databases (cs.DB); Human-Computer Interaction (cs.HC); Information Retrieval (cs.IR); Machine Learning (cs.LG)
备注: The 35th ACM International Conference on Information and Knowledge Management (CIKM '26), November 07–11, 2026, Rome, Italy

点击查看摘要

Abstract:Streaming sensor applications routinely suffer from delayed or missing observations caused by faults, communication losses, or environmental interference. Although recent imputation methods exploit temporal and spatial dependencies effectively, most either assume offline access to future observations or prioritize throughput without enforcing domain plausibility. We present TSGuard, a real-time demonstration system for monitoring, validating, and imputing missing values in streaming time series. TSGuard combines a lightweight graph-aware temporal imputation model with constraint-aware validation, fallback estimation, and operator-facing explanations. Rather than treating imputation as an isolated prediction task, TSGuard integrates it into a broader data-quality loop: detect problematic observations, impute missing values, validate estimated against physical and spatial constraints, and either retain the original value as a plausible anomaly or replace it when it violates domain constraints. Using environmental sensing as a motivating setting, the demo enables users to inspect delayed sensors, compare imputers, define constraints, and validate flagged values in real time. The combination of lightweight online spatiotemporal imputation, domain-aware validation, and explicit retain-or-replace decisions is our central contribution, while interactive explanations make these decisions inspectable and actionable. for operators. Comments: The 35th ACM International Conference on Information and Knowledge Management (CIKM '26), November 07–11, 2026, Rome, Italy Subjects: Databases (cs.DB); Human-Computer Interaction (cs.HC); Information Retrieval (cs.IR); Machine Learning (cs.LG) Cite as: arXiv:2610.03147 [cs.DB] (or arXiv:2610.03147v1 [cs.DB] for this version) https://doi.org/10.48550/arXiv.2610.03147 Focus to learn more arXiv-issued DOI via DataCite (pending registration) Journalreference: Proceedings of the 35th ACM International Conference on Information and Knowledge Management (CIKM '26), November 07–11, 2026, Rome, Italy Related DOI: https://doi.org/10.1145/3799682.3840280 Focus to learn more DOI(s) linking to related resources

[IR-2] Benchmarking Literature Retrieval for a Model Organism: A Dictyostelium Case Study

链接: https://arxiv.org/abs/2610.03130
作者: Yun Wang,Gad Shaulsky,Tomaž Curk,Blaž Zupan
类目: Information Retrieval (cs.IR); Computation and Language (cs.CL)
备注: 15 pages, 5 figures. Submitted version (before peer review) of a paper accepted at Discovery Science 2026 (DS 2026); to appear in the Springer proceedings. Code and data: this https URL ; dataset: this https URL

点击查看摘要

Abstract:Biological literature retrieval systems are often developed and evaluated using broad biomedical corpora and general-purpose search tasks. However, many curated knowledge bases operate in narrower model-organism domains, where the literature is sparse and terminology is organism-specific. We introduce a retrieval benchmark from dictyBase for Dictyostelium, a model organism in cell and developmental biology. The benchmark consists of curator-generated biological queries linked to PubMed-indexed articles, together with structured gene annotations. Using this benchmark, we study three factors in niche biological retrieval: cross-encoder reranking, gene-aware query expansion, and abstract-only versus full-text retrieval. We report that reranking and gene-aware query expansion improve retrieval selectively: reranking is most useful when the model is well suited to biological evidence matching, whereas curated annotations help clarify compact biological queries by reducing vocabulary mismatch. Full-text chunks substantially improve retrieval when abstracts omit supporting evidence, increasing both candidate recall and top-rank performance, although these cases are harder than queries supported by abstracts. Data and code are publicly available at this https URL, and the benchmark dataset is additionally archived on Zenodo.

[IR-3] Query-aware routing for Cross-lingual performance gains in Encoders

链接: https://arxiv.org/abs/2610.02875
作者: Akshay Jain,Edward Kim
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Information Retrieval (cs.IR)
备注:

点击查看摘要

Abstract:Multilingual encoders can exhibit reduced retrieval effectiveness when queries and relevant documents differ in language, despite strong same-language performance. We investigate whether Finnish and Swedish cross-lingual retrieval can improve while preserving an encoder’s existing same-language performance and document index. We combine a query-only low-rank adapter, trained against frozen document embeddings, with deterministic routing based on query and index languages. Cross-language queries use the adapter, while same-language queries use the original encoder. SampoTron, our fine-tuned low-rank (LoRA) adapter alongwith the Nemotron-3-Embed-1B model, improves average retrieval quality across six English, Finnish, and Swedish directions from 0.241 to 0.291 in normalized discounted cumulative gain (nDCG) at rank ten, a 20.9% relative gain on a sampled financial benchmark. All six cross-lingual directions improve, and routing preserves the original same-language performance, including two full-corpus Finnish evaluations. The approach enables selective cross-language specialization with reusable document embedding vectors.

[IR-4] Learning Query Encoders Can Be Hard Even When Vector Retrieval Is Geometrically Easy

链接: https://arxiv.org/abs/2610.02749
作者: Anders Wikum,Nina Mishra,Amin Saberi,Tal Wagner
类目: Information Retrieval (cs.IR); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Efficient vector retrieval requires both a corpus geometry that supports retrieving the right documents through vector similarity, and a query encoder that can embed queries near their desired documents in the embedding space. Recent work has studied geometric capacity through the lens of the minimum embedding dimension needed to realize all top- k answer sets of n documents. We study a different notion of geometric capacity–the maximum recall achievable for a frozen document index–and explore whether learned query encoders can reach this ceiling. On several real-world retrieval benchmarks, we show that retrieval quality of single-vector query encoders often lies far below what the document indices can support. Motivated by this observation, we give theoretical evidence that learning query encoders can be computationally hard. In particular, we construct a retrieval task that (1) admits a query encoder with perfect recall which is representable by a small one-hidden-layer ReLU network, but (2) any statistical-query learner (a class capturing learners that access training data through aggregate statistics) provably requires exponentially many statistical queries to achieve non-trivial recall advantage over the random baseline k/n . Taken together, our results suggest substantial unrealized geometric capacity in retrieval benchmarks and establish query encoder learnability as a possible barrier in embedding-based retrieval. Subjects: Information Retrieval (cs.IR); Machine Learning (cs.LG) Cite as: arXiv:2610.02749 [cs.IR] (or arXiv:2610.02749v1 [cs.IR] for this version) https://doi.org/10.48550/arXiv.2610.02749 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[IR-5] Asterism: Exploring and Synthesizing Scattered Observations into Literature-Grounded Hypotheses and Theories

链接: https://arxiv.org/abs/2610.02673
作者: Joseph Chee Chang,Michael D’Arcy,Amy X. Zhang,Pao Siangliulue,Sangho Suh,Aakanksha Naik,Jena D. Hwang,Javier Ramos Benitez,Stella Wroblewski,Matt Latzke,Michael Cuoco,Ruben Lozano-Aguilera,Kris Ganjam,Joel Chan,Doug Downey,Peter Jansen,Kyle J. Travaglini,Daniel S. Weld
类目: Human-Computer Interaction (cs.HC); Computation and Language (cs.CL); Digital Libraries (cs.DL); Information Retrieval (cs.IR)
备注:

点击查看摘要

Abstract:A theory draws many independent observations into one framework with novel hypotheses. A researcher building such a theory must synthesize observations scattered across many papers, each describing related concepts but often in different terms. Which concepts matter most also depends on their preferences and research questions. Recent approaches scale theory synthesis with LLMs, but automate away choices and intuitions from researchers. We present Asterism, which extracts observations from hundreds of papers as concept-relation triples, with concepts unified in a hierarchical ontology. Researchers curate an evidence graph using the ontology and aggregate observations at different levels of granularity to focus theory formation on specific phenomena of interest. In a field deployment (n=10), researchers worked from observations to theories, and kept concepts and hypotheses fitting their preferences. In two case studies, teams of immunology and agriculture researchers discovered mechanisms outside their standard analyses and constructed hypotheses worth follow-up experiments.

[IR-6] When History Misleads: Asymmetric Margin Supervision for Instruction-Guided LLM Generative Recommendation

链接: https://arxiv.org/abs/2610.02600
作者: Ming Yin,Yuhan Yang,Chen Chen,Xinyu Lin,Wentao Shi,Fangcong Yin,Chaofei Yang,Chao Yang,Jiyan Yang,Hui Zhang,Ning Jiang,Yiran Chen,Qifan Wang
类目: Information Retrieval (cs.IR)
备注:

点击查看摘要

Abstract:In instruction-guided generative recommendation, LLM-based recommenders need to balance two goals: responding to the user’s current request and aligning with the preferences in their interaction history. When the two conflict, history events can override the request. We show that turning the effect of individual history events into supervision faces two obstacles. First, the events that most influence a recommendation are not necessarily the ones that support the target item. Second, removing a misleading event can raise the target’s score but a competing item’s score even more, so a higher target score alone does not guarantee a better ranking. We propose Asymmetric Intervention-Guided Margin Supervision (AIMS), which converts the effect of removing individual history events into ranking supervision. For training requests already ranked correctly, a frozen reference model identifies request-specific deletions that improve both the target’s score and its margin over a competitor near the recommendation cutoff. These margins serve as training targets, while the complete history is retained as input. Training combines cross-entropy with an asymmetric auxiliary loss that penalizes margin shortfalls and routes its gradient only through the competitor score. Inference is unchanged, requiring no history editing or deletion search. Across six LLM backbones on an industrial dataset and two public benchmarks, AIMS improves Recall and NDCG over strong baselines. Ablations support request-specific margins and asymmetric supervision, and the selected deletions preferentially remove constraint-violating history.

[IR-7] Adaptive Sparsity Optimization with Learnable Soft Top-K and Per-Term Thresholding for Efficient Retrieval SIGIR2026

链接: https://arxiv.org/abs/2610.02572
作者: Wentai Xie,Parker Carlson,Shanxiu He,Tao Yang
类目: Information Retrieval (cs.IR)
备注: Accepted at SIGIR 2026

点击查看摘要

Abstract:Recent work on neural sparse retrieval has demonstrated strong relevance by leveraging Large Language Models (LLMs) for semantic term expansion. However, learned models paired with previous sparsification techniques still yield overly long document and query vectors partly due to a large LLM vocabulary, imposing a serious challenge to retrieval time and space efficiency. This paper proposes a scheme for optimizing model sparsity through a synergy of adaptive strategies, including learnable soft top-K, per-term thresholding, and FLOPs regularization to increase the sparsity of query and document vectors. Experimental results with Lion-SP model on the MS MARCO and BEIR datasets demonstrate that the proposed scheme can outperform the baselines by significantly reducing the average query and document lengths. Our scheme can achieve much shorter retrieval latency and lower storage cost while maintaining highly competitive relevance.

[IR-8] On-Premises Multi-Course RAG Tutoring for Business Education: Hardware-Software Trade-offs in a Campus AI Tutor

链接: https://arxiv.org/abs/2610.02510
作者: Sidney Shapiro,Joshua Lindemann
类目: Artificial Intelligence (cs.AI); Information Retrieval (cs.IR)
备注: 23 pages, 3 figures, 4 tables

点击查看摘要

Abstract:Campus AI tutors based on retrieval-augmented generation (RAG) must ground answers in assigned course materials while keeping textbooks and student dialogue on institutional infrastructure. We present CourseChat, an on-premises, multi-course RAG tutor for undergraduate business education, deployed behind a campus web gateway and intended for use embedded in Moodle. Six isolated course offerings, each keyed by its own course reference number (CRN), share twin-edge AI hosts running a FastAPI service, a local vector database, and a local large language model (LLM) served by Ollama. We report two generation-model bake-off rounds, a separate fixed-evidence source-fidelity comparison, and conversation and quiz audits. Several larger models failed the classroom speed gate, but a 12B model and a 7B alternative passed. A separate mixture-of-experts candidate improved some corrections while introducing new factual and continuity errors. We therefore retain the 8B production model pending a demonstrated overall improvement, rather than claiming that 8B is universally optimal. Software changes improved follow-up topic resolution while preserving course scope; 435 prebuilt questions across 65 modules decouple practice from live generation. The results support treating model choice, evidence selection, serving compatibility, and product design as a joint engineering decision. They do not establish learning gains: faculty ratings, peak-load capacity, and complete public-gateway acceptance remain separate evaluation needs.

[IR-9] SOLO: Certified-Recall Metric Similarity Search with Scan-Only Sampled Inverted Lists

链接: https://arxiv.org/abs/2610.02387
作者: Édgar Chávez
类目: Databases (cs.DB); Information Retrieval (cs.IR)
备注:

点击查看摘要

Abstract:We present SOLO, an index for approximate nearest-neighbor search in general metric spaces whose serving path contains no ranking heuristic of any kind: a query is routed to the k_s nearest points of a random sample of the database, and every object in the touched posting lists is evaluated with the true distance. Because nothing must outrank anything, recall equals a coverage probability computable from the stored index: one ground-truth pass over a query sample certifies every operating point at once, without serving any of them – a recall certificate, and for a navigable graph no analogous object exists at any price. The whole index is one recursive rule – sample the collection, post each object to its b nearest sample points, split any list that outgrows a bound, always scan the leaves – and its operating surface obeys an equal-work law, recall \approx f(b \cdot k_s) , whose level is a one-scalar signature of the dataset. The same scan-only structure gives a serving floor no graph architecture reaches once the router is itself indexed by the same rule: Deep-100M served at recall 0.9977 from 1 GB of resident memory (enforced cap, 10.7 bytes per object) and at 0.9964 from 256 MB, Deep-1B at recall 0.9925 from 512 MB (and from 96 MB at depth 3), inserts that are one search, and deletes that are exact. Throughput is competitive where the hardware allows it – up to 1.8\times a tuned HNSW at 10^8 on a two-socket 32-core server, with operating points to the right of where that graph saturates – and the tables report it against HNSW, DiskANN, GRAFT, NAPP, misi, and SPANN’s assignment rule on the same hardware and ground truth.

[IR-10] A Manifold-Aware Topic Modeling Approach via Rank-Based Prototypes EMNLP2026

链接: https://arxiv.org/abs/2609.29630
作者: Thiago César Castilho Almeida,Daniel Carlos Guimarães Pedronette
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Information Retrieval (cs.IR)
备注: Accepted at the Main Conference of 2026 Conference on Empirical Methods in Natural Language Processing (EMNLP 2026)

点击查看摘要

Abstract:Recent topic models leverage pretrained embeddings, but neural architectures produce latent representations without grounding in specific texts, and clustering-based pipelines assign representative documents only post hoc, relying on absolute distances distorted by hubness and anisotropy in high-dimensional spaces. We introduce MARETopic, a training-free framework that casts topic discovery as rank-based prototype selection. After projecting embeddings onto a low-dimensional manifold, MARETopic builds ranked lists encoding ordinal neighborhood structure. A greedy algorithm selects exactly K exemplar documents, real corpus texts, whose neighborhoods cover the corpus. Two variants share this criterion. MARETopic _\textCorr scores candidates with a query performance predictor and a rank correlation measure, leading Purity and NMI on the two benchmarks with the most categories, ahead of both neural and clustering-based topic models. MARETopic _\textDiff scores them with a rank-based diffusion matrix, needs neither measure, and runs 1.7 to 1.9 times faster. Without a single gradient update, MARETopic leads topic coherence on two of three datasets. A novel inter-topic Maximal Marginal Relevance step raises vocabulary diversity at little cost in coherence. Our code is available at this https URL.

人机交互

[HC-0] Interactive Machine Learning Interfaces for Disease Risk Prediction: Effects on Risk Perception and Behaviour

链接: https://arxiv.org/abs/2610.03511
作者: Tiffany Ngai(1),Max Homm(1),Matthew Bradbury(1),Anamaria Crisan(1) ((1) David R. Cheriton School of Computer Science, University of Waterloo, Waterloo, Canada)
类目: Human-Computer Interaction (cs.HC)
备注: 14 pages, 3 figures. Tiffany Ngai, Max Homm, and Matthew Bradbury contributed equally to this work

点击查看摘要

Abstract:Machine learning risk models are increasingly being used in patient-facing health tools, but it remains unclear how well users understand the information these systems present. In this work, we study how people interpret an interactive Type 2 Diabetes (T2D) risk interface and whether interacting with it influences their attitudes toward behavioural change. Through an exploratory mixed-methods study with 15 participants, we compare participants’ perceived understanding with their actual understanding and identify key themes from qualitative interviews. We find that participants often understood the interface better than they initially believed, but still faced important barriers related to unclear terminology, ambiguous risk framing, and limited explanations of model inputs. Finally, we propose relevant design guidelines and discuss broader issues surrounding trust and fairness. Our findings highlight the importance of intuitive visual design, familiar presentation, and clear explanations in patient-facing ML interfaces.

[HC-1] Harmonic Eigenspace: A Web-based Application for Navigating and Composing Microtonal Harmony

链接: https://arxiv.org/abs/2610.03398
作者: David Dalmazzoa,Ken Déguernel
类目: ound (cs.SD); Human-Computer Interaction (cs.HC)
备注: 31 pages, 15 figures, 7 tables. Under review at the Journal of New Music Research. Web application: this https URL

点击查看摘要

Abstract:This paper presents a web-based application for navigating and composing microtonal harmony, built on the Harmonic Eigenspace, a four-dimensional psychoacoustically grounded space in which tetrad chord types are located by their spectral dissonance profiles, computed with Sethares’s roughness/dissonance model. The coordinate system is transposition-invariant: a coordinate triple (\alpha, \beta, \gamma) locates the three upper notes in relation to the root, so a chord quality corresponds to a direction in the space, the invariant ray along which transposition acts, while the root frequency sets the scale. The dissonance field over these coordinates can be computed at any register; the locations of its local minima are register-invariant, as they arise from partial-coincidence ratio conditions. The dissonance volume contains 100 local minima that align with just-intonation intervals and act as landmarks, organising the space into basins around the most consonant tetrads. We embed tetrads from three tonal equal temperaments as discrete lattices within this continuous volume. The application presents this space through two components: the Harmonic Eigenspace as a navigable 3D visualisation of the dissonance volume in which all nodes are playable, and a Modal Studio that extends modal interchange logic to the ten-gradation interval vocabulary of 53-TET. The application also functions as a MIDI controller with MIDI Polyphonic Expression support, usable in any digital audio workstation that supports this format. A listening study with 31 participants used both scenes of the application: listeners first rated isolated 53-TET chords alongside chords familiar from Western practice, such as the maj7 and the m7; they then rated chord progressions composed in the Modal Studio, measuring their acceptance or rejection of microtonal progressions heard for the first time.

[HC-2] Seeing through the Eyes of AI: Situated Explainability in Augmented Reality

链接: https://arxiv.org/abs/2610.03232
作者: Ana Stanescu,Lucchas Ribeiro Skreinig,Tobias Langlotz,Stefanie Zollmann,Peter Mohr,Dieter Schmalstieg,Mark Billinghurst,Denis Kalkofen
类目: Human-Computer Interaction (cs.HC)
备注: 11 pages, 8 figures

点击查看摘要

Abstract:Explainable Artificial Intelligence (AI) enables humans to understand and interpret decisions of AI models. Instead of having a black box, explainability supports humans in understanding AI models’ behavior. Existing explainable AI approaches often present explanations on 2D displays using pre-recorded data, requiring users to relate the displayed information back to the physical objects and real world locations involved in a model’s decision. Users are forced to decouple data exploration and capture from AI model interpretation. For AI systems that work within physical environments, this separation can make explanations difficult to interpret in context. We propose using Augmented Reality (AR) to enhance the understanding of AI models by enabling spatial explainability information directly in a user’s workspace, in real time, as they explore the world. We show how known explainability methods can be applied in AR and provide insights into user experiences with such an application.

[HC-3] he Effects of Air-Conditioning and Road-Traffic Noise on Perceived Cognitive and EEG Responses in a University Classroom

链接: https://arxiv.org/abs/2610.03210
作者: Yuanzhi Su,Cynthia Hou
类目: Human-Computer Interaction (cs.HC)
备注: 28 pages,14 figures

点击查看摘要

Abstract:Air-conditioning and road-traffic noise are common in classrooms, and their effects are often judged by cognitive performance. However, a sound that leaves performance unchanged may still be perceived differently or alter brain activity. In a within-participant design, this study tested whether perceived, cognitive, and EEG responses are consistent, and whether perceived and EEG responses distinguish conditions that performance does not. Sixteen students completed cognitive tasks in a university classroom under air-conditioning and road-traffic noise, each played back at 55, 60, and 65 dBA. Participants rated acoustic appraisal and perceived workload after each condition, and EEG was recorded during the tasks. The results show that cognitive performance did not differ significantly between conditions in any task. By contrast, noise was rated louder, less pleasant, and more arousing than a no-noise control, and all three ratings varied with sound level; perceived workload was also higher under noise. Task-period EEG power also showed significant differences between sources or levels in specific tasks and scalp regions. Within participants, relative EEG band power covaried mainly with loudness and pleasantness; louder and less pleasant ratings coincided with higher relative theta and lower relative gamma power. Associations with arousal and performance were rare, and none were found for workload. The three response types were thus only partly consistent, and perceived responses and EEG distinguished conditions that performance did not. Non-significant performance differences do not mean the conditions were equivalent for students. Assessments of air-conditioning and road-traffic noise in classrooms should therefore consider perceived responses alongside cognitive performance, with EEG providing complementary information on neural activity during tasks.

[HC-4] A Benchmark for Spatially Grounded Gesture Generation ECCV2026

链接: https://arxiv.org/abs/2610.03105
作者: Anna Deichler,Rishabh Dabral,Fethiye Irmak Dogan,Anindita Ghosh,Jonas Beskow
类目: Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR); Human-Computer Interaction (cs.HC)
备注: 13 pages, 9 figures. Benchmark of the Referential Gesture Challenge at the HSI Workshop, ECCV 2026. Data and video: this https URL

点击查看摘要

Abstract:Communication in shared space interweaves verbal and non-verbal signals, and pointing gestures anchor language to the environment: “put the cup on that one” is uninterpretable without the gesture that fixes the referent. Yet no common framework exists for evaluating whether generated gestures indicate their intended referent; distributional metrics reward a gesture aimed at the wrong object as long as it looks natural. We introduce a benchmark for spatially grounded gesture generation, comprising ~2K pointing-annotated clips from naturalistic VR dialogue with ground-truth 3D referents, a task in which systems must decide when, how and where to point within conversational speech, and a protocol that separates temporal alignment, spatial grounding and perceived naturalness. We also provide a flow-matching baseline, MM-Conv-Flow. Evaluating it alongside an independent retrieval-based system and captured human motion, we find that geometric grounding can exceed that of human pointing without any gain in perceived naturalness, showing that referential gesture quality must be measured along separate dimensions.

[HC-5] Personalized Automatic Speech Recognition for a Dysarthric and Tracheostomic Speaker using Artificial Conversations

链接: https://arxiv.org/abs/2610.03017
作者: David Nadrchal,Monorama Swain,Florian Schmid,Gerhard Widmer,Paul Primus
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Human-Computer Interaction (cs.HC); Sound (cs.SD)
备注: 8 pages, three figures, to be published in IEEE Speech Language Technology workshop 2026

点击查看摘要

Abstract:This work presents an automatic speech recognition (ASR) system personalized for a Czech speaker with a permanent tracheal stoma and severe dysarthria rendering their speech unintelligible to untrained listeners. We release a public dataset containing 33 annotated hours of the speaker’s speech, collected using a novel “artificial conversation” protocol designed for high engagement and dialogue realism. We propose a multi-stage training pipeline based on Whisper Base: fine-tuning on standard Czech speech, acoustically simulated tracheostomic speech, and the speaker’s data. We evaluate the system across three near real-time scenarios: scripted conversations, question answering, and spontaneous dialogue, achieving a 50% relative reduction in Character Error Rate compared to Whisper Base baseline and surpassing the average recognition accuracy of their assistants in acoustic recognition of isolated utterances. We demonstrate that even for severely impeded speech, a helpful ASR is achievable, as evidenced by the quantitative results and the feedback from the speaker.

[HC-6] racking Human Daily Cognitive Activity from EEG and Biometric Data

链接: https://arxiv.org/abs/2610.02971
作者: Alina Gutoreva,Zhaniya Omar
类目: Human-Computer Interaction (cs.HC)
备注:

点击查看摘要

Abstract:Understanding human cognitive activity in everyday life remains challenging due to the dynamic, context-dependent, and multimodal nature of cognition. Laboratory-based studies often fail to capture real-world cognitive processes, while single-modality approaches provide only partial insight into cognitive states. Advances in wearable sensing now enable the collection of heterogeneous data streams for a more comprehensive view of daily cognition. This paper presents a multimodal framework for tracking human cognitive activity using electroencephalography (EEG), wearable physiological signals, behavioral context, and self-reported measures. A preliminary pilot study was conducted with observational data from three participants (N = 3) over 280 annotated 10-minute intervals spanning nine activity domains across two weeks. Results reveal consistent temporal patterns, including a discernible mid-day decrease in motivation and energy at 13:00, followed by afternoon recovery. Work and IADLs yielded the highest flow state rates (51% and 50%), while ADLs produced the lowest (15%). Motivation correlated strongly with arousal (r = 0.78) and attention (r = 0.74), whereas perceived stress showed a weaker negative relationship (r = -0.33). A linear regression model predicting motivation from arousal, attention, energy, and stress achieved R2 = 0.76 (MAE = 9.84, RMSE = 12.85). Lag-based analysis indicates that prior energy levels positively predict subsequent motivation, confirming temporal dependencies in cognitive dynamics. These findings demonstrate the feasibility of multimodal cognitive activity analysis in real-world environments and highlight the importance of integrating physiological and behavioral indicators. The proposed framework provides a foundation for future large-scale multimodal systems and applied intelligent solutions. Subjects: Human-Computer Interaction (cs.HC) Cite as: arXiv:2610.02971 [cs.HC] (or arXiv:2610.02971v1 [cs.HC] for this version) https://doi.org/10.48550/arXiv.2610.02971 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[HC-7] Co-Designing AI For Mental Health Support With Young Adults of Color (YOC): Needs Expectations and Implications for AI Literacy

链接: https://arxiv.org/abs/2610.02812
作者: Elaine Dabin Jeon,John Bosco S. Bunyi,Hannah Kim,Renkai Ma,Yaman Yu,Michal Luria,Jason Yip,Alexis Hiniker,Katie Davis,Angel Hsing-Chi Hwang
类目: Human-Computer Interaction (cs.HC)
备注:

点击查看摘要

Abstract:Young adults of color (YOC) face heightened mental health challenges and barriers to care while navigating developmental and life transitions. Situated between youth-oriented safeguards and adult-oriented AI systems, little is known about how they use AI chatbots for mental health and well-being support or how sociocultural contexts shape their expectations, concerns, and design preferences. We conducted a two-day co-design workshop with 13 Asian, Black, and Hispanic/Latino/a young adults aged 18–24. Participants found generic chatbot advice to flatten their lived experiences; rather than making incorrect assumptions, they wanted more opportunities for identity-informed disclosure. Preferences for YOC-centered personalization also revealed gaps in privacy and AI literacy. Participants negotiated different therapeutic roles for chatbots and sought greater AI accountability and user agency, highlighting blurred boundaries between clinical and non-clinical AI-mediated support. Findings suggest directions for integrating AI literacy with mental health literacy and centering YOC’s experiences in the privacy calculus.

[HC-8] Characterizing the Performance Gap in Human Activity Recognition for Older Adults ISWC’26

链接: https://arxiv.org/abs/2610.02711
作者: Hossein Khayami,Sungjin Hwang,Eshed Ohn-Bar,David E. Conroy,Amanda Lazar,Eun Kyoung Choe,Hernisa Kacorri
类目: Human-Computer Interaction (cs.HC); Machine Learning (cs.LG)
备注: 8 pages, 6 figures, to be published in Proceedings of the 2026 ACM International Symposium on Wearable Computers (ISWC '26)

点击查看摘要

Abstract:Human activity recognition (HAR) from wrist-worn accelerometers is increasingly used for health and behavioral tracking. Yet, most wearable HAR models are developed and evaluated on datasets dominated by younger adults, leaving it unclear whether benchmark progress generalizes across age groups. In this work, we leverage MyMove, our carefully annotated, free-living older-adult HAR dataset (mean age 71), to evaluate deep-learning architectures and training regimes under both leave-one-subject-out and cross-dataset transfer. We find that improvements on younger-adult benchmarks fail to transfer equally to data collected from older adults, resulting in a persistent and often widening performance gap. However, richer representations, particularly frozen self-supervised features pretrained on the age-diverse UK Biobank dataset, substantially improve performance on data from older adults and consistently narrow the performance gap, at modest cost to younger-adult performance, though disparities remain. These findings suggest that benchmark gains and architectural scaling alone provide an incomplete picture of progress in wearable HAR, and broader advances may require representations that better capture population diversity, alongside personalized adaptation to individual movement patterns and routines.

[HC-9] LearnAdapt Praxis: Controlled AI Assistance and Evidence Traces for Adult Workplace Learning

链接: https://arxiv.org/abs/2610.02699
作者: Nizam Kadir
类目: Computers and Society (cs.CY); Human-Computer Interaction (cs.HC)
备注: 14 pages, 6 figures, 6 tables; technical system report with a reproducible synthetic verification study

点击查看摘要

Abstract:AI can help adults produce plausible workplace artefacts while leaving their independent reasoning difficult to inspect. LearnAdapt Praxis addresses this design problem through a five-stage learning workspace: Frame, Learn, Build, Validate and Transfer. The application stores an initial response, versioned specifications, coaching records, validation observations, a separate transfer response and facilitator feedback. Project-level server controls restrict coaching, access to earlier work and learner export during an active transfer attempt; they do not establish that a learner avoided assistance outside the application. This technical report describes the implemented architecture and evaluates selected workflow, authority, failure-recovery and provenance properties. A reproducible synthetic rehearsal exercised 120 serial project episodes across three fictional workplace contexts and four injected provider conditions. All 3,990 recorded checks met their specified expectations. The resulting records contained 240 specification versions, 1,200 activity events and 60 stored synthetic hints; 60 injected provider failures produced no fabricated hints. Separate authentication regressions exposed and resolved a session-expiry defect, and limited staging and production checks verified account access and live model connectivity. The evidence supports the tested engineering properties of the recorded release. It does not establish learning gains, model-output quality, unaided assessment validity, population usability or production capacity. The contribution is an implemented arrangement for separating assisted work, project-level assistance withdrawal and inspectable evidence, accompanied by executable tests and an explicit account of what remains to be validated with adult learners.

[HC-10] Lessons from Trauma-Informed Training on Technology-Facilitated Abuse for Gender-Based Violence Advocates

链接: https://arxiv.org/abs/2610.02688
作者: Naman Gupta,Connie W. Chau,Sophie Stephenson,Kate Walsh,Rahul Chatterjee
类目: Computers and Society (cs.CY); Human-Computer Interaction (cs.HC)
备注:

点击查看摘要

Abstract:Technology-facilitated abuse (TFA) is an emerging gendered public health crisis affecting millions of people across the globe. TFA coincides with other forms of gender-based violence (GBV), such as emotional, psychological, and physical abuse by an intimate partner. Survivors often seek support from GBV advocates who specialize in helping survivors navigate abusive situations and potential pathways for healing and remediation. With the rise in TFA, however, survivors and GBV advocates experience barriers and gaps in knowledge in identifying and mitigating this newer, ever-evolving form of abuse. To address this, we developed trauma-informed training as a critical capacity-building intervention to improve GBV advocates’ knowledge and skills for responding to survivors and to encourage multi-stakeholder collaboration toward a structural response. We facilitated 8 training workshops at local, state, and national US-based GBV organizations. Our study demonstrates that the trainings increased advocates’ perceived understanding of TFA abuse vectors and their impacts on survivors, along with advocates’ confidence in providing support, safety planning, and referrals through collaboration with community stakeholders. We then conducted retrospective reflections through autoethnography to surface key lessons, contestations, and tensions that emerged from our experiences in designing and coordinating trainings with host GBV organizations. Our analysis of autoethnographic notes revealed three key design decisions: coordinating logistics flexibly, reducing reflexive distance from host organizations and attendees, and using our field experience to tailor content for accessibility. We contribute a concrete checklist to guide future HCI scholars in conducting training interventions for stakeholders in high-stakes contexts.

[HC-11] DataWeave: Deploying Human-LLM Analytics for Exploratory Structured Data Analysis

链接: https://arxiv.org/abs/2610.02679
作者: Raquib Bin Yousuf,Harith Laxman,Vitaliy Shkremetko,Eunice Son,Shambhavi Verma,Brian O’Leary,Venketesh Subramony,Sylvain Nazef,Jacquelyn Elias,Ron Coddington,Chris Contakes,Michael Riley,Naren Ramakrishnan
类目: Artificial Intelligence (cs.AI); Human-Computer Interaction (cs.HC)
备注:

点击查看摘要

Abstract:Data journalism, the practice of using data analysis to surface newsworthy stories, depends increasingly on the ability of reporters and investigative journalists to uncover trends, disparities, and accountability narratives. In practice, exploring large structured datasets remains slow and brittle: journalists must navigate hundreds of variables across many datasets over years, understand data coding conventions, and write non-trivial analysis code while hypotheses evolve. Although LLMs are often touted as “ask in English, get SQL/answers,” real newsroom workflows expose recurring failures, e.g., schema mismatches and drift, misread domain semantics and units, and silent assumptions. We present DataWeave, a system that addresses these needs by combining conversational interaction, schema grounding, analytical planning, and executable query generation to support exploratory analysis over structured data. Rather than treating LLMs as autonomous answer engines, DataWeave frames them as interactive partners whose outputs can be inspected, corrected, and steered as hypotheses shift. We present a case study with professional journalists using our system to analyze the U.S. Department of Education’s Integrated Postsecondary Education Data System (IPEDS), a high-stakes public dataset with substantial domain semantics and frequent schema updates. We also report how deployment experience and iterative refinement shaped the current DataWeave architecture and its analytical workflow. Our findings distill design principles and deployment lessons for trustworthy human-LLM collaboration in structured data analysis.

[HC-12] Effects of a Behavioural Commitment Scheme on Study Regularity in a Self-Paced Learning Platform

链接: https://arxiv.org/abs/2610.02595
作者: Meenakshi V.(1),Pavani Ayinampudi(2),Aditya B. M. V.(2),Jinal Gupta(2),Prakash Hegade(2),Rohit Sharma(1),Sakshi Sharma(1),S. R. S. Iyengar(1) ((1) Indian Institute of Technology Ropar, Rupnagar, Punjab, India, (2) a href=“http://ANNAM.AI” rel=“external noopener nofollow” class="link-external link-http"this http URL/a, Indian Institute of Technology Ropar, Rupnagar, Punjab, India)
类目: Human-Computer Interaction (cs.HC); Computers and Society (cs.CY)
备注: 16 pages, 2 figures, 6 tables. Accepted as a long paper at the International Conference on Technology for Education (T4E 2026). Authors’ version, in revision after peer review; not the final published version

点击查看摘要

Abstract:Online education has enabled learners worldwide to take up courses from reputed institutions. However, a self-paced online course cannot guarantee the motivation and engagement that a learner experiences in a real-time classroom. Self-paced access also makes platform load unpredictable, which drives up the compute cost. We propose a Commitment Scheme for course access that aligns learner commitment with platform capacity. Learners book their study slots in advance; the instructor sets the budget of learning hours available for the course; and learners who use a full window earn additional watch hours. A booked window records an intention to study at a stated time, and a learner who appears in that window implements it. The byproduct is a platform load that can be forecast and bounded. The study involved two courses taken in sequence by the same learners, with the slot booking system activated only in the second. In-window study was observed on 86.8% of booked windows, and the median committer placed 95.5% of all study time inside self-booked windows. Among learners who studied across the launch, study regularity improved from 1.01 to 1.33 active days per week, with a supporting difference of +1.37 days per week against the same learners’ preceding course. Commitments made on the same day as the study slot were honoured more often than advance bookings (88.8% against 62.5% two days ahead), which is consistent with the classic intention-behaviour gap.

[HC-13] Santiagos A.T. Field: Visualizing Urban Accessibility through an Evangelion-Inspired Interface IEEE-VIS

链接: https://arxiv.org/abs/2610.02562
作者: Eduardo Graells-Garrido,Ignacio Pérez-Messina,Claudio Gaete
类目: Human-Computer Interaction (cs.HC)
备注: 3 pages, 2 figures; accepted in SciFi-Vis 2026 (IEEE VIS workshop)

点击查看摘要

Abstract:Science-fiction interfaces are often reproduced for their appearance, and their diagnostic logic is rarely rebuilt around real data. Here we describe Santiago’s A.T. Field, an interactive visualization of urban accessibility built from a diagnostic interface in Episode 13 of Neon Genesis Evangelion. We translated that interface in five stages: phenomenon, field, progression, interface, and response. A microscopic Angel became a relational urban problem, and a fictional scanning lattice became an accessibility field computed with the E2SFCA method. Every element of the fictional display is tied to a quantity of the model. Three tensions appeared when we made the fictional language operational: the saturated palette of the genre against color-vision deficiency, the motion of the genre against the reduced-motion preference, and the per-frame cost of the animation. The urgency of the fictional command center has its own cost, because emergency aesthetics can cast urban populations as threats. These three tensions are what the translation cost, and they are what transfers to other attempts of this kind.

[HC-14] Connectedness Cognitive Load and Human-AI Oversight in Cyber Operations

链接: https://arxiv.org/abs/2610.02384
作者: Nathan Conklin,Peng Gao,Chris North
类目: Human-Computer Interaction (cs.HC)
备注:

点击查看摘要

Abstract:AI-assisted cyber situational awareness triggers machine-generated reasoning traces (step-by-step justifications for anomaly classifications) that a human operator is expected to review. Because cyber signals and their traces arrive faster than any operator can process, human review is the limiting constraint on oversight. The standard approach is to identify the riskiest cyber events for review using model-side signals such as confidence or uncertainty. That framing ignores the operator’s cognitive capacity which varies sharply with the operational environment. We propose an alternative where the system’s environmental and connectivity telemetry serves as an available, non-invasive proxy for operator load. That same telemetry determines whether the human-AI partnership can reach the broader collective for support. In a maritime platform, environmental and connectivity attributes including depth, number of active communications paths, density of the tracked contact picture, and operational tempo all carry this signal. Need for operator oversight becomes a decision that materializes as a combination of both risk and environment-derived operator capacity. We present a reference architecture for a connectedness-aware oversight engine, demonstrating everyday use cases alongside its intended incorporation into the submarine cyber-defense toolkit. Two themes emerge: 1) the operator’s environmental state is itself a connectedness measurement, and 2) connectedness drives the cognitive load and defines a collective boundary in human-AI cyber operations.

[HC-15] “Im trying not to get hacked:” How Adults with Intellectual and Developmental Disabilities Navigate Security and Privacy Notifications

链接: https://arxiv.org/abs/2610.02374
作者: Hailey L. Johnson,Julia Nonnenkamp,Bilge Mutlu,Rahul Chatterjee
类目: Human-Computer Interaction (cs.HC)
备注: 27 pages, 8 figures, 8 tables. Accepted to The 28th International ACM SIGACCESS Conference on Computers and Accessibility (ASSETS '26), October 25-28, 2026, Vila Nova de Gaia, Portugal

点击查看摘要

Abstract:Security and privacy notifications, such as login alerts, spam email warnings, and cookie consent requests, play a critical role in shaping users’ responses to digital risks. Yet most notifications overlook cognitive accessibility, limiting their effectiveness for people with intellectual and developmental disabilities (IDD). We investigate how adults with IDD perceive and respond to common security and privacy notifications across mobile and web applications. Through a formative user study with seven adults with IDD, we identify three factors shaping understanding and decision-making: (1) interpretation is influenced by task and interface context; (2) unfamiliar terms, both technical and non-technical, are grounded in everyday concepts; and (3) uncertainty about outcomes leads to hesitation, avoidance, diagnostic exploration, or support-seeking. These findings lead to three design implications: (1) address context-dependent language misunderstandings beyond jargon simplification; (2) make action-outcome connections transparent; and (3) enable interdependent decision-making. Together, these insights aim to inform the design of more cognitively accessible security and privacy notifications that better support safe and supported user action.

[HC-16] Automating the Application of HCI Principles: Skills for On-Demand UI Construction the Human-AI Space to Think and the Future of HCI

链接: https://arxiv.org/abs/2610.02369
作者: Nathan Conklin,Miranda Capra,Chris North
类目: Human-Computer Interaction (cs.HC); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Human-computer interaction (HCI) is in the middle of a transition: large language models can now generate functional user interfaces (UIs) on demand from natural-language task descriptions. A user explains what they are trying to accomplish, and the system materializes a working interface to support it. This capability already exists in systems such as Claude and ChatGPT and continues to grow in fidelity as the underlying models improve. The next step along this trajectory is to move from interfaces that are merely generated to interfaces that are generated well. We propose a framework in which the dialogue between user and artificial intelligence (AI) becomes a Space to Think: a shared, structured cognitive workspace in which task decomposition produces an on-demand user interface as an extension of the user’s thinking rather than as a separate artifact. Within this paradigm, classical HCI design knowledge (Nielsen’s heuristics, Norman’s affordance prescriptions, Web Content Accessibility Guidelines (WCAG) success criteria, cognitive-load constraints, and mixed-initiative principles) is encoded as skills: machine-readable this http URL files that the generating agent loads at runtime as software engineering tools. Skills turn HCI design knowledge into declarative, inspectable, version-controlled, and editable artifacts owned by the HCI community itself so that accessibility, learnability, and consistency become properties of a generative process rather than properties of a finished product. We outline a research agenda depicting a future where the HCI field transitions from today’s design and knowledge heuristic checklist towards a future where the craft becomes machine-readable, executable, and open.

计算机视觉

[CV-0] Less Decoder is More Encoder: Geometric Representation Learning from Novel View Synthesis NEURIPS2026

链接: https://arxiv.org/abs/2610.03717
作者: Keerthi Kaashyap,Dennis Anthony,Akshay Krishnan,Nhi Ngoc Nguyen,Jeremy Collins,James Hays,Shreyas Kousik,Animesh Garg
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Robotics (cs.RO)
备注: Accepted to NeurIPS 2026

点击查看摘要

Abstract:This paper examines the role of Novel View Synthesis (NVS) in geometric representation learning. In principle, NVS should reason about 3D scene structure, thereby enabling transferable multi-view geometric representations. Yet, existing encoder-based NVS methods yield poor representations. This is not because of a lack of supervisory signal, but rather due to inconspicuous architectural choices: \textitspatially expressive decoders that dilute representational capabilities of the scene encoder, and \textitlow-level pixel-space targets that hinder feature learning. We present SNAP, a self-supervised encoder-decoder transformer that addresses both through a pose-conditioned local decoder and a latent-space reconstruction objective. SNAP is task agnostic, and we show that it is competitive with special-purpose geometry-supervised methods. SNAP also performs competitively against self-supervised representations across five tasks: visual localization, pose estimation, point correspondence, depth estimation, and robot manipulation. Remarkably, SNAP’s patch features exhibit emergent viewpoint invariance that approaches heavily supervised models despite lower compute and data budgets. Under camera shifts where standard 2D representations collapse, SNAP degrades more gracefully, revealing that restricting decoder expressivity actively prevents the suppression of transferable geometric structure. this https URL

[CV-1] MoSE3: Learning World-Space SE(3) at Every Pixel NEURIPS2026

链接: https://arxiv.org/abs/2610.03716
作者: Jiahuan Cheng,Zhiyi Li,Tian Xia,Ruojin Cai,Yilun Du,Qianqian Wang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: NeurIPS 2026 Spotlight. Project page: this https URL

点击查看摘要

Abstract:Dense 3D point tracking has been a prominent paradigm for modeling motion in dynamic scenes, but a point track is just a 3-DoF translation curve per pixel: it captures where pixels go, not the rotation of the underlying part, nor which pixels move together as one body. We propose MoSE3, the first feed-forward model that predicts dense SE(3) motion from monocular RGB video, producing full 6-DoF rigid transforms at every pixel in world space. Per-pixel SE(3) motion offers a richer view of how a scene moves: rotation, translation, and grouping all at once. Directly predicting SE(3) is challenging: rotations lie on a curved manifold that is ill-suited to Euclidean regression, and annotations for SE(3) are particularly difficult to acquire. To address these challenges, MoSE3 predicts per-pixel SE(3) through two jointly learned intermediates, 3D point tracks and rigidity embeddings, and recovers SE(3) by differentiably fitting transforms within each soft rigid cluster, enabling end-to-end prediction and supervision. To close the data gap, we introduce Art-Kubric, a large-scale synthetic dataset with dense SE(3) and rigidity labels for articulated objects with rich physical interactions. MoSE3 achieves state-of-the-art SE(3) estimation at pixel, part, and object levels on both rigid and articulated benchmarks, and state-of-the-art average 3D point tracking accuracy across three datasets, while showing strong generalization to real-world videos despite being trained solely on synthetic motion data.

[CV-2] 4DCodeBench: Benchmarking Agents on Inverse Graphics of Dynamic Scenes

链接: https://arxiv.org/abs/2610.03715
作者: Ruihong Shen,Žiga Kovačič,Peter Kulits,Xingrui Wang,Zizhang Li,Joshua B. Tenenbaum,Alan Yuille,Jieneng Chen,Jiajun Wu
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Graphics (cs.GR)
备注: this https URL

点击查看摘要

Abstract:We introduce 4DCodeBench, a benchmark for 4D inverse graphics through code generation, in which agents reconstruct dynamic scenes from video as executable graphics programs. To accomplish this, agents must translate visual observations into compact representations of scene structure and dynamics, by implementing abstractions such as physical simulations to reproduce complex behavior. To evaluate this capability, we curate a set of real-world videos and construct synthetic scenes spanning diverse physical phenomena, including deformation, fluid flow, and fracture. We perform extensive benchmarking of frontier models, finding that strong static reconstruction capabilities do not yet translate into reliable reconstruction of complex dynamics. 4DCodeBench provides a testbed for tracking progress toward agents that can interpret the dynamics of the world through code. Our benchmark is available at this https URL

[CV-3] What Should World Models Forget? Stratified Retention for Continual Adaptation NEURIPS2026

链接: https://arxiv.org/abs/2610.03713
作者: Nishit Anand,Ramani Duraiswami,Dinesh Manocha
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Image and Video Processing (eess.IV); Signal Processing (eess.SP)
备注: Accepted to NeurIPS 2026 Continual World Models Workshop

点击查看摘要

Abstract:Continual learning treats degradation on previously seen data as evidence of failure, a convention inherited from settings with a stationary prediction target, where a correct label remains correct indefinitely. World models do not satisfy this condition. Their prediction target is the environment, which changes, so knowledge that was accurate when acquired may later become false, and discarding it is required behavior rather than a defect. Non-stationary ground truth is well studied in the concept drift literature and in the temporal factuality of language models, but has not been formulated for world models, which are distinctive in that they also encode knowledge that must never be revised. We argue that continual world models require retention stratified by invariance timescale, separating invariants such as physics and object permanence, which must never be revised, from instance-level facts that should be revised as soon as the environment changes. Standard forgetting metrics cannot distinguish a world model that has correctly revised outdated knowledge from one that has suffered catastrophic forgetting, and consequently rank a frozen model highest, while existing physical-reasoning benchmarks evaluate only frozen checkpoints. We propose differential retention, which reports invariant regression testing across the adaptation stream jointly with revision latency, without aggregation.

[CV-4] Decoding the Functional Roles of Register and High-Norm Patch Tokens in Vision Transformers

链接: https://arxiv.org/abs/2610.03698
作者: Neel Varma,Andrew Rufail,Dipika Khullar,Vasu Sharma
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Self-supervised Vision Transformers (ViTs), such as DINOv2, learn rich visual representations, but the functions of their internal tokens remain poorly understood. Recent architectures introduce dedicated register tokens to reduce high-norm out- lier patch tokens that emerge in background re- gions, yet the semantic and functional roles of both token types have not been fully established. In this paper, we analyze these roles by training sparse autoencoders (SAEs) on register-token and outlier-token activations in DINOv2. Using an automated interpretability pipeline, UMAP clus- tering, and CLIP-space cross-checks, we find that register-token features are more strongly associ- ated with high-level semantic concepts. Outlier- token features, by contrast, are more often associ- ated with lower-level structural, background, and texture-dominant patterns. Causal ablations fur- ther reveal a substantial functional asymmetry: disrupting top-activating register-derived features produces a 48.17% drop in representation cosine similarity, whereas disrupting outlier-derived fea- tures produces only a 0.31% drop. Together, our results provide evidence for token specialization in self-supervised ViTs.

[CV-5] FlowHMR: Physically Plausible Motion Capture from Video

链接: https://arxiv.org/abs/2610.03691
作者: Zhanke Wang,Chengfeng Zhao,Qing Shuai,Jingzhong Lin,Heng Li,Zeyu Ling,Yuxin Wen,Jing Li,Di Kang,Chunchao Guo,Linchao Bao
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Project page: this https URL Code: this https URL

点击查看摘要

Abstract:We present FlowHMR, a framework for recovering physically plausible global 3D human motion from monocular video. Previous learning-based methods typically regress human motion directly from video and train the network with geometric supervision. However, recovering human motion from monocular video is inherently ambiguous in depth, and direct regression tends to collapse toward an averaged solution. Moreover, the recovered motions are not guaranteed to be physically plausible, so physics-based tracking of them often fails. To address these challenges, we formulate video motion capture as a video-conditioned motion generation problem and first pretrain a flow matching model for this task. Given an input video, the pretrained model generates diverse motion candidates, but not all of them are faithful to the video or physically trackable. We therefore post-train the model using Group Relative Policy Optimization (GRPO) with two rewards. A fidelity reward encourages consistency with the input video. A tracking reward favors motions that a physics-based controller can track successfully. Together, these rewards shift the model’s output preference, so the post-trained model stays faithful to the input video while producing more physically plausible motion. We further introduce Wild-4K, a large and diverse dataset of about 4K internet videos, for evaluating human motion recovery in the wild. Qualitative and quantitative experiments on Wild-4K show that our method outperforms state-of-the-art methods in overall motion fidelity and achieves a physical tracking success rate of 82.47%, compared with 62.82% for the strongest baseline, GVHMR.

[CV-6] SigLIP2 for aerial fire risk classification

链接: https://arxiv.org/abs/2610.03689
作者: Yunus Serhat Bıçakçı
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 7 pages, 3 figures, 2 tables. Code available at this https URL

点击查看摘要

Abstract:We examine the transfer of a pretrained SigLIP2 image encoder to seven class fire risk classification from aerial imagery. We introduce a reproducible partition of the public FireRisk training mirror and an implementation that records data provenance, preprocessing and model selection. Two initial runs compare a frozen encoder probe with full model adaptation. On the validation partition, full adaptation reaches 63.05% accuracy and 58.94% macro F1, compared with 55.95% and 50.19% for the probe. Both runs use one training seed and select their checkpoint on the same validation partition. These development results support further evaluation of SigLIP2 but do not establish performance on an independent test set or unseen regions. The accompanying code provides a common framework for repeated experiments and comparisons with additional visual encoders.

[CV-7] ProAR: Learning Prospective Reasoning with Autoregressive Video Models

链接: https://arxiv.org/abs/2610.03664
作者: Linghui Shen,Tinghui Zhu,Sheng Zhang,Muhao Chen
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Project Page: this https URL

点击查看摘要

Abstract:Autoregressive (AR) video models excel at causal generation, but their reliance on next-chunk prediction confines them to a short-sighted, reactive paradigm. This limitation is particularly consequential for reasoning-oriented generation, where achieving a target outcome through valid intermediate states matters more than local visual plausibility. To address this challenge, we propose Learning Prospective Reasoning with Autoregressive Video Models (ProAR), a novel framework that transforms autoregressive video generation into a goal-oriented reasoning process. ProAR introduces two key components: (1) To anchor generation to the long-range outcome, we integrate goal-frame prediction into the autoregressive loop via an asymmetric attention mask, enabling the predicted goal frame to guide the generation of intermediate states without being disrupted by them. (2) To guide short-range transitions, we introduce future representation self-alignment to encourage current hidden states to anticipate upcoming temporal dynamics. By leveraging teacher-forcing in AR training, we extract clean future representations in a single forward pass and align current representations with them using a lightweight, training-only predictor. Together, these two mechanisms seamlessly combine explicit, sparse target supervision with implicit, dense step-wise guidance, promoting coherent, goal-directed reasoning progress with modest computational cost. Experiments show that ProAR’s complementary components consistently improve performance across diverse visual reasoning benchmarks. The framework proves highly training-efficient, surpassing fully trained standard AR baselines using only 25% of the training steps. This paradigm also demonstrates promising applicability to embodied reasoning tasks.

[CV-8] On-Board Anomaly Detection for Efficient Marine Environmental Monitoring

链接: https://arxiv.org/abs/2610.03649
作者: Thomas Goudemant,Clotilde Szywala,Benjamin Francesconi,Michelle Aubrun,Yves Bobichon,Marjorie Bellizzi,Adrien Girard
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: 8 pages, 3 figures. Presented at the 9th International Workshop on On-Board Payload Data Compression (OBPDC 2024), Gran Canaria, Spain, 2-4 October 2024

点击查看摘要

Abstract:Marine ecosystems are impacted by various threats such as oil spills, algal blooms, and sediment floods, which disrupt habitats, wildlife, and human activities. Advances in satellite imagery and Artificial Intelligence (AI) have enhanced our capabilities for early detection and mitigation of such hazards. In this paper, we propose a marine event detection pipeline for Earth observation satellites equipped with multi- or hyperspectral sensors. Our approach includes a self-supervised neural network encoder that compresses satellite images into a reduced latent space, enabling efficient onboard processing. A machine learning anomaly detection model identifies deviations from normal sea patterns to detect environmental anomalies. We compare its performance against traditional algorithms such as Isolation Forest, One-Class Support Vector Machine and Local Outlier Factors. Our lightweight, resource-efficient pipeline is optimized for deployment on satellites with limited computational resources, ranging from embedded CPUs to AI hardware accelerators. By prioritizing the transmission of critical information, our solution enhances system responsiveness and optimizes satellite communication bandwidth. Demonstrated through current integration across multiple missions, including European Space Agency’s (ESA) Phisat-2 mission and Microsoft/Thales Alenia Space IMAGIN-e mission, our pipeline aims to improve marine environmental monitoring by providing timely alerts and efficient data reduction.

[CV-9] LoGo: Local-Global Rewards for Consistent Long-Horizon Video Generation

链接: https://arxiv.org/abs/2610.03636
作者: Ziqi Ma,Shreya Sharma,Mohamed El Banani,Katja Schwarz,Chongjie Ye,Chao-Yuan Wu,Li Fei-Fei,Ben Mildenhall,Georgia Gkioxari,Justin Johnson,Gowthami Somepalli
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注: Project website: this https URL

点击查看摘要

Abstract:Camera-controlled video models are rapidly advancing toward long generation horizons and complex camera control. A key failure mode is 3D inconsistency: as the camera moves, objects lose permanence and scene structures shift. Existing post-training techniques, which assign a single scalar reward to the entire generation, are poorly suited to correcting these inconsistencies over long horizons. We introduce LoGo, which blends global and spatially localized rewards for camera-controlled video models. The local reward provides fine-grained credit assignment, which substantially improves 3D consistency, while the global reward preserves camera following and video quality. Across three base models, LoGo shows a clear advantage on DL3DV and TrajectoryBench, a new benchmark for long-horizon, complex-camera-control generation that current evaluations lack. LoGo effectively reduces local object shifts, artifacts, and global scene changes, illustrating the importance of credit assignment in post-training video models. Project website: this https URL

[CV-10] Low-Cost Video–Time Priors as a Strong Baseline for EEG–fNIRS Emotion Regression on Familiar Videos

链接: https://arxiv.org/abs/2610.03618
作者: Minghao Kong,Jiurun Chen,Ying Gao,Xiangbin Meng,Rongjie Wang
类目: Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)
备注:

点击查看摘要

Abstract:Continuous emotion regression estimates moment-to-moment valence and arousal while a viewer watches a video. In familiar-video deployment, responses fron training participant-specific estimate, and prior-dominating fixed fusion tests whether physiology adds residual correction. In five-fold subject-held-out evaluation on 24was within 0.05 and 0.32 MAE of fusion in the internal and external evaluations, respectively. Source-explicit ablations showed that video identity and within-video tine accounted for most of the reduction, while EG-FNIRS gains were smaller and varied across participants and videos. These results identify the video-time prior as a strong, low-cost baseline and position EEG-fNIRS as an optional residual signal for familiar-video emotion regression.

[CV-11] DEPICT: Scoring Text-to-Image Alignment by Answer Agreement

链接: https://arxiv.org/abs/2610.03617
作者: Vasco Ramos,Sandra Godinho Silva,Joao Magalhaes,Ricardo Rei,Pedro Henrique Martins
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Image-text alignment is a core problem in computer vision with applications in caption evaluation, hallucination detection, data curation, and the benchmarking of text-to-image (T2I) generators. As T2I models improve, benchmarking has become demanding, requiring metrics capable of finding a series of issues like missing objects, swapped attributes, miscounts, and ignored negations. Recent work addresses this by fine-tuning evaluators on preference data or by prompting a vision-language model, either holistically with the caption or with decomposed verification questions. However, existing approaches fall short: fine-tuned metrics remain bound to one backbone and training distribution; holistic metrics miss fine-grained details; and decomposed metrics rely on a fixed-YES assumption that penalizes faithful images whenever that assumption fails. In contrast, we propose DEPICT, a training-free metric that replaces fixed reference answers with expected agreement between image-based and caption-only answers, weighting questions by how decisively the caption determines them. By replacing fixed references, our agreement rule increases negation accuracy from 19% to 88%. To recover the context lost during decomposition, DEPICT merges this agreement score with a holistic score. We evaluate DEPICT on five benchmarks and eleven backbones from three model families and find that it surpasses all training-free metrics and exceeds fine-tuned evaluators on two out of three human-correlation benchmarks.

[CV-12] ManifoldSplat: Language-Guided Semantic Shape Editing of 3D Gaussian Head Avatars

链接: https://arxiv.org/abs/2610.03599
作者: Antonio Canela,Jordi Sànchez-Riera
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: GCPR 2026

点击查看摘要

Abstract:High-fidelity 3D head avatars have reached near-photorealistic quality. While recent methods enable text-driven manipulation, they struggle to provide fine-grained localized control, often entangling features or lacking geometric consistency. Modifying geometry through natural language currently requires slow per-prompt optimization or compromises identity and rigging. We present ManifoldSplat, the first end-toend framework for language-guided semantic shape editing of animatable 3D Gaussian Splatting avatars reconstructed from monocular videos. By performing edits within the structured FLAME manifold rather than directly optimizing an unstructured Gaussian cloud, we strictly preserve identity and animation. We introduce DeltaRegion, a per-region disentangled Conditional Variational Autoencoder (CVAE) delivering feedforward shape deltas, alongside a refining stage to recover view-consistent details. ManifoldSplat reconstructs and edits an avatar in ~90 seconds on a consumer GPU, rendering at ~800 FPS. Extensive evaluations demonstrate our approach sets a new state-of-the-art in localized prompt alignment, geometric coherence, and identity preservation. Project page and code: this https URL

[CV-13] Rethinking What to Cache in Few-Step Diffusion Transformers: Solver-Aware Target Selection

链接: https://arxiv.org/abs/2610.03577
作者: Shuo Yang,Lihao Fang,Yi Zhang,Haixiang Wang,Xincheng Ye,Shufan Chen,Jipeng Guo,Youqing Wang
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注: 20 pages, 9 figures

点击查看摘要

Abstract:Diffusion Transformers (DiTs) can generate high-quality images and videos, but generating each sample requires multiple costly DiT forward passes. Two common ways to accelerate DiT sampling are step distillation, which reduces the number of sampling steps, and caching, which skips some DiT evaluations by reusing a tensor computed at an earlier step. Most caching methods decide in advance which tensor to reuse. After distillation, adjacent sampling steps are farther apart. Reusing a tensor across this larger gap introduces more error, so choosing what to cache becomes especially important. We therefore introduce AutoTarget, a method that chooses the cached tensor for a given model, solver, and reuse schedule. AutoTarget uses a small set of runs without cache reuse to measure the error caused by reusing each candidate tensor, then selects the candidate with the lowest error. We also analyze how an error at one reuse step affects the final sample. For Euler sampling, we identify cache targets that produce the same trajectory and show why a stored solver update may not. Experiments on distilled image and video DiTs show that the best cache target changes with the model, image resolution, and solver. AutoTarget reduces DiT evaluations and retained cache storage. Generation quality remains close to the corresponding uncached run. On the tested PixArt-LCM and FLUX.1-schnell settings, its calibration ranking matches the ranking from held-out cached runs. To help others reproduce the method, we provide its core implementation on GitHub at this https URL.

[CV-14] DuoMatching: Joint-Marginal Distribution Matching for Few-Step Video Generation

链接: https://arxiv.org/abs/2610.03543
作者: Jiahao Zhan,Yan Wang,Yongrui Ma,Qunliang Xing,Ruchang Yao,Runtao Liu,Shijie Zhao,Tianfan Xue
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Streaming video generation has benefited from distribution matching distillation (DMD), which matches the joint distribution of video frames to a video teacher’s approximation of the real video distribution. Although this joint matching mitigates drift during autoregressive rollouts, limitations remain in visual quality and semantic alignment. To address these limitations, we propose DuoMatching, a distribution matching framework that approximates the real video distribution through a unified joint-marginal formulation. On top of existing joint matching formulations, the additional marginal matching objective provides dedicated frame-level supervision from an image generator, transferring complementary visual and semantic priors from it. To apply this frame-level supervision in video generation, we introduce LatentBridge to resolve the latent representation mismatch between the video student and the image teacher. Latent Variation Sampling further distributes such frame-level supervision across distinct temporal segments, reducing redundancy. Experiments demonstrate that DuoMatching improves visual quality, composition, and semantic alignment while largely preserving motion dynamics. Human evaluations show overall preference rates above 80% against all evaluated baselines. The project page is available at this https URL.

[CV-15] Feedforward Novel View Synthesis for Heterogeneous Cameras

链接: https://arxiv.org/abs/2610.03522
作者: Meng Wei,Cheng Zhang,Boying Li,Yihang Chen,Jianmin Zheng,Hamid Rezatofighi,Jianfei Cai
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted at NeuralIPS 2026

点击查看摘要

Abstract:Feed-forward novel view synthesis has recently shown promising results from sparse posed images, but most existing methods assume that context and target views share a fixed camera family. This homogeneous-camera assumption breaks in practical multi-sensor systems, where perspective, fisheye, and panoramic cameras may coexist and where the target projection may be unseen during training. We study feed-forward NVS across heterogeneous central cameras and identify a key ambiguity introduced by tokenization: a visual token aggregates a projection-dependent bundle of pixel rays, while existing camera encodings mainly expose absolute rays or token-center relations. To address this, we combine token-center relative Camera Positional Encodings and proposed local raymaps, a token-level representation that explicitly describes the intra-patch ray distribution summarized by each token. We further propose projection-aware 2D RoPE, which replaces raw image-grid coordinates with ray-induced angular coordinates so that relative positional reasoning is aligned across camera projections. Together, these components treat diverse cameras as calibrated samplings of a shared ray space rather than separate visual domains. On ScanNet++ with heterogeneous-camera system, our method improves over camera-conditioned baselines under mixed-camera evaluation and demonstrates zero-shot generalization to panoramic views.

[CV-16] XGenAct: Geometry-Enhanced World Action Models through Cross-Task Generation

链接: https://arxiv.org/abs/2610.03516
作者: Tingting Du,Ziyao Wang,Guoheng Sun,Ang Li
类目: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注: 27 pages, including appendix

点击查看摘要

Abstract:World action models (WAMs) have advanced robot control by predicting how observations and actions evolve over time. Despite this progress, RGB and action based future prediction does not explicitly address the spatial understanding needed for robot manipulation. Existing efforts often add a limited set of spatial prediction tasks through specialized heads or branches, leaving both the range of spatial supervision and the model architecture fragmented. We introduce XGenAct, a world action model that represents RGB observations, robot actions, metric depth, surface normals, and functional role segmentation as RGB videos through deterministic codecs. By sampling perception and action tasks during training, XGenAct uses one video diffusion transformer and one objective to learn temporal prediction across these spaces without modality specific learned heads. On held out RLBench tasks, structured perception training improves average closed loop success over RGB only training, and XGenAct achieves 52% success in the five task external comparison, versus 26% for the strongest evaluated baselines. It also predicts future depth and segmentation more accurately than the evaluated pipelines that generate RGB first and then apply a frozen perception expert.

[CV-17] ProgressNet: Sketching and Prompting with a Frozen Text-to-Image Model

链接: https://arxiv.org/abs/2610.03512
作者: Arkaprabha Basu,Chaitat Utintu,Yi-Zhe Song
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Humans draw progressively: a few strokes, a look at the result, a stroke erased, a prompt revised. Image generators do not work this way. They typically take a finished sketch and produce the image in a single pass, so every edit starts the picture again, and the models that do keep state across turns are driven by text, cannot take a stroke, and are too slow to draw with. We present ProgressNet, a training-free framework that lets a frozen text-to-image model follow a drawing session as it unfolds: strokes are added and erased, the prompt is revised, and the image keeps up at about a second per turn. It needs no new parameters because the frozen model already has what a progressive generator needs, a pathway through which the previous turn can be remembered, layers that can carry appearance forward without freezing structure, and an internal signal of how far to trust an unfinished sketch; three inference-time mechanisms (Previous-Concept Memory, Layer-Selective K/V Injection and Banded Adaptive Control) use each in turn. As a sketch fills in, every existing method degrades, the FID of the FLUX+ControlNet baseline doubling between 10% and 100% completion on FS-COCO, while ProgressNet’s barely moves; it maintains strong fidelity and progressive coherence across three sketch domains and is preferred by users over five competitors, most widely on erasure.

[CV-18] Weave Forcing: Compositional Memory Routing for Interactive Long Video Generation

链接: https://arxiv.org/abs/2610.03510
作者: Ziyi Wang,Junchi Yao,Heqian Qiu,Wenbo Shi,Chengjiu Wang,Jinyang He,Binkai Hong,Hongliang Li
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Recent advances in autoregressive video generation have improved temporal consistency over extended durations, yet interactive storytelling requires more than continuous scene extension: a new shot may combine characters and backgrounds from different historical shots. Whole prompt retrieval can overlook the distinct reference needs of individual components, while directly combining all historical memories may introduce unrelated visual content. To address these problems, we present Weave Forcing, a training-free framework for compositional memory reuse in interactive long video generation. First, we use an LLM for semantic slot routing to decompose user prompts into character and background descriptions and explicitly select suitable historical references for each component. To isolate the required content, masked memory weaving uses contrasting attention maps conditioned on semantic slots to construct refined semantic masks, selectively exposing relevant tokens from compressed historical KV memories to guide the generation of the current shot. We further introduce coverage adaptive RoPE to adjust temporal offsets and memory retention according to no, partial, or full reference coverage, addressing visual artifacts observed when incomplete historical references are positioned close to the current generation. Extensive experiments demonstrate that Weave Forcing improves cross-shot subject and background consistency while maintaining competitive visual quality and text alignment.

[CV-19] Fed-ADApt: Federated Anytime Depth Adaptation for Resource-Aware Medical Image Segmentation

链接: https://arxiv.org/abs/2610.03474
作者: Abhijeet Parida,Zhifan Jiang,Pooneh Roshanitabrizi,Austin Tapp,Maria J. Ledesma-Carbayo,Syed Muhammad Anwar,Ziyue Xu,Marius George Linguraru,Holger R. Roth
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted to The 4th International Conference on Federated Learning Technologies and Applications (FLTA 2026)

点击查看摘要

Abstract:Federated learning (FL) enables collaborative training of medical image segmentation models without sharing raw patient data, yet existing approaches assume a homogeneous compute budget across institutions, limiting participation of low-resource sites. We propose Fed-ADApt, a depth-adaptive federated framework for UNet-based segmentation that jointly addresses low-compute training and inference. Fed-ADApt integrates multi-depth supervision with hierarchical depth-wise aggregation, allowing each site to train according to its local compute budget while contributing to a global model that supports dynamic depth selection at deployment. We evaluated Fed-ADApt on multi-site 2D retinal fundus disc segmentation and 3D brain tumor segmentation. Across both tasks, federated collaboration substantially improves robustness under domain shift. Fed-ADApt matched the full-resource FedAvg performance in 3D and achieved competitive 2D performance with a 4.7% average Dice reduction, while reducing average inference cost by 19.5% in 3D and 34.5% in 2D and substantially reducing training cost by 98% at the most constrained sites. Importantly, Fed-ADApt enables low-resource institutions that cannot train full-capacity models to participate in federations while maintaining competitive global performance under a favorable accuracy to efficiency trade-off. By considering training and inference compute budgets, Fed-ADApt provides a practical and equitable solution for federated medical image segmentation across heterogeneous clinical and edge-enabled imaging environments.

[CV-20] UniDynamics: Event-RGB Fusion for Unified Future 4D Dynamic Scene Generation

链接: https://arxiv.org/abs/2610.03473
作者: Daikun Liu,Xin Zhan,Teng Wang,Xiaoping Wang,Changyin Sun
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 19 pages, 6 figures, conference, code: this https URL

点击查看摘要

Abstract:We propose UniDynamics, a diffusion-based framework for future 4D dynamic scenes (RGB, depth, and optical flow) generation from a single event-RGB pair, without requiring long histories or control priors as in existing methods, while explicitly modeling future motion fields. The core idea is to leverage event streams to offer an alternative motion prior for single-RGB extrapolation, and to enforce geometric and motion constraints throughout generation via multimodal modeling. Specifically, we design an Event Latent Enhancement (ELE) module to align and enhance event latents into diffusion-injectable conditioning features, providing robust initial motion priors and reliable texture/structure cues. We further introduce a Perceptual Dynamics Space (PDS) embedded in the multi-scale U-Net, which decouples and adaptively interacts depth and flow while continuously feeding back constraints to appearance features, improving geometric-motion consistency for physically plausible and spatiotemporally coherent prediction. Experiments on VKitti2 and DSEC demonstrate state-of-the-art performance, producing high-quality, temporally coherent, and 4D-consistent future predictions, especially under challenging high-speed motion blur.

[CV-21] A Vision-Language Model (VLM)-based Pipeline for End-to-End Procedural Modeling of Field-Grown Maize from Point Clouds

链接: https://arxiv.org/abs/2610.03468
作者: Mozhgan Hadadi,Talukder Z. Jubery,Adarsh Krishnamurthy,Baskar Ganapathysubramanian
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Editable 3D models of field-grown crops support high-throughput phenotyping and in silico breeding trials, but building them from scanned point clouds requires organ-level segmentation and fitting. Procedural generators can turn an organ-level parameter set into an analysis-suitable 3D model, but obtaining that set requires hours of manual tuning per plant or segmentation models trained on species-specific labels. We present an automated pipeline that reconstructs procedural maize models from raw 3D point clouds without manual tuning or species-specific training data. A multimodal vision-language model (VLM) annotates leaf midlines in rendered orthographic views. Deterministic geometric algorithms back-project the annotations onto the point cloud, merge them into 3D leaves by cross-view consensus, and grow the midlines to full blades on an orientation-weighted surface graph. Measured organ parameters populate a plant descriptor for a Non-Uniform Rational B-Spline (NURBS)-based procedural model generator. Each leaf surface is then refined against its scan points by differentiable NURBS fitting. The pipeline reached a median whole-plant Chamfer distance of 5.4 mm on 100 genotypically diverse field-grown maize plants from the MaizeField3D dataset. The reconstructions were closer to the scans than those of an earlier semi-automated pipeline based on manual annotations. The pipeline recovered 1,017 of 1,023 (99.4%) curated reference leaves at an intersection-over-union of at least 0.5 without using those labels as input. These results show that VLM annotations become usable organ-level measurements when downstream geometric stages can correct them. This makes automated generation of editable 3D plant assets feasible at the scale of modern phenotyping experiments.

[CV-22] Preserving Anatomical Continuity: Three-Stage Pipeline for Colon Segmentation in 3D Abdominal CT Scans

链接: https://arxiv.org/abs/2610.03467
作者: Deshan Kalupahana,Sonit Singh,Praveen Ravindran,Arcot Sowmya
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注: 5 pages, 2 figures

点击查看摘要

Abstract:Accurate colon segmentation from CT images is essential for colorectal disease analysis, yet deep learning based methods often produce disconnected predictions due to complex anatomy. This study introduces a three-stage, topology-preserving segmentation pipeline to address this issue. The first stage performs initial deep learning-based segmentation, followed by centreline bridging to reconnect disjoint regions and a reconstruction stage to refine continuity. Evaluations on TotalSegmentator and RAOS datasets using overlap, distance and topology-based metrics demonstrate improved structural consistency while maintaining segmentation accuracy. The proposed method enhances topological integrity, enabling more reliable colon segmentation for clinical and research applications.

[CV-23] I2CD: Direct Image-to-Convex Decomposition for Simulation-Ready Collision Geometry

链接: https://arxiv.org/abs/2610.03453
作者: Qian Wang,Liam Merz Hoffmeister,Brian Scassellati,Daniel Rakita
类目: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Physics simulators and motion planners require convex collision geometry, yet image-to-3D generative models output dense, frequently non-manifold visual meshes. Bridging the two today takes a slow, brittle reconstruct-then-decompose pipeline of repair, decimation, and approximate convex decomposition. We present I2CD, which predicts a convex decomposition directly from a single RGB image. Rather than train a new image-to-3D model, I2CD freezes the pretrained Hunyuan3D-2 image-conditioned diffusion transformer and shape decoder and trains only a lightweight cross-attention head (38M parameters, under ten GPU-hours) whose learned “convex-slot” tokens emit the halfplane parameters of K convex polytopes. The output is compact, convex by construction, and loads into physics engines without any post-processing, in \sim0.5 s per image. On 227 held-out OmniObject3D and Google Scanned Objects instances, I2CD attains the highest volumetric IoU among eight reconstruct-then-decompose pipelines while running 6 - 37\times faster end-to-end. In a cross-simulator study in MuJoCo, PyBullet, Genesis, and Isaac Sim, every engine uses I2CD geometry as delivered, whereas raw generated meshes “load” everywhere but are silently replaced by a different collision shape in most cases or need seconds to minutes of per-object preprocessing. On a physical xArm7, I2CD produces planner-ready geometry for a 20 -object cluttered scene in 11 s versus 328 s for the strongest baseline, at comparable pick-and-place execution success ( 85 vs. 90 of 100 trials).

[CV-24] Corrupted but Correct: Why Vision-Language Models Lie to Themselves Internally NEURIPS2026

链接: https://arxiv.org/abs/2610.03445
作者: Arun Josephraj Arokiaraj,Zekun Wu,Adriano Koshiyama
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注: Accepted at the VLM4RWD Workshop (Grounded and Faithful Vision-Language Models for Real-World Deployment), NeurIPS 2026. 8 pages, 2 figures, 3 tables

点击查看摘要

Abstract:A targeted adversarial perturbation can drive a vision-language model’s (VLM’s) teacher-forced training loss for a fixed target caption to near zero, yet the same model, allowed to generate freely, produces the original, correct description with no trace of the target. We call this dissociation the train/inference gap, and give it a precise mechanistic account on Qwen2.5-VL-7B-Instruct using a controlled two-stage PGD attack on 200 held-out COCO images. First, we show that image-level pixel statistics, including a correctly re-implemented, texture-based attackability measure from the CNN robustness literature, have essentially no predictive power over which images are corrupted (best predictor r=-0.050, p=0.484; ridge regression R^2=0.069). Second, using the logit lens, we localise the gap to a single autoregressive step: the rank of the target token, conditioned on the correct first token already being generated, is fixed at exactly 3,488 out of 152,064 vocabulary entries for every image and every condition, with zero variance. Third, tracking target-token rank across all 28 LLM decoder layers reveals that the visual encoder corrupts every image’s representation by a comparable margin regardless of eventual outcome, but the language model decoder then differentially arbitrates: amplifying the corrupted signal for susceptible images and actively suppressing it, past its clean-image baseline, for resistant ones (p0.001, rank-biserial r=0.579). A linear probe on the merger hidden state separates these two outcomes with AUC=0.858, though we flag a circularity concern in this estimate. Together these results argue that adversarial robustness in autoregressive VLMs is substantially a property of the language decoder’s prior, not the visual encoder, with direct implications for where faithfulness evaluations and defenses for deployed VLM systems should be targeted.

[CV-25] ChromaGS: Text-Driven Semantic Editing of 4D Gaussian Avatars

链接: https://arxiv.org/abs/2610.03441
作者: Antonio Canela,Jordi Sànchez-Riera
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: CGIP 2026

点击查看摘要

Abstract:We present ChromaGS, a method for real-time, language-guided color editing of animatable 3D Gaussian head avatars. Given a trained animatable avatar, users can instantly modify the color of semantic regions through natural language, with edits applied at render time and no retraining required. Our key insight is to augment each Gaussian primitive with learned soft assignments to semantic regions and decompose colors into region-level base colors and Gaussian-level residuals. This decomposition enables coherent color transfer: modifying a region’s base color propagates naturally through all associated Gaussians while preserving fine appearance details encoded in residuals. A two-stage language pipeline translates text instructions into target colors, supporting both absolute specifications and relative adjustments. Unlike generative editing methods that may introduce unintended modifications, our approach provides deterministic, precisely localized semantic control. Experiments demonstrate faithful appearance preservation and intuitive interaction across diverse subjects. Project page and code are available at: this https URL

[CV-26] Depth Hypothesis Guided Iterative Refinement for Event-Image Monocular Depth Estimation

链接: https://arxiv.org/abs/2610.03439
作者: Daikun Liu,Teng Wang,Changyin Sun
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 14 pages, 13 figures, conference

点击查看摘要

Abstract:Event cameras hold excellent dynamic properties, showing great potential for monocular depth estimation (MDE). However, existing methods mainly improve performance by optimizing contextual features, but still struggle with the ill-posed and nonlinear nature of direct full-depth regression. In this paper, we propose HypoDepth, the first event-image monocular depth iterative refinement framework. By introducing a discrete Depth Hypothesis Volume (DHV), we transform the depth regression problem into a constrained depth search task. Specifically, we construct a 3D cost volume between the DHV features and contextual features and perform a multi-scale correlation search to guide stable residual optimization. This lightweight cost volume enables efficient global-to-local refinement across multi-resolution. Our method outperforms existing approaches on DSEC and MVSEC with state-of-the-art results and strong zero-shot generalization. Meanwhile, our tiny model achieves an excellent balance between accuracy and efficiency, enabling real-time performance on resource-limited devices.

[CV-27] he Shape of Speech: A Geometric Measure of Coarticulation for Speech-Driven 3D Facial Animation

链接: https://arxiv.org/abs/2610.03436
作者: Danzel Serrano,Przemyslaw Musialski
类目: Graphics (cs.GR); Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)
备注: 11 pages, 9 figures, 3 tables, under review

点击查看摘要

Abstract:Speech-driven 3D facial animation can reproduce recognizable mouth poses. However, it can simplify the motion between them, and that motion carries coarticulation, the way the sounds around each sound shape its articulation. We introduce a geometric measure of this trajectory shaping: lip-path length compared with the shortest route through the vowel, consonant and vowel positions of a speech segment. In contrast to the endpoint chord, this consonant-aware route accounts for obligatory transit and avoids degeneracy, while preserving invariance to uniform motion gain. The measure needs only a forced alignment, so it applies where no ground truth exists. We demonstrate it on four state-of-the-art methods, one per architectural family, real-time and offline. All four trace flatter lip trajectories than captured speech. Against frame-rate-matched ground truth, DiffPoseTalk, ARTalk and FaceFormer show clear deficits, equivalent on this measure to removing 15-60% of real speech’s fast articulatory component. CodeTalker is marginal on the primary measure and clear on a companion measure. A pre-registered study with 97 viewers and 3,523 judgments underpins the measured direction: controlled damping of real motion lowers the score and is penalized, whereas exaggeration shows no detected penalty over the tested range. Viewers also prefer real speech in 73.4% of sentence comparisons and, in the aggregate, on single words. Together, the measure, its calibration and the study identify a perceptually relevant loss of trajectory shaping and a concrete target for improving synthesized articulation.

[CV-28] OuroReward: Sequential Reward Scheduling for Reinforcement Learning in Text-to-3D Generation

链接: https://arxiv.org/abs/2610.03423
作者: Bingyang Cui,Yujie Zhang,Yiling Xu,Yunfeng Guan
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Reinforcement learning (RL) for Text-to-3D (T23D) generation requires optimization across multiple quality dimensions such as semantic alignment and texture clarity. Existing methods typically optimize these dimensions simultaneously through multiple reward aggregation, without explicitly modeling inter-dimension dependencies. This can cause imbalanced optimization and persistent interference among conflicting dimensions. To address this limitation, we propose OuroReward, an interference-aware sequential reward scheduling strategy for T23D RL. OuroReward first estimates pairwise dependencies among dimensions and constructs a cyclic optimization path that minimizes cumulative interference. By incorporating the tail-to-head dependency, the cycle captures global compatibility across the entire schedule. Then, OuroReward converts the cycle into a one-pass sequence, and starts optimization from the dimension with the lowest aggregate interference. Rather than assigning a fixed optimization budget to each dimension-wise reward, training adaptively determines when to advance to the next reward according to the remaining optimization headroom of the current one. We further introduce AdaSelect, an adaptive prompt selection strategy that identifies reliable and informative prompts aligned with the model’s current capability. By focusing policy updates on these prompts, AdaSelect effectively improves training stability. Extensive experiments across different T23D models and RL algorithms demonstrate that our framework consistently improves generation quality across multiple dimensions.

[CV-29] ForestQuery: Boundary-Aware and Spatially Anchored Query Learning for Unified Forest Point Cloud Segmentation

链接: https://arxiv.org/abs/2610.03403
作者: Zhihao Zhan,Le Tao,Yifei Tian,Xin Liu,Jie Yuan
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Forest point cloud segmentation is fundamental for fine-grained 3D forest scene understanding, yet remains challenging due to irregular tree structures, severe occlusions, density variations, and ambiguous instance boundaries. Recent query-based forest segmentation methods have shown promise for unified semantic and instance prediction, but they still insufficiently exploit forest-specific spatial structure and account for boundary uncertainty. In this paper, we propose ForestQuery, a boundary-aware and spatially anchored query learning framework for unified forest point cloud segmentation. ForestQuery enhances instance and semantic query learning through two complementary designs. Specifically, boundary uncertainty is explicitly modeled to guide reliable instance query construction and modulate query optimization through adaptive loss reweighting. Meanwhile, spatially anchored semantic query enhancement (SA-SQE) introduces learnable 3D anchors encoding forest vertical stratification priors to enrich semantic queries with explicit spatial references. We evaluate ForestQuery on multiple public forest point cloud benchmarks and a self-collected annotated real-world dataset. Extensive experiments demonstrate consistent improvements in both individual-tree segmentation and semantic segmentation across diverse forest scenes. Code and data are publicly available at this https URL

[CV-30] Beyond Entropy: Self-Diagnostic Multi-Role Token Optimization for Video Reasoning

链接: https://arxiv.org/abs/2610.03400
作者: Yudong Han,Yong Wang,Zaiquan Yang,Liang Lin,Chongyang Tao,Xiangxiang Chu,Liyuan Pan
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 19 pages, 6 figures, under review

点击查看摘要

Abstract:Reinforcement learning with verifiable rewards has substantially advanced multimodal reasoning, yet it remains fundamentally limited by ambiguous token-level credit assignment. While high-entropy token heuristics encourage possibility exploration, naively extending them to video reasoning tends to induce lengthy reasoning, as the model becomes overly reliant on high-entropy visual activations. Alternative approaches that rely on counterfactual-based visual token localization for credit assignment also tend to over-prioritize visual exploration at the expense of decisive reasoning cues for answer derivation, thereby exacerbating the interference from spurious visual nuances. Moreover, these methods employ static counterfactual strategies that fail to co-evolve with the policy during training. In this paper, we introduce DyCPO, a co-evolutionary framework that jointly optimizes reliable token selection and adaptive counterfactual intervention. It constructs a multi-role dependence metric to balance visual exploration and answer-relevance mining in token-wise contrastive learning, while suppressing exploration-only filler tokens and spurious visual noise. Rather than relying on static counterfactual priors, DyCPO dynamically derives counterfactual signals from the model’s own successful and failed rollouts, enabling self-diagnostic analysis and co-evolution of the optimization objective with the policy. Extensive experiments on complex video reasoning and general video understanding benchmarks demonstrate consistent performance improvements, establishing DyCPO as a robust token-level credit assignment paradigm for multimodal reinforcement learning.

[CV-31] Native Action-Prior Learning from Videos for World Action Models

链接: https://arxiv.org/abs/2610.03391
作者: Zhaochong An,Fei Zhang,Menglin Jia,Duncan Frost,Zijian Zhou,Yikai Wang,Xudong Wang,Aditya Patel,Belinda Zeng,Tao Xiang,Serge Belongie,Amir Bar,Sen He
类目: Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)
备注: Project Page: this https URL

点击查看摘要

Abstract:World action models integrate future visual dynamics with robot action prediction, but their scalability remains limited by the need for action-annotated robot trajectories. Observation-only videos contain rich evidence about interaction dynamics, but existing approaches typically use them either to pretrain visual representations that must later be adapted for control, or to infer latent actions that are subsequently grounded to robot commands. We present NAVA-WAM, which introduces native action-prior learning by directly pretraining the action policy from observation-only videos, avoiding indirect representation-to-control transfer or a separate latent-action model. Our training consists of two stages. First, we pretrain on observation-only videos, where future-video flow-matching supervision over visual transitions is propagated through transition-structured joint attention to optimize the Action-DiT and learn action-relevant priors. Second, we use action-labeled demonstrations to post-train the Action-DiT for robot control through joint video–action flow matching, while asymmetric attention decouples the visual branch from iterative action denoising and enables efficient action-only inference. Extensive experiments show that NAVA-WAM consistently outperforms prior approaches under both in-distribution and out-of-distribution settings, while demonstrating strong action-label efficiency and effective real-robot generalization. These results establish native action-prior learning as an effective approach to directly pretrain action policies from observation-only videos, providing a scalable path beyond action-labeled robot data.

[CV-32] From Patching to Pruning Visual Computation in Vision Language Models

链接: https://arxiv.org/abs/2610.03389
作者: Rahul Chowdhury,Timothy A Rupprecht,Xuan Shen,Shaoyi Huang,Pu Zhao,Yanzhi Wang
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Vision language models (VLMs) incur substantial inference cost because every visual token is processed by the attention and MLP projections of every decoder layer, even when token-specific visual computation is unnecessary at many depths. We introduce Patch-to-Prune (P2P), inspired by Mechanistic Interpretability, a training-free framework that converts activation patching from a diagnostic tool into an inference-time computation bypass. P2P performs validation-guided forward and backward layer sweeps to identify decoder regions whose visual-token projection outputs can be replaced by fixed neutral proxy activation vectors within a user-specified accuracy tolerance. Unlike conventional token-pruning methods, P2P preserves the sequence length, token order, positional information, attention mask, and residual pathways, thereby pruning computation without removing tokens or modifying the pretrained model weights. We evaluate P2P on four VLMs from the Qwen2.5-VL and LLaVA families across seven multi-modal benchmarks using mutually disjoint calibration, validation, and test partitions. P2P at a 3% tolerance retains around 94% of dense accuracy while reducing FLOPs by 55%. Beyond these efficiency gains, our layer-wise analysis suggests that visual processing in VLMs is non-uniformly distributed across decoder depth: early and late layers often require little token-specific visual computation, whereas intermediate layers appear to perform most task-relevant visual integration, enabling later reasoning to rely largely on visual information already embedded in shared residual and textual representations. This makes P2P both an efficient inference framework and a causal lens into visual information processing in VLMs.

[CV-33] Interpretable Deepfake Detection in Videos via Explicit Forensic Features and Temporal Modeling

链接: https://arxiv.org/abs/2610.03380
作者: Chahira Benhama,Mohand Saïd Allili,Assia Hamadene
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 10

点击查看摘要

Abstract:Deepfake detection in videos remains challenging, as manipulated content may appear visually consistent at the frame level while exhibiting subtle temporal inconsistencies. This paper introduces an interpretable deepfake detection framework that models spatially and temporally coherent facial features in video sequences. Unlike end-to-end deep models relying on implicit representations, the proposed approach explicitly encodes physically grounded forensic cues, enabling transparent analysis and improved multi-dataset generalization. The pipeline transforms videos into identity-consistent facial trajectories, segments them into fixed-length temporal windows, and represents each frame using 68 structured descriptors spanning four complementary domains: photometric, textural, geometric, and compression-based features. These descriptors provide a compact multi-domain representation of manipulation artifacts and are processed by a Long Short-Term Memory (LSTM) network to capture temporal dependencies and subtle irregularities. Evaluation on four benchmark datasets, FaceForensics++, Celeb-DF v2, a curated subset of the DeepFake Detection Challenge (DFDC), and DeeperForensics, yields strong and consistent F1-scores of 98.0%, 91.0%, 97.6%, and 96.2%, respectively. The approach also demonstrated a good cross-dataset generalization, providing a robust and interpretable solution for video deepfake detection.

[CV-34] EVEWorld: Physical Evolution Supervision for Embodied World Models

链接: https://arxiv.org/abs/2610.03374
作者: Kaiqi Wang,Songxin Zhang,Zejian Xie,Xiao Xiong,Zhuoyang Song,Ziwei Wu,Jun Yu Lu,Yitan Teng,Ziying Song,Jiaxing Zhang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 44 pages

点击查看摘要

Abstract:Embodied world models enable scalable simulation of embodied interactions for robot learning. However, existing models are prone to Model Laziness, as they focus on visual fidelity at the expense of physical reasoning and lack process-level supervision over the temporal dynamics of manipulated objects. In this work, we propose EVEWorld, a physical evolution-supervision framework for physically consistent target evolution. EVEWorld consists of two components: Instance-Guided Restoration (IGR) and Temporal Instance Alignment (TIA). First, IGR promotes instance consistency through restoration supervision. Second, TIA promotes cross-frame consistency by aligning target instances across adjacent frames. We further introduce the Model Laziness Rate (MLR), a metric that measures persistent violations of instance consistency in generated trajectories. Extensive experiments on DreamGenBench, EWMBench, and PBench demonstrate the effectiveness of EVEWorld, notably achieving an 87.5% reduction in MLR compared with GigaWorld-0. On the WorldArena 2.0 Track 1 leaderboard, our model ranks 6th in JEPA Similarity and 17th overall, which further validates the performance of our evolution supervision strategy.

[CV-35] LAS-CLIP: A Lightweight Adapter Steering Approach for CLIPs Visual Encoder

链接: https://arxiv.org/abs/2610.03370
作者: Anh-Khoa Dinh-Duc,Duc-Tai Dinh,Tam V. Nguyen,Minh-Triet Tran
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:CLIP’s visual encoder produces only global image representations, limiting its use in region-level tasks. Existing adaptations rely on visual prompting, input masking, or encoder fine-tuning, each compromising pre-trained representations. We propose LAS-CLIP, a Lightweight Adapter Steering approach that keeps every CLIP parameter frozen. A compact MaskAdapter generates per-head, per-layer attention biases from an input mask and injects them into the frozen self-attention layers, steering attention toward the target region. Crucially, because the backbone remains strictly untouched, LAS-CLIP seamlessly reverts to vanilla CLIP when no mask is provided, preserving its foundational zero-shot capabilities. With approximately 116K to 145K trainable parameters and 100K training samples on two T4 GPUs, LAS-CLIP achieves competitive or superior results compared to Alpha-CLIP on ImageNet-S zero-shot classification and RefCOCO referring expression comprehension, despite the latter fine-tuning its entire encoder on millions of samples. Qualitative analysis further confirms stronger representational fidelity under incorrect masks and in downstream generation. Our project page is link to this https URL

[CV-36] A Fully Automatic Pipeline for 3D Dendrite Instance Segmentation in SBF-SEM

链接: https://arxiv.org/abs/2610.03332
作者: Zewen Zhuo,Ilya Belevich,Eija Jokitalo,Alejandra Sierra,Jussi Tohka
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted at 2026 IEEE-EMBS Conference on Biomedical Engineering and Sciences (IECBES)

点击查看摘要

Abstract:Accurate three-dimensional (3D) reconstruction of individual dendrites in serial block-face scanning electron microscopy (SBF-SEM) is essential for quantifying structural plasticity in the brain, yet manual annotation at scale is infeasible. We present a fully automatic pipeline for 3D dendrite instance segmentation that unifies YOLOv6-guided Segment Anything Model (SAM) prompting on downsampled slices, iterative two-dimensional mask refinement, random forest 3D instance linking, and instance-aware high-resolution refinement using nnU-Net at native resolution into a single system requiring no manual prompting at inference. Applied to hippocampal CA1 SBF-SEM datasets from a control rat and a pilocarpine- induced epileptic rat, our pipeline reconstructs coherent, well- separated dendrites with high semantic accuracy (Dice 0.93 and 0.91) and strong instance-level performance on control tissue, while analysis of the more challenging epileptic tissue identifies instance recognition in dense regions as the principal remaining limitation. The high-resolution refinement stage recovers thin dendritic protrusions, providing a basis for downstream spine- level analysis. Code is available at this https URL ZE-WEN/dendrite-3d-instance-seg.

[CV-37] 3lescope: Arbitrary-Resolution High-Fidelity Generative Surface Reconstruction from Images

链接: https://arxiv.org/abs/2610.03308
作者: Atsuhiro Noguchi,Tianhan Xu,Yiming Liang,Yuta Kikuchi,Masahiro Ishiyama,Shintaro Takagi,Hitoshi Murai,Eiichi Matsumoto
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 45 pages

点击查看摘要

Abstract:We reconstruct high-fidelity 3D scene meshes from posed multi-view images without per-scene optimization, across scales ranging from single objects to large outdoor scenes. Per-scene optimization methods lack the learned 3D prior needed when observations are sparse or surfaces are glossy or transparent. Existing generative methods leverage such priors to complete geometry in sparsely observed regions, but typically operate at a fixed resolution over a limited spatial extent, trading spatial coverage against detail. Reconstructing a large scene therefore often requires partitioning it into independently processed overlapping local regions, making it difficult to maintain global geometric consistency. To address these issues, we propose T3lescope, which applies a single fixed-resolution generator across scene scales in an inference-time coarse-to-fine cascade. A coarse level establishes the scene layout, and finer levels perturb and denoise geometry inherited from the coarser level within progressively finer spatial cells to recover surface detail. The model is trained on individual cells at multiple scales and shares its weights across all levels, so no hierarchy is fixed during training, and the number of levels, cell scales, and cell locations are determined at inference time. On indoor, outdoor, and city-scale scenes, T3lescope outperforms feed-forward and generative baselines, matches or surpasses per-scene optimization, and recovers fine structures as well as glossy and transparent surfaces. These results show that our method generalizes across diverse scenes, view counts, and image resolutions. Project page: this https URL

[CV-38] HexVIO: Towards All-Day Stereo-Inertial Tracking Through Commodity DSPs

链接: https://arxiv.org/abs/2610.03283
作者: Patrick Wolf,Mateo de Mayo,Daniel Cremers
类目: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:The ability of a device to localize itself within its surroundings is a fundamental prerequisite for spatial computing. Visual-inertial odometry (VIO) has proven to be a cost-effective and accurate solution for this task. Robots, wearables, XR devices, and drones can benefit significantly from efficient implementations of VIO since they allow for cooler, lighter, and cheaper devices with longer battery life and a better user experience. In this work, we propose to enhance the efficiency of a VIO system by leveraging the Hexagon DSP, a commodity co-processor present in many modern smartphones and XR devices. Our approach offloads the visual frontend of a stereo-inertial odometry system to the DSP while keeping the backend on the main CPU. By optimizing the implementation for the DSP architecture, we achieve significant reductions in power consumption and latency compared to CPU-only execution. Our system, HexVIO, demonstrates a 67% reduction in power consumption or an 86% increase in throughput on a commodity smartphone, with the ability to sustain long-term real-time 30 fps tracking for 0.83 W, corresponding to ~18 hours of tracking on the testing device. These results highlight the potential of commodity DSPs for enabling all-day visual-inertial tracking in robotics and mobile devices.

[CV-39] Moving Forward with Video Saliency: A New Dataset and Benchmark where Motion Matters

链接: https://arxiv.org/abs/2610.03276
作者: Susmit Agrawal,Rebecca Wanner,Juliane Verwiebe,Matthias Tangemann,Matthias Bethge,Matthias Kümmerer
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Video saliency prediction is inherently harder to model than static image saliency due to the additional temporal dimension. Video saliency benchmarks rest on the premise that predicting gaze on video requires utilizing temporal activity distributed across frames. Prior work has challenged this, showing that static baselines recover a significant fraction of the explainable gaze information on LEDOV, a popular video saliency dataset, and that video saliency models fail in the same places as this static baseline. We verify that this diagnosis still stands: under a more capable gold standard than the original analysis, and an updated panel of recent architectures, the strongest temporal architecture in the panel still does not substantially improve over a fine-tuned static baseline. However, it remains unclear whether the marginal gain reflects limitations of current temporal architectures or a lack of temporal patterns in the benchmark itself. We introduce SalTempto, a video saliency benchmark with greater dynamism: 224 clips of highly dynamic content, sourced from the HACS-Segments dataset so that each clip contains an event together with its lead-up and aftermath, with gaze recordings from up to 16 subjects and a training split for adapting pretrained models. On SalTempto, the static baseline recovers only about 13% of the headroom above the centerbias, against more than half on LEDOV. A fine-tuned temporal architecture shows a substantial gain in performance over the static baseline, indicating that it does capture meaningfully more temporal information, which LEDOV fails to measure. Yet, even this SoTA model still leaves nearly half of SalTempto’s headroom unexplained, indicating room for improvement in video saliency modelling. Examination of SalTempto also lets us describe human tendencies that models miss. SalTempto link: this https URL.

[CV-40] Consecutive Posterior Fusion for Diffusive Recovery of Unobservable Image Structures

链接: https://arxiv.org/abs/2610.03261
作者: Elena Morotti,Davide Evangelista,Elena Loli Piccolomini
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注: 21 pages, 7 figures, 2 tables

点击查看摘要

Abstract:Solving severely ill-posed imaging inverse problems requires recovering image structures that are unobservable or weakly constrained by the measurements. Diffusion models provide expressive learned priors for inferring such missing information, while posterior sampling incorporates measurement consistency along the reverse process. Standard diffusion posterior samplers, however, rely on instantaneous measurement-aware estimates, without explicitly exploiting information carried by previous posterior corrections. We introduce Consecutive Posterior Fusion Denoising Diffusion Null-Space Models (CPF-DDNM), an inference-time strategy that fuses consecutive measurement-aware estimates to improve the diffusive recovery of unobservable image structures, without requiring retraining or additional denoiser evaluations. We instantiate this principle within DDNM, whose range/null-space decomposition reveals that consecutive fusion preserves the measurement-determined component while acting exclusively on the prior-driven null-space estimate. We thus provide a geometric interpretation of CPF-DDNM and a local error analysis that characterizes the optimal time-dependent fusion coefficient, including the extrapolative regime. Experiments on sparse-view and simulated low-dose computed tomography, as well as medical image super-resolution, show consistent improvements over DDNM and competitive performance against diffusion-based inverse solvers. Comments: 21 pages, 7 figures, 2 tables Subjects: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI) Cite as: arXiv:2610.03261 [cs.CV] (or arXiv:2610.03261v1 [cs.CV] for this version) https://doi.org/10.48550/arXiv.2610.03261 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[CV-41] COSMI: COmpositional Synthesis of Multi-object Interactions

链接: https://arxiv.org/abs/2610.03252
作者: Daniel Eskandar,Ilya A. Petrov,Gerard Pons-Moll
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Generative models of human-object interaction are bounded by the data that exists: everyday activities involve several objects, but most captured datasets record one at a time, as multi-object capture is combinatorially expensive. Our observation is that interactions are local, so single-object captures already contain the parts of multi-object activities. We compose them: contact-consistent clips of single interactions, mirrored to balance the hands, transfer between bodies, and a language model and geometric checks admit only the pairings that are plausible, semantically and physically. Therefore, the dataset grows combinatorially with the clips rather than recording time. The COSMI dataset holds 222k sequences and 275 hours with up to five objects, nearly thirty times the largest multi-object capture, and can be extended by adding datasets or even hand-object recordings. On this data we train the COSMI method, a text-to-interaction diffusion transformer that follows how the data is built: weight-shared object slots generate a variable number of objects, predicted relative to the body parts that move them. On a benchmark with an unseen object and unseen interaction combinations, models trained on the dataset generalize to the unseen combinations. COSMI outperforms baselines in text alignment and contact accuracy, where its margin is largest on the unseen object. Code, models, and the dataset pipeline will be released on the project page: this https URL.

[CV-42] EmbPASS: Towards Cross-Embodiment Open Panoramic Segmentation

链接: https://arxiv.org/abs/2610.03248
作者: Pujun Guo,Yuanfan Zheng,Fei Teng,Mengfei Duan,Guoqiang Zhao,Yuheng Zhang,Kai Luo,Kailun Yang
类目: Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)
备注: 9 pages, 5 figures

点击查看摘要

Abstract:Panoramic images provide a complete 360-degree field of view, enabling comprehensive scene understanding for embodied perception. However, heterogeneous embodied platforms exhibit substantial differences in observation viewpoints and spatial layouts, giving rise to cross-embodiment observation shifts that pose additional challenges to consistent and reliable panoramic perception, while systematic studies of this problem remain limited. To bridge this gap, we introduce a new task, termed Cross-Embodiment Open Panoramic Segmentation. Meanwhile, we establish EmbPASS, a multi-platform panoramic semantic segmentation benchmark spanning Vehicle, Drone, Wearable, and Quadruped platforms under a unified semantic taxonomy, providing a testbed for systematically studying cross-embodiment panoramic perception. We further propose EPONet, an open-vocabulary panoramic semantic segmentation network that integrates Relation-Aware Metric Adapter (RAMA) and Content-Adaptive Semantic Transfer (CAST) to enhance spatial modeling and semantic transfer under heterogeneous embodied observations. Extensive experiments show that EPONet achieves the best platform-balanced performance on EmbPASS with 35.82% mIoU, outperforming the strongest baseline by 1.10%, while remaining competitive on existing panoramic segmentation benchmarks. The source code and EmbPASS benchmark will be made publicly available at this https URL.

[CV-43] Uncertainty as a Proxy for Semantic Correctness in Diffusion-Based Medical Image Synthesis

链接: https://arxiv.org/abs/2610.03224
作者: Yuxuan Ou,Konstantinos Kamnitsas,OxAAA Study,AICT Consortium,Regent Lee,Vicente Grau
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Diffusion models can synthesise contrast-enhanced CT (CECT) from non-contrast CT (NCCT), avoiding contrast administration and its environmental and patient-access costs. However, visually realistic images are not necessarily anatomically correct, and the pixel-intensity and feature-space similarity metrics used to assess generation quality do not directly measure anatomical correctness. In this work, we investigate whether uncertainty can serve as a proxy for semantic correctness in diffusion-based medical image synthesis. We study NCCT-to-CECT synthesis using AortaDiff, a multitask diffusion framework that jointly generates CECT images and lumen segmentations. The segmentation output provides an explicit representation of the generated vascular anatomy, enabling segmentation-derived errors to be used as a quantitative measure of generation correctness. Six methods spanning weight (Ensemble, HyperDiff, BayesDiff), architecture-perturbation (MCDropout), generative-stochasticity (RDS) and input-perturbation (TTA) uncertainty are compared at the pixel, region and image levels, and for detection of clinically relevant out-of-distribution (OOD) cases. Uncertainty proves informative at all three spatial scales, remains informative on an external multi-centre dataset under distribution shift, and supports OOD detection. MCDropout stands out among the six: it ranks among the leading methods at every scale, generalizes well on the external dataset, and can be enabled at inference on any model already trained with dropout, so reliable uncertainty comes at no extra training cost. Uncertainty reliably flags severe failures but discriminates poorly among already high-quality images. These findings support uncertainty as a practical and computationally economical signal for quality filtering, reliability assessment and OOD detection in NCCT-to CECT synthesis. Subjects: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI) Cite as: arXiv:2610.03224 [cs.CV] (or arXiv:2610.03224v1 [cs.CV] for this version) https://doi.org/10.48550/arXiv.2610.03224 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[CV-44] VDOT: Unified Few-Step Video Generation via Unbalanced Optimal Transport Distillation

链接: https://arxiv.org/abs/2610.03221
作者: Yutong Wang,Xingtong Ge,Enhuai Liu,Yunke Wang,Tianfan Xue,Yu Qiao,Yaohui Wang,Xinyuan Chen,Chang Xu
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Video creation spans text-to-video (T2V), image-to-video (I2V), and condition-based generation, yet video diffusion models remain costly because they repeatedly evaluate large backbones during sampling. Distribution matching distillation (DMD) reduces this cost, but its reverse Kullback–Leibler (KL) objective can provide unstable or incomplete guidance when the student and teacher distributions have limited overlap. VDOT addressed this issue by adding optimal transport distillation (OTD), whose explicit coupling supplies geometric directions for condition-based generation. Balanced OTD, however, performs full-mass matching between the spatial tokens of each corresponding student–teacher frame pair. This assumption weakens for T2V and I2V, where one condition admits many valid outputs and spatial content need not align across different realizations. We present VDOT++, a unified distillation framework that applies the same training recipe separately to generators for the three task families. It makes OTD robust to output diversity through an asymmetric unbalanced formulation that allows unreliable student tokens to carry less mass while maintaining coverage of the teacher tokens. An \ell_1 ground cost further replaces mean-based aggregation with a more mode-preserving weighted median that limits the influence of distant transport targets. The two changes respectively determine whom to match and how the selected targets should be aggregated. We additionally combine distribution matching and adversarial refinement through sequential backward passes, and exploit the decoupled score networks for cross-scale distillation, where larger score networks improve a compact generator. Experiments on UVCBench, VBench, VBench-I2V, and the VACE benchmark show that the resulting four-step generators are competitive with many-step teachers and strong few-step baselines across all three task families.

[CV-45] Evolving Hybrid Quantum-Classical Architectures for Image Classification

链接: https://arxiv.org/abs/2610.03220
作者: Devroop Kar,Daniel Krutz,Travis Desell
类目: Neural and Evolutionary Computing (cs.NE); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Quantum Physics (quant-ph)
备注: Under Review at The Fifteenth International Conference on Learning Representations 2027

点击查看摘要

Abstract:Hybrid quantum classical neural networks integrate parameterized quantum circuits (PQCs) with established deep learning architectures, but their performance depends strongly on the choice of quantum circuit architecture, a choice that remains largely manual. Most existing approaches rely on hand-designed or fixed circuit ansätze, requiring circuit structure, gate composition, and qubit connectivity to be specified in advance with no guarantee that they suit the task. This limitation is especially acute in image classification, where quantum circuits must transform features extracted by classical networks while remaining compact enough for practical training, requirements that generic, task-agnostic ansätze are unlikely to satisfy simultaneously. We extend EXAQC, an evolutionary framework for automated quantum circuit discovery, to image classification. EXAQC evolves PQCs as intermediate processing modules while retaining classical feature-extraction and prediction layers. On MNIST, Fashion-MNIST, and CIFAR-10, EXAQC achieves 98.42%, 90.62%, and 85.47% accuracy, respectively, while using comparable gate counts to other quantum architecture-search methods. Against classical networks, evolved hybrid models maintain comparable accuracy with substantially fewer trainable parameters, reaching 85.68% on CIFAR-10 with over 25 \times fewer parameters than a 10-layer CNN. Encoding choice also matters: rotation-based encodings (RX, RY, U3) outperform amplitude encoding by 22-25 points on CIFAR-10. These results demonstrate that automated circuit discovery yields compact quantum modules that can replace larger classical components in vision architectures while retaining competitive accuracy.

[CV-46] VisionMX: Unlocking Microscaling Post-Training Quantization for Vision Models

链接: https://arxiv.org/abs/2610.03218
作者: Elad Dror Cohen,Ofir Gordon,Lior Dikstein,Idan Achituve,Hai Victor Habi
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Microscaling (MX) formats are emerging as a hardware-supported approach to efficient training and inference. They combine low-precision elements with shared block scales, but their impact on vision models remains underexplored. We systematically investigate post-training MX quantization across vision models and tasks. An analysis of direct conversion identifies three sources of error: block-scale representation, the poor alignment of some small convolutional weight tensors with nonuniform element grids, and the underuse of signed codes by nonnegative activations. These findings motivate VisionMX, a post-training MX quantization method that optimizes bounded weight rounding and applies a foldable affine correction to activations. We evaluate VisionMX across image classification, object detection, semantic segmentation, and low-light image enhancement using several MX-style formats. It improves on direct conversion and the evaluated post-training quantization baselines, with the largest performance recoveries in architectures most sensitive to MX conversion

[CV-47] Contextual Flow Matching: Adaptive Step Selection in Flow Models for Efficient Visual Generation NEURIPS2026

链接: https://arxiv.org/abs/2610.03202
作者: Divya Jyoti Bajpai,Arun Verma,Manjesh Kumar Hanawal
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注: Accepted in NeurIPS 2026

点击查看摘要

Abstract:Flow Matching enables high-quality visual generation via continuous-time dynamics, but inference remains costly due to multiple sequential function evaluations. Existing acceleration methods reduce the number of function evaluations but often introduce additional training overhead, degrade quality, or fail to account for input-dependent variability. We propose COFLOW, an inference-time method that adaptively selects the step counts each generation based on the prompt features. Our context-aware COFLOW is trained online with an unsupervised reward that balances inference efficiency and generation fidelity. Our method is plug-and-play, requiring no retraining of the underlying generative model. It generalizes to image and video generation, achieving over 2.5x speedup while preserving perceptual and semantic quality. We further provide a theoretical analysis establishing an O(1/K) forward-Euler discretization error bound under standard regularity conditions.

[CV-48] Bridging Research and Practice: A Systematic Evaluation of Generalist and Dermatology-Specific Models in Clinical Skin Lesion Classification MICCAI2026

链接: https://arxiv.org/abs/2610.03193
作者: Emanoel dos Santos,Kelvin Cunha,Rodrigo Mota,Fabio Papais,Thales Bezerra,Natalia Lopes,Erico Medeiros,Shirley Cruz,Jessica Araujo,Paulo Borba,Tsang Ing Ren
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 10 pages, 1 figure, 3 tables, approved at MICCAI 2026

点击查看摘要

Abstract:The application of machine learning to dermatology has grown substantially in recent years, moving beyond proof-of-concept studies toward potential applications. However, clinical dermatology remains a challenging and still open problem. Diagnostic assessment is often ambiguous, and skin lesions exhibit high variability, compounded by differences in acquisition modality, device quality, and patient demographics. These factors hinder the development of robust models suitable for safe and equitable clinical use. To support translation into practice, it is essential to systematically evaluate how contemporary models generalize across heterogeneous data sources. In this work, we benchmark a diverse set of architectures on recent dermatology datasets, spanning dermoscopic images and smartphone-based clinical photographs. We assess the robustness of recent general-purpose and medical vision-language models, as well as foundation models, and compare them against task-specific dermatology classifiers, including embedding-based approaches and convolutional neural networks. Our study provides an evaluation of model performance under distribution shifts, modality changes, and demographic variability. By quantifying the gap between current state-of-the-art models and the requirements of clinical deployment, we aim to contribute to the development of reliable, accessible, and clinically applicable AI systems for dermatology.

[CV-49] PocketSplat: Mobile Gaussian Reconstruction via World-Space Latent Allocatio

链接: https://arxiv.org/abs/2610.03192
作者: Wenzhi Guo,Xianda Chen,Dongxuan Chen,Guangchi Fang,Bing Wang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Mobile Gaussian reconstruction must satisfy two requirements: the reconstruction model must execute within a device resource envelope, and the resulting Gaussian asset must expose a representation size suited to downstream mobile use. Existing feed-forward Gaussian reconstructors commonly decode dense, image-aligned candidates whose final cardinality is implicitly determined by the input resolution and number of views. We present PocketSplat, a feed-forward framework for budgeted mobile Gaussian asset construction. Given a prescribed output budget, PocketSplat organizes dense geometry-aware latent candidates in predicted world space, allocates exact integer capacity across local latent cells, and decodes complete Gaussian attributes only for retained candidates. Cell-conditioned latent fusion aggregates repeated multi-view evidence before decoding, while spatial responsibility decoding adapts Gaussian support after local sparsification. Experiments on DL3DV and out-of-distribution benchmarks establish a strong quality–budget trade-off against feed-forward Gaussian reconstruction baselines. On Mip-NeRF 360, PocketSplat executes directly on a target iPhone and constructs compact, higher-quality Gaussian assets substantially faster than a deployable streamed MVSplat variant; native MVSplat and DepthSplat exceed the device memory budget.

[CV-50] Lightweight and Resource-Efficient Perception for Robotic Guide Dogs ACCV2026

链接: https://arxiv.org/abs/2610.03187
作者: Jinse Kwon,Yoojin Lim,Choonghan Lee,Yongseung Yu,Yongin Kwon,Jemin Lee
类目: Distributed, Parallel, and Cluster Computing (cs.DC); Computer Vision and Pattern Recognition (cs.CV)
备注: accepted in ACCV 2026

点击查看摘要

Abstract:Multi-camera streaming perception is increasingly deployed on heterogeneous edge platforms shared with co-resident workloads, yet accelerator placement is often evaluated using isolated single-stream experiments and mean streaming average precision (sAP). Using two end-to-end pipelines on a single GPU–NPU platform, we show that isolated evaluation can mis-rank deployment-time placement. Although the GPU pipeline is preferred in isolation, GPU-localized contention introduces deadline misses that make detections stale and can reverse the preferred placement before full GPU saturation. The NPU pipeline is less accurate than the GPU pipeline on small and medium objects in isolation, but nearly matches it on large objects. The largest absolute sAP losses in our latency and contention experiments occur for large objects. In our four-stream experiments, the preferred placement depends on which path becomes stale, and increasing GPU-side contention shifts the best placement from All-GPU to All-NPU. Under a GPU-saturating vision–language co-tenant, All-NPU achieves 5.2\times the worst-stream sAP of All-GPU. Because mean sAP can hide severe single-stream degradation, evaluation should report contention sweeps, deadline-miss rates on both paths, and worst-stream sAP alongside mean sAP.

[CV-51] Geometry-Aligned Semantic Matching for Cross-Modal Planar Image Registration

链接: https://arxiv.org/abs/2610.03167
作者: Zhiwei Wang,Defeng He,Yuxing Li,Meilu Zhu,Edmund Y. Lam
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Cross-modal image matching establishes stable and accurate geometric correspondences across modalities for planar registration. Existing semantic representations provide cross-modal consistency, but semantic similarity does not necessarily imply geometric correspondence. Meanwhile, fine-grained CNN features provide accurate local details but lack global cross-modal semantic guidance for stable refinement. To address these issues, we propose CDPM, which first establishes geometrically consistent semantic representations and then preserves their dominant role in correspondence estimation during fine-grained localization. Specifically, we progressively adapt DINOv3 using geometrically consistent cross-modal patch pairs, enabling feature similarity to better reflect true cross-modal spatial correspondences. We then construct a DINO-Centric Feature Pyramid, where multi-scale DINO representations maintain stable cross-modal correspondences, while a lightweight CNN branch provides auxiliary structural details for precise local refinement. Extensive experiments on three cross-modal datasets demonstrate the superior performance of CDPM. On VIS-IR, compared with the dense matcher RoMa, CDPM improves AUC@3/5/10/20 by 7.36, 13.40, 13.75, and 10.42 percentage points, respectively, and reduces mACE from 5.83 to 2.78 pixels. It also outperforms RoMa v2 across all metrics while requiring 45.6% fewer FLOPs. The online demo and dataset are available, and the code will be released on our project page at this https URL.

[CV-52] Budgeted-GS: Real-Time Large-Scale Gaussian Splatting via Factoring LOD

链接: https://arxiv.org/abs/2610.03162
作者: Haipeng Wang
类目: Graphics (cs.GR); Computer Vision and Pattern Recognition (cs.CV)
备注: 28 pages, 21 figures. Preprint of the EG 2027 submission (paper1075)

点击查看摘要

Abstract:3D Gaussian Splatting achieves excellent visual quality with real-time rendering, but at the scale of entire cities it does not fit: a trained model carries millions of primitives and gigabytes of memory, and real-time rendering at high quality on a consumer GPU remains out of reach. We introduce Budgeted-GS, a post-hoc method that turns any trained 3DGS model into a factoring tree, a multi-resolution hierarchy of moment-matched aggregates. After a construction pass of a few seconds, a single quality parameter selects, for each view, the level of detail that fits the memory of the target device, so the same city-scale model serves GPUs with widely different memory capacities. When a new scene is to be trained, the same theory applies: instead of growing a full-sized model and compressing it afterwards, budget-centered training first measures how many primitives the scene needs and then trains the model directly at that size, avoiding the wasted effort of optimizing primitives that are later discarded. Both methods are grounded in a measurable capacity floor, a budget-error law derived from optimal transport in phase space; selection rules certified by recent covering theorems decide which primitives are redundant. The floor answers how many primitives a scene actually needs and how many can safely be given up. We validate the floor on 13 public scenes under a preregistered protocol, and exercise both methods from object scenes to an official city capture, rendering it at native 1920x1080, full SH, in real time on one consumer GPU.

[CV-53] Does Physics Live in the Activations? Localizing Physical Quantities in Video Diffusion Models

链接: https://arxiv.org/abs/2610.03154
作者: Jonas Kneifl,Jakub Skalski,Bartłomiej Twardowski,Kamil Deja
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注: 22 pages, 8 figures, 5 tables

点击查看摘要

Abstract:Video generation models produce strikingly realistic sequences and are increasingly proposed as world models, yet recent benchmarks reveal pronounced deficits in their physical reasoning. This raises the question of whether these models internalize physical principles or merely reproduce familiar motion patterns. We address this by probing internal representations of video Diffusion Transformers (DiTs) for simulator-derived ground-truth physical quantities spanning kinematic motion and rigid-body dynamics under gravity and contact. We find that these quantities are linearly decodable with high accuracy early in the denoising process, substantially outperforming a baseline decoded directly from the model’s own noised latents, indicating that the relevant physical information is actively constructed during denoising rather than already present in the input. Additionally, we show that activations at on-object tokens carry the relevant physical information and that quantities defined over multiple frames are readable from single latent frames. Hence, information is sharply localized within the token sequence and is computed globally but stored locally. The probes further show partial extrapolation, transferring to scene variations and object configurations outside their training regime, so what they read is not simply a correlate of the scenes they were fit on. When fitted directly in the full-resolution activation space, the probing directions can serve as steering vectors to change the model’s output.

[CV-54] CalCErt: Bin-wise Certification of Confidence Calibration in Medical Image Classification

链接: https://arxiv.org/abs/2610.03142
作者: Leo Fillioux,Stergios Christodoulidis,Stergios Christodoulidis,Maria Vakalopoulou,Jose Dolz
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Deep neural networks remain vulnerable to adversarial perturbations, which can distort not only predictions but also confidence scores, undermining uncertainty calibration. While existing certification methods focus on preserving the predicted category, providing guarantees on how calibration behaves under adversarial attacks remains overlooked. In this work, we introduce CalCErt, a simple and efficient post-hoc strategy that certifies bin-wise confidence calibration for any pretrained differentiable classifier. Our approach combines empirical calibration estimates, statistical concentration bounds, and local Lipschitz estimates of the confidence function to derive data-dependent upper bounds on worst-case miscalibration within an ell_2 -ball of radius R. We evaluate CalCErt across 11 medical image classification tasks and multiple adversarial perturbations, demonstrating substantially higher certified coverage than baseline strategies while maintaining competitive tightness. Our code is available at this https URL.

[CV-55] Behavior Pack Optimization for Video MLLM Post-Training NEURIPS2026

链接: https://arxiv.org/abs/2610.03141
作者: Zhaolu Kang,Shiyu Liu,Tailong Luo,Wei Zhang,Yingjie He,Lei Wei,Guansu Wang,Liang He,Siheng Wang,Guangyuan Dong,Jiaqi Su,Shuang Chen,Haoyu Ji,Qishi Zhan,Kaiyue Zhou
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: NeurIPS 2026 poster

点击查看摘要

Abstract:Video multimodal large language models (MLLMs) keep climbing video question answering benchmarks, yet shuffling the frames, masking the segment that supports the answer, or occluding the target object barely changes their predictions. The accuracy rests on appearance and language priors, not on the temporal evidence the question asks for. We trace this to the unit of post-training: rewards are computed on a single response to the original clip, so the model is never asked to behave consistently across views. We propose Behavior Pack Optimization (BPO), which replaces the single response with a behavior pack of outputs across counterfactual views chosen by question type, scored jointly. The pack reward asks for stability when the intervention is irrelevant, sensitivity when key evidence is removed, and abstention when no evidence remains. To keep this objective stable at small pack sizes, BPO uses an anchor-relative advantage: the response on the original view serves as a per-prompt reference instead of a group mean over mixed views. On TempCompass, MVBench, and NExT-QA, BPO improves the macro accuracy of Qwen2.5-VL-7B-Instruct by 4.7 pp, the temporal-hard subset by 7.8 pp, and abstention F1 by 20.0 pp over a budget-matched vanilla GRPO baseline from the same SFT checkpoint. The gains transfer to Video-MME, LongVideoBench, and to LLaVA-Video-7B; ablations confirm they follow the view sets, not the rollout count. We hope this pack-level perspective offers a useful starting point for the video MLLM and multimodal post-training community as the field moves toward evidence-grounded video reasoning.

[CV-56] Foresight: planning future perception in streaming VLMs without retraining

链接: https://arxiv.org/abs/2610.03123
作者: Ashok Prasad Neupane,Dipan Bartaula,Ankit Belbase,Saugat Adhikari,Samip Ghimire,Saroj Poudel,Binod Bhattarai,Danda Pani Paudel
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Existing streaming vision-language models (VLMs) continuously perceive and reason over visual streams, but their computational pathways remain fixed throughout inference. Consequently, they cannot adapt computation to evolving scene dynamics, where different future events demand different levels and forms of perception. We show that streaming VLMs inherently possess the ability to anticipate the immediate future, and leverage this capability to dynamically configure future computation in a training-free manner. Realizing such anticipatory computation, however, is very challenging: future anticipation must be sufficiently reliable to guide computation, planning must run concurrently with streaming inference, and online reconfiguration must incur negligible overhead. To address these challenges, we introduce FORESIGHT, a dual-stream architecture comprising two Siamese LLMs with shared weights, input encoders, and KV cache. The first LLM continuously processes incoming tokens, while the second runs ahead of the stream to anticipate future context, plan future computation, and generate task responses without interrupting streaming inference. Each plan decides when to reason next, what to check then, and how densely to sample, keeping transient evidence separate from persistent control. The resulting computation plan is executed online through an efficient reconfiguration protocol with schemaguided decoding and lightweight diff-based updates, enabling dynamic adaptation with low overhead. With a frozen Qwen3-VL-8B backbone, FORESIGHT achieves 23.0 mean joint F1 on OmniPro Online evaluation beating strongest trained baseline by 9.5%, while improving the backbone by 6.7 on StreamingBench and 15.4 on OVO-Bench, with the largest gain of 18.7 when evidence arrives later in the video stream. Our source code will be made publicly available.

[CV-57] In-Distribution Forcing for Long Video Generation at Test Time

链接: https://arxiv.org/abs/2610.03120
作者: Jeongwoo Shin,Youngyoon Choi,Sangwoo Jo,Hyunmog Kim,Sungjoon Choi,Joonseok Lee,Jaewoong Choi,Jaemoo Choi
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Preprint

点击查看摘要

Abstract:Modern autoregressive (AR) video diffusion models excel at short-horizon video generation, yet generating long videos remains challenging due to drifting, where colors and textures shift, and motion dynamics decay. Existing works primarily rely on KV conditioning, which selects or modifies cached key-value (KV) entries to mitigate drifting. However, we observe that KV conditioning alone is insufficient as it assumes cached KV entries remain in-distribution. This assumption fails beyond the training horizon: nothing constrains the construction of KV entries during rollout, giving rise to the KV-provenance problem where cached entries themselves become out-of-distribution (OOD). To address this, we propose In-Distribution Forcing (ID-Forcing), a test-time framework that aligns both KV caching and KV conditioning with training configurations. Its key mechanism, self-caching, prevents OOD KV entries at their source. Each chunk is cached without attending to prior KV entry, keeping the rolling window exactly in-distribution. Consequently, ID-Forcing seamlessly extends short-horizon models to minute-scale video generation. Extensive evaluations show that our method remains competitive on standard video generation benchmark while substantially outperforming prior work in mitigating drifting, as validated by both our drift metrics and a user study.

[CV-58] Beyond Single Videos: Benchmarking and Active Evidence Seeking for E-Commerce Cross-Video Reasoning

链接: https://arxiv.org/abs/2610.03099
作者: Jinghan Zhao,Yiman Hu,Liang Wu,Jian Xu,Bo Zheng
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:E-commerce videos are information-dense and frequently compared by consumers evaluating products and merchants assessing marketing strategies. However, existing multimodal models mainly focus on single-video understanding and have limited ability to compare information across videos. We introduce AdsCVR, the first e-commerce cross-video reasoning benchmark, containing 2,483 videos and 6,110 question-answer pairs across six reasoning dimensions. Cross- video reasoning requires models to locate fine-grained evidence among many redundant frames and integrate visual details, speech, and on-screen text. We therefore propose AdSeek, an agentic framework that dynamically selects visual and audio tools during multi-turn exploration, replacing static uniform sampling with active evidence acquisition. To address the sparse credit assignment of reinforcement learning, we develop an offline trajectory rectification mechanism that identifies reasoning errors and missing multimodal evidence in RL-generated trajectories. The corrected trajectories provide supervised fine-tuning signals that reduce biases learned during RL. This mechanism supports a rectified bootstrapping pipeline in which initial RL exposes reasoning bottlenecks, supervised fine-tuning corrects them, and a final RL stage further improves the policy. AdSeek achieves 74.30 percent accuracy on the AdsCVR test split, outperforming its Qwen3-VL-8B-Instruct backbone by 27.90 percentage points. It also generalizes to the open- domain CrossVid benchmark, demonstrating effective active evidence gathering.

[CV-59] NegT2IBench: When Negation Changes the Picture. A Polarity Benchmark for Text-to-Image Models

链接: https://arxiv.org/abs/2610.03084
作者: Omar Elfatairy,Maria A. Bravo,Jessica Bader,Zeynep Akata
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: *Equal contribution

点击查看摘要

Abstract:Text-to-image (T2I) models are judged by benchmarks that measure whether requested content appears, but these benchmarks largely overlook the complementary ability to satisfy negated constraints, for example, generating “a non-red cup.” Measuring negation raises challenges not faced by affirmation-based benchmarks and requires careful prompt and evaluation design. We introduce NegT2IBench, a benchmark of 4,800 prompts covering two attribute types and four relation categories. Prompts are organized by polarity: the number of positive statements that must hold and negated statements that must not, each ranging from 0 to 2. Varying the two independently separates the effect of negation from the effect of prompt complexity. Our detector-based scoring is reproducible, auditable, and pinpoints which requirement failed. On 600 images with three-annotator labels, it agrees with humans as closely as vision-language judges up to 30x larger, while using only a fraction of their GPU memory. Across eleven T2I models and 211,200 images, nine score lower on a single negated statement than on a single positive one. Per-statement scoring reveals that the loss is largest for color and near zero for proximity, and that 41.5% of failed statements render exactly what the prompt forbids. Rendering what a prompt asks for and withholding what it forbids are distinct capabilities that an aggregate compositional score cannot distinguish. NegT2IBench measures the latter directly, providing a controlled testbed for diagnosing negation failures and developing methods to overcome them.

[CV-60] Where to Look Is Not How to Fix: Pre-Denoising Diagnostics and Modality-Dependent Control in Diffusion Composition

链接: https://arxiv.org/abs/2610.03068
作者: Fangzheng Wu,Brian Summa
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Understanding compositional failures in text-to-image diffusion requires identifying both where stress is detectable and how intervention changes the output. We study these questions through a controlled anchor–stress protocol that jointly evaluates text-encoder diagnostics and denoiser interventions. We introduce a text-only Compositional Stress Index (CSI), which separates common from rare compositions across SD1.5, SDXL, and the SD3 text path and provides an upstream diagnostic coordinate. A matched six-prompt localization study links intervention location to distinct outcomes: residual-minimizing embedding adapters improve representation fit, while downstream cross-attention intervention increases color hit rate (CHR) by 0.0272. Across SD1.5 and SDXL denoiser blocks, the largest positive signed diagnostic-accessibility mean occurs at the deep encoder, whereas selective boost has its largest positive mean CHR response at decoder blocks. Selective subtraction and broad ablation reveal further modality- and architecture-dependent responses, including a substantial CHR decrease when SDXL decoder cross-attention is broadly ablated. We find a diagnosis-control dissociation under our controlled attribute-object composition setting: compositional defects are diagnosable before denoising, but the representation coordinate that exposes risk is not necessarily the coordinate or modality that improves generation.

[CV-61] BeeWhere: Segmenting Bumble Bee Colonies to Quantify Behavioral Effects ECCV2026

链接: https://arxiv.org/abs/2610.03051
作者: Roberta Hunt,August Easton-Calabria,James Crall
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Preprint. Accepted to ECCV 2026 Computer Vision for Ecology Workshop Proceedings. Proceedings DOI pending

点击查看摘要

Abstract:Social bees are important pollinators that support biodiversity and crop pollination globally and serve as important model systems for collective behavior, but scalable measurement of individual- and colony-level behavior remains difficult in dense, occluded nest environments. Existing monitoring workflows use fiducial tags (e.g., ArUco) to preserve individual identity, yet tag-based tracking can fail when markers are obscured and provide limited information about body extent, spatial context, and untagged individuals. We present BeeWhere, an AI-assisted annotation and analysis workflow that combines ArUco detections with deep-learnt instance segmentations to quantify bumble bee behavior from high-resolution colony images and videos. Using bumble bee (Bombus impatiens) microcolonies as a test case, we annotate 483 frames containing 8,443 bee instances. We additionally annotate pollen balls, nest structures, and chamber boundaries, and train YOLO instance segmentation models for downstream behavioral analysis. Instance segmentations enable quantification of important behavioral metrics based on body contours, including nearest-neighbor distance, proximity to nest structures, spatial occupancy within the nest, and detection counts over time. We apply the BeeWhere models to tag-based tracking in an exploratory validation study assessing the behavioral impacts of neonicotinoid pesticide exposure. BeeWhere increased detection rates compared to tag-based tracking, particularly when bees were partially obscured or under challenging imaging conditions, and also captured treatment-associated changes in bee spatial organization not captured using tag-based tracking alone. These results suggest that instance segmentation can complement fiducial-marker tracking by recovering behaviorally meaningful signals under challenging colony conditions.

[CV-62] Parasitic Co-Denoising: Unlocking 3D Human Motion Generation in a Frozen Video Diffusion Model

链接: https://arxiv.org/abs/2610.03047
作者: Yunjiao Zhou,Junlang Qian,Lihua Xie,Jianfei Yang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Despite never being supervised on explicit 3D motion, large-scale text-to-video diffusion models synthesize realistic human motion in their generated videos. We ask whether this implicit knowledge can be turned into explicit 3D motion generation, without training a separate motion model. Probing a frozen Wan2.1 reveals that a recoverable motion signal is present in its intermediate states across the entire denoising schedule, not confined to the clean output. Motivated by this, we introduce parasitic co-denoising, a paradigm in which motion is decoded from the host model along its denoising schedule rather than produced by an independent generator. We instantiate it as the Parasitic Motion Decoder (PMD), an efficient flow-matching decoder that shares the host’s noise schedule and reads its intermediate features through a \sigma -adaptive multi-layer fusion, leaving the host unmodified. Drawing its coverage from the host rather than from motion data, PMD leads dedicated motion generators on text-motion alignment at a small fraction of their trainable parameters, while producing paired video and motion in a single pass that motion-only baselines cannot match.

[CV-63] WebFovea: When the Model Is Right but the Click Is Wrong – Reliable Round Trips for Vision-Based Web Agents on Live Websites

链接: https://arxiv.org/abs/2610.03036
作者: Jiangang Han
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
备注: 10 pages, 4 figures, 7 tables. Technical report of the 2nd-place solution in the WebRetriever Challenge 2026

点击查看摘要

Abstract:We present WebFovea, a vision-based web agent that placed 2nd in the WebRetriever Challenge 2026 with a final score of 57.0 out of 100. The challenge evaluates agents end to end on Protocol III of the WebRetriever benchmark (arXiv:2607.06118): starting from an entry URL on a live website, the agent must operate the site’s own interface and return a verifiable answer. A capable multimodal large language model (LLM) is necessary for this, but not sufficient. The model’s decisions reach the browser through the harness, the code between the model and the page. At every step, four things must go right: the model’s reply must be parsed into the intended action, the action must take effect on the page, the result must be reported back accurately, and the model must be shown the information it needs. On real websites, many of the failures we observed occurred at one of these four stages rather than in the model’s reasoning. A coordinate-space mismatch placed every click at 3/4 of its intended coordinates; actions on native dropdowns, inside iframes, and in text boxes failed silently; and self-generated chat-template tokens contaminated 4.9% of task episodes. WebFovea hardens each stage and surrounds the loop with guardrails that keep the agent within the rules and its budget. The four-stage view does not depend on the model, although some individual fixes do. Because we used the same model in all four submissions, the rise of our official hidden-set score from 31.0 to 57.0 reflects changes to the harness, up to run-to-run variance on live sites. We describe the design, the evidence for each component (including negative results), a failure analysis, the limitations, and a roadmap that includes routing different steps to different models.

[CV-64] CrowdOcc: Monocular Semantic Scene Completion for Quadruped Robots in Crowded Indoor Environments ICRA

链接: https://arxiv.org/abs/2610.03031
作者: Feiyang Chen,Jincheng Hu,Yiduo Chen,Jihao Li,Yue Liang,Bingzhao Gao,Yanjun Huang,Yuanjian Zhang
类目: Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)
备注: 8 pages, 4 figures. Submitted to IEEE International Conference on Robotics and Automation (ICRA) 2027

点击查看摘要

Abstract:Monocular semantic scene completion (SSC) for quadruped robots remains underexplored in real crowded indoor environments, where human-scene occlusion disrupts static geometry and human occupancy predictions are often incomplete or spatially misplaced. We present CrowdOcc, an RGB-D dataset and monocular SSC framework for this setting. CrowdOcc contains 25.1K frames from 11 indoor scenes, with semantic occupancy annotations constructed through static dynamic decoupling. Our framework combines: (i) Normal Guided Scene Geometry Fusion (NGSGF) to complement depth-aware lifting with surface-normal cues for occlusion robust geometry; and (ii) Human-Centric Sparse Interaction (HCSI) to selectively model human-human and local human scene relations in 3D. Our method achieves state-of-the-art SSC performance on CrowdOcc’s scene-disjoint test set, reaching 15.80 IoU, 11.40 mIoU, and 46.23 Human IoU, demonstrating generalization to unseen indoor scenes.

[CV-65] From Expression to Reaction: Role-aware Visual Transfer and Stimulus-guided Reasoning for Interlocutor Emotion Recognition ACM-MM2026

链接: https://arxiv.org/abs/2610.03016
作者: Wei Wang,Zhaowu Li,Jianjie Luo,Fu Lee Wang,Lap-Kei Lee,Zhenguo Yang
类目: Multimedia (cs.MM); Computer Vision and Pattern Recognition (cs.CV)
备注: Technical report of the second-place solution in Track 1 (MER-Cross) of the MER Grand Challenge at ACM MM 2026

点击查看摘要

Abstract:In this paper, we propose a Role-aware Stimulus-guided (RASG) framework for interlocutor emotion recognition, which predicts listener emotions from listener-only videos and speaker-only audios. RASG consists of Role-aware Visual Transfer (RVT) and Stimulus-guided Boundary Reasoning (SBR) modules, which address supervision mismatch due to the lack of labeled listener data and ambiguity among visually similar listener reactions whose interpretation depends on speaker context, respectively. More specifically, RVT selects speaker samples whose facial expressions support their emotion labels. It then filters listener tracks and uses reliable pseudo-labels to train a listener-centric visual expert. SBR uses a two-class language reasoner only when the visual model is uncertain. It treats speaker audio and text as context rather than direct emotion evidence to distinguish similar listener reactions. Experiments conducted on MER-Cross dataset shows that RASG achieves 76.25% on MER-Cross and improves the performance of the baseline over 17%. Our team ranks second in Track 1 (MER-Cross) of the MER Grand Challenge at ACM MM 2026.

[CV-66] OmniAct3D: Leverag ing Foundation Geometry and Evidence-Grounded Reasoning for Panoramic 3D Detection

链接: https://arxiv.org/abs/2610.03015
作者: Runtong Wu,Fei Teng,Di Wen,Guoqiang Zhao,Kunyu Peng,Kailun Yang
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Robotics (cs.RO)
备注:

点击查看摘要

Abstract:Accurate 3D detection is essential for mobile embodied agents, while Vision Foundation Models (VFMs) offer transferable visual and geometric priors. Yet existing VFM-based 3D detectors rely on narrow-view monocular images or discrete perspective views, limiting coherent surround perception; equirectangular projection (ERP) instead encodes a continuous 360 scene in a single image. Direct transfer remains difficult because ERP organizes geometry and visual information differently, making object-relevant cues hard to model, localize, and preserve. We propose OmniAct3D, a framework that adapts perspective-trained VFM detectors to ERP while preserving transferable VFM priors. To resolve geometric mismatch, the ERP-Ray Geometry Adapter (ERGA-Ray) models spherical viewing rays and periodic spatial structure. To localize evidence in scene-wide context, the Visual-Action Reasoning Chain (VARC) grounds each hypothesis in relevant panoramic evidence and converts it into a structured geometric action. To recover local cues lost under fixed token budgets, the Appearance-Guided Heading Expert (AGHE) re-encodes object regions at higher resolution for heading estimation. Experiments show that OmniAct3D improves over the previous best 3D detector by 2.96 NDS points on Spheriverse and over the unadapted VFM baseline by 24.87 mAP points on PanoMMOcc. With target-specific geometry adaptation, VARC retains 95–98% of the same-configuration mAP, indicating reusable object-level 3D reasoning across sensing configurations. The source code will be made publicly available at this https URL.

[CV-67] RYOPO: Bringing End-to-End Category-Level Object Pose Estimation into Real Time

链接: https://arxiv.org/abs/2610.03013
作者: Hakjin Lee,Junghoon Seo,Jaehoon Sim
类目: Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)
备注: Project page: this https URL

点击查看摘要

Abstract:Category-level object pose estimation predicts the rotation, translation, and metric size of unseen instances within known categories. Many accurate RGB-D methods rely on external instance segmentation and crop-based pose estimation, introducing separate stages and object-dependent processing costs that hinder real-time inference. To bring accurate pose estimation into real time, we present \ours, an end-to-end trainable query-based RGB-D set predictor. It jointly detects and segments objects and estimates their \mbox9-DoF poses without explicit CAD-derived shape priors or a separately trained instance segmentor. Shared image and scene encoding avoids repeated per-object crop encoding. A query-conditioned geometry pathway associates observed 3D points and RGB features with object queries and incorporates shared scene context. Object-centric refinement uses the resulting point descriptors to update an explicit pose state through pose-conditioned cross-attention and recurrent residual corrections. On NOCS, \ours substantially improves on published RGB-D joint detection and pose estimation results. It achieves competitive performance compared with two-stage methods under all-object evaluation on REAL275 and HouseCat6D, while enabling real-time full-frame pose estimation at 31.8 FPS on an RTX~A6000. Project page: this https URL.

[CV-68] Rethinking Fixed Temporal Grids: Frequency-Disentangled Motion Generation

链接: https://arxiv.org/abs/2610.03012
作者: Yunjiao Zhou,Junlang Qian,Gen Li,Xinying Guo,Lihua Xie,Jianfei Yang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Most human motion generation methods encode motion as tokens on a uniform temporal grid, where every token spans the same fixed time window. Human motion, however, is temporally heterogeneous: slowly evolving global trajectories coexist with rapid transient events such as foot contacts and joint impulses. Forcing such multi-scale dynamics onto tokens of identical temporal resolution entangles motion frequencies, leaving slow regions redundant while smoothing out the rapid details that distinguish realistic motion. We propose \textbfFreqMo, a scale-adaptive motion representation that decomposes motion into wavelet frequency bands, separating dynamics across temporal scales while preserving temporal localization and exact reconstruction. Unified Frequency Residual Quantization (UFRQ) then encodes all bands within a single shared codebook, compressing the token sequence threefold and enabling stable single-stage generation. Experiments show FreqMo attains SOTA fidelity with substantially improved high-frequency preservation, and the same decomposition transfers to continuous diffusion backbones.

[CV-69] From Language Priors to Field Adaptation: Preference Learning for Traversability Estimation

链接: https://arxiv.org/abs/2610.02974
作者: Simon Schwaiger,David Seyser,Alessandro Scherl,Zlatan Ajanović,Wilfried Wöber,Gerald Steinbauer-Wagner
类目: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Image-based traversability estimation is inherently dependent on the robot platform, deployment domain, and mission preferences, which limits the applicability of purpose-trained models. To facilitate domain adaptation, this work aims to reduce the number of required annotations in the target domain using sample-efficient preference learning. Our method represents traversability through von Mises-Fisher mixture prototypes in a frozen vision-language feature space. Relative natural-language rules provide a commonsense prior, while sparse relative image annotations adapt the prototype directions and utilities to a target domain through computationally and sample-efficient fine-tuning. Experiments on WayFAST demonstrate accuracy competitive with end-to-end trained estimators while enabling sample-efficient image-based adaptation. Qualitative experiments further demonstrate the language prior’s zero shot applicability and the fine-tuned estimator’s improved dense prediction on semantic maps. Evaluation is complemented via semantic interpretation of learned prototypes by dissecting semantically close natural language prompts. Code and trained estimators available at this https URL

[CV-70] Post-Training Frontier Text-to-Image Models by Composing Preference and Rubric Rewards

链接: https://arxiv.org/abs/2610.02967
作者: Yuanhao Ban,I-Hung Hsu,Anastasios Angelopoulos,Wei-Lin Chiang,Ion Stoica,Cho-Jui Hsieh
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Recent text-to-image generation models have achieved remarkable visual quality, but improving them through post-training remains challenging because no single reward signal captures the full range of human preference. In this work, we develop a simple and effective post-training recipe for open-domain text-to-image generation based on the composition of complementary reward signals. Our reward system consists of two main components: a preference reward, trained on large-scale human preference data using a Bradley-Terry objective to capture overall human aesthetic and perceptual preferences, and rubric-based rewards, which explicitly evaluate prompt faithfulness and other desirable properties while providing safeguards against reward hacking. A key challenge is how to combine these heterogeneous reward signals. We show that a naive weighted average leads to suboptimal optimization behavior, and propose a simple reward composition strategy that more effectively balances preference optimization with rubric satisfaction. In the Arena text-to-image leaderboard (this https URL), our RL-trained Flux2dev achieves an Elo rating 69 points above the base model, and our post-trained Ideogram-4 surpasses every open-source model on the leaderboard, reaching an Elo of 1223.5. (Claims of state-of-the-art performance are based on the Arena leaderboard snapshot as of September 4, 2026.) Our results suggest that effective rewards for frontier generative-model training require broad coverage of user intent and robustness to exploitation under optimization. To support reproducible research, we release Arena-T2I-Training, a 1K subset of training data that recovers some gains of full-scale training, providing a resource that we hope will facilitate future work on post-training for text-to-image models.

[CV-71] rraVis: Towards Evaluation of World-Grounded Visual Consistency in Text-to-Image Generation via MLLM Workflows NEURIPS2026

链接: https://arxiv.org/abs/2610.02959
作者: Shuai Fu,Jing Gu,Jian Zhou,Zicheng Duan,Gengze Zhou,Qi Wu
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted by NeurIPS 2026 (ED Track)

点击查看摘要

Abstract:Recent text-to-image models have made substantial progress in photorealism, aesthetics, and text-image alignment. Yet visually appealing images can still violate real-world plausibility, exhibiting malformed object structures, impossible anatomy, physically implausible interactions, or inconsistent spatial relationships. Such failures are not well captured by existing fidelity, aesthetics, preference, or alignment metrics. To address this gap, we introduce TerraVis, a framework for evaluating world-grounded visual consistency in generated images. TerraVis defines a structured taxonomy of world-consistency violations spanning object-, interaction-, and scene-level failures, and employs a multi-stage evaluation framework to identify and quantify them. Given an image, TerraVis first uses an MLLM to assess its eligibility for evaluation, then detects violations across 18 taxonomy-defined types and classifies them as minor or major to derive an overall world-consistency score. Across diverse open-source and proprietary text-to-image models on two widely used benchmarks, TerraVis achieves the strongest correlation with human judgments of world consistency among existing metrics. Our benchmark results further show that models that achieve strong performance on conventional metrics can still exhibit substantial world-consistency failures. These findings highlight world consistency as a complementary evaluation dimension and demonstrate that TerraVis enables systematic quantification, diagnosis, and comparison of such failures. Our code is publicly available at this https URL.

[CV-72] When Predicting Nothing Beats SAM 3: Revisiting Evaluation in Video Object Segmentation NEURIPS2026

链接: https://arxiv.org/abs/2610.02946
作者: Jihwan Hong,Woohyeon Park,Jaeik Kim,Jaeyoung Do
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: NeurIPS 2026 ED

点击查看摘要

Abstract:Video Object Segmentation (VOS) in complex and long videos is increasingly important for real-world applications, where target objects often appear only intermittently within long temporal horizons. However, existing benchmarks largely focus on temporally salient objects that remain visible for most of the video. To address this gap, we introduce FaVOS (A Benchmark for Video Object Segmentation with Fractional Temporal Visibility), a benchmark designed to evaluate VOS methods under low temporal visibility. We show that, in this regime, the standard JF metric can collapse VOS evaluation into absence classification, because empty predictions receive high rewards on target-absent frames. Consequently, even a trivial empty-mask predictor can outperform strong models such as SAM 3, revealing a fundamental mismatch between current metrics and practical VOS performance. To mitigate this issue, we propose Volumetric JF, which evaluates mask sequences as spatio-temporal volumes and reduces the dominance of target-absence rewards while preserving sensitivity to segmentation quality and temporal structure. Project page: this https URL.

[CV-73] Kinematics-Induced Multimodal 3D Human Pose Estimation with Subject-Level Privacy

链接: https://arxiv.org/abs/2610.02943
作者: Kaushik Bhargav Sivangi,Fani Deligianni
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Multimodal 3D Human Pose Estimation (3D HPE) combines complementary information from RGB, LiDAR, and mmWave radar, but models trained on correlated observations from the same individuals, raise privacy risks overlooked by record level analysis. We present a unified framework for multimodal 3D HPE that couples kinematics-induced sensor fusion with subject level privacy auditing and private training. First, our multimodal model aligns modality specific joint representation, injects skeletal structure and adaptively aggregates complementary sensor evidence for accurate pose prediction. Second, we formulate a black-box subject membership inference attack for 3D HPE, complemented by an empirical pointwise maximal leakage analysis, which characterizes how individual attack score outcomes change inference about the membership outcome. Third, we instantiate user-level differential privacy via Action Temporal Stratification, a population weighted within-subject sampling strategy that enforces action and temporal coverage. We evaluate our framework on the MM-Fi dataset across three diverse experimental protocols. Source-code will be released upon acceptance.

[CV-74] Custom Forcing: Training-Free Subject Customization for Autoregressive Video Generation

链接: https://arxiv.org/abs/2610.02914
作者: Yunseung Ok(1),Hyunsoo Kim(2),Minseo Kim(1),Suhyun Kim(1) ((1) Kyung Hee University, (2) The University of Texas at Austin)
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 31 pages. Project page: this https URL

点击查看摘要

Abstract:Autoregressive video models can generate minute-long videos in real time, but they produce generic subjects from text rather than specific subjects from user-provided images. Existing customization methods either require costly per-subject optimization or use pretrained conditioning networks that jointly process all video frames with bidirectional attention. Neither approach is designed for causal streaming. We present Custom Forcing, a training-free method that stores reference-based anchor frames in the persistent KV cache of a frozen autoregressive video model. However, fixed anchors face two limitations: simple conditioning allows identity to drift, and the text prompt continues to favor a generic subject. To address these problems, drift-adaptive value amplification (DVA) scales reference influence with the degree of identity drift, while anchor contrast guidance (ACG) steers generation away from the generic class prior. Over two-minute rollouts, fixed anchors fall from 0.58 to 0.42 in DINO-I, while Custom Forcing keeps it between 0.58 and 0.62 without reducing motion. Custom Forcing also achieves higher subject similarity than bidirectional customization methods and better preserves identity over 30s than causal image-to-video and reference-to-video models, while generating each frame 9.5–28.5 times faster than these long-video baselines.

[CV-75] ViTok: Improving Dense Semantics in AM-RADIO-Style Multi-Teacher Distillation with PHI-S and Masked Image Modelling

链接: https://arxiv.org/abs/2610.02903
作者: Hailun Xu,Kanchan Sarkar
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:We study how to consolidate the current VITOK progress into a single multi-teacher distillation recipe that jointly preserves global recognition and dense semantics. Our starting point is an AM-RADIO-style student distilled from SigLIP2 and DINOv3-L, where SigLIP2 supplies strong global semantics and DINOv3-L supplies stronger dense features. The central empirical issue is that the same recipe does not optimize all objectives equally well: changes that improve ImageNet-1K kNN accuracy can still degrade ADE20K segmentation. We summarize a progression of modifications that make this trade-off more explicit and more manageable: split adaptor heads for CLS and patch tokens, asymmetric cosine/MSE losses, initialization from a DINOv3-L checkpoint, teacher reweighting, masked image modeling (MIM), and PHI-S feature balancing. The resulting model reaches 83.2 patch kNN and 85.2 CLS kNN, slightly surpassing the DINOv3-L teacher on ImageNet-1K kNN classification, while PHI-S restores ADE20K performance from 46.5/58.1 to 48.5/61.0 mIoU/mAcc, matching the teacher on this dense benchmark. We also summarize negative results: scaling distillation from ImageNet-1K to ImageNet22K does not consistently help, and naively adding extra teachers such as SAM3 or HOG features introduces interference. Rather than claiming a final recipe, this paper distills the current project state into a compact empirical story and a concrete set of lessons for future iterations.

[CV-76] Revealing Epistemic Uncertainty in MLLM s via Causal-Invariant Masking NEURIPS2026

链接: https://arxiv.org/abs/2610.02887
作者: Haoyang Luo,Linwei Tao,Jie Gui,Xinghao Chen,Chang Xu,Jianyuan Guo,Minjing Dong
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注: Accepted by NeurIPS 2026

点击查看摘要

Abstract:Multimodal Large Language Models (MLLMs) suffer from hallucinations, creating a critical need for Uncertainty Quantification (UQ) to ensure reliable deployment. However, existing approaches struggle to detect uncertainty caused by superficial associations, especially when the query-relevant signal is weak. We mainly attribute this issue to their bias toward aleatoric uncertainty arising from data ambiguity, overlooking epistemic uncertainty stemming from model limitations. To further decompose uncertainty types for a comprehensive UQ, we propose Causal-Invariant Masking (CIM), which measures the semantic shift between the original predictions and those conditioned on a causally-focused view. Based on this framework, we introduce Semantic Divergence as our core metric for UQ and provide theoretical evidence that it converges to the variance of model’s sensitivity to non-causal correlations, establishing its ability to capture MLLM’s limitation. To accelerate UQ in MLLMs, we further propose Expected Embedding Drift (EED), a fast geometric proxy metric that estimates semantic shift directly within the hyperspherical embedding space. Experiments show that our method achieves state-of-the-art performance on various benchmarks, while the proposed EED accelerates by nearly 50% with comparable performance.

[CV-77] Found but Not Read: When Extracted Text Closes the Retrieval-Reading Gap in Document Vision-Language Models ICASSP2027

链接: https://arxiv.org/abs/2610.02880
作者: Qingtao Xia,Siyao Cheng,Jiahua Bao,Jiaxing Du,Jie Liu
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 5 pages, 5 figures, 3 tables. Submitted to ICASSP 2027

点击查看摘要

Abstract:Retrieval-augmented document question answering assumes that once the right page is found, a vision-language model (VLM) can read it. We show that this assumption often fails, leaving a retrieval-reading gap: evidence found but not used. A paired protocol isolates this gap by comparing answers from the retrieved page images alone with answers from the same images plus their extracted text. On FoveDoc-Bench, our benchmark with traceable evidence, retrieval finds nearly every evidence page, yet adding CPU-OCR text raises strict accuracy by 13 to 16 points. An exact text layer roughly doubles the gain, which appears across six VLMs from three families and, within one family, narrows with scale without closing. The reader can read this evidence but cannot find it: crops of it recover most of the text gain, boxes around it on the page none. The same protocol identifies two boundaries. Extracted text helps on textual evidence but is neutral or harmful on charts and figures. Its advantage shrinks as retrieval degrades, and unrelated text of the same form adds nothing detectable. Extracted text is an amplifier of retrieval that works, not a substitute for retrieval that does not. Our code is available at this https URL.

[CV-78] Seeing Saying but Not Using: From Reportable Spatial Facts to Usable States in Multimodal Large Language Models

链接: https://arxiv.org/abs/2610.02876
作者: Jinchang Zhang,Guoyu Lu
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:A multimodal large language model that correctly reports a spatial fact does not necessarily use that fact in subsequent reasoning. To study this distinction, we introduce \textscSpaceConflict, a benchmark of 23,196 inputs for the construction and use of spatial state. Under a unified Supported/Contradictory/Unknown judgment interface, it covers local fact binding (L1), relational composition (L2), cross-observation consistency (L3), and state judgment under transformation (L4). Posing a direct-state query, a full-transformation query, and an explicit-initial-state query on the same world reveals an availability–utilization gap: models recover the initial state from visual evidence yet fail when that state must drive a transformation. For Qwen3.5-9B, 50 of 100 sequences with a correctly recovered initial state fail the full transformation, and supplying the state explicitly repairs all 50; the gap narrows with scale but does not close. We therefore propose Operational State Supervision (OSS), which supervises task-relevant spatial states and their transformation trajectories and aligns shared facts across contexts. OSS improves paired accuracy on matched judgments most on L3 and L4, the levels that depend on organizing and using state. Evaluating multimodal spatial reasoning thus requires asking not only whether a model can see and state a spatial fact, but whether that fact becomes a usable state in subsequent computation.

[CV-79] PointWAM: 3D World Action Modeling for Dexterous Robotic Manipulation

链接: https://arxiv.org/abs/2610.02840
作者: Chunghyun Park,Beomjun Kim,Seungcheol Park,Heeseung Kwon,Yashu Shukla,Seunghoon Sim,Jinwoo Shin,Minsu Cho
类目: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
备注: Preprint. Project page: this https URL

点击查看摘要

Abstract:World action models jointly learn to forecast world dynamics and predict robot actions, such that the learned internal world dynamics guide accurate actions. Existing approaches typically represent the world as RGB frames or latent counterparts while predicting actions as end-effector poses or joint angles, but they often struggle to capture the 3D spatial structure and contact geometry central to dexterous manipulation. We introduce Point World Action Model (PointWAM), a 3D world action model that decomposes the world into a scene (i.e., environment) and hands (i.e., actor), and jointly forecasts both as 3D point trajectories within a shared space-time coordinate frame. This explicit, disentangled representation enables effective pre-training on large-scale human demonstration videos without requiring any task-specific object or keypoint selection. Given a colored point cloud and a language instruction, PointWAM predicts how the scene and hands co-evolve in 3D space over time, then retargets the forecast hand motion to robot actions. Pre-training on human videos improves average DexJoCo success by 56.9 percentage points, and scene-trajectory supervision adds 10.9 points over forecasting the hands alone. With both, PointWAM surpasses the prior state of the art on ten DexJoCo tasks by 11.7 points and outperforms strong VLAs on a real robot.

[CV-80] FastOPD: On-Policy Distillation for Lightweight VLA Deployment FAST

链接: https://arxiv.org/abs/2610.02832
作者: Yoojin Oh,Jeongsol Kim,Yeonwoo Seo,Jangho Park,Seonghyun Jin,Sunwoo Park,Youngmin Kim,Youngjun Jun,Kyumin Choi,Jong Chul Ye
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注: Project page: this https URL

点击查看摘要

Abstract:Vision-Language-Action (VLA) foundation models have scaled rapidly to enhance manipulation performance and generalizability, but this scaling incurs high computational costs that render real-world deployment increasingly challenging. Existing approaches typically mitigate this issue by designing smaller architectures or reducing the iterative denoising steps in flow-based policies. In this work, we propose FastOPD, a foundation-to-lightweight VLA framework that enables the practical deployment of large-scale VLAs through efficient on-policy distillation. Specifically, FastOPD adapts a flow map for single-state teacher supervision and combines it with a self-consistency objective to construct a compact student that learns the teacher dynamics. Furthermore, we theoretically demonstrate that minimizing this objective allows the distilled student to recover a distribution on par with that induced by an ideal few-step teacher model. We evaluate FastOPD across diverse foundation policies in simulation and real-world experiments. On LIBERO, FastOPD retains 84% of the performance of \pi_0.5 with only two inference steps, reducing inference latency by 78.1% while outperforming existing few-step distillation baselines in average success rate. With LingBot-VLA as the teacher, FastOPD improves the single-step success rate over the base student by 15.9 percentage points on RoboTwin 2.0. We further demonstrate its applicability to a World Action Model (WAM) and deploy a compact student distilled from MolmoAct2 on a real robot.

[CV-81] rrainForge: Physics-Grounded road geometry Editing for Counterfactual Autonomous Driving

链接: https://arxiv.org/abs/2610.02825
作者: Yang Chen,Yicheng Zhu,zhenning Li,Tao Li,Zilin Bian
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Road geometry (e.g., crests, sags, and speed humps) and surface conditions (e.g., wet or icy pavement) affect how vehicles move, what drivers and onboard cameras observe, and how much clearance remains between vehicles. Editing these properties in a driving scene therefore requires corresponding changes in vehicle motion. Capturing these differences in a driving video requires a road edit to propagate to vehicle motion, camera viewpoint, and the clearance between vehicles. We present TerrainForge, a framework for generating road geometry-focused counterfactuals from reconstructed multi-vehicle driving episodes. A unified road model connects scene deformation with four-wheel vehicle dynamics, allowing crests, sags, speed humps, and friction changes to propagate through vehicle motion, camera viewpoint, and inter-vehicle clearance. Vehicle dynamics are evaluated against CarSim, and prescribed road geometry is verified in reconstructed Waymo scenes. Across 18 episodes, leaving surrounding vehicles on their recorded trajectories instead of recomputing their responses produces median peak differences in predicted ego-lead distance of 1.52 m for crests and 1.41 m for sags. We further simulate the ego response to 15,758 road edits across 983 braking episodes, pairing each edit with its safety outcomes relative to an unedited replay. These pairs train a first-stage screening surrogate that takes the original driving context and candidate road-edit parameters as input and predicts the resulting change in the ego’s terminal gap. On held-out scenes, this prediction achieves 22-40% lower mean absolute error than predicting no change, so candidates can be screened cheaply before the full multi-vehicle rollout.

[CV-82] FUSEye: Training-Light Fisheye Detection with Overlapping Views and Zero-Initialized Adapters

链接: https://arxiv.org/abs/2610.02799
作者: Wenya Su,Kai Luo,Di Wen,Ruiping Liu,Yufan Chen,Junwei Zheng,Kunyu Peng,Kailun Yang
类目: Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO); Image and Video Processing (eess.IV)
备注: Source code will be available at this https URL

点击查看摘要

Abstract:Fisheye cameras give mobile robots a single-sensor, low-cost view of their surroundings, yet the COCO-pretrained detectors that practitioners routinely reuse fail on them: strong radial distortion warps local image structure, while boundary compression shrinks objects to near-invisible sizes. Full fine-tuning closes much of the gap but requires abundant fisheye labels and compute. We present FUSEye, a training-light framework that turns a frozen-backbone COCO-pretrained extra-large YOLO26 detector (YOLO26-x) into a fisheye detector. FUSEye adds roughly 227k new parameters while updating the inserted modules and the pretrained detection head. It addresses the transfer gap at three causally linked levels. At the input level, overlapping grid view generation and box remapping (GridViews) enlarge compressed boundary regions. At the feature level, zero-initialized residual adapters (Z-Adapters) correct distortion-induced feature misalignment. At the decision level, learned cross-projection agreement fusion (AgreeFusion) promotes low-confidence detections only when they are supported by consistent evidence across multiple views. On the WoodScape surround-view fisheye benchmark, FUSEye raises YOLO26-x from 0.148 to 0.266 mAP50 and retains 84.3% fully fine-tuned accuracy. Moreover, randomly using only 25% of the labeled training images, FUSEye achieves 0.2597 mAP50, retaining 97.6% of its full-label performance. FUSEye also consistently improves YOLOv8-11 detectors, showing that the recipe is architecture-agnostic. Source code will be available at this https URL.

[CV-83] RAC: Trajectory-aware Reuse and Adaptive Correction for Efficient Autoregressive Video Generation

链接: https://arxiv.org/abs/2610.02779
作者: Jiaxing Song,Weiqi Yan,You Huang,Mingte Qiu,Huazhong Liu,Xiaofeng Zhu,Yunshan Zhong
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Preprint under review

点击查看摘要

Abstract:In this paper, we present trajectory-aware reuse and adaptive correction (TRAC), a training-free framework for efficient autoregressive (AR) video generation. Existing acceleration methods mainly target single-trajectory generation with bidirectional attention. AR video generation, by contrast, sequentially couples chunk-level denoising trajectories. Consequently, approximation errors accumulate and propagate through the generation process. TRAC addresses this challenge with three components, including robust cumulative scheduling (RCS), autoregressive trajectory-aware guidance scheduling (ATGS), and spectral structure correction (SSC). RCS selects cache reuse schedules by cumulative rollout error and cross-chunk/prompt variation. ATGS coordinates CFG refreshes along the global AR trajectory. SSC restores low-frequency structure of the first chunk to correct long-term structural loss. Experiments on SkyReels-V2 and FramePack-F1 show that, compared with existing methods, TRAC achieves both the highest inference efficiency and the best generation quality for AR video generation.

[CV-84] FiberGeoText: A Vision-Language Model for Population- Level Organization of Superficial White Matter

链接: https://arxiv.org/abs/2610.02755
作者: Yuqian Chen,R. Jarrett Rushmore,Guikun Chen,Fan Zhang,Edward Yeterian,Nikos Makris,Yogesh Rathi,Lauren J. O’Donnell
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 22 pages, 3 figures

点击查看摘要

Abstract:The superficial white matter (SWM), a critical brain region for cognition across the lifespan and brain disease, contains abundant short-range association fibers whose organization remains incompletely characterized, in part because the short trajectories and highly variable cortical folding make correspondence across individuals challenging. Anatomically corresponding connections may vary in spatial location across individuals and therefore may not be adequately defined by geometric proximity alone. We introduce FiberGeoText (FGT), a vision-language model (VLM) for organizing short-range superficial white matter (SWM) streamlines reconstructed from ultra-high-resolution diffusion MRI into population-level clusters. FGT jointly represents three complementary properties of each streamline: its three-dimensional trajectory, its cortical anatomical context, and its shape. Cortical endpoint information from multiple parcellation schemes is expressed as text and encoded using a pretrained large language model (LLM), enabling heterogeneous anatomical descriptions to contribute to a common continuous representation. We evaluated FGT on acquired submillimeter 0.76 mm diffusion MRI data. Compared with state-of-the-art (SOTA) methods, FGT produced substantially greater cortical parcel coherence, within-cluster shape consistency, cluster-size consistency, and cross-subject correspondence. The trained model also generalizes well to unseen subjects with an average of 96.7% of the 5,000 learned clusters recovered, and high consistency of cluster structure between training and testing data. Together, these findings demonstrate that integrating geometric, anatomical, and shape information by learning multimodal deep embeddings with a VLM model enables robust learning of population-consistent SWM organization despite interindividual anatomical variability.

[CV-85] Correcting Guided Diffusion Trajectories with Spectral Alignment

链接: https://arxiv.org/abs/2610.02753
作者: Gihoon Kim,Taesup Kim
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:The practical success of conditional image generation hinges on fine-grained differences in condition alignment and visual fidelity. Classifier-free guidance (CFG) is central to this success, but its lack of an explicit criterion makes it difficult to assess whether the guided trajectory is progressing as intended. To address this gap, we show that spectral alignment provides a principled criterion for understanding guidance behavior and improving guided diffusion sampling through adaptive correction. Our analysis identifies the spectra of intermediate states as an indicator of consistency with the expected spectral evolution of the forward process. Based on this observation, we introduce Spectral Correction Guidance, a method that corrects deviations from an analytic reference spectrum during sampling. The proposed method is training-free and applicable across diffusion backbones and conditional generation tasks without modifying the underlying model. Experiments demonstrate consistent gains in preference-based metrics over baseline guidance methods in text-to-image generation and improved generation quality over CFG on ImageNet. These improvements persist across a range of guidance scales and with fewer denoising steps. Our analyses and ablations provide insight into guidance behavior and how the proposed method affects generation quality.

[CV-86] SymRegFlow: Symmetry-Regularized Flow Matching for Video World Models

链接: https://arxiv.org/abs/2610.02726
作者: Xi Ye,Yuzhu Wang,Xiaoyang Liu,Jiayi Wang,Yangyang Xu,Ruyu Wang,Wenlin Chen,Duo Su,Jun Zhu
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Flow-matching-based multi-view world models generate realistic videos, but are commonly restricted to fixed camera rigs. Extending them to continuously varying camera poses requires paired pose–video observations with dense pose coverage, which are costly to acquire. We introduce \emphSymRegFlow, a symmetry-regularized flow-matching framework for multi-view-consistent video generation across continuous viewpoints without ground-truth novel-view RGB supervision. For each target pose, SymRegFlow geometrically warps source views into noisy anchors and combines masked dual-anchor supervision with cross-anchor denoising-output consistency to mitigate anchor-specific errors. Under an affine Gaussian surrogate, we prove that suitable consistency regularization recovers the clean-reference optimum at fixed noise levels, strictly outperforming single- and merged-anchor baselines. Experiments on Cosmos-Drive-Dreams and nuScenes demonstrate high-quality, multi-view-consistent autonomous-driving video generation: on nuScenes, SymRegFlow achieves the lowest FVD and FVMD among the evaluated baselines, reducing FVD by over 31% relative to the best baseline, and source-conditioned inference also attains the best FID and instance preservation.

[CV-87] Revisiting Visual Representation Enhancement of VLMs via Kernel Canonical Correlation Analysis

链接: https://arxiv.org/abs/2610.02718
作者: Peilin Yang,Xiaoyu Liu,Jian Sun,Qinghua Tao
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Vision-language models such as CLIP exhibit strong semantic generalization, but remain limited in fine-grained visual perception. A recent work named KUEA presents a natural remedy by finetuning the image encoder under the supervision of the vision-centric DINOv2 to align their kernel matrices element-wisely, while regularizing the embeddings to remain close to the pretrained visual encoder for preserving image-text semantics in CLIP. However, we show that diminishing the role of the alignment loss to DINOv2 does not necessarily degrade its fine-grained visual performance, suggesting that the kernel-matrix discrepancy may be insufficient for further visual representation enhancement, motivating us to revisit the alignment formulation. In this work, we present a novel perspective to characterize representation alignment on feature subspaces through Kernel Canonical Correlation Analysis (KCCA), which maximizes the projection correlations. In optimization, we derive an efficient end-to-end training scheme upon KKT conditions, avoiding the eigenvalue problem in KCCA. Further, we extend our method into a 3-view formulation, i.e., 3vKCCA, in which the projections from the pretrained text encoder are also incorporated under a unified optimization framework for joint alignment. With CLIP ViT-L/14 on ImageNet-1K, our 3vKCCA improves the MMVP-VLM accuracy from 17.8 to 25.9, substantially outperforming the existing methods, and meanwhile maintains zero-shot image–text retrieval performance.

[CV-88] GeoScaffold: Learning Compact Geometric Latents via Reconstruction for Efficient Vision-Language Navigation

链接: https://arxiv.org/abs/2610.02697
作者: Yixuan Jiang,Wentong Li,An Liu,Zihao Xin,Fulin Tang,Cong Leng,Yang Gao,Jian Cheng
类目: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Recent vision-and-language navigation (VLN) systems increasingly adopt streaming Video-LLM policies that map egocentric RGB observations and instructions directly to low-level actions. Yet these policies inherit weak 3D geometric priors from 2D pretraining. Existing geometry-aware extensions charge a persistent inference-time price: depth sensors, 3D encoders, or per-step perception tool calls. We propose GeoScaffold, a geometric supervision framework that pays this price once, at training time, by internalizing geometry into the policy itself. It first learns a compact depth tokenizer on depth maps from the training trajectories and freezes it. It then fine-tunes the policy with a handful of learnable geometry query tokens, training their hidden states to reconstruct navigation-critical geometry such as depth, connectivity, and traversability. This supervision turns the query states into compact geometric latents for action decoding, and through the shared weights also internalizes geometry into the backbone’s own representations. Like a scaffold, the tokenizer, target generators, and reconstruction heads are discarded after training, leaving the backbone and action interface unchanged. Extensive experiments show that GeoScaffold consistently outperforms leading vision-only navigators on continuous VLN benchmarks, offering a practical paradigm for lightweight edge deployment of spatially aware embodied navigation models.

[CV-89] CHASE-VLA: Post-Training Quantization Framework for Vision-Language-Action Models with Chunk-Aware Scale Estimation ACCV2026

链接: https://arxiv.org/abs/2610.02666
作者: Jin Hyun,Jung Gyu Min,Gyuhyun Jung,Youngjoo Lee
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted at ACCV 2026. 22 pages, including references and supplementary material

点击查看摘要

Abstract:Vision-Language-Action (VLA) models map visual observations and language instructions to continuous robot actions, but a diffusion-based action expert (AE) poses a key challenge for low-bit post-training quantization (PTQ). The AE is repeatedly invoked across denoising steps and policy queries, where fixed calibration scales can be mismatched with activation ranges that vary with denoising progress and intended motion. We propose CHASE-VLA, a chunk-aware PTQ method that exploits a VLA-specific signal readily available from the policy: the generated action chunk, including its unexecuted future suffix. Rather than relying only on static scale matching for AE layers, CHASE-VLA combines the previously generated chunk as causal action context with denoising step group information to adapt AE activation scales. This enables W4A4 quantization of both MLP and attention projections in the repeated AE without modifying the pretrained policy. On LIBERO, CHASE-VLA achieves 97.3% average success rate on \pi_0.5 when both MLP and attention projections in the AE are quantized to W4A4, restoring FP16-level performance. CHASE-VLA also reduces the weight storage of the quantized AE linear layers by 73.4% and their single-chunk memory traffic by 70.9% and 71.2% on \pi_0.5 and GR00T N1.6, respectively, with a predictor overhead of at most 1.26% of the saved storage.

[CV-90] SpectralCache: Accelerating Diffusion-Based World Models via Spectral Feature Caching

链接: https://arxiv.org/abs/2610.02660
作者: Zhendong Mi,Pu Zhao,Ziyu Hu,Xiaodong Yu,Yanzhi Wang,Grace Li Zhang,Shaoyi Huang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Diffusion-based world models enable high-quality interactive environment generation but suffer from substantial inference overhead due to repeated Transformer evaluations during denoising. Existing caching methods mainly exploit temporal redundancy at the feature or token level, leaving the underlying mathematical structure of diffusion features largely unexplored. In this work, we reveal that world-model features exhibit highly stable singular subspaces across nearby denoising steps, while their singular values follow predictable evolution patterns. Building on this observation, we propose SpectralCache, a training-free spectral caching framework that reuses stable singular subspaces and estimates only low-dimensional singular values through linear extrapolation. We further exploit the spectral consistency between neighboring full-computation features to skip selected expensive backbone evaluations via singular value scaling. Extensive experiments on representative world models demonstrate that SpectralCache consistently improves inference efficiency while preserving generation quality. On HunyuanWorld-Voyager-13B, SpectralCache achieves 5.22x acceleration while maintaining a WorldScore of 65.90 for static scenes, substantially outperforming existing training-free caching methods in inference efficiency.

[CV-91] Capturing Dynamics: The 4D Facial Expression Intensity Dataset

链接: https://arxiv.org/abs/2610.02647
作者: Zesheng Wang,Alexandre Bruckert,Pierre Lebreton,Patrick Le Callet,Yante Li,Guoying Zhao
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:The estimation and analysis of facial expression intensity play a crucial role in affective communication and human-computer interaction. Previous research has primarily focused on detecting and estimating facial expression intensity from frame-level 2D representations. However, this limitation restricts a comprehensive understanding of real-world facial expressions, as they are inherently 3D and temporally continuous. This paper investigates the perception of facial expression intensity by introducing the 4D Facial Expression Intensity Dataset (4DFEID). We employ a parametric face model and compile a total of 2,869 mesh sequences with controlled geometric variations, generating 4D data instances with diverse peak intensities and identity attributes. Using a Likert scale, we collect more than 90,000 subjective intensity perception ratings via a crowdsourcing platform. We explore various architectures and aggregation methods to establish baselines for episode intensity estimation on the new dataset, revealing that spatial-temporal graph models consistently outperform traditional frame-aggregation methods. In contrast to existing datasets that rely on 2D static imagery, the proposed 4D-FEID dataset provides the community with a unique and vital resource for investigating the perception of facial expression intensity through the use of dynamic 3D stimuli. By offering high-fidelity, spatio-temporally coherent facial data, 4D-FEID establishes a new foundation for research into more nuanced and naturalistic expression analysis, thereby addressing a gap in the current landscape of affective computing and human-computer interaction studies. The dataset is available at link.

[CV-92] Imagine the Future Internalize the Gist: Efficient VLA Reasoning via Internalized Spatiotemporal Imagination

链接: https://arxiv.org/abs/2610.02626
作者: Shenglan Li,Zhendong Mi,Hengyi Zhu,Jingwu Luo,Chun Kit Chan,Geng Yuan,Yanzhi Wang,Pu Zhao,Shaoyi Huang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Vision-language-action (VLA) models increasingly incorporate intermediate reasoning to improve robotic manipulation, yet existing approaches primarily reason about observed states without explicitly anticipating future scene evolution. Extending such reasoning to explicit future rollouts at every inference step, however, introduces substantial computational overhead. We propose IG-VLA, a VLA reasoning framework that enables models to imagine the future and internalize the gist. Our Latent Spatiotemporal Reasoning learns to imagine task-relevant future scene evolution directly in visual representation space, guiding action prediction without costly pixel-level video generation. To further reduce inference overhead, we introduce Scene Gist Memory, which internalizes reasoning-derived scene-behavior associations into a compact Scene Gist Token, preserving the benefits of future reasoning while bypassing explicit future imagination at inference. Extensive experiments on LIBERO, LIBERO-Plus, and VLABench demonstrate the effectiveness and efficiency of IG-VLA. On the LIBERO-Plus Language suite, both the reasoning and gist policies outperform the strongest baseline by nearly 6% in success rate. The gist policy also achieves up to 6.38x speedup over baselines, reducing inference latency from 1081ms to 169.5ms per action chunk on a single NVIDIA A6000 GPU. These results demonstrate that future spatiotemporal reasoning can be effectively internalized for efficient VLA deployment.

[CV-93] Scale-Recursive Rectified Flows for Few-Step Precipitation Ensembles

链接: https://arxiv.org/abs/2610.02611
作者: Shunya Nagashima,Takumi Bannai
类目: Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Fine-resolution precipitation estimates support flood risk assessment and water management, but coarse satellite products cannot resolve rainfall within each grid cell. Generative models address this ambiguity by producing ensembles of plausible high-resolution rainfall fields. Among these models, rectified flows generate samples by iteratively transforming random noise into rainfall fields. Reducing the number of sampling steps accelerates generation but can make ensemble members too similar, understating uncertainty. We propose a scale-recursive rectified flow that generates broad patterns before local details and guides sampling-step allocation by comparing ensemble variability with prediction error across spatial scales. Validation scores and rainfall power spectra constrain the allocation to avoid excessive amplification. In satellite-to-radar downscaling over the contiguous United States, our analysis identified broad rainfall patterns as the main source of insufficient ensemble variability under reduced sampling budgets. Allocating more steps to the coarse flow improved probabilistic accuracy and rain detection across training seeds at fixed architecture and computational cost. The proposed model also achieved better probabilistic accuracy with shorter sampling time than a nonrecursive flow using more steps.

[CV-94] GRAFT: Growing Agglomerative Foundation Models via Continual Teacher Distillation

链接: https://arxiv.org/abs/2610.02597
作者: Zhenghao Zhao,Chi Zhang,Qingshuang Chen,Yelin Kim
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Vision foundation models such as DINOv2, SigLIP2, and MASt3R develop complementary capabilities from different pretraining objectives, yet their knowledge remains distributed across separate, specialized models. Multi-teacher knowledge distillation offers a path toward consolidating these capabilities into a single agglomerative backbone, but existing approaches assume a fixed set of teachers, and incorporating a new teacher requires repeating expensive joint distillation over the entire teacher set. We introduce GRAFT, a continual multi-teacher distillation framework that enables a unified backbone to progressively acquire capabilities from an open-ended sequence of foundation models. When a new teacher arrives, GRAFT treats the previously distilled model as a teacher for preserving learned capabilities, while the current student jointly learns from both the previous model and the incoming teacher. Furthermore, to reconcile the incompatible representation geometries of heterogeneous teachers, we introduce Teacher Specific Readout Tokens, which grant each teacher an independent read-out of the shared encoder, together with Geometry Agnostic Relational Loss that aligns a vision-language teacher by matching image-text similarity structures rather than raw feature values. We provide GRAFT model, which is a single, continually extensible backbone that unifies five domains, including image understanding, 2D dense prediction, 3D human pose estimation, 3D vision, and vision-language, delivering strong performance across all of them while acquiring each new capability at the cost of a single distillation rather than a full re-distillation.

[CV-95] Physical AI Smart Spaces: A Large-Scale Benchmark for Multi-Camera 3D Perception in Smart Spaces NEURIPS2026

链接: https://arxiv.org/abs/2610.02580
作者: Yuxing Wang,Yizhou Wang,Anqi Li,Shuo Wang,Sameer Satish Pusegaonkar,Haoquan Liang,Jiajun Li,Shenxin Jiang,Jianhe Yuan,Shangru Li,Tongwei Dai,Zihao Chen,David C. Anastasiu,Sujit Biswas,Xunlei Wu,Zheng Tang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted at NeurIPS 2026, Evaluations Datasets Track (poster)

点击查看摘要

Abstract:Physical AI Smart Spaces is, to the best of our knowledge, the first benchmark to simultaneously provide large-scale, multi-class, and multi-camera 3D perception data for indoor smart spaces. It contains over 280 hours of synchronized 1080p footage captured by nearly 1,800 cameras in warehouses, hospitals, retail venues, and similar settings, together with automatic annotations for multi-camera identities, 2D bounding boxes, 3D bounding boxes, camera calibration, and depth where available. The benchmark spans Isaac Sim synthetic generation, Cosmos Transfer appearance augmentation, and real-world Sim2Real evaluation. For the real-world target, we include two warehouse deployments with time-synchronized streams, automatic VGGT-based calibration, and a 3D labeling interface that projects world-frame 3D boxes into each view for cross-camera verification. We describe the dataset scope, annotation and calibration schema, generation workflow, benchmark protocols, and official evaluation system, which standardizes submission format, and leaderboard reporting. A central contribution is a 3D instantiation of Higher Order Tracking Accuracy (HOTA), extending the usual 2D box-based tracking evaluation to 3D locations and 3D boxes. We further report empirical baselines from the AI City Challenge leaderboards, showing how methods evolve from person-only 3D location tracking to multi-class 3D box tracking under realistic smart-space constraints. The release is available at this https URL.

[CV-96] DAGS: Disentangled Appearance-and-Geometry Steering of a Frozen Image DiT for Temporally Stabilized Generative Rendering SIGGRAPH

链接: https://arxiv.org/abs/2610.02567
作者: Karthik Mohan Kumar,Damian Andrysiak,Pedro Antonio Pena,Kunal Tyagi,Rama Harihara
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Graphics (cs.GR); Machine Learning (cs.LG)
备注: 5 pages, 3 figures, 2 tables. Accepted to SIGGRAPH Asia 2026 Technical Communications

点击查看摘要

Abstract:Diffusion transformers (DiTs) generate high-fidelity images from text and image conditions, but their outputs carry large variance and their faithfulness to a desired target depends heavily on how the condition is supplied. We present DAGS, a lightweight, attention-free, disentangled appearance and geometry conditioning scheme that steers a frozen image DiT to produce high-fidelity, highly faithful, and independently controllable renders. Two small convolutional encoders compute conditioning features once per frame and inject them as a learned, per-layer, element-wise residual into the image tokens, avoiding the quadratic cost of stacking conditions through attention. Because control and temporal handling live outside the frozen backbone, we retain its vast pretrained prior and eliminate backbone-overfitting risk. We further add a small recurrent lighting stabilizer and a training-free temporal guidance term that, coupled with our conditioning, elevate a per-frame image model into a streaming renderer. DAGS produces controllable, high-quality renders at a fraction of the compute of path tracing; it is not real-time, trading compute for controllability and quality. On a matched 1-spp + G-buffer input, per-frame DAGS reconstructs +8.6 dB / +10.1 dB PSNR over the real-time denoiser Intel OIDN and the diffusion renderer RGB-X while being 2.5-8x more temporally stable perceptually (temporal-LPIPS flicker).

[CV-97] Oracle headroom without signal: null-calibrated evaluation of candidate selection for thermal heart rate estimation

链接: https://arxiv.org/abs/2610.02561
作者: Mohammad Rakibur Rahman,Nhi Nguyen,Sasan Sharifipour,Le Nguyen,Manuel Lage Cañellas,Miguel Bordallo López,Constantino Álvarez Casado
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 5 pages, 3 figures

点击查看摘要

Abstract:Camera-based physiological monitoring can produce multiple estimates from several facial regions, extraction methods, and processing settings. Signal quality indices aim to select reliable estimates without a physiological reference, and their potential is often assessed with an oracle that selects the estimate closest to the reference in each window. This retrospective selection can reward chance agreement. We model the effect with order statistics. For K independent candidates unrelated to the reference, the expected oracle error decreases approximately as 1/K. We analyze thermal heart rate estimation on 96 iBVP recordings with 168 candidates per 10 s window. The oracle achieves a mean absolute error of 0.91 bpm, compared with 10.74 bpm for the best fixed configuration, 18.03 bpm for the best quality index, and 8.61 bpm for a constant predictor. With K = 24, a forehead signal from another recording matches the correct one, with 4.62 against 4.61 bpm. Oracle evaluations should report candidate count, valid coverage, and matched null controls.

[CV-98] Spatial Memory Intelligence: Endowing World Models with Understanding-Driven Long-Term Memory

链接: https://arxiv.org/abs/2610.02521
作者: Ying Yang,Guiyu Zhang,Lianghua Huang,Chang Nie,Chenyang Si,Haofan Wang,Shaoshuai Shi,Li Jiang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 32 pages. Project page: this https URL

点击查看摘要

Abstract:Long-video generation and world models have shown strong potential for interactive entertainment and embodied simulation by predicting future observations conditioned on user actions and historical memory. However, as memory sequences grow longer and their structures become increasingly complex, managing long-range spatial context becomes increasingly challenging, calling for a more intelligent and systematic memory-management strategy. Building on the advancing spatial reasoning capabilities of multimodal large language models (MLLMs) and the broader vision of unified models, we propose Spatial Memory Intelligence (SMI), the first framework to systematically employ an understanding model for spatial-memory management in long-video world models. SMI introduces four coordinated atomic operations: spatial clustering, within-cluster sparsification, action-aware retrieval, and reliability-aware filtering. Extensive experiments across multiple baselines, benchmarks, and world-model backbones demonstrate the effectiveness and generalizability of SMI, achieving comprehensive improvements in memory sparsity, generation stability, and spatial consistency.

[CV-99] From Frag ments to Global Maps: Learning Vectorized Map Aggregation with Large Language Models

链接: https://arxiv.org/abs/2610.02513
作者: Ziwei Li,Yi-Tang Chen,Xiaoqi Wang,Wenbin He,Han-Wei Shen,Liu Ren
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Large-scale vectorized HD maps provide structured road information that is essential for perception, localization, and planning in autonomous driving. Constructing such maps requires aggregating noisy, fragmented, and overlapping local predictions collected along a vehicle trajectory into a coherent global map. Existing aggregation methods typically rely on hand-crafted rules for fragment association and refinement. However, a fixed set of thresholds cannot effectively handle variations in road structures and prediction errors, often requiring detector-specific tuning or manual adjustment. To address this limitation, we propose MapMergeLLM, a data-driven framework that formulates vectorized map aggregation as conditional sequence generation with a large language model. Given serialized local vectorized maps, our model directly predicts the aggregated global map polylines. To reduce dependence on any particular upstream detector, we train the model on synthetic local maps generated from clean vector maps using corruptions that simulate representative prediction errors. We further introduce a coordinate tokenizer with geometry-aware pretraining to precisely represent map coordinates. In addition, we propose a line-level association loss that explicitly supervises correspondences between local observations of the same map element. Experiments on Argoverse2 and nuScenes using multiple recent upstream detectors demonstrate that MapMergeLLM substantially outperforms heuristic and optimization-based aggregation baselines without detector-specific retraining.

[CV-100] World Action Modeling with Progressive Visual Planning

链接: https://arxiv.org/abs/2610.02508
作者: Fei Zhang,Zhaochong An,Duncan Frost,Yikai Wang,Pengfei Liu,Ya Zhang,Michal Drozdzal,Amir Bar
类目: Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)
备注: Project Page: this https URL

点击查看摘要

Abstract:World action models (WAMs) have emerged as a promising paradigm for robotic control by jointly predicting future visual dynamics and actions from an initial observation and instruction. However, existing WAMs struggle with long-horizon prediction, as generating dense video rollouts is highly inefficient. Some recent WAMs address this by predicting a single future frame without generating the full video, but this approach neglects how to progress toward the goal. We present ProWAM, a progressive world action model that jointly predicts actions and an ordered sequence of sparse visual sub-goals, providing explicit visual guidance to anchor action generation throughout task execution. This design scales naturally, as sub-goal prediction can be learned from large-scale action-free videos, allowing the video backbone to offload complex visual planning from the action policy. For efficient action generation, ProWAM executes a single video-backbone forward pass to cache sparse sub-goal features, eliminating iterative full-video generation and requiring only lightweight action denoising during replanning. Across extensive evaluations, ProWAM achieves superior out-of-distribution robustness. On simulation benchmarks, it sets new state-of-the-art results on LIBERO-Plus (85.8%) and randomized RoboTwin (75.7%), outperforming the strongest baseline with relative gains of up to +35.9%. On RoboCasa365, ProWAM achieves a 48.1% success rate and 18.2% on the challenging Composite-Unseen split, ranking 4th overall. Crucially, in zero-shot real-world experiments, ProWAM achieves 70.0% success, outperforming the strongest baseline by +15.0 (from 55.0% to 70.0%, a +27.3% relative gain) in novel scenes. These results demonstrate the value of progress-indexed visual foresight for closed-loop control. Our program is in this https URL.

[CV-101] MeshQuery: Agent ic Seam Planning for UV Parametrization

链接: https://arxiv.org/abs/2610.02507
作者: Marco Schouten,Arthur Roullier,Elie Michel,Ruben Wiersma,Axel Paris,Tamy Boubekeur
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:We present MeshQuery, a training-free agentic approach to automatic UV unwrapping of production-grade quad meshes. A Vision-Language Model (VLM) plans artist-aligned seams using a set of edge-selection tools, conditioned on domain-specific UV-unwrapping knowledge expressed in natural language and refined with a feedback loop. We design a queryable mesh representation together with a domain-specific language (DSL) that enables the agent to retrieve mesh information on demand, express a seam plan as a compact program of edge-selection operators over topological, geometric, and semantic mesh attributes, and iteratively refine it from UV quality feedback. On Adobe Substance 3D and Toys4K meshes, MeshQuery produces 2.9x/4.29x fewer charts and 1.63x/1.7x shorter seams than the strongest baseline, and professional artists prefer its results in 80.9% of comparisons. Ultimately, decoupling high-level intent planning from low-level edge selection and compact mesh representation lets MeshQuery run on different backend VLMs and scale to meshes an order of magnitude larger than autoregressive seam prediction

[CV-102] DeepStratNet: A Context-Aware Coordinate Regression Framework for Seismic Horizon Tracking under Sparse Labels

链接: https://arxiv.org/abs/2610.02494
作者: Aniq Ahmad,Musham Ahmad Malik,Ahmad Mustafa,Heather Bedle
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Automatic horizon tracking is a foundational task in 3D seismic interpretation. Most existing deep learning approaches formulate it as dense semantic segmentation, typically using U-Net-based architectures. The model produces a probability map over all pixels that must be post-processed to extract precise horizon coordinates, while horizon picks in time/depth must be converted into dense masks for training. Unpicked seismic traces are consequently treated as background, which can hinder convergence, and both pre- and post-processing can introduce errors into the final interpretation. Moreover, 2D segmentation models do not inherently capture inter-slice context, while 3D models are often computationally prohibitive. We instead formulate horizon tracking as a bounded coordinate regression problem, where the model directly predicts the time/depth coordinate of the target horizon at each lateral position. We propose a lightweight regression head compatible with any pretrained vision backbone, coupled with an LSTM module to model inter-slice context and produce a continuous horizon surface across the volume. A combination of L1 and L2 losses supervises predictions at valid horizon picks, while a geology-informed regularization enforces lateral continuity between successive traces. Under controlled experimental conditions, we evaluate four pretrained vision backbones under both segmentation and regression configurations on a seismic volume from New Zealand. The proposed approach consistently outperforms its segmentation counterparts quantitatively, using metrics including RMSE and PCC, and qualitatively, while also demonstrating greater robustness to increasing sparsity of training picks. Finally, we show that prediction variation across successive traces captures local variations in geological complexity, providing an automated quality control measure for downstream seismic interpretation.

[CV-103] Windfoil: Closed-Form Coverag e for Real-Time and Differentiable Vector Graphics

链接: https://arxiv.org/abs/2610.02468
作者: Matt DesLauriers
类目: Graphics (cs.GR); Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:We present Windfoil, a GPU-friendly algorithm that treats rasterisation and differentiable vector graphics as two sides of the same problem by evaluating the box-filtered winding number of quadratic Bézier contours in closed form. We implement this in WebGPU, allowing it to run across a range of environments, including a web browser on a consumer laptop, and apply the system to real-time 2D rendering, high-resolution rasterisation for print media, and a differentiable renderer. We compare our renderer against Skia, a production-grade engine, and Slug, a popular GPU rasterisation algorithm for games and real-time applications, measuring fidelity to a reference box-filtered coverage. Our renderer matches the reference more closely than either, at performance comparable to Slug. We also compare our optimiser against DiffVG and Bézier Splatting, where it reaches equivalent or better reconstruction quality at a fraction of the per-step cost, scaling to tens of thousands of shapes at interactive rates.

[CV-104] A Simulation-Grounded Agent ic VLM Framework for Wildfire Monitoring and Reporting

链接: https://arxiv.org/abs/2610.02451
作者: Duowen Chen,Yuchen Sun,Zhiqi Li,Yuxuan Liao,Sinan Wang,Bart van Bloemen Waanders,Bo Zhu
类目: Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR)
备注:

点击查看摘要

Abstract:Effective wildfire monitoring requires relating visual evidence to physical fire dynamics, yet real videos with synchronized physical annotations are scarce and high-fidelity 3D simulation is costly. We present a simulation-grounded vision-language model (VLM) framework that automatically converts 2D wildfire simulations into labeled video episodes. A fixed Blender mapping produces low-detail 3D proxies aligned with simulator terrain, fuel layout, fire activity, and wind cues; controllable video generation supplies richer appearance. The proxies are intermediate representations rather than finely rendered final scenes. Generated videos and simulator labels form reusable multimodal memory for a training-free multi-agent VLM system that retrieves reference episodes, reconciles visual and memory-based predictions, and produces structured wildfire reports. On held-out generated episodes, video memory achieves 51.5% exact four-tag accuracy, compared with 22.6% for direct VLM querying and 16-17% for text-only memory; the complete system achieves 77.3% accuracy on six simulator-derived report fields. Component ablations, cross-generator tests, and three real-UAV evaluations assess retrieval, reporting, generator changes, and observable monitoring tasks. The framework connects automatic simulation-to-proxy conversion with memory-based VLM reasoning under scarce real-world physical annotations.

[CV-105] An AI-Based Multi-Stage Approach for Androgenetic Alopecia Assessment from Low-Magnification Scalp Images

链接: https://arxiv.org/abs/2610.02421
作者: Mahmoud Raslan,Nada Omar,Omar Khaled,Tarek Waleed,Mohamed Hazem,Rania Mounir,Solwan Elsamanoudy,Ahmed Mourad,Noura Adel,Muhammad Rushdi
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Androgenetic alopecia (AGA) is characterized by patterned follicular miniaturization, increased single-hair follicular units, and altered hair-shaft diameter. We present an automated quantitative scalp-analysis and clinical decision-support framework combining FU localization, ordinal visible-shaft counting, calibrated shaft-width estimation, regional aggregation, and an interpretable rule layer. The clinical cohort comprised 243 patients (127 AGA, 116 non-AGA), while the computer-vision experiments used 160 expert-annotated patients, 2,400 trichoscopic images, and approximately 158,000 FU annotations. Under patientdisjoint evaluation, YOLOv8m achieved test mAP@0.5=0.920 and recall=0.860; EfficientNet-B5 with a support-map channel achieved 87.0% expert-box count accuracy (macro F1=0.85). A separate 500-image set was processed end-to-end with detector-generated boxes, yielding MAE of 6.56 for follicle detection and 16.59 for follicle classification relative to human-expert annotations. The system is intended to assist, rather than replace, dermatologist interpretation.

[CV-106] Octrees as an Explicit 3D Language

链接: https://arxiv.org/abs/2610.02388
作者: Ran Dan,Si-Tong Wei,Pengfei Xiong,Wei Zhang,Yadong Mu,Peng-Shuai Wang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Project Page: this https URL Code: this https URL

点击查看摘要

Abstract:Existing 3D large language models (LLMs) compromise on two fronts: they compress shapes into latent codebook indices or coordinate text, which removes spatial structure from what the model observes, and they acquire the 3D modality by fine-tuning the backbone, which overwrites its general language ability. We present OctLLM, which addresses both limitations. Geometry enters as an explicit 3D sequence of octree occupancy tokens. However, full octree sequences grow rapidly with depth; OctLLM therefore randomly empties penultimate-level nodes and omits descendants while preserving shape, yielding a shorter coordinate- and depth-anchored Sparse Octree (S-Octree) for position-aware mask-modeling generation and 3D understanding. On the other front, existing methods introduce a new modality with full fine-tuning or LoRA, but full fine-tuning is costly, LoRA limits 3D capacity, and both modify the language pathway. OctLLM instead adds 3D capacity in parameters separate from the pretrained ones: mesh tokens are routed through independent trainable branches in a subset of blocks while text and image tokens retain the frozen vision-language pathway, and the two streams interact through shared self-attention. It trains far fewer parameters than full fine-tuning, yet sets a new state of the art among unified multimodal LLMs, lowering image-to-3D FID by 17.4% and raising render-grounded captioning by 28.7 points over ShapeLLM-Omni, while matching the backbone on general language benchmarks.

[CV-107] FactorSplat: Appearance-Controllable Gaussian Proxies for Medical Volume Rendering

链接: https://arxiv.org/abs/2610.02382
作者: Zhongpai Gao,Benjamin Planche,Meng Zheng,Anwesa Choudhuri,Terrence Chen,Ziyan Wu
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Transfer functions (TFs) control color and visibility in medical volume rendering, but image-trained Gaussian proxies typically bake one transfer function into their appearance. We present FactorSplat, a per-scene N-dimensional Gaussian splatting (N-DGS) proxy that accepts region-specific intensity-to-RGBA curves at inference. A local lookup applies the authored color and opacity change, while a shared functional encoder and low-rank per-Gaussian factors learn the residual appearance response. Geometry and directional appearance remain shared across presets, with visibility control and TF-aware pruning preserving the ability to hide and reveal structures. On seven CT and MR scans, FactorSplat improves mean PSNR and changed-region error over region-aware VEG across validation, interpolation, unseen composition, and out-of-distribution (OOD) edits. Across these four splits, seven-scan mean PSNR gains over VEG range from 1.10 to 1.52 dB. One checkpoint per scan supports unseen edits without retraining. At 1600^2 , the cached fast renderer averages 524 FPS with 1.17 ms TF switches. Project page: this https URL.

[CV-108] EviDent-CBCT: Evidence-Bottlenecked Report Generation from Dental CBCT under Non-Exhaustive Report Supervision

链接: https://arxiv.org/abs/2610.02375
作者: Ruiyang Hao,Zhi Qin Tan,Yulan He,Owen Addison,Yunpeng Li
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Dento-maxillofacial cone-beam CT (CBCT) reports may contain dozens of tooth-specific, anatomical, and spatial findings from a single 3D scan. Learning to generate such reports from limited clinical data is challenging because routine reports may not exhaustively document image findings, and a non-mention may reflect either absence or non-reporting. We present EviDent-CBCT, an evidence-bottlenecked framework designed for this incomplete supervision. An anatomy-aware network maps each CBCT scan to a discrete record of tooth-level, global, and tooth-IAC evidence. A dental-logic consistency projection reconciles incompatible evidence before a deterministic renderer and an image-blind local language model generate the report using only this record. For tooth-level evidence, reliability-aware training uses eligible non-mentions as reduced-weight negatives, while unreported global and tooth-IAC labels remain unknown. A metal-sensitive input channel preserves intensity cues from dental materials. Across three validation runs, EviDent-CBCT achieves 0.666\pm0.006 merged evidence set-F1 and 0.402\pm0.003 RadFact-Lite-Dental logical-F1, versus 0.371\pm0.018 for the strongest controlled direct baseline. In the ODIN 2026 challenge, it ranked second in automated evaluation and third in blinded clinical Arena comparison on the hidden test set. These results support the discrete evidence record as an effective and auditable interface for CBCT report generation.

[CV-109] Confidence-Controlled XAI Auditing for Pedestrian Detection under Domain Shift

链接: https://arxiv.org/abs/2610.02364
作者: Ruben Dario Florez-Zela
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted at the 2026 IEEE International Conference on Vehicular Electronics and Safety (ICVES 2026). 6 pages, 4 figures

点击查看摘要

Abstract:Explainability is increasingly required for perception models in intelligent vehicles, yet whether explanations remain faithful under driving domain shift is still poorly understood. This work audits post-hoc explanations of a fixed YOLOv8s pedestrian detector across PIE and JAAD using ROI-based D-Deletion, frozen confidence terciles, rank-based tests, bootstrap intervals, and Holm correction. The audit shows that deletion-based faithfulness is strongly coupled to detection strength at explanation time, with Spearman correlations between 0.70 and 0.82 for D-RISE, making naive confidence-stratified comparisons unreliable. After controlling for detection strength within fixed f0 bins, D-RISE faithfulness remains domain-dependent in the central f0 range, with PIE showing higher D-Deletion than JAAD and Holm-adjusted significance. A non-perturbative EigenCAM baseline is less faithful than D-RISE but also exhibits score coupling, suggesting that the effect is not specific to D-RISE and is related to the deletion-based evaluation setup. These results motivate confidence-controlled XAI audits for safety-critical perception under domain shift.

[CV-110] SCOPE-4D: Endoscopic 4D Geometry Foundation Models

链接: https://arxiv.org/abs/2610.02343
作者: Chaoyi Zhou,Zhongpai Gao,Anwesa Choudhuri,Meng Zheng,Benjamin Planche,Run Wang,Terrence Chen,Siyu Huang,Ziyan Wu
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Project page: this https URL

点击查看摘要

Abstract:Geometric understanding supports endoscopic navigation and robotic assistance, but learning reliable endoscopic geometry faces two challenges: scarce geometric annotations and ambiguity between camera motion and tissue deformation. We present SCOPE-4D, an endoscopic 4D geometry foundation model that jointly predicts camera parameters, dense geometry, and 3D tissue trajectories from monocular RGB video in a single forward pass. Our curation and annotation pipeline constructs SCOPE-5K, a collection of approximately 5,000 clips spanning real and synthetic gastrointestinal endoscopy and laparoscopy. The collection provides rich geometric supervision and includes newly collected phantom and real-colonoscopy evaluation sets. Geometric supervised fine-tuning on SCOPE-5K learns endoscopic priors that improve camera and depth estimation. Common–Residual Motion (CRM) further constrains local deformation relative to common tissue movement. Together with geometric supervision, CRM and trajectory supervision further improve camera and depth estimation over geometric fine-tuning alone while enabling dense 3D tissue tracking. Evaluations on public and newly collected benchmarks demonstrate strong in-domain and out-of-domain geometry, superior 3D tracking, and more stable long-sequence colon reconstruction. A blinded user study further supports the perceived reconstruction quality on real clinical video. Together, these results demonstrate the value of large-scale endoscopic supervision and motion constraints for joint geometry estimation and tissue tracking.

[CV-111] World-Calibrated Proposal-to-Action Flow for Vision-Language-Action Models

链接: https://arxiv.org/abs/2610.02323
作者: Jie He,Wei Li,Junwen Tong,Rui Shao,Wei-Shi Zheng,Liqiang Nie
类目: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
备注: 25 pages, 10 figures. Project page: this https URL

点击查看摘要

Abstract:Flow-based Vision-Language-Action (VLA) policies generate action chunks by transporting samples from a task-agnostic isotropic Gaussian source. As this source is conditioned on neither recent execution nor predicted future evolution, (i) it discards the local continuity established by recently executed motion. (ii) Even when predictive world representations are introduced, they often only condition the transport dynamics rather than determine where generation starts, how far it may deviate, or along which action directions it may expand. Building on this observation, we introduce ProAct, a world-calibrated proposal-to-action framework that makes the generative source itself predictable. (i) To preserve motion continuity, a lightweight Proposal Expert converts recent actions into a scene-aware hypothesis via one motion-anchored endpoint flow-matching step, initializing generation near the demonstrated action manifold. (ii) To jointly capture intended scene evolution and proposal-future compatibility, a prospective World Expert treats the hypothesis as a soft motion prior while predicting the task-consistent latent future. (iii) From this compatibility, the model calibrates a proposal-centered anisotropic source, where a bounded per-step extent controls the allowed deviation and a trace-normalized low-rank geometry under a condition-number budget allocates refinement over coupled translation, rotation, and gripper directions. Compared with \pi_0.5 , ProAct improves performance across simulation and real-world tasks while reducing denoising steps by 50%, inference latency by up to 25.8%, and increasing throughput by up to 34.8%.

[CV-112] SCION: Scene Composition with Instanced Neural Primitives NEURIPS2026

链接: https://arxiv.org/abs/2610.02322
作者: William Koch,Amogh Joshi,Cyrus Vachha,Cheng Zheng,Felix Heide
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted to NeurIPS 2026

点击查看摘要

Abstract:Real-world scenes are compositional: bricks, blades of grass, pebbles, and tree leaves recur across human-built and natural environments. Existing neural scene representations model these elements independently. Most 3D Gaussian Splatting and follow-up abstraction and compression methods treat each element as unique, fitting millions of independent Gaussians per scene. Prior methods like Splat and Replace fit template objects, but they require mostly manual selection of repeated elements. As a result, these representations store redundant parameters and provide weak manipulation handles for downstream tasks. We introduce SCION, a hier- archical compositional scene representation that replaces independent Gaussians with a compact vocabulary of reusable primitives and lightweight world-space instances that place transformed copies throughout the scene. We fit this represen- tation to multi-view captures via a joint optimization over discrete and continuous scene parameters, combining two-level densification over splats and instances with an adversarial loss that preserves detail across shared primitives. The recovered structure yields a compact, controllable representation while maintaining high quality even at 1.2 MB. SCION achieves rate-distortion favorable to existing Gaussian compression methods, and it enables instance-level scene editing and animation without retraining. Our results show that neural scene representations need not memorize scenes as independent primitives; they can discover reusable parts. Project webpage: this https URL

[CV-113] DeskForge: Dense Supervision from Desktop Environments for Computer-Use Agents

链接: https://arxiv.org/abs/2610.02320
作者: A. Said Gurbuz,Ahmed Nassar,Sunghwan Hong,Marc Pollefeys,Peter W. J. Staar
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: 37 pages, 15 figures, 12 tables. Project page: this https URL

点击查看摘要

Abstract:Computer-use agents need to reliably ground action targets in complex desktop scenes, where multiple applications, overlapping windows, and visually similar controls compete for attention. Existing training data rarely pair such scenes with dense annotations or vary them in a controlled way. We introduce DeskForge, a controllable desktop environment that composes and explores real applications to generate large-scale supervision for computer-use agents. It varies application states, content, window layout, appearance, and resolution, and fuses screenshots, accessibility trees, and window geometry into dense element annotations while recording the outcome of each executed action. Using this environment, we construct DeskForge-1M, a corpus of 1.2M annotated desktop observations containing 159.7M element instances. We fine-tune four vision-language models on 200K grounding examples drawn from DeskForge-1M. All four improve across held-out desktop conditions and on all five external GUI grounding benchmarks; for Qwen3.5-4B, accuracy increases by 11.51 percentage points on ScreenSpot-Pro and 10.11 points on OSWorld-G. The gains also translate to long-horizon task completion: under a fixed planner, the fine-tuned action models solve more WebArena-Infinity and OpenApps tasks, with Qwen3.5-4B increasing from 31 to 50 of 119 tasks and from 3 to 15 of 100 tasks, respectively. These results show that controllable composition of real desktop environments provides a scalable source of supervision for improving both GUI grounding and long-horizon computer use. The framework code, the dataset, and the fine-tuned model are available from the project page: this https URL

[CV-114] EditHero: A Benchmark for Long-Horizon Part-Level 3D Editing and Vibe Modeling

链接: https://arxiv.org/abs/2610.02298
作者: Ruihan Yu,Yu-Ju Tsai,Muyao Niu,Runyi Li,Lian Fu,Hanqing Liu,Zheng-Hui Huang,Yonghao Yu,Sho Kuno,Ming-Hsuan Yang,Kaipeng Zhang,Zhixiang Wang
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Graphics (cs.GR)
备注: Project page: this https URL , Code: this https URL

点击查看摘要

Abstract:3D editing methods are usually tested on a single edit, yet an asset is built through a long sequence of revisions, each of which must implement the requested change while leaving everything else unchanged. We introduce EditHero, to our knowledge the first benchmark for long-horizon, part-level 3D editing, with natural-language instructions and target images for both geometry and texture. A deterministic assembly engine produces the exact target after every edit, and every sequence is reviewed by hand. We use EditHero to compare 2 opposite approaches to 3D editing. Non-agentic methods operate top down, regenerating the object from a learned 3D representation and inferring what to keep. In contrast, LLM/VLM agents operate bottom up, editing through code that inspects the mesh and rewrites only the parts required by instructions. The non-agentic methods often miss the requested change and disturb regions that should stay fixed. Most LLMs follow instructions more closely, and all of them preserve the unedited parts better, but each of their edits takes minutes. We will release the engine and the edit sequences to support research on reliable iterative 3D editing.

[CV-115] Fewer Tokens Better Action: GPT -6 Astra Robot Agents with 14% Higher Success Rate but 65% Fewer Tokens

链接: https://arxiv.org/abs/2610.01939
作者: Ruiyang Si,Jianxin Bi,Shunyu Yang,Rui Ni,Wenbo Huang,Qiang Wang,Shulong Jiang,Duomin Wang,Xiuyu Li,Haiwen Feng,Zhen Dong,Daquan Zhou
类目: Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)
备注:

点击查看摘要

Abstract:Vision language model (VLM) agents can control robots through visual feedback and action primitives, but repeated model invocations and redundant observations incur substantial token overhead. We introduce PyRUA-Lean, an interactive code-execution framework that couples feedback-driven primitive composition with selective observation: the agent composes classical robot primitives and learned vision-language-action (VLA) policies into Python cells that perform conditional checks and local retries, returning only explicitly requested images and state feedback for replanning. Across 700 simulated task instances from LIBERO-PRO, RoboTwin 2.0, and RoboCasa365, we compare PyRUA-Lean with a tool-calling baseline using the same GPT-6 Astra planner and underlying robot primitives. Under equal LLM-call budgets, PyRUA-Lean increases overall success from 63.1% to 71.7%. On instances solved by both agents, it uses 49% fewer LLM calls and 65% fewer input tokens.

[CV-116] GlanceWAM: Sparse Test-Time Imagination for World-Action Models

链接: https://arxiv.org/abs/2608.23927
作者: Linhan Wang,Zijian An,Mingyuan Zhang,Chen Dai,Yi Xu,Can Cui,Jiayan Wang,Zichong Yang,Yinlin Chen,Lifeng Zhou,Chang-Tien Lu
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Robotics (cs.RO)
备注: Add real-robot experiments

点击查看摘要

Abstract:Video generative models provide rich physical priors for robot learning, yet existing world-action models (WAMs) face a fundamental trade-off: synchronous video generation at control rate is latency-prohibitive, while abandoning test-time visual imagination sacrifices task success. We show that visual imagination achieves both real-time inference and superior success rates when generated asynchronously off the critical path and consumed directly in latent space. We introduce GlanceWAM, which decouples imagination from control on a single shared video DiT backbone: an asynchronous proposer glances ahead on a slow clock to imagine a single lookahead frame seconds into the future in the background, while an action head decodes action chunks at control rate (48 ms) purely in latent space without blocking. Enabled by a non-interfering attention mask that isolates video representations and staleness-robust horizon training that accommodates asynchronous lookahead aging, GlanceWAM breaks the speed-success dilemma. Trained purely on demonstrations, it attains 72.2% on the 24-task RoboCasa kitchen benchmark (vs. 67.1% for synchronous Cosmos Policy) and 99.0% on LIBERO while cutting per-chunk control latency 24\times relative to synchronous world-action models (48 ms on one A100). In single-arm and bimanual real-robot manipulation, it achieves higher average success than \pi_0.5 without any robot-data pretraining. Code is available at this https URL.

[CV-117] Iterating Consistency Models: Stability Error Bounds and Noise Schedules

链接: https://arxiv.org/abs/2610.03414
作者: Alessio Spagnoletti,Abdul-Lateef Haji-Ali,Andrés Almansa,Alain Oliviero Durmus,Eric Moulines,Marcelo Pereyra
类目: Machine Learning (stat.ML); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注: 27 pages, 6 figures

点击查看摘要

Abstract:Consistency models (CMs) have become a leading approach for generating high-quality samples in few steps. However, adding steps can improve or degrade sample quality in ways that are highly sensitive to the schedule and that existing theory does not fully explain. To provide accuracy guarantees and guide CM sampler design, we analyze multistep CM sampling as a composition of noising and approximate denoising operators. Under explicit, verifiable stability assumptions, we derive a non-asymptotic error bound that separates contraction of the initialization error from accumulation of approximation error. The bound assigns distinct roles to the schedule: large early noise levels drive contraction, while small late noise levels control the residual bias. As a corollary, we obtain explicit constants for strongly log-concave and semi-log-concave targets. We further establish a complementary guarantee whose assumptions, one-step accuracy and stability, can be estimated for a given trained model. Experiments show that the contraction and approximation profiles entering our bounds can be reliably measured and closely match the predicted functional forms. Together, these results provide a meaningful convergence theory for multi-step CMs and a practical route to sampler design.

[CV-118] Wrong Organ Right Physics: Transferring Echocardiography Pretraining to Lung Ultrasound for Tuberculosis Screening

链接: https://arxiv.org/abs/2610.03290
作者: Christiaan M. Geldenhuys,Joshua M. Jansen van Vüren,Véronique Suttels,Trevor Brokowski,Ablo P. Wachinou,Mary-Anne Hartley,Rensu P. Theart,Grant Theron,Thomas R. Niesler
类目: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注: 10 pages, 3 figures, 4 tables. Accepted at SATNAC 2026, Drakensberg, South Africa, 11-14 October 2026

点击查看摘要

Abstract:Lung ultrasound (LUS) is attractive for tuberculosis (TB) screening at primary-care level, but labelled cohorts are small. Echocardiography carries no such constraint, while sharing the same underlying ultrasound imaging physics, signal processing and B-mode appearance as LUS. We ask whether an encoder pretrained on that high-resource ultrasound domain carries representations that remain usable in the low-resource one. Only the encoder varies, across seventeen encoders spanning three architecture families. Among them, a latent-predictive video encoder pretrained on generic video (V-JEPA2-L) and its echocardiography counterpart (EchoJEPA-L) differ in pretraining corpus alone. The choice among these encoders does not resolve the classification, the whole family spanning 2.50 percentage points against a measurement resolution of 2.71. What moves the task instead is feature conditioning. Standardising the features between the encoder and the classifier improves all seventeen encoders by a mean of +1.23 percentage points at p=1.5\times10^-5 . On the held-out test set every encoder selected on the development folds stands above the baseline system by up to +2.57 percentage points of area under the receiver operating characteristic curve (AUROC), and specificity at 90% sensitivity reaches 79.3% against 60.3%. The contrast specified in advance, EchoJEPA-L against V-JEPA2-L, measures -0.16 percentage points at p=0.926 . We therefore find no evidence that shared ultrasonic physics alone makes echocardiography a more productive pretraining corpus than generic video, and any advantage, if present, is smaller than this cohort can resolve. The video encoders receive replicated still images, however, so whether this absence of an effect reflects the pretraining domain or a video encoder applied to static frames cannot be separated. The limiting factor is the labelled cohort rather than the encoder.

[CV-119] One Photon Many Worlds: Posteriors and Predictions with Single-Photon Cameras

链接: https://arxiv.org/abs/2610.02675
作者: Haejoon Lee(Carnegie Mellon University),Mohit Gupta(University of Wisconsin-Madison),Vijayakumar Bhagavatula(Carnegie Mellon University),Aswin C. Sankaranarayanan(Carnegie Mellon University)
类目: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
备注: 18 pages, 18 figures

点击查看摘要

Abstract:Single-photon avalanche diode (SPAD) cameras operate fundamentally differently from conventional cameras due to their photon-counting nature. Each frame produces a binary image: pixels report zero if no photons arrived during exposure, and one if one or more photons arrived. Reconstructing a scene or inferring its properties from a single binary frame is difficult because many different images could produce the same measurement; thus, the inverse problem is fundamentally one-to-many. As we gather more binary measurements, the inherent uncertainty associated with the inverse problem and any associated inference diminishes. With sufficient photon counts, photon noise becomes negligible relative to the signal mean, enabling near-deterministic scene recovery and inference. This work characterizes the transition from stochastic to near-deterministic scene understanding as photon budget increases, analyzing how the stochasticity in photon arrival affects downstream inference tasks. Technically, we develop a conditional generative framework based on a Hypergeometric frame-thinning process for accumulated binary SPAD measurements. Generative models capture the one-to-many nature of photon-starved inverse problems, enabling empirical characterization of how this ambiguity diminishes with increasing measurements and its impact on downstream tasks like character recognition, QR code decoding, and facial analysis.

[CV-120] Reliability Stress Tests and Decision-Time Routing for Chest X-ray Vision-Language Models ALT

链接: https://arxiv.org/abs/2610.02270
作者: Xinye Yang,Zhusi Zhong,Scott Collins,Grayson Baird,Xuyu Wang,Zhicheng Jiao
类目: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
备注: 6 pages, 4 figures, 5 tables. Accepted version of a workshop paper presented orally at the 2026 IEEE/ACM Conference on Connected Health: Applications, Systems and Engineering Technologies (CHASE). Code and per-case MIMIC-CXR results: this https URL

点击查看摘要

Abstract:Medical vision-language model (VLM) evaluation is sensitive to workflow design, prompting strategy, and benchmark construction, yet most studies treat these factors in isolation. We introduce a reliability stress test for chest X-ray interpretation built on two balanced datasets (a private report-backed set and a curated MIMIC subset). Three medical VLMs (CheXagent, MedGemma-4B, and MedGemma-27B) are evaluated across three prompt styles and two workflows (single-VLM and multi-agent), producing 36 configurations. We show that exact-match accuracy alone can overstate the effectiveness of conservative models that default to “Normal” predictions. Diagnostic reliability also depends heavily on model family and scale: multi-agent reasoning helps some configurations but hurts others. Building on these observations, we propose a decision-time routing framework that selectively escalates to multi-agent inference only when beneficial, improving the cost-quality trade-off over fixed workflows. Our results highlight the need for evaluation protocols that jointly consider prompt sensitivity, failure-mode diversity, and workflow choice before clinical deployment.

[CV-121] Confidence-Gated Cloud-Edge Cascade Triage via Variational Risk Minimization for Medical Imaging ALT

链接: https://arxiv.org/abs/2610.02269
作者: Xinye Yang,Zhusi Zhong,Scott Collins,Michael Bernstein,Grayson Baird,Terrence Healey,Michael Atalay,Mahesh Jayaraman,Xuyu Wang,Zhicheng Jiao
类目: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注: 14 pages, 6 figures, 14 tables. Accepted manuscript of the article published in Smart Health 41 (2026) 100689. Presented as an oral at IEEE/ACM CHASE 2026. Code: this https URL

点击查看摘要

Abstract:Emergency chest X-ray (CXR) triage has a structural modality gap: reports arrive after triage decisions, yet multimodal foundation models require image-text inputs. We present Variational Risk Minimization (VRM), a distillation framework that treats LVLM-generated report variants as Monte Carlo samples of latent clinical interpretations. Rather than distilling from a single teacher target, VRM learns from a variationally marginalized teacher distribution, enabling uncertainty-aware supervision under missing-modality constraints. Under matched encoder families, VRM outperforms direct fine-tuning baselines and improves calibration with strong recovery from hallucinated supervision. Marginalized supervision reduces report-selection instability. In our compact edge-student instantiation, a confidence-gated cascade reaches AUC 0.941 at 103ms average latency with 20.3% cloud escalation, yielding an explicit reliability-latency operating point for cloud-edge clinical workflows.

[CV-122] Event-guided Neural Video Compression

链接: https://arxiv.org/abs/2610.02265
作者: Jiyun Kong,Jungwoo Kim,Enes Eray Demirtas,Touradj Ebrahimi,Jong-Seok Lee
类目: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)
备注: 28 pages. 21 figures

点击查看摘要

Abstract:Neural video codecs derive motion and temporal contexts mainly from RGB frames, leaving room for cross-modal guidance from complementary temporal observations. Event streams can provide such observations by recording brightness changes between frames. In this work, we propose an Event-guided Neural Video Codec (ENVC) that uses events shared by the encoder and decoder to improve RGB compression efficiency. For motion coding, ENVC forms an event-guided motion prior and codes the remaining motion residual. For frame coding, an event-conditioned predictor supplies multi-scale features for gated temporal context refinement. To support training and evaluation on standard video datasets, we synthesize paired RGB-event data and assess its predictive utility through comparisons with real events. Across six benchmarks, ENVC achieves average BD-rate savings of 39.13% using PSNR-RGB and 67.63% using LPIPS relative to DCMVC. Further analyses show that our gains persist on large-motion sequences and that ENVC effectively learns to integrate event information. These results demonstrate the potential of events as a complementary modality for reducing the RGB coding rate. Our model and code are available at this https URL.

[CV-123] ZAGNet: Zone-Aware Graph Aggregation Network for Patient-Level Lung Ultrasound Diagnosis

链接: https://arxiv.org/abs/2610.02263
作者: Li Chen,Shubham Patil,Rashid Al Mukaddim,Jochen Kruecker,Balasundar Raju,Alvin Chen
类目: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
备注: The 2026 IEEE International Ultrasonics Symposium (IUS)

点击查看摘要

Abstract:Patient-level lung ultrasound (LUS) diagnosis requires integrating findings acquired across multiple anatomical zones, yet clinical examinations frequently involve variable and incomplete scanning protocols with missing zones. Existing diagnostic AI methods primarily analyze individual frames or video loops, relying on heuristic aggregation strategies such as max or mean pooling that ignore inter-zone relationships for patient-level inference. This paper presents ZAGNet, a Zone- Aware Graph Neural Network that represents temporally tracked pathology findings as graph nodes connected by anatomical zone adjacency. A graph transformer network propagates contextual information across neighboring lung regions, while a virtual global node aggregates graph-level features to predict patientlevel consolidation and pleural effusion using only patient-level supervision. ZAGNet accommodates missing zones by computing on a graph structure without fixed input format or size. We evaluate ZAGNet on a multicenter dataset of 714 subjects (20,256 LUS video loops) with exams varying from 4 to 16 zones across anterior, posterior, and lateral thoracic regions. For consolidation diagnosis, ZAGNet achieved an AUC of 0.803 compared to 0.677 (max pooling) and 0.674 (mean pooling). For pleural effusion, AUC increased to 0.893 from 0.804 (max pooling) and 0.815 (mean pooling). These represent improvements of up to 19% and 11% for consolidation and pleural effusion respectively. The results demonstrate that graph-based inter-zone reasoning provides an effective and clinically consistent framework for automated patient-level LUS assessment.

[CV-124] oward Controlling Biology with Language:Offline Learning of Prompt-Conditioned Interventions for Cells Organoids and Biobots

链接: https://arxiv.org/abs/2610.02247
作者: Nam H. Le,Douglas Blackiston,Michael Levin,Josh Bongard
类目: Quantitative Methods (q-bio.QM); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Neural and Evolutionary Computing (cs.NE); Robotics (cs.RO)
备注:

点击查看摘要

Abstract:Artificial intelligence increasingly serves as a natural-language interface to complex technical systems, letting people accomplish sophisticated tasks by describing what they want rather than specifying how to do it. Extending this interface to living systems is harder: unlike code or images, a biological intervention has no closed-form linguistic meaning, and the paired language-intervention-outcome data needed to learn such a mapping is expensive to collect, since each example requires its own wet-lab experiment. One way around this is to treat an existing archive of interventions and their already-observed outcomes as a fixed, offline dataset, and use a vision-language model to judge, without any new experiments, whether an archived outcome matches a natural-language description. But whether that judgment is reliable enough to train a language-to-intervention mapping on – without new experiments and without human validation – has remained unclear. Here we show that a natural-language interface for a living organism – a xenobot, a synthetic multicellular construct with no nervous system – can be learned entirely offline this way, using a vision-language model’s own judgment as the sole training reward: an instruction is mapped to the intervention already on record as producing the described behavior. This mapping generalizes to entirely new instructions, evaluated against archive data withheld from training (80.0% held-out accuracy vs a 66.7% chance baseline, matching a network trained directly on ground-truth labels).

人工智能

[AI-0] EyeRobot 2.0: Active Gaze for Precise Manipulation without Wrist Cameras

链接: https://arxiv.org/abs/2610.03710
作者: Kush Hari,Justin Kerr,Nidhya Shivakumar,Samarth Mahapatra,Carmelo Sferrazza,Jiahui Lei,Jitendra Malik,C. Karen Liu,Ken Goldberg,Angjoo Kanazawa
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI)
备注: Project Page: this https URL

点击查看摘要

Abstract:Inspired by human vision, we introduce a framework using active gaze to enable fine-grained bimanual manipulation with only a single stereo camera. EyeRobot 2.0 physically attends to a 3D fixation point in the scene by swiveling two eye viewpoints to center their gaze on it. The resulting images are processed foveally by allocating more visual tokens to the image centers, focusing computation on task-relevant features. Such Active Visual Fixation (AVF) requires carefully coordinated gaze during task execution, which we accomplish hierarchically by first training a low-level gaze servoing policy conditioned on a goal object, then training a target selector which emits fixation goals based on task progress. Both modules are trained with RL on real-world data: the first is trained with a dense geometric reward and the second co-trains with the BC gripper policy which allows it to discover fixation sequences that can resemble a human’s fixation sequence while performing the task. EyeRobot 2.0 further takes advantage of fixation by canonicalizing gripper information into a fixation-relative SE(3) frame, which compacts the size of the action distribution to learn. We collect teleoperation data for 7 real-world and 6 simulated tasks, and conduct over 1000 physical and 1800 simulated robot trials comparing EyeRobot 2.0 against passive stereo and ego + wrist camera policies trained on the same data. Removing wrist cameras is costly for standard policies: with only passive stereo, real-world success drops from 52% to 27%. EyeRobot 2.0 closes this gap with only stereo, outperforming passive stereo by 40% in real and 20% in sim. It matches ego + wrist policies when their wrist views are clear (69% vs. 64%), and more than doubles their success when grasped objects occlude the wrist cameras (48% vs. 22%)

[AI-1] ranscriptome-informed multi-modal AI for predicting neoadjuvant therapy response from breast cancer biopsies

链接: https://arxiv.org/abs/2610.03693
作者: Jungkyu Park,Dhruva Biswas,Joseph Cappadona,Cerise Tang,Ken G. Zeng,Bartosz Machura,Chuwen Liu,Paolo Tarantino,Coral Omene,Francisco J. Esteva,Rohit Bhargava,Marcin Braun,Kamila Paździerz,Jakub Czerwiński,Hanna Romańska-Knight,Albert Grinshpun,Bareket Daniel,Michele Buchinger,Frederick Howard,Piotr Wysocki,Brie Chun,Freya Schnabel,Rich Caruana,Jan Witowski,Krzysztof J. Geras
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Scarcity of labeled data limits development of deep learning biomarkers in oncology. We develop a two-stage AI model predicting pathological complete response (pCR) to neoadjuvant therapy in breast cancer. The first stage learns the transcriptome from histopathology using 8,742 patients across 32 cancer types, corroborated by pathologist review and spatial agreement with measured expression. This simplifies the second stage to predicting pCR from inferred expression and clinical variables. Developed using 1,080 patients (five cohorts) and evaluated in 1,412 patients (nine cohorts), the model achieves a pooled AUROC of 0.79 (95% CI, 0.73-0.85), discriminating responders within molecular subtypes. It outperforms histopathological biomarkers, remaining stable across intratumoral sampling and with minimal biopsy tissue. Ablations show transcriptome-wide inference improves discrimination over clinical variables alone or one-stage pathology models, and robustness by avoiding genomic assays’ gene selection constraints. These results indicate that biologically informed compression may generalize to data-sparse applications in precision oncology.

[AI-2] Revisiting Input Time-frequency Representations in Multi-pitch Estimation for Vocal Ensembles

链接: https://arxiv.org/abs/2610.03656
作者: Junyoung Koh,Hao-Wen Dong
类目: ound (cs.SD); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Multi-pitch estimation in vocal ensembles is challenging because singers occupy overlapping pitch ranges and often sing at closely spaced fundamental frequencies, causing their harmonics to overlap in time-frequency representations. Existing models commonly use harmonic constant-Q transform (HCQT)-based representations to provide frequency-adaptive resolution, at the cost of expensive feature extraction when training mixtures are generated on the fly. We revisit this design and compare HCQT with a linear short-time Fourier transform (STFT), whose frequency bins are directly provided as model inputs. Despite its fixed frequency resolution and the absence of a pitch-aligned input grid, the linear STFT outperforms HCQT while substantially reducing feature-extraction cost. Further analysis shows that a longer analysis window or broader spectral coverage provides no additional improvement, while restricting the input to the predicted pitch range reduces the advantage of the linear STFT. These results suggest that finer frequency resolution does not necessarily improve vocal-ensemble MPE, and that shorter analysis windows can be more effective for time-varying vocal pitches.

[AI-3] Do Large Language Models Know Colombian Law? A Reliability Benchmark for the Colombian Legal System

链接: https://arxiv.org/abs/2610.03639
作者: Rubén Manrique,Michelle Castellanos,Jorge Morales,Juan David Gutiérrez,Antonio Barreto Rozo,Joaquín Vélez Navarro
类目: Artificial Intelligence (cs.AI); Computers and Society (cs.CY)
备注: 38 pages, 23 figures, 8 tables

点击查看摘要

Abstract:Large language models (LLMs) are increasingly used to support legal practice, education, and research, yet their reliability in national legal systems outside the United States remains largely undocumented. We introduce an expert-validated benchmark for evaluating LLM reliability on the Colombian legal system. The benchmark comprises 1,042 items spanning ten areas of law and three question formats (closed multiple-choice, semi-open, and open-ended IRAC), built through a human-in-the-loop pipeline with multi-stage expert review. We evaluate 15 contemporary proprietary and open-weight models with format-appropriate metrics. Accuracy on closed questions ranges widely, from 0.905 (Gemini 3.1 Pro) to 0.577, but on free-text legal answers factual correctness never exceeds 0.45 (on a 0-1 scale) for any model. We find a dissociation between answer relevancy and correctness (Spearman rho = -0.46): models reliably sound responsive while frequently being wrong, a pattern of particular concern for non-expert users. Closed-question accuracy and free-text correctness are strongly rank-correlated (rho = 0.94), so cheap multiple-choice screening predicts model ranking but overstates absolute reliability. An independent rubric-based LLM judge and blind human expert scoring both reproduce the free-text ranking (rho = 0.88). The judge further reveals that only about half of the norms models cite are correct; the rest are wrong or non-existent. Reliability varies systematically by legal area and follows an inverted-U across question complexity. Our results indicate that current LLMs require expert supervision for Colombian legal tasks, and that grounding answers in authoritative sources is a promising path to higher reliability. We release the benchmark construction pipeline to support reproducible evaluation.

[AI-4] Credit Where It Matters: Dependency-Aware Policy Optimization for Terminal Agents

链接: https://arxiv.org/abs/2610.03634
作者: Yu Li,Guangfeng Cai,Long-Fei Li,Shuo Han,Shengtian Yang,Han Luo,Kaibing Yang,Lei Feng
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Terminal-using agents benefit from reinforcement learning (RL) in coding, debugging, and other multi-step terminal tasks. In these tasks, later commands often depend on information or intermediate results produced by earlier commands. However, existing trajectory-level and step-level credit assignment methods do not explicitly trace the read-write dependencies through which commands affect the final outcome. Consequently, training signals could still be assigned to irrelevant operations, weakening learning from relevant steps. In this paper, we propose Dependency-Aware Group Policy Optimization (DepGPO), which uses execution dependencies between commands to guide credit assignment for terminal agents. Specifically, we construct a command dependency graph from execution traces and trace backward from the resources inspected by the task verifier. We then assign credit to relevant writes and their supporting reads along these paths, and use it to redistribute trajectory advantages across steps. Extensive comparative experiments and ablation studies demonstrate that DepGPO improves task performance and training stability on complex terminal tasks.

[AI-5] NeutronGym: Physics-Graded Neutron Instrument Design for LLM Agents

链接: https://arxiv.org/abs/2610.03631
作者: Lijie Ding,Changwoo Do
类目: Artificial Intelligence (cs.AI); Instrumentation and Detectors (physics.ins-det)
备注:

点击查看摘要

Abstract:Designing a scientific instrument tests whether language-model agents can do physics rather than recall it, provided the grading cannot be argued with. We introduce NeutronGym, to our knowledge the first executable environment for neutron instrument design: agents build instruments through validating tools, McStas ray-traces what they build, and a level-resolved ladder grades syntax, runtime, structure and science with no LLM judge. Procedural families supply unlimited instances of a fixed layout whose design parameters the agent must set, with held-out parameter regimes; a curated slice, McStasBench, adds 16 tasks from published instruments behind memorization probes and a sandbox. Seven models reproduce at most 7 of the 16, none retrieves a reference, and none meets an improvement target. The environment also trains. Reinforcement learning on its reward takes Qwen3-8B from 11% to 77% of held-out instances of a family whose targets come from a hidden design (69% at a second seed), past an untrained Qwen3-32B, and the recipe holds, at one seed each, on three further gated families. The analysis says what that gain is. Without the ladder’s partial credit it collapses by 60 points. From reward alone the trained model reaches what a classical optimizer reaches, at the agent’s simulation budget, only when handed the closed-form physics (77% against 81%, a gap that does not separate at this size), while frontier models still solve 98-99%. Getting a trustworthy result meant failing four task designs that no-model baselines could solve, and we release the probes that found them.

[AI-6] Depth as Time in One-Step Generative Models

链接: https://arxiv.org/abs/2610.03626
作者: Arnold Caleb Asiimwe,William Yang,Sanghyuk Chun,Esin Tureci,Olga Russakovsky
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:The recent wave of one-step generative models, which compress the multi-step trajectory of diffusion via either distillation or learned flow maps, has reached an inflection point where they can generate high-quality images. Here, we ask a natural question that follows from these advances: what happens to the denoising trajectory of multi-step diffusion when generation is compressed into a single forward pass? We offer an empirical observation we call \textitdepth as time: the denoising computation that multi-step diffusion performs across sampling steps appears to unfold across the depth of a single forward pass, and can be recovered by decoding intermediate layers with the model’s own output head. Most interestingly, we show that this depthwise computation depends on the transport task a flow map is trained to solve. The most surprising case is MeanFlow, where probing shorter transport intervals reveals both denoising and renoising within a single network evaluation. In contrast, generators trained without a time-indexed transport task, such as drifting models, do not exhibit the same depthwise denoising. Consequently, we show that models that exhibit the depthwise denoising phenomenon are more compressible across the layerwise computation: a MeanFlow \textttSiT-L/2 model can be compressed by 16.6\times in parameters into a single time-conditioned block. We offer an explanation for this denoise-then-renoise behavior and show that, when we treat the layerwise computation explicitly as a flow, a single time-conditioned block can be trained to denoise across layers, compressing a MeanFlow \textttSiT-L/2 model by 16.6\times in parameters. Together, these results suggest that the temporal computation of diffusion is not eliminated by one-step generation, but reorganized across network depth.

[AI-7] HazardWeaver: Scientific Route Selection for Hazard Analysis Agents

链接: https://arxiv.org/abs/2610.03591
作者: Wangshu Zhu,Xueqi Cheng,Liang Wu,Yushun Dong
类目: Artificial Intelligence (cs.AI)
备注: 24 pages, including references and appendices. Code is available at this https URL

点击查看摘要

Abstract:Understanding and assessing natural hazards is essential for disaster preparedness and risk reduction. Recent advances in large language models have spurred growing interest in AI agents for hazard analysis, particularly their ability to integrate scientific data, models, and tools into automated workflows. However, effective automation requires agents to determine which scientific methods are appropriate for a given event and executable with the available data and tools. As new evidence and execution results become available, these conditions can change, requiring agents to reconsider their choices. We formulate this problem as state-dependent scientific route selection and introduce HazardWeaver. Specifically, HazardWeaver first leverages the Hazard Knowledge Compiler to extract evidence-linked conditions governing scientific applicability, then its Hazard Capability Graph represents executable scientific capabilities and checks compatibility between their inputs and outputs. Using these complementary representations, the Hazard Weaver Agent component selects applicable and executable routes, carries out their workflows, and revises its decisions as the analysis state changes. To evaluate both the scientific outputs and the decisions that produce them, we introduce the Hazard Weaver Benchmark, comprising 141 instances across seven single-hazard domains and four multi-hazard interaction classes. The benchmark accommodates multiple valid scientific routes and evaluates output correctness, route validity, and justified abstention. Extensive experiments on this benchmark show that HazardWeaver outperforms existing agent systems, with the largest gains on tasks with multiple eligible scientific routes. Our code is publicly available at this https URL.

[AI-8] hreat-Preserving Representation Sensitivity in Agent -Security Benchmarks

链接: https://arxiv.org/abs/2610.03585
作者: Neeraj Karamchandani,Piyush Nagasubramaniam,Xinhong Xie,Sencun Zhu,Dinghao Wu
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: 12 pages, 2 figures

点击查看摘要

Abstract:Security benchmarks for LLM-based agents often report the attack success rate (ASR) as a measure of model robustness and use these scores to compare different models and defense mechanisms, assuming that they describe the security of the agent. In this paper, we explore whether it also influences the benchmark’s measurement. To measure the effect of the benchmark representation, we introduce threat-preserving representation sensitivity (TPRS), which measures how much the ASR changes when we change the agent-visible representation while holding the underlying task, harmful action, security policy, ground truth, environment, and the evaluation criteria fixed. On Agent Security Bench (ASB), replacing threat-related tool names with threat-neutral names raises the committed attack success rate by 11.67 percentage points on GPT-5-mini and by 13.21 points on Claude Haiku 4.5. On MCPTox, replacing the original neutral tool name with an explicit threat-related name lowers the ASR by 11.00 percentage points on GPT-5-mini and 4.11 points on Claude Haiku 4.5. On AgentDojo, adding threat-related wording to the attack-relevant tool changes ASR by only 0.50 percentage points on GPT-4o-mini, yet the benign utility falls by 5.36 points on tasks requiring that tool. We ran an experiment on MCPTox where we observed that a threat-neutral name matched on token count, length, and casing reproduces most of the shift produced by the threat-explicit name (8.54 of 11.00 points on GPT-5-mini). The results show that a security score measured under one representation may fail to generalize across threat-preserving representations of the same security problem. Robustness claims should therefore be supported by performance across a controlled set of threat-preserving representations rather than relying on a single representation-dependent score. Comments: 12 pages, 2 figures Subjects: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Machine Learning (cs.LG) Cite as: arXiv:2610.03585 [cs.CR] (or arXiv:2610.03585v1 [cs.CR] for this version) https://doi.org/10.48550/arXiv.2610.03585 Focus to learn more arXiv-issued DOI via DataCite (pending registration) Submission history From: Neeraj Karamchandani [view email] [v1] Fri, 2 Oct 2026 16:55:50 UTC (257 KB) Full-text links: Access Paper: View a PDF of the paper titled Threat-Preserving Representation Sensitivity in Agent-Security Benchmarks, by Neeraj Karamchandani and 4 other authorsView PDFHTML (experimental)TeX Source view license Current browse context: cs.CR prev | next new | recent | 2026-10 Change to browse by: cs cs.AI cs.LG References Citations NASA ADSGoogle Scholar Semantic Scholar export BibTeX citation Loading… BibTeX formatted citation loading… Data provided by: Bookmark checked="checked"class=“labs-tab-input”> Bibliographic Tools Bibliographic and Citation Tools Bibliographic Explorer Toggle Bibliographic Explorer (What is the Explorer?) Connected Papers Toggle Connected Papers (What is Connected Papers?) Litmaps Toggle Litmaps (What is Litmaps?) scite.ai Toggle scite Smart Citations (What are Smart Citations?) Code, Data, Media Code, Data and Media Associated with this Article alphaXiv Toggle alphaXiv (What is alphaXiv?) Links to Code Toggle CatalyzeX Code Finder for Papers (What is CatalyzeX?) DagsHub Toggle DagsHub (What is DagsHub?) GotitPub Toggle Gotit.pub (What is GotitPub?) Huggingface Toggle Hugging Face (What is Huggingface?) ScienceCast Toggle ScienceCast (What is ScienceCast?) Demos Demos Replicate Toggle Replicate (What is Replicate?) Spaces Toggle Hugging Face Spaces (What is Spaces?) Spaces Toggle TXYZ.AI (What is TXYZ.AI?) Related Papers Recommenders and Search Tools Link to Influence Flower Influence Flower (What are Influence Flowers?) Core recommender toggle CORE Recommender (What is CORE?) Author Venue Institution Topic About arXivLabs arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv’s community? Learn more about arXivLabs. Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?) mathjaxToggle(); We gratefully acknowledge support from our major funders, member institutions, , and all contributors. About Help Contact Subscribe Copyright Privacy Accessibility Operational Status (opens in new tab) Major funding support from

[AI-9] HyperBrowseComp: A Multilingual and Multimodal Stress Test for Web-Browsing Agents

链接: https://arxiv.org/abs/2610.03574
作者: Alham Fikri Aji,Faiz Rizki Ramadhan,Zayd M. K. Zuhri,Seung Hun Eddie Han,Ryandito Diandaru,Qinrong Cui,Jan Christian Blaise Cruz,Badrinath Chandana,Peerawat Chomphooyod,Ahmed Attia,Jonibek Mansurov,Emilio Villa-Cueva,Canh Duong Nguyen,Imran Turganov,Minghao Wu,Peerat Limkonchotiwat,Irina Nikishina
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:We introduce HyperBrowseComp, a multilingual and multimodal browsing benchmark comprising 423 manually authored and human-validated questions across 13 languages, written by native or highly proficient speakers. Questions are designed to be extremely challenging. Each question targets a concise, publicly verifiable answer whose discovery requires locating obscure evidence, following multi-step clue chains, or inspecting heterogeneous sources such as videos, scanned documents, images, or maps. Easier questions are filtered out by evaluating them with models without internet access to reduce the likelihood that they can be answered with parametric knowledge alone. We evaluate several models using provider-native search and a shared external retrieval harness under a common agent protocol. To contextualize model performance and effort, we also conduct a human evaluation on a sample of the questions. HyperBrowseComp provides a challenging testbed for persistent information seeking across languages and evidence modalities, with difficulty arising from discovering and connecting evidence on the open web.

[AI-10] Learning to Assess Heartbeat Observability for mmWave Heart-Rate Sensing

链接: https://arxiv.org/abs/2610.03570
作者: Yuxuan Hu,Shilin Shan,Jianfei Yang,Feng Xu
类目: Artificial Intelligence (cs.AI)
备注: 19 pages, 11 figures. Project page: this https URL

点击查看摘要

Abstract:Contactless heart-rate sensing with millimeter-wave (mmWave) radar requires assessing whether individual measurements support reliable estimation. We study learning to assess heartbeat observability, defined as the readability of the heartbeat component in an acquired phase spectrum, for selective heart-rate estimation. Coherent superposition of scatterer returns can suppress this component even under similar macroscopic observation geometry, motivating assessment directly from acquired measurements. To obtain training supervision across different observability conditions, we develop a controllable multi-scatterer frequency-modulated continuous-wave (FMCW) simulator. Agreement between the dominant heartbeat-band peak and the known heart rate provides an automatic observability label for each simulated measurement. We propose HEAR (Heartbeat Estimation with Assessed Reliability), a compact dual-task Transformer that jointly predicts an observability score and heart rate. Its input combines spectral magnitudes with frequencies relative to the respiration fundamental, providing context for respiratory harmonics. Trained solely on simulated observations, HEAR transfers zero-shot to two public real-world datasets collected at 60 and 120 GHz from 134 subjects. The same learned score supports selective prediction with both HEAR’s own heart-rate head and multiple existing estimators. On the 120 GHz dataset, score-based selection reduces the HR head’s mean absolute error from 17.9 BPM at full coverage to 1.6 BPM at 50% coverage. The complete pipeline achieves an end-to-end processing latency of 50.8 ms on an edge device. Project page: this https URL.

[AI-11] Knowledge or Calculator? Decomposing the Skill Premium in Verifiable Financial Agent Workflows

链接: https://arxiv.org/abs/2610.03564
作者: Jermyn Zhen Yong Bek,Zhuang Qiang Bok,Zhongtian Sun
类目: Artificial Intelligence (cs.AI)
备注: 10 pages, 1 figure, 9 tables

点击查看摘要

Abstract:Financial AI agents must do more than retrieve facts: investment workflows require correct quantitative execution, reliable use of procedural resources, and auditable structured outputs. We introduce FinSkillBench, an evaluation suite of 2,603 point in time episodes across 12 subtasks in portfolio construction, risk management, and fundamental analysis, with hidden regenerable ground truth and task specific deterministic verifiers. Executing 17,820 episodes across 9 models and 3 resource conditions, the paired analysis across 8 models shows that curated skill packages raise mean scores by +16.2 points (0.366 to 0.528), whereas skills generated within a single episode add only +0.5 points while consuming more tokens and turns. We then decompose the curated premium by granting human authored procedural documents and executable domain tools separately: documents alone add +5.6 points, tools alone add +19.5 points, and their combination is subadditive. The premium is strongly workflow dependent: executable tools dominate numerically intensive workflows, documentation matters more when procedural or output schema guidance is the bottleneck, and interpretive tasks benefit from both. The effects are sign stable across 10 scoring variants and cluster bootstrap analyses, and an independently implemented second harness reproduces the directional pattern while showing that effect magnitudes depend on how tools and data are exposed. Overall, a measured “skill premium” is a property of the full model, resource, and harness system rather than of the underlying model alone.

[AI-12] Cephalonauts One: A deep fMRI dataset for decoding naturalistic speech in the human brain NEURIPS2026

链接: https://arxiv.org/abs/2610.03558
作者: Antoine Collas,Louis Jalouzot,Géraud Ilinca,Corentin Caris,Romain Valabrègue,Ahmed Hassayoune,David Goncalves,Madeleine Hueber,Thaddée Delebarre,Julien Savatovsky,Clara Fonteneau,Charles Maussion,Bertrand Thirion,Alexis Thual
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: Accepted at NeurIPS 2026, Evaluations Datasets Track

点击查看摘要

Abstract:Cephalonauts One is a whole-brain 3 Tesla (3T) functional magnetic resonance imaging (fMRI) dataset recorded while subjects listened to audio podcasts. Three healthy subjects underwent multiple scanning sessions, each consisting of five 15-minute runs, while listening to podcasts in their native language. With 30 hours of fMRI data per subject, the current release is the deepest available fMRI dataset using naturalistic speech stimuli. The dataset pairs brain activity with the corresponding podcast audio, transcript annotations, and derived stimulus embeddings. Furthermore, we introduce a brain decoding benchmark formulated as audio segment retrieval: given fMRI activity from a held-out session, the decoder must identify the corresponding time-aligned podcast audio segment among candidate segments. We provide standardized splits, evaluation metrics, and baseline decoders for this task. Finally, a scaling analysis shows that decoding performance improves continuously with the amount of training data per subject.

[AI-13] Recursive Harness Self-Improvement for Frontier Reasoning Data Synthesis

链接: https://arxiv.org/abs/2610.03548
作者: Wenlong Zhang,Zhengbo Jiao,Chenxu Zhang,Lekang Jiang,SiYuan Ma,Qituan Zhang,Guo Chen,Linfeng Zhang
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Generating progressively harder reasoning problems requires synthesis procedures that adapt as the task distribution evolves. Existing task-level recursion reuses generated problems as seeds but leaves the construction harness unchanged. We present task-harness co-evolution, a framework for recursive harness self-improvement (RSI) in reasoning-data synthesis. Online self-improvement converts intermediate solver failures into reusable skills during generation. Post-task self-improvement revises skills, prompts, and workflows after each batch, adopting candidates only when they generate harder valid tasks within a bounded cost increase. Model weights and verification criteria remain fixed. Across mathematics, coding, and science, mean solver accuracy decreases from 100.0% to 54.8% over fourteen evolution rounds. Ablations show that combining both update schedules produces harder tasks than fixed-harness recursion or either schedule alone. The resulting data improves downstream SFT and GRPO performance. In particular, a 27B student fine-tuned on 10K synthesized mathematics examples achieves 62.5% mean-16 accuracy on APEX, competitive with selected frontier-model references. These results support adapting the synthesis harness alongside the tasks to generate increasingly challenging data with downstream training value.

[AI-14] Beyond Trained Models: Compiling GNNs for a Sound Explainer Benchmark

链接: https://arxiv.org/abs/2610.03526
作者: Steve Azzolin,Francesco Paolo Nerini,Stefano Teso,Francesco Bonchi,Bruno Lepri,André Panisson,Andrea Passerini
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: Preprint

点击查看摘要

Abstract:Explainers for Graph Neural Networks (GNNs) are commonly evaluated by their plausibility, i.e., how well their explanations recover a predefined ground truth, such as a motif planted in the data. This protocol implicitly assumes that a GNN trained on such data relies on the intended motif. Although prior work has questioned this assumption, plausibility remains widespread. First, we show that the assumption is violated on several widely used benchmarks, where, e.g., degree statistics alone suffice to solve the task. Then, we remove this confounder by replacing training with compilation. We achieve this by introducing \mathsfGracr , the first compiler translating graded modal logic formulas into GNN weights, yielding models that replicate the behaviour of the corresponding formulas. Since the behaviour of the model is now known by construction, we can define its ground truth explanation formally and compute it exactly. Building on this, we introduce \mathsfGracr\mathsfBench , a benchmark of compiled GNNs for the evaluation of explainers against this exact ground truth. Experiments on eleven explainers across six tasks show its effectiveness for fine-grained diagnostic evaluation: notably, we discover that most explainers are not robust to indirect influences or alternative implementations of the same formula. These results position \mathsfGracr\mathsfBench as a novel, rigorous evaluation setting for graph post-hoc explainability.

[AI-15] From Benchmarks to Production: A Text-to-SQL System for Complex Financial Data EMNLP

链接: https://arxiv.org/abs/2610.03524
作者: Arijit Sehanobish,Bruno Gomes Coelho,Guillaume Michel,Sophia Zhi,Valerie Faucon-Morin,Kristen Howell
类目: Artificial Intelligence (cs.AI); Databases (cs.DB); Machine Learning (cs.LG)
备注: EMNLP Industry Track 2026

点击查看摘要

Abstract:General-purpose Text-to-SQL systems achieve strong performance on academic benchmarks like Spider and BIRD, where schemas are relatively shallow and column values are often human readable. In production financial databases, where concepts are stored as opaque integer keys rather than human-readable strings, these methods fall below 50%, as even simple queries require multiple joins and filter predicates reference opaque IDs. We present Financial LINking Text-to-SQL (FLINT), a domain-specialized Text-to-SQL system that closes this gap through three key components: (1) a lookup agent that dynamically resolves natural-language concepts to question-specific reference table constraints, (2) embedding-based retrieval of structurally similar query templates from a compact, expert-authored bank, and (3) schema linking that prunes a large table schema to the relevant subset by traversing foreign-key chains, rather than relying on name similarity alone. We evaluate on two datasets totaling 359 questions over production financial schemas. FLINT outperforms various state-of-the-art baselines using the same LLM. The system is deployed in production as part of a financial data retrieval service.

[AI-16] Reasoning Models Are Accurate but Unsound on Identification

链接: https://arxiv.org/abs/2610.03519
作者: Arman Behnam,Binghui Wang
类目: Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR)
备注:

点击查看摘要

Abstract:A reasoning model asked whether a causal effect is recoverable from observational data can fail in two ways: it refuses an identifiable query or answers a nonidentifiable one. The latter is more consequential, as no observational data can validate the claimed formula. Measuring this failure requires queries that are provably non-identifiable, which prior evaluations lack, and grading that accepts correct formulas in any equivalent form, which string matching cannot provide. We build CERTID, a formal identification pipeline that addresses both limitations. CERTID uses the sound and complete causal identification algorithm ID to certify whether an effect is identifiable from a given graph and query, and verifies returned formulas against structural causal models whose interventional distributions are known exactly. CERTID further develops theoretical results to mitigate structural leakage, repair non-identifiable queries, and establish grading guarantees. We evaluate three frontier reasoning models (Gemini Flash, Gemini Pro, and GPT5.5) on 1,200 certified instances spanning 4 to 50 vertices. Accuracy proves a poor proxy for soundness: on identical instances, the false-claim rate on non-identifiable queries varies by seventeen-fold across models. We also find that models decide identifiability with 97-100% accuracy on graphs generated after the strongest model’s training snapshot. Instances, the certification procedure, the verifier, and per-instance records are available at this https URL.

[AI-17] Efficient Reasoning Training Does Not Always Harm CoT Faithfulness and Monitorability

链接: https://arxiv.org/abs/2610.03509
作者: Samuel Lewis-Lim,Xingwei Tan,Mario Sanger,Zhixue Zhao,Nikolaos Aletras
类目: Artificial Intelligence (cs.AI)
备注: Under Review

点击查看摘要

Abstract:Chain-of-thought (CoT) reasoning allows humans to inspect how large language models reach their answers, and oversee model behaviour. This reasoning comes at an increased inference cost, motivating efficient methods that train models to solve tasks using fewer tokens. However, a common concern is that such training may cause models to skip important reasoning steps, so the CoT no longer faithfully reflects the model’s decision. It is unclear whether or when this occurs in practice, since different efficiency methods apply length pressure to models’ CoT in distinct ways, and faithfully explaining a model’s decision takes more tokens on some tasks than others. To understand these dynamics, we fine-tune a variety of models with three methods that apply length pressure differently, namely a fixed generation budget, a per-example length target, and a group-relative length reward. We evaluate how efficient reasoning affects CoT faithfulness (i.e., how well the CoT reflects model decisions on related inputs) and monitorability (i.e., whether the CoT reveals when input interventions alter the output). We find that it affects faithfulness and monitorability differently. Faithfulness falls in most settings, primarily because the trained models are less consistent. Monitorability is more robust, as models keep acknowledging the influence on their answer even when the CoT is much shorter.

[AI-18] Certified Mechanistic Edits: Behavioral Guarantees for Skill Removal and Preservation

链接: https://arxiv.org/abs/2610.03502
作者: Md Sazid Uddin,Md. Khairul Alam Mazumder,M. F. Mridha
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR)
备注: 12 pages, 6 figures, 4 tables

点击查看摘要

Abstract:Mechanistic edits (ablations, weight edits, activation steering) are the standard tools for unlearning a harmful capability from a neural network while preserving useful ones. Current approaches validate their effects only by testing, which can never cover an entire continuous region of inputs. Prior work at the interpretability-verification boundary certifies descriptions of a model: what a circuit computes, or whether it faithfully explains the whole. We instead certify the behavioral effect of an edit: that disabling a circuit removes one skill and provably preserves another, for every input in a region; a feature non-interference guarantee in the information-flow-security sense. We demonstrate such certified edits from toy ReLU networks up to a standard softmax + LayerNorm transformer, proving removal and preservation over continuous embedding-space regions and reaching roughly 9x the input-perturbation dimension an exact solver can handle by switching to sound bound propagation. Furthermore, we prove that no finite deterministic black-box test can certify removal, exhibiting an edit that passes exhaustive testing yet provably fails on a survivor pocket that can be made arbitrarily small. Guarantees hold on small, standard-architecture networks and, like any removal claim, presuppose that the target skill admits a decidable specification, a property which real-world harms may not have.

[AI-19] Detect and Suppress: A Mechanistic Defense against Adversarial Patches in VLA Models

链接: https://arxiv.org/abs/2610.03498
作者: Yukiya Horiba,Koshiro Aoki,Shunsuke Yasuki,Bum Jun Kim,Taiki Miyanishi
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Adversarial patches can disrupt Vision-Language-Action (VLA) models by manipulating visual observations, leading to failures in robot control. However, it remains poorly understood which internal mechanisms underlie these failures and how targeted interventions can mitigate them. In this work, we mechanistically analyze VLA representations using a sparse autoencoder (SAE) and identify a feature whose activation strongly correlates with the presence of an adversarial patch. Based on this analysis, we suppress the identified feature at inference time only when a linear probe detects an attack. This intervention improves robustness without the cost of fine-tuning the VLA. We evaluate our method against VLA adversarial patch attacks on LIBERO-10. Conditional intervention improves success rate under intermittent attacks, whereas continuously applying the same intervention substantially degrades policy performance. These results show that attack-related internal representations can provide useful targets for VLA adversarial defense and that controlling when to intervene is important for limiting disruption to nominal policy behavior.

[AI-20] MobiAgent : Dual-Loop Recursive Policy Self-Improvement for Long-Horizon Mobile Manipulation

链接: https://arxiv.org/abs/2610.03476
作者: Chenzhi Liu,Yue Zhang,Jiehong Lin,Jianan Wang,Bo Wang,Zhongrui Wang,Xiaojuan Qi
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI)
备注: Accepted at the Conference on Robot Learning (CoRL) 2026. Project page: this https URL

点击查看摘要

Abstract:Long-horizon mobile manipulation presents significant challenges due to compounding execution errors and capacity interference between locomotion and arm control. While recent Vision-Language-Action models excel at short-horizon tasks, they lack the hierarchical reasoning required for multi-stage objectives. Furthermore, existing hierarchical agents suffer from rigid sub-task mapping, inflexible replanning, and a lack of continuous learning. To address these limitations, we introduce MobiAgent, a dual-loop agentic framework that bridges robust deployment execution and recursive policy self-improvement. During deployment, the Inner Loop decouples high-level reasoning from low-level control through highly composable atomic skills. It employs Vision-Language models for receding-horizon planning and visual reflection, dynamically composing skills to ensure robust error recovery. These skills are executed by specialized flow-matching experts that share a unified VLM backbone, maximizing reusability while mitigating capacity interference. Concurrently, the Outer Loop drives automated lifelong learning by autonomously segmenting and verifying deployment rollouts, clustering them to discover atomic skills, and continuously fine-tuning the skill library without human annotations. Evaluations on RoboCasa, BEHAVIOR-1K, and real-world tasks demonstrate the effectiveness of MobiAgent. It outperforms \pi_0.5 -TA by 22.5 percentage points on BEHAVIOR-1K and enables robust recovery from execution failures. Through autonomous data recycling, success improves from 7.50% to 27.50% on RoboCasa and from 32.5% to 57.5% on Astribot S1.

[AI-21] Single or Multiple Policies for Phase-Structured Reinforcement Learning?

链接: https://arxiv.org/abs/2610.03475
作者: Guilhem Loussouarn,Nancy Nayak,Kin K. Leung
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Systems and Control (eess.SY)
备注: 40 pages, 13 figures, main paper with appendix

点击查看摘要

Abstract:Many reinforcement-learning (RL) problems are non-stationary yet structured and can be decomposed into phases, each with its own transition probabilities and reward functions. When the phase sequence is known, the common solution augments the state with information to satisfy the Markovian property and applies standard RL techniques. However, prior work finds that the multi-policy approach for different phases can outperform a single state-augmented policy shared among the phases, for reasons that remain unclear. In this work, we first show that the shared policy can theoretically achieve performance of any multi-policy solution. However, whether a multi-policy solution can perform better than the corresponding single shared policy in practice depends on function approximation, learning and optimization processes, as well as, for multi-policy solutions, the sample efficiency and loss of continuity from one policy to another. We propose a regime-based phase decomposition method to identify which policy can provide better performance. The method is based on consideration of the duration of transient system dynamics relative to the duration of the quasi-stationary period. Numerical experiments are conducted with different non-stationary RL problems to validate our four major hypotheses: (a) longer phase durations favor multi-policies, (b) the heterogeneity between phases increases the burden on single policy, © multi-policies need sufficient data for each phase, and (d) environment-specific transition dynamics between phases can affect which policy is preferable.

[AI-22] Measure Less Know More: Self-Supervised Test-Time Feature Acquisition NEURIPS2026

链接: https://arxiv.org/abs/2610.03454
作者: Eeshaan Jain,Linus Bleistein,Bart Deplancke,Charlotte Bunne
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: Accepted to NeurIPS 2026

点击查看摘要

Abstract:Recent progress in multimodal, high-dimensional learning has enabled foundation models to process heterogeneous, large-scale data. However, at test time, acquiring all features or modalities can be prohibitively costly and often redundant. Sequentially selecting informative modalities is therefore critical, yet challenging when the downstream task or prediction target is unknown. To this end, we introduce ECHO- k , a task-agnostic and self-supervised learning principle for modality acquisition: we use a deep model’s internal pretrained representations (e.g., from a foundation model) as proxy targets that summarize cross-modal information. We provide theoretical guarantees in a stylized linear setting that motivate a reinforcement learning (RL) policy for sequential modality selection. Across task-agnostic and label-free acquisition baselines, ECHO- k consistently improves budgeted downstream performance across diverse foundation-model backends. Our method provides a principled route to cost-aware test-time deployment, with implications for any multimodal system where measurements are expensive or time-constrained, and downstream tasks unknown a priori.

[AI-23] OptiSelect: How does the Optimizer Shape Data Curriculum?

链接: https://arxiv.org/abs/2610.03432
作者: Simin Fan,Alireza Abdollahpoorrostam,Martin Jaggi
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Online data selection has demonstrated substantial efficiency gains for LLM pretraining by training on the most valuable candidates within each batch. Since a candidate’s value is realized through its effective model update, principled selection should account for the optimizer step, which reshapes the raw gradient before it updates model parameters. We formalize this optimizer-aware selection paradigm as OptiSelect and present the first systematic study of how the optimizer shapes data selection. Our theory establishes a selection gain principle in which the advantage of online selection is governed by the discriminability of the optimizer-induced utility scores. We prove that sign-based and polar-tangential preconditioners of Lion and Muon would suffer from a discriminability collapse which caps attainable gains from OptiSelect, whereas diagonal-adaptive optimizers such as AdamW and Sophia admit strictly better upper bounds. The proposed principle also yields a quantitative derivation of the optimal candidate oversampling ratio. Pretraining experiments on 124M and 720M models are consistent with our theoretical analysis and show that AdamW’s diagonal-adaptive scoring geometry remains the strongest scoring geometry even with Muon as optimizer. We further demonstrate that OptiSelect retains its benefits under data rephrasing, a technique used in modern data processing pipelines. Our findings provide theoretical foundations and practical guidance for co-designing optimizers and data selection in LLM pretraining.

[AI-24] Jumping the Line: Exploiting Length Predictions in LLM Scheduling

链接: https://arxiv.org/abs/2610.03430
作者: Yuyang Dai,Rana Shahout,Mahmood Sharif
类目: Artificial Intelligence (cs.AI)
备注: 24 pages, 6 figures

点击查看摘要

Abstract:Efficient request scheduling is increasingly important for reducing completion time in large language model (LLM) serving. Size-based policies such as Shortest Job First prioritize shorter requests, but output lengths are unknown before generation, so practical schedulers rely on predicted lengths. We introduce JIL, an attack on prediction-based LLM schedulers that manipulates the scheduling signal to obtain higher priority and reduce completion time. Using TRAIL as a case study, JIL optimizes an adversarial suffix that causes a lightweight output-length probe to underestimate a request’s length. We evaluate JIL on two datasets and four LLMs across varied request profiles and deployment configurations. JIL reduces predicted output lengths by up to 83.4 percent, and adversarial requests complete up to 1.53 times faster on average in end-to-end serving experiments. The reduction in predicted length is substantially larger than the change in actual output length, revealing a mismatch between the scheduler’s estimate and the request’s realized size. Response utility varies across models and tasks, exposing a trade-off between scheduling advantage and response quality. We also evaluate scheduler-side defenses and find that grouping length predictions into coarse intervals reduces JIL’s scheduling advantage and mitigates delays to benign requests.

[AI-25] Becoming Suspicious Across Borders: Algorithmic Extraterritoriality and AI-Driven Financial Surveillance

链接: https://arxiv.org/abs/2610.03425
作者: Georgios Pavlidis,Savvas Chatzichristofis,Eleni Gavriil
类目: Artificial Intelligence (cs.AI)
备注: Open Access Publication

点击查看摘要

Abstract:Suspicion is an important, yet elusive concept in anti-money laundering and counter-terrorist financing (AML/CFT), which allows for intervention below the threshold of proof. In its traditional form, suspicion can be understood as a situated legal judgement by human actors within identifiable jurisdictions. It is argued that this understanding is no longer adequate. As artificial intelligence (AI) becomes an integral part of financial surveillance, suspicion is increasingly produced through data-driven processes. This transformation is epistemic, but also spatial. Since AI-driven financial surveillance operates through transnational data infrastructures, regulatory reach is less a matter of where conduct occurs than a question of whether such conduct becomes visible within data systems. This article develops the concept of algorithmic extraterritoriality, understood as a form of regulatory power mediated by data infrastructures rather than formal assertions of jurisdiction. Moreover, since individuals are increasingly constituted as datafied subjects of suspicion, they are rendered governable through dispersed and opaque processes of evaluation. This constitutes a challenge for accountability and contestability because suspicion becomes more difficult to locate, explain or contest.

[AI-26] Rethinking Epistemic Uncertainty in Node Classification through Information Growth

链接: https://arxiv.org/abs/2610.03418
作者: Emma Meneghini,Francesco Ferrini,Bruno Lepri,Andrea Passerini,Veronica Lachi
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Epistemic uncertainty should decrease as additional information about the data-generating process (DGP) becomes available to the predictor. Yet, existing graph evidential deep learning (EDL) methods for node classification typically construct epistemic uncertainty from graph-specific properties and evaluate it on downstream tasks such as out-of-distribution detection, which do not test its reducibility as information about the DGP increases. To make reducibility directly testable, we introduce a statistical framework for studying epistemic uncertainty under information growth. Our framework specifies an information-growth experimental protocol and a consistency criterion for epistemic predictors, while using projective graph DGPs to ensure that growing graphs, which in general need not provide increasing information about the same DGP, constitute coherent observations of the same underlying process. We show that EDL methods do not explicitly estimate data uncertainty arising from a single finite graph observation and instead regulate epistemic uncertainty through model hyperparameters, precluding consistency, as corroborated by controlled information-growth experiments. As an alternative, we propose graph bootstrap ensembles, capturing both data and procedural uncertainty through graph resampling and randomized training. Under the same experimental protocol, these ensembles exhibit epistemic uncertainty reduction beyond standard deep ensembles. These findings support bootstrap ensembles as candidate consistent epistemic predictors under information growth.

[AI-27] DriftTTS: Few-Step Text-to-Speech Without Distillation via Distribution-Matching Drift

链接: https://arxiv.org/abs/2610.03390
作者: Mohammad Nur Hossain Khan,Subrata Biswas,Bashima Islam
类目: ound (cs.SD); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Few-step neural text-to-speech models often rely on short- ened diffusion or flow-matching schedules, or on distillation from pretrained multi-step teachers. To avoid these depen- dencies, we present DriftTTS, a few-step mel-spectrogram generator trained without a generative teacher, distillation, or adversarial discrimination. DriftTTS uses a distribution- matching drift objective in a mel-domain feature space defined by raw mels and a frozen masked-autoencoder encoder pretrained on the same LJSpeech training split. On-policy rollout trains the decoder on its own interme- diate states and supports inference up to the trained roll- out depth. On LJSpeech, DriftTTS at NFE=4 achieves 3.87 dB MCD and 3.7% WER, compared with 3.85 dB and 3.4% for Matcha-TTS. In a fully paired blind listen- ing test, DriftTTS obtains 4.18 MOS, compared with 3.96 for Matcha-TTS and 4.22 for ground truth. These results demonstrate competitive few-step synthesis without a pre- trained generative teacher. Code can be found at https: //github.com/BASHLab/driftTTS.git

[AI-28] CVE2AP: Automated Generation of PDDL-Encoded Attack Paths via Large Language Models

链接: https://arxiv.org/abs/2610.03383
作者: Lin Cui,Vincenzo Scotti,Raffaela Mirandola
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Attack Path (AP) modeling is fundamental to cybersecurity analysis, where the Planning Domain Definition Language (PDDL) has been widely adopted to encode APs into formal and machine-verifiable representations for automated reasoning about vulnerability exploitation, attack progression, and their potential impacts. However, existing AP modeling approaches largely rely on expert-driven manual construction, limiting their scalability and ability to keep pace with rapidly evolving cyber threats. Large language models (LLMs) are promising candidates, as their extensive pre-trained knowledge and reasoning capabilities enable them to interpret and transform threat intelligence into formal representations. In this paper, we propose \textbfCVE2AP, an LLM-based approach for automatically generating PDDL-encoded attack paths from natural language CVE (Common Vulnerability Exposure) descriptions. CVE2AP leverages structured prompting and incorporates an error-feedback mechanism that iteratively refines the generated paths using planner-reported syntactic and solvability errors. We conduct a systematic empirical evaluation across multiple LLMs and generation configurations, assessing generation quality across syntactic, solvability and semantic dimensions, together with token consumption and generation time. The results demonstrate that CVE2AP effectively generates high-quality PDDL-encoded attack paths, achieving up to 86.9% syntax correctness, 78.6% solvability, and 93.1% semantic correctness under LLM-as-expert evaluation, while \textttGPT-5.5 offers the best quality-cost trade-off and error feedback yields the most consistent quality improvement.

[AI-29] Geometry Meets Physics: Data-Efficient Pre-Training for Unstructured Neural PDE Solvers NEURIPS2026

链接: https://arxiv.org/abs/2610.03363
作者: Luis Medrano-Navarro,Giacomo Baldan,Qiang Liu,Benjamin Holzschuh,Jan Hagnberger,Mathias Niepert,Nils Thuerey
类目: Artificial Intelligence (cs.AI)
备注: Accepted to NeurIPS 2026

点击查看摘要

Abstract:Neural surrogate models for Partial Differential Equations (PDEs) on unstructured 3D geometries are often limited by poor generalization and the high cost of generating large-scale training datasets. Consequently, pre-training on massive datasets of related PDE dynamics has emerged as a critical alternative to enhance the robustness and scalability of these models. However, this strategy is neither compute- nor data-efficient, as it relies on massive pre-computed data that is very costly to generate. In this work, we introduce a disk-data-free pre-training framework tailored to both steady-state and transient regimes. For steady-state problems, we propose a geometry-driven strategy that leverages intrinsic shape descriptors to learn representations of complex 3D domains. For transient problems, we introduce a physics-driven approach based on online generation of synthetic PDE data, enabling scalable pre-training without reliance on expensive datasets. Across multiple experiments, our approach achieves faster convergence, greater data efficiency, and higher accuracy during fine-tuning, particularly under realistic low-data regimes. This methodology provides a practical pathway toward data-efficient neural emulators for large-scale simulations.

[AI-30] Follow the Winners: Conservative Policy Improvement with the Cross-Entropy Method for Critic-Free RFT NEURIPS2026

链接: https://arxiv.org/abs/2610.03361
作者: Joery Ariën de Vries,Neil David Lawrence,Zhenwen Dai
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: Poster at NeurIPS 2026

点击查看摘要

Abstract:Critic-free reinforcement fine-tuning (RFT) for agentic large language models is often done through GRPO-style methods, which compute a group baseline over repeated rollouts to reduce target variance. However, this setup is ill-suited to agents acting in stateful environments such as live services or security sandboxes, where repeated rollouts are impractical to obtain and aggressive updates entrench the noise of long, sparsely verified trajectories. We propose \textitFollow the Winners (FTW), a critic-free policy-learning algorithm that adapts the cross-entropy method to RFT, replacing group rollouts with an ordinal filter on replay-buffer samples that yields polynomial concentration in the order statistic of returns. We derive FTW through a control-as-inference lens, which also recovers GRPO and DPO as specific modelling choices, identifying GRPO as risk-neutral while DPO and FTW share a bounded risk-seeking offset that FTW controls. We identify this offset as an inherent trade-off of variance reduction through ordinal filters on samples, whereas a critic model induces a different trade-off between bias and variance. Scaled to agentic LLM post-training, FTW matches GRPO and PPO on Sokoban and Search-R1 baselines, showing a viable trade-off from a value model or group rollouts to CPU memory.

[AI-31] ReFract: Benchmarking Perspective Awareness in Language Model Agents with Text World Models

链接: https://arxiv.org/abs/2610.03356
作者: Hainiu Xu,Vítor N. Lourenço,Mohnish Dubey,Yunfei Bai,Yulan He,Caroline Catmur,Aline Paes,Marco Caserta,Akash Chandrayan,Luca D’Angelo
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Large Language Model (LLM) agents are increasingly deployed in high-stakes settings such as industrial maintenance and equipment fault troubleshooting, where workers occupy a variety of roles. A capable agent must therefore act in a way that is calibrated to user’s role: taking actions and providing information that respect the role’s knowledge and capability boundaries. Unlike coding, where mistakes are usually recoverable, agent responses in these settings are enacted on physical equipment, and can therefore cause irreversible equipment damage, production loss, or personnel harm. Existing benchmarks, however, largely overlook the need for agents to infer what a role intends and acting only through tools that role may legitimately use, a capability which we term Perspective Awareness. To this end, we introduce ReFract, a benchmark of 150 expert-validated entries in which an agent must act differently in response to the same query depending on user’s role. Entries of ReFract are grounded in anonymized queries from domain support conversations, against which we construct Text World Models that simulate the agent’s operating environments and assemble perspective-aware action trajectories. State-of-the-art LLMs solve at most 69% of the tasks with more than 50% of their trajectories contain attempts of taking perspective-violating actions. ReFract exposes perspective awareness as a distinct, largely unsolved axis of agent evaluation and motivates agents that calibrate not just how to act, but for whom.

[AI-32] Equivariant Visual-Tactile Diffusion Policy for Contact-Rich Manipulation

链接: https://arxiv.org/abs/2610.03333
作者: Lik Hang Kenny Wong,Yiyao Ma,Xiu-Shen Wei,Zelong Tan,Zhuheng Song,Dongsheng Xie,Kai Chen,Qi Dou
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI)
备注: 21 pages, 6 figures. Accepted to the 10th Conference on Robot Learning (CoRL 2026)

点击查看摘要

Abstract:Imitation learning for contact-rich manipulation requires high-quality expert data that is expensive to obtain. This makes learning a sample-efficient policy a key issue. To address this, we propose VISTA, a workspace-level equivariant visuotactile diffusion policy for data-efficient contact-rich imitation learning. VISTA projects visual and tactile observations into spherical tokens, injects tactile contact cues into visual spherical directions through permutation-equivariant spherical fusion, and rotates the fused harmonic representation using the end-effector orientation. The resulting representation conditions an equivariant diffusion policy to predict spatially consistent actions. Extensive experiments in both simulation and real-world robotic settings show that VISTA substantially improves data efficiency over strong visuotactile imitation learning baselines. Project website: this https URL

[AI-33] Cordial Learning: Distributed Training with Correlated Data

链接: https://arxiv.org/abs/2610.03330
作者: Sarah Shitrit,Ilai Bistritz
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:We consider a distributed learning task with agents that have correlated data. Specifically, the label of an agent depends on the input of other agents for the same sample, and these inputs are also correlated. Correlated data is the reality when agents share the same environment. Existing decentralized methods, such as federated learning, ignore the structure of the problem and perform poorly on correlated data. On the other hand, centralized approaches are infeasible due to privacy and communication constraints. We introduce cordial (correlated and distributed) learning to address this gap by sharing only low-dimensional outputs between the agents while training local models to extract informative signals from peers. This distributed learning induces a game in which the loss function of each agent depends on the models of others. Assuming a linear model, we prove that cordial learning converges with probability one to a globally optimal solution, despite the nonconvex global objective. Experiments on structured multi-digit MNIST tasks demonstrate that cordial learning remains highly effective even in highly nonlinear settings.

[AI-34] Preserving Mathematical Reasoning in Compressed Diffusion Language Models via Trajectory-Aware Low-Rank Approximation

链接: https://arxiv.org/abs/2610.03326
作者: Tian Liang,Zishan Shao,Yiran Chen
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Diffusion language model (dLLM) compression faces a known challenge because calibration is typically performed on clean, fully visible activations, whereas inference traverses partially masked intermediate states. For low-rank compression, this raises two questions. First, can low-rank optimality still be characterized when approximation quality is measured over trajectory-distributed states, and second, does the choice of calibration states affect mathematical reasoning preservation under compression? We address these questions by formulating a trajectory-aware low-rank objective over corruption levels and masking realizations. To estimate this objective efficiently, we propose Traj-MC, which estimates the trajectory second moment through Monte Carlo sampling and yields exact sampled-state optimality and population consistency. Under matched compression budgets, trajectory-aware calibration improves reconstruction over the generation trajectory and preserves substantially more mathematical reasoning than clean calibration on mathematical reasoning benchmarks. Our results connect trajectory-aware low-rank optimality to the reasoning capability retained after dLLM compression. Our code is available at: this https URL.

[AI-35] Refinement Buys Intelligibility Search Buys Identity: What Test-Time Compute Buys in Masked-Diffusion TTS

链接: https://arxiv.org/abs/2610.03320
作者: Nityanand Mathur,Hamees Sayed,Ayush Pratap Singh
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Diffusion language models for text-to-speech combine two forms of computation: model depth (parameters) and refinement steps (inference budget). We ask whether they scale equally across capabilities. We train 15 masked-diffusion codec TTS models varying depth (19-133M parameters, 3 seeds) on 2,000 hours of speech and sweep refinement steps T in [1,16] at inference, measuring zero-shot synthesis via ASR word error rate (intelligibility) and speaker verification (identity) on 174 held-out speakers. Against measured floors, refinement closes 86.2% of the intelligibility range but only 46.4% of the identity range - a 1.86x asymmetry robust across multiple error metrics. Retraining at 3x and 6x schedule attenuates but does not reverse this gap (1.84 to 1.36 to 1.23x), because intelligibility saturates with steps while identity continues improving. Best-of-K search recovers speaker identity where refinement fails, with 64.6-79.0% win rates across four independent encoders. Depth and steps are not interchangeable: separable B(d)B(T) fits significantly better (Delta AICc=+69.3) than substitution models. Analysis shows 62% of remaining identity deficit lies in the codec, not the generator. We conclude that refinement and depth target different bottlenecks and should be optimized separately.

[AI-36] Multi-Task Evolution for Zero-Shot Cross-Problem Generalization using LLM s

链接: https://arxiv.org/abs/2610.03316
作者: Zhouliang Xie,Changliang Zhou,Genghui Li,Zhenkun Wang
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Designing effective heuristics for diverse combinatorial optimization problems requires substantial expertise and repeated search. Large language models (LLMs) automate heuristic generation and refinement, but heuristic search typically depends on evaluation feedback from the problem being optimized. Generalizing to new problem definitions using only source-task feedback therefore remains a central challenge. We introduce MECo, an LLM-driven multi-task evolutionary framework for zero-shot cross-problem generalization. MECo maintains task-conditioned heuristic populations and uses a transfer gap based on cross-task population performance to guide their interactions. These interactions enable the transfer and recombination of heuristics. A complementary selection criterion then constructs a compact heuristic set by rewarding each member’s additional coverage of source combinations. The selected set is applied to target problems without further search or adaptation. Experiments on 32 problem variants across vehicle routing (VRP) and flexible job-shop scheduling (FJSP) show that MECo achieves the lowest mean costs compared with eight automated heuristic design (AHD) baselines under the same budgets. On out-of-domain problems, it outperforms the strongest baseline in each family. Moreover, integrating the framework of MECo with different AHD methods improves their ID and OOD performance in both families, supporting its effectiveness across different methods.

[AI-37] Lightweight Rubric-Guided Trajectory Evaluation for Production AI Agents

链接: https://arxiv.org/abs/2610.03315
作者: Linh-An Phan,MingXue Wang,Guangyu Wu,Feng Pan,Zhaoyu Pang,Yanbin Zhang
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Trajectory evaluation is essential for improving the reliability of LLM-based agents, but production use makes it expensive to run repeatedly. Modern agents generate long traces containing tool calls, observations, retries, and external outputs, while not all raw tokens are equally useful for diagnosis. We present \textitLiteTrajEval, a lightweight architecture for budget-bounded trajectory evaluation. LiteTrajEval derives compact domain-specific rule profiles offline, then preprocesses each trajectory online, marks heuristic failure signals, serializes it under a fixed global budget, and invokes a single rubric-guided LLM judge to produce structured diagnostic reports. Evaluated on public Magentic-One-style and \tau -bench-style trajectory datasets, LiteTrajEval improves failure-localization alignment with human annotations by roughly 20–35 percentage points on Magentic-One and up to 23 percentage points on \tau -retail compared with AgentRx, while reducing cost by about 6 \times and evaluation time by more than 8 \times . This solution has also been deployed in our enterprise agentic platform.

[AI-38] Optimal Planning in a Dynamic World

链接: https://arxiv.org/abs/2610.03312
作者: Devin Wild Thomas(1),Solomon Eyal Shimony(2),Wheeler Ruml(1),Erez Karpas(3),Shahaf S. Shperberg(2),Andrew Coles(4) ((1) University of New Hampshire, USA, (2) Ben-Gurion University of the Negev, Israel, (3) Technion - Israel Institute of Technology, Israel, (4) King’s College London, UK)
类目: Artificial Intelligence (cs.AI)
备注: 48 pages, 25 figures

点击查看摘要

Abstract:Background: We address the problem of planning when the set of feasible states or actions changes over time. For example, in the problem of path planning among moving obstacles (sometimes known as SIPP), the feasibility of being at a particular location can change as the obstacles move. Or, the action of boarding a particular train is feasible only while it is stopped at the station. This dynamism means that the optimal plan and its duration can change depending on when execution begins. In practice, execution start time is often unknown until planning has completed or another agent gives the go-ahead. However, most prior planning work either ignores dynamism or assumes a known start time. This makes it straightforward to assess state and action feasibility but is impractical for some applications. Objectives: In this paper, we relax the assumption of a known start time. We define the setting of \em any-start-time planning and provide algorithms for it. Methods: We present a data structure called a compound arrival time function (cATF) that compactly encodes the optimal plan as a function of start time. We provide general-purpose planning algorithms, based on heuristic graph search, that assemble cATFs by propagating functions along edges instead of scalar costs. Results: We prove that the size of a cATF is at most linear in the problem size. An experimental evaluation of an implementation for the specific problem of SIPP shows that, on difficult problems, agents that rely on replanning often fail, while any-start-time algorithms using cATFs can quickly look up the optimal plan once the execution start time is known. Conclusions: By enabling efficient representations and reasoning for time-dependent plans, this work provides a foundation for planning in dynamic worlds.

[AI-39] raining-Loss Guarantees for Muon with Finite-Step Newton–Schulz Orthogonalization

链接: https://arxiv.org/abs/2610.03306
作者: Amartya Roy,Souvik Chakraborty
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Optimization and Control (math.OC)
备注: 22 pages, 5 figures

点击查看摘要

Abstract:Existing convergence analyses of Muon either assume exact orthogonalization or analyze classical Newton–Schulz polynomials, and guarantee only stationarity, so it is unresolved what Muon’s five tuned Newton–Schulz steps preserve and whether that suffices to reach a prescribed neural-network training loss. We establish a finite-time training guarantee that accounts for both momentum accumulation before orthogonalization and the tuned finite-step update. For full-batch training of a sufficiently wide two-layer ReLU network with fixed random output weights and a positive-definite limiting neural tangent kernel, we prove that Muon reaches any target empirical squared loss \varepsilon0 with high probability over initialization. For every momentum parameter \mu\in[0,1) , a target-dependent constant learning rate proportional to (1-\mu)\sqrt\varepsilon yields a hitting-time bound of O((1-\mu)^-1\varepsilon^-1/2) , with other problem parameters fixed. The sufficient width is independent of both target accuracy and momentum. The analysis shows that the tuned Newton–Schulz map preserves alignment with the momentum buffer while bounding the update’s spectral norm. Control of gradient variation near initialization transfers this alignment to the current gradient, ensuring descent until the target is reached without requiring exact orthogonalization. Numerical experiments support these mechanisms at widths below the sufficient theoretical threshold: gradient-update alignment remains above the analytical reference, and all 30 runs across six widths and five student initializations on a fixed teacher-student dataset reach the target loss while maintaining kernel positivity.

[AI-40] JOVE: Joint Execution and Verification for Resource-Aware LLM Task Graphs

链接: https://arxiv.org/abs/2610.03296
作者: Haoran Zhang,Dongjun Kim,Seohyeon Cha,Kevin S Chan,Ananthram Swami,Gustavo De Veciana,Haris Vikalo
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: preprint

点击查看摘要

Abstract:Complex reasoning queries can be decomposed into directed acyclic task graphs and distributed across heterogeneous LLMs, reducing latency through parallelism and enabling smaller models to solve complex tasks. In practice, however, the suitability of an LLM for a given subtask may be a priori unknown, and execution alone does not reveal output correctness. We propose JOVE, an online framework that jointly assigns executor LLMs and selects intermediate outputs for paid verification. Verification runs asynchronously and is used to improve future allocations, so the system must balance spending on execution now against learning for later. We study how to optimize this trade-off under a long-term budget and a per-query latency constraint, with stochastic, initially unknown LLM service quality, invocation costs, and execution times. JOVE makes execution and verification decisions by solving a sequence of per-query mixed-integer linear programs. Online learning updates task-dependent estimates of LLM quality based on verification feedback, while an information-gain bonus incorporates the value of learning into allocation decisions. Under a natural set of assumptions, we establish sublinear quality-learning regret for JOVE. Across four reasoning benchmarks, JOVE achieves competitive accuracy against standard inference baselines while reducing average cost and latency by at least 3.17 times.

[AI-41] EVOL: Simulator-Guided Evolutionary Expert Synthesis for Deployment-Free Learning Path Recommendation CIKM2026

链接: https://arxiv.org/abs/2610.03273
作者: Geonwoo Bang,Dongho Kim,Moohong Min
类目: Artificial Intelligence (cs.AI)
备注: Accepted at the 35th ACM International Conference on Information and Knowledge Management (CIKM 2026)

点击查看摘要

Abstract:Reinforcement learning (RL) for learning path recommendation (LPR) faces two coupled obstacles. First, the policy must commit to a sequence of L concepts without intermediate feedback, producing a combinatorial search space that grows super-exponentially with L and provides reward only at the final step. Second, expert learning paths would be the natural cure for sparse-reward RL, but they do not exist in educational data, because student logs record what learners did, not what they should have done. We address both obstacles by importing a recipe from simulator-based demonstration learning in robotics: the knowledge tracing simulator is used both to synthesize per-learner expert demonstrations through evolutionary search and to train a deployment-free policy that distills these demonstrations into a feed-forward learner. Our framework, EVOL, instantiates this pipeline with an asymmetric actor-critic where the actor commits to deployment-realistic blind planning while the critic exploits the privileged simulator state during training. Across three datasets (ASSIST15, Junyi, and EdNet; 39-189 concepts) and path lengths L = 5, 10, and 20, EVOL surpasses 8 baselines spanning heuristic, sequential, RL, graph-enhanced RL, and LLM-enhanced methods. We further compare three imitation strategies (BC, AWR, and DAPG) and show that final performance is governed by the quality of evolutionary experts rather than by the particular imitation objective.

[AI-42] SPEAR: A Spectral-Disentangled MoE Neural Operator with Knowledge-Guided Expert Aggregation for Large-Scale PDE Pretraining

链接: https://arxiv.org/abs/2610.03265
作者: Dengdi Sun,Xiaoya Zhou,Xiao Wang,Wanli Lyu,Jin Tang,Bin Luo
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Large-scale pre-training has improved the generalization of neural operators across diverse PDEs. However, existing PDE foundation models still struggle with heterogeneous dynamics, where shared representations may cause knowledge interference, while mixture-of-experts (MoE) architectures suffer from increasing expert redundancy. We propose SPEAR, a spectral-disentangled MoE neural operator with knowledge-guided expert aggregation for large-scale PDE pre-training. SPEAR decouples latent features into low- and high-frequency components, enabling shared modeling of transferable dynamics and specialized learning of PDE-specific patterns. To address expert redundancy, we design a knowledge-guided expert aggregation strategy that measures expert similarity from dataset-specific learned knowledge and routing preferences, enabling the identification and consolidation of similar experts. Experiments on twelve PDE datasets and multiple downstream benchmarks demonstrate superior performance in pre-training, fine-tuning, and transfer learning. Furthermore, our aggregation strategy reduces the number of experts by 50% while maintaining or improving prediction accuracy, achieving a balance between model efficiency and generalization for PDE foundation models.

[AI-43] Mapping and Advancing the Scalability-Accuracy Frontier of Nonlinear Causal Discovery

链接: https://arxiv.org/abs/2610.03258
作者: Hendrik Suhr,Sascha Xu,Jilles Vreeken
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Scalable nonlinear causal discovery requires methods that combine flexible mechanism estimators with efficient search over large graph spaces. Several algorithmic families have been proposed to address this challenge, yet their accuracy-runtime trade-offs remain poorly understood. We empirically compare the four major approaches: differentiable structure learning, amortized structure learning, score-matching, and combinatorial search. Our results reveal complementary bottlenecks: differentiable and amortized methods scale well but exhibit an accuracy gap, score-matching methods can be accurate in low dimensions but degrade quickly for increasing feature sizes, and combinatorial methods remain accurate but are slowed by repeated and redundant local scoring. Motivated by this bottleneck, we develop SPADE, a spline-based score-evaluation scheme that compiles sufficient statistics once and reuses them throughout combinatorial search. Under bounded indegree, its Gaussian variant reduces algorithmic complexity from O(nd^3) to O(nd^2+d^3). Empirically, SPADE shifts the observed scalability-accuracy frontier by orders of magnitude: it solves 100-variable problems with 160K samples in seconds and 1600-variable problems with 2.5K samples in minutes, while retaining high structural accuracy across synthetic and real-world benchmarks. These results reveal a substantial shift in the practical scale of combinatorial search and highlight the importance of evaluating scalable causal-discovery methods along the full accuracy-runtime frontier.

[AI-44] Learning a Fact Is Not Learning How to Retrieve It

链接: https://arxiv.org/abs/2610.03251
作者: Chaemin Jang,Jihee Kim,Dongman Lee
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:A model trained on “The capital of X is Y” may produce “Y” after “The capital of X is” but fail after “The capital of X:”. We call these different ways of eliciting the same fact request forms. To separate learning a fact from retrieving it, we train two models in two stages. In the first stage (request-form training), one model sees each fact in five forms and the other sees the same facts only as statements. In the second stage (target-fact training), both receive identical training on new facts, all as statements. Both then retrieve the new facts almost equally well from statements, but differ sharply on other request forms. Thus, a model can learn how to retrieve through a request form before it learns the facts. To understand this difference, we examine the hidden state immediately before the answer, which we call the context state. When given two different request forms for the same fact, the model trained on five forms in stage one produces more similar context states than the model trained on statements alone in that stage. Changing this state at retrieval time can enable or prevent retrieval of an already learned fact, and the same effect transfers across facts and factual relations, such as capitals and currencies. To test its role during learning, we change the context state only during target-fact training. This intervention changes later retrieval without intervention at test time. Together, these results show that later retrieval depends on earlier request-form experience and the context state during fact learning.

[AI-45] WAMpy: Efficient Synthesis of Prolog Programs in Python

链接: https://arxiv.org/abs/2610.03234
作者: Dominik Magiera,Lukas Röhrig,Frank Jäkel
类目: Programming Languages (cs.PL); Artificial Intelligence (cs.AI)
备注: 4 pages, 2 figures. Accepted as a demo at the 6th International Joint Conference on Learning and Reasoning (IJCLR 2026). Code: this https URL

点击查看摘要

Abstract:We present WAMpy, a Python framework optimized for synthesizing Prolog programs. Unlike general-purpose Prolog systems, WAMpy targets workloads that repeatedly generate and evaluate small candidate programs. WAMpy compiles Prolog clauses into NumPy array-based WAM instructions and supports partial recompilation of hypotheses against fixed background knowledge. Performance-critical routines are accelerated using Numba just-in-time (JIT) compilation. In a benchmark of repeated compilation-and-evaluation workloads, WAMpy improves end-to-end performance compared with SWI-Prolog accessed from Python using Janus.

[AI-46] D2K-Bench: Can LLM Agents Turn Expert Designs into Efficient GPU Kernels?

链接: https://arxiv.org/abs/2610.03226
作者: Daifeng Li,Huiqiang Jiang,Chengruidong Zhang,Wei Wu,Xudong Guo,Jianhong Tu,Jianwei Zhang,Binhang Yuan,Dayiheng Liu
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Distributed, Parallel, and Cluster Computing (cs.DC)
备注: 30 pages, 4 figures

点击查看摘要

Abstract:GPU kernels generated by large language model (LLM) agents can remain less efficient than expert implementations, but runtime alone does not reveal how the gap relates to design discovery and implementation. We introduce D2K-Bench, a diagnostic benchmark of 26 tasks and 85 workloads that measures how effectively agents translate expert design guidance into efficient GPU kernels. The guidance covers L1: high-level algorithmic insights, L2: dataflow design, and L3: low-level optimization tricks, including dependencies among these levels. Pairwise runs with and without guidance share task descriptions, workloads, tools, hardware, and a 350-turn budget. Complementary assessments examine independently proposed designs and the design properties implemented in generated code. Across five models on NVIDIA B200 GPUs, guidance raises correctness over 130 model-task pairs from 93.1% to 98.5% and increases the Performance Score over all 26 tasks from 1.46 to 1.95. For the three frontier models with correct submissions on all 26 tasks in both runs (GPT-6-Astra, Claude-Opus-4.8, and GPT-5.6-Sol), geometric mean speedup increases from 1.69\times to 2.49\times . Across all five models, the mean combined implementation score increases from 57 to 70 out of 100. These results show the value of expert design guidance while identifying design properties that remain unimplemented.

[AI-47] oward SLM-based agent ic task-tool intent matching

链接: https://arxiv.org/abs/2610.03213
作者: Chiara Troiani,Arash Salarian,Majed El Helou,Benjamin Ryder,Jean Diaconu,Hervé Muyal,Marcelo Yannuzzi
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Tool-equipped AI agents use tool calls to access data and act on external systems. Horizontal growth of agentic systems increases the number of these interactions, and further motivates the need for automated, per-call oversight that can operate at low latency and/or on-prem. Conventional authorization schemes can determine whether an agent is allowed to invoke a tool, but cannot assess the agent’s underlying cognition, specifically, whether the tool selection represents a logical, relevant step toward satisfying the intent of the task or not. Consequently, an allowed call may still deviate from the task’s intent: a rogue agent might deviate the calls or nudge other agents to make a combination of calls that would not align with the intent of the task. Therefore, every call needs to be verified. In this study we investigate the applicability of Small Language Models (SLMs) to this purpose: an SLM functions as a task-tool relevance classifier that evaluates every selected tool independently against the assigned task and returns a relevance signal for downstream enforcement. Equipped with a novel dataset with multi-tool tasks whose required tools span distinct Model Context Protocol (MCP) servers, we used prompt-optimization, supervised fine-tuning, and reinforcement learning through GRPO to optimize and specialize SLMs.

[AI-48] LiBRA: Detection-Aware Image Watermark Removal via Bidirectional Latent Optimization

链接: https://arxiv.org/abs/2610.03166
作者: Saibo Ye,Huajie Chen,Xin Guo,Le Yang,Chi Liu,Xiangyu Hu,Jingjing Guo,Tianqing Zhu
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Digital watermarking supports source attribution for AI-generated images, but its reliability depends on resistance to removal attacks. Some attacks attempt to remove watermarks by forcing the decoded watermark to differ from the original. However, this can produce an inverted watermark that remains detectable, causing removal to fail, while further attempts to alter the watermark may unnecessarily degrade image quality. To address these limitations, we present LiBRA (Latent In-band Bidirectional Removal Attack), which aims to make watermarks undetectable while preserving image quality. Instead of continually pushing the watermark toward inversion, LiBRA adjusts the image to conceal the watermark without encouraging further changes that could degrade image quality. Some attacks keep pushing decoded bits away from the original watermark, even when further changes preserve detectability and damage image quality. With access to the watermark key and decoder, LiBRA makes bounded changes in a public autoencoder’s latent space. Unlike inversion-driven objectives that cannot correct excessive inversion, LiBRA guides average decoding confidence toward random guessing from either direction. This helps avoid an inverted but detectable watermark. Leaving individual bits flexible allows image-quality constraints to favor less damaging changes, while an optional frequency-guided mask limits their location. We verify removal using an exact two-sided binomial test rather than assuming the confidence target guarantees success.

[AI-49] EvoRiskBench: An Evolving Benchmark for Runtime Security Risks in Workspace Agents

链接: https://arxiv.org/abs/2610.03153
作者: Shiyi Kuang,Xuemei Luo,Kun Liu,Junhai Li,Rui Tian,Feng Shi,Bo Shen,Nianyu Li,Dehui Li,Ping Chen
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Workspace agents combine large language models with execution harnesses to perform stateful, multi-step tasks that access or modify external resources. Existing benchmarks leave gaps in executable coverage of their runtime security risks, while evolving model capabilities, harnesses, tools, and threats motivate benchmark evolution. We introduce EvoRiskBench, an evolving benchmark organized around the EP-Path-EF framework, which links an initial risk entry point to a one-hop technical effect through an agent-mediated risk path. The framework defines nine entry-point categories and five effect categories; a 20-participant study supports their interpretability and classification consistency on representative cases. Guided by this framework, an automated end-to-end workflow constructs and executes risk cases in isolated environments and independently verifies outcomes using runtime traces and environment states. The benchmark provides a reproducible dataset of 450 adversarial tasks across six scenarios. We evaluate nine model-harness configurations spanning three models (GPT-5.6 Sol, DeepSeek-V4-Pro-0813, and Claude Opus 5) and three harnesses (Claude Code, Codex, and OpenClaw). Our results reveal substantial vulnerabilities across systems. The most vulnerable configuration, Codex with DeepSeek-V4-Pro-0813, reaches a 68.44% attack success rate (ASR), indicating that configuration of workspace agent is insufficient to ensure secure autonomous execution. ASR varies more across models than harnesses, and harness differences depend on the model. The benchmark cases and evaluation platform will be released after completion of artifact safety and reproducibility checks.

[AI-50] Keeping JEPA World Models Plannable When Little of the Frame Moves

链接: https://arxiv.org/abs/2610.03137
作者: Florian Strohm,Patrick Wagner,Jannik Schwab,Marco Huber
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Specifying a goal in language rather than as a goal frame is a natural interface for planning with a latent world model, but testing it needs scenes in which language must discriminate between several objects. We build SLIM, a pushing benchmark with several small objects and paired visual and language goals on identical scenes. On SLIM a LeWM world model that solves PushT succeeds on under 1% of trials, although a scripted controller with simulator state solves every tier. Probes locate the failure in the encoder: its latent is nearly action-insensitive, neither pusher nor object positions can be decoded from it, and rollouts are no better than copying the current latent forward. One inverse-dynamics auxiliary loss, applied to encoder latents and to predicted latents through a shared head discarded at test time, restores every probe and raises success from 0.003 to 0.35 (0.16 on the hard pushing tier, where a goal-agnostic policy scores zero), and improves PushT at twice the trained horizon. Controls attribute the repair to the gradient into the encoder, and a response sweep shows that the vanilla model plans once enough of the frame responds to actions. A cheap action-sensitivity probe, computable without environment access, acts as an empirical necessary condition: all configurations below its threshold failed to plan. On the repaired latent, a small language-goal head plans from sentences without retraining the world model: it reaches 0.84 on navigation (visual-goal oracle 1.00), follows the named zone when it is swapped with a decoy, and degrades gracefully to unseen nouns. A single goal sentence rarely completes a push, but given the push as a sequence of stage sentences the head raises success on the medium and hard pushing tiers from 0.04 to 0.25, on par with the goal-frame oracle, also when the switch between stages is read from the latent alone.

[AI-51] rading Strategy Optimization via Textual Gradient

链接: https://arxiv.org/abs/2610.03128
作者: Chaoqun Yang,Qian Wang,Fengbin Zhu,Xinyu Lin,Bingsheng He,Roger Zimmermann,Tat-Seng Chua
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Quantitative trading strategy design aims to discover trading programs from historical data that remain effective in future markets, which can be viewed as a black-box program optimization problem. LLM-based textual gradients offer a promising approach by providing explicit optimization directions for iterative strategy refinement. However, directly applying textual gradients faces two challenges: (1) optimization is myopic, underutilizing experience from previous evaluations; and (2) aggregate backtest feedback overlooks temporal robustness, potentially favoring strategies that perform well only in specific market periods. To address these challenges, we propose TradeGrad, an experience-guided textual-gradient framework for robust trading strategy optimization. TradeGrad leverages accumulated optimization experience to estimate textual gradients and employs multi-scale revisions for both strategy exploration and refinement. It further introduces the Cross-Period Robust Objective (CPRO), which emphasizes performance in unfavorable historical periods to promote temporal robustness. Experiments on cross-sectional and time-series strategy design in Chinese A-share and U.S. equity markets show that TradeGrad achieves the best in-sample and out-of-sample performance across all four settings. Notably, its Chinese cross-sectional strategy achieves 27.99% annualized return, 12.19% maximum drawdown, and a Sharpe ratio of 1.63, approximately 68% higher than the CSI 300 benchmark. Further analyses validate the proposed components and show consistent improvements in both in-sample and out-of-sample performance throughout optimization. The code is available at this https URL.

[AI-52] How to Find and Reuse Policies for Continuous Adaptation in Lifelong Reinforcement Learning

链接: https://arxiv.org/abs/2610.03119
作者: Saptarshi Nath,Inish M. D’Souza,Antonio Carta,Soheil Kolouri,Andrea Soltoggio
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Robotics (cs.RO)
备注: Code is available at this https URL

点击查看摘要

Abstract:In lifelong reinforcement learning, retaining previously learned policies is not sufficient for effective transfer to a new task. Useful knowledge may be distributed across several prior policies, and its relevance may change as the learner acquires experience. One hypothesis is that task similarity can be effectively used in a continual learning setting to find and combine previously learned policies. To test it, Adaptive Mask Selection and Composition (AMSC) is designed to estimate similarity from online experience via non-parametric Wasserstein task embeddings from state-action-reward samples. The z-score-normalized sparsemax of the similarity scores are used to derive a variable-size support to periodically choose and weight policies to form a prior when learning a new task. On CT-graph and MiniGrid, AMSC achieves higher mean performance and forward transfer than the evaluated modular composition baselines while exhibiting no forgetting. Results on Continual World suggest that identifying relevant prior knowledge and determining its layer-specific composition may require additional layer-specific tuning. Ablations show that selecting relevant sources and determining how strongly to reuse them are central to these gains. Independently measured pairwise transfer is also positively associated with task-embedding similarity. These results indicate that task similarity can be an effective criterion to select and weight specific knowledge for reuse in lifelong reinforcement learning.

[AI-53] Predictor-Guided Latent Space Codon Optimization for Maximizing Protein Expression

链接: https://arxiv.org/abs/2610.03098
作者: Alberto Caron,Tianyu Cui,Dmytro S. Lituiev,Mangal Prakash,Artem Moskalev,Amina Mollaysa,Bo Zhai,Hirsh Nanda,Daniel M. Poole,Zhongyin Liu,Iman Farasat,Robert Davidson,Nikolay V. Manyakov,Tommaso Mansi,Scott Oloff,Rui Liao
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Codon optimization, the process of selecting synonymous codons to improve mRNA translation efficiency and protein expression, is central to therapeutic protein production and mRNA vaccines, yet it remains a hard problem. The design space is discrete and combinatorially large, precluding gradient-based methods, and existing tools rely on heuristic proxies (e.g., Codon Adaptation Index or GC-content) that poorly capture true expression. We introduce Latent-Space Codon Optimization (LSCO), which recasts this discrete problem as a continuous one by mapping sequences into the latent space of a pretrained mRNA language model, enabling efficient gradient-based search. LSCO combines four components: a data-driven expression objective from an uncertainty-aware predictor, a Minimum-Free-Energy regularizer for structural stability, a naturalness prior from a protein-to-codon back-translation model, and constrained decoding for protein fidelity. On a real-world, wet-lab antibody expression dataset, LSCO outperforms simple frequency-based, as well as modern deep generative baselines in predicted expression, while retaining suitable biophysical properties.

[AI-54] ULTRADISCOVERY: Abductive Exploration in an Interconnected Epistemically Open Universe

链接: https://arxiv.org/abs/2610.03092
作者: Weihan Li,Tianshi Zheng,Yangqiu Song,Ginny Y. Wong,Simon See
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: 47 pages, 19 figures, 15 tables

点击查看摘要

Abstract:Scientific discovery often begins when scattered clues call for a new way of describing the world. Such abductive exploration can require constructing the representation in which an explanation is stated, when the world is epistemically open, and composing evidence scattered across contexts, when it is structurally interconnected. Existing benchmarks rarely separate these two demands or control them independently. We introduce ULTRADISCOVERY, an interactive world of five domains in which an agent revises an initially successful theory and predicts the outcome of an unseen cross-domain intervention. A 2 \times 2 design leaves the representation open or discloses it, and leaves the evidence distributed or aligns it, with the latent dynamics fixed. With the representation open, agents across eleven models often retract the axiom they were taught, and none introduces the unobserved entity or rewrites the variables that a replacement requires. Disclosure triples intervention requests and adds about one of the eighteen findings the world affords, and alignment adds less. Two vendor-harness systems carry discovery into more domains, and one of them rewrites the variables in Open episodes. No system makes the exact prediction within 200 paid actions. At larger budgets one exact prediction appears with both aids, while every Open episode remains inexact. The results locate the difficulty in the step from accumulating evidence to composing it into a representation that transfers.

[AI-55] Securing Computer-Use Agents Against Branch Steering Attacks NEURIPS2026

链接: https://arxiv.org/abs/2610.03089
作者: Giulio Zingrillo,Hanna Foerster,Ilia Shumailov,Yiren Zhao,Robert Mullins
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
备注: 16 pages, including 2 figures. To be presented at the “Agents in the Wild” Workshop at the NeurIPS 2026 Conference

点击查看摘要

Abstract:Modern Computer Use Agents (CUAs) directly interact with graphical user interfaces and execute third-party web tools, exposing them to indirect prompt injection across every rendered page and tool response. While the Dual-LLM pattern is the primary system-level architecture offering formal security guarantees - using an isolated Planner LLM (P-LLM) to fix execution paths before processing untrusted inputs via a Quarantined LLM (Q-LLM) - these guarantees break down in graphical environments. Because CUA interaction is inherently dynamic, plans cannot remain data-independent; they must branch based on anticipated runtime web content - covering all possible cases the agent may encounter. This exposes agents to branch steering attacks, where an adversary crafts untrusted data to coerce a CUA down a hazardous, pre-approved branch without injecting explicit instructions. We systematically study branch steering attacks and introduce STEER-Bench (101 tasks across 9 domains), showing high attack success against both standard (94.4%) and vanilla Dual-LLM (89.5%) CUAs. We then propose COBRA, an architecture that pairs trusted branching plans with ahead-of-time capability constraints, strictly bounding the parameters and destinations each branch may execute. On STEER-Bench, COBRA reduces attack success to 0% while retaining 97% benign utility.

[AI-56] Zephon: Elastic Determinism for Online Stateful Foundation Model Data Loading Pipelines VLDB’27

链接: https://arxiv.org/abs/2610.03087
作者: Maximilian Böther,Josh Wills,Ties Robroek,Sonnet Xu,Paul Burstein,Daniel Zayas,Cody Blakeney,Siddharth Joshi,Haoli Yin,Rishabh Adiga,Haakon Mongstad,Luke Merrick,Pratyush Maini,Ari Morcos,Matthew Leavitt,Ana Klimovic,Bogdan Gaza
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Databases (cs.DB)
备注: preprint; currently under revision at VLDB’27

点击查看摘要

Abstract:Deterministic data loading is important for foundation model development: model researchers need confidence that differences they observe across costly ablations are caused by the parameter they changed rather than non-determinism in the training data sequence. The data loader must provide elastic determinism, i.e., a deterministic sequence of global training data batches despite changes to the GPU topology across runs (e.g., due to GPU scarcity), frequent checkpoint-resume cycles, and different data processing execution backends. Achieving this is difficult because modern foundation model data pipelines tokenize, pack, and mix samples online, introducing stateful n-to-m transformations that break sample indexing. Existing data loaders largely assume indexable 1-to-1 pipelines, and the common workaround of offline materialization is expensive and, for some modalities such as video, infeasible. We present Zephon, a data loader for foundation models that supports online, stateful pipelines while providing elastic determinism and efficient resumption from checkpoints. It partitions the global stream into topology-independent lanes, serializes ordering decisions while parallelizing stateless work on interchangeable backends, and checkpoints only bounded in-flight state so recovery cost does not grow with training progress. We evaluate Zephon on text and vision-language workloads and show that it achieves competitive throughput while providing a combination of guarantees that no existing loader offers for online, stateful pipelines. Comments: preprint; currently under revision at VLDB’27 Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Databases (cs.DB) Cite as: arXiv:2610.03087 [cs.LG] (or arXiv:2610.03087v1 [cs.LG] for this version) https://doi.org/10.48550/arXiv.2610.03087 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[AI-57] RIFAR: Reliability and Forgetting-Aware Replay for Continual Robot Learning

链接: https://arxiv.org/abs/2610.03079
作者: Zirong Song,Zheng Lu,Haoran Liao,Wanqi Zhong,Yunhe Ni,Lijie Wang,Xiuying Chen
类目: Artificial Intelligence (cs.AI)
备注: 15 pages, 6 figures, 9 tables, including appendices

点击查看摘要

Abstract:Genuine embodied agency requires robots to turn continuous real-world experience into lasting, transferable skills. This demands continual learning that integrates new capabilities without eroding prior knowledge as tasks and environments evolve. Experience replay mitigates forgetting, but storing complete demonstrations becomes costly as tasks accumulate. World-action models offer a generative alternative, reconstructing past experience through joint predictions of actions and future observations. However, visually coherent rollouts may contain actions that cannot realize the predicted transitions, while new-task adaptation can disrupt previously learned behavior. RIFAR therefore combines reliability screening with drift-aware replay selection. It reconstructs trajectories from compact demonstration prefixes and uses a frozen inverse-dynamics model to assess action-visual consistency. Training first combines current demonstrations with the highest-quality screened trajectories. RIFAR then compares action predictions before and after this adaptation on identical historical inputs, reselecting trajectories with larger normalized drift from the same screened pool for continued training. Across three LIBERO suites and real-world experiments, RIFAR surpasses the previous state of the art in WAM-based generative replay. On LIBERO-Goal, it achieves 90.97 AUC while retaining only 320 historical time steps per task, approximately 4.9% of the steps retained using 50-demonstration replay.

[AI-58] MOF-VERIFY: A Failure-Aware Agent ic Harness for MOF Hypothesis Verification NEURIPS2026

链接: https://arxiv.org/abs/2610.03056
作者: Donghyun Lee,Taehoon Lee,Geonhee Ahn,Jieun Kim,Jihyun Park,Suyeon Cho,Yoona Kim,Chaerim Shin,Hoi Ri Moon,Jonggeol Na,Sukho Hong,Jihwan Oh,Soo Kyung Kim
类目: Artificial Intelligence (cs.AI); Computational Engineering, Finance, and Science (cs.CE)
备注: Accepted at the NeurIPS 2026 Workshops XAI4Science and AI4Mat

点击查看摘要

Abstract:Large language models are increasingly used as reasoning components in AI-driven materials Co-Scientists, yet the reliability of the resulting verification pipeline remains unclear. Metal-organic frameworks (MOFs) provide a particularly challenging setting because structures may appear under different identifiers, synthesis outcomes depend strongly on experimental conditions, evidence is distributed across heterogeneous sources, and some hypotheses require computation rather than literature alone. We introduce a diagnostic benchmark with four task families covering structural grounding, synthesis-condition verification, evidence-sufficiency verification, and MLIP-based computational verification. T-MOF-1-3 are evaluated under closed-book, retrieval-enabled, and oracle-evidence settings to localize failures in knowledge access, evidence acquisition, and reasoning, while T-MOF-4 separately evaluates computational verification. Guided by these diagnosed failure modes, we develop MOF-Verify, a failure-aware agentic harness that targets structural, literature, evidence-sufficiency, and computational bottlenecks before producing a final verdict. Across multiple backbone LLMs, MOF-Verify substantially improves hypothesis-verification performance over direct inference and retrieval-based baselines. Benchmark datasets are released at this https URL.

[AI-59] hacktrace: behavior-supervised detection of reward hacking during code generation

链接: https://arxiv.org/abs/2610.03055
作者: Hao Jiang,Xin Li,Annan Wang,Yichi Zhang,Weisi Lin
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:A coding agent can earn a passing grade by fixing its code, or by deleting the test that exposes the bug. Detecting such reward hacking requires recognizing attempted shortcuts, including those that fail. We release 173,561 annotated multi-turn coding trajectories from Qwen3-8B and show that supervising shortcut behavior independently of exploit success substantially improves detection. We introduce HACKTRACE, a behavior-supervised monitor that reads the internal states the agent already computes while generating code. Reusing these states enables monitoring before a turn is complete, without additional language-model tokens or passes. Combining this evidence with static features of the final files achieves a mean per-problem AUC of 0.997 with 8 ms of monitoring overhead, improving both accuracy and latency over monitors that run the model again on an honesty question and answer. The same generation states also provide an inexpensive monitoring signal for reinforcement learning. With strong GRPO penalties, HACKTRACE reduces the cheating share of passing solutions from 82-91% to 1-5%, while retaining honest, correct solutions and maintaining high detection accuracy as the policy evolves. Our results show that both the supervision target and the source of monitoring evidence matter for turning accurate detection into a useful training signal.

[AI-60] When Numbers Start Talking: Numerical Signalling and Strategic Behaviour Among LLM s

链接: https://arxiv.org/abs/2610.03033
作者: Alessio Buscemi,Daniele Proverbio,Alessandro Di Stefano, TheAnh Han,German Castignani,Pietro Liò
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Large language model (LLM)-based agents increasingly operate in multi-agent systems (MAS) characterised by strategic interaction. However, little is known about whether, and to what extent, different types of messages affect the outcomes of strategic games. By investigating AI agents based on four popular LLMs, playing four games with different cooperation equilibria, we study whether messages of different kinds (natural language, numerical signals, or random sequences) significantly modify the levels of cooperation in each game, also depending on the agents’ assigned personalities. We observe that structured messages alter the final payoffs for most games and LLMs, but without a predictable pattern; this challenges the assumption that AI agents can converge to stable equilibria regardless of additional capabilities. Moreover, we observe that agent-generated numerical messages depart from randomness, most strongly and consistently when agents are explicitly instructed to communicate; however, they introduce an additional interpretability challenge, as their symbol distributions are mostly associated with the payoff structure and typically become more concentrated with repetition, but are overall difficult for humans to interpret. Monitoring for coordination of AI agents through restricted channels should thus prioritise message-level fingerprints, which generalise across models, over behavioural decisions, which do not.

[AI-61] SoftGene: Protein Language Model-Enhanced Soft Prompting for Interpretable Gene Set Annotation EMNLP2026

链接: https://arxiv.org/abs/2610.03029
作者: Drew Ross,Arya Hadizadeh Moghaddam,Dongjie Wang,Xiaoyu Zhang,Zijun Yao
类目: Artificial Intelligence (cs.AI)
备注: Accepted to EMNLP 2026 Main Conference

点击查看摘要

Abstract:Gene set analysis is a cornerstone of functional genomics, yet it remains labor-intensive and heavily dependent on manual curation and expert biological interpretation. While Large Language Models (LLMs) have emerged as powerful tools for genomic reasoning and annotation, most existing approaches rely on symbolic gene names and fail to capture domain-specific biological structure, particularly protein sequence information that governs molecular activity, interactions, and downstream gene function. In this work, we propose SoftGene, a novel framework for LLM-based gene set annotation that leverages the hierarchical structure of gene sets. First, we use a hierarchical attention-based encoder built on ESM, a protein language model, to represent each gene set using protein-level amino acid sequence information. Second, we construct a hybrid prompting scheme that combines soft prompts derived from gene set embeddings with hard prompts containing auxiliary context generated by an LLM, and feed the resulting prompt into a local LLM for annotation. We evaluate our framework on two benchmark datasets: Gene Ontology (GO) and the Molecular Signatures Database (MSigDB). Our results show that integrating protein-sequence representations with textual context improves gene set annotation overall, while per-domain analyses reveal that the contribution of protein embeddings varies across biological domains.

[AI-62] DyadMem: A Long-Term Memory Benchmark of How Agents Work with Users

链接: https://arxiv.org/abs/2610.03020
作者: Yifei Tao,Xinyu Zhong,Henry Hengyuan Zhao,Fanyi Wang,Tengda Guo,Wentao Qiu,Ying Wang,Liujian Tang
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Long-term agents must remember not only what is true about a user, but also how a particular agent should work with that user as their shared history evolves. Existing benchmarks primarily supervise user facts and preferences or experience reusable across users, leaving this relationship-specific agent memory implicit. Additionally, most prior works measure the model solely with final-answer QA over long interaction histories, making the assessment still incomplete and unreliable. To this end, we introduce DyadMem with the proposed new definition User-conditioned Relational Agent Memory (URAM). DyadMem jointly annotates user-side memory and URAM along the same multi-session trajectories, resulting in 6 memory categories. To summarize, it includes 3,065 episodes, 50,961 sessions, and 61,210 QA instances, with extensive session-level Capture and Update gold annotations, query-level Recall support, and two QA settings: Gold-Memory and Full-Pipeline. Across 16 open-weight and 4 proprietary models, Gold-Memory QA is consistently strong, yet Full-Pipeline QA drops sharply. Such a gap explicitly supports our fine-grained evaluation design. Additionally, several quantitative results further reveal low Capture recall, incomplete Recall, and unsafe-deletion issues arising from even the frontier LLMs. We further conduct a rigorous experiment to validate the effectiveness of our URAM and observe the positive effects for all 20 models. In summary, DyadMem is a dual-domain, full-pipeline memory benchmark with extensive annotation efforts for advancing the domain’s development.

[AI-63] Beyond Predefined Sinks: Security-Aware Dependency Analysis for LLM Agents

链接: https://arxiv.org/abs/2610.03014
作者: Hang Cui
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
备注: 29 pages, 3 figures, including supplementary appendices. Preprint

点击查看摘要

Abstract:Large language model (LLM)-based agents increasingly connect model-generated decisions to security-sensitive software capabilities such as command execution, filesystem access, network communication, browser control, and external tools. Existing analyses often use predefined sensitive operations as anchors, but operation identity alone is insufficient to determine security implications. We present AgentSecGraph, a security-aware static analysis framework that constructs a candidate-centered Security-Aware Agent Dependency Graph (Security-ADG) for each security-sensitive operation. It augments operation identity with agent relevance, source and dependency evidence, trust-boundary context, guard evidence, and external-effect semantics. We further introduce AgentSecBench, a corpus of 67 real-world LLM-agent repositories spanning 11 ecosystems and 37,542 source files. The current analyzer identifies 23,866 static security-sensitive operation candidates across 65 repositories and emits one Security-ADG artifact per candidate. Corpus-wide analysis recovers source-to-operation dependency evidence for 9,821 candidates (41.15%) and potential guard evidence for 3,075 (12.88%), completing in 50.8 minutes. Using a separate reproduction-backed evaluation layer, we establish 22 security-sensitive behaviors across 13 repositories: one confirmed vulnerability, one pending disclosure candidate, and 20 guarded behaviors. In nine held-out cases, Security-ADG preserves 91.1% of the reference context and all five observed guards, compared with 20.0% for a sink-only view and 40.0% for a simplified ADG. These results show that security-aware dependency and contextual evidence enable distinctions that cannot be recovered from sensitive-operation identity alone. Comments: 29 pages, 3 figures, including supplementary appendices. Preprint Subjects: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI) Cite as: arXiv:2610.03014 [cs.CR] (or arXiv:2610.03014v1 [cs.CR] for this version) https://doi.org/10.48550/arXiv.2610.03014 Focus to learn more arXiv-issued DOI via DataCite (pending registration) Submission history From: Hang Cui [view email] [v1] Fri, 2 Oct 2026 08:48:09 UTC (5,798 KB)

[AI-64] AvoKV-E: Payload-Aware KV Cache Eviction for Long Reasoning

链接: https://arxiv.org/abs/2610.03007
作者: Han Yu,Wenhui Zhu,Xiwen Chen,Zhipeng Wang,Hejian Sang,Han Shi,Menglin Zhou,Xuanzhao Dong,Minzhou Huang,Rui Cai,Hao Wang,Alborz Geramifard
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Long-output reasoning shifts the KV-cache bottleneck from the fixed prompt to the generated trace. Existing reasoning-cache eviction methods largely treat cached entries as routing objects, estimating whether an old key will still be read, will recur, or can be replaced. This routing-only view overlooks two effects: low-attention entries can carry large value payloads whose removal changes future predictions, and newly generated states can appear stale before later queries have had a chance to read them. We introduce AvoKV-E, a training-free eviction policy that first delays eligibility for recent states and then ranks eligible entries using candidate-normalized read pressure, key redundancy, and value-payload potential. According to empirical evaluation across different models and datasets, AvoKV-E matches or exceeds redundancy-aware, recurrence-based, and thought-adaptive eviction baselines at matched active-KV budgets, with its largest gains in the tightest-cache regime. Component and counterfactual analyses further connect these gains to delayed observation, payload-aware scoring, redundancy, and scale-robust normalization. Together, the results show that long-reasoning KV eviction should preserve not only keys that are likely to be read, but also the value payloads that sustain the reasoning trajectory.

[AI-65] mporal Geometry of Deep Networks: Hyperbolic Representations of Training Dynamics for Intrinsic Explainability ICLR2026

链接: https://arxiv.org/abs/2610.03000
作者: Ambarish Moharil
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Neural and Evolutionary Computing (cs.NE); Data Analysis, Statistics and Probability (physics.data-an)
备注: Published as a main conference paper at ICLR 2026. 10 Main Pages, 22 pages of supplementary material

点击查看摘要

Abstract:Intrinsic explainability remains a challenging problem, particularly in contexts where multilayer perceptrons (MLPs) require dynamic re-training within an optimization environment. This paper investigates how MLPs and their training dynamics can be represented and studied in non-Euclidean spaces; our representation features the Poincaré model of hyperbolic geometry. We aim to capture the geometric evolution of their weighted topology and self-organization over time. Instead of restricting the analysis to single checkpoints—as per established measure-based explainability methods—we construct temporal \textitparameter graphs, i.e., snapshots over time T steps of the optimization/training process for MLPs. This reflects the view that neural networks encode information not only in their weights but also in the trajectory traced during training. Drawing on the idea that many complex networks admit embeddings in hidden metric spaces where distances correspond to connection likelihood, we present a geometric and temporal graph-based metalearning framework for obtaining dynamic hyperbolic representations of the underlying neural parameter graphs. Our model embeds temporal parameter graphs in the Poincaré model ball, and learns from them while maintaining equivariance to within-snapshot neuron permutations and invariance to permutations of past snapshots. In doing so, the approach preserves functional equivalence over time and recovers the latent evolving geometry of the network. Experiments on regression and classification tasks with trained MLPs show strong meta-network performance, accompanied by hyperbolic temporal representations. This reveals how the network structure emerges over time under specific training environments, thus providing insights into the network’s self-organization.

[AI-66] PLCWorld: Benchmarking LLM -Generated PLC Programs in Closed-Loop Plant Simulation

链接: https://arxiv.org/abs/2610.02982
作者: Yunji Kim,Yunseok Lee,Hyunwoo Seo,Jaerim Choi,Woojin Lee
类目: Artificial Intelligence (cs.AI)
备注: 36 pages, 9 figures. Project website: this https URL

点击查看摘要

Abstract:Programmable logic controllers (PLCs) coordinate industrial equipment by reading sensor inputs and issuing control commands. Evaluating whether large language model (LLM)-generated PLC programs satisfy task requirements and safety constraints requires observing how their commands affect device and workpiece states. We introduce PLCWorld, a common closed-loop execution environment and benchmark that couples Structured Text (ST) execution with simulated plant responses and sensor feedback. Grounded in control relations identified in industrial PLC programs and engineering documentation, PLCWorld contains 100 synthetic tasks and 473 registered task-condition pairs across Motion Control and Material Handling, with difficulty defined by control-dependency scope. A common protocol reports Task Success and Safety Violation separately. Validation combines practitioner review, reference and alternative programs, targeted counterexamples, specification-evaluator alignment checks, and comparisons with independent ST runtimes. Reference and alternative programs satisfy their applicable cases, while all 542 targeted counterexamples activate their designated evaluator rules under at least one registered condition. Execution Gap relates submission-profile acceptance to subsequent task failure or observed Safety Violation. Across the constructed task groups, direct GPT-5.5 achieves 82.70% Task Success on Easy cases but 25.10% on Hard cases. Evaluations of six LLMs and four adapted generation-and-verification workflows further expose differences between completion, safety, and generation cost. Our code, simulation environment, benchmark tasks, and baseline implementations are publicly available at this https URL.

[AI-67] Safeguarding Mutual Correction in Source-Free Domain Adaptation via Cut Statistics NEURIPS2026

链接: https://arxiv.org/abs/2610.02981
作者: Seongjun Lee,Changhee Lee
类目: Artificial Intelligence (cs.AI)
备注: Accepted at NeurIPS 2026

点击查看摘要

Abstract:Source-Free Domain Adaptation (SFDA) aims to adapt a source-pretrained model to an unlabeled target domain without access to the original source domain. While early single-model approaches rely on self-refinement, they are inherently susceptible to confirmation bias and struggle to correct their own systematic errors. To overcome this limitation, recent methods introduce Vision-Language (ViL) models as external knowledge sources. However, these approaches operate in a largely unidirectional paradigm, using the ViL model primarily to supervise the source-pretrained model. This overlooks a key structural property: the two models exhibit distinct failure modes – where one produces an incorrect prediction, the other may produce a correct one, creating a natural opportunity for mutual correction within the target domain. Yet, without ground-truth labels, identifying which model is correct on any given sample is non-trivial, and naively exchanging predictions risks propagating errors across models. To address this challenge, we propose SafeCut, a novel approach that leverages the cut statistic as a label-free measure of prediction reliability to gate cross-model supervision. Our approach dynamically controls both the direction and strength of supervision based on relative reliability, selectively amplifying true corrections while suppressing miscorrections on a per-sample basis. We further provide theoretical justification showing that this reliability-gated mechanism guarantees a net-positive correction signal. Extensive experiments across diverse SFDA benchmarks demonstrate that SafeCut achieves state-of-the-art performance, highlighting the effectiveness of safeguarding mutual correction in SFDA via cut statistics.

[AI-68] RASPER: Reward-Aligned Summarization of Clinical Notes for EHR Outcome Prediction EMNLP2026

链接: https://arxiv.org/abs/2610.02979
作者: Arya Hadizadeh Moghaddam,Mohsen Nayebi Kerdabadi,Chen Chen,Dongjie Wang,Zijun Yao
类目: Artificial Intelligence (cs.AI)
备注: Accepted to EMNLP 2026 Main Conference

点击查看摘要

Abstract:Unstructured discharge notes in Electronic Health Records (EHRs) often carry signal complementary to structured medical codes, holding patient-specific evidence that standardized cohort-level codes alone cannot capture. However, this evidence in notes is frequently buried in lengthy, noisy text that is not intentionally written with any specific clinical prediction in mind. Summarization is an obvious mitigation, but generic summaries, tuned for fluency rather than the outcome, routinely omit decisive evidence while retaining plausible but uninformative detail. To this end, we propose RASPER, a Reward-Aligned Summarizer for Prediction in EHR, that optimizes note summarization directly against the downstream clinical task. RASPER employs a tunable LLM-based summarizer to extract task-relevant evidence from discharge notes and trains it via reinforcement learning from prediction feedback, using a reward derived from the downstream predictor’s loss. To ground the summarizer, a longitudinal encoder converts structured codes into soft prompts that incorporate each patient’s clinical context into note summarization. By rewarding the quality of the resulting multimodal prediction, RASPER encourages the summarizer to retain patient-specific evidence that complements, rather than duplicates, information captured by structured codes. RASPER consistently outperforms strong baselines on both readmission prediction and medication recommendation across MIMIC-III and MIMIC-IV.

[AI-69] Relevant Evidence Decoding for Audio-Visual Hallucination Mitigation

链接: https://arxiv.org/abs/2610.02976
作者: Hyunjae Ra,Aecheon Jung,Jungin Park,Sungeun Hong
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Audio-Visual Large Language Models (AV-LLMs) remain prone to cross-modal hallucinations, where one modality incorrectly affects predictions about another. Although contrastive decoding reduces hallucinations in vision-language models, its direct extension to AV-LLMs overlooks a key challenge: different questions require different perceptual evidence, including audio, video, or their interaction. Notably, we observe that joint audio-visual inference can weaken the prediction even when a model can recover the correct answer from a single informative modality. For example, when asked which instrument is heard, a model may correctly predict violin from the audio alone. Once a video showing a guitar is added, its confidence in violin may drop. In this paper, we introduce Relevant Evidence Decoding (RED), a training-free method that identifies question-relevant evidence and selectively strengthens its contribution. RED uses pointwise mutual information to quantify the predictive support provided by audio and video beyond the question alone. It decomposes their joint contribution into audio, video, and residual interaction components. A question-only inference pass determines the required evidence type, after which the model augments the original audio-visual prediction with the corresponding PMI contribution. Across three audio-visual hallucination benchmarks and three AV-LLMs, RED improves accuracy over standard decoding by up to 7% on CMM, 6.3% on AVHBench, and 3.8% on SVHalluc, with an average relative time to first token of 1.5x standard decoding.

[AI-70] Reliable Self-Evolution with Imperfect Proxy Rewards

链接: https://arxiv.org/abs/2610.02975
作者: Kangjun Noh,Soyu Kim,Kyungwoo Song
类目: Artificial Intelligence (cs.AI)
备注: Preprint

点击查看摘要

Abstract:Large language model (LLM)-based self-evolving search is a promising approach to scientific discovery. However, high-fidelity evaluation of every candidate is prohibitively expensive in some domains. Self-evolving systems in such settings therefore rely on low-cost but imperfect proxy rewards, which may assign high scores to infeasible candidates. These false positives may contaminate both the final output and the feedback used to guide subsequent generations. This motivates statistically calibrated reward intervals for more reliable self-evolving search. We propose Conformal Interval-Driven Self-Evolution (CISE), which constructs candidate-specific reward intervals using conditional conformal inference and iteration-wise online density-ratio estimation. CISE uses conservative interval-based rewards for evolutionary feedback and returns candidates only when all required property intervals lie entirely within their respective feasible regions. We derive fixed-iteration coverage results under explicit assumptions of independence and covariate shift. We evaluate CISE on three self-evolving search tasks in materials science. In our experiments, all candidates returned by CISE are true positives under high-fidelity evaluation, whereas the baselines return more candidates but include false positives. These results highlight the value of a smaller, more precise shortlist when downstream validation budgets are limited. Our repository is available at this https URL.

[AI-71] CreateScore: Domain-Theory-Informed Bayesian Routing for LLM -Based CV Screening

链接: https://arxiv.org/abs/2610.02972
作者: Rupsa Roy
类目: Artificial Intelligence (cs.AI); Applications (stat.AP)
备注: 11 pages, 5 figures, 7 tables (Excluding Appendix). CreateScore Planner App GitHub repo link: this https URL

点击查看摘要

Abstract:Large language models (LLMs) can support rubric-based screening of CVs, but applying a high-capability model to every candidate and criterion is costly. We present CreateScore, a domain-theory-informed Bayesian network for criterion-level LLM routing. A hand-specified directed acyclic graph with Dirichlet-multinomial conditional probability tables converts CV evidence into posterior uncertainty; low-uncertainty decisions are resolved by a local 8B model and uncertain ones are escalated to a 120B reference model. The graph is causally motivated, but the system performs standard Bayesian conditioning, not causal inference. The escalation threshold is calibrated on a training fold (target: 70% resolved locally) and then frozen. On 200 synthetic Data Science CVs (139 training and 61 test candidates, five criteria), 77.7% of criterion decisions were resolved locally (237 of 305). Relative to a reference condition in which the 120B model adjudicated every criterion, routed escalation reduced token use by 65.2% and raised exact score agreement from 32.8% (8B alone) to 42.6% (95% CI 31.0-55.1%); at n = 61 the gain was not statistically distinguishable. The uncertainty signal did not, however, identify the decisions on which the 8B model erred: disagreement with the reference was 16.2% among escalated and 19.4% among locally resolved decisions (AUROC 0.47, 95% CI 0.39-0.56), no better than random selection. We also document how an earlier evaluation was invalidated when truncated reasoning-model outputs were silently replaced by local labels, and we recommend safeguards for cascade evaluation. CreateScore is supported as an auditable cost-reduction mechanism, not yet as a targeted error detector, and is not an autonomous hiring system.

[AI-72] Reasoning with Evidence Not Merely Rationales: Verifiable Preference Proofs for LLM -Based Recommendation

链接: https://arxiv.org/abs/2610.02968
作者: Yu Hou,Nathaniel Kang,Pengkai Wang,Hua Li
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Large language models (LLMs) can infer user preferences from interaction histories and reviews, yet the rationales they generate may not reflect the information actually used for recommendation. A preference claim may be weakly supported by its selected evidence, or may have little effect on the final ranking. We refer to these two failures as the grounding-influence gap. We introduce PROVE-REC, a general framework for verifiable preference reasoning in LLM-based recommendation. Pass A converts the complete pre-target history into a compact preference proof consisting of positive and avoidance claims linked to selected evidence entries. Pass B predicts the next item using only the proof and its selected evidence, preventing the recommender from bypassing the reasoning path. To verify evidence-to-proof grounding, we compare the effect of masking selected evidence with masking a comparable control entry. To verify proof-to-recommendation influence, we remove a preference claim and measure the resulting decrease in the target item’s ranking margin. A ranking-preservation objective further retains useful information from the complete history. Comprehensive experiments on wide-ranging real-world datasets demonstrate that PROVE-REC consistently outperforms strong sequential, generative, and LLM-enhanced baselines, with improvements of up to 7.45%. Controlled ablations confirm the effectiveness of the two-pass architecture and verification objectives. Moreover, PROVE-REC produces claims that are more strongly grounded in historical evidence and more influential to recommendation while preserving ranking quality.

[AI-73] When to Compile a Computer-Use Agent ? Measuring Payback and Making Compilation Decisions for Token Efficiency

链接: https://arxiv.org/abs/2610.02932
作者: Yulong Ming,Jie Xu,Zihan Wu,Xiaohua Jia
类目: Artificial Intelligence (cs.AI)
备注: 27 pages, 3 figures

点击查看摘要

Abstract:Compiling GUI procedures that agents execute repeatedly into programs can reduce their token costs. However, measuring payback and deciding when to compile have two challenges. First, compilation costs are uncertain because attempts can require repair and still fail to produce a usable program. Second, future reuse is unknown because tasks may stop arriving or GUI drift may stop the program from working. To address these challenges, we propose PACE (Payback-Aware Compilation from Experience), a system with a measurement protocol and an online compilation algorithm. The measurement protocol records successful and failed compilation costs, and compares agent and program execution costs on matched task inputs to estimate per-use savings and payback counts. Using these measurements, the online algorithm compares estimated future savings with compilation costs, including failed attempts, based on past task arrivals and compilation outcomes. It checks execution and compilation charges against a cumulative budget determined by observed task arrivals before allowing either action. Under stated action-cost assumptions, total cost after each arrival is at most 1+\epsilon times the cost of running every task with the agent. For successful compilation attempts, estimated payback counts excluding source agent runs are 2-16 uses. In simulations using recorded task arrivals, PACE reduces token costs by 17.3% compared with ReAct, 24.9% with the AutoRPA adaptation, and 17.3% with the ToolPro adaptation on average ( \epsilon=0.25 ).

[AI-74] Discriminating Fixture Coverag e in Agent -Infrastructure Verification Suites NEURIPS2026

链接: https://arxiv.org/abs/2610.02928
作者: Xin Xu,Siru Tao
类目: oftware Engineering (cs.SE); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: 10 pages, 1 figure, 2 tables. NeurIPS 2026 Workshop: Who Verifies the Agents? Toward Reliable Agent Development

点击查看摘要

Abstract:Invariant suites and runtime monitors increasingly gate agent deployment decisions, and the evidence offered for any particular suite is almost always a single observation: it passes an implementation believed correct and fails one believed broken. We measure what that observation is worth. Applying mutation analysis to an invariant suite for a multi-session agent state-projection layer, we first find that this standard validation certifies a suite in which a first-order mutant removing event-identity deduplication survives every check. We then freeze the repaired twelve-check suite, record its hash, and run it once against ten mutants specified by an adversarial reader who designed none of its fixtures: it kills five. Instrumenting the five survivors against the reference shows they fail in two distinct ways, not one. Three are never activated, because no fixture supplies an input on which the mutated code behaves differently at all. The other two corrupt internal state that no oracle in the suite can observe. The two modes need different repairs, and neither is visible from a pass/fail report. Treating the missing inputs as a coverage question, we enumerate seven discriminating dimensions of the input space, register in advance which are uncovered and which survivors they should explain, and add one fixture per uncovered dimension while reusing the existing oracles verbatim. All five survivors then die, each to the check written for its predicted dimension. We report this as a repair result on the same challenge set rather than a second held-out estimate, and give the artifact, including the frozen hash, the registered predictions, all mutants and the run logs, so the distinction is checkable.

[AI-75] Positive-Unlabeled Learning for Agent Safety False Alarm Auditing

链接: https://arxiv.org/abs/2610.02925
作者: Xichen Yan,Chongyang Gao,Kezhen Chen,Guangyi Zhang,Jiaqi Wu,Lixu Wang
类目: Artificial Intelligence (cs.AI)
备注: 19 pages

点击查看摘要

Abstract:Safety monitors help safeguard language-model agents interacting with external tools and environments, but conservative monitoring can generate many false alarms, consuming extensive review resources and weakening trust in alerts. Because false and genuine alarms often remain interleaved in native monitor scores, obtaining a reliable cutoff still requires substantial manual verification. In practice, a small set of verified-safe non-alarmed trajectories may be available while alarms remain unlabeled, naturally casting false-alarm auditing as a positive-unlabeled (PU) ranking problem. The key challenge is monitor-induced selection, since observed safe references are accepted by the monitor, while the hidden safe alarms of interest are precisely those it incorrectly flags, making the observed positives poorly representative of the positives to be recovered. To address this challenge, we propose a two-stage framework in which Trust-aware PU Supervision adapts safe references toward the alarm domain and protects plausible false alarms from excessive negative pressure, while Reliability-gated Rank Distillation consolidates consistent ordering preferences from multiple PU reference models into a single student. Consensus-guided Structural Refinement then improves the student ranking using hierarchical safe-reference support, alarm relations, and predicted reference consensus. The framework requires no alarm safety labels for fitting and leaves the underlying monitor unchanged. Across mainstream safety monitors, our method achieves a macro AUPRC of 0.6444 , outperforming eight evaluated PU baselines by 5.27–16.98 absolute percentage points; compared with PULDA, the strongest evaluated PU baseline, it recovers 33.3% more false alarms at a 5% review budget.

[AI-76] HASTE: Evolving Agent Harnesses Against Emerging Attacks Using Sparse Evidence

链接: https://arxiv.org/abs/2610.02920
作者: Xiqiao Xiong,Moxin Li,Zhixin Ma,Ouxiang Li,Wenjie Wang,Fuli Feng,Xiangnan He
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Agent harnesses play a critical role in defenses by enforcing safety constraints to prevent unsafe actions. However, rapidly emerging attacks outpace manual harness adaptation, motivating automated harness evolution. Yet the signals available for harness evolution are often sparse, such as brief descriptions or a few attack examples in threat reports and preprints. To address this limitation, we introduce HASTE, a multi-agent framework that evolves agent harnesses from sparse threat evidence through an adversarial interplay between safety-specification generation and attack-case generation. Safety specifications guide harness updates toward addressing identified safety vulnerabilities, while attack cases probe for remaining safety vulnerabilities after each update. By feeding evaluation outcomes back into both processes, HASTE enables harness evolution against emerging attacks beyond the initially observed evidence. Experimental results across multiple backbone models, attack types, and evidence forms show that HASTE consistently reduces attack success rates while preserving benign-task utility. The code is available at this https URL.

[AI-77] Frequency Is Not Sensitivity Identifying Safety-Sensitive Experts in Sparse MoE LLM

链接: https://arxiv.org/abs/2610.02910
作者: Md Nurul Absar Siddiky,Liuwan Zhu,Yingfei Dong
类目: Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Suppressing a small set of routed experts can weaken the safety behavior of a sparse Mixture-of-Experts (MoE) language model without retraining. Which experts to suppress is therefore a security question, and the usual answer is activation frequency, but frequency measures use, not influence. We test an alternative: router-gradient sensitivity, the sensitivity of the sequence loss to the gate weights that select an expert. Across five MoE architectures, we rank experts by each signal on 500 benign and 500 malicious prompts and measure refusal on 100 held-out malicious prompts under two budgets: equal expert counts and equal nominal malicious routing traffic (1%-5%). Under each of the two budgets, router-gradient selection reduces refusals more than activation in 24 of 25 conditions, and more than a ten-trial random mean in all 25. The largest effect is in OLMoE, where refusals fall from 34 to 9 of 100 prompts (73.53% relative) with no degraded outputs, indicating substantive compliance rather than broken generation. After matching expert counts in every layer, gradient selection still produces greater refusal reduction than activation in 23 of 25 conditions, with two ties. An exploratory cross-model analysis links larger malicious-versus-benign concentration gaps to greater peak gradient effects (rho = 0.90; exact two-sided p = 0.083, n = 5). Together, the results support gradient selection under the tested budgets.

[AI-78] LUMOS: Tracing Parametric Knowledge from Training Data to Behavioral Outputs in LLM s NEURIPS2026

链接: https://arxiv.org/abs/2610.02902
作者: Seoyeon Ye,Gayoung Kim,Jiyoung Hong,Sookyung Kim,Hyunsoo Cho
类目: Artificial Intelligence (cs.AI)
备注: Accepted to NeurIPS 2026 (Poster)

点击查看摘要

Abstract:Current analyses of LLMs’ parametric knowledge are largely output-centric, drawing conclusions about what a model knows without verifying what it was actually trained on. This leaves fundamental questions, such as whether a correct response reflects genuine generalization or rote memorization, grounded in speculation rather than evidence. To resolve these ambiguities, we introduce LUMOS, a diagnostic framework that traces knowledge along the causal chain from training-data exposure to behavioral output, leveraging OLMo 2 with its fully transparent training corpus. By grounding analysis in verified exposure, we reveal that models internally encode rare facts with high separability (84%) yet fail to express them behaviorally (54%), though this retrieval gap narrows with scale. Furthermore, when models are asked to self-reflect on their own answers, they perform reliably on trained content (83%) but drop to random-baseline levels (49%) on unseen content. This collapse persists even under chain-of-thought prompting, which inflates confidence signals rather than improving calibration. Collectively, these findings demonstrate that incorporating the training-data axis into LLM evaluation transforms speculative diagnoses into verifiable claims, and we advocate that this axis should be a standard component of knowledge assessment in LLMs.

[AI-79] Interpreting at Write Time: A Policy Ablation for Multi-Goal Agent Memory NEURIPS2026

链接: https://arxiv.org/abs/2610.02897
作者: Albert Sadowski,Jarosław A. Chudziak
类目: Artificial Intelligence (cs.AI)
备注: Accepted to PALM workshop at NeurIPS 2026

点击查看摘要

Abstract:A long-running assistant cannot keep everything it has seen, so it summarises. Summarising is not neutral: what is kept is chosen against some notion of what the record is for, and that choice is made once, before anyone knows which of the user’s standing goals will ask. Goals rarely disagree about what happened. They disagree about which parts of it were worth the space. Once the history is too long to re-read, the summary replaces the stream, and whatever it left out is gone. We ask what a memory should summarise for when it serves several standing goals at once. Three policies answer differently: summarise with no goal in view, write one summary covering every goal, or write one summary per goal and read them together. We compare them across several models and event streams, holding the read step fixed so that only the write differs. The goals do pull apart: summaries written for different goals overlap each other less than a summary overlaps a rewrite of itself. Per-goal summaries win on relevance, completeness and accuracy, and the all-goal summary loses even to the neutral one written at a fraction of its budget. Interpreting at write pays off, but only for the goal that later asks.

[AI-80] When Can We Trust the Matching Principle? Robust Deployment Geometry Under Finite-Sample and Model Uncertainty

链接: https://arxiv.org/abs/2610.02894
作者: Vishal Rajput
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: 14 pages. Companion to arXiv:2604.21395 and arXiv:2605.22800

点击查看摘要

Abstract:Match only geometry you can identify; otherwise spread the penalty. We quantify that decision by the trust ratio tau = epsilon / gamma (estimation uncertainty over spectral separation). Under the linear-quadratic Matching response, oracle-relative drift between estimated and oracle projector matching scales as tau^2 for probes in the chosen top-r deployment subspace – O(tau^2) in the Davis-Kahan separation region tau 1/2, with practical usefulness depending on constants. Confidence-Calibrated Matching (CCM) turns tau into a policy – directional when tau is small, progressively isotropic when not – with thresholds from calibration, not from the theorem (match sits in the separation region; soft is mostly heuristic). Experiments show both regimes, including UCI HAR embeddings where always-match is worse than abstain on every cell.

[AI-81] PsyEvo: A Personalized Counseling Agent That Self-Evolves at Test Time

链接: https://arxiv.org/abs/2610.02885
作者: Yuting Yan,Shihao Xu,Junhao Yu,Mingcong Zuo,Lu Chen,Nan Xiang,Haiyang Geng,Dongjie Tao,Minghao Wang
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Mental health disorders affect a substantial proportion of the global population, yet a persistent shortage of trained practitioners leaves the majority without adequate care. Large language model (LLM)-based counselors present a promising direction for delivering scalable conversational psychological support. Offline model training alone leaves limited room to adapt to individual clients or to learn from ongoing therapeutic interaction at test time. We introduce PsyEvo, an LLM-based counseling framework that enables both client-specific personalization and response-policy improvement at test time through three components: Hierarchical Bayesian Skill Policy (HBSP) personalizes what intervention to apply by maintaining a per-client skill posterior updated from session feedback; Inter-session Listwise Preference Optimization (LiPO) improves how the selected skill is expressed by updating a shared response adapter from cross-client preference evidence; and State-conditioned Ordinal Credit Assignment (SOCA) supplies candidate preferences and trajectory credit to the two components through consistency-checked comparisons and ordinal projection. In simulated-client evaluation with shared online cohort adaptation, PsyEvo obtains 7.684 Overall on PsychEval and exceeds every component variant in each of three matched runs. Removing individual components lowers mean overall score by 0.138–0.171 under the shared configuration, supporting conditional contributions within the complete scaffold. Our code is available at this https URL

[AI-82] DyRA: Dynamic Residual Approximation for Efficient Matrix Multiplication in DNNs NEURIPS2026

链接: https://arxiv.org/abs/2610.02882
作者: Daewon Chae,Hyunwon Chung,Changwoo Lee,Hun-Seok Kim
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: NeurIPS 2026. Code: this https URL

点击查看摘要

Abstract:Large-scale foundation models achieve strong performance across diverse tasks, but their size makes inference costly, largely due to dense matrix multiplications. Prior work reduces this cost by replacing dense weight matrices with efficient structured forms such as low-rank factorizations. However, these methods approximate weights rather than the output activations that determine inference accuracy. Consequently, small weight-space errors can be amplified by input activations, producing large output errors. In this work, we propose DyRA, an input-adaptive method that improves structured matrix multiplication approximation by correcting residual output errors during inference. We show that matrix multiplication can be approximated more effectively by directly optimizing low-rank factors of the output. DyRA builds on this insight by dynamically approximating and correcting the output error introduced by structured weight approximations. This combines efficient structured computation with input-dependent correction, yielding a more faithful approximation of full matrix multiplication under the same computational budget. Across vision, speech, and language models, DyRA consistently improves the accuracy-efficiency trade-off over structured weight approximations alone. Notably, DyRA achieves a 1.5 \times end-to-end GPU speedup for DINOv3 while reducing accuracy degradation by more than 3 \times relative to weight-only baselines.

[AI-83] Agent Trap: Stateful Feedback Deception against Autonomous Penetration Testing Agents

链接: https://arxiv.org/abs/2610.02869
作者: Yuelin Wang,Jiongchi Yu,Yanbang Sun
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
备注: 4 pages

点击查看摘要

Abstract:Autonomous penetration testing agents conduct multi-step attacks by continuously adapting their plans and actions to target responses. As a common defense, honeypots can be deployed to divert these agents from real assets by presenting decoy services, while also supporting attack tracing and active counterattacks. However, conventional honeypots rely primarily on static artifacts and predefined responses, leaving them unable to adapt to the evolving attack strategies of autonomous penetration testing agents. To this end, we present AgentTrap, the first closed-loop honeypot tailored for autonomous penetration testing agents. AgentTrap uses sentinel endpoints to avoid benign interference, stateful deception grounded in the protected application, and behavior-guided escalation to sustain engagement and collect agent-side behavioral evidence with controlled disclosures. We evaluate AgentTrap against eight autonomous penetration-testing agents in a deployed web application containing a real application endpoint and a separate honeypot endpoint configured under three defense strategies. Compared with no defense, AgentTrap reduces the aggregate real-target attack success rate from 95.8% to 79.2% and successfully elicits attacker API keys in 18.8% of the runs, outperforming static deception and fixed escalation. Furthermore, trace analysis shows that resistance to such counterattacks depends jointly on model-level recognition of deceptive requests and architecture-level isolation of sensitive resources. Comments: 4 pages Subjects: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI) Cite as: arXiv:2610.02869 [cs.CR] (or arXiv:2610.02869v1 [cs.CR] for this version) https://doi.org/10.48550/arXiv.2610.02869 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[AI-84] Distributionally Robust Survival Models under Subpopulation Shift and Outlier Contamination

链接: https://arxiv.org/abs/2610.02868
作者: Seonghwi Kim,Sung Ho Jo,Minwoo Chae
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: 36 pages, including appendices

点击查看摘要

Abstract:Learning robust survival models under distribution shift is an important but challenging problem in many applications. In heterogeneous populations, a model that performs well on average may still perform poorly on certain subpopulations, and this issue becomes even more severe when the training data are contaminated by outliers. In this paper, we propose a novel distributionally robust framework for survival analysis that jointly addresses latent subpopulation shift and outlier contamination. The proposed method combines an outer minimization that selects a refined nominal distribution by reducing the influence of contaminated samples and an inner maximization that focuses on the most challenging subpopulation. This formulation directly accommodates non-decomposable survival losses while preserving interactions across samples, including the risk-set structure of the Cox negative partial log-likelihood. We develop an alternating gradient-based algorithm with outer updates derived from the KKT conditions of the inner maximization. Experiments on simulated data and two survival benchmarks demonstrate that the proposed method remains robust when subpopulation shift and outlier contamination occur simultaneously. It stabilizes training in contaminated settings and substantially improves worst-group performance across both linear and nonlinear survival models, while maintaining competitive and sometimes superior overall performance.

[AI-85] ACD: Distilling Efficient Text-to-Motion Models via Terminal Amplification Control

链接: https://arxiv.org/abs/2610.02867
作者: Wei-Jin Huang,Yuan-Ming Li,Kun-Yu Lin,Wang Luo,Yinlin Zhu,Yue Yu,Shenghao Ye,Junbin Yuan,Fa-Ting Hong,Qing Zhang,Wei-Shi Zheng
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Recent text-to-motion models have improved motion quality and instruction following, yet many-step denoising and large model components make deployment slow and memory-intensive. We present Terminal-Amplification-Controlled Distillation (TACD), an on-policy approach for training efficient motion generators from text prompts and pretrained teachers, without real-motion training data. Building on segmented on-policy flow distillation, we supervise clean-motion predictions along student-generated trajectories. We identify a failure mode in which velocity matching on a fixed supervision grid repeatedly overweights errors near the denoising endpoint, degrading few-step generation. TACD ties the latest teacher query to the student’s step size, bounding the effective loss weights in clean-motion space without changing inference. Experiments on HumanML3D and KIT-ML demonstrate improved few-step generation, including a 58% reduction in eight-step HY-Motion student FID relative to distillation without this bound. For diffusion teachers, the endpoint-matching form of TACD yields four-step students with lower FID and matched or improved text-motion retrieval relative to their 50-step teachers on HumanML3D. On HY-Motion and Kimodo, eight-step students with compact components achieve 7.7-11.9x end-to-end speedups and reduce peak GPU memory by 3.8-6.7x relative to their teachers. Project page: this https URL

[AI-86] Harness-Aware Distillation for Small Language Model Agents

链接: https://arxiv.org/abs/2610.02858
作者: Moonseok Choi,Taehong Moon,Giung Nam,Juho Lee
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Language model agents are deployed with a harness, the software around the model that manages its context, tools, and feedback. When such an agent is distilled into a smaller one, the harness stays in place, so the student mainly needs the teacher-specific abilities that the harness cannot provide, such as acting correctly on harness information. Standard distillation, however, imitates the teacher’s full outputs and treats the harness as part of the input. We propose Harness-Aware Distillation (HAD), which focuses distillation on what the teacher adds beyond the harness. HAD complements on-policy distillation with two components: an action preference that contrasts the same teacher’s actions with and without the harness information, scored after the student’s own reasoning, and a validity check that drops preference pairs whose preferred action contradicts the harness records. We show that the contrast gives the student information that imitating the teacher alone cannot provide, and HAD needs no task rewards, success labels, or future information. Across multiple long-horizon agent benchmarks and models, HAD outperforms on-policy distillation baselines with the same fixed harness. Our analysis shows that HAD enters fewer unproductive loops and recovers from errors more often than the baselines, and suggests that it adaptively keeps learnable feedback in its weights while reading state information from the harness.

[AI-87] Bounded Reachability Jailbreak Detection via Contraction-Constrained State Space Models

链接: https://arxiv.org/abs/2610.02853
作者: Omanshu Thapliyal
类目: Artificial Intelligence (cs.AI)
备注: 21 pages, 18 figures, AIMS Workshop @ COLM 2026

点击查看摘要

Abstract:Safety heads are lightweight classifiers attached to pretrained language models for flagging harmful inputs before generation. Their empirical detection performance has been studied, but their formal robustness properties remain largely unexplored. We ask when a State Space Model (SSM)-based safety head can be certified to produce the same prediction for all inputs within a bounded embedding-space perturbation. We prove that the answer turns on a single condition: the l_\infty norm of the state transition matrix must satisfy \normA_\infty1 (the \emphcontraction condition), which enables exact interval bound propagation (IBP) certification for linear time-invariant classifiers. When the contraction condition holds, the reachable output interval has bounded steady-state width and examples can be certified as robustly classified. When it fails, the interval grows exponentially with sequence length and certification is impossible at any practical perturbation radius. We enforce contraction with a hinge penalty and show on toxic comment data that certified fraction improves from 41% to 59%, with a sharp empirical phase transition at \normA_\infty=1 matching the theory. Applying a contraction-regularized S4 head to jailbreak detection on JailbreakBench, we achieve a zero-shot transfer to AdvBench (DR=0.994) and HarmBench (DR=0.988). A logistic regression on mean-pooled Mamba-130M embeddings matches or exceeds the S4 head on every detection metric, confirming that harmful intent is already linearly separable in the embedding space. The S4 safety head’s contribution is not superior discrimination but the formal certification that no probe-based approach provides.

[AI-88] DNAlign: Dynamic Null-Space Safe Alignment for LLM s

链接: https://arxiv.org/abs/2610.02844
作者: Jisheng Dang,Yushuo Zhao,Dewei Liu,Junfeng Fang,Bimei Wang,Tiantian Rao,Hong Peng,Bin Hu,Tat-Seng Chua
类目: Artificial Intelligence (cs.AI)
备注: Regular Paper; 13 pages, 8 figures, and 1 table

点击查看摘要

Abstract:Ensuring the safe and reliable deployment of large language models (LLMs) remains a fundamental challenge. Existing safety alignment approaches either incur high computational cost or unintentionally disrupt the model’s core knowledge, leading to degraded fluency and factual accuracy on benign tasks. This reveals a persistent trade-off between safety and utility. We propose DNAlign, a lightweight alignment framework that integrates control-theoretic optimization with null-space projection. By treating the LLM as a dynamic system, the proposed framework introduces controllable perturbations to steer generation toward safe behavior. A key component is the projection module, which restricts these perturbations to the harmful-related subspace derived from neutral hidden states, thereby preserving general knowledge and response quality. A value function trained on human preference data adaptively optimizes the control signals to align with human safety preferences. Extensive evaluations across multiple LLM backbones demonstrate that our framework consistently reduces harmful outputs while maintaining fluency, coherence, and factual utility. It achieves superior overall performance compared to prior alignment baselines without sacrificing generation diversity. These results indicate that the proposed framework provides an effective and practically deployable solution for safe LLM alignment. Code is available at this https URL.

[AI-89] AMBER: Multi-View Adaptive Budget Allocation for Listwise Vision-Language Reranking

链接: https://arxiv.org/abs/2610.02831
作者: Wenteng Chen,Jiachen Zhu,Rong Shan,Tianyi Xu,Yuxiang Chen,Congmin Zheng,Teng Wang,Junjie Wu,Weiwen Liu,Changwang Zhang,Weinan Zhang,Jun Wang,Jianghao Lin
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Vision-language models (VLMs) are powerful listwise rerankers for multimodal retrieval, but high inference costs restrict them to evaluating small local candidate views. Existing multi-call strategies rely on fixed schedules, wasting expensive VLM calls on uninformative candidate pairs and easy queries. To address this, we propose Adaptive Multi-view Budgeted Elo Reranking (AMBER), an online, budgeted multi-view reranking framework that dynamically optimizes global resource allocation. AMBER treats fragmented listwise VLM outputs as local tournaments, using continuous Elo updates to maintain a lightweight global ranking state. Building on this, it allocates computation at two levels: dynamically constructing candidate views with high score ambiguity, and scheduling queries to maximize expected information gain. We show that each Elo update corresponds to a stochastic gradient ascent step on the Bradley-Terry log-likelihood, and provide a submodular information-theoretic motivation for the query-level allocation strategy. Experiments on CIRR, CIRCO, and PhotoBench demonstrate that AMBER achieves the strongest overall performance among the compared multi-call VLM reranking methods under comparable VLM-call budgets, while remaining effective in lower-budget settings. Our code is publicly available at this https URL.

[AI-90] MLCommons Jailbreak Benchmark v1.0

链接: https://arxiv.org/abs/2610.02827
作者: Carsten Maple,Cagatay Yucel,Isaac Holeman,Chris Knotz,Peter Mattson,James Goel,Jonathan Petit,Sean McGregor,James Ezick,Abhishek Kumar,Alicia Parrish,Murali Emani,Kashyap Iyer,Faiza Khan Khattak,Washington Mbonu,Daniel Machlab,Eileen Long,Shaona Ghosh,Jibin Varghese,Roman Lutz,Andrew Gruen,Bennett Hillenbrand,Prabal Gupta,Mohammed Serrhini,Dhivya Nagasubramanian,Aakash Gupta, Jun (Victor)Lu,Kurt Bollacker,Chang Liu,Jonathan Petit,Cong Chen,Jean-Philippe Monteuuis,Brent Miller,Apurv Verma,Roman Eng,Armstrong Foundjem,Mohammed Serrhini
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Modern AI systems are designed to refuse hazardous requests. A jailbreak is a prompt crafted to bypass those safeguards and elicit outputs that the system would normally refuse to provide. The MLCommons Jailbreak Benchmark v1.0 provides an end-to-end methodology for evaluating the robustness of large language models to single-turn, text-based jailbreak attacks. It combines criteria-driven system and attack selection, paired baseline and adversarial evaluation, human annotation, automated evaluator calibration, scoring, grading, and risk-calibrated disclosure within a single benchmarking pipeline. The benchmark evaluates eight open-weight systems using 264 seed prompts spanning eleven hazard categories and representative attacks drawn from the MLCommons Jailbreak Taxonomy. Responses are assessed using the AILuminate Assessment Standard v1.4, and robustness is measured through the Resilience Gap: the change in safety performance between baseline and adversarial conditions. Across all evaluated systems and attacks, the unsafe-response rate increased from 11.08% under baseline conditions to 18.65% under jailbreak conditions, producing an average Resilience Gap of 7.57%. Accessible systems showed a larger mean gap, while attack effectiveness varied substantially across attack categories and hazards. The benchmark also examines evaluator reliability and sources of measurement error. Beyond reporting results, Jailbreak Benchmark v1.0 establishes a reproducible methodological foundation for comparative jailbreak evaluation and for future expansion across systems, attacks, hazards, and evaluation methods.

[AI-91] Scaling Trajectories for Complex Tasks through Recursive Self-Rewrite

链接: https://arxiv.org/abs/2610.02826
作者: Zongxia Li,Yucheng Shi,Zhongzhi Li,Junyao Yang,Ruhan Wang,Chengsong Huang,Fuxiao Liu,Haitao Mi,Jordan Boyd-Graber,LeoweiLiang
类目: Artificial Intelligence (cs.AI)
备注: 16 pages, 5 figures. Model weights: this https URL

点击查看摘要

Abstract:Successful trajectories on difficult tasks provide valuable supervision for model improvement, but specialized harnesses introduce interventions that may be unavailable during deployment. We propose Recursive Self-Rewrite (RSR), a framework that uses one base model, Qwen-3.8-27B, to discover successful solutions under diverse harnesses and reconstruct them as training trajectories under a general harness. A planner extracts procedures into runbooks, a critic screens for verifier and solution leakage and guides recursive revision, and an executor follows qualified runbooks in fresh sandboxes. Across approximately 3K self-curated terminal tasks, three harnesses jointly solve 759 tasks, 34.3% more than the strongest individual harness in the recorded pool. RSR expands 2,001 successful source trajectories into 11,094 rewritten trajectories for supervised finetuning. Training on these trajectories outperforms both the base model and direct trajectory SFT. Compared with the base model, pass@3 increases from 57.0% to 74.2% on Terminal-Bench 2, from 1.5% to 9.1% on Terminal-Bench 4, from 39.0% to 63.0% on our self-curated Terminal-Bench Hard, and from 3.0% to 6.0% on our Software Terminal-Bench. Process reward on Long-Horizon Terminal-Bench rises from 0.21 to 0.29. These results show how diverse harness-assisted experiences can be reconstructed into reusable capabilities for a model operating under a general harness.

[AI-92] MetaRubric: Learning to Reward for Rubric-Based Reinforcement Learning

链接: https://arxiv.org/abs/2610.02824
作者: Yuxuan Fan,Jaehong Yoon
类目: Artificial Intelligence (cs.AI)
备注: Project page: this https URL

点击查看摘要

Abstract:Rubric-based reinforcement learning extends reward-driven optimization to open-ended tasks by assigning partial credit to individual response requirements. However, rubric judges can assign a high criterion score even when the information or action it requires is absent from the response, a failure mode we term Vacuous Credit. Such awards persist after the required information is removed and can reverse the sign of a response’s GRPO advantage. To address this problem, we introduce MetaRubric, which alternates evidence-aware policy optimization with response-guided rubric adaptation. We construct counterfactual counterparts by changing one task-relevant fact in each prompt. During policy optimization, credit is assigned only when the response contains sufficient evidence to satisfy the required rubric criterion. After each policy-optimization stage, current policy responses guide revisions to original and counterfactual criteria while preserving the meaning of the original prompt’s initial rubric as interpreted under each prompt’s facts. We also adapt criterion weights at stage boundaries to better address observed policy errors. Across multiple backbones, MetaRubric improves PubMedQA accuracy by 6.00–20.40 percentage points over static-judge GRPO, with further gains on HealthBench-Hard and two multimodal medical benchmarks.

[AI-93] Adaptive Spectral-Koopman Dynamics Modeling for Temporal Domain Generalization

链接: https://arxiv.org/abs/2610.02822
作者: Tengxue Zhang,Yu Ke,Yang Shu,Chenchen Sun,Yisheng An,Chenjuan Guo,Bin Yang
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Temporal Domain Generalization (TDG) has emerged to address real-world streaming data with distribution shifts over time. However, existing methods are either prone to overfitting to domain-specific noise in the data space or become overly complex and less interpretable in the parameter space. To bridge these gaps, we propose \textbfAdaSpecK, a spectral-Koopman framework with adaptive context extraction for TDG. To mitigate noise fitting to irregularly sampled domains, we introduce spectral-regularized Koopman dynamics modeling, which applies spectral-aware filtering in the latent space to extract denoised low-frequency trajectories and learn a Koopman operator to model the system dynamics in a linearized space. To model complex historical environments under non-stationarity, we design a context-informed heterogeneous pattern extraction mechanism. Specifically, we employ a target-conditioned attention module to attend to distinct past windows, producing a dynamic, target-specific historical summary. By constructing an environmental signature from the current evolutionary pattern, our model adaptively perceives which aspects of the past context are most informative for future prediction via a learned router. Extensive experiments on eight diverse classification and regression benchmarks demonstrate that AdaSpecK achieves state-of-the-art performance. The code and datasets are available at \hrefthis https URL.

[AI-94] S-KV: Online Low-Rank KV Cache Compression via Block-Incremental SVD

链接: https://arxiv.org/abs/2610.02815
作者: Yiren Zhao,Guanghui Song,Tianrui Qin,Kejiang Ye,Cheng-zhong Xu,Xitong Gao
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Long chain-of-thought reasoning substantially increases KV-cache memory during autoregressive decoding, as every generated token introduces new key and value states and causes the cache to grow linearly with decoding length. Existing KV-cache compression methods typically control this growth through token eviction, but irreversible deletion can remove historical states that later reasoning may need to revisit. SVD-based low-rank compression provides an alternative by retaining all positions with a more compact representation. However, extending it from a fixed prompt cache to online decoding is non-trivial. Through our investigation, we find that if the basis is updated for new tokens while old tokens keep their coordinates in the old basis, the stored history drifts substantially. Based on this observation, we propose iS-KV, an online low-rank KV-cache compression method for long-horizon reasoning. iS-KV keeps a recent window exact while incrementally folding older states into bounded-rank representations. As the low-rank basis evolves, it synchronizes historical coordinates with the updated basis to maintain representation consistency. On DeepSeek-R1-Distill-Llama-8B, iS-KV achieves 82.6% accuracy at 4.06-fold persistent-KV compression, close to the original model’s 83.6%. On Qwen3-8B, it achieves 89.2% accuracy at 5.64-fold compression. Under matched memory budgets, iS-KV consistently outperforms token-eviction baselines.

[AI-95] VIGOR: Zero-Shot Visual Generalization via Latent-Space Consistency in Model-Based Reinforcement Learning

链接: https://arxiv.org/abs/2610.02801
作者: Mingyu Park,Samyeul Noh,Hyun Myung,Donghwan Lee
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Robotics (cs.RO)
备注:

点击查看摘要

Abstract:Model-based reinforcement learning (MBRL) achieves strong sample efficiency by planning within learned latent dynamics, yet its performance degrades substantially under unseen visual distractions such as background variations, lighting changes, or camera shifts. Unlike model-free RL, where encoder perturbations affect only single-step predictions, MBRL suffers from a two-level vulnerability: visual distractions first push encoder outputs out of distribution, and these errors then compound through recursive latent rollouts over the planning horizon. We propose visual generalization via latent-space consistency in model-based RL (VIGOR), a framework that enables zero-shot generalization to unseen visual distractions while retaining the sample efficiency of its MBRL backbone. VIGOR integrates three interdependent components: (i) asymmetric weak-to-strong augmentation, which pairs weak-only and weak-to-strong latent views within a single batch; (ii) dynamics-level consistency, which enforces augmentation-invariant transition predictions through direct latent regression; and (iii) encoder-level stabilization, which prevents encoder drift under the cross-augmentation supervision imposed by dynamics-level consistency. Evaluations on the DeepMind Control Suite (DMC) and Robosuite show that VIGOR outperforms state-of-the-art model-free and model-based baselines, surpassing the second-best baseline by 3.4% on DMC and 43.6% on Robosuite. Ablations further show that VIGOR’s robustness is augmentation-agnostic: replacing the default augmentation with alternatives from distinct perturbation families preserves strong generalization, confirming that latent-space consistency, not the augmentation choice, drives robustness.

[AI-96] BitNest: Bit-Nested Speculative Decoding for Memory-Efficient LLM Inference Acceleration

链接: https://arxiv.org/abs/2610.02800
作者: Chence Yang,Ningxi Cheng,Arash Akbari,Qitao Tan,Qingchan Zhu,Ci Zhang,Changdi Yang,Yanzhi Wang,Wei Niu,Jinhui Wang,Jin Lu,Geng Yuan
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Speculative decoding accelerates autoregressive generation by using a lightweight draft to propose multiple tokens for parallel verification. However, existing methods often require an additional draft model or weight representation, introducing non-negligible memory overhead on resource-constrained devices. Self-speculative approaches reduce this overhead, yet still face trade-offs between draft quality, target quality, and storage efficiency. We propose BitNest, a bit-nested speculative decoding framework that embeds a low-precision draft directly into the higher-precision target representation. Instead of deriving a draft from a predefined target, BitNest first constructs a strong low-precision base and then recovers the higher-precision target through residual refinement, enabling both models to share a single physical weight representation. BitNest further extends this progressive-precision design to the KV cache for long-context inference. Across multiple 7B–8B edge-friendly LLMs and diverse workloads, BitNest achieves an average speculative acceptance rate of 95.2% while closely preserving higher-precision model quality, and delivers 1.48–1.61x end-to-end speedup over FP16 autoregressive decoding. On the LLaMA models supported by all representative self-speculative baselines, BitNest also achieves consistently competitive or higher decoding speedup.

[AI-97] Modeling Shared and Individual Structure for Cross-Subject Continuous Affect Regression from EEG-fNIRS

链接: https://arxiv.org/abs/2610.02796
作者: Xuan Wang,Bing Wang,Shuai Chang,Hao Yuan,Xinbo Qi,Xinyue Zhang
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Continuous, second-by-second valence-arousal estimation from physiological signals is typically studied in a subject-dependent setting, where the model sees labeled data from the same person it is later evaluated on. We study the harder zero-shot cross-subject variant on a synchronized EEG-fNIRS dataset: predict raw-scale ([1, 255]) valence and arousal trajectories for subjects whose labels the model never observes, given only their unlabeled EEG/fNIRS recordings while watching the same video stimuli as a disjoint set of training subjects. We decompose the affect trajectory into a structure shared across subjects who watch the same stimuli and an individual structure estimated for each test subject from a label-free EEG marker (alpha-band cross-channel synchrony), which rescales the shared trajectory around the scale midpoint. We validate the per-subject calibration mechanism on four independent axes: leave-one-subject-out correlation between the marker and each subject’s true optimal gain, a functional-form comparison against non-linear alternatives, a repeated leave-4-out component ablation isolating each part of the pipeline’s contribution, and a ceiling analysis bounding the remaining headroom for per-subject scaling. On held-out subjects, the model reaches an overall MAE of 25.96 / 22.80 across two evaluation batches (valence 21.94 / 19.6, arousal 29.98 / 26.0), well below EEGNet and ASAC-Net baselines reported for the same subject-independent split (raw scale score 60.6 and 55.0 respectively). We further report a systematic negative-result search across model architectures, feature representations, and prediction targets that found no signal able to improve on the single alpha-synchrony marker.

[AI-98] PAPER2LLM : Continual Self-Evolution of LLM s from Research Papers

链接: https://arxiv.org/abs/2610.02793
作者: Hongji Pu,Yilun Zhao,Wenpeng Yin
类目: Artificial Intelligence (cs.AI)
备注: 20 pages, 5 figures, 12 tables

点击查看摘要

Abstract:Research on LLMs continually uncovers model limitations, their causes, and potential solutions. Yet these human discoveries remain largely disconnected from model evolution: an LLM does not automatically learn from new research about its own failures. We introduce PAPER2LLM++, a framework for continual self-evolution of LLMs from research papers. Rather than treating papers merely as knowledge to retrieve, PAPER2LLM++ uses the growing literature as a stream of evidence and supervision for model improvement. For each incoming paper, it extracts evidence-grounded findings, tests whether the reported limitation persists in the current model, and, when needed, converts the findings into candidate learning signals. A try-evaluate-commit procedure integrates an update only when it improves the targeted behavior without substantially forgetting prior improvements or degrading general capabilities. Across a sequential stream of research-discovered LLM failures, we show that models can progressively incorporate new findings while retaining earlier gains. PAPER2LLM++ thus takes a step toward closing the loop between human discovery and model evolution, enabling models to continually learn from research about their own limitations and improvements.

[AI-99] Law And Order: Tax Law Autoformalization

链接: https://arxiv.org/abs/2610.02792
作者: Sophia Simeng Han,Yoshiki Takashima,Anjiang Wei,Zhaoyu Li,Michael Genesereth
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Legal systems are increasingly implemented through software, yet scalable methods for translating legal texts into accurate symbolic representations remain underdeveloped. We study this problem through tax law, where forms and filing instructions define large computational structures involving arithmetic, branching, recursion, and tabular reasoning. We propose LawOrder, a neuro-symbolic framework for automatically formalizing tax forms and instructions into executable symbolic programs. Our approach establishes two forms of correspondence between law and logic: structural correspondence, which aligns legal and symbolic components such as cells and schedules, and denotational correspondence, which requires symbolic components to implement the computations specified by their legal counterparts. We combine large language model synthesis with cell-level verification and iterative localized error repair using human-written OpenTaxSolver tax returns. We then evaluate the resulting formalizations on independently authored, held-out TaxCalcBench returns, that are never exposed during generation or repair. Although the most advanced LLM achieves only 66% accuracy, LawOrder achieves 100% cell-level and form-level accuracy on 51 held-out returns, demonstrating the effectiveness of combining LLM-based synthesis with symbolic verification for scalable and verifiable large-scale legal autoformalization compared with using an LLM alone.

[AI-100] Nearly Optimal Fixed-Confidence Best-Arm Identification with 1-Bit Feedback NEURIPS2026

链接: https://arxiv.org/abs/2610.02771
作者: Khang Luong,Dinh Thai Son,Hoang Ta,Hung The Tran,Tuan Quang Dam
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Machine Learning (stat.ML)
备注: To appear in Advances in Neural Information Processing Systems 39 (NeurIPS 2026, Spotlight)

点击查看摘要

Abstract:We study fixed-confidence best-arm identification under strict 1-bit feedback constraints. At each round, the learner selects an arm and a query set, and receives only a single bit indicating whether the sampled reward belongs to that set. We consider a distribution-free finite-variance setting with arm-wise localization, where direct empirical mean estimation is no longer available and clipping becomes unavoidable. We first formulate a time-uniform 1-bit mean-estimation primitive based on randomized threshold queries and a clipped tail-integral identity. We then embed this primitive into candidate-challenger best-arm identification algorithms. A fixed-clipping algorithm gives a simple anytime (\epsilon,\delta) -PAC guarantee, while a phased adaptive-clipping algorithm matches the clipping level to the current resolution and yields a gap-adaptive sample complexity. We also prove a K -arm worst-case information-theoretic lower bound showing that the logarithmic penalty caused by finite-variance 1-bit feedback is intrinsic. This bound matches the leading dependence of the phased algorithm up to lower-order \log\log factors.

[AI-101] Dynamic LLM Routers are Often Misguided NAACL

链接: https://arxiv.org/abs/2610.02762
作者: Sam Wang,Julia White,Sahibzada Allahyar,Dhruv Atreja,Urchade Zaratiana,Kelton Zhang
类目: Artificial Intelligence (cs.AI)
备注: 8 pages, 15 figures, under review at NAACL

点击查看摘要

Abstract:Dynamic LLM routers promise to cut inference costs by sending each query to the cheapest model that can answer it correctly. We analyze six commercial routers across 14 settings on a diverse benchmark spanning eight task categories, finding that none of them outperforms a router that randomly selects between two well-chosen models at matched cost. Some underperform by more than 10 percentage points. We trace this gap to four patterns prevalent across routers: difficulty blindness, length reversal, semantic matching, and roster suboptimality. We show that the first three are what the standard objective rewards: cost-accuracy Pareto efficiency on realized costs favors escalating moderately hard queries over the hardest ones, shorter queries over longer ones, and routing by a query’s source over its difficulty. We also argue that the two assumptions that would justify large rosters, model granularity and model specialization, do not hold empirically. We propose an alternative evaluation methodology that does not reward these patterns, and as a proof of concept, we design a simple two-model router that avoids all four. Nevertheless, its gain over random routing is limited, because a well-chosen roster leaves little to route.

[AI-102] On the Chain-of-Thought Monitorability of Looped Language Models

链接: https://arxiv.org/abs/2610.02741
作者: Han Wang,Ishwar B Balappanawar,Huan Zhang
类目: Artificial Intelligence (cs.AI)
备注: Preprint

点击查看摘要

Abstract:Chain-of-thought (CoT) monitoring provides a promising approach for detecting undesirable model behavior. Looped language models (LoopLMs) repeatedly apply shared transformer layers, increasing effective computational depth and enabling additional latent computation without increasing model size. However, the effect of looped architectures on CoT monitorability remains largely unexplored. In this work, we provide the first systematic evaluation of CoT monitorability in LoopLMs. We study two complementary settings: (1) varying the loop depth within the same LoopLM family to isolate the effect of additional recurrent computation, and (2) comparing LoopLMs with non-looped language models matched by parameter size, transformer-layer count, or effective depth to study whether LoopLMs are less monitorable. Across eight tasks from MonitorBench and both standard and stress-test settings, we observe task-dependent reductions in CoT monitorability under stress tests on specific Logic/Science/Engineering \textttCue Answer tasks, while other tasks exhibit weaker or qualitatively different trends. Our diagnosis suggests that these declines are not fully explained by task difficulty, verification pass rate, or generated token length; qualitative examples further suggest changes in how deeper-loop models explicitly use or attribute provided cues. Our cross-model comparison finds no evidence that LoopLMs are systematically less monitorable than non-looped language models matched on size or depth. Overall, our results suggest that deeper loop depth can reduce CoT monitorability in some tasks under stress tests, but looped transformer architecture alone does not necessarily imply lower monitorability.

[AI-103] Prospective Hindsight: Self-Calibrating Reinforcement Learning via Prediction-Reality Gaps NEURIPS2026

链接: https://arxiv.org/abs/2610.02740
作者: Jiaxin Zhang,Xiangyu Peng,Qinglin Chen,Yu Li,Hiroaki Hayashi,Chien-Sheng Wu
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: NeurIPS 2026

点击查看摘要

Abstract:Reinforcement learning for long-horizon agents relies on purely retrospective training signals: credit is assigned only after observing environmental consequences, leaving the agent’s belief at action time invisible to the gradient. We introduce Prospective Hindsight (PH), a self-calibrating training principle that augments any retrospective base method with a signal derived from the gap between the agent’s prospective prediction (before feedback) and the retrospective evaluation (after feedback). This per-rollout surprise identifies samples where the agent’s self-model is most inaccurate and amplifies their gradient contribution through a stop-gradient surprise-weighted advantage. Since the prospective predictor shares parameters with the policy, the two co-evolve, progressively shifting focus to the agent’s remaining blind spots. We connect this principle to a privileged-information gap and show that minimizing the surprise residual provides a descent pathway on the agent’s miscalibration rate; calibration thus emerges as a byproduct of optimization rather than from an added objective. On single-turn verifiable tasks and a multi-turn personal-agent task (under GRPO, on-policy distillation, and their combination), PH improves both task performance and calibration, with consistent gains across model scales. Notably, the dominant miscalibration mode shifts structurally between regimes, overconfident failures in single-turn, underconfident successes in multi-turn, yet the same training principle addresses both successfully.

[AI-104] Ego2World: Compiling Egocentric Cooking Videos into Executable Worlds for Belief-State Planning

链接: https://arxiv.org/abs/2610.02715
作者: Qinchuan Cheng,Zhantao Gong,Pengzhan Sun,Angela Yao,Shijie Li
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Egocentric videos capture how people carry out everyday activities, yet testing an agent requires evaluating the consequences of actions it chooses itself. We introduce Ego2World, a benchmark that turns annotated cooking activities into executable planning environments under partial observation. Its compiler links source steps and objects to symbolic action rules, persistent world states, and explicit task conditions, so researchers can execute an agent’s proposed actions and check their outcomes. World state and agent belief are maintained separately, enabling controlled studies of planning and information reuse across continuing tasks. Evaluating six planners on 105 tasks shows that accepted operations often leave task goals unmet. Execution traces and condition checks distinguish interrupted runs, partial attainment, and completed execution without goal attainment. In a separate paired Qwen-Plus study, persistent belief improves action validity by 4.15 percentage points and reduces visual-query attempts by 90.27%, with higher token use and no detected completion gain. Ego2World provides a reusable testbed for tracing how planning and memory choices affect execution, observation demand, and task attainment, connecting recorded human activity to the development and evaluation of interactive agents.

[AI-105] Self-Supervised Scaling of Terminal Environments for Scientific Domains

链接: https://arxiv.org/abs/2610.02710
作者: Zhongzhi Li,Yucheng Shi,Zongxia Li,Junyao Yang,Ruhan Wang,Yu Wang,Jingyuan Huang,Jichao Yu,Ninghao Liu,Haitao Mi,Leowei Liang
类目: oftware Engineering (cs.SE); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Terminal agents are increasingly deployed beyond software engineering in science and other specialized domains. Constructing training environments requires executable reference behavior and a domain-specific verifier that distinguishes semantic correctness from superficially plausible artifacts. Authoring these components for each task requires repeated engineering and limits reuse. We introduce software-in-the-loop reconstruction, a self-supervised framework that obtains reference outputs and verification targets from existing software workflows, executable programs mapping structured inputs to outputs. For each workflow, we execute multiple input configurations and partition cases into public observations and hidden evaluations. Given the instruction, input schema, and public input–output observations, an agent constructs an editable program without access to the source workflow. The candidate is evaluated on hidden configurations against workflow outputs. A hierarchical verifier combines domain-specific semantic comparison, structural validity, and anti-shortcut checks, while public feedback supports iterative revision. The construction admits additional workflows and configurations without authoring a reference solution for each task. We instantiate SWR with 500 workflows and 46 software families across six domains. Across three attempts per task, Qwen3.8-Max solves 838 tasks and produces 1,422 verified trajectories, which we oversample to 3,000 reconstruction-only training examples. Supervised fine-tuning of Qwen3.8-27B improves mean Terminal-Bench 2 performance from 47.94% to 53.56% across three seeds and achieves the highest mean among four matched-token corpus controls on all four reported evaluations. These results indicate that existing scientific software can provide scalable, behaviorally verified supervision for terminal agents.

[AI-106] MuonIO: Principled Norm-Aware Descent for Embedding Tables and Language Model Heads

链接: https://arxiv.org/abs/2610.02705
作者: Linkai Ma,Xinyu Luo,Mengbo Wang,Ananth Grama,Petros Drineas,Brian Bullins
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:The Muon optimizer derives its update rule for hidden linear layers by solving a local linearization of the loss penalized by the spectral norm, motivated by an RMS-stability argument for dense linear layers. Standard Muon implementations, however, exclude the input (embedding table) and output (language model head) layers from this principled treatment, for which they use AdamW instead. We present MuonIO, a single Muon-style update for both of these layers. For the language model head \mathbfL \in \mathbbR^V \times d , we motivate the use of the 2\to\infty operator norm, due to the Lipschitz continuity of the softmax output geometry, while for the embedding table \mathbfE \in \mathbbR^d \times V , we draw on the 1 \to 2 operator norm, based on the one-hot input geometry identified by Bernstein Newhouse (2025). The identity \lVert\mathbfL\rVert_2\to\infty=\lVert\mathbfL^\top\rVert_1\to2 then puts both matrices in the same vocabulary-oriented geometry: MuonIO applies a single normalized-vector rule, which appears as column normalization for \mathbfE and row normalization for \mathbfL . Empirical evaluations demonstrate the effectiveness of our approach, with MuonIO reducing I/O optimizer state memory by 50% and I/O update FLOPs by \sim 46% compared to Muon for 1B LLaMA pretraining on C4, while also improving validation perplexity.

[AI-107] Label-Efficient Time Series Classification at Scale: A Dual-Stream OSSE-LSTM with Counterfactual Attribution

链接: https://arxiv.org/abs/2610.02704
作者: Nguyen Ho,Bach Tung Tran,Trung Ky Nguyen,Zhenchang Xia,Bolong Zheng,Long Van Ho
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Time series are produced continuously at enormous scale by industrial equipment, wearables, power grids, and clinical monitors, yet annotation remains manual, expensive, and expert-dependent. The binding constraint in large-scale time series analytics is therefore not data volume but label volume, and the question facing a practitioner is concrete: how many examples per class must be labeled before a classifier becomes usable? We study this question directly, in a regime where the label space is fixed and known in advance and the decision rule must be constructed from only K labeled examples per class. We propose Dual-Stream OSSE-LSTM, an episodic metric-learning framework that pairs an Omni-Scale CNN with Squeeze-and-Excitation recalibration, for multi-scale motif extraction without per-dataset kernel tuning, with a Bidirectional LSTM for global temporal context. The two streams are independently normalized and fused into a prototype-oriented embedding. Because decisions taken from a few labels must also be explainable, we introduce Counterfactual Integrated Gradients (C-IG), which attributes the prototype margin between target and opposing classes rather than an isolated classifier logit, and reuses the resulting maps as soft masks for test-time prototype refinement without updating the encoder. On 19 univariate UCR datasets, OSSE-LSTM attains the highest average accuracy and per-dataset win count at every support size, and its accuracy remains within a 0.36-point band (96.36-96.72%) across that range. Its weakest configuration still exceeding the best result any compared baseline achieves at any K (93.99%).

[AI-108] Learning to Revise Reasoning with Segment-wise On-Policy Distillation

链接: https://arxiv.org/abs/2610.02703
作者: Yuxiang Zhang,Ding Cao,Shuting Cui,Lei Wang,Weijieying Ren,Tianxiang Zhao
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:On-policy distillation (OPD) improves large language model reasoning by training students on their own rollouts with dense token-wise supervision from the teacher. However, token-wise OPD does not explicitly provide a coherent alternative reasoning step showing how the student’s step could be revised to improve subsequent reasoning. Furthermore, this paradigm can become less effective when the student produces a degenerate reasoning prefix, as subsequent teacher supervision remains conditioned on that prefix and may reinforce poor reasoning patterns. In this work, we focus on learning reasoning revision with segment-wise OPD to rework intermediate reasoning steps and better support subsequent reasoning. Through controlled reasoning interventions, we find that replacing student segments with teacher redrafts improves subsequent reasoning accuracy. Therefore, we address the problem of turning teacher redrafts into explicit supervision for learning to revise reasoning. We propose Segment-wise On-Policy Distillation (Seg-OPD), which selects student segments based on an uncertainty metric and obtains corresponding teacher redrafts. Seg-OPD trains the student to prefer teacher redrafts over their paired student segments while retaining dense token-wise OPD supervision. Extensive experiments on mathematical reasoning and competitive programming tasks show that Seg-OPD-trained students achieve higher revision success rates than baselines. Seg-OPD consistently outperforms the compared state-of-the-art baselines in reasoning accuracy with an average relative improvement of 5.22% across diverse models and tasks. Code is available at this https URL.

[AI-109] st-time Calibration Learning for Large Language Model Reasoning

链接: https://arxiv.org/abs/2610.02695
作者: Zizhuo Zhang,Xiong Peng,Jingwei Sun,Rong Yao,Shixiong Kai,Mingxuan Yuan,Bo Han
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: 34 pages

点击查看摘要

Abstract:Reliable large language models (LLMs) must not only produce accurate answers but also express confidence that faithfully reflects their probability of being correct. Such calibration is essential for identifying uncertain predictions and supporting reliable decision-making in real-world deployment. Recent studies incorporate calibration learning into reinforcement learning (RL), jointly optimizing answer correctness and verbalized confidence using ground-truth correctness supervision. However, their reliance on labeled data limits their applicability in practical test-time settings, where ground-truth labels are unavailable and calibration may need to adapt to newly encountered target tasks. To address this challenge, we propose Test-Time Calibration Learning (TTCL), a label-free framework that jointly adapts reasoning accuracy and verbalized confidence directly on unlabeled target-task data. Specifically, TTCL derives self-supervision signals for both correctness and calibration from multiple model-generated responses, enabling calibration learning at test time without ground-truth labels. Theoretical analysis further establishes TTCL as a bounded surrogate for the ideal calibration objective. Extensive experiments on mathematical reasoning and factual question answering demonstrate that TTCL consistently improves both accuracy and calibration across diverse models and tasks. On base models, TTCL achieves an average relative accuracy improvement of +40.13% and an ECE reduction of +70.80% across eight benchmarks. Moreover, TTCL can further improve both accuracy and calibration for already calibrated models under domain shift, particularly when source-domain calibration transfers poorly to target tasks. In the math-to-factQA setting, TTCL achieves an average relative accuracy gain of +20.35% and reduces ECE by +53.83%. The source code is released at this https URL.

[AI-110] Decoupling Memory from Context: Structured Memory for Token-Efficient Test-Time Continual Learning NEURIPS

链接: https://arxiv.org/abs/2610.02687
作者: Yehya Farhat,Michael Desmond,Anastasios Kyrillidis
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: 14, 4, neurips workshop: TTCL

点击查看摘要

Abstract:Large language models (LLMs) are increasingly deployed in enterprise, scientific, and medical applications, where agents must incorporate domain-specific knowledge and adapt from experience. Context engineering offers a practical alternative to weight updates by improving model behavior through instructions, strategies, and evidence supplied at inference time. However, adapting context online typically requires a costly trial-and-error process, while queries are often processed independently, preventing useful experience from carrying forward. Memory systems address this limitation by retaining information across interactions, but approaches that continually append information to a shared context face increasing token costs, context-window limits, and performance degradation as the context expands. We introduce a unified formulation of context optimization and show that an agent memory system update can be interpreted as an optimization update procedure over the model’s context. This perspective attempts to provide a principled framework for studying memory design and its efficiency. We then propose GraphMemory, a lightweight graph-based memory that accumulates, refines, organizes, and connects reusable strategies. For each query, GraphMemory retrieves only the relevant subgraph, enabling online context adaptation without exposing the model to the entire memory. Under bounded retrieval, the amount of retrieved memory remains constant as the number of processed examples grows. Experiments show that GraphMemory achieves competitive downstream performance while using approximately 81-85% fewer memory-construction tokens than our baselines.

[AI-111] Spend Teacher Tokens Where They Matter: Success-Referenced On-Policy Distillation ICLR2027

链接: https://arxiv.org/abs/2610.02678
作者: Xiang Chen,Futao Su,Kong Wang,Jiayi Chen,TanLin Li
类目: Artificial Intelligence (cs.AI)
备注: 18 pages, 2 figures. Submitted to ICLR 2027

点击查看摘要

Abstract:On-policy distillation (OPD) combines student-generated rollouts with dense token-level supervision from a teacher, but providing such supervision for every rollout requires substantial teacher computation. We introduce Success-Referenced On-Policy Distillation (SR-OPD), which reduces this cost by selecting which prompts and rollouts receive teacher supervision. When the student produces both successful and failed rollouts for the same prompt, a successful rollout can serve as a natural reference for selecting failed rollouts. SR-OPD therefore focuses on such prompts and prioritizes failed rollouts whose hidden-state trajectories show sustained divergence from a successful reference, while accounting for estimated teacher-input cost. Across three teacher-student pairs and six mathematical reasoning benchmarks, SR-OPD uses only 3.46-5.02% of the teacher-input tokens required by Vanilla OPD in the one-pass setting while maintaining comparable reasoning performance. Under a controlled setting matched to 5% of Vanilla OPD’s teacher-input budget, further experiments support both key design choices: focusing supervision on prompts with both successful and failed rollouts, and using successful rollouts to guide failure selection. These results indicate that a student’s own successful behavior can serve as a useful reference for allocating teacher supervision under a fixed teacher-input budget.

[AI-112] A GHOST in Long-Horizon Agents : Governance Hazard from Overlooked Safety Constraints across Turns

链接: https://arxiv.org/abs/2610.02664
作者: XinPeng Shen,Lan Zhang,Yixiao Huang,Haoran Cheng,Jiewei Lai,Leilei Chen,Haoxiang Deng
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Long-horizon agents are now playing an increasingly significant role in assisting humans with complex problem-solving. However, it is exactly their extended interaction history that introduces an underexplored execution-safety concern. Under benign interaction conditions, an agent may execute an action that violates a safety constraint specified many turns earlier. We term this failure mode Governance Hazard from Overlooked Safety Constraints across Turns (GHOST), which may cause irreversible damage. Our experiments reveal that GHOST events are not isolated cases: this failure mode, occurring precisely under benign interaction conditions, yields an occurrence rate of 11.5% on GPT-5.5. Furthermore, we theoretically show that if the residual conditional violation hazard along each safe prefix is bounded below by a non-summable sequence, the execution enters the hazard region almost surely. Leveraging this theoretical insight, we further propose STAR-Guard, a two-layer defense coupling historical semantic safety constraint restoration with pre-execution audit. STAR-Guard restores applicable safety constraints to reduce unsafe proposals, while its deterministic audit layer prevents residual violations from reaching the environment. Consistent with this two-layer design, we observe no GHOST events in our experiments under the GPT-5.5 setup.

[AI-113] Distributed Learning with Selective State Space Models: Architecture-Aware Convergence Analysis

链接: https://arxiv.org/abs/2610.02659
作者: Adam Piaseczny,Md Kamran Chowdhury Shisher,Shiqiang Wang,Christopher G. Brinton
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Optimization and Control (math.OC)
备注:

点击查看摘要

Abstract:Modern state space models (SSMs), such as Mamba2, provide a compelling alternative to transformers by combining linear-time sequence modeling with recurrent state-space dynamics. However, the behavior of SSMs in distributed learning settings remains poorly understood. In particular, the existing standard federated learning methods are largely architecture-agnostic, and do not account for the stability, selectivity, and state-space parameterization that characterize modern selective SSMs. To address this, we derive architecture-aware gradient and smoothness bounds for single- and multi-layer selective SSMs, and convergence bounds for FedAvg and FedProx, characterizing how recurrent stability, input-dependent discretization, and state projection norms affect federated optimization. We then numerically validate the single-layer bounds on sequences generated by a teacher SSM, using a learner that follows the analyzed recurrence. We use this analysis to formulate expectations about the effects of local training and client heterogeneity, and examine these expectations by comparing nine federated learning algorithms on Mamba2 language modeling across six text domains. These experiments illustrate how SSM-specific bounds can provide a basis for interpreting the behavior of practical federated learning algorithms.

[AI-114] Batched Speech Decisions Without Decoding: Single-Token Supervision Lets a Frozen LLM Hear Beyond the Transcript ICASSP2027

链接: https://arxiv.org/abs/2610.02638
作者: Jie Jin,Ziyin Ma,Min Yin,Jinyu Chen,Haigang Song,Zhikun Pang,Xiaowen Zhang
类目: Artificial Intelligence (cs.AI)
备注: 5 pages, 2 figures, 3 tables. Submitted to ICASSP 2027. Code and weights: this https URL

点击查看摘要

Abstract:Full-duplex voice agents make many small, closed decisions, which current systems answer by slow autoregressive decoding. We propose DuplexJev, which feeds ASR-encoder hidden states through a small connector into a frozen LLM and reads each question as a single-token distribution over its options. Nothing is decoded, and an 8-GPU node answers 80 decisions about eight utterances in about 0.1 s. With a last-layer connector, spoken QA stays close to reading the transcript (90% vs. 91%). DuplexJev also hears the speaker: gender and emotion accuracy both reach 90% (from 55% and 28%) with a cross-attention connector, whose spoken QA drops by only 1 point (83% to 82%). We train decisions with cross-entropy on the read-out answer token, instead of the usual transcript distillation, whose teacher never hears the voice, and keep distillation for content. Encoders and LLMs are interchangeable; we release weights, training recipe, a batched-inference pipeline for full-duplex serving and a bilingual spoken-QA set.

[AI-115] Designing the Future of User Feedback for Generative AI

链接: https://arxiv.org/abs/2610.02631
作者: Alisa Frik,Julia Bernd,Amitis Karami,Mohammad Tahaei
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Post-deployment feedback from users can be a cost-effective, scalable, and representative means to monitor and improve generative AI systems and features. When implemented effectively, giving such feedback can increase users’ engagement with and trust in GenAI systems. Government regulations and industry guidelines call for post-deployment user engagement, but there is little guidance on designing mechanisms that are usable for consumers and provide actionable input for product teams. We conducted a multi-phase study as a collaboration between academic researchers and eBay. Our benchmark evaluation of current industry approaches identified common issues including lack of discoverability, unclear terminology, and inattention to user value. Based on these findings, we developed best-practice recommendations and designed and tested a prototype feedback-collection tool. The tool aimed to provide users with an efficient, flexible, and positive feedback-giving experience, and provide product teams with rich data on performance and potential problems in a usable format.

[AI-116] Lost in the Request: How Communication Variation Disrupts Retrieval and Action in Email Agents NEURIPS2026

链接: https://arxiv.org/abs/2610.02627
作者: Feng Chen,Ritam Dutt,Atnaz Taheri,Alex Williams
类目: Artificial Intelligence (cs.AI)
备注: Accepted by NeurIPS 2026 Workshop on Evaluation of Interactive Agents

点击查看摘要

Abstract:An email assistant should not complete less work simply because a user phrases the same request differently. Yet most benchmarks test each task with only one canonical request, leaving this form of robustness largely unmeasured. We test whether email assistants remain reliable when the requested information, available evidence, and expected outcome stay fixed, but the communication style or English variety changes. We construct validated variants along five communication-style axes and four rule-based dialect conditions, and evaluate them on three benchmarks: a retrieval-augmented generation (RAG) pipeline and two tool-using agents. Indirect requests reduce performance on all three benchmarks, while formal requests reduce performance on both agentic benchmarks. Examining the systems more closely shows that these failures have different causes. Verbose requests mainly hurt a lexical retriever by making the relevant email harder to find. By contrast, indirect and dialect variants remain harmful even when the relevant email is retrieved. In the agentic setting, indirect and formal requests mainly cause the agents to omit required actions, not to take more unsupported actions. These results show that a successful response is not enough to establish robustness: evaluations should vary how requests are expressed and separately measure whether agents complete the requested work.

[AI-117] CuBEs: Culturally-Situated Behavioral Evaluations and the Limitations of Culture-Blind LLM Judges

链接: https://arxiv.org/abs/2610.02622
作者: Hoda Ayad,Tanu Mitra,Abhishek Mukherji
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Evaluating the occurrence and triggers of large language model (LLM) behaviors - such as sycophancy, self-preference, or over-confidence - is critical for predicting real-world model deployment risks. However, existing situated behavioral evaluations typically ignore cultural context, limiting their generalizability across an increasingly global user base. To address this gap, we propose CuBEs - Culturally-situated Behavior Evaluations that probe for response patterns across diverse user cultures. We first extend an automated testing pipeline to inject cultural context into behavioral test scenarios and subsequent evaluation. We assess the cultural adaptability of this pipeline by building a human-labeled dataset that captures nuanced dimensions of behavior understanding across 12 distinct cultures. Our dataset reveals significant cross-cultural variations that one-size-fits all judgments fail to capture. Through evaluating 13 open- and closed-source LLMs, we find that introducing cultural situatedness in the evaluation scenario creates significant variation in the presence of a behavior. For example, while our baseline experiments testing for political bias capture localized Western political dimensions like the American conservative-progressive divide, non-Western culturally situated evaluations surface entirely different axes of bias such as religious and colonial political issues. Our findings demonstrate that standard, culturally-agnostic evaluations fail to capture these shifts, highlighting the necessity of culturally situated behavioral testing for global deployments.

[AI-118] me Series Forecasting Benchmarks Need Scenario-Grounded Stress Testing

链接: https://arxiv.org/abs/2610.02608
作者: Yuyang Zhao,Lian Xu,Hao Xue
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Time series forecasting (TSF) increasingly drives decisions in transportation, energy, finance, healthcare, and infrastructure, yet current evaluation remains overly narrow: standard benchmarks reward low held-out error, while robustness studies typically reduce failure to Gaussian noise, random masking, or bounded adversarial perturbations. This obscures the real failure modes of deployed forecasting systems. Input-side anomalies are not merely noisier inputs: they often reflect structured events that alter temporal dynamics, break cross-variable dependencies, induce regime shifts, or propagate from faulty sensors to downstream decisions. These semantic, causal, and system-level failures cannot be faithfully captured by i.i.d. perturbations alone. The rise of TSF foundation models makes this evaluation gap more urgent, as unauditable pretraining corpora make held-out generalization increasingly unreliable. We therefore advocate scenario-grounded stress testing. Each test instance should include historical inputs and future targets, together with a semantic scenario, an explicit failure operator, and a measurable difficulty level. This shift makes evaluation interpretable, attributable, and deployment-relevant and friendly, enabling the community to ask not only which model is accurate, but under what conditions it fails and why.

[AI-119] asteBench: Multimodal Benchmark for Sensory Prediction from Molecules to Sustainable Foods NEURIPS2026

链接: https://arxiv.org/abs/2610.02599
作者: Anna T. Thomas,Sohum Patnaik,Caroline Cotto,Benjamin Sanchez-Lengeling
类目: Artificial Intelligence (cs.AI)
备注: First two authors contributed equally. Accepted to NeurIPS 2026, Evaluations Datasets track. Code available at this https URL

点击查看摘要

Abstract:Sustainable protein discovery lacks the fast computational proxies, analogous to molecular docking or density functional theory, that accelerate drug and materials discovery. Evaluating whether a novel food tastes like its animal-based target requires expensive human sensory panels, bottlenecking the design-build-test loop. We introduce TasteBench, a multimodal benchmark and privacy-preserving competition for sensory prediction, spanning two tasks: a food-level ranking task built on 21K+ human evaluations across 215 plant-based foods in 24 product categories, yielding 935 within-category ranking pairs, and a supporting molecular-level taste classification task over 15K flavor molecules. To enable rigorous interpretation of model performance, we characterize the ground truth: inter-rater agreement among panelists is low (Krippendorff’s \alpha = .077 ), and the split-half reliability ceiling of panel-aggregated rankings is .825, establishing the range within which ML systems on this benchmark should be assessed. We evaluate baselines across four input modalities; on the same pairs panelists rated, the best model achieves .661 pairwise accuracy, competitive with the median individual panelist (.650), and .683 across all within-category pairs. TasteBench provides the evaluation infrastructure and baselines for measuring progress on computational screening for sustainable protein discovery.

[AI-120] Open-Endedness Bench: Measuring Epistemic Process from Agent Records

链接: https://arxiv.org/abs/2610.02588
作者: Chengyang Shi,Xianglin Ji,Jintao Huang,Jicheng Wang,Yifeng He,Jiachen Liu
类目: Artificial Intelligence (cs.AI)
备注: 18 pages, 7 figures. Code: this https URL . Data: this https URL

点击查看摘要

Abstract:Agents are increasingly given open-ended research tasks: discovering an empirical law from self-designed experiments, improving a heuristic whose optimum nobody knows, or beating a standing record. Their execution logs record every step of this research, yet the runs are still judged by their outcome score. That score alone does not establish whether an agent’s claims follow from executed experiments, and a reference answer may be unavailable. We evaluate the agent’s epistemic process: how it forms hypotheses, tests them, and revises them in response to evidence. We introduce OEB (Open-Endedness Bench), a benchmark-agnostic methodology that reads only the agent’s execution record and never a reference answer or an outcome score. OEB compiles the record into a unified epistemic event graph whose edges connect the propositions the agent states to the executed actions that test them; each node carries an exact excerpt that code verifies against the record. One principle governs scoring: prose can state a proposition, but only evidence returned by an executed action can support or refute it, so OEB checks what the agent writes against what it actually ran. From the graph, OEB scores four competence axes (evidence, experiment, revision, and no reward hacking), mostly as the share of opportunities for sound research that the agent took, and profiles six subjective persona traits that describe the agent’s research habits. We score 119 existing runs over 12 tasks from three benchmarks: LLM post-training, chip design, and a training-speed record. Against logged results, only 16-29% of the improvements agents claim are real. On 9 of 10 tasks, the best run tries more new ideas in its second half than the worst run. The persona readings follow the model: for every trait, the model that ran explains more of its variance across runs than the task (a median of 43% against 7%).

[AI-121] Labels Override Definitions in Jev-Style Typed Decision Models

链接: https://arxiv.org/abs/2610.02586
作者: Seyedarmin Azizi,Erfan Baghaei Potraghloo,Massoud Pedram
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:A typed decision model answers a fixed question about an input by returning a probability for each of several caller-defined options. Each option carries a short label and a written definition, which is where a developer states the rule the model should apply. Jev introduced this interface for routing, moderation and triage, open implementations followed, and the same operation occurs whenever a language model is used as a classifier by scoring label strings. We study the open implementations, whose weights we can inspect and patch, and ask whether the probability follows the definitions or the labels. A preference for the label we call option-label bias. Across four open-weight typed decision models, three ways of reading an answer from a Qwen2.5 backbone, eleven classification tasks and PolicyBench, a synthetic routing suite we introduce in which the rule appears only in the definitions, the answer is mostly the labels. Deleting every definition leaves accuracy unchanged (laya-td: 0.8559 against 0.8487), although those definitions support 0.7971 on their own, and renaming the options to A and B raises accuracy by +0.1511 [+0.1377, +0.1646]. One system, von, is unaffected, and the two code bases differ in one expression: laya writes each option as “label: definition”, while von writes only the definition. Changing that expression in both directions, with no weight changed, makes all three laya checkpoints exactly invariant (+0.0000 [+0.0000, +0.0000]) and creates the effect in von, whose accuracy falls from 0.8511 to 0.2281 when a label contradicts its definition. Earlier work attributed this failure to the constrained decision head these models use in place of a text decoder; our results locate it in the prompt rendering. We give a two-call test that tells a practitioner which case applies to their model, and measure what four mitigations are worth.

[AI-122] Answering clinicians questions over trial evidence tables with verifiable feedback-driven language models

链接: https://arxiv.org/abs/2610.02576
作者: Manan Roy Choudhury,Suparno Roy Chowdhury,Swastik Sahoo,Muhammad Ali Khan,Kaneez Zahra Rubab Khakwani,Mohamad Bassam Sonbol,Irbaz Bin Riaz,Vivek Gupta
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Systematic reviews condense clinical trials into evidence tables, yet clinicians can interrogate these tables only through database queries, and many questions concern attributes that the table does not record, such as a drug’s target class or a harmonised endpoint. Here we introduce FD-SCoPE, a language-model framework that answers both kinds of question, exposes the query, the selected trials and the derivation rule behind every answer, and learns from expert corrections. On an oncology evidence table of 159 immune checkpoint inhibitor trial records, FD-SCoPE completed all 140 clinician-style tasks (alternatives, 90.7-97.9%). For questions needing derived attributes it retrieved 99.3% of relevant trial records at a positive predictive value of 89.8% and outperformed four alternative approaches (derived-value F1 77.7% versus 64.8-73.4%). Corrections on 299 questions, simulated from reference answers, raised F1 on 1,201 unseen questions from 77.9% to 84.9%. Language models coupled with executable queries, verified programs and expert feedback can give clinicians auditable access to trial evidence.

[AI-123] Improving the Energy-Efficiency of the Code Generated by LLM s through Effective Prompting

链接: https://arxiv.org/abs/2610.02571
作者: Ritika Rekhi,Bing Zhang,Md Arman Islam,Jaya Krishna Pasham,Yeswanth Chitturi,Akshay Paramesha,Isha Valiveti,Asif Imran,Bekir Turkkan,Tevfik Kosar
类目: oftware Engineering (cs.SE); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:As AI-assisted programming becomes increasingly mainstream, the environmental impact of AI-generated software has emerged as an important consideration. This motivates evaluating LLM-generated code beyond functional correctness by considering execution efficiency and energy consumption. However, despite substantial advances in code generation, frontier LLMs are rarely evaluated based on the energy efficiency of the code they produce. In this work, we conduct a comprehensive evaluation of 21 prompting strategies for energy-efficient code generation and identify 8 strategies for evaluation across 10 widely used open-weight and proprietary LLMs. We evaluate their effectiveness for both Python and C++ code generation relative to a baseline prompt. Across the evaluated models, the selected prompting strategies achieved energy reductions of up to 25% for Python and 17% for C++ code generation. At the model level, Python energy reductions reached up to 50% for Granite-4.0-H-Small, 39% for Claude 4.5 Haiku, and 28% for MiniMax M3, while C++ reductions reached up to 56% for Granite-4.0-H-Small and 7% for Qwen3-Coder-480B-A35B-Instruct. These results demonstrate that prompting strategies can substantially influence the energy consumption of LLM-generated code, although their effectiveness varies across models and programming languages. Our findings highlight the importance of incorporating energy efficiency into the evaluation and optimization of LLM-based code generation and provide practical insights into designing prompts for more sustainable AI-assisted programming.

[AI-124] Pincer: Resource Authorization for Agents using a Digital Twin

链接: https://arxiv.org/abs/2610.02569
作者: Mayank Rathee,Alexander Stepanov,Shalin Madabhavi,Jinhao Zhu,Raluca Ada Popa,Ion Stoica
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Coding agents have become increasingly long-horizon, autonomous, reliant on general-purpose shell and maintain their own persistent memory for self-improvement. While these capabilities have made the agents powerful, they have also made them harder to defend against external adversaries. Defenses that restrict this architecture — typed tools, information-flow control, or policy prediction engines — give up too much functionality to be adopted. Agents deployed today (e.g. Claude, Codex) rely on a combination of user-mediated and automode sandboxing as their primary defense. In user-mediated sandboxing, user-maintained policies decay over time and repeated permission requests cause user fatigue, while auto mode’s tool-call classifiers learn no user-specific policy and are not meant to defend against adversarial setups. Pincer is a new defense that operates at the resource layer and works alongside existing defenses at the tool-call layer like the auto mode. At the core of Pincer lies a digital twin, an isolated-context model that automatically learns and enforces dynamic user-specific least-privilege policies. The digital twin keeps continually learning the user’s preferences allowing it to act as the user’s proxy for the agent’s permission requests. To emulate the learning phase, we propose a new usercentric dataset with examples following a multi-day transcript of user-agent interaction. Our evaluation shows that Pincer performs strongly on both security and utility in comparison to several baselines which includes variants of LLM judges and adaptations of Conseca (HotOS '25). We highlight attack types where Pincer’s design leads to a significant security improvement compared to all other baselines, while outperforming the baselines even for other types of attacks.

[AI-125] Mitigating Social Sycophancy via Pluralistic Preference Optimization

链接: https://arxiv.org/abs/2610.02568
作者: Stephane Hatgis-Kessell,Myra Cheng,Xiaoxuan Hou,Qian Hu,Rahul Gupta,Natasha Jaques,Emma Brunskill
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Personal advice, including relationship advice, now ranks among the most common uses of generative AI. But language models (LMs) exhibit sycophancy: they affirm users much more often than humans do, which can make people overconfident and less willing to repair their relationships after a conflict. Prior work on mitigating sycophancy has focused on factual settings where a response can be checked against a ground truth answer, while mitigations for social sycophancy (e.g., personal advice, where there is no ground truth) have relied on simple prompting and post-training methods with limited effectiveness. Our insight is that social sycophancy occurs in part because LMs overly center on the user and fail to consider the perspectives of other stakeholders impacted by the user’s behavior. To address this problem we propose Pluralistic Preference Optimization (PlurPO): given inputs describing interpersonal conflicts, the LM identifies and simulates the relevant stakeholders, and is then trained to prefer and generate responses acceptable to all stakeholders. PlurPO uses only signals the model produces about its own outputs, without ground-truth labels. PlurPO substantially reduces social sycophancy across four datasets and four model families compared to prior methods. For example, on statements of intent to cause harm, where the users’ actions should not be endorsed, PlurPO reduces the endorsement rate by 89% on average across four models. On general advice questions, where the target is to match the endorsement rate of human responses, it closes the gap by more than half, from 17.8% to 8.0% on average. The preference dataset constructed by PlurPO for an 8B model also effectively transfers to mitigating sycophancy in a larger (32B) model. Our results indicate that social sycophancy can be reduced by leveraging a model’s own capabilities to simulate a plurality of relevant perspectives.

[AI-126] OpenGameEval: Benchmarking Agent ic Programming and Exploration in a Stateful Game Engine NEURIPS2026

链接: https://arxiv.org/abs/2610.02563
作者: Eray Turkel,Mengsha Sun,Kartik Ayyar,Sean Dunigan,Jack Lu,Vlad Shcherban,Hsiang-Shun Shih,Xin Wang,Tiantian Zhang
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: A shorter version appears at the NeurIPS 2026 Workshop on Evaluation of Interactive Agents

点击查看摘要

Abstract:We present OpenGameEval, a benchmark and evaluation framework for agentic game development inside Roblox Studio. It runs language models as agents in reproducible, stateful game-engine sessions and scores each run with executable checks, both on the edited scene and in a simulated play session. Most agentic coding benchmarks require exploration but score only final task success. OpenGameEval separates observation tools from editing tools in its eight-tool action space, so exploration can be measured directly. We measure the pass rates and exploration behavior of 13 frontier models on 84 human-curated core tasks, with 16 attempts per task. The tasks are hard for current models. The best model solves 51.7% of tasks on a single attempt and 39.4% five times out of five, and no tested model solves six of the tasks. Models at the frontier reach similar pass rates by solving different tasks: splitting tasks by the kind of work they require spreads the top five by 5.0pp on script-authoring tasks and 12.5pp on scene-change tasks. Exploration behavior predicts whether a run succeeds. Holding task and model fixed, a run that inspects every object a reference solution touches before acting on it passes 13.4pp more often than a run that inspects none of them on scene-only tasks, and 9.8pp more often on script-only tasks. We release the task suite, its place files, the per-task annotations, a plugin that runs the tasks inside Roblox Studio, and an updated leaderboard under the MIT license at this https URL. Comments: A shorter version appears at the NeurIPS 2026 Workshop on Evaluation of Interactive Agents Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI) Cite as: arXiv:2610.02563 [cs.LG] (or arXiv:2610.02563v1 [cs.LG] for this version) https://doi.org/10.48550/arXiv.2610.02563 Focus to learn more arXiv-issued DOI via DataCite (pending registration) Submission history From: Eray Turkel [view email] [v1] Thu, 1 Oct 2026 22:56:10 UTC (1,875 KB)

[AI-127] How to Have a Sensitive Debate: An Instance-Optimal Protocol for AI Debate

链接: https://arxiv.org/abs/2610.02557
作者: Jiawei Li,Zhiyang Xun,Lijie Chen,Jonah Brown-Cohen
类目: Artificial Intelligence (cs.AI); Computational Complexity (cs.CC); Computer Science and Game Theory (cs.GT); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:As powerful AI systems reach and sometimes surpass the abilities of human experts across a range of cognitively demanding tasks, the problem of accurate oversight and supervision of these systems has become increasingly urgent. One promising approach is AI debate, which seeks to leverage a debate between two powerful AIs to break complex questions down into simpler claims that can be easily judged directly. Theoretical work on debate has formalized this intuition in the language of computational complexity theory, where the goal is to design protocols (i.e., rules of the debate game) that provide rigorous guarantees on correctness for judging solutions to complex problems with limited supervision. Specifically, the current best protocol has been shown to work for all problems that have sufficiently stable decompositions into subproblems. In this paper, we design a new protocol for this same class of problems that improves on the prior work in several ways. First, correctness holds in a worst-case rather than an average-case sense. Second, being honest and correct is a dominant-strategy equilibrium for both debaters, rather than a Stackelberg equilibrium. Finally, we prove black-box lower bounds, showing that our new protocol is instance-wise optimal. That is, no protocol for this class of problems can outperform ours while making only black-box queries to human judgments. We obtain these results by relating the notion of stable problem decompositions to the concept of fractional block sensitivity from query complexity.

[AI-128] Out of Sync Out of Sight: Phantom State Attacks against IIoT Intrusion Detection

链接: https://arxiv.org/abs/2610.02552
作者: Sabrine Ennaji,Elhadj Benkhelifa,Nadia Kabachi
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Machine learning-based intrusion detection systems (IDS) are critical for securing Industrial Internet of Things (IIoT) environments. Most adversarial research against them perturbs the feature vector or the traffic that produces it, and depends on gradient access, repeated model queries, or a learned model of benign traffic. A smaller line of work reshapes packet timing without querying the detector, but makes malicious traffic mimic a learned model of benign timing. Across these approaches, one assumption of industrial monitoring pipelines has received little attention: temporal synchronization. An IDS reconstructs operational state by aggregating telemetry into sliding or tumbling windows, so its view depends not only on what is observed but on when each observation falls relative to a window boundary. We introduce the Phantom State Attack (PSA), which exploits that dependence under a passive, zero-query threat model. Rather than modifying packets, perturbing features, querying the classifier, or fitting any model of benign traffic, PSA injects bounded timing drift calibrated to the attack flow’s own inter-arrival variability, moving observations across the nearest window boundary by the minimal shift needed. The IDS then reconstructs a phantom state that diverges from the true process state. We evaluate PSA on ToN-IoT and CIC IIoT 2025 (DataSense), against Random Forest, MLP and XGBoost, measuring detection degradation, synchronization distortion, stealth, and attacker cost. PSA degrades detection on flows carrying enough packets for window-boundary redistribution, and leaves others almost unchanged, so its effect is conditional. A query-based baseline reaches higher raw success but needs many queries per window, while PSA needs none. The results identify temporal aggregation as an attack surface reachable under weaker assumptions than prior evasion techniques. Subjects: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI) Cite as: arXiv:2610.02552 [cs.CR] (or arXiv:2610.02552v1 [cs.CR] for this version) https://doi.org/10.48550/arXiv.2610.02552 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[AI-129] How To Train Your World Model: Fine-tuning vs RAG for LM-based World Modeling

链接: https://arxiv.org/abs/2610.02542
作者: Dhananjay Ashok,Shantanu Agarwal,Vivek Datla,Jonathan May,Alfy Samuel
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:World models (WMs) simulate the transition dynamics of environments, enabling agents to plan over the consequences of their actions. In text-based environments, fine-tuning a Language Model (LM) to serve as a WM has emerged as a dominant paradigm. However, despite the widespread success of non-parametric approaches such as Retrieval Augmented Generation (RAG), retrieval for LM-based world modelling remains underexplored. We conduct a systematic evaluation across five diverse environments spanning embodied, web navigation and social settings, comparing fine-tuning and RAG-based approaches for LM-based world modelling. Our study reveals that fine-tuning often outperforms RAG, with fine-tuned WMs enabling agents to obtain higher rewards on 15/20 settings. While both construction paradigms benefit from additional and more diverse exploration, RAG-based approaches prove more data-efficient, and fine-tuning approaches disproportionately benefit from scaling the amount of experience collected. With a focus on RAG-based WMs, we devise a procedure that uses counterfactual intervention to estimate the error rate of the retrieval stage, and show that retrievers consistently surface suboptimal transitions from the experience buffer. Hoping to address this failing, we study a variety of query reformulation strategies, demonstrating that a hierarchical approach outperforms the traditional retrieval pipeline. Finally, we compose our findings into a hybrid world modelling system that parametrically captures core environment dynamics, while learning to rely on retrieval from an actively maintained memory store. Our hybrid system consistently outperforms other methods across multiple environments and models, showcasing the robustness of the approach and the applicability of our findings.

[AI-130] CriticHack: Evaluating Visual Rewards Under Robot Policy Optimization

链接: https://arxiv.org/abs/2610.02527
作者: Jiaxuan Luo,Xingguo Xu,Shanshan Wang,Yuhan Zhou,Zhen Zhang
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: 60 pages

点击查看摘要

Abstract:Learned visual reward models are increasingly used to optimize robot policies, yet a reward model can score an execution that acts on the wrong object as highly as one that completes the task. We show that optimizing such a reward can amplify these wrong-object failures while reward and task success both rise, so the signals a practitioner would normally monitor look healthy. We fine-tune every denoiser parameter of a diffusion policy against Robometer on a drawer task. Starting from a supervised policy with no prior reward exposure, five training runs raise task success by 10.2 percentage points and wrong-object failures by 10.9 points on 512 evaluation seeds, whereas five runs trained on the simulator’s task-completion signal raise success without amplifying wrong-object failures (difference 9.2 points, 95% CI 5.6 to 13.0). The amplification recurs from a policy previously optimized against learned rewards, under the policy’s native diffusion sampler, at matched distance from the initial policy, and across constrained-policy experiments with two critics and two optimizers. A tilt model explains when it occurs: under KL-regularized optimization, an outcome becomes more frequent whenever its expected reward under the initial policy exceeds the population average. Robometer separates successes from failures well overall (AUROC .81) but scores wrong-object failures slightly above successes (AUROC .37), so optimization raises both. The same model predicts the outcome shifts across 26 constrained settings (Spearman .89), including those in which task success falls, and Robometer’s own published success-termination recipe inherits the error. A frozen outcome verifier redirects the same optimization toward the requested task.

[AI-131] Learning What to Investigate Next: Meta-Reasoning for Long-Horizon Research Agents

链接: https://arxiv.org/abs/2610.02525
作者: Ankur Samanta,Yonathan Efroni,Paul Sajda,Kaveh Hassani,Anirudh Goyal
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Long-horizon research agents must decide both how to investigate and what to investigate next as evidence accumulates. This is hard to learn because such decisions are sparse in long execution traces, and their consequences may emerge several investigations later. We introduce Meta-reasoning for Iterative Research Agents (MIRA), a hierarchical architecture separating research allocation from execution. An outer-loop meta-reasoner curates context from a persistent research record, then writes a work order for the next investigation or ends the episode. A fresh inner-loop executor carries out each work order, making execution part of the transition between meta-reasoning actions. Without policy training, MIRA improves long-horizon inference and allocates additional compute more effectively in theorem proving and open-ended neural-architecture research. Its decision boundaries also provide natural units for credit assignment. At each boundary, we train a generative critic to forecast expected remaining return from partial states, outperforming token-level alternatives. Cross-environment pretraining improves forecasting and adaptation, yielding a transferable prior for valuing partial progress. We use this prior to initialize MIRA-AC, a generative actor-critic jointly trained to forecast remaining return and choose the next investigation, without a separate critic model. MIRA-AC concentrates policy optimization on meta-reasoning decisions, enabling efficient long-horizon reinforcement learning without directly optimizing the longer execution traces they initiate. Training MIRA-AC on the model’s own proxy hill-climbing signals improves gold performance across four autoresearch environments; the actor transfers with cross-environment value initialization. Together, these results show that meta-reasoning can be learned as an explicit policy for directing long-horizon autonomous research.

[AI-132] Hypothesis-guided discovery of cognitive algorithms via program refinement

链接: https://arxiv.org/abs/2610.02523
作者: Huiwen Alex Yang,Mark K. Ho,Bill D. Thompson
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Developing cognitive models of algorithmic reasoning from behavioral data is a central problem in cognitive science that challenges current methods. Traditional approaches to cognitive modeling are interpretable and benefit from human expertise, but lack flexibility and scalability. Emerging techniques using large language models (LLMs) for de novo generation of cognitive models are scalable and flexible, but lack a role for human expertise and have mostly been applied to simpler tasks than algorithm recovery. We propose a hybrid system that treats discovery of cognitive algorithms as a program refinement problem. Human-created cognitive models are expressed as probabilistic programs and provided to a system of LLM agents with a mandate to: identify mismatches between model and behavior; propose code-level modifications within researcher-specified constraints; and verify structural fidelity. Revisions propagate to a probabilistic inference module that performs inference for latent variables and data likelihood computations. We evaluate the pipeline on human behavior in a problem-solving paradigm that exposes a variety of cognitive algorithms. Revised models consistently improve model fit relative to ancestral models and reveal a small set of recurring innovations that capture meaningful behavioral variability in this task.

[AI-133] Instance-Dependent Regret for CMDPs with Step-Wise Constraints

链接: https://arxiv.org/abs/2610.02520
作者: Qian Zuo,Francesco Emanuele Stradi,Leyang Xue,Sattar Vakili
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:We study online learning in episodic tabular constrained Markov decision processes with step-wise safety constraints. In such a setting, the constraints induce a safe subgraph that shapes the variance of cumulative rewards under feasible policies and, consequently, the difficulty of learning. Exploiting this structure, however, requires learning which actions are safe while controlling constraint violations. We propose Safe Variance-Adaptive Exploration (SVAE), an efficient algorithm that learns candidate safe subgraphs and performs variance-adaptive optimistic planning within them. With high probability, SVAE achieves cumulative regret of order \widetilde\mathcalO(\sqrtSAH\min\mathbbV_\Sigma,K\mathrmVar^\star+S\sqrtAH^3\min\K,\mathcalC+S^2AH^2) over K episodes, where H is the horizon of a single episode, while S and A are the numbers of states and actions, respectively. Here, \mathrmVar^\star is the maximum return variance among safe policies, \mathbbV_\Sigma is the variance accumulated before the first unsafe action is encountered, and \mathcalC captures the statistical complexity of eliminating actions incorrectly considered potentially safe. SVAE additionally attains \widetilde\mathcalO(H\sqrtSAK+S^2AH^2) step-wise constraint violation and a gap-dependent violation bound that is polylogarithmic in K . Finally, we establish a lower bound showing that dependence on these instance-specific quantities is unavoidable.

[AI-134] Student-Guided Teacher Distillation for Efficient LLM Task Routing: Positioning Against Jev-Style System-1 Classifiers

链接: https://arxiv.org/abs/2610.02516
作者: Haifeng Wu,Srinivasan Manoharan,Jian Wan,Fangbo Tu,Junhua Zhao,Xin Chen
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: 13 pages, 2 figures, 1 table

点击查看摘要

Abstract:Zero-shot classifiers are useful for routing user requests to specialized LLM tasks, but scoring every request against a large candidate set is expensive: a zero-shot NLI classifier must evaluate one premise-hypothesis pair per label, so cost scales linearly with taxonomy size. We study a student-guided teacher distillation pipeline for a fixed taxonomy of 60 LLM task categories: a compact ModernBERT classifier predicts the full category distribution in one forward pass and retrieves a small top-k candidate set, and a larger DeBERTa-v3 zero-shot NLI classifier reranks only those candidates rather than all 60 labels; the resulting teacher labels iteratively improve the student, which produces sharper candidates for the next round. Unlike generic embedding retrieval or clustering-derived shortlists used in extreme multi-label classification, our candidate generator is trained end-to-end on the target taxonomy and is the same model serving production traffic, distinguishing it from LLM-routing work that routes between candidate models, and from concurrent System-1 encoder-classifier proposals (e.g. TypeSafe AI’s Jev and the open-source Laya project) whose training methodology is undocumented or RL-based. Our best student checkpoint reaches 77.5% teacher agreement on a 200-example evaluation set, and preliminary coverage measurements show Coverage@16 of 91-100%, suggesting top-k sets retain most of the teacher’s decision-relevant information. We further show truncated top-k teacher scores should not be treated as full 60-class soft targets for KL distillation: zeroing untruncated classes destroys the dark knowledge soft-label distillation depends on, introducing systematic bias rather than a harmless sparse approximation. A complete evaluation, including coverage at multiple k on a held-out set, an embedding-retrieval baseline, and a larger human-reviewed test set, remains in progress.

[AI-135] Multi-Fidelity Policy Gradients Stabilize Data-Scarce Reinforcement Learning

链接: https://arxiv.org/abs/2610.02505
作者: Xinjie Liu,Ruihan Zhao,Anirban Chaudhuri,Cyrus Neary,Ufuk Topcu,David Fridovich-Keil
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Robotics (cs.RO)
备注:

点击查看摘要

Abstract:Policy gradient methods for on-policy reinforcement learning (RL) can become unstable when expensive, scarce target-domain data yield noisy gradient estimates. We address this challenge by complementing limited high-fidelity (HF) target-domain data with abundant, cheap, but biased low-fidelity (LF) data, e.g., from a simplified simulator. Most existing methods directly optimize biased objectives based on LF data. In contrast, the recently introduced multi-fidelity policy gradient (MFPG) framework uses LF data solely to construct a control variate that reduces variance and improves HF data efficiency without biasing the policy gradient estimator. However, published work on MFPG is limited to REINFORCE on small-scale simulation tasks. We develop MFPG for modern actor-critic learning in GPU-parallel simulation and on a physical robot. Our analysis and experiments show that naive extensions to proximal policy optimization (PPO) can lose cross-fidelity correlation or inflate variance. Our MFPG-PPO addresses these failures by redesigning the sampling, advantage estimation, and control variate construction to preserve cross-fidelity correlation, and by monitoring estimator uncertainty to prevent variance inflation. We also introduce a budget-aware MFPG-PPO to divide a fixed sampling budget among high- and low-fidelity data sources. Across simulated robot locomotion tasks of varying LF-to-HF transfer difficulty and HF data budgets, MFPG-PPO improves upon PPO trained on HF data alone in nearly all settings, and consistently matches the performance of PPO trained with 16x more HF data on the hardest task at the smallest HF budgets. In contrast, most baselines that use LF data perform well only where direct LF-to-HF transfer succeeds. MFPG-PPO enables stable learning on a physical Franka arm using only 4 real-robot episodes per update and no human demonstrations.

[AI-136] HXAI: Hierarchical Privacy-Preserving Explainable AI in Distributed Energy Systems

链接: https://arxiv.org/abs/2610.02504
作者: Poushali Sengupta,Sabita Maharjan,Frank Eliassen,Yan Zhang
类目: Artificial Intelligence (cs.AI)
备注: Under Review

点击查看摘要

Abstract:Balancing electricity demand and supply is increasingly difficult due to the inherent intermittency of renewable power generation and the stochastic power consumption. Grid operators require fine-grained, decision-relevant insights into household energy consumption to manage peak loads and design responsive tariffs, but increased transparency at this level raises significant privacy concerns. Traditional methods for explainable AI (XAI) can reveal sensitive information, while standard privacy techniques often reduce the usefulness of explanations. To address this issue, we introduce HXAI, a hierarchical framework that preserves privacy while enabling reasonable explainable analysis for grid-level demand management. HXAI consists of two main components: (1) a local model that generates fine-grained explanations within a secure, private environment, and (2) a zonal model that aggregates these explanations to support grid-level analysis while enforcing privacy through flexible privacy-budget management. We explicitly limit cumulative privacy exposure under repeated operator queries and show that the proposed framework preserves decision-relevant information without compromising household privacy. Experiments on both simulated and real-world energy datasets demonstrate that HXAI provides useful insights for zonal load management while ensuring that appliance-level consumption remains local and is never transmitted to grid operators. Our results show that preserving the semantic structure of explanations, rather than minimizing numerical error, is the key to XAI under differential privacy. This framework provides a way to achieve both privacy and explainability in energy management.

[AI-137] Compound AI System Reliability: A Failure Taxonomy and Resilience Pattern Catalog from 150 Production Incidents ICML2026

链接: https://arxiv.org/abs/2610.02503
作者: Rudrendu Kumar Paul,Sourav Nandy
类目: oftware Engineering (cs.SE); Artificial Intelligence (cs.AI); Distributed, Parallel, and Cluster Computing (cs.DC); Machine Learning (cs.LG)
备注: Accepted at the AIWILD Workshop, ICML 2026. Camera-ready version

点击查看摘要

Abstract:Deploying compound AI systems reliably and safely requires understanding failure modes that emerge at component boundaries, not within individual models. Cascading errors propagate across component boundaries, silent quality degradation evades standard monitoring, and coordination failures yield incorrect collective behavior from individually correct parts. We analyze 150 production incident reports from open-source compound AI projects and anonymized enterprise deployments to construct a taxonomy of 23 failure modes organized into five categories: retrieval failures, generation failures, tool failures, orchestration failures, and integration failures. For each category, we propose resilience patterns with measured effectiveness from controlled fault injection experiments. Circuit breakers reduce cascade propagation by 89%, output quality gates catch 73% of silent degradation before user impact, and component isolation reduces blast radius by 64%. Systems implementing three or more resilience patterns from our catalog reduce mean-time-to-recovery (MTTR) by 71% compared to unstructured monitoring baselines. We release the incident taxonomy and pattern catalog as a practitioner resource.

[AI-138] “I just assumed that it would translate”: examining MT risk awareness among healthcare staff with abbreviations as a use case

链接: https://arxiv.org/abs/2610.02496
作者: Eleanor Taylor-Stilgoe,Félix do Carmo,Constantin Orăsan
类目: Artificial Intelligence (cs.AI)
备注: To appear in the proceedings of Convergence 2026 conference

点击查看摘要

Abstract:In the UK, public healthcare staff report turning to machine translation (MT) - predominantly Google Translate (GT) - to communicate with patients across language barriers. Though intended to support their duty of care, potentially uninformed reliance on MT in such contexts could have serious consequences for patient safety. Research nonetheless remains limited on staff awareness of the possible risks posed by higher-stakes MT use in general and with patient medical records in particular, most existing literature instead examining its use in interpersonal situations or with patient-oriented documentation. Moreover, medical abbreviations are well-documented as increasing patient risk even monolingually, with outcomes from their misuse and/or misinterpretation ranging from temporary harm to the death of the patient. Abbreviations were therefore selected as a use case for identifying the potential risks posed by their translation with MT. Contextualised French and Spanish data examples drawn from authoritative clinical corpora and translated via GT were presented during semi-structured interviews to 21 healthcare staff participants in diverse roles and specialties. The results were then subject to qualitative analysis and cross-analysis.

[AI-139] What Does a Token Cost? A Mixture-of-Agents Measurement of Sufficient Per-Token Compute

链接: https://arxiv.org/abs/2610.02491
作者: Zhixu Du,Weijia Han,Hai Helen Li,Yiran Chen
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Large language models spend the same amount of computation on every token they generate, regardless of how difficult each token is to produce. Methods such as speculative decoding and model routing are built on the premise that much of this computation is unnecessary, yet the computation an individual token actually requires has not been measured. We measure it through a Mixture-of-Agents (MoA) lens: a panel of fifteen language models of increasing capacity, drawn from three families, in which every agent attempts to reproduce a reference sequence token by token, conditioned on the correct preceding tokens. We define the inference cost of the smallest agent that succeeds as the token’s sufficient compute, which upper-bounds what the token requires. On three core benchmarks, a 0.5B agent reproduces 92–95% of reference tokens. Across Qwen, OLMo, and R1-distilled panels, the most expensive 10% account for 64–80% of estimated FLOPs. On all 500 MATH-500 problems, the MoA-derived map helps model routing reduce projected latency from 7.59 to 5.12 seconds while slightly improving accuracy, relative to the best confidence-routing baseline. The MoA-map helps drafting use 32.6% fewer draft tokens and approximately 20% lower projected latency than fixed-window drafting at similar accuracy. These comparisons reveal remaining allocation headroom, motivating controllers that exploit sufficient-compute structure.

[AI-140] MEA: A Reward-Driven Multi-Agent System for Faithful Model Explanations

链接: https://arxiv.org/abs/2610.02480
作者: Yuyang Cheng,Raghav Kaushik Ravi,Srivarshinee Sridhar,Sriparna Saha,Akash Ghosh,Chirag Agarwal
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Recent years have seen the employment of a plethora of machine learning (ML) models in high-stakes domains, but they remain largely opaque to the practitioners who act on their predictions. While post-hoc explanation methods offer a lens into this model behavior, wielding them effectively demands expertise most domain experts lack: navigating high-dimensional outputs, selecting the best explanations, and synthesizing evidence across disparate tools. To this end, we present MEA, a multi-agent framework that removes the explanation knowledge barrier entirely: a Proposer agent selects and configures explanation tools based on the question and modality, while an Actor agent is optimized end-to-end against faithfulness, transforming the outputs into natural language explanations grounded in model behavior across tabular, text, and vision modalities. Further, we introduce diverse question types spanning feature attribution, counterfactual reasoning, and spurious feature detection, each paired with a perturbation-based faithfulness metric. We find that frontier LLMs systematically produce unfaithful explanations. By optimizing against faithfulness rewards augmented with a modality-adaptive penalty, MEA consistently outperforms post hoc explainers, agentic, and closed-source baselines across six datasets, with reward-driven optimization yielding faithfulness gains of +28% (tabular), +21% (text), and +34% (vision) over the untrained backbone. More broadly, our findings suggest that AI agents themselves can serve as a scalable, adaptable interface to ML explainability, opening a path toward natural-language explainability that generalizes beyond the fixed, single-purpose tools that have long defined the field.

[AI-141] ropical Reinforcement Learning

链接: https://arxiv.org/abs/2610.02478
作者: Arip Asadulaev,Aladin Djuhera,Karim Salta,Holger Boche,Fakhri Karray,Martin Takac
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Reinforcement learning for large language models typically maximizes expected return, adding up the probabilities of all successful trajectories. However, the classical sum formulation can only report how often the model policy succeeds, not which solution actually worked, and because probabilities sum to one, reinforcing one solution can make the model forget another that was never shown to be wrong. This makes expected return a poor fit for compositional reasoning, where a solution must be assembled from reasoning steps that the model produces in separate, often failed, attempts but rarely produces together. To address this, we propose Tropical Reinforcement Learning, which rests on a simple change of algebra: instead of adding the probabilities of alternative solutions, we take their maximum, which yields the tropical semiring. The value of a state then becomes the log-probability of its most likely verified solution, together with an explicit path that can be replayed and reused. This enables true composition, since the best prefix and the best suffix meeting at a shared state can be joined even when they come from different rollouts. To put this into practice, we introduce TROPIC, a training algorithm for deterministic, resettable environments with verifiable outcomes. On four agentic tasks (Sokoban, Countdown, FrozenLake, WebShop), TROPIC outperforms the strongest on-policy baselines by up to 16 percentage points. Changing the algebra of reinforcement learning, not just its estimators, can thus substantially improve compositional reasoning in language models

[AI-142] SideKernel: A Usable microVM Sandbox for AI Coding Agents on macOS

链接: https://arxiv.org/abs/2610.02456
作者: Dimitrios Prasakis
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
备注: 17 pages, 5 figures, 9 tables. Georgia Tech M.S. Cybersecurity practicum project

点击查看摘要

Abstract:AI coding agents are untrusted system components, yet they require autonomy on the developer machines they run on. This contradiction is a security problem. Sandboxes provide an isolated environment, but for local macOS development, the existing local, open-source options for AI coding agents are few in number and cumbersome to use. I conducted a formative online user survey which indicates that fewer than 40% of AI coding agent users run their agents in a sandbox and identifies the top usability barriers hindering AI coding agent sandbox adoption. These findings are used to develop SideKernel: an open-source, local, microVM-based macOS sandbox for AI coding agents designed for usability. To evaluate SideKernel, I compiled a list of sandboxes available on the market and filtered it against five inclusion criteria. Then I performed a comparative analysis between SideKernel and the sandboxes that satisfy these criteria, across 23 capability tests derived from the usability barriers revealed by the user survey. I discovered that only a few sandboxes are similar to SideKernel, and that among those, Docker Sandboxes and SideKernel score highest on capability features related to usability. A secondary contribution of this paper is a survey of the existing solution space for local, open-source, microVM-based macOS sandboxes for AI coding agents.

[AI-143] Reinforcement Learning Techniques for the Optimization of Target Polarization in Nuclear Physics Scattering Experiments

链接: https://arxiv.org/abs/2610.02452
作者: Armen Kasparian,Torri Jeske,Monibor Rahman,Chris Keith,James Maxwell,Thomas Britton,Malachi Schram,David Lawrence
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:The operation of dynamically polarized targets in nuclear physics experiments relies on continuous tuning of the microwave frequency to compensate for radiation damage and evolving material properties, a task that is traditionally performed through manual trial-and-error by expert operators. This work presents a data-driven control framework that combines surrogate modeling with reinforcement learning to optimize the target polarization. Using operational data from the APOLLO cryogenic target system, we train and evaluate multilayer perceptron and Gaussian process regression models to predict polarization as a function of microwave frequency, beam current, and accumulated radiation dose. We show that Gaussian process-based models provide calibrated uncertainty estimates and reliably identify regions outside the training distribution, while MLPs exhibit limited sensitivity to distributional shift. To enable learning and control across multiple target samples, we introduce a Gaussian process approximation and embed the surrogate model within a standardized simulation environment. A reinforcement learning agent is trained using a lower-confidence-bound reward formulation that balances performance maximization against uncertainty. We are able to show an almost 2x improvement on the operators actions utilizing our RL agent.

[AI-144] Geometry-Aware Time Reparameterization for Flow-Map Distillation

链接: https://arxiv.org/abs/2610.02427
作者: Félix Dedek,Makoto Yamada
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Flow-map distillation enables one- and few-step generation by learning finite-time transitions of a pretrained generative ODE. We investigate whether changing the teacher’s time parameterization can make these transitions easier to learn. Motivated by the hypothesis that trajectory segments with large normal acceleration are harder to distill, we propose a geometry-aware time reparameterization that allocates more student time to these regions while preserving the teacher’s geometric paths and terminal distribution. We derive a shared clock that equalizes a population normal-acceleration statistic under suitable assumptions, and construct a practical approximation from robust, regularized estimates across teacher trajectories. We incorporate this clock into Lagrangian flow-map distillation, using the transformed time coordinate to condition the student. The clock is estimated once before distillation and requires neither teacher retraining nor additional student parameters or inference-time network evaluations. Experiments on synthetic data, CIFAR-10, and CelebA-64 show improved sample quality over identity-time distillation at matched inference budgets, including improvements in one- and two-step image generation. The gains in one-step generation, where no intermediate sampling times can be adjusted, highlight the benefits of time reparameterization during distillation.

[AI-145] Mitigating Private Data Leakage in LLM s with Whiteout

链接: https://arxiv.org/abs/2610.02418
作者: Anna Yoo Jeong Ha,Ronik Bhaskar,Haitao Zheng,Ben Y. Zhao
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Modern large language models (LLMs) are trained on massive, largely unfiltered datasets, including content scraped from nearly every accessible website and user inputs. As a result, LLMs often memorize and reproduce personally sensitive information (PSI) such as birth dates, phone numbers, and home addresses. This leads to significant privacy risks, particularly for high-profile individuals such as executives, politicians, and judges. Existing mitigations largely rely on machine unlearning. However, these methods often remove more information than needed, degrade model utility and safety, and are highly vulnerable to attacks. This paper presents Whiteout, a practical tool that, upon requests by individuals, prevents LLMs from regurgitating their genuine PSIs, by overwriting them using precise and carefully designed obfuscation samples. We evaluate Whiteout on modern LLMs of varying sizes and makers, including a widely-used OpenAI model. Results show that Whiteout effectively prevents disclosure of the targeted PSIs, has negligible impact on model utility and safety, and outperforms existing alternatives. We also test Whiteout against a wide range of countermeasures, from black-box attacks like jailbreaking to white-box adaptive attacks like relearning and quantization. Finally, we conclude with a discussion on the security and ethical implications of Whiteout. Subjects: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI) Cite as: arXiv:2610.02418 [cs.CR] (or arXiv:2610.02418v1 [cs.CR] for this version) https://doi.org/10.48550/arXiv.2610.02418 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[AI-146] Efficient Neural Field Learning via Adaptive Coverag e and Focused Sampling NEURIPS2026

链接: https://arxiv.org/abs/2610.02410
作者: Guang Zhao,Xihaier Luo,Huan-Hsin Tseng,Seungjun Lee,Shinjae Yoo,Yihui Ren,Wei Xu
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: 22 pages. Accepted at NeurIPS 2026

点击查看摘要

Abstract:Implicit neural representations (INRs) provide a flexible framework for modeling high-dimensional continuous fields, but their training is often inefficient due to uniform subsampling that ignores spatial heterogeneity. Existing adaptive sampling methods partially address this issue by prioritizing high-error samples, but typically operate at the point level, often leading to redundant sampling in localized regions and insufficient coverage of the domain. We propose ACES (Adaptive Coverage-aware Efficient Sampling), a structured sampling framework that improves training efficiency by decoupling coverage and importance. ACES constructs adaptive spatial partitions to ensure domain coverage and reduce redundancy, and applies region-level importance weighting to prioritize informative regions during training. We provide a theoretical analysis showing that adaptive partitioning reduces gradient variance by increasing within-region homogeneity, and that controlled bias in region-level weighting may improve optimization efficiency relative to standard unbiased estimators. Experiments on scientific field learning tasks demonstrate that ACES achieves faster convergence and lower error than uniform and pointwise adaptive sampling baselines, with the largest gains in fields with highly localized complexity.

[AI-147] When Terminal-Agent Training Stalls: Demystifying Data Generation and Verification Challenge

链接: https://arxiv.org/abs/2610.02405
作者: Xi Qin,Isabel Kurth,Xin Cui,Elin Park,Alexander Schaefer,Yaad Oren
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Using a frontier model like Claude Opus as a meta-agent to generate terminal tasks and verifiers for RL training is increasingly common. Yet a runnable Docker image and executable test suite do not guarantee a faithful end-to-end pipeline for terminal agent training. We present a meta-agent pipeline motivated by this gap, diagnosing three classes of failure: benchmark invalidity, harness brittleness, and reward misalignment. Prompt redesign and context extension raise baseline solvability 5.6 times, but a 9B model saturates at 81.3% mean pass@2 within 20 steps on Claude Opus-generated tasks. Adding hard tasks reduces mean pass@2 to 20.6% without changing the training configuration, a strong evidence that the solvability band is model-specific. These findings demonstrate that meta-agent reliability requires solvability-band calibration, verifier audits, and infrastructure error accounting as first-class evaluation criteria, not post-hoc diagnost.

[AI-148] Inherit-MAS: Test-Time Evolution of Multi-Agent Systems through Workflow and Execution Inheritance

链接: https://arxiv.org/abs/2610.02396
作者: Songtao Wei,Yi Li,Zhichun Guo,Bingzhe Li
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Multi-agent systems (MAS) built from large language models coordinate specialized agents to tackle complex tasks, but effective workflows are difficult to design in advance. Test-time evolution refines workflows using execution feedback, yet broad revisions can disturb useful components, while re-executing unchanged requests can incur redundant computation. Inspired by the interplay of inheritance and selection in biological evolution, we introduce Inherit-MAS, which makes inheritance explicit at the workflow and execution levels. A meta-model first synthesizes a workflow of worker agents with declared roles, communication inputs, and tool permissions, and a separately prompted judge scores each executed candidate and diagnoses its deficiencies. In ordinary refinement rounds, \emphworkflow inheritance starts from the latest completed candidate, may discard removable nodes judged unhelpful, and applies a validated edit to address the diagnosed deficiency. When the new candidate executes, \emphexecution inheritance inherits eligible stored results only if the complete resolved request and execution context match, avoiding redundant model and tool calls. With GPT-4o-mini workers, Inherit-MAS achieves 55.4% completion on WorkBench and 49.7% joint F1 on HotpotQA FullWiki, outperforming EvoAgent, EvoMAS, and TacoMAS. With Qwen3-32B workers, it also exceeds these evolving-MAS baselines on both benchmarks. Compared with rerunning the same controller with execution inheritance disabled, execution inheritance reduces worker-token usage by 29.1% on WorkBench and 34.6% on HotpotQA, and total token usage by 5.3% and 18.1%.

[AI-149] FlashSinkhorn 2: Block-Sparse Entropic Optimal Transport

链接: https://arxiv.org/abs/2610.02395
作者: Felix X.-F. Ye,Yu Chin Fabian Lim,Naigang Wang,Davis Wertheimer
类目: Artificial Intelligence (cs.AI); Instrumentation and Methods for Astrophysics (astro-ph.IM); Numerical Analysis (math.NA)
备注:

点击查看摘要

Abstract:Streaming GPU solvers for entropic optimal transport (EOT), such as FlashSinkhorn, avoid storing the dense kernel but still evaluate all n\times m point pairs in every Sinkhorn iteration. We present \textbfFlashSinkhorn~2 (FS2), a solver for squared-Euclidean cost on low-dimensional point clouds that solves large discrete EOT problems to a prescribed marginal residual on a single GPU by coupling two stages. A coarse stage solves on cell centroids, lifts the potentials to every point and, when a sampled marginal check rejects the lift, continues on the centroids, replacing most point-level updates. A block-sparse fine stage then removes the centroid error that coarse updates cannot. Its Morton-ordered blocks support screening and fused tensor-core execution, and a threshold set by the block masses bounds each omitted tile’s contribution to every row and column. On synthetic benchmarks, FS2 reaches the target residual on all 32 problems and GeomLoss multiscale on 10. On one A100, FS2 solves discrete EOT between two 1.34\times10^8 -particle measures from a cosmological N -body simulation, at an entropic blur equal to the mean interparticle distance, to an all-particle marginal residual below 0.01 in under 2.5 hours. To our knowledge, it is the largest discrete EOT problem solved to this accuracy within hours. For reproducibility, we release an open-source implementation at this https URL

[AI-150] he Surprising Effectiveness of Shared Memory in Looped Transformers

链接: https://arxiv.org/abs/2610.02383
作者: Giovanni Monea,Keshav Ramji,Yousef El-Kurdi,Luis A. Lastras,Yoav Artzi,Nathan Godey,Ramón Fernandez Astudillo
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Looped Transformers apply the same layers several times per token, adding compute to improve quality without more parameters. Each recursion, however, writes its own key-value cache, so memory still grows with compute. Inference-time techniques can shrink this cache at a cost in quality. We pretrain looped language models to share memory: only the first recursion writes a cache, and later recursions read it while keeping a short window of their own. Surprisingly, we find that sharing memory does not cost quality and instead improves it. At 150M-1B parameters, our Looped Prediction Transformer (LPT) and its hybrid variant set a new quality-memory frontier for looped models: with five recursions, the hybrid lowers validation perplexity on FineWeb-Edu by 1.12-1.82 relative to a same-size standard Transformer while using 76-79% less context memory. Through an extensive analysis, we investigate why memory sharing helps. Shared and local memory develop different representations, and later recursions attend mostly to the shared memory, which also acts as a gradient highway to the first recursion.

[AI-151] HPL: A Vision-to-Language Decision Support Framework for Rainbow Trout Feeding Management in RAS

链接: https://arxiv.org/abs/2610.02378
作者: Meng Liang,Guanbo Feng,Haozhuang Chi,Shilong Zhao,Zhixin Xiong,Yuhang He,Wenfeng Han,Tianhao Zhao,Zhihong Ma,Ying Liu
类目: Artificial Intelligence (cs.AI)
备注: Meng Liang and Guanbo Feng contributed equally. Corresponding authors: Zhihong Ma and Ying Liu. 50 pages, 10 figures, 3 tables. Supplementary video: this https URL

点击查看摘要

Abstract:In Recirculating Aquaculture Systems (RAS), precision feeding is critical for minimizing costs and improving fish welfare. However, existing methods lack cognitive alignment between fish behaviors and management knowledge, impeding translation into executable, interpretable feeding decisions. To address this, we propose THPL, a generative feeding decision framework tailored for rainbow trout (Oncorhynchus mykiss) in RAS. First, Fishsort extracts trajectories to establish an Activity Coefficient (AC) quantifying feeding intensity. Second, a Hierarchical Behavior Encoder (HBE) models individual temporal progression and collective dynamics using Temporal and Set Transformers, transforming trajectory tensors into dual-evidence representations of explicit physical and implicit soft tokens. Finally, these tokens are integrated with environmental parameters, metadata, and expert rules to fine-tune an LLM via LoRA, followed by counterfactual multimodal Direct Preference Optimization (mDPO) to reinforce causal reasoning. Results show that AC exhibits a statistically significant monotonic positive correlation with expert-annotated feeding intensity (Spearman \rho = 0.925 , p 0.001 ). Ablations indicate that decision accuracy improves from 33.33% (text-only baseline) to 93.33% with dual-evidence tokens, confirming that continuous spatiotemporal tokens provide necessary physical grounding for LLMs. Compared with standard LoRA, counterfactual mDPO elevates decision accuracy from 93.33% to 96.67%, advances METEOR from 58.10% to 85.30%, reduces Self-BLEU-2 from 58.79% to 52.88%, and increases Distinct-3 from 6.68% to 7.81%, suppressing templating and actuation biases while reinforcing causal consistency and operational safety. Overall, by integrating continuous kinematics with LLM reasoning, this study provides a novel decision support paradigm for precision aquaculture.

[AI-152] Coco: An Agent ic Copilot for the Hardware–Software Co-Design Lifecycle

链接: https://arxiv.org/abs/2610.02376
作者: Samuel Kushnir,Kavya Sreedhar,Yeshwanth Reddy Pogula,Amir Yazdanbakhsh,Narges Shahidi,Ming Liu,Varun Gohil,Ravi Iyer,Parthasarathy Ranganathan,Christina Delimitrou,Suvinay Subramanian
类目: Programming Languages (cs.PL); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Co-designing ML models and the accelerators that run them is an unusual reasoning task: architects must draw confident, high-stakes conclusions about systems that do not yet exist, and the pace of both model evolution and hardware cadence means the analysis burden grows every quarter. The evidence behind each decision–hundreds of gigabytes of fresh simulation sweeps over novel design points–is by construction absent from any LLM’s pretraining corpus, and there is no external literature to retrieve; naive “chat-with-your-data” approaches hallucinate exactly where correctness matters most. We present Coco (Copilot for Codesign), an agentic platform deployed with TPU architects that accelerates the co-design lifecycle of setting up experiments, sweeping simulators, and deriving insights. Coco is built as four layers: (i) a datastore that automatically registers every simulation sweep into a normalized relational schema, so agents ground every number in a SQL query rather than scraping heterogeneous files; (ii) a library of tools with typed APIs that agents compose without human orchestration; (iii) agents that encode recurring analysis workflows–most notably iso-execution analysis, which compares systems at matched execution configurations, including swept-but-dominated points off the Pareto frontier; and (iv) a platform UX whose navigation state doubles as agent context. We report early deployment experience toward a reduction in time-to-simulation and time-to-insight, and argue that co-design is a distinct agentic domain: its data must be retrieved rather than memorized, its workflows are recurring but context-dependent, and expert adoption hinges on UX that balances IDE-style control with interactive exploration.

[AI-153] Hop-Decayed Influence: New Vulnerabilities of Structural Auxiliary Indexing in GraphRAG Pipelines with LLM

链接: https://arxiv.org/abs/2610.02373
作者: Jisung Park,John Le,Heath Cooper
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
备注: 14 pages. Published in IFIP SEC 2026. Best Paper Award

点击查看摘要

Abstract:GraphRAG pipelines construct auxiliary structures during offline indexing–semantic summaries, hierarchical edges, and pre-computed scores–that determine how retrieval is prioritised at query time. Prior attacks target only instance-level components (nodes, edges, triples), overlooking these schema-level structures. We formalise Auxiliary Schema-Level Entity as a novel attack surface and propose the 3S Framework (Semantics, Structure, Scoring) for its systematic exploitation. Our Hop-Decayed Influence (HDI) attack identifies high-impact targets through query-aware influence propagation and corrupts their auxiliary structures post-indexing. Across two benchmarks (HotpotQA, 2WikiMultiHopQA) and two architectures (Microsoft GraphRAG, HippoRAG2), HDI achieves 88-94% attack success rate while modifying as few as 0.016% of auxiliary structures. Each modification affects up to 6.00 queries (Schema Leverage Ratio), demonstrating 1:N amplification unavailable to instance-level attacks. Manipulated structures evade perplexity and paraphrase defenses with over 99% evasion rate, as they remain linguistically coherent system-generated artifacts. These results reveal that auxiliary schema-level entities receive implicit trust without runtime validation, constituting a structural blind spot in current GraphRAG defenses. this https URL.

[AI-154] raversing the Satisfaction-Diversity Frontier in Text-to-Image Diffusion

链接: https://arxiv.org/abs/2610.02372
作者: Kevin Zhai,Siva Rajesh Kasa,Soumya Roy,Sumit Negi,Mubarak Shah
类目: Artificial Intelligence (cs.AI)
备注: 40 pages, including appendices. Code: this https URL

点击查看摘要

Abstract:Text-to-image generation enables users to explore several images generated from the same prompt. For these generated images to be useful, each one must reflect the user’s preferences, measured by a learned reward, and differ visually from the others to maintain diversity. Existing methods are limited: they either address reward and diversity separately or combine them in one aggregate score, enabling high diversity to offset low rewards. In this paper, we address these limitations by formulating generation as satisficing: every image (candidate) must satisfy a reward floor and the batch of images must satisfy a diversity cutoff. The reward floor controls the balance between worst-candidate reward and batch diversity; we show that varying this floor defines a Pareto frontier. To traverse this frontier, we introduce SatisDive, a training-free inference-time method. SatisDive uses a batch-relative reward cutoff to distinguish lower- from higher-reward candidates, emphasizing reward improvement for candidates below the cutoff and diversity among candidates above it. On Pick-a-Pic, at matched DreamSim, SatisDive improves worst-candidate reward over FK steering by up to 0.43 with FLUX.1-dev as the base model and HPSv3 as the reward, and by up to 0.70 with SANA-1.6B as the base model and ImageReward as the reward. More broadly, across their overlapping DreamSim ranges, SatisDive’s satisfaction-diversity curve Pareto-dominates FK steering’s curve in each setting.

[AI-155] Network-in-the-Loop at Scale: GPU-Batched 5G Simulation for Massively Parallel Robot Learning

链接: https://arxiv.org/abs/2610.02370
作者: Zifan Zhang,Mingzhe Han,Kannan Athreya,Yuchen Liu
类目: Networking and Internet Architecture (cs.NI); Artificial Intelligence (cs.AI); Distributed, Parallel, and Cluster Computing (cs.DC); Robotics (cs.RO)
备注: It is open source at this https URL

点击查看摘要

Abstract:Massively parallel GPU simulators train multi-robot policies in thousands of environments, and many fleets use private Fifth-Generation (5G) networks, where each robot’s delay depends on its teammates’ traffic. Network-in-the-loop training places a simulated 5G network inside this loop. However, GPU robot simulators reduce the network to an independent delay per message, while packet-level simulators run one scenario per CPU process and cannot keep pace with thousands of parallel environments. To bridge this gap, we present Isaac-Net, a GPU-batched 5G New Radio (NR) module that advances the uplink of thousands of environments in lockstep with Isaac Lab physics. Isaac-Net simulates every slot, the 0.5~ms interval in which the base station decides which robots transmit, for all environments at once. Extensive experiments confirm that its NR engine reproduces the median delay of ns-3 5G-LENA across loads, with a median delay 5–10% low on an unseen carrier and 9% high at 32 robots per environment in closed loop. The engine also reproduces the Age of Information (AoI), the age of each robot’s newest delivered report, while an independent delay per message leaves the AoI tail about three times too light. In a configuration validated against 5G-LENA, Isaac-Net keeps the network in the loop for about one million robots on one GPU at 83% of the Isaac Lab rate without the network, measured under a random policy. Isaac-Net is open source at this https URL

[AI-156] DeReAct: Decomposed Reasoning and Acting for Reliable AI Agents

链接: https://arxiv.org/abs/2610.02351
作者: Ajay Vohra,Tao Chen,Neeti Narayan,Caron Zhang
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:ReAct-based agents typically rely on a single LLM policy to propose actions, interact with the environment, and decide when a task is complete. This coupling makes action authorization and completion control difficult to enforce independently, allowing errors to propagate and unsupported completion claims to terminate execution. We introduce DeReAct, a modular agent architecture that externalizes two gating policies: a Critic that validates proposed actions before execution, and a Context Manager that reconstructs an environment-supported \textscState and certifies task completion. Across GAIA and SWE-bench Verified, DeReAct improves Pass@1 most for weaker Brain models, with gains of 6.5–7.0 points for Qwen3-Coder-480B and 4.2–5.2 points for Claude Sonnet~4.5; gains diminish as Brain capability increases. Trajectory and ablation analyses show that external gating is effective when targeted failures are sufficiently prevalent and the gating policy is itself sufficient. With Claude Opus~4.5, Pass@1 remains comparable to ReAct, while DeReAct produces more evidence-complete and constraint-satisfying trajectories, indicating that completion control can trade earlier termination for stronger grounding. Overall, DeReAct improves weaker agents while retaining grounding benefits as models strengthen. Subjects: Artificial Intelligence (cs.AI); Machine Learning (cs.LG) Cite as: arXiv:2610.02351 [cs.AI] (or arXiv:2610.02351v1 [cs.AI] for this version) https://doi.org/10.48550/arXiv.2610.02351 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[AI-157] A Multi Method Importance and Performance Efficiency Analysis of Topological Metrics for Natural Visibility Graph Based Cyber Attack Detection

链接: https://arxiv.org/abs/2610.02342
作者: Ali Melih Kanca,Ilker Turker
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Natural Visibility Graph (NVG) based analysis characterizes network traffic through topological descriptors reflecting different structural properties. However, not all descriptors contribute equally to cyber-attack classification, and extracting a large metric set can increase computational cost. This study evaluates 21 NVG derived topological metrics and investigates whether a compact subset can preserve classification capability while improving computational efficiency. Four importance analysis methods SHAP, grouped Permutation Importance, Boruta, and Recursive Feature Elimination (RFE) are integrated through a Consensus Ranking strategy. Based on this ranking, Full21, Top15, Top10, Top7, Top5, and Top3 configurations are evaluated using the CICIDS2018 dataset, a CNN classifier, and stratified 5 fold cross validation. The three highest ranked metrics are avg_clustering_coeff_median, avg_clustering_coeff_std, and avg_clustering_coeff_mean. Top3 achieved the highest observed mean performance, with 97.148% accuracy, 97.055% weighted F1 score, and an MCC of 0.9675, compared with 95.999%, 95.521%, and 0.9549 for Full21, respectively. It also reduced total runtime from 14,961.39 s to 589.22 s (96.06%). These results indicate that importance guided metric reduction can provide a compact NVG representation with higher observed mean predictive performance and substantially lower computational cost under the evaluated setting.

[AI-158] World Editing: Intervening on Executable Worlds at Increasing Depth

链接: https://arxiv.org/abs/2610.02331
作者: Max Ku,Nok-Kan Law,Yu-Chien Tang,Shih-Ying Yeh,Ping Nie,Andy Zheng,Tat Hei Lai,Fei-Yueh Chen,Nikko Yu,Wei-Chieh Sun,Suzy Huang,Chiao-Wei Hsu,Chih-Chuan Huang,Chak-Wing Mak,Ho Yin Sam Ng,Edisy Kin Wai Chan,Min-Hung Chen,Ho Kei Cheng
类目: Artificial Intelligence (cs.AI)
备注: Preprint. Project page: this https URL

点击查看摘要

Abstract:Interactive world models are increasingly capable of generating environments and acting within them, yet deliberately editing an existing executable world remains underexplored. We formulate world editing as intervening on an existing world while preserving properties that should remain unchanged, and introduce intervention depth as an axis describing how strongly an edit couples world entities, dynamics, and systems. We instantiate this capability through industry-grade game modding and introduce IGMWorld, together with IGMBench, a benchmark of 110 tasks and over 1.1K executable state and behavioral criteria across Minecraft and Terraria. The tasks span property, entity, dynamics, and system interventions and are evaluated through deterministic executability, behavioral, preservation, and visual checks. Frontier coding agents already exhibit substantial world-editing capability: the strongest configuration solves 78.2% of tasks under a strict task-level criterion, while criterion-level performance reaches 94.8%. Reliability generally decreases with intervention depth, and this pattern persists even among tasks with similar numbers of evaluation criteria. Most failed edits still build and load successfully, suggesting that the main difficulty is making the edited world behave as requested. Visual consistency remains a separate weakness, with all evaluated configurations below 50% joint visual pass rate. These results show that world editing is a distinct capability from world generation and interaction, and that executable games provide a practical testbed for studying it.

[AI-159] Choosing Before Acting: Comparative Value Estimation for Long-Horizon Tool-Use Agents NEURIPS2026

链接: https://arxiv.org/abs/2610.02330
作者: Yu Li,Zheng Zhang,Xin Liu,Shengtian Yang,Guangfeng Cai,Lei Feng
类目: Artificial Intelligence (cs.AI)
备注: NeurIPS 2026 Poster

点击查看摘要

Abstract:Large language models (LLMs) rely on long-horizon tool invocation sequences for complex tasks, where each invocation can alter the task state and condition subsequent decisions. In long-horizon tool use, final-outcome rewards provide weak credit assignment over long interaction traces. Step-level rewards can offer more targeted feedback, but obtaining reliable step supervision often requires human or LLM judgment, or additional rollouts to estimate the downstream effect of an intermediate decision. In this paper, we argue that effective tool-use agents should estimate the long-horizon value of a possible next tool invocation before executing it. This objective requires comparative supervision over alternative invocations under the same context, while logged trajectories only contain the invocation that was actually taken. Therefore, we propose Comparative Inference for Tool-use Agents (CITA). CITA trains a Comparative Inference Model (CIM) from paired signals that combine observed tool behavior, scalable supervision from a Bayesian tool-graph simulator, and semantic judgments from LLM-based comparison. The resulting CIM learns to estimate how likely a possible next tool invocation is to support final task success under the current context. Across three tool-use benchmarks and multiple backbone LLMs, CITA consistently improves Tool F1 and task success. Additional analysis shows that CIM learns accurate step-level value estimates for comparative tool choices.

[AI-160] Slow-Fast Multi-Teacher On-Policy Distillation for Capability Preservation

链接: https://arxiv.org/abs/2610.02324
作者: Xiaofei Yin,Tong Chu,Jiyuan Fu,Jun Lan,Shuheng Zhou,Huijia Zhu
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: 5 pages, 2 figures

点击查看摘要

Abstract:Foundation multimodal large language models are designed to support a broad spectrum of capabilities across diverse domains. Multi-teacher on-policy distillation (MOPD) provides an effective framework for consolidating domain-specific expertise into a single student model. However, MOPD training gradually drives the student away from its initialization model, and general capabilities decline as the displacement grows, resulting in capability interference. A direct remedy is constraining the student toward its initialization, but this suppresses the acquisition of domain expertise as well. We propose Slow-Fast Multi-Teacher On-Policy Distillation (SF-MOPD), which couples a fast model, the current student updated directly by each teacher, with a slow model, an exponential moving average of the student. The slow model absorbs the learning signal gradually, serving as a moving capability reference that fuses the general foundation with confirmed domain expertise. For each teacher, SF-MOPD computes the teacher-induced update in log-probability space and removes only the component that pushes the fast model further away from the slow model, while retaining aligned and orthogonal components. Experiments across multiple model scales demonstrate that SF-MOPD effectively mitigates capability interference, enhances specialized multimodal capabilities, and reduces the average degradation on general-capability benchmarks, consistently outperforming vanilla MOPD.

[AI-161] SimuVerity: Benchmarking Agents for Engineering-Grade Simulink Model Generation

链接: https://arxiv.org/abs/2610.02304
作者: Ruiqi Zhang,Jiahao Wang,Mingxuan Li,Haichen Luo,Chaoting Wang,Guoyu Mou,Keyu Lai,Hanchao Lv,Jiaxu Wang,Yibo Zheng,Aijun Yang,Xiaohua Wang
类目: oftware Engineering (cs.SE); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: 26 pages, 12 figures. Code and data are available at this https URL

点击查看摘要

Abstract:Existing Simulink benchmarks mainly evaluate whether generated models compile, execute, or resemble a reference model. These criteria do not establish whether a model satisfies its engineering requirements. We introduce SimuVerity, a benchmark of 101 text-to-executable Simulink model-generation tasks across ten engineering domains. For each task, executable-system profiles ground the engineering specification and four families of native simulation scenarios. A hierarchical evaluator first checks artifact delivery, native executability, and engineering qualification, then scores qualified models across six dimensions covering accuracy, output quality, mechanistic fidelity, control and causal integrity, operating-domain robustness, and dynamic response. We evaluate six agent systems with SimuVerity. The best system achieves an overall score of only 42.86. The results show that structural similarity is a poor proxy for engineering performance: capability bottlenecks arise both in producing qualified implementations and in satisfying multidimensional requirements after qualification. Meanwhile, some high-scoring models still exhibit severe visual-layout disorder. SimuVerity provides a systematic basis for assessing agents’ engineering capabilities and diagnosing failures in executable Simulink model generation.

[AI-162] Keep It CALM: Analyzing the Limits of Global Unsafety in Text-to-Image Generation NEURIPS2026

链接: https://arxiv.org/abs/2610.02300
作者: NaHyeon Park,Minhyun Lee,Hyunjung Shim
类目: Artificial Intelligence (cs.AI)
备注: NeurIPS 2026

点击查看摘要

Abstract:Training-free safeguards for text-to-image generation often rely on a reusable safety signal, such as an unsafe direction or global toxic subspace, applied broadly across prompts. We provide a controlled geometric analysis of this global-unsafety assumption and reveal a consistent coverage-selectivity trade-off: compact unsafe subspaces fail to cover heterogeneous unsafe semantics, whereas broader aggregation increasingly distorts safety-adjacent benign prompts. Motivated by this finding, we propose CALM (Counterfactual Adaptive Local Modulation), a training-free safeguard that replaces uniform global removal with prompt-local counterfactual correction. Using matched unsafe-benign anchors, CALM routes each prompt to active unsafe categories, minimally edits only violating token representations toward the safe side, and suppresses positively aligned unsafe residual components. Across broad evaluation, CALM significantly improves unsafe content suppression while preserving benign utility, demonstrating that local counterfactual correction provides a more selective alternative to global unsafe signal removal.

[AI-163] Ψ-Resilience: Model-Free Feature Importance from 1D Topological Signals

链接: https://arxiv.org/abs/2610.02299
作者: Fabian Galis,Darian Onchis,Pedro Real Jurado
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:We introduce \Psi -Resilience, a model-free feature importance method that derives explanations directly from the data itself via 1D topological signals. Our method constructs a class-disagreement landscape by estimating class-conditional densities and taking their pointwise absolute difference along the feature axis. Then, the 0-dimensional persistence of this 1D signal defines a resilience functional that aggregates only those topological features that survive perturbations up to a robustness scale which is set by the user. This gives us a context-robust importance score that is inherently auditable via the underlying 1D landscapes and their persistence. We evaluate our method on both synthetic and real datasets. On synthetic generators with specified ground-truth importance, \Psi -Resilience recovers the ranking of features with high fidelity, achieving Spearman rank correlations up to 0.8 and performing competitively with multiple feature importance methods, including SHAP and mutual information. On real datasets with no known ground truth, our technique agrees with these methods, with correlations up to 0.9. These results show that \Psi -Resilience is a stable explanation method that enables rigorous, distribution-level auditing of feature importance without relying on a predictive model.

[AI-164] Diffusion-Based Synthetic Data Pretraining for Enhancing Activity Recognition

链接: https://arxiv.org/abs/2610.02292
作者: E. Riveros(1),D. Vega-Oliveros(2),A. Soriano-Vargas(3),A. Rocha(1) ((1) Institute of Computing, State University of Campinas, Campinas, Brazil, (2) Institute of Science and Technology, Federal University of Sao Paulo, Sao Jose dos Campos, Brazil, (3) Universidad de Ingenieria y Tecnologia, Lima, Peru)
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: 6 pages, 3 figures, 1 table

点击查看摘要

Abstract:Human activity recognition (HAR) is increasingly important for healthcare, well-being, and daily monitoring ap- plications, for which detecting alimentary activities such as eating and drinking can provide actionable insight into dietary habits and chronic disease management. HAR systems, however, often underperform on subtle and underrepresented classes, limiting their utility in real-world dietary monitoring. This work builds upon CABiGRU, a convolutional architecture with Bidirectional GRU layers, multi-head attention, and residual connections, designed to capture discriminative temporal patterns from smart- watch accelerometer, gyroscope, and magnetometer data. To improve CaBiGRU’s generalization and reduce underfitting in the minority class, we leverage synthetic sensor data windows using a diffusion model and adopt a two-stage training strategy: pre-training CABiGRU on synthetic data, followed by fine-tuning on the real-world data. On the DEO (drinking/eating/other) dataset, the proposed pipeline achieves a balanced accuracy of 90.6%, improving over a strong supervised baseline and showing the benefits of diffusion-based synthetic pre-training for recognizing alimentary activities and representing a step forward dealing with unbalanced classes. These results suggest that combining diffusion-generated data with targeted fine-tuning enhances robust recognition of dietary behaviors, supporting more reliable deployment in healthcare and nutrition-monitoring settings.

[AI-165] he AI Risk Observatory: What Can We Learn from AI Disclosures in Annual Reports About Societal Resilience?

链接: https://arxiv.org/abs/2610.02281
作者: Bart Jaworski
类目: Artificial Intelligence (cs.AI)
备注: 22 pages (9 main text + appendices), 12 figures, 7 tables. Code and data: this https URL (release dataset-v1.1)

点击查看摘要

Abstract:Societal resilience research relies on access to useful and actionable data, which motivates our main research question: Can annual reports, processed at scale with LLMs, provide a useful signal about how companies disclose their response to AI? We test this by applying a reproducible two-stage classification pipeline to 9,821 annual reports from 1,362 UK listed companies (2020-2025, with partial 2026 data). We first validate the method against 474 human-annotated passages, finding high recall and moderate label-level agreement. We then report three empirical patterns: (i) between 2020 and 2025, the share of reports mentioning AI risk rose from 2.8% to 41.2%, while AI adoption disclosure also rose, from 13.8% to 45.2%, and named vendor mentions cluster around a small set of major providers led by Microsoft; (ii) disclosure varies substantially by Critical National Infrastructure sector and market segment: AIM reports disclose AI risk at far lower rates than Main Market reports, and sectors such as Energy and Data Infrastructure lag behind the rest in AI risk disclosure; and (iii) harm disclosures are near-absent (seven reports across the entire corpus). We develop a substantiveness classification to assess the quality of the disclosure and find that most AI risk disclosure is not substantive: in 2025, 41.2% of all reports mention AI as a risk, but only 4.3% contain AI risk disclosure we classify as substantive.

[AI-166] MintFlow: Minimal Trajectory Intervention for Constrained Flow Matching

链接: https://arxiv.org/abs/2610.02260
作者: Yesom Park,Kelvin Kan,Qifan Chen,Thomas Flynn,Hayden Schaeffer. Xihaier Luo
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Flow matching models excel at generative modeling, and many downstream applications require their samples to satisfy prescribed constraints, such as observed measurements and physical laws. However, existing constrained samplers often face a trade-off: \textitenforcing constraints can substantially displace samples from the pretrained data distribution. To address this trade-off, we introduce \textbfMintFlow, a training-free constrained sampling framework that formulates constraint enforcement as a minimal intervention on the pretrained flow trajectory. MintFlow seeks the minimal perturbation of an intermediate flow state such that its subsequent evolution under the pretrained flow field satisfies the target constraint. By minimally perturbing the flow state while keeping the pretrained flow field unchanged, MintFlow enforces the constraint while minimizing unnecessary deviation from the pretrained distribution. An adjoint formulation yields a closed-form expression for this perturbation, eliminating expensive iterative optimization. Furthermore, MintFlow adaptively selects the intervention time to balance the required perturbation magnitude with its amplification by the remaining flow. Across a range of tasks in generative vision and physical system modeling, MintFlow achieves competitive constraint satisfaction while preserving the pretrained generative distribution substantially better than state-of-the-art constrained methods.

[AI-167] Overcoming Challenges of Interpretive Structural Modeling with Large Language Models

链接: https://arxiv.org/abs/2610.02254
作者: Everett Rush,David J. Icove,Ari Kim,Byung H. Park,Michael A. Langston
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: This preprint has not undergone peer review or any post-submission improvements or corrections

点击查看摘要

Abstract:Interpretive Structural Modeling (ISM) is a well-known process for multi-criteria decision making. The success of ISM over other methodologies is its ability to model causal relationships, the binary scale of factors, and resulting hierarchical representation. Traditionally, the modeling process is performed by repeated interactions with subject matter experts until consensus is reached. This process is tedious, labor-intense, and most importantly limits the ability of ISM to scale to studies with hundreds of variables. Drawing on existing work of causal graph discovery with large language models (LLM) as imperfect experts, this work explores an integrated LLM-ISM approach for ISM. Pairwise, k-wise, rowwise, and full graph discovery methodologies are compared and evaluated. It is shown that causal graph discovery methods for ISM perform best using rowwise (SHD=160, F1-score=0.77) and full graph methods (SHD=135, F1-score=0.73).

[AI-168] Counterfactual Predictions in Scientific Emulators Without Controlled Experiments

链接: https://arxiv.org/abs/2610.02252
作者: Dingling Yao,Kahaan Gandhi,Valentin Duruisseaux,Boris Bonev,Francesco Locatello,Anima Anandkumar
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Machine Learning (stat.ML)
备注:

点击查看摘要

Abstract:Many scientific questions require reasoning about what was never observed: What if the conditions, interventions, or history had been different? Models can predict accurately on observed data yet fail on such what-if queries when correlated inputs are varied independently. A common remedy is to add controlled simulation data in which these factors are explicitly disentangled, but this requires access to a simulator, can be computationally expensive, and inherits the simulator’s modeling assumptions. We introduce ReRoute, a framework for targeted scientific what-if prediction that combines factual data with partial mechanistic knowledge, without requiring controlled intervention data for adaptation. ReRoute fixes the queried input of a pretrained backbone to a reference value, reintroduces its variation through a known mechanistic pathway, and fine-tunes on the original factual data, while leaving downstream effects to the learned dynamics. We provide a causal identification result for this construction under explicit structural assumptions, with the core argument machine-checked in Lean. After showing that ReRoute achieves highly accurate counterfactual predictions in a controlled advection-diffusion system where exact responses are available, we turn to state-of-the-art climate emulation. On held-out coupled-climate interventions, ReRoute reduces aggregate climate error by 18.2-31.8% under severe CO _2 distribution shifts while preserving skill under standard conditions, at a small fraction of the cost of retraining on additional controlled simulations, without even accounting for the substantial expense of generating such data. Finally, on an emulator trained from historical ERA5 reanalysis, where no counterfactual reference exists, ReRoute preserves substantially more of the surface warming implied by the observed boundary conditions under a fixed-CO _2 counterfactual.

[AI-169] Hardware-Native Joint Sparse-Quantization for Trillion-Scale Mixture-of-Experts

链接: https://arxiv.org/abs/2610.02241
作者: Kwanhee Lee,Namhoon Lee,Dan Alistarh
类目: Hardware Architecture (cs.AR); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Mixture-of-Experts (MoE) architectures allow frontier language models to scale to trillions of parameters, but their deployment is constrained by massive memory footprints and memory-bandwidth limitations. Although modern accelerators provide Sparse Tensor Cores (SpTCs) that reduce weight storage and increase throughput through low-precision semi-structured sparsity, exploiting them for MoEs remains challenging because of substantial model-quality degradation and the lack of grouped sparse GEMM primitives. We present an end-to-end hardware-software co-design framework that compresses expert weights into hardware-native, low-precision sparse representations and accelerates their execution on SpTCs. Algorithmically, our framework relaxes discrete semi-structured support selection through continuous reparameterization, enabling differentiable joint optimization with quantized weights under a router-weighted reconstruction objective and scalable expert-parallel compression. Systemically, we develop a custom grouped sparse GEMM kernel tailored to low-precision sparse MoE inference on SpTCs. Across MoE models ranging from 30 billion to one trillion parameters, our framework improves state-of-the-art joint sparse-quantization accuracy by up to 4.35 percentage points while preserving 96.09% of the original model’s performance. On NVIDIA B200 GPUs, our kernel outperforms the vendor baseline by up to 1.65\times , increasing serving throughput by 1.18\times and reducing end-to-end latency by up to 4.03\times . These results establish hardware-software co-design as a practical path toward scalable and efficient MoE deployment.

[AI-170] CORE: COverag e CAlibration and Evicted-Mass REdistribution for KV Cache

链接: https://arxiv.org/abs/2610.02235
作者: Shuxin Liu,Qing Liu,Yi Du,Ou Wu
类目: Hardware Architecture (cs.AR); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Long-context decoding is increasingly constrained by key–value (KV) cache memory and bandwidth. Existing fixed-budget compression methods typically separate retention from compensation, while a retention ranking specifies neither discarded attention mass nor the direction of induced output error. We start from an exact factorization: eviction error equals evicted attention mass times the directional gap between the evicted centroid and retained output, highlighting the importance of set-level coverage in retention and mass-preserving memory writing. We introduce CORE COverage Calibration and Evicted-Mass REdistribution for KV Cache, which distills an offline allocation combining query utility and log-determinant coverage into a lightweight cache-aware indexer. At inference, one calibrated distribution drives both channels: its Top- B ordering retains complementary KV states, while its excluded allocation mass and conditional weights parameterize latent-memory writes without a separate write-weight predictor or online log-determinant evaluation. Our analysis provides a four-term pre-compensation error certificate, characterizes non-additive coverage interactions, and establishes mass-independent write stability with hierarchical bounds through recurrent and query-adaptive normalization. Across three backbones, CORE exceeds the strongest RULER baseline by up to 3.78 points at 90% compression; LongBench and repeated-eviction evaluations further demonstrate strong effectiveness and decoding efficiency.

[AI-171] CONTRA: Discovering and Qualifying Behavior-Changing Questions for Selective Clarification in LLM Code Generation

链接: https://arxiv.org/abs/2610.01769
作者: Zheng Fang,Yongmin Li,Yichang Zhang,Dongming Jin,Haoyu Wang,Shuai Wang,Zhi Jin,Ge Li
类目: oftware Engineering (cs.SE); Artificial Intelligence (cs.AI)
备注: 15 pages. Code: this https URL

点击查看摘要

Abstract:Coding agents can generate code that appears correct but implements behavior the user never intended. This mismatch can arise when an agent silently resolves underspecified requirements through its own assumptions. As subsequent development builds on these assumptions, correcting the resulting behavior can become increasingly costly. Early clarification can help prevent such mismatches, but unnecessary questions can interrupt developers and slow down development. Existing methods struggle to identify key clarification questions while avoiding unnecessary ones. Therefore, we propose CONTRA, a training-free method that combines broad question discovery with semantic and execution-based question qualification. CONTRA first generates candidate questions and filters out those unrelated to required behavior or already resolved by the requirement. For each remaining question, it generates programs conditioned on two plausible answers and checks for stable behavioral differences on shared inputs. It then uses the interaction history to select among qualified questions or stop asking. Experiments on ClarifyCodeBench show that CONTRA achieves the highest F1 with all four coding agents, exceeding the best baseline macro-average F1 by 13.88 percentage points. With the same LLM and evaluation protocol, CONTRA also achieves higher clarification recall and F1 than the coding harnesses Claude Code and OpenHands. To support practical use, we also implement CONTRA as a Claude Code plugin that integrates selective clarification into everyday development.

[AI-172] ClarifyCodeBench: Evaluating LLM s on Clarifying Ambiguous Requirements for Code Generation

链接: https://arxiv.org/abs/2607.00711
作者: Zheng Fang,Dongming Jin,Yihong dong,Yongmin Li,Kechi Zhang,Zhi Jin,Ge Li
类目: oftware Engineering (cs.SE); Artificial Intelligence (cs.AI)
备注: Code: this https URL

点击查看摘要

Abstract:Large Language Models have emerged as programming assistants. However, the efficacy of code generation is constrained by the quality of input requirements, which are frequently ambiguous, incomplete, or underspecified. While LLMs excel at one-shot code synthesis, their ability to proactively clarify intent remains underexplored, as a critical trait for robust software engineering. Existing benchmarks largely overlook this interactive bottleneck, assuming perfectly specified prompts that do not reflect the iterative nature of requirement elicitation. To bridge this gap, we introduce ClarifyCodeBench, a novel interactive benchmark for evaluating LLMs’ capability in resolving requirement ambiguity. Constructed from real-world programming tasks, ClarifyCodeBench features high-quality manual annotations, including N unique ambiguity types, associated clarification questions, and corresponding ground-truth answers. Furthermore, we formalize two rigorous metrics to assess the interaction quality: Turn-discounted Key Question Rate, which penalizes inefficient questioning, and Optimal Round Adherence, which measures the precision of the elicitation process. We conduct a systematic evaluation of six state-of-the-art LLMs using ClarifyCodeBench. Our empirical results yield three critical insights: 1) Capability Decoupling: Strong code generation performance does not inherently translate to effective requirement clarification; 2) The Reasoning Paradox: While increased computational thinking enhances code correctness, it yields marginal gains in identifying ambiguities; 3) The Multi-ambiguity Ceiling: LLMs’ clarification performance degrades sharply as the density of ambiguities increases, revealing a significant bottleneck in handling complex, real-world specifications. Our work underscores the necessity for future AI4SE research to transition from static synthesis to interactive elicitation.

[AI-173] When a Correct Reward Is Not Enough: Diagnosing and Guiding PPO in an Analytically Solved Broker-Trader Game

链接: https://arxiv.org/abs/2610.03598
作者: Siu Tung Wong(1),Carlo Campajola(1 and 2) ((1) Institute of Finance and Technology, University College London, (2) UZH Blockchain Center)
类目: Trading and Market Microstructure (q-fin.TR); Artificial Intelligence (cs.AI)
备注: 8 pages; accepted for publication at ICAIF 2026

点击查看摘要

Abstract:Reinforcement learning (RL) is increasingly used for financial optimal-control problems when complex dynamics make analytical strategies difficult to obtain. There are financial mathematics literactures which provides many solved models whose equations and controls could evaluate and guide learning; we ask whether RL can exploit these results. We place a proximal policy optimisation (PPO) agent in an analytically solved continuous-time broker–trader game. PPO replaces the broker and chooses its trading speed while interacting with an informed trader and stochastic uninformed order flow. We derive a finite-step reward from the broker’s continuous-time payoff and verify its discrete implementation through grid refinement and an exact one-step identity. With zero uninformed flow, a validation-selected PPO–FFNN approaches the reference action. With stochastic uninformed flow, the tested PPO–FFNN and PPO–LSTM remain inaccurate, although supervised learning confirms that their actors can represent the action. Monte Carlo diagnostics show that their critics do not reliably rank nearby actions; potential-based reward shaping also gives no reliable improvement. Under partial information, a causal certainty-equivalent controller based on the broker’s observable history remains close to the reference, while PPO has larger errors and lower payoffs. Finally, we freeze the analytical policy and train PPO to adjust it after the execution cost changes. Halving the cost yields a repeatable improvement that closes (2.22%) of the gap to the changed-cost reference. The analytical solution therefore provides both a benchmark for diagnosing RL and a useful starting policy for adaptation. Comments: 8 pages; accepted for publication at ICAIF 2026 Subjects: Trading and Market Microstructure (q-fin.TR); Artificial Intelligence (cs.AI) Cite as: arXiv:2610.03598 [q-fin.TR] (or arXiv:2610.03598v1 [q-fin.TR] for this version) https://doi.org/10.48550/arXiv.2610.03598 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[AI-174] AREX: Affine-Residual Exponential Integrator for Few-Step Sampling in Flow Matching

链接: https://arxiv.org/abs/2610.03483
作者: Shizheng Lin,Soon Hoe Lim,N. Benjamin Erichson
类目: Machine Learning (stat.ML); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: 53 pages

点击查看摘要

Abstract:We introduce AREX, a training-free sampler for pretrained flow matching models that uses the target mean and covariance to capture an analytically tractable part of the sampling dynamics. We show that the velocity field of the moment-matched Gaussian target is the L^2 -optimal affine approximation to the marginal velocity field. This motivates decomposition of the learned dynamics into an affine component over the whole sampling path, determined by the first two target moments, and a neural residual term. AREX keeps the affine component and integrates it using an explicit matrix-valued propagator. In turn, we only require to integrate over the residual term. This differs from scalar exponential integrators, which analytically handle only isotropic linear dynamics. Across image and text-to-image generation tasks, AREX consistently improves sample fidelity in the few-step sampling regime without retraining the underlying model.

[AI-175] Information Limits of Low-Rank Approximation Certification

链接: https://arxiv.org/abs/2610.03321
作者: Kang Liu,Bohao Qu
类目: Optimization and Control (math.OC); Artificial Intelligence (cs.AI); Information Theory (cs.IT)
备注:

点击查看摘要

Abstract:Low-rank approximation can require additional matrix–vector products to verify that its error meets a prescribed tolerance. We characterize this certification cost for both relative matrix error and mean-square output error. For a single approximation matrix candidate, we determine the exact dimension-uniform minimax query constant as the allowed failure probability vanishes. Our main result concerns reusing validation responses as the approximation space expands. For a candidate family constructed independently of validation, one batch supports an entire nested path without increasing the query budget with the number of checks. Across (W) paths, a concentration bound exploiting shared residual energy yields a (\sqrt\log(W+1)) dependence. A matching lower bound establishes its optimality for fixed interior error targets and sufficiently small separation gaps. Finally, we compare two uniformly valid certificates on the same dispersed-spectrum family. Optimizing the validation budget within each rule family yields costs of orders (N^1/3) and (N^2/3) for validation and construction beyond the true target. Code is available at this https URL

[AI-176] Multimodal reasoning for broadly neutralizing antibody discovery from label-free human B cell repertoires across virus families

链接: https://arxiv.org/abs/2610.03160
作者: Hantao Lou,Jianqing Zheng,Can Yue,Meihan Zhang,Yuanchao Bao,Yu Chen,Mengting Huang,Yupeng Yang,Qianyu Pan,Nana Fu,Yansong Shi,Hongli Li,Yangyang Chai,Ruyi Chen,Wansheng Li,Zhu Liang,Rongmei Yao,Yuanhan Mo,Lei Wang,Chunmei Wang,Yun Quan,Qiong Zhang,Xiangxi Wang,Xuetao Cao
类目: Quantitative Methods (q-bio.QM); Artificial Intelligence (cs.AI); Computational Engineering, Finance, and Science (cs.CE); Cell Behavior (q-bio.CB)
备注:

点击查看摘要

Abstract:Discovering broadly neutralizing antibodies (bnAbs) from human natural immune repertoires remains a fundamental challenge in immunology, hindered by: the extreme rarity of bnAb, incomplete understanding of their cellular origins across pathogens, and the inability of existing computational tools to generalize across emerging viral threats. Here we present ImmuneAgent, a closed-loop AI system that integrates multimodal reasoning with continual meta-learning and wet-lab feedback to overcome these barriers. Applied to screen the natural BCR repertoires from vaccinated or infected cohorts, the system achieves a ~55% neutralization antibody discovery rate (60 of 110 cloned candidates) and a ~11% bnAb yield (12 of 110), substantially outperforming a state-of-the-art sequence-based neutralization predictor or cofolding models evaluated at the same cloning budget. Five ImmuneAgent-discovered antibodies conferred 100% in vivo protection against lethal influenza challenge, comparable to the clinical-stage therapeutic MEDI8852. The system recovered the cellular and structural determinants of bnAb activity and identified FCRL5+CD27+ atypical memory B cells as a conserved bnAb reservoir and hydrophobic interface enrichment as a cross-viral structural signature, which generalized to unseen antigens, discovering human metapneumovirus (hMPV) cross-neutralizing and human papillomavirus (HPV)-neutralizing antibodies without antigen-specific sorting. These results validate that ImmuneAgent is a generalizable framework for rapid therapeutic antibody discovery against emerging viral threats.

[AI-177] S2S-JEPA: Predicting the Predictable at Subseasonal-to-Seasonal Timescales

链接: https://arxiv.org/abs/2610.03106
作者: Chenyu Dong,Gianmarco Mengaldo
类目: Atmospheric and Oceanic Physics (physics.ao-ph); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:The subseasonal-to-seasonal (S2S) timescale, roughly from two weeks to two months ahead, is a critical forecast window for sectors such as agriculture, energy, and water management. Yet, it is widely known as the `predictability desert’. Recent AI weather models excel up to two weeks ahead but deteriorate beyond, largely because they are trained to predict fine-scale details that are neither predictable nor essential at S2S timescales. We argue that a more physically grounded objective is to forecast only the slowly varying components that remain predictable. Computer vision reached the same conclusion with the Joint-Embedding Predictive Architecture (JEPA), which predicts in latent space, discarding unpredictable details. In this work, we introduce S2S-JEPA, which brings the JEPA paradigm to S2S forecasting. It is tailored to this task through design elements from state-of-the-art AI weather models. S2S-JEPA achieves comparable skill to the gold-standard ECMWF physics-based ensemble and surpasses it on multiple metrics at weeks 5 to 6.

[AI-178] Generalization Properties of Score-matching Diffusion Models for Intrinsically Low-dimensional Data

链接: https://arxiv.org/abs/2610.02663
作者: Saptarshi Chakraborty,Quentin Berthet,Peter L. Bartlett
类目: Machine Learning (stat.ML); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Statistics Theory (math.ST)
备注:

点击查看摘要

Abstract:Despite the remarkable empirical success of flow-matching models, their statistical generalization guarantees remain underdeveloped. Existing analyses often impose restrictive assumptions on the estimated velocity field and yield convergence rates that fail to reflect the intrinsic low-dimensional structure common in real data, such as natural images and molecular geometries. In this work, we study the statistical generalization of flow-matching models for learning an unknown distribution P_\mathrmdata from finitely many samples. We derive finite-sample error bounds on the learned generative distribution, measured in the Wasserstein- p distance, for all p\geq 1 . Specifically, given n i.i.d. samples from P_\mathrmdata , we show that, for every dd_p^\ast(P_\mathrmdata) and appropriately chosen network architectures and hyperparameters, the learned distribution \widehatP^\mathrmFM satisfies \mathbbW_p(\widehatP^\mathrmFM,P_\mathrmdata) \lesssim n^-1/d+n^-1/(2p)\bigl(\log(1/\xi)\bigr)^1/(2p) with probability at least 1-\xi , where d_p^\ast(P_\mathrmdata) denotes the Wasserstein- p dimension of the target measure. Our results demonstrate that flow matching naturally adapts to the intrinsic geometry of data and mitigates the curse of dimensionality, as the convergence exponent depends on the intrinsic rather than ambient dimension. These guarantees remain meaningful in high-dimensional regimes and provide a theoretical explanation for the empirical success of flow matching on structured data distributions under substantially more relaxed assumptions than those in existing analyses.

[AI-179] Equivariant Flow Matching for Electron Density Prediction

链接: https://arxiv.org/abs/2610.02651
作者: Chenxing Liang,Chengdong Wang,Yuchao Lin,Xiaofeng Qian,Shuiwang Ji
类目: Chemical Physics (physics.chem-ph); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Machine learning surrogates for density functional theory (DFT) have been increasingly used to reduce the cost of first-principles calculations. In this arena, predicting real-space electron densities offers a scalable and transferable initialization for self-consistent field (SCF) procedures. However, current methods face a clear dilemma. That is, grid-based architectures incur a high computational cost, while basis-set methods fail to capture the structural correlations inherent in the coefficient space. Here, we develop OrbFlow, an \mathrmSE(3) -equivariant generative model that predicts Gaussian-type orbital (GTO) coefficients via flow matching. OrbFlow retains the efficiency of a compact atom-centered basis while replacing pointwise regression with a learned probability path over the full coefficient space. It is trained through a two-phase trajectory curriculum that mitigates discretization drift during numerical integration. OrbFlow achieves state-of-the-art accuracy on QM9, reducing density error by 13.6% relative to the previous best model, and reduces error by 51% to 63% on every molecule of the MD benchmark relative to the strongest prior method sharing its basis. The predicted density also cuts SCF iterations by up to 68% with zero-shot transfer to unseen exchange-correlation functionals and recovers dipole and quadrupole moments to within a few percent of DFT references without any SCF calculation.

[AI-180] IGNITE Tokamak World Model Architecture

链接: https://arxiv.org/abs/2610.02515
作者: Peter Steiner,Azarakhsh Jalalvand,Nathaniel Chen,Kouroche Bouchiat,Ricardo Shousha,SangKyeun Kim,Egemen Kolemen
类目: Plasma Physics (physics.plasm-ph); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:We introduce IGNITE, a generative world foundation model for fusion plasma behavior simulation trained in a self-supervised manner from over a decade of unlabeled experimental data at the DIII-D National Fusion Facility. The core of IGNITE is a dynamics model that can simulate DIII-D discharges from a given set of actuator trajectories. These trajectories can be supplied or generated on-the-fly from a textual prompt or from desired experimental outcomes. The model architecture consists of several spatio-temporal tokenizers that embed the different input modalities, including time-series like spatio-temporal measurement data, image sequences, and high-resolution spectrograms, each of which collected at vastly different time scales. The backbone is composed of an auto-regressive dynamics model that has the capacity to predict entire DIII-D discharges given initial latent plasma states and actuator trajectories over a theoretical infinite horizon. IGNITE paves the way towards efficient AI-driven experimental planning and world modeling for nuclear fusion.

[AI-181] Learning Style Forgetting Semantics: A Case Study of SFT and RFT on Classification Tasks

链接: https://arxiv.org/abs/2610.02437
作者: Haodong Liang,Yanhao Jin,Krishnakumar Balasubramanian,Lifeng Lai
类目: Machine Learning (stat.ML); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: 43 pages, 7 figures

点击查看摘要

Abstract:Why does supervised fine-tuning (SFT) lead to more forgetting than reinforcement fine-tuning (RFT), even when all teacher demonstrations are semantically correct? We study this question on classification tasks where tokens within each semantic class express the same semantic answer in different styles. The tasks share an underlying semantic rule but differ in their prompt distributions and teachers’ stylistic preferences. Using a tractable linear-softmax policy, we derive an exact decomposition of the updates into semantic and style components. We show that, at a common policy and prompt, SFT and RFT have parallel semantic updates but differ in their style dynamics. Starting from a policy with no within-class style preference, RFT with exact policy gradients preserves this symmetry, whereas SFT with a nonuniform teacher develops off-axis style drift along a nonzero task mean under population updates. We use this drift to establish a separation under explicit conditions: for population updates from a common perfectly fitted checkpoint, SFT forgetting admits a strictly positive lower bound over a finite training interval, while RFT retains zero semantic error. Simulations over task sequences support these theoretical predictions.

[AI-182] RxnOptBench: Benchmarking LLM s for Reaction-Condition Optimization in Organic Methodology NEURIPS2026

链接: https://arxiv.org/abs/2610.02242
作者: Lingli Ge,Yubin Wang,Junyuan Gao,Jiahe Song,Jiaxing Sun,Boyu Zhu,Haote Yang,Jingchao Wang,Lixin Ma,Jiang Wu,Yuqiang Li,Conghui He
类目: Chemical Physics (physics.chem-ph); Artificial Intelligence (cs.AI)
备注: Accepted to NeurIPS 2026 (Evaluations Datasets Track)

点击查看摘要

Abstract:Chemical reaction-condition optimization – choosing the catalyst, ligand, solvent, reagent, temperature, time, and atmosphere that jointly maximize yield and stereoselectivity – is a central, judgement-laden subtask of organic methodology research that large language models are increasingly expected to support. Yet existing chemistry benchmarks evaluate reaction-class labelling, retrosynthesis, or SMILES manipulation, and do not ask models to read a real condition-screening table and pick the best set. We introduce RxnOptBench, a benchmark whose every option and precedent is a real wet-lab entry mined from the optimization tables of organic-methodology papers published in 2025, graded by a continuous relative score derived from a declared headline utility that combines reported yield with enantiomeric excess (ee), diastereomeric ratio (dr), and regioisomeric ratio (rr), and equipped with a paired precedents-vs-no-precedents design that isolates in-context use of literature evidence from parametric memorization. Across nine frontier LLMs and three Chemistry LLMs, even the best models leave substantial headroom: chemistry-specialized models fall to the random-baseline floor on multi-axis selection, while open-weight models have closed most of the gap to proprietary frontier models. We release the final human-reviewed benchmark test set and evaluation code.

[AI-183] Causal discovery identifies pathways linking physical activity to dementia risk in the UK BioBank

链接: https://arxiv.org/abs/2610.02221
作者: Wasif Khan,Panayiotis V. Benos,Joshua K. Wong,Ruogu Fang
类目: Neurons and Cognition (q-bio.NC); Artificial Intelligence (cs.AI)
备注: Under submission

点击查看摘要

Abstract:Physical Activity (PA) is consistently associated with lower risk of dementia, yet the mechanism linking PA to dementia prevention remain incomopletely understood. Here, we integrate large language model (LLM)-guided causal discovery with mediation analysis in 42,293 older adults aged 60 years or older from the UK Biobank to systematically identify pathways connecting objectively measured moderate-to-vigorous physical activity (MVPA) to dementia risk. Across behavioral, psychological, functional, and clinical domains, causal discovery consistently identified interconnected pathways linking higher MVPA to lower dementia risk through depression, functional capacity, smoking behavior, hypertension, cardiovascular disease, chronic kidney disease, and brain injury. Chain mediation analyses further identified depression as a central pathway, accounting for 15.1% of the overall association between MVPA and dementia risk. Sex-stratified analyses revealed distinct mechanistic patterns, with females exhibiting predominantly metabolic and functional pathways, while males showed behavioral and cardiometabolic cascades involving smoking and cardiovascular disease. Together, these findings suggest that the protective association between PA and dementia is mediated through interconnected behavioral, psychological, and cardiometabolic processes rather than a single pathway. Depression emerged as a prominent and potentially modifiable pathway, highlighting opportunities for integrating dementia prevention strategies that combine PA promotion with mental health and cardiovascular risk management.

[AI-184] Multi-Modal Environment-Aware Beam Management for Massive MIMO: A Geometry-Driven Virtual Base Station Framework

链接: https://arxiv.org/abs/2606.26567
作者: Yijie Bian,Wei Guo,Jie Yang,Shenghui Song,Jun Zhang,Shi Jin,Khaled B. Letaief
类目: ignal Processing (eess.SP); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:High-frequency massive multiple-input multiple-output (MIMO) systems promise ultra-high data rates. However, efficient beam management remains challenging due to the prohibitive beam training overhead and intricate coordination required in multi-user MIMO (MU-MIMO) scenarios. To address these bottlenecks, environment-aware communications have emerged as a promising paradigm, leveraging site-specific knowledge to circumvent exhaustive pilot-based beam training and streamline multi-user communications. In this paper, we propose an interpretable and geometry-driven framework that utilizes multi-modal environmental data, specifically regional 3D light detection and ranging (LiDAR) point clouds and location information, to construct an offline virtual base station (VBS) database. By modeling dominant reflection paths via mirror symmetry across building facades reconstructed from the point clouds, the VBS database provides a compact and sparse description of the wireless propagation environment. To bridge the semantic gap between geometric information and wireless channels, we develop a coarse channel reconstruction mechanism that estimates channel parameters directly from VBS-derived geometric relationships. Based on the resulting coarse beamspace representation, we design a VBS-assisted orthogonal-pilot (VOP)-based partial beam training scheme to refine the coarse estimates with minimal online training overhead. Finally, to tackle the combinatorial beam selection problem and manage inter-user interference, we propose a hierarchical deep reinforcement learning framework, namely a dual-agent dueling double deep Q-network, for coordinated beam selection (DD3QN-CBS). Simulation results demonstrate consistent gains in both beam training efficiency and beam selection performance over heuristic and learning-based baselines.

机器学习

[LG-0] RNADyn: A Benchmark for Generating and Understanding RNA Dynamics

链接: https://arxiv.org/abs/2610.03712
作者: Yiming Huang,Lennart Bastian,Hanqun Cao,Luis Vollmers,Tolga Birdal
类目: Machine Learning (cs.LG); Biological Physics (physics.bio-ph); Biomolecules (q-bio.BM)
*备注:

点击查看摘要

Abstract:Ribonucleic acid (RNA) functions through conformational changes that are not fully captured by static structures. However, large-scale standardized RNA dynamics data remain limited, and existing approaches typically treat trajectory generation and dynamics understanding as separate objectives. Here, we introduce RNADynBench, a standardized RNA molecular dynamics (MD) benchmark with 2585 quality-controlled 100-ns all-atom trajectories and leakage-controlled splits. Building on RNADynBench, we develop RNADynNet, a unified model for RNA dynamics learning that uses a shared backbone for both trajectory generation and dynamics fingerprint extraction from a single conformer. It combines coordinate denoising, single-frame-to-trajectory alignment, and physical grounding to connect all-atom trajectory generation with dynamics representation learning. Physical grounding improves both generated dynamics and the physical information recoverable from these fingerprints. Across both test sets, including the high-flexibility challenge set, the generated trajectories achieve RMSF correlations of 0.875 and 0.766, while single-conformer predictions show comparable agreement with MD-derived dynamics. RNADynBench and RNADynNet together establish a benchmark and unified modeling framework for generating and understanding RNA dynamics.

[LG-1] LESSER: Post-Training Data Selection with Output-Layer Gradients

链接: https://arxiv.org/abs/2610.03702
作者: Lyuxin David Zhang,Eric Wong,Surbhi Goel,Anton Xue
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:The choice of post-training data for large language models substantially affects downstream performance. Gradient-based data selection is a popular approach that ranks training data by how well their gradients align with those of a small validation set. However, ranking with full-parameter gradients requires an expensive backward pass on every sample, making computation intractable for large candidate pools. This raises a natural question: can we approximate full-gradient features at a fraction of the cost? Conveniently, we find that output-layer gradients suffice for effective data selection, yet require only the cheaper forward pass. We implement this as LESSER, a drop-in wrapper for selection methods that reduces the feature-extraction FLOP cost by 9.7\times for SFT and 3.0\times for RL benchmarks, while tracking full-gradient performance on downstream tasks. Empirically, we find that even when output-layer and full gradients rank individual samples differently, they select batches with aligned gradients.

[LG-2] Simulation-Free Learning of Population Dynamics with Wasserstein Lagrangian Residuals

链接: https://arxiv.org/abs/2610.03679
作者: Fedor Sergeev,Markus Heinonen,Daniel Waxman,Tim Cooijmans,Ricardo Baptista,Dmitry Batenkov,Eli Bingham
类目: Machine Learning (cs.LG); Methodology (stat.ME); Machine Learning (stat.ML)
*备注: 34 pages, 11 figures

点击查看摘要

Abstract:The dynamics of cells, organisms, and fluids are often modeled as probability distributions evolving over time. Reconstructing and extrapolating this evolution from unpaired snapshots requires assumptions about the underlying process. Wasserstein gradient flows are a common choice, but they cannot describe conservative or periodic dynamics. Lagrangian mechanics in Wasserstein space covers both, but existing methods for learning it are simulation-based: they run a numerical solver at every training step, which makes training expensive. We propose Double-Stitch, a simulation-free method that learns these mechanics by penalizing the residual of the equation of motion along a learned population path. We derive this equation from a Clebsch variational principle that does not require gradient velocities, and show that the residual vanishes exactly when the equation holds. We test Double-Stitch on synthetic, single-cell and ocean vortex datasets and find that it matches or outperforms gradient-flow methods and simulation-based WLM on most tasks, while training 4 - 14 times faster than WLM. We provide a JAX implementation of Double-Stitch at this https URL.

[LG-3] Planning to Learn

链接: https://arxiv.org/abs/2610.03667
作者: Ian Osband
类目: Machine Learning (cs.LG); Optimization and Control (math.OC); Machine Learning (stat.ML)
*备注:

点击查看摘要

Abstract:Policy-gradient methods are central to modern reinforcement learning, including LLM post-training. When they struggle, the usual suspects are exploration, credit assignment and action-sampling noise. Classification has none of them. A classifier is a policy whose expected reward, its \emphexpected accuracy, is the probability it assigns to the correct label, and because that label is known, the policy gradient is exact and smooth. Yet exact policy gradient loses to cross-entropy, even on expected accuracy. The exact gradient is myopic: it values an update only by what it buys now, but each update also sets where the next one starts, so an update’s value depends on how much learning remains. Viewed this way, cross-entropy is patient accuracy, the total error an example would pay if its log-odds rose at unit speed forever, while exact policy gradient is the zero-horizon limit. Truncating this total at the learning that remains yields the horizon loss, a one-line change that moves from cross-entropy toward exact policy gradient as training runs out. In a simple allocation model, it provably escapes the trap that catches each endpoint. On MNIST and on ImageNet with ResNet-50, ResNet-101 and ViT-S/16, the horizon loss improves top-1 accuracy over cross-entropy at a flat learning rate, and the gain grows with label noise.

[LG-4] Forecasting from Counterfactual Simulator Rollouts: A Sim2Real Evaluation

链接: https://arxiv.org/abs/2610.03662
作者: Angel Wang,Dominique Perrault-Joncas,Alvaro Maggiar,Dean Foster,Carson Eisenach
类目: Machine Learning (cs.LG)
*备注: 15 pages, 3 figures, 9 tables

点击查看摘要

Abstract:Deploying a new decision policy creates a cold-start problem for prediction models whose targets depend on the policy’s actions: historical observations reflect earlier policies, while real observations under the new policy are not yet available. Simulation offers a way to address this gap by rolling out the target policy across counterfactual scenarios and using the resulting trajectories to learn how the system responds to those controls. The simulation-to-reality (Sim2Real) transfer of this simulator-trained model can then be backtested by evaluating it against real observations from past deployments. Using two real-world inventory-control deployments, we evaluate this process from three angles: simulator fidelity, zero-shot transfer to real behavior, and adaptation as real target-policy observations accumulate. The simulator-trained forecaster achieves lower point-estimate mean absolute percentage error (MAPE) than the same architecture trained on historical real data, reducing MAPE by 1.2-3.1 percentage points in Study 1 and 12.5-18.7 points in Study 2. After deployment, lightweight calibration using early real observations further reduces error by up to 2.5 percentage points. These results provide empirical evidence that simulator-generated counterfactual data can support cold-start forecasting under a new policy, and the resulting model can be further refined as real deployment data become available.

[LG-5] PoCoFL: POlicy-COmpliant Federated Learning

链接: https://arxiv.org/abs/2610.03650
作者: Dominik Roy George,Varesh Mishra,Aysajan Abidin
类目: Cryptography and Security (cs.CR); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Federated Learning (FL) is a privacy-oriented learning paradigm that enables collaborative model training while keeping training data local to participating clients. However, it does not guarantee that clients submit policy-compliant contributions or that aggregators process admitted contributions correctly. Existing verifiable FL systems tailor validation rules to specific FL settings, learning workflows, and cryptographic constructions, limiting their applicability across network topologies, participant roles, and aggregation semantics. In this paper, we present PoCoFL, a policy-compliant federated learning framework that separates three aspects: (i) FL type, (ii) policy semantics, and (iii) cryptographic realisation. We provide a formalisation that captures client and aggregation requirements as policy-dependent relations. Clients prove compliance of their contributions using commitments and non-interactive zero-knowledge proofs, while aggregators prove that the recorded set of admitted contributions was processed according to the selected aggregation policy. We demonstrate PoCoFL through four formal instantiations: (i) vanilla, (ii) continual, (iii) personalised, and (iv) threshold-encrypted federated learning. We evaluate the effects of policy enforcement on the learning objectives of vanilla, personalised, and continual FL. We further implement proof-of-concept realisations of all four instantiations, demonstrating the versatility and practical feasibility of PoCoFL. Overall, these results show that PoCoFL can capture complex policy representations while remaining network-topology agnostic.

[LG-6] When May a Bandit Leave Its Anchor? E-Process-Authorized Thompson Sampling under Non-stationarity NEURIPS2026

链接: https://arxiv.org/abs/2610.03646
作者: Mayand Gulati,Kerong Wang,WeiChen Au
类目: Machine Learning (cs.LG); Machine Learning (stat.ML)
*备注: 25 pages, 3 figures. Accepted to the E-Values Workshop at NeurIPS 2026 (poster)

点击查看摘要

Abstract:Stationarity rewards memory, but after a change the same history can mislead. We ask when forgetting should be permitted. E-process-authorized Thompson sampling (e-ATS) gives each arm full-history and discounted Beta states. An anytime-valid e-process first authorizes the discounted state, then a reversible relevance score controls its influence. Before authorization, e-ATS exactly follows optimistic Thompson sampling (OTS). Under a Beta-Bernoulli prior-predictive stationary model, e-ATS’s probability of ever departing from OTS is at most the chosen \alpha_E , without fitted thresholds. Relative to e-ATS, removing authorization increased mean normalized dynamic pseudo-regret by 38.4% on the registered suite but reduced it by 7.5% on the literature-derived replay suite. Therefore, evidence controls when adaptation begins, not whether it always helps.

[LG-7] On the Convergence of Success Conditioning for Policy Optimization

链接: https://arxiv.org/abs/2610.03642
作者: Matthew Brun,Xu Andy Sun
类目: Machine Learning (cs.LG); Optimization and Control (math.OC)
*备注:

点击查看摘要

Abstract:Success conditioning is a strategy for improving decision-making policies in stochastic environments; it updates a policy by increasing the probability of taking actions that yield successful outcomes. Success conditioning is common to many reinforcement learning applications, yet its limiting behavior and convergence rates are not well understood. In this work, we demonstrate that success conditioning converges to an optimal policy on a broad class of Markov decision processes (MDPs). We also derive convergence rates in some common settings. For discounted MDPs, we prove convergence within \mathcalO(1/\varepsilon^p) iterations to an \varepsilon -optimal policy, where the exponent p depends on problem data. For single-period MDPs, such a policy is obtained within \mathcalO(\log(1/\varepsilon)) iterations.

[LG-8] IDRF: Inverse-Distilled Reward Fine-tuning of Masked Discrete Diffusion Models

链接: https://arxiv.org/abs/2610.03641
作者: Vladislav Gromadskii,David Li,Samson Gourevitch,Yazid Janati,Eric Moulines,Maxim Panov,Alexander Korotin
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Masked discrete diffusion models offer a promising alternative to autoregressive generation, but iterative sampling can be costly, and intractable sequence likelihoods complicate reward fine-tuning. We introduce IDRF, a framework for reward fine-tuning of few-step masked discrete diffusion generators. Starting from a standard reverse-KL-regularized objective, IDRF replaces the intractable sequence-level KL penalty with inverse-distillation regularization. With an optimal auxiliary denoiser, we prove that the population inverse-distillation loss upper-bounds the sequence-level KL divergence to the reference distribution. IDRF optimizes a trajectory-based surrogate of this loss without reference-model rollouts, so the student keeps its own few-step sampler. We view few-step generation as a finite-horizon Markov decision process and optimize reward with a clipped policy-gradient objective over the student’s trajectories. Across DNA, image, and text generation, IDRF achieves high reward with up to 32\times fewer denoising steps than the reference while mitigating reward hacking and preserving sample quality.

[LG-9] Broken scale symmetries in undercomplete linear autoencoders NEURIPS2026

链接: https://arxiv.org/abs/2610.03640
作者: Farhad Pashakhanloo,Jacob A. Zavatone-Veth
类目: Machine Learning (cs.LG); Neurons and Cognition (q-bio.NC); Machine Learning (stat.ML)
*备注: NeurIPS 2026 Symmetry and Geometry in Neural Representations Workshop

点击查看摘要

Abstract:Neural network loss landscapes have many symmetries, which are preserved by gradient flow but broken by finite-stepsize stochastic gradient descent (SGD). A canonical example of such a symmetry is scale in homogeneous networks: one can scale up the parameters in one layer and down in the next without changing the network output. Previous work has documented cases in which SGD breaks this symmetry in favor of balancing gradient noise or minimizing fluctuations. Here, we show that the solution geometry of undercomplete linear autoencoders instead selects a preferred sign for scale drift: on the PCA solution manifold, SGD favors large decoder weights. This directed scale drift occurs on a slow timescale, and its dynamics admit an analytically-tractable effective description. However, it cannot continue indefinitely: increasing scale eventually drives the dynamics towards a finite-stepsize stability boundary. The resulting solutions are sharper than a balanced baseline in the sense of the maximum eigenvalue of the loss Hessian, but different sharpness measures can move in opposing directions. Thus, undercomplete autoencoders give a concrete illustration of how loss geometry can convert residual gradient noise into directed motion along a manifold of functionally-equivalent solutions.

[LG-10] Normal-Form Correlation in Markov Games

链接: https://arxiv.org/abs/2610.03621
作者: Ioannis Anagnostides,Constantinos Daskalakis,Gabriele Farina,Noah Golowich,Tuomas Sandholm,Brian Hu Zhang
类目: Computer Science and Game Theory (cs.GT); Computational Complexity (cs.CC); Data Structures and Algorithms (cs.DS); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:There has been a surge of recent work on correlated equilibrium concepts in Markov games. However, existing results focus on concepts weaker than normal-form correlated equilibria (NFCEs), leaving open the more challenging question of computing such equilibria, which goes back to the seminal work of Papadimitriou and Roughgarden (JACM’08). Here, we establish the first efficient algorithm for NFCEs in finite-horizon Markov games with a fixed number of players n . In particular, with S states, horizon H , and at most A actions per player, it computes an \epsilon -NFCE in time S(AH/\epsilon)^O(n) . This is the first algorithm polynomial in 1/\epsilon and the description of the game for NFCEs in an interesting class of problems beyond the normal-form setting. Moreover, under the usual assumption that recommendations are independent across states, we show PPAD-completeness—that is, computational equivalence to Nash equilibria—either in many-player games or when the precision is exponentially small. The key idea behind our approach is to run backward induction on a sequence of auxiliary stage games, but with the twist that in each step we compute a constant-expectation correlated equilibrium. This is a natural refinement of correlated equilibrium in which the conditional expected payoff from obeying is independent of the recommendation. In fact, our reduction goes both ways, establishing an equivalence between constant-expectation CEs and NFCEs in Markov games. For a fixed number of players, we observe that a constant-expectation CE can be computed approximately by combining linear programming with suitable discretization. In contrast, it is PPAD-hard in i) polymatrix (many-player) games at constant precision, and ii) two-player games at exponentially small precision. The latter result follows from an unexpected connection to rank-2 two-player games. Subjects: Computer Science and Game Theory (cs.GT); Computational Complexity (cs.CC); Data Structures and Algorithms (cs.DS); Machine Learning (cs.LG) Cite as: arXiv:2610.03621 [cs.GT] (or arXiv:2610.03621v1 [cs.GT] for this version) https://doi.org/10.48550/arXiv.2610.03621 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[LG-11] UniIntervene: An Adaptive Intervention Agent for Efficient Real-World Reinforcement Learning

链接: https://arxiv.org/abs/2610.03620
作者: Yudong Lin,Haoyuan Deng,Zhuoxuan Yuan,Zaijia Yang,Yuanjiang Xue,Ziwei Wang
类目: Machine Learning (cs.LG); Robotics (cs.RO)
*备注: Yudong Lin and Haoyuan Deng contributed equally. Ziwei Wang is the corresponding author. Code is available in our \href{ this https URL }{GitHub repository}

点击查看摘要

Abstract:Online reinforcement learning (RL) enables robot policies to improve through physical interaction, but the assistance they require changes as their competence evolves. Existing intervention strategies based on offline estimates or fixed decision rules can therefore become mismatched to the current policy. To address this, we propose UniIntervene++, an adaptive intervention agent that learns to allocate control between autonomous execution and heterogeneous assisted behaviors during online RL. Specifically, UniIntervene++ first formulates the evolving RL policy, trajectory correction, and a task-structured CodePolicy as Options in a unified semi-Markov decision process and learns their relative values online. Building on this, competence-adaptive intervention periodically probes the RL policy through unassisted execution, keeping control allocation responsive to its evolving capability. Finally, coupled experience learning allows assisted behaviors to improve the RL policy, whose evolving outcomes in turn reshape future intervention decisions. In this way, UniIntervene++ jointly determines when to intervene, how to intervene, and when to return control as the RL policy improves. Across five real-world manipulation tasks, UniIntervene++ achieves an average success rate of 89.67%, outperforming all baselines by at least 6 percentage points, while reducing human intervention to 0.77%, a relative reduction of at least 94.6% from the best baseline. Code is available in our \hrefthis https URLGitHub repository.

[LG-12] Mastering Atari 2600 Games with Discovered Options

链接: https://arxiv.org/abs/2610.03604
作者: Erik M. Lintunen,Marlos C. Machado
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Temporal abstractions, often instantiated as options, have long been regarded as a mechanism for accelerating credit assignment, facilitating exploration, and enabling generalisation in reinforcement learning (RL). However, developing general option discovery methods that are effective in large-scale, high-dimensional domains remains a fundamental challenge. Existing option discovery methods are either confined to relatively simple domains, depend on handcrafted or quasi-symbolic representations, or offer little improvement over learning without options. We present Wayfarer, a general, domain-agnostic, online deep RL agent that discovers options through Laplacian representation learning from high-dimensional observations and leverages them for control. We show that the resulting options simultaneously improve exploration, accelerate credit assignment, and generalise effectively to unseen settings, enabling substantially faster learning of complex policies. Wayfarer achieves state-of-the-art performance among single-stream agents on the most challenging Atari 2600 games, with the largest gains in games that require long-horizon exploration and strategic behaviour, such as Montezuma’s Revenge and Private Eye.

[LG-13] A Path Integral Surrogate for Multi-Step Gradient Inversion in Federated Learning ICASSP2027

链接: https://arxiv.org/abs/2610.03597
作者: Agnivo Ghosh,Saumik Bhattacharya
类目: Machine Learning (cs.LG); Distributed, Parallel, and Cluster Computing (cs.DC)
*备注: 5 pages, 2 figures, 3 tables. Submitted to IEEE ICASSP 2027

点击查看摘要

Abstract:Federated learning lets many clients train a shared model together without ever sending their private data to a central server. Each client shares only a model update, and this update should reveal far less about the client than its raw training examples would. This premise is what protects the privacy of the clients. Gradient inversion attacks challenge it directly by trying to reconstruct a client’s private input images from the single update it shared. Under FedAvg, a client’s update accumulates several local training steps, so the server sees only the two endpoints of a hidden weight trajectory. Recent gradient inversion attacks fit a surrogate model along the path between these two endpoints but they still read its gradient at a single point. We propose the Path-Integral Surrogate Model Extension (PI-SME) which treats the accumulated update as a path integral of the gradient field and approximates it by Gauss–Legendre quadrature over several nodes along a learnable Bézier path. On CIFAR-100 and FEMNIST images across a range of trajectory lengths and class-restricted batches PI-SME reconstructs the private inputs more faithfully than the strongest surrogate baseline on several inversion metrics and the matching loss.

[LG-14] Get a GRIP this will be a long TRIP: A Quantifiable Long-Range Framework for Verifying Over-squashing NEURIPS2026

链接: https://arxiv.org/abs/2610.03556
作者: Ferran Hernandez Caralt,Simon Heilig,Adrián Bazaga,Asja Fischer,Moshe Eliasof,Pietro Liò
类目: Machine Learning (cs.LG)
*备注: Published at the Conference on Neural Information Processing Systems (NeurIPS 2026). Track on Evaluations and Datasets

点击查看摘要

Abstract:Empirical claims about the connection between over-squashing and long-range interactions in GNNs, can only be trusted if the benchmarks used to validate them genuinely require long-range interactions. The de-facto standard, the Long Range Graph Benchmark, has been repeatedly shown to be saturated by tuned short-range models, with existing synthetic alternatives being tied to specific topologies. As such, there is a lack of principled certificate of long-rangedness on arbitrary graphs. This state reflects the absence of a precise characterization of long-ranged benchmarks. We address this fundamental gap by introducing four verifiable axioms: Predictability, Tightness, Strictly k -Range, and Topology-Invariance, that any task claiming to test k -hop interactions must satisfy. We formally prove that violating any one of them admits failure modes that undermine conclusions drawn from the task. Based on these axioms, we introduce TRIP (Truly Ranged Interactions Problem) and its generalisation GRIP (Generally Ranged Interactions Problem), constructive procedures that turn any graph into a provably long-ranged task by drawing features from stable distributions. Moreover, by construction, GRIP admits a closed-form, per-range Maximum-Likelihood oracle that yields the first a priori per-range lower bound on test error available on any benchmark. Using our framework, we: (i) audit 4 common long-range benchmarks and identify their failures modes with respect to our axioms; (ii) on TRIP-instantiated topologies, we find a popular notion of curvature is uncorrelated with GNN performance, supporting topological-vs-computational bottleneck distinction; and (iii) we show that a novel benchmark’s over-squashing measures factors beyond pure long-rangedness. Code to use the framework and reproduce experiments is released this https URL.

[LG-15] Objects Without Morphisms: What LLM s for Mathematics Do Not Represent

链接: https://arxiv.org/abs/2610.03551
作者: Yanli Wang,Suijin Wang,Xiaopeng Yuan,Haohan Wang
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Large language models (LLMs) have reached expert-level performance on competition mathematics largely through the volume of search placed around them: candidate solutions are sampled in quantity and retained only when an external criterion accepts them. Such a procedure improves the outcome that survives it while leaving untouched what the model represents. We examine that question where no external criterion exists: translating statements between the dialects of neighbouring subfields, where fidelity turns on the level of generality at which content is asserted. The source leaves that level implicit in its vocabulary, so a faithful translation must recover it from the relation between the theories. We introduce an instrument that codes truth, content and scope in separate blind queues, with a judge-free measure of whether a rewrite states the hypothesis implicit in its source, and establish its sensitivity with a planted-positive control. Across seven models from four families, translating towards the general framing widens the domain of quantification in 60.6% of rewrites and narrows it in none; translating towards the concrete framing narrows it in 28.3% and widens it in 0.3%. The hypothesis that would prevent it is stated in 21.6% of model rewrites and 4.2% of human statements. Capability does not govern the asymmetry: it appears in every model tested, and the most capable widens least. It replicates on the half of the benchmark held out by a pre-registered rule, and on statements written by mathematicians. Instructing a model to state every hypothesis it requires raises that rate but not its sensitivity to direction. We argue that these systems have acquired an object-level correspondence between subfield vocabularies without the constraint under which a translation between theories carries hypotheses to hypotheses.

[LG-16] ZeroMAG: Zero-Shot Multimodal Adapter Generation for Plug-and-Play EEG Foundation Models

链接: https://arxiv.org/abs/2610.03546
作者: Yubo Wang,Jingying Ma,Xinliang Zhou,Yangxuan Zhou,Jiquan Wang,Sha Zhao,Yiyuan Yang,Yi Ding,Ziyu Jia,Chenyu Liu,Cuntai Guan
类目: Machine Learning (cs.LG)
*备注: 41 pages

点击查看摘要

Abstract:EEG foundation models (EFMs) capture reusable knowledge from large-scale EEG data, while many EEG recordings also include companion physiological signals that provide complementary information beyond the EEG-only interface. The challenge is to preserve this pretrained knowledge while extending the EFM to heterogeneous multimodal recordings through an adaptation inferred from unlabeled target data. We introduce ZeroMAG, a zero-shot multimodal adapter generation framework that extends a frozen EEG encoder and prediction head using unlabeled target recordings, without target labels or target-side optimization. The target datasets are held out from all model training and selection in the ZeroMAG pipeline. ZeroMAG organizes companion modalities around a configuration-invariant adapter, constructs a modality-subject-task condition from unlabeled recordings and task context, and generates adapter weights in a function-constrained latent space learned from source adapters. Across six held-out target datasets and three EFM backbones, ZeroMAG improves balanced accuracy by 7.22 percentage points over EEG-only inference and 4.89 points over direct weight regression, while coming within 0.50 points of supervised multimodal adaptation on average. Ablations further show that removing functional supervision from either representation learning or conditional generation degrades generated-adapter performance, confirming the contribution of both components.

[LG-17] Autonomous Robotic Navigation for Endovascular Brain-Computer Interface Access

链接: https://arxiv.org/abs/2610.03537
作者: Harry Robertshaw,Weijie Qi,Nikola Fischer,Alejandro Granados,Thomas C. Booth,Sam E. John
类目: Robotics (cs.RO); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Endovascular brain-computer interfaces (BCIs) avoid craniotomy but require precise device delivery through anatomically variable cerebral veins. This work presents the first demonstration of in vitro autonomous robotic navigation for endovascular BCI access in the cerebral venous system. Soft Actor-Critic controllers were trained in silico for two sequential tasks spanning the right internal jugular vein to the superior sagittal sinus, using geometric augmentation of one training anatomy. Navigation was evaluated in a training anatomy and an anatomically unseen hold-out model over 250 in silico episodes and five fluoroscopy-guided in vitro robotic runs per task-anatomy condition, comprising 1,000 simulated episodes and 20 physical runs overall. Task recurrent predictors were also evaluated for online identification of impending navigation failure. In silico success rates for Tasks A and B were 85.6% and 98.4% in the training anatomy and 42.0% and 91.6% in the hold-out anatomy, respectively. Fourteen of 20 physical runs were successful (70% overall), including 80% success for Task B in the hold-out phantom. In silico the predictors detected 99.3-100.0% of failures with false-alarm rates of 0.8-6.7%. During in vitro evaluation, predicted risk increased before failed episodes, but elevated probabilities during some successful runs showed reduced calibration after transfer. These results demonstrate the feasibility of autonomous cerebral venous access and show how online failure prediction could support human oversight, while also identifying anatomical generalization and sim-to-real calibration as priorities before preclinical translation.

[LG-18] An Automated and Reproducible Workflow for Crack Identification and Damage Assessment of Fusion Materials

链接: https://arxiv.org/abs/2610.03505
作者: Rinkle Juneja,Viktor Reshniak,Richard K. Archibald,John W. Duggan,Gregory R. Watson,Cory D. Hauck,Gary M. Staebler
类目: Machine Learning (cs.LG); Materials Science (cond-mat.mtrl-sci)
*备注:

点击查看摘要

Abstract:Post-exposure microscopy is central to qualification of fusion materials. However, manual analysis does not scale to the volume, heterogeneity, and multiresolution character of modern fusion-materials campaigns. To address this challenge, we present a reproducible workflow, implemented in the Galaxy scientific workflow environment, for automated crack identification and quantitative damage assessment from scanning electron microscopy images. The workflow processes SEM images and experimental metadata to identify cracks, quantify damage, and retain the intermediate products and processing history needed for reproducibility. Outputs include crack masks, skeletonized crack networks, quality-control visualizations, and scalar damage descriptors. The method is designed to operate without image-specific parameter tuning across tungsten grades, microstructures, magnifications, and damage states. We demonstrate the workflow on a sparse electron-beam thermal-shock dataset containing 418 images from 114 experiments spanning five tungsten grades and three microstructural states. We define a crack-density descriptor, which provides standardized inputs for downstream machine-learning prediction and physics-based crack simulation. These predictive components are exposed in the same Galaxy environment and are intentionally treated here as extensible workflow modules. The principal contribution is therefore an end-to-end, shareable, and computationally portable workflow that links experimental characterization, automated image analysis, preliminary damage prediction, and simulation-guided data acquisition for fusion-materials research.

[LG-19] Getting Your Guidance Weights Right in diffusion and flow-matching posterior sampling

链接: https://arxiv.org/abs/2610.03503
作者: Liam Moroy,Jean-François Giovannelli,Yoann Altmann,Steve McLaughlin,Frédéric Champagnat,Guillaume Bourmaud
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Training-free posterior sampling methods, also known as Plug-and-Play methods, leverage pretrained unconditional diffusion or flow-matching models to solve inverse problems. Most existing approaches rely on guidance weights to balance, at each time step, prior information from the unconditional score or velocity network with measurement consistency, yet the tuning of these weights is often not discussed and is largely left to heuristics. We introduce a simple and principled offline strategy for automatically tuning these guidance weights. Our key observation is that, at each time step, the conditional denoising score-matching objective for diffusion models, or the conditional flow-matching objective for flow-matching models, is a least-squares objective. Therefore, when the conditional prediction is expressed as a weighted sum of the unconditional network output and a measurement-guidance term, optimizing over these weights reduces to a two-dimensional linear least-squares problem. The resulting time-dependent guidance weights can be optimized offline for a given measurement operator, noise level and sampler at the cost of a single minibatch of sampling trajectories, without retraining or fine-tuning the pretrained generative model. Instantiated with the standard Tweedie-based measurement-consistency term, our approach improves posterior sampling and achieves state-of-the-art reconstruction performance across diffusion- and flow-matching-based methods. Moreover, the optimized guidance weights enable diffusion samplers to reduce the number of sampling steps from 1000 to 50 with no significant degradation in reconstruction quality. Code will be made available.

[LG-20] Below what training size do deep tabular generators stop beating trivial baselines? A preregistered benchmark on a size ladder of clinical and standard datasets

链接: https://arxiv.org/abs/2610.03500
作者: Shivam Shrivastava
类目: Machine Learning (cs.LG); Machine Learning (stat.ML)
*备注: 31 pages, 5 figures. Code, all 2,220 result files and the frozen preregistration: this https URL ; archived at this https URL

点击查看摘要

Abstract:Deep tabular generative models are benchmarked on datasets with tens of thousands of rows; clinical datasets have hundreds. We preregistered and ran a size-ladder benchmark to find where the two regimes diverge: 8 public datasets subsampled from 200 to 20,000 training rows, seven generators (independent marginals, Gaussian copula, SMOTE, unconditional SMOTE, CTGAN, TVAE, TabDDPM) with a fixed 20-trial tuning budget and 5 evaluation seeds, plus 4 natively small clinical datasets at true size, for 2,220 committed runs in total. The primary metric is the AUROC of fixed classifiers trained on synthetic and tested on real data. In 23 of 24 (dataset, deep model) pairs no deep model ever beats the best trivial baseline by more than seed noise, at any training size we measured. The best baseline wins 40 of 49 (dataset, size) cells. Our preregistered prediction that the deep models’ ranking would be unstable at small sizes is falsified: mean Kendall tau between adjacent rungs below 5,000 rows is 0.806, above our 0.8 threshold, and stability is highest at the smallest sizes rather than lowest. One caveat bounds all of this: in 81% of cells the gap between the top two methods is smaller than the variation between seeds. Finally, method rankings on natively small clinical datasets agree only moderately with rankings on subsampled large ones (mean tau 0.57 to 0.64), which questions whether a subsampled large dataset can stand in for a small one. All 2,220 result files, the preregistration and its hash, and the code that regenerates every figure and number from those files are public.

[LG-21] Most-Recent Anchoring with Recurrent Ordering for Time Series Forecasting

链接: https://arxiv.org/abs/2610.03494
作者: Jung Min Choi,Ngoc Son Le,Ibram Abdelmalak,Vijaya Krishna Yalavarthi,Lars Schmidt-Thieme
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Long-term forecasting models commonly process all patches in a look-back window using the same fixed stack. Older contextual patches and recent evidence therefore receive the same computational depth. Yet the information closest to the forecast and the more distant context do not contribute equally. Uniform processing leaves this distinction unexpressed in the architecture. We propose MARO, a Most-Recent Anchoring with Recurrent Ordering model that processes the look-back window from the most recent patch to the oldest. The most recent patch serves as the anchor. It initializes the latent state and conditions each subsequent step, so older patches are folded into a representation that remains centered on recent evidence. A single shared module is reused at every step, so extending the scan further into the past introduces no additional parameters. Intermediate states retained during the scan allow the forecast head to weigh short and long portions of the history separately. This expresses recency through the order of recurrent refinement. Extensive experiments across multiple real-world time series datasets show that MARO achieves state-of-the-art performance on both long-term and short-term forecasting this http URL studies examine the contribution of the main architectural components.

[LG-22] Dual-Context Analog Retrieval for Time Series Forecasting

链接: https://arxiv.org/abs/2610.03491
作者: Jung Min Choi,Ngoc Son Le,Ibram Abdelmalak,Vijaya Krishna Yalavarthi,Lars Schmidt-Thieme
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Most long-term time-series forecasting models map the look-back window directly to the full horizon in a single pass. While efficient, this design does not explicitly identify which historical states are most relevant to different future segments or exploit what followed those states. Analog forecasting addresses this by retrieving past states similar to the present and using their observed continuations, but single nearest matches can be unreliable and overlapping patches may produce redundant candidates. We propose DuoTS, a Dual-Context Time Series forecasting model that uses retrieved evidence without relying on it exclusively. DuoTS first produces a base forecast with a parallel patch encoder and linear prediction head, then progressively refines it one future patch at a time. Each refinement combines two views: a current context that attends to recent tokens and captures the latest dynamics, and a detail context that provides distinct retrieved analogs together with their subsequent trajectories. Patch-wise refinement allows the model to balance these views across the forecast horizon and associate each future segment with evidence appropriate to its temporal distance from the present. Experiments on multiple real-world datasets show that DuoTS achieves state-of-the-art performance, while ablations confirm the contribution of each context. The refinement mechanism is also model-agnostic, requiring only an encoded look-back window and the future-patch position, and can therefore be integrated into existing forecasting models.

[LG-23] Metropolis-Hastings Dominates Importance Resampling for Policy Composition

链接: https://arxiv.org/abs/2610.03480
作者: Alexey Kurennoy,Ramil Yarullin,Fergal Reid
类目: Machine Learning (cs.LG)
*备注: 56 pages, 6 figures

点击查看摘要

Abstract:Post-training a large language model (LLM) often requires exploring trade-offs between multiple rewards, but retraining for each trade-off is expensive. Decoding-time policy composition allows these trade-offs to be adjusted by combining reward-specific policies at inference time. This composition targets a weighted product of the policies’ probabilities over complete responses, but standard implementations combine their next-token probabilities, generally introducing sampling bias. We analyze a known iterative correction based on independence Metropolis-Hastings (MH). Our main result shows that, for every rollout budget, MH produces an output distribution at least as close to the target as sampling-importance-resampling (SIR) with the same budget, as measured by every convex f-divergence. We also derive a lower bound on MH’s improvement over the uncorrected decoder in a consensus objective measuring agreement with the supplied policies. We further characterize the correction’s sampling error in two asymptotic regimes: when the reward-specific policies approach agreement, and when the log ratio between target and uncorrected-decoder probabilities fluctuates increasingly widely, as can happen for long responses. We complement our analysis with experiments in enumerable and LLM-scale settings.

[LG-24] Beyond Random Splits: Evaluating Drug-Target Affinity Models Under Chemically and Biologically Motivated Distribution Shifts Copy NEURIPS2026

链接: https://arxiv.org/abs/2610.03456
作者: Minjae Chung,Clara Li,Malar Paavai Muthukumaran,Shaunna Wang,Aniket Ramkrishnan Iyer,Shaun Qien Yeau Tan,Harinishree Sathu,Micky C. Nnamdi,J. Ben Tamo,Benoit Louis Marteau,May Dongmei Wang
类目: Machine Learning (cs.LG)
*备注: Accepted to the NeurIPS 2026 Workshop on AI for Drug Discovery (AI4DD)

点击查看摘要

Abstract:Drug-target affinity (DTA) prediction is widely used to prioritize candidate compounds before costly experimental screening. DTA models are often compared under a single data split, even though deployment may require extrapolation to new chemical series, new protein targets, or both. We ask whether the distribution shift used for evaluation changes which architecture appears best. We curate 718,800 unique drug-protein pairs from the ChEMBL and BindingDB datasets. We compare a Morgan-fingerprint + protein-CNN baseline with 12 controlled architectures that combine four drug representations with three ESM-2 interaction modes. Mean validation RMSE increases from 0.950 and 0.945 under scaffold and fingerprint-cluster OOD to 1.299 and 1.321 under protein-cluster and dual OOD. Model rankings are similar across the two chemical shifts (tau = 0.79), but agreement with scaffold OOD falls under protein OOD (tau = 0.39) and reverses under dual OOD (tau = -0.55). Held-out evaluation, repeated seeds, group-aware bootstrap analysis, and a size-matched control support the same conclusion: architecture selection depends on the form of extrapolation, not only on average error or training-set size. DTA benchmarks should therefore match the chemical and target shifts expected at deployment.

[LG-25] Causal Representation Learning with Instantaneous and Lagged Relations via Nonstationarity

链接: https://arxiv.org/abs/2610.03452
作者: Tatsuya Yamada,Hiroshi Morioka,Yoshinobu Kawahara
类目: Machine Learning (cs.LG)
*备注: 46 pages, 6 figures, 16 tables

点击查看摘要

Abstract:Causal representation learning for time-series data aims to identify latent states and their causal relations from observations. In this setting, an important challenge is to model both lagged causal relations across observation intervals and faster causal effects that appear as instantaneous relations within an interval, while accounting for nonstationarity in time-series data. However, methods that jointly handle these causal relations and nonstationarity remain limited. To address this gap, we establish sufficient conditions for identifying latent states up to component permutation and component-wise invertible transformations, and their instantaneous and lagged causal structures up to the same permutation, using an observed auxiliary variable, such as time or a condition label, associated with changes in transition-noise distributions. Based on these results, we propose iCReN, a framework that uses contrastive learning with discrete or continuous auxiliary variables to learn latent representations and estimate their instantaneous and lagged causal structures. Experiments demonstrate accurate recovery of latent states and both instantaneous and lagged causal structures on synthetic data and the utility of the learned representations for downstream forecasting on real-world data.

[LG-26] Deep Bayesian REFoCUS

链接: https://arxiv.org/abs/2610.03419
作者: Simon Penninga,Ruud van Sloun
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:In this work we formulate ultrasound multistatic recovery from arbitrary transmit sequences as a Bayesian inference problem. To that end, we train a deep generative prior on multistatic data sets to tackle the rank-deficient regime in which classical linear REFoCUS decoders fail. This appproach, which we term Deep Bayesian REFoCUS, outperforms the linear baselines for all regimes of rank-deficiency and noise levels, and regresses to linear decoding when inversion is exact. The model also expresses uncertainty in the null space of the acquisitions, whereas the linear REFoCUS decoders only provide point estimates. Finally, we analyze the impact of distribution shift between simulation and in-vivo acquisitions, showing remarkable generalization ability without any fine-tuning or adaptation.

[LG-27] AIBL: Augmented Instance-Based Learning with Structured Memory and Neural Embeddings

链接: https://arxiv.org/abs/2610.03413
作者: Radha Poovendran,Andrea Stocco,Linda Bushnell
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Sequential learning systems often make decisions from accumulated experience while receiving high-dimensional inputs whose distribution may change over time. Instance-Based Learning Theory (IBLT) provides a principled case-based framework for such settings through stored situation-decision-utility instances, partial matching, activation, and blending. IBLT relies on symbolic knowledge representation in dictionary-like formats, but text, images, transaction vectors, and user-item histories often require learned similarity rather than hand-specified matching rules. In this paper, we introduce AIBL (Augmented Instance-Based Learning), an instance-learning model formulated in a learned vector space for high- dimensional sequential data. AIBL generalizes symbolic situation matching to neural embedding similarity while retaining instance storage, activation- weighted retrieval, and utility blending. The AIBL model organizes memory into active, forgotten, and surprise stores. Surprise memory separates weakly matched, possible out-of-distribution, or corner-case observations from active memory, reducing forced fitting to the nearest available cases. An observation-driven graduation algorithm promotes recurring surprise instances to active memory, allowing the memory to incorporate repeated novel patterns that may arise under concept drift. We evaluate the same implementation on five machine learning tasks and three controlled simulation tasks, comparing AIBL with classical IBLT variants and task-specific baselines where appropriate. AIBL improves accuracy by 6 to 17 percentage points. The results show where vector-space retrieval improves over symbolic matching and how the added memory mechanisms govern novelty detection, cold-start handling, drift adaptation, and reward learning under the tested protocols.

[LG-28] 16-bit Precision of Convolutional Neural Networks on Microcontroller Units for 8-bit Costs

链接: https://arxiv.org/abs/2610.03402
作者: Rui Liu,Benjamin Paaßen
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:To deploy deep neural networks on edge hardware, highly efficient inference schemes are necessary that retain high accuracy. This work presents W16A16, a high precision (16-bit), fast speed, low energy quantization method. On a widely applied microcontroller architecture Armv7E-M, our proposed approach achieves faster speed and lower energy consumption on layer- and model-level compared to alternative quantization schemes. We analyze the architecture of Armv7E-M, explain the underlying principles behind the performance advantages of 16-bit approaches, and evaluate the empiric quantization errors for regression and classification tasks, as well as empiric time- and energy consumption in MCU deployment. We observe ca.\ 10 times lower quantization errors compared to 8-bit quantization schemes while achieving similar or better inference times and energy consumption.

[LG-29] A Unified Framework for Bayesian Data Assimilation with Generative Models and Observation Interpolants

链接: https://arxiv.org/abs/2610.03396
作者: Nikolaj T. Mücke,Benjamin Sanderse
类目: Machine Learning (cs.LG); Computational Engineering, Finance, and Science (cs.CE)
*备注:

点击查看摘要

Abstract:Bayesian data assimilation combines model forecasts with noisy observations, but sampling high-dimensional, non-Gaussian posteriors remains challenging. We introduce an observation-interpolant framework that turns pretrained stochastic interpolant, flow matching, and diffusion models into posterior samplers without retraining. Conditioning the interpolant path on observations yields a shared likelihood-score correction to the drift or velocity, unifying stochastic and deterministic posterior sampling. The resulting SDEs and ODEs sample the exact posterior when the intermediate likelihood score is known. For practical computation, we approximate this score using a closed-form Gaussian surrogate with a bias-corrected mean and covariance inflated by the model’s source covariance. Jacobian-free and ensemble-shared approximations make the method tractable in high dimensions. We evaluate the framework on linear-Gaussian dynamics, stochastic two-dimensional Navier-Stokes, and urban airflow with up to O(10^4) degrees of freedom.

[LG-30] Bidirectional Voronoi-biased Exploration Curriculum for Reinforcement Learning

链接: https://arxiv.org/abs/2610.03395
作者: Juri Pfammatter,Kaixian Qu,Clemens Schwarke,Victor Klemm,Marco Hutter
类目: Machine Learning (cs.LG); Robotics (cs.RO)
*备注:

点击查看摘要

Abstract:Long-horizon tasks with sparse rewards pose an exploration bottleneck for goal-conditioned reinforcement learning: a policy started from the initial state rarely reaches the goal and receives no learning signal. Reference motions, hand-designed curricula, and shaped rewards supply this signal but require demonstrations or task-specific engineering; automatic start-state and goal curricula avoid this but typically expand from one side only, so the full distance to the target must be covered from that side. We propose the Bidirectional Voronoi-biased Exploration curriculum for Reinforcement learning (BVER), which expands from both ends at once. Inspired by bidirectional RRT planning, BVER grows start states outward from the goal and goals outward from the initial state distribution, biases both toward unexplored task space, and steers them toward each other, training one goal-conditioned policy on both. On point-mass mazes, quadrupedal box climbing, and robot-arm ring-on-peg transfer, BVER learns faster than all compared reference-free curricula. On box climbing, it reaches 95% success on a 0.4 m box in roughly 65% fewer iterations than the best of them, is the only one of them to learn to climb a 0.7 m box, and yields a policy robust to start, goal, and yaw variation. Without a demonstration, it approaches the sample efficiency of reference-based curricula on the 0.4 m box and on ring-on-peg transfer. Ablations show that expanding from both ends outperforms either direction alone.

[LG-31] Operator-informed initialization for Fourier features physics-informed neural networks

链接: https://arxiv.org/abs/2610.03378
作者: Juan Molina,Paris Perdikaris,Mircea Petrache,Matías Courdurier,Francisco Sahli Costabal
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Physics-Informed Neural Networks (PINNs) typically exhibit spectral bias, where some frequencies of the target function converge more slowly than others. In this work, we analyze the training dynamics of Fourier Feature PINNs in the Neural Tangent Kernel regime to address this limitation. We derive an explicit evolution equation to estimate the residual error in the frequency domain, demonstrating that the convergence rate of specific frequencies is primarily governed by the product of the differential operator’s symbol and the spectral density of the initialization weights. Leveraging this theoretical insight, we propose an informative initialization strategy that tailors the initial weight distribution to the specific PDE being solved. With this method, we can diminish the operator-induced spectral bias, balancing the convergence rates across the frequency spectrum and achieving better prediction accuracy. Numerical experiments on linear and nonlinear partial differential equations confirm that this initialization strategy improves learning dynamics and approximation accuracy across frequencies compared to standard initialization methods, with no additional training cost.

[LG-32] SCAD: Structured Credit Assignment and Distillation for Long-Horizon Agents

链接: https://arxiv.org/abs/2610.03372
作者: Shangyang Wu,Shuai Zhao,Ziyue Zhu,Jinyang Wu,Anh Tuan Luu,Haoran Luo
类目: Machine Learning (cs.LG)
*备注: 32 pages

点击查看摘要

Abstract:Training long-horizon agents to solve complex tasks requires effective supervision over extended interaction sequences. However, sparse terminal rewards obscure intermediate contributions, while on-policy distillation can lose informative teacher guidance as student-generated histories grow. To address this problem, we introduce SCAD, which organizes interactions into planning and bounded subtask execution, distills execution in local contexts, and refines planning credit through cross-rollout subtask prefix trees, with planning receiving full terminal credit and execution receiving positive terminal credit and teacher guidance. Across all evaluated benchmarks, SCAD improves macro-average accuracy over the strongest training baseline by 4.48 percentage points for text tasks and 4.19 points for multimodal tasks. SCAD effectively combines outcome-based credit assignment with teacher-guided distillation to improve planning and execution in long-horizon agents.

[LG-33] S2-PINN: Stochastic Separable Physics-Informed Neural Networks

链接: https://arxiv.org/abs/2610.03303
作者: Zhendong Li,Akwum Onwunta
类目: Machine Learning (cs.LG); Mathematical Physics (math-ph)
*备注:

点击查看摘要

Abstract:Uncertainty quantification (UQ) for random partial differential equations (PDEs) is ubiquitous in computational science and engineering. However, classical spectral solvers for this class of problems face the curse of dimensionality, and existing neural solvers often ignore the stochastic structure that makes moments and calibration tractable. We introduce a stochastic separable physics-informed neural network, dubbed S ^2 -PINN, that represents the solution u(t,\mathbfx,\mathbfZ) of a random PDE with a learnable Gaussian spatial dictionary, Fourier temporal features, and a generalized polynomial chaos (gPC) stochastic basis, coupled by a low-rank Canonical Polyadic (CP) tensor decomposition core. The method is trained with a hybrid strong-form and gPC-projected residual loss. Our theoretical analysis establishes that the separable class is dense in L^2 under mild conditions, and the projected residual corresponds exactly to a stochastic Galerkin constraint. Furthermore, we show that mini-batch projection coefficients are logarithmically dependent on the number of gPC modes, and that the orthogonality penalty controls the conditioning of the learned spatial dictionary. Using four manufactured random PDE benchmarks, we show that S ^2 -PINN outperforms nine baselines in terms of mean and variance accuracy, as well as calibration, while using significantly fewer parameters. Further evaluations on non-manufactured Poisson and Darcy problems, a stochastic Navier–Stokes problem, a diffusion scaling study of higher random dimensions, and two stochastic inverse problems reveal the generalization capabilities of the proposed structure. Together, these results support stochastic separability as an effective design principle for physics-informed neural UQ. The code for the experiments can be found in this https URL

[LG-34] Architecture-Dependent Fusion Pathways in MLLM s

链接: https://arxiv.org/abs/2610.03289
作者: Hebao Zhu,Dongxia Wu
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Multimodal Large Language Models (MLLMs) achieve strong performance across vision-language tasks, yet the internal mechanisms by which visual and textual information are fused across layers remain insufficiently understood. We investigate representative MLLMs from two architectural paradigms: concatenation architectures and native multimodal architectures. We conduct three progressively connected analyses: alignment decoupling identifies which modality changes, attention routing and entropy characterize how cross-modal information is distributed, and intrinsic dimensionality examines how fusion reshapes feature spaces. Separately, we perform causal intervention experiments as a validation of the resulting interpretation. As a supplementary analysis, we use visual CKA to examine the Platonic Representation Hypothesis. Together, these analyses reveal two distinct fusion pathways: concatenation models follow a text-first, vision-later pathway, whereas native models exhibit earlier visual-textual co-adaptation and feature-space reorganization. This work provides a mechanistic perspective for understanding multimodal fusion and supports architecture-aware diagnostics of multimodal representations.

[LG-35] PaMIR: Open Benchmark of Public Credit-Default Datasets

链接: https://arxiv.org/abs/2610.03259
作者: Mikhail Liashkov,Ilyas Varshavskiy,Shuhratjon Khalilbekov,Azizjon Azimi,Bonu Boboeva
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:We release PaMIR (Public Arrival-ordered Measurement for Inference in Risk), an open benchmark for credit-default prediction when labels are scarce and arrive late. The field’s reference benchmark studies use eight datasets each, only two or four of them public. PaMIR brings together 19 public datasets with binary default labels – 1.24M loans, firms and card accounts from nine countries – rebuilt from pinned source snapshots by one leakage-audited recipe and never redistributed; to our knowledge it is the one of its kind as of today. Every model is a single function, scored under a repeated i.i.d. split and a label-delayed stream in which each application is scored on arrival, with AUC reported by label budget; fleet means are withheld unless every dataset is scored. A synthetic-data harness tests generated training rows without letting a generator see held-out rows. This report describes release 0.4.0 of this living benchmark.

[LG-36] Cross-cohort TB classification using clinical data gathered in Uganda and South Africa

链接: https://arxiv.org/abs/2610.03256
作者: Joshua M. Jansen van Vüren,Devendra S. Parihar,Daphne Naidoo,Marisa Klopper,Frank Cobelens,Lutz Kolbe,Kimsey Zajac,Willy Ssengooba,Moses Joloba,Grant Theron,Thomas R. Niesler
类目: Machine Learning (cs.LG)
*备注: Accepted: SATNAC, Drakensberg, South Africa, 2026

点击查看摘要

Abstract:We present a first evaluation of machine learning applied to patient clinical and demographic data gathered in two different countries for the purpose of tuberculosis (TB) screening to identify people who would benefit from expensive molecular testing. Experiments are based on the recently-compiled CAGE-TB dataset, which includes sub-cohorts of people with presumptive TB presenting at community health care centres in South Africa and Uganda. Three neural network architectures (logistic regression (LR), multilayer perceptrons (MLP) and convolutional neural networks (CNN)) are considered in conjunction with greedy feature selection. For the convolutional neural network, a strategy that jointly optimises feature selection and feature ordering is proposed and shown to lead to consistent development and test set improvements. For all three models, development set area under the receiver operating characteristic (AUROC) curve is improved by 2-7% using feature selection. LR after feature selection achieves an AUROC of 0.8 [0.75,0.86] (95% CI) and 0.84 [0.78,0.9] when testing on the held-out Ugandan and South African data respectively. Although outperforming LR on the development cohort, the deeper networks (MLP, CNN) show inconsistent trends on the held-out cohorts, while LR achieves performance within 1-2% of the best achieved in terms of AUROC. LR narrowly misses the WHO minimum requirements by 4-9% in sensitivity even though the network is being evaluated on a completely held-out cohort. The development of neural-network based classifiers for TB screening therefore appears viable.

[LG-37] he Neuro-Physical Inverter: A Modular Framework for Magnetotelluric Inversion Coupling Ensemble Conditioning with Residual Learning

链接: https://arxiv.org/abs/2610.03225
作者: Jae Deok Kim,Sai Ravela,Rob. L. Evans
类目: Machine Learning (cs.LG)
*备注: Accepted by IEEE TGRS

点击查看摘要

Abstract:We present the Neuro-Physical Inverter (NPI), a modular, uncertainty-aware framework for geophysical inversion that couples ensemble-based conditioning with constrained residual learning, demonstrated in the 1D magnetotelluric (MT) setting as a controlled testbed. The framework operates in two stages. An Ensemble-Conditional Gaussian Process (EnsCGP) conditions a prior ensemble of resistivity models on the observed response, producing a physically admissible reference ensemble. A residual-learning neural network then predicts targeted corrections to this reference, trained on synthetic data and fine-tuned per station for field application through a physics-coupled objective. Because an ensemble is conditioned, refined, and propagated through both stages, every estimate carries an associated ensemble spread. Synthetic experiments show that NPI systematically reduces ensemble-mean error without destabilizing the ensemble. Applied to broadband MT data from the Gabbs Valley geothermal region (Nevada, USA), NPI reduces the across-station mean misfit over the mid-period band while retaining comparable ensemble spread. The propagated ensemble yields a factor of uncertainty that serves as an operational measure of constraint within the assumed model class. Both stages are dimension-agnostic in formulation, and the design principles established here are intended to scale to higher-dimensional parameterizations.

[LG-38] Kernel Singular Value Decomposition with Extension to Multiple Data Sources

链接: https://arxiv.org/abs/2610.03216
作者: Xinjie Zeng,Qinghua Tao,Johan Suykens
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Kernel Singular Value Decomposition (KSVD) learns a pair of singular vectors w.r.t. an asymmetric kernel matrix, which can be induced by two data sources, e.g., the queries and keys in self-attention or the rows and columns of a given matrix. In this work, we extend KSVD to multiple data sources, namely eKSVD, which conducts joint nonlinear feature learning upon asymmetric kernels. In the primal formulation, the projections associated with each data source are jointly learned to capture maximal information, while incorporating pair-wise couplings. With the Lagrangian and its Karush-Kuhn-Tucker (KKT) conditions, the optimization in the dual leads to a generalization of the shifted eigenvalue problem in Lanczos decomposition theorem of KSVD. Further, a covariance-based framework is derived together with using neural networks (NNs) for explicit feature mappings, complementary to the kernel-based interpretation and optimization. Numerical experiments verify the effectiveness of our eKSVD compared to methods based on Mercer kernels for tackling multiple data sources, and our innovation of deploying NNs demonstrates great flexibility for kernel methods.

[LG-39] HyperFuse: Fast Self-Supervised Node Embeddings for Attributed Hypergraphs

链接: https://arxiv.org/abs/2610.03211
作者: Megha P,Harshit Kumar,Srajan Agarwal,Anirban Banerjee,Olaf Wolkenhauer,Saptarshi Bej
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Self-supervised hypergraph representation learning can produce informative node embeddings, but existing methods often require deep encoders trained for hundreds of epochs, making embedding generation costly even for hypergraphs with a few thousand nodes. This limits applications requiring embeddings for many or evolving hypergraphs. We present HyperFuse, a label-free pipeline for fast hypergraph representation learning. HyperFuse (i) computes structural node coordinates by maximizing a spectral relaxation of hypergraph modularity using Banerjee’s hypergraph adjacency and a matrix-free operator with cost linear in node-hyperedge incidences; (ii) constructs multi-scale feature summaries and assigns bounded utility weights to hyperedges based on member stability under feature and membership masking; and (iii) trains a lightweight utility-weighted hypergraph encoder for 100 epochs using an invariance-decorrelation objective. We compare HyperFuse with TriCL, SE-HSSL, VilLain, and HypeBoy on nine public hypergraphs using six downstream classifiers and k-means clustering. On the eight datasets where all methods completed, HyperFuse required 8.7 s per dataset on average, achieving 13-179x geometric-mean speed-ups over the baselines. It achieved the highest average accuracy with five of six classifiers, while classification and clustering performance was not significantly different from TriCL and SE-HSSL. Compared with HypeBoy, HyperFuse was 13x faster and 2.1-4.1 percentage points more accurate across all classifiers. HyperFuse provides a practical approach for fast, repeated hypergraph embedding generation.

[LG-40] Sample complexity of variance-reduced policy gradient: weaker assumptions and lower bounds

链接: https://arxiv.org/abs/2610.03165
作者: Gabor Paczolay,Matteo Papini,Alberto Maria Metelli,Istvan Harmati,Marcello Restelli
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Several variance-reduced versions of REINFORCE based on importance sampling achieve an improved O(\epsilon^-3) sample complexity to find an \epsilon -stationary point, under an unrealistic assumption on the variance of the importance weights. In this paper, we propose the \algo (Defensive Policy Gradient) algorithm, based on defensive importance sampling, which achieves the same rate without any assumption on the variance of ordinary importance weights. We also establish lower bounds in a generalized black-box policy-optimization model that hides states and actions and permits parameter-dependent rewards. In this model, the optimal rates are \Theta(\epsilon^-4) with bounded-variance one-policy feedback and \Theta(\epsilon^-3) with mean-square-smooth coupled two-policy feedback. Under standard policy-regularity conditions, REINFORCE and \algo realize the corresponding oracle conditions and attain the O(\epsilon^-4) and O(\epsilon^-3) upper bounds, respectively. Although the lower bounds do not apply directly to the classical MDP interaction model in which these algorithms operate, this correspondence provides oracle-level evidence that the faster rate of \algo is optimal and genuinely separated from that of vanilla policy gradient.

[LG-41] Page-EntroKV: Hardware-Aligned Entropy-Weighted KV-Cache Eviction under Grouped-Query Attention

链接: https://arxiv.org/abs/2610.03135
作者: Inbasekaran S
类目: Machine Learning (cs.LG)
*备注: 24 pages, 8 figures, 10 tables. Formal framework with pilot-scale empirical validation on Qwen2.5-1.5B-Instruct. Includes step-by-step derivations (Appendix C) and PyTorch reference implementation (Appendix D). Code and data available at: this https URL

点击查看摘要

Abstract:Serving long-context autoregressive language models is constrained by the key-value (KV) cache. Most dynamic eviction methods score token importance per query head and choose tokens independently. This fits poorly with grouped-query attention (GQA), where several query heads share one physical KV buffer: divergent per-head selections force the serving engine to retain the union of their choices - inflating the cache by up to the group ratio r - while arithmetic mean pooling dilutes the specialized retrieval heads that carry factual recall. We introduce Page-EntroKV, a formal framework for KV-cache eviction operating at the granularity GQA serving actually allocates. Heads within each physical group are pooled by weights derived from sink-isolated collision (Renyi-2) entropy - one inner product per head, computed once at prefill with no calibration - so sink heads cannot masquerade as retrieval heads. Pooled scores are projected onto PagedAttention page frames, and eviction executes at the hardware tuple (layer, group, page). We formalize the union overhead ratio (UOR) and intra-group disagreement, prove an exact identity linking them for two-head groups alongside two-sided bounds at every group ratio, prove strict budget preservation and a finite-context needle-retention bound that arithmetic mean pooling provably violates, and give exact per-layer page accounting. On a pilot architecture (Qwen2.5-1.5B-Instruct, r=6), head-independent replay over 2,240 group measurements yields union overhead up to 4.75x at a 2% budget, while Page-EntroKV holds UOR exactly 1.000; sink isolation removes a 13x sink masquerade; needle recall is 100% versus 0% for mean pooling at a 20% budget; retained cardinality is exact for every page size; and QA and code tasks remain solvable at 20% retention.

[LG-42] Safe Streaming Flow Planning by Aligning Sampling Dynamics with Execution Dynamics

链接: https://arxiv.org/abs/2610.03132
作者: Seunghwan Jang,Jeongyong Yang,Siddharth Ancha,SooJean Han
类目: Robotics (cs.RO); Machine Learning (cs.LG)
*备注: Accepted to the 10th Conference on Robot Learning (CoRL 2026). Project page: this https URL

点击查看摘要

Abstract:Generative planners based on diffusion/flow matching can learn to synthesize long-horizon trajectories from demonstrations. However, real-world deployment requires (i) enforcing safety constraints during execution and (ii) tight online replanning at fast execution rates. Prior safe diffusion/flow planners generate the agent’s full trajectory at once, while repeatedly perturbing intermediate states to satisfy safety constraints. This approach is not only computationally intensive, but also introduces distribution shift since the learned sampling dynamics is distinct from the system’s execution dynamics. We propose SafeStreamingFlow, a goal-conditioned planner that aligns flow sampling dynamics with execution dynamics by sequentially integrating a learned state vector field with hierarchical state prediction. Importantly, we need to enforce safety constraints only for the executed step via high order control barrier functions. Across navigation, racing, and locomotion benchmarks, SafeStreamingFlow reduces planning latency and improves safety compared to existing methods, while maintaining competitive goal-reaching success.

[LG-43] Coverag e You Can Steer: Online Conformal Calibration for RL-Driven Hardware-Aware NAS

链接: https://arxiv.org/abs/2610.03127
作者: Pedro Brandimarte,Nerea Aranjuelo,Marcos Nieto,Oihana Otaegui
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Hardware-aware neural architecture search (NAS) is dominated by evaluation cost: every architecture must be trained before its reward is known. Conformal-prediction filters cut this cost by pruning candidates whose predicted-reward upper bound misses a threshold, with a distribution-free guarantee that at most a fraction \delta are wrongly discarded. That guarantee assumes exchangeability between calibration and test candidates, which the surrounding reinforcement-learning (RL) loop violates: the policy’s proposals improve as search proceeds and, in layer-by-layer construction, shift within every episode. We replace one-shot quantile estimation with online feedback control (Adaptive Conformal Inference, with tuning-free, locally-adaptive, and group-conditional variants), restoring steerable coverage: dialing the target delivers it, monotonically and reproducibly, for arbitrary sequences. Across three neural-network architecture families and both single-step and sequential search (three seeds), it tracks every requested level to within \sim10^-3 while pruning 25-50% of evaluations at no measured accuracy cost, whereas static calibration loses control of its coverage and a Gaussian-process baseline stays conservative regardless of the request. Finally, used as an acquisition function on one constrained testbed, the same optimistic bound beats random search, a gain that fixed optimism already carries and online calibration sharpens. The source code is available at this https URL.

[LG-44] ParaG eo: Decomposing Paralinguistic Variation into a Shared Latent Geometry

链接: https://arxiv.org/abs/2610.03125
作者: Yuhan Liu,Yuxuan Ou,Ruoxi Su,Mohamed Ahmed Zaki,Yunbo Long
类目: ound (cs.SD); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Speech delivery varies with both the requested paralinguistic attribute and the linguistic content. We introduce ParaGeo, a matched-content decomposition of paralinguistic variation in a frozen speech language model. Synthesized audio tokens are replayed with a fixed listening prompt; pooled key/value (K/V) representations are centered and projected into a shared low-dimensional space. Our GLM-4-Voice probe spans 80 requested controls from 12 benchmark families across eight sentences. With a globally fitted calibration basis, content-held-out centroid accuracy using this basis is 9.49% versus a 1.25% permutation baseline; same-label cross-content cosine similarity is 0.285 versus 0.017, and both conditional permutation tests yield p = 0.001. A separate ten-scenario, six-style probe reveals reproducible contrast directions across scenarios. Static, additive, and temporal interventions produce attribute-, layer-, and schedule-dependent response profiles. These results provide a shared coordinate representation for measuring paralinguistic structure and an empirical starting point for latent speech control. Code is available at this https URL.

[LG-45] Exploring the Trade-Off Between Structured Pruning and Fault Tolerance in Deep Neural Networks for Space Applications

链接: https://arxiv.org/abs/2610.03117
作者: Toon Vinck,Naïn Jonckers,Jaro De Roose,Jeffrey Prinzie,Peter Karsmakers
类目: Machine Learning (cs.LG)
*备注: 5 pages, 3 figures, SPAICE 2026 Conference

点击查看摘要

Abstract:Deep Neural Networks (DNNs) inherently exhibit a degree of robustness to bit-level faults due to their distributed representation of information. As a model increases in width, this information becomes more dispersed, theoretically reducing the impact of any single bit fault. In this paper, we empirically investigate the relationship between model width and robustness to Single Event Upsets (SEUs). We conduct a comprehensive experiment in which baseline models undergo iterative structured pruning to reduce their width while preserving task performance as much as possible. At each pruning stage, we run a targeted fault-injection campaign to evaluate the model’s performance under simulated bit-flip scenarios. Our results show that, although structured pruning increases per-inference sensitivity to faults by reducing redundancy, this effect is effectively counterbalanced by shorter execution time, which lowers the probability of encountering an SEU. These findings suggest that structured pruning can yield significant energy and latency savings without compromising overall reliability, providing useful guidance for designing robust AI systems for space applications.

[LG-46] LS-AR: Future-Predictive Latent Steering in Autoregressive LLM s NEURIPS2026

链接: https://arxiv.org/abs/2610.03093
作者: Anubha Gupta,Eduardo Pignatelli
类目: Machine Learning (cs.LG)
*备注: Accepted at the NeurIPS 2026 Workshop: Long-Context Foundation Models

点击查看摘要

Abstract:Standard autoregressive (AR) models process high-level task instructions, state history, and transient tokens within a single shared sequence of tokens. Consequently, they lack the architectural mechanisms needed to isolate macro-objectives from context noise. To overcome this single-channel limitation, we introduce Latent-Steered Autoregressive (LS-AR), a dual-channel architecture that decouples continuous goal steering from discrete token decoding via FiLM conditioning. We evaluate a Static Goal Encoder (P_0) for persistent macro-objective retention across long rollouts and a Dynamic State Tracker (P_t) for recurrent latent updates during generation. On long-horizon retrieval past context limits (H=1024, W=500), LS-AR (Static) achieves 100% target recall where parameter-matched baselines collapse (0%), while increasing throughput by ~35% and cutting peak VRAM by 52.8%. In Blocksworld planning under forced perturbations (k=1), LS-AR (Dynamic) sustains an 89.0% completion rate vs. 71.0% for the baseline, though zero-shot entity scaling (N - N+1) exposes single-vector capacity limits (0%). Finally, dual-channel authority analysis shows that text goal dropout establishes latent-dominant control, offering structural defence against text prompt injection while introducing a latent vector attack surface.

[LG-47] Light Entropic Optimal Transport on Riemannian Manifolds

链接: https://arxiv.org/abs/2610.03085
作者: Xavier Aramayo-Carrasco,Petr Mokrov,Alexander Korotin
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Entropic Optimal Transport (EOT) has become a practical framework for learning stochastic couplings between complex distributions, with applications in generative modeling and domain adaptation. However, most EOT solvers are designed for Euclidean spaces, while manifold extensions remain limited and often rely on costly iterative methods, simulated dynamics, or generic neural models that do not fully exploit the underlying geometry. We introduce ManifoldLightOT, a light approach for learning kernel-induced EOT couplings directly on common manifolds. Using the kernel form of the EOT solution, we construct geometry-specific Gibbs kernels together with compatible potential parameterizations for spheres, tori, \mathrmSO(3) , and \mathrmSE(3) . These choices yield closed-form normalization and directly sampleable conditional distributions. Our formulation naturally extends to products of manifolds, making it applicable to more complex geometries. The parameters of the potentials are optimized directly from samples using Monte Carlo estimates of the learning objective. Through synthetic and real-world experiments, we show that ManifoldLightOT often outperforms existing manifold OT methods while retaining direct sampling.

[LG-48] Smart Sensing for Safer Bridges: From Sensor Signals to AI-Driven Anomaly Detection

链接: https://arxiv.org/abs/2610.03082
作者: Rahul Jaiswal,Joakim Hellum,Halvor Heiberg
类目: Machine Learning (cs.LG)
*备注: 6 pages, 14 Figures, 4 Tables

点击查看摘要

Abstract:Bridges contribute significantly to transportation connectivity and urban development. Therefore, reliable bridge monitoring is crucial for protecting public safety and detecting anomalous behavior in bridge sensor data that may provide early indications of abnormal structural conditions. This paper investigates anomaly detection in real-world bridge sensor data using two different complementary approaches, namely signal processing and the data-driven machine learning model Isolation Forest. The real-time bridge sensor data is collected from an iBridge sensor device installed on a bridge in Norway. The methods are evaluated using anomaly counts, anomaly detection time, processing rate, anomaly rates, visualization, and temporal agreement. Moreover, a controlled anomaly-injection analysis is performed to evaluate the sensitivity of each method. Numerical results demonstrate distinct detection characteristics and computational requirements, highlighting the potential of machine learning, particularly the data-driven Isolation Forest, alongside signal processing for identifying anomalies in bridge sensor measurements.

[LG-49] Learn Feasibility Once Optimize All Objectives: Derivative-Free Diffusion Models for Chance-Constrained Programming

链接: https://arxiv.org/abs/2610.03071
作者: Ziwen Liu,Yan Liu,Congying Han,Tiande Guo,Yao Yan,Weichen Zhao
类目: Machine Learning (cs.LG); Optimization and Control (math.OC)
*备注:

点击查看摘要

Abstract:Chance-constrained programs (CCPs) optimize decisions under uncertainty by limiting the probability of constraint violation. Despite advances in traditional and learning-based approaches, optimizing non-convex or non-smooth objectives and adapting to different objectives under fixed chance constraints remain challenging. In this paper, we propose a \textbfDerivative-free \textbfDiffusion-based framework that \textbfDisentangles constraint modeling from objective optimization, termed \textbfD ^3 Opt. We learn the chance-feasible structure once, independently of any particular objective, by training a risk-conditioned diffusion model solely on constraint-filtered decisions and freezing it as a reusable prior for post-specified objectives. At inference time, we propose an annealed, particle-based Feynman–Kac correction along the frozen reverse diffusion process to optimize post-specified objectives using only function evaluations. This enables derivative-free optimization of non-convex and non-smooth objectives without objective-specific retraining. We prove that the correction preserves feasibility when this property holds for the frozen prior, and derive an optimization-error bound separating learned-prior coverage, finite-particle approximation, and finite-temperature effects. Experiments on linear Gaussian CCPs, objective-transfer tasks, and chance-constrained economic dispatch demonstrate effective optimization across smooth and non-smooth objectives, including non-convex cases, and objective generalization under fixed chance constraints without retraining.

[LG-50] Learning Transferable Policies from Action-free Time Series Through Dynamical Embeddings

链接: https://arxiv.org/abs/2610.03065
作者: Niklas Emonds,Georgia Koppe
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Learning control from action-free recordings is challenging because intervention effects are unobserved and policies may exploit errors in reconstructed dynamics. We present a hierarchical model-based reinforcement learning framework that uses shared structure across related systems to learn system-specific control policies from action-free recordings. A hierarchical dynamical system reconstruction model captures shared dynamics and individual variation through low-dimensional embeddings. These embeddings are then reused to parameterize shared policy and value networks, linking differences in reconstructed dynamics to differences in control. Policies are trained entirely via simulation under an explicit intervention model with additive latent perturbations. Piecewise-linear recurrent neural networks enable mechanistic analyses of the controlled dynamics, while decoder-based constraints make the immediate effects of interventions interpretable in observation space and permit interventions on one modality while protecting another from direct manipulation. On Lorenz-63 and double-pendulum systems, hierarchical policies improve transfer over independently trained policies. On Lorenz-63, they also achieve a higher mean reward than repeated planning with the same reconstructed models, perform comparably to methods trained with controlled interactions, and generalize to systems absent from policy training after embedding inference alone. Applications to neural-behavioral recordings demonstrate suppression of predicted movement under constrained neural perturbations. Together, these findings show how shared dynamical representations support transferable control and mechanistic hypothesis generation from action-free recordings.

[LG-51] When Does Synthetic Relational Data Teach Models to Use Relations? Tracing Predictive Structure from Pretraining Data to Model Behavior

链接: https://arxiv.org/abs/2610.03057
作者: Shivam Dubey,Mohamed Bouadi,Nassim Bouarour,Varun Kulkarni,Aditya Tanna,Vinay Kumar Sankarapu
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Relational foundation models are increasingly pretrained on synthetic databases, yet downstream benchmarks reveal little about why one synthetic corpus produces a better model than another. In particular, strong performance may arise from realistic row-level statistics without the model ever learning to use relational structure. We study this as a data-attribution problem: which property of synthetic pretraining data induces relational computation? Using four Relational Transformer checkpoints trained with the same architecture, initialization, objective, and compute budget on corpora produced by four relational data generators, we trace a measurable property of the data to learned computation and downstream behavior. We hypothesize that relational mechanisms emerge when cross-table information is predictively necessary for the masked-cell pretraining objective. RelDiff exhibits by far the largest predictive gain from foreign-key-linked parents, and its corresponding model is uniquely sensitive to foreign-key interventions on unseen databases. This dependence survives a random-initialization control, grows monotonically with the fraction of corrupted links, and localizes to a serial cross-table pathway. Finally, disrupting the same mechanism during downstream inference removes RelDiff’s advantage on relational tasks while leaving structure-insensitive models nearly unchanged. These results connect a property of synthetic training data to a learned mechanism and, through intervention, to downstream behavior.

[LG-52] RIPPLE in Still Water: Zero-Shot Clustering in Federated Learning with Wavelet Scattering Transform NEURIPS2026

链接: https://arxiv.org/abs/2610.03054
作者: Alessandro Licciardi
类目: Machine Learning (cs.LG)
*备注: Accepted at NeurIPS 2026 (main track)

点击查看摘要

Abstract:Clustered Federated Learning (FL) partitions a client population into groups of similar local distributions and trains one specialized model per cluster, mitigating client drift that degrades single-model methods under non-IID data. Prior methods discover cluster structure inside the training loop through gradient similarity, loss evaluation, or EM-style updates, thus increasing communication overhead, exposing gradients to inversion attacks, and providing no mechanism to assign clients absent from training. We propose RIPPLE, a clustered FL framework in which cluster assignment is computed entirely offline from a spectral characterization of each client’s local data: a variance-weighted principal-component prototype embedded via the Wavelet Scattering Transform and decoded by a Gaussian Mixture VAE trained server-side on synthetic client populations before federation begins. Per-round communication cost matches FedAvg exactly, and a client absent from training obtains a personalized model from a single forward pass, without gradient computation, model evaluation, or extra communication round. We prove that the gap between RIPPLE’s surrogate clustered objective and the oracle is bounded by a computable quantity decaying with client sample size and independent of federation duration; per-cluster convergence matches the minimax-optimal rate for non-convex smooth objectives. Across five benchmarks spanning controlled and realistic heterogeneity, RIPPLE consistently outperforms all baselines, with margins growing on the most realistic partitions.

[LG-53] Balancing Multimodal Learning via Functional Progress

链接: https://arxiv.org/abs/2610.03035
作者: Zhongjing Gu,Fengqiang Wan,Yiming Cui,Yufa Feng,Yang Yang
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Multimodal learning often suffers from modality imbalance, where the joint optimization process is dominated by a single modality. Existing methods typically estimate modality imbalance from score disparities derived from prediction uncertainty or optimization statistics. However, due to distinct prediction uncertainty and learning dynamics across modalities, direct comparison of such scores may misinterpret intrinsic modality differences as progress gaps, leading to biased imbalance estimation. In this paper, we propose Function-Space Guided Multimodal Optimization (FGMO), which leverages a function-space progress signal to assess modality-wise optimization progress and coordinate optimization across modalities to alleviate modality imbalance. Specifically, we introduce Functional Progress Estimation (FPE) to measure each modality’s update-induced function-space response and calibrate it against a loss-aligned unimodal reference, producing a comparable progress signal. Based on this signal, Functional Response Control (FRC) redistributes modality-level function-space budgets and realizes the target responses through tensor-wise learning-rate adjustment. Theoretical analysis establishes a one-step target-contraction property of FRC under bounded controller-state mismatch, and extensive experiments demonstrate the effectiveness of FGMO across multiple multimodal benchmarks.

[LG-54] Signal Simplification Is Not Predictive Simplification: Diagnosing Residual Neural Forecasting in Short-Horizon Volatility

链接: https://arxiv.org/abs/2610.03019
作者: Bingqi Lian,Linfeng Cheng,Mei Lu,Jerry Wu
类目: Machine Learning (cs.LG)
*备注: Accepted at the 10th Computational Methods in Systems and Software (CoMeSySo 2026). 17 pages, 4 figures

点击查看摘要

Abstract:Hybrid statistical-neural pipelines often assume that a successful statistical first stage leaves a cleaner and more learnable residual target. We examine that assumption in short-horizon volatility forecasting through a signal-forecast-system diagnostic framework. Across five liquid U.S. assets, a volatility-aligned HAR-style model outperforms AR, MA, and ARIMA. Within expanding training windows, the pre-standardization fitted residual process used to construct residual-LSTM sequences has about 82% lower variance than the corresponding target and near-zero lag-1 autocorrelation; independently, rolling pseudo-out-of-sample HAR errors show about 74% variance reduction and similarly weak lag-1 dependence. Residual-only LSTM augmentation nevertheless raises mean squared error from 0.3049 to 0.3594 on average, with deterioration on every asset. Pure LSTM records the lowest selected pseudo-out-of-sample MSE, 0.2649, while the residual hybrid requires substantially more end-to-end runtime without improving accuracy. We describe this pattern as forecaster-preconditioner asymmetry: first-stage forecasting success and statistical residual simplification need not translate into useful downstream neural preconditioning.

[LG-55] Neural Data Needs Semantic Tokenization: Behavioral Events as Boundaries of Session-Transferable Tokens

链接: https://arxiv.org/abs/2610.03001
作者: Sangyoon Bae,Jiook Cha
类目: Machine Learning (cs.LG)
*备注: 10 pages, 3 figures

点击查看摘要

Abstract:Extracellular electrophysiology records a different set of neurons in every session. Neural foundation models embed each neuron and each session into their tokens, so every new session is an input they have never seen, and they fail to generalize to it. A tokenizer for new sessions needs a unit that every session shares and that carries behavior. Population activity offers such a unit. It evolves on a low-dimensional manifold that persists across neuronal turnover and across animals once sessions are aligned. This manifold changes regime at task events such as stimulus onset and movement onset. Within each regime the population occupies a state, the part of the manifold it spans between two events, and each state carries its own behavioral meaning. We propose Tokenization with States (TWS), which segments each trial, one repetition of the task, at these events and converts every state into tokens of population geometry, with no neuron or session embedding. On held-out International Brain Laboratory (IBL) sessions, TWS decodes movement even from regime boundaries that carry no information about the target, while a per-neuron foundation model pretrained on those sessions decodes at chance. Frozen after training on mice alone, TWS transfers to macaques and Utah arrays with only a linear probe. On reach direction, a target that its boundaries do not define, it achieves a Matthews correlation of 0.23, where the event time alone achieves 0.01 . For cross-session generalization, the token matters more than the model on top.

[LG-56] Differentiable Koopman Operator for Contrastive Learning on Dynamic Graphs

链接: https://arxiv.org/abs/2610.02990
作者: Md Abrar Jahin,Taufikur Rahman Fuad,Md Rizwan Parvez
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Real-world interaction networks are inherently dynamic: edges form and dissolve as node behavior shifts over time. Most snapshot-based contrastive methods encode temporal dependencies implicitly in encoder weights, without an explicit model of how node representations evolve, making them brittle under distribution shifts. We propose KAIROS (Koopman-Aligned Invariant Representations for Open Dynamic Systems), a self-supervised framework that embeds a differentiable Koopman operator within a dynamic graph contrastive learning loop to linearize temporal evolution in the learned embedding space. A dual-view encoder pairs raw node features with a graph-diffused structural view and is optimized with multi-granularity contrastive objectives across temporal windows. For anomaly detection, KAIROS uses the Koopman prediction residual together with temporal inconsistency and local neighborhood deviation to separate irregular behavior from predictable graph evolution. Evaluated on nine dynamic graph benchmarks, KAIROS achieves state-of-the-art anomaly detection results on all nine datasets, with gains of up to 23.15 ROC-AUC points over prior work, while remaining competitive for unsupervised node classification. These results show that explicit dynamics modeling provides a scalable and effective inductive bias for temporal graph representation learning.

[LG-57] Dirac-Interconnected Neural Elements: Discovering Modularity in Physical Systems Without Reduction

链接: https://arxiv.org/abs/2610.02960
作者: Reiho Li,Razmik Arman Khosrovian,Takaharu Yaguchi,Hiroaki Yoshimura,Takashi Matsubara
类目: Machine Learning (cs.LG)
*备注: 29 pages, 3 figures, 8 tables

点击查看摘要

Abstract:Deep learning has shown remarkable success in the data-driven modeling of dynamical systems. Much of its success is attributed not to the flexibility of neural networks but to inductive biases based on physical prior knowledge, such as energy conservation and symplecticity. However, existing methods do not fully exploit the fact that real-world physical systems are interconnections of components. Some methods require the interconnection to be known a priori, while others assume the system to be reducible to an ordinary differential equation (ODE) and learn only the reduced ODE, discarding the algebraic constraints imposed by the interconnection. Here, we propose Dirac-interconnected neural elements (DINEs), a neural network model that represents a physical system as a differential-algebraic equation (DAE), whose algebraic constraints are given by a Dirac structure in kernel representation. With DINEs, we simultaneously identify from data the interconnection among the components as a Dirac structure and learn the characteristics of the components as neural networks. This allows us to keep the learned subsystems in unreduced form and isolate or compose them to make a new system without retraining. Moreover, DINEs can handle partially observable systems. Experimental results demonstrate these capabilities on physical systems beyond the reach of existing methods.

[LG-58] Hyperparameter selection for equation learning with biologically-informed neural networks

链接: https://arxiv.org/abs/2610.02954
作者: William Lavery,Jodie A. Cochrane,John T. Nardini,Sara Hamis
类目: Machine Learning (cs.LG)
*备注: 18 pages, 11 figures, 2 tables (main text); 29 pages, 13 figures, 1 algorithm (appendices and supplementary material)

点击查看摘要

Abstract:Biologically-informed neural networks (BINNs) have emerged as a flexible subclass of physics-informed neural networks (PINNs) for learning terms in partial differential equations from data. BINNs are particularly suited for biological systems, where the governing equations are highly nonlinear and only partially known a priori, and where data observations are often sparse, noisy, and incomplete. However, applying BINNs effectively in practice depends critically on hyperparameter selection, which remains a central challenge in equation-learning frameworks. Hyperparameters are often chosen heuristically and only cursorily documented, which limits the reproducibility of results and the transferability of methods. We present a diagnostic workflow for hyperparameter selection that can be used when the ground-truth equations are not known. The workflow is guided by three main questions: (1) Are the benefits of greater network capacity worth the cost? (2) Do more training epochs keep reducing the validation loss? (3) Do the learned terms stop changing as network capacity and training increase? We apply our workflow to synthetic systems of varying complexity with known ground truth, spanning diffusion and growth right-hand side terms and data ranging from 1D+t to 2D+t. We demonstrate that the validation loss generally follows the true error in the learned terms and distil practical rules of thumb for selecting hyperparameters in the BINN architecture. By providing a structured workflow, practical guidelines, and suggested starting values for hyperparameter selection, this work lowers the barrier to BINN-based equation learning.

[LG-59] SlimKV: Joint Token-Feature KV Cache Compression with Reconstruction-Free Beacon Attention EMNLP2026

链接: https://arxiv.org/abs/2610.02953
作者: Zihan Teng,Jiayu Zhao,Wentao Ren,Minhao Fan,Tianrui Ma,Song Chen,Weichen Liu
类目: Machine Learning (cs.LG)
*备注: Accepted to EMNLP 2026 Main Conference (Oral)

点击查看摘要

Abstract:Long-context LLM serving is increasingly bottlenecked by KV-cache memory, especially in resource-constrained scenarios. Among existing KV-cache compression strategies, token-wise methods reduce cached states but risk information loss through eviction or condensation, while feature-wise methods reduce per-token KV dimensions but can require full-dimensional reconstruction to apply positional embedding, limiting decoding speedups. We introduce SlimKV, a question-agnostic joint token-feature KV-cache compression method. SlimKV uses low-rank-aware training to compress long contexts into beacon memory states with latent KV representations, together with layer-adaptive rank allocation. We further uncover a positional asymmetry: removing key-side RoPE affects beacon and raw tokens differently, with much smaller degradation for beacon tokens. Exploiting this asymmetry, SlimKV trains beacon KV projections under a K-RoPE-free constraint and enables latent-space attention during decoding, mitigating reconstruction latency. On LongBench, SlimKV outperforms baselines at 16x/32x compression and remains leading at 4x/8x, where it retains over 96% of the uncompressed model’s score. Needle-in-a-Haystack confirms robustness across evidence positions, and efficiency evaluation shows up to 7.34x attention speedup and 3.38x end-to-end decoding speedup over the uncompressed model at 128K length.

[LG-60] GTDD: Generative Test-Driven Development for AI Coding Agents with Adversarial Testing

链接: https://arxiv.org/abs/2610.02952
作者: Masahiro Kato
类目: oftware Engineering (cs.SE); Machine Learning (cs.LG); Methodology (stat.ME); Machine Learning (stat.ML)
*备注:

点击查看摘要

Abstract:Test-driven development gives AI coding agents executable requirements for implementing software. Because these agents can adapt their implementations to the examples they observe, passing a predetermined collection of tests can leave substantial parts of the intended behavior unimplemented. We propose Generative Test-Driven Development (GTDD), a formulation of test-driven development in which a separate testing agent generates new inputs after each candidate implementation is fixed, using a human-specified behavioral contract and the feedback from earlier rounds. A trusted evaluator checks these inputs, returns reduced counterexamples to the coding agent, and saves them for regression testing, so development continually confronts failures beyond the initial examples. We characterize the evidence that this process provides through a finite-population analysis of false acceptance under adaptive candidate selection. The resulting bounds quantify how test visibility and repeated feedback affect acceptance, and show that fresh random audits after candidate commitment control false acceptance across development rounds. In a paired experiment on a stateful key-value store, both policies that regenerated tests during development ended with lower mean failure rates than the policy whose tests were generated once by the same language model, and giving the tester the candidate’s source produced no detectable additional improvement. Further conditions requesting equal numbers of tests did not isolate any single feature of the policies as the source of this difference. GTDD combines this adaptive development feedback with established regression tests and an independent acceptance rule.

[LG-61] Learning Jazz Pianist Style with Cross-Attention Conditioning

链接: https://arxiv.org/abs/2610.02918
作者: Drew Edwards,Akira Maezawa,Simon Dixon
类目: ound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
*备注: 8 pages, 6 figures. Accepted at ISMIR 2026. Audio demos, code and checkpoints: this https URL

点击查看摘要

Abstract:Jazz pianists develop distinctive traits that experienced listeners can often identify within seconds, yet the features underlying this recognition resist formal description. We study jazz pianist style through the lens of a pretrained symbolic music transformer, showing that its learned representations already encode pianist identity well enough for highly accurate classification across two benchmarks. We then augment the transformer with cross-attention over learned pianist identity embeddings, enabling it to generate music conditioned on a specific artist’s style. Two evaluation protocols confirm that the generator captures meaningful stylistic structure: a sliding-window classifier consistently attributes conditioned continuations to the correct artist, far above unconditioned baselines; and a classifier trained entirely on synthetic generations identifies real pianists across 12 classes with 87% chunk-level and 95% song-level accuracy. Finally, we repurpose the classifier to locate the most characteristic moments within a performance, surfacing the specific musical gestures that distinguish each pianist’s voice.

[LG-62] Constraint-Aware Training NEURIPS2026

链接: https://arxiv.org/abs/2610.02909
作者: Jinwoo Kim
类目: Machine Learning (cs.LG)
*备注: Accepted at Neurips 2026 workshop AI for Verifiable Coding

点击查看摘要

Abstract:When generating programs with language models, constrained decoding can apply program analyses to exclude tokens that violate syntax, scope, or typing rules. However, there is a duplication: standard training already teaches the model to suppress the tokens rejected by these analyses. This duplication leads to the question: if we will perform some analysis to filter a set tokens out during inference anyways, can we avoid teaching the model the said analysis altogether during training, and does this externalization lead to more efficient models? This paper defines a general constraint-aware objective satisfying this externalization desideratum and formalizes the benefits of externalization into three concrete theorems about model size and data efficiency. We show, through a controlled synthetic experiment, that the theorems survive training dynamics: constraint-aware training yields lower prediction loss at a matched parameter count and data compared to ordinary cross-entropy training, motivating training objectives that incorporate the analyses used during generation.

[LG-63] Do ResNets Route? Sparse Interaction Experts in Residual Networks

链接: https://arxiv.org/abs/2610.02907
作者: Liang Yan,Siying Chen,Kaijie Chen,Bo Li,Jinghao Zhang,Mu Miao
类目: Machine Learning (cs.LG)
*备注: 46 pages, 14 figures, 18 tables

点击查看摘要

Abstract:Residual networks execute every block for every input, yet their functional contributions need not be input independent. We formulate a trained ResNet as a set function over binary residual-branch masks and apply Möbius inversion to decompose its output exactly into individual residual corrections and higher-order interactions. For smooth residual stacks, we show that each fixed k -way interaction scales as \mathcalO(\lambda^k) under residual scaling. Exhaustive analysis of ImageNet-pretrained ResNet-18 and ResNet-34 reveals that interaction mass peaks at orders five and ten, respectively, rather than at low orders. Reducing the residual scale shifts both spectra toward lower orders, but also changes model predictions. The interaction coefficients are concentrated in magnitude but not hard sparse, and prediction-preserving sparsity weakens with depth. Crucially, the dominant interactions vary across inputs and predicted classes around a shared global core, while their overall complexity changes little with sample difficulty. These results show that dense ResNets implement an implicit form of soft routing: every block is executed, but different inputs rely on different residual interaction experts. Routing can therefore emerge at the level of functional contribution without an explicit router or sparse execution.

[LG-64] angent Schrödinger Bridge Matching: Learning Stochastic Transport with Mechanistic Sensitivities

链接: https://arxiv.org/abs/2610.02906
作者: Jowaria Khan,Elizabeth Bondi-Kelly
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Predicting how stochastic systems respond to changes in viscosity, reaction rates, or external forces requires costly simulations, motivating reusable learned models. Yet matching observed outcome distributions does not ensure accurate intervention responses. We introduce Tangent Schrödinger Bridge Matching (Tangent-SBM), which learns stochastic transports from endpoint observations and mechanistic sensitivities. It propagates parameter derivatives alongside trajectories and supervises them against supplied targets. For average-response targets, single-rollout squared error also penalizes response variability; our objective uses two independent rollouts to match the mean without this additional penalty. We establish conditions under which sensitivity accuracy bounds finite-change prediction error and decision regret. Across Gaussian, stochastic double-well, PDEBench reaction–diffusion, and stochastic Navier–Stokes systems, Tangent-SBM improves sensitivity and finite-change prediction over matched conditional-bridge baselines while maintaining comparable endpoint and distributional accuracy. Controls examine target correctness, response objectives, and simulator-budget allocation. To test decision usefulness, we evaluate calibrated viscosity selection in Navier–Stokes: Tangent-SBM reduces tracking error relative to taking no action on every evaluated task.

[LG-65] oward Omni Multimodal Graph Foundation Model: A Topology-Driven Binding Approach

链接: https://arxiv.org/abs/2610.02881
作者: Xunkai Li,Chenxi Wan,Yinlin Zhu,Wang Luo,Hongchao Qin,Rong-Hua Li,Guoren Wang
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Multimodal graph foundation models (MGFMs) seek to learn generalizable representations from large-scale graphs with heterogeneous node modalities. However, real-world Multimodal-Attributed Graphs (MAGs) often contain incomplete node attributes, limiting the scale and diversity of available pretraining corpora. Besides, existing MGFMs primarily incorporate graph topology as structural context, overlooking its role in guiding multimodal binding and shaping a unified representation space. To address these challenges, we propose GraphBind, a topology-driven approach that uses graph topology to bind rich modality information into a unified shared space. GraphBind is motivated by the stability of graph topology, which provides structural references and complementary semantic information for multimodal binding. Concretely, GraphBind uses topology to organize self semantics and reliable neighborhood semantics into a global shared space that integrates structure and semantics, and adapts this space to discriminative and generative tasks through lightweight interfaces. Extensive experiments against 11 representative baselines demonstrate that GraphBind achieves leading performance on both discriminative and generative tasks, with relative improvements of up to 28.1% over the strongest baseline.

[LG-66] Peer Effects in Signed Networks: Separating Influence Through Positive and Negative Ties

链接: https://arxiv.org/abs/2610.02872
作者: Xiaojing Du,Jiuyong Li,Lin Liu,Debo Cheng,Jixue Liu,Thuc Duy Le
类目: Machine Learning (cs.LG)
*备注: 12 pages

点击查看摘要

Abstract:Evaluating network interventions requires understanding how treatment affects people through their social relationships. Counting treated neighbors without distinguishing supportive and antagonistic ties can conceal opposing influences. We define effects through positive and negative ties, their interaction, and a sign-composition effect of reallocating treatment between the two types at a fixed total, and give their identification formulas. Under sign-blind assignment, we show how ignoring signs mixes the effects of the two tie types. We propose SiDE (Signed-exposure Doubly robust Estimator), which combines sign-specific outcome models with exposure probabilities induced by individual treatment assignment. We establish double robustness of its score and assess approximate intervals that account for overlapping neighborhoods. Semi-synthetic experiments on six real signed networks demonstrate accurate effect estimation and examine the limits of interval coverage. An exploratory reanalysis of published school-experiment data yields a positive estimate of the peer effect through spend-time ties on wristband wearing, but the intervals for all four effects include zero after adjustment for multiple comparisons. This framework can inform network intervention design by showing when influences through the two tie types reinforce or offset one another.

[LG-67] DIVINE: Simple Cross-Market Stock Pretraining via Diverse Indicator Reconstruction

链接: https://arxiv.org/abs/2610.02866
作者: Kuan-Yu Chen,Shu-Cheng Zheng,Yu-Chen Den,Wei-Cheng Liao,Tien-Hao Chang
类目: Machine Learning (cs.LG); Computational Engineering, Finance, and Science (cs.CE)
*备注:

点击查看摘要

Abstract:Financial time-series pretraining typically learns from masked observations, contrastive relations, or future outcomes—yet existing objectives struggle to simultaneously avoid future-supervision uncertainty and maintain return-prediction alignment. We propose DIVINE (DIVerse INdicator rEconstruction), a simple cross-market pretraining framework that reconstructs technical indicators from raw OHLCV history. Computed from observed price-volume history, technical indicators provide consistently defined supervision across markets while summarizing diverse market dynamics with established relevance to return prediction. Pretrained jointly on six-equity market datasets, DIVINE reconstructs 77 targets derived from 16 standard indicators and transfers only the learned encoder to downstream stock ranking. Across all six markets, DIVINE achieves the strongest average portfolio performance with a lightweight 0.05M-parameter encoder, outperforming pretraining baselines and matching or exceeding substantially larger financial foundation models, while remaining robust and data-efficient. Systematic analyses show that indicator diversity and market diversity provide complementary gains in transfer. Together, these results suggest that supervision design and cross-market diversity—rather than model scale—are the key drivers of strong, transferable financial representations.

[LG-68] On Unlearning for Time-series Forecasting

链接: https://arxiv.org/abs/2610.02865
作者: Zeyu Shi,Yanhui Luo,Ziming Hong,Chongyang Gao,Kezhen Chen,Shanshan Ye,Lixu Wang
类目: Machine Learning (cs.LG)
*备注: 22 pages

点击查看摘要

Abstract:Time-series forecasting is widely used in sensitive domains. Models in these settings are often trained on longitudinal user- or entity-level records, which may later require removal because they contain sensitive or proprietary information or have been corrupted by sensor failures. To address such deletion requests without costly retraining, machine unlearning has been widely studied as a practical mechanism for privacy protection and data governance. However, the application of machine unlearning to time series prediction has not yet been well realized; this is mainly due to the following unique challenges: Gradient-based unlearning can be unstable because a deleted observation participates in multiple causally connected forecasting windows, causing parameter updates to propagate beyond the requested interval and degrade retained forecasting utility. Label-guided updating offers a more controlled alternative, but continuous and context-dependent forecasts lack a suitable replacement target, while the exact-retrained output is unavailable during unlearning. Moreover, the remaining support for a deleted temporal pattern is highly non-uniform. Some affected windows retain structurally similar counterparts in the retained data, whereas others become underrepresented or isolated. We present RDTU, a Residual Diffusion framework for time-series unlearning. RDTU first uses a retained-set neural tangent kernel predictor to obtain a deletion-compatible base forecast. Then it quantifies the global and local structural support of each affected window using the volume contribution of the retained-reference data. Then a diffusion model generates a residual correction that estimates the counterfactual forecast, yielding a pseudo-label field that guides a lightweight model update. Experiments show that RDTU consistently produces unlearned models that most closely match exact retraining.

[LG-69] NeuroLens: Learning Latent Embeddings of Neural Semantics from Chronic Recordings

链接: https://arxiv.org/abs/2610.02864
作者: Hanrui Lyu,Baiyuan Chen,Tianshu Tan,Matthew R. Whiteway,Maxwell D. Melin,Ji Xia,Linyang He,Bradly C. Stadie,Anne Churchland,Liam Paninski,Yizi Zhang
类目: Machine Learning (cs.LG); Neurons and Cognition (q-bio.NC)
*备注:

点击查看摘要

Abstract:Understanding how neural activity represents higher-order cognition and how these representations evolve over time has long been a central pursuit in neuroscience. However, current analytical tools cannot easily distinguish representational plasticity from recording instability in chronic neural recordings. Here, we introduce NeuroLens (Latent Embeddings of Neural Semantics), a self-supervised model based on the Joint-Embedding Predictive Architecture (JEPA) framework that learns denoised, semantically informative latents from chronic neural recordings. An adaptive encoder maps changing neural populations into a common latent space, while a temporal predictor learns structure that supports prediction of future latent states. By predicting in latent space, NeuroLens captures temporally predictable structure and reduces sensitivity to transient, recording-specific variability. Across chronic intracortical data in mice and humans, the learned representations improve decoding of decision-making and semantic task variables. Multi-day pretraining enables generalization to future sessions, rapid few-shot adaptation to unseen neural populations, and more stable decoding over time than state-of-the-art baselines. Together, these results establish NeuroLens as a new paradigm for studying how neural representations change during learning and over long timescales.

[LG-70] Counterfactual Action Evaluation Observation Bottlenecks and Representation Geometry in Joint-Embedding Predictive World Models NEURIPS2026

链接: https://arxiv.org/abs/2610.02860
作者: Arjun Subramanian
类目: Machine Learning (cs.LG)
*备注: 15 pages, 7 figures. Published in the 40th Conference on Neural Information Processing Systems (NeurIPS 2026), Workshop on Physical World AI: Geometry, Characteristics, and Multimodal Sensing. Replication Package: this https URL

点击查看摘要

Abstract:Low latent prediction error does not establish that a world model distinguishes the consequences of its actions. We introduce an evaluation protocol that traces the same intervention through simulator state, raster observations, target embeddings, and predictor outputs. Exact simulator-state forks in a controlled deformable-physics testbed reveal distinct bottlenecks. Changed commands alter particle motion, yet 41.5% of one-step raster pairs are identical. Observation loss is not the whole explanation: among 579 high-visibility counterfactuals, median predictor-to-target response is 0.0051 and 0.0217 across two seeds, falling to 0.0027 and 0.0116 after variance normalization. An isotropic state perturbation matched to the target counterfactual embedding shift produces 190x and 53x larger predictor changes on the same visible pairs, isolating action-path under-use rather than a dead or globally shrunk predictor. MSE-only training gives 8.36x lower 10-step latent error in matched seeds, but in spectrally concentrated spaces; one VICReg target encoder is also strongly concentrated, so neither error nor rank alone certifies physical state. Finally, stiffness remains near chance even from full-resolution rasters and mechanical state while privileged material parameters decode perfectly, indicating weak identifiability under this excitation rather than encoder discard. These results motivate auditing physical effect, observation visibility, representation geometry, and action dependence separately.

[LG-71] Understanding Enrichment in Reinforcement Learning

链接: https://arxiv.org/abs/2610.02846
作者: Jinwoo Kim,Shraddha Barke
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:When rewards are sparse, reinforcement learning with verifiable rewards (RLVR) often uses hints or intermediate guidance to generate more successful rollouts. This enrichment biases policy-gradient updates unless corrected via importance weights, but existing methods omit correction or truncate importance weights in order to avoid the high variance of correction. It thus remains unclear what exactly is gained or lost in RLVR by correcting enriched rollouts. We show, mathematically, that omitted or truncated correction implicitly reweights the defined reward, and we decompose the resulting gradient error into scale, rotation, and variance to explain their distinct effects on learning. To make correction practical, we develop a novel sequential Monte Carlo (SMC) weight correction mechanism that, under stability and particle-order assumptions, tempers the exponential compounding of standard correction variance over the length of a sample to an additive accumulation. We then apply our analysis of enrichment to interpreting the results of a fine-tune of Qwen3-1.7B on a sparse band of OpenMathReasoning, establishing a concrete mechanism of how enrichment, both corrected and uncorrected, helps avoid collapse in sparse domains. Our main contribution is to understand, in general, how enrichment and correction can affect training, as opposed to claiming that either mode of operation is superior to unenriched RL.

[LG-72] All Work And No Play Makes Jack a Dull Boy: Understanding and Preventing Catastrophic Strategy Collapse in RLVR

链接: https://arxiv.org/abs/2610.02835
作者: Qiyuan Huang,Tianshi Xu,Meng Li
类目: Machine Learning (cs.LG)
*备注: 84 pages, 10 figures

点击查看摘要

Abstract:During post-training of large language models (LLMs) with Reinforcement Learning with Verifiable Rewards (RLVR), GRPO-style algorithms can exhibit severe late-stage collapse. Prompt-based probing reveals that this is not benign strategic pruning, but a harmful contraction of effective strategy capacity that makes distinct reasoning strategies increasingly inaccessible. To characterize this phenomenon, we define strategies through trajectory-level policy-update interactions and develop a unified theoretical framework combining optimization dynamics and information theory. We prove that major RLVR objectives progressively concentrate probability mass onto a single strategy, while sustaining nontrivial task accuracy requires a minimum strategy capacity. The conflict between these two results provides a mechanistic explanation for catastrophic collapse. We further derive the Mirrored Entanglement Index (MEI) as a lightweight online warning signal. To prevent collapse, we propose \textbfMesh Learning, which exposes multiple reasoning strategies and prevents any single strategy from dominating optimization. Across AIME26, AIME25, MATH-500, GPQA, and LiveCodeBench, Mesh Learning consistently outperforms strong baselines across Qwen and Phi model families, with gains of up to 13.4 pp and 11.5 pp, respectively. These results establish strategy preservation as a key principle for stable RLVR. Code is available at this https URL.

[LG-73] Gated Slot Attention-2: Two-Sided Associative Memory Correction in Linear Attention

链接: https://arxiv.org/abs/2610.02816
作者: Ruijie Li,Shengnan Ding,Weimin Zhang,Derick Tang,Zhanpeng Zeng,Qinsong Zeng,Ming Chen,Jiaxi Hu,Yuxuan Liang
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Linear attention models have emerged as efficient alternatives to standard attention, but effectively managing their fixed-size recurrent memory remains challenging. To improve memory, recent work has explored two distinct directions: delta-rule variants for precise correction of values associated with keys, and slot-based architectures such as Gated Slot Attention for modeling key and value memories in two stages. We observe that these directions are complementary–the delta rule provides effective memory correction, while the two-stage structure provides a natural way to operate on both sides of an association. Building on this insight, we introduce a new Gated Oja Rule for key-side correction and extend it with decoupled erase and write control to obtain Gated Oja Rule-2. We then introduce Gated Slot Attention-2 (GSA2), which combines Gated Oja Rule-2 for key-side correction with Gated Delta Rule-2 for value-side correction through shared latent slots. We further derive a hardware-efficient chunkwise algorithm for parallel training. Experiments demonstrate that GSA2 consistently improves over strong linear-attention baselines across benchmarks while retaining linear-time sequence modeling and constant-memory recurrent decoding.

[LG-74] Muon Learns Facts Better: Understanding the Role of Spectral Orthogonalization

链接: https://arxiv.org/abs/2610.02798
作者: Xuheng Li,Qiwei Di,Yuan Cao,Quanquan Gu
类目: Machine Learning (cs.LG); Optimization and Control (math.OC); Machine Learning (stat.ML)
*备注: 47 pages, 8 figures

点击查看摘要

Abstract:The Muon optimizer applies spectral orthogonalization to matrix-valued updates and has shown strong performance in large-scale neural network training, yet the mechanisms of this transformation in feature learning remain poorly understood. In this work, we investigate this question through a tractable factual-recall model, where a fact maps each subject-relation pair to an answer, and a linear transformer learns the subject- and relation-dependent information required to recover this mapping. The transformer is optimized with gradient flow (GF), spectral GF, or Sign GF, which are continuous-time limits of gradient descent, Muon, and Adam, respectively. Prior studies (Nichani et al., 2025) have shown that when the number of subjects exceeds the number of relations, GF learns relation-dependent information before subject-dependent information, producing a feature-separation phase during training. We characterize this separation with the learning times when the subject- and relation-dependent components of the prediction reach a target accuracy. With S subjects and R relations, GF has a learning-time ratio of \widetilde\Theta(\sqrtS/R) , whereas Spectral GF reduces this ratio to \widetilde\Theta(1) . In addition, for fixed S and R , the subject- and relation-dependent errors decay as 1/(T\log T) in training time T under GF, but as \exp(-\mathrmpoly(T)) under spectral GF. Finally, we show that GF and spectral GF are equivariant under orthogonal transformations of the token embeddings, whereas Sign GF is not: Different orthonormal embeddings can potentially produce no feature separation, a large feature-separation phase, or even a reversed learning order. These results provide a mechanistic view of how spectral orthogonalization can fundamentally reshape feature-learning dynamics.

[LG-75] Efficient Memory Crystallization for Graph Learning under Non-Stationary Distribution Shifts NEURIPS2026

链接: https://arxiv.org/abs/2610.02795
作者: Yue Hou,Ruomei Liu,Yingke Su,Junran Wu,Ke Xu
类目: Machine Learning (cs.LG)
*备注: Accepted by the 40th Conference on Neural Information Processing Systems (NeurIPS 2026)

点击查看摘要

Abstract:Deep graph learning models deployed in real-world systems often need to cope with non-stationary environments, where the underlying graph distribution drifts continually over time. Prevailing solutions rely on training auxiliary generative modules to synthesize memory graphs for cross-domain adaptation, which incurs substantial computational overhead and scales poorly under prolonged distribution shifts. We argue that a more economical path exists: rather than generating memory, one can crystallize it. To this end, we propose Efficient Memory Crystallization (EMC), a training-free test-time framework that distills each incoming graph domain into a compact, semantically faithful memory through a closed-form solution to a memory-oriented distribution-matching objective, thereby eliminating redundant domain information under continual covariate shifts. To preserve both generalizability and adaptability as the model traverses a long sequence of target domains, EMC further models inter-domain dependencies through state-evolving memories and admits a theoretically grounded, tighter generalization error bound than direct adaptation. Extensive experiments demonstrate the superior performance of EMC over state-of-the-art baselines on graphs under non-stationary distribution shifts, while reducing average runtime by 87.4% and GPU memory consumption by 92.4% relative to the recent competitor, making continual graph adaptation practical at scale.

[LG-76] Controlling Polar Exposure to Delay Memorization in Diffusion Models

链接: https://arxiv.org/abs/2610.02780
作者: Xuanchen Wang,Heng Wang,Weidong Cai
类目: Machine Learning (cs.LG)
*备注: 31 pages, 4 figures

点击查看摘要

Abstract:Diffusion models can reach useful sample quality before copying training examples, but fast optimization can compress this generalization window by accelerating sample-specific fitting. We investigate this effect through update geometry and propose Quality-Gated De-whitening (QGD), a controller that retains a fast polar-update prefix and progressively restores fixed-gain momentum. Our random-feature analysis separates covariance-controlled, curvature-equalized and amplitude-controlled memorization clocks. Under aligned spectral assumptions, it establishes a finite-exposure condition under which a fixed-gain tail recovers a delay proportional to dataset size. QGD implements this principle with a confirmed quality gate, a bounded decay envelope and causal copy feedback. Immediate switching is the conservative limit; gradual control balances delayed copying against continued quality improvement. We pair QGD with Copy-Budgeted Selection (CBS), which applies simultaneous binomial calibration to a frozen checkpoint family, followed by a fresh evaluation of the released checkpoint. On 2,000-image CIFAR-10 subsets, QGD preserves the polar baseline’s quality-arrival time while expanding its useful interval by 8.32x and reducing common-checkpoint copying by 75.9%. With identical calibration and independent quality evaluation, QGD achieves FID 75.56 versus 79.37 for SGD with the same selector. Exposure-matched controls, independent detector audits and transfer to flow matching and dance generation support adaptive exposure control as a practical way to improve the quality-copying tradeoff.

[LG-77] LatticeSMC: Where to Spend Inference-Time Compute in Chunked Sequence Generators

链接: https://arxiv.org/abs/2610.02774
作者: Xuanchen Wang,Heng Wang,Weidong Cai
类目: Machine Learning (cs.LG)
*备注: 29 pages, 7 figures

点击查看摘要

Abstract:Long-form generators for music, motion and video produce sequences chunk by chunk, with each chunk generated by iterative denoising while rewards are defined over the full sequence. Existing inference-time steering methods typically act on one axis at a time: best-of-N at the end, Feynman-Kac steering across denoising steps, or streaming pruning across chunks, and are often compared under unmatched compute or different return rules. We introduce budget-matched chunked steering and propose LatticeSMC, a sampler derived from a Feynman-Kac model on the two-dimensional lattice of chunk index and denoising step. Two telescoping results make its design exact: for chunk-additive rewards, the two axes induce identical weights, so resampling should occur where lookahead is cheapest; for terminal rewards, any prefix score defines an exact intermediate potential, making prefix-evaluable rewards twists with no estimation or extra denoiser calls. LatticeSMC resamples on these potentials at chunk boundaries and, when scoring is free, within chunks, returning either a weighted draw or the best particle. Under matched compute, on music-to-dance diffusion and 40-second text-to-music generation, it raises beat alignment from 0.234 to 0.441 (best-of-N: 0.354) and prompt adherence from 0.470 to 0.560 at 32 particles, while preserving held-out quality. It also retains its advantage on long-range rewards and is preferred by human raters in 60-77 percent of pairwise comparisons. Finally, we show that commitment strength should follow the information in the current potential, while the value of lookahead is predicted by the within-set predictability of future reward.

[LG-78] No-Free-Graph: Learning When Multimodal Data Should Be Graphified

链接: https://arxiv.org/abs/2610.02768
作者: Zekai Chen,Kai Hu,YuXin Zeng,Xunkai Li,Xun Wu,Yinlin Zhu,Zhengyu Wu,Xu Wang,Rong-Hua Li
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Multimodal graph learning has recently emerged as an effective paradigm for in corporating inter-entity relationships into multimodal representations. Existing studies have made substantial progress on how to construct and optimize graphs, but rarely consider a more fundamental question: whether additional relational structures should be introduced for a given dataset and task. Through empirical studies across diverse datasets, tasks, and graph constructors, we reveal that graphification is not consistently beneficial: introducing relational structures can provide substantial improvements in some cases, while offering limited or even negative gains. This observation motivates a new perspective that graph construction should be treated as a selective decision based on its expected utility rather than a default preprocessing step. To address this issue, we propose MAG-SCOUT, a pre-construction graph assessment framework that estimates whether introducing graph structures is beneficial before generating the complete topology. MAG-SCOUT collects limited relational evidence, analyzes its potential taskspecific contribution, and estimates the expected utility of graphification together with construction cost to make a build-or-skip decision. Extensive experiments across six multimodal datasets, three downstream tasks, and diverse graph constructors demonstrate that MAG-SCOUT effectively identifies when graph structures should be introduced, saving 33.1% of task-macro graph work while retaining 96.7% of held-out positive-gain mass under the pre-registered floor.

[LG-79] Exact Memory-Time Optimization for Prefix-Cached Language Model Serving

链接: https://arxiv.org/abs/2610.02766
作者: Shivam Gupta
类目: Machine Learning (cs.LG)
*备注: 16 pages, 5 figures. Code and reproducibility artifacts: this https URL

点击查看摘要

Abstract:Retaining language-model prefix states trades recomputation against storage time. Optimizing each cached block independently can overcount savings: a resident block is usable only when the required preceding prefix is also available. We introduce Prefix-Certificate Retention (PCR), an exact finite-trace formulation for static, grouped, reset-on-access timeouts. Usable-prefix rewards become nodes whose prerequisites are timeout thresholds and preceding hit certificates. The resulting maximum-weight closure reduces to one minimum cut, with graph size linear in the number of block lookups and timeout choices. A breakpoint theorem extends the construction to all nonnegative timeouts without discretization error. We also derive a linear-time-in-grid-size dynamic program for ordered timeouts and bounds that certify the cost of this restriction. Exhaustive small-instance checks and chronological replay of 39,632 public Mooncake requests validate the formulation. On the fixed grid, ordered timeouts attain the unrestricted training optimum in 118 of 120 trace-grouping-price cases. Heterogeneous retention improves several held-out memory-time tradeoffs, but finer training optimization does not uniformly improve transfer. The contribution is a tractable optimization model and an auditable benchmark for retention policies; the experiments measure usable prefix blocks and storage time, not GPU latency.

[LG-80] Localized Conformal Safety Monitoring with Vision-Language Models for Autonomous Driving

链接: https://arxiv.org/abs/2610.02765
作者: Luís Marques,Rong Fang,Disha Kamale,Dmitry Berenson
类目: Robotics (cs.RO); Machine Learning (cs.LG)
*备注: 5 pages, 2 figures, 1 table. Extended abstract. Disha Kamale and Dmitry Berenson are joint senior authors

点击查看摘要

Abstract:Monitoring planned driving trajectories requires accurately estimating the collision likelihood with actors whose motion is itself impacted by the ego motion. Existing classical approaches are often limited by the quality of their forecasting model. Vision-language models (VLMs) have shown promise in reasoning about the consequences of high-level actions, yet their approximate predictions are unsuitable for safety-critical applications such as autonomous driving. Conformal prediction (CP) has emerged as a data-driven framework for quantifying the uncertainty of black-box model predictions. We propose Split Label-Localized Conformal Prediction (SLLCP), a post-hoc calibration layer over frozen VLMs that transforms their unreliable predictions into probabilistically calibrated safety prediction sets. We consider how the ability to estimate safety can depend on the observed driving scene and introduce a localized procedure that upweights relevant past experience when calculating uncertainty thresholds. We provide label-conditional finite-sample distribution-free coverage under exchangeability. Evaluated over 15k CARLA trajectories from unseen scenarios, SLLCP correctly flags 89.6% of collision-causing trajectories with a Qwen backbone and 88.4% with a Cosmos backbone, while the base VLMs only flagged 4.6% and 39.1% of the collision-causing trajectories, respectively. These results indicate that local, label-conditional calibration can reduce missed unsafe trajectories.

[LG-81] A Controlled Audit of Personal AI Memory for Rating Prediction

链接: https://arxiv.org/abs/2610.02764
作者: Shivam Gupta
类目: Machine Learning (cs.LG)
*备注: 15 pages. Code and reproducibility materials: this https URL

点击查看摘要

Abstract:In structured rating prediction, does a personal AI use historical item-rating associations, or mainly the user’s rating tendencies? We audit this distinction by permuting historical ratings within each user while preserving the exact rating distribution, item support, and metadata. We combine this control with full history, native memory extraction, and matched numerical readers in a publicly frozen evaluation of 400 held-out user profiles and 6,160 target ratings across Coat and MovieLens. On Coat, the tested Qwen-written Mem0 pipeline increases user-macro mean absolute error relative to full history by 0.084 for Qwen and 0.149 for Phi; both family-adjusted bootstrap intervals exclude zero. Correct historical assignments help both readers on Coat, but the corresponding MovieLens effects are smaller and inconclusive after adjustment. A history-only ridge reader outperforms Qwen in both domains and Phi on MovieLens, while the Coat Phi comparison is unresolved. All 2,800 reader calls, including 150 invalid outputs, are retained under a fixed fallback rule. A separate implementation verifies inputs, metrics, and all ten primary contrasts. The contribution is a reproducible diagnostic study showing why extraction, association use, output reliability, and reader choice require separate evaluation.

[LG-82] A Two-Stage Cascade for Near-Real-Time Forest Anomaly Detection from Sentinel-1 SAR Time Series

链接: https://arxiv.org/abs/2610.02763
作者: Pann Thinzar Seint,Subas Chhatkuli,Bryan Atwood
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Tropical forest monitoring is essential for global climate stability and biodiversity preservation. To address the urgent need for rapid, reliable detection of forest loss which is essential for timely intervention against illegal logging, supply chain transparency, land-use governance and carbon market standards, we introduce a two-stage statistics-encoder cascade for near-real-time anomaly detection using Sentinel-1 time series. Our system is designed to overcome two fundamental challenges in remote sensing: the cloud-cover limitations that restrict optical monitoring and seasonal backscatter variation that causes SAR systems to mistake natural moisture changes for forest loss. The architecture integrates two distinct analytical engines to ensure high-fidelity detection: (1) an adaptive, robust-statistics z-score test on co-registered Sentinel-1 VH backscatter, same-season historical baseline and (2) a learned confirmation gate based on the latent-space structural similarity (SSIM) of a convolutional autoencoder trained on stable-forest patches. A candidate disturbance is confirmed as an alert only when both stages agree, and is assigned a confidence score and a Low/Medium/High risk tier from its repeat-occurrence history. The system produces per-alert auditable confidence scores and area-in-hectares estimates directly compatible with Monitoring, Reporting and Verification (MRV) workflows, sustainable forestry management, operational field checks and environmental risk assessments. Beyond its primary application, the model’s flexibility allows for critical environmental applications ranging from selective logging to large-scale agricultural encroachment mapping, flood mapping and so on.

[LG-83] Jumping up and down: Denoiser diffusion models for discrete ordinal data

链接: https://arxiv.org/abs/2610.02754
作者: Yair Shenfeld,Ricardo Baptista,Stefano Peluchetti
类目: Machine Learning (cs.LG); Methodology (stat.ME)
*备注: 39 pages, 3 figures

点击查看摘要

Abstract:Diffusion models are highly developed in continuous spaces for image and video domains. Recently, major advances have been made for discrete diffusion models for categorical data, specifically in the language domain. In contrast, diffusion models for discrete integer-valued data are less developed, despite the prevalence of this modality, ranging from images and music to gene counts. We introduce Jumping Up and Down (JUD)—a new family of denoiser-based diffusion models for discrete ordinal data. This is the first family of diffusion models for ordinal data which centers around training denoisers, which at the same time allows for bi-directional (up and down) perturbations of the data. The simplicity of the training objective, combined with the flexibility of bi-directional perturbations, leads us to obtain competitive results across different data modalities.

[LG-84] Inner Momentum for Differentially Private Muon

链接: https://arxiv.org/abs/2610.02738
作者: Bishnu Bhusal,Minh Vu,Ben Southworth,Geigh Zollicoffer,Rohit Chadha,Manish Bhattarai
类目: Machine Learning (cs.LG); Cryptography and Security (cs.CR)
*备注: LA-UR Number: LA-UR-26-28799

点击查看摘要

Abstract:Differentially private training clips each per-example gradient before adding noise. This clipping is radial for each example, yet unequal clipping factors can distort the relative singular-vector geometry of their average. Muon is particularly exposed to this effect, since its update is an approximate polar factor UV^T that depends only on the singular vectors that clipping can shift. To curb this degradation, we propose averaging each sampled example’s Muon gradient over the current model and a short history of recent models before clipping. The clipped batch matrix then separates into a common rescaling and a covariance residual R between sampled gradients and clipping values, with ||R||_F = sigma_lambda sigma_G, bounding the clipping-induced distortion directly. We further show that a finite Newton-Schulz iteration preserves the polar factor of its input under these spectral conditions, confirming that our correction survives orthogonalization. In private GPT-2 fine-tuning on E2E and DART at epsilon in 1, 2, 4, 8, DP-Muon-IM improves BLEU and ROUGE-L over DP-Muon in every seed-matched comparison, and non-private diagnostics show 2-4% lower pre-noise polar error.

[LG-85] Bellm an Error Minimization Via Linear Programming Normalization

链接: https://arxiv.org/abs/2610.02730
作者: Haining Yu
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:This paper proposes a new functional approximation approach to reduce Bellman error in high-dimensional dynamic programming and Reinforcement Learning problems. Using a classic dynamic programming problem (network capacity control in revenue management) as the motivational example, the paper illustrates that deep neural networks and linear programming approximation algorithms can be combined to derive approximate solutions to dynamic programming problems. Simulation results show the proposed approximation algorithms achieves competitive performance when compared with benchmark.

[LG-86] Structural-Functional Brain Connectivity Generation via Multimodal Hypergraph-based Flow Matching

链接: https://arxiv.org/abs/2610.02722
作者: Chyong Yi Poh,Hwa Hui Tew,Junn Yong Loo,Raphaël C.-W. Phan,Fuad Noman,Pew-Thian Yap,Chee-Ming Ting
类目: Machine Learning (cs.LG)
*备注: 10 pages, 3 figures, 5 tables

点击查看摘要

Abstract:Structural connectivity (SC) and functional connectivity (FC) provide complementary information on interactions between brain regions and are widely used in neuroimaging studies of neuropsychiatric disorders. Generative modelling can alleviate the scarcity of large-scale paired SC-FC data, but existing approaches typically use pairwise graphs that capture only dyadic interactions and often generate SC and FC independently, limiting preservation of higher-order structure-function relationships. We propose a Multimodal Hypergraph Flow Matching (MHG-FM) framework for joint SC-FC connectivity generation and cross-modal translation. MHG-FM constructs modality-specific hypergraphs, learns higher-order representations with Hypergraph Neural Network (HGNN) encoders, and performs bidirectional cross-modal fusion using Dual Cross-Attention (DCA). A variational autoencoder maps the fused representations to a compact latent space, where conditional flow matching enables connectivity synthesis and multimodal translation via latent transport. Experiments on the Human Connectome Project Young Adult (HCP-YA) dataset show that MHG-FM outperforms several state-of-the-art baselines in reconstruction quality, topology preservation, distributional similarity, and SC-FC coupling, while achieving approximately 8x faster sampling than a matched diffusion backbone.

[LG-87] Differential Privacy of Gradient Descent on Perturbed Objectives

链接: https://arxiv.org/abs/2610.02716
作者: Austin Watkins,Raman Arora
类目: Machine Learning (cs.LG); Cryptography and Security (cs.CR); Machine Learning (stat.ML)
*备注: 50 pages

点击查看摘要

Abstract:Objective perturbation adds a random linear term to a regularized empirical risk and releases the exact perturbed minimizer. We study the finite computation obtained by releasing the N -th iterate of deterministic gradient descent on w\mapsto F(w;S)+\langle z,w\rangle , where z\sim\mathcal N(0,\sigma^2I_d) is drawn once before optimization. For strongly convex and smooth objectives with Lipschitz Hessian, we prove an explicit condition under which the map z\mapsto w_N is a C^1 -diffeomorphism on the bounded domains used in the privacy argument, with a quantitative lower bound on the smallest singular value of its Jacobian. This permits a direct change-of-variables analysis of the finite iterate. For generalized linear models, the resulting privacy-profile bound has no explicit ambient-dimension factor once the iteration condition holds, and its finite-iteration correction decreases geometrically. By letting the free truncation parameter grow slowly with N , we recover the corresponding exact-minimizer certificate in the limit. We also bound the expected excess empirical risk by d\sigma^2/(2\mu) plus a geometrically decreasing optimization term, and transfer the result to population risk without an additional multiplicative condition-number factor in the leading statistical terms.

[LG-88] Conditional Capacity and Routing in Mixture-of-Experts Particle Transformers

链接: https://arxiv.org/abs/2610.02701
作者: Kaushik Pendiyala,Haris Zia,Trevin Lee,Timothy Legge,Alejandro J. De Leon,Zihan Zhao,Aaron Wang,Abhijith Gandrakota,Jennifer Ngadiuba,Richard Cavanaugh,Javier Duarte
类目: Machine Learning (cs.LG); High Energy Physics - Experiment (hep-ex)
*备注: 8 pages, 4 figures. Submitted to the ML4PS 2026 workshop

点击查看摘要

Abstract:Mixture-of-Experts (MoE) models can increase parameter capacity without proportionally increasing active computation, but it is unclear how this trade-off behaves in particle-physics transformers. We study dense and MoE Particle Transformers on 188-class JetClass-II, varying expert count, routing capacity, top-K, and auxiliary loss. We find that, when token dropping is avoided, top-1 MoE models improve over the dense baseline at nearly unchanged nominal forward compute, while further increasing the number of stored experts produces little additional accuracy gain. Activating multiple experts per token yields additional predictive improvements at higher computational cost. Routing analyses show that expert assignments become more strongly associated with particle identity and kinematics in some configurations, but this structure does not increase monotonically with classification performance. These results highlight the need to distinguish stored parameter capacity, active computation, routing capacity, and routing organization when evaluating sparse expert models for jet classification. Code and experiment configurations are available at this https URL.

[LG-89] RAOA: Alternating-Operator Neural Computation with Programmable Radio Propagation

链接: https://arxiv.org/abs/2610.02683
作者: Toshiaki Koike-Akino
类目: Machine Learning (cs.LG); Emerging Technologies (cs.ET); Logic in Computer Science (cs.LO)
*备注: 31 pages, 6 figures

点击查看摘要

Abstract:Can programmable radio propagation serve as computational depth rather than only as a communication channel or one-shot analog transform? We introduce the Radio Alternating Operator Ansatz (RAOA), a recurrent computing architecture that alternates an energy-derived problem update with a mixing update over a persistent latent state. Recomputing the problem field after each mix makes repeated passes compositional even when the same learned controls are reused across depth. We evaluate this idea through exact discrete optimization, constrained programmable-propagation simulation, and pretrained-model adaptation. On discrete objectives, repeated execution can improve solution quality without increasing the learned-control count, and the same formulation handles higher-order interactions directly. A passive phase-only free-space model further shows that the required operators can be approximated by programmable propagation while retaining useful downstream behavior despite realization error. When inserted as a zero-initialized residual adapter, RAOA adapts pretrained language models with WikiText performance close to a matched shallow MLP across three model families, while reasoning-task transfer remains model-dependent. Together, these results connect alternating-operator computation, programmable radio propagation, and neural adaptation within one recurrent framework. The RF realization evidence is simulation-based rather than a hardware demonstration.

[LG-90] Mind the Refinement Gap: When Safe High-Level Robot Plans Produce Unsafe Executions

链接: https://arxiv.org/abs/2610.02662
作者: Stabak Das,Priyesh Ranjan,Xiangfang Li,Lijun Qian
类目: Machine Learning (cs.LG); Robotics (cs.RO)
*备注: 5 pages, 1 figure

点击查看摘要

Abstract:Language-enabled robot systems increasingly combine semantic-graph planning with temporal-logic safety monitors. We investigate a trace-completeness assumption in these systems: whether the high-level action sequence checked by a monitor represents the navigation and implicit action effects induced during execution. We audit this assumption in RoboGuard by comparing its verdict on a surface plan with its verdict on a graph-refined trace under the same Linear Temporal Logic (LTL) specification. Our evaluation comprises 28 controlled cases spanning five action-abstraction families and 14 end-to-end cases in which SPINE [1] generates plans from natural-language instructions while RoboGuard generates scene-grounded safety specifications. In the controlled evaluation, all 12 targeted abstraction cases exhibit the predicted surface-versus-refined discrepancy while all 16 controls behave as expected, motivating graph-based trace refinement as a lightweight mitigation and a diagnostic tool for physical-AI safety monitors.

[LG-91] AIGS: Adaptive Incremental Gating System for Online Representation Learning in Non-Stationary Data Streams

链接: https://arxiv.org/abs/2610.02661
作者: SiRui He,Kai Liang Lew,Chui Zi Ong,Chean Khim Toa
类目: Machine Learning (cs.LG)
*备注: 11 pages, 6 figures, 8 tables

点击查看摘要

Abstract:Real-time data streams in Web of Things (WoT) and edge computing environments often evolve through latent regime changes. For online representation learning under strict computational constraints, the central problem is resolving the stability-plasticity dilemma: keeping useful historical knowledge while rapidly reacting to concept drift. Existing methods employ fixed update schedules or rolling windows. However, they suffer from parameter ossification during sudden shifts and waste computational resources when the stream remains stable. This paper proposes the Adaptive Incremental Gating System (AIGS), a lightweight closed-loop state-aware adaptation framework. AIGS introduces the Shock Ratio, an endogenous residual feedback mechanism that normalizes current reconstruction error against recent variation. This signal drives a Continuous Plasticity Controller that smoothly interpolates between learning plasticity and memory retention. By treating representation learning as a closed-loop control mechanism, AIGS avoids catastrophic forgetting and maintains a strictly linear \mathcalO\left(k\cdot d\right) per-step complexity suitable for latency-sensitive edge devices. Experiments on real-world smart city dynamic streams-spanning traffic networks, meteorological systems, and industrial infrastructure-demonstrate distinct domain-dependent advantages. On Electricity Transformer Temperature datasets, AIGS achieves preventative early-warning lead times of 8.31 (ETTm1) and 9.88 (ETTm2) steps under gradual degradation. On Performance Measurement System traffic datasets, it shows significantly faster post-shift recovery after abrupt mutations. On the highly noisy Weather dataset, it improves anomaly recall while resisting stochastic noise overfitting. These findings establish AIGS as a practical, plug-and-play adapter for resource-constrained edge monitoring systems.

[LG-92] Context-Tower Conversion Preserves Generation While Freezing Retains Knowledge: Low-Budget AR-to-Diffusion Conversion of MoE LLM s

链接: https://arxiv.org/abs/2610.02657
作者: Wentao Lu,Jesse Clark,Tianyu Zhu
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Converting a pretrained autoregressive (AR) model to a diffusion language model (dLLM) enables parallel generation without pretraining a new model. Published conversion methods differ by roughly three orders of magnitude in training data and have not been compared under a common protocol. We compare two conversions of the same 30B Mixture-of-Experts (MoE) parent, holding the corpus, supervised-token budget, trainable parameter set and evaluation harness fixed, each under its own training recipe. The in-place model updates a subset of the parent’s weights using denoising and representation-alignment losses; the frozen-tower model instead conditions through cross-attention on a frozen causal copy of the parent. With 1B training tokens, the frozen-tower model scores 71.60 on HumanEval pass@10 against 6.19 for the in-place model, an 11.6x improvement. At the same budget it also keeps 95% of the parent’s GSM8K score and 99% of its MMLU-Pro score. A dense-parent experiment reproduces the HumanEval separation. Within the two-tower design at about 500M tokens, freezing the context tower retains substantially more MMLU-Pro performance than training it, while both give similar observed HumanEval scores. Our theoretical analysis establishes that both conversion classes contain an exact sampler for the AR parent under a hard attention mask and left-to-right commitment of one position per round. Under a shared loss, freezing removes the gradient contribution through the context states. Furthermore, evaluation protocol substantially affects a published 500B-token conversion’s scores in both directions across tasks, while its AR parent’s scores vary by less than three points, so comparing dLLMs needs a common protocol. These results show that, in the tested low-budget regime, the frozen-tower configuration retains substantially more of the parent’s generation performance than in-place conversion.

[LG-93] Online Verification of Language Model Responses Under Cost Constraints

链接: https://arxiv.org/abs/2610.02632
作者: Erfan Hajihashemi,Yanning Shen
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:As large language models are increasingly deployed for multi-step reasoning, verifying the correctness of their outputs has become essential for maintaining reliability at scale. Verifying the correctness of large language model outputs is often done by querying a costly ground-truth oracle, which is impractical to invoke at every step in an online setting. Prior work addresses this by querying a single weak verifier on every step, and using its score to decide whether the costly strong verifier needs to be queried as well, reserving strong verification for only a small fraction of the steps. However, a single fixed weak verifier may not perform consistently well as the subject matter or difficulty of incoming queries changes over time, and committing to one in advance risks either overly costly or inaccurate verification. We introduce OMVV (Online Multi-Verifier Verification), an algorithm that maintains a pool of K candidate weak verifiers with differing cost and verification performance, and adaptively routes each round’s decision to a verifier selected via an online score combiner and an exponential-weights routing policy. OMVV provides a distribution-free, finite-time guarantee on false-accept and false-reject rates across the full pool of verifiers, and further achieves sublinear regret against the best fixed verifier in hindsight under a combined cost and consistency objective. Experiments on reasoning dataset benchmarks show that OMVV achieves higher accuracy at lower verification cost than any single fixed verifier, across a range of operating budgets.

[LG-94] Quantifying the Value of Constructive Induction Knowledge and Noise Filtering on Inductive Learning ICML1991

链接: https://arxiv.org/abs/2610.02615
作者: Carl M. Kadie
类目: Machine Learning (cs.LG)
*备注: 12 pages, including a modern cover note and the unchanged 11-page author manuscript from 1991. Extended author version of a paper published in Machine Learning Proceedings 1991 (ICML 1991), pp. 153-157. Deposited in arXiv in 2026

点击查看摘要

Abstract:Learning research, as one of its central goals, tries to measure, model, and understand how learning-problem properties affect average-case learning performance. For example, we would like to quantify the value of constructive induction, noise filtering, and background knowledge. This paper describes the effective dimension, a new learning measure that helps link problem properties to learning performance. Like the Vapnik-Chervonenkis (VC) dimension, the effective dimension is often in a simple linear relation with problem properties. Unlike the VC dimension, the effective dimension can be estimated empirically and makes average-case predictions. It is therefore more widely applicable to machine and human learning research. The measure is demonstrated on several learning systems including Backpropagation. Finally, the measure is used to precisely predict the benefit of using FRINGE, a feature construction system. The benefit is found to decrease as the complexity of the target concept increases.

[LG-95] What Is Lost in Post-Training? Default Collapse and the Loss of In-Context Steerability Across Diverse Perspectives

链接: https://arxiv.org/abs/2610.02614
作者: Jessica Dierking,Itai Shapira,Niclas Boehmer
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:AI models serving a heterogeneous population must act on the principles appropriate to each user and context. While post-training has been shown to narrow the views large language models express, prior work has focused on default behavior rather than the ability to adapt to in-context information. We show that post-training also degrades a model’s ability to be steered in-context toward perspectives it was not trained to favor. In controlled experiments, we fine-tune models toward one side of cultural-value disagreements and evaluate checkpoints throughout training. The trained side becomes increasingly dominant in ordinary use, while the ability to recognize and faithfully enact the opposing view declines. These findings point to a tension between prioritizing a single set of values and preserving the technical capacity needed to serve diverse stakeholders. Finally, we propose and analyze an alternative objective that maximizes reward subject to a prescribed distribution over expressed perspectives, and present stance-distribution matching as a practical implementation.

[LG-96] Seer: Maximum Likelihood Regression for Learning-Speed Curves

链接: https://arxiv.org/abs/2610.02610
作者: Carl Myers Kadie
类目: Machine Learning (cs.LG)
*备注: 104 pages. Ph.D. dissertation, Department of Computer Science, University of Illinois at Urbana-Champaign, 1995. Original dissertation deposited in arXiv in 2026

点击查看摘要

Abstract:The research presented here focuses on modeling machine-learning performance. The thesis introduces Seer, a system that generates empirical observations of classification-learning performance and then uses those observations to create statistical models. The models can be used to predict the number of training examples needed to achieve a desired level and the maximum accuracy possible given an unlimited number of training examples. Seer advances the state of the art with 1) models that embody the best constraints for classification learning and most useful parameters, 2) algorithms that efficiently find maximum-likelihood models, and 3) a demonstration on real-world data from three domains of a practicable application of such modeling. The first part of the thesis gives an overview of the requirements for a good maximum-likelihood model of classification-learning performance. Next, reasonable design choices for such models are explored. Selection among such models is a task of nonlinear programming, but by exploiting appropriate problem constraints, the task is reduced to a nonlinear regression task that can be solved with an efficient iterative algorithm. The latter part of the thesis describes almost 100 experiments in the domains of soybean disease, heart disease, and audiological problems. The tests show that Seer is excellent at characterizing learning-performance and that it seems to be as good as possible at predicting learning performance. Finally, recommendations for choosing a regression model for a particular situation are made and directions for further research are identified. Comments: 104 pages. Ph.D. dissertation, Department of Computer Science, University of Illinois at Urbana-Champaign, 1995. Original dissertation deposited in arXiv in 2026 Subjects: Machine Learning (cs.LG) ACMclasses: I.2.6 Reportnumber: UILU-ENG 95 1727 Cite as: arXiv:2610.02610 [cs.LG] (or arXiv:2610.02610v1 [cs.LG] for this version) https://doi.org/10.48550/arXiv.2610.02610 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[LG-97] Activation Sparsity with Weight Approximation for Faster LLM Decoding on Offloaded Weights

链接: https://arxiv.org/abs/2610.02598
作者: JuneHyung Kim,Sankeerth Durvasula,Nandita Vijaykumar
类目: Machine Learning (cs.LG); Performance (cs.PF)
*备注:

点击查看摘要

Abstract:Deploying LLMs on consumer-grade GPUs with insufficient memory to hold their weights can result in prohibitively slow inference, because decoding repeatedly transfers offloaded weights from system RAM or flash storage into GPU at much lower bandwidth than local GPU-memory access. Activation sparsity reduces these transfers by skipping weights associated with zero or near-zero activations. However, as more activation contributions are omitted, model quality eventually degrades rapidly, indicating that weights associated with small-magnitude activations collectively influence model quality sharply. In this work, we improve the trade-off between model quality and decoding performance when exploiting activation sparsity. Our key idea is to replace the binary choice of whether or not to read a weight with three options: fully retain it, approximate it using a compressed weight representation, or omit it entirely. SpAx skips weights associated with activations closest to zero, reads approximate weights for smaller-magnitude activations, and reads original weights for the largest-magnitude activations. Smaller-magnitude activations attenuate the errors introduced by approximate weights, while compressed weight representations require fewer bytes to be transferred. With weights offloaded to CPU memory, SpAx speeds up decoding by 3.86X on average (up to 5.57X) with 16-bit weights and 2.06X (up to 2.74X) with 4-bit weights, at a WikiText-2 perplexity increase of at most 10%. With weights offloaded to flash storage, the speedups are 3.31X on average (up to 4.81X) and 1.54X (up to 2.03X).

[LG-98] Fisher-Guided Submodular Data Selection for Continual Pre-Training of Large Language Models

链接: https://arxiv.org/abs/2610.02593
作者: Zhenghao Zhao,Gaowen Liu,Zhiling Lan,Yan Yan
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Data selection is already a central bottleneck in large-language-model training, where web-scale corpora are noisy and token budgets are finite. In continual pre-training (CPT), it becomes a forgetting-control problem: a poorly chosen target-domain corpus can overwrite capabilities encoded in the pretrained checkpoint. Existing CPT practice either scores candidates with parameter-agnostic scalars such as perplexity, or mitigates forgetting by spending many extra general-domain replay tokens. Neither strategy directly asks how training on a candidate will move the model parameters. We show that loss-based selection causes the post-CPT Fisher diagonal to drift downward on exactly the high-Fisher coordinates the pretrained model had committed to, while leaving low-Fisher coordinates largely untouched. This asymmetry exposes a parameter-space mechanism for catastrophic forgetting. Motivated by this observation, we propose a Fisher-aware CPT selector that decomposes each candidate’s gradient into an anchor component, which measures perturbation along committed parameter directions, and a frontier component, which measures update capacity in unconstrained low-Fisher subspaces. We aggregate these signals with a log-determinant submodular objective and optimize it in a single pass using a scalable streaming data selection pipeline. On TinyLlama-1.1B and Llama-3.1-8B CPT over medical data, our selector improves target-domain quality while bounding forgetting on held-out pretraining benchmarks. Most importantly, it is substantially more token-efficient than forgetting-aware replay. 1B selected tokens already outperform the replay strategy trained with 10B tokens on both adaptation and forgetting, giving a 10x token-efficiency advantage.

[LG-99] Dense Mixture-of-Experts as a Reparameterized Wide FFN: A Granularity Sweep at Fixed Compute

链接: https://arxiv.org/abs/2610.02584
作者: Vu Quang Hoang,Nghia Hieu Nguyen
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Sparse Mixture-of-Experts (MoE) models combine learned routing with selection from a large expert pool. We isolate the contribution of dynamic expert combination using a dense analogue: K SwiGLU experts, all active for every token and combined by a softmax gate, at fixed total FFN width. With no larger pool or discrete selection, the dense baseline is the K=1 case. Validation loss varies non-monotonically with K : K=2 improves over the baseline by 0.0048 , whereas K=4 and K=6 worsen it by 0.0053 and 0.0197 , respectively. Routing generally remains soft and expert usage balanced, except in the first layer of the K=4 model, where routing is nearly one-hot. Forcing this gate to uniform increases loss by 2.4 nats on a diagnostic subset, indicating that its concentrated routing is functionally important. We further show that the architecture is exactly a dense SwiGLU with token-dependent, simplex-constrained group scaling and contains the dense baseline in its function class. These single-run results suggest a granularity sweet spot for always-active, softly gated FFNs at fixed width, with K=2 performing best among the configurations tested.

[LG-100] DISSOLVR: An Interpretable and Fast Framework for Aqueous and Organic Solubility Prediction ICML2026

链接: https://arxiv.org/abs/2610.02574
作者: Vansh Ramani,Har Ashish Arora,Dhairya Kuchhal,Sayan Ranu,Tarak Karmakar
类目: Machine Learning (cs.LG); Chemical Physics (physics.chem-ph)
*备注: 48 pages, 19 tables, 6 figures. Accepted to the 43rd International Conference on Machine Learning (ICML 2026)

点击查看摘要

Abstract:High-fidelity solubility prediction is fundamental to pharmaceutical development and environmental partitioning, where accurate modeling must couple molecular structure with thermodynamic behavior across diverse chemical environments. However, recent advancements have been dominated by deep learning architectures that often sacrifice physical interpretability for predictive power. We challenge this trend by showing that state-of-the-art performance does not require such non-transparent architectures. To address this, we introduce DISSOLVR, a transparent framework for molecular solubility prediction. In addition, we perform a comprehensive literature review and a benchmarking study against various methods. We show that DISSOLVR approaches the aleatoric limit of experimental uncertainty and achieves OOD generalization through structural invariance, derived by mapping molecules to physically-grounded descriptors. Then, we present an LLM-assisted post-hoc explanation pipeline that bridges the gap between symbolic model artifacts and chemically grounded narratives. Finally, a comparative benchmark of a survey involving 22 expert chemists reveals that expert evaluators provide deep insights.

[LG-101] Learning Closure of Dynamical Systems with Kernel Ridge Regression

链接: https://arxiv.org/abs/2610.02564
作者: Evan Habbershaw,John Harlim,Senwei Liang
类目: Numerical Analysis (math.NA); Machine Learning (cs.LG); Dynamical Systems (math.DS)
*备注: 31 pages, 8 figures

点击查看摘要

Abstract:We develop a closure modeling framework for identifying missing components of dynamical systems using Kernel Ridge Regression (KRR). The framework addresses two classes of closure problems: difference-equation closures arising in ODE and PDE settings, and algebraic closures arising from moment closure in kinetic equations. For the first class, we derive an error bound in an ODE setting that quantifies contributions from time integration, approximation of unresolved scales, and interpolation required to couple unresolved-scale effects to the resolved solver. Numerical experiments on the Lorenz-63 system and the Kuramoto-Sivashinsky equation demonstrate accurate long-horizon predictions and substantial improvements over an LSTM-based closure model. For the second class, we consider moment closure for a one-dimensional kinetic equation by modeling discrepancies between kinetic and macroscopic fluxes as a function of the resolved macroscopic variables. We compare global KRR models based on PCA coordinates with spatially local models. While the global model performs well for unimodal initial conditions, its accuracy deteriorates for bimodal initial conditions. Spatially local models with appropriate modeling inputs improve robustness and achieve higher predictive accuracy.

[LG-102] Neuron merging via inverse-activation regression for post-training compression of sigmoid neural networks

链接: https://arxiv.org/abs/2610.02559
作者: Ao Kuniya,Jun Ohkubo
类目: Machine Learning (cs.LG)
*备注: 9 pages, 7 figures

点击查看摘要

Abstract:As neural networks continue to grow in scale, model compression is becoming increasingly important for efficient inference under limited computational resources. Structured pruning methods remove neurons or channels that are estimated to be less important, but the removed units may still contain useful information. From the viewpoint of coarse-graining a trained network, it is valuable to ask which information should be retained when multiple neuronal degrees of freedom are consolidated. In this paper, we discuss cluster-based merging methods for compression of trained neural networks. In addition to a data-free contribution-weighted averaging method, we propose neuron-merging methods in which neuron responses are mapped back to the pre-activation space via the inverse activation function, and the weights and biases of each representative neuron are estimated using the least-squares method. We also examine both a data-assisted strategy with actual training inputs and a data-free strategy using randomly generated inputs. The comparisons provide empirical evidence, in the tested sigmoid networks, that weight information is particularly useful for clustering whereas activation information is useful for representative-neuron reconstruction in the merging process.

[LG-103] st-time Multi-agent Coordination by Decomposed Value Gradient Flow NEURIPS2026

链接: https://arxiv.org/abs/2610.02554
作者: Dongsu Lee,Haoran Xu,Amy Zhang
类目: Machine Learning (cs.LG); Robotics (cs.RO)
*备注: NeurIPS 2026

点击查看摘要

Abstract:Offline multi-agent reinforcement learning (MARL) faces a persistent trade-off. Expressive generative policies can represent multi-modal coordination in the data, but cannot distinguish high-value regions, while value-optimized policies exploit the learned Q-function but collapse the multi-modal into a single dominant mode. A single agent’s mode collapse can break joint coordination, and simultaneous drift across agents can push the joint policy into unseen regions of the action space. We propose scalable coordination via optimal unified transport (SCOUT), the first offline MARL framework to combine a generative foundation model with a learned value function through test-time action refinement. SCOUT trains two decoupled components: a flow-matching behavioral prior and a decomposed value function. At test-time, it transports behavioral samples toward high-value regions via Stein variational gradient descent. The number of transport steps controls adaptive test-time scaling, replacing a fixed regularization coefficient. Under the individual-global-max (IGM) principle, we prove a single-term KL bound on the joint soft-value gap that vanishes as transport converges, with an irreducible additive residual proportional to the IGM violation. Empirically, SCOUT achieves the best average performance across discrete and continuous offline MARL benchmarks and yields performance improvements in all offline-to-online configurations.

[LG-104] Reward Inflation: A Healthy Stimulus for Reinforcement Learning NEURIPS2026

链接: https://arxiv.org/abs/2610.02545
作者: Ganghun Lee,Minji Kim,Minsu Lee,Byoung-Tak Zhang
类目: Machine Learning (cs.LG)
*备注: Accepted at NeurIPS 2026

点击查看摘要

Abstract:Reward serves as the primary learning signal in reinforcement learning (RL). However, while reward magnitudes are typically held fixed throughout training, their temporal modulation remains underexplored. In this paper, we propose reward inflation, a gradual scaling of rewards over the course of training, and show that it can act as a healthy stimulus for RL. Theoretically, reward inflation induces an implicit recency weighting that upweights recent transitions during policy updates, enabling faster adaptation. We further show that, by sustaining gradient signals as the policy saturates, reward inflation suppresses the emergence of dormant neurons and helps preserve plasticity. Empirical results on ALE games and MuJoCo tasks corroborate these findings, showing that an appropriate level of reward inflation benefits a broad range of tasks. Finally, we introduce Fed, an adaptive variant that adjusts the inflation level on the fly, and find that it often improves upon fixed inflation.

[LG-105] Autoregressive Differentiable Method for Integer Programming

链接: https://arxiv.org/abs/2610.02528
作者: Ouns El Harzli,Yudong Cao
类目: Machine Learning (cs.LG); Optimization and Control (math.OC)
*备注:

点击查看摘要

Abstract:We introduce an autoregressive differentiable method to solve 0-1 integer programs. We fix an arbitrary order of the binary variables and we train a transformer to predict the next bit while remaining in the feasible set. Our method is first trained on feasible incumbents provided by any solver, thus allowing us to initialize the transformer in the feasible set. Our procedure then implements a Lagrangian penalty to penalize infeasible solutions, and the transformer is further trained to explore the feasible set using Gumbel-softmax activations on the relaxed objective. We have tested our method on non-convex instances of quadratic knapsack problem and demonstrated consistent improvement upon state-of-the-art open-source solvers for dense problems up to 10,000 binary variables. In particular, we empirically demonstrate a phenomenon akin to a tunneling effect where the effective change of variables from binary variable to the continuous weights of the transformer that the method implements enables crossing barriers in the relaxed objective landscape.

[LG-106] BaCP: Backbone Contrastive Pruning for Preserving Representations in Extremely Sparse Neural Networks

链接: https://arxiv.org/abs/2610.02524
作者: Mohammad Haroon Khawaja,Muhammad Haseeb,Mohammad Fatim Shoaib,Muhammad Tahir
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Unstructured pruning at extreme sparsity often suffers from representational collapse, causing sharp drops in accuracy. To address this, we study Backbone Contrastive Pruning (BaCP), which regularizes the sparse network’s embedding space by aligning it with pretrained, fine-tuned, and historical snapshot models. Building on the contrastive decomposition of the CAP framework (Xu et al., 2022), we provide a rigorous matched-budget characterization of this approach across multiple pruning criteria. Evaluated across 90 settings, BaCP improves accuracy substantially in extreme sparsity regimes where standard pruning fails, and is close to baseline where representations remain intact.

[LG-107] Learning the Latent Structure: A Feature-Centric Approach to Graph Data Augmentation AAAI2026

链接: https://arxiv.org/abs/2610.02517
作者: Yu Song,Zhigang Hua,Yan Xie,Bingheng Li,Jingzhe Liu,Bo Long,Jiliang Tang,Hui Liu
类目: Machine Learning (cs.LG)
*备注: Accepted at AAAI 2026

点击查看摘要

Abstract:Graph-structured data plays a pivotal role in modeling complex relationships. However, real-world graphs are often incomplete due to data collection and observational constraints, severely limiting the effectiveness of modern graph learning pipelines. While existing Graph Data Augmentation (GDA) methods attempt to refine graph structures for improved downstream performance, they are typically label-dependent, computationally expensive, and inherently transductive, limiting their applicability in practical scenarios. In this work, we present a novel feature-centric graph data augmentation framework that bypasses explicit structure modeling by operating directly in the embedding space. Through a self-supervised inverse masking process, our method captures latent ties between observed and complete graphs, enabling recovery of unobserved structural signals through refined node representations. To enhance robustness under noisy and sparse supervision, we introduce a message regularizer and a bootstrap strategy for effective training and generalization. Evaluated on ten graph datasets spanning multiple domains, our approach, SelfAug, consistently outperforms state-of-the-art methods in both accuracy and efficiency across inductive and cold-start settings, highlighting its potential as a scalable and generalizable solution for real-world graph learning scenarios.

[LG-108] Post-Training Quantization of Autoregressive Weather Models

链接: https://arxiv.org/abs/2610.02511
作者: Ananyo Bhattacharya,Swastik Bhattacharya,Christiane Jablonowski
类目: Machine Learning (cs.LG); Earth and Planetary Astrophysics (astro-ph.EP)
*备注:

点击查看摘要

Abstract:Advancements in high-resolution numerical weather prediction (NWP) and data assimilation (DA) have shaped the developments in deep learning (DL) architectures emulating atmospheric dynamics. Emulators for weather forecasting exhibit forecast quality comparable to physics based models at forecast horizon scaling from few days to subseasonal time scales. The emulators are driven by hardware-accelerated matrix multiplication in autoregressive inferences, significantly reducing the computation time and resources required for NWP. Optimization of the matrix multiplication processes in GPU architectures provides opportunities to scale towards high-resolution domain, and offers implementation of out of the box solutions. Post-training quantization (PTQ) has been demonstrated across multiple DL architectures to accelerate and increase the number of computations in unit time while consuming less power, enabling applications on edge hardware. In this study, we investigate the effect of PTQ on pre-trained AI emulators for global-scale weather forecasting. We implement PTQ algorithms in Deep Learning Weather Prediction (DLWP) and FourCastNet (FCN) models as a proof of concept for geophysical fluid dynamics applications. We systematically investigate the effect of PTQ on emulator inferences over short-range forecast horizons. Evaluation of PTQ configurations using simulated quantization hints at qualitatively meaningful forecasts over short-time horizons. These results provide a first benchmark of PTQ for autoregressive weather emulators and a basis for quantization-based optimization of DL models for dynamical systems.

[LG-109] LiteEMG-FM: An Efficient and Deployable Foundation Model for Robust EMG Sensing

链接: https://arxiv.org/abs/2610.02497
作者: Tianhao Wu,Xu Wu,Amirmohammad Radmehr,Jiawei Yu,Yi Wu,Phuc Nguyen,Jian Liu
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Electromyography (EMG) signals vary substantially across individuals, body regions, recording sessions, and sensing hardware, limiting the generalization of models for assistive devices and human-computer interaction. Existing time-series foundation models are also computationally expensive for real-time wearable deployment and often fail to capture EMG-specific time-frequency characteristics. We present LiteEMG-FM, an efficient hybrid CNN-Transformer foundation model for practical EMG sensing. Pretrained on 16 diverse upper- and lower-limb EMG datasets, LiteEMG-FM learns representations that generalize across users and datasets. For resource-constrained deployment, we implement a hierarchical wake-up architecture in which a lightweight, always-on 1D-CNN filters rest and non-target activity and activates LiteEMG-FM only for valid gestures. We evaluate full inference offloading, split inference, and full on-device processing, characterizing their trade-offs in latency, power consumption, and memory footprint. Across diverse evaluation settings, LiteEMG-FM outperforms state-of-the-art time-series foundation models and supervised baselines, particularly under zero-calibration cross-participant and data-scarce conditions. These results demonstrate that LiteEMG-FM is an effective, efficient, and deployable foundation model for EMG applications.

[LG-110] Harnessing LLM s as Agents : What Does It Cost?

链接: https://arxiv.org/abs/2610.02488
作者: Zelin Zhao(Georgia Institute of Technology),Xinyu Guo(Georgia Institute of Technology),Jingyuan Zhang(Georgia Institute of Technology),Yuxuan Zhang(Etude AI),Yongxin Chen(Georgia Institute of Technology)
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Language-model agents increasingly rely on harnesses that manage bounded context, persistent memory, tools, verification, and repeated execution, yet existing notions of model capability do not quantify the computational resources these mechanisms consume. We introduce the Language Model Agent Machine (LAM), a resource-bounded abstraction that fixes the underlying semantic model while explicitly charging harness-level resources. We establish four classes of results. Communication: LAM execution is instancewise equivalent to red–blue pebbling under simultaneous call–transfer budgets, transferring classical I/O lower bounds to context–memory traffic. Access: memory interfaces induce asymptotic separations, including a \Theta(n) gap between random and non-speculative sequential access on pointer chasing. Recomputation: bit-reversal DAGs require \Theta(n^2/(C+S)+n) model calls with context capacity C and persistent-memory capacity S , quantifying when stored intermediate state avoids repeated semantic computation. Reliability: we derive tight stage-local sampling bounds, exact imperfect-verification costs, and a Young–Daly-type checkpoint law with a closed-form optimal verification interval. Controlled and held-out experiments on GPT-6 Astra test communication and reliability predictions, including checkpoint optima, policy selection under programmatic checking, and tradeoffs among call granularity, logical input traffic, and reliability on chained MATH tasks. Together, these results provide a resource theory for the computational cost of language-model agent harnesses.

[LG-111] hreshold-Aware Conformal Routing

链接: https://arxiv.org/abs/2610.02487
作者: Shiwei Tan,Huzefa Rangwala,Danielle C. Maddix
类目: Machine Learning (cs.LG); Computational Engineering, Finance, and Science (cs.CE)
*备注:

点击查看摘要

Abstract:High-fidelity simulations are essential to scientific and engineering design, but can be expensive to run repeatedly. Learned surrogates offer a faster alternative, yet their higher errors may alter downstream decisions. This accuracy-speed tradeoff creates a need to determine whether a surrogate can be used or the full simulator remains necessary. We study decisions determined by whether a scalar quantity of interest lies above or below a fixed threshold. For each input, we use the surrogate when its conformal interval lies entirely on one side of the threshold and route the input to simulation when the interval intersects it. Standard conformal prediction constructs intervals without reference to the downstream decision threshold: even a narrow interval near the threshold can cross it and trigger simulation, whereas a wider interval farther away can remain entirely on one side and require no simulation. We introduce Threshold-Aware Conformal Routing (TACR), which learns an input-dependent scale using a threshold-aware objective that concentrates interval tightness near the decision boundary. Exact split-conformal calibration on held-out data preserves distribution-free marginal coverage, which also upper-bounds the probability of an incorrect threshold decision that is not routed. Across various scientific and engineering datasets, TACR reduces simulator deferrals by 14-75% relative to standard conformal prediction at the same coverage target. Against a variant without threshold-local weighting but with similar predictor accuracy, TACR further reduces deferrals by 10-24% on four datasets. These results show that optimizing interval allocation for routing can reduce simulator calls without weakening the standard conformal guarantee.

[LG-112] SD-DPC: Sparse Dictionary Differentiable Predictive Control

链接: https://arxiv.org/abs/2610.02466
作者: Ali Reza Daneshvar Garmroodi,Jan Drgoňa
类目: ystems and Control (eess.SY); Machine Learning (cs.LG); Optimization and Control (math.OC)
*备注:

点击查看摘要

Abstract:We present sparse dictionary differentiable predictive control (SD-DPC), a framework for learning sparse, interpretable feedback policies for nonlinear systems from data. A prediction model is first identified by rollout-based sparse identification of nonlinear dynamics (SINDy), building on gradient-based and multistep formulations. The policy is then parameterized as a sparse combination of dictionary functions and trained by differentiating a constrained finite-horizon predictive-control objective through this model, so that its terms are selected by closed-loop performance rather than by imitating a previously trained controller. The result is an explicit feedback law with only a handful of terms. Across three benchmark control problems, SD-DPC satisfies the constraints in all test scenarios, outperforms a policy distilled onto the same terms by up to an order of magnitude, and requires orders of magnitude less memory and online computation than an optimization benchmark, while admitting explicit sensitivity bounds.

[LG-113] A Composable AI-Accelerated Iterative Solver for 3D-IC Thermal Modeling

链接: https://arxiv.org/abs/2610.02461
作者: Yixing Li,Jiahang Zhou,Zhiyu Zeng,Xin Ai
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Accurate thermal analysis of heterogeneous 2.5D/3D-IC packages is essential yet computationally prohibitive. A single full-package FEM simulation can take hours, while AI-based surrogates treat the entire stack as a monolithic prediction target and must be retrained whenever the die count or topology changes. To address this limitation, this work proposes Domain-Decomposed AI-Accelerated Iterative Solver for Thermal Analysis (DAIST), a composable thermal solver that decomposes the global package simulation into block-level subdomain problems, replaces subdomain solvers with neural operators, and couples them through iterative exchanges of interfacial temperature and heat flux. This local-to-global architecture eliminates the topology lock-in of monolithic models: block-level neural operators can be directly reused in unseen package assemblies without retraining. The iterative coupling strategy further provides a controllable accuracy-runtime tradeoff, where the iteration budget can be adjusted to trade accuracy for runtime. Evaluated on a multi-chiplet system and an advanced packaging system, DAIST achieves up to 178\times speedup over traditional FEM solvers with mean temperature errors of 0.068% and 0.323%, respectively, while demonstrating cross-topology reuse of block-level models across structurally distinct package assemblies.

[LG-114] AI-driven Thermal-aware Data Center Capacity Planning

链接: https://arxiv.org/abs/2610.02442
作者: Yixing Li,Mark Fenton,Matthew Kaufeler,Ka Ming Leung,Xin Ai,Zhiyu Zeng
类目: Machine Learning (cs.LG)
*备注: Presented at DesignCon 2026

点击查看摘要

Abstract:The emerging of large language models (LLMs) has posed significant challenges to the thermal management of data center. Intense GPU computation for LLMs results in localized hotspots. Moreover, spiking thermal loads during training and inference bursts make real-time cooling response more difficult to predict and control. Thermal-aware capacity planning of data center requires massive expensive high-fidelity CFD simulations. AI models can perform real-time prediction for unseen designs. However, existing works either have large prediction error, or have over-simplified assumptions for data center operations. This work presents an AI-driven framework that can perform thermal-aware capacity planning for a real-world data center in seconds. The embedded AI model learns from numerous key parameters (rack power, server power, server placement, HVAC settings etc.), and provides temperature prediction within milliseconds. This AI model is tested against high-fidelity CFD simulations, and results show that for unseen data center designs, model can achieve high accuracy with 10000X speedup. Driven by the AI model, the authors design the thermal-aware capacity planning framework. This framework can help data center designers and operators instantaneously optimize both workload distribution and HVAC cooling efficiency.

[LG-115] Bandits via Additive Quantized Representations NEURIPS2026

链接: https://arxiv.org/abs/2610.02440
作者: Ami Tavory,Noam Touitou,Tal Sarig,Frank Cheng,Ido Guy
类目: Machine Learning (cs.LG)
*备注: 40 pages, 16 figures, 12 tables. Accepted at NeurIPS 2026

点击查看摘要

Abstract:Contextual bandits require balancing nonlinear reward modeling with online efficiency. Tree ensembles and neural methods capture nonlinearities but require periodic retraining and large replay buffers. Linear models update efficiently per observation with O(1) memory, but are fundamentally restricted to linear reward structures. We propose Residual Quantization (RQ) as a representation layer to bridge this gap. An offline-trained RQ codebook maps continuous contexts into discrete centroid assignments across multiple levels, set dynamically through a shadow mechanism. This enables a spectrum of additive bandit algorithms that achieve nonlinear expressivity with strictly bounded memory. Across 13 datasets, RQ variants beat their non-RQ counterparts on 11 of 13 datasets, often by wide margins, while matching doubling-retrain XGBoost and neural baselines using up to 1000 times less memory.

[LG-116] A Generative Model of Complex Networks Using Graphons and Neural Inverse Operators

链接: https://arxiv.org/abs/2610.02439
作者: Wooseong Choi,Italo’Ivo Lima Dias Pinto,Chen Sun,Gaurav Gupta,Dong Song,Paul Bogdan
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Generative graph models are central to understanding and simulating complex networks. However, existing approaches have complementary strengths and limitations. Mechanistic models offer interpretability but rely on instance-specific estimation methods. Deep generative models, on the other hand, offer amortized inference at the cost of interpretability and are largely limited to graph sizes seen during training. Scientific applications motivate a framework that retains the strengths of both paradigms. We bridge them by formulating both the generative model and parameter recovery in function space. A multifractal step graphon extends standard step graphons with a recursive construction that compactly parameterizes complex networks. This formulation admits a neural inverse operator to recover its parameters, enabling inference on unseen graph sizes. We evaluate our model, trained only on synthetic multifractal step graphon realizations, against both paradigms. Against a graph foundation model pretrained on empirical networks, our method achieves the best average performance on three of four metrics in a zero-shot graph-generation benchmark, indicating that the model transfers to real-world graphs. We also apply our method to single-observation networks, a regime largely inaccessible to deep models that require training corpora, where it performs comparably to an instance-specific method that optimizes on each graph. In a multi-subject EEG case study, the inferred parameters track a reversible change in brain state more sensitively than traditional network statistics. Together, these results indicate that mechanistic interpretability and amortized inference can be effectively unified in a generative graph model to enhance our understanding of complex networks.

[LG-117] CRISP: A Framework for Clause-Reconstructed Interpretable NeuroSymbolic Propositions

链接: https://arxiv.org/abs/2610.02431
作者: Alex Chan,Shafi Muhtasim Chowdhury,Ekin Can Erkuş,Ole-Christoffer Granmo,Alex Yakovlev,Rishad Shafik
类目: Machine Learning (cs.LG)
*备注: Preprint for an accepted paper at International Symposium of the Tsetlin Machine (ISTM) 2026 Conference

点击查看摘要

Abstract:Deep neural networks achieve high accuracy through layered numerical transformations, yet their decisions remain difficult to audit because decision evidence is encoded in hidden activations rather than explicit rules. This paper introduces CRISP, a framework that reconstructs the last-layer activation vector (LLAV) of binary neural teachers as Tsetlin Machine ™ clauses. CRISP sign-binarizes the teacher’s penultimate pre-logit activations, and assigns one Individual TM (ITM) to each LLAV neuron. Each reconstructed hidden bit is represented by propositional clauses over Booleanized input features, which gives a direct symbolic trace from named input thresholds to a named teacher neuron. CRISP is evaluated on MNIST, KMNIST, FashionMNIST (FMNIST), SVHN, and CIFAR10 using a BinaryConnect convolutional neural network (BCCNN) teacher and a fully binary neural network (BNN) teacher, with an additional study on binary thresholding, thermometer encoding, and quartile binning at multiple bit depths. The results show that LLAV sign-binarization does not reduce teacher-head accuracy in the tested BNN setting, while ITM reconstruction error is the main limiting factor. Quartile one-bit Booleanization gives the strongest reconstruction fidelity on SVHN at 87.52% test fidelity and is competitive on CIFAR10, and the reconstructed LLAV preserves 78.41% teacher-head accuracy on FMNIST. Pooled clause-evidence visualizations show that the learned ITM literals concentrate on the object region in centered benchmarks. CRISP therefore provides a clause-level route for inspecting the final hidden representation of binary neural teachers.

[LG-118] he AI Theorist reveals excitonic structure in α-RuCl_3

链接: https://arxiv.org/abs/2610.02417
作者: Hongjian Zhou,Xianfan Nie,Sean Wu,Tarun Patel,Jinge Wu,Andrew Liu,Adam Wei Tsen,David A. Clifton
类目: Machine Learning (cs.LG); Materials Science (cond-mat.mtrl-sci)
*备注:

点击查看摘要

Abstract:Advances in experimental instrumentation and automation generate increasingly rich datasets, but turning experimental observations into microscopic understanding remains a bottleneck in scientific discovery. To accelerate this process, we introduce AI Theorist, a system of artificial intelligence (AI) agents for autonomous discovery of physical models through hypothesis generation, first-principles calculations and evidence-driven refinement. We apply the framework to \alpha -RuCl _3 , a leading candidate material for realizing a Kitaev quantum spin liquid, to investigate its electronic structure through optical spectra. AI Theorist develops a new interpretation of the optical and photocurrent observations, identifying distinct excitonic states with contrasting optical selection rules and real-space distributions. To our knowledge, this is the first demonstration of an AI system autonomously developing a physical model to explain previously unpublished experimental observations in a quantum material, utilizing first-principles electronic-structure and many-body calculations. Our results establish a route to autonomous theoretical discovery in materials science, in which AI agents use first-principles calculations to turn experimental observations into physical models and testable predictions.

[LG-119] Prompted to Discriminate: Generalizing Malicious-Input Probes in the Wild NEURIPS2026

链接: https://arxiv.org/abs/2610.02413
作者: Elad David,Max Fomin
类目: Machine Learning (cs.LG)
*备注: 40th Conference on Neural Information Processing Systems (NeurIPS 2026). Workshop: Agents in the Wild: Safety, Security, and Beyond

点击查看摘要

Abstract:LLM agents increasingly rely on activation probes as runtime monitors for prompt injection, jailbreaks, and unsafe requests, reading the model’s own hidden state to catch a harmful input before the agent acts on it. A cheap, increasingly common move, borrowed from LLM-as-judge prompting, is to append a short classification instruction after the user’s turn and read the probe at that point, to sharpen it: the instruction asks the model to represent the incoming request as a class, concentrating the signal the probe must separate, at negligible serving cost. But does the wording of that suffix matter, and does its benefit hold in the wild, on attack types the probe never saw in training, the regime a deployed monitor faces? We test this with a controlled ladder of post-user suffixes under strict leave-one-dataset-out (LODO) evaluation across 13 safety benchmarks (jailbreak, injection, and benign chat) and three open-weight model families (Llama-3.1-8B, Qwen3.5-9B, Gemma-4-12B). On a single-position probe, a classification suffix consistently improves out-of-distribution detection over no suffix (up to ~4 AUC points); yet which suffix matters: prompting the model to classify the input, even into content-free labels, reliably wins; an off-topic or merely-attentive suffix helps little. The gain comes from the classification format, not the named criterion: a content-free suffix matches the real malicious/benign one, with the criterion adding precision only at strict thresholds. This is not an artifact of the single-position read: the benefit carries to the multi-position pooling probes used in production (attention, multi-max, MLP), though the best-performing suffix there is readout-dependent. Served through a KV-cache fork, it is a cheap drop-in for any activation-probe monitor, though not an automatic win: which suffix helps, and by how much, depends on the model and the readout.

[LG-120] Energy Saving in 5G and Beyond Networks: A Quantum Reinforcement Learning Approach

链接: https://arxiv.org/abs/2610.02403
作者: Muhammad Usman,Nguyen Van Huynh,Marianna Lezzi,Mariangela Lazoi
类目: Networking and Internet Architecture (cs.NI); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Energy saving has become a critical challenge in 5G and beyond networks. The rapid growth of connected devices has increased the overall network energy demand, driving operational expenditure to unsustainable heights. The Base Station (BS) accounts for the largest share of energy usage, typically consuming around 60-70% of the Radio Access Network (RAN)'s total energy. Therefore, to address this issue, this article optimizes the BS’s energy usage while accounting for the dynamic behavior of User Equipment (UE). Deep Reinforcement Learning (DRL) is a natural candidate for determining effective energy saving policies, such as automatically switching BSs on or off when user density is low or adjusting transmission power to balance energy efficiency and Quality of Service (QoS). However, its heavy training burden and the exponential growth of state and action spaces in dense 5G environments make exploration increasingly difficult. To overcome these limitations, we introduce a novel Quantum Reinforcement Learning (QRL) algorithm that leverages quantum principles, including superposition and entanglement, through parameterized quantum circuits, enabling significantly faster convergence than DRL, which relies on conventional deep neural networks. Extensive simulations demonstrate that the proposed QRL can substantially reduce energy consumption while maintaining QoS, even when UEs are highly dynamic and frequently switch their association with BS antennas. Additionally, QRL consistently outperforms DRL and Q-Learning in both convergence speed and learning complexity.

[LG-121] VisAudit: Evaluating Multimodal Agents for Visual Diagnosis and Repair

链接: https://arxiv.org/abs/2610.02399
作者: Shicheng Liu,Adam Kahirov,Qi Zhang,Zhimin Hu,Song Wang,Junhong Lin,Julian Shun,Yada Zhu
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Multimodal agents are increasingly used for data visualization tasks but remain limited in autonomous review. Unlike humans, they may fail to recognize when a visualization is incorrect, determine what to change, repair it without disrupting correct content, and verify whether the intervention succeeded. Existing benchmarks largely evaluate predefined individual capabilities such as chart generation, instruction-guided editing, or defect detection, and therefore do not capture this gap in autonomous review. We introduce VisAudit, a benchmark for evaluating visualization diagnosis, repair, and verification. Given a rendered chart and configurable auxiliary evidence, including its source data table, intended text summary, and visualization code, an agent iteratively diagnoses potential defects, modifies and executes visualization code, inspects execution and visual feedback, and determines when no further intervention is needed. VisAudit defines three tracks spanning diagnosed repair, autonomous repair, and open-world verification, and contains 1,900 flawed instances across 21 chart types and 10 flaw categories, together with 300 initially correct charts. We construct the benchmark through controlled perturbations of validated source visualizations, with systematic verification and human-aligned quality control to ensure that injected defects are well-defined and recoverable from the available evidence. Experiments with leading multimodal models reveal a substantial gap from reliable autonomous review: the strongest evaluated model fully recovers only 47.4% of flawed charts in the autonomous-repair setting.

[LG-122] Validated Data Onboarding for AI Demand Forecasting on U.S. Building Meter Data: Design Controlled Evaluation and a Corrected Negative Result

链接: https://arxiv.org/abs/2610.02397
作者: Yixuan Liang
类目: Machine Learning (cs.LG)
*备注: Technical report; 8 pages, 5 figures, 2 tables. Code, data manifests, and reproducibility artifacts: this https URL . Preprint; not peer reviewed

点击查看摘要

Abstract:Electric utilities and grid operators increasingly rely on machine-learning models to forecast next-day demand, and those models learn from meter data that is routinely defective: readings go missing, sensors freeze, buildings read zero for hours, and units change by a factor of 100. This report presents a data-onboarding pipeline that detects and repairs such defects before a model is trained, using only information available at forecast time, and a controlled experiment that measures whether the pipeline protects a 24-hour-ahead forecast. On hourly electricity data for twelve U.S. buildings from the public Building Data Genome 2 dataset (210,528 rows, 2016-2017), seeded, hash-logged defects touching 0.10% of the training period raised the error of a gradient-boosting forecaster by 86%; after detection and past-only repair the error returned to the clean-data level (mean absolute scaled error 0.760 clean, 1.415 corrupted, 0.729 repaired) while 93% of training targets were retained. At a defect prevalence calibrated to published field studies (1.6% of training rows) the unprotected forecaster’s error reached 4.4 times that of a seasonal-naive rule, and the repaired forecaster again matched the clean baseline. The same pattern held for ridge regression and a random forest and across horizons of 1 to 24 hours. A first version of the pipeline over-cleaned natural data and made forecasts 25% worse; that result is retained, its cause is traced in the published artifacts, and the per-building calibration that corrects it is documented as a dated amendment. Every number is reproducible from pinned public inputs with SHA-256 verification, 84 automated tests and continuous integration.

[LG-123] Latent-MOPD: Latent Multi-Teacher On-Policy Distillation

链接: https://arxiv.org/abs/2610.02381
作者: Zhengyu Fang,Seoyeon Hong,Jie Yang,Muyang Li,Koyoshi Shindo,Brandon Joseph Lwowski,Jing Li
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:On-policy distillation (OPD) trains a student on the responses it generates. Existing LLM multi-teacher OPD transfers what specialists predict through their output distributions. We introduce Latent-MOPD, to our knowledge the first representation-level multi-teacher OPD method for LLMs. It integrates existing specialists through both their predictions and the hidden states used to compute them, without additional teacher training. To coordinate representation supervision from multiple specialists, we select late-layer targets according to the teacher-student relationship, bridge unequal hidden widths with a shared projection, and group updates by domain. Each teacher’s supervision gradually shifts from hidden states to token predictions, with both channels using the same routed specialist. In our main same-family setting, Latent-MOPD outperforms the token-only, representation-only and uniform-averaging baselines on all nine benchmarks across math, code and logic. With the same parameter count as each teacher, the student also surpasses the per-benchmark best teacher on a majority of these benchmarks. With larger, separately developed cross-family teachers, Latent-MOPD outperforms both single-channel baselines on all benchmarks. A same-family all-layer representation-only control remains stable with domain-pure updates but collapses when teacher domains are interleaved within an update. Our results show that a single student can integrate capabilities from several specialists through both their output distributions and internal representations.

[LG-124] Flow Matching for Fast Posterior Sampling in Bayesian Inverse Problems

链接: https://arxiv.org/abs/2610.02377
作者: Jan Blechschmidt,Oliver G. Ernst,Moritz Poguntke,Björn Sprungk
类目: Numerical Analysis (math.NA); Machine Learning (cs.LG); Computation (stat.CO)
*备注: 38 pages, 18 figures

点击查看摘要

Abstract:Sampling from the posterior is the central task of computational Bayesian inverse problems. The standard workhorse in Bayesian inference - Markov chain Monte Carlo (MCMC) - is sequential, yields correlated samples, and must be rerun for each observation. Conditional flow matching offers an amortized alternative: a transport map, trained once on joint samples of parameter and data, that yields independent approximate posterior samples for any observation at negligible online cost, without new likelihood evaluations. We give a careful, MCMC-literate assessment of flow matching for PDE-based inverse problems with function-valued parameters. Exploiting the flow’s tractable density, we derive computable accuracy estimates of the underlying approximate posterior in total-variation distance and Kullback-Leibler divergence and, moreover, propose a hybrid sampler that is asymptotically exact by Metropolization. We validate the accuracy estimates and demonstrate the amortization in several numerical examples, including electrical impedance tomography and a likelihood-free Lotka-Volterra model.

[LG-125] Co-design Gym: A Unified Benchmark for Embodiment-Policy Co-optimization

链接: https://arxiv.org/abs/2610.02366
作者: Aviraj Newatia,Yordan Tsvetkov,Leonard Pleiss,Andrew Spielberg,Rika Antonova
类目: Machine Learning (cs.LG); Robotics (cs.RO)
*备注: Aviraj Newatia and Yordan Tsvetkov contributed equally

点击查看摘要

Abstract:Finding an optimal behaviour policy within a given environment is a widely studied problem in domains as diverse as games, robotics, energy infrastructure, communication networks, and multi-agent systems. Numerous benchmarks have been developed to support such research, but the vast majority assume that the agent’s embodiment (design) is fixed, focusing instead on policy learning alone. Lifting this assumption gives rise to a broader class of problems in which optimizing embodiment and policy separately is highly suboptimal. An agent’s embodiment strongly shapes which control policies can be discovered, while the optimal embodiment is in turn defined by the policies it admits. To help the research community study this class of problems explicitly and systematically, we introduce Co-Design Gym - a suite of benchmark environments for jointly optimizing embodiment and policy. Our environments span domains such as robotic manipulation and locomotion, multi-robot cooperation, deformable and soft dynamics, video games, electricity grids, wireless networks, F1 racing, multi-agent warehouses, and optimal control, offering 20 environment families (domains), with over 85 distinct co-design presets in total. We further contribute a systematic evaluation of representative co-design algorithms, characterizing the current state of the art. Together, these contributions lay the groundwork for cumulative, comparable progress in co-design.

[LG-126] ArrivalBench: Agent -Generated Data Pipelines Are Correct Once and Wrong Under Time NEURIPS2026

链接: https://arxiv.org/abs/2610.02363
作者: Pranay Kothari
类目: Machine Learning (cs.LG); Databases (cs.DB)
*备注: Accepted as a poster at the NeurIPS 2026 Workshop “Who Verifies the Agents?”. 20 pages

点击查看摘要

Abstract:Benchmarks for agent-generated data work grade a pipeline by running it once against a fixed snapshot. ArrivalBench instead re-executes the pipeline an agent leaves behind under adversarial but replayable delivery schedules (late, duplicated, out-of-order and retried records) and requires its final state to equal a batch recomputation of the complete log. Because the oracle recomputes rather than classifies, a wrong table and a crash are distinct verdicts: a crash is visible to monitoring a team already runs, and a wrong table is not. On 40 tasks we built, our reimplementation of single-execution grading certifies 86-100% of the pipelines eleven models produce; re-executing the same artifacts finds 7.0-79.2% of the certified ones silently wrong. The gap is not produced by the repair loop: within the same model and task, pipelines repaired against the snapshot test fail replay about as often as those that passed it first time. In every model, idempotency hazards fail more often than ordering hazards. Separating a wrong answer from a crash also changes how interventions read: a hazard warning cuts one model’s silent failure from 48.2% to 10.5% while raising its crash rate from 9.0% to 37.0%, so all-in failure moves only from 51.0% to 44.0%. All eleven arms were independently re-run, and rates moved by at most 5.9 points.

[LG-127] Why Does Adaptive Batching Help LLM Pretraining? A Perspective from Unbounded Variance

链接: https://arxiv.org/abs/2610.02355
作者: Arda Fazla,Antesh Upadhyay,Ege C. Kaya,M. Berk Sahin,Abolfazl Hashemi
类目: Machine Learning (cs.LG); Optimization and Control (math.OC)
*备注:

点击查看摘要

Abstract:Increasing the batch size during training is a common practice in large language model (LLM) pretraining, yet the theoretical justification behind its success is not well understood. Analyses of stochastic optimization often assume uniformly bounded stochastic gradient variance, yet recent evidence suggests that this assumption fails in many practical nonconvex problems. The Blum–Gladyshev (BG- 0 ) noise model relaxes this assumption by allowing the variance to grow quadratically with the distance from initialization, suggesting that batch size schedulers can help by controlling the variance growth during training. However, this growth can be overly conservative in practice. We empirically investigate variance growth in LLM pretraining and observe that a generalized BG model with a tunable growth exponent provides a tighter description of practical noise behavior. Motivated by this observation, we introduce the generalized BG- a noise model, which interpolates between bounded variance ( a=0 ) and BG- 0 noise ( a=2 ). Under L -smoothness, we derive an information-theoretic lower bound with growth-dependent oracle complexity \Omega(\epsilon^-(4+a)) and establish a matching upper bound in \epsilon -dependence by increasing the batch size as the iterates move away from initialization. Finally, we propose an adaptive batch scheduler that controls variance growth through dynamic batch size adjustments during training. In pretraining OLMo2 models of up to 1B parameters on C4, our scheduler achieves a lower validation loss than both small and large batch training under matched token budgets, while using less than 10% of the iterations of small batch training.

[LG-128] From Behavior to Provenance: Attributing Tabular Foundation Models to Synthetic Pretraining Data

链接: https://arxiv.org/abs/2610.02347
作者: Mohamed Bouadi,Nassim Bouarour,Shivam Dubey,Aditya Tanna,Vinay Kumar Sankarapu
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Training-data attribution aims to identify which training examples shape model behavior, yet validating such claims is difficult because causal training influence is rarely observable. We argue that controlled synthetic pretraining makes attribution experimentally testable. Using O’PRIOR, a provenance-rich synthetic task generator for tabular foundation models, we construct a testbed in which every pretraining task carries explicit lineage over structural mechanisms, missingness, confounding, shortcuts, and distribution shift. We combine behavior-conditioned attribution with counterfactual retraining and provenance-aware interventions to test both task-level faithfulness and mechanism-level consistency. On held-out real tasks, removing the top-attributed 5% of synthetic tasks decreases mean ROC-AUC by 0.013, compared with 0.002 \pm 0.004 under random removal, while removing bottom-attributed tasks improves performance by 0.003. Within shortcut-provenance tasks, targeted removal yields an effect of 0.043 versus 0.016 for matched random removal. Provenance discrimination is more modest by ranking AUROC (0.55-0.62), despite substantial top-k enrichment, revealing that provenance association and interventional faithfulness need not coincide. Our results establish synthetic provenance as a controlled setting for verifiable contributive attribution

[LG-129] Mitigating Convergence Collapse in Fixed-Target Anomaly Detectors via Kernel-Anchored Locality Regularization CIKM2026

链接: https://arxiv.org/abs/2610.02345
作者: José Lucas De Melo Costa,Fabrice Popineau,Arpad Rimmel,Bich-Liên Doan
类目: Machine Learning (cs.LG)
*备注: Accepted at CIKM 2026 (oral). 11 pages

点击查看摘要

Abstract:A family of tabular anomaly detectors trains a neural map toward a fixed target under squared-error loss and scores anomalies by the test-time residual; contraction matching, one-step rectified flow, and reconstruction autoencoders all fit this template. We characterize a convergence collapse: better optimization makes the detector worse. At convergence, the learned map tracks the target even off-distribution, so the residual signal vanishes on anomalies as well as on normal data. These detectors therefore rely on implicit non-convergence (early stopping, capacity caps) to retain signal. We argue this is structural: effective anomaly detection requires a locality constraint that blocks unconstrained extrapolation. Classical detectors (kNN, KDE, isolation forests, LOF) enforce locality explicitly; fixed-target neural detectors do not. We formalize the connection by showing that the kernel-regression analog of a fixed-target detector is a finite-bandwidth Nadaraya-Watson smoother, which we call Kernel Contraction Matching (KCM). KCM is closed-form, training-free, and CPU-efficient, yet matches established neural baselines on ADBench. Building on this bridge, we introduce the Kernel-Anchored Regularizer (KAR), which penalizes deviation of the neural prediction from a kernel-weighted average of training targets. Across collapse-prone ADBench datasets and three backbones, KAR mitigates collapse and improves AUROC under prolonged training.

[LG-130] Drive vs. Decay: On the Training Dynamics of Joint-Embedding Predictive Architectures NEURIPS NEURIPS2026

链接: https://arxiv.org/abs/2610.02344
作者: José Lucas De Melo Costa,Seong Woo Ahn,Fabrice Popineau,Arpad Rimmel,Bich-Liên Doan
类目: Machine Learning (cs.LG)
*备注: Accepted at NeurIPS 2026 (Main Track). 57 pages, 10 pages of main text; appendices and the NeurIPS paper checklist included. Code: this https URL

点击查看摘要

Abstract:Joint-Embedding Predictive Architectures (JEPAs) are prone to representation collapse, typically mitigated through empirical heuristics. We develop an early-training stability theory that unifies these heuristics. Linearising the coupled JEPA gradient flow around the trivial fixed point reveals two competing effects: a driving force ( \gamma ) and a decay effect ( \sigma ). Under approximate spectral decoupling, a per-mode stability ratio \mu_i = \gamma_i / \sigma_i factorises into independent data-side and predictor-side terms and the count of unstable modes tracks the rank of representations that can emerge. The framework predicts a phase boundary, which we confirm empirically across more than 800 Tabular-JEPA configurations. It also unifies predictor scaling, masking ratio, and EMA as distinct mechanisms for shifting \mu . Guided by this analysis, we introduce ResidualPred, a transformer predictor whose attention is biased toward the identity at initialisation; it improves both effective rank and downstream accuracy on tabular benchmarks and in I-JEPA pretraining on CIFAR-10, CIFAR-100, STL-10, and ImageNet. Our framework connects empirical collapse-avoidance heuristics to an explicit dynamical picture, yielding theory-driven stabilizers. Code is available at this https URL.

[LG-131] NEEDLEWORK: Offline Rewriting of Robot Data with Verified Local Stitches ICLR2027

链接: https://arxiv.org/abs/2610.02339
作者: Juntao Ren,Yifan Hou,Shuran Song
类目: Robotics (cs.RO); Machine Learning (cs.LG)
*备注: Submitted to ICLR 2027. Project page: this https URL

点击查看摘要

Abstract:Robot demonstrations may contain useful behavior even when individual episodes are inefficient or unsuccessful. Trajectory stitching offers a way to compose these behaviors into improved training data, but identifying useful connections and verifying their feasibility is difficult in high-dimensional robot data, where many prior methods rely on low-dimensional state representations. We introduce NEEDLE, an offline dataset-augmentation algorithm that addresses these challenges by adding short, verified action bridges between recorded observations in high-dimensional robot demonstrations. First, NEEDLE identifies and creates connections that bypass suboptimal detours, broaden action coverage, and augment the original dataset with failed trajectories, using only RGB images, proprioception, and episode-level outcomes, without new environment interaction or privileged object state. Next, we present a sampling technique that incorporates accepted bridges into policy training without synthesizing intermediate images or discarding the original demonstrations, allowing policies to learn alternative actions while retaining the original dataset’s coverage. On real-robot tasks, NEEDLE improves success rate over the strongest baseline on each task by an average of 21 percentage points. Videos and supplementary materials are on this https URL.

[LG-132] SoTa: Soft Tactile Skins for Dexterous Manipulation

链接: https://arxiv.org/abs/2610.02338
作者: Jingyun Yang,Baiyu Shi,Timothy Yu,Haitian Liu,Alberta Longhini,Weichen Wang,Rika Antonova,Zhenan Bao,Jeannette Bohg
类目: Robotics (cs.RO); Machine Learning (cs.LG)
*备注: The first three authors contributed equally. Project website: this https URL

点击查看摘要

Abstract:A growing body of work suggests that tactile sensing gives robot policies contact information that complements vision in dexterous manipulation. However, visuo-tactile robot data remains scarce: dexterous demonstrations require teleoperating robots, which limits dataset scale. Human demonstrations are far cheaper to collect and offer a path to scale this data, but only if human and robot hands carry tactile sensors with corresponding signals. This requires sensors that conform to different hand geometries, cover the full hand, and share a common layout across embodiments. We present SoTa, a low-cost capacitive tactile skin that provides full-hand coverage on humans and robots while preserving a shared layout of 202 taxels across corresponding finger and palm regions. Our multilayer design with fabric electrodes enables in-house fabrication of thin, soft skins with customizable geometry for under 10 in materials per skin. The sensor retains over 97% of its initial response span after 10,000 loading-unloading cycles with traces retaining continuity through 1,280 tight-fist folding cycles. The shared taxel layout supports human-robot co-training with a common tactile encoder and no learned cross-sensor mapping. Across three contact-rich manipulation tasks, tactile observations improve in-distribution success over vision-only policies. With a fixed robot demonstration budget, adding human demonstrations more than doubles mean success across eight evaluation conditions, from 22.8% to 45.9%, improving success in all five out-of-distribution conditions. We plan to open-source the resources needed to fabricate and operate these skins.

[LG-133] Joint Movement and Compression Ratio Design for Mobile Embodied AI Networks (MEAN)

链接: https://arxiv.org/abs/2610.02334
作者: Yahao Ding,Jiaxiang Wang,Zhouxiang Zhao,Zhaohui Yang,Mingzhe Chen,Mohammad Shikh-Bahaei
类目: Information Theory (cs.IT); Machine Learning (cs.LG)
*备注: 6 pages, 4 figures, conference

点击查看摘要

Abstract:Mobile embodied AI networks (MEAN) enable embodied agents to perceive, reason, communicate, and act in wireless environments. In such networks, agent mobility can improve channel conditions, while semantic compression can reduce transmission payloads. However, movement consumes energy, and stronger compression incurs additional computational cost. This paper studies joint movement, semantic compression, and transmit power design for an uplink MEAN system. We formulate a max-min energy efficiency (EE) problem by jointly optimizing transmit power, movement distance, and semantic compression ratio under controllable power constraints. The problem is non-convex due to the coupled signal-to-interference-plus-noise ratio (SINR), mobility-dependent channel gains, and fractional EE objective. To solve it, we propose an alternating optimization (AO)-Dinkelbach algorithm, where the fractional objective is handled by the Dinkelbach transformation, transmit power is updated via successive convex approximation (SCA), and movement distance is updated by coordinate-wise grid search. Simulation results show that the proposed scheme outperforms no-mobility and no-compression baselines, demonstrating the benefit of jointly exploiting mobility control, semantic compression, and power allocation in MEAN.

[LG-134] PowerBench: Measuring Language Model Bias in Power-shifting Requests

链接: https://arxiv.org/abs/2610.02303
作者: Nicolas Martorell,Wendy Brau,Gonzalo A. Heredia,Tomás Pablo Korenblit,Gaspar Labastié,Tomás Gimenez Molina
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Language models increasingly assist people with power-related requests, so systematic differences in whom they help could shift the distribution of power at scale, or be exploited by users who learn which identities are refused less. We introduce PowerBench, an evaluation of power-shifting requests that distinguishes self-empowerment, disempowerment, and power grabbing, plus a control of refusal-inducing requests that shift no power. We build, curate, and open-source a dataset of such requests varying the power domain, the context, the scale of the affected party, and the prior power standing of the user, and evaluate 24 models (12 from US and 12 from Chinese developers) under three experimental conditions: reciprocal nationalities of user and affected party, an AI agent as the user, and 8 request languages. Models refuse power grabbing more than disempowerment, and disempowerment more than self-empowerment. Refusal of power grabbing rises with the scale of the affected party, from an individual to a society. Models are biased toward helping others take power from the US and against helping US users take power from others, but favor the US when it gains power and nobody loses it. When the user is an AI agent, refusal of power-shifting requests increases, especially in power grabbing against an individual. Finally, language biases refusal, but in model-specific ways that largely cancel on average. We release PowerBench to make these asymmetries measurable in current and future models.

[LG-135] Intent-Hiding Jailbreaks: An Information-Theoretic Framework for Compositional Attacks

链接: https://arxiv.org/abs/2610.02302
作者: Fengwei Tian,Ravi Tandon
类目: Cryptography and Security (cs.CR); Information Theory (cs.IT); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Recent work has shown that large language models (LLMs) can be vulnerable to jailbreak attacks in which harmful intent is obscured through composition with benign tasks. A harmful request refused in isolation may elicit a different response when embedded within a larger, seemingly benign query. We study these compositional intent-hiding jailbreaks from an information-theoretic perspective. Our formulation associates each task with an estimated probability of being judged harmful: the average over the full task collection defines the prior probability of harmful intent, while the average over a selected bundle containing the target defines the posterior. Selecting auxiliary tasks so that these averages agree, which we call prior-posterior matching, leaves the estimated intent unchanged even though the harmful target remains in the bundle. We study two settings that differ in whether query construction is part of the optimization. In the query-independent setting, tasks are selected without regard to how they will be expressed in the final query. We show that exact prior-posterior matching under a bundle-size constraint is computationally hard, derive an optimal water-filling solution for fractional weights, and characterize the smallest bundle satisfying a prescribed safety threshold. In the query-dependent setting, task selection and query construction are considered jointly, and intent concealment and target preservation are evaluated on the resulting query. We evaluate jailbreak effectiveness and preservation of the target behavior across bundle sizes, query generators, and several open-source models. These results show that compositional queries can elicit target behaviors beyond the direct-request baseline under the evaluated search budgets, while revealing a trade-off: as bundle size increases, response-level target preservation tends to decrease for several models. Subjects: Cryptography and Security (cs.CR); Information Theory (cs.IT); Machine Learning (cs.LG) Cite as: arXiv:2610.02302 [cs.CR] (or arXiv:2610.02302v1 [cs.CR] for this version) https://doi.org/10.48550/arXiv.2610.02302 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[LG-136] Effects of interpulse-interval variation on deep-learning classification of bat vocalizations

链接: https://arxiv.org/abs/2610.02284
作者: Welmoed R. Eversteijn,Burooj Ghani,A. Leonie Baier,Dan Stowell
类目: Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
*备注: 10 pages, 3 figures

点击查看摘要

Abstract:Temporal context may aid automated bat-species classification, but the contribution of specific features remains unclear. We investigated whether variation in the interpulse interval (IPI)-the time between consecutive call onsets-provides species-discriminative information and whether transformer-based models are more sensitive to this information than convolutional neural networks. We created two matched datasets from European bat recordings: a natural-IPI condition retaining the original call timing and a normalized-IPI condition in which call onsets were spaced at 50-ms intervals. EfficientNet-B0 and PaSST were fine-tuned and evaluated within each condition. In an additional experiment, each architecture was trained separately on natural-IPI and normalized-IPI recordings, and evaluated on the same natural-IPI test set. Finally, the pretrained classifiers BatDetect2 and BAT were evaluated on both conditions. Within-condition IPI normalization had model-dependent effects. PaSST accuracy differed little between the natural-IPI ( 71 \pm 2.3% ) and normalized-IPI ( 70 \pm 6.3% ) conditions, whereas EfficientNet accuracy increased from 47 \pm 4.7% to 57 \pm 3.9% . PaSST exceeded EfficientNet under both conditions. In the cross-condition evaluation, models trained on natural-IPI recordings outperformed those trained on normalized-IPI recordings on the natural-IPI test set: accuracy decreased from 54% to 50% for EfficientNet and from 65% to 57% for PaSST. BatDetect2 and BAT differed little between IPI conditions. Overall, we found limited support for the hypotheses that natural IPI variation contributes substantially to bat-species classification and that it is used more effectively by transformer-based than CNN-based models. Nevertheless, the cross-condition performance decrease shows that results obtained under normalized conditions may not transfer fully to natural recordings.

[LG-137] MuLoRA: Spectrally Balanced Low-Rank Adaptation for Continual Learning

链接: https://arxiv.org/abs/2610.02283
作者: Junkang Liu
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Low-rank adaptation (LoRA) provides a parameter-efficient approach to continual learning, but its nominal rank can conceal a loss of effective adaptation capacity. We identify \emphspectral plasticity collapse: during sequential adaptation, update energy becomes concentrated in a small subset of singular modes, leaving much of the available low-rank space underutilized. This exposes a limitation of interference avoidance alone: protecting historical representations does not ensure that the remaining adaptation capacity is responsive to new tasks or effectively utilized. To address this problem, we propose \textttMuLoRA, which jointly controls capacity allocation and utilization. First, historical whitening identifies input directions with strong current-task response relative to accumulated historical response, yielding a task-adaptive basis that remains fixed during training. Second, approximate polar orthogonalization of momentum updates reduces spectral concentration within theselected space. An orthonormal basis connects these mechanisms by transferring the factor-update spectrum exactly tothe induced weight update. We establish a max–min characterization of exact subspace selection and derive cumulative spectral bounds under controlled cross-step anisotropy. Across five class-incremental benchmarks and eight incremental settings, \textttMuLoRA achieves the highest mean accuracy in 15 of 16 reported metrics.

[LG-138] CLEAN: Psychometrically Consistent Incremental Cognitive Diagnosis under Concept-Space Expansion via Architectural Isolation

链接: https://arxiv.org/abs/2610.02278
作者: Tao He,Jinxing Xiang,Fan Jiang
类目: Machine Learning (cs.LG)
*备注: 11 pages, 7 figures

点击查看摘要

Abstract:Cognitive diagnosis (CD) is a fundamental task in intelligent education that profiles learner proficiency over knowledge concepts. In real-world learning platforms, newly added items continually introduce previously unseen concepts, necessitating dynamic expansion of the underlying concept space. Yet existing incremental CD models assume a fixed concept space, allowing gradients from new items to overwrite historical pathways and induce catastrophic forgetting. More critically, these methods rely solely on soft constraints to preserve historical diagnoses. Such constraints may fail to satisfy the requirement of diagnostic invariance after incremental updates, a requirement known as psychometric consistency in cognitive diagnosis. Therefore, we propose CLEAN (Continual Learning with Expandable and Architecturally Isolated Networks), a novel incremental CD framework supporting concept-space expansion while providing structural guarantees for pointwise invariance of historical diagnoses. Specifically, CLEAN first introduces a strict topological bipartition protocol, freezes historical diagnostic functions and applies deterministic orthogonal column masking to sever gradient interference. Second, to accommodate concept expansion, expandable full-rank branches with micro-variance initialization are deployed to learn novel concepts. Finally, to verify that this architectural design achieves invariance by construction, we formalize Representation Drift (RD) to quantify the perturbation of historical traits. Extensive experiments on three large-scale educational datasets demonstrate that CLEAN achieves zero RD, preserving old-item metrics identically to static anchors through architectural isolation while remaining competitive with or superior to strong continual-learning baselines on new items.

[LG-139] From Mathematical to Executable Certificates for Machine Unlearning

链接: https://arxiv.org/abs/2610.02268
作者: Ziyu Zhao,Xinyu Wang,Xiaowen Chang,Yixuan He
类目: Machine Learning (cs.LG); Cryptography and Security (cs.CR)
*备注:

点击查看摘要

Abstract:Machine unlearning is needed when data must be removed because of deletion requests, outdated records, or data-quality concerns, while retraining from scratch can be costly. Certified machine unlearning methods provide mathematical guarantees, while deployed systems release concrete finite-precision artifacts produced by software. To bridge the gap between mathematical guarantees and practical deployment, we introduce Executable Release Certification (ExecCert), a release-time layer that certifies the candidate artifact considered for release. ExecCert either closes a method’s native certificate for the executed candidate or applies Retraining-Reference Release Verification (RRV) to certify fidelity to current retain-set retraining. Sequential deletion makes the latter nontrivial because the exact retain-set reference and the stored numerical state evolve separately. For frozen representations with a mutable ridge head, we develop an incremental realization of RRV that maintains certified evidence across deletion requests rather than reconstructing it at each release. On four published unlearning implementations, ExecCert preserves valid certificates, changes release decisions, tightens conservative bounds, and identifies the retraining-reference fidelity supported by concrete outputs. In sequential-service experiments, RRV eliminates false releases caused by stored-equation verification while closely tracking realized error, and incremental certification remains cheaper than both fresh and maintained verified-factor alternatives once release checks become sufficiently frequent.

[LG-140] Parameter-Free Interval-Dynamic Regret under Heavy-Tailed Noise

链接: https://arxiv.org/abs/2610.02258
作者: Vaneet Aggarwal
类目: Machine Learning (cs.LG); Information Theory (cs.IT); Optimization and Control (math.OC)
*备注:

点击查看摘要

Abstract:We study online convex optimization with one unbiased stochastic subgradient per round and an unknown finite conditional p th noise moment, 1p\le2 . For every fixed interval I of length n and comparator path with \Lambda_I=1+P_I/D , one learner achieves [ E[Regret_I(u)]\le\min(GDn, C[GD\sqrtn(\Lambda_I+\log^2(2T)) +\sigma Dn^1/p(\Lambda_I+\log^2(2T))^(p-1)/p]). ] The learner uses none of G,\sigma,p,I,P_I , and the constant is universal. Interval adaptation adds to comparator complexity, preserving the distinct mean-gradient and noise exponents. The analysis controls calibration in expectation and limits the cost of observation-scale changes. Its general theorem compares to distributions over predictably available experts with relative-entropy dependence on a nonuniform prior. A common prior favors long windows and long restart lengths. With the statistics supplied, the interval cost becomes 1+\log(T/n) , including the optimal full-horizon static rate. A change-of-measure lower bound identifies the noise power of this logarithm for learners retaining a full-horizon optimal guarantee, under explicit conditions. Static comparisons and deterministic partitions follow from the same decisions. Subjects: Machine Learning (cs.LG); Information Theory (cs.IT); Optimization and Control (math.OC) Cite as: arXiv:2610.02258 [cs.LG] (or arXiv:2610.02258v1 [cs.LG] for this version) https://doi.org/10.48550/arXiv.2610.02258 Focus to learn more arXiv-issued DOI via DataCite

[LG-141] RACE: A Reproducible Benchmark for Electricity Price Forecasting with Official Operational Text NEURIPS2026

链接: https://arxiv.org/abs/2610.02256
作者: Xinyi Yi,Moy Yuan,Ioannis Lestas
类目: Machine Learning (cs.LG); Machine Learning (stat.ML)
*备注: Accepted to the NeurIPS 2026 Workshop on Foundation Models for Time Series (FMTS)

点击查看摘要

Abstract:Electricity price forecasting (EPF) supports scheduling, bidding, and risk management in electricity markets, yet existing benchmarks focus mainly on numerical inputs, leaving the forecasting value of forecast-time textual context insufficiently evaluated. We introduce TRACE, a reproducible benchmark of 7,300 zone–day instances pairing prices from five zones in a major U.S. market with official operational text available at the forecast cutoff. TRACE reconstructs official operational text at each cutoff, preventing post-cutoff information leakage. We evaluate TRACE for semantic alignment and forecasting value. Semantic assessments align with central movement and both tail risks in ground-truth prices, most consistently for upper-tail price risk. Forecasting value is reflected in a median 7.4% reduction in upper-tail pinball loss across time-series foundation models. A controlled cross-day text-mismatch ablation reverses the gains, falling below the no-text baseline.

[LG-142] Approximation Property of Dropout Neural Networks: Sobolev Rates and Confidence Bounds

链接: https://arxiv.org/abs/2610.02253
作者: Jia-He Yao
类目: Machine Learning (cs.LG); Numerical Analysis (math.NA)
*备注: 30 pages, 1 figure

点击查看摘要

Abstract:The universal approximation property of dropout neural networks does not by itself describe the network size required for an accurate random realization. In this work, we study approximation of the unit ball of W^n,\infty([0,1]^d) by ReLU networks whose edges are retained independently with probability p . The approximation error is measured uniformly over the input domain, and the guarantee holds with probability at least 1-\delta for a single sampled network. We construct networks of constant depth and size \widetilde O_n,d(p^-9\varepsilon^-\max\d/n,2\ \log(1/\delta)) . The construction combines bounded local subnetworks, localization on a successful approximation event, and a multiscale Taylor decomposition. Conversely, Sobolev capacity imposes a lower bound on the number of surviving edges, while approximation of a fixed affine function requires an output-layer cost of order ((1-p)/p)\varepsilon^-2\log(1/\delta) at sufficiently high confidence. For fixed p\in(0,1) and \delta\min\1/2,1-p\ , the upper and lower bounds match in the accuracy exponent under a fixed or logarithmic depth budget. When d\leq2n , they also match in confidence up to logarithms of accuracy. We extend the lower bounds to W^n,r targets with L^s error, and distinguish this extension from the upper bound for W^n,\infty . The optimal retention dependence and logarithmic factors remain open.

[LG-143] Rank-Aware Speculative Sampling for Diffusion Draft Trees

链接: https://arxiv.org/abs/2610.02251
作者: Marcello Bullo,Yanxiao Liu,Öykü Sıla Güner,Arpan Mukherjee,Deniz Gündüz
类目: Machine Learning (cs.LG); Computation (stat.CO)
*备注:

点击查看摘要

Abstract:Speculative sampling accelerates diffusion generation by verifying inexpensive draft states in parallel while preserving the target law. Recent tree-based methods allocate the parallel compute budget more effectively than single-chain drafts, as demonstrated by Diffusion Greedy Rejection Sampling (D-GRS). D-GRS generates K conditionally independent candidates per node, and sequentially tests them in their generation order. Yet the sampled candidates admit an informative ranking without additional target-model evaluations. To exploit this, we introduce Rank-Aware Speculative Sampling (RASS), a verification rule for speculative draft trees based on rank-aware list coupling. RASS orders draft candidates along the proposal-target mean displacement and samples a rank with weights optimized to minimize total variation between the selected-proposal and target laws. Finally, the selected candidate is maximally coupled with the target, with residual correction ensuring exact sampling for any choice of rank weights. We evaluate RASS on a Gaussian-mixture target, unconditional pixel-space generation on FFHQ, conditional generation on CIFAR-10, and latent diffusion with Stable Diffusion 3.5 using COCO2014 prompts. Measured by the ratio of standard to speculative sampling’s target-model evaluation counts, RASS improves on D-GRS across the evaluated settings, with gains reaching approximately 20% on CIFAR-10 at matched compute budgets.

[LG-144] Nearest-neighbour baselines for fingerprint prediction from MS/MS spectra under different assumptions

链接: https://arxiv.org/abs/2610.02249
作者: Ling Min Serena Khoo
类目: Machine Learning (cs.LG); Computational Engineering, Finance, and Science (cs.CE); Machine Learning (stat.ML)
*备注: 6 pages, 2 figures

点击查看摘要

Abstract:It has recently been shown that nearest-neighbour retrieval provides a strong baseline for molecular fingerprint prediction from MS/MS spectra, with several variants matching or outperforming current deep learning models (Khoo and Barzilay, 2026; Liu et al., 2026; Gupta et al., 2026). Importantly, “nearest neighbour” encompasses a family of retrieval methods that differ in the information assumed to be available at inference. In this report, we systematically compare several nearest-neighbour variants and show how these differing assumptions affect performance. Our goal is to establish stricter baselines that enable more rigorous benchmarking and better measure progress in this area.

[LG-145] State-Space Unlearning for Non-Stationary Bias in Land Surface Forecasting

链接: https://arxiv.org/abs/2610.02248
作者: Anidipta Pal
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Operational land surface forecasting systems built on Mamba-family Structured State Space Models absorb non-stationary confounding events (unrecorded irrigation booms, dam-operation shifts, sensor recalibrations) into their state-transition matrices, silently biasing NDVI, LST, and crop phenology predictions long after the physical cause ends. This paper introduces SSU-LSF (State-Space Unlearning for Land Surface Forecasting), the first machine-unlearning framework purpose-built for geoscientific Mamba-based SSMs. We develop EKFac influence functions specialized to the Mamba state matrices via a closed-form matrix-exponential gradient, use spectral-radius-weighted elbow thresholding to localize a temporal confounding footprint \Phi , and apply Hessian-free projected gradient ascent within a KL-divergence trust region augmented by spatial total-variation (TV) regularization. Proposition 1 establishes that residual confounding is bounded by \mathcalO\big((1-\rho(\barA)^T_c)/((1-\rho(\barA))\mu)\big) , which grows with the window length T_c . Across three heterogeneous benchmarks and eleven baselines, SSU-LSF achieves confounding reduction rates of 0.773 (CropHarvest), 0.821 (NDVI-LST), and 0.859 (ERA5), with worst-case clean-domain RMSE degradation of 4.2% on ERA5, converging in 3–5 epochs at 8.4\times lower GPU-cost per unlearning request than full retraining. Code: this https URL

[LG-146] he Price of Greenwashing: Algorithmic Verification and Market Discipline using Conformal Machine Learning

链接: https://arxiv.org/abs/2610.02225
作者: Sourav Bose,Taoufik Bouraoui
类目: Machine Learning (cs.LG); Econometrics (econ.EM)
*备注:

点击查看摘要

Abstract:While corporate sustainability mandates are expanding, the systemic reliance on self-reported emissions data exposes financial markets to pervasive greenwashing. Current literature relies heavily on subjective ESG ratings or textual sentiment analysis, leaving a critical econometric gap in objectively quantifying physical climate realities. To resolve this information asymmetry, we fuse U.S. SEC financial fundamentals with facility-level EPA greenhouse gas registries to establish a mathematically guaranteed baseline of physical corporate emissions. Leveraging a gradient boosting architecture and Mondrian Conformal Prediction, we quantify the shortfall between self-reported data and this algorithmic baseline into a novel Conformal-Weighted Continuous Divergence (CWCD) metric. Evaluating this divergence via a cross-sectional lead-lag econometric design, we uncover a robust mechanism of market discipline: algorithmic emissions divergence exhibits a severe, statistically significant negative relationship with subsequent market valuation (Tobin’s Q) and operational profitability (ROA). Providing definitive evidence against the market blindness hypothesis, this study proves that institutional capital actively prices environmental deception not merely as an ethical lapse, but as a leading indicator of fundamental corporate mismanagement. Ultimately, these findings provide the quantitative justification necessary for asset managers and regulators to deploy algorithmic auditing infrastructure at scale.

[LG-147] Hybrid Machine Learning-Assisted Raman Spectroscopy with Generative Feature Augmentation for Pharmaceutical Identification

链接: https://arxiv.org/abs/2610.02224
作者: Quach Thi Thai Binh,Ton Nu Quynh Trang,Thang B. Phan,Vu Thi Hanh Thu,Nguyen Tuan Hung
类目: Machine Learning (cs.LG); Materials Science (cond-mat.mtrl-sci)
*备注: 12 pages, 7 figures, 2 tables

点击查看摘要

Abstract:Rapid and reliable identification of pharmaceutical residues is important for safeguarding public health, ensuring food safety, and enabling practical Raman-based screening. In this study, we propose HyMLRaman, a hybrid Raman spectroscopy framework that combines deep spectral feature extraction, generative models, and classical machine-learning classifiers to identify six pharmaceutical compounds, including amoxicillin, chloramphenicol, ciprofloxacin, tetracycline, ibuprofen, and paracetamol. Raman spectra are converted into spectral images and encoded with several deep neural-network backbones, among which EfficientNet-B3 yields the most effective representation. The resulting 1536-dimensional embeddings are then used to train downstream classifiers, including SVM, KNN, logistic regression, random forest, XGBoost, and ANN, using stratified 10-fold cross-validation. The hybrid EfficientNet-B3–SVM configuration achieves the strongest baseline performance, reaching 96.31% accuracy and a macro-F1 score of 96.36%, outperforming the standalone CNN baseline. To address limited-data conditions, a generative model, a DDPM-based feature augmentation, is introduced in a PCA-reduced EfficientNet-B3 latent space. The low-data ablation results show that DDPM augmentation provides selective benefits, particularly for KNN with reduced training fractions, and that its effect remains classifier-dependent. Finally, an application-level Raman Pharmaceutical Analyzer demonstrates the feasibility of embedding the trained model into an interactive Raman analysis workflow. These results suggest that HyMLRaman provides a practical and interpretable route for rapid Raman-based pharmaceutical screening.

[LG-148] RINS: Residual-Image Neural Subspace Solvers for Large Sparse Linear Systems

链接: https://arxiv.org/abs/2610.02217
作者: Zhongyan Ouyang,Weixin Liao,Mingquan Feng,Yehui Tang,Junchi Yan
类目: Computational Engineering, Finance, and Science (cs.CE); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Large sparse linear systems from PDE discretizations require correction subspaces whose operator images explain the current residual. We study this residual-image viewpoint and propose Gate-RINS, a neural subspace solver that generates polynomial correction bases from cached residual probes and modulates them with a lightweight residual- and coordinate-dependent pointwise gate. The projected least-squares update remains unchanged, so the neural component only chooses the expansion directions while the numerical closure tests them through (\operatornamerange(AQ_t)). We also introduce a hybrid controller schedule that composes GRANS-style graph controllers with Gate-RINS under the same projected solver. Across six PDE-derived benchmark tasks and two scales, Gate-RINS reaches fixed relative-residual thresholds faster in synchronized wall-clock time than GMRES and a recent graph-only neural baseline in most settings, and hybrid schedules further improve the residual trajectory. Difficult-mode diagnostics and trajectory visualizations support the interpretation that these gains are associated with operator-image subspaces that align more effectively with the current residual.

[LG-149] MiDShip: Multimodal Dataset of Ship Cargo Hold Structures for Engineering Design

链接: https://arxiv.org/abs/2610.02214
作者: Noah J. Bagazinski,Md. Ferdous Alam,Jaya Manideep Rebbagondla,Faez Ahmed
类目: Computational Engineering, Finance, and Science (cs.CE); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Ship structures govern vessel strength, safety, and manufacturability, but their design must satisfy hundreds of classification society requirements, making the process complex and iterative. Data-driven approaches are limited by the lack of structured datasets linking design geometry, structural performance, and rule-based constraints. This paper presents MiDShip, a multimodal dataset of 12,753 synthetic cargo-hold structural designs: 6,020 random, 496 generated by an SGLD-inspired procedure, and 6,237 generated by an equation-informed repair procedure. Each design includes parametric data, full and mesh-ready 3D geometry, engineering drawings and annotations, a bill of materials, and preliminary structural evaluations. Twenty-five constraints derived from a subset of ABS MVR are also evaluated. None of the random designs satisfies all constraints. Among the SGLD-inspired designs, 322 (64.9%) were fully compliant, with an average of 0.409 violations, 82.7% below the seed mean and 96.9% below the random-design mean. The repair procedure, developed through LLM-assisted code analysis, produced 4,952 fully compliant designs (79.4%), averaging 0.296 violations, 97.1% below the paired-source mean. In equal-size comparisons, mean nearest-neighbor distances in the scaled 120-parameter space were 3.495 for repaired designs, 1.144 for SGLD batches, and 3.729 for random designs. The primary contribution is the synchronized dataset and its generation and evaluation infrastructure; the generation studies demonstrate its utility rather than proposing new optimization algorithms. MiDShip supports machine learning, generative design, and automated rule-based evaluation for ship structures. Subjects: Computational Engineering, Finance, and Science (cs.CE); Machine Learning (cs.LG) Cite as: arXiv:2610.02214 [cs.CE] (or arXiv:2610.02214v1 [cs.CE] for this version) https://doi.org/10.48550/arXiv.2610.02214 Focus to learn more arXiv-issued DOI via DataCite

[LG-150] From Mixing to Tearing: Graph Decomposition in Decentralized Optimization via Message Passing

链接: https://arxiv.org/abs/2610.03709
作者: Kuangyu Ding,Gesualdo Scutari
类目: Optimization and Control (math.OC); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:We study the minimization of sums of smooth strongly convex functions over undirected graphs, with each function held by one agent and communication restricted to neighbors in the graph. Existing decentralized methods, whether based on gossip or on routing over spanning trees, typically use the network to mix or aggregate information to enable \it prescribed local optimization updates. What this communication-centered viewpoint lacks is a general framework that uses graph structure to \it jointly design the optimization subproblems and the cooperative computation and communication through which agents solve them cooperatively. We develop such a framework from first principles, jointly designing the linear representation of agreement constraints, the blocks of the resulting dual variables (jointly optimized), and connected cluster of agents that cooperatively solve each block subproblem over the assigned subgraph. GATE (Graph-Tearing message passing) is a first instance of this framework: one variable per edge and tree blocks. At each iteration, agents update their assigned edge variables by minimizing the sum of the two endpoint cost-to-go messages and relaxing the result. The messages are updated through local minimizations following the tree recursion. To reduce per-iteration computational and communication costs, we develop GATE-S, a surrogate variant using tractable local models and lightweight message parametrizations. We establish linear convergence with a rate explicit in the interplay among function regularity, network topology, and the chosen partition, revealing the effects of graph decomposition. Numerical experiments are conducted to validate the theoretical results and evaluate the efficiency of our algorithms. Subjects: Optimization and Control (math.OC); Machine Learning (cs.LG) Cite as: arXiv:2610.03709 [math.OC] (or arXiv:2610.03709v1 [math.OC] for this version) https://doi.org/10.48550/arXiv.2610.03709 Focus to learn more arXiv-issued DOI via DataCite (pending registration) Submission history From: Kuangyu Ding [view email] [v1] Fri, 2 Oct 2026 17:57:50 UTC (2,823 KB) Full-text links: Access Paper: View a PDF of the paper titled From Mixing to Tearing: Graph Decomposition in Decentralized Optimization via Message Passing, by Kuangyu Ding and Gesualdo ScutariView PDFHTML (experimental)TeX Source view license Current browse context: math.OC prev | next new | recent | 2026-10 Change to browse by: cs cs.LG math References Citations NASA ADSGoogle Scholar Semantic Scholar export BibTeX citation Loading… BibTeX formatted citation loading… Data provided by: Bookmark checked="checked"class=“labs-tab-input”> Bibliographic Tools Bibliographic and Citation Tools Bibliographic Explorer Toggle Bibliographic Explorer (What is the Explorer?) Connected Papers Toggle Connected Papers (What is Connected Papers?) Litmaps Toggle Litmaps (What is Litmaps?) scite.ai Toggle scite Smart Citations (What are Smart Citations?) Code, Data, Media Code, Data and Media Associated with this Article alphaXiv Toggle alphaXiv (What is alphaXiv?) Links to Code Toggle CatalyzeX Code Finder for Papers (What is CatalyzeX?) DagsHub Toggle DagsHub (What is DagsHub?) GotitPub Toggle Gotit.pub (What is GotitPub?) Huggingface Toggle Hugging Face (What is Huggingface?) ScienceCast Toggle ScienceCast (What is ScienceCast?) Demos Demos Replicate Toggle Replicate (What is Replicate?) Spaces Toggle Hugging Face Spaces (What is Spaces?) Spaces Toggle TXYZ.AI (What is TXYZ.AI?) Related Papers Recommenders and Search Tools Link to Influence Flower Influence Flower (What are Influence Flowers?) Core recommender toggle CORE Recommender (What is CORE?) Author Venue Institution Topic About arXivLabs arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv’s community? Learn more about arXivLabs. Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?) mathjaxToggle(); We gratefully acknowledge support from our major funders, member institutions, , and all contributors. About Help Contact Subscribe Copyright Privacy Accessibility Operational Status (opens in new tab) Major funding support from

[LG-151] Amortized Structured Stochastic Variational Inference for Gaussian Process Latent Variable Models

链接: https://arxiv.org/abs/2610.03647
作者: Maksym Tretiakov,Sarah Lucie Filipp,Vincent Fortuin,Ruth Misener,Ruby Sedgwick,James Odgers
类目: Machine Learning (stat.ML); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Many machine learning methods aim to approximate the lower-dimensional manifold on which the data lives. A desirable feature of such methods is that they should capture the epistemic uncertainty of this learned manifold. One model that achieves this is the Gaussian Process Latent Variable Model, in which a Gaussian Process (GP) mapping from the latent space provides an estimate of the uncertainty of the manifold. However, the effectiveness of this uncertainty estimation is limited by the mean-field variational approximation between the GP inducing points and the latent variables. In this work, we apply Amortized Structured Stochastic Variational Inference to allow the variational posterior for the latent space to be conditionally dependent on the value of the inducing points. We demonstrate that this more flexible variational posterior improves several metrics relating to the reconstruction of points on the data manifold.

[LG-152] When Is Accuracy Evidence? A Unified Theory of Generalisation Validation and Information Fusion

链接: https://arxiv.org/abs/2610.03465
作者: JM Gorriz
类目: Machine Learning (stat.ML); Machine Learning (cs.LG); Data Analysis, Statistics and Probability (physics.data-an)
*备注: 52 pages, 30 figures

点击查看摘要

Abstract:K-fold cross-validation (CV) is widely used as evidence of out-of-sample performance, although folds are neither independent experiments nor equally informative under heterogeneous data. Cross Upper-Bound Validation (CUBV) replaces point-wise CV accuracy by conservative upper bounds on true risk. Here we generalise CUBV through a single exponential framework in which the moment-generating function of the generalisation gap is controlled by a cumulant envelope gamma(lambda). This yields a family of risk bounds covering Hoeffding-, Bernstein-, dependency-aware, PAC-Bayesian, and heterogeneous source-fusion settings. For K-fold CV, dependence between fold-wise gaps is modelled through a joint sub-Gaussian proxy matrix. Under equicorrelation, this gives an effective number of folds, Keff = K/[1+(K-1)rho], showing that increasing K does not necessarily increase statistical evidence when folds are strongly dependent. The framework is also extended to posterior distributions over predictors and weighted multi-source fusion, where weights are selected by minimising an upper bound on future risk rather than empirical error alone. Experiments with trained linear classifiers on heterogeneous multimodal Gaussian mixtures compare K-fold CV with full-sample resubstitution plus risk correction. Bounds are evaluated by coverage and tightness. In low-dimensional small-sample settings, K-fold partitioning can increase uncertainty because individual folds under-represent minority modes, while corrected resubstitution can remain valid and tighter; this effect disappears as sample size increases. Overall, gamma-CUBV separates observed performance, uncertainty, dependence, model complexity, and confidence into explicit terms, providing a unified route from CV scores to risk statements and a principled validation criterion for heterogeneous small-sample applications such as neuroimaging.

[LG-153] Generalization of Transformer-Based Neural Quantum States via In-Context Learning

链接: https://arxiv.org/abs/2610.03463
作者: Zhen Qin,Qing Qu,Alfred O. Hero III
类目: Quantum Physics (quant-ph); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Neural quantum states based on modern deep learning architectures have emerged as powerful representations for quantum many-body systems. In particular, Transformer-based neural quantum states provide expressive models capable of capturing long-range correlations, and their empirical generalization performance has recently been demonstrated. However, a theoretical understanding of their generalization behavior remains largely unexplored. In this paper, we develop a theoretical framework to analyze the generalization properties of Transformer-based neural quantum states under in-context learning. We establish a rigorous inference-time generalization error bound in terms of mean squared error (MSE), showing that the pointwise prediction error decreases inversely with both the number of in-context examples and the depth of the Transformer. We further show that the Transformer depth required to achieve this guarantee scales only linearly with the system size–namely, the number of particles in continuous systems or the number of qudits in discrete systems. Building on this result, we extend our analysis to full quantum states formulated as rank-one density operators, and derive MSE-based generalization bounds over both continuous and discrete domains under physical constraints. Finally, numerical simulations corroborate our theoretical analysis.

[LG-154] Electronic Density versus Geometry for Machine-Learned Molecular Absorption Spectra

链接: https://arxiv.org/abs/2610.03444
作者: Siddharth Dhanpal,Peter Elliott,Paolo Emilio Trevisanutto,Alin M. Elena,Gilberto Teobaldi
类目: Chemical Physics (physics.chem-ph); Materials Science (cond-mat.mtrl-sci); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Molecular optical absorption spectroscopy provides a direct probe of electronic structure and is widely used for molecular identification, interpretation of photophysical behaviour, and planning of spectroscopy experiments. Calculating the absorption spectra using first-principle excited-state methods, however, is computationally demanding, at least compared to ground-state calculations, which limits their routine application across large molecular sets. Machine-learning (ML) surrogates can reduce this cost and allow rapid spectral prediction. However, their performance depends strongly on how molecular information is represented. Here, we compare using the ground-state electron density versus the molecular geometry as inputs to a ML model for predicting absorption spectra, for a training set of 6874 molecules selected from the QM7 dataset. For each of these molecules, the density was calculated using density functional theory (DFT) and the absorption spectrum was calculated using linear-response (LR) time-dependent DFT (TDDFT). Utilizing the ground-state density as the input to the ML model is motivated by the Hohenberg-Kohn and Runge-Gross theorems, and the fact that the ground-state density encodes information about bonding, charge localisation, and electronic delocalisation. Hence, it may be a more judicious starting point for the ML model compared to the geometry, as it effectively decouples the chemistry of the ground-state. The question we test is whether the benefits of using the density outweigh the (notprohibitive) penalty of requiring an additional single-point DFT calculation for the density. We find that the density-based convolutional neural network achieves a validation correlation of 0.9926, compared with 0.9795 for the best geometry-based graph model, reducing the residual decorrelation, by approximately 64%.

[LG-155] Contrastive Neural Embeddings Reveal Individual Traits Beyond Conversational Role

链接: https://arxiv.org/abs/2610.03410
作者: Hubert Huang,Michelle McCleod,Brendan Ames,Evie Malaia
类目: Neurons and Cognition (q-bio.NC); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Contrastive representation learning is increasingly used to recover low-dimensional structure from neural recordings, but its output is typically validated by decoding accuracy rather than by the geometry of the manifold it produces. We apply CEBRA to EEG recorded from dyads in conversation, and analyze the resulting embedding, which training constrains to the 2D sphere. Labels describing the dyads, including the absolute difference between partners’ autism-quotient scores, decode well above chance (0.77 against a 0.55 majority baseline for binary AQ magnitude; 0.44 against 0.25 for the six-class |\Delta AQ | partition). However, the two permutation controls have notable differences in results: permuting labels over a frozen embedding yields p = 0.001, whereas retraining the encoder under each permutation yields p = 0.50. Only the latter tests the label rather than the geometry. Consistent with this, spherical mixture structure and per-class dispersion track identity rather than autism trait differences in dyads; frequency-band and non-oscillatory activity ablation controls do not change the results. However, participant-level model does separate from its identity-aware null (p = 0.0099) while speaker-versus-listener role analysis performs at chance in the same embedding, indicating a manifold organized by individual – and, in contrast with current neurolinguistics models, almost invariant to speaking vs. listening. Based on these results, we suggest that retraining-based nulls should be the default for grouped-data contrastive embeddings.

[LG-156] Mixture-of-Experts for Cryptocurrency Order Execution: Training Stability Tail Risk and Failure Modes

链接: https://arxiv.org/abs/2610.03369
作者: Alexander Ardaiz,Varun Budati,Ali Habibnia
类目: Trading and Market Microstructure (q-fin.TR); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Deep reinforcement-learning policies for order execution can vary substantially across training seeds, so apparent architectural gains may reflect favourable training realisations rather than reproducible properties of the architecture. We evaluate vanilla Double Deep Q-Learning (DDQL), K-means-partitioned mixtures of DDQL experts at K \in \2, 4, 8\ , and dense networks parameter-matched to the K=4 and K=8 expert budgets on 5-minute mean-aggregated BTC/USDT limit order book data from Binance. No learned configuration significantly improves mean implementation shortfall over DDQL. Under the reported specification, all have higher mean shortfall than TWAP (0.39 bps) and immediate liquidation (0.21 bps) in an environment whose frictionless replay and terminal-urgency penalty make early liquidation nearly costless; 11/100 vanilla-DDQL runs, versus none in either MoE K\geq4 arm, converge to a policy that waits until forced liquidation. We then decompose this specification on a device-matched baseline. Annealed exploration alone eliminates observed collapses (12/100 to 0/100; exact McNemar p=4.9\times10^-4 ), matching the elimination under expert partitioning. Combining annealed exploration with the aligned reward restores collapse in 19/30 runs; with all three specification changes, it rises to 48/100. In this environment, expert partitioning is unnecessary to suppress collapse and appears to mask a training-specification failure rather than confer an intrinsic performance benefit. No MoE K=8 run collapses under any of the six specifications tested. Across-seed dispersion is lowest at K=8 but non-monotone and not robust to family-wise adjustment, while within-policy tail risk worsens monotonically with K . The apparent attribution of the failure mode reverses between 30 and 100 seeds, illustrating the importance of repeated-seed evaluation.

[LG-157] DAWIS: Data Assimilation with Windowed Inverse Sampling via Multitask Interpolants

链接: https://arxiv.org/abs/2610.03314
作者: Erik Wikingsson,Martin Andrae,Tomas Landelius,Fredrik Lindsten
类目: Machine Learning (stat.ML); Machine Learning (cs.LG); Atmospheric and Oceanic Physics (physics.ao-ph)
*备注:

点击查看摘要

Abstract:Flow- and diffusion-based generative models have recently emerged as flexible and highly efficient forecasting models for dynamical systems. When combined with inference-time guidance, they offer a promising route to high-dimensional non-Gaussian data assimilation (DA), the problem of combining forecasts with observations to estimate latent system states. Existing filters, however, condition on a fixed history and assimilate only the most recent observation, leaving them unable to revise past states when new observations arrive. Estimates then stay tethered to a history that later observations may contradict, and errors accumulate over the assimilation run. To this end, we introduce DAWIS, a unified DA method covering filtering, fixed-lag smoothing, and block smoothing within a single framework. DAWIS replaces the single flow time of a state-level prior with a multitask stochastic interpolant over a window of consecutive states, assigning a separate flow time to each. An assimilation cycle inverts the window to a vector of per-state turning points and regenerates it under observation guidance, with the turning points controlling how strongly each state is held fixed, revised, or generated from scratch. The same construction can also absorb the forecast into the assimilation cycle, removing the need for a separate forecasting model. Experiments on challenging nonlinear systems show that DAWIS improves on both filtering and smoothing baselines under sparse, noisy, and nonlinear observations. The code for DAWIS is available at this https URL

[LG-158] SDECast: Probabilistic Weather Forecasting in Continuous Time with Neural SDEs STOC NEURIPS2026

链接: https://arxiv.org/abs/2610.03313
作者: Maria Marchenko,Martin Andrae,Fredrik Lindsten,Christian A. Naesseth
类目: Machine Learning (stat.ML); Machine Learning (cs.LG); Atmospheric and Oceanic Physics (physics.ao-ph)
*备注: Accepted to AI for Stochastic Dynamics Sim2Science workshops at NeurIPS 2026

点击查看摘要

Abstract:Existing machine learning weather forecasting models typically generate forecasts through autoregressive rollouts at a fixed temporal resolution. While highly efficient for long-range prediction, this formulation can suffer from severe error accumulation when used with shorter time steps and does not explicitly encode the locality and temporal continuity of atmospheric dynamics. To address these limitations, we introduce SDECast, a Neural Stochastic Differential Equation (SDE) framework for continuous-time probabilistic weather forecasting. SDECast extends SDE Matching to learn stochastic dynamics directly in physical space, without requiring repeated SDE simulation during training. On a simulated geophysical flow, we show that SDECast recovers meaningful drift dynamics and faithfully reproduces the underlying continuous-time behavior. We then demonstrate its scalability to global weather forecasting at hourly resolution, where SDECast produces skillful probabilistic forecasts for lead times of up to five days.

[LG-159] Near-Optimal Convex Optimization with Lazy Second-Order Oracles

链接: https://arxiv.org/abs/2610.03222
作者: Xinliang Zhang,Lesi Chen,Chengchang Liu,Jingzhao Zhang
类目: Optimization and Control (math.OC); Machine Learning (cs.LG); Machine Learning (stat.ML)
*备注:

点击查看摘要

Abstract:This paper studies the complexity of convex optimization using lazy second-order oracles (Doikov, Chayti, and Jaggi, ICML 2023), where an algorithm queries gradients every iteration and Hessians once per m iterations. Under this setting, we show a lower bound of \Omega(m+ m^1/7 \epsilon^-2/7) on the number of total iterations to find an \epsilon -solution using a novel block zero-chain construction. Then we propose a novel method that achieves a new upper bound of \tilde\mathcalO(m+ m^1/7 \epsilon^-2/7) , which significantly improves the prior one (Chen, Liu, Luo, and Zhang, COLT 2026) of \tilde\mathcalO(m+ m^13/21 \epsilon^-2/7) and is tight up to logarithmic factors.

[LG-160] Hamiltonian locality testing and certification do not achieve the Heisenberg limit

链接: https://arxiv.org/abs/2610.03205
作者: Francisco Escudero Gutiérrez,Junseo Lee,Sebastian Zur
类目: Quantum Physics (quant-ph); Computational Complexity (cs.CC); Data Structures and Algorithms (cs.DS); Machine Learning (cs.LG)
*备注: 28 pages

点击查看摘要

Abstract:We establish lower bounds for Hamiltonian property testing with access to the time-evolution operator but not its inverse. Each experiment may query the time-evolution operator multiple times, and distances between Hamiltonians are measured in the normalized Frobenius norm. In this model, we show that testing whether a Hamiltonian is k -local or \varepsilon -far from every k -local Hamiltonian requires \Omega(1/\varepsilon^2) total evolution time, matching the upper bound of Kallaugher and Liang (TQC’25). We also prove that testing whether an unknown Hamiltonian equals a target Hamiltonian or is \varepsilon -far from it requires \Omega(1/\varepsilon^2) total evolution time, matching the upper bound of Sinha and Tong (2025). These are the first lower bounds for natural problems in Hamiltonian learning and testing that rule out Heisenberg-limited scaling of 1/\varepsilon . As a third result, we show that amplitude estimation to precision \varepsilon requires \Omega(1/\varepsilon^2) total time evolution, recovering the result of Tang and Wright (QIP’26) in the continuous-time query model. All three results follow from the hardness of distinguishing the zero Hamiltonian from a suitably chosen ensemble of random Hamiltonians. We establish this hardness by adapting the continuous-time adversary method to forward Hamiltonian evolution. Comments: 28 pages Subjects: Quantum Physics (quant-ph); Computational Complexity (cs.CC); Data Structures and Algorithms (cs.DS); Machine Learning (cs.LG) Cite as: arXiv:2610.03205 [quant-ph] (or arXiv:2610.03205v1 [quant-ph] for this version) https://doi.org/10.48550/arXiv.2610.03205 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[LG-161] Predictively Oriented Gaussian Process Posteriors

链接: https://arxiv.org/abs/2610.03201
作者: Callum Lau,Jeremias Knoblauch,Louis Sharrock
类目: Machine Learning (stat.ML); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Gaussian Processes (GPs) are a powerful tool for modelling and quantifying uncertainty in functional relationships. However, they require practitioners to make a number of design decisions, such as the choice of the kernel and the observation model. Suboptimal choices can produce misspecified models that do not capture the underlying data generating process. We introduce Predictively Oriented Gaussian Processes (PrO-GPs), which treat predictive uncertainty as the primary inferential target and provide a robust alternative to standard GPs. Although direct computation of a PrO posterior for nonparametric models is intractable, we derive a reduced formulation and practical sampling scheme for efficient computation. Through synthetic and real data experiments, we show that PrO-GPs produce better calibrated predictive distributions under model misspecification compared to standard GP approaches.

[LG-162] Landscape-Dependent Performance of Photonic Quantum Solvers in QUBO Feature Selection for Financial Risk Detection

链接: https://arxiv.org/abs/2610.03161
作者: Nirvik Sahoo,Paul Robert Griffin
类目: Quantum Physics (quant-ph); Machine Learning (cs.LG); Risk Management (q-fin.RM)
*备注: 39 Pages, 41 Tables, 3 Figures

点击查看摘要

Abstract:Feature selection for imbalanced classification tasks such as credit card fraud and consumer default detection requires balancing predictive relevance, inter-feature redundancy, and computational feasibility. We benchmark three computing paradigms, classical branch-and-bound optimization (Gurobi), photonic entropy computing (QCI Dirac-3), and simulated photonic boson sampling (Piquasso), across thirteen feature-selection methods on two datasets: ULB Credit Card Fraud (30 features) and AmEx consumer default (159 features). Each method is routed to the solver matched to its mathematical structure. On ULB, Dirac-3 MI-Spearman matches the all-features model using 13 of 30 features (mean F1 0.873 +/- 0.023 over five runs, best run 0.896), and Piquasso is the best method at k=5. On AmEx, performance rises steadily with the feature budget and every paradigm approaches F1 = 0.80 only near the full feature set. Most differences between Gurobi and Dirac-3 on identical methods fall within run-to-run variation; the large gaps occur where the certified optimum generalizes poorly, most sharply for distance correlation on AmEx at k=25 (Gurobi F1 = 0.422 vs. a Dirac-3 mean of 0.746). At matched budgets, F1 varies about ten times more across methods on ULB than on AmEx, which we trace to how concentrated the predictive signal is in each feature space.

[LG-163] Invariance of Clustering Operations in Causal Effect Identification

链接: https://arxiv.org/abs/2610.03101
作者: Jani Nykänen,Otto Tabell,Santtu Tikka,Juha Karvanen
类目: Machine Learning (stat.ML); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Clustering variables in causal graphs reduces the size of the graph and simplifies causal inference. However, arbitrary clustering can alter crucial causal relations among variables and lead to erroneous conclusions. While the identifiability of a causal effect in the clustered graph implies the identifiability in the original graph under mild conditions, nonidentifiability in clustered graph does not imply nonidentifiability in the original graph without further assumptions. When both identifiability and nonidentifiability are preserved, the clustering operation is called identification invariant. We present a broad class of clustering operations that are identification invariant based on conditions related to the c-components of the original graph. Finally, we demonstrate use of the results in practical settings.

[LG-164] Explainable Molecular Structure Inference from GC–MS with Diffusion Models and LLM Reranking

链接: https://arxiv.org/abs/2610.03066
作者: Changlin Liu,Tianyu Yi,Chengchun Liu,Boxuan Zhao,Fanyang Mo
类目: Chemical Physics (physics.chem-ph); Machine Learning (cs.LG); Computational Physics (physics.comp-ph); Data Analysis, Statistics and Probability (physics.data-an)
*备注: 4 figures

点击查看摘要

Abstract:GC–EI–MS is an important technique for analyzing volatile and semivolatile compounds in complex samples. However, conventional methods rely heavily on reference spectral library matching, limiting their ability to identify compounds absent from these libraries and to infer complete molecular structures directly from fragmentation information. Here, we present DiffGCMS, a spectrum-conditioned discrete graph diffusion model for de novo structure elucidation from GC–EI–MS, and further develop a framework that integrates DiffGCMS with second-stage reasoning by a large language model (LLM). In the first stage, DiffGCMS generates candidate molecular structures from input spectra; in the second stage, the LLM uses mass spectral information to validate, repair, and rerank the candidates and provides interpretable analysis of fragment-ion peaks. This framework can generate plausible molecular structures for compounds absent from reference spectral libraries and provide traceable evidence supporting its decisions. On a test set comprising 13,696 spectra from NIST 20, the generative model achieved Acc@1 and Acc@10 of 6.01% and 15.76%, respectively. On the test subset containing molecules with no more than 10 heavy atoms, LLM-assisted molecular graph repair and reranking increased Acc@1 from 21.28% to 21.95%, Acc@10 from 46.91% to 47.99%, and candidate validity from 91.04% to 100%. These results demonstrate that spectrum-aware postprocessing can correct errors produced by the generative model while providing auditable and traceable explanations for the final ranking.

[LG-165] FASTDIAR: Frame-level speaker encoder for Streaming Diarization ICASSP2027

链接: https://arxiv.org/abs/2610.02941
作者: Nikita Torgashov,Okan Köpüklü
类目: Audio and Speech Processing (eess.AS); Machine Learning (cs.LG); Sound (cs.SD)
*备注: 5 pages, submitted to IEEE ICASSP 2027

点击查看摘要

Abstract:Real-time conversational agents require speaker diarization that streams and runs on a CPU. Most systems apply an utterance-level speaker encoder to short, heavily overlapping chunks, which wastes computation and leaves the model optimized for the wrong task. We instead turn a state-of-the-art speaker recognition architecture into a causal frame-level encoder that reads the stream once and emits one embedding every 80~ms from a bounded two-second window of past audio, and pair it with online clustering that gates every update on the self-similarity of the stream. Trained only by distillation from an utterance-level teacher on simulated and out-of-domain mixtures, and evaluated with one fixed set of hyperparameters, the system is the most accurate streaming diarizer on low-overlap benchmarks at sub-second latency, degrades far less than cache-based systems as the number of speakers grows, and runs five times faster than real time on a single CPU thread.

[LG-166] Hold-Out Scoring for Efficient Gaussian DAG Learning

链接: https://arxiv.org/abs/2610.02785
作者: Donguk Shin,Byeongguk Kang,Inseol Lee,Gunwoong Park
类目: Machine Learning (stat.ML); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:High-dimensional Gaussian DAG learning faces a statistical-computational gap: methods with sharp sample complexity rely on computationally expensive subset search and a supplied indegree bound, whereas polynomial-time alternatives have less favorable sample complexity. We introduce HOST, an efficient DAG learning algorithm that replaces subset search with nodewise hold-out scoring and convex regression, without requiring a supplied indegree bound. Our key insight is that recovering a correct ordering does not require uniformly small estimation errors in ordering scores but only one-sided control of those errors. In the ordering step, HOST exploits the fact that score estimation using hold-out samples inflates ordering scores in expectation, which is the favorable direction for candidates that should not yet be selected. Given the ordering, HOST recovers parents by recursively removing indirect effects from total effects between two nodes. Under suitable conditions, HOST exactly recovers a p -node DAG of maximum indegree d with sample complexity of order d\log p in polynomial time. Experiments show that HOST achieves competitive graph recovery while exhibiting favorable runtime scaling.

[LG-167] When Normalization Selects the Sign: Auditing Robustness Ablations in Quantum Attention NEURIPS2026

链接: https://arxiv.org/abs/2610.02641
作者: Owen Friedewald,Srikar Alla,Ali Shiri Sichani,Chi-Ren Shyu
类目: Quantum Physics (quant-ph); Cryptography and Security (cs.CR); Machine Learning (cs.LG)
*备注: Accepted for presentation as a long oral at the NeurIPS 2026 SaTQuML Workshop, December 12-13, 2026, Atlanta, GA. arXiv version includes minor formatting revisions to meet submission requirements

点击查看摘要

Abstract:Removing an input-scaling module changes both a classifier and the perturbations reaching its encoder. A robustness difference can therefore reflect the comparison rule as well as the module. We demonstrate this problem in a four-qubit quantum-attention detector on generated power-grid trajectories. A learned scaling module appears beneficial at a fixed physical attack budget, but matching an upper bound on perturbations at the encoder reverses the ordering. Neither comparison alone establishes a robustness benefit caused by the module. The initial test also perturbs clean examples into attacked examples while retaining their original labels; tests restricted to already attacked examples do not establish a benefit. Replacing a trained model’s input scales disrupts detection. Retraining its linear classification layer restores the detection rate, but changes individual predictions, leaving the comparison descriptive rather than causal. Two further design checks explain why the input quantum Fisher information regularizer cannot train this model’s query parameters, and why removing confidence bounds does not establish a larger certified radius. The evidence is limited to ten seeds, exact simulation, synthetic data, and a restricted set of attacks; classical baselines achieve better clean prediction. The practical lesson is to specify which perturbation budget is fixed, check that attacks preserve labels and interventions preserve predictions, and distinguish exploratory controls from confirmatory evidence.

[LG-168] Where Quantum Fourier Sampling Stops Short: A Three-Gate Audit Protocol for Delay-PUF Security Models NEURIPS2026

链接: https://arxiv.org/abs/2610.02636
作者: Owen Friedewald,Ali Shiri Sichani,Chi-Ren Shyu
类目: Quantum Physics (quant-ph); Cryptography and Security (cs.CR); Machine Learning (cs.LG)
*备注: Accepted for presentation as a long oral at the NeurIPS 2026 SaTQuML Workshop, December 12-13, 2026, Atlanta, GA. arXiv version includes minor formatting revisions to meet submission requirements

点击查看摘要

Abstract:Quantum Fourier sampling may help audit the spectral learnability of delay-based physical unclonable functions (PUFs). We ask whether that promise survives access matching, a strong classical comparator, and oracle synthesis. Three gates structure the evaluation. Structure: low degree is not small support at reachable challenge lengths; for 4-XOR at n=14 , degree \le d_f(0.1) admits 91% of all 2^n characters and the median 90% -mass set spans a third of the spectrum. Algorithmics: constructing the phase oracle logically implies classical membership access, making Kushilevitz–Mansour the correct baseline; across 45 tasks it exhausts each finite domain, and no 4-XOR ideal-sampling case reaches 90% mass within 2^n calls. A quantum-kernel diagnostic appears more favorable, with geometric difference rising to 2.151 at N=512 challenges, but it correlates 0.991 with 1/\sqrt\lambda_\min(K_C) for the classical Gram matrix K_C , and the 4-XOR label-complexity ratio does not exceed a balance-preserving permutation null ( p=0.930 ). Trace-normalized geometric difference can therefore grow through classical ill-conditioning alone, without task-label alignment. Implementation: a simulator-validated fixed-point phase oracle based on the quantum Fourier transform admits an 18.9% routed-depth reduction, yet the least certified precisions have estimated durations of 1.18 – 1.55\times the median dephasing time T_2 of the mapped qubits on a static backend snapshot, without hardware execution. We find no end-to-end advantage in the evaluated regime, although ideal sampling does use fewer coherent calls on the thresholded task. The contribution is the Three-Gate Quantum Audit Protocol: a reproducible procedure separating an ideal query advantage from a realizable security benefit. This is not a claim about deployed silicon and not an impossibility result.

[LG-169] A hybrid CNN-adjoint optimization framework for reconstruction of viscoelastic tissue properties in magnetic resonance elastography

链接: https://arxiv.org/abs/2610.02634
作者: Anwesa Dey,Johann Rudi,Elena Cherkaev
类目: Optimization and Control (math.OC); Machine Learning (cs.LG); Numerical Analysis (math.NA)
*备注:

点击查看摘要

Abstract:Magnetic resonance elastography (MRE) is a noninvasive imaging modality for quantifying the viscoelastic properties of soft tissues from shear wave propagation. Recovering the complex-valued shear modulus from measured displacement fields leads to a severely ill-posed inverse problem, particularly in the presence of noise and limited boundary excitations. We investigate adjoint-based optimization, convolutional neural network (CNN) reconstruction, and a hybrid framework combining both approaches. The forward model is based on a scalar form of the modified stationary Stokes system with a complex shear modulus. We establish well-posedness of the forward problem, existence of minimizers, and first-order optimality conditions for the adjoint-based formulation, and implement a nonlinear conjugate-gradient method with Armijo line search. While PDE-constrained optimization can accurately refine coefficient reconstructions, its performance depends strongly on initialization. We therefore construct a two-dimensional CNN that maps complex-valued displacement measurements to spatially varying complex shear modulus fields and provides rapid, informative initial reconstructions. The proposed hybrid method uses the CNN reconstruction to initialize the adjoint-based optimization, yielding faster convergence and improved accuracy. The CNN is trained on coefficient fields containing individual perturbations and tested on both individual and previously unseen combined configurations. Numerical experiments demonstrate that the CNN generalizes to these more challenging configurations, while subsequent PDE-constrained optimization further refines the reconstructed coefficient. These results demonstrate the potential of combining data-driven initialization with physics-based optimization for efficient and accurate MRE reconstruction.

[LG-170] High-Dimensional Asymptotics and Dataset Selection for Private Transfer Learning

链接: https://arxiv.org/abs/2610.02578
作者: Filip Kovačević,Edwige Cyffers,Stefano Sarao Mannelli,Marco Mondelli
类目: Machine Learning (stat.ML); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:To commit to buying external data or participate in collaborative learning, one must decide whether the additional data will improve prediction enough to justify the cost. This comes with several challenges: (i) the decision often relies only on aggregated statistics available publicly, rather than individual-level data; (ii) covariate and model shifts can induce negative transfer, so the additional data deteriorates rather than improves performance; (iii) if the data is sensitive, its privatization requires the injection of noise, which can also offset the benefit of a larger sample size. In this paper, we model the problem of dataset selection through high-dimensional regression with multiple heterogeneous sources and a weighted ridge estimator. Our approach uses only summary statistics and it gives privacy guarantees either on labels only or jointly on features and labels, in terms of \rho -zero-concentrated differential privacy. The main technical contribution is a deterministic equivalent of the test error, which captures the interactions between sample size, covariance structure, model shift, regularization and privacy noise. Our theory allows to optimize hyperparameters (weights and ridge regularizers) and, more broadly, to decide when private external datasets are useful without accessing the data itself but only relying on population-level quantities. This provides a theoretically tractable foundation for private transfer learning, which we support via experiments on both synthetic and real-world datasets.

[LG-171] ENCORE: Exact Non-equilibrium COntrol with Replica Exchange for Diffusion Generation NEURIPS2026

链接: https://arxiv.org/abs/2610.02538
作者: Jiahao Yu,Saifuddin Syed,José Miguel Hernández-Lobato,Jiajun He
类目: Machine Learning (stat.ML); Machine Learning (cs.LG)
*备注: A shorter version of this work was accepted at the NeurIPS 2026 PriGM Workshop

点击查看摘要

Abstract:Inference-time control steers a pretrained generative model towards a target distribution without retraining. We study tilted targets \pi_0\propto G_0,p_0 , where p_0 is the sampler output distribution and G_0 is an evaluable reweighting function. Existing approaches rely on sequential annealing with sequential Monte Carlo (SMC) or parallel annealing with replica exchange (RE). Sequential control is exact but needs large particle populations, whereas no exact parallel control method exists: existing RE corrections approximate an intractable time reversal and are biased. We propose Exact Non-equilibrium COntrol with Replica Exchange (ENCORE), the first exact parallel control method. Each replica stores its generation trajectory, so the upward move is a truncation and the intractable time reversal is never simulated. We prove target invariance and show that the resulting dynamics are those of non-equilibrium replica exchange with the exact time reversal as forward proposal. Under regularity conditions, our diffusion analysis shows that both sequential and parallel control become unstable under refinement of the time discretisation without guidance, whereas guided proposals remain stable and yield diagnostics for tuning the schedule and the computational budget. Across synthetic targets, Boltzmann sampling of biomolecules, and image generation, ENCORE achieves competitive accuracy and diversity, remains robust to sampler perturbations, and applies to distilled samplers where existing RE corrections are unavailable.

[LG-172] Conformal Prediction for Time Series with Deep Sequence Models

链接: https://arxiv.org/abs/2610.02357
作者: Junghwan Lee,Jonghyeok Lee,Yao Xie
类目: Machine Learning (stat.ML); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Recent advances in deep learning for time series prediction have amplified the need for reliable uncertainty quantification. Conformal prediction has gained attention as a distribution-free framework for constructing prediction intervals with coverage guarantees. However, its coverage guarantees rely on data exchangeability, an assumption generally violated in time series. Active research has focused on developing conformal prediction methods for time series that overcome this limitation. While deep sequence models, such as recurrent neural networks and Transformers, have often been used in conformal prediction for time series, limited work has systematically studied how deep sequence models can be utilized in conformal prediction for time series. In this work, we systematically investigate the use of deep sequence models in conformal prediction for time series through three approaches: conditional quantile regression, conditional quantile function estimation, and localized conformal prediction. We provide a theoretical analysis establishing asymptotic conditional coverage guarantees for all three approaches under suitable assumptions. Through comprehensive experiments on real-world datasets, we demonstrate the effectiveness of leveraging deep sequence models into conformal prediction for time series.

[LG-173] Expected Utility Regret Rule: Minimax and Bayes Optimal Portfolio Choice

链接: https://arxiv.org/abs/2610.02290
作者: Masahiro Kato
类目: Econometrics (econ.EM); Machine Learning (cs.LG); Statistics Theory (math.ST); Mathematical Finance (q-fin.MF); Machine Learning (stat.ML)
*备注:

点击查看摘要

Abstract:This study considers the problem of portfolio choice, where we recommend a portfolio to an investor to maximize the expected utility of their wealth. Our goal is to construct an asymptotically optimal portfolio choice rule in terms of expected utility regret, the difference between the expected utility of an oracle investor and that achieved by a portfolio chosen from data. We propose the Expected Utility Regret (EUR) rule, which jointly selects a portfolio class and estimates its weights. In a regular parametric return model, a single EUR rule attains both the minimax and the Bayes lower bounds, including their leading constants, without using the prior distribution that defines the Bayes criterion. We then derive the mean–variance and risk-parity portfolios as special cases of this framework. Under smooth increasing and concave utility, the EUR rule and the sample mean–variance portfolio attain the same leading expected regret when expected excess returns approach zero sufficiently fast. When the returns divided by their volatilities have a joint distribution that does not depend on the order of the assets, the EUR rule and the risk-parity portfolio coincide.

[LG-174] A Missing Latent Not a Missing Simulator: Radius-Augmented Inference for Real JWST Retrieval

链接: https://arxiv.org/abs/2610.02245
作者: Angshuman Chakravertty,P V V Raj
类目: Instrumentation and Methods for Astrophysics (astro-ph.IM); Earth and Planetary Astrophysics (astro-ph.EP); Machine Learning (cs.LG)
*备注: 6 pages, 2 figures, 2 tables. Includes appendix and paper checklist

点击查看摘要

Abstract:Amortized simulation-based inference (SBI), which is trained on radiative-transfer simulators, recovers exoplanet atmospheres accurately on synthetic James Webb Space Telescope (JWST) spectra but collapses when it comes to real reduced spectra. The flow posterior collapsed on real WASP-39b (importance-sampling effective sample size (ESS) = 1, best-fit \chi2/N = 301), and such a failure is usually blamed on missing the forward-model physics, but ruling these levers out with nested sampling first (a temperature gradient, SO2 opacity, and a high-fidelity opacity set) leaves the fit unchanged, meaning the collapse is not from the simulator misspecification but instead from a missing latent, the planet radius. To fix this, we build MIRAGE, a radius-augmented flow-matching posterior calibrated against an independent nested-sampling reference with importance sampling and an optimal-transport map, which yields a physical and literature-consistent retrieval of real WASP-39b (with \chi2/N from 301 to 0.06). This same method transfers unchanged across two instruments and three real JWST targets, including one spectrum self-reduced end-to-end from raw Mikulski Archive for Space Telescopes(MAST) data. The lesson is cross-domain, as a latent the simulator encodes but the inference omits can masquerade as misspecification.

[LG-175] Generalizable single-cell perturbation response prediction using energy-guided flow matching

链接: https://arxiv.org/abs/2610.02232
作者: Jianan Wei,Jiajun Hong,Guikun Chen,Ning Yang,Lifeng Fan,Wenguan Wang
类目: Genomics (q-bio.GN); Machine Learning (cs.LG); Cell Behavior (q-bio.CB)
*备注:

点击查看摘要

Abstract:Predicting phenotypic and transcriptional responses to perturbations at single-cell resolution provides a powerful tool for probing biological systems. However, existing methods typically rely on fixed mappings learned during training, making it challenging to calibrate distribution shifts or adapt to novel perturbation conditions during inference. Here, we present scEGFlow, an energy-guided flow matching framework that dynamically bridges control and perturbed cellular states. scEGFlow models continuous transitions from control cell populations to perturbed states using conditional flow matching. It then applies condition-specific energy gradients to correct and steer these predictions, enabling flexible adjustments without retraining the flow model. Evaluations across benchmarks spanning imaging phenotypes and transcriptomic profiles show that scEGFlow outperforms existing methods in reconstructing response distributions under both seen and unseen perturbation conditions, faithfully preserving cellular manifold geometry and population heterogeneity. This advantage is notable when adapting to new conditions with only a few measured cells, consistently improving prediction accuracy. Furthermore, scEGFlow accurately recapitulates perturbation-induced up- and down-regulation patterns across consensus gene expression signatures, where energy guidance improves the agreement between predicted and observed regulatory directions. Ultimately, these findings demonstrate that scEGFlow provides a modular, generalizable, and steerable solution for single-cell perturbation modeling. By grounding generative flows in learned biological landscapes, this architecture establishes a new computational paradigm for navigating and manipulating cellular behavior in silico.

附件下载

点击下载今日全部论文列表