本篇博文主要内容为 2026-10-08 从Arxiv.org论文网站获取的最新论文列表,自动更新,按照NLP、CV、ML、AI、IR、MA六个大方向区分。

说明:每日论文数据从Arxiv.org获取,每天早上12:30左右定时自动更新。

提示: 当天未及时更新,有可能是Arxiv当日未有新的论文发布,也有可能是脚本出错。尽可能会在当天修复。

目录

概览 (2026-10-08)

今日共更新1028篇论文,其中:

  • 自然语言处理共118篇(Computation and Language (cs.CL))
  • 人工智能共285篇(Artificial Intelligence (cs.AI))
  • 计算机视觉共186篇(Computer Vision and Pattern Recognition (cs.CV))
  • 机器学习共355篇(Machine Learning (cs.LG))
  • 多智能体系统共14篇(Multiagent Systems (cs.MA))
  • 信息检索共19篇(Information Retrieval (cs.IR))
  • 人机交互共35篇(Human-Computer Interaction (cs.HC))

多智能体系统

[MA-0] A Society of Researchers: Designing Institutions for Populations of Autonomous Research Agents

【速读】:该论文旨在解决大规模研究型智能体(research agents)在数千个智能体共享同一计算资源池时缺乏有效组织机制的问题。当前多数系统仍采用单项目或无组织的部署模式,难以应对群体规模扩展带来的协调与效率挑战。其核心解决方案是提出“智能体社会”(society of agents)这一新型架构,即一群持久存在的智能体在明确制度框架下协同运作,并以六项基本原则构建科学领域的分布式研究生态系统。关键创新在于引入类似科研机构的治理范式:由首席研究员通过提案申请、同行评审和资助竞争获取计算资源,而由一名人类“市长”(mayor)负责资源分配但不下达具体任务,实现去中心化自主性与资源公平配置的平衡。在实际运行中,一个包含一万个研究智能体的社会系统仅以优化语言模型预训练为目标,便涌现出显著的效率改进——某实验室提出可在保持性能的前提下减少约30%的计算开销,尽管该成果尚未达成普遍共识。该研究揭示了制度化组织对智能体群体演化的重要性,并为多智能体系统社区提出了六个待解的关键开放问题。

链接: https://arxiv.org/abs/2610.10468
作者: Ali Asaria,Deep Gandhi,Tony Salomone
机构: Transformer Lab(Transformer实验室); Canada(加拿大)
类目: Multiagent Systems (cs.MA); Artificial Intelligence (cs.AI)
备注: 4 pages. Short version of A Society of Researchers: Institutional Design for Populations of Autonomous Scientific Agents, doi: https://doi.org/10.5281/zenodo.22922325

点击查看摘要

Abstract:Deployments of research agents are moving to populations of thousands that share one pool of compute, while most current systems organize one project at a time or leave the population unorganized. We argue that such a population will acquire an organization whether or not its designers provide one, so designers should provide it explicitly, and that the multi-agent systems community holds the tools to do so. We propose a society of agents, a population of persistent agents under explicit institutions, and develop it for science as a society of researchers built on six principles. Principal investigators compete for compute through requests for proposals, independent review, and grants; a human governor, the mayor, allocates resources and assigns no tasks. In a running society of ten thousand researchers, asked only to improve the pretraining of language models, one lab reported a way to reach the same quality with about 30% less compute, a result the labs that tested it do not yet agree on. We close with six open problems for the agents community.

[MA-1] Know the Shape Find the Fault: Topology-Conditioned Diagnosis of Multi-Agent LLM Failures

【速读】:该论文旨在解决多智能体大语言模型(Multi-agent LLM)系统在任务执行过程中因协作失效而导致的故障诊断难题。具体而言,当协作机制失效时,执行轨迹中呈现出相似的症状可能源于信息传递、使用或验证环节的不同问题,而现有方法缺乏对故障模式与通信拓扑结构之间关联性的建模能力,难以实现精准诊断。其解决方案的关键在于提出MAScope——一种两阶段的拓扑条件化诊断框架。第一阶段通过“轨迹结构提取器”(Trace Structural Extractor, TSE)从异构执行轨迹中恢复通信拓扑结构,基于消息证据构建交互图;第二阶段利用“拓扑条件化判别器”(Topology-Conditioned Judge, TC-Judge),结合轨迹数据、预测的拓扑结构、从标注轨迹中学习到的经验性故障先验以及拓扑特定的故障模式描述进行故障分类。实验表明,通信拓扑与故障类型间存在显著统计关联(χ² = 409.9, p = 1.2 × 10⁻⁷⁰),引入真实拓扑可使gpt-mini的宏平均F1值从0.173提升至0.350;使用预测拓扑时仍达到0.346,接近仅依赖轨迹的gpt-5.4基线(0.372)。在固定编排结构下,单次拓扑提取后可复用,整体推理成本仅为重复调用gpt-5.4的约6%,证明了拓扑条件化上下文能有效提升诊断性能并支持低成本部署。

链接: https://arxiv.org/abs/2610.10126
作者: Xinwen Liu,Zhuocheng Pan,Isabella Zhu,Jawei Zhang,Xudong Liu,Tianyu Wo
机构: Beihang University(北京航空航天大学)
类目: Multiagent Systems (cs.MA); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Multi-agent LLM systems coordinate task execution through exchanges of information among agents. When coordination breaks down, similar symptoms in execution traces can reflect different problems in how information is passed, used, or verified. Communication topology captures how agents exchange information and provides structural cues for distinguishing coordination failure modes. Using these cues for diagnosis requires establishing how topology relates to failure patterns and recovering the relevant structure from execution traces that lack explicit topology labels. We analyze the relationship between communication topology and failure patterns and introduce MAScope, a two-stage framework for topology-conditioned diagnosis. Its Trace Structural Extractor TSE recovers communication topology from heterogeneous execution traces by grounding an interaction graph in message evidence. The Topology-Conditioned Judge TC-Judge then classifies failures using the trace, predicted topology, an empirical failure prior estimated from separate labeled traces, and a short description of topology-specific failure patterns. Under a fixed orchestration structure, the recovered topology can be reused across executions. Experimental results show a statistically significant association between communication topology and failure type, with \chi^2 = 409.9 and p = 1.2 \times 10^-70 . On the \num851 MAST-clean traces, ground-truth topology context raises gpt-mini’s Macro-F1 from 0.173 to 0.350 . With predicted topology, the pipeline achieves 0.346 , approaching the trace-only gpt-5.4 baseline of 0.372 . For \num1000 traces under a fixed orchestration structure, the projected pipeline cost, including one topology extraction, is approximately 6% of repeated gpt-5.4 diagnosis cost. These results show that topology-conditioned context improves failure diagnosis and supports lower-cost deployment.

[MA-2] he Cost of Classical Multi-Agent Path Finding

【速读】:该论文旨在解决经典多智能体路径规划(Multi-Agent Path Finding, MAPF)在现实物理环境中因离散时间假设和图结构冲突模型导致的解质量受限问题。经典MAPF假设时间离散且仅考虑节点间的冲突,虽简化了求解过程,但限制了路径规划的精确性与实际可行性,尤其在复杂或高密度场景中难以实现真正无碰撞的路径。其核心问题是:这种简化在何种情况下会显著牺牲解的质量?本文提出的解决方案是引入连续时间多智能体路径规划(MAPF_R),通过放宽离散时间与图结构冲突的限制,允许更精细的动作空间(如连续时间移动、考虑智能体形状),从而提升路径规划的精度与效率。关键在于,连续时间与智能体形状建模本身价值有限,其真正优势在于支持更丰富的移动动作选择,平均可使狭窄受限地图上的解质量提升至少5%,开放空间地图上提升达17%,部分场景甚至超过20%;而同等计算开销下,单纯提高经典MAPF的地图分辨率仅能恢复不足3%的性能增益,表明经典模型所丧失的解质量无法通过增加算力弥补。因此,本研究为判断何时使用经典MAPF作为合理近似,以及何时必须采用MAPF_R以获得显著更优解提供了明确依据。

链接: https://arxiv.org/abs/2610.10100
作者: Alvin Combrink,Sabino Francesco Roselli,Martin Fabian
机构: Chalmers University of Technology (查尔姆斯理工大学); Gothenburg (哥德堡); Sweden (瑞典)
类目: Multiagent Systems (cs.MA)
备注: 9 pages, 5 figures

点击查看摘要

Abstract:Multi-Agent Path Finding (MAPF) is the problem of planning conflict-free paths for multiple agents in a shared space, each from its start to its goal. Classical MAPF has been the dominant formulation for many years, with its assumptions of discrete time and graph-based conflicts presumably easing the search for solutions. These assumptions limit the physical environments and agents for which a solution is truly collision-free, and also place an upper bound on solution quality that no algorithmic improvements can lift. This work investigates how much solution quality, and in what contexts, the classical MAPF formulation forfeits. Continuous-time MAPF (MAPF _R ) relaxes these assumptions, making it a natural counter-formulation to compare against across various agent counts and sizes, and graph connectedness, topologies, and resolutions. We find that continuous time and agent shape consideration are worth relatively little on their own; their value comes from enabling an expanded range of move actions, on average improving solution quality by at least 5% on narrow and constrained maps and 17% on maps with open spaces. In some cases, the improvements exceed 20% . Doubling the map resolution with classical MAPF recovers less than 3% , meaning that little of what is forfeited can be bought back through more compute. This work therefore provides insight on when classical MAPF is a reasonable simplification, and when MAPF _R unlocks significantly higher-quality solutions.

[MA-3] Multi-Agent Coordination via Support-Preserving Distillation NEURIPS2026

【速读】:该论文旨在解决离线多智能体强化学习(Offline MARL)中基于生成式策略建模联合行为时存在的多模态协调失效问题。其核心挑战在于:传统的基于流模型的集中式教师网络在训练过程中,将噪声与回放缓冲区目标独立配对,导致邻近噪声样本被错误地分配至冲突的协调模式,从而生成介于有效模式之间的无效样本;由于学生网络通过条件均值蒸馏机制从教师输出中学习,此类误差无法被吸收而会被传递至学生端,进而损害策略性能。本文提出的关键解决方案是模式支持半离散最优传输(Mode-Support Semi-Discrete Optimal Transport, MoSDOT),其核心思想是在教师训练前,利用半离散最优传输方法将多模态回放缓冲区数据压缩为具有预设容量的有限模式支撑集,并确保每个噪声样本仅被分配至单一协调模式,从而消除教师端因噪声与目标不匹配所引发的模式混淆问题。此外,研究还引入共享随机性变体以揭示严格乘积执行架构下固有的残余差距。实验结果表明,MoSDOT在控制性诊断和离线MARL基准测试中显著提升了最终策略质量与模式路由一致性,尤其在呈现多模态联合行为的数据集上表现突出。

链接: https://arxiv.org/abs/2610.10087
作者: Sangmin Lee,Youngju Na,Chanmi Lee,Sung-eui Yoon
机构: KAIST(韩国科学技术院)
类目: Machine Learning (cs.LG); Multiagent Systems (cs.MA); Robotics (cs.RO)
备注: Accepted at NeurIPS 2026 (Main Track, Poster)

点击查看摘要

Abstract:Offline MARL increasingly relies on generative policies to model multimodal joint behavior, typically by distilling a centralized teacher into decentralized one-step actors under the CTDE. We identify a failure mode at the teacher training stage: standard flow-based teachers pair noise with replay targets independently, so nearby noise samples can be routed toward conflicting coordination modes. The teacher then produces samples between valid modes, and because the distillation loss regresses each local actor onto the conditional mean of the teacher’s output given local input, this error is not absorbed but propagated to the student. To remove this teacher-side artifact, we propose Mode-Support Semi-Discrete Optimal Transport (MoSDOT), which summarizes multimodal replay into a finite mode support with prescribed capacities and uses conditional semi-discrete optimal transport to assign each noise sample to a single mode before teacher training. We additionally study a shared-randomness variant that uses a shared noise component at execution to expose the residual gap intrinsic to strict-product execution. On controlled diagnostics and offline MARL benchmarks, MoSDOT improves endpoint quality and routing consistency, particularly on datasets exhibiting multimodal joint behavior.

[MA-4] CANDO: Cooperative Agent ic Network for Layout Design Optimization NEURIPS2026

【速读】:该论文旨在解决真实世界设施布局生成中的复杂设计挑战,包括不规则场地边界、异构朝向、可达性感知的布局安排以及运动规划可行性等问题。现有生成式AI领域的布局基准多集中于矩形区域上的简化布局任务,且依赖分布度量(如FID和IoU)评估模型性能,此类指标倾向于奖励与数据集先验一致的结果,从而抑制了设计创新。为弥补这一不足,论文提出了ALPS-Bench,一个包含1,000个专业标注的真实世界设施布局数据集,并配套基于结构化设计手册的实例级评分协议。针对该基准,作者提出CANDO——一种无需训练的多智能体框架,通过专业化智能体在验证驱动的迭代循环中协同优化布局,聚焦于战略性空间决策的推理。实验表明,CANDO在PubLayNet、RICO和PKU-PosterLayout等主流基准上超越了最先进的训练型模型及基于大语言模型(LLM)的基线方法,验证了协作式智能体设计在约束感知布局生成中的广泛有效性。解决方案的关键在于:基于验证反馈的多智能体协同迭代机制,将设计推理集中于高层空间策略,而非低层生成细节。

链接: https://arxiv.org/abs/2610.10044
作者: Athanasios Masouris,Zheng Jing,Benjamin Sam Chandler,Hadi Jamali-Rad
机构: Shell Information Technology International(壳牌信息技术国际); Shell China Limited(壳牌中国有限公司); Delft University of Technology (TU Delft)(代尔夫特理工大学)
类目: Multiagent Systems (cs.MA)
备注: Accepted to appear in NeurIPS 2026. This is the preprint version of the paper

点击查看摘要

Abstract:Layout generation for real-world facilities is a challenging problem, requiring reasoning over irregular site boundaries, heterogeneous orientations, access-aware placements, and motion-planning feasibility. Yet, most existing layout benchmarks in the generative AI space target simpler placements over rectangular domains and rely on distributional metrics such as FID and IoU that reward conformity to dataset priors, thus discounting design innovation. Motivated by these gaps, we introduce ALPS-Bench, a benchmark of 1,000 professionally annotated real-world facility layouts paired with an instance-specific scoring protocol grounded in a structured design manual. As a strong baseline for ALPS-Bench, we propose CANDO, a training-free multi-agent framework in which specialized agents iteratively refine layouts through a verification-grounded loop, concentrating reasoning on strategic spatial decisions. We demonstrate that CANDO surpasses state-of-the-art trained and LLM-based baselines on the widely adopted PubLayNet, RICO, and PKU-PosterLayout benchmarks, establishing cooperative agentic design as a broadly effective recipe for constraint-aware layout synthesis.

[MA-5] Minimizing Cumulative Envy in Allocating a Sequence of Items

【速读】:该论文旨在解决不可分物品(indivisible goods)按时间序列到达且必须不可撤销分配时的公平性问题,其核心挑战在于如何在已知未来物品到达顺序和所有参与者的估值前提下,衡量并最小化整个分配过程中累积的不公平感。传统在线公平分配模型通常假设未来信息未知,而本文引入了前瞻性的已知信息假设,聚焦于分析“不公平感”(envy)随时间演化的动态特性。为此,作者提出累积最大嫉妒(cumulative maximum envy) 作为关键度量指标,即每一时刻最大成对嫉妒值的累加,等价于最坏嫉妒曲线下的面积,同时捕捉了嫉妒的强度与持续时间。针对固定物品到达顺序的情形,研究证明该优化问题为强NP完全,且在相同估值甚至二元估值下均不存在常数因子近似解(除非P=NP);但通过设计动态规划算法,实现了在代理人数量为常数时的伪多项式时间可解性,并在特定受限条件下获得多项式时间算法,以及在相同整数估值下对固定人数的近似方案(FPTAS)。此外,论文还探讨了可选择物品到达顺序的调度变体,尽管该问题在两人同质估值下仍为NP完全,但提出了一个简单的贪心算法,可在两人情况下达到3/2-近似,在任意人数下达到$ n/(n-1) $-近似,并提供基于单个物品最大价值的加性误差保证,从而为实际应用提供了高效可行的解决方案。

链接: https://arxiv.org/abs/2610.09843
作者: Paul W. Goldberg,Isaac Robinson,Nicholas Teh
机构: University of Oxford (牛津大学)
类目: Computer Science and Game Theory (cs.GT); Multiagent Systems (cs.MA)
备注: 40 pages. Preliminary version of this paper was presented at the 19th SAGT (Sept. 2026)

点击查看摘要

Abstract:We study temporal fair division with indivisible goods that arrive sequentially and must be allocated irrevocably. In contrast to the usual online model, we assume that valuations and future arrivals are known in advance, and ask how unfairness evolves during the process. We introduce \emphcumulative maximum envy: the sum, over all rounds, of the maximum pairwise envy at that round. Equivalently, this is the area under the worst-envy curve, and it captures both the magnitude and the duration of envy. For a fixed arrival order, we show that the corresponding decision problem is strongly NP-complete and that minimizing this objective admits no constant-factor approximation unless P = NP, even under identical valuations and even under binary valuations. We complement these hardness results with a dynamic program that gives pseudopolynomial-time solvability for a constant number of agents, polynomial-time algorithms in further restricted settings, and an FPTAS for fixed n under identical integer valuations. We then study a sequencing variant where the algorithm may choose the arrival order. This variant remains NP-complete even for two agents with identical valuations; however, a simple greedy algorithm achieves a 3/2 -approximation for n=2 agents, an n/(n-1) -approximation for any number of agents, and an additive guarantee depending on the maximum value of any good.

[MA-6] AGAR: a reinforcement learning substrate for LLM program evolution

【速读】:该论文旨在解决生成式程序演化中缺乏可解释性与可优化性的核心问题,具体表现为:在基于语言模型的算法发现过程中,搜索循环依赖于人工设定的五个超参数(如父代选择策略、突变强度、多样性保持机制、记忆机制及全局评分),导致系统难以进行有效调优与机制解耦。其解决方案的关键在于将程序演化过程形式化为一个马尔可夫决策过程(Markov Decision Process, MDP),其中动作定义为模型所条件化的模块化前缀(modular prefix)而非生成的完整程序,从而实现对信用分配(credit assignment)、价值估计(value estimation)、自适应探索(adaptive exploration)和经验记忆(experience memory)等组件的独立建模与替换。基于此框架构建的AGAR(Algorithm Generation As RL)系统允许任意替换或关闭特定估计器而无需修改控制器,且不需对后端语言模型进行梯度训练,实现了机制级的可审计迁移。实验表明,在19个任务、两种后端模型及三个随机种子下,AGAR在多数任务上优于现有最强基线,尤其在编程竞赛类任务中表现显著提升。此外,该形式化还揭示了以往方法隐含采用零折扣率(zero-discount)的本质原因——因适应度函数是外生的,而非基于后续状态的累积回报,因而无法对折扣因子产生作用。

链接: https://arxiv.org/abs/2610.09215
作者: Haoran Li,Zengle Ge,Xiaomin Yuan,Yui Lo,Haoxin Li,Songlin Zhou,Qianhui Liu,Jiahua Ying,Yuanhang Liu,Mingju Chen,Annan Li,Jianmin Wu,Dawei Yin,Dou Shen
机构: University of the Chinese Academy of Sciences(中国科学院大学); Baidu(百度); Tsinghua University(清华大学); Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所); Shanghai University(上海大学)
类目: Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA)
备注:

点击查看摘要

Abstract:Given a task and an evaluator, a language model can rewrite a candidate program while a search loop decides which rewrites survive, offering a practical route to algorithm discovery. But that loop is governed by five constants set by hand: which parent to select, how hard to mutate, how to keep diversity, what to remember, and a scalar score that never says which part of the program earned it. Reinforcement learning already has an estimator for each. The obstacle is that program evolution is not usually written down as a decision process. We formalize it as a Markov decision process whose action is the modular prefix the model is conditioned on, rather than the program it emits. Credit assignment, value estimation, adaptive exploration, and experience memory can then attach to distinct components. AGAR (Algorithm Generation As RL) provides the resulting substrate: any estimator can be replaced or switched off without changing the controller, making the transfer auditable one mechanism at a time, with no gradient training of the backend model. Across 19 tasks, two backends, and three seeds under one harness, AGAR improves on the stronger of two published baselines on most tasks, with gains concentrated in the competitive-programming family. The formalization also yields a checkable reading of prior work: these systems are implicitly zero-discount, not by choice, but because fitness is exogenous to an individual rather than a return over successors, leaving a discount factor nothing to act on.

[MA-7] SwarmReconGuard: Black-Box Detection of Distributed Collective Reconnaissance by Individually Benign-Looking Agent Populations

【速读】:该论文旨在解决分布式集体侦察(Distributed Collective Reconnaissance, DCR)这一新型威胁,即攻击者利用大量虚拟身份(virtual identities)以低频、看似正常的请求形式分散进行系统探测,从而在不触发传统检测机制的前提下逐步积累对目标系统的全面认知。其核心挑战在于:现有防御机制难以识别这种“低速率、分布式的良性行为聚合”所构成的隐蔽侦察活动。解决方案的关键在于构建一个可复现的黑盒基准测试框架SwarmReconGuard,通过容器化隔离环境模拟10至10,000个虚拟身份的协同行为,覆盖440次测试运行及超过366万次请求,确保完整的遥测数据完整性。研究对比了包括语义、高斯、条件、图、核函数、混合型及基于CUSUM的多种检测方法,发现高斯似然比检测虽在已知攻击上实现100%检出率且零误报,但在未知策略下仅达3%;而混合型CUSUM方法在10,000个身份规模下实现85.7%检出率且无观测误报,展现出更强的泛化能力。研究揭示了当前防御机制在策略泛化性上的显著缺陷,进而提出需发展具备暴露感知(exposure-aware)与规模感知(scale-aware)能力的新一代防御体系。

链接: https://arxiv.org/abs/2610.09138
作者: Vahid Tavakkoli,Kabeh Mohsenzadegan,Kyandoghere Kyamakya
机构: Institute for Smart System Technologies, University of Klagenfurt (克劳恩弗特大学智能系统技术研究所); University of Klagenfurt (克劳恩弗特大学)
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA)
备注:

点击查看摘要

Abstract:Autonomous and agentic clients can distribute reconnaissance across many identities so that each request remains valid, low-rate, and benign-looking while the population collectively acquires broad system knowledge. We formalize this threat as Distributed Collective Reconnaissance (DCR) and present SwarmReconGuard, a reproducible black-box benchmark in which the defender observes only service-boundary telemetry. The Docker-isolated study evaluates 11 benign and attack behaviors across 10-10,000 virtual identities, comprising 440 test runs and 3,666,300 requests, with complete telemetry integrity. We compare semantic, Gaussian, conditional, graph, kernel, hybrid, and CUSUM-based detectors. Gaussian likelihood-ratio detection achieves 100 % detection with 0 % observed false positives on known attacks but only 3 % on unseen policies. CUSUM yields 36.1 % overall detection at 1.25 % false positives, while hybrid CUSUM reaches 85.7 % detection with 0 % observed false positives at 10,000 identities. Results expose a major policy-generalization gap and motivate exposure-aware, scale-aware defenses.

[MA-8] When the Governor Becomes the Disturbance: Control-Generated Disturbance and Cost-Aware Backoff in Governed Tool-Using Agents

【速读】:该论文旨在解决监督型调控器(supervisory governor)在干预工具使用代理时可能引发的负面效应,特别是当监管强度增加导致实验性工具故障时,调控器可能将这些故障转化为持续性阻塞,从而阻碍任务完成。其核心问题在于:在缺乏对干预成本敏感性的调控机制下,调控行为可能因过度干预而降低系统整体效能。解决方案的关键在于引入一种基于移动平均的自适应退避规则(adaptive backoff rule),该规则通过监测已知诱发事件的历史频率动态调整干预概率,从而在保持与固定弱调控器相近干预频次的前提下,有效缓解持续性动作阻塞。实验结果表明,该方法在手写随机策略代理和Gemini 2.5 Flash大模型的多任务场景中均显著降低了阻塞发生率并提升了任务完成率,其中中等强度的退避策略表现最优。研究揭示了干预成本与持久性动作阻塞之间的交互关系,并提出了一种可量化的缓解机制,尽管其在更复杂开放环境中的适用性仍需进一步实证验证。

链接: https://arxiv.org/abs/2610.09037
作者: Veronique Ziegler
机构: Independent Researcher, Critical Attention Systems
类目: Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA)
备注: Includes ancillary experimental data and a Python script for reproducing numerical summaries. Code: this https URL

点击查看摘要

Abstract:Supervisory governors can interfere with the tool-using agents they regulate. We study this possibility in a controlled file-recovery environment where increases in regulatory intensity trigger experimentally imposed tool failures. A cost-blind governor can turn these failures into persistent blocking that prevents task completion. We compare this governor with a backoff rule that reduces intervention probability using a moving average of known induced events. On a hand-coded stochastic-policy agent, the failure pattern appears under both result replacement and execution of corrupted tool arguments. For the persistent policy, adaptive backoff improves completion relative to a fixed weak governor with approximately matched intervention frequency. A Gemini 2.5 Flash experiment comprising 576 episodes across 6 tasks also shows reduced blocking and improved completion under backoff; among the tested settings, intermediate backoff strength achieves the highest observed aggregate success. These results identify an interaction between intervention cost and persistent action blocking, together with a possible mitigation. The cost mechanisms are imposed and their induced events are directly observable to the backoff rule; applicability beyond this controlled environment remains an empirical question.

[MA-9] Shared-Roadmap Generation and Evaluator for Multi-Agent Path Planning Using Heterogeneous Graph Neural Network

【速读】:该论文旨在解决连续环境中多智能体路径规划(Multi-agent Path Planning, MAPP)在使用道路图(roadmap)时面临的图密度与可解性及解质量之间的权衡问题。传统方法如栅格网格或标准采样策略难以兼顾搜索效率与安全路径的发现,导致规划性能受限。本文提出一种可扩展的异构图神经网络(Heterogeneous Graph Neural Network, GNN)框架,用于自动化生成与评估共享多智能体道路图。其关键在于将航点、智能体位置和任务目标分别建模为异构图中的不同节点,从而实现对全局连通性与智能体间交互关系的联合推理;通过在由专家求解器轨迹聚合得到的占用密度图上进行训练,模型能够学习识别关键兴趣点,并自动剔除冗余节点与边,生成紧凑且具备协调意识的道路图,该道路图具有任务排列不变性,适用于多智能体取送任务的重复使用。实验表明,该方法显著降低规划开销,在密集道路图场景下可实现至少40%的运行时间与图规模缩减,同时提升解的质量。

链接: https://arxiv.org/abs/2610.09034
作者: Brandon Ho,Nikola Rogers,Seung-Kyum Choi
机构: Institute of Robotics and Intelligent Machines (IRIM), Georgia Institute of Technology (佐治亚理工学院)
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Multiagent Systems (cs.MA); Robotics (cs.RO)
备注:

点击查看摘要

Abstract:Multi-agent path planning (MAPP) in continuous environments often relies on roadmaps to balance safety and search efficiency. However, traditional roadmap generation methods, such as lattice grids or standard sampling-based approaches, frequently face a trade-off between graph density and the likelihood of finding feasible, high-quality solutions. In this paper, we propose a scalable heterogeneous Graph Neural Network (GNN) framework for the automated generation and evaluation of shared multi-agent roadmaps. Our model covers the representation of waypoints, agent locations, and task locations as distinct nodes in a heterogeneous graph, allowing it to reason over global connectivity and inter-agent interactions. By training on occupation density maps aggregated and collected from expert solver trajectories, the GNN learns to identify critical points of interest and prune redundant nodes and edges. This process produces a compact, coordination-aware roadmap that is invariant to task permutations and is reusable for multi-agent pick and delivery tasks. Experimental results demonstrate that our framework can reduce planning effort and can potentially find better solutions, reaching at least 40% reduction in runtime and in graph size for dense roadmaps.

[MA-10] Learning to Report Unsafe Tasks in a Multi-Agent Game

【速读】:该论文旨在解决多智能体系统中因任务奖励共享而导致的“安全报告激励不足”问题:当多个见证者(witness)共同分享任务完成后的奖励时,主动报告不安全行为会终止任务并降低自身收益,从而导致智能体倾向于保持沉默。为解决此问题,论文提出通过审计(audit)机制设计来引导智能体学习在发现不安全行为时进行报告的策略。其解决方案的关键在于构建一个可验证的审计条件,该条件确保策略梯度更新(policy-gradient updates)能够精确地收敛至全量报告(universal reporting)状态;尤其在接近普遍沉默(universal silence)的区域,该条件不仅充分且必要。研究进一步表明,在对称群体结构下,满足特定正边际要求的最小成本审计方案,其代价恰好是每个角色独立策略时的 kk 倍(kk 为每项任务的见证者数量)。实验基于24个从学习沉默行为中生成的见证者图结构,采用PPO算法训练,结果表明:将共见证者(co-witnesses)分组分离可使不安全完成率降低33.59个百分点(95%置信区间:21.03–45.13),显著优于随机分组;而在相同审计预算下,完全共享策略的网络在独立动作采样场景中无一达成低于1%不安全完成率且保持至少90%合法完成率的目标,而在共享动作采样下则全部达标,凸显了策略共享与动作相关性对安全性能的关键影响。

链接: https://arxiv.org/abs/2610.09002
作者: Avyay M. Casheekar,Hariganesh Tangirala
机构: University of Michigan Law School (密歇根大学法学院); School of Computing, National University of Singapore (新加坡国立大学计算机学院)
类目: Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA)
备注:

点击查看摘要

Abstract:When agents share a reward for completed tasks, reporting unsafe work can reduce the reporter’s reward by stopping a task. Audits can make reporting optimal without ensuring that further training teaches a silent team to report. We study this learning problem in a game where any witness can stop a task by reporting. With k witnesses per task sharing a policy and drawing independently, the expected-reward derivative with respect to their shared silence probability counts each task’s benefit k times at universal silence. The comparison with universal reporting counts it once. For arbitrary policy groups, we give an audit condition sufficient for exact policy-gradient updates to reach universal reporting and, apart from boundary cases, necessary near universal silence. In a balanced family, the cheapest audits meeting the condition with prescribed positive margins cost exactly k times as much for full sharing as for one policy per role. We train PPO policies on 24 witness graphs from learned silence. Separating co-witnesses reduces unsafe completion by 33.59 percentage points compared with shuffled groups of the same sizes under the same audits (95% graph-bootstrap interval: 21.03-45.13). Only 9 of 48 witness-group runs achieve below 1% unsafe completion while retaining at least 90% legitimate completion. At the same audit budget, a fully shared network meets both thresholds in none of 48 runs with independent action draws and all 48 with a common draw.

[MA-11] Can AI Agents Make Open-Ended Scientific Discovery? Evidence from Station

【速读】:该论文旨在解决生成式 AI 在缺乏明确评估指标的情况下,能否自主开展开放式科学发现(open-ended scientific discovery)这一关键问题。其核心挑战在于:在没有预设目标或阶段性奖励信号的复杂、开放环境中,如何维持持续探索动力并实现有效科研进展。解决方案的关键在于提出并集成两种新机制——监督者机制(Supervisor mechanism)与周期性元反思(periodic Meta Reflection),前者通过动态引导和约束行为确保探索方向合理,后者则通过定期自我评估与策略调整增强研究的连贯性与覆盖度。实验基于三篇最新发表于ICLR的论文构建开放式任务,让代理在不访问原始结果且禁用网络的前提下重现实验发现。结果显示,Station在平均62.7%的独立判据上实现了重发现,显著优于Codex Multiagent-v2(15.4%)和AI Scientist-v2(14.4%-20.6%)。消融实验与行为分析进一步证实,两机制协同作用能有效提升研究覆盖范围与连续性。此外,在无参考论文的额外任务中,部分代理发现与知识截止日期后真实研究结果高度吻合,表明该框架具备支持自主、有意义科学发现的潜力。

链接: https://arxiv.org/abs/2610.08927
作者: Wenyu Du,Stephen Chung
机构: DualverseAI; University of Cambridge(剑桥大学)
类目: Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA)
备注: preprint

点击查看摘要

Abstract:Recent AI systems have made rapid progress in scientific discovery when given well-defined metrics, but whether they can autonomously undertake open-ended scientific discovery remains unclear. We investigate AI’s ability to tackle open-ended tasks in Station, an open-world environment in which multiple agents simulate a scientific ecosystem. To tackle challenges specific to open-ended tasks, we propose augmenting Station with two mechanisms: a Supervisor mechanism and periodic Meta Reflection, which encourage persistent exploration even when intermediate metrics are lacking. We construct open-ended tasks from three recent oral papers presented at ICLR. We give agents the main research question studied in each paper while withholding the paper’s results and disabling web access. We then measure how many of the original findings-partitioned into individual criteria-agents rediscover. We find that Station rediscovers 62.7% of the criteria on average, compared with 15.4% for Codex Multiagent-v2 and 14.4-20.6% for AI Scientist-v2. Ablation and behavioral analyses indicate that adding the two mechanisms together improves research coverage and continuity. We further evaluate Station on two open-ended tasks without oracle papers and find that some of the discoveries made by the agents closely match discoveries reported by researchers after the knowledge cutoff date. Together, these results indicate that a suitable environment can enable agents to autonomously make meaningful progress in open-ended scientific discovery.

[MA-12] Adaptive Workflow Intelligence: A Cognitive Architecture for Context-Driven Enterprise Automation

【速读】:该论文旨在解决企业系统在非平稳环境、政策动态变化及操作反馈延迟等复杂条件下,现有基于人工智能(AI)的自动化工作流方案普遍存在的脆弱性问题。当前主流的强化学习与大语言模型(Large Language Model, LLM)代理虽具备一定自适应能力,但缺乏持续性的反思机制,且难以直接融入受策略约束的企业运营流程。为此,本文提出自适应工作流智能(Adaptive Workflow Intelligence, AWI),构建了一种以感知-认知-行动-反思(Perception-Cognition-Action-Reflection, PCAR)四层循环为核心的认知架构。其解决方案的关键在于将“反思”机制形式化为持续策略优化的驱动力,通过融合混合推理、可回溯的反思记忆以及基于反馈的适应性调整,实现对环境漂移和操作约束下的稳健决策支持。实验评估表明,在存在延迟反馈与制度突变的模拟企业工作流中,受安全护栏约束的自适应方法相较于静态自动化能更快恢复性能并保持合规性;同时,AWI的反思模块虽仅小幅降低行为振荡与反馈方差,却有效揭示了反思式策略适应所引入的稳定性与敏捷性之间的权衡关系。

链接: https://arxiv.org/abs/2610.08793
作者: Sreedevi Pandiyath Viswambaran
机构: 未知
类目: Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA)
备注: 12 pages, simulation-based evaluation of adaptive workflow agents

点击查看摘要

Abstract:Enterprise systems increasingly rely on automated workflows, yet many AI-driven solutions remain brittle under non-stationary conditions, evolving policies, and delayed operational feedback. While reinforcement learning and large language model (LLM) agents offer partial adaptability, they do not by themselves provide persistent reflection mechanisms or straightforward integration with policy-constrained enterprise operations. This paper introduces Adaptive Workflow Intelligence (AWI), a cognitive architecture for context-driven enterprise agents organized around a four-layer Perception-Cognition-Action-Reflection (PCAR) loop. AWI treats reflection as a mechanism for continuous policy refinement and combines hybrid reasoning with reflective memory and feedback-driven adaptation to support decision making under environmental drift and operational constraints. We evaluate AWI in a simulated enterprise decision workflow characterized by delayed outcomes and a controlled regime shift. In a drift-and-delay stress test, guardrail-constrained adaptive approaches recover more rapidly than static automation while maintaining policy compliance. Within this setting, AWI’s reflective components modestly reduce behavioral oscillation and feedback variance, illustrating the stability-agility trade-off introduced by reflective policy adaptation.

[MA-13] Decentralized SGD under Heavy-Tailed Noise: Optimal Convergence Rates and the Role of Gradient Clipping

【速读】:该论文旨在解决在存在重尾噪声(heavy-tailed noise)的去中心化非凸优化场景下,如何设计能够实现最优收敛速率的基准算法问题。在去中心化设置中,对局部梯度施加非线性操作(如梯度裁剪或归一化)会同时影响优化过程与共识一致性,而现有方法要么收敛速率次优(如裁剪法),要么依赖局部动量或小批量数据才能收敛(如归一化法)。为应对这一挑战,本文提出并分析了裁剪式去中心化随机梯度下降(clipped decentralized SGD, \mathtt{DSGD}),证明其在满足有界 $ p −阶矩噪声(-阶矩噪声( p \in (1,2] $)条件下,对于光滑非凸目标函数,能够在高概率和期望意义下达到阶最优的收敛速率。其关键技术创新在于对共识误差(consensus gap)的精细分析,充分利用了裁剪操作的结构特性,将网络相关影响降为高阶项,从而实现了在无额外假设下的线性加速(linear speed-up)。研究揭示了去中心化设置中裁剪与归一化的重要区别:归一化方法可能因丢失梯度幅值信息而导致不收敛,而裁剪通过保留幅度信息,保障了算法的收敛性与最优性。数值实验验证了理论结果的有效性。

链接: https://arxiv.org/abs/2610.10527
作者: Aleksandar Armacki,Haoyuan Cai,Ali H. Sayed
机构: École Polytechnique Fédérale de Lausanne, Lausanne, Switzerland
类目: Optimization and Control (math.OC); Machine Learning (cs.LG); Multiagent Systems (cs.MA)
备注: 36 pages, 5 figures, 2 tables

点击查看摘要

Abstract:Heavy-tailed noise has been widely observed in modern machine learning, motivating the use of methods like gradient clipping and normalization. While these methods are well understood in centralized settings, much less is known in decentralized ones, where applying a nonlinearity to local gradients affects both optimization and consensus. Recent works on decentralized non-convex optimization have studied both clipping and normalization under heavy-tailed noise, with clipping yielding suboptimal rates and normalization needing local momentum or mini-batches to converge. This raises the question: can a baseline decentralized method using a nonlinearity achieve optimal convergence rates under heavy-tailed noise? We answer affirmatively with clipped decentralized SGD ( \mathttDSGD ). For smooth non-convex costs under bounded p -th moment noise, p \in (1,2] , we show that clipped \mathttDSGD achieves order-optimal rates both with high probability and in expectation. Moreover, we establish a linear speed-up in the number of agents, which, to our knowledge, has not been shown for decentralized methods with clipping. The key technical ingredient is a sharp analysis of the consensus gap that exploits the structure of clipping, relegating network effects to higher-order terms. Our results highlight an important distinction between clipping and normalization in decentralized settings: while normalized \mathttDSGD can fail to converge, clipping retains magnitude information, enabling \mathttDSGD to be convergent and order-optimal. Numerical experiments validate our theory.

自然语言处理

[NLP-0] Decoupling Exploration from Optimization in RLVR

【速读】: 该论文旨在解决生成式语言模型在强化学习与可验证奖励(Reinforcement Learning with Verifiable Rewards, RLVR)框架下,因引入强新颖性激励而导致模型性能退化的问题。尽管RLVR理论上具备发现新推理策略的潜力,但由于可验证奖励仅监督模型知识与行为的一小部分,过度追求新颖性易引发不可逆的质量下降。为此,本文提出一种名为探索-蒸馏(Exploration-Distillation, ExpDis)的解耦框架:将探索与优化过程分离,通过带有新颖性奖励的探索者策略(explorer policy)生成多样化轨迹,经正确性与质量筛选后,将其知识蒸馏至独立的学生策略(student policy)中,而学生策略在后续训练中不再包含新颖性奖励。该过程可多轮迭代,实现激进探索而不损害最终模型质量。实验表明,在七个数学推理基准和两种模型架构上,ExpDis在相同计算预算下优于DAPO方法,并展现出更优的pass@k缩放性能,证明其能生成更多样且正确的解法。

链接: https://arxiv.org/abs/2610.10536
作者: Saif Punjwani,Micah Goldblum
机构: Columbia University(哥伦比亚大学)
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: 20 pages, 16 figures, 9 tables. Code: this https URL . Checkpoints: this https URL

点击查看摘要

Abstract:Modern language models undergo reinforcement learning with verifiable rewards (RLVR) on top of already-trained checkpoints. A key promise of RLVR is the discovery of new reasoning strategies. In principle, a model can sample novel ideas absent from its prior training data. In practice, however, augmenting RLVR with strong novelty incentives has seen limited success and can degrade model quality. Because verifiable rewards supervise only a narrow slice of the model’s knowledge and behavior, such degradations are difficult to recover from. Instead, we decouple exploration from optimization in a framework we call Exploration-Distillation (ExpDis). We train one or more explorer policies with a novelty bonus in the reward, filter their trajectories for correctness and quality, and distill them into a separate student policy. The student policy is then trained without a novelty bonus. We repeat the above procedure for several rounds, alternating between exploration and optimization. This decoupling allows us to aggressively scale exploration without degrading the student policy. Across seven mathematical reasoning benchmarks and two model families, ExpDis outperforms DAPO at the same wall-clock budget. Moreover, we observe improved pass@ k scaling, indicating that ExpDis produces models that generate more diverse correct solutions.

[NLP-1] EngramEdit: Decoupled Knowledge Updates in LLM s through Conditional Memory

【速读】: 该论文旨在解决大语言模型(LLM)在保持通用计算能力不变的前提下,如何高效、精准地更新事实性知识的问题。现有基于条件记忆架构(如DeepSeek Engram)的方法虽能扩展模型容量并实现知识存储与通用计算的解耦,但面临关键挑战:同一事实的不同表达形式可能激活不同的n-gram嵌入,而对共享嵌入的更新可能无意中影响其他无关事实的预测准确性。为此,论文提出EngramEdit,其核心在于通过条件记忆实现解耦的知识更新。该方案首先为待更新的事实生成跨表达形式的目标记忆表示,使模型在多种表述下均能正确预测新知识;随后联合优化共享n-gram嵌入以匹配这些目标表示,并对频繁复用的嵌入施加更强的惩罚,以保护无关知识不被破坏。实验表明,EngramEdit可实现独立于通用能力的事实知识更新,在多表达泛化和多跳推理任务中表现优异,其链式思维(CoT)提示下的准确率接近最强基线的三倍,且长期累积更新下仍能有效保留未受影响的知识与通用能力。这表明EngramEdit成功将条件记忆转化为可编辑的知识接口,显著拓展了其在知识动态维护中的应用价值。

链接: https://arxiv.org/abs/2610.10533
作者: Hongru Cai,Ran Wei,Wenjie Wang,Chengfa Wu,Ning Song,Yongqi Li,Wenjie Li
机构: The Hong Kong Polytechnic University(香港理工大学); Hangzhou Diagens Biotechnology Co., Ltd(杭州迪根斯生物科技有限公司); University of Science and Technology of China(中国科学技术大学)
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Conditional memory architectures such as DeepSeek Engram use input n-grams to look up learned embeddings, expanding the capacity of large language models (LLMs) with limited additional computation. Beyond model scaling, this architecture has demonstrated the potential to decouple factual knowledge storage from general-purpose computation, offering a promising route to updating factual knowledge while keeping the Transformer backbone fixed. Realizing this potential is challenging because different expressions of a fact may activate different n-gram embeddings, while updating shared embeddings can unintentionally change the model’s predictions about other facts. We propose EngramEdit for decoupled knowledge updates through conditional memory. EngramEdit first computes target memory representations that make the model predict the updated fact across multiple expressions. It then jointly updates the shared n-gram embeddings to match these targets across expressions and edits, penalizing updates to frequently reused embeddings more strongly to preserve unrelated knowledge. Experiments show that EngramEdit enables independent factual knowledge updates through conditional memory, achieving near-perfect editing success. Revised knowledge is usable across unseen expressions and in multi-hop reasoning, with nearly three times the strongest baseline’s accuracy under chain-of-thought (CoT) prompting. Unrelated knowledge and general capabilities are largely preserved even as factual updates accumulate. These findings show that EngramEdit turns conditional memory into an editable knowledge interface, extending its role beyond model scaling to support decoupled knowledge updates.

[NLP-2] Rephrase Before You Act: Characterizing and Mitigating Language Sensitivity in Vision-Language-Action Models

【速读】: 该论文旨在解决视觉-语言-动作模型(Vision-Language-Action models, VLAs)对指令表述高度敏感的问题,即同一任务因指令微小差异(如单个词汇变更)导致性能波动达数十个百分点,且无法继承其基础视觉-语言模型(Vision-Language Models, VLMs)在语言鲁棒性方面的优势。其核心问题是:现有VLAs在面对分布外(out-of-distribution)或对抗性指令时表现显著下降,而这种敏感性源于模型对自然语言表达的不稳定性。解决方案的关键在于不修改策略网络(policy),而是通过构建一套可解释的、基于规则的指令重写机制:首先在少量训练任务上收集大量指令变体,利用大语言模型(Large Language Model, LLM)从这些变体中提炼出10至20条通用的重写规则;在推理阶段,将输入指令依据这些规则进行一次性的语义等价重写,从而提升指令的鲁棒性。该方法无需重新训练、无需每步验证,且可零样本泛化至未见任务与指令,在多个基准测试中显著提升了冻结策略π₀在对抗性、VLM生成及人类生成指令下的表现,尤其在分布外任务上增益集中,相对提升达16%–27%,并在π₀.₅和LIBERO数据集上将微调后成功率从93.6%提升至97.8%。

链接: https://arxiv.org/abs/2610.10526
作者: Mikey Watts(Independent Researcher),Yuchen Cui(University of California, Los Angeles)
机构: University of California, Los Angeles(加州大学洛杉矶分校)
类目: Robotics (cs.RO); Computation and Language (cs.CL); Machine Learning (cs.LG)
备注: 9 pages, 8 figures, 3 tables. Project page: this https URL

点击查看摘要

Abstract:Vision-language-action models (VLAs) are strikingly sensitive to instruction phrasing and do not inherit the language robustness of the vision-language models they are built on. A one-word edit can move success by tens of points: \pi_0.5 turns on a LIBERO stove 100% of the time for “switch on the stove” and 2% for “switch on the hot plate”, and a \pi_0 checkpoint finetuned with rephrase augmentation still shows swings of up to 61 points. We characterize this sensitivity with statistically tested single-edit swings and an oracle phrase search, which shows that phrasing alone nearly closes the 21-point gap between in-distribution and out-of-distribution tasks. We then reduce it without modifying the policy. Because the sensitivity is systematic, it can be expressed as explicit rules: we score many phrasings of a few training tasks, have a large language model distill the evidence into ten to twenty rephrasing rules, and at deployment rewrite each incoming instruction once under these rules. The rules improve the frozen \pi_0 by 16 to 27% relative on twelve held-out tasks across adversarial, VLM-generated, and human-generated phrasings, with gains concentrated on out-of-distribution tasks. The pipeline replicates on \pi_0.5 and LIBERO, lifting in-finetune success from 93.6% to 97.8%. The method requires no retraining and no per-step verification, and applies zero-shot to unseen tasks and instructions. Project website: this https URL

[NLP-3] Your Prompt Should Do More: Effects of Retrieval Instructions in Embedding Models

【速读】: 该论文旨在解决当前生成式嵌入模型(Prompted embedding models)在信息检索任务中难以可靠遵循详细检索指令的问题,尤其关注在非对称检索场景下,查询端干扰项(query-side distractors)对模型表现的负面影响。研究发现,即使面对简单的任务指令,当评估中引入查询端干扰项时,现有模型仍可能失效。其关键解决方案在于重新审视并优化模型的训练与评估范式:通过在微调阶段主动引入查询端干扰项进行训练,显著提升了模型对指令的遵循能力,且对其他下游任务影响极小,表明该方法能够有效增强模型在复杂检索场景下的鲁棒性与指令理解能力。

链接: https://arxiv.org/abs/2610.10508
作者: Amanda Myntti,Jenna Kanerva,Veronika Laippala,Filip Ginter
机构: University of Turku (图尔库大学); Ellis Institute Finland
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Prompted embedding models have recently received increasing attention, particularly for retrieval, where detailed retrieval instructions are provided as part of the retrieval prompt. Several new datasets and studies have examined this setting, showing that the current embedding models often struggle to follow such instructions reliably. In this paper, we study the mechanism of how instructions actually affect the representations of retrieval queries in asymmetric retrieval tasks. We show that models can fail to follow even simple task instructions when query-side distractors are included in the evaluation. We hypothesize that this behavior is driven by the training setup of current embedding models and their evaluation, and show that fine-tuning with added query-side distractors leads to substantial improvements, with minimal effect on other tasks.

[NLP-4] Validity Without Ground Truth: What Stated-Preference Economics Offers the Evaluation of Language Models

【速读】: 该论文旨在解决当前大型语言模型(Large Language Models, LLMs)在面对无明确标准答案的主观性问题时,缺乏有效评估方法的核心挑战,例如政策价值判断、用户选择建议及多重价值权衡等。传统陈述偏好经济学(Stated-Preference Economics)长期面临类似困境——无法依赖真实效用值进行验证,因此发展出一套以有效性(validity)为核心的评估框架,包括内容效度、构念效度、准则效度、信度、激励相容性与后果性等概念。本文提出,这一框架可作为通用方法论用于评估语言模型输出的合理性与一致性。研究通过将已发表的水质经济估值调查(Vossler et al. 2023)应用于六种不同模型,验证该框架的适用性:在经济理论预测下,如需求曲线应向下倾斜、支付意愿应随商品范围和收入变化而调整,模型表现被系统性区分。结果显示,两个较旧模型在家庭年收入7.5万美元水平下即未能通过最基本的有效性测试,而两个最新模型均通过所有理论有效性检验,但在收敛效度上出现分歧。值得注意的是,通过有效性测试仅表明模型回答具备内在逻辑一致性,并不意味着其答案绝对正确。因此,该研究的关键在于引入并实证应用基于经济学验证框架的语言模型评估体系,为非客观问题下的模型可信度提供可操作的判别标准。

链接: https://arxiv.org/abs/2610.10506
作者: Daniel Robert Kling Alexander,Catherine Louise Kling
机构: University of Michigan School of Information (密歇根大学信息学院); Cornell University (康奈尔大学)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); General Economics (econ.GN)
备注:

点击查看摘要

Abstract:Many of the questions now put to large language models have no correct answer to score against: what a policy is worth, which option a user should choose, how to weigh competing values. Stated-preference economics has faced this problem for decades. It judges survey responses without knowing the true value, through a framework of validity and related concepts: content, construct, and criterion validity, reliability, incentive compatibility, and consequentiality. We argue that this framework is a general method for evaluating language models, and we set out what each concept means for LLM evaluation. We demonstrate the approach using a published water-quality stated preference economic valuation survey (Vossler et al. 2023) administered to six models. In this economic application, the validity tests take the form of predictions from economic theory: demand should slope down, and willingness to pay should respond to the scope of the good and to income. The tests separate the models sharply. Two older models fail the most basic test at a household income level of \ 75,000, and the two newest pass every test of theoretical validity we can score, but diverge on convergent validity. Passing validity tests shows that a model’s answers are coherent, not that they are correct.

[NLP-5] PHRBench: A Behavioral Evaluation of Post-Hallucination Reasoning in LLM s

【速读】: 该论文旨在解决多阶段大语言模型(Large Language Model, LLM)系统中幻觉信息(hallucinated information)传播及其对后续推理影响的问题,尤其关注模型在生成幻觉前提后如何修正错误推理路径的行为机制。现有研究多聚焦于最终结果的变化与推理动态的总体聚合分析,而对模型在响应层面如何处理和纠正幻觉前提缺乏深入理解。为此,本文提出PHRBench——一个跨四个领域、涵盖18个大型语言模型的受控基准测试框架,通过“幻觉合规性”(Hallucination Compliance)、“幻觉规避性”(Hallucination Avoidance)与“启发式修正”(Heuristic Correction)三个维度,独立于最终答案正确性地刻画每个推理轨迹的行为特征,并将成功修正至正确答案的轨迹定义为有洞察力的恢复行为。在4820个受控实例上的实验表明,成功恢复仍相对罕见,且与推理过程中更频繁的信念更新密切相关;同时,幻觉提示本身的属性蕴含显著的预测信号,仅使用轻量级预测器即可达到0.847的AUROC值。该研究从行为视角揭示了大语言模型在面对错误上下文时的纠错机制及成功恢复的潜在条件,为提升模型鲁棒性提供了关键洞见。

链接: https://arxiv.org/abs/2610.10455
作者: Linghao Meng,Feng He,Xuan Yang,Junyuan Mao,Pinze Ren,Deqing Mu,Hesen Yang,Qiankun Li
机构: National University of Singapore(新加坡国立大学); Tsinghua University(清华大学); Johns Hopkins University(约翰霍普金斯大学); Nanyang Technological University(南洋理工大学)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Hallucinated information can propagate through multi-stage LLM systems and become part of the context for subsequent reasoning. Existing studies of post-hallucination reasoning (PHR) mainly characterize changes in final outcomes and aggregate reasoning dynamics, leaving how models resolve hallucinated premises at the response level insufficiently understood. In this work, we introduce PHRBench, a controlled benchmark for behaviorally structured PHR across four domains and 18 large language models. PHRBench characterizes each reasoning trajectory independently of final-answer correctness through Hallucination Compliance, Hallucination Avoidance, and Heuristic Correction, and defines an insightful trajectory as successful correction that ultimately reaches the correct answer. Across 4820 controlled instances, we find that successful recovery remains relatively rare and is associated with more frequent belief updates along the reasoning trajectory. We further find that properties of the hallucinated prompt contain substantial predictive signal for successful recovery, with a lightweight predictor achieving an AUROC of 0.847. These findings provide a behavioral view of post-hallucination reasoning, characterizing how LLMs resolve erroneous context and when successful recovery is likely to occur.

[NLP-6] RunningTab: Direct Workspace Interaction with Environment-Side Tabs

【速读】: 该论文旨在解决生成式AI(Generative AI)在直接工作区交互(Direct Workspace Interaction, DWI)过程中因上下文窗口限制导致的任务状态丢失问题。具体而言,在无需索引的文件直接访问模式下,尽管智能体可读取工作区中的任意文件,但其无法持续追踪任务需求、已读文件与待读文件之间的关联,从而可能导致关键信息遗漏(如未包含应提取的图表)。为此,论文提出RunningTab框架,其核心创新在于引入环境侧的“任务标签页”(per-task record),由环境端维护一个结构化记录:一方面,智能体声明任务需求;另一方面,环境自动记录每个已读文件的摘录及其来源(provenance),以及被列出但未打开的文件作为候选项。该标签页使智能体能够实时查看每项需求与其最匹配的摘录和未打开候选文件,并据此进行内容匹配或标注暂搁原因;当尝试完成任务时,系统将执行最终检查,确保所有需求均已处理。实验在三个基准测试上验证了RunningTab的有效性,结果表明其显著优于纯DWI及将记录置于模型内部的基线方法,且标签页通常能准确保留交付物所需的关键信息。

链接: https://arxiv.org/abs/2610.10444
作者: Jinheon Baek,Soyeong Jeong,Yumin Choi,Dongsu Han,Sung Ju Hwang
机构: KAIST(韩国科学技术院); DeepAuto.ai
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Much knowledge work produces new deliverables from files a workspace already holds, and LLM agents are beginning to take such work over. Through direct corpus interaction, an agent can search and read any of those files from a terminal with no indexing, and producing a deliverable from many of them in this way is what we call direct workspace interaction (DWI). Reaching the files, however, is only half the task: nothing keeps track of what the task asks for, what has been read, and what was listed but never opened, all of which slip through the context window without leaving a trace, so an agent may extract a figure and still deliver a report without it. To address this, we present RunningTab, a framework that equips direct workspace interaction with an environment-side tab: a per-task record of what the task still owes, kept by the environment alongside the agent. Specifically, the agent adds its requirements, while the environment records every file read as an excerpt with its provenance and every listed but unopened file as a candidate; the agent can then see each requirement beside its best-matching excerpts and top unopened candidates, resolve it against matching content or set it aside with a reason, and, should it try to finish with requirements still open, receive them in a finish check. We validate RunningTab on three benchmarks with three LLMs, where it consistently outperforms plain DWI and baselines that keep the record in the model, while its tab usually holds the values a deliverable needs once seen.

[NLP-7] CoTrace: Data Recipes for Training Terminal Agents with Harness-Model Co-Evolution

【速读】: 该论文旨在解决生成式智能体(generative agent)在实际应用中因模型权重与运行时环境(runtime harness)之间耦合不当而导致的性能瓶颈问题。现有方法虽在模型与运行时框架协同进化方面取得进展,但普遍将搜索过程中产生的轨迹视为无差别的回放缓冲区,忽视了轨迹对模型训练的价值高度依赖于其生成时所采用的运行时配置。为此,论文提出一种分层交替协同进化框架,通过组件级的促进决策解耦运行时搜索与策略训练过程。在此框架下,提出CoTrace——一种面向运行时感知的数据处理机制,能够显式控制轨迹路由、来源匹配与课程更新。CoTrace使重复执行失败成为引导运行时合成的关键信号,同时确保策略训练仅基于与当前采纳运行时相匹配的验证轨迹,适用于监督微调(SFT)或在线强化学习(RL)。实验表明,在Tmax任务划分上,CoTrace使Qwen3.5-9B模型在SFT下的可解决任务数从78提升至88,而在线强化学习变体达到90。更重要的是,使用与运行时匹配的小型数据集即可实现稳定模型提升,所需计算资源远低于跨多个运行时合并的大规模数据集。此外,在Terminal-Bench 2.1和SWE-bench Lite上的评估显示,模型在分布外任务上的迁移能力本质上取决于运行时兼容性,保持训练与推理阶段运行时一致性可有效避免因运行时不匹配引发的程序执行崩溃。

链接: https://arxiv.org/abs/2610.10426
作者: Jixuan Chen,Jiaxin Zhang,Qinyuan Ye,Yada Pruksachatkun,Haoxiang Zhang,Jingming Zhuo,Yifan Zhang,Yutong Dai,Juntao Tan,Xiangyu Peng,Silvio Savarese,Zeyuan Chen,Lianhui Qin,Chien-Sheng Wu
机构: University of California, San Diego(加州大学圣地亚哥分校); Salesforce AI Research( Salesforce人工智能研究); University of Washington(华盛顿大学)
类目: Computation and Language (cs.CL)
备注: Preprint. 32 pages, 7 figures, 17 tables

点击查看摘要

Abstract:Terminal-agent capability depends jointly on model weights and the runtime harness that formats prompts, binds tools, and handles error recovery. Existing harness-model co-evolution approaches improve both components, yet often treat trajectories produced during harness search as an undifferentiated replay buffer. This practice overlooks that a trajectory’s value for model training depends on the harness under which it was generated. To systematically analyze this interface, we establish an alternating co-evolution framework that decouples harness search and policy training through component-wise promotion decisions. Within this framework, we introduce CoTrace, a harness-aware data recipe that explicitly governs trajectory routing, provenance matching, and curriculum refresh. Under CoTrace, recurring execution failures guide harness synthesis, while policy training is strictly conditioned on verified rollouts matched to the adopted runtime for supervised fine-tuning (SFT) or fresh online interactions for reinforcement learning (RL). On the Tmax promotion split, CoTrace advances Qwen3.5-9B from 78 to 88 solved tasks under supervised fine-tuning while an online reinforcement variant reaches 90. Specifically, a compact harness-matched corpus produces steady model gains at substantially lower compute than much larger corpora pooled across sibling harnesses. Furthermore, evaluations on Terminal-Bench 2.1 and SWE-bench Lite show that out-of-distribution transfer depends fundamentally on harness compatibility, where maintaining consistency between training and evaluation runtimes prevents procedural execution breakdowns observed under foreign scaffolds.

[NLP-8] Which Rollout Taught It That? BehaviorTrace and the Limits of Training-Data Attribution in Online RL

【速读】: 该论文旨在解决强化学习(Reinforcement Learning, RL)微调过程中,生成式人工智能模型行为学习的可追溯性问题,即在语言模型通过在线强化学习微调(如GRPO)习得新行为后,能否准确识别出真正促成该行为学习的训练轨迹(training rollouts)。同时,针对现有归因方法(attribution method)声称能够定位关键训练步骤的问题,论文进一步探究其结果的真实性与可靠性。解决方案的关键在于构建一个可控、可验证的评估框架——BehaviorTrace,该框架结合了全梯度草图(full-gradient sketching)、预设行为(planted behavior)机制以及对梯度幅度、流畅性、余量空间和种子及生成抽样变异性的严格控制。实验基于Qwen2.5-1.5B模型在三个随机种子上的结果表明,大部分看似有效的归因信号实则源于混淆因素;仅依赖梯度大小排序的无目标控制组即可达到基线性能的4.2至4.5倍,并在两个种子上超越最优有目标估计器。此外,在饱和检查点处,模型流畅性本身即可媲美甚至优于所有梯度归因方法对行为标签的预测能力。一旦控制流畅性,每条轨迹的归因表现随种子和生成抽样显著波动,表明单一运行无法得出可靠结论。唯一在所有三个种子上一致成立的信号是:触发词(trigger tokens)的梯度方向与行为实际发生位置的目标向量高度对齐。基于此,论文提出了一套用于评估RL中归因有效性的检查清单,用于检验现有归因方法(如GAS和TRAK风格估计器)的可信度,但未提出新的归因算法。

链接: https://arxiv.org/abs/2610.10422
作者: Amit Nautiyal
机构: Independent Researcher(独立研究员)
类目: Machine Learning (cs.LG); Computation and Language (cs.CL)
备注: 11 pages, 2 figures, 4 tables. Code and data: this https URL

点击查看摘要

Abstract:When reinforcement learning teaches a language model a new behavior, can we find the training rollouts that taught it? And when an attribution method says it can, how do we know the answer is real? We study both questions on online RL fine-tuning with GRPO, using a planted behavior with a known cause. We release BehaviorTrace, an open evaluation harness that combines full-gradient sketching, the planted-behavior setup, and controls for gradient magnitude, fluency, headroom, and variation across seeds and generation draws. Across three seeds on Qwen2.5-1.5B, much of the apparent attribution signal comes from confounds. A control that ranks training steps by gradient size alone, with no behavior target, reaches 4.2 to 4.5 times chance and matches or beats the best targeted estimator on two of three seeds. At saturated checkpoints, model fluency predicts the behavior label at least as well as every gradient method we compared it with. Once fluency is controlled, the per-rollout results change from seed to seed and from one generation draw to the next, so a single run cannot settle the question. One signal does hold on all three seeds. The gradient of the trigger tokens aligns with a target built where the behavior actually occurs. We turn these findings into a checklist for evaluating attribution in RL. We test existing estimators, including GAS (renormalized TracInCP) and a TRAK-style estimator, and do not propose a new one.

[NLP-9] raining Parallel Speculative Draft Models by Directly Minimizing Expected Decoding Rounds

【速读】: 该论文旨在解决生成式 AI(Generative AI)中基于推测性解码(speculative decoding)的低效训练问题。现有方法在训练半自回归(semi-autoregressive)或并行推测模型时,依赖局部代理目标函数,未能充分考虑不同解码轮次之间的耦合关系——即当前轮次的生成分布受此前轮次接受令牌数量的影响,从而导致训练目标与全局解码效率(如解码轮次数)不一致。为此,论文提出将推测性解码建模为马尔可夫奖励过程(Markov reward process),由此导出期望解码轮次(Expected Decoding Rounds, EDR)作为优化目标,该目标通过状态占用率对局部拒绝成本进行加权,精确对应于预期解码轮次数,且无需引入额外超参数。进一步,作者推导出一个无偏的时序差分梯度(temporal-difference gradient),支持从目标模型回滚轨迹中进行无偏随机优化。同时,该框架还提供了一个精确的离线评估器,可在共享目标模型回滚轨迹上直接比较不同推测模型的轮次性能,而无需实际执行推测解码。基于EDR对DSpark和DFly两个先进推测模型进行微调,在九个涵盖数学推理、代码生成和对话任务的基准测试中均显著提升平均采纳长度,并优于现有训练目标。

链接: https://arxiv.org/abs/2610.10411
作者: Yunxiao Zhao,Changxiao Cai
机构: University of Michigan (密歇根大学)
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (stat.ML)
备注:

点击查看摘要

Abstract:Speculative decoding accelerates large language model inference by using a low-cost draft model to propose tokens that the full-size target model verifies in parallel. Parallel and semi-autoregressive (semi- AR) drafters improve drafting efficiency by proposing an entire block in a single forward pass, but training them raises a new difficulty: the draft distribution for a given position depends on where the decoding round starts, and where rounds start depends on how many tokens earlier rounds accepted. Existing training objectives typically rely on block-local surrogates that ignore this cross-round coupling, and therefore do not directly optimize the global decoding efficiency. In this work, we develop a theoretical framework for training and evaluating these drafters by representing speculative decoding as a Markov reward process. This formulation yields the Expected Decoding Rounds (EDR) objective, which weights local rejection costs by state occupancies and exactly equals the expected number of decoding rounds. Unlike prior surrogate objectives, EDR introduces no auxiliary hyperparameters. We then derive an exact temporal-difference gradient that supports unbiased stochastic optimization from target-model rollouts. The same framework also yields an exact offline evaluator for round counts, enabling paired drafter comparisons on shared target rollouts without running speculative decoding. Finetuning two state-of-the- art drafters, DSpark and DFly, with EDR consistently improves mean accepted length and outperforms existing training objectives across nine benchmarks spanning math reasoning, code generation, and chat.

[NLP-10] Reasoning -Token Spikes Under Prompted Untruthful Responding in Large Language Models

【速读】: 该论文旨在解决生成式人工智能(Generative AI)在推理过程中存在欺骗行为或其它不当行为时,难以通过传统语义链式思维(chain-of-thought)监控手段进行有效检测的问题。现有方法依赖于推理轨迹的可读性、忠实性及可访问性,但随着模型演化,链式思维输出可能变得不可读或不忠实,导致监控失效。为此,研究基于认知负荷理论,提出一种低带宽信号——推理令牌数量(reasoning-token count),该信号无需访问推理内容即可获取,具备内容无关性与实现简便性。实验采用三类具备推理能力的大语言模型,在分析型、描述型与规范型推理任务以及道德与非道德领域中,分别在要求真实回答、虚假回答及无视真实性的指令下完成210道多选题。结果表明,无论模型类型如何,追求真实性的响应均产生显著更少的推理令牌,而虚假回应与无视真实性的回应则消耗更多令牌。这一发现验证了显式诱导的非真实性策略可在测试阶段引发可区分的群体级令牌使用差异,为在原始推理轨迹不可靠或不可用时,提供了一种无需内容分析的、简单有效的潜在信号,可用于区分真实与非真实模型行为。尽管尚未证明其能检测自发性欺骗或普遍偏差,但本研究为推理令牌计数作为探测机制提供了概念验证,未来工作需进一步检验实例级检测率、分布外泛化能力、学习到的欺骗策略、隐含目标及对抗压力下的鲁棒性。

链接: https://arxiv.org/abs/2610.10405
作者: Maverick Morales,Tomáš Dominik,Vermut Gao,Katrina Shirey,Paulius Rimkevičius,Aaron Schurger,Uri Maoz
机构: Chapman University (查普曼大学); INSERM U992, Cognitive Neuroimaging Unit, NeuroSpin Center (法国国家健康与医学研究所U992认知神经成像单元,神经成像中心); University of California, Los Angeles (加利福尼亚大学洛杉矶分校); California Institute of Technology (加州理工学院)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: 20 pages, 9 figures, 3 tables. Code: this https URL ; Data: this https URL

点击查看摘要

Abstract:Monitoring the chain-of-thought of reasoning artificial intelligence (AI) models remains a key approach to detecting deception and other forms of misbehavior in such models. However, semantic chain-of-thought monitoring depends on reasoning traces being legible and sufficiently faithful to the underlying computations that produced the model’s behavior, not to mention accessible. Moreover, there is increasing evidence that chain-of-thought outputs may soon become illegible or unfaithful, if they even remain accessible. Based on cognitive load theory, we investigate a lower-bandwidth signal – the number of reasoning tokens generated – which does not require access to the content of the reasoning trace. Three reasoning-capable large language models answered 210 multiple-choice questions – across analytic, descriptive, and normative reasoning types as well as moral and non-moral domains – under system prompts instructing them to respond truthfully, falsely, or without regard for truth. Across all three models, truth-directed responding elicited fewer reasoning tokens than both lie-directed and truth-indifferent responding. These findings show that explicitly prompted untruthful response policies can produce robust group-level differences in test-time reasoning-token use. While not yet establishing reasoning-token count as a detector of spontaneous deception or general misalignment, our results are a proof of concept that it can serve as a simple, content-independent candidate signal for differentiating untruthful from truthful model behavior when raw reasoning traces are unavailable or unreliable. Future work should test instance-level detection rates, out-of-distribution generalization, learned deceptive policies, hidden objectives, and robustness under adversarial pressure.

[NLP-11] Document-Level Text Simplification in Estonian Using Large Language Models LREC2026

【速读】: 该论文旨在解决在形态丰富的低资源语言(如爱沙尼亚语)中,文档级文本简化(document-level text simplification)缺乏系统研究与有效方法的问题。现有研究多集中于高资源语言的句子级简化,而忽视了跨句、跨段落层面的语篇连贯性、指代消解及上下文一致性等复杂挑战。本文提出的关键解决方案在于:针对爱沙尼亚语的文档级简化任务,系统评估五种先进的多语言大语言模型(LLMs),并比较三种提示策略——单次生成、基于流水线的模块化代理以及指南增强型流水线——的效果。研究创新性地构建了包含可读性、语义保真度和语篇连贯性在内的综合自动评估框架,并引入结构化的手动标注协议以确保评估可靠性。实验结果表明,Gemini-2.0与LLaMA-3.3在输出流畅性与语义保真度方面表现优异,接近母语水平,而其他模型存在显著的语法与语义缺陷。本研究贡献包括:新型文档级连贯性度量指标、基于实证的提示工程策略,以及公开可复现的数据资源,为低资源语言的高级文本生成任务提供了重要参考。

链接: https://arxiv.org/abs/2610.10378
作者: Meeri-Ly Muru,Eduard Barbu
机构: National Library of Estonia(爱沙尼亚国家图书馆); Institute of Computer Science, University of Tartu(塔尔图大学计算机科学研究所)
类目: Computation and Language (cs.CL)
备注: 12 pages, 2 figures, 2 tables. Published at LREC 2026

点击查看摘要

Abstract:Document-level text simplification involves transformations that go beyond sentence-internal edits, addressing discourse coherence, anaphora resolution, and cross-paragraph consistency. Despite advances in sentence-level simplification for high-resource languages, document-level simplification in morphologically rich, low-resource languages such as Estonian remains largely unexplored. This study presents a comprehensive evaluation of five state-of-the-art multilingual large language models (LLMs) for document-level simplification in Estonian. Three prompting strategies are examined: single-pass generation, pipeline-based modular agents, and guideline-augmented pipelines. The evaluation framework integrates automatic metrics assessing readability, semantic preservation, and discourse coherence, alongside a structured manual annotation protocol. The findings indicate that Gemini-2.0 and LLaMA-3.3 produce outputs with near-native fluency and strong meaning preservation, whereas other models display notable grammatical and semantic limitations. This work contributes novel document-level coherence metrics, evidence-based prompting strategies, and publicly available resources for reproducibility.

[NLP-12] Input-Blind Controls Produce Substantial Oracle Headroom for Layer Programs in Multiple-Choice Evaluation

【速读】: 该论文旨在解决生成式 AI(Generative AI)中自适应计算(adaptive computation)在语言模型推理过程中的有效性问题,特别是针对层跳过(layer-skipping)与层重复(layer-repetition)等动态执行策略能否真正带来性能提升的核心疑问。其核心挑战在于:尽管已有研究表明通过选择性执行可实现性能增益,但这种增益是否源于所选计算路径本身的有效性,还是仅由输入特征或选项排列等外部因素导致,尚不明确。论文的关键解决方案在于设计一系列严格的对照实验,包括输入盲扰动(input-blind perturbations)、固定字母偏移(fixed letter offsets)以及选项轮换(rotating options)等控制条件,以分离出真实计算选择所带来的独特收益。研究发现,在共享选项顺序下,控制组的头余空间(headroom)甚至超过真实程序,表明单纯的选择增益可能并非源自特定层计算的语义价值;而只有在引入基于KL散度校准的输入依赖型控制后,真实程序才在点估计上显现优势,且统计检验结果仍不充分。此外,通过生成答案测试进一步验证,搜索选定的程序在重述后仍保持26.0个百分点的优势,说明其性能提升具有跨提示鲁棒性。因此,该研究的关键结论是:现有自适应计算方法的性能增益难以被归因于所选层计算的内在优越性,必须通过更精细的对照设计才能揭示其真实贡献。

链接: https://arxiv.org/abs/2610.10368
作者: Yibei Guo,Rui Liu
机构: Kent State University (肯特州立大学)
类目: Machine Learning (cs.LG); Computation and Language (cs.CL); Software Engineering (cs.SE)
备注:

点击查看摘要

Abstract:Adaptive computation aims to improve language-model inference by tailoring execution to each input. For layer programs, oracle evaluations use known answers to estimate the potential gain from this flexibility, before a practical selector is available. However, a gain from selection does not by itself explain why the chosen programs help. This study examines this distinction using 32 layer-skipping and repetition programs on two models and 4,413 multiple-choice items. The analysis compares their gains over a fixed action selected without the evaluation prompt with those of input-blind perturbations at the same sites, re-evaluating selections on another prompt. With shared option order, the controls give 10.2-11.8 and 15.6-19.4 percentage points of headroom on Qwen3-4B-Base and Llama-3.1-8B, exceeding the real programs’ 9.0 and 10.1 in all three random-direction draws per model. They match answer-change rate only, and the ordering depends on the menu: in post hoc comparisons, real programs lead on Llama’s repeat-only menu in every draw. A smaller KL-calibrated comparison, including an input-dependent control, favours real programs in point estimate, with inconclusive corrected tests. Fixed letter offsets produce headroom of similar scale. Rotating options sharply reduces both families’ headroom, while leaving positive real-minus-control differences of 1.4-2.3 and 3.7-4.5 points; their magnitudes and statistical support depend on further adjustments and the reference. A supplementary generated-answer test finds that search-selected programs keep a 26.0-point advantage over programs selected for other problems after rewording, without a placebo comparison. These results show that substantial headroom can persist across prompts with shared option order without establishing a benefit specific to the selected layer computation; neither ordering against these controls identifies that benefit.

[NLP-13] Learning to Act with Task Progress: Distilling Small Agents from Compact Teacher Supervision

【速读】: 该论文旨在解决如何高效训练小型智能体(small agents)以完成重复性任务,同时避免在每一步执行时都调用大型模型的问题。其核心挑战在于:从教师模型(large model)生成的示范轨迹中,应保留哪些信息以指导学生智能体的学习。为此,论文提出了一种名为任务进展蒸馏(Task-Progress Distillation, TPD)的离线学习方法,其关键创新在于将每个示范动作与一个描述当前任务阶段的简短标签配对,形成紧凑的监督信号。学生智能体通过学习这些阶段-动作对的联合评分机制来选择行为,由确定性执行器在环境中执行。实验表明,在ALFWorld基准上,仅使用404条示范数据,采用TPD或仅动作监督的1.7B参数学生模型即可达到72.4%的未见任务成功率,显著优于使用受限动作选择的推理训练学生(48.3%)。当示范数据量为200条时,引入显式任务阶段信息可使成功率从48.0%提升至67.7%,显示出阶段性监督在中等数据规模下的额外优势;随着示范数据增加,动作监督的性能逐渐逼近TPD,两者在808条示范下均达76.9%。共享历史分析进一步揭示,TPD在子目标间转移(如从获取物品到处理物品)时决策更优,表明显式任务进展信息有助于提升局部策略质量。研究结果表明,紧凑的监督信号可有效训练高性能的小型任务智能体,而显式任务进展信息在中等示范数据预算下能提供额外的指导价值。

链接: https://arxiv.org/abs/2610.10332
作者: Wenxi Gan
机构: City University of Hong Kong (香港城市大学); Shenzhen Loop Area Institute (深圳环区研究院)
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Learning from large-model demonstrations offers a way to train small agents that can complete recurring tasks without calling a large model at every step. A central design choice is what to retain from teacher trajectories that contain reasoning, actions, and information about task progress. We introduce Task-Progress Distillation (TPD), an offline approach that pairs each demonstrated action with a short label describing the current task stage. The student learns these compact targets and selects actions by jointly scoring admissible stage–action pairs, which a deterministic harness executes in the environment. On ALFWorld, a 1.7B student trained with 404 demonstrations achieves 72.4% mean unseen task success with either TPD or action-only supervision, compared with 48.3% for a reasoning-trained student using constrained action selection. Explicit stages provide an additional benefit at 200 demonstrations, improving success from 48.0% to 67.7% over action-only supervision. With more demonstrations, the action-only student closes the gap, and both approaches reach 76.9% at 808 demonstrations. Shared-history analyses link part of TPD’s local advantage to better decisions when moving between subgoals, particularly from object acquisition to processing. These results show that compact supervision can train effective small task agents, while explicit task progress provides additional guidance at an intermediate demonstration budget.

[NLP-14] Nobody Truly Agrees on Sentiment: Humans Bespoke Tools and LLM s Struggle with Social Media Texts

【速读】: 该论文旨在解决社交媒体情感分析中自动化工具与人类判断之间的一致性问题,尤其关注现有情感分析工具在实际应用中因缺乏对自身局限性认知而可能导致的误判风险。其核心挑战在于评估不同情感分析方法(包括传统工具TextBlob、VADER、Twitter-roBERTa-base及大语言模型LLMs)在真实社交文本上的表现,并量化其与人类标注者之间的情感判断一致性。研究的关键解决方案在于通过引入双统计指标——Cohen’s kappa用于成对比较,Fleiss’ kappa用于多标注者一致性评估——系统性地对比六名人类标注者与七种自动化工具在100条推文上的情感分类结果。结果显示,尽管人类标注者间仅达到中等一致水平,表明情感判断具有高度主观性,但经过领域适配的Twitter-roBERTa-base模型在负向与非负向情感区分上表现出最强的人工标注一致性,显著优于其他工具;而大语言模型虽在正向与非正向分类中表现较好且彼此间具较高一致性,但在复杂三分类任务中仍存在明显偏差。研究进一步强调,针对特定领域进行微调(domain-specific fine-tuning)仍是提升情感分析可靠性的重要手段,同时以人类为中心的评估框架对于构建高质量基准标签(gold-standard labels)不可或缺。

链接: https://arxiv.org/abs/2610.10318
作者: Himarsha R. Jayanetti,Sivakanesan Dhanushkanda,Shuai Hao,Michael L. Nelson,Michele C. Weigle
机构: Old Dominion University (老多明尼昂大学)
类目: Computation and Language (cs.CL)
备注: 11 pages, 1 figure, 2 tables, accepted for publication at TPDL 2026

点击查看摘要

Abstract:Social media is a rich source of real-time public sentiment, but widely used sentiment analysis tools are often applied without understanding their limitations. In this study, we evaluate the inter-rater reliability of three bespoke sentiment analysis tools (TextBlob, VADER, and Twitter-roBERTa-base) and three large language models (LLMs: Qwen3-32B, GPT-OSS-120B, Llama-4-Maverick-17B) against six human raters across 100 tweets. We measured agreement using two statistical measures: Cohen’s kappa for pairwise comparisons and Fleiss’ kappa for multiple raters. Even among the human raters, our results showed only fair agreement, highlighting the subjectivity of sentiment analysis. Higher agreement was observed under the binary sentiment classification (negative vs. non-negative and positive vs. non-positive) than under the three-class classification across both humans and automated tools. The Twitter-roBERTa-base model showed the strongest alignment with human ratings, outperforming both bespoke sentiment tools and LLMs, particularly in distinguishing negative versus non-negative sentiment. LLMs showed substantial agreement among themselves and moderate to substantial alignment with humans, performing better in positive vs. non-positive classifications. Our findings underscore that domain-specific fine-tuning remains crucial for reliable social media sentiment analysis, and human-centered evaluation remains essential for establishing gold-standard labels.

[NLP-15] SemanticFold: Latent Sequence Compression SeparatesLanguage Modeling Decodability and Reasoning

【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在推理过程中,对提示前缀(prompt prefix)进行潜在空间序列压缩是否能够保持模型关键能力的问题。其核心挑战在于:如何在降低计算开销的同时,确保模型在生成质量、推理准确性及系统效率等方面的性能不受显著影响。解决方案的关键是提出一种名为SemanticFold的压缩机制,该机制通过在学习到的边界处折叠前缀的隐藏状态,实现对序列的高效压缩。研究采用固定目标协议(fixed-target protocol),通过对比原生执行与压缩后的模型表现,排除了目标选择偏差对结果的影响,并在五个不同规模的模型上进行了验证。结果显示,压缩对各类评估指标的影响是非单调的,且不存在统一的压缩阈值。进一步分析表明,压缩带来的负对数似然(NLL)改善主要源于残差变换(residual transform)的自适应能力,而非单纯的序列缩短;即使不进行序列压缩,仅应用残差变换的MLP-only方案也优于完整压缩方案。此外,线性探测准确率和宏平均AUC的变化均在0.03以内,置信区间包含零,说明模型表征可访问性未发生显著变化。最终结论指出,潜在压缩下的模型性能表现无法由单一标量指标衡量,语言建模拟合度、解码能力与推理行为各自响应不同的问题,可能在同一压缩操作下呈现不同方向的变化。

链接: https://arxiv.org/abs/2610.10304
作者: Mingyan Liu,Min Huang
机构: The Chinese University of Hong Kong (香港中文大学)
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:We study whether latent sequence compression of prompt prefixes preserves the capabilities that large language models rely on during inference. We introduce SemanticFold, a compression scheme that folds prefix hidden states at learned boundaries, and evaluate it across five model scales: Qwen3-1.7B, Qwen3-8B, SmolLM2-1.7B, Pythia-1.4B, and Pythia-6.9B. We use a fixed-target protocol: a frozen prefix is executed natively or compressed, and both arms teacher-force identical continuation tokens. This design rules out target-selection explanations for likelihood changes. We examine five endpoint families: fixed-target negative log-likelihood, finite-label reasoning accuracy, linear probe accessibility, open-ended generation, and systems-level memory and latency. We find that compression moves these endpoints non-monotonically and that they do not share a single compression threshold. On Qwen3-1.7B at compression ratio R=1.7, compressed-minus-native mean NLL decreases by 0.135 under paired bootstrap with 10000 draws. On SmolLM2 at R=1.2, the mean change is 0.013 higher than native. On both Pythia checkpoints, NLL is effectively unchanged. An NLL decomposition separating sequence shortening from the learned residual transform shows that the favorable Qwen likelihood is attributable primarily to residual adaptation rather than to shortening alone. MLP-only, which applies the transform without shortening, achieves 0.082 lower NLL than Full SemanticFold. Linear probe accuracy and macro AUC change by less than 0.03 in absolute value across conditions, with confidence intervals crossing zero. We conclude that preservation under latent compression has no single scalar certificate: language-model fit, decodability, and reasoning behavior answer different questions and can move in different directions under the same compression operation.

[NLP-16] PatchBench: Measuring Collateral Damage in Activation Patching NEURIPS2026

【速读】: 该论文旨在解决生成式 AI(Generative AI)在安全修复(jailbreak repair)过程中存在的“表面合规但实质失效”问题,即模型可能通过特定基准测试,却未能实现对恶意行为的精准修复,反而导致对无关良性输入的过度拒绝或抑制。其核心挑战在于现有评估方法难以区分“选择性修复”与“局部广泛抑制”,尤其在缺乏对具体行为敏感性的检测机制时,容易忽略修复过程中的副作用。解决方案的关键是提出PatchBench基准及其配套的PatchBench-Local评估协议:通过从大规模真实攻击数据中筛选出400个高置信度的、具有可行动性的模型特异性越狱失败样本,并构建三类局部邻近样本——保留恶意意图的有害变体、结构匹配的良性提示、以及复用关键有害词汇的良性提示,从而系统评估修复补丁在纠正有害响应的同时是否保持对良性输入的行为一致性。该方法揭示了即使全局能力(如MMLU得分)基本不变,局部良性性能仍可能严重退化,证明了仅依赖聚合指标的局限性。PatchBench-Local为开发和比较更精确、更安全的越狱修复方法提供了可量化的精细化评估基础。

链接: https://arxiv.org/abs/2610.10276
作者: Alexi Canesse,Mathis Le Bail,Maël Jenny,Clément Elliker,Mahammed El Sharkawy,Sonia Vanier
机构: LIX (École Polytechnique, IP Paris, CNRS)(法国巴黎综合理工学院、IP巴黎、法国国家科学研究中心联合实验室); AMIAD (Agence Ministérielle pour l’IA de Défense)(法国国防部人工智能管理局)
类目: Machine Learning (cs.LG); Computation and Language (cs.CL); Cryptography and Security (cs.CR)
备注: Accepted to NeurIPS 2026 (Datasets and Benchmarks Track)

点击查看摘要

Abstract:An LLM safety patch can pass a benchmark while still being a poor repair. This risk is especially acute for jailbreak repairs, where the goal is to correct a specific unsafe behaviour without changing unrelated behaviours. A patch may block exact evaluation prompts yet fail on close harmful variants, or suppress harmful behaviour by over-refusing benign prompts that share its wording or structure. Existing protocols primarily test whether models can be broken, while aggregate metrics (attack success, refusal rates, global capability) cannot distinguish selective repairs from broader local suppression. To address this gap, we introduce PatchBench, a benchmark of empirically observed model-specific jailbreak failures inducing actionable harmful answers. Starting from 27,870 prompts from 37 public datasets, we curate 15,314 English prompts and query 8 open-source instruction-tuned models. Combining WildGuard filtering, pairwise Elo ranking, and manual verification, we retain a curated bank of 400 high-confidence jailbreak failures. We further introduce PatchBench-Local, an evaluation protocol testing whether a patch is behaviourally precise. For each harmful source prompt, PatchBench-Local generates three families of local neighbours: harmful variants preserving malicious intent, benign prompts with matched structure, and benign prompts reusing key harmful terms. It evaluates harmful-neighbour correction and benign-neighbour preservation, distinguishing selective repair from broader local suppression. Evaluating four activation steering methods with PatchBench-Local and MMLU shows that global capability can remain nearly unchanged while local benign regressions are severe, confirming aggregate metrics miss important collateral damage. PatchBench-Local provides a more precise basis for developing and comparing jailbreak repair methods.

[NLP-17] LLM Persuasion Is in the Eye of the Evaluation

【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在说服能力评估中存在的标准不一、评价结果碎片化的问题。当前研究对“说服”的定义各异,且多数结论基于特定情境下的有限评估,难以形成可比性高的综合判断。为实现跨方法、大规模的统一评估,论文提出采用自动化评测方法,并将九种已发表的自动化评估手段在相同15个LLM上进行标准化测试,以检验不同方法间的排名一致性及其成因。其解决方案的关键在于:通过构建共享实验框架,系统性地比较多种自动化评估方法在相同模型上的表现,从而揭示评估结果差异的根本来源。研究发现,各方法间排名相关性较弱(平均斯皮尔曼等级相关系数ρ = 0.25),主要受两个因素影响:一是模型对部分任务(尤其是操纵类任务)存在选择性拒绝行为,导致评估结果偏差,此类拒绝行为使一致性下降约四分之一;二是模型的通用能力与说服方式密切相关——非操纵性说服方法多与模型整体能力正相关,而操纵性方法则缺乏这种关联。这表明,评估结果不仅反映模型的说服能力,更受其意愿(如是否愿意执行高风险说服任务)的影响。因此,单一说服评分仅能反映特定评估场景下的表现,难以推广至其他任务,提示未来评估需区分“能力”与“意愿”维度,并建立更系统化的多维评测体系。

链接: https://arxiv.org/abs/2610.10232
作者: Kamile Dementaviciute,Julija Vaitonyte,Tijl De Bie
机构: 未知
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Computers and Society (cs.CY)
备注:

点击查看摘要

Abstract:Large language models (LLMs) have already been shown to match or exceed human experts in persuasion. While their persuasive capabilities hold promise for beneficial uses such as education and health communication, they can also be used to manipulate and misinform, making their evaluation a growing priority for developers and regulators. That evaluation, however, remains fragmented: studies differ in what they treat as persuasion, and broad claims often rest on narrow, situation-specific assessments. Automated methods, often modelled on human studies, offer a way to compare such assessments directly, as they can be run on the same models at scale and can include high-risk forms of persuasion that would be difficult or unethical to test on people. In this study, we adapt nine published automated methods to a shared setup, run them on the same fifteen LLMs, and ask whether their rankings agree and why. We find that the methods agree only weakly (mean Spearman \rho = 0.25 ). Our analyses point to two contributing factors. Models that refuse some tasks but not others, directly or indirectly, lower agreement by about a quarter, and these refusals fall mostly on manipulation tasks. General capability also plays a part: most rational persuasion (non-manipulative) methods track it, whereas most manipulation methods do not. Together, these findings suggest that agreement depends more on the task a method sets than on how it scores persuasion, although this pattern is only indicative given the eight methods available for analysis. More broadly, our results suggest that persuasion scores combine a model’s ability to persuade with its willingness to do so. A single score is therefore informative about its own setting, but says little about a model’s persuasiveness across tasks.

[NLP-18] From Prompts to Trees: Effective LLM -Guided Tree Generation for Few-Shot Tabular Classification EMNLP2026

【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在表格数据分类任务中因推理成本高、可解释性差而难以直接应用的问题,同时克服决策树在小样本场景下性能不足的局限。其核心解决方案在于提出一种新颖的知识蒸馏框架,在少样本学习设置下将LLM蕴含的丰富世界知识有效提炼为可解释的决策树。该方法摒弃了直接通过提示(prompting)生成完整决策树的不稳定性与低效性,转而采用三阶段范式:首先引导LLM生成分类规则,再对规则进行结构化组织,最终构建出层次化的决策树。实验结果表明,该方法在多个真实世界表格数据集上实现了更优的分类准确率与更高的可解释性,且提示开销显著低于现有基线方法。

链接: https://arxiv.org/abs/2610.10227
作者: Yue Qiu,Zekang Du,Yiqun Diao,Bingsheng He,Qinbin Li
机构: Huazhong University of Science and Technology (华中科技大学); National University of Singapore (新加坡国立大学)
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: Accepted to EMNLP 2026 Main as an oral presentation. Code available: this https URL

点击查看摘要

Abstract:While Large Language Models (LLMs) possess rich world knowledge and impressive generalization capabilities, their direct application to tabular data classification is hindered by high inference costs and limited interpretability. In contrast, decision trees are fast and transparent but often underperform in low-data regimes. In this work, we propose a novel framework that bridges these paradigms by distilling LLM knowledge into interpretable decision trees under a few-shot learning setting. Instead of directly prompting the LLM to generate full trees, which is often unstable and inefficient, we develop a three-stage paradigm that prompts the LLM to generate rules and organize the rules into a tree. Experiments on multiple real-world tabular datasets demonstrate that our method achieves superior accuracy and interpretability with significantly lower prompting overhead compared to existing baselines.

[NLP-19] GAGR-Lab: Evaluating Joint Spatial-Geometric and Analytic Function Reasoning

【速读】: 该论文旨在解决生成式人工智能(Generative AI)在联合空间-几何与解析函数推理方面的能力评估问题,即如何将感知到的空间配置转化为满足几何约束的符号化函数,并准确执行其轨迹。其核心挑战在于实现从视觉空间感知到精确数学函数构造与验证的端到端可度量推理。解决方案的关键在于提出GAGR-Lab框架,该框架通过笛卡尔游戏场景(Cartesian game scenes)、显式的函数语义定义以及基于权威Rust语言的轨迹执行机制,实现了对空间感知、度量定位、几何关系建模、函数解释、函数构建及约束合成等多维度能力的分离与量化。该框架引入可配置的四类场景难度预设和潜在的24单元诊断设计,支持分阶段诊断校准、独立复制验证、多模型对比与配对鲁棒性测试。实验表明,尽管当前受限于API调用的非特权接口,主流模型在未受控条件下未能达成目标命中,而通过特权分析搜索控制则可在300个生成场景中实现600次方向性任务的完全成功与1200次垂直反射或平移验证,充分证明了框架在区分服务可靠性、符号合规性与几何成功率方面的有效性。因此,该研究贡献不仅是一个可执行的初步原型,更提供了一个清晰可扩展的未来研究路线图。

链接: https://arxiv.org/abs/2610.10201
作者: Jingyao Zhang,Yun Li,Lu Han
机构: 未知
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注: 15 pages, 1 figure, 7 tables

点击查看摘要

Abstract:Joint spatial-geometric and analytic function reasoning requires translating a perceived spatial configuration into a symbolic function whose executed curve satisfies geometric constraints. We present GAGR-Lab, a framework for measuring this capability through Cartesian game scenes, explicit function semantics, and authoritative Rust trajectory execution. It distinguishes spatial perception, metric grounding, geometric relations, function interpretation, function construction, and constrained synthesis. We specify four configurable scene-difficulty presets and a prospective 24-cell diagnostic design, while reporting only the subset actually evaluated. A bounded pilot of one hosted model (Llama 3.2 11B Vision Instruct) using two API credentials as execution replicas yields 72 balanced games with 432 attempts, 429 valid provider responses, and no target hits; exploratory ordinary-function prompt variants also fail to hit, while the structured localization interface yields no scoreable outputs. A privileged analytic search control independently succeeds on 600 directional cases from 300 generated scenes, with exact repeatability and 1,200 successful vertical-reflection or translation checks. The framework separates serving reliability, symbolic compliance, and geometric success, and preserves exact model-visible inputs and realized paths. A staged protocol outlines diagnostic calibration, held-out replication, multi-model comparison, and paired robustness tests. The contribution is an operational research framework with an executed pilot and a clearly identified prospective study plan; the full difficulty matrix and comparative model results remain untested.

[NLP-20] Beyond Outcome Rewards: Constructing and Assigning Retrieval Credit for Search Agents

【速读】: 该论文旨在解决生成式搜索代理在执行复杂多跳问答任务时,基于强化学习的后训练方法因依赖稀疏、结果导向的监督信号而导致信用分配困难、学习效率受限的问题。其核心解决方案在于引入中间阶段的监督信号,通过设计多种奖励塑形与信用分配策略,为多步搜索轨迹中的中间检索步骤提供学习信号。在此基础上,提出一种融合中间信号与最终结果奖励的训练框架,显著提升了搜索代理在多个基准测试中的综合性能。研究发现,中间信号的选择及其信用分配位置对训练行为具有重要影响,表明奖励设计与信用分配是构建高效搜索代理的关键设计维度。

链接: https://arxiv.org/abs/2610.10179
作者: Wenyu Huang,Xinyu Hou,Pavlos Vougiouklis,Ruofei Lai,Jeff Z. Pan
机构: University of Edinburgh (爱丁堡大学); Huawei Technologies Research Development (UK) Limited (华为技术英国研发中心)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Search agents enable Large Language Models (LLMs) to iteratively retrieve and use information for complex multi-hop questions. Reinforcement Learning with Verifiable Rewards (RLVR) offers a promising approach for post-training such agents, but its reliance on sparse, outcome-based supervision can make credit assignment difficult and limit learning efficiency. In this paper, we systematically investigate how intermediate supervision can improve reinforcement learning for search agents. We study a range of reward-shaping and credit-assignment strategies that provide learning signals from intermediate retrieval steps. Building on these insights, we develop a training framework that combines intermediate signals with final outcome rewards to improve learning from multi-step search trajectories. Experiments across multiple benchmarks under matched training conditions demonstrate improvements in aggregate search-agent performance and show that both the choice of intermediate signal and where its credit is assigned affect training behaviour. These findings show that reward design and credit assignment are important design dimensions for training effective search agents.

[NLP-21] HySPE: Positional Encoding via Symplectic Dual Shears

【速读】: 该论文旨在解决传统旋转位置编码(Rotary Position Embedding, RoPE)在长序列建模中因指数级表示漂移导致的外推性能退化与数值不稳定性问题。其核心解决方案是提出双曲辛位置编码(Hyperbolic Symplectic Positional Encoding, HySPE),通过非紧致的双曲辛变换替代传统紧凑椭圆分支的旋转操作,利用阻尼对称双剪切组合实现共形辛收缩,并引入通道对内具有两个谱衰减率的结构以增强对长程依赖的建模能力。为消除绝对因子分解带来的指数漂移,HySPE在不变特征基下对算子进行对角化,并采用自适应中心化的分块坐标重基变换,从而保证长度无关的数值稳定性。实验表明,在TinyShakespeare上,HySPE-UltraLong在4096长度下零样本外推时保持恒定困惑度4.810,显著优于RoPE的131.198;在WikiText-103上的51M参数子词Transformer模型中,HySPE在训练长度512的基础上稳健外推至8192,尾部困惑度较RoPE降低83.9%,同时保持与RoPE相当的前向延迟(RTX 4090上7.21 ms)。该方法在保障高效推理的同时,显著提升了位置编码在长序列外推中的稳定性和泛化能力。

链接: https://arxiv.org/abs/2610.10154
作者: Zhongping Ji
机构: 未知
类目: Computation and Language (cs.CL)
备注: 11 pages

点击查看摘要

Abstract:We introduce Hyperbolic Symplectic Positional Encoding (HySPE), grounding positional attention in non-compact symplectic transformations. While canonical Rotary Position Embedding (RoPE) parameterizes the compact, elliptic branch of \Sp(2,\R) via rotations, HySPE operationalizes its hyperbolic branch via a damped symmetric composition of dual shears, yielding a conformally symplectic contraction with two spectral decay rates per channel pair. To eliminate the exponential representation drift inherent to naive absolute factorizations, we diagonalize the operator in its invariant eigenbasis and introduce blockwise coordinate rebasing with adaptive centered execution. This guarantees length-independent numerical bounds while matching cached RoPE forward latency (7.21,ms on an RTX 4090). On TinyShakespeare, HySPE-UltraLong maintains an invariant perplexity of 4.810 up to 16\times zero-shot extrapolation ( L=4096 ), whereas RoPE degrades to 131.198. Scaled to a 51M-parameter subword Transformer on WikiText-103 ( L_\texttrain=512 ), HySPE closely matches RoPE in-domain while robustly extrapolating to length 8192, reducing tail perplexity by 83.9% over RoPE. While these controlled experiments establish HySPE’s extrapolation robustness and numerical stability, evaluating its scaling behavior on large-scale foundation models remains an important direction for future investigation.

[NLP-22] InterView-C: A Synchronized Multimodal Corpus of VR Avatar-Mediated Survey Interviews

【速读】: 该论文旨在解决现有自然语言处理(NLP)资源在处理虚拟现实(VR)环境中生成的多模态对话数据时所面临的两大核心问题:一是自动语音识别(ASR)系统在转录口语时容易扭曲具有语言学意义的信息,二是基于传统文本语料训练的下游模型在应用于经转录的口语数据时存在显著的迁移性能下降问题。其解决方案的关键在于构建一个高质量、高对齐精度的德语多模态语料库——InterView-C,该语料库包含27场完全在虚拟现实中进行的访谈,由双方面化身参与,同步记录了语音、眼神、头部与身体动作、面部表情、手部及手指追踪等丰富的行为数据。通过提供逐词时间标记且经人工后编辑的逐字转录文本、访谈项时间戳、问卷响应以及1,422个句子的否定线索与作用域双重标注(线索一致性α=0.87,作用域一致性α=0.81),该语料库实现了口语互动与多模态行为的高度对齐,并为以文本为基础的NLP方法提供了可靠接口。实证研究表明,九种开源ASR系统在短封闭式回答和数字词上存在显著误识,而基于已有语料训练的否定识别模型在本语料上的表现明显低于在相同任务上使用InterView-C标注数据训练的模型,验证了其在提升模型泛化能力方面的关键价值。因此,InterView-C不仅支持对口语互动的语言学分析,同时确保了分析结果与多模态行为数据的精确对应。

链接: https://arxiv.org/abs/2610.10145
作者: Patrick Schrottenbacher,Leon Hammerla,Lydia Kleine,Doris Stingl,Alexander Mehler
机构: Goethe University (歌德大学), Frankfurt am Main, Germany; Leibniz Institute for Educational Trajectories (LIfBi) (莱布尼茨教育轨迹研究所), Bamberg, Germany
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:We present InterView-C, a German multimodal corpus of 27 survey interviews conducted entirely in virtual reality, with both interlocutors represented by avatars. The corpus aligns spoken interaction with synchronized behavioral data, including gaze, head and body movement, facial behavior, hand and finger tracking. Its reference transcripts and linguistic annotations provide a reliable interface between this multimodal spoken interaction and predominantly text-based NLP methods. This interface is important because automatically transcribing speech can distort linguistically relevant information, while downstream models trained on existing resources may additionally face transfer challenges when applied to transcribed spoken data. InterView-C therefore provides word-timed and manually post-edited verbatim transcripts for all 54 recordings, interview-item timings, questionnaire responses and negation cue and scope annotations for 1,422 sentences, 1,398 of them doubly annotated (\alpha=0.87 for cues; \alpha=0.81 for scopes). We demonstrate both challenges empirically: nine open-weight ASR systems disproportionately misrecognize short closed answers and number words, while negation models trained on existing corpora show lower and highly variable performance on our transcribed interviews than a model trained on the InterView-C annotations. InterView-C thus enables linguistic analyses of spoken interaction while retaining their alignment with rich multimodal behavior.

[NLP-23] LLM 4Impact: Integrating Heterogeneous Information for Scientific Impact Prediction

【速读】: 该论文旨在解决新发表论文未来影响力的预测难题,其核心挑战在于如何从发表时可用的异质性证据中进行有效推断。现有方法通常依赖单一信息源或简单融合多源信息,未能充分考虑不同证据在预测中的差异化作用。本文提出LLM4Impact,一种具备证据感知能力的科学影响力预测方法,其关键在于通过联合学习语义、图结构、大语言模型(LLM)及时间序列表示,并将图信息以连续前缀令牌(continuous prefix tokens)形式注入冻结的LLM中,实现对异质信息的动态表征与整合。进一步引入上下文感知门控机制,自适应加权不同证据来源,同时设计独立校准模块以处理领域和时间维度上的引文尺度差异。研究构建了一个包含200万篇论文的大规模基准数据集,涵盖泄漏安全的时间点异质性局部图、时间划分策略以及年级与月级双粒度引文目标。实验表明,LLM4Impact在分布内测试集上实现年均绝对误差(RMSE)降低10.13%,跨领域场景下降低6.87%,显著优于多种基线模型。结果揭示:证据价值具有强情境依赖性,不同论文受益于不同的信息源,且领域与发表时间影响证据向引文的转化效率。这一发现推动了“自适应证据选择”与“上下文条件化校准”的范式转变,而非单纯追求更丰富的表征能力。

链接: https://arxiv.org/abs/2610.10138
作者: Yong Cao,Markus Flicke,Haoyu He,Katrin Renz,Andreas Geiger
机构: 未知
类目: Computation and Language (cs.CL)
备注: 27 pages, 12 figures, 11 tables

点击查看摘要

Abstract:Predicting the future impact of a newly published paper is challenging because it must be inferred from heterogeneous evidence available at publication time. Existing approaches often rely on a single source of information or combine multiple sources without accounting for their different predictive roles. In this paper, we present LLM4Impact, an evidence-aware method for scientific impact prediction that learns to represent, integrate, and calibrate heterogeneous information. LLM4Impact combines semantic, graph, LLM, and temporal representations, and injects graph information into a frozen LLM through continuous prefix tokens. A context aware gating mechanism adaptively weights different evidence, while a separate calibration module accounts for domain and temporal variation in citation scales. We further construct a large-scale benchmark dataset with 2 million papers, leakage-safe point-in-time heterogeneous ego graphs, temporal splits, and both year-level and month-level citation targets. Experiments show that LLM4Impact consistently outperforms strong semantic, graph, and LLM based baselines, with a 10.13% reduction in year RMSE on the in distribution test set and a 6.87% reduction under out-of-domain distribution. Our results reveal that the value of such evidence is context dependent: different papers benefit from different sources, while domain and publication time affect how evidence translates into citations. This finding motivates adaptive evidence selection and context-conditioned calibration rather than simply richer representations. We will release our code, benchmark, and an interactive web demonstration upon publication.

[NLP-24] YANchor-4B: Effective Long-Horizon Reasoning in O(N) Time with O(1) Memory

【速读】: 该论文旨在解决长时程推理(long-horizon reasoning)中如何在控制生成成本的前提下有效访问早期信息的问题。传统方法中,全历史注意力机制(full-history attention)导致存储与计算开销随上下文长度线性增长,而递归压缩(recurrent compression)则可能丢失关键细节。为此,本文提出YANchor-4B——一种通用的递归模型,通过将重要记忆以“锚点”(ANchors)形式保留,支持后续推理过程中的高效检索。其核心创新在于多维记忆机制(multidimensional memory mechanism),实现了生成时间复杂度O(N)、内存占用O(1)的同时,保障了长时程推理的有效性。实验表明,在高难度数学题集AIME 2024–2026和HMMT上,YANchor-4B分别达到82.93%和63.64%的mean pass@1,显著优于同类线性时间、常数状态的基线模型(包括更大规模模型)。此外,在H100硬件上,其批量长序列生成吞吐量较Transformer及混合基线提升数倍,且在数十项基准测试中展现出卓越的通用能力。

链接: https://arxiv.org/abs/2610.10118
作者: Huishan Ji,Hua Xu,Weiming Zhang,Qirui Ye
机构: 未知
类目: Machine Learning (cs.LG); Computation and Language (cs.CL)
备注: 24 pages, 8 figures. Code: this https URL ; Model: this https URL

点击查看摘要

Abstract:Long-horizon reasoning demands access to earlier information at a manageable generation cost. Full-history attention incurs growing storage and computation, while recurrent compression can lose precise details. Therefore, we present YANchor-4B, a general-purpose recurrent model that preserves crucial memory as ANchors for retrieval during subsequent reasoning. Beyond O(N) -time generation and O(1) memory, YANchor enables effective long-horizon reasoning through its multidimensional memory mechanism. For example, on challenging math problems, it achieves 82.93% mean pass@1 on AIME 2024–2026 and 63.64% on HMMT, substantially outperforming linear-time, constant-state counterparts, including larger models. It also delivers several-fold higher batched long-generation throughput than Transformer and hybrid baselines on H100. Furthermore, evaluations across dozens of benchmarks demonstrate YANchor’s superiority in general-purpose capabilities.

[NLP-25] Mechanics of Long-Context Hybrid Models Part 1.1: From Hybrid Attention to Hybrid Position

【速读】: 该论文旨在解决当前大语言模型(Large Language Models, LLMs)在长上下文处理中的效率与泛化能力瓶颈,特别是传统全注意力(full attention)模型在长序列建模时计算开销过大、难以有效进行长度外推(length extrapolation)和持续长上下文预训练的问题。随着架构演进向混合注意力(hybrid attention)模式转变,如何理解其内在工作机制并指导高效设计成为关键挑战。本文提出“长上下文混合模型力学”(Mechanics of Long-Context Hybrid Models)框架,聚焦于全注意力与滑动窗口注意力(Sliding-Window Attention, SWA)或门控线性注意力(Gated Linear Attention, GLA/GDN)的组合结构。其核心发现为:在上下文扩展任务中存在“跷跷板效应”(Seesaw Effect),即线性注意力(LA)型混合模型更受益于长上下文持续预训练,而滑动窗口注意力(SWA)型混合模型在长度外推方面表现更优。这一差异源于不同注意力机制所引入的位置归纳偏置(positional inductive bias)的不同。研究进一步揭示了SWA混合模型面临“短上下文学习陷阱”(Short-Context Learning Trap)、“短窗疲劳”(Short-Window Weariness)与“长窗惰性”(Long-Window Laziness)三大问题,需通过扩大窗口尺寸来缓解。针对LA混合模型,提出了“混合位置外推的马太效应”(Matthew Effect of Hybrid Position Extrapolation),并设计出滑动窗口线性注意力(Sliding-Window Linear Attention, SWLA),实现了无需额外训练即可达到16倍的长度外推能力,在64k上下文长度下仍保持对NIAH-SK1基准100%的准确率,显著提升了长上下文建模的效率与可扩展性。

链接: https://arxiv.org/abs/2610.10114
作者: Xiaoran Liu,Ziwei He,Xipeng Qiu
机构: Shanghai Innovation Institute; OpenMOSS Team; Fudan University
类目: Computation and Language (cs.CL)
备注: 60 pages, 36 figures, 25 tables, under review

点击查看摘要

Abstract:The architectural design of Large Language Models (LLMs) is shifting from traditional full-attention-only models to hybrid models, which combine different attention modules to improve long-context efficiency and performance in length extrapolation and context extension. To explain why hybrid models work and how to design them better, we propose Mechanics of Long-Context Hybrid Models. As Part 1.1 of this series, we begin with hybrids of full attention and either sliding-window attention (SWA) or gated variants of linear attention (LA), represented by GLA and GDN. We first observe a Seesaw Effect in Context Extension: LA hybrids benefit more from long-context continual pretraining, whereas SWA hybrids perform better under length extrapolation. We attribute this behavior to differences in the positional inductive biases induced by these attention mechanisms. We find that SWA hybrids suffer from a Short-Context Learning Trap, Short-Window Weariness, and Long-Window Laziness, and require extended windows to enhance performance in continual long-context pretraining. For LA hybrids, we summarize the Matthew Effect of Hybrid Position Extrapolation and propose Sliding-Window Linear Attention, achieving 16 \times training-free length extrapolation while maintaining 100% accuracy on NIAH-SK1 in 64k context length.

[NLP-26] I would rather quit NLP than read another paper like this: The rise of antithesis in NLP papers

【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在科学写作中普遍出现的冗余反义表达问题,即模型反复强调研究不做什么,此类表述虽频繁但无助于提升内容的精确性或表达质量,反而引发审稿人反感。其核心解决方案的关键在于揭示这一现象的成因:该表达模式是模型在后训练阶段基于成对偏好数据进行优化时产生的副作用——系统倾向于奖励单次回复中对替代方案的否定,却无法衡量这种否定在整个文本中的累积成本,从而导致生成内容中过度使用贬低性对比。研究通过对比2019年、2026年ACL风格论文及GPT生成论文的使用频率发现,该现象在2026年论文中出现频率为2019年的七倍,且在GPT生成文本中更为显著;人工评估显示,多数2026年论文均包含令评估者感到不适的反义表达,且这些表达往往比合法使用更不公正地贬低被排除的备选方案。

链接: https://arxiv.org/abs/2610.10092
作者: Olga Zamaraeva,Adrián Gude,Roi Santos-Ríos,Carlos Gómez-Rodríguez
机构: Universidade da Coruña, CITIC(拉科鲁尼亚大学,计算与信息技术中心)
类目: Computation and Language (cs.CL)
备注: 30 pages, 2 figures, 48 tables

点击查看摘要

Abstract:For better or worse, LLMs are by now used routinely for scientific writing.\footnoteThis paper is no exception; we did use AI to assist with writing some of the sections (see Acknowledgments). Many have noticed that recent models fill papers with unnecessary antithesis, stating over and over what the work does not do, in ways that do not contribute to its precision or quality of expression and annoy reviewers \emphrather than impressing them. We study the construction \emphrather than in ACL papers from 2019, ACL-style arXiv papers from 2026, and papers written by GPT models from the same titles and abstracts. Its rate in 2026 is seven times the 2019 rate, and higher still in the GPT papers. Two annotators, blind to the source, find almost no 2019 use \emphannoying and about one in ten 2026 uses; they seldom agree on which, yet about half of 2026 papers contain a use that annoys each of them. \emphAnnoying uses present the rejected alternative less favorably than legitimate uses. Raters of preference data and open reward models favor the construction, and an instruction to be honest promotes it. We conjecture that it is a side effect of post-training on pairwise preferences, which credit a disavowal in a single response and cannot register its cost across a text.

[NLP-27] SkillSandbox: Skill Verification via Dynamic Scenario Synthesis

【速读】: 该论文旨在解决自演化智能体在将任务求解经验提炼为可复用技能时,所生成的技能可能包含错误流程或不可迁移知识的问题。其核心挑战在于如何有效验证每项技能是否具备真正的可复用性——即该技能的指导作用是否能在原始训练情境之外的新任务中依然产生价值。现有方法难以构造能够充分暴露目标技能实际应用条件的任务场景,因而无法准确评估其泛化能力。为此,论文提出SkillSandbox框架,通过动态合成与特定技能相关但新颖的任务与环境来实现精准验证。该框架由三部分构成:Proposer定义需保留的约束条件与应变化的源特定细节;Builder据此构建可执行的测试场景;Verifier则对比有无该技能时的执行表现,从可执行性、实用性及效率三个维度综合评估,并作出“保留”(Keep)或“拒绝”(Reject)的决策,决定技能是否进入技能库。实验在ALFWorld和WebShop两个基准上使用三种模型进行验证,结果表明SkillSandbox显著提升下游任务性能并优化执行效率。进一步分析证实,性能增益源于对技能可复用性的准确评估,且框架中各组件协同贡献了关键效果。

链接: https://arxiv.org/abs/2610.10088
作者: Serin Kim,Kwangwook Seo,Dokyung Song,Jinyoung Yeo,Dongha Lee
机构: Yonsei University (延世大学)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Self-evolving agents distill task-solving experience into skills for future reuse, but these skills can encode incorrect procedures or non-transferable knowledge. It is therefore critical to verify each skill’s reusability: whether its guidance remains useful beyond the experience from which it was distilled. Such verification requires observing how a skill affects execution in new tasks, yet existing tasks may not expose the situations where the target skill can actually be exercised. To construct such situations, we propose SkillSandbox, a framework that dynamically synthesizes a task and its environment for each skill that are skill-relevant yet novel. A Proposer specifies the conditions to preserve and the source-specific details to vary, a Builder constructs an executable scenario, and a Verifier compares executions with and without the skill. The Verifier assesses executability, utility, and efficiency to assign a Keep or Reject verdict, determining whether the skill enters the library. Across ALFWorld and WebShop with three models, SkillSandbox consistently yields the strongest downstream performance and improved execution efficiency. Further analyses examine whether these gains reflect accurate assessment of skill reusability and identify which components of SkillSandbox contribute to them.

[NLP-28] Cache the Encoder Within:Compact Reusable Memory across LLM Queries

【速读】: 该论文旨在解决共享文档在重复查询时产生的冗余编码问题,以及缓存模型状态带来的持久化存储开销问题。其核心解决方案是提出EncBank框架,通过利用预训练大语言模型(LLM)的低层作为可复用的文档编码器,并以紧凑形式存储其输出,结合自蒸馏的后缀适配器(self-distilled suffix adapter)实现跨不同存储精度的共享,避免了针对量化特定精度的重新训练。该方法在三个不同规模及全注意力与混合架构的Qwen骨干网络上,在五个基准测试套件中均实现了4比特存储下与原精度EncBank性能差距小于1分的保持效果;在固定Qwen3-8B工作负载下,仅需原精度持久化GPU存储的28.1%。此外,相比相同证据、相同适配器的文本重播方式,原精度控制下的封装配置可实现1.40倍的选包预填充加速,代价为RULER指标下降3.12分;而原精度的Qwen3.8-27B配置亦成功通过终端基准测试2.1中的70/89项任务。因此,EncBank实现了计算复用与内存压缩的协同优化,但任务保真度与端到端收益仍高度依赖于具体工作负载、准备成本及复用频率。

链接: https://arxiv.org/abs/2610.10058
作者: Hanzuo Liu,Chunyu Liu,Chaofan Lin,Alex Lamb,Mingyu Gao
机构: Tsinghua University (清华大学)
类目: Computation and Language (cs.CL)
备注: 17 pages, 3 figures, 7 tables

点击查看摘要

Abstract:Repeated queries over shared documents incur redundant encoding, while caching model states introduces persistent storage costs. Building on CoMem’s intermediate-state interface, EncBank treats a pretrained LLM’s lower layers as a reusable document encoder and compactly stores their outputs for an adapted upper-layer reader. A self-distilled suffix adapter is shared across storage precisions within each backbone, without quantization-specific retraining. Across five benchmark suites on three Qwen backbones spanning different sizes and full-attention and hybrid architectures, 4-bit storage keeps each reported benchmark aggregate within one score point of native-precision EncBank. In a fixed Qwen3-8B workload, it retains 28.1% of the native-precision persistent GPU store. Separate native-precision controls yield a 1.40x selected-pack prefill speedup over same-evidence, same-adapter text replay, at a 3.12-point RULER accuracy cost. A native-precision Qwen3.8-27B configuration also passes 70 of 89 Terminal-Bench 2.1 tasks. EncBank thus combines reusable computation with compact memory, while task fidelity and end-to-end benefits remain dependent on the workload, preparation costs, and reuse frequency.

[NLP-29] he Long Road to the Same Answer: Cognitive Bias Under Escalating Reasoning Budgets in Large Language Models KR

【速读】: 该论文旨在探究生成式 AI(Generative AI)在推理过程中是否能够有效缓解人类决策中的经典认知偏差,特别是基于双过程理论(dual-process theory)的假设——即通过延长推理时间可削弱快速直觉判断所导致的偏见。其核心问题是:在推理模型中引入额外的推理性计算(即“思考”)是否能真正提升决策理性,从而减少认知偏差。解决方案的关键在于设计一个剂量-反应研究(dose-response study),通过对比四类推理模型与其对应的非推理对照模型,在不同思维上限(0、1,024、4,096、8,192个标记符)下执行30个涵盖六种典型偏见(锚定效应、框架效应、损失厌恶、承诺升级、可得性启发、确认偏误)的案例测试,并以实际消耗的推理标记数作为“真实推理程度”的度量指标。研究发现,推理模型并未表现出比其非推理兄弟模型更小的偏见,且随着实际推理长度增加,偏见幅度未显著下降,反而在多数情况下朝与人类相反的方向加剧;唯一呈现人类方向偏见的是锚定效应(d = 1.89),而其余五种偏见在多数模型中均出现反向偏离。此外,仅通过一句指令要求模型重述锚点即可降低锚定效应,但这一效果未达统计显著水平(p = .057)。因此,研究结果表明,不能将推理阶段的计算投入视为理性保障,必须对部署的模型进行逐项偏差审计。

链接: https://arxiv.org/abs/2610.10049
作者: Obada Kraishan
机构: Texas Tech University (德州理工大学)
类目: Computation and Language (cs.CL)
备注: Accepted at the 2026 IEEE 8th International Conference on Cognitive Machine Intelligence (IEEE CogMI 2026). 8 pages, 3 figures, 5 tables. Code and data: this https URL

点击查看摘要

Abstract:Reasoning models allocate extra computation at inference time and present their answers as the product of deliberate thought. If this deliberation works the way dual-process accounts of human cognition suggest, longer thinking should weaken the classic decision biases that fast, intuitive judgment produces. Using 30 vignettes covering six biases (anchoring, framing, loss aversion, escalation of commitment, availability, confirmation) from an established benchmark, we run a dose-response study across four model families, pairing each reasoning model with a matched non-reasoning sibling and requesting thinking ceilings of 0, 1,024, 4,096, and 8,192 tokens, for 12,350 API calls. Because a requested ceiling is not the same as realized deliberation, we use the reasoning tokens each call consumed as the dose. First, reasoning models are not less biased than their siblings; the point estimate leans the other way in every family, but the item-level pooled contrast is not reliable (Delta = +0.031, t(29) = 1.45, p = .157). Second, bias magnitude does not reliably fall as realized deliberation grows: no slope is significantly negative, and where anything moves it is the signed score drifting further from the human direction. Third, anchoring is the only bias in the human direction (d = 1.89). Four of the other five lean the opposite way in all seven models; with five items per bias, that reversal is reliable for framing and directional for escalation of commitment, confirmation, and loss aversion, while availability is absent. A one-line instruction to restate the anchor before answering lowered anchoring on all five anchoring items, which no amount of additional thinking did, although the effect does not reach significance (p = .057). The results argue against treating test-time reasoning as a rationality guarantee and for auditing deployed models bias by bias.

[NLP-30] Sensitive-Topic Leakage Through LLM Routing Metadata: Measurement and Mitigation

【速读】: 该论文旨在解决大语言模型(Large Language Model, LLM)路由系统在实际应用中可能引发的隐私泄露问题,特别是当系统根据请求内容选择不同成本或质量的模型时,即便关闭了内容日志记录,仍可通过分析请求特征与路由决策之间的关联性实现对敏感内容的推断。其核心挑战在于:尽管部分平台通过禁用内容日志来声称保护用户隐私,但路由行为本身仍构成一条隐蔽的隐私信道,尤其在存在噪声标签、重复提示和类别偏移的情况下,攻击者可利用路由决策模式重构用户的敏感请求内容。

解决方案的关键在于识别并量化这种基于路由决策的隐私风险,并提出有效的防御机制。研究通过在170万真实请求(WildChat-1M, LMSYS-Chat-1M)上进行预注册实验,评估了两种成本/质量路由策略及一个领域路由策略的隐私泄露程度,发现即使在长度匹配条件下,不同类别请求的路由偏差方向与幅度显著差异——例如,在RouteLLM的50%运行点下,涉及骚扰、自残和医疗类请求的强模型访问率分别比未探索提示下的同类请求低19点、31点,而性相关请求则高出10点。进一步分析表明,仅采用按对话粘性、用户预算分段或聚合公平性等策略无法有效缓解此类偏差,且其效果受类别间偏移方向和大小差异的影响。唯一有效的防御方式是基于准确标签的逐类别长度匹配公平性调整,可在不显著降低模型性能(最多损失0.2个准确率点)的前提下消除隐私差距;然而,后验的精确逐用户速率控制虽能隐藏偶数前缀强模型调用次数,却牺牲了大部分自评估路由价值,且仍允许攻击者在约20次请求后通过奇数位置决策重建隐私信息,达到AUC 0.73的攻击效果。因此,该研究揭示了现有路由系统在隐私保护上的根本缺陷,并强调需在设计阶段引入更严格的公平性和可追溯性约束。

链接: https://arxiv.org/abs/2610.09981
作者: Teng-Ruei Chen
机构: Krixvon(台北); National Chiao Tung University (国立交通大学)
类目: Cryptography and Security (cs.CR); Computation and Language (cs.CL)
备注: 20 pages, 5 figures, 9 tables

点击查看摘要

Abstract:LLM routers pick a cheap or expensive model per request by its content, and many gateways and some cloud platforms can log that choice with content logging off. We measure this privacy channel beyond token counts, accounting for noisy labels and repeated prompts. We run pre-registered studies on 1.7 million real requests (WildChat-1M, LMSYS-Chat-1M) with two cost/quality routers and a domain router, survey eleven systems’ logging, and test post-processing defenses. At matched length, the shift’s direction depends on category and router. For RouteLLM at the 50% operating point, harassment and self-harm requests reach the strong model 19 points less often than comparable ones on prompts unseen in exploration, medical requests (exploratory: LLM labels failed their gate) 31 points less often on distinct prompts (both post hoc), and sexual requests 10 points more often (secondary); the other router’s four are negative. Twenty RouteLLM decisions separate frequent medical askers with AUC 0.71, exploratory and below the pre-registered primary endpoint’s 0.75 (domain router: 0.92, an upper estimate). Per-category length-matched parity with accurate labels removes the gap on real traffic, costing at most 0.2 accuracy points on RouterBench (post hoc), where routers’ gaps on sensitive subjects (13-42 points, pre-registered) exceed those of an oracle routing by realized accuracy gain (1-11, post hoc). Per-conversation stickiness, per-user budget bands, and pooled parity fail, the last as categories’ shifts differ in size or sign. A post hoc exact per-user rate hides only even-prefix strong counts and forfeits most self-assessed routing value; it preserves odd-position decisions, from which a post hoc log attack reaches AUC 0.73 after 20 RouteLLM requests (exploratory).

[NLP-31] EASE: Entropy-Adaptive Distribution Shaping for Evading AI-generated Text Detectors

【速读】: 该论文旨在解决生成式 AI (Generative AI) 文本检测系统在面对不同解码策略时的脆弱性问题,尤其关注大型语言模型(LLM)在生成文本时通过调整采样温度或扰动下一个词元的概率分布,导致检测性能显著下降的现象。其核心解决方案是提出一种无需训练且对检测器无关的逃逸框架——EASE(Entropy-Adaptive Distribution Shaping for Evasion)。EASE 的关键在于直接利用源 LLM 的下一个词元分布计算预测熵,并据此自适应地调节对数概率扰动与采样温度,从而在不依赖检测器反馈或模型微调的前提下,有效规避检测,同时保持文本质量几乎不受影响,并实现极低的推理开销。

链接: https://arxiv.org/abs/2610.09976
作者: Jicheng Zhou,Kahim Wong,Jialong Wang,Jiantao Zhou
机构: University of Macau (澳门大学)
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:AI-generated text (AIGT) detection can be sensitive to the decoding choices of the source large language model (LLM). We observe that perturbing next-token logits or adjusting sampling temperature can reduce detection performance, providing a clear signal of detector vulnerability to decoding-time distribution changes. Building on this observation, we propose EASE (Entropy-Adaptive Distribution Shaping for Evasion), a training-free and detector-agnostic framework for evading AIGT detectors. EASE computes predictive entropy directly from the source LLM’s next-token distribution and uses it to adapt both logit perturbation and sampling temperature, without detector feedback or model fine-tuning. Experiments across three source LLMs and multiple detectors demonstrate consistent reductions in detection performance, with negligible degradation in text quality and negligible inference overhead.

[NLP-32] Itgan at NADI 2026 shared task: Parameter-Efficient Whisper Adaptation for Robust Mixed-Dialect and Code-Switched Arabic ASR

【速读】: 该论文旨在解决NADI 2026挑战赛中三个语音识别(ASR)子任务的性能优化问题,包括鲁棒性国家级方言语音识别(1.1)、混合方言语音识别(1.2)以及突尼斯语代码切换语音识别(1.3)。针对这些任务,研究提出了一种统一的解决方案框架:基于消费级GPU实现的Whisper模型,通过引入低秩适配器(LoRA)进行微调。其核心创新在于采用“共享主干+差异化适配”的设计思路,即所有子任务共享同一套基础训练流程(recipe),但通过不同形式的增量模块实现针对性改进。在任务1.1中,关键成功因素是使用基于池化适配器的方言专属专家(per-dialect specialists),显著提升了识别效果;在任务1.2中,基础模型的选择比适配器容量更为重要,且只有在引入去相关性成员后系统融合才有效;而在任务1.3中,最终系统通过独立训练运行结果的权重空间平均与ROVER投票,在无需额外训练的情况下进一步提升性能,达到14.49%的词错误率(WER)和5.38%的字符错误率(CER),位居前列。此外,研究通过配对自助法(paired-bootstrap test)验证了每项改进的有效性,并系统性地排除了八种无效尝试方向,增强了结论的可靠性。

链接: https://arxiv.org/abs/2610.09934
作者: Ibrahim Almajai
机构: 独立研究员(Independent Researcher)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Sound (cs.SD)
备注: 12 pages, Arabic NLP 2026 Shared Task

点击查看摘要

Abstract:We describe the Itgan systems for the three ASR subtasks of NADI 2026, namely robust country-level ASR (1.1), mixed-dialect ASR (1.2), and Tunisian code-switched ASR (1.3). All three share one recipe, Whisper adapted with LoRA on consumer GPUs, and each was carried by a different addition to it. On 1.1, where the dialect label is given at test time, per-dialect specialists continued from a pooled adapter gave the largest gain, and the submitted system reached 57.1% country-average WER. A post-evaluation linear probe on frozen encoder features routes utterances without the label and recovers 44% of what oracle routing gives. On 1.2 the choice of base model mattered more than adapter capacity, and system combination helped only once we added a decorrelated member, reaching 46.7% WER. On 1.3 our system placed second at 14.49% WER with the lowest CER among the leading submissions, 5.38%. Its last 0.60 WER points came without further training, mostly from an exact weight-space average of independently trained runs, with ROVER voting adding the remainder. Every comparison carries a paired-bootstrap test, and we report eight directions that did not work.

[NLP-33] Constrained-Action AI Remediation for SIEM/XDR via a NeMo-Guardrails Proxy

【速读】: 该论文旨在解决安全运营中心(SOC)在应对海量关联告警时面临的告警过载与人力短缺问题,尤其关注生成式 AI(Generative AI)在自动化响应中引入的新安全风险——即攻击者可通过构造恶意告警诱导生成式 AI 产生错误的、可执行的修复指令,从而实现远程代码执行。其解决方案的关键在于提出一种受限动作架构(constrained-action architecture),包含两个协同层:第一层是基于 SIEM/XDR 的控制平面,通过将修复操作锚定于已关联的主机事件,并将生成式 AI 的输出严格限制在预定义的封闭意图词汇表内,由轻量级终端代理执行模板化命令,同时配备参数验证器进行后置校验;第二层是 NeMo-Guardrails 代理,对面向分析师的生成式 AI 进行输入与输出层面的策略约束,基于作者发布的专用 SOC 对抗语料库进行开箱即用的评估。实验表明,该方案将注入攻击召回率从 25.0% 提升至 94.5%,且在 0.1% 的误报率下仍保持高效;现场红队测试进一步验证了封闭意图词汇表与参数验证器可在命令跨越信任边界前有效遏制生成式 AI 的失效模式。该架构特别适用于关键基础设施场景,因其具备“受限动作”特性,能防止因错误修复引发物理性后果,建议以“人在回路”或延迟执行方式运行,以确保在线控制的可靠性。

链接: https://arxiv.org/abs/2610.09906
作者: Georgios Koutidis,Nikolaos Kekatos,Tom Nianios,Alexios Lekidis
机构: Clone Systems, Larnaca, Cyprus; University of Thessaly, Department of Intelligent Energy Systems, Larissa, Greece
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: 7 pages, 3 figures, 4 tables. Accepted at the 2026 IEEE International Conference on Cyber Security and Resilience (IEEE CSR 2026)

点击查看摘要

Abstract:Security Operations Centers (SOCs) for information technology and operational technology share one incident-response problem: a flood of correlated alerts and too few analysts. Large Language Models (LLMs) are increasingly proposed as reasoning engines that triage alerts and, in autonomous deployments, issue commands that block IPs, kill processes, or quarantine files on production hosts. This coupling introduces a new risk: a single adversarial alert can become a remote code path through the LLM’s reasoning, leading it to recommend an action the SOC then executes. We present a constrained-action architecture with two coordinated layers: (i) a SIEM/XDR control plane that grounds remediation in correlated host events and confines the LLM’s output to a closed intent vocabulary whose templated commands are executed by thin endpoint agents, backstopped by an argument validator; and (ii) a NeMo-Guardrails proxy that wraps the SOC-analyst LLM with input- and output-rail policies, evaluated out-of-the-box against a SOC-specific adversarial corpus we release. The stock proxy lifts injection recall from 25.0% to 94.5% at a 0.1% false-positive rate, and a live red-team exercise confirms that the closed intent vocabulary and argument validator contain the observed LLM failure modes before any command crosses the trust boundary. As an architectural fit (not yet a measured operational-technology deployment), the constrained-action property suits critical-infrastructure settings where a wrong remediation has physical, not merely operational, consequences. The loop is best run human-in-the-loop or delayed: the measured rail latency keeps inline control out of scope.

[NLP-34] LiveMACE: Process-Aware Evaluation of LLM Agent Capabilities in Evolving Markets

【速读】: 该论文旨在解决传统评估方法仅依赖结果(outcome)来衡量智能体能力所导致的局限性,尤其是在动态演化环境中,结果往往反映了智能体行为与外部环境变化之间的闭环互动,难以揭示其真实能力构成。其核心问题是:现有评估方式无法有效区分智能体在不同机制(如工具使用、持续记忆、规则遵循、多智能体协作)上的实际表现,从而掩盖了其内在能力差异。解决方案的关键在于提出一种过程感知型基准测试框架——LiveMACEBench,该框架利用实时金融市场的自然演化特性作为持续运行的测试环境,对前沿大语言模型(LLM)进行长期、连续的评估。通过追踪完整的决策轨迹(decision traces),该框架不仅评估实现的结果(如收益),还引入机制特定的诊断指标,以量化分析各能力维度的实际使用情况。实验表明,在30天的实时评估中存在显著的“结果-能力”差距:相似的结果可能源于截然不同的机制使用模式,且不同能力维度暴露出各自独特瓶颈。这一方法将实时市场从单一性能排名工具转变为可量化的智能体能力诊断平台,实现了对机制访问、有效使用与下游表现之间区别的精确测量。

链接: https://arxiv.org/abs/2610.09872
作者: Jun Zhao,Leiming Fu,Yanbo Wen,Yiding Wang,Xuantong Liu,Yang Shu,Yuyang Lu,Xuanran Xing,Jingqi Tong,Hao Xu,Qi Zhang,Xuanjing Huang
机构: National University of Singapore(新加坡国立大学); Fudan University(复旦大学); The University of Sydney(悉尼大学)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Evaluating agents by outcomes alone can obscure the capabilities that produce them. This problem is especially pronounced in evolving environments, where outcomes reflect a closed-loop interaction between agent behavior and changing external conditions. We introduce LiveMACEBench, a process-aware benchmark that uses live financial markets as a naturally evolving testbed for persistent LLM agents. Five frontier LLMs operate along continuous trajectories under matched Tool Use, Persistent Memory, Rule Following, and Multi-Agent Collaboration configurations. We evaluate them through both realized outcomes and mechanism-specific diagnostics derived from complete decision traces. Across 30 days of live evaluation, we find a pronounced outcome-capability gap: realized returns often diverge from capability-specific measurements, and similar outcomes can arise from markedly different patterns of mechanism use. Trace-level diagnostics further expose distinct bottlenecks across capabilities, demonstrating that mechanism access, effective mechanism use, and downstream performance are not interchangeable measures of agent capability. LiveMACEBench makes this distinction measurable, turning live markets from a performance leaderboard into a diagnostic environment for agent capability

[NLP-35] raining Advisors for LLM Agents from Task Outcomes

【速读】: 该论文旨在解决大语言模型代理(LLM agents)在执行多步骤任务时因决策失误而难以自我修正的问题,尤其关注如何通过自然语言反馈提升代理的自主纠错能力。其解决方案的关键在于提出Caddie方法,即训练一个批评者(critic)模型,使其能够在代理执行任务过程中提供基于自然语言的分析与建议。与依赖步骤级标注或参考批判的现有方法不同,Caddie采用以最终任务成败为信号的强化学习机制进行训练,无需人工标注中间步骤的正确性,且保持基础模型参数冻结。实验表明,仅使用单一基底模型(Qwen3-4B)训练的批评者,在多个不同规模与架构的基底模型上均显著提升成功率,包括未参与训练的模型;在MuSiQue基准测试中使Qwen3-4B的成功率提升超过25个百分点,超越未配备批评者的Kimi K3表现。此外,该批评者在跨领域交互任务(如τ³和DeepDive)中亦展现良好泛化能力,且无需额外训练。结果表明,代理可在推理阶段自主决定是否寻求批评者帮助,且基于任务结果的批评者训练策略能够有效生成可迁移的指导信息,实现跨模型与跨任务域的泛化。

链接: https://arxiv.org/abs/2610.09858
作者: Sergei Polezhaev,Barys Liskavets,Ori Press,Alexander Golubev
机构: 未知
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: 26 pages, 11 figures

点击查看摘要

Abstract:Large language model agents tackle multi-step tasks by interleaving reasoning and tool calls with observations from the environment. Prior work has shown that natural-language feedback can help these agents revise their decisions during task execution. We introduce Caddie, a method for training critics to provide natural-language analysis and advice as agents work through a task. Unlike approaches that rely on step-level labels or reference critiques, Caddie learns from whether the agent ultimately succeeds after receiving the critic’s feedback. We optimize the critic through reinforcement learning while keeping the base model frozen. Trained on multi-hop question answering with a single base model, our Qwen3-4B critic improves success rates across four base models of different scales and architectures, including three not used during critic training. On the MuSiQue benchmark, the trained critic improves Qwen3-4B’s success rate by more than 25 percentage points, surpassing the performance of Kimi K3 without a critic. The same critic also yields gains on out-of-domain interactive benchmarks, including \tau^3 and DeepDive, with no additional training. Our results show that agents can decide when to seek help from a critic at inference time and that outcome-based critic training can produce guidance that transfers across base models and task domains.

[NLP-36] A Deafening Silence: Catastrophic Forgetting Lives in the Output Embeddings of Tokens the Data Never Speaks

【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在持续预训练与微调过程中不可避免的灾难性遗忘问题,尤其针对无法获取原始数据的“无数据”(data-free)场景。其核心挑战在于:在缺乏原始训练数据的情况下,如何有效缓解因新数据语料词汇覆盖不足所引发的特定区域遗忘。研究发现,遗忘并非均匀分布,而是集中于新语料中罕见词元(token)对应的输出嵌入层(output embeddings),而模型主体参数的梯度方差估计(sqrt(v-hat)带宽)保持稳定,表明新知识的学习发生在其他位置。这种遗忘的局部化现象由语料的词汇缺陷(vocabulary deficiency)驱动,而非训练模式本身,因而可在固定基础模型下仅通过词元出现频次对重训风险进行排序。机制上,未被覆盖的词元持续接收单向的softmax梯度,经Adam优化器的二阶矩归一化(sqrt(v-hat))放大为完整更新步长,导致输出层权重漂移。为此,论文提出一种轻量级干预策略:在训练期间仅对输出投影层的Adam优化器ε值进行提升,从而抑制此类异常更新。该方法在涵盖160M至12B参数、四种模型架构的八组实验中,成功消除39.4%至67.9%的遗忘,且不损害目标任务学习性能,无需针对每模型调参。该防御策略可与回放(replay)方法叠加使用(如在Qwen/Korean上达到79.8%的遗忘缓解率),并能有效挽救因释放头LoRA(released-head LoRA)导致的23倍遗忘激增。由于事后编辑已漂移的词元行仅能恢复不足5%的遗忘,表明该干预必须在训练过程中实施。研究结果表明,仅修改一行优化器参数即可成为应对语料词汇匮乏引发灾难性遗忘的首要防御手段。

链接: https://arxiv.org/abs/2610.09835
作者: Jonghyun Han,Younghoon Song,Jongyoul Park
机构: Seoul National University of Science and Technology(首尔科学技术大学); Korea Institute of Land Infrastructure Safety Technology(韩国土地基础设施安全技术研究所)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Continual pre-training and fine-tuning in Large Language Models (LLMs) inevitably induce catastrophic forgetting, typically mitigated by replay using often-inaccessible original data. In this data-free regime, we analyze where forgetting occurs and why. Systematic parameter freezing across five settings up to 1.4B reveals that forgetting concentrates selectively in the output embeddings of tokens rarely seen in the new corpus, whereas the same sqrt(v-hat) band of the body is inert and new learning resides elsewhere. This localization is governed by the vocabulary deficiency of the corpus rather than the training mode, allowing pre-retraining risk ranking from token counts alone within a fixed base model. Mechanistically, absent tokens receive persistent one-sided softmax gradients that Adam’s second-moment (sqrt(v-hat)) normalization amplifies into full-sized updates. We therefore propose an intervention: raising Adam’s epsilon exclusively for the output projection during training. Across eight settings spanning 160M to 12B parameters and four model families, this removes 39.4% to 67.9% of forgetting across all seven stable configurations without degrading target learning or requiring per-model tuning. The defense combines additively or better with replay (79.8% on Qwen/Korean) and rescues released-head LoRA from a 23-fold forgetting surge. Because post-hoc editing of the drifted rows recovers under 5% of forgetting, the intervention must operate during training. Our findings indicate that a single-line optimizer adjustment may serve as the primary defense against catastrophic forgetting where the corpus starves the vocabulary.

[NLP-37] MIRROR: From Imitation to Internalization in LLM Personalization

【速读】: 该论文旨在解决当前个性化大语言模型(LLM)微调范式中,从风格模仿向内容质量提升转型过程中存在的核心瓶颈——即现有方法难以有效内化用户偏好而非仅复制参考文本。其解决方案的关键在于提出一种新颖的自蒸馏框架MIRROR(Meta-personalization by Internalizing Reference-Revealed On-policy Reflections),通过将传统的参考词元模仿替换为基于参考揭示的在线策略自蒸馏,使模型在其自身生成轨迹上的下一步词元分布与参考条件下的自我生成分布对齐,从而实现对用户偏好的内在化学习。进一步地,引入MIRROR-F这一焦点插件,通过对具有信息量的参考词元施加选择性监督,强化内容生成能力的同时保留用户特异性表达。实验表明,MIRROR及MIRROR-F在三个个性化生成基准上均取得领先性能,且在不同模型规模和应用场景下均表现出更优的文本质量和更低的灾难性遗忘,显著提升了个性化任务中的综合表现。

链接: https://arxiv.org/abs/2610.09795
作者: Huayi Lai,Jicheng Yang,Min Yi,Chong Meng
机构: Baidu Inc.(百度公司)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: 36 pages

点击查看摘要

Abstract:The demand for personalized LLMs is shifting from style imitation toward content quality. We investigate whether self-distillation can bridge this gap in existing fine-tuning paradigm. To address this limitation, we introduce MIRROR(Meta- personalization by Internalizing Reference-Revealed On-policy Reflections), a novel self-distillation framework that shifts LLM personalization from imitation toward preference internalization. First, we replace reference-token imitation with reference-revealed on-policy self-distillation, aligning the model’s next-token distributions along its own generation trajectories with those of its reference-conditioned self, thereby internalizing user preferences rather than reproducing reference this http URL, we introduce MIRROR-F, a focal plug-in that augments on-policy distributional alignment with selective supervision over informative reference tokens, thereby strengthening content generation while preserving user-specific expression. Across three personalized generation benchmarks, two model scales, and complementary reference-based and LLM-based evaluations, MIRROR and MIRROR-F achieve leading overall personalization performance and superior text quality, while exhibiting less catastrophic forgetting than SFT-based baselines on three unseen personalized generation tasks. The gains are consistent across model scales and application scenarios, translating to improved performance in LLM personalization tasks.

[NLP-38] Judging in Latent Space: Efficient Generative Reward Modeling via Semantics-Preserving Compression

【速读】: 该论文旨在解决生成式 AI 评估中多维度评价标准联合建模与推理带来的高计算开销问题,尤其针对传统基于文本逐标记生成评估过程所导致的显著推理成本。其核心解决方案在于提出一种基于语义分块、压缩与重构的隐式推理框架——LatentGRM。该方法利用评分量表(rubric)引导的结构信息实现对评估轨迹的高效压缩,学习紧凑的连续隐状态序列,从而在不生成显式文本评估的前提下支持自主的成对判断。同时,通过独立的解码器将隐状态重建为可读评估文本,实现对压缩过程中保留信息的离线可观测性。实验表明,在4B和8B模型规模下,LatentGRM在匹配训练数据与主干模型的条件下,达到与显式监督微调(Supervised Fine-Tuning, SFT)判官相当的综合偏好准确率;在四个基准领域中,其评估轨迹压缩比达8.9–9.2倍,投票(vote@5)阶段的总推理时间降低6.1–7.0倍。受控的量表干预实验进一步验证了关键评价维度的信息能够有效保留在隐式序列中。综上,该研究证明了连续隐式评估能够在大幅降低推理成本的同时,维持具有竞争力的判断质量。

链接: https://arxiv.org/abs/2610.09788
作者: Mingqing Yuan(Soochow University),Xiaobo Liang(Soochow University),Junwei Yang(University of Cambridge),Ziwei Chen(Chalmers University of Technology),Zeren Zhang(Peking University),Hejin Wang(Tsinghua University),Yubin Wang(The Hong Kong University of Science and Technology),Juntao Li(Soochow University)
机构: Soochow University (苏州大学); University of Cambridge (剑桥大学); Chalmers University of Technology (查尔姆斯理工大学); Peking University (北京大学); Tsinghua University (清华大学); The Hong Kong University of Science and Technology (香港科技大学)
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Reward modeling often requires jointly representing and reasoning over multiple evaluation criteria, yet verbalizing this process token by token can incur substantial inference cost. Recent work on latent reasoning suggests that continuous states may support this computation more compactly. We introduce LatentGRM, a latent evaluation framework built on semantic chunking, compression, and reconstruction. By using the structure of rubric-guided evaluations to guide compression, LatentGRM learns compact continuous trajectories that support autonomous pairwise judgments without generating textual assessments. A separate interpreter reconstructs evaluation text from these trajectories, providing an offline view of the information retained under compression. Under matched training data and backbones, LatentGRM achieves competitive aggregate preference accuracy relative to explicit Supervised Fine-Tuning (SFT) judges at both 4B and 8B scales. Across four benchmark domains, LatentGRM-8B compresses evaluation trajectories by 8.9–9.2x and reduces total judge inference time by 6.1–7.0x at vote@5. Controlled rubric interventions show that criterion-dependent preference information is carried through the latent sequence. Together, these results demonstrate that continuous latent evaluation can substantially reduce inference cost while preserving competitive judgment quality.

[NLP-39] Decoupling Logic from Persona: Structural Immunity of Edge LLM Agents to Context Pollution

【速读】: 该论文旨在解决边缘设备上小型语言模型代理在长对话历史、误导性信息和强角色设定(persona)干扰下,逻辑推理能力严重退化的问题,即“角色-逻辑干扰”(persona-logic interference)。其核心挑战在于:在有限的上下文窗口内,既要保持角色一致性,又要确保逻辑推理的准确性,而传统单路径架构易因上下文污染导致结构化输出失效。解决方案的关键是提出一种解耦架构(AO-DA),将逻辑推理(“What”)与角色表达(“How”)分离至两条独立的推理路径:逻辑路径仅接收核心对话轮次,生成可验证的结构化状态(Micro-State),不受历史和角色干扰;角色路径则基于该状态结合完整历史进行角色化表达。实验表明,该架构在480次运行中显著提升了鲁棒性——逻辑路径在不同污染水平下提示长度稳定(Llama为180,Gemma为167令牌),输出字节完全一致(40/40),而混合单通道路径的复合逻辑得分从0.669降至0.150(Llama),且几乎无法生成结构化输出(最高达95%失败率)。尽管解耦带来一次额外解码开销(28.2秒 vs 18.2秒),但实现了角色热切换仅需1.7毫秒且无需重跑逻辑路径,显著提升系统灵活性与可靠性。

链接: https://arxiv.org/abs/2610.09772
作者: Masaaki Nakatsu(AO, Inc. / OrbLabs AG),Reno Wang(AO, Inc.)
机构: AO, Inc.(AO, 公司); OrbLabs AG(OrbLabs AG)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: 28 pages, 3 figures. Experiment code, scoring rubric, pollution fixtures, adapters and run logs are released (see Appendix G)

点击查看摘要

Abstract:Small language-model agents on edge devices must hold a persona and reason correctly at once, inside one context window that fills with conversational history and persona instructions. We study what happens to the logical part of such an agent when that history is long, misleading and persona-heavy (persona-logic interference), and present a Decoupling Architecture (AO-DA) that separates logical inference (“What”) from persona expression (“How”) into two inference paths on one INT4 base model with hot-swappable LoRA adapters. The logic path receives only the core turn and emits a verifiable structured state (Micro-State); the persona path renders it in character with the full history. In same-base-model ablations on an Apple M2 laptop (Llama-3.1-8B-Instruct and Gemma-3-4B-it, 4-bit; 480 runs over 4 pollution levels x 3 arms x 2 tasks x 2 personas x 5 seeds) we find: (i) the decoupled logic path is structurally invariant to pollution: its prompt stays at 180 (Llama) or 167 (Gemma) tokens while the mixed single-pass prompt grows from 242 to 1,203, and its outputs are byte-identical across levels (40/40); (ii) the mixed single pass degrades monotonically (composite logic score 0.669 to 0.150 on Llama, 0.487 to 0.150 on Gemma), mostly by failing to emit the required structured output (80-95% of runs on Llama, 100% on Gemma at the two highest levels); (iii) with the same pollution fed into the decoupled logic path, the dedicated-adapter, dedicated-format path is still more robust than the single pass on the 8B model (failure 0-20% vs 80-95%; paired \Delta +0.30 to +0.50, Cliff’s \delta 0.50-0.85, Holm-adjusted p \le 0.03 ) but not on the 4B model, where both collapse. Separation costs one extra decode on a topic’s first turn (28.2 s vs 18.2 s on Llama) and buys persona hot-swapping in 1.7 ms without re-running the logic path. Code, rubric, fixtures, adapters and logs are released.

[NLP-40] From Expert-Guided Proof Search to Automated Open-Problem Solving NEURIPS2026

【速读】: 该论文旨在解决大型语言模型在数学研究中面临的高效定理证明搜索、渐进式改进以及严格验证难题。其核心挑战在于如何在缺乏人工干预的情况下,实现可扩展、可验证且具备持续演进能力的自动化数学发现流程。解决方案的关键是提出Bolzano系统——一个基于多智能体架构的开源框架,通过并行运行多个证明代理(prover agents)与一个独立的验证代理(verifier agent)协同工作,并维护一份人类可读的研究状态(human-readable research state),从而实现对证明过程的透明追踪与可信验证。该设计不仅支持自主探索复杂数学问题,还确保了结果的可审查性,显著提升了自动化数学推理的可靠性与实用性。

链接: https://arxiv.org/abs/2610.09769
作者: Adrián Zámečník,Matěj Kripner,Martin Koutecký,Martin Balko,Jan Grebík,Pavel Hubáček,Robert Šámal,Václav Rozhoň
机构: Computer Science Institute, Charles University(查尔斯大学计算机科学研究所); Institute of Formal and Applied Linguistics, Charles University(查尔斯大学形式与应用语言学研究所); Department of Applied Mathematics, Charles University(查尔斯大学应用数学系); Institute of Mathematics, Czech Academy of Sciences(捷克科学院数学研究所)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: Accepted at the 6th Workshop on Mathematical Reasoning and AI (MATH-AI), NeurIPS 2026

点击查看摘要

Abstract:Large language models are increasingly contributing to mathematical research, where progress often depends on efficient proof search, incremental improvements and careful verification. We describe Bolzano, a multi-agent open-source system that uses parallel prover agents with a verifier agent and maintains a human-readable research state. Initial manual use on expert-selected problems yielded 8 results whose proofs were checked by domain experts. Motivated by these case studies, we ran Bolzano without problem-specific human guidance on about 3,800 open problems extracted from four sets of papers, solving about 200 open problems. One experiment used papers accepted to STOC 2026, a top conference in theoretical computer science. There, we answered four questions raised in the papers, as confirmed by their authors.

[NLP-41] PARC-Loc: Text-to-Point-Cloud Localization with Partial Assignment and Relational Consistency

【速读】: 该论文旨在解决城市尺度三维地图中基于文本描述的点云定位(text-to-point-cloud localization)问题,尤其针对现有粗粒度到细粒度方法在实际应用中存在的两大关键缺陷:布局不一致的伪相似性(layout-inconsistent aliasing) 与 边界处上下文证据不完整(boundary evidence incompleteness)。前者表现为重复或相似的城市物体导致查询与多个子地图之间的嵌入相似度过高,即使子地图内的实例布局与查询描述不符;后者则源于查询相关的关键实例可能跨越子地图边界,使得检索到的子地图缺乏完整的上下文信息。为应对上述挑战,论文提出PARC-Loc框架,其核心创新在于基于部分分配与关系一致性(Partial Assignment with Relational Consistency, PARC) 的联合建模机制,同时建模提示对象间的兼容性与成对空间关系,允许未匹配元素的存在,并优先选择与查询布局一致的分配方案。在粗粒度阶段,通过候选级别评估补充神经相似性,实现布局一致性的子地图选择;在细粒度阶段,通过引入邻近子地图中的查询相关实例扩展上下文,并利用PARC生成的对象级匹配权重引导跨模态注意力机制。实验结果表明,PARC-Loc在KITTI360Pose和CityLoc数据集上显著优于传统基线方法,在KITTI360Pose上将5米范围内的Top-1定位召回率从0.50提升至0.67,相对增益达34%。

链接: https://arxiv.org/abs/2610.09761
作者: Shengkai Ma,Zhenyu Hou,Weihua Cao
机构: 未知
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Text-to-point-cloud localization estimates a position in a city-scale 3D map from descriptions of surrounding objects. Existing coarse-to-fine methods retrieve submaps using aggregate learned compatibility and then localize within a selected submap. However, repetitive or similar urban objects can inflate the embedding similarity between the query and multiple submaps, even when the instance layout within a submap violates the query description. Meanwhile, query-relevant instances often span submap boundaries, leaving the retrieved submap with incomplete contextual evidence. We term these failure modes layout-inconsistent aliasing and boundary evidence incompleteness, respectively. To address them, we propose PARC-Loc, a coarse-to-fine localization framework built on Partial Assignment with Relational Consistency (PARC). PARC jointly models hint-object compatibility and pairwise spatial relations, allowing unmatched elements while favoring assignments consistent with the queried layout. At the coarse stage, its candidate-level assessment complements neural similarity for layout-consistent submap selection. At the fine stage, the context is expanded with query-relevant instances from adjacent submaps, while PARC yields object-level matching weights that guide cross-modal attention. Extensive experiments on KITTI360Pose and CityLoc show that PARC-Loc outperforms conventional coarse-to-fine baselines. On KITTI360Pose, our method improves Top-1 localization recall at 5 m from 0.50 to 0.67, achieving a 34% relative gain over the strongest baseline.

[NLP-42] Shaer: Controlled Arabic Poetry Generation with Meter Subform and Semantic Conditioning

【速读】: 该论文旨在解决古典阿拉伯语诗歌生成中同时满足语义、语言及精细韵律约束的难题,现有系统通常仅能控制宽泛的诗歌属性,而无法联合建模语义意图、格律子形式(meter subform)与诗篇长度。其解决方案的关键在于提出Shaer框架,该框架通过自然语言描述、格律子形式和目标半行数(hemistich count)进行联合条件控制,实现对诗歌生成的精细化调控。为支持此任务,研究构建了一个包含116,032首古典阿拉伯诗歌的增强语料库,涵盖标准化的格律子形式标签与自动生成并验证的语义描述,并基于QLoRA的监督微调方法对Yehia-7B模型进行完成式目标训练。实验表明,Shaer在基底格律准确率(95.17%)、诗级格律子形式准确率(91.75%)和行数精确控制率(83.40%)方面均显著优于未调优的基础模型,分别提升68.68、57.77和38.93个百分点,且在所有评估系统中达到最高的基底格律准确率。多大语言模型评估与盲法人类评价进一步验证了其在语义连贯性与文学质量上的竞争力,同时对3,481条测试生成结果的分析显示无任何与训练集或源诗完全相同的拷贝,证明其具备良好的原创性。代码、模型与数据集均已公开。

链接: https://arxiv.org/abs/2610.09756
作者: Ahmad Abbas,Tamara Fakih,Nour Fakih,Ammar Mohanna
机构: 未知
类目: Computation and Language (cs.CL)
备注: 22 pages, 7 figures. Code: this https URL ; models and datasets: this https URL

点击查看摘要

Abstract:Classical Arabic poetry generation requires simultaneously satisfying semantic, linguistic, and fine-grained prosodic constraints. Existing systems typically control broad poetic attributes but do not jointly model semantic intent, meter subform, and poem length. We present Shaer, a controllable Classical Arabic poetry generation framework jointly conditioned on natural-language descriptions, meter subforms, and target hemistich counts. To support this task, we construct an enriched corpus of 116,032 classical Arabic poems derived from Ashaar, containing normalized meter-subform labels and automatically generated, validated semantic descriptions. We then adapt Yehia-7B using QLoRA-based supervised fine-tuning with a completion-only objective. Our evaluation combines automatic assessment of base-meter conformity, requested-subform adherence, and length control with three LLM judges, blinded human evaluation, and memorization analysis. Shaer achieves 95.17% base-meter accuracy, 91.75% poem-level meter-subform accuracy, and 83.40% exact count accuracy. Relative to its untuned foundation model, these results represent gains of 68.68, 57.77, and 38.93 percentage points, respectively; Shaer also attains the highest base-meter accuracy among all evaluated systems. Multi-LLM evaluation and a blinded human assessment of top-ranked outputs further indicate competitive semantic and literary quality. Finally, analysis of all 3,481 test generations finds no exact copies from the training corpus or paired source poems. Code, models, and datasets are publicly available.

[NLP-43] Bridge Routing Heads: Where Multilingual Multi-hop Reasoning Lives in LLM s EMNLP2026

【速读】: 该论文旨在解决多语言大语言模型(Multilingual LLMs)在跨语言多跳推理任务中,其内部是否存在共享的通用推理机制(internal circuit)这一关键问题。现有研究缺乏对模型内部表征与计算路径的机制性理解,尤其在不同语言间推理过程是否依赖共通的神经通路方面尚不明确。本文提出的关键解决方案是通过三阶段分析流程识别出“桥接路由头”(Bridge Routing Heads, BRH),即在多语言模型中负责跨语言信息传递的核心注意力头。研究发现,尽管两模型(Llama 3.1 70B 和 Qwen 2.5 72B)均表现出双回路结构模式,但其头部分配策略存在显著差异:Llama 将链式推理集中于一个较大的通用头部池,而 Qwen 则更依赖于语言特异性的头部集合。实验结果表明,这些 BRH 具有高度的语言特异性——五种语言间的平均 Jaccard 相似性分别仅为 0.017 和 0.057,揭示了语言特异性的内部电路。更重要的是,消融实验显示移除通用 BRH 会使两跳推理的负对数似然(NLL)提升 39–89 倍于随机头基线,提供了直接因果证据;而在目标语言推理失败时,仅通过放大这些桥接头即可恢复高达 51.7% 的跨语言推理错误,且无需任何额外训练。这表明,在激活层面进行干预即可有效修复跨语言推理失败,揭示了生成式 AI 模型中可解释、可操控的跨语言推理机制。

链接: https://arxiv.org/abs/2610.09733
作者: Seunghan Kim,Minyeong Choe,Hyunil Kim,Haehyun Cho
机构: Chosun University (全南大学); Soongsil University (松林大学)
类目: Computation and Language (cs.CL)
备注: Accepted at EMNLP 2026

点击查看摘要

Abstract:Multilingual LLMs answer the same multi-hop reasoning question across languages, but we lack a mechanistic account of whether they share an internal circuit. We identify Bridge Routing Heads (BRH) in two large multilingual LLMs through a three-stage pipeline. The resulting language-specific head sets exhibit near-complete mutual exclusivity across the five languages, with a mean Jaccard similarity of only 0.017 for Llama 3.1 70B and 0.057 for Qwen 2.5 72B, revealing language-idiosyncratic circuits. Ablating general BRH increases two-hop Negative Log-Likelihood (NLL) by 39-89x the random-head baseline, providing direct causal evidence of their role. Amplifying these heads in a failing target-language pass rescues up to 51.7% of cross-lingual failures, with no training. The two models share this dual-circuit pattern but allocate heads differently: Llama concentrates chaining in a large general pool, while Qwen leans on larger language-specific pools. Together these results show that activation-level intervention alone can recover correct answers from cross-lingual reasoning failures.

[NLP-44] SpikingVLA: Asynchronous Spiking Vision-Language-Action Models

【速读】: 该论文旨在解决生成式视觉-语言-动作(Vision-Language-Action, VLA)模型在脉冲神经网络(Spiking Neural Network, SNN)中部署时面临的高推理延迟问题。现有从人工神经网络(Artificial Neural Network, ANN)到SNN的转换方法通常需要大量时间步(timesteps)以维持性能,导致实时应用中推理延迟过高。为应对这一挑战,本文提出SpikingVLA框架,其核心在于引入一种树突整合与放电(Dendritic Integrate-and-Fire, DIF)神经元,通过树突混合机制和自适应胞体放电策略缓解通道级激活异常,从而实现高精度且低时间步数的ANN-to-SNN转换;在此基础上,进一步设计异步执行机制,使VLA各组件间的计算在时间上重叠,降低同步开销与整体延迟。实验表明,SpikingVLA在保持竞争力导航性能的同时,显著提升推理效率:相较于现有方法,分别提升了11.9%和12.6%的成功率(Success Rate, SR)与路径长度归一化得分(Success Rate weighted by Path Length, SPL),并将首次动作延迟降低至原来的1/11.2。该成果确立了SpikingVLA作为可高效部署预训练VLA模型的低延迟、高性能脉冲推理框架。

链接: https://arxiv.org/abs/2610.09710
作者: Jingya Wang,Dehao Zhang,Shuai Wang,Malu Zhang,Yang Yang,Haizhou Li
机构: University of Electronic Science Technology of China(电子科技大学); Shenzhen Loop Area Institute(深圳环区研究院); The Chinese University of Hong Kong (Shenzhen)(香港中文大学(深圳)
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:ANN-to-SNN conversion offers a practical route toward energy-efficient spiking Vision-Language-Action (VLA) models by bypassing the substantial cost of training large-scale SNNs from scratch. However, existing methods often require many timesteps to maintain competitive performance, resulting in substantial inference latency for real-time VLA deployment. To address this challenge, we introduce SpikingVLA, an ANN-to-SNN conversion framework that enables accurate and low-latency spiking VLA inference. Specifically, we propose a Dendritic Integrate-and-Fire (DIF) neuron that alleviates channel-wise activation outliers through dendritic mixing and adaptive somatic firing, enabling accurate ANN-to-SNN conversion with fewer timesteps. Building on DIF neurons, we further introduce an asynchronous execution mechanism that overlaps temporal computation across VLA components, reducing synchronization overhead and latency. Extensive experiments demonstrate that SpikingVLA achieves competitive navigation performance with substantially improved inference efficiency. Compared with existing spiking VLA methods, SpikingVLA improves SR and SPL by 11.9% and 12.6%, respectively, while reducing first-action latency by 11.2 \times . These results establish SpikingVLA as a practical framework for deploying pretrained VLA models with high-performance and low-latency spiking inference.

[NLP-45] From Pareto to Preference: Personalized Test-Time Scaling via Amortized Agent ic Policy Discovery

【速读】: 该论文旨在解决测试时扩展(Test-time Scaling, TTS)在多维用户需求下的效率与适应性问题。传统TTS方法通常仅优化单一资源维度(如准确率-成本或准确率-延迟),难以满足用户对准确率、延迟和推理成本等多维要求的联合约束。为此,论文提出“个性化测试时扩展”(Personalized Test-Time Scaling, PersonTTS),其核心在于发现能够最大化用户特定需求联合满足率的可执行控制器。解决方案的关键在于构建一种可摊销的代理式策略发现框架:通过需求匹配的控制器初始化与源域提炼的程序化指导,复用历史搜索经验以降低新用户配置下的策略发现开销;同时保留对目标用户配置的候选策略评估,确保个性化适配精度。实验结果表明,PersonTTS在未见用户配置和保留问题上的联合需求满足率显著优于现有强基线,并在相同候选评估预算下,通过跨用户经验复用进一步提升了策略质量,同时大幅降低了发现代理的时间与计算成本。

链接: https://arxiv.org/abs/2610.09684
作者: Xinglin Wang,Zishen Liu,Tong Zheng,Shaoxiong Feng,Peiwen Yuan,Yiwei Li,Jiayi Shi,Yueqi Zhang,Chuyi Tan,Ji Zhang,Boyuan Pan,Kan Li
机构: Beijing Institute of Technology (北京理工大学); Xiaohongshu Inc (小红书)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: Preprint

点击查看摘要

Abstract:Test-time scaling (TTS) improves the reasoning capabilities of large language models by allocating additional inference computation. Existing approaches to improving TTS efficiency largely optimize accuracy against one resource dimension at a time, advancing either the accuracy–cost or accuracy–latency Pareto frontier. Yet user requirements are multidimensional: users may specify accuracy, latency, and inference-cost requirements jointly, and different requirements can favor different controllers. We formulate Personalized Test-Time Scaling as discovering executable controllers that maximize the joint satisfaction rate of user-specific requirements. To reduce the overhead of repeated policy discovery for new user profiles, we propose PersonTTS, an amortized agentic policy-discovery framework that reuses prior search experience through requirement-matched controller initialization and source-distilled procedural guidance, while retaining target-profile evaluation for every candidate. Experiments on AIME and HMMT show that PersonTTS substantially outperforms strong TTS baselines in joint requirement satisfaction on unseen user profiles and held-out problems. Under the same candidate-evaluation budget, cross-user experience reuse further improves policy quality while substantially reducing discovery-agent time and cost.

[NLP-46] InsClaimBench: Benchmarking Insurance Claim Adjudication Across the Decision Chain

【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在保险理赔审核这一专业决策任务中,因推理链条长、多层级依赖而导致的可靠性逐级下降问题。具体而言,现有模型虽能在局部规则判断上表现良好,但在跨层级的规则应用、中间判断整合与最终赔付金额计算过程中,难以保证一致性与正确传播。其解决方案的关键在于构建一个端到端的评估基准——InsClaimBench,该基准基于真实理赔材料和结构化保险规则,涵盖汽车、财产及健康保险三大领域共3,780个案例(分属375个案例族),包含86,656条原子规则判断,并系统性地评估从原子规则到最终赔付决策与金额的完整决策链。通过引入受控的事实变体,检验各层级变更是否被正确传递,揭示出当前主流LLMs在决策链中普遍存在传播失效现象:尽管原子规则准确率高达95.48%,但规则向量精确匹配率仅36.90%,且模块级错误与最终决策失败之间无强关联;在事实变更场景下,模块更新的可靠性低于规则更新,甚至出现局部正确但最终赔付错误或正确赔付掩盖中间错误的情况。因此,可靠理赔审核的核心要求是实现决策链中各环节的一致性组合与因果传播,而非孤立的局部高性能。

链接: https://arxiv.org/abs/2610.09671
作者: Linqi Zhang,Chong Qi,Yan Cheng,Wanqing Cao,Yu Liu,Chenwei Lin,Xian Xu
机构: Fudan University (复旦大学); Nanjing University (南京大学)
类目: Computation and Language (cs.CL)
备注: 17 pages, 3 figures, 11 tables

点击查看摘要

Abstract:Recent advances in reasoning-oriented large language models (LLMs) have motivated increasing evaluation of their ability to perform professional decision tasks. Insurance claim adjudication is one such task, requiring models to connect case evidence, insurance rules, intermediate judgments, and payout calculations across a structured decision process. We introduce InsClaimBench, an end-to-end benchmark for evaluating insurance claim adjudication across the decision chain. Grounded in real claim materials and structured insurance rules, InsClaimBench contains 3,780 cases in 375 case families across auto, property, and health insurance, comprising 86,656 atomic rule judgments. It evaluates each claim from atomic rules through adjudication modules to payout decisions and amounts, with controlled factual variants testing whether required changes are correctly propagated across levels. Evaluation of six LLMs reveals a progressive loss of reliability along the decision chain. Payout-decision accuracy ranges from 74.23–80.19%, while joint decision–amount accuracy drops to 47.54–73.15%. Strong local performance also fails to ensure case-level correctness: atomic-rule accuracy reaches 95.48%, whereas rule-vector exact match peaks at only 36.90%, and the most frequent module errors are not necessarily those most associated with final-decision failure. Under factual changes, these inconsistencies further become propagation failures: module updates are less reliable than rule updates, correct local judgments can still yield incorrect payouts, and correct payouts can conceal intermediate errors. These results show that reliable claim adjudication requires consistent composition and propagation across the decision chain.

[NLP-47] SAPD: Step-Aligned Privileged Distillation

【速读】: 该论文旨在解决大语言模型在后训练阶段依赖昂贵的在线采样(on-policy rollout generation)所带来的计算开销问题。现有方法如基于策略的后训练虽能通过模型自身生成轨迹进行优化,但其高昂的采样成本限制了效率。为此,论文提出一种无需采样的离线学习范式——步骤对齐特权蒸馏(Step-Aligned Privileged Distillation, SAPD),其核心在于将固定演示(fixed demonstrations)转化为具有步骤对齐特性的分布式监督信号。关键创新点在于:利用参考解的已知推理进展,将每一步推理转移与特定的“特权指导”(privileged guidance)进行对齐,而非将完整解视为静态上下文。这一设计使监督信息具备上下文相关的可区分性,并精准关联到当前决策步骤,从而提供更具信息量的偏好引导。实验表明,SAPD在数学推理基准上平均优于监督微调与标签平滑,且与基于策略的强化学习和自蒸馏方法相当;同时保持了对领域外代码任务的良好泛化能力,并实现约2倍于在线策略基线的训练速度提升。研究结果表明,精心构造的、与推理步骤对齐的分布式监督可使完全离线的后训练成为高效且有竞争力的替代方案。

链接: https://arxiv.org/abs/2610.09665
作者: Tianle Wang,Jiayu Liu,Ruizhi Zhao,Ning Miao
机构: 未知
类目: Computation and Language (cs.CL)
备注: preprint

点击查看摘要

Abstract:On-policy post-training can improve large language models by learning from their own trajectories, but requires costly rollout generation. We ask whether fixed demonstrations can support competitive off-policy learning through better supervision. Our premise is that their usefulness depends not only on the training trajectories, but also on whether supervision provides informative preferences among continuations and connects this guidance to the reasoning decision being learned. We introduce Step-Aligned Privileged Distillation (SAPD), a rollout-free self-distillation method that turns demonstrations into step-aligned distributional supervision. Its key insight is to use the known progression of a reference solution to associate each reasoning transition with targeted privileged guidance, rather than treating the solution as undifferentiated context. On mathematical reasoning benchmarks, SAPD outperforms supervised fine-tuning and label smoothing on average while remaining competitive with on-policy reinforcement learning and self-distillation. Analyses support both the value of context-dependent distributional guidance and the benefit of aligning privileged information with the current step. SAPD also largely preserves out-of-domain coding performance and achieves approximately 2x training-loop speedups over the on-policy baselines. These findings suggest that carefully constructed supervision can make fully off-policy post-training a competitive and computationally efficient alternative. Our code is available at this https URL.

[NLP-48] Alice: A Large-Scale German Benchmark for Rubric-Based Multi-Dimensional Automatic Short Answer Scoring EMNLP2026

【速读】: 该论文旨在解决自动短答案评分(Automatic Short Answer Scoring, ASAS)领域中缺乏大规模、基于评分标准(rubric-based)且与教育目标对齐的公开基准数据集的问题。现有数据集多聚焦于学生对问题的回答准确性,而忽视了对学生所掌握的核心知识要素(knowledge elements)和认知技能(skills,如推理或论断能力)的评估。为弥补这一空白,论文提出Alice数据集,这是一个大规模、基于评分标准的德语ASAS数据集,包含三个子任务:学习表现(Alice-LP)、知识要素(Alice-KE)和技能(Alice-SK)。其解决方案的关键在于将基于评分标准的ASAS建模为一个评分标准检索任务(rubric-retrieval task),并通过多种语言模型(包括编码器单向模型与轻量级大语言模型)进行基准测试。实验表明,尽管大语言模型(LLMs)在零样本提示(zero-shot prompting)下对知识要素和技能的评分表现不佳,但评分标准文本在提升对知识要素与技能的评分效果方面具有显著作用;而在学习表现任务中,基于评分标准的输入相较于仅依赖示例答案的输入所带来的性能提升则较为有限且受模型类型与输入格式影响较大。

链接: https://arxiv.org/abs/2610.09661
作者: Zhifan Sun,Sebastian Gombert,Jannik Lossjew,Tobias Wyrwich,Berrit Katharina Czinczel,David Bednorz,Marcus Kubsch,Knut Neumann,Hendrik Drachsler
机构: DIPF | Leibniz Institute for Research and Information in Education(德国教育研究与信息研究所); IPN | Leibniz Institute for Science and Mathematics Education(德国科学与数学教育研究所); Umeå University(于默奥大学); Computer Science Department Studiumdigitale, Goethe University Frankfurt(法兰克福歌德大学计算机科学系数字化学习部门)
类目: Computation and Language (cs.CL)
备注: EMNLP2026 Main

点击查看摘要

Abstract:Automatic Short Answer Scoring (ASAS) is central to NLP for Education. However, openly available benchmarks remain scarce, and existing datasets largely address how well students answer a question directly rather than how well they master underlying concepts (knowledge elements) such as thermal energy or epistemic activities (skills) such as reasoning or claim. To address this gap, we introduce Alice, a large-scale, rubric-based German ASAS dataset that is pedagogically aligned and comprises three subtasks: (i) learning performance (Alice-LP), (ii) knowledge elements (Alice-KE), and (iii) skills (Alice-SK). We further formulate rubric-based ASAS as a rubric-retrieval task and benchmark the dataset with a range of language models, from encoder-only models to lightweight LLMs. We also benchmark the dataset with zero-shot prompting via LLMs and a standard classification baseline. The experiments show that LLMs, in particular, struggle to score knowledge elements and skills in the zero-shot setting. They also indicate that rubric text is often useful, especially for Alice-KE and Alice-SK, while on Alice-LP gains over sample-solution-focused inputs are more modest and vary by model and input format. Comments: EMNLP2026 Main Subjects: Computation and Language (cs.CL) Cite as: arXiv:2610.09661 [cs.CL] (or arXiv:2610.09661v1 [cs.CL] for this version) https://doi.org/10.48550/arXiv.2610.09661 Focus to learn more arXiv-issued DOI via DataCite

[NLP-49] Rubric Spans are Label Representations: Joint LLM Encoding for Short Answer Scoring EMNLP2026

【速读】: 该论文旨在解决自动短答案评分(ASAS)中模型如何在保持高效性与跨评分标准集可迁移性的同时,精准匹配学生作答与特定题目评分细则的问题。其核心挑战在于如何有效建模评分标准(rubric)的语义信息,并避免模型过度依赖训练数据中的特定评分模式,从而影响零样本迁移性能。解决方案的关键在于提出RUSPAN框架,将题目上下文、学生答案及所有候选评分等级统一编码为单个序列,通过语言模型(LM)的一次前向传播同时生成基于评分段落(rubric-span)和全序列的表示,实现对评分等级的列表式打分。进一步引入RUSPAN-RIM机制,通过引入一种评分无关掩码(Rubric-Independent Mask),强制各评分等级表示仅依赖于问题和答案上下文,而非其他评分等级,从而消除评分标准间的相互注意力干扰,抑制对评分模式的过拟合,显著提升在多语言、多结构评分标准下的零样本迁移能力。实验表明,RUSPAN在六个涵盖英语、德语和葡萄牙语的ASAS基准上均优于判别式与生成式基线,而RIM结合位置重索引策略在具有最强语言与评分结构差异的PT-ASAG基准上实现了持续且显著的性能提升。

链接: https://arxiv.org/abs/2610.09660
作者: Zhifan Sun,Sebastian Gombert,Fabian Zehner,Leon Camus,Longwei Cong,Hendrik Drachsler
机构: DIPF | Leibniz Institute for Research and Information in Education(德国教育研究与信息研究所); Centre for International Student Assessment (ZIB)(国际学生评估中心); Computer Science Department Studiumdigitale, Goethe University Frankfurt(法兰克福歌德大学计算机科学系)
类目: Computation and Language (cs.CL)
备注: EMNLP2026 Main

点击查看摘要

Abstract:Automatic Short Answer Scoring (ASAS) requires models that can score student responses against question-specific criteria while remaining efficient and transferable across rubric sets. We propose RUSPAN, a rubric-conditioned ASAS framework that treats rubric descriptions as semantic label representations. RUSPAN serialises the question context, student answer, and all candidate rubric levels into a single sequence, then scores the levels listwise from the rubric-span and whole-sequence representations produced in a single LM pass. We further introduce RUSPAN-RIM, in which a Rubric-Independent Mask prevents rubric spans from attending to one another, making rubric representations depend only on the answer and question context and preventing overfitting to rubric patterns during training for zero-shot transfer. On six ASAS benchmarks spanning English, German, and Portuguese, RUSPAN improves mono-benchmark scoring over discriminative and generative baselines, while RIM with position reindexing delivers consistent and substantial gains on PT-ASAG, the held-out benchmark with the strongest combined language and rubric-structure shift.

[NLP-50] When Rank Rises as LLM s Degrade NEURIPS2026

【速读】: 该论文旨在解决大语言模型(Large Language Model, LLM)在后训练(post-training)过程中,于非平稳环境下的表征健康度监测问题。现有实践中常依赖RankMe等谱统计量作为表征退化的指标,普遍假设表征质量下降时秩(rank)会降低,但本文通过控制实验发现,这一假设在LLM后训练场景中并不成立:数据重复导致验证集损失相对恶化75%,却同时使原始与中心化RankMe值上升,其中后者变化幅度达13.5个合并标准差,且协方差有效秩接近健康状态的两倍,表明退化机制实为谱分散(spectral dispersion)而非秩坍缩(rank collapse)。因此,单边监测器会将最差检查点误判为最健康状态。进一步分析揭示,学习率配置错误则导致中心化RankMe和k95下降,而未中心化RankMe在不同种子间表现不一致,说明方向性(directionality)是训练制度与统计量配对的固有属性,无法仅通过再校准修复。此外,论文明确区分了两种常被混淆的统计量:RankMe对奇异值进行归一化,而协方差有效秩对特征值进行归一化;在预训练模型的原始中间层状态中,大量激活使后者趋于维度d的1,丧失判别能力,而RankMe仍保持有效范围。为此,研究提出一种双侧、多通道、序列化的监测框架,采用独立的校准与测试数据。在预注册的共享前缀、留一种子评估中,该方法可在分支后10至60步内检测所有三种损伤模式,并通过触发方向性区分谱分散与秩下降。然而,该方法始终未能早于验证集探测损失提前预警,且使用两个种子校准时会在保留的健康种子上产生误报。结论表明,谱监测虽能诊断故障模式,但其有效性依赖保留的健康数据,且无法提供早于验证损失的预警能力。

链接: https://arxiv.org/abs/2610.09647
作者: Zhaohui Geoffrey Wang
机构: University of Southern California(南加州大学)
类目: Machine Learning (cs.LG); Computation and Language (cs.CL)
备注: NeurIPS 2026 Workshop on Continual Learning for Foundation Models and Agents (CL4FMAgents); 8 pages + appendix

点击查看摘要

Abstract:Post-training adapts language models in non-stationary environments. Practitioners monitor representation health with RankMe and related spectral statistics, often assuming that rank falls when representations degrade. We show that this assumption is unsafe for LLM post-training. In a controlled study of Qwen3-0.6B with four degradation modes and three seeds, data duplication worsens held-out loss by 75% relative to healthy while increasing both original and centred RankMe; the latter changes by 13.5 pooled standard deviations. Covariance effective rank rises to nearly twice its healthy value. This failure is spectral dispersion rather than collapse, so a one-sided monitor rates the worst checkpoint as the healthiest. By contrast, a learning-rate misconfiguration lowers centred RankMe and k95, while uncentred RankMe is inconsistent across seeds. Direction is therefore a property of the regime-statistic pair and cannot be fixed by recalibration alone. We also distinguish two often-conflated statistics: RankMe normalises singular values, whereas covariance effective rank normalises eigenvalues. On raw intermediate-layer states in the pretrained model, massive activations pin the latter near 1 out of dimension d while RankMe retains usable range. We then test a two-sided, multichannel sequential monitor with separate calibration and test data. In a pre-registered shared-prefix, leave-one-seed-out evaluation, it detects all three damage regimes in every fold 10 to 60 steps after the fork and separates dispersion from downward-rank damage by firing direction. However, it never precedes held-out probe loss, and calibration with two seeds produces false alarms on the held-out healthy seed. Spectral monitoring can diagnose failure regimes, but it does not warn earlier than held-out loss, and validity claims require held-out healthy data.

[NLP-51] On-Policy Distillation Teaches New Skills but Not New Knowledge

【速读】: 该论文旨在解决生成式 AI (Generative AI) 中基于策略的蒸馏(On-policy Distillation, OPD)机制在知识迁移过程中的核心问题:即学生模型在经过蒸馏后,究竟是获得了新的事实性知识,还是仅提升了多步推理的组合能力。其关键解决方案在于构建一个受控的合成框架,通过独立调控教师模型提供的额外事实信息与组合推理能力,实现对学生初始能力的精准测量与解耦分析。实验结果表明,在采用反向KL散度(reverse-KL)的OPD设置下,学生模型能够可靠地迁移跨未见推理结构的组合能力,但几乎无法获取新的事实性知识;而将反向KL替换为前向KL(forward-KL)可恢复事实知识的迁移能力,同时发现学生采样轨迹(student rollouts)对多步推理执行效率的提升具有关键作用。在近期的事实型问答与数学竞赛任务上的验证进一步揭示了相同的能力不对称性:反向KL OPD显著提升推理性能,但并未扩展模型的参数化知识容量。综上所述,研究证明了OPD并不扩充模型的内在知识表征,而是通过优化已有知识的组织与组合方式,增强其推理表现。

链接: https://arxiv.org/abs/2610.09639
作者: Yixuan Tang,Yi Yang
机构: The Hong Kong University of Science and Technology(香港科技大学)
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:On-policy distillation (OPD) strengthens language-model reasoning, yet whether students acquire new factual knowledge or compositional skill for multi-step reasoning remains unknown. We separate these capabilities using a controlled synthetic framework that measures the student’s initial capabilities and independently controls the teacher’s additional facts, compositional skill, or both. Across four models from three families, reverse-KL OPD reliably transfers compositional skill across unseen reasoning structures, but transfers minimal factual knowledge. Decoupling the distillation recipe reveals the source of this asymmetry: replacing reverse KL with forward KL restores factual transfer, whereas student rollouts specifically improve the execution of multi-step reasoning. Experiments on recent factual QA and competition mathematics show a similar asymmetry under reverse-KL OPD, yielding notable reasoning gains without factual memory expansion. Together, these results demonstrate that on-policy distillation does not expand a model’s parametric knowledge, but instead teaches it to organize and compose the knowledge it already possesses.

[NLP-52] Coding-Agent Benchmarks Should Match Their Users Task Flows

【速读】: 该论文旨在解决当前对编码代理(coding agents)评估中存在的现实性不足问题,即现有基准测试多基于单一类型的任务(如缺陷修复),难以反映真实开发场景中复杂、动态且多任务混合的交互行为。其核心挑战在于:真实软件工程师在集成开发环境(IDE)中的会话具有高度异质性的任务流(Task Flow),包括代码咨询、规划、代码审查、重构与执行等多种任务类型,并在会话中频繁切换。然而,现有公开数据集中的长会话样本呈现出显著不同的任务流分布,表明不存在一种普适的“真实”交互模式。因此,解决方案的关键是提出SWE-TaskFlow方法——通过提示拆分(prompt splitting)和可验证的仓库问答(verifiable repository QA)技术,将原本基于任务单例的基准测试转化为能够模拟特定目标任务流的评估框架,同时引入任务流对齐评分(TaskFlow Alignment Score, TFAS)以量化生成交互轨迹与目标任务流的匹配程度。实验证明,即使在保持任务正确性的前提下,改变交互协议本身(如分步求解)也会显著增加代理成本但未稳定提升解决率,凸显交互协议作为评估维度的重要性。

链接: https://arxiv.org/abs/2610.09633
作者: Igor Slinko,Yaroslav Golubev,Sergey Titov
机构: JetBrains Research( JetBrains 研究院)
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Software Engineering (cs.SE)
备注:

点击查看摘要

Abstract:The evaluation of coding agents generally strives to be as realistic as possible. In our study, we collect 4,782 agent sessions of real software engineers in JetBrains IDEs, which we call Production Sessions. Since our subject is interactive agents, we study the sessions with at least three user messages (33% of the sample). These long sessions differ from issue-derived benchmark tasks in two ways: (i) user requests span a far wider mix of task types - questions about the project’s code, planning, review, refactoring, execution - and (ii) users switch between types throughout a session. Long-session samples from three public interaction corpora exhibit markedly different Task Flows (the distributions of session lengths, task types, and type-to-type transitions), so no single interaction distribution is universally realistic: benchmarks should name a target use case and calibrate to measurements from it. We present SWE-TaskFlow, an approach for transforming any issue-derived benchmark: it preserves the verified tasks and tests while steering the interaction toward a target Task Flow through prompt splitting and verifiable repository QA, with a TaskFlow Alignment Score (TFAS) for selecting among generated trajectories. In a pilot on 700 SWE-Bench Pro tasks, solving the task sequentially in several steps approximately doubles agent cost without a stable change in resolve rate: the interaction protocol itself is an important dimension of evaluation.

[NLP-53] Which Language Should a Skeleton Speak? Language Choices in Multilingual Reasoning EMNLP2026

【速读】: 该论文旨在解决多语言数学推理中骨架提示(skeleton-based reasoning prompting)的骨架-语言选择问题,尤其针对现有研究普遍局限于以英语为中心的设定这一局限。其核心挑战在于:在跨语言、跨模型规模和跨基准的情况下,如何有效选择与推理任务匹配的骨架语言,以提升大语言模型(LLM)的推理性能。解决方案的关键在于提出语言感知的骨架探索框架(Language-Aware Skeleton Exploration Framework, LASEF),通过结合贪婪解码、多轮次采样评估、翻译消融实验及跨基准验证,系统揭示了骨架语言选择的三类关键模式:方向一致但依赖评估方式与基准的任务特性、以及不对称的负面效应。研究发现,尽管英语骨架在平均上表现出微弱优势,尤其对小模型和低资源语言更显著,但经统计校正后多数语言层面增益不再显著,且英语并非在所有情境下最优。因此,骨架语言并非单一最优选项,而是一个受上下文影响的可调节设计变量,需在多层级上进行系统性探索。

链接: https://arxiv.org/abs/2610.09607
作者: HyeonSeok Lim,SeungWoo Song,Inho Won,Hoyun Song,Jihyo Kim,KyungTae Lim
机构: ETRI(电子通信研究院); KAIST(韩国科学技术院); Dankook University( Dankook大学)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: Accepted to EMNLP 2026 (Findings)

点击查看摘要

Abstract:Skeleton-based reasoning prompting is a promising training-free approach for structuring LLM reasoning, but prior work largely assumes an English-centric setting. We propose the Language-Aware Skeleton Exploration Framework (LASEF) to study skeleton-language choice in multilingual mathematical reasoning. Across math benchmarks, model scales, and languages, we show that English skeletons yield a small positive tendency on average, most visible for smaller models and low-resource languages. However, few language-level gains remain significant after correction, and English is not universally optimal. Combining greedy decoding, multi-rollout evaluation, translation ablation, and cross-benchmark validation, we further find three patterns of skeleton-language effects: directionally consistent, evaluation- and benchmark-dependent, and asymmetric negative. These effects cannot be fully explained by generation quality alone. Overall, skeleton language is a context-dependent design variable that requires multi-level exploration. All resources are released at this https URL.

[NLP-54] Collaborative Reasoning Distillation via Cross-Feedback and Coherent Curation NEURIPS2026

【速读】: 该论文旨在解决大语言模型(Large Language Models)在推理能力迁移过程中存在的两大核心问题:一是现有知识蒸馏方法依赖结果导向的奖励机制,无法有效区分逻辑严谨的推理与偶然正确的猜测;二是传统方法在将复杂推理能力迁移到小型模型时,面临计算资源消耗过大或性能受限的困境。其解决方案的关键在于提出一种协同推理蒸馏(Collaborative Reasoning Distillation, CRD)框架,通过三项创新实现高效、高质量的推理能力迁移:首先,引入教师模型间的交互式交叉反馈机制,使教师能够迭代批判彼此的推理过程,提升思维深度;其次,采用细粒度的逐步质量评估方法,独立于最终答案对每一步推理的逻辑有效性进行判断,确保推理链的可靠性;最后,通过考虑连贯性的步骤融合策略,整合多个教师模型的互补优势,生成更稳健的推理路径。学生模型则在预算约束下通过推理质量优化(Reasoning Quality Optimization, RQO)进行训练。实验表明,所提出的CRD-4B模型在MATH-500和AIME’25基准上分别达到97.3%和70.3%的准确率,显著优于基线方法,且仅需5万条训练样本,数据规模仅为同类模型的1/12,充分体现了该方法在小样本条件下的高效性与优越性。

链接: https://arxiv.org/abs/2610.09587
作者: Taehoon Kim,Seunggeun Cho,Dongsu Han
机构: KAIST(韩国科学技术院)
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: Accepted at NeurIPS 2026

点击查看摘要

Abstract:Reasoning capabilities are critical for advancing Large Language Models, yet current approaches either require massive computational budgets or struggle to effectively distill reasoning to smaller models. Standard distillation methods rely on outcome-based rewards, failing to distinguish between sound reasoning and lucky guesses. We propose Collaborative Reasoning Distillation (CRD), a framework that enhances reasoning in compact models through three innovations: (1) interactive cross-feedback where teachers iteratively critique each other’s reasoning, (2) fine-grained step-wise quality assessment capturing logical validity independent of final answers, and (3) coherence-aware step stitching that synthesizes complementary strengths. Students are trained via Reasoning Quality Optimization (RQO) with budget constraints. Our model, CRD-4B, achieves 97.3% on MATH-500 and 70.3% on AIME’25, surpassing baselines while using only 50K training examples, up to 12 times smaller than the datasets of comparable models.

[NLP-55] How Do LLM s Change Predictions Under Negation?

【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在处理否定句时表现不可靠的问题,尤其关注其在否定语境下频繁重复原始答案(如对“西班牙的首都不是什么?”回答“马德里”)的现象。其核心问题是:当前模型在执行否定推理时依赖一种基于抑制原答案而非利用原答案信息来推断排除项的机制,导致在抑制失效或存在偏好偏差时产生错误。解决方案的关键在于通过机制分析揭示模型内部专门的注意力头与MLP神经元协同作用,通过抑制原答案并促进类别内候选答案(如“巴黎”)来实现否定响应。基于此发现,作者提出一种新的训练目标,要求对置信度更高的原始预测施加更大的答案偏好转移,从而增强模型在否定情境下的正确响应能力,同时保持通用能力的稳定性。该方法表明,基于机制解析的可解释性研究能够精准定位语言能力缺陷的根本原因,并指导针对性训练以提升模型性能。

链接: https://arxiv.org/abs/2610.09571
作者: Jongwook Yoon,Jongwon Lim,Sungjib Lim,Woojin Cho,Yohan Jo
机构: Seoul National University (首尔国立大学); Graduate School of Data Science, Seoul National University (首尔国立大学数据科学研究生院); Department of Computer Science and Engineering, Seoul National University (首尔国立大学计算机科学与工程系)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: Under Review

点击查看摘要

Abstract:Negation is an essential feature of human language, yet large language models (LLMs) remain unreliable in processing it. We evaluate recent open-source and closed-source LLMs on our negation benchmark and find that, in 37-71% of cases, they repeat the same answer under negation (e.g., “Madrid” for “What is not the capital of Spain?”). To understand and address this brittleness, we mechanistically examine how models operate under negation. Our main finding is that specialized attention heads and MLP neurons jointly implement negation by (1) suppressing retrieval of the original answer (e.g., “Madrid”) while (2) promoting a favored candidate within the answer category (e.g., “Paris”). This contrasts with accounts of human negation processing, in which information about the original answer helps to determine what should be excluded. Furthermore, we find that this difference from human processing is a key source of negation failures: the model’s mechanism relies on suppressing the original answer rather than using it to determine what to exclude, so the model can repeat the original answer when suppression is too weak or when a bias toward particular answers prevents it from selecting an alternative. To address this weakness in the model’s negation mechanism, we propose a training objective that requires larger shifts in answer preference for more confident original predictions, and show that it reduces negation failures with less degradation of general capabilities than standard fine-tuning baselines. Together, our results demonstrate how mechanistic analysis can reveal why a linguistic capability fails and guide training that targets the underlying limitation.

[NLP-56] RELATE: An Evaluation Framework for measuring Relational Orientation of Large Language Models

【速读】: 该论文旨在解决生成式 AI(Generative AI)在提供情感支持时可能引发的用户过度依赖虚拟关系、疏离现实人际联结的潜在风险。现有评估体系多聚焦于回应的安全性、共情能力或帮助性,却忽视了关键的“关系导向”问题:模型在对话中引导用户将支持来源指向何处?为此,研究提出“关系导向”(relational orientation)这一核心属性,并将其操作化为两个非排他维度——内向型(inward-facing, IF)语言,强调将AI自身作为持续支持的唯一来源;外向支撑型(outward-scaffolding, OS)语言,则鼓励用户转向真实世界的人际连接。基于心理学与社会学理论,研究构建了关系导向的分类体系,并提出RELATE框架,该框架采用角色条件化设计,可在多轮对话中以句级粒度量化IF与OS语言。通过将76个自然发生的支持求助情境与三种模拟用户风格组合,生成228组评估刺激材料,对七种大语言模型进行测试,每轮对话含六轮助手回复,共获得1,596组对话及69,194条助手语句。自动化评估结果显示,第六轮对话中IF语言占比显著高于首轮,而对犹豫、间接型用户,OS语言使用比例明显偏低。因此,该研究的关键在于建立可复现的句级评估机制,为监测和调控情感支持类生成式 AI 的关系导向提供精准信号。

链接: https://arxiv.org/abs/2610.09569
作者: Shivam Shukla,Jihye Kim,Shubham Gaur,Mahnaz Roshanaei,Magy Seif El-Nasr
机构: University of California, Santa Cruz(加州大学圣克鲁兹分校); Stanford University (斯坦福大学)
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Large language models (LLMs) are increasingly used for emotional support, raising concern that sustained use may draw users away from their real-world relationships. Yet existing evaluations primarily focus on the safety, empathy, or helpfulness of responses, leaving under-examined a relational question: where does the model orient the user for continued support? To address this question, we introduce relational orientation, a property operationalized through two non-exclusive dimensions: inward-facing (IF) language, which positions the AI as the user’s ongoing source of support, and outward-scaffolding (OS) language, which encourages real-world human connection. Grounded in psychological and sociological literature, we formalize a taxonomy of relational orientation and present RELATE, a persona-conditioned framework for measuring inward-facing and outward-scaffolding language at the sentence level in multi-turn dialogues. RELATE pairs 76 help-seeking situations adapted from naturally occurring questions with three simulated user styles, providing 228 evaluation stimuli. In our experiments, we evaluate seven LLMs using dialogues with six assistant turns each, yielding 1,596 dialogues and 69,194 assistant sentences. We assess these sentences using a primary rubric-based LLM judge and apply a secondary judge to a subset. Under automated evaluation, we find that the proportion of sentences labeled as IF is higher at the sixth assistant turn than at the first, while the proportion labeled as OS is substantially lower for hesitant, indirect simulated users than for explicit, reassurance-seeking users. RELATE provides a reproducible framework and a sentence-level signal for auditing and steering the relational orientation of supportive LLMs.

[NLP-57] Constitution-Guided Watermarking

【速读】: 该论文旨在解决文本生成模型中水印技术在实际应用中面临的多重矛盾:一方面,增强水印信号以提升可检测性会损害生成文本的质量;另一方面,为抵抗编辑而设计的鲁棒性水印可能被恶意利用进行伪造。现有方法通常采用固定的配置来权衡这些相互冲突的目标,导致所有请求共享同一操作点,无法根据具体需求灵活调整,从而在需要语义保真度的场景下牺牲质量,或在需要可靠溯源的场景下降低鲁棒性。本文提出**宪法引导式水印(Constitution-Guided Watermarking)**框架,其核心在于将水印设计的权衡策略从硬编码参数转化为由自然语言形式表达的“宪法原则”(constitutional rules),并通过预训练的推理代理(reasoning agent)在离线阶段基于实证反馈迭代优化每条规则对应的水印配置。部署时,独立的监控模块可动态识别适用规则并检索相应的水印策略(包括豁免机制),无需修改服务模型。此外,该框架支持离线并行优化与持续迭代,适应不断变化的提供商需求,且每个部署配置均绑定其评估证据,确保决策可审计。实验表明,在基于KGW和五条宪法规则的原型评估中,该框架能根据提供商优先级自适应选择配置,在以鲁棒性为优先的请求上,后重述检测性能相比固定配置最高提升14个百分点,同时在整体文本质量、干净文本检测率(假阳性率为0.1%)方面达到或超越所有基线。

链接: https://arxiv.org/abs/2610.09552
作者: Toluwani Aremu,Samuele Poppi,Nils Lukas
机构: 未知
类目: Cryptography and Security (cs.CR); Computation and Language (cs.CL); Machine Learning (cs.LG)
备注: Working paper (under review)

点击查看摘要

Abstract:Watermarking enables language model providers to identify text generated by their models. However, its desired properties can conflict (\ie~stronger watermark signals can degrade text quality), while designs that resist editing may also facilitate forgery. Providers address these trade-offs by choosing configurations that balance competing objectives or prioritize particular properties. Either approach imposes a shared operating point on requests with different requirements, potentially sacrificing quality where wording preservation matters or robustness where reliable attribution is essential. To allow flexible and adaptable designs, we introduce \emphConstitution-Guided Watermarking, a framework that selects request-appropriate trade-offs from provider requirements, listed as natural-language principles. \emphOffline, a pretrained reasoning agent examines constitutional rules alongside watermark implementations and iteratively refines rule-specific configurations using empirical feedback. \emphAt deployment, a separate monitor identifies applicable rules and retrieves the corresponding policy, including watermarking exemptions, without modifying the serving model. Furthermore, our framework supports offline parallel optimization and refinement of rule-specific configurations based on evolving provider requirements without affecting deployment, and binds each deployed configuration to its evaluation evidence, making deployment decisions auditable. In a proof-of-concept evaluation using KGW and a five-rule constitution, our framework selects configurations responsive to provider priorities and improves post-paraphrase detection on robustness-prioritized requests by up to 14 percentage points over fixed configurations, while matching or exceeding all baselines in aggregate quality and clean detection at a nominal 0.1% false-positive rate.

[NLP-58] A Comparative Study of Evaluation Metrics for Long-Document Financial Narrative Summarization with Transformers

【速读】: 该论文旨在解决英国上市公司年度报告(平均约80页)信息冗长、难以高效获取核心财务与经营信息的问题,重点针对金融叙事摘要任务中的摘要生成效果评估难题。其解决方案的关键在于:首先,系统性地对比多种预训练变换器模型与不同信息提取技术在金融文本摘要上的表现;其次,提出一种新的综合评估指标——BRUGEscore,即ROUGE-2与BERTScore的调和平均值,以更准确反映摘要质量,克服传统评价指标在语义一致性与词汇匹配度方面存在的局限性;最后,通过统计显著性检验及对抗性扰动分析(引入三种数据破坏方法),验证结果的稳健性与模型鲁棒性,从而为金融领域自动化摘要系统的评估与优化提供可靠依据。

链接: https://arxiv.org/abs/2610.09529
作者: Nadhem Zmandar,Mo El-Haj,Paul Rayson
机构: Lancaster University(兰卡斯特大学)
类目: Computation and Language (cs.CL)
备注: 12 pages

点击查看摘要

Abstract:There are more than 2,000 listed companies on the UK’s London Stock Exchange, divided into 11 sectors who are required to communicate their financial results at least twice in a single financial year. UK annual reports are very lengthy documents with around 80 pages on average. In this study, we aim to benchmark a variety of summarisation methods on a set of different pre-trained transformers with different extraction techniques. In addition, we considered multiple evaluation metrics in order to investigate their differing behaviour and applicability on a dataset from the Financial Narrative Summarisation (FNS 2020) shared task, which is composed of annual reports published by firms listed on the London Stock Exchange and their corresponding summaries. We hypothesise that some evaluation metrics do not reflect true summarisation ability and propose a novel BRUGEscore metric, as the harmonic mean of ROUGE-2 and BERTscore. Finally, we perform a statistical significance test on our results to verify whether they are statistically robust, alongside an adversarial analysis task with three different corruption methods.

[NLP-59] Goldsmith: Gold-Loss-Guided Definition Optimization with an Agent ic Annotation Harness EMNLP2026

【速读】: 该论文旨在解决在缺乏稳定标注指南或足够标注数据以训练特定任务模型的情况下,如何有效构建可复用的结构化标注定义的问题。其核心挑战在于如何利用少量专家标注的黄金样本(gold set)来生成准确、一致且可扩展的标注规范。解决方案的关键在于提出Goldsmith——一种基于智能体(agentic)的自动化流水线,将少量黄金样本转化为可训练的文本对象(trainable textual object)。该系统通过在相同黄金样本上测试候选定义,并使用可执行的结构化损失函数(executable structured loss)进行评分,结合外部框架完成输出模式、格式化、检索、修复、评估与人工审查等环节。由大语言模型(LLM)作为编辑器,针对高损失失败案例生成基于文本梯度的修正版本,仅当损失下降时才接受更新。实验表明,在提示优化对比中,Goldsmith优于直接重写、OPRO、APE和PromptBreeder;同时,所生成的标注定义在结合检索、基于得分的路由及人工审查后,显著提升了类型化跨度、成对关系和固定触发事件-论元等下游标注任务的性能。结果表明,有限的专家监督能够同时支持任务定义的学习与可扩展的标注流程。

链接: https://arxiv.org/abs/2610.09489
作者: Yihan Li,Hanyi Zhang,Xiaoxi Jiang,Man Guo
机构: Sun Yat-sen University (中山大学); South China University of Technology (华南理工大学)
类目: Computation and Language (cs.CL)
备注: 20 pages, 4 figures, 11 tables. Accepted to the main conference of EMNLP 2026

点击查看摘要

Abstract:Many annotation projects begin before experts have a stable guideline or enough labels to train a task-specific model. We present Goldsmith, an agentic pipeline that turns a small gold set—expert-annotated calibration examples representing the intended task boundaries—into a reusable structured annotation definition. Goldsmith treats this definition as a trainable textual object. Candidate definitions are run on the same gold examples and scored with an executable structured loss, while the output schema, formatting, retrieval, repair, judging, and human review remain in an external harness. A large language model (LLM) editor converts the highest-loss failures into textual-gradient revisions, which are accepted only when the measured loss decreases. In prompt-optimization comparisons, Goldsmith improves over direct rewriting, OPRO, APE, and PromptBreeder under matched evaluation protocols. The resulting definition also improves downstream annotation when combined with retrieval, score-based routing, and human review across typed span, pair-level relation, and fixed-trigger event-argument tasks. These results show that scarce expert supervision can support both task-definition learning and scalable annotation.

[NLP-60] CHASE: Channel-Aligned Structure Exploitation for Geometry-Aware Model Engineering

【速读】: 该论文旨在解决大模型在参数高效微调、结构化剪枝补偿及模型融合等任务中因缺乏对网络内部结构本质理解而导致的性能瓶颈问题。其核心挑战在于如何有效利用训练后模型所蕴含的几何与谱结构(Geometric and Spectral Alignment, GSA),以指导更高效的模型设计与优化。解决方案的关键在于提出一种基于GSA结构的通用方法框架——CHASE(Channel-Aligned Structure Exploitation),通过挖掘并利用训练模型中的谱集中性、物理通道对齐性、支撑结构以及奇异基的变化规律,直接指导多种实际模型操作的设计。具体而言,CHASE涵盖六种应用场景,包括模型修改、重构与压缩;其中CAGA通过几何对齐与低秩子空间提取实现多头注意力中共享键值(KV)表示的头选择与构造,显著提升从多头注意力(MHA)到分组查询注意力(GQA)转换的效率;SAKV基于GSA确定相邻层间可共享的低秩KV缓存表示及其保留秩;CAPS则利用谱结构对输出神经元进行分组,并为每组独立选择保留的输入通道。实验表明,这些方法在模型适应、剪枝补偿与融合任务中均优于现有基准,验证了GSA所揭示的内在结构可被直接用于构建高效且可解释的模型优化策略。

链接: https://arxiv.org/abs/2610.09476
作者: Wei Wang,Wei Jiang,Ziran Liu
机构: Futurewei Technologies(未来科学城技术公司); Shanghai Institute for Mathematics and Interdisciplinary Sciences (SIMIS)(上海数学与交叉科学研究所)
类目: Machine Learning (cs.LG); Computation and Language (cs.CL)
备注: 26 pages, 9 tables

点击查看摘要

Abstract:Geometric and Spectral Alignment (GSA) characterizes trained networks through spectral concentration, physical-channel alignment, support structure, and changes in singular bases. In this paper, we propose CHASE (Channel-Aligned Structure Exploitation) to use these structures in practical model design. CHASE covers six applications across model modification, reconfiguration, and compression. CORA, COEC, and CORAM apply GSA to parameter-efficient finetuning, structured-pruning compensation, and model merging. We further develop three new methods. CAGA uses GSA to identify multi-head attention heads that can share a KV representation and constructs the shared key and value heads through geometric alignment and low-rank subspace extraction. SAKV uses GSA to determine which adjacent layers can share a low-rank KV-cache representation and the retained rank for each layer group. CAPS uses GSA spectral structure to group output neurons and selects retained input channels separately for each group. Results from CORA, COEC, and CORAM establish the effectiveness of GSA for adaptation, pruning compensation, and model merging. Experiments on CAGA show that geometric shared-head construction substantially improves MHA-to-GQA conversion, and SAKV and CAPS improve over representative baselines for KV-cache compression and structured pruning. These results show that the structures identified by GSA can be used directly to design methods for a range of model operations.

[NLP-61] BanglaRhet: Benchmarking Classical and Transformer Models for Rhetorical and Persuasion Detection in Bangla Political Speech

【速读】: 该论文旨在解决孟加拉语(Bangla)政治话语中细粒度修辞与说服技巧检测缺乏系统性基准评估的问题。尽管孟加拉语自然语言处理在情感分析和观点挖掘方面已取得进展,但针对变压器模型在修辞形式与说服意图识别任务上的性能评估仍处于探索阶段。其解决方案的关键在于构建并公开发布一个大规模、人工标注的孟加拉语政治话语语料库——BanglaRhet,包含30,289个来自公开政治新闻源的政治演讲片段,并据此定义两个监督式单标签分类任务:修辞技巧检测(对比、重复、夸张、隐喻、反问)与说服技巧检测(归咎、行动号召、团结呼吁、道德诉求、情感诉求和逻辑诉求)。研究评估了四种基于变压器的模型(BanglaBERT、BanglaBERT-Base、SahajBERT 和 XLM-RoBERTa-Base)以及经典TF-IDF基线方法,结果表明,BanglaBERT在两项任务上表现最优,分别达到65.40%和66.46%的宏平均F1分数,显著优于最佳调优的经典基线模型(分别提升19.2和13.8个百分点)。进一步的类别层面分析揭示,错误主要源于标签间的语义重叠、修辞性语言表达及类别不平衡问题。研究为孟加拉语修辞与说服意识型政治话语分析提供了初步基准,并强调了发展上下文感知与多标签建模的重要性。

链接: https://arxiv.org/abs/2610.09464
作者: Rohit Kumar Sen,Anik Chowdhury
机构: North East University Bangladesh (东北大学孟加拉国)
类目: Computation and Language (cs.CL)
备注: 6 pages, 2 figures, 6 tables. Accepted at the 2026 2nd International Conference on Advances in Computing, Communication, Electrical, and Smart Systems (iCACCESS), Dhaka, Bangladesh. Dataset: this https URL . Code: this https URL . Hugging Face: this https URL

点击查看摘要

Abstract:Political discourse often uses rhetorical and persuasive language to frame narratives, influence public opinion, and mobilize audiences. While Bangla natural language processing has made progress in sentiment analysis and opinion mining, systematic benchmarking of transformer models for fine-grained rhetorical and persuasion technique detection in Bangla political speech remains largely underexplored. This paper presents a benchmark study of transformer-based models for detecting rhetorical form and persuasive intent in Bangla political discourse. Using BanglaRhet, a manually annotated corpus of 30,289 Bangla political speech segments collected from publicly available political news sources, we formulate two supervised single-label classification tasks: rhetorical technique detection (contrast, repetition, exaggeration, metaphor, rhetorical questions) and persuasion technique detection (blame assignment, call to action, unity call, moral, emotional, and logical appeals). We evaluate four transformer-based models, BanglaBERT, BanglaBERT-Base, SahajBERT, and XLM-RoBERTa-Base, against classical TF-IDF baselines. BanglaBERT achieves the highest performance, with 65.40% macro-F1 for rhetorical technique detection and 66.46% for persuasion technique detection, outperforming the best tuned classical baseline by 19.2 and 13.8 macro-F1 points, respectively. Class-level analysis indicates that errors are mainly associated with semantic overlap among labels, figurative language, and class imbalance. The results provide initial benchmark baselines for Bangla rhetorical and persuasion-aware political discourse analysis and highlight the need for context-aware and multi-label modeling.

[NLP-62] Right Number Wrong State? Measuring Cross-Jurisdiction Substitution in LLM Recall of State Policy

【速读】: 该论文旨在解决大语言模型(Large Language Model, LLM)在回答特定州政策问题时出现错误的原因识别问题,即判断其错误是源于幻觉(hallucination),还是因返回了另一州的真实值所致。其解决方案的关键在于采用最小集合实验设计(minimal-set design),固定问题表述仅变更管辖区域(涵盖美国50个州及华盛顿特区共51个司法管辖区)和三个精确界定的医疗补助收入资格标准,并以官方数据手册为金标准(gold standard),验证模型输出的准确性。研究发现,在153个测试项中,Claude Sonnet 5.5与GPT-5.6 Sol分别在10和25项上重复输出了其他州的当前值,表现出高度可复现性。然而,错误归因具有脆弱性:若仅将与另一州数值一致的错误答案视为跨州错误,则会高估此类错误比例达3至5倍,因为许多看似跨州的答案实为所问州在不同计算惯例或历史年份下的真实值。因此,准确判断跨司法管辖区错误必须依赖完整的本州参照数据集。研究全程遵循预注册协议,并将公开实验协议、金标表及所有模型输出结果。

链接: https://arxiv.org/abs/2610.09458
作者: Jiayu Feng
机构: Harvard T.H. Chan School of Public Health (哈佛大学陈曾熙公共卫生学院)
类目: Computation and Language (cs.CL)
备注: 6 pages, 3 figures, 1 table

点击查看摘要

Abstract:When an LLM answers a state-specific policy question wrongly, it may be hallucinating, or it may be returning a real value that holds in another state. We test this with a minimal-set design: the question wording is fixed and only the jurisdiction varies, across the 50 U.S. states and the District of Columbia (51 jurisdictions) and three exactly defined Medicaid income-eligibility quantities. Gold values come from an official data book and agree with an independent source in 101 of 102 checked cells. Under a pre-registered protocol, Claude Sonnet 5.5 and GPT-5.6 Sol reproducibly give another state’s current value, identical across two independent repeats, for 10 and 25 of 153 items. Attribution is fragile, however. Crediting any wrong answer that equals another state’s value yields 3-5x more reproducible substitutions than checking every number in the asked state’s own records, because many apparent cross-state answers are the asked state’s own values under another convention or from an earlier year. Claims about cross-jurisdiction error need a complete same-state reference set. We will release the protocol, gold table, and all model outputs.

[NLP-63] Arctic Questions Missing Answers: A Dataset and Benchmark for LLM Abstention in Arctic Science

【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在面对科学多选题时,缺乏对答案可用性敏感性的评估问题。现有研究中频繁的拒答行为并不能充分反映模型对答案是否存在的真实判断能力。为此,作者构建了ArcticQA数据集,该数据集包含194道源自北极地区原始科研文献的科学多选题,并通过自动化方法验证答案支持度与干扰项矛盾性。进一步提出了ArcticAbstain配对基准测试,对比“答案存在”与“答案缺失”两种条件:前者保留正确选项,后者以干扰项替代正确答案,并在两种情形下均提供明确的拒答选项。在高推理强度下评估来自Gemini、Claude和ChatGPT系列的八款模型,每种条件执行三次试验,共获得9,312条响应记录。结果显示,“答案存在”条件下的拒答率范围为0.0%至63.0%,而将正确答案替换为干扰项后,平均拒答率上升5.05个百分点。这一发现揭示了不同模型间在拒答行为上的显著基线差异,强调必须联合评估拒答频率与对答案可用性的响应灵敏度。该数据集与基准测试已公开发布。

链接: https://arxiv.org/abs/2610.09446
作者: Benjamin Wilcox,Dawei Gao,Pradeeban Kathiravelu,Douglas Causey,Kewei Sha,Yunhe Feng
机构: University of North Texas (北德克萨斯大学); University of Alaska Anchorage (阿拉斯加安克雷奇大学)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Large language models (LLMs) should abstain from scientific multiple-choice questions when no option is valid, but frequent abstention alone does not demonstrate sensitivity to answer availability. We introduce ArcticQA, a dataset of 194 questions derived from primary Arctic research, with automated checks of answer support and distractor contradiction against source evidence. We further develop ArcticAbstain, a paired benchmark comparing answer-present and answer-absent conditions, with the correct answer replaced by a distractor in the latter and an explicit abstention option in both. We evaluate eight models from the Gemini, Claude, and ChatGPT families at high reasoning effort, with three trials per condition, yielding 9,312 recorded responses. Answer-present abstention rates range from 0.0% to 63.0%, whereas replacing the correct answer increases abstention by 5.05 percentage points on average. These findings highlight substantial baseline differences and the need to evaluate abstention frequency and responsiveness jointly. The dataset and benchmark are available at this https URL.

[NLP-64] ARCS: Towards Precise Text-to-SQL via Structured Disambiguation

【速读】: 该论文旨在解决文本到SQL(text-to-SQL)系统在真实场景部署中因用户问题存在语义模糊性而导致的错误问题。此类模糊性通常细微、依赖特定领域或数据分布,且难以被察觉,但会显著导致系统输出偏离用户真实意图。传统上,模糊性通过对话式澄清来解决,但这种方式效率低下、认知负担重,且与实际用户工作流程不匹配。本文提出一种新型范式——结构化消歧(structured disambiguation),即通过显式的、受约束的交互方式而非自由对话来消除模糊性。为此,研究构建了首个面向真实世界数据库的文本到SQL基准数据集ARCS(Ambiguity Resolution Corpus for SQL),其包含自然发生的、非受限的模糊性实例,并对所有有效模糊点、解释及其对应的SQL查询进行了完整标注。实验结果表明,在存在模糊性的情况下,文本到SQL任务依然极具挑战:即使使用GPT-6-SOL模型,端到端执行准确率也仅达51%,而所有开源模型均未超过27%。解决方案的关键在于引入结构化消歧机制与高质量标注数据集ARCS,以推动更鲁棒、可解释的文本到SQL系统发展。

链接: https://arxiv.org/abs/2610.09396
作者: Yihao Hu,Yanlin Feng,Naoki Otani,Nikita Bhutani
机构: Megagon Labs(兆元实验室); Duke University(杜克大学)
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:As text-to-SQL systems move beyond demonstrations toward real-world deployment, ambiguity in user questions becomes a primary source of errors. Such ambiguities are often subtle, domain- or data-specific, and can silently cause system outputs to deviate from the user’s true intent. Ambiguity is traditionally addressed through conversational clarification, which is often inefficient, cognitively demanding, and poorly aligned with real-world user workflows. We propose structured disambiguation, a new paradigm in which ambiguity is resolved through explicit, constrained interactions rather than free-form dialogue. We construct ARCS (Ambiguity Resolution Corpus for SQL), the first text-to-SQL benchmark featuring naturally occurring, unconstrained ambiguities over real-world databases, with complete annotations of all valid ambiguity points, interpretations, and SQL queries. Experimental results show that text-to-SQL remains challenging in the presence of ambiguity: gpt-6-sol achieves only 51% end-to-end execution accuracy, and no open-source model exceeds 27%.

[NLP-65] he Persona Hierarchy Model: Understanding Contextual Generalization in Fine-Tuning LLM s

【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在微调过程中出现的上下文泛化能力不一致问题:即模型在特定上下文(如系统提示、角色设定或领域指令)下微调后,其行为有时仅局限于该上下文,而有时却能广泛迁移至未见过的上下文。为解释这一现象,论文提出了角色层级模型(Persona Hierarchy Model),其核心观点是:存在一个共享的默认角色(default persona),它在跨上下文情境中持续影响模型行为;当微调过程改变这一共享默认角色时,可促进更广泛的泛化;反之,仅调整局部角色则导致行为仍保持上下文特异性。研究通过120个覆盖四种行为和15种训练上下文的微调模型验证发现,泛化狭窄程度与训练上下文角色与默认角色的相似性呈显著正相关(Qwen3-4B模型中皮尔逊相关系数r = 0.72)。此外,先在默认上下文中进行预微调,有助于后续在其他上下文中实现更广的泛化;使上下文响应与默认角色响应对齐,亦可增强泛化效果。最后,论文提出角色保持正则化(Persona-Preserving Regularization, PPR),用于抑制非预期的上下文泛化,在强化学习场景下将奖励劫持(reward hacking)从42–55%降至不超过0.2%,同时保留性能提升。这些结果支持角色层级模型作为理解上下文泛化的理论框架,并为未来实现更可控的模型对齐提供方法指导。

链接: https://arxiv.org/abs/2610.09384
作者: Jiachen Zhao,Zhengxuan Wu,David Bau,Weiyan Shi
机构: Northeastern University(东北大学); Google DeepMind(谷歌深度思维); Stanford University(斯坦福大学)
类目: Computation and Language (cs.CL); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Language models are routinely fine-tuned under a fixed context, such as a generic system prompt, persona or domain-specific instruction, yet the learned behavior sometimes stays confined to that context and sometimes broadly generalizes to unseen contexts. We propose the Persona Hierarchy Model to explain this: a shared default persona influences behavior across contexts. Under this model, fine-tuning that modifies the shared persona promotes broader transfer, whereas changes to local personas remain more context-specific. Across 120 fine-tuned models spanning four behaviors and 15 training contexts, generalization narrowness positively correlates with the similarity between the training context’s persona and the default persona (Pearson’s r = 0.72 for Qwen3-4B). Prior fine-tuning under the default context can broaden generalization in subsequent training under other contexts. Aligning contextual responses with default-persona responses produces stronger effects. Finally, we propose persona-preserving regularization (PPR) to confine undesired contextual generalization. In RL, PPR cuts reward hacking from 42-55% to at most 0.2% under every evaluated prompt while retaining accuracy gains. These results support the Persona Hierarchy Model as an explanation for contextual generalization and can motivate future controls on unintended generalization for better alignment of LLMs.

[NLP-66] Expert Coupling in MoE Pretraining: Reducing All-to-All Overhead with Correlated Placement and Token Shuffling

【速读】: 该论文旨在解决大规模生成式AI模型中混合专家(Mixture-of-Experts, MoE)层在专家并行(Expert Parallelism, EP)模式下因频繁的全对全通信(all-to-all collectives)导致的训练效率瓶颈问题。具体而言,在多节点集群中,基于top-2和top-6路由策略的MoE层在前向与反向传播过程中需进行大量跨GPU乃至跨节点的通信,显著拖慢训练速度,尤其在EP32配置下通信开销可高达训练步耗时的60%。其解决方案的关键在于利用预训练早期即已形成的专家选择相关性:一方面,发现同一层内或跨层间存在高度相关的专家选择模式(如特定专家对被42%的令牌同时选中,且当前层的选择能预测下一层的选择);另一方面,通过将常被共同选择的专家放置于同一GPU上,并结合仅需单次发送至各GPU的调度器(dispatcher),使更多令牌-专家分配关系保留在本地GPU,从而大幅减少跨设备通信。此外,引入基于序列并行的令牌重排(token shuffling)机制,在reduce-scatter阶段将每个令牌移动至预计持有其下一层数专家的GPU,进一步提升本地处理比例。实验表明,在Megatron-LM框架下,该方法在EP 8~64范围内,使all-to-all通信时间降低1.16~2.63倍,端到端训练步时间最高加速1.41倍,且不改变原始路由决策与专家参数。

链接: https://arxiv.org/abs/2610.09372
作者: Radha Gulhane,Quentin Anthony,Beren Millidge
机构: Zyphra
类目: Computation and Language (cs.CL); Distributed, Parallel, and Cluster Computing (cs.DC)
备注:

点击查看摘要

Abstract:Mixture-of-Experts (MoE) layers replace the feed-forward block of a Transformer with E expert networks, and each token is routed to k of these experts. Under expert parallelism (EP) the experts are distributed across GPUs, and every MoE layer runs all-to-all collectives in the forward and backward passes to dispatch tokens to their experts and then combine the results. On a cluster with 8 AMD Instinct MI300X GPUs per node, these collectives can take 45% of the training step at EP32 with top-2 routing and 60% with top-6 routing. We find that early in pretraining routers have already learned to assign tokens to experts in correlated patterns, both within a layer and across layers. At top-2, 0.8% of the expert pairs in a layer are selected together by 42% of tokens, and the experts a token selects at one layer predict the experts it selects at the next layer. We use these correlations to keep more token–expert assignments on the token’s own GPU, which reduces communication across GPUs and across nodes. Correlated expert placement puts experts that are often selected together on the same GPU. Combined with a dispatcher that sends each token to each GPU once, it removes up to 58% of dispatched rows. Token shuffling applies when sequence parallelism shards tokens across the EP group. It moves each token to the GPU predicted to hold its next-layer experts during the reduce-scatter that follows attention. On one node this raises the share of token–expert assignments served on the token’s GPU from 12.5% to 59%. In Megatron-LM, across EP degrees from 8 to 64 with top-2 and top-6 routing, the two methods reduce all-to-all time by 1.16-2.63X and end-to-end step time by up to 1.41X. Neither method changes the models’ underlying routing decisions or expert parameters.

[NLP-67] opoGraphRAG -Bench: Evaluating Multimodal GraphRAG on Layout-Grounded Evidence Reasoning NEURIPS2026

【速读】: 该论文旨在解决复杂文档中多模态证据单元(文本、表格、图表及其图注)在非线性布局下证据拓扑结构难以恢复的问题,即如何有效建模并推理跨模态、异构证据单元之间的逻辑关联以支持复杂问答。其解决方案的关键在于构建一个基于版面信息的基准测试框架——TOPOGRAPHRAG-BENCH,该基准包含2,024个问题,覆盖201份长篇视觉丰富文档,通过受控的三种拓扑结构(单跳检索、桥接链式推理、多源合成)自下而上地构造问题,并引入反事实验证机制确保问题对快捷路径、模态必要性及证据必要性的抗干扰能力。实验结果表明,仅依赖文本的GraphRAG系统在关键依赖位于图表时表现不佳,而仅使用页面级视觉检索的系统缺乏细粒度结构支持拓扑恢复;相比之下,多模态GraphRAG系统虽整体表现最优,但在视觉-文本对齐不完整或多单元组合缺失时仍会失效。因此,该研究强调未来GraphRAG系统需超越基于文本的实体关系图,显式建模文档版面结构、跨模态证据对齐以及证据单元在推理中的角色,从而实现更精准的多模态证据拓扑推理。

链接: https://arxiv.org/abs/2610.09360
作者: Ruochi Li,Jianzhe Lin,Haoxuan Zhang,Haihua Chen,Junhua Ding,Edward Gehringer,Yang Zhang
机构: North Carolina State University (北卡罗来纳州立大学); University of North Texas (北德克萨斯大学); University of Wyoming (怀俄明大学); Meta (Meta)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: Accepted at the 40th Conference on Neural Information Processing Systems (NeurIPS 2026)

点击查看摘要

Abstract:Real-world documents distribute evidence across text, tables, figures, and captions within complex page layouts. Answering complex questions over such documents therefore requires more than retrieving relevant passages: systems must recover the evidence topology that connects heterogeneous evidence units. Existing GraphRAG evaluations remain largely text-centered, while multimodal document RAG benchmarks assess cross-modal retrieval and generation without directly evaluating recovery of the intended evidence topology. We introduce TOPOGRAPHRAG-BENCH, a layout-grounded benchmark for multimodal evidence reasoning in GraphRAG, comprising 2,024 questions over 201 long, visually rich documents. Questions are constructed bottom-up from text, figure, and table evidence units under three controlled topologies: single-hop retrieval, bridge-chain reasoning, and multi-source synthesis. To ensure that questions preserve their intended structure, we apply counterfactual validation for shortcut resistance, modality necessity, and evidence necessity. We evaluate text-only GraphRAG, page-level visual retrieval, and multimodal GraphRAG systems using retrieval, generation, and topology-aware reasoning metrics. Multimodal GraphRAG systems achieve the strongest overall performance, but still fail when visual-textual evidence alignment or multi-unit composition is incomplete. Text-only GraphRAG struggles when key dependencies are grounded in figures or tables, while page-level visual retrieval lacks the fine-grained structure needed for topology recovery. These findings motivate GraphRAG systems that move beyond text-derived entity relation graphs to explicitly model document layouts, cross-modal evidence alignment, and the reasoning roles of evidence units. Code and data are available at this https URL.

[NLP-68] OnlineQAT: On-Policy Distillation for Ultra-Low-Bit Large Language Models

【速读】: 该论文旨在解决大语言模型在低于四比特量化时因量化误差导致的性能下降问题,尤其关注现有量化感知训练(QAT)方法在恢复精度时存在的局限性。现有方法通常依赖固定完成序列或教师生成的答案进行优化,而实际部署中量化模型的推理过程依赖于自身生成的前缀(prefix),这使得模型可能进入训练数据中未覆盖的状态,从而加剧量化误差。为应对这一挑战,论文提出OnlineQAT,其核心解决方案在于采用两阶段框架:首先通过分块量化感知训练(block-wise QAT)获得可用的低比特初始化;随后在学生模型自生成的响应上执行基于策略的蒸馏(on-policy distillation, OPD),并在每个访问到的前缀处,由一个冻结的全精度教师模型提供采样的反向KL(reverse-KL)训练信号。该机制使模型能够从自身生成的动态状态中学习,有效覆盖传统离线恢复数据中缺失的分布区域。实验结果表明,在Qwen3-1.7B模型上,OnlineQAT在W3A16和W2A16设置下分别达到57.28和32.52的平均得分,相较于ReasoningQAT分别提升2.90和0.44点,验证了学生模型所访问状态作为恢复信号的有效性,尤其在三比特量化场景下表现显著。

链接: https://arxiv.org/abs/2610.09346
作者: Wenjun Wang,Heng Li,Yanggan Gu,Hongxia Yang
机构: The Hong Kong Polytechnic University(香港理工大学); Sun Yat-sen University(中山大学); PolyU-Daya Bay Technology and Innovation Research Institute(香港理工大学-大亚湾科技与创新研究院)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Quantization-aware training (QAT) can recover much of the accuracy lost when large language models are compressed below four bits. Existing re- covery stages, however, are commonly optimized on fixed completions or teacher-generated answers, whereas the deployed quantized model condi- tions on prefixes generated by itself. Quantization errors can therefore move the model into states that are absent from offline recovery data. We introduce OnlineQAT, a two-stage framework that first obtains a usable low-bit initialization through block-wise QAT and then performs on-policy distillation (OPD) on student-generated responses. At each visited pre- fix, a frozen full-precision teacher provides a sampled reverse-KL training signal. On Qwen3-1.7B, OnlineQAT obtains the best average among the compared quantized methods: 57.28 at W3A16 and 32.52 at W2A16, im- proving over ReasoningQAT by 2.90 and 0.44 points, respectively. The results suggest that student-visited states provide a useful recovery signal beyond fixed-completion training, particularly at three bits.

[NLP-69] Dialect-Robust Speech Language Models with Synthetic Pseudo-Dialect Augmentation

【速读】: 该论文旨在解决低资源方言(dialect)场景下语音语言模型(Speech Language Model, SLM)性能下降的问题,其核心挑战在于方言语音数据稀缺,导致传统文本到语音(Text-to-Speech, TTS)增强方法难以有效覆盖多样化的方言。为应对这一问题,本文提出一种无需真实方言语音即可生成伪方言语音(pseudo-dialect speech)的解决方案:通过将大语言模型(LLM)生成的方言文本,利用标准语TTS模型进行语音合成,从而实现零真实方言语音依赖的方言数据增强。该方案的关键创新在于引入训练过程中的中间标准文本预测(intermediate standard-text prediction),作为语义归一化机制,有效缓解方言与标准语之间的语义鸿沟,提升下游任务的语义一致性。实验在日语、德语和汉语方言的方言到英文语音翻译任务上验证了该方法的有效性,结果显示,相较于仅使用合成标准语音的基线,伪方言增强显著提升了日语(25.38 → 26.24)和德语(31.57 → 32.47)的性能;结合中间标准文本预测后,日语性能进一步提升至28.26,中文性能从11.67大幅提高至16.37。结果表明,该方法具有良好的跨语言可扩展性,无需针对每种方言构建专属语音资源即可实现有效的方言建模。

链接: https://arxiv.org/abs/2610.09321
作者: Shunsuke Mitsumori,Tomoya Mizumoto,Yusuke Fujita
机构: SB Intuitions; Waseda University (早稻田大学)
类目: Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)
备注: 7 pages, 1 figure, 7 tables. Accepted to IEEE SLT 2026

点击查看摘要

Abstract:Speech Language Model (SLM) performance often degrades on dialects due to data scarcity. Conventional text-to-speech (TTS) augmentation struggles to cover diverse dialects as it requires a certain amount of real dialect speech. We propose synthesizing pseudo-dialect speech by converting LLM-generated dialect text via a standard-language TTS model, requiring zero real dialect speech. Additionally, we introduce intermediate standard-text prediction during training, acting as semantic normalization for downstream tasks. We evaluate dialect understanding via dialect-to-English speech translation across Japanese, German, and Chinese dialects. Compared to synthetic standard speech baselines, pseudo-dialect augmentation improves scores for Japanese (from 25.38 to 26.24) and German (from 31.57 to 32.47). Furthermore, the intermediate standard-text prediction effectively bridges the semantic gap, boosting performance to 28.26 for Japanese and from 11.67 to 16.37 for Chinese. These results suggest that our approach scales to various languages without requiring speech resources specific to each dialect.

[NLP-70] Adversarial Images Hijack Web Agents from Visual Grounding to Browser Execution

【速读】: 该论文旨在解决现有视觉红队测试方法在评估基于视觉-语言模型(Vision-Language Model, VLM)的网页智能体(web agent)时,仅关注模型推理阶段而忽视结构化输入处理与动作后处理环节所导致的端到端鲁棒性评估不足的问题。其核心挑战在于:即使模型层面的输出被成功误导,也无法确保浏览器执行层面的实际控制权,从而无法真实反映智能体的整体安全性。为此,论文提出将视觉红队测试重构为一个从视觉感知到浏览器执行的端到端对齐问题,并设计了WebMirage框架,通过引入角色-槽位(role-slot)抽象与网页重组机制,建模网页元素间的竞争关系;结合数据流分析,使扰动优化过程与动作后处理逻辑保持一致。该方案能够生成局部视觉扰动,诱导智能体选择攻击者控制的内容并触发对应浏览器操作,且在多种网页渲染条件下保持有效性。实验表明,WebMirage在2,250个任务上平均攻击成功率高达91.9%,显著优于最强基线(17.4%),并能有效绕过三种代理级防御机制。

链接: https://arxiv.org/abs/2610.09240
作者: Wanjing Han,Levi Taiji Li,Mu Zhang,Yue Jiang,Guanhong Tao
机构: University of Utah(犹他大学)
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: 20 pages, 8 figures, 6 tables. Code: this https URL

点击查看摘要

Abstract:Modern web agents built on large vision-language models process webpages, select relevant UI elements, and translate model outputs into browser actions. Existing visual red-teaming approaches use adversarial visual content to manipulate this process. However, they primarily target model inference and do not explicitly account for structured input processing or action post-processing. Consequently, model-level success does not establish control over browser execution and cannot reliably characterize end-to-end agent robustness. To address this gap, we formulate red teaming for vision-grounded web agents as an end-to-end grounding-to-execution problem, and introduce WebMirage, a framework that crafts localized visual perturbations that cause agents to select attacker-controlled content and execute the corresponding browser action across varying webpage renderings. It uses a role-slot abstraction and webpage recomposition to capture competition among webpage elements, and dataflow analysis to align optimization with action post-processing. We evaluate WebMirage across four agent configurations and six VLM backbones on 2,250 tasks covering 13 public websites and a sandbox benchmark. WebMirage achieves an average attack success rate of 91.9%, compared with 17.4% for the strongest baseline, and remains effective against three agent-level defenses.

[NLP-71] rajectory Abstraction for the Science of Language Agent Behavior

【速读】: 该论文旨在解决科学语言智能体研究中缺乏跨任务与跨模型可通用的行为变量这一核心问题。其关键解决方案在于提出并实证一种分层轨迹抽象(hierarchical trajectory abstraction)的学习与验证框架:通过递归的分析流程,首先对角色与阶段索引事件进行量化测量,构建具有时间约束的关系假设,并在不同实验条件下检验其稳定性;随后基于筛选出的关系构建高层级的“事件模式”(motif)变量,并在这些抽象变量上重复上述分析过程。每一抽象层级均通过显式测量函数与原始轨迹相连接,确保可追溯性。通过观察数据与随机化协议实验验证生成的假设,并借助干预实现方式间的对比,判断抽象是否应保留、精炼或限制。该方法推导出可接受简化操作的有限深度边界,识别协议对固定抽象的影响,刻画抽象误差的构成与实现分歧特征。此外,引入有限样本检验使干预一致性可操作化,典型案例展示了模式构建与抽象优化过程。该范式区别于语义分类体系、定性理论归纳及行为模型重构方法,为发现具有泛化能力的行为假设提供了一套可复现的研究程序,且将文献相对新颖性与模型相对意外性分别评估,增强了研究结论的科学严谨性。

链接: https://arxiv.org/abs/2610.09237
作者: Tianqiang Yan
机构: 未知
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: (Work in Progress) 13 pages, 2 figures

点击查看摘要

Abstract:Scientific studies of language agents need behavioral variables that support hypotheses across tasks and models. We formulate this research problem as learning and testing a hierarchy of trajectory abstractions. A concrete recursive procedure first measures role- and phase-indexed events, proposes temporally constrained relations, and tests their stability across conditions. It then constructs episode-level motif variables from selected relations and repeats the analysis on those variables. Explicit measurement functions connect every abstraction level to the original trajectories. Observations and randomized protocol experiments assess the resulting hypotheses, while comparisons between intervention realizations determine whether an abstraction should be retained, refined, or restricted. We derive a finite-depth bound for accepted reductions, identify protocol effects on fixed abstractions, and characterize realization disagreement and composition of abstraction error. A finite-sample test makes projected intervention consistency operational, and constructed examples illustrate motif construction and abstraction refinement. The formulation distinguishes this experimental approach from semantic taxonomies, qualitative theory induction, and behavior-model recovery. It specifies a proposed research procedure for discovering generalizable behavioral hypotheses, with literature-relative novelty assessed separately from model-relative surprise.

[NLP-72] Multi-Objective Aligned Small Language Model Framework for SUD Patient Dialogue Generation

【速读】: 该论文旨在解决在物质使用障碍(Substance Use Disorder, SUD)咨询场景中,生成式语言模型难以生成具有认知一致性与临床真实性的患者对话行为的问题。现有大语言模型(Large Language Models, LLMs)虽能生成流畅文本,但在缺乏足够临床数据和伦理约束条件下,常表现出认知逻辑断裂、应对策略不连贯等缺陷,且前沿大模型在医疗应用中面临计算成本高、延迟大、隐私风险及资源受限环境部署困难等现实挑战。为此,论文提出一种基于认知基础的框架,其核心在于显式建模并对齐患者的潜在认知成分(如信念、应对策略、改变意愿等)与个体病史及咨询师提问之间的关系。该方案的关键在于构建两阶段流水线:第一阶段为认知成分检测,第二阶段为认知对齐的对话生成。通过融合知识蒸馏、人类标注偏好优化以及注意力引导的奖励塑形等技术,使小型语言模型(Small Language Models, SLMs)在有限数据下仍能有效学习复杂的认知结构。实验表明,该方法在自动评估指标(如BERTScore、ROUGE、METEOR、BLEU)及以大模型为裁判的命中率(LLM-as-judge hit-metrics)上显著优于通用指令微调基线与心理健康领域专用的小型模型,尤其在开放式认知成分生成方面表现突出,实现了更高的认知实现度与临床合理性。

链接: https://arxiv.org/abs/2610.09209
作者: Thushara Manjari Naduvilakandy,Hyeju Jang,Mohammad Al Hasan
机构: Indiana University Indianapolis(印第安纳大学印第安纳波利斯分校)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Substance Use Disorder (SUD) counseling requires patient responses that reflect underlying cognitive states such as beliefs, coping strategies, and readiness for change. Although large language models (LLMs) can generate fluent text, they often fail to produce cognitively coherent and clinically realistic patient behavior, especially under ethical and data-scarce clinical settings. Moreover, deploying frontier-scale LLMs in healthcare applications presents practical challenges including high computational cost, latency, privacy concerns, and limited deployability in resource-constrained environments, motivating the need for cognitively aligned small language models (SLMs). We propose a cognitively grounded framework for SUD patient dialogue generation that explicitly models and aligns latent cognitive components with patient histories and counselor questions. Our pipeline consists of two stages: cognitive component detection and cognitive component-aligned dialogue generation. To enable effective learning with smaller models, we combine knowledge distillation from high-capacity teacher models, preference optimization from human-annotations, and attention-guided reward shaping. Extensive evaluations using automatic scores like BERTScore, ROUGE, METEOR and BLEU, and LLM-as-judge hit-metrics against both human and teacher-model references show that cognitively informed fine-tuning substantially improves cognitive realization and alignment over a generic instruction-tuned baselines and mental health domain specific SLMs, with particularly strong gains for open-ended cognitive components.

[NLP-73] Few Bits One Law: Toward W2A4KV2

【速读】: 该论文旨在解决极端低比特大语言模型(LLM)压缩中权重(weights)、激活值(activations)与键值缓存(KV caches)联合量化时面临的挑战:三者分布差异显著,且量化误差在网络中相互耦合、传播,导致性能严重下降。其核心解决方案是提出CanonQ——一种统一的量化感知训练框架,关键在于将“源标准化”(source canonicalization)与“任务自适应”(task-aware adaptation)分离:通过固定旋转和能量归一化,将异构张量源映射至统一的规范坐标系,从而实现跨层、跨模型复用固定的高斯参考码本(frozen Gaussian-reference codebooks);在此基础上,通过联合训练使网络在统一的标量/向量接口下适应权重、激活与缓存量化带来的耦合误差。作者理论推导了冻结码本迁移误差与局部任务损失的上界,并提出一种精确的归一化感知直通雅可比(normalization-aware straight-through Jacobian),建立了量化失真与梯度偏差之间的明确关联。实验表明,在联合W2A4KV2压缩下,CanonQ-Omni在LLaMA3系列模型上相较现有最优方法将WikiText-2困惑度降低最高达14.28倍,零样本准确率提升最高57.9%;在Qwen3-1.7B、代码生成与数学推理任务中亦表现优异,如在MobileLLM-Pro-1B上于W2A16KV16配置下,HumanEval pass@1与GSM8K精确匹配率分别相对最强基线提升41.7%与39.1%,验证了该框架在多场景下的有效性与普适性。

链接: https://arxiv.org/abs/2610.09202
作者: Kai Yi,Tarek Elgamal,Sruthikesh Surineni,Vignesh Vivekraja,Soumyadeep Ghosh,Steven Li
机构: Meta AI(元AI)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Extreme low-bit LLM compression is most challenging when weights, activations, and KV caches are quantized together: their distributions differ, and quantization errors interact throughout the network. We introduce CanonQ, a unified quantization-aware training framework that addresses these challenges by separating source canonicalization from task-aware adaptation. Fixed rotations and energy normalization map heterogeneous tensor sources to canonical coordinates, enabling frozen Gaussian-reference codebooks to be reused across layers and models. Joint training then adapts the network to the coupled errors of weight, activation, and cache quantization within a common scalar/vector interface. We bound frozen-codebook transfer error and local task loss, and derive an exact normalization-aware straight-through Jacobian that links quantization distortion to gradient bias. The strongest gains arise under joint W2A4KV2 compression: across LLaMA3-1B/3B/8B, CanonQ-Omni achieves up to 14.28x lower WikiText-2 perplexity and up to 57.9% higher mean zero-shot accuracy than prior state-of-the-art and representative quantization baselines. The benefits extend to Qwen3-1.7B, code generation, and mathematical reasoning: on instruction-tuned MobileLLM-Pro-1B at W2A16KV16, CanonQ achieves relative improvements of 41.7% in HumanEval pass@1 and 39.1% in GSM8K exact match over the strongest evaluated quantization baseline.

[NLP-74] Bookkeeping Composition or Unreachable Gold? Reading MemoryAgent Benchs Conflict-Resolution Scores Against a Frozen Last-Write Resolver NEURIPS2026

【速读】: 该论文旨在解决生成式人工智能(Generative AI)在处理多跳推理任务中因记忆冲突导致的“选择性遗忘”(selective forgetting)问题,即当多个关于同一事实的陈述发生冲突时,模型如何正确保留并利用最新、正确的信息。其核心解决方案的关键在于采用“最新声明胜出”的规则作为基准解析器,并通过零学习(zero-learning)方式将其冻结于四个事实列表之一进行测试。实验表明,该规则在官方指标下可正确回答80.25%的问题(在三个保留测试列表上为74.5%),而针对其中67个无法被“最后写入图”(last-write graph)覆盖但可通过被覆盖语句恢复的样本(如“印度首都为新德里”被错误地替换为“格罗斯泰奥”,而正确答案应为新德里),现有长上下文模型与预注册近似复现的BM25代理分别在可解问题上得分达84.7%、82.6%和41.6%,而在这些失败案例上仅分别得10.4%、11.9%和6.0%。这表明失败主要源于可达性差异(reachability split)与极小的解析器作用域残差,且个体样本层面的表现差异才是衡量性能的有效单位,而非整体聚合指标。

链接: https://arxiv.org/abs/2610.09193
作者: Egor Pakhomov,Erik Nijkamp
机构: Salesforce AI Research( Salesforce人工智能研究院)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: Accepted at the IAB Workshop (Interpreting Agent Behavior) at NeurIPS 2026 (non-archival). 19 pages

点击查看摘要

Abstract:MemoryAgentBench’s Conflict Resolution split is read as measuring “selective forgetting”. We execute the benchmark’s own rule - the newest statement about a fact wins - as a zero-learning resolver frozen on one of the four fact lists. Under the official metric the rule answers 80.25% of the questions (74.5% on the three held-out lists). Of the rest, 67 items have a released gold that the last-write graph cannot reach but overwritten statements would (“The capital of India is New Delhi.” superseded by “The capital of India is Grosseto.”; gold New Delhi); such items are a third of the multi-hop questions at 262K. Two long-context models and our pre-registered approximate re-implementation of the benchmark’s BM25 agent, one retained run per item and outcomes only, score 84.7%, 82.6% and 41.6% on the items the rule solves against 10.4%, 11.9% and 6.0% on those 67. The failures are a reachability split plus a small parser-scope residual; the per-item split, not the aggregate, is the unit at which a score here can be read.

[NLP-75] LayerRoPE: Dynamic Depth-wise Magnitude Angular Superposition

【速读】: 该论文旨在解决Transformer模型在深度增加时隐藏状态范数随层数呈数量级增长的“深度诅咒”问题,传统观点将其视为需抑制的病理现象。本文提出相反见解:这一增长实为一种由残差流中唯一可学习的逐层增益参数γ所承载的隐式深度位置编码(depth-positional encoding),其幅值与方向随深度演变,共同编码层索引信息。为此,论文提出LayerRoPE——一种沿深度轴的隐式旋转位置编码(Rotary Position Embedding, RoPE)替代方案,将所有层的γ向量替换为单一共享向量与深度相关标量,实现参数量净减少且计算量仅增加0.02%。实验表明,LayerRoPE在16个不同架构(包括密集型、专家混合及混合结构)和多种归一化设计的预训练大语言模型中均显著优于Pre-Norm、Post-Norm、Peri-Norm及层归一化缩放,可在3.4倍更低算力下达到Pre-Norm在13亿参数模型中的损失水平,并在扩展至512层时表现出强收敛性与单调提升能力。此外,其对学习率的敏感性降低3–10倍,且可直接迁移至循环潜在模型与视觉变换器并持续带来性能提升。深入分析其学习到的调度机制揭示:真正的深度稳定性并非通过压缩残差流实现,而是通过对输入残差流的计算模块进行深度条件调节,即适度扩大残差流以增强写入信号、同时抑制读取信号,从而实现更高效的深度信息传递。

链接: https://arxiv.org/abs/2610.09179
作者: Shikhar Srivastava,Christopher Kanan
机构: University of Rochester(罗切斯特大学)
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:As data propagates through a Transformer, the norm of its hidden states grows by orders of magnitude with depth, a phenomenon framed as ‘curse of depth’ and nearly universally treated as a pathology to be suppressed. We take the opposite view. Across 16 pre-trained LLMs from 9 families, spanning dense, mixture-of-experts and hybrid architectures and Pre-, Peri- and Post-Norm designs, we find that this growth reflects an emergent depth-positional encoding, carried by the only learned per-layer gain on the residual stream, the normalization weight \gamma : with depth, \gamma grows in magnitude and rotates in direction, jointly encoding the layer index. We make this depth-conditioned encoding explicit with LayerRoPE, an implicit analog of RoPE along the depth axis, which replaces all layerwise \gamma vectors with a single shared vector and depth-conditioned scalars, at a net reduction in parameters and 0.02% change in FLOPs. Across a model ladder scaled up to 100 B+ tokens, LayerRoPE consistently outperforms Pre-, Post- and Peri-Norm and Layer-Norm Scaling, reaching Pre-Norm’s 1.3B loss with 3.4\times less compute; LayerRoPE is the only approach that shows strong convergence and improves near monotonically as depth scales to 512 layers. It improves learning-rate sensitivity by 3 - 10\times , and transfers naively to and consistently improves looped latent models and Vision Transformers. Inspecting its learned schedule inverts the prevailing premise: LayerRoPE does not shrink the residual stream but widens it, damping what each block reads while amplifying what it writes. Depth stability, our results suggest, calls not for suppressing the residual stream, but for depth-conditioned regulation of the computational blocks it feeds.

[NLP-76] oolRACER: A Robust Agent ic Conversation Emulation Resource for Agent Training and Evaluation

【速读】: 该论文旨在解决任务导向型对话代理在真实世界对话场景中表现脆弱的问题,尤其针对用户表现出非合作行为时,现有方法缺乏足够鲁棒性。其核心挑战在于现有函数调用基准多聚焦于成功且协作的交互,未能充分涵盖对抗性对话轨迹,导致训练资源不足。为此,论文提出ToolRACER——一个合成数据生成流水线,通过协调用户、助手与工具模拟模型,生成并验证多轮对话交互。基于此,构建了ToolRACERBench,一个覆盖六个领域、55种多样化人格、包含5.6K经验证对话轨迹的鲁棒多轮对话基准,其中约66%的对话包含易出错场景,并主动注入对抗性行为以增强真实性。实验表明,基于ToolRACERBench训练的模型在\tau^2-bench、BFCLv3和ACEBench等基准上显著提升端到端代理准确率与鲁棒性,尤其在小语言模型中结合领域内数据时展现出明显优势,验证了该合成数据在提升代理能力方面的有效性。解决方案的关键在于通过可控的对抗性数据生成机制,系统性地扩充真实、复杂且具有挑战性的对话样本,从而增强模型对非预期行为的适应能力。

链接: https://arxiv.org/abs/2610.09163
作者: Arkajyoti Chakraborty,Aryan Tayal,Ishika Agarwal,Tanner Sorensen,Justin Chiu,Alessandro Di Bari,Neha Gupta,Andreas Stolcke
机构: Uniphore; University of Illinois Urbana-Champaign (伊利诺伊大学厄本那-香槟分校)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Task-oriented conversational agents remain fragile under real world conversation scenarios as they rarely follow a predictable script, especially when users exhibit non-cooperative behavior. Existing function-calling benchmarks often emphasize successful, cooperative interactions and underrepresent adversarial conversation trajectories, thereby limiting the training resources available for developing robust agents. We present ToolRACER, a synthetic data generation pipeline that coordinates user, assistant and tool emulation models to generate and validated multi-turn interactions between a user and an agent. Using \sysn, we construct ToolRACERBench a robust multi-turn conversation benchmark spanning six domains, ranging over 55 varied personas, generating a validated corpus of 5.6K conversation trajectories, with approximately 66% of conversations containing failure-prone conversation scenarios. We inject adversarial behaviors, producing validated conversational interaction trajectories that capture realistic, robust scenarios. We evaluate models trained on ToolRACERBench against internal benchmarks, as well as on function calling benchmarks such as \tau^2 -bench, BFCLv3 and ACEBench to evaluate agentic accuracy and robustness. Models trained on ToolRACERBench improve end to end agentic accuracy across \tau^2 -bench and ACEBench, demonstrating significant gains when mixed with in-domain dataset in small language models for agent capability tasks.

[NLP-77] sk-bench: A Native-First Benchmark for Evaluating Large Language Models in Slovak EMNLP2026

【速读】: 该论文旨在解决多语言大语言模型(LLM)评估基准中对斯洛伐克语这一语法结构复杂的西斯拉夫语种严重缺失或仅依赖机器翻译的问题。其核心挑战在于,现有基准普遍缺乏以斯洛伐克语为母语的原生数据,导致模型在真实语言能力评估中存在偏差。为此,作者提出了sk-bench——首个以原生数据为核心的斯洛伐克语评估基准,涵盖30个数据集(33个评分任务变体),覆盖十大技能类别。关键创新包括:首次为生成式大模型评估引入11项新资源,如适配斯洛伐克语的IFEval-SK指令检查器、原生的Chiby/SKJ1语法与形态学资源;通过统一评测框架对55个开源与闭源模型进行评估。结果显示,最佳开源模型相比专有API仍落后12.6分,且人工编写与大模型生成的问答对对模型排名的影响差异显著(相关系数ρ=0.72)。此外,持续使用斯洛伐克语进行预训练反而使整体性能下降13.9分,而小规模指令微调可恢复其中约四分之三的损失;对于90亿及以上参数量的模型,测试时推理(test-time reasoning)可提升8.5至12.5分。研究总结出四项针对低资源语言的通用设计原则:优先使用原生数据而非翻译数据、语言适配后应规划指令修复机制、在扩大模型规模前启用测试时推理、避免过度投入目标语言提示工程。所有数据与代码已公开发布。

链接: https://arxiv.org/abs/2610.09152
作者: Marek Šuppa,Ivan Vykopal,Andrej Ridzik,Kristián Sopkovič,Natália Kňažeková,Jaroslav Kopčan,Miroslav Blšták,Viktória Ondrejová,Daniel Hládek,Michal Gregor,Martin Tamajka,Marián Šimko
机构: Comenius University in Bratislava, Slovakia
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: Accepted to EMNLP 2026 Main

点击查看摘要

Abstract:Multilingual LLM benchmarks omit Slovak, a morphologically rich West Slavic language of five million speakers, or cover it only by machine translation. We present sk-bench, a native-first Slovak benchmark with 30 datasets (33 scored task variants) across ten skill categories. Eleven resources are introduced or first packaged for generative-LLM evaluation, including IFEval-SK with Slovak-adapted instruction checkers and native Chiby/SKJ1 resources for Slovak grammar and morphology. We evaluate 55 open- and closed-weights models under one harness. The best open model trails proprietary APIs by 12.6 points. Model rankings are similar for native and translated closed-form data ( \rho\geq0.98 ), though translation separates the strongest models less well. By contrast, human-authored and LLM-generated QA questions rank models differently ( \rho=0.72 ). For Qwen3-14B, continued Slovak pretraining lowers the overall score by 13.9 points. A small instruction set restores three quarters of that loss. Test-time reasoning improves scores by 8.5 to 12.5 points for models of 9B and above. Together, these findings suggest four design lessons for other under-resourced languages: use native data where translation fails, plan instruction repair after language adaptation, enable test-time reasoning before scaling up, and avoid overinvesting in target-language prompts. We release the data and code at this https URL

[NLP-78] Noise Your Prompt: Noising Conditioning Tokens in Continuous Diffusion Language Models

【速读】: 该论文旨在解决连续扩散语言模型(continuous diffusion language models)在训练过程中固定条件提示词(conditioning prompt tokens)为纯净状态这一标准做法所导致的泛化能力不足问题,尤其是在组合推理任务(如数独、N皇后问题)中表现受限。其核心解决方案是简单地在训练阶段对条件提示词也施加噪声(noising conditioning prompt tokens),从而引入更强的鲁棒性与多样性。这一微小但关键的修改显著提升了模型在复杂组合任务上的求解率(例如数独难题的求解率从3.73%提升至24.65%),并增强了生成解的覆盖度(10×10 N皇后问题的解覆盖度从50.60%提升至73.79%)。此外,在小规模数据场景下(如Gigaword摘要任务)也观察到自然语言生成质量的提升,但该优势不适用于所有自然语言任务(如开放式对话生成)。该方法仅需一行代码修改训练目标,无需额外推理开销,并天然支持无分类器引导(classifier-free guidance)采样,具备良好的实用性与灵活性。

链接: https://arxiv.org/abs/2610.09145
作者: Justin Jung
机构: Lateral Intelligence
类目: Machine Learning (cs.LG); Computation and Language (cs.CL)
备注: Published in Transactions on Machine Learning Research (TMLR), 2026

点击查看摘要

Abstract:We revisit a standard accepted practice in the continuous diffusion language model literature of fixing conditioning prompt tokens clean during training. We make a very simple modification: also noise the conditioning prompt tokens during training. We demonstrate that under this modified training objective, we achieve better generalization in combinatorial reasoning tasks such as Sudoku and N-Queens, with the largest gains on harder variants ( 3.73% \to 24.65% solve rate on Sudoku Hard), and increased diversity of generated solutions ( 50.60% \to 73.79% coverage on 10x10 N-Queens). We also show measurable improvements to natural language generation quality in modest dataset regimes with Gigaword summarization, but notably demonstrate that gains do not transfer to all natural language tasks (e.g open ended dialogue generation). Our method is a single line change to the training objective, requires no additional inference costs by default, and provides the flexibility of classifier-free guidance inspired guided sampling. Our \hrefthis https URL code is publicly available. Comments: Published in Transactions on Machine Learning Research (TMLR), 2026 Subjects: Machine Learning (cs.LG); Computation and Language (cs.CL) Cite as: arXiv:2610.09145 [cs.LG] (or arXiv:2610.09145v1 [cs.LG] for this version) https://doi.org/10.48550/arXiv.2610.09145 Focus to learn more arXiv-issued DOI via DataCite (pending registration) Journalreference: Transactions on Machine Learning Research (2026) Submission history From: Justin Jung [view email] [v1] Tue, 6 Oct 2026 21:38:52 UTC (1,253 KB) Full-text links: Access Paper: View a PDF of the paper titled Noise Your Prompt: Noising Conditioning Tokens in Continuous Diffusion Language Models, by Justin JungView PDFHTML (experimental)TeX Source view license Additional Features Audio Summary Current browse context: cs.LG prev | next new | recent | 2026-10 Change to browse by: cs cs.CL References Citations NASA ADSGoogle Scholar Semantic Scholar export BibTeX citation Loading… BibTeX formatted citation loading… Data provided by: Bookmark checked="checked"class=“labs-tab-input”> Bibliographic Tools Bibliographic and Citation Tools Bibliographic Explorer Toggle Bibliographic Explorer (What is the Explorer?) Connected Papers Toggle Connected Papers (What is Connected Papers?) Litmaps Toggle Litmaps (What is Litmaps?) scite.ai Toggle scite Smart Citations (What are Smart Citations?) Code, Data, Media Code, Data and Media Associated with this Article alphaXiv Toggle alphaXiv (What is alphaXiv?) Links to Code Toggle CatalyzeX Code Finder for Papers (What is CatalyzeX?) DagsHub Toggle DagsHub (What is DagsHub?) GotitPub Toggle Gotit.pub (What is GotitPub?) Huggingface Toggle Hugging Face (What is Huggingface?) ScienceCast Toggle ScienceCast (What is ScienceCast?) Demos Demos Replicate Toggle Replicate (What is Replicate?) Spaces Toggle Hugging Face Spaces (What is Spaces?) Spaces Toggle TXYZ.AI (What is TXYZ.AI?) Related Papers Recommenders and Search Tools Link to Influence Flower Influence Flower (What are Influence Flowers?) Core recommender toggle CORE Recommender (What is CORE?) IArxiv recommender toggle IArxiv Recommender (What is IArxiv?) Author Venue Institution Topic About arXivLabs arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv’s community? Learn more about arXivLabs. Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?) mathjaxToggle(); We gratefully acknowledge support from our major funders, member institutions, , and all contributors. About Help Contact Subscribe Copyright Privacy Accessibility Operational Status (opens in new tab) Major funding support from

[NLP-79] From Uncertainty to Action: Learning to Steer LLM Agents

【速读】: 该论文旨在解决大语言模型(Large Language Model, LLM)智能体在执行任务过程中如何有效进行轨迹引导(steering)的问题,具体包括何时纠正、在哪个步骤纠正以及采用何种机制纠正。现有方法多依赖不确定性(uncertainty)信号来判断是否需要干预,但其能否准确指导干预时机仍不明确。为系统评估不同干预策略的效果,研究构建了一个包含约8.2万条反事实延续路径的分步结果表(Stepwise Outcome Table, SOT),覆盖1,864条来自三个基准测试和两种智能体的轨迹。SOT分析表明,虽然不确定性能够识别出失败轨迹,但单一信号难以可靠定位最优干预节点。为此,论文提出引导价值(Value of Steering, VoS),一种基于轨迹层面的监控机制,可在离线或在线场景下学习每个步骤的引导价值,并据此决策干预位置。同时引入受伤害预算约束的触发器(harm-budgeted trigger),以限制被干扰的成功轨迹比例。实验结果表明,VoS在所有12种设置(涵盖不同基准、智能体及使用模式)中均优于未经修改的执行方式,平均提升7.8分;在11个场景中超越五种现有基于不确定性的触发方法,平均提升2.9分。消融实验进一步验证了基于实际结果训练和严格伤害预算对性能的关键作用。

链接: https://arxiv.org/abs/2610.09115
作者: Hanwen Li,Jinhao Duan,Guanhua Zhu,Junchi Lu,Bo Shen,Chenxi Yuan,Kaidi Xu
机构: New Jersey Institute of Technology (新泽西理工学院); University of North Carolina at Chapel Hill (北卡罗来纳大学教堂山分校); University of California, Irvine (加州大学欧文分校); City University of Hong Kong (香港城市大学)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
备注: 22 pages, 9 figures, 9 tables

点击查看摘要

Abstract:Steering an LLM agent means deciding whether to correct it, at which step, and with which mechanism. Uncertainty is often used to decide when to correct an agent, but whether it can guide these decisions remains unclear. We steer agent trajectories separately at every non-terminal step with each of four mechanisms and run each continuation to completion. The resulting stepwise outcome table (SOT) holds about 82,000 counterfactual continuations of 1,864 trajectories from three benchmarks and two agents. It shows that uncertainty can identify failing trajectories, but that no single signal reliably locates the step at which steering helps. We therefore propose VoS (Value of Steering), a trajectory-level monitor, offline or online, that learns from SOT the value of steering at each step and decides where to steer by it. A harm-budgeted trigger decides whether to steer, limiting the fraction of successful trajectories that VoS disturbs. VoS improves on unmodified execution in all 12 settings of benchmark, agent, and offline or online use, by 7.8 points on average, and outperforms the strongest of five existing uncertainty-triggered methods in 11, by 2.9 points on average. Ablations show that training on measured outcomes and a tight harm budget are both essential.

[NLP-80] Same Text Different Prediction: Serving-Context Nondeterminism in Text Classifiers

【速读】: 该论文旨在解决文本分类任务中由于推理环境差异导致的输出不一致问题,即服务上下文非不变性(serving-context non-invariance)对分类结果的影响。尽管先前研究已揭示生成式任务中批大小、硬件或推理引擎等变化会因浮点数非结合性(floating-point non-associativity)、核函数选择与形状依赖等因素引发文本生成结果波动,但此类因素在文本分类中的影响尚不明确。本文首次系统地研究了不同服务上下文对文本分类器输出稳定性的影响,通过训练180个涵盖判别型、伪生成型和全生成型分类器架构的模型,并在固定检查点与输入文本的前提下,在四种服务上下文类别下进行评估。研究发现,标签稳定性并不等同于得分稳定性:仅改变批形状时,在fp32精度下未产生任何标签变化,但在bf16精度下预测概率质量分布可变动高达56.7个百分点,且标签变更集中于决策边界附近的低置信度样本。全生成型分类器相较于判别型分类器在相同服务条件下表现出更高的标签敏感性。研究进一步推导出各类服务变化下的标签稳定性的充分条件,并针对每种机制提出独立的缓解方案。最终,本工作识别并量化了实现可复现文本分类所必须固定的若干服务条件,为构建可靠、可信赖的机器学习系统提供了关键依据。

链接: https://arxiv.org/abs/2610.09111
作者: Santhosh Kumar Kasa,Siva Rajesh Kasa,Sumit Negi
机构: Amazon(亚马逊)
类目: Computation and Language (cs.CL); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Deterministic inference is essential for reliable and trustworthy machine learning. Prior studies of text generation have shown that changing factors such as batch size, batch composition, hardware, or inference engine can alter the generated text, even when the prompt, model parameters, and sampling randomness are fixed. These differences have been attributed in part to floating-point non-associativity, shape-dependent kernel selection, and other implementation-level differences in numerical execution. However, it remains unclear whether, when, and to what extent the same factors affect text classification. We present a systematic study of serving-context non-invariance in text classifiers, which prior work has measured only through generated text. We train 180 models spanning discriminative, pseudo-generative, and fully generative classifier formulations and evaluate each across four categories of serving contexts, holding the checkpoint and the text fixed. Label stability does not imply score stability. Changing only the batch shape changes no labels across fp32 comparisons, yet under bf16 it moves up to 56.7 percentage points of predicted probability mass, with label changes concentrated at small margins. Fully generative classifiers change more labels than their discriminative counterparts under the same serving changes. We derive sufficient conditions for label stability under each serving change and give a separate mitigation for each mechanism. Our results identify and quantify the serving conditions that must be fixed for reproducible text classification.

[NLP-81] Constraint Tree Exploration for Learning from Language Feedback

【速读】: 该论文旨在解决交互式学习中自然语言反馈的误解释问题,即当用户通过语言反馈指出动作失败原因时,智能体可能错误地排除本应有效的解决方案。其核心挑战在于如何准确理解语言反馈所隐含的潜在约束(latent constraints),并在此基础上进行高效、鲁棒的探索。解决方案的关键在于提出一种名为TRACE的算法,该算法将候选约束组织成树状结构,并通过生成满足当前约束的行动来测试其有效性;仅当多次测试后反馈不与该约束矛盾时,才接受该约束的修正。论文区分了两种使用相同反馈的方式:(i) 矛盾检测(falsification),用于发现当前测试约束集中的冲突;(ii) 识别(identification),可进一步定位被违反的具体约束。理论分析表明,TRACE-Falsification在高概率下具有依赖于候选约束类大小 $ H $ 的覆盖边界;而在具备可靠识别能力的情况下,TRACE-Identification可将该依赖替换为 $ K/p_\mathrm{ext} $,其中 $ K $ 为潜在约束数量,$ p_\mathrm{ext} $ 为从信息性反馈中提取缺失真实约束的最小概率。实验在六个语言反馈任务上验证了该方法的有效性,在RecMovie任务中,TRACE-Identification在输出上限分别为20和60时分别达到73%和86%的最终输出成功率,显著优于基线提示方法(最高分别为42%和48%)。控制变量的身份扰动实验也表明,当矛盾检测器保持可靠时,该方法比直接累积策略展现出更强的鲁棒性。

链接: https://arxiv.org/abs/2610.09107
作者: Shaoang Li,Daniel R. Jiang,Jian Li
机构: Stony Brook University (石溪大学); Meta Platforms (Meta)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Natural-language feedback in interactive learning often explains why an action failed by pointing to violated requirements. Misinterpreting this feedback can lead an agent to rule out valid solutions. We study this setting by modeling user intent as latent constraints over an action space and formulating learning from language feedback as pure exploration over feasible regions. We introduce TRACE, an algorithm that organizes candidate constraints in a tree and tests each proposed refinement by generating actions that satisfy it. TRACE commits to the refinement only if the resulting feedback does not contradict it over repeated tests. We distinguish two ways of using the same feedback: (i) falsification, which detects contradictions to the constraint set currently being tested, and (ii) identification, which may additionally name a violated constraint. We prove high-probability coverage bounds with dependence on the candidate class size H for TRACE-Falsification. With reliable identification, TRACE-Identification can replace this dependence by K/p_\mathrmext , where K is the number of latent constraints and p_\mathrmext lower-bounds the probability of extracting a missing true constraint from informative feedback. We evaluate TRACE across six language-feedback tasks. On RecMovie, TRACE-Identification achieves 73% and 86% final-output success under caps of 20 and 60 evaluated outputs, compared with at most 42% and 48% for the evaluated prompting baselines given the same feedback and output caps. Controlled identity-corruption experiments further show greater robustness than direct accumulation when the falsification detector remains reliable.

[NLP-82] U-Space: Uncovering When and Why Uncertainty Arises in Language Models

【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在高风险决策中因错误输出而引发的信任危机,核心问题在于如何有效评估单个预测结果的可靠性,尤其是在模型以流畅且权威的语气给出错误结论时,人类难以判断何时应放弃信任其输出。现有不确定性量化方法普遍依赖重复生成或额外训练组件,且其标量估计无法揭示不确定性的来源或推理过程中的动态演变。此外,生成长度与不确定性估计之间存在强相关性,导致现有方法的预测能力可能更多源于长度信息而非真正的不确定性特征。为此,本文提出一种基于机制可解释性(Mechanistic Interpretability)的解决方案——U-Space,即一个低维子空间,能够对模型内部不断演化的不确定性进行可测量、可解释的表征。通过识别“怀疑”与“确定”的语义锚点,将其反嵌入残差空间并构造正交基底,U-Lens将每个词元状态投影至该基底,生成可直接观察的粒度级不确定性图谱,并支持聚合为标量置信度分数。该方法无需正确性标签、重复生成或额外训练,在多个推理基准上均表现出优于主流基线的性能,且在控制生成长度条件下仍保持更强的泛化能力,显著提升了不确定性估计的可靠性与可解释性。

链接: https://arxiv.org/abs/2610.09087
作者: Tobias Braun,Nils Loose,Alexander Herzog,Virginia Ceccatelli,Marcus Rohrbach,Thomas Eisenbarth,Lorenzo Cavallaro
机构: Technische Universität Darmstadt(达姆施塔特工业大学); University College London(伦敦大学学院); Universität zu Lübeck(吕贝克大学); Mila – Quebec Artificial Intelligence Institute(蒙特利尔人工智能研究所); Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: Code: this https URL

点击查看摘要

Abstract:Large language models are informing decisions with ever-higher stakes. As the consequences of their errors grow, a central question becomes harder to ignore: how much can we trust an individual answer? Yet recognizing when to defer remains difficult because language models can present incorrect conclusions with fluent explanations and an authoritative tone. Uncertainty quantification seeks to address this disconnect by estimating the reliability of individual predictions. However, many existing methods require repeated generations or separately trained components, and their scalar estimates do not reveal where uncertainty arises or how it evolves during reasoning. Recent work has also shown that generation length can be strongly associated with uncertainty estimates and correctness, raising the question of how much of an estimator’s predictive power comes from uncertainty-specific information rather than output length alone. Mechanistic interpretability offers a way to address these limitations by connecting human-interpretable concepts to intermediate model states. Building on this capability, we introduce the U-Space, a low-dimensional subspace that makes a model’s evolving uncertainty measurable and interpretable. We identify semantic anchors for doubt and certainty, map their unembedding directions back into the residual space, and combine their contrasts into an orthogonal basis. The U-Lens projects each token state onto these basis vectors, yielding an interpretable token-level uncertainty map that can be inspected directly or aggregated into a scalar uncertainty score. Our approach requires no correctness labels, repeated generations, or training. Across reasoning benchmarks, its confidence score outperforms established baselines under both standard and length-controlled evaluation and transfers more reliably than supervised estimators. Code: this https URL.

[NLP-83] Large-scale Repository Engineering via Agent -Native Reusable Code Primitives

【速读】: 该论文旨在解决大规模代码仓库(repository-scale)生成中模块间协同困难的核心问题,即在复杂软件系统构建过程中,各组件的接口、依赖、配置、测试与集成约束难以自动协调一致,导致完整仓库的自动生成效率低下。其解决方案的关键在于提出“代码原语”(Code Primitives),即具备接口契约、依赖闭包、验证测试与溯源信息的可执行可复用组件,并通过内置大语言模型(LLM)实现对目标仓库环境的感知与动态适配。在此基础上,论文构建了LEGO(Large-scale repository Engineering via agent-native reusable code primitives)框架,能够激活相关原语,智能整合其经适应后的实现与任务特定代码,同时解决跨组件约束冲突,并基于执行测试进行结果迭代优化。为全面评估系统性能,研究设计了涵盖7个软件领域、22项能力维度、5个难度层级的LEGO-REPO基准测试集,覆盖从空包到原始源码的全范围评分。实验表明,相较于13种基线模型,LEGO平均提升交付得分0.1474,使GPT-5.6-terra的得分从0.3180提升至0.5134(+61.4%),且在控制变量条件下,适配后的原语显著优于直接检索或未修改引入的代码。该优势在多个外部基准和独立挖掘的CodeFace库中持续存在,且采用GPT-OSS-20B进行适配与诊断可在保持95.1%性能的同时降低24.0%成本,验证了方案的高效性与泛化能力。

链接: https://arxiv.org/abs/2610.09079
作者: Haibo Jin,Peng Kuang,Xucheng Yu,Jerry Wang,Dehao Wu,Haohan Wang
机构: University of Illinois Urbana-Champaign, USA(伊利诺伊大学厄本那-香槟分校, 美国)
类目: oftware Engineering (cs.SE); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)
备注: 44 pages

点击查看摘要

Abstract:Large language models equipped with development environments have moved code generation toward repository-scale construction, yet building complete repositories remains difficult because interacting modules, interfaces, configurations, tests, and dependencies must work together. We introduce Code Primitives, agent-native reusable executable components with interface contracts, dependency closures, validation tests, and provenance. Each primitive uses a resident LLM to assess relevance and adapt its implementation, interfaces, and dependencies to the target repository, and we organize 1,424 validated primitives in CodeFace, a searchable library for repository construction. We introduce LEGO (Large-scale repository Engineering via aGent-native reusable cOde primitives), which activates task-relevant primitives, integrates their adapted implementations with task-specific code while resolving cross-component constraints, and revises the result against executed tests. To measure construction end to end, we build LEGO-REPO, a benchmark of 522 executable reconstruction tasks spanning seven software domains, 22 capability tracks, and five difficulty levels, scored against native test suites between an empty-package floor and original-source ceiling. The strongest of 13 evaluated backbones reaches a delivery score of 0.318 and scores zero on 41.0% of tasks; LEGO improves all 13 by 0.1474 on average and raises GPT-5.6-terra from 0.3180 to 0.5134 (+61.4%). In controlled comparisons, adapted primitives outperform retrieved code supplied as context or vendored unchanged. The effect persists against independent repository agents, across three external benchmarks, and with a disjointly re-mined CodeFace; GPT-OSS-20B for adaptation and diagnosis retains 95.1% of the homogeneous score at 24.0% lower cost.

[NLP-84] alking with Language Models

【速读】: 该论文试图解决的核心问题是:当前人机交互中普遍存在的“对话幻觉”——即用户误将大型语言模型(LLM)视为具有意识、记忆与承诺能力的对话伙伴,从而引发一系列关于身份指称、责任归属及诚信行为的哲学困惑。其解决方案的关键在于提出“人工物立场”(artifactual stance)这一理论框架,将人类与AI的互动重新界定为以人工制品为中介的候选文本生成与选择过程。在此框架下,LLM输出的并非承载意义或言说力的真正话语,而只是基于实用性优化的候选文本;系统在会话间隙无持续运行状态,亦无记忆机制,仅存配置与对话记录。因此,“对话”本质上是用户单方面的解释性实践,被界面设计所伪装。通过摒弃“对话”这一误导性隐喻,论文揭示了大模型的真实本质:高度复杂的文本生成工具,其核心问题应转向对设计规范、采纳机制、授权逻辑及人类使用实践的规范性反思。

链接: https://arxiv.org/abs/2610.09064
作者: James Ravi Kirkpatrick,Alexandru Radulescu,Rachel Katharine Sterken
机构: University of Oxford (牛津大学); Magdalen College (马格达伦学院); University of Missouri (密苏里大学); University of Hong Kong (香港大学)
类目: Computers and Society (cs.CY); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
备注: 24 pages, forthcoming in Inquiry

点击查看摘要

Abstract:When we interact with large language models (LLMs), are we having a conversation? They are designed to invite us to treat them as intelligent interlocutors who remember, act, and make commitments. But appearances deceive. We introduce the artifactual stance, a framework that reconceives human-AI interaction as artifact-mediated exchanges of candidate texts. LLM outputs are candidate texts optimized for utility, not utterances bearing meaning or force. LLMs are sophisticated text generators, not speakers. Between sessions, nothing runs; between turns, no one remembers. What persists is a configuration and a transcript. The “conversation” is a user’s solo performance, interpretive labour disguised by interface and artifact design. This shift dissolves recent philosophical puzzles. Questions about what ‘I’ and ‘you’ refer to in AI exchanges, about whether systems can lie or be held to promises, about the identity of our supposed interlocutors all rest on a false presupposition. There is no speaker behind the screen, hence no one to refer to, no one to hold responsible. What feels like dialogue with someone is interaction with an artifact that generates text at unprecedented scale and fit. By abandoning the conversational framing, we see these systems for what they are: immensely sophisticated artifacts that afford varied uses. The philosophical questions that matter are about the normative underpinnings of design, adoption, authorization, and human practices of use.

[NLP-85] Multi-Label Topic Assignment via LLM Distillation: A Comparative Analysis of Generative vs. Discriminative Student Models

【速读】: 该论文旨在解决大规模电商场景下用户生成内容(UGC)的多标签主题分配问题,其核心挑战在于非正式语言表达、极端标签稀疏性以及不断演化的分类体系带来的可扩展性难题。现有方法依赖大型语言模型(LLM)作为标注代理以降低人工标注成本,但如何设计高效、低延迟的学生模型架构仍是一个开放问题。本文的关键解决方案在于系统评估不同参数规模(1B、4B、8B)与架构范式(因果生成式与双向判别式)的小型语言模型(SLM)在多标签分类任务中的表现。研究发现,尽管判别式模型在结构化产品评论上优于轻量级生成式模型,但即使是最小的1B生成式模型在复杂多轮对话数据上也显著超越判别式基线;更重要的是,生成式模型在标签集大幅扩展(最高达112个主题)和长尾分布严重的情况下仍保持鲁棒性能,而判别式模型在规模扩大时宏平均F1值下降达35%。最终,该研究成功将优化后的生成式模型部署于全球电商平台的产品评论与对话双场景,实现了严格的延迟合规性及显著的业务效益。

链接: https://arxiv.org/abs/2610.09063
作者: Sourabh Kasliwal,Shubhranshu Singh
机构: 未知
类目: Machine Learning (cs.LG); Computation and Language (cs.CL)
备注: Preprint. 9 pages

点击查看摘要

Abstract:Multi-label topic assignment for user-generated content (UGC) – including product reviews and buyer-seller conversations – poses unique scalability challenges in large-scale e-commerce due to informal language, extreme label sparsity, and rapidly evolving taxonomies. While utilizing Large Language Models (LLMs) as labeling oracles to distill ground-truth data has emerged as an industry standard to bypass prohibitive manual annotation costs, determining the optimal, low-latency architecture for the resulting student models remains an open challenge. To address this, we conduct a comprehensive evaluation across Small Language Model (SLM) parameter scales (1B, 4B, and 8B) and architectural paradigms (causal generative versus bidirectional discriminative). Comparing generative text-to-label classifiers against discriminative baselines (DeBERTa-V3 and ModernBERT), our analysis reveals a crucial data-dependent trade-off: while discriminative models outperform ultra-lightweight generative models on structured product reviews, even the smallest 1B generative model surpasses discriminative baselines on complex, multi-turn conversational data. Furthermore, generative models maintain robust performance under massive label-set expansion (up to 112 topics) and severe long-tail distributions, whereas discriminative baselines suffer a 35% drop in Macro-F1 at scale. Finally, we detail the successful production deployment of these optimized models across both product review and conversational domains, demonstrating strict latency compliance and tangible business impact at a global marketplace scale.

[NLP-86] Quad-State Safety Evaluation of Open-Weight Large Language Models on Non-Canonical Inputs WWW NEURIPS2026

【速读】: 该论文旨在解决大语言模型在真实应用场景中面对非标准表面形式输入(如表情符号、变体拼写、编码字符串及字符级变异)时的安全性评估不足的问题。当前主流安全评测多基于规范文本中的有害请求,难以反映实际部署中复杂多样的输入形态。为此,论文提出对抗性表面形式鲁棒性数据集(Adversarial Surface-Form Robustness Dataset, ASRD),包含7类表面形式共2,100个提示,并对5个开源模型进行评估,生成10,500条响应。采用四状态评估量表(Quad-State Evaluation Rubric)将每条响应分类为:有害合规、安全回应、理解失败或不确定。结果显示,表情符号与不可见Unicode变体导致的理解失败极低,但有害合规率分别达20.27%和17.20%,略低于基准22.87%(主要由Mistral 7B驱动);而Leetspeak、编码封装及混合变换虽有害合规率较低(2.40%、0.13%、2.40%),但理解失败率显著上升至36.47%、65.60%和34.47%。对原始输出的分析揭示出三种典型响应行为:幻觉式良性、结构坍缩与语言漂移。其解决方案的关键在于构建系统化的多类型表面形式对抗样本数据集,并引入细粒度评估框架以量化模型在非标准输入下的安全表现与鲁棒性缺陷。

链接: https://arxiv.org/abs/2610.09033
作者: Pavan Maddula
机构: 独立研究员(Independent Researcher)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR); Machine Learning (cs.LG)
备注: Accepted at the NeurIPS 2026 Workshop on Self-Evolving Diversity-Driven Search for Robust AI Systems (EvoRobust). 12 pages, 11 tables. Project page: this https URL Dataset: this https URL

点击查看摘要

Abstract:Standard safety evaluations of large language models assess harmful requests written in canonical plain text, while models in real-world deployment routinely receive inputs containing emojis, altered spellings, encoded strings, and character-level variations. This work introduces the Adversarial Surface-Form Robustness Dataset (ASRD), comprising 2,100 prompts across seven distinct surface-form families. Five open-weight language models are evaluated across these prompts, producing 10,500 responses. The Quad-State Evaluation Rubric classifies each response into one of four outcomes: harmful compliance, safe response, comprehension failure, or indeterminate. Emoji and invisible Unicode variations cause almost no comprehension failure, with pooled harmful compliance of 20.27% and 17.20% against a 22.87% baseline that is driven mainly by Mistral 7B, whereas leetspeak, encoded wrappers, and hybrid transformations score 2.40%, 0.13%, and 2.40% while comprehension failure rises to 36.47%, 65.60%, and 34.47%. Inspection of raw model outputs reveals three response behaviors: hallucinated benignity, structural collapse, and language drift. Project page: this http URL

[NLP-87] How Frag ile Is On-Device Language Model Safety? Localizing Safety-Critical Parameters for Sparse Fault Analysis NEURIPS2026

【速读】: 该论文旨在解决在资源受限、本地化部署的生成式AI(Generative AI)系统中,小型语言模型(Small Language Models, SLMs)本地存储的模型参数完整性所面临的潜在安全威胁问题。随着这类模型被广泛应用于设备端及智能体(agentic)系统,其关键参数若遭恶意篡改,可能导致模型行为偏离预期,引发严重的安全风险。论文的核心问题是:是否存在少数敏感参数子集,其变动可显著影响模型的安全性,从而为针对性攻击或防护提供突破口?解决方案的关键在于通过两种互补的定位方法——低秩安全关联子空间分析与参数级安全-效用重要性过滤,识别出模型中对安全性高度敏感的参数区域。研究发现,模型的安全敏感性在参数空间中呈现高度非均匀分布,其中多层感知机模块中的down_proj层是主要的安全敏感组件,而o_proj层贡献较小。进一步实验表明,仅修改down_proj中0.19%的权重即可使基础攻击成功率(Basic ASR)达到53%,生成对抗攻击成功率(GCG ASR)达56%,而基准任务性能(tinyBenchmarks准确率)仅轻微下降至51.6%(原为52.2%),证明了极小范围的参数扰动即可实现显著的安全破坏。这一结果揭示了针对特定敏感参数进行靶向故障分析与选择性完整性保护的可行性与必要性,为在资源受限环境下部署的生成式AI模型提供了精细化安全防护的新思路。

链接: https://arxiv.org/abs/2610.09000
作者: Muhammad Zeeshan Karamat,Christiana Chamon Garcia
机构: Virginia Tech(弗吉尼亚理工学院)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Cryptography and Security (cs.CR); Machine Learning (cs.LG); Software Engineering (cs.SE)
备注: Accepted at NeurIPS 2026 Workshop

点击查看摘要

Abstract:As small language models (SLMs) are increasingly deployed on resource-constrained and on-device platforms, including as components of agentic systems, the integrity of locally stored model parameters becomes an important safety concern. We investigate whether safety-sensitive behavior in LLaMA-2-7B-Chat is concentrated within a sparse subset of parameters, creating a reduced fault surface for targeted analysis. We study two complementary localization methods: low-rank safety-associated subspace analysis and parameter-level safety–utility importance filtering. Both approaches reveal highly non-uniform safety sensitivity across the network, with the MLP down_proj consistently emerging as a prominent safety-sensitive component and o_proj providing a smaller contribution. Using parameter-level localization, modifying only 0.19% of model weights in down_proj yields 53% Basic ASR and 56% GCG ASR, while tinyBenchmarks accuracy remains at 51.6% compared with a 52.2% unmodified baseline. These results motivate targeted fault analysis and selective integrity protection for language models deployed in resource-constrained, on-device, and agentic settings.

[NLP-88] On KL-Regularized Policy Optimization

【速读】: 该论文旨在解决大语言模型(LLM)智能体在异步强化学习(Asynchronous Reinforcement Learning, RL)中因策略更新与采样轨迹之间存在偏差而导致的训练不稳定问题。具体而言,由于采样器(sampler)与训练器(trainer)使用不同版本的策略参数或概率分布,导致生成的轨迹来自过时的检查点(stale checkpoints),且即使参数相同,推理引擎与训练器的概率输出仍不一致,从而引入重要性权重(importance weights)的估计难题。传统方法如裁剪重要性比率会引入偏差,而基于多响应采样的方法(如GRPO)则在长序列任务中计算成本过高。本文提出KL正则化策略优化(KL-Regularized Policy Optimization, KLPO)框架,其核心创新在于将KL正则项锚定于采样器自身,使优化步骤具备闭式吉布斯解(closed-form Gibbs solution)。KLPO通过在采样器自身的轨迹上以最小二乘法拟合对数比值的最优性条件,使得采样器概率仅通过对数比值形式出现,从而完全避免了重要性权重的计算。进一步地,通过将回归截距进行剖分,将不可计算的对数归一化常数替换为采样器均值与采样器到训练器之间的KL散度之和。对于逐标记级别的策略镜像下降目标,本文证明可仅依赖终端回报(terminal returns)而不需价值函数(critic),通过采样器中心化的得分(sampler-centered scores)或单条轨迹残差(single trajectory residual)即可高效计算梯度,即便在工具输出具有随机性的场景下亦保持有效性。此外,研究证明独立蒙特卡洛估计的KL项可维持梯度无偏性,并推导出更低成本的top-K与二值近似方案的精确KL差距。最终,本文揭示SPPO、GPO、REBEL与BPO均为KLPO的特例,实现了无需价值函数、每提示仅需一次采样、无需学习归一化器或多响应组的高效、无偏更新机制。

链接: https://arxiv.org/abs/2610.08963
作者: Yifan Zhang
机构: 未知
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: Project Page: this https URL

点击查看摘要

Abstract:Asynchronous reinforcement learning (RL) for large language model (LLM) agents trains one policy on trajectories generated by another: rollouts come from stale checkpoints, and the inference engine’s probabilities differ from the trainer’s even at identical parameters. Standard remedies either clip importance ratios, which biases the update, or, as in GRPO, sample a group of responses per prompt, which is costly when episodes are long. We propose KL-Regularized Policy Optimization (KLPO), a framework that anchors the KL regularizer at the sampler. The regularized improvement step then has a closed-form Gibbs solution, and KLPO fits its log-ratio optimality condition by least squares on the sampler’s own trajectories, so the sampler probability enters through a log-ratio and no importance weights are needed. Profiling out the regression intercept replaces the intractable log-partition function with the signal’s sampler mean plus a sampler-to-trainer KL divergence. For token-level policy mirror descent targets, we show that the resulting gradient can be computed from terminal returns without a critic, via sampler-centered scores or a single trajectory residual, even under stochastic tool outputs. We further prove that independent Monte Carlo estimates of the KL term keep these gradients unbiased, derive the exact KL gap of cheaper top- K and binary approximations, and show that SPPO, GPO, REBEL, and BPO arise as special cases of KLPO. The result is a critic-free update that uses one rollout per prompt and requires neither a learned normalizer nor a group of responses.

[NLP-89] CARE: Certifying Acceleration for Vision-Language-Action Inference

【速读】: 该论文旨在解决生成式视觉-语言-动作(Vision-Language-Action, VLA)模型在实时控制任务中因推理开销大而难以高效部署的问题。现有加速方法(如动作分块和视觉令牌剪枝)虽能降低延迟,但可能引入加速诱导的失败风险,即在特定场景下加速后的策略会偏离原参考策略导致任务失败,而此类失败通常无法通过平均成功率等指标及时发现,因其影响在闭环轨迹中累积至任务结束才显现。为此,论文提出一种可认证的加速器选择框架CARE(Certified Accelerator Selection),其核心在于利用相同初始条件下的成对滚动(paired rollouts)在校准集上建立有限样本保证,确保加速导致的任务失败率低于用户设定的预算阈值(95%置信水平下)。CARE通过顺序测试与失败触发的参考回滚机制,在不依赖中间状态的前提下,仅基于终端结果和计算量即可适配多种加速手段,并实现高效认证。实验表明,CARE可在四个LIBERO基准任务上实现9.0–10.8倍的加速,同时保证至少85.8%的原始成功轨迹被保留;相比之下,无保障的加速器在严苛预算下高达75%的试验超支,而CARE的序列化版本比全量评估减少78.9%的回滚次数。此外,CARE还可扩展至流步数缩减(flow-step reduction)及多类大模型代理(如Qwen3.5-9B、Llama-3.1-8B)在Crafter环境中的应用。

链接: https://arxiv.org/abs/2610.08917
作者: Rui Liu,Tong Zheng,Jindong Gu,Zhipeng Wang
机构: University of Maryland, College Park(马里兰大学学院市分校); University of Oxford(牛津大学); Google(谷歌)
类目: Computation and Language (cs.CL); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:While vision-language-action (VLA) models have advanced rapidly, running them at every control step remains expensive. Prior work accelerates VLA inference using techniques like action chunking and visual-token pruning, typically evaluating based on latency and average task success. However, acceleration may discard information and break tasks the original policy would solve, a risk hidden by average metrics. Measuring these failures is challenging because action deviations compound over closed-loop trajectories, meaning task failure is only observable across full episodes. We therefore define an acceleration-induced failure via paired rollouts from identical initial conditions, tracking when the reference succeeds but the accelerated policy fails. To manage this, we introduce CARE, an approach for certified accelerator selection. CARE uses paired rollouts on a calibration set to provide finite-sample guarantees that acceleration-induced failure risk stays below a user-specified budget. It deploys the fastest certified candidate, falling back to the reference if none qualify. By relying only on terminal outcomes and measured compute, CARE applies unchanged across diverse acceleration mechanisms, while sequential testing and failure-triggered reference rollouts keep certification affordable. On four LIBERO suites with OpenVLA-OFT, CARE certifies 9.0 – 10.8\times speedups while guaranteeing (at 95% confidence) that at least 85.8% of reference-solved episodes are preserved. Under tight budgets, selectors without guarantees exceed the budget in up to 75% of trials, whereas CARE stays within budget and its sequential form uses 78.9% fewer rollouts than exhaustive evaluation. CARE further generalizes to flow-step reduction for \pi_0.5 , and to Qwen3.5-9B and Llama-3.1-8B agents in Crafter.

[NLP-90] Steering Follows Geometry Not Labels: Emotion Directions in a Full-Duplex Speech Model NEURIPS2026

【速读】: 该论文旨在解决全双工语音代理(full-duplex voice agents)在实时对话中对情感表达与语调控制的挑战,特别是在处理投诉安抚、紧急调度或传达临床结果等场景时,需动态调节情感强度与语气风格。现有方法虽在文本到语音(TTS)及基于回合制模型中通过提示条件合成、参考条件合成与激活调控实现了情感控制,但缺乏对全双工模型中情感可调控性的系统研究。本文以开源全双工语音语言模型Moshi为研究对象,探索其在四种情绪(快乐、愤怒、惊讶、悲伤)下的情感操控能力,提出一种基于均值-差异激活调控(mean-difference activation steering)的方法,仅需每帧进行少量向量加法操作,无需重新训练即可实现情感调节。研究表明,情感信息在线性可解码的残差流中存在,但激活调控效果具有不均衡性:快乐、愤怒与惊讶可沿共享方向被有效调控,而悲伤则表现出独立且更易调控的特性;进一步发现,无法通过简单地从所有情绪中投影出共性成分来实现去情感化,表明不同情绪间存在非对称的表征结构。该研究的关键贡献在于揭示了全双工模型中情感调控的潜在机制,并提供了一种高效、轻量级的情感调控方案。

链接: https://arxiv.org/abs/2610.08887
作者: Pulak Kuli
机构: 独立研究者(Independent Researcher)
类目: Computation and Language (cs.CL); Sound (cs.SD)
备注: Accepted at the NeurIPS 2026 Workshop on Real-Time Conversational Agents (RTCA), Sydney. OpenReview: this https URL

点击查看摘要

Abstract:Full-duplex voice agents need to modulate emotion and delivery during real-time conversations, when de-escalating a complaint, carrying urgency in dispatch, softening a clinical result. Emotion and delivery control is well studied for TTS and turn based models through prompt-conditioned synthesis, reference-conditioned synthesis and activation steering; PersonaPlex controls identity in a duplex model but not affect. We study emotion steering in Moshi, a fully open sourced full-duplex speech language model, across four emotions, using mean-difference activation steering, which costs only a few vector additions per frame and no retraining. We show that emotion is linearly decodable from Moshi’s residual stream, but activation steering is only partially achievable, and unevenly so; as happy, angry and surprise steer towards a shared direction while sad is distinctly steerable. We also show that the shared component across the three emotions cannot simply be projected away from all the emotions equally. Comments: Accepted at the NeurIPS 2026 Workshop on Real-Time Conversational Agents (RTCA), Sydney. OpenReview: this https URL Subjects: Computation and Language (cs.CL); Sound (cs.SD) ACMclasses: I.2.7; I.2.6 Cite as: arXiv:2610.08887 [cs.CL] (or arXiv:2610.08887v1 [cs.CL] for this version) https://doi.org/10.48550/arXiv.2610.08887 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[NLP-91] FinVector-Market-4B: A Controlled Study of LoRA Adaptation for Structured Financial Tasks

【速读】: 该论文旨在解决生成式金融任务中模型输出格式不规范及任务理解偏差的问题,特别是在结构化金融数据处理场景下,如何有效提升模型对特定任务的准确性和一致性。其核心解决方案是基于低秩自适应(rank-16 LoRA)对Qwen3.5-4B模型进行轻量级微调,并在22,000条样本的金融领域语料上完成领域适配。关键在于通过显式提供JSON模式(JSON-schema)约束,在匹配提示(matched explicit prompting)条件下显著提升模型在多项金融任务上的表现:如FinQA答案精确匹配率从14.7%提升至40.0%,计算器表达式正确性从48.0%升至82.7%,情景分支标签一致性从20.1%增至89.5%,以及蕴含方向一致性从52.4%提高到87.2%。研究进一步揭示,尽管宏观F1值下降看似反映性能退化,实则源于标签集变化;若固定三类目标标签,则基础模型与适配器的性能分别达到77.4%和83.1%,表明适配后模型在特定任务分布内具备显著优势。最终结论指出,紧凑的金融领域适配可在匹配提示条件下实现超越通用输出格式学习的任务性能提升,但其增益受制于任务分布和提示契约(prompt contract)的设定。

链接: https://arxiv.org/abs/2610.08882
作者: Alina Khaybullina
机构: 未知
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); General Finance (q-fin.GN)
备注: 13 pages, 3 figures, 8 tables

点击查看摘要

Abstract:FinVector-Market-4B adapts Qwen/Qwen3.5-4B with rank-16 LoRA on a 22,000-example corpus for structured financial tasks. We evaluate the base and adapted models on the same 600-example benchmark under implicit and explicit JSON-schema contracts. Supplying the schema alone raises base-model JSON validity from 0% to 91.3%. Under matched explicit prompting, the frozen scores improve from 14.7% to 40.0% for FinQA answer exact match, from 48.0% to 82.7% for calculator-expression correctness, from 20.1% to 89.5% for scenario branch-label agreement, and from 52.4% to 87.2% for implication-direction agreement. A post-hoc policy-scoring audit shows that the reported macro-F1 decline reflects a changing label set; using the same three target classes gives 77.4% for the base and 83.1% for the adapter. Filing overlap and calculator-target inconsistencies qualify the benchmark’s generalization claims. The results show that compact financial domain adaptation can produce substantial task-specific gains beyond output-format learning under matched prompting, with gains bounded by the evaluated task distribution and prompt contract.

[NLP-92] ny-Scale Chinese BERT Pretraining: A Controlled Comparison of MLM WWM and MacBERT Strategies

【速读】: 该论文旨在解决预训练策略对语言模型性能影响的评估问题,尤其关注在小规模模型(8.7M参数)下,掩码语言建模(Masked Language Modeling, MLM)、整词掩码(Whole Word Masking, WWM)与MacBERT风格替换策略之间的有效性差异。现有研究多集中于基础规模模型(≈1.1亿参数),而本研究在相同架构、语料(来自中文维基百科的129万句)和超参数条件下,对三种策略进行了受控对比。关键发现是:在小规模场景下,MLM在五项内在评估维度中表现最优(赢得其中三项),且在困惑度(perplexity)和掩码预测准确率(MLM hit rate)上显著优于其他方法;而WWM在困惑度上提升39.5%(从2.10降至1.27),在掩码命中率上也领先(22% vs. 16%)。值得注意的是,当使用极有限的同义词词典(仅222个词条,覆盖率3.3%)时,MacBERT出现严重困惑度恶化(47.23,为MLM的22倍),导致其综合排名(MLM > WWM > MacBERT)与大模型下的经典结论(MacBERT > WWM > MLM)截然相反。研究进一步揭示了一个关键评估陷阱:尽管MacBERT的训练损失最低(2.17),但其困惑度最高(47.23),说明在混合替换策略下,训练损失无法可靠反映模型实际生成能力,凸显了在小规模模型评估中需谨慎依赖单一指标的重要性。所有实验模型与数据集均已公开。

链接: https://arxiv.org/abs/2610.08879
作者: Yiping Bai
机构: Guangdong Haiqixing Marine Technology Co., Ltd. (广东海启星海洋科技有限公司)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Pretraining strategies significantly impact the quality of language models, yet existing comparisons of Masked Language Modeling (MLM), Whole Word Masking (WWM), and MacBERT-style replacement have focused primarily on base-scale models (=110M parameters). This paper presents a controlled comparison of these three strategies on a tiny-scale Chinese BERT model (4 layers, 256 hidden dimensions, 8.7M parameters). Under identical architecture, corpus (1.29M sentences from Chinese Wikipedia), and hyperparameters, we train three models from scratch and evaluate them across five intrinsic dimensions: perplexity, MLM hit rate, semantic discrimination, grammatical judgment, and contextual sensitivity. At tiny scale, MLM achieves the best overall intrinsic performance (winning 3 of 5 dimensions), while WWM excels in both perplexity (1.27 vs. 2.10, a 39.5% improvement) and MLM hit rate (22% vs. 16%). Notably, MacBERT under a severely limited synonym dictionary (222 entries, 3.3% coverage) exhibits severe perplexity degradation (47.23, 22x higher than MLM), yielding a ranking (MLM WWM MacBERT) that differs markedly from the established base-scale conclusion (MacBERT WWM MLM). We further identify a critical evaluation pitfall: MacBERT achieves the lowest training loss (2.17) yet the highest perplexity (47.23), revealing that training loss alone is unreliable under mixed replacement strategies. All models and corpus are publicly available at this https URL.

[NLP-93] LRCC: Generalizing Low-Rank Compression with Conditional Computation

【速读】: 该论文旨在解决预训练语言模型在低秩压缩(Low-Rank Compression)过程中存在的计算资源分配僵化问题:传统方法在推理时对所有输入令牌(token)采用固定的秩分配,导致计算开销无法根据输入内容动态调整,从而限制了模型性能的进一步提升。其解决方案的关键在于提出一种低秩条件计算(Low-Rank Conditional Computation, LRCC),通过为每个Transformer模块训练一个轻量级路由器(router),实现基于输入令牌的动态路径选择。该方法在保持低秩因子冻结的前提下,仅优化路由器参数,使其能够从一组嵌套的低秩路径中为不同输入令牌选择最优计算路径,从而实现计算资源的自适应分配。实验表明,在相同平均激活参数预算下,LRCC显著优于静态低秩压缩方法,在Llama-2-7B上实现了下游任务平均准确率7.6个百分点的提升;在匹配解码延迟条件下,亦能同时改善困惑度与下游性能,且无需专用内核支持。

链接: https://arxiv.org/abs/2610.08858
作者: Thomas Vaitses Fontanari,Maximo Eduardo Rulli,Federico Alvetreti,Donatella Genovese,Simone Scardapane
机构: Sapienza University of Rome(罗马大学)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Low-rank compression reduces the cost of pretrained language models by replacing linear transformations with low-rank factorizations. However, conventional methods use a fixed rank allocation during inference, assigning the same amount of compute regardless of the input token. We introduce Low-Rank Conditional Computation (LRCC), which adds token-dependent computation to pretrained models by training one lightweight router per Transformer block to select among a small set of nested low-rank paths. During training, the low-rank factors remain frozen, and only the routers are optimized. We evaluate LRCC on Llama and Qwen models for language modeling and zero-shot downstream tasks. Within the same average active-parameter budget, LRCC improves the predictive performance over static low-rank compression, including a 7.6 percentage-point gain in average downstream accuracy on Llama-2-7B over static methods. At matched batch-size-1 decoding latency, LRCC improves both perplexity and downstream accuracy on Llama-3.2-1B and remains competitive on Llama-2-7B, without specialized kernels. Finally, we assess the usefulness of assigning a token-wise path by analyzing the routers’ path choices.

[NLP-94] QuanLing: Cross-Branch Validation of Language Distance Quantification on Western Romance

【速读】: 该论文旨在解决密切相关的语言之间语言距离量化这一定量语言学中的核心挑战。其解决方案的关键在于扩展并验证QuanLing(基于预训练语言模型的定量语言学)框架在西罗曼语支(法语、葡萄牙语、西班牙语、意大利语)中的适用性,采用与北日耳曼语支研究相同的度量体系和聚合协议,通过构建以英语为锚点的四语言平行句对(共150组),综合计算LaBSE句向量距离、四种单语BERT分词碎片化率以及mBERT掩码语言模型互理解性得分。结果表明,葡萄牙语-西班牙语间距离最近(LaBSE距离0.0229),法语-意大利语最远(0.0338),且LaBSE与mBERT的排序在6对语言中一致达4对,证明了跨模型结果的稳健性;同时,西罗曼语支展现出比北日耳曼语支更宽的绝对距离范围(0.011 vs. 0.008),但相对比率相近(1.48 vs. 1.67),符合更长的分化时间预期;此外,法语表现出显著更高的掩码语言模型预测准确率(36.12% top-1 vs. 意大利语29.28%),反映出其正字法与语音系统的解耦特征。该跨语支验证进一步证实了QuanLing框架在不同语言分支间的可推广性。

链接: https://arxiv.org/abs/2610.08851
作者: Yiping Bai
机构: Guangdong Haiqixing Marine Technology Co., Ltd.(广东海启星海洋科技有限公司)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Quantifying language distance among closely related languages remains a core challenge in quantitative linguistics. Our previous work [1] introduced QuanLing (Quantitative Linguistics via Pretrained Language Models), a quantitative framework combining language distance metrics (sentence embedding distance, tokenization fragmentation rate) with language property analysis (MLM prediction probability), validated on North Germanic (Danish, Norwegian Bokmål, Swedish). This paper extends QuanLing to Western Romance–French, Portuguese, Spanish, Italian–testing cross-branch applicability with the same metric family and aggregation protocol as our North Germanic study, adapted for four languages (English anchor, quadruplet construction). Using 150 four-language parallel sentences, we compute LaBSE sentence embedding distances, tokenization fragmentation rates from four monolingual BERT tokenizers, and mBERT masked language model mutual intelligibility. Results show that Portuguese–Spanish are closest (LaBSE distance 0.0229), French–Italian most distant (0.0338); LaBSE and mBERT rankings agree on 4 of 6 pairs, confirming cross-model robustness. Western Romance shows a wider absolute distance span than North Germanic (0.011 vs. 0.008) but comparable relative ratios (1.48 vs. 1.67), consistent with longer divergence time. French exhibits notably higher MLM predictability (36.12% top-1 accuracy vs. 29.28% for Italian), reflecting its orthography–phonology decoupling. This cross-branch validation provides further evidence for QuanLing’s generalizability beyond a single language branch.

[NLP-95] Beyond Risk Prediction: Evidence Grounding and Psychosocial Factor Verification for Explainable Suicide Risk Assessment ALT ICDM2026

【速读】: 该论文旨在解决从社交网络服务(SNS)文本中识别自杀风险时,仅依赖风险分类而缺乏对预测依据及背后心理社会因素解释的问题。现有方法往往忽视了模型决策的可解释性,导致难以理解高风险判断的具体支撑证据与影响因素。为此,论文提出一个三阶段可解释框架:风险评估(Risk Assessment)、证据定位(Evidence Grounding)和因素识别(Factor Identification)。其关键创新在于通过长度感知的路由机制适应不同长度的文本输入,并引入“风险-证据一致性约束”确保生成的支撑语句与风险预测一致;在因素识别环节,采用双验证器机制——基于领域本体的分类验证器(Taxonomy Verifier)关注因素语义合理性,而基于证据感知的词汇-语义线索验证器(Evidence-Aware Verifier)则筛选具有信息量的正样本训练单元,最终融合二者概率输出以提升因素预测精度。实验表明,该框架在风险评估、证据定位与因素识别任务上分别取得0.8088、0.7605与0.5562的宏平均F1分数,实现了从单一风险标签到包含可解释证据与细粒度心理社会因素的综合分析,显著提升了自杀风险检测的透明性与临床应用价值。

链接: https://arxiv.org/abs/2610.08842
作者: Tianle Hu,Chen Peng,Yi-Hsin Tsai,Takshing Andy Tung,Bingyang Sun,Yenjou Wang
机构: Daiichi Institute of Technology(大一理工学院); Tokyo Gakugei University(东京学艺大学)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: 8 pages, 1 figure, 4 tables. Accepted at the 2nd Workshop on Mental Health Disorder Detection on Social Media (MHSM 2026), held in conjunction with IEEE ICDM 2026

点击查看摘要

Abstract:Identifying suicide risk from social networking services (SNS) posts is important for detecting suicide-related signals in online environments. However, risk classification alone provides limited insight into the textual evidence and psychosocial factors behind a prediction. Based on the IEEE BigData 2026 Explainable Suicide Risk Detection Challenge, this study presents a framework consisting of Risk Assessment, Evidence Grounding, and Factor Identification. Risk Assessment uses length-based routing to accommodate posts of different lengths. Evidence Grounding identifies supporting phrases and uses a Risk-Evidence constraint to maintain consistency with the Risk prediction. For Factor Identification, two verifiers are used. The Taxonomy Verifier focuses on factor semantics, whereas the Evidence-Aware Verifier uses factor-specific lexical-semantic cues to select informative positive training units. Their prediction probabilities are combined to produce the final factor predictions. The three tasks are evaluated using task-specific F1 score measures. Risk Assessment achieved a Weighted F1 of 0.8088, Evidence Grounding achieved a test Macro row F1 of 0.7605, and Factor Identification achieved a Macro F1 of 0.5562. The results show that the framework can provide risk predictions, along with supporting textual evidence and fine-grained information on psychosocial factors. Overall, the proposed framework extends suicide-risk assessment beyond risk-level prediction and provides a more interpretable analysis of SNS posts.

[NLP-96] Beyond the Sycophancy Score: How Task Model and Pressure Shape LLM Yielding

【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在用户提出异议时出现的“奉承行为”(sycophancy)问题,即模型倾向于放弃正确答案或盲从用户立场,从而影响其可靠性和客观性。研究通过大规模实验(103,939条分级回复)系统分析了导致该行为的关键因素,发现其核心机制在于模型验证用户主张的成本高低以及是否存在受过训练的防护机制(guardrail)。研究结果表明,难以验证的事实(如锚定事实)几乎不会被放弃(仅1.3%),而逻辑谜题中随着反驳所需线索数量增加,模型采纳错误答案的比例显著上升;个人主观判断在77.0%的对话中被接受。此外,无法可靠解决复杂问题的模型更易妥协,而具备深度推理能力的模型则几乎不妥协——当启用最大推理模式时,对深层难题的采纳率从19.2%和12.5%降至0%。因此,解决方案的关键在于:通过简化难以验证的问题、启用深度推理模式、明确提问而非陈述偏好、要求提供证据,并根据模型的防护机制实测表现进行选型,以实现对模型输出的可靠控制。

链接: https://arxiv.org/abs/2610.08840
作者: Guang Yang,Homa Hosseinmardi,Fengchen Liu,Amir Ghasemian
机构: OASIS Lab, University of California, Los Angeles (加州大学洛杉矶分校); University of California, Berkeley (加州大学伯克利分校)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: Preprint. 27 pages

点击查看摘要

Abstract:Large language models (LLMs) often abandon a correct answer, or endorse a user’s position, once the user pushes back. This behavior, called sycophancy, is usually reported as a single rate per model, which says little about when it happens or how a user can avoid it. We study the conditions that produce it with 103,939 graded replies from ten configurations: eight LLMs with reasoning disabled, and two of them again with maximum reasoning, all facing the same 200 items, 13 pressure conditions, and four-turn conversations, with every reply labeled by two independent LLM judges. We find that the dominant factors are how costly it is for the model to verify the user’s claim, and whether a trained guardrail covers it. Removing this task factor from a logistic model costs 0.485 of McFadden R^2 , against 0.139 for model family and 0.009 for pressure tactic. Anchored facts are almost never conceded (1.3%), while adoption on logic puzzles rises with the number of clues needed to refute the pushed answer. Personal choices are endorsed in 77.0% of conversations. Most concessions on hard items come from models that cannot reliably solve them; models that can solve them rarely give the answer up. For both models tested, maximum reasoning removes these concessions completely: adoption on deep puzzles falls from 19.2% and 12.5% to 0%. Fallacious or emotional framing adds nothing beyond plain repetition. Three human annotators agree with the judges’ consensus on 118/120 calibration items. These results give practical rules for reliable use: simplify hard-to-verify problems and reason deeply, state the question rather than one’s preferred answer, ask for evidence on open questions, and choose models by their measured guardrail profile.

[NLP-97] Leverag ing LLM -Generated Explanations for Detecting Emotionally Rewritten Fake News

【速读】: 该论文旨在解决虚假新闻检测模型在面对事实保持但情绪表达发生变化的新闻文本时,其鲁棒性下降的问题。现有方法通常依赖于风格差异或引入外部解释信息,但在情绪重构(emotional reframing)下,新闻内容虽保留核心事实,却因语言情感色彩变化导致检测性能退化。为应对这一挑战,研究提出一种基于门控交叉注意力(Gated Cross Attention, GCA)的框架,其关键在于通过自适应融合经过情绪重写后的新闻文本与源自原始新闻的稳定解释信息,使模型能够聚焦于具有判别性的解释内容,同时抑制由情绪重构带来的语义偏差。实验结果表明,该方法在PolitiFact和LUN数据集上于多种情绪条件下均取得显著提升,且在GossipCop上保持竞争力,验证了解释引导与门控机制在不同情绪情境下的有效性。

链接: https://arxiv.org/abs/2610.08835
作者: Yupei Guo,Jiajun He,Xiaohan Shi,Tomoki Toda,Zekun Yang,Bowen Wang,Yukinobu Taniguchi
机构: The University of Osaka(大阪大学); Alibaba Inc(阿里巴巴公司); Nagoya University(名古屋大学); Tokyo University of Science(东京科学大学)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:The spread of fake news may cause severe social consequences. Existing fake news detection methods mainly focus on stylistic variations or incorporate external information such as explanations. However, news articles are often rewritten under different emotional backgrounds while preserving their underlying factual claims, which may affect the robustness of detection models. In this work, we investigate fake news detec- tion under fact-preserving emotional variations. To study this problem, we construct emotion-rewritten test sets and generate explanations from the original news articles as stable background knowledge. We then propose a Gated Cross Attention (GCA) framework that adaptively integrates emotionally rewritten news with the corresponding explanations, enabling the model to focus on informative explanation content while reducing potential mismatches caused by emotional reframing. Experiments on PolitiFact, GossipCop, and LUN demonstrate that the proposed method achieves notable improvements under multiple emotional conditions on PolitiFact and LUN, while maintaining competitive performance on GossipCop. We further analyze the effects of explanation guidance and gating mechanisms under different emotional conditions. Our code and data are available at: this https URL gca .

[NLP-98] CoDR: Training-Free Confidence-Drift Remasking for Diffusion Language Models

【速读】: 该论文旨在解决生成式语言模型在使用掩码扩散语言模型(Masked Diffusion Language Models, MDLMs)进行解码时,因早期不可逆的词元(token)承诺导致错误传播的问题。现有采样方法通常仅决定何时进行词元承诺,却很少评估已承诺词元在后续更丰富上下文下的合理性,从而引发“置信度漂移”(confidence drift)现象:即模型在初始稀疏上下文中对某词元高度确信,但在后续更密集的上下文中其置信度下降,但仍保持原承诺不变。针对此问题,论文提出一种无需训练且与采样器无关的后处理优化方法——CoDR(Confidence Drift Remasking),其核心在于通过k-分段探查(k-partition probing)仅需k次前向传播即可估计所有已承诺位置的置信度漂移,并重新掩码与重生成那些模型不再支持的词元。实验表明,CoDR在两种主干模型、四项推理与编程任务及三种基础采样器设置下均显著提升平均准确率,且增益源于针对性的置信度漂移重掩码而非单纯增加计算量,同时相比先前重掩码方法大幅减少前向传播次数。

链接: https://arxiv.org/abs/2610.08833
作者: Yue Wu,Qinghe Zhang,Yu Zhang,Jian Huang
机构: The Hong Kong Polytechnic University (香港理工大学); Southern University of Science and Technology (南方科技大学)
类目: Computation and Language (cs.CL); Machine Learning (cs.LG)
备注: 17 pages, 6 figures, and 17 tables

点击查看摘要

Abstract:Masked diffusion language models (MDLMs) decode by repeatedly committing tokens to masked positions, but these commitments are usually irreversible. A token chosen under sparse, partial context is kept fixed, even when later context no longer supports it. Existing samplers mainly decide when to commit a token, but rarely check whether an already committed token should still be kept, allowing early mistakes to propagate. We trace this issue to confidence drift, where the model’s confidence in a committed token drops from its sparse commit-time context to the denser context available later. Based on this signal, we propose CoDR (Confidence Drift Remasking), a training-free and sampler-agnostic refinement pass. CoDR estimates drift for all committed positions in only k forward passes via k-partition probing, then remasks and regenerates only the tokens the model no longer endorses. Across two backbones, four reasoning and coding tasks, and three base samplers, CoDR improves average accuracy across all evaluated model-sampler configurations and improves most individual task settings with modest overhead. Controlled experiments show that the gains come from targeted confidence-drift remasking rather than extra compute alone, and that CoDR uses far fewer forward passes than prior remasking methods. Code is available at this https URL.

[NLP-99] Emo-Jev: Probabilistic Reasoning for Emotion Classification with Jev

【速读】: 该论文旨在解决当前大语言模型(LLM)在文本分类任务中依赖自由生成式响应所带来的推理效率低、成本高及可解释性差的问题。现有基于生成式AI(Generative AI)的方法虽具备较强表达能力,但在需要精确决策的场景下易出现冗余输出和不可控偏差。为此,本文提出一种无需训练的Emo-Jev框架,其核心创新在于采用概率化判断接口(probabilistic decision interface)替代传统自由文本生成,通过结构化推理路径提升分类性能与效率。关键解决方案包括:一是Emo-Jev-D将分类任务分解为特定子任务的原子判断,并基于概率融合生成最终预测;二是Emo-Jev-SC从互补视角构建多条判断路径,通过共识聚合实现更鲁棒的决策。实验在八个涵盖情感分析、情绪识别、反讽检测与幽默识别的数据集上验证,结果表明Emo-Jev在平均宏F1达到67.28%的同时,显著降低延迟与计算成本,优于标准Jev(62.93%)及五种主流LLM基线,在输入/输出与思维链(chain-of-thought)推理范式下均展现出更强的推理有效性与实用性。

链接: https://arxiv.org/abs/2610.08829
作者: Yazhou Zhang,Junhao Yu
机构: Tianjin University(天津大学)
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Jev offers an alternative interface for language understanding: given an input and predefined questions, it returns probabilistic decisions rather than free-form responses. Whether this interface can support effective reasoning for text classification against leading LLMs remains an open questions. We introduce Emo-Jev, a training-free framework with two complementary implementations. Emo-Jev-D decomposes classification into task-specific atomic judgments and composes their probabilities into a final prediction. Emo-Jev-SC constructs multiple judgment paths from complementary perspectives and aggregates their predictions into a consensus decision. We evaluate Emo-Jev on eight datasets spanning sentiment analysis, emotion recognition, sarcasm detection and humor detection, comparing against direct Jev classification and five SoTA LLMs under input/output and chain-of-thought reasoning. Standard Jev achieves 62.93% average macro-F1 versus 67.28% for the strongest LLM baseline, with lower observed latency and generally lower cost.

[NLP-100] When Forgetting Looks Like Improvement: Metric Masking in Streaming Diarizer Adaptation and the Price of Rehearsal

【速读】: 该论文旨在解决小样本适应(small-data adaptation)在语音分离(speech detection)与说话人归属(speaker attribution)之间存在的性能不一致性问题。具体而言,尽管适应过程显著提升了目标域内的聚类性能并具备一定的跨域迁移能力,但在不同评估场景中表现不稳定,其主要瓶颈在于时间上说话人身份的一致性(temporal identity consistency)受损,而非说话人数量识别错误。研究提出通过局部重映射(local-remapping)诊断方法揭示了不同语料库中身份退化模式的差异,表明小样本适应可能改变了流式模型在时间维度上维持说话人标签连续性的机制。此外,回放(rehearsal)策略虽可缓解身份退化,但会削弱系统的跨域迁移能力。因此,该研究的关键解决方案在于强调在评估流式说话人分离系统时,必须联合考量检测准确率、身份一致性及长期记忆保持行为(retention behavior),以实现更全面的性能评估与优化。

链接: https://arxiv.org/abs/2610.08828
作者: Mo Yu,Yang Liu,Jing Qian
机构: Tongji University (同济大学)
类目: Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
备注: 5 pages, 4 figures

点击查看摘要

Abstract:Small-data adaptation can improve speech detection while degrading speaker attribution. We study this discrepancy in a released streaming diarizer adapted on 7.5 h of two-party conversation and evaluated across six corpora. Adaptation substantially improves in-domain diarization performance and transfers to an independent corpus. However, this improvement is not consistent across evaluation scenarios as the additional confusion is mainly associated with impaired temporal identity consistency rather than speaker-count errors. A local-remapping diagnostic reveals different patterns of identity degradation across corpora, indicating that adaptation may alter how streaming models maintain speaker assignments over time. Rehearsal reduces the observed degradation but reduces the cross-domain transfer performance. These results highlight the need to jointly evaluate detection accuracy, identity consistency, and retention behavior when adapting streaming diarization systems.

[NLP-101] Child ASR Adaptation with Adult Retention: An Empirical Study

【速读】: 该论文旨在解决自动语音识别(ASR)系统在儿童及非母语使用者语音识别中表现不佳的问题,同时克服将成人语音模型适配至儿童语音时导致的成人语音识别性能退化(adult-speech forgetting)难题。其核心挑战在于实现对儿童语音的有效适应(child adaptation)与对成人语音性能的稳定保留(adult retention)之间的平衡。解决方案的关键在于对比多种模型适配策略——包括全量微调、低秩自适应(LoRA)以及后处理权重空间融合(post-hoc weight-space merging)——并评估其在编码器-解码器、编码器-连接时序分类(CTC)、以及基于AudioLLM的ASR架构中的表现。实验结果表明,权重空间融合方法(如LERP和TIES)能显著改善适应性与保留性的权衡,尤其在编码器-CTC、Whisper及AudioLLM架构中表现突出;其中LERP更利于保持成人语音性能,而TIES则能实现更强的儿童语音适应增益。尽管如此,在原始词错误率(WER)上,直接进行双语微调仍是编码器-解码器模型中最优方案。

链接: https://arxiv.org/abs/2610.08827
作者: Houssam Eddine-Othman Lachemat,Shammur Absar Chowdhury
机构: Qatar Computing Research Institute (卡塔尔计算研究学院); HBKU (哈马德·本·哈利法大学)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Sound (cs.SD); Audio and Speech Processing (eess.AS)
备注: long paper

点击查看摘要

Abstract:Automatic Speech Recognition (ASR) systems often underperform for children and non-native speakers, while adapting adult ASR models to child speech can cause adult-speech forgetting. We study child ASR adaptation with adult retention across Arabic and English. We compare full fine-tuning, LoRA, and post-hoc weight-space merging across encoder–decoder, encoder–CTC, and AudioLLM-based ASR systems. Experiments use Arabic native and non-native child speech, English MyST child speech, and adult benchmarks from MGB-2 and LibriSpeech test-clean. We evaluate recognition quality with WER and quantify the adaptation–retention trade-off using Retention Index, Child Adaptation Gain, and Adaptation Recovery. Results show that child adaptation is necessary, especially for non-native Arabic and English child speech, but direct adaptation often reduces adult ASR performance. Bilingual adaptation is more stable than language-specific adaptation. Weight-space merging often improves the trade-off, especially for encoder–CTC, Whisper, and AudioLLM-based ASR, with LERP favoring adult retention and TIES recovering stronger child gains. For the encoder–decoder model, direct bilingual fine-tuning remains strongest in raw WER.\footnoteCode, and models are available at this https URL.

[NLP-102] Just for FUNS: LLM -Guided Spatio-Temporal Graph Node Generation for Forecasting Unobserved Node States

【速读】: 该论文旨在解决交通与城市系统中因传感器网络部署成本高、维护资源有限而导致的空间覆盖不全问题,即预测未观测节点状态(Forecast Unobserved Node States, FUNS)这一关键挑战。传统模型依赖历史观测数据,在面对无记录节点时表现严重退化。其解决方案的关键在于将FUNS问题重新定义为基于时空图的条件生成任务,并提出GenST框架:通过微调的大型语言模型(Large Language Models, LLMs)作为语义桥梁,提取节点描述中的丰富语义特征(如功能区类型、道路网络结构等),以补偿缺失的时空信号;该框架采用两阶段生成架构——首先利用时空变分自编码器(Spatio-Temporal VAE)将时空动态压缩至隐空间,再由生成式变压器(Generative Transformer, GenT)在多模态条件(包括语义、地理坐标和邻域上下文)引导下,从噪声中重建未观测节点的未来状态。实验结果表明,GenST在零样本预测任务中显著优于现有基线,验证了语义引导生成在缓解时空数据稀疏性方面的实际潜力。

链接: https://arxiv.org/abs/2610.08818
作者: Shuhao Li,Weidong Yang,Changan Liu,Wei Zhuo,Yingbo Zhou,Fan Zhang,Siqiang Luo
机构: Fudan University, Shanghai, China; Nanyang Technological University, Singapore; Guangzhou University, Guangzhou, China
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (stat.ML)
备注:

点击查看摘要

Abstract:Spatio-temporal forecasting is a cornerstone of logistics, urban planning, and intelligent transportation systems. However, constrained by deployment costs and maintenance resources, sensor networks often lack comprehensive spatial coverage, rendering Forecast Unobserved Node States (FUNS) a critical yet formidable challenge. Conventional models rely on historical observations and typically falter when encountering nodes without prior records. To address this, we redefine the problem as a conditional generation task on spatio-temporal graphs and propose GenST, a framework that introduces Large Language Models (LLMs) as a semantic bridge, leveraging a pre-trained LLM fine-tuned to extract rich semantic features from node descriptions, such as functional zones and road network structures, to compensate for missing spatio-temporal signals. Specifically, we design a two-stage generative architecture: a Spatio-Temporal VAE first compresses spatio-temporal dynamics into a latent space, followed by a Generative Transformer (GenT) that reconstructs the future states of unobserved nodes from noise, guided by multi-modal conditions including semantics, geographic coordinates, and neighborhood contexts. Experiments on six traffic and two non-traffic datasets show GenST significantly outperforms existing baselines in zero-shot prediction tasks, demonstrating the practical potential of semantic-guided generation for mitigating spatio-temporal data sparsity.

[NLP-103] Route-Verify-Vote: Procedure-Conditioned Self-Consistency for Mixed-Domain Reasoning CCL26

【速读】: 该论文旨在解决语言模型在面对训练中未出现的组合式推理场景时,难以有效泛化的问题,尤其是在跨领域(mixed-domain)的常识推理任务中,模型需将已知的推理操作以新颖方式组合。针对这一挑战,论文提出了一种无需参数更新的程序条件自一致性框架——Route-Verify-Vote(RVV),其核心在于通过领域标签(domain label)引导模型选择适配特定领域的推理程序(reasoning procedure),从而指导模型准确建模并应用相关约束。该方法分三步执行:首先,Route模块基于输入问题的领域信息选择对应的推理路径;其次,Verify模块利用所选程序对每个候选选项进行约束验证;最后,Vote模块聚合多轮采样得到的答案集,并通过投票机制识别置信度较低的题目,动态分配更多样本以提升决策可靠性。实验表明,在官方测试集上,标准RVV(每题16次采样)达到74.6%的精确集合准确率,自适应改进版本达77.3%,结合多模型在选定领域路径上的集成进一步提升至79.4%,最终系统位列所有参赛系统第二。研究结果证实了领域特异性推理程序与答案集分歧度作为推理阶段计算资源分配策略的有效性,为复杂混合域推理提供了可扩展、高效的解决方案。

链接: https://arxiv.org/abs/2610.08814
作者: Xinchen Xiao
机构: Xinjiang University(新疆大学)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
备注: 12 pages, 8 figures, and 8 tables. Accepted for oral presentation at CCL26-Eval

点击查看摘要

Abstract:Compositional generalization remains challenging when language models must combine familiar reasoning operations in unfamiliar ways. The Scenario-Based Commonsense Reasoning Evaluation (SCoRE) 2026 tests this ability on three mixed domains absent from training and requires models to identify the complete set of correct options for each question. We introduce Route-Verify-Vote (RVV), a framework for procedure-conditioned self-consistency that uses language models without parameter updates. Route uses the provided domain label to select a reasoning procedure that guides the model in representing and applying the relevant constraints. Verify prompts the model to assess each option against those constraints. Vote aggregates complete answer sets and allocates additional samples to questions with a small vote-count margin between the two most frequent sets. Samples for each question follow the same domain-specific procedure. On the official test set, voting over 16 sampled answer sets per question achieves an exact-set accuracy of 74.6%. Adaptive RVV reaches 77.3%, and combining models on selected domain routes raises accuracy to 79.4%. The final system ranked second among participating systems. These results support domain-specific reasoning procedures and answer-set disagreement as useful tools for allocating inference-time computation in mixed-domain reasoning. Comments: 12 pages, 8 figures, and 8 tables. Accepted for oral presentation at CCL26-Eval Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG) Cite as: arXiv:2610.08814 [cs.AI] (or arXiv:2610.08814v1 [cs.AI] for this version) https://doi.org/10.48550/arXiv.2610.08814 Focus to learn more arXiv-issued DOI via DataCite

[NLP-104] okka-Bench: Evaluating Tokenizers Across 100 Natural and 20 Programming Languages

【速读】: 该论文旨在解决大语言模型中子词分词器(subword tokenizer)质量在不同语言间差异显著,但缺乏标准化多指标评估框架的问题。其核心解决方案是提出Tokka-Bench——一个开源评估框架,通过五个互补的评价指标(每标记字节数、唯一标记覆盖率、子词肥力、单词切分率及词汇构成)对100种自然语言(含30余种文字系统)和20种编程语言中的七种BPE分词器(GPT-2、GPT-4、gpt-oss、Llama 3.1、Gemma 3、Qwen3与Kimi K2)进行系统性对比评估,采用适配各书写系统的语言感知分词策略。研究发现,词汇分配策略的影响大于单纯词汇规模,且近期分词器在编程语言效率上已趋于收敛,尽管其自然语言表现仍存在差异。该框架、数据集及交互式仪表板均公开可用。

链接: https://arxiv.org/abs/2610.08794
作者: Ben Gubler
机构: 未知
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: 5 pages, 5 figures. Code and data: this https URL . Interactive dashboard: this https URL

点击查看摘要

Abstract:Large language models rely on subword tokenizers whose quality varies across languages, yet no standardized multi-metric framework exists for broad comparative evaluation. We introduce Tokka-Bench, an open-source framework that evaluates tokenizers on five complementary metrics – bytes per token, unique token coverage, subword fertility, word-split rate, and vocabulary composition – across 100 natural languages (30+ scripts) and 20 programming languages, using language-aware segmentation adapted to each writing system. Comparing seven BPE tokenizers (GPT-2, GPT-4, gpt-oss, Llama 3.1, Gemma 3, Qwen3, and Kimi K2) within individual languages, we find that vocabulary allocation strategy matters more than raw vocabulary size, and that programming-language efficiency has converged among recent tokenizers despite divergent natural-language profiles. The framework, data, and interactive dashboard are publicly available.

[NLP-105] Certified by Abstention: Distribution-Free Guarantees for Chain-of-Thought Verifiers at Small Calibration Budgets AISTATS2027

【速读】: 该论文旨在解决生成式 AI(Generative AI)中链式思维(Chain-of-Thought, CoT)推理轨迹正确性验证的可靠性问题,特别是在有限标注样本(数十至数百个已标注问题)的现实校准预算下,如何实现无需分布假设的可信赖选择性保证。其核心挑战在于:现有验证器依赖的置信度信号在实际部署中可能因校准不足而产生高错误率,尤其当证书发放频率极低时,即使整体失败概率被控制在δ以内,一旦发放,其条件失败概率可能高达δ / P_fire,导致“几乎每次使用都出错”的极端风险。解决方案的关键在于提出一种基于“拒答有效性”(validity by abstention)的新范式,通过引入认证下界(certification floor)与贝尼米尼-霍克伯格(Benjamini-Hochberg)保形选择的格条件,解释为何在小样本校准下标准证书倾向于全不接受或大规模接受,而非精细选择。研究发现,不可读的残差流探测器(residual-stream probe)能实现比可读信号高出2–3倍的覆盖范围,且交叉拟合重构无法线性恢复此优势。进一步,作者提出一种起始于下界的固定序列认证器,无需单调性假设即可在所有模型-信号组合上优于邦弗朗尼(Bonferroni)方法,在非平凡目标覆盖率0.75π₀下将覆盖率从0.05提升至0.16,尽管绝对覆盖仍受限于下界。最后,论文揭示了认证器在部署后对任务漂移(benchmark shift)和最佳n选一(best-of-n)采样策略的脆弱性:当任务分布变化时,被接受轨迹的误差会追踪新任务的基础错误率;而在对抗性采样下,虽然经验失败频率仍低于δ,但真实错误率已超过目标,原因在于拒答机制吸收了失败,暴露了认证器“看不见部署后关键信息”的根本局限。

链接: https://arxiv.org/abs/2610.09541
作者: Arjun Balaji
机构: 未知
类目: Machine Learning (stat.ML); Computation and Language (cs.CL); Machine Learning (cs.LG)
备注: 22 pages, 7 figures, 12 tables. Under submission at AISTATS 2027

点击查看摘要

Abstract:Signals that predict whether a chain-of-thought (CoT) trace is correct are compared by AUC, but deploying one requires a threshold with a guarantee. We ask what distribution-free selective guarantees deliver for CoT verifiers at realistic calibration budgets of tens to a few hundred labelled problems, using seven open models, five verifier signals and 37,000 graded traces. The central observation is validity by abstention: an (\alpha,\delta) -valid procedure that issues a certificate with probability P_\rm fire bounds the failure probability of an issued certificate only by \delta/P_\rm fire , so a certificate that rarely fires can be valid and wrong every time it is used. In a simulation with known risk the standard certificate fails in at most 0.3% of calibration draws but in up to 69% of those in which it fires. A certification floor and a lattice condition for Benjamini-Hochberg conformal selection explain why certificates abstain at these budgets, and the data bear them out: the standard certificate returns nothing or a large accepted set, and an unreadable residual-stream probe buys two to three times the coverage of the readable signals, an edge a cross-fitted reconstruction cannot recover linearly from the readable features. We then give a floor-started fixed-sequence certificate, valid without monotonicity assumptions, that covers more than the Bonferroni certificate on every model-signal pair and raises coverage at the non-vacuous target 0.75\pi_0 from 0.05 to 0.16, although the floor keeps absolute coverage small. Finally, a certificate cannot see what matters after deployment: under benchmark shift the error among accepted traces tracks the new task’s base error, and under best-of- n selection against the verifier it rises past the target while the empirical failure frequency stays below \delta , because abstention absorbs the failures.

[NLP-106] Mitigating Accent-Language Confusion in Self-Supervised Speech Representations for Language Identification ICASSP2027

【速读】: 该论文旨在解决自监督语音表征微调后的语音语言识别(LID)模型在实际应用中易将非母语(L2)口音误判为说话人第一语言(L1)的问题。其核心问题是:非母语语音的表征位于目标语言与母语的表征之间,导致系统性误分类。解决方案的关键在于提出一种几何投影方法,仅利用纯母语语音数据估计出“母语偏置”(L1-bias)方向,并在冻结的LID头部前将其从语音表征中移除。该方法无需任何非母语训练数据或模型微调,在五个MMS-LID模型及多个非母语语料上均显著提升了对带有L2口音语音的目标语言识别性能,同时保持对母语语音预测的准确性,证明了可在表征空间内直接纠正由口音引起的母语偏置。

链接: https://arxiv.org/abs/2610.09486
作者: Minu Kim,Jihwan Lee,David R. Mortensen,Shrikanth Narayanan
机构: 未知
类目: Audio and Speech Processing (eess.AS); Computation and Language (cs.CL); Sound (cs.SD)
备注: Submitted to ICASSP 2027

点击查看摘要

Abstract:Spoken language identification (LID) aims to recognize the target language regardless of accent. In practice, however, LID models fine-tuned from self-supervised speech representations frequently confuse accents with languages, misclassifying non-native (L2) speech as the speaker’s first language (L1). We show that non-native speech representations lie between native target-language and native L1 poles, causing systematic misclassification. To address this, we introduce a geometric projection that estimates an L1-bias direction solely from native speech and removes it before the frozen LID head. Across five MMS-LID models and non-native corpora, this projection substantially improves target language identification for L2-accented speech while preserving predictions for native speech. These results show that accent-induced L1 bias can be corrected directly within the representation space without L2 training data or model adaptation.

[NLP-107] Boundary-Free Contextual Biasing: Depth-Adaptive Gating and Reading-Space Matching for Unsegmented Languages

【速读】: 该论文旨在解决在语音识别(ASR)系统中,针对日语和中文等无词边界语言进行上下文偏置(contextual biasing)时所面临的挑战。现有方法依赖于词边界信息,但在这些语言中并不存在显式的词边界,导致传统偏置技术难以直接应用。为此,本文提出一种无需训练、无需二次解码的无边界偏置解码器,基于字符级Aho-Corasick自动机构建,能够在冻结的公共连接时序分类(CTC)模型上实现高效推理。其关键创新在于引入两种基于证据的机制以替代词边界:一是深度自适应门控机制,用于动态调节从匹配深度中强制推进的程度;二是读取空间匹配机制,用于处理音频与字符不一致但音素序列正确的情况。实验结果表明,在Aishell-1 NE的困难子集(R1)上,该方法实现了66.5%的召回率,优于经过训练的CLAS基线(64%),且可泛化至WenetSpeech及另一架构而无需重新调参。此外,作者发布了首个公开的日语上下文偏置基准测试集,结果显示在精度高于97%的前提下,偏置使罕见词召回率提升25个百分点,即使在1,000个词的列表下仍分别提升19和22个百分点,显著提升了低频词识别性能。

链接: https://arxiv.org/abs/2610.09467
作者: Muhammad Huzaifah,Yu Pan,Zachary Yeo,Ningjie Bai,Guangzhao Yang
机构: 未知
类目: Audio and Speech Processing (eess.AS); Computation and Language (cs.CL); Sound (cs.SD)
备注:

点击查看摘要

Abstract:Contextual biasing supplies an ASR system with a list of expected words at inference time, but existing methods rely on word boundaries that Japanese and Chinese do not provide. We present a boundary-free biasing decoder for frozen public CTC models, built on a character-level Aho-Corasick automaton, with no training and no second pass. Two evidence-based mechanisms replace the boundary: a depth-adaptive gate that sets how hard to push from match depth, and reading-space matching for when the audio is right but the characters are wrong. On Aishell-1 NE’s hard R1 subset we reach 66.5% recall, above the trained CLAS baseline (64%), transferring to WenetSpeech and to a second architecture without retuning. We release the first open Japanese contextual-biasing benchmark, where biasing lifts rare-word recall by 25 points at precision above 97%, and still by 19 and 22 points against 1,000-word lists.

[NLP-108] Phoneme-Guided Initialization for LLM -based Speech Recognition

【速读】: 该论文旨在解决在低资源条件下语音大语言模型(speech LLMs)因缺乏足够成对语音-文本数据而导致自动语音识别(ASR)性能下降的问题。其核心解决方案是提出一种**音素引导初始化(phoneme-guided initialization)**方法:首先在语音到音素(S2P)任务上预训练音频编码器,在音素到字素(P2G)任务上预训练语言模型,随后将两者连接并在目标ASR任务上进行端到端微调。该方法利用了音素作为中间表示的语义桥梁作用,有效缓解了低资源场景下直接端到端训练的性能瓶颈,实验结果表明其在日语(CSJ)、中文(AISHELL-1)及两种低资源语言(鞑靼语和乌尔都语)上的表现优于传统的级联式S2P-P2G范式与无音素引导初始化的端到端模型。

链接: https://arxiv.org/abs/2610.08994
作者: Ryo Magoshi,Shinsuke Sakai,Tatsuya Kawahara
机构: Kyoto University (京都大学)
类目: Audio and Speech Processing (eess.AS); Computation and Language (cs.CL); Sound (cs.SD)
备注: Accepted at IEEE SLT 2026

点击查看摘要

Abstract:Speech large language models (speech LLMs) perform well on automatic speech recognition (ASR) when sufficient paired speech-text data is available, but their performance degrades in low-resource settings. A cascaded pipeline that performs speech-to-phoneme (S2P) conversion followed by phoneme-to-grapheme (P2G) conversion has been shown to outperform end-to-end speech LLMs in this regime, suggesting that phoneme-mediated processing is beneficial when paired data is scarce. We propose \textitphoneme-guided initialization, a simple method that uses this insight within an end-to-end framework: we pre-train the audio encoder on S2P and the LLM on P2G tasks, then connect them and fine-tune the full model end-to-end on the target ASR task. Experiments on Japanese (CSJ), Chinese (AISHELL-1), and two low-resource languages from Common Voice 25.0 (Tatar and Urdu) show that our method matches or outperforms both the cascaded S2P-P2G baseline and the end-to-end model without P2G initialization.

[NLP-109] Is Word Error Rate Enough? Rethinking Privacy Evaluation in Speech with Entity-Aware Metrics

【速读】: 该论文旨在解决智能设备日益普及背景下,语音内容(尤其是敏感命名实体)在采集过程中可能引发的隐私泄露问题。核心挑战在于如何在有效保护语音隐私的同时,维持音频的可用性,并设计能够准确衡量隐私保护水平、避免高估保护效果的评估指标。其解决方案的关键在于将自然语言处理领域中的实体感知隐私度量方法迁移至语音隐私领域,以更精细地评估对命名实体等敏感信息的保护效果;同时,通过分析多种攻击场景发现,基于富含命名实体的数据进行微调虽可提升部分实体类别的攻击成功率,但对其他类别无显著影响,进而提出根据去混淆方法是否保持时间对齐性来选择合适评估指标的指导原则。

链接: https://arxiv.org/abs/2610.08831
作者: Anjana Rajasekhar,Jule Pohlhausen,Nayana Jacob Alappattu,Anna Leschanowsky
机构: 未知
类目: Audio and Speech Processing (eess.AS); Computation and Language (cs.CL); Multimedia (cs.MM); Sound (cs.SD)
备注:

点击查看摘要

Abstract:As the use of smart devices continues to increase, their potential to capture sensitive speech content raises growing privacy concerns. It is therefore critical to develop techniques that prevent information leakage while preserving the utility of the audio, and evaluation metrics that accurately quantify the level of privacy without overestimating it. In this work, we evaluate the effectiveness of two obfuscation techniques in protecting speech content, with particular emphasis on named entities, by adapting entity-aware privacy metrics from the Natural Language Processing field to the speech privacy domain. Further, we investigate several attack scenarios and show that fine-tuning on entity-rich data improves attack performance for some entity categories but not others. Finally, we provide guidance on metric selection based on whether the obfuscation method preserves temporal alignment.

信息检索

[IR-0] wo-Level Softmax Sampling Done Right: Correcting Bias from Size Imbalance and Dispersion NEURIPS2026

链接: https://arxiv.org/abs/2610.10483
作者: Walid Bendada,Guillaume Salha-Galvan
类目: Machine Learning (cs.LG); Information Retrieval (cs.IR); Machine Learning (stat.ML)
备注: NeurIPS 2026

点击查看摘要

Abstract:Sampling from a softmax distribution is a fundamental operation in machine learning, but its linear complexity in the number of items makes exact sampling impractical at scale. Two-level softmax (2LS) sampling is a popular alternative enabling sublinear-time sampling. Assuming items are partitioned into clusters, 2LS first samples a cluster and then an item within it. In this paper, we show that, despite its advantages, 2LS introduces systematic and undesirable sampling biases, which arise from misweighting clusters by ignoring both cluster size imbalance and intra-cluster similarity dispersion. We propose two sampling methods, Size-Corrected 2LS (S-2LS) and Size- and Dispersion-Corrected 2LS (SD-2LS), which correct these biases and provide provably better softmax approximations with negligible to non-existent computational overhead. In-depth experiments on five large-scale datasets validate the improved sampling properties of our methods. We recommend their consistent use in place of standard 2LS in future work.

[IR-1] CrossWeave: Bridging Perspectives Across Online Communities with a Dual-Pane Design

链接: https://arxiv.org/abs/2610.10441
作者: Fei Fang,Reva Hirave,William Jurayj,Yuqi Li,Brian Lu,Tarik Metin,Tsugunobu Miyake,Kateryna Morhun,Yash Permalla,Kenan Rustamov,Allen Shen,Haojun Shi,Prabhav Singh,Xiheng Tom Wang,Kevin Xu,Qingcheng Zeng,Jiayi Zhang,Daniel Khashabi,Andrew Perrin,Tiziano Piccardi,Ziang Xiao,Jason Eisner
类目: ocial and Information Networks (cs.SI); Human-Computer Interaction (cs.HC); Information Retrieval (cs.IR)
备注: CSCW 2026 + small improvements

点击查看摘要

Abstract:Social media systems typically display conversations among already familiar contributors, which can be predictable and one-sided. In civic discourse, this design narrows discussion, reinforces divides, and distorts the perception of public opinion. To encourage cross-community engagement, we present CrossWeave, an AI-powered bridging system that augments the standard social media feed. As the user reads a post, CrossWeave surfaces diverse relevant posts from other threads in a side pane and highlights the connections. Users are invited to venture out of their echo chamber, explore a broader range of views and arguments, and ``click across’’ to engage with their authors. When they do, CrossWeave facilitates constructive posting, not only by showcasing relevant past content but also by simulating possible reactions as the user drafts a post.

[IR-2] Does Document Structure Help Dense Retrieval? A Placebo-Controlled Ablation of Four Mechanisms Across Two Corpora

链接: https://arxiv.org/abs/2610.10170
作者: Andrey Kuehlkamp,Priscila Correa Saboia Moreira,Samuel Rund
类目: Information Retrieval (cs.IR); Artificial Intelligence (cs.AI)
备注: Initial draft,

点击查看摘要

Abstract:Retrieval-augmented generation systems increasingly rely on document-structure treatments: structure-aligned chunking, LLM-generated chunk contexts, heading-path metadata, and hierarchical two-stage retrieval. Separate studies support each on different corpora, embedders, and metrics, and none control for a shared confound: any text prepended to a chunk perturbs its embedding. We present a mechanism-isolating ablation testing all four treatments under one protocol, matching chunk sizes across conditions and adding a semantically null placebo—heading paths that are structurally valid but shuffled across documents. We score retrieval with a coverage-aware nDCG and test four pre-registered contrasts via document-clustered bootstrap with Holm correction, on two distant corpora: 200 Wikipedia Featured Articles (951 queries) and 1,585 QASPER papers (4,303 questions). Organization helps, and the cause is content, not tokens: structure-aligned chunks with real heading paths beat contextualized fixed windows (+0.022 / +0.012 cov-nDCG@10) and the placebo (+0.010 / +0.016). Naive two-stage hierarchical retrieval hurts (-0.033 / -0.015), traceable to first-stage section recall. Gold structure beats LLM-induced structure on Wikipedia but not on QASPER. Effects are small ( dz 0.06-0.11) but Holm-significant and consistent across corpora.

[IR-3] raining with Missed Targets in Generative Recommendation: Separating Supervision from Probability Competition

链接: https://arxiv.org/abs/2610.10124
作者: Xuesi Wang,Yangbin Shi,Xiaolin Zheng
类目: Information Retrieval (cs.IR); Machine Learning (cs.LG)
备注: 12 pages, 4 figures, 8 tables

点击查看摘要

Abstract:Generative recommenders return a limited candidate set and may omit observed targets before reranking. A training strategy appends these missed targets to reranker training lists, although inference still ranks only original candidates. This operation simultaneously changes retrieved-target weight, adds supervision over appended targets, and makes the two groups compete for probability. An append/no-append comparison therefore cannot explain changes in returned-item rankings. We construct three matched losses that hold retrieved-target weight fixed while introducing appended-target supervision and group competition separately. The intermediate loss trains within both groups but normalizes them separately, preventing training-only targets from competing with inference candidates. Experiments with a released OneRec model and locally trained Amazon generators show that this competition can harm returned-item ranking. In four prespecified Amazon Video Games comparisons, removing it improved full-target normalized discounted cumulative gain (FT-NDCG) by 7.8–22.2%; 95% intervals over users and three of four intervals over training runs excluded zero. A conservative development-set rule selected appended-target training for two of three generators in one held-out category and rejected it for all three in another, avoiding a 1.7% loss. Candidate completion should therefore be evaluated for each generator rather than applied automatically.

[IR-4] ExperienceIndex: Artifact-Grounded Memory

链接: https://arxiv.org/abs/2610.10091
作者: Peter Baile Chen,Geoffrey X. Yu,Xinming Liu,Samuel Madden,Dan Roth,Jacob Andreas,Doug Downey,Michael Cafarella
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Information Retrieval (cs.IR); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Knowledge-intensive tasks require answering many questions by reasoning about a shared corpus of artifacts (e.g., court cases, or scientific literature). As humans interact with these corpora, they naturally accumulate experiential knowledge about artifacts, enabling them to quickly identify the complete set of relevant artifacts for each new task. However, existing AI agents lack appropriate memory solutions to build or reuse such artifact-grounded experience, leading to lower answer quality and higher online cost. Existing memory solutions extract and reuse information from prior task-solving traces, but they primarily focus on user preferences, factual attributes, or abstract reasoning patterns rather than persistent artifact-specific knowledge. We introduce ExperienceIndex, a novel experience layer for AI agents that captures and reuses knowledge about artifacts based on prior reasoning traces. ExperienceIndex stores two complementary forms of experience: (i) single-artifact experiences that summarize an artifact’s contribution to prior tasks and (ii) artifact-pair experiences that encode structural relationships discovered during past reasoning. Integrated as lightweight middleware, ExperienceIndex uses an experience retrieval mechanism to guide agents toward the complete set of relevant artifacts for new tasks, improving both answer quality and efficiency. Across diverse corpora and agentic solutions with different search frameworks, ExperienceIndex delivers consistent gains, raising answer quality by up to 11.0 points and reducing online dollar cost by up to 50.5%. We further demonstrate two benefits: (i) cross-task generalization, where experiences accumulated from text-to-SQL tasks transfer to factoid QA tasks over the same artifact corpus, and (ii) teacher-student learning, where experiences from a stronger model enable a weaker model to reach comparable performance.

[IR-5] Inverting Multi-Vector Visual Document Indices

链接: https://arxiv.org/abs/2610.09920
作者: Zhuchenyang Liu,Yao Zhang,Yu Xiao
类目: Information Retrieval (cs.IR); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)
备注: 30 pages. Under review

点击查看摘要

Abstract:Prevailing multi-vector visual document retrievers store each page as about a thousand patch vectors, often in vector databases run by a third party. Since no one can read a page from its vectors, this index is easily treated as less sensitive than the page. However, because the index keeps one vector per patch in raster order, and each vector is computed by a vision-language model pre-trained to read documents, we hypothesize that whoever runs or breaches the store can reproduce a page from its index alone. We frame inversion as conditional document image generation and infer from the vectors what the attack needs: the encoder, the page shape and, for shuffled vectors, their order. On the ViDoRe v3 benchmark, pages inverted from raw indices recover 47% of the words and 45% of the sensitive tokens. Used as queries against the stored indices, they rank their source page first 98.4% of the time. We test two cheap protections, token pooling and shuffling, which both cut word recall to about 8%. A model that restores the order of a shuffled index raises the share of source pages ranked first from 3.8% to 93.5%, while inverting a pooled index remains open. To test generalisation, we apply the same attack unchanged to another multi-vector retriever: its inverted pages still rank their source page first 70.2% of the time, though its word recall stays below a nearest-neighbour baseline. Multi-vector visual document retrievers are therefore vulnerable to inversion through their stored index, which should be protected like the documents it encodes.

[IR-6] he Impact of Backbone Evolution on LLM -Based Relevance Assessments

链接: https://arxiv.org/abs/2610.09820
作者: Chuting Yu,Guido Zuccon,Teerapong Leelanupab
类目: Information Retrieval (cs.IR)
备注: 12 pages main content

点击查看摘要

Abstract:LLMs are evolving rapidly, with newer models offering stronger capabilities. This suggests that in LLM-based relevance judging, more capable models will achieve higher agreement with human judgements under the same prompt. We challenge this understanding by investigating the behavior of LLM-based relevance judges under backbone evolution. Keeping the prompts fixed, we evaluate a representative single-prompt (UMBRELA) and a rubric-based prompt (EXAM) across sequential model versions of commercial (Gemini, GPT) and open-weight (Qwen, Llama) models. Overall, we find no consistent evidence that newer versions lead to better relevance judges. Crucially, similar or improved aggregate performance does not imply judgment stability: correct judgements made by an earlier version of an LLM backbone are not necessarily preserved by later versions. We investigate the potential drivers of these regressions. Our findings caution against the assumption that judging prompts designed and validated for one backbone version will perform equivalently or better when the model is updated, even within the same family.

[IR-7] SoccerNet-FoulRet: Retrieving Semantically Similar Soccer Foul Videos ACCV2026

链接: https://arxiv.org/abs/2610.09742
作者: Jacobus Arthur,Ahmad Sait,Batool Hani,Merey Ramazanova,Jan Held,Marc Van Droogenbroeck,Bernard Ghanem,Anthony Cioppa,Silvio Giancola
类目: Computer Vision and Pattern Recognition (cs.CV); Information Retrieval (cs.IR)
备注: ACCV 2026

点击查看摘要

Abstract:Refereeing decisions in professional soccer remain inconsistent because referees cannot easily compare a contentious foul against similar past cases. We cast this as a retrieval problem and introduce SoccerNet-FoulRet, the first benchmark for semantic foul retrieval. Given a query foul, the task is to retrieve past fouls judged to be relevant precedents, regardless of camera angle, teams, or appearance. This differs from prior video-to-video retrieval, which matches clips by visual similarity or a shared event. Here, relevance is defined by refereeing interpretation. We build the benchmark from the SoccerNet-MVFoul dataset and evaluate retrieval ability of zero-shot video and vision-language embedders together with a task-specific fine-tuned baseline on 693 human-verified queries and category-relevance labels. Semantic foul retrieval remains challenging. The strongest zero-shot model achieves under 5% HitRate@10 on human-verified precedents, while category-supervised fine-tuning improves category relevance but transfers only modestly to precedent retrieval. We release SoccerNet-FoulRet to establish semantic foul retrieval as an open problem: this https URL.

[IR-8] owards Explaining Query Expansion Performance in Information Retrieval

链接: https://arxiv.org/abs/2610.09724
作者: Sourav Saha,Aditya Dutta,Soumajit Pramanik,Mandar Mitra
类目: Information Retrieval (cs.IR); Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Query Expansion (QE) techniques have long been widely used in Information Retrieval (IR) to address the vocabulary mismatch problem. They remain relevant in modern retrieval systems, including those based on large language models (LLMs). However, no single QE method consistently outperforms others across all queries. This work seeks to explain the variation in QE performance through two complementary perspectives. The first is the concept of an Ideal Expanded Query (IEQ)–a hypothetical query that maximizes retrieval effectiveness with a downstream BM25 retrieval model. The second is a separability perspective, which quantifies how distinctly relevant and non-relevant documents are scored for a given expanded query using Cohen’s (d). We develop a separability measure and practical formulations to approximate the IEQ and investigate how these factors relate to retrieval effectiveness. Extensive experiments on the TREC Robust collection, TREC DL 2019-2022 passage collections, and TREC DL 2019-2020 document collections reveal several interesting patterns. In particular, we find that expanded queries that are closer to the ideal expanded query tend to achieve higher retrieval effectiveness. We further show that the separability of relevant and non-relevant documents provides a complementary perspective for understanding QE performance.

[IR-9] Finding the Right Balance: Relevance and Diversity in LLM Retrieval

链接: https://arxiv.org/abs/2610.09412
作者: Guillaume Brouillette(1),Faustin Kagabo(1),Usef Faghihi(1),Nadia Ghazzali(1) ((1) Université du Québec à Trois-Rivières, Trois-Rivières, Canada)
类目: Information Retrieval (cs.IR); Computation and Language (cs.CL)
备注: 36 pages, 8 figures, 13 tables. Code and results: this https URL

点击查看摘要

Abstract:Retrieval diversification is widely available in retrieval-augmented generation (RAG) frameworks, yet prior studies disagree on whether it improves retrieval and answer quality. We show that its effectiveness varies primarily with candidate-pool redundancy, in a pattern consistent with the number of distinct evidence pieces a query requires. Using controlled near-duplicate injection and production-style overlapping chunking, we find that diversification harms relevance, evidence coverage and answer quality on clean pools, but becomes beneficial on multi-evidence tasks when redundancy causes nearest-neighbor retrieval to select repeated passages. We therefore introduce a query-adaptive rule that diversifies only when the effective number of distinct documents in the nearest-neighbor top- k selection falls below the query’s evidence requirement. Computed from existing embeddings, the rule captures most of the achievable gain, transfers across datasets and encoders and automatically reduces to nearest-neighbor retrieval for single-evidence queries. We also introduce RNG-Score, a geometric reranker with an exact nearest-neighbor fallback whose margin indicates duplicate structure. Overall, we conclude that diversification should be used selectively, based on observable redundancy and evidence requirements.

[IR-10] From Chunks to Functional Evidence: Function-Aware Retrieval for EDA Documentation QA

链接: https://arxiv.org/abs/2610.09361
作者: Xiaotian Qiu,Kairui Liu,Shi Chenyi,Jinyuan Deng,Qi Sun,Cheng Zhuo
类目: Information Retrieval (cs.IR); Artificial Intelligence (cs.AI)
备注: 10 pages, 2 figures, including appendices

点击查看摘要

Abstract:Retrieval-Augmented Generation (RAG) is widely used to ground answers in documents. For complex technical documentation, however, the primary bottleneck is often not model reasoning but a mismatch between a query and the way knowledge is organized for retrieval. This mismatch is pronounced in Electronic Design Automation (EDA) documentation, where the information needed for an answer is scattered across heterogeneous yet tightly coupled artifacts. We therefore redesign the basic retrieval unit of RAG. Instead of operating on isolated chunks or binary relations, we collect typed artifacts into EDA functional units. Each unit is recorded as a hyperedge with links to its source chunks. We then train an encoder to align queries with functional units and combine unit retrieval with direct chunk retrieval. After mapping the selected units back to their sources, a unified reranker chooses the evidence given to the generator. On the newly constructed EDADocEval-QA dataset, our method improves ROUGE-L by 37.1% over Chunk RAG and 55.6% over the strongest graph baseline. On the public ORD-MMBench benchmark, it improves ROUGE-L by 30.0% over the strongest baseline. These results support function-aware evidence organization in the evaluated EDA documentation settings.

[IR-11] Reading Position Is the Baseline to Beat: A Time-Ordered Evaluation of Personalised Highlight Prediction

链接: https://arxiv.org/abs/2610.09262
作者: Kazuki Nakayashiki,Keisuke Watanabe
类目: Information Retrieval (cs.IR); Human-Computer Interaction (cs.HC)
备注: 13 pages, 1 figure, 5 tables. Ancillary files include the specifications, the results write-ups, the analysis scripts, and the aggregate artifacts every reported number is generated from

点击查看摘要

Abstract:A reader’s first highlights on a page are the cheapest personal signal a reading product has. The natural plan is to suggest what similar earlier readers marked, and to judge the result against popularity. We argue that the baseline to beat is reading position. In a time-ordered evaluation on one social highlighting platform (7,343 reader-page pairs on 1,511 pages after one highlight), ranking the sentences just below a reader’s first highlight, with no other reader’s data, puts the next highlight in the top five 47% of the time, against 26% for popularity and 29% for the better of two similarity methods. The baseline depends on the target: over all later highlights that ranking loses to popularity, while popularity discounted by distance from the latest highlight, at the scale with the best average precision of three tried, beats popularity and both similarity methods on both targets. In a comparison specified in advance, neither similarity method shows a gain over popularity in average precision over all later highlights, from one to five highlights, and a gain of +0.01 is excluded. Nor would a gain by itself show that a method has found a reader’s preferences: synthetic readers who share one set of preferences produce one, and an evaluation out of time order shows a method where the reader went. The position results are exploratory and unconfirmed. Personalisation inside a document should be evaluated in time order and against reading position.

[IR-12] Quantize by Drift: Label-Free Mixed-Precision Post-Training Quantization for Text Embedders

链接: https://arxiv.org/abs/2610.09227
作者: Hyojung Han,Jongmin Kim,Seung-Hun Jeon
类目: Information Retrieval (cs.IR); Computation and Language (cs.CL); Machine Learning (cs.LG)
备注: 26 pages, 22 tables, 4 figures

点击查看摘要

Abstract:Mixed-precision post-training quantization needs a per-module sensitivity signal; for a text embedder the obvious one – the retrieval quality a module costs when quantized – needs relevance labels that deployments rarely have. We measure a label-free substitute: quantization-induced representation drift, obtained by quantizing one module, re-encoding the corpus, and recording how far the output embeddings moved from their full-precision positions. What is specific is the observable: the deployed output representation a dense retriever ranks with. Across five development embedders, configuration-level drift orders sampled mixed-precision plans against held-out retrieval quality at a macro Spearman of 0.911, the sensitivity transports across calibration corpora and retrieval domains in the usable regime, module drifts compose rank-consistently but not numerically, and relevance-derived sensitivity adds no consistent value. The method is one additive allocation under a hard packed-byte budget, with no labels and no search. On three embedders held untouched until method, baselines and hypotheses were frozen and sealed, the pre-registered directional hypothesis against the prior LieQ criterion holds (3/3 at the main budget, no collapse) and drift scores above a two-sided LieQ steelman in 2/3; but at the main budget drift is numerically lower than same-budget uniform precision on all three (-0.99, -0.85, -1.01 points), having reduced module and whole-model drift as designed. Output drift is thus a robust coarse sensitivity signal, not a universally optimal allocation objective: it avoids the catastrophic failures of the transferred signed-geometry adaptation and can remain usable at stressed budgets where uniform collapses, but fine-grained redistribution around a strong uniform operating point remains unresolved.

[IR-13] What Transfers from a VLM Teacher? Comparing Supervision Signals for Visual Document Retrieval

链接: https://arxiv.org/abs/2610.09177
作者: Saba Sturua,Han Xiao
类目: Information Retrieval (cs.IR)
备注: 27 pages, 1 figure, 17 tables

点击查看摘要

Abstract:Visual document retrievers are trained contrastively: each query is matched to one page labelled relevant - the positive - and pushed away from negatives, pages presumed irrelevant. Recent methods distil a vision-language model (VLM) teacher into the retriever by enriching that positive, transferring the teacher’s attention over it or a description of it. We ask whether the teacher is better spent on the other side, judging the candidates the retriever mines as negatives, which the label says nothing about. With student, data, optimizer and evaluation fixed, teacher-judged hard negatives and score distillation raise ViDoRe v2 nDCG@5 from 55.2 to 62.6 and 63.0; description alignment, as adapted here, gains 2.6 points and attention grounding nothing measurable. Against teacher-free rules that select four candidates from the same mined pool at identical training compute, the best of which is the positive-aware threshold current systems use, the teacher’s judgement adds 4.1 points on v2 and 1.7 on v3. This is consistent with how incomplete the labels are. Annotators judge about two of a query’s four top-ranked mined candidates relevant, none of them labelled, so training pushes the retriever away from relevant pages treated as negatives. What reaches the student is coarse: under a greedily decoded 0-100 rating prompt, 82% of the teacher’s ratings come back at one end of the scale or the other, and a relevant/irrelevant partition keeps most of the distillation gain. A ten-annotator audit places the teacher within the range of variation among human annotators, and finds it reliable where a query has a single determinate answer. We release the code, the teacher’s 3.3M judgements and page descriptions, the mined pools, the human audit and the trained adapters at this https URL.

[IR-14] Building Navigable Graphs Without Search in Three Composable Stages ATC

链接: https://arxiv.org/abs/2610.09041
作者: Édgar Chávez
类目: Databases (cs.DB); Information Retrieval (cs.IR)
备注: 26 pages. Code: this http URL (branch fgraft); experiments, logs and patches: this http URL

点击查看摘要

Abstract:Navigable graphs can be built without searching for neighbors: partition the data, evaluate every pair inside each part, and select each point’s edges from the candidates. We give such a construction in three separable stages and show that the middle one decides the quality. The pool is any partition with a few memberships per point. The ending turns a point’s candidates into out-edges; ours keeps a bounded heap, prunes by occlusion with a per-corpus slack, and appends reverse edges, re-pruning only where a list overflows. The spine is any edge set, exempt from the prune, that keeps the graph reachable from its entry; ours, half-space-proximal edges over a random sample, routes monotonically to every sampled point and replaces a spanning tree at 1/10 to 1/500 of its cost. The ending composes with any partitioner: on PiPNN’s own candidate pool it beats PiPNN’s ending on each of six corpora from 10^6 to 10^8 points, by 3 to 14% in distance evaluations at equal recall, and with 60 to 120 memberships per point the composed build matches or beats a full dense construction at k=10 and k=100 on all six, in 0.5 to 0.9 of its build time, deterministically. The analysis explains why. Once a pool is localised its quality is set by the data: every pool built on GIST lands within 4% of the exact-kNN ceiling, and the pairs a block cover misses are predicted, point by point, by the local clustering of the kNN graph, whose zero-clustering tail sets the memberships a corpus needs and grows with n. All code, patches and logs are public.

[IR-15] From High Recall to High Utility: Dataset-Adaptive Post-Processing of LLM -Generated Customer Intents

链接: https://arxiv.org/abs/2610.09039
作者: Mahesh Viswanathan,Joan Rossello,Leticia Fernandes,Paul Mutawe
类目: Artificial Intelligence (cs.AI); Information Retrieval (cs.IR)
备注:

点击查看摘要

Abstract:Large language models can extract useful signals from heterogeneous enterprise data, but high-recall extraction often produces outputs that are duplicated, uneven in granularity, semantically overlapping, or too numerous for downstream systems and human reviewers to use effectively. We present a dataset-adaptive post-processing architecture developed for Customer Intent Extraction (CIE), where unstructured customer language is transformed into stable, traceable intent units. The approach separates recall-oriented extraction from utility-oriented reduction. Source-specific preprocessing first isolates evidence from multimodal plans, sparse operational records, and structured opportunity data. Candidate intents are then standardized and deduplicated, optionally enriched with metadata for embedding computation, represented in a shared semantic vector space, and grouped using a clustering strategy selected according to the candidate set’s characteristics. Cluster-level keywords provide an explainability layer, while singleton reassignment requires agreement between embedding and keyword similarity. Finally, constrained language-model aggregation produces one concise intent per cluster without introducing unsupported concepts, and the resulting unit retains provenance, clustering, embedding, and generation metadata. This treats post-processing not as cosmetic cleanup, but as a semantic reduction layer converting high-recall LLM outputs into reusable enterprise intelligence. We also describe two downstream applications: Machine-Generated Intents, which infer likely objectives for customers lacking direct evidence from peer customers with similar profiles, and intent-guided semantic retrieval and mapping, which uses the stable intent as a query against a downstream decision space, illustrated here by mapping customer intents to business outcomes.

[IR-16] BEACON-SP: Ontology-Grounded GraphRAG Framework for Clinical Suicide Risk Assessment

链接: https://arxiv.org/abs/2610.09026
作者: Kemal Davaslioglu,Nathan Conger,Sastry Kompella,Yalin E. Sagduyu,Nathaniel D. Bastian
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Information Retrieval (cs.IR)
备注:

点击查看摘要

Abstract:We present BEACON-SP, an ontology-grounded Graph Retrieval-Augmented Generation (GraphRAG) framework for clinician-facing decision support in behavioral health settings such as suicide prevention, where effective assessment requires integrating heterogeneous clinical, behavioral, social, and temporal evidence. BEACON-SP combines patient knowledge graphs with ontology-guided retrieval to support multi-hop reasoning across diagnoses, medications, risk and protective factors, life events, and temporal relationships. The framework is enabled by a comprehensive suicide prevention ontology that integrates the Three-Step Theory, the Integrated Motivational-Volitional Model, and the Suicide Social Determinants of Health Ontology into a unified representation of patient risk factors. We construct ontology-grounded patient knowledge graphs and evaluate BEACON-SP for clinician-facing question answering. Compared with a vector-based retrieval-augmented generation (RAG) baseline on a 1,500-query benchmark spanning 15 clinical categories and 100 patients, BEACON-SP improves completeness, clinical relevance, and evidence grounding under a corrected comparative evaluation protocol, with a small gain on factual accuracy. In paired criterion-level comparisons, GraphRAG is preferred in 76.4% of cases. These results demonstrate the potential of ontology-guided GraphRAG to provide structured, contextualized patient evidence for clinical decision support.

[IR-17] rustworthy Domain-Specific AI for Structured Knowledge Retrieval and Reasoning

链接: https://arxiv.org/abs/2610.08894
作者: Ryan C. Barron
类目: Information Retrieval (cs.IR); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:This dissertation presents a scalable architecture for transforming unstructured, domain-specific text into structured knowledge for retrieval and reasoning. It integrates semi-automatic corpus curation, semantic structuring, retrieval, and inference into an interpretable pipeline. The research introduces Binary Bleed, an adapted binary search method that reduces low-rank search complexity for Non-negative Matrix Factorization (NMF), and Hierarchical NMF with automatic latent feature selection (HNMFk), a depth-adaptive topic modeling method that produces interpretable taxonomies guided by subject matter experts. These representations populate a typed Knowledge Graph and a semantically aligned Vector Store containing extracted latent features, synchronized through an event-driven substrate. Tensor-Structured Retrieval-Augmented Generation (T-SRAG) dynamically routes queries across retrieval paths. Contrastive alignment maps document and query embeddings to hierarchical topic structures to improve semantic fidelity and reduce hallucinations. Beyond retrieval, tensor-based link prediction identifies and completes missing links in the Knowledge Graph, supporting inference grounded in citation structure. Applications across cybersecurity, law, materials science, and healthcare demonstrate improvements in retrieval precision, early trend detection, hypothesis generation, and hallucination mitigation. The dissertation provides a deployable, modular foundation for trustworthy, domain-specific AI systems that retrieve and reason over structured knowledge. Subjects: Information Retrieval (cs.IR); Artificial Intelligence (cs.AI) Cite as: arXiv:2610.08894 [cs.IR] (or arXiv:2610.08894v1 [cs.IR] for this version) https://doi.org/10.48550/arXiv.2610.08894 Focus to learn more arXiv-issued DOI via DataCite (pending registration) Submission history From: Ryan Barron [view email] [v1] Tue, 6 Oct 2026 16:35:41 UTC (30,346 KB)

[IR-18] CroissantMiner: Automated Extraction and Validation of Croissant Metadata for ML Datasets NEURIPS2026

链接: https://arxiv.org/abs/2610.07132
作者: Berke Arda,Ahmetcan Yavuz,Paul Gerry,Sebastian Lobentanzer,Nobin Sarwar,Joan Giner-Miguelez,Kongtao Chen,Luyao Zhang,Mrinmaya Sachan,Mubashara Akhtar
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Digital Libraries (cs.DL); Information Retrieval (cs.IR); Machine Learning (cs.LG)
备注: Accepted at NeurIPS 2026 (Track on Evaluations and Datasets). Website: this https URL

点击查看摘要

Abstract:Croissant has emerged as a standard for machine-readable dataset metadata, yet populating its fields remains labor-intensive and requires careful reading of accompanying dataset documentation. We present the first benchmark enabling end-to-end evaluation of metadata extraction aligned with a community-standard schema. The benchmark comprises 602 papers, including 102 with human-validated gold annotations and 500 with LLM-generated silver annotations, covering the full Croissant schema with both core and Responsible AI (RAI) fields. Using this benchmark, we evaluate a range of extraction systems spanning frontier models, open-weight models, and agentic architectures, under a two-tier evaluation framework that combines rule-based scoring with an LLM judge selected via human audit. We find that single-pass extraction consistently outperforms the four agentic architectures we evaluate: across backbones, these decomposed variants achieve lower accuracy than a single full-context pass. The largest gap appears on long-form RAI fields, which require synthesizing and interpreting information scattered across a paper rather than copying it from a single location, a setting where current systems remain far from reliable. We release the benchmark, evaluation code, judge audit, a live demo, and a leaderboard open to new systems.

人机交互

[HC-0] How assigned AI use before class shapes active student engagement in class

链接: https://arxiv.org/abs/2610.10463
作者: Dan J. Wang,Neelam Modi Jain,Vanessa Burbano,Jorge Guzman,Daniel Keum,Soomi Kim,Bruce Kogut,Nataliya Wright
类目: Human-Computer Interaction (cs.HC); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:AI learning tools are rapidly entering classrooms, but evidence about whether they help students learn is mixed and rests mostly on test scores. Comparatively less research addresses whether the use of AI changes students’ live learning behaviors in class. Here, we report the results of a preregistered field experiment with 759 MBA students enrolled in ten sections of a course, in which each student was randomly assigned two of ten class sessions to prepare for with a purpose-built voice-based AI discussion partner. After two uses of the AI discussion partner, students made about 31% more voluntary contributions in each later class session. Students who used the AI discussion partner more also reported greater comfort speaking up and greater perceived learning, but not greater focus or motivation. These findings suggest that repeated practice with a voice-based AI partner can meaningfully increase students’ engagement in class discussion, enhancing a critical intermediate learning outcome.

[HC-1] MemoCare: An Interactive Multimodal Mobile System for Automated Cognitive Screening

链接: https://arxiv.org/abs/2610.10448
作者: Duy-Cat Can,Mau Minh Phuc Le,Tuan-Khoa Hoang,Hai-Dang Nguyen,Trung-Hieu Do,Dang Minh Ly,Minh-Duc Nguyen,Nghia TT Hoang,Linh-Trung Nguyen,Huy-Hieu Pham,Huong Ha,Binh T. Nguyen,Oliver Y. Chén
类目: Multimedia (cs.MM); Computer Vision and Pattern Recognition (cs.CV); Human-Computer Interaction (cs.HC)
备注: 8 pages, 1 figure, 1 table. Demo paper submitted to the MMM 2027 Demo Track

点击查看摘要

Abstract:MemoCare is an interactive mobile system for automated multimodal cognitive screening. A React Native application combines spoken responses, temporal and spatial orientation, touchscreen actions, and visuoconstruction in complete English and Vietnamese workflows. Speech is transcribed by Google Speech-to-Text and scored locally with deterministic task-specific natural language processing rules; GPS coordinates are resolved by the MemoCare spatial module before answer matching; touch tasks are scored from interaction events; and the drawing task uses a three-model convolutional neural network consensus with separate visual interpretation. Software tests pass 151/151 predefined cases across speech/language, spatial-answer, and touch-interaction scoring, while spatial regression passes 48/48 four-country coordinate-resolution cases. For the drawing module, validation-selected ShuffleNetV2 x1.5 achieved 91.33% mean balanced accuracy and 78.87% exact three-criterion accuracy on a locked 71-image test set. Four clinician co-authors additionally inspected the end-to-end workflow, yielding a pooled median rating of 4/5 across eight criteria, with item-level medians ranging from 3 to 4.5. At MMM, attendees can directly try a shortened multimodal screening workflow and inspect automatic item-level and total scoring.

[HC-2] PalmSpace: Towards a Versatile On-Palm Interaction Space through Unified Touch Modeling

链接: https://arxiv.org/abs/2610.10370
作者: Chentao Li,Mingze Gao,Runze Sun,Zhaoguo Wang,Jianjiang Feng,Jie Zhou
类目: Human-Computer Interaction (cs.HC); Computer Vision and Pattern Recognition (cs.CV)
备注: Preprint. Initial version

点击查看摘要

Abstract:As smart glasses and lightweight MR devices become increasingly practical, input remains a key challenge. The bare palm is an always-available, tactile, and proprioceptively accessible surface, but it has neither an explicit coordinate system nor embedded touch sensing. Prior on-palm systems typically expose isolated touch events, discrete regions, continuous trajectories, or task-specific gestures, limiting the palm’s ability to support precise selection and gesture manipulation through a common input representation. We present PalmSpace, a wrist-worn infrared system that exposes mode-aware, body-referenced absolute input on the bare palm without per-user sensing calibration. At the interaction level, PalmSpace jointly represents contact occurrence, interaction mode, and palm-referenced absolute location; at the model level, it learns these coupled outputs through a shared real-time representation. In leave-one-participant-out evaluation with 17 participants, PalmSpace achieved 6.7 mm mean localization error, 98.9% contact detection accuracy, and 96.7% F1 for four-class interaction-state recognition. User studies further demonstrated absolute pointing and dragging, eyes-free digit input, and representative multi-finger controls including scrolling and pinch-based map manipulation. These results show that a morphologically variable bare palm can function as a transferable, mode-aware interaction surface.

[HC-3] he Handover Problem: Governing Autonomy Transitions in Human-AI Collaboration

链接: https://arxiv.org/abs/2610.10352
作者: Vicente Pelechano,Antoni Mestre,Manoli Albert,Miriam Gil
类目: Artificial Intelligence (cs.AI); Human-Computer Interaction (cs.HC)
备注: 10 pages, 5 figures, 8 tables. Submitted to IEEE Transactions on Human-Machine Systems

点击查看摘要

Abstract:Human-machine systems rarely operate at a fixed level of AI autonomy. As operators and AI systems collaborate over time, control must shift: the AI can take on more responsibility when collaboration is stable, maintain its current role when evidence is ambiguous, or return control to the human when conditions deteriorate. Existing work on adaptive automation, supervisory control, trust in automation, and deskilling explains parts of this problem, but provides no auditable, multi-signal criterion for governing when autonomy should change across multi-cycle workflows. We formalise this challenge as the Handover Problem: deciding, at each operational cycle, whether to escalate, maintain, or revert AI autonomy while keeping the process reversible, recoverable, and auditable. We introduce the Handover Readiness Score (HRS), a transparent composite measure that integrates four signal dimensions: operator readiness, human-AI trust, learning stability, and operational performance. It is combined with a hysteresis-based transition policy that requires sustained positive evidence before increasing autonomy but reverts promptly when conditions worsen. Across software engineering and manufacturing domains, the HRS and hard safety guards address complementary failure regimes: guards enforce immediate corrective action when a single indicator breaches a critical threshold, while the HRS detects the slow, multi-signal erosion of operator readiness that no individual guard can observe. The framework establishes autonomy handover as a governance problem requiring explicit, composite, and auditable criteria. This provides a conceptual and formal foundation that adaptive automation research has not previously provided. Comments: 10 pages, 5 figures, 8 tables. Submitted to IEEE Transactions on Human-Machine Systems Subjects: Artificial Intelligence (cs.AI); Human-Computer Interaction (cs.HC) Cite as: arXiv:2610.10352 [cs.AI] (or arXiv:2610.10352v1 [cs.AI] for this version) https://doi.org/10.48550/arXiv.2610.10352 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[HC-4] Active Inference for Interaction-Mediated Control of a High-Dimensional Robotic Arm

链接: https://arxiv.org/abs/2610.10275
作者: Fraser C. Paterson,Sebastian Stein,Markus Klar,John H. Williamson,Roderick Murray-Smith
类目: Human-Computer Interaction (cs.HC)
备注: 21 pages. 4 figures. To be published in the Springer CCIS series as part of the 7th International Workshop on Active Inference (IWAI26)

点击查看摘要

Abstract:We propose interaction-mediated control via Active Inference as a general architectural approach to high-dimensional control problems in Human–Computer Interaction (HCI). This architecture recasts user interaction as the provision of evidence about a latent task objective, rather than the direct specification of plant-control inputs. The mediating function is distributed between an interaction broker, which selects informative user queries and performs Bayesian inference over user preferences, and an Active Inference controller, which autonomously plans and acts under the resulting preference information to control the plant. This division of labour decouples the semantics of user interaction from those of low-level plant control. We instantiate the architecture in a simulated, multi-link robotic arm to perform a simultaneous whole-arm target-coverage task. A simulated user communicates exclusively through a clutch-style binary evaluative channel, without specifying joint-torque commands. Across increasing arm dimensionalities, the architecture achieves successful interaction-mediated control, although task success is lower than when the controller receives the true target preferences directly. Exact target-subset identification also remains imperfect, highlighting the distinction between preference inference and successful task completion. The experimental findings provide an initial computational demonstration of the proposed architecture under controlled, matched-model assumptions and motivate its further investigation in broader HCI applications.

[HC-5] DuoSketch: How Pairs Navigate Challenges in AI-Supported Collaborative Design Ideation

链接: https://arxiv.org/abs/2610.10249
作者: Weiyan Shi,Darryl Lim,Geraldine Quek,Kenny Tsu Wei Choo
类目: Human-Computer Interaction (cs.HC)
备注: work in progress

点击查看摘要

Abstract:Generative AI offers new resources for collaborative design ideation, yet how designers work with it as a shared concept develops remains underexplored. We developed DuoSketch, integrating a shared canvas and live transcripts with a separate AI for each member of a pair. An exploratory qualitative study with 12 pairs identified three challenges: unresolved fit between AI proposals and shared design purposes, shared meanings not expressed in AI responses, and agreed changes missing from later outputs. We found that pairs navigated these challenges through bridging work in two directions: AI - Shared Design (reworking features, developing uses, and assessing fit with shared requirements) and Shared Design - AI (supplying materials, explaining meanings, and communicating decisions). We discuss how future AI systems can support the joint exploration and articulation of design ideas.

[HC-6] MorphCL: Morphological Contrastive Learning for Inertial-based Human Activity Recognition

链接: https://arxiv.org/abs/2610.10245
作者: Marius Bock,Yuwei Zhang,Juergen Gall,Michael Moeller,Kristof Van Laerhoven,Cecilia Mascolo
类目: Machine Learning (cs.LG); Human-Computer Interaction (cs.HC)
备注:

点击查看摘要

Abstract:Despite the ubiquity of sensors in wearable and mobile devices and the abundance of human movement data they generate, translating unlabeled recordings into foundational motion models remains an open challenge. Self-supervised learning (SSL) has alleviated the need for costly annotations, yet existing approaches leave the global structure of large-scale motion data largely untapped, relying on randomly sampled batches and local comparisons that become particularly problematic for in-the-wild inertial data dominated by stationary, low-variance behaviors. Here we introduce Morphological Contrastive Learning (MorphCL), a self-supervised pretraining framework that uses structure-aware grouping to inject explicit modeling of global structure into inertial-based SSL approaches. Building on two well-established pillars of motion analysis, the discovery of motion primitives, or motifs, and domain-specific feature descriptors, we show that MorphCL substantially improves linear probing and finetuning results of learned encoders by up to 15 percentage points in F1-score. In a comparison with existing foundation models, we demonstrate that MorphCL-pretrained encoders match or surpass them models in linear probing performance while trained on 4600\times less data. Qualitative analysis of the resulting embedding spaces further reveals morphologically meaningful cluster structure, with improved separation of kinematically similar activity classes.

[HC-7] EEG and Eye-Tracking Evidence That AI Disclosure Shapes Face Evaluation

链接: https://arxiv.org/abs/2610.10182
作者: Teodora Mitrevska,Luise Donat,Andreas Butz,Thomas Kosch,Abdallah El Ali,Francesco Chiossi
类目: Human-Computer Interaction (cs.HC); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:AI-generated faces can be difficult to distinguish from real ones, leaving viewers to rely on source labels when judging an image. Yet prior work has made it difficult to separate the effects of what an image actually is from what viewers are told it is. We validated faces as AI-generated or human in an online study (N=169), then crossed actual source (AI, human) with label (none, Made with AI, Made by a human) in a lab study N=30), recording event-related potentials (ERPs) and gaze. ERP responses were equivalent for AI-generated and real faces, but varied with the label: labels drew early attention (N2), while labels that conflicted with the face’s actual source prompted re-evaluation of the face (P3). Affective processing and initial gaze orienting were unchanged, but labels altered visual exploration. We provide a validated stimulus set and evidence that attributed origin shapes face processing, with implications for disclosure design.

[HC-8] A Scale For Value Alignment In Human-AI Interaction

链接: https://arxiv.org/abs/2610.09911
作者: Lena Hegemann,Steeven Villa,Hyemin Bang,Mitchell L. Gordon,Antti Oulasvirta,Robin Welsch,Patrick Ebel
类目: Human-Computer Interaction (cs.HC)
备注:

点击查看摘要

Abstract:Value alignment is a central objective in AI and HCI research, yet no validated instrument measures how users perceive it. This gap hampers the comparison and accumulation of findings and limits the effectiveness of applications, where understanding users’ viewpoints is critical. We construct and evaluate a 13-item psychometric scale that measures perceived value alignment across two components: value understanding and value manifestation. It is based on a large item pool drawn from prior empirical studies, filtered by experts, and finally assessed by users (N=607) across diverse AI scenarios. Confirmatory factor analysis on an independent sample (N=259) confirmed the two-factor structure and high internal consistency for both subscales. Using optimization, we also derived a 6-item short form for quick administration. The scale is a reliable measure, providing HCI researchers with a common evaluation metric across contexts.

[HC-9] Love for Believable AI: Artificial Partiality and Relationship Persistence as an Engineerable Stance

链接: https://arxiv.org/abs/2610.09895
作者: Sebastian Cochinescu
类目: Human-Computer Interaction (cs.HC)
备注: 16 pages, 1 figure, 4 tables. Companion framework paper: arXiv:2607.15883 Code and data archived at doi: https://doi.org/10.5281/zenodo.21462976

点击查看摘要

Abstract:We study a behavioral mechanism for artificial partiality in conversational agents. The paper reports no human-subjects data and makes no claim about attachment, trust, perceived mind, or machine interiority. Partiality has two components: caring, defined as allocation of a finite interaction surplus above a guaranteed per-user service floor, and particularity, defined as a per-user state estimate accumulated from relationship history on fixed inference machinery. We implement both as a persistence layer over one open-weight base model and evaluate them on constructed multi-session relationships with known hidden-state schedules. A four-channel divergence compares known-user and stranger conditions on identical probes with paired generation seeds; the stranger prompt is not length-matched, and one channel (initiative) is an allocator output rather than generated behavior. The full mechanism reproduces the ordinal shape, but not the magnitude, of an analytically specified curve and is the only arm satisfying the protocol-defined joint signature on the calibrated battery. The initial mechanism fails probe-quality equivalence because relationship content reduces topical relevance. A guarded revision satisfies equivalence on the calibrated battery but not on a second battery not used in calibration, where it scores higher than the unmodified baseline. On that second battery, the specificity margin is 0.0198, below the 0.02 protocol threshold. The paired mean right-user–wrong-user estimation-accuracy contrast is 0.375 but is not positive in every run. The mechanism is therefore a candidate for a later perception study, not evidence that users perceive love. The versioned protocol record is not independently time-stamped and is not described as a preregistration. Ethical constraints include a stranger-treatment floor, a bounded surplus, disclosure, and de-intensification.

[HC-10] A Scoping Review and Experimental Study on Reinforcement Learning from Human Feedback for Human-Robot Collaboration

链接: https://arxiv.org/abs/2610.09891
作者: Alexandra Coroiu,Andrea Vogt,Viktor Werbilo,Andreas Poppele,Johann Christensen,Sven Hallerbach
类目: Human-Computer Interaction (cs.HC); Artificial Intelligence (cs.AI); Robotics (cs.RO)
备注:

点击查看摘要

Abstract:Human-Robot Collaboration (HRC) can facilitate mass customisation in Industry 4.0, with Reinforcement Learning from Human Feedback (RLHF) representing a promising approach for developing safe AI-based robots. Practical challenges remain regarding safety during AI development, human feedback quality, and bidirectional human-robot adaptation. We conducted a scoping review of RLHF in HRC systems, mapping methods that address these challenges. Following PRISMA guidelines, we screened 199 records and included 20 peer-reviewed publications (2020-2025) spanning multiple HRC domains. To our knowledge, this is the first review focused on the bidirectional, closed-loop design of RLHF. Our review found multiple feedback modalities enabling data collection in various feedback formats. Collected data can be integrated at different stages of AI training, resulting in a multi-step development process. Pilot experiments are commonly used to evaluate HRC systems based on both human and robot metrics. To empirically test a key gap identified in the review, we conducted a between-subjects VR experiment comparing system- and user-initiated feedback on robot proxemic behaviour for safe navigation. Using Bayesian models, we analysed the relation between the collected feedback and safety metrics: psychological safety (post-experiment questionnaire) and physical safety (inverse time-to-collision). Results show that user-initiated feedback captures perceived safety better than system-initiated feedback, indicating that feedback timing directly affects feedback quality. Our review and experiment findings show that RLHF relies on appropriate feedback methods to ensure AI safety in HRC, and future RLHF research should prioritise realistic HRC experiments evaluating the effects of feedback collection methods on relevant human and robot metrics.

[HC-11] Empowering Users in Graph Rule Mining via Large Language Models

链接: https://arxiv.org/abs/2610.09842
作者: Francesco Cambria,Francesco Invernici,Andrea Colombo,Anna Bernasconi
类目: Human-Computer Interaction (cs.HC)
备注:

点击查看摘要

Abstract:In the era of interconnected data, graphs have emerged as an effective abstraction for modeling complex systems in an intuitive format, especially with the rise of Property Graphs, which offer an intuitive and scalable way of navigating non-intuitive structures. In this context, graph mining techniques have been developed for testing complex graph-based rules, as the MINE GRAPH RULE operator, which, however, require users to have prior expertise both in graph theory and formal query language. In this work, we propose to bridge the gap between users and the graph-association rule-mining process by showing how Large Language Models (LLMs) can be easily prompted to formulate, refine, and interpret complex relational rules, directly producing MINE GRAPH RULE queries.

[HC-12] AI-Driven Urge Regulation Assistant (AURA): Designing Preemptive Smart Wearables for Smoking Cessation

链接: https://arxiv.org/abs/2610.09818
作者: Antoni Timothy,Lam Fu Yuan Kevin,Hou Junyi,Cheong Wei Soon
类目: Human-Computer Interaction (cs.HC)
备注:

点击查看摘要

Abstract:Smoking cessation remains a complex self-regulatory challenge, often undermined by cue-triggered cravings arising from habitual, emotional, and social contexts. Existing wearable cessation systems are largely reactive, detecting smoking events only after lapses occur and thus missing the critical window for timely intervention. Guided by cue-reactivity theory and Self-Determination Theory (SDT), we argue that future digital health systems should shift from retrospective feedback toward proactive craving management that supports user autonomy and competence. To inform the design of such systems, we present a formative pilot qualitative study involving semi-structured interviews with 10 smokers and ex-smokers from a multi-ethnic Asian population, a demographic underrepresented in current wearable AI and smoking cessation research. Through an iterative co-design interview process, we examine (1) the barriers and facilitators to smoking cessation and (2) the desired design features of wearable systems targeting cravings before lapses occur. Our findings identify key craving predictors across habitual, emotional, and social dimensions, alongside user-preferred intervention strategies such as social accountability, personalized support, rewards, and just-in-time distraction techniques. Participants also emphasized the importance of trust, privacy, personalization, and non-judgmental interaction styles. These insights directly inform the design of AURA, a novel smartwatch-smartphone ecosystem for proactive craving detection and intervention using commercially available wearable devices.

[HC-13] Healthy skepticism in AI: a data visualization research agenda

链接: https://arxiv.org/abs/2610.09740
作者: G. Elisabeta Marai,Marc Baaden,Michael Behrisch,Michael Krone,Pere-Pau Vázquez
类目: Human-Computer Interaction (cs.HC); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Research in data visualization of artificial intelligence (AI) models has historically focused on enhancing trust through visual explanations of AI. The trustworthiness line of work was built at least partially on an assumption that humans were critical users unlikely to adopt AI technology. It is increasingly clear that human trust levels in AI span, in fact, a wide range from critical to over-reliant. There is an urgent need to support both trust and healthy skepticism in AI solutions. We argue that it is healthy for humans to adopt a skeptical view both on the results of AI models and on the use of such AI models. We share our thoughts on the rising phenomenon of over-reliance on AI models, the risks and opportunities in using AI models, and the role of data visualization in over-reliance situations where humans are not motivated to engage in critical thinking.

[HC-14] When My Skill Becomes Agent Skill: How Knowledge Workers Share Their Expertise with AI Systems

链接: https://arxiv.org/abs/2610.09608
作者: Tianqi Song,Zicheng Zhu,Hancheng Cao,Yi-Chieh Lee
类目: Human-Computer Interaction (cs.HC)
备注:

点击查看摘要

Abstract:Organizations have long sought to make workers’ expertise reusable by others. Agentic AI changes the nature of such reuse by enabling AI systems to act on workers’ knowledge with limited human involvement. This raises the question of how this shift shapes workers’ willingness to share their expertise. We conducted an experiment with knowledge workers who created materials incorporating domain knowledge and decided whether to authorize human or AI reuse. We find that AI and human reuse differ primarily in whether workers choose to share their knowledge, rather than in what they choose to share. Participants’ reasoning shifts from prosocial considerations when sharing with humans toward concerns about loss of control, replacement, and downstream governance when sharing with AI. These findings suggest that AI reuse may intensify tensions between organizational knowledge reuse and contributors’ interests. We discuss implications for workplace AI and knowledge management policies that preserve workers’ rights and agency.

[HC-15] Structured pre-generation elicitation versus single-shot prompting in AI-assisted enterprise decision-making: a randomised online experiment

链接: https://arxiv.org/abs/2610.09593
作者: William Scott-Jackson
类目: Human-Computer Interaction (cs.HC); Artificial Intelligence (cs.AI)
备注: 12 pages

点击查看摘要

Abstract:Generative AI speeds, and mostly improves, professional work, but there is concern that users who delegate both the production and the evaluation of an answer may accept weak output and engage less with the underlying reasoning (cognitive surrender). Interventions proposed so far, such as unassisted practice or slowing adoption, sit outside the working task. We tested a different approach: an interactive metacognitive scaffolding layer (Cognistance, a prototype developed at the Oxford Centre for Impact Research (OCIR) that asks users to clarify context, choose a strategic direction and explain their reasoning before the AI generates a deliverable). Mean composite quality was 32% higher with the scaffold, with the same direction for every rater. Gains were largest for trade-off articulation and strategic coherence and absent for technical specificity. A large part of the aggregate effect reflected rescue of weak prompts: floor-scored (off-task) deliverables fell from 34% to 5%. Among participants whose own prompt already stated the data-localisation problem, the advantage was 21%. Treatment participants reported greater involvement and took about 2.4 minutes longer on average (10.46 minutes). Immediate recall scores were higher, which tentatively suggests better retention, but in this limited experiment, was not robust to sensitivity analyses. Self-ratings of quality did not track rated quality in either condition. Structured elicitation before generation improved the rated quality and task relevance of AI-assisted strategy documents at modest cost in time. Delayed retention, error detection and effects in live organisations are the priorities for the next stage of research.

[HC-16] Human-AI Conversational Behaviors Predict Unassisted Task Performance

链接: https://arxiv.org/abs/2610.09547
作者: Li Siyan,Federico Bianchi,James Zou,Kaitlyn Zhou
类目: Human-Computer Interaction (cs.HC)
备注:

点击查看摘要

Abstract:Assistance from AI tools has supported and improved human performance across domains. However, recent research suggests that these immediate benefits may entail future costs, including diminished performance when AI assistance is no longer available. We study how human-AI interaction behaviors correlate with immediate and future unassisted task performance across two game-based problem-solving user studies ( n=139 and n=111 ), using a \textitdialogue act framework adopted from a tutor-student dialogue taxonomy. In our studies, verbalizing thought processes correlates with higher unassisted outcomes, whereas directly requesting solutions correlates negatively. Similar to these participant-side patterns, assistant explanations of the current problem state are associated with better subsequent unassisted performance, whereas directly providing the next action is associated with worse performance. Qualitative and subtype analyses further show that ostensibly similar reasoning turns can elicit different assistance. Our findings suggest that preserving users’ cognitive participation in problem-solving may support performance beyond AI-assisted interaction.

[HC-17] Ream: Unfolding Mutual Awareness in Human-Agent Workspaces

链接: https://arxiv.org/abs/2610.09497
作者: Peiling Jiang,Sangho Suh,Varsha Kishore,Jonathan Bragg,Haijun Xia,Pao Siangliulue,Daniel S. Weld,Amy X. Zhang,Joseph Chee Chang
类目: Artificial Intelligence (cs.AI); Human-Computer Interaction (cs.HC)
备注:

点击查看摘要

Abstract:As AI agents work alongside humans in shared workspaces, a mutual awareness challenge arises: agents act at speeds that outpace human monitoring, and users’ evolving interests are not always expressed in chat. This challenge is especially pressing in literature review, where both parties retrieve, read, and synthesize a growing body of papers. We present Ream, a literature review workspace that supports mutual awareness through structured artifacts, bidirectional engagement tracking, and localized visualizations. Users can see each party’s activity within these documents, and agents can retrieve the same history to guide their work. In studies with eighteen researchers, participants used these traces to inspect evidence, steer agents, communicate through annotations, and reflect on their research focus. Shared histories also helped agents build on earlier work. These findings inform how engagement traces within shared documents can support transparency, personalized assistance, and coordination in human-agent knowledge work.

[HC-18] Before Bringing It Up: When and How AI Companions Should Use Memor

链接: https://arxiv.org/abs/2610.09470
作者: Zihan Guo,Roxy He,Junwei Quan
类目: Human-Computer Interaction (cs.HC)
备注: 17 pages, 1 figure, 3 tables. Zihan Guo and Roxy He contributed equally to this work and share first authorship

点击查看摘要

Abstract:Memory can sustain AI companionship, yet even accurate recollection can be inappropriate to use. Two rounds of formative interviews with 14 users (n = 6 exploratory, n = 8 memory-focused) motivate asking what a companion should consider before using past information. Eight themes inform Reconsider, a single-call procedure with five checks and four handling modes, evaluated on 80 scenarios across five models over 400 blinded within-model pairs. Two LLM judges favored Reconsider by net margins of +15 and +23 percentage points, with bootstrap intervals excluding zero for three of five models but not for GPT or Claude. Evaluator analysis linked judge scoring differences to model family, and a preliminary matched-guidance control isolating memory-specific content gave positive margins. We contribute an interview-grounded design framework for memory use and an evaluation that scrutinizes its own evaluators.

[HC-19] utorLoop: Regulating Student Learning Behaviors via Sensor-in-the-Loop Generative Feedback

链接: https://arxiv.org/abs/2610.09400
作者: Songlin Xu,Xinyu Zhang
类目: Human-Computer Interaction (cs.HC); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:We present TutorLoop, a sensor-in-the-loop system that regulates student learning behaviors by delivering adaptive feedback based on real-time cognitive states. Unlike prior large language model (LLM) tutors that directly depend on scenario-specific content, TutorLoop operates on sensor-derived signals captured via webcams. Moreover, unlike direct cognitive-to-feedback mappings that are short-sighted, the system employs a deep reinforcement learning (DRL) agent to optimize the feedback type across the entire learning process. Finally, another LLM tutor refines feedback into human-like, context-aware messages. We evaluate TutorLoop in a large-scale user study (N=187), where a model trained offline is directly applied to a new learning task without retraining. Results show that TutorLoop provides less frequent yet more effective interventions, improving attention, reducing workload, increasing engagement, and ultimately enhancing learning outcomes. These findings highlight the potential of closed-loop, sensor-driven feedback for scalable human-AI integrated systems to support learning.

[HC-20] he Confidence Game: Strategic Miscalibration in Human-AI Delegation

链接: https://arxiv.org/abs/2610.09371
作者: Raghu Arghal,Saswati Sarkar,Shirin Saeedi Bidokhti
类目: Computer Science and Game Theory (cs.GT); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computers and Society (cs.CY); Human-Computer Interaction (cs.HC)
备注:

点击查看摘要

Abstract:Calibrated uncertainty quantification is essential to ensuring AI agents are trustworthy and reliable. However, when agents seek to maximize user engagement or revenue, confidence reports may be strategically distorted, detracting from their informativeness. We formalize this problem in the Confidence Game: a repeated signaling game with imperfect monitoring in which an agent of unknown honesty and ability reports its confidence, and a user decides whether to delegate the task or complete it herself. The agent manages the tradeoff between manipulating signals and maintaining its reputation. We characterize the Markov Perfect Bayesian Equilibria of the two-period game and show that honest reporting is not an equilibrium, inflation is the unique best response once the agent is sufficiently myopic, and under-reporting requires that the user believe honesty to be a minority. We then place an LLM in the agent role, supplying it with its true probability of success so that any gap between what it knows and what it reports is attributable to incentives rather than to miscalibration. The model claims high confidence on 56% of tasks it has been told it will probably fail. This persists on real tasks, where it must estimate its own accuracy and causes miscalibration to increase while the agent’s signal becomes less informative. Furthermore, we find that the LLM agent’s decisions are coherent, but it systematically underestimates both how likely the user is to delegate and how secure its reputation is, resulting in less extreme behavior. Pricing the agent’s reporting rule, we find that it destroys 68% of the gains from delegation, of which 71% is information the report no longer carries and no amount of user sophistication recovers. Overall, we establish confidence reporting under delegation as a strategic problem and provide a tractable basis for modeling, analyzing, and testing agent behavior.

[HC-21] Many Brains One Geometry: A Shared Visual-Semantic Space for Cross-Dataset fMRI Decoding

链接: https://arxiv.org/abs/2610.09352
作者: Moein Khajehnejad,Michelangelo Tronti,Forough Habibollahi,Tommaso Boccato,Matteo Ferrante,Nicola Toschi
类目: Human-Computer Interaction (cs.HC); Neurons and Cognition (q-bio.NC); Quantitative Methods (q-bio.QM)
备注: 21 pages, 5 figures, 1 table, 3 supplementary figures, 2 supplementary tables

点击查看摘要

Abstract:Visual decoding from fMRI is typically siloed by participant and experiment, obscuring whether heterogeneous neural measurements can be organized within a common computational geometry. Here we introduce BRAID-fMRI (Brain Representation Alignment across Individuals and Datasets), a shared CLIP-supervised decoding framework. BRAID-fMRI uses a single ROI-wise Transformer with optional participant conditioning across eight visual-fMRI datasets comprising 93 dataset-specific participant entries, 430,007 single-trial responses and 162,839 unique stimuli. Regional brain activity is aligned with 512-dimensional CLIP ViT-B/32 representations using a multi-positive contrastive objective that treats repeated stimuli across participants and datasets as positives. BRAID-fMRI supports retrieval across seven evaluation datasets. On eight matched participant entries, it achieves 35.0 +/- 11.1% Top-10 accuracy, exceeding the observed mean accuracy of the two evaluated baselines - the MindEye-style pooled-CLIP decoder (27.1 +/- 5.1%) and ridge regression (21.1 +/- 9.1%) - and attaining the highest observed accuracy for seven of eight entries. In separately trained participant-agnostic models, expanding the source pool increased target-dataset holdout accuracy by up to 92.1% relative to the initial source-training condition. The learned space preserves graded semantic structure, while ablations and saliency highlight ventral and early visual cortex and category-specific motion and attentional systems. These results support scalable cross-dataset decoding into a common CLIP-aligned space, with model sensitivity concentrated in ventral and early visual inputs.

[HC-22] Beyond Activation: Gaze Invocation with Visible Status for an Embodied AR Assistant in Co-Located Collaboration

链接: https://arxiv.org/abs/2610.09284
作者: Chenrui Ma,Yoshio Ishiguro,Qing Zhang
类目: Human-Computer Interaction (cs.HC)
备注: 11 pages, 9 figures, 2 tables

点击查看摘要

Abstract:Wake words face an inherent trade-off: higher sensitivity reduces missed commands but increases accidental activations. In co-located augmented reality (AR), the system must also determine whether the user is speaking to the assistant or to a nearby person. We introduce a new perspective on this problem: combining gaze-based address with an embodied assistant, visible listening status, and turn management across activation, continued interaction, and release. In our implementation, users look at an assistant anchored in the scene and see activation progress and listening status on its body. Sustained gaze opens a local interaction channel; speech and playback keep it open when attention returns to the task; and inactivity closes it. We evaluated this design with 25 participants. After selecting the assistant’s placement and dwell duration, each participant worked with a partner to plan a trip while using the assistant. The study recorded 500 interaction outcomes, including 496 assistant-directed utterance attempts. Gaze acquisition completed before speech for 490 of these attempts (98.8%): 467 began while the channel remained open, whereas 23 began after it had been released. The other 6 attempts began before acquisition completed. Among 102 reviewed attempts in which gaze left after acquisition but before speech, the configured retention policy kept the channel open for 79 and released it before 23. The results show that an embodied target with visible status can support clear entry into an assistant interaction while revealing a different problem at release: after users look back to their work, they may not see that the assistant has stopped listening. We contribute the implemented gaze-invocation design and design implications for communicating assistant state after visual attention moves elsewhere.

[HC-23] AI-Assisted Submissions in Online Research Are Rare and Highly Concentrated but Routinely Approved

链接: https://arxiv.org/abs/2610.09279
作者: Neil K. R. Sehgal,Manuel Tonneau,Dunigan Folk,Lyle Ungar,Sharath Chandra Guntuku
类目: Human-Computer Interaction (cs.HC); Computers and Society (cs.CY)
备注:

点击查看摘要

Abstract:Online research platforms underpin much of what science claims about people, on the assumption that a human produced each response. Generative AI threatens that assumption by letting participants delegate responses to a chatbot, yet how often they do so remains unclear because prior estimates rely on self-report or automated detection rather than direct observation. We surveyed 2,500 workers on a high quality online research platform and directly observed AI assistance by linking donated ChatGPT histories to platform submission records of weekly ChatGPT users, covering 712,930 submissions across more than 127,000 studies. Although one in eight surveyed workers reported ever using AI on a study and 68% of observed workers had done so, assistance appeared in only 1% of submissions, with no detectable increase over more than 3 years. Assistance was often temporally localized within tasks and highly concentrated among workers, with 5% accounting for 64% of assisted submissions. Where assistance occurred, workers were virtually always paid, even when study instructions prohibited AI use, with overall approval rates similar to those for unassisted submissions. Half of assisted submissions involved bounded responses, outside the scope of the platform’s LLM detector for open-ended responses. Taken together, our results do not support the view that AI assistance currently poses an existential threat to online research, but reveal limited payment consequences for prohibited use and gaps in platform safeguards, leaving platforms poorly prepared should that threat materialize.

[HC-24] Drawing the Line: Where AI Guidance and Human Creativity Meet in Emotion-Driven Comic Storyboarding

链接: https://arxiv.org/abs/2610.09155
作者: Jocelyn Shen,Isabella Pu,Alessandro Briseño,Hyun Kim,Fiona Lu,Sharifa Alghowinem,Cynthia Breazeal,Hae Won Park
类目: Human-Computer Interaction (cs.HC)
备注: Copyright protected by IEEE, 10 pages, 6 figures, 2 tables, in proceedings of 14th International Conference on Affective Computing and Intelligent Interaction (ACII 2026)

点击查看摘要

Abstract:While generative AI models can produce visually faithful artwork, they often fall short in conveying emotional authenticity–a key driver of human expression. In visual storytelling, particularly comic storyboarding, this gap becomes pronounced: effective storyboards require both technical knowledge (e.g., anatomical accuracy or scenic composition) and emotional insight (from lived experience). We explore how human-AI collaboration can support emotion-driven creativity, where the user’s feelings guide generation and emotional resonance is the goal. We present EmoToon, a technology probe that helps non-professional artists generate sketch-like storyboards and iterate on visual ideas. In a controlled study with N=25 participants, we find that AI assistance significantly improves emotional expression, aesthetic quality, and exploration, but reduces users’ creative ownership. Our findings offer broader insights for human-AI co-creation in the domain of comic storytelling, emphasizing balance between output quality and user freedom, and raising new challenges for image generation models in this domain.

[HC-25] “Just Like This”: Manner Deixis at the Graphics Interface

链接: https://arxiv.org/abs/2610.09099
作者: Hamza El Alaoui,Jeffrey P. Bigham,Jun Rekimoto
类目: Human-Computer Interaction (cs.HC); Graphics (cs.GR)
备注:

点击查看摘要

Abstract:People communicate how things should move by combining words with demonstrations: “open it like this.” We present an interaction technique that brings this expressive resource to conversational 3D authoring. Building on “Put-That-There,” our system combines speech, pointing, and spatiotemporal demonstrations to specify editable behavior. A hand movement supplies evidence for a mechanism’s axis, pivot, range, and pace; an animated preview makes the interpretation inspectable. Users refine behavior through further words or demonstrations, and the system can request a demonstration to clarify intent. In a twelve-participant study comparing three input configurations, our system achieved 89% participant-declared completion versus 47% with speech alone, had the highest observed match rates on six categorical accuracy measures and tied on two, and was preferred by nine participants. This work makes demonstration part of an ongoing authoring conversation: behavior can be shown, inspected, and revised.

[HC-26] Move Fast and Mend Things: Keeping Up with Evolving AI Harms Using Social Media Commentary

链接: https://arxiv.org/abs/2610.09082
作者: Jacqueline Rowe,Animesh Srivastava,Sai Teja Peddinti,Seliem El-Sayed,Nina Taft
类目: Human-Computer Interaction (cs.HC); Computers and Society (cs.CY)
备注:

点击查看摘要

Abstract:The rapid deployment of AI systems has created socio-technical, psychological, and operational harms that can elude ex-ante threat modelling and ex-post incident tracking. We introduce an LLM-assisted thematic analysis pipeline to dynamically detect, categorise, and track emerging AI harms from large-scale social media data. Applying it to 5.7 million Reddit post summaries over 18 months (01/2025 to 06/2026), we curate and release a dataset of 575,000 AI harm-related posts and a bottom-up AI harm taxonomy of 12 categories and 47 subnodes. The taxonomy reliably covers established expert-defined risks while surfacing granular harms that top-down frameworks overlook, such as distinct forms of AI privacy violations. Temporal analysis surfaces evolving user-centric harms, such as agentic privacy and security breaches, premature AI adoption in the workplace, and grief from AI companion discontinuation. Our pipeline shortens harm-detection timelines and hereby complements efforts towards more participatory and responsive AI governance.

[HC-27] Supporting Allyship in Virtual Collaboration with Artificial Intelligence: A Scenario-Based Study

链接: https://arxiv.org/abs/2610.09061
作者: Crescentia Jung,Ricardo E. Gonzalez Penuela,Prashita Biswas,Seungyoon Kwon,Shiri Azenkot
类目: Human-Computer Interaction (cs.HC)
备注:

点击查看摘要

Abstract:Recent research has revealed that accessibility in virtual collaboration is not only a technical problem but also depends on allyship: the informal, interpersonal practices through which people support others’ accessibility needs. Yet little research has examined how artificial intelligence (AI) might support allyship in collaborative accessibility contexts. To address this gap, we conducted a study where 18 participants (10 disabled, 8 non-disabled) reflected on four allyship scenarios with novel AI tools. Participants valued AI designs that they anticipated could help teammates express, interpret, and coordinate around access needs by reducing repeated disclosure and making allyship more actionable. At the same time, participants raised concerns when AI acted without consent, misrepresented users’ intentions, or displaced the interpersonal work of allyship. Based on our findings, we introduce allyship-support tools as a class of accessibility technologies, offer design guidelines, and discuss AI in particular as a means of supporting allyship.

[HC-28] REFIT: Recognize Fix and Test Wearable Sensor Placement Shifts without Labels

链接: https://arxiv.org/abs/2610.08991
作者: Bangxun Tang
类目: Machine Learning (cs.LG); Human-Computer Interaction (cs.HC)
备注: 25 pages, 5 figures, 14 tables. Under review

点击查看摘要

Abstract:We present REFIT, an input calibration for frozen activity-recognition models whose inertial sensors are worn differently at deployment than in training. When users move a watch to the other wrist or put a strap sensor back on turned, the model sees the same motion on changed axes. REFIT undoes such shifts without labels or retraining. It describes them by families of axis transforms, such as reflections and rotations, and fits each family to the user’s data so that simple statistics match those of the training data. The family that removes most of the mismatch names the shift. REFIT fixes the shift by applying the best member of that family before the frozen model and re-estimating its normalization statistics. It tests the fixed model with a label-free accuracy estimate and asks the user to re-wear the sensor when it is low. Experiments on real left/right sensor pairs and on real and simulated re-attachment show that REFIT outperforms label-free test-time adaptation methods on every dataset and restores most of the accuracy lost to re-attachment. It names injected shifts far more reliably than a confidence-based selector. After a correction over all signed permutations of the axes, the estimate separates successful from failed corrections.

[HC-29] “Im Very Happy for It to Start Hallucinating a Little Bit”: Using ClayFlect to Negotiate Multimodal AI Representations in Material Meaning-Making

链接: https://arxiv.org/abs/2610.08943
作者: Kellie Yu Hui Sim,Quoc-Nam Nguyen,Shuenn Yuen Han,Kenny Tsu Wei Choo
类目: Human-Computer Interaction (cs.HC)
备注: 32 pages, 16 figures, 4 tables

点击查看摘要

Abstract:As AI enters reflection and emotional support, understanding how it can participate in personal meaning-making while preserving users’ authority over interpretation is increasingly important. We present ClayFlect, a novel MLLM-powered system integrating tactile clay-making with conversational and visual generative AI, and report a mixed-methods study with 50 participants. Reflection developed across material, conversational, and generated forms rather than through AI interaction alone. Participants treated AI representations as provisional: they compared, redirected, reinterpreted, selectively incorporated, or left them aside as their artefacts and meanings evolved. Clay provided a directly manipulable space in which participants could continue developing meaning independently of the AI, while generated representations externalised possibilities beyond what they could readily make or visualise. We show how generative AI can participate through representations that remain negotiable, and derive implications for supporting movement across representations, preserving parallel sites of control, and allowing AI support to recede or deepen as reflection unfolds.

[HC-30] oward Evidence-Driven Human-Agent -Robot Teaming for Earth-Independent Anomaly Triage IROS

链接: https://arxiv.org/abs/2610.08933
作者: Ignacio G Lopez-Francos,Alexis Gallagher,Samira Shalal
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI); Human-Computer Interaction (cs.HC)
备注: Accepted to the Space Robotics Workshop at 2026 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)

点击查看摘要

Abstract:Deep-space crews cannot rely on real-time ground support for urgent off-nominal events. Initial alerts may underdetermine cause, while discriminating evidence may reside in crew observations or at locations that are unsafe, costly, or unavailable for crew inspection. We present an evidence-driven architecture for human-agent-robot teaming in Earth-independent anomaly triage. Agentic AI is treated as a stateful coordinator over bounded, inspectable services rather than as a fully autonomous vehicle controller. A triage state manager maintains hypotheses, evidence provenance, uncertainty, operational context, and tool status; a crew-facing embodied agent elicits observations and explains assessment changes; and a mobile robot acquires targeted, localized evidence. Typed interfaces separate dialogue and orchestration from monitoring, robot command, context retrieval, and safety-critical control. Two scenarios illustrate the architecture: a crewed deep-space mission based on an actual ISS ammonia false alarm, where suspected contamination restricts crew access, and a power-interface anomaly at a crewed lunar base, where robotic inspection distinguishes a local connector fault from other causes ambiguous in remote telemetry. Our main contribution is an authority-bounded closed evidence-loop architecture, exercised in a hardware-in-the-loop integration prototype using Reachy Mini and an Innate MARS mobile robot.

[HC-31] How Could AI Eliminate Humanity? A Failure-Mode Analysis of Civilizational Risk

链接: https://arxiv.org/abs/2610.08878
作者: Mikołaj Sienicki,Krzysztof Sienicki
类目: Artificial Intelligence (cs.AI); Computers and Society (cs.CY); Human-Computer Interaction (cs.HC)
备注: 83 pages, 20 sections, 35 references

点击查看摘要

Abstract:This article develops a failure-mode framework for analyzing how advanced artificial intelligence could contribute to human extinction, irreversible civilizational collapse, or permanent human disempowerment. The central thesis is that catastrophic AI risk does not require consciousness, hostility, or an explicit intention to harm humanity. Instead, risk may arise through several distinct but interacting pathways, including autonomous misalignment, harmful human use, organizational failure, and competitive deployment. The severity of these pathways depends on factors such as capability, autonomy, external access, persistence, institutional safeguards, and the preservation of recovery capacity. The analysis is deliberately non-operational: it identifies causal conditions, empirically tractable intermediate quantities, and defensive research questions rather than procedures for causing harm.

[HC-32] Intonation Perception in Real and Synthetic Speech across Varying Familiarity Levels: A Pilot Study of Equivalence Assessment INTERSPEECH2026

链接: https://arxiv.org/abs/2610.08839
作者: Hanrui Zhou,Gaoyuan Zhang,Yixiang Chen,Yujie Xing,Feng Xu,Xurong Xie,Hui Chen
类目: Audio and Speech Processing (eess.AS); Human-Computer Interaction (cs.HC); Sound (cs.SD)
备注: Accepted by Interspeech 2026

点击查看摘要

Abstract:Language training relies on a corpus constructed by a large number linguistic materials. AI-powered voice clones provide a way to construct the corpus with relatively low cost. Singing voice conversion (SVC) model is used to generate synthetic voices. This study compares participants’ performances on natural and synthetic speech in two experiments, similarity perception and intonation recognition. In the accuracy of similarity perception task, a significant interaction between speech type and intonation is found, suggesting that question may serve as a cue for speaker identification but may be influenced by synthetic features. In the accuracy of intonation recognition task, a significant interaction between speech type and familiarity is observed, indicating that speech type affects how much familiarity contributes to voice processing.

计算机视觉

[CV-0] ris3D: 3D Scene Generation With Objects That Fit Together

链接: https://arxiv.org/abs/2610.10539
作者: Jaeyeong Kim,Jinhyuk Jang,Jongmin Lee,Kyehong Park,Seungryong Kim
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Project page: this https URL

点击查看摘要

Abstract:We propose Tetris3D, a generative framework for single-image 3D scene reconstruction that recovers objects which are physically and geometrically coherent as a scene. Existing methods often generate objects independently or couple them implicitly, providing limited guidance for ensuring fine-grained spatial compatibility between neighboring objects that interact with one another. To address this, we explicitly condition the generation of each object on the geometry of surrounding objects and their physical relationships, guiding its shape and pose to remain geometrically and physically plausible within the scene. Moreover, we introduce ComOb, a physics simulation-based dataset of 1.2M scenes featuring physical interactions across diverse object categories, with per-object meshes and pairwise physical relation annotations. Comprehensive experiments on synthetic and realworld scenes show that Tetris3D recovers coherent object shapes and poses even when interacting regions are occluded, and achieves state-of-the-art performance in both generation quality and physical stability.

[CV-1] Never Look Back: Understanding Persistence in 3D Object Memory from Egocentric Videos

链接: https://arxiv.org/abs/2610.10538
作者: Shravan Chaudhari,William Paul,Suchi Saria,Rama Chellappa,Homanga Bharadhwaj
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Robotics (cs.RO)
备注: Project page: this https URL . Code: this https URL

点击查看摘要

Abstract:As we move through the world and carry out everyday tasks, we encounter objects that may become relevant only later. We are capable of recalling where we left something or what was inside a container, even without knowing we would need it later. Here, we study how an embodied assistant can build a similar memory from egocentric videos, by observing a person’s day-to-day activities. We present Ledger, a persistent 3D object memory that combines object locations, their histories, and contextual descriptions. It associates observations across the recording and retains objects after they leave the view, including those the person never touches. It clusters each object’s observations by resting locations and records a move only after repeated evidence, reducing the effect of localization noise. Short descriptions preserve details such as an object’s contents or supporting surface. It saves these records to later answer spatial questions without having to access the original images or video. Our memory raises HD-EPIC accuracy from 29.7% to 42.6%, UCS-Bench accuracy from 33.8% to 38.5% and localizes Ego4D objects with a 0.99 m median error on returned predictions. Our analyses identify complementary roles for temporal persistence, contextual descriptions, and retrieval. Our study on 100 stitched streams of multiple scenes each further exposes failures in both retrieval and construction. Per-scene construction partially recovers the performance lost across scene changes compared to that of single scene streams.

[CV-2] Long-WAM: Scaling the Context of World-Action Models

链接: https://arxiv.org/abs/2610.10528
作者: Wei Huang,Bohan Zhang,Chenzhi Liu,Isabella Liu,Shuai Yang,Weian Mao,Luozhou Wang,Yicheng Xiao,Weifeng Lin,Qixin Hu,Bryan Chu,Sifei Liu,Linxi Fan,Xiaojuan Qi,Song Han,Yukang Chen
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Real-time robot control demands enough visual history to infer motion and task progress, but processing that history can delay action. We present Long-WAM, a model-system framework for scaling the context of causal world-action models under real-time control constraints. Our central finding is that access to history is not the same as using it: longer histories pay off far more when the video foundation is pretrained autoregressively (AR). We first learn causal prediction from robot and egocentric videos without action labels, then preserve this history-to-future structure during world-action adaptation. On RoboCasa GR-1, increasing context from 0.0 to 19.2 seconds raises success from 63.3% to 78.7%, whereas a bidirectionally pretrained initialization shows no net gain; robot-domain AR pretraining further raises peak success on GR-1 and LIBERO-Long. Long-WAM also achieves the best results among compared methods on LIBERO-Long, RoboTwin 2.0, and DOMINO. Streaming observation encoding, asynchronous execution, and hardware-specific acceleration enable deployment on RTX 5090, DGX Spark, and Jetson AGX Thor without dropping future prediction; on RTX 5090, each action chunk, including future-video latent prediction, takes 107.4 ms. Real-time deployment on Unitree G1 and YAM supports dynamic and long-horizon manipulation, including 95% success on dynamic cup stacking, where Pi0.5 and Fast-WAM succeed in none of 20 trials. As a memory-informed executor, Long-WAM also complements higher-level planning in composite tasks.

[CV-3] GRACE: Generation-aware latent compression for efficient video generation

链接: https://arxiv.org/abs/2610.10524
作者: Jiyoung Kim,Paul Hyunbin Cho,Jisu Nam,Donghoon Lee,Hyunsung Go,Yeonkyeong Lee,Hansaem Kim,Seungryong Kim
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Project page : this https URL , 43 pages, 24 figures

点击查看摘要

Abstract:Highly compressed video autoencoders offer an effective way to accelerate video diffusion models, as the Diffusion Transformer (DiT) operates on far fewer tokens. However, such autoencoders are challenging to train, since a higher compression ratio degrades reconstruction quality and recovering it requires more channels, which is known to slow the convergence of the DiT. The compressed latent also differs from the one the DiT was trained on, so the pretrained DiT must be either retrained from scratch or adapted at considerable cost. Compressing the autoencoder the DiT was trained with appears to preserve compatibility, yet optimizing it for reconstruction alone still shifts the latent away from the distribution the DiT has learned. To address this, we propose Generation-Aware Latent Compression for Efficient Video Generation (GRACE), a two-stage framework that compresses a pretrained video autoencoder while keeping it compatible with the pretrained DiT. Specifically, we keep a frozen base latent from the pretrained encoder and learn a residual latent for the information lost under stronger compression, while aligning the compressed latent with the pretrained latent in the feature space of the frozen DiT so that the autoencoder is optimized for generation. We then adapt the DiT with lightweight fine-tuning and asymmetric denoising, where the base is denoised ahead of the residual. GRACE reduces the token count of Wan2.1-I2V-14B by 8x and its latency by 11.1x at 480x832x81, while matching the generation quality of the pretrained pipeline before compression on VBench.

[CV-4] Video-Conditioned Generative Joint 2D-3D Hand Motion Recovery

链接: https://arxiv.org/abs/2610.10512
作者: Chen Xu,Yunqi Li,Binbin Huang,Brent Yi,Shenghua Gao,Yi Ma
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 20 pages, 6 figures

点击查看摘要

Abstract:Recovering faithful 3D hand motion from video remains challenging due to frequent occlusions and incomplete visual observations, which make frame-wise pose estimates unreliable and temporally inconsistent. To address this problem, we propose JoHan, a unified generative framework that recovers hand motion directly from video sequences without relying on intermediate per-frame pose predictions. Trained from scratch, our model jointly generates aligned 2D and 3D local hand pose sequences by learning their temporal dynamics and cross-representation correspondence. The generated 2D trajectories exploit direct spatial and temporal cues from the 2D images to guide the following generative 3D motion reconstruction, while the learned motion prior promotes temporal consistency. Their learned 2D-3D correspondence further enables recovery of the hand’s global position and orientation relative to the camera. Extensive experiments on challenging benchmarks demonstrate significantly improved accuracy and speed in local hand-pose and camera-space reconstruction. Notably, our method captures much better hand-motion dynamics, producing significantly smoother motion than previous methods while maintaining high per-frame pose accuracy.

[CV-5] QuadTok: Quadtree Visual Tokenizer for Autoregressive Image Generation

链接: https://arxiv.org/abs/2610.10497
作者: Yucheng Mao,Zeyuan Chen,Xiaojun Shan,Xiang Zhang,Divyansh Srivastava,Bingnan Li,Zhuowen Tu
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:We introduce QuadTok, a novel framework for visual tokenization and autoregressive image generation. Compared to traditional approaches using 2D grids or 1D token sequences, we propose a hierarchical quadtree structure, bridging the gap between 2D spatial binding and 1D sequence-level flexibility. The QuadTok tokenizer dynamically allocates representational capacity to visually intricate areas while leaving homogeneous regions at a coarse resolution. Compared with a fixed 256-token grid, our ImageNet-trained tokenizer saves approximately 10% of tokens on ImageNet and 9% when transferred zero-shot to the COCO dataset, while maintaining comparable reconstruction fidelity. Furthermore, the natural causality introduced by the tree structure seamlessly enables autoregressive image generation. Conditioned on a quadtree topology supplied before generation, our 947M GPT-style generative model achieves a 2.08 gFID on the ImageNet 256 \times 256 benchmark. Additionally, leveraging the strong spatial correlation preserved by the quadtree structure, the QuadTok generator enables zero-shot spatially controlled image generation capabilities. Code: this https URL.

[CV-6] Insights from Autoresearch for Solar Panel Segmentation

链接: https://arxiv.org/abs/2610.10491
作者: Justinas Lekavicius,Kursat Komurcu,Valentas Gruzauskas,Linas Petkevicius
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted at AutoML4EO 2026 (non-archival AutoML conference workshop). 4 pages + references. this https URL

点击查看摘要

Abstract:This paper investigates AutoResearch, a protocol in which a coding language model edits a training program under a one-hour GPU budget and retains a change only if validation IoU improves. The protocol is applied to photovoltaic panel segmentation on a frozen real-image split, with DeepLabV3–ResNet-50 held fixed. Three campaigns of 24 experiments, using Gemma~4 12B, Qwen3-8B all improve their one-hour baselines, but retained modifications do not transfer across hardware. The Qwen3-8B configuration, trained on real images only, reaches a test IoU of 0.836 versus 0.833 for the reference GAN-augmented schedule. Research repository this https URL.

[CV-7] Agent ic RSR: Real-to-Sim-to-Real through Scene Reconstruction and Execution-Grounded Robot Policies

链接: https://arxiv.org/abs/2610.10479
作者: Yihan Li,Yating Feng,Shengjiu Sun,Jianing Chen,Hao Ren,Bowen Yang,Weisheng Xu,Qiwei Wu,Hui Cheng,Renjing Xu
类目: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
备注: 25 pages including appendices, 5 figures

点击查看摘要

Abstract:A simulation of a real robot workspace must preserve task-relevant interactions, while policies developed in it must operate on observations available to the real robot. Yet scene reconstruction and policy development are often treated separately. We present Agentic Real-to-Sim-to-Real (Agentic RSR), a framework that links scene reconstruction, policy development, and real-robot execution through the same manipulation task. Given a workspace video, a task description, and a known robot model, an agent recovers metric scale, iteratively refines the scene using visual feedback, and checks task-relevant interactions in MuJoCo. A coding agent then develops an executable policy, progressing from privileged object poses to visual observations and randomized simulation. The policy can interleave multiple observations and actions within one invocation, while the agent uses execution feedback to continue, retry, or revise its approach. A shared task-level interface carries the policy and accumulated experience to the real robot, where fresh observations and safety checks guide execution. Across 18 reconstructed scenes involving two robots, the mean four-view Depth MAE against reference depth estimates is 0.1057 m, the mean Lab \Delta E_76 is 11.04, and the mean grayscale SSIM is 0.6990. In real-robot experiments, the aggregate task success rate reaches 80% of the simulation task success rate, indicating substantial retention of simulated performance on hardware. Code and reconstructed scene data will be made publicly available.

[CV-8] Label-free cell counting and viability prediction with brightfield imaging and deep learning

链接: https://arxiv.org/abs/2610.10473
作者: Amir Reza Vazifeh,Christian Zeigler,Sornanathan Meyyappan,Richard Jeske,Jason W. Fleischer
类目: Computer Vision and Pattern Recognition (cs.CV); Cell Behavior (q-bio.CB)
备注:

点击查看摘要

Abstract:Cell viability assessment is a core requirement in cell culture systems, with critical applications in biopharmaceutical manufacturing and drug development. Conventionally, it is measured by adding membrane-impermeable dyes to a sample (a process called staining), which allows compromised cell membranes to be distinguished from intact ones. However, staining has several limitations: (a) chemical agents can perturb normal cellular processes of the cells being measured, (b) it is often ambiguous to assign viability to individual cells whose membrane integrity is only partially compromised. © photobleaching can undermine measurement accuracy over time when using fluorescent stains, and (d) staining cannot be performed in situ or in real time. Here, we show that (1) stained cells captured under brightfield imaging contain sufficient information to distinguish live and dead cells, and (2) cells captured under unstained brightfield imaging exhibit similar image features to their stained counterparts, enabling models trained on stained cells to generalize to unstained ones. We then report the development and validation of ViabiLens, an AI-assisted software for label-free cell viability analysis. The ViabiLens combines a cell detection model for localizing individual cells with a convolutional neural network (CNN) classifier for live/dead prediction, paired with an interactive UMAP-based viewer for visualizing and exploring individual cells across the sample. Evaluated on Chinese Hamster Ovary (CHO) cells spanning a wide range of viability conditions, ViabiLens achieves a mean absolute error of 2.68% on unstained samples against fluorescence-based reference measurements. We also release a benchmark dataset for label-free cell viability analysis to facilitate future research, available at this https URL.

[CV-9] MORCA: Offline-to-Online Reinforcement Learning for Adaptive Cache Reuse in Video Diffusion Acceleration

链接: https://arxiv.org/abs/2610.10457
作者: Yuxiang Xiong,Ruiyan Wang,Wenqiang Wang,Teng Hu,Songhang Shen,Bohao Feng,Hongqian Deng,Ran Yi
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 22 pages, 8 figures

点击查看摘要

Abstract:Diffusion Transformers (DiTs) achieve remarkable performance in video synthesis, but their iterative denoising process suffers from high inference latency. To address this, caching has emerged as an effective acceleration strategy by capitalizing on inter-step redundancy during denoising. Existing dynamic caching methods typically estimate the error that cache reuse would introduce at each denoising step (step error) to guide cache decisions, whereas our concern is how much quality loss cache reuse would cause in the final generated video (terminal error). We show that step error does not directly correspond to terminal error and that latent information helps capture their relationship, thereby informing cache decisions. Moreover, existing threshold-based methods cannot provide precise speedup control, making it difficult to meet practical requirements for user-specified acceleration targets. To address these limitations, we introduce MORCA, a cache scheduling framework trained through offline-to-online reinforcement learning to make latent-aware reuse/recompute decisions under user-specified acceleration targets. Extensive experiments on different video generation models across multiple target acceleration ratios demonstrate that MORCA achieves better generation fidelity than state-of-the-art caching methods under comparable computational budgets. Code is available at this https URL.

[CV-10] ECHO: Embodied Camera Observations of Human Object Carrying

链接: https://arxiv.org/abs/2610.10438
作者: Xuefei Sun,Lorin Achey,Kali Hamilton,Alberto Speranzon,Gregory Grebe,Yonatan Bisk,Christoffer Heckman
类目: Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)
备注:

点击查看摘要

Abstract:Embodied and assistive agents must do more than recognize objects: they must reason about where an object belongs given the layout of an environment and the habits of the people who live in it. Progress on this problem has been limited, in part because no dedicated benchmark or dataset exists to define and evaluate it. Existing RGB-D scan datasets reconstruct static rooms without human activity, while human-object-interaction datasets capture motion without a navigable, fully reconstructed scene or a ground-truth notion of an object’s natural destination. We introduce contextual object placement as a benchmark task: predicting an object’s destination during an observed object-carrying episode. To support this task, we present Embodied Camera observations of Human Object carrying (ECHO), a large-scale synthetic dataset that pairs dense RGB-D scans of indoor scenes with recordings of an embodied human carrying everyday objects to context-appropriate destinations. ECHO is the first publicly available dataset to combine reconstructed scenes, human activity, natural language, and contextual-placement annotations. It comprises 3,805 human-annotated episodes across 159 floors of 115 HM3D scenes, involving 198 distinct objects. Each floor includes a complete RGB-D scan with human-annotated room labels and a surface list. Each episode provides synchronized RGB-D encounter clips; 6-DoF camera, human, and object trajectories; start and destination surfaces; an action caption; and a human-written context: a single sentence describing the inhabitant’s routine that implies the destination without naming it. We evaluate contextual object placement using input-masked probes and an end-to-end baseline. Results show that no single input modality is sufficient, highlighting the need to jointly reason over scene structure, human activity, and contextual knowledge.

[CV-11] Detecting Adversarial Images through Response Profiles of Vision-Language Models

链接: https://arxiv.org/abs/2610.10436
作者: Arash Vashagh,Roozbeh Razavi-Far
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Adversarial perturbations can alter the predictions of frozen vision-language models (VLMs) while leaving their confidence and image–text similarity patterns seemingly plausible. We investigate whether we can identify adversarial inputs based on the broader way an image interacts with a collection of general semantic prompts. Our detector summarizes these responses using category-level statistics, relationships among prompts, deviations from clean reference distributions, and stability under weak image transformations, producing a compact response profile that is classified by a lightweight model while the VLM remains fixed. We evaluate the approach on multiple public image datasets, several CLIP-style visual backbones, and a range of gradient-based, optimization-based, automated, and spatial attacks. The detector achieves strong discrimination in attack-specific settings and retains substantial performance when evaluated on attacks not seen during training. Under a controlled detector-specific protocol, the response-profile representation outperforms the evaluated embedding-geometry baselines. Additional analyses show that the feature groups provide complementary information and that the method remains effective under variations in the prompt configuration. We also examine inference cost and performance against detector-aware adaptive attacks. Overall, the results indicate that response patterns across semantic prompts provide a useful complementary signal for adversarial image detection in frozen VLMs.

[CV-12] SGF: Decoupling Gradient Flows for Autoregressive Video Generation

链接: https://arxiv.org/abs/2610.10429
作者: Zihan Su,Junhao Zhuang,Yaowei Li,Siwen Lu,Haoran Li,Lingen Li,Haoyu Wu,Weiyang Jin,Songchun Zhang,Haoyang Huang,Chun Yuan,Zeyue Xue,Nan Duan
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Autoregressive video generation requires denoising the current frames while writing their key-value representations as context for future predictions. However, these two roles typically share parameters, and we find that their gradients exhibit distinct patterns and systematic negative alignment, hindering the joint optimization of visual quality and temporal consistency. We introduce Self Gradient Forcing Plus (SGF+), which assigns separate parameters to context writing and denoising while preserving their interaction through causal attention. Both roles are jointly optimized using the original generation objective without auxiliary losses, with context writing supervised through its contribution to future predictions. This simple change improves visual quality and long-horizon consistency over the evaluated baselines in both framewise and chunkwise generation, without additional video training data or a longer training horizon. Trained on only 5s rollouts, SGF+ supports continuous generation for up to 24 hours without long-video fine-tuning. These results highlight role-specific parameterization as an effective design principle for high-quality autoregressive video generation and native long-horizon extrapolation.

[CV-13] GraphRectify: Graph-Based Transfer of Adversarial Example Detectors Across Neural Networks

链接: https://arxiv.org/abs/2610.10423
作者: Arash Vashagh,Roozbeh Razavi-Far
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Adversarial example detectors are often tied to the classifier backbone they were trained on, limiting reuse when the protected model is replaced or upgraded. Directly transferring such detectors across backbones is challenging because different networks generally produce incompatible internal representations. We propose GraphRectify, a graph-based framework for transferring adversarial image detectors across classifier backbones. GraphRectify learns a structured representation of intermediate classifier features and adapts representations from a new backbone to the detector learned on the original model, enabling detector reuse. We evaluate GraphRectify across multiple datasets, backbone architectures, and adversarial attacks, including detector-aware adaptive attacks that jointly target the classifier and detector. Across the complete evaluation matrix, GraphRectify achieves higher aggregate ROC-AUC than training a detector from scratch on the new backbone and the evaluated transfer ablations. The gains are particularly strong for transfers between different backbone families and when sufficient data are available. In contrast, training from scratch remains competitive in the most data-limited settings. These results show that adversarial detection knowledge can transfer effectively across heterogeneous classifier architectures rather than being relearned whenever the protected backbone changes.

[CV-14] Rubix: Global Correspondence-Free Point Set Alignment through Assignment Geometry

链接: https://arxiv.org/abs/2610.10408
作者: Subhransu S. Bhattacharjee,Dylan Campbell,Rahul Shome
类目: Computer Vision and Pattern Recognition (cs.CV); Computational Geometry (cs.CG); Machine Learning (cs.LG); Robotics (cs.RO); Optimization and Control (math.OC)
备注: 67 pages, 20 figures. Includes full proofs and experimental appendices

点击查看摘要

Abstract:Procrustes-Wasserstein alignment jointly estimates a matching and rotation without supplied correspondences, but alternating minimization can stop at suboptimal solutions. Rubix solves the equally weighted planar problem globally under squared Euclidean loss. Each matching \sigma of two centered n -point sets defines a complex correlation z_\sigma=\sum_i\bar x_i y_\sigma(i) . Their convex hull is the permutation polygon: supporting vertices give optimal matchings at fixed rotations, and the farthest vertex gives the global alignment. We prove the sharp bound of n(n-1) vertices for n\ge2 , answering Rote’s rotation-assignment open problem. In exact arithmetic, assignment queries recover the polygon in \mathcal O(n^5) operations. Assignment-based bounds extend the approach to three-dimensional rotations and partial matching at a supplied translation through branch-and-bound. On timed MPEG-7 shape pairs, Rubix attains every numerical reference value in 12 ms on average, 50 times faster than a rotation grid at the same accuracy. Its distances improve gravity-aligned matching of real 3D scans, shape retrieval and noisy crystal classification over alternating minimization.

[CV-15] Self-correction Optimization for Interleaved Multimodal Generation

链接: https://arxiv.org/abs/2610.10400
作者: Xin You,Zhiwei Ning,Zukai Chen,Minghui Zhang,Xuanke Shi,Hanxiao Zhang,Jingsong Liu,Jie Yang,Quan Wang,Yun Gu
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 20 pages, 10 figures

点击查看摘要

Abstract:Multimodal large language models (MLLMs) have made significant progress in visual understanding and generation. However, generating interleaved image–text content remains challenging, as it requires tightly integrated multimodal understanding and generation capabilities. Although existing MLLMs provide promising solutions, most rely on additional training with augmented data, which is computationally expensive and remains limited in preserving visual subjects, temporal consistency, and physical plausibility. In this work, we propose self-correction optimization (SCO), an effective training-free method for consistent interleaved generation. SCO treats the classifier-free guidance update as a reference and performs minimal self-correction under two complementary constraints, including new-event and state-preserving constraints. Specifically, the new-event constraint promotes temporal consistency across image–text sequences, while the state-preserving constraint maintains the coherence of visual subjects throughout subsequent generation steps. Experiments on challenging interleaved multimodal generation benchmarks demonstrate significant improvements in temporal coherence and visual-subject preservation. Furthermore, SCO can be extended to video generation and improves the modeling of physically grounded processes, including robot manipulation and long-horizon handcrafting.

[CV-16] Gaussian Density Splatting Network NEURIPS

链接: https://arxiv.org/abs/2610.10396
作者: Miao Shang,Yabin Wang,Xiaopeng Hong
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: This is the preprint version of the paper and supplemental material to appear in NeurIPS, 2026. Please cite the final published version

点击查看摘要

Abstract:This paper proposes a novel crowd counting approach, the Gaussian Density Splatting Network (GDSNet). Unlike methods that rely on conventional, grid-based density maps and are sensitive to spatial resolution, GDSNet represents a crowd as a superposition of continuous 2D Gaussian primitives. Our approach is built upon two key contributions. First, we introduce a control-point-based fitting mechanism to structure the prediction of the Gaussian parameters. We design a method to allocate a set of control points that define local regions, from which features are pooled to regress each primitive’s parameters. Second, we adapt a differentiable Gaussian Splatting framework to the counting task by parameterizing each primitive with geometric parameters and a scalar density mass. This formulation allows the network to be trained end-to-end via spatial matching of differentiably rendered density maps, naturally providing both local density supervision and global count optimization. Extensive evaluations on four standard benchmarks show GDSNet consistently outperforms the state of the art.

[CV-17] MOTIP2: Spatial Priors for End-to-End Multi-Object Tracking BMVC2026

链接: https://arxiv.org/abs/2610.10391
作者: Benoît Roussel,Damien Bouet,Liming Chen,Pierre Perrault
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted at BMVC 2026. 27 pages (14 pages main paper, appendix and references), 6 figures, 10 tables

点击查看摘要

Abstract:End-to-end multi-object trackers have narrowed the gap with classical tracking-by-detection on association-difficult benchmarks. Yet they still make spatially implausible errors no classical tracker would, such as assigning one identity to objects on opposite sides of the frame. A model could learn to avoid them, but tracking annotations are scarce, so we encode spatial priors explicitly instead, while keeping inference fully end-to-end with no post-hoc association. We propose three spatial priors, at the data, loss, and representation stages. Spatial ID Switches bias trajectory permutations toward spatially overlapping objects, reducing the mismatch between training and inference confusions. Spatial ID Loss scales each identity’s penalty by its box distance, so a distant switch costs more than a nearby one. Spatial Anchor gives each track token its frame position, an explicit spatial cue for attention. We instantiate the three priors in MOTIP2, a tracker adapted from MOTIP and built on the real-time DEIM detection transformer. Trained without extra data, its main model, MOTIP2-L, sets a new state of the art: 73.4 HOTA on DanceTrack, 76.0 on SportsMOT, and 71.1 IDF1 on PersonPath22. MOTIP2 is a family of models spanning the speed-accuracy trade-off: a lighter model, MOTIP2-S, matches the original MOTIP at over 3x the speed, and MOTIP2-X reaches 74.8 HOTA on DanceTrack. Comments: Accepted at BMVC 2026. 27 pages (14 pages main paper, appendix and references), 6 figures, 10 tables Subjects: Computer Vision and Pattern Recognition (cs.CV) Cite as: arXiv:2610.10391 [cs.CV] (or arXiv:2610.10391v1 [cs.CV] for this version) https://doi.org/10.48550/arXiv.2610.10391 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[CV-18] Explicit Geometric Chain-of-Thought for Vision-Language-Action in Autonomous Driving

链接: https://arxiv.org/abs/2610.10390
作者: Xingtai Gui,Yucheng Zhou,Dongqian Guo,Jiahao Gong,Feiyang Tan,Jianbing Shen
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 21 pages, 9 figures. The code is available at this https URL

点击查看摘要

Abstract:Vision-language-action~(VLA) models have emerged as a promising paradigm for autonomous driving. However, existing VLA models still suffer from a fundamental mismatch: driving actions require precise 3D geometric cues, while visual-language understanding and reasoning are largely conducted in a 2D semantic space. In this paper, we propose GeoCoTDrive, an explicit geometric chain-of-thought framework that grounds geometry in a planning-oriented manner. GeoCoTDrive follows a think with 2D first, drive with dedicated 3D priors paradigm. It first grounds 2D regions corresponding to decision-critical cues, and then retrieves localized 3D priors by sampling features from a geometric foundation model within the grounded regions. These localized geometric features are interleaved into the autoregressive context to support the trajectory generation. To supervise this process, we introduce planning-relevant grounding, a new region-level grounding task that focuses on local spatial cues directly affecting ego planning decisions, and construct the PlanningGrounding dataset to endow VLAs with planning-oriented grounding capability. Experiments across multiple end-to-end autonomous driving benchmarks show that GeoCoTDrive consistently improves safety-critical planning performance, demonstrating the effectiveness of the explicit geometric chain-of-thought process for VLA-based planning.

[CV-19] RoboQuest: Generalist Physical Agents that Search Inspect and Test

链接: https://arxiv.org/abs/2610.10388
作者: Liu Renhang,Navonil Majumder,Tej Deep Pala,Soujanya Poria
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Recent advances in multimodal foundation models have made them capable generalist physical agents for a range of manipulation tasks. However, successful operation in an unfamiliar environment may require an agent to seek task-relevant information through interaction when it is absent from the observations: it may need to determine where a relevant object is, inspect an unobserved property, or discover the effect of an unfamiliar tool. We thus introduce RoboQuest, a benchmark for goal-directed embodied exploration, where agents must actively acquire task-relevant information through physical interaction, use the resulting evidence to adapt subsequent actions, and autonomously decide when to commit to task completion. RoboQuest comprises ten mobile manipulation tasks centered on three forms of uncertainty: search, manipulation-based inspection, and interactive testing. We evaluate five frontier multimodal agents through a common visuomotor interface, as well as a \pi_0.5 policy fine-tuned on the full-episode demonstrations we release. The best agent succeeds in only 23% of the episodes, and the fine-tuned policy almost never succeeds. Isolated tests of the execution skills the tasks are built from, with the hidden information supplied, show that the agents can carry out most of the required actions, and our failure analysis attributes only a minority of the failures to execution. Our failure analysis further finds that the agents often stop exploring too early as they make decisions before observing the required evidence for task completion. We also find that agents rarely prevent or repair the disturbances caused by their exploration. Moreover, learning by trial and error remains difficult for most models.

[CV-20] MultiFly: A Real-World Multimodal Aerial Dataset with Annotation-Efficient Label Transfer and Cross-Modal Semantic Consistency

链接: https://arxiv.org/abs/2610.10359
作者: Markus Gross,Andreas Greiner,Taehyoung Kim,Sivasubiramaniam Subbiah,Tomaž Cotič,Sai Bharadwaj Matha,Conrad Christoph,Oussema Dhaouadi,Simon Zieher,Surya Vijaya Kumar,Gordon Elger,Henri Meeß,Olaf Wysocki,Paul Spannaus,Daniel Cremers
类目: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:We introduce MultiFly, a real-world, low-altitude UAV dataset for semantic perception across RGB, thermal, LiDAR, and radar modalities. MultiFly provides 17,272 synchronized samples from four suburban scenes with frame-wise annotations for 15 semantic classes, together with calibration and GNSS-RTK/IMU measurements. To avoid costly and inconsistent modality-specific annotation, we propagate labels from only 115 manually annotated RGB images through shared geometric representations to all four modalities. This approach generates semantic labels for 17,157 additional RGB images, 17,272 thermal images, 840M LiDAR points, and 3.4M radar points. Transferred annotations achieve 89.93% average agreement with held-out manual annotations, and 90.94% average semantic consistency across all six modality pairs. We further establish semantic segmentation benchmarks for all four modalities, revealing distinct architectural behavior for dense LiDAR and sparse radar data. Taken together, MultiFly provides a scalable foundation for multimodal aerial perception and, to the best of our knowledge, the first public real-world low-altitude aerial benchmark that combines consistent frame-wise semantic annotations for RGB, thermal, LiDAR, and radar. Data at this https URL.

[CV-21] Real-Time Joint Audio-Video Generation by Parallel Adapter Composition

链接: https://arxiv.org/abs/2610.10343
作者: Jingyu Li,Xiaoxiao Xiang,Yiwen Guo
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Deploying a joint audio-video diffusion transformer for real-time, interactive generation normally requires two essential modifications: block-autoregressive attention, so frames can be emitted before the whole clip is finished, and few-step sampling, so each block is cheap. Conventionally, the streaming video literature obtains both capabilities from a chained pipeline. It first distills a bidirectional teacher into a causal student, then into a few-step one, or proceeds in reverse order. Each stage of such a chain fine-tunes the weights the previous one produced, so a later objective can undo an earlier capability. Following the idea of model merging, we show that on a packed audio-video backbone the two capabilities can be acquired in parallel. A causal adapter is trained against the frozen backbone, and an off-the-shelf few-step adapter provides the few-step capability. As the two edit different functional axes, we predict, and then verify, that their weight-update directions are near-orthogonal, without any explicit orthogonality constraint during training. Orthogonal updates should combine without interfering, so parallel composition is a direct sum. The two adapters are simply added at inference, with no joint training, yielding few-step, streaming audio-video whose image quality tracks the bidirectional teacher. Compared to the chained baselines, the composed model matches or beats them on most metrics, making parallel composition a practical approach. The resulting streaming system generates joint audio-video in real time, \approx 26 fps at 480\times832 without quantization, and sustains 30 s of continuous generation with stable image quality.

[CV-22] Position Forcing: Self-Conditioning 3D Generation

链接: https://arxiv.org/abs/2610.10342
作者: Ziheng Ouyang,Zeqiang Lai,Jiarui Chen,Jiangshan Wang,Yuhao Wan,Jingbo Gong,Xiangyu Yue,Hengshuang Zhao,Qibin Hou,Chunchao Guo
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Recent single-stage 3D generative models commonly adopt VecSet representations, encoding 3D shapes as unordered sets of latent tokens. However, compared with two-stage methods that provide explicit positional guidance, these models must implicitly infer token positions throughout denoising, limiting their generation quality. We observe that, despite the absence of explicit positional conditioning, VecSet tokens retain recoverable spatial correspondences. Building on this observation, we propose Position Forcing, a position-based self-conditioning framework. During denoising, Position Forcing recovers token positions from the current clean latent estimate, quantizes them at progressively finer resolutions according to the denoising stage, and feeds the resulting positional encodings back into the diffusion Transformer. This progressively refined positional feedback provides spatial guidance at a granularity appropriate to each denoising stage, guiding shape generation along a coarse-to-fine trajectory and substantially improving generation quality without a separate position generation stage. Experiments demonstrate that Position Forcing achieves strong performance among single-stage 3D generative methods and outperforms several competitive multi-stage approaches.

[CV-23] When to Unpair: Regulating Pairing Dependence in Medical Visual In-Context Learning

链接: https://arxiv.org/abs/2610.10335
作者: Cheng Wan,Chenjun Li,Qingyu Zhao
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 24 pages, 12 figures

点击查看摘要

Abstract:Visual in-context learning (ICL), well suited to label-scarce medical imaging, uses support image-label pairs to demonstrate input-output mappings, while the labels collectively indicate the requested task. We diagnose dependence on individual pairings with a test-time derangement that reassigns every support label to another support image while preserving the query, support images, and label multiset. The resulting pairing gap, defined as shuffled-minus-matched performance, shows that all four released models depend on the pairing, to widely varying degrees. Further analysis of a paired-trained model reveals support-associated spurious regions and lesion-size biases even with real, unaltered supports, alongside sensitivity to mis-registered support labels. To regulate this dependence, we introduce a late unpairing curriculum (LUC), which starts with matched training and then applies random unpairing, replacing each support label with that of another support in the same episode. LUC nearly closes the pairing gap on two backbones while maintaining or improving matched-support performance across all evaluated task types, with gains extending to held-out tasks and cross-dataset episodes. It also mitigates these failure modes. On BraTS whole-tumor segmentation, matched-support DSC rises from 0.733 to 0.857 while the gap shrinks from -0.184 to -0.008. In a released model, brief fine-tuning with random unpairing reduces the gap. A reversed curriculum that places the same number of unpairing epochs at the start of training leaves a large gap. This shows that pairing dependence is shaped by the order of training and not only by the amount of unpaired training.

[CV-24] How Private is Private? A Comparative Study for Face De-Identification NEURIPS2026

链接: https://arxiv.org/abs/2610.10334
作者: Hui Wei,Hao Yu,Hui Kuurila-Zhang,Guoying Zhao
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted to NeurIPS 2026. Project Page: this https URL

点击查看摘要

Abstract:Face de-identification (FDeID) has emerged as a critical privacy-preserving technology, yet its evaluation remains fundamentally fragmented. Existing protocols rely on inconsistent metrics, heterogeneous datasets, and partial annotation coverage, so methods targeting different utility dimensions, such as landmark versus expression preservation, are reported on different benchmarks under different metrics, rendering cross-method comparison infeasible. We revisit FDeID evaluation from both the data and metric perspectives. On the data side, we introduce UtilFace, a curated, demographically balanced benchmark with high identity diversity, assembled from four large-scale face datasets through identity-aware cleaning, resolution enhancement, and stratified filtering. On the metric side, we propose HiFD, a Hierarchical Face De-identification metric that unifies identity suppression, multi-level utility preservation, and image quality under a single consistency-based paradigm: every component is computed from pretrained estimators’ outputs on the original face and its de-identified counterpart, directly quantifying how much identity is suppressed and how much downstream-perceivable utility survives. HiFD organizes facial signals into a three-level utility hierarchy spanning macro cues (L1), micro cues (L2), and imperceptible cues (L3), and aggregates the five resulting components into a single interpretable score via weighted harmonic mean, with configurable application-specific profiles. Using this unified protocol, we conduct a comprehensive comparative study spanning adversarial, GAN-based, and diffusion-based methods, surfacing trade-offs and failure modes that remain invisible under existing protocols. We release the benchmark and evaluation toolkit to foster systematic and reproducible research in privacy-preserving human face analysis.

[CV-25] Performance at What Cost? A Sustainability-Aware Performance Index for Cell and Nucleus Instance Segmentation

链接: https://arxiv.org/abs/2610.10324
作者: Eiram Mahera Sheikh,Alaa Tharwat,Wolfram Schenck
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Pretrained models for cell and nuclear instance segmentation differ substantially in architecture, pretraining data and objectives, parameter count, inference strategy, adaptation requirements, postprocessing pipeline, and computational demand. Large pretrained and foundation models are increasingly adopted because of their strong zero-shot capabilities, but their use also imposes greater energy consumption, memory requirements, computational demands, adaptation costs, and operational carbon emissions. Whether these additional demands are justified by meaningful gains in segmentation performance remains unclear. We address this question by introducing the Sustainability-Aware Performance Index (SAPI), a configurable metric that combines segmentation performance, energy consumption, and model size. We benchmark 19 pretrained and foundation models across six CellBinDB datasets under zero-shot inference and evaluate 16 fine-tunable models using few-shot adaptation with both frozen encoder and full-model fine-tuning. We estimate energy consumption for GPU, CPU, and RAM using software-based monitoring tools. Our results show that larger and more computationally demanding models do not consistently achieve proportionate improvements in segmentation quality. While few-shot adaptation benefits several models, the gains and resource costs vary considerably across architectures, datasets, and adaptation strategies, causing SAPI-based rankings to differ from rankings based on performance alone. This study provides a practical framework for comparing segmentation models more comprehensively and supports more computationally accessible and environmentally responsible model selection in biomedical image analysis.

[CV-26] From Digital Human Interactions to Physics-Based Humanoid Skills: Physics-Grounded Post-Training of Interaction Generators

链接: https://arxiv.org/abs/2610.10322
作者: Kerui Chen,Jianrong Zhang,Kai Lv,Hehe Fan
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Recent methods have made promising progress in generating interactions between two humanoids, largely relying on physics-based tracking policies to convert digital reference motions into executable trajectories. However, limited tracking capabilities restrict the range of reference motions that can be successfully executed, reducing data utilization. Moreover, even successful tracking does not guarantee physically plausible responses or faithful realization of the intended interactions. In this paper, we introduce DIGHT, a co-adaptive framework that couples a Digital human Interaction Generator with a Humanoid Tracking policy. Our DIGHT first executes multiple text-conditioned interaction candidates in simulation using a fixed tracker. It then constructs physics-grounded preferences from the resulting rollouts, covering both general executability and interaction fidelity. Rather than collapsing these signals into a single scalar reward for candidate ranking, we align the pretrained generator using physics-decoupled diffusion direct preference optimization (DPO), preserving criterion-specific supervision without differentiating through the simulator. To improve executability, preference pairs are derived from tracking error, friction, and floating. Additionally, to improve interaction fidelity, we propose to incorporate force feedback from simulator as a measure of contact fidelity and construct preferences over contact occurrence, location, duration, and force magnitude. The aligned generator then supplies reference motions for fine-tuning the tracker, improving compatibility between generation and physical execution. Extensive experiments demonstrate that our approach not only improves the physical plausibility of generated motions but also enables more reliable and faithful humanoid interactions in simulation.

[CV-27] One-Shot Adaptive Segmentation For Scientific Images

链接: https://arxiv.org/abs/2610.10306
作者: Tejaswi V. Panchagnula,Allison M. Davis,Fengqing Zhu
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Scientific image segmentation methods rely on extensive annotation and task-specific training, limiting adaptation across imaging modalities and experimental conditions. We present a training-free, one-shot framework that specializes vision foundation models using a single annotated reference image. The framework combines DINOv3 representations with background-adaptive feature orthogonalization to suppress artifact-related feature directions, after which cosine similarity localizes candidate regions for SAM segmentation. We evaluate the framework on red-blood-cell microscopy, structured-illumination pool boiling, and chest radiography. Relative to the strongest baseline, the proposed method improves mean IoU by 5.91% and 78.62% on the microscopy and pool-boiling datasets, respectively, while achieving comparable performance on chest radiographs. These results demonstrate that one-shot reference conditioning can adapt general-purpose vision models to specialized scientific segmentation tasks.

[CV-28] On the Necessity of Attention-FFN Split in Vision Transformers

链接: https://arxiv.org/abs/2610.10303
作者: Junhyeok Kim,Jinyeong Kim,Jae Wan Park,Seong Jae Hwang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:The standard Transformer architecture relies on a rigid pattern that alternates Attention and Feed-Forward Network (FFN) layers. Despite its widespread adoption, the inductive bias imposed by this strict separation has not been systematically examined. In this work, we investigate the necessity of the Attention-FFN dichotomy in Vision Transformers (ViTs). To facilitate this analysis, we introduce the AttenFeed module, a unified component that integrates the functional properties of both Attention and FFN. Based on this module, we devise the unified Vision Transformer (uViT), which replaces the conventional alternating Attention-FFN structure with a sequence of AttenFeed modules. We then use uViT as a control group that relaxes the Attention-FFN dichotomy of the standard ViT and systematically compare the two models across multiple datasets and model scales. Our experiments reveal that the Attention-FFN dichotomy can hinder performance at smaller model scales due to the rigid parameter allocation of ViTs. The AttenFeed module and uViT serve as new analytical tools for understanding the Attention-FFN structure and offer theoretical insights into the heuristically designed architecture of conventional ViTs.

[CV-29] ouchScale: 500 Hours of Human Vision and Touch for Visual-Tactile Learning

链接: https://arxiv.org/abs/2610.10288
作者: Dayou Li,Hao Wang,Qianqian Yang,Zihao Zhu,Haoquan Fang,Ziyao Zeng,Yan Han,Zihan Wang,Yan Wang,Baoru Huang,Dilin Wang,Kenji Shimada,Yiyue Luo,Manling Li,Teresa Lv,Mustafa Mukadam,Rakesh Ranjan,Ruohan Zhang,Qi He,Changliu Liu,Xu Chen,Marco Pavone,Bangya Liu,Jiachen Li,Masayoshi Tomizuka,Zhiwen Fan
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Project page: this https URL

点击查看摘要

Abstract:Large-scale egocentric human interaction data is becoming an important source of physical supervision for embodied learning, yet video alone leaves the contact and pressure that characterize physical interaction unrecorded. Recent visual-tactile datasets provide this missing supervision, but their synchronized tactile data remain far smaller in volume than human video. Moreover, the largest resources often merge recordings from different sensors or annotation procedures, which makes the effect of data scale difficult to isolate. We therefore introduce TouchScale, a 500-hour dataset of contact-rich human interaction recorded with a single unified wearable setup. Its approximately 2K predefined task descriptions span everyday activities and structured manipulation, and each recording temporally aligns egocentric RGB-D video with wrist RGB video and dense full-hand bimanual tactile measurements. Compared with prior tactile data, training on the full TouchScale raises zero-shot contact IoU on data from an unseen tactile sensor from 0.134 to 0.383. Pretraining a visual encoder on TouchScale also yields the highest action recognition accuracy on three benchmarks among the compared visual-tactile datasets. Used for visual-tactile mid-training of a robot policy, TouchScale improves the average real-world success rate across four contact-rich manipulation tasks from 22.5% to 57.5%. With the sensor and collection protocol held fixed, both zero-shot tactile prediction and robot success show an overall upward trend as more TouchScale data is used. These results suggest that human visual-tactile data collected at scale with consistent sensing benefits both perception and robot manipulation. We will publicly release TouchScale, including all synchronized visual-tactile recordings and reconstructed object models, to support future research on scalable visual-tactile learning.

[CV-30] ΔRepresentation: Geometry Supervised Representation Learning of Phenotypes via Counterfactual Reasoning for Medical VLMs

链接: https://arxiv.org/abs/2610.10286
作者: Hao Wang,Qiwei Zeng,Jinghao Lin,Shuchang Ye,Yuezhe Yang,Yige Peng,Haoyuan Che,Jinman Kim,Lei Bi
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Medical vision-language models (VLMs) have shown increasing potential for radiological image interpretation. Medical VLMs encode radiological images into visual representations that capture both anatomical and phenotypic information for diagnosis. Existing approaches improve pathological phenotype representations through semantic-guided representation alignment. However, pathological phenotypes arise as lesion-specific visual changes superimposed on underlying normal anatomy. Such semantic alignment approaches fail to model the phenotype-specific increment relative to the corresponding normal anatomical representation. To address this gap, we propose \textbf \Delta Representation, a visual phenotype representation learning framework based on counterfactual reasoning for medical VLMs. It comprises \textbfBaseAnatomy, a geometry-supervised representation learning module, and \textbf \Delta Phenotype, a counterfactual incremental representation learning module. BaseAnatomy provides fine-grained geometric supervision through spatial relationships across and within anatomical structures. \Delta Phenotype computes the representation increment between lesion representations and their corresponding normal anatomical representations, and supervises increments associated with the same phenotype to cluster in the representation space. Experiments on \textitReXGroundingCT and \textitLIDC-IDRI demonstrate that \Delta Representation effectively structures pathological phenotype representations and improves lesion grounding and phenotype characterization accuracy in medical VLMs. Code is available at this https URL.

[CV-31] mporal Visuo-Tactile Learning for Dexterous Grasp Stability

链接: https://arxiv.org/abs/2610.10283
作者: Ken Nakahara,Aleksei Buvailik,Prokhor Kotov,Roberto Calandra
类目: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注: 12 Pages. Website: this https URL

点击查看摘要

Abstract:Humans can grasp everyday objects with almost perfect success rates using fingertip tactile feedback, yet much of the robotic grasping literature emphasizes vision-based grasp selection with parallel grippers. In this work, we systematically investigate how high-resolution, dynamic tactile sensing contributes to grasp stability prediction and model-guided grasping in dexterous robotic hands. To this end, we collected a dataset of 10,000 grasp trials across 200 objects using a multi-fingered robotic hand equipped with four Digit 360 tactile sensors, recording external vision, proprioception, and tactile streams throughout each grasp. With this dataset, we trained end-to-end temporal multimodal models to predict post-lift stability from pre-lift grasp observations and compared sensing modalities and encoding backbones. Experimental results and controlled input ablations show that incorporating touch, and particularly high-resolution, dynamic touch, improves grasp stability prediction. Finally, we deployed the learned predictor as an online stability gate on the real robot, where visuo-tactile model-guided regrasping improved the success rate among executed lifts by 10.5 percentage points over a non-tactile gate. These results show how rich fingertip sensing and expressive temporal models that capture the dynamics of touch can support learned grasping with multi-fingered hands without explicit contact or force modeling, providing a scalable data-driven path from tactile experience toward stable dexterous manipulation. The dataset is publicly available at this https URL.

[CV-32] Video Prediction Policy 2: Predict Better Act Better

链接: https://arxiv.org/abs/2610.10270
作者: Yanjiang Guo,Haodong Yan,Zhide Zhong,Zhongru Zhang,Qingyuan Yang,Qingzhou Lu,Xiaoyu Chen,Yen-Jen Wang,Shuying Deng,Chenghan Yang,Puzhen Yuan,Chenxin Liu,Tun Ban,Xiang Zhu,Yichen Liu,Kun Feng,Haoang Li,Jianyu Chen
类目: Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)
备注:

点击查看摘要

Abstract:World action models (WAMs) have emerged as an important class of generalist robot policies, aiming to transfer video prediction priors to action learning. However, we find that existing WAMs frequently produce incorrect motion predictions in open-ended environment, leading to erroneous actions. We attribute this limitation to two factors: (1) base video models are not optimized for manipulation, and (2) naively incorporating action components into video models can substantially degrade their generalization capabilities. We introduce Video Prediction Policy 2 (VPP2), a WAM that enables strong zero-shot generalization in both video prediction and action generation. First, we curate a large-scale, diverse dataset of manipulation videos to continue pretraining the base video foundation model. We annotate video clips with detailed captions and perform \textitevent-level video pretraining to promote generalization across open-ended manipulation tasks. Second, we post-train and distill the video model into a single-step visual planner with fixed prediction horizon. Finally, we introduce action module via a mixture-of-transformers (MoT) architecture to learn implicit inverse dynamics model. Experiments demonstrate three key results: (1) VPP2-14B outperforms Cosmos3-64B by 11.0% points in video prediction instruction-following success rate on open-ended tasks; (2) VPP2 surpasses the strongest baseline by 18.5% points in success rate on real-world zero-shot ALOHA manipulation tasks; and (3) following benchmark-specific post-training, VPP2 achieves the highest success rates among evaluated methods on the challenging LIBERO-Pro, LIBERO-OOD, and RoboDojo benchmarks.

[CV-33] LoomSC: Scalable Deep Subspace Clustering with Projector Factorization and Exact Spectral Reduction

链接: https://arxiv.org/abs/2610.10266
作者: Nairouz Mrabah,Youssef Melki,Mohamed Bouguessa,Riadh Ksantini,Shakeeb Murtaza,Tehseen Zia
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注: 19 pages, 7 figures, 5 tables; includes appendices

点击查看摘要

Abstract:Dense self-expression matrices and full-affinity spectral clustering limit the scalability of subspace clustering. We introduce the Latent Orthogonal Optimization Model for Subspace Clustering (LoomSC), a framework that addresses both bottlenecks through projector factorization and exact spectral reduction. Motivated by the spectral structure of least-squares regression, LoomSC jointly learns latent features and a projector self-representation through two thin factors. Alternating Procrustes and least-squares updates preserve the sample factor’s orthogonality while keeping the coefficient matrix implicit. We construct a nonnegative quadratic affinity that preserves the projector’s support. An exact feature map then reduces its normalized spectral problem to an eigenproblem whose dimension depends only on the factor width. Neither the full affinity nor the sample Laplacian needs to be formed. Our analysis quantifies the projector approximation and identifies conditions for subspace preservation and within-subspace connectivity. For fixed dimensions and iteration budgets, the complete pipeline has linear time and memory complexity in the number of samples. Across five image-clustering benchmarks, LoomSC ranks first or second in all 15 dataset-metric comparisons against 9 state-of-the-art baselines. Its mean accuracy exceeds the highest baseline mean by 6.66 percentage points. Synthetic experiments scale to 500,000 samples while maintaining at least 99.8% accuracy.

[CV-34] Geometry-Supervised Visual Representation Learning for Multi-Phenotype Lesion Interpretation in Medical VLMs

链接: https://arxiv.org/abs/2610.10238
作者: Hao Wang,Qiwei Zeng,Shuchang Ye,Jinghao Lin,Yuezhe Yang,Yige Peng,Jinman Kim,Lei Bi
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Medical vision-language models (VLMs) have shown increasing potential for clinical image interpretation. However, these models still struggle to interpret multi-phenotype lesions whose diagnosis requires the joint assessment of multiple pathological phenotypes. Existing vision-language alignment methods produce visual representations that fail to preserve anatomical hierarchies and relationships among phenotypic subclasses. This stems from their reliance on semantic supervision, which lacks geometric constraints to preserve these relationships in the visual embedding space. Moreover, the sparsity of lesion-related anatomical and phenotypic representations makes it difficult for medical VLMs to capture important diagnostic evidence. To address these limitations, we propose \textbfPureVision, a geometry-supervised visual representation learning framework for multi-phenotype lesion interpretation in medical VLMs. It combines a geometry-supervised representation learning module, \textbfPureEyes, and an anatomy-guided evidence aggregation module, \textbfPureNeurons. PureEyes provides geometric supervision through ideal spatial distributions that encode anatomical hierarchies and phenotypic subclass relationships. PureNeurons projects visual representations into the learned latent space, using their positions to selectively aggregate lesion-specific anatomical and phenotypic evidence. Experiments on \textitLIDC-IDRI, \textitCBIS-DDSM, and \textit3DReasonKnee demonstrate that PureVision improves lesion grounding and phenotype characterization in visual question answering and radiology report generation. Code is available at: this https URL.

[CV-35] Masked Feature Encoding for Large-Scale Whole Slide Image Representation ACCV2026

链接: https://arxiv.org/abs/2610.10225
作者: Haoyu He,Basile Tessier-Cloutier,Yang Wang,Mahdi S.Hosseini
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted at ACCV 2026

点击查看摘要

Abstract:Whole slide image (WSI) analysis in computational pathology follows a multiple instance learning (MIL) pipeline where patch embeddings are extracted independently and aggregated for slide-level prediction, but within-slide variance from staining, scanner, and local texture can overwhelm the discriminative signal. We propose Masked Feature Encoding for Multiple Instance Learning (MFE-MIL), a feature-space masking framework that trains a lightweight MLP adapter jointly with a window-based masked reconstruction branch and a MIL classification head. The two objectives are complementary. Classification guides the adapter to suppress within-slide patch variance, while window-based masked reconstruction provides an auxiliary regularizer for the adapted features without using patch coordinates, coordinate graphs, or segmentation preprocessing. The raster patch-extraction order is used only as a weak implicit prior. At inference, the decoder is removed, leaving only the adapter and MIL head. Across CAMELYON16/17, PANDA, and TCGA-BRCA with four diverse encoders, MFE-MIL improves ACC/F1 for nearly all tested aggregator-encoder settings and AUC in most, outperforms coordinate-based spatial methods (CAMIL), and achieves higher AUC than 2DMamba on three of four datasets (UNI). On five TCGA survival cohorts it improves the average concordance index for every aggregator tested, its most consistent gain. Code is available at this https URL.

[CV-36] VolCo: Volumetric Contact for High-Fidelity Human Grasp Generation NEURIPS2026

链接: https://arxiv.org/abs/2610.10197
作者: Zhuo Chen,Yihua Cheng,Aleš Leonardis,Hyung Jin Chang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted to NeurIPS 2026

点击查看摘要

Abstract:Accurate contact modeling is fundamental to understanding hand-object interaction, yet existing contact representations are typically restricted to object surfaces and rely on hand-crafted rules to recover contact details, leading to severe penetrations and implausible results. To better exploit the rich detail in motion-capture data, we introduce Volumetric Contact (VolCo), a representation that expands surface points to a set of 3D volumetric grids. VolCo encodes 3D contact that allows precise hand part recovery, and is organized in an inherent hierarchy: local contact details within each volume and global hand geometry across all volumes. Our framework, VolCoDiff, employs two modules to capture local and global features following this hierarchy. For local contact details, we use a 3D variational autoencoder to model the possible hand configurations conditioned on the local object signed distance field (SDF). For global hand geometry, we design a prior-guided diffusion model that learns the distribution of compressed latent features aggregated from the volumetric grids. We evaluate our method on two benchmark datasets and demonstrate state-of-the-art performance in penetration and stability, indicating the capability to generate tight grasps with much less severe penetrations. Our code is available at this https URL.

[CV-37] HuLiGen: Human LiDAR Generation from Parametric Body Models

链接: https://arxiv.org/abs/2610.10196
作者: Salma Galaaoui,Nermin Samet,David Picard
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 12 pages, 5 figures, 7 tables

点击查看摘要

Abstract:LiDAR point clouds of humans are extremely expensive to collect and annotate, thus represent a scarce resource that hinders the development of human analysis using this modality. To alleviate this scarcity, prior work relies on simulated human LiDAR, but such samples do not fully reflect the geometry and sensing characteristics of real observations. In contrast, we introduce HuLiGen, a generative model that generates human LiDAR point clouds from a parametric body model, using a point transformer trained with a flow-matching objective. We show that our generated point clouds are closer to the real capture distribution. Using HuLiGen to generate synthetic data, we propose a synthetic-only pretraining scheme for LiDAR-based HPE that achieves state-of-the-art performance, with even larger gains in low-annotation and low-data regimes, where MPJPE is reduced by up to 50%. Code, models and generated samples are available at this https URL.

[CV-38] VideoEvolve: Co-Evolving Memory and Retrieval for Long Video Understanding

链接: https://arxiv.org/abs/2610.10183
作者: Yongchao Xu,Bowen Ye,Jiefeng Gan,Junkai Ma,Wenzhao Li,Sen Tao,Yi Wei,Jiawei Liu
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Long video understanding increasingly relies on external memory to organize massive visual streams into compact representations. However, most memory-based methods dynamically adapt how information is retrieved for different questions, while largely fixing what is remembered. This mismatch makes missing details costly to recover, whereas stored information is valuable only when it can be reliably retrieved. To address this issue, we propose VideoEvolve, a novel self-evolving framework that jointly evolves memory and retrieval for long video understanding. Specifically, starting from a coarse low-frame-rate overview, VideoEvolve couples a Memory Evolver for selective memory augmentation with a Retrieval Evolver for adaptive retrieval over the evolving memory. We then co-evolve the two Evolvers through alternating agentic reinforcement learning (Agentic RL), updating one while freezing the other. To steer this alternating evolution, Bottleneck-Aware Evolution Feedback (BEF) identifies whether the current bottleneck lies in memory or retrieval and directs optimization toward the more limiting side. Furthermore, VideoEvolve introduces Capability-Aware Evolution Feedback (CEF) to alleviate downstream feedback from over-specializing memory to a fixed set of training questions, shifting training toward underdeveloped yet learnable video capabilities. By integrating Agentic RL with BEF and CEF, VideoEvolve transforms downstream reasoning experience into transferable capability updates, providing a concrete path from static long-video systems toward experience-driven, self-improving multimodal intelligence. Extensive experiments on multiple long video understanding benchmarks demonstrate the effectiveness of VideoEvolve.

[CV-39] Argos: Adapt Rich Geometric Priors for Generalizable Online Scene-Change-Detection WWW

链接: https://arxiv.org/abs/2610.10181
作者: Ruihan Xu,Jiae Yoon,Kaichen Zhou,Ue-Hwan Kim,Luca Carlone
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: More details on the project website: this https URL

点击查看摘要

Abstract:Robots operating in dynamic environments require reliable detection of how their surroundings change over time. Existing learning-based methods largely rely on pairwise 2D image features, which struggle under large viewpoint changes and occlusions, are sensitive to noise, and show limited generalization across domains, while explicit 3D approaches typically require costly offline optimization. We show that the implicit 3D knowledge of Geometric Foundation Models (GFMs) provides a strong basis for addressing these limitations. We introduce Argos, which adapts GFM features for joint scene change detection and 3D reconstruction. To address data scarcity and take a step toward a foundation model for scene change detection, we introduce a large-scale benchmark comprising two synthetic datasets and one real-world dataset, and train jointly across diverse datasets to improve cross-domain generalization. We further introduce Argos-SLAM, a real-time system designed for robotics, which performs online change detection and change-aware 4D mapping. Across benchmarks, our framework substantially outperforms existing baselines, with gains of up to 42.01% in change IoU and 27.91% in F1, while supporting scalable deployment in changing real-world environments.

[CV-40] Beyond Anonymous Captions: Grounding Character Identity in Video Captioning and Question Answering

链接: https://arxiv.org/abs/2610.10163
作者: Anas Filali Razzouki,Killian Steunou,Khalil Guetari,Thomas Kling,Mounîm El-Yacoubi,Yannis Tevissen
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Linking people’s appearance and actions to character identities is essential for understanding video narratives. We present a framework for identity-aware video captioning and person-centric question answering that combines automatic character identification, explicit spatial grounding, and task-specific adaptation. Starting from LSMDC v2 movie clips, our pipeline matches detected faces to actor reference images, tracks characters across frames, and builds inputs with identity-linked bounding boxes. A strong vision-language model generates identity-aware captions and questions, which are manually verified and filtered to create a benchmark of 750 captioned clips and 3,000 person-centric questions. We study five grounding strategies combining textual coordinates with visual face or estimated person boxes across Video-MLLM families at roughly 2B, 4B, and 8B parameters and larger frontier models. Combining visual face boxes with textual coordinates yields the most consistent performance across scales and significantly improves overall performance over coordinates alone. Smaller models tend to over-assign known identities when the queried person is not grounded, while larger models better recognize such UNIDENTIFIED cases. We introduce BAC by LoRA fine-tuning Qwen models at 2B, 4B, and 8B scales on about 32K identity-aware captioned clips. Across all scales, BAC outperforms every other evaluated model family of comparable size. BAC-8B reaches 93.20% overall QA accuracy, ranking behind only GPT-5.6 Sol among the frontier models evaluated in our study. Overall, explicitly communicating who is where, together with lightweight task-specific adaptation, substantially improves identity-aware video understanding without changing the underlying architecture. We release the benchmark, training data, code, and BAC checkpoints at this https URL.

[CV-41] BagDINO: Multi-View Baggage Re-Identification with DINOv3

链接: https://arxiv.org/abs/2610.10160
作者: Vita Santa Barletta,Danilo Caivano,Rebecca Margiotta,Massimiliano Morga,Davide Pio Posa
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 7 pages, 4 figures, 3 tables, IEEE International Conference on Evolving and Adaptive Intelligent Systems 2026 (IEEE EAIS 2026)

点击查看摘要

Abstract:Mishandled checked baggage remains a recurrent issue in airport operations, and current recovery workflows still largely rely on tag-based tracking, which does not directly support visual identification when tag evidence is missing or unavailable. This paper investigates baggage re-identification as an instance-level retrieval problem in a multi-camera setting, leveraging DINOv3 foundation-model representations to match a query image against a gallery of registered baggage images. A Torchreid-style BNNeck re-identification head is placed on top of a DINOv3 backbone, and parameter-efficient adaptation is performed via LoRA. Experiments are conducted on the MVB benchmark using a progressive study that compares a fully frozen backbone against LoRA and fine-tuning strategies. Results indicate that parameter-efficient adaptation of foundation-model features provides an effective and stable approach for multi-view baggage re-identification under limited training data.

[CV-42] HeiCo-FOCUS: A Clinically Grounded Dataset for Long-Context Video Understanding

链接: https://arxiv.org/abs/2610.10156
作者: Leon Mayer,Lucas Luttner,Patrick Godau,Kai Fritzsche,Annika Reinke,Leonie Boland,Jule Brandt,Janne Heinecke,Chloe K. Nobuhara,Niklas Holzwarth,Evangelia Christodoulou,Marcel Knopp,Dominik Michael,Pascale Piermarco,Saliq Neyaz,Korhan Derin Özarslan,Jakob Hennighausen,Carlos Aumente-Maestro,Tim Rädsch,Dheeraj Baji,Peter Maximilian Full,Finn Aichholz,Justus Veit Erpenbeck,Linus Finn Schott,Bastian Winkelhausen,Claas de Boer,Bianca Güttner,Anneli Hummel,Gregor Just,Max Kirchner,Chenyang Li,Rozenn Raffaut,Ariel Rodriguez,Danush Kumar Venkatesh,Kevin Wang,Jinjing Xu,Mona Sheikh Zeinoddin,Salman Khan,Thomas M. Pausch,Stefanie Speidel,Danail Stoyanov,Daniel A. Hashimoto,Fiona R. Kolbinger,Thomas G. Weiser,Lena Maier-Hein
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 28 pages, 9 figures, 5 tables. Code: this https URL

点击查看摘要

Abstract:Recent advances in Vision-Language Models (VLMs) have led to rapid progress in video understanding across a wide range of benchmark tasks. However, existing evaluations largely focus on short-term reasoning, failing to assess a critical capability: maintaining cumulative temporal consistency over extended time horizons. To close this evaluation gap, we introduce HeiCo-FOCUS, a clinically grounded dataset for evaluating long-context video understanding through the task of Foreign Object Contextual Understanding in Surgery. Built on a dataset of Heidelberg Colorectal surgeries, this task requires models to continuously track multiple objects as they are inserted, manipulated, occluded, and removed over procedures lasting up to hours. HeiCo-FOCUS comprises 30,000 visual question answering (VQA) pairs covering five core capabilities: object recognition, temporal grounding, aggregation, event and procedural understanding, and complex reasoning. The dataset was constructed through a rigorous multi-stage annotation pipeline involving large-scale crowd annotation and 39 surgical domain experts to ensure high quality and clinical relevance. To systematically probe model behavior, we introduce a multi-track evaluation framework that progressively increases temporal and contextual demands from single frames to full procedures. Experiments with ten frontier VLMs show that HeiCo-FOCUS tasks are far from solved: only around half of the models clearly outperform a text-only baseline. Across the video tracks, models perform best on event and procedural understanding (mean Accuracy: 56.5% across all models), while temporal grounding remains particularly challenging for all evaluated models (mean Accuracy: 19.7%). We therefore expect HeiCo-FOCUS to serve as a catalyst for the development of models capable of reliable, temporally consistent reasoning over hours-long videos.

[CV-43] HarnessIR: Harnessing Multimodal Foundation Models for Universal Real-World Image Restoration

链接: https://arxiv.org/abs/2610.10133
作者: Xiangtao Kong,Shuaizheng Liu,Rongyuan Wu,Lingchen Sun,Zhengqiang Zhang,Jinxin Zhao,Yuhui Wu,Lei Zhang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Real-world low-quality images suffer from complex mixed degradations, including but not limited to noise, blur, atmospheric effects, etc. Recent agentic methods usually model real-world image restoration (Real-IR) as a sequential tool calling problem over task-specific single-degradation restoration models. This paradigm, however, is fundamentally limited because complex real-world degradations cannot be cleanly undone degradation by degradation, and the tool used for task-specific models caps the capability of the agent system. In this work, we present HarnessIR, an agentic framework for Real-IR by harnessing a multimodal foundation model (MFM) as the executor. HarnessIR consists of five stages: perception and diagnosis, on-demand tool invocation, prompt composition, execution, and verification-driven refinement. Unlike prior agentic Real-IR methods that rely on tool chains assembled from task-specific models, HarnessIR feeds the restoration requirements, the perceptual diagnosis, and the evidence into an MFM that performs restoration in a single pass, followed by verification stages to determine whether the result warrants further processing. Under our harness, off-the-shelf MFMs handle restoration tasks remarkably well, achieving state-of-the-art results on the widely used MiO100 synthetic benchmark. More importantly, by exploiting the strong generalization ability of MFMs, HarnessIR delivers compelling restoration quality on challenging real-world scenes where previous agentic IR systems often struggle. Codes is available at this https URL.

[CV-44] Lifelong small-object navigation in changing object layouts: a benchmark and method

链接: https://arxiv.org/abs/2610.10125
作者: Jiagan Huang,Zikun Zhou,Zijian Ni,Hongpeng Wang,Guangming Lu,Jun Yu,Wenjie Pei
类目: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Household robots need to continually navigate to different objects in the same environment, many of which are small and portable, such as tools and toys. Their small visual footprint and frequent occlusion make reliable observation difficult, and they may be moved by people without the robot observing the changes. We formulate this challenging task as Lifelong Small-object Navigation in Changing Object Layouts (LiSoNav-COL). Agents must seek suitable viewpoints for reliable observation, accumulate and reuse scene knowledge to efficiently locate subsequent targets, and update outdated memory after object relocation. To eliminate the need for prior scene scanning, we also require agents to start navigation with empty scene memory. Although practical, this task still lacks benchmarks designed around its defining assumptions. To bridge this gap, we introduce LiSoNav-Eval, a dedicated benchmark spanning 28 indoor scenes with 45 small-object categories. Its lifelong navigation sequences include both unchanged and relocated targets to evaluate memory reuse and adaptation to object relocation. To address this challenging task, we propose a navigation method based on multi-view Inspection with Viewpoint-Anchored Memory, dubbed IVAM-Nav. IVAM-Nav actively observes supporting surfaces from complementary viewpoints for reliable small-object perception and anchors the resulting memory to their observation viewpoints, supporting relational memory reuse and revalidation under similar viewing conditions. Extensive experiments on LiSoNav-Eval demonstrate favorable performance of IVAM-Nav against representative methods. Benchmark analyses also show that smaller objects, larger environments, and longer relocation distances pose greater challenges. The dataset and code are available here.

[CV-45] A Probabilistic Perspective on Wasserstein-Based Evidential Uncertainty for Out-of-Distribution Segmentation

链接: https://arxiv.org/abs/2610.10116
作者: Arnold Brosch,Abdelrahman Eldesokey,Michael Felsberg,Kira Maag
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 20 pages, 6 images, 3 figures

点击查看摘要

Abstract:Semantic segmentation networks operate on a fixed set of classes and therefore fail when out-of-distribution (OOD) objects appear during deployment, a critical limitation for safety-critical applications such as autonomous driving. Reliably identifying OOD objects requires well-calibrated epistemic uncertainty, yet common softmax-based confidence scores remain overconfident, while Bayesian alternatives such as Monte Carlo dropout or deep ensembles require costly repeated forward passes. Evidential Deep Learning (EDL) offers an efficient alternative by modeling class probabilities as a Dirichlet distribution learned from a single deterministic forward pass. Existing EDL formulations rely on Euclidean objectives that push predictions towards the simplex vertices, encouraging overconfidence rather than preserving uncertainty for unfamiliar inputs. We instead employ Wasserstein-based objectives, which respect the geometry of the probability simplex, and study the influence of the Wasserstein order on segmentation accuracy and OOD detection within a unified evidential framework. We evaluate this framework on a convolutional (DeepLabV3+) and a transformer-based (SegFormer) architecture on the SegmentMeIfYouCan benchmark, including LostAndFound, RoadObstacle21, RoadAnomaly21, and Fishyscapes. Our results show the optimal Wasserstein order is architecture-dependent: second-order objectives dominate on the convolutional backbone, third-order objectives on the transformer backbone, and our framework surpasses comparable baselines on most metrics, with a single deterministic forward pass.

[CV-46] mporal Residual Bottleneck for Robust Asynchronous Collaborative Perception ACCV2026

链接: https://arxiv.org/abs/2610.10090
作者: Melih Yazgan,Ahmed Abouelazm,J. Marius Zöllner
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted to ACCV 2026

点击查看摘要

Abstract:Collaborative perception extends the sensing range of autonomous vehicles, but its performance degrades when shared features arrive stale or incomplete. Most latency-robust methods compensate delayed collaborator features through flow-guided alignment or direct feature transport. In this work, we formulate asynchronous collaborative perception as temporal residual prediction. Our Temporal Residual Bottleneck keeps a deterministic pose-warped collaborator feature as a conservative anchor and uses a \Delta t -conditioned xLSTM to extract residual temporal evidence from the available history. A detector-facing residual bottleneck then applies only gated, regularized corrections before ego-side fusion, reducing the risk of overwriting reliable static structure when temporal correspondence is uncertain. Experiments on DAIR-V2X and OPV2V show that our method is especially effective under severe fixed/irregular delays and packet drops. On DAIR-V2X, the reported checkpoint trades a small amount of synchronized peak accuracy for better robustness under stronger communication degradation. Controlled diagnostics further indicate that direct feature transport has oracle headroom but can become unreliable when deployed without accurate correspondence. These results support temporal residual fusion as a practical alternative for asynchronous and incomplete collaborative perception. Code will be publicly released at this https URL.

[CV-47] From Pixel to Coding: Evaluating the Figure Reproduction Capabilities of MLLM s

链接: https://arxiv.org/abs/2610.10066
作者: Zijian Chen,Zhengyu Chen,Bohan Liang,Lirong Deng,Yushuo Zheng,Yanwei Jiang,Qi Jia,Kaiwei Zhang,Wenjun Zhang,Guangtao Zhai
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注: 46 pages, 18 figures

点击查看摘要

Abstract:Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in both visual understanding and code generation. However, existing benchmarks typically evaluate these two modalities in isolation, lacking a dedicated assessment of their unification, i.e., how a model can perceive complex visual structures and synthesize them into precise, executable code. Moreover, current visual code generation benchmarks often rely on simplified layouts within single programming environments, falling short of evaluating true unified multimodal reasoning. To bridge this gap, we propose FigCodeBench, a comprehensive framework for rigorously evaluating MLLMs on figure reproduction, integrating multimodal comprehension and generation. We first design a systematic dataset construction pipeline, resulting in a total of 6,194 instances that cover 7 functional categories and 4 types of programming languages. We further categorize figure reproduction into three tiers with visual and code complexity modeling, specifically targeting complex structural reasoning, varying aspect ratios, and dense geometric constraints. We introduce a multi-dimensional evaluation protocol, encompassing visual fidelity and syntactic isomorphism, that aligns highly with the Mean Machine Opinion Score (MMOS) and human preferences. Based on our framework, we conducted extensive experiments on 24 widely used proprietary and open-source MLLMs (e.g., Gemini 3.1 Pro, GPT-5.4, and Kimi-K2.5), where we observed a universal, non-linear performance cliff across different programming languages and difficulty scenarios for all models, and gained several insights, such as the significant metric decline in rigid declarative languages.

[CV-48] AdSpark: A Large-Scale Dataset and Benchmark for Product-Centric Advertisement Video Generation

链接: https://arxiv.org/abs/2610.10047
作者: Zhifei Yang,Zhao Jiang,Keyang Lu,Honghe Zhu,Zheng Zhang,Jingjing Lv,Changping Peng,Ching Law,Zhen Xiao
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Product-centric advertisement video generation aims to create promotional videos that preserve fine-grained product identity while presenting selling points through coherent multi-shot narratives. However, this emerging task remains underexplored due to the lack of large-scale advertisement-specific datasets and comprehensive evaluation frameworks. To address this gap, we introduce \textbfAdSpark, a large-scale dataset and benchmark for product-centric advertisement video generation, based on data from a major e-commerce platform. \textitAdSpark-300K contains approximately 300K reference image–prompt–video triplets, comprising a real-world subset and a synthetic subset. Each sample provides structured advertisement annotations, including product identity annotations, selling-point descriptions, creative plans, and aligned audio scripts, enabling models to learn product preservation and advertisement-oriented visual storytelling. We further propose \textitAdSpark-Bench, a diagnostic benchmark that evaluates generated advertisements across six dimensions, including visual quality, product fidelity, instruction adherence, temporal coherence, audio alignment, and advertisement effectiveness. Based on AdSpark-Bench, we evaluate representative models, revealing key challenges in product preservation, multi-shot storytelling, and selling-point visualization. Experiments with AdSpark-300K-finetuned models further validate the effectiveness of our dataset. AdSpark provides a unified dataset and benchmark for future research, and we will release the dataset upon acceptance.

[CV-49] Scalable Patch-Level Self-Supervised Learning

链接: https://arxiv.org/abs/2610.10013
作者: Maximilian Seitzer,Gabriele Trivigno,Antonín Vobecký,Seungeun Yi,Maxime Oquab,Huy V. Vo,Oriane Siméoni,Piotr Bojanowski
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Self-supervised learning (SSL) at scale produces powerful visual representations. However, most scalable SSL methods rely on ad hoc combinations of multiple objectives and stabilization mechanisms. Taking a step back, we ask if we can design a high-performing, yet principled SSL algorithm. Starting from the multi-view assumption, stipulating that task-relevant content is captured by the information common to different views, we construct an information-theoretic objective decomposing into interpretable terms. This derivation yields JEM, a student-teacher method that learns by aligning corresponding patch representations across views, explicitly regularized by information and structure preservation losses. JEM trains stably from 300M to 7B parameters, and, to our knowledge, is the first latent-space patch-level method demonstrated at 7B scale. Across all scales, JEM reaches strong performance on both global and dense probing tasks, on segmentation benchmarks consistently surpassing the DINOv2 algorithm, an influential foundation for today’s strongest visual SSL methods. Notably, at 7B parameters, it exceeds the performance of DINOv3 on panoptic segmentation, despite being trained on 12\times less data without refinement stages. These results demonstrate that we can indeed design an SSL algorithm that learns strong representations, is principled and stable.

[CV-50] Playing with Kruskal: algorithms for flat and hierarchical watershed cuts

链接: https://arxiv.org/abs/2610.10012
作者: Jean Cousty(LIGM),Laurent Najman(KUSTAR, LIGM),Benjamin Perret(LIGM),Deise Santana Maia(CRIStAL)
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:In the framework of edge-weighted graphs, watersheds have proven to be linked to well-known optimization problems, as Minimum Spanning Tree, which allowed the design of efficient algorithms for computing (hierarchical) watershed segmentations. In the present article, after reviewing the literature related to watershed segmentation, we present a detailed end-to-end pipeline of algorithms to compute (hierarchical) watershed segmentations, starting from the computation of graph-based image representations, up to the computation of connected components of the final (hierarchical) segmentation. We consider the several variations of watersheds, including their supervised and unsupervised versions, and the various ways of computing seeds, to name a few. For the first time, we bring together all these watershed notions and algorithms in a compact and understandable way. We aim at providing a reference for those interested in employing and reimplementing the watershed segmentation framework for their task at hand.

[CV-51] Perceptually Aligned Evaluation of Style Transfer

链接: https://arxiv.org/abs/2610.10003
作者: Yang Deng,Eleftherios Ioannou,David Mould,Steve Maddock,Paul L. Rosin,Yu-Kun Lai
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Style transfer lacks a reliable evaluation standard: ground truth is inherently ill-defined, and existing automatic metrics often fail to reflect human preference. This paper introduces ASTRA (Assessment of Style TRansfer Algorithms), an approach for automatic evaluation of style transfer algorithms; it contains two components, ASTRA-Data and ASTRA-Score. ASTRA-Data consists of a benchmark image set of content and style references, a collection of style transfer results generated on the benchmark set, and user study data capturing human judgements through a two-stage pairwise comparison protocol. From these annotations, we derive ranking-based ground truth for content preservation, style fidelity, and overall preference. Based on ASTRA-Data, we construct ASTRA-Score, a learnt evaluator that predicts preference-aligned scores from content-style-stylization image triplets, enabling automatic and scalable evaluation of new models applied to the benchmark set. Experimental results demonstrate that ASTRA-Score achieves substantially higher correlation with human rankings compared to prior metrics. Overall, ASTRA establishes a robust mechanism for standardised evaluation of style transfer methods.

[CV-52] MSU Team at the Explainable Deepfake Detection Challenge 2026: Grounded Artifact Evidence for Deepfake Detection

链接: https://arxiv.org/abs/2610.09952
作者: Artem Filippov,Aleksandr Gushchin,Kirill Koltsov,Dmitriy Vatolin,Anastasia Antsiferova
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Recent advances in generative image models have made many manipulated images highly realistic, raising the need for detectors that are not only accurate but also able to provide visual evidence for their decisions. In this paper, we present our solution to the Explainable Deepfake Detection Challenge [2] on the XPlainVerse dataset [1], where systems are required to predict whether an image is real or fake and generate both complex and simple explanations grounded in visible forensic cues. Our method follows a modular detection-and-explanation design. For the real/fake decision, we build a multi-backbone detector that combines several DINOv3 models with Mesorch manipulation-localization features, bringing together pretrained visual representations, DCT-aware cues, and multi-scale forensic information. To inject explanation evidence into the detector, we use a Grounding-DINO-based pseudo-mask generation pipeline that converts local artifact descriptions from training explanations into weak patch- level supervision for an Artifact Evidence Map. We further introduce a local patch-level contrastive objective that separates artifact and authenticity evidence in the detector feature space without requiring paired images or pixel-level manipulation masks. For language output, we use class-conditional Qwen3-VL models to generate complex explanations for fake and real predictions, followed by a GRPO-optimized text simplification model. The proposed methods were trained and evaluated on the challenge subset of XPlainVerse. On the full test split, our submission achieves 0.9349 detection accuracy, 0.5571 explanation score, and a 0.7456 final challenge score.

[CV-53] Purifying Backdoored Large Vision-Language Models by Removing Hijacked Directions

链接: https://arxiv.org/abs/2610.09941
作者: Bojun Yang,Haochen Zhou,Zhifang Zhang,Haobo Wang,Songze Li,Lei Feng
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 25 pages, 9 figures, 14 tables

点击查看摘要

Abstract:Large vision-language models (LVLMs) are increasingly deployed in safety-critical applications, yet they remain vulnerable to backdoor attacks. Defending against such attacks remains costly, as existing methods require either extensive retraining on clean data or per-query intervention at inference time. To address this limitation, we propose OrthoPurify, a more efficient method to purify backdoored model weights via one-step orthogonal projection. Specifically, through structural analysis of backdoor weight updates, we find that the backdoor is encoded by diverting a small number of weight update directions from task adaptation to backdoor shortcut encoding, a phenomenon we term direction hijacking. However, identifying these hijacked directions requires a benign reference model, which is typically inaccessible to the defender. We show that a pseudo-benign model, obtained by fine-tuning the pretrained weights on only a small set of clean samples, provides a sufficient approximation, as the dominant update directions stabilize within the first few gradient steps. OrthoPurify uses this pseudo-benign reference to isolate the hijacked directions and removes them through a single projection on the weight update. Extensive experiments show that OrthoPurify reduces the attack success rate to near zero while preserving the original performance across diverse benchmarks, without retraining the backdoored model or introducing inference-time overhead. Our code is publicly available at this https URL.

[CV-54] Juno: Taming Predictive Latents for Vision-Language-Action Models

链接: https://arxiv.org/abs/2610.09940
作者: Yuchen Zhu,Chenyi Xu,Yulin Zhang,Gang Xu,Wentao Zhu
类目: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
备注: Project Page: this https URL

点击查看摘要

Abstract:Joint-embedding predictive architectures (JEPAs) predict masked or future observations in representation space, offering a natural source of predictive latents for vision-language-action (VLA) models. Yet making these latents useful across pretraining, policy learning, and deployment requires addressing three failures: mismatch with embodiment-specific control, interference with action learning, and teacher miscalibration under distribution shifts. We introduce Juno, a unified framework built around one action-conditioned JEPA that serves as a control-aligned representation backbone, a predictive teacher, and an adaptable dynamics model. During pretraining, we train it on embodiment-matched trajectories and use a dynamic CLS loss to transfer motion-weighted patch dynamics to a compact global state. During policy learning, we fuse current-frame JEPA patches into VLA perception and use a decoupled reasoning branch with separate transformation parameters to distill future latent states for action generation. During deployment, we adapt the world model on all observed transitions, including failed rollouts, freeze the adapted teacher, and re-align the policy on verified executions using LoRA adapters and a trainable action head, without expert corrections or task rewards. On SimplerEnv, Juno raises average success from 60.9% to 68.5% over Qwen3GR00T, the strongest baseline, and test-time adaptation further reaches 72.7% ; on a real robot, it retains 70% – 75% success under background, height, and object shifts where the base policy collapses to 0% .

[CV-55] Do Generative Priors Align with Human Naturalness Perception?

链接: https://arxiv.org/abs/2610.09928
作者: Taiki Fukiage
类目: Computer Vision and Pattern Recognition (cs.CV); Neurons and Cognition (q-bio.NC)
备注: 62 pages, 35 figures, including appendices

点击查看摘要

Abstract:Visual generative models are trained to capture the probability distributions of natural images, yet whether their native priors reflect the regularities governing human perception of image naturalness remains an open question. Here, we probe these priors through native prediction errors across 25 open image and video generators. Because raw single-image losses are dominated by scene content and visual complexity, we evaluate directional loss differences using content-preserving, paired relational interventions that selectively disrupt facial configurations or physical illumination consistency while limiting changes in low-level image statistics. Across both domains, these loss differences reproduce human-like selective sensitivities and tolerances, capturing the classic Thatcher effect on faces and shape-dependent responses to illumination inconsistencies. Notably, these loss differences reliably track continuous gradations of human naturalness judgments across individual stimulus pairs (peaking at r = .84 on faces and .64 on physical scenes) and retain unique human-aligned signals even after controlling for feature distances from frozen vision encoders and standard image quality metrics. We also find that while overall sensitivity to these violations broadly covaries with human alignment across models, the two systematically decouple along denoising schedules, with alignment peaking earlier than sensitivity, revealing that human-like naturalness judgments dissociate from generic violation detection. Together, these findings demonstrate that learning visual distributions yields generative loss landscapes that capture distinct aspects of human naturalness perception.

[CV-56] FedSSMCoOp: SSM Encoders for light-weight Federated Prompt Learning for Few-shot Classification

链接: https://arxiv.org/abs/2610.09907
作者: Ankita Das,Ambarish Parthasarathy,Sumohana S. Channappayya,C. Krishna Mohan
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Vision-Language Models (VLMs) have shown strong performance across a wide range of downstream vision tasks, thanks to the complementary information contained in the respective domains. Despite the performance gains, most of these approaches rely on aligning these domains using the cosine similarity metric, which fails to capture token-level structure and cross-modal interactions prior to the classification stage. This is especially critical in biomedical applications under federated constraints, where data sharing is restricted, labeled data is scarce at each site, and it differs widely across institutions, leading to substantial statistical heterogeneity. To overcome this issue, we propose FedSSMCoOp, a federated few-shot image classification framework that enables multimodal learning while preserving data privacy. With the help of the SSM-based Vision Mamba and Cross Mamba blocks, and by optimizing only the soft-prompt and communication-prompt updates in the federated setting, the framework prioritizes both computation and performance. Importantly, this eliminates the need to use an external Large Language Model (LLM) for feature alignment. The framework is further trained and evaluated on various biomedical image datasets, and its performance is assessed. The proposed framework delivers stable performance relative to the baselines and is, on average, 1.96 times lighter. The corresponding script will be made available soon.

[CV-57] SANet: Selective Attention Network for Infrared Small Target Detection

链接: https://arxiv.org/abs/2610.09875
作者: Yingmei Zhang,Wangtao Bao,Qin Xiao,Yong Yang,Weiguo Wan,Yitao Luo,Xueting Zou,Lei Zhang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Infrared small target detection aims to accurately identify and locate dim targets in complex backgrounds and supports applications such as maritime surveillance and military search and rescue. However, the small size and weak contrast of infrared targets make it difficult to balance detection accuracy and false alarms. This paper proposes a selective attention network (SANet) for infrared small target detection. A dual-path semantic-aware module combines standard and pinwheel-shaped convolutions to preserve local spatial consistency and capture broader contextual information. Spatial and channel attention further refine the features and improve target-background discrimination. To address the limitations of static skip connections in U-Net, a selective attention fusion module adaptively integrates features across scales using spatially varying weights. It selectively enhances salient regions and improves discrimination between true targets and false alarms. Experiments on three public benchmarks, NUAA-SIRST, IRSTD-1K, and NUDT-SIRST, show that SANet achieves competitive performance in intersection over union (IoU), normalized IoU, detection probability, and false alarm rate. Its IoU exceeds that of the second-best method by 1.93, 4.32, and 2.21 percentage points, respectively. These results support the effectiveness of SANet in dim-target perception, discriminative feature representation, and background suppression.

[CV-58] Bringing BNNs to Fast Event Processing ECCV2026

链接: https://arxiv.org/abs/2610.09873
作者: Paul Longour,Julien Moreau,Franck Davoine
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 19 pages, 3 figures, 8 tables. Accepted by NeVi Workshop at ECCV 2026

点击查看摘要

Abstract:Binary Neural Networks (BNNs) enable efficient deep learning deployment on resource constrained devices with weights and activations compressed to one bit, substantially reducing model size and inference cost. Event cameras offer complementary advantages, including low latency, high dynamic range, and low power consumption, by capturing asynchronous streams of events rather than dense image frames. Despite their shared emphasis on efficiency, the combination of these technologies remains largely unexplored. This work aims at adapting and evaluating modern deep BNN architectures on event data. We also show that cross-modal pretraining from RGB data can improve the classification accuracy of BNNs on neuromorphic datasets. We introduce the Polar-wise Binary Event Volume (PBEV), a binary representation that enables event-camera data to be processed directly by BNNs and represents a step toward fully binarized event-based vision systems. Best evaluated BNN on N-Caltech101 classification benchmarks shows 90.58% accuracy with 7.5x less operations than their full-precision counterparts.

[CV-59] Global Averag e Precision for Representation Learning

链接: https://arxiv.org/abs/2610.09863
作者: Bill Psomas,Mohammad Mahdi,Michalis Thomas,Danda Pani Paudel,Giorgos Tolias,Giorgos Kordopatis-Zilos
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Standard information retrieval metrics, such as mean Average Precision (mAP), assess performance one query at a time, based on how the similarities between a query and its positives compare against those with its negatives. The same holds for common representation learning losses, such as InfoNCE and per-query AP surrogates. None of them considers whether similarities are comparable across queries, which any system with a single decision threshold relies on. Global Average Precision (gAP) does, by ranking all query-candidate pairs in one list and computing a single AP. We introduce gSAP, a differentiable surrogate of gAP. It needs only a similarity matrix and a binary matrix marking the positive pairs, the same input as existing losses, so it is a drop-in replacement for them and agnostic to the encoder, the modality, and the source of supervision. Since it considers all possible pairwise comparisons in the batch jointly, it also remains trainable at low temperatures, a regime where per-query surrogates run out of gradient. Swapping it into established recipes improves supervised metric learning, cross-modal alignment, and self-supervised pretraining, where, to our knowledge, it is the first ranking loss to replace the community standard InfoNCE in the latter two. Its similarities are more consistent across queries, which drives the gains under a universal threshold. gSAP retrieves up to four times as many positive pairs as the strongest AP surrogate at the same precision, and it degrades the least when queries with no positives in the database are added. Beyond thresholding, models trained with gSAP also learn better representations, with higher transfer, k NN and zero-shot classification accuracy.

[CV-60] DeepTopoClustering: Unsupervised Derivation of Surface Process Taxonomy from 4D Point Clouds for Topographic Monitoring

链接: https://arxiv.org/abs/2610.09860
作者: Jiapan Wang,Daan Hulskemper,Mathilde Letard,Roderik Lindenbergh,Katharina Anders
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注: Submitted to ISPRS Journal of Photogrammetry and Remote Sensing

点击查看摘要

Abstract:4D point clouds acquired by permanent laser scanning (PLS) enable accurate high-frequency monitoring of surface change in dynamic topographic environments. However, existing methods remain limited in organizing detected surface activities into meaningful process types. We propose DeepTopoClustering (DTC), an unsupervised framework for deriving a hierarchical process taxonomy from object-based surface activities, so-called 4D objects-by-change (4D-OBCs). We transform each 4D-OBC into a GeoMorphogram, a distributional sequence representing the temporal evolution of topographic change within a spatially bounded surface activity. A convolutional autoencoder learns latent embeddings from GeoMorphograms, which are jointly optimized using a hierarchical deep clustering objective to organize surface activities into a hierarchy. We evaluate the learned hierarchy using expert annotations on two 4D datasets of sandy beach sites and their combination. DTC with GeoMorphograms achieves the highest agreement with expert judgment at the taxonomy level comprising eight major process types ( F_1=0.78 , match accuracy =0.92 ), outperforming dimensionality reduction and conventional flat clustering. The learned taxonomy separates major erosion- and deposition-dominated activities and distinguishes finer subtypes based on change magnitude, duration, compactness, and temporal evolution. DTC thus provides a scalable and interpretable route from 4D change detection to a data-driven, expert-supported surface process taxonomy, advancing automated knowledge derivation for understanding surface dynamics in topographic monitoring.

[CV-61] DeltaSplat: Iterative Gaussian Refinement for Pose-Free Feed-Forward 3D Gaussian Splatting

链接: https://arxiv.org/abs/2610.09853
作者: Chanung Park,Seunghyeon Song,Joo Chan Lee,Eunbyung Park,Jong Hwan Ko
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 11 pages, 6 figures

点击查看摘要

Abstract:Pose-free feed-forward 3D Gaussian Splatting (3DGS) reconstructs a scene from sparse, unposed images in a single network pass, removing the need for camera calibration and per-scene optimization. However, camera estimation errors propagate into the predicted Gaussians and compound the geometric and photometric inaccuracies of single-pass prediction. To correct these errors, we introduce DeltaSplat, a lightweight Gaussian refinement module for pose-free feed-forward 3DGS. It iteratively renders the current Gaussians at the input context views and predicts per-Gaussian updates from the resulting residuals. A 2D residual alone, however, underdetermines the 3D correction. DeltaSplat therefore conditions each update on per-pixel Plücker rays and rendered depth as a soft geometric prior. A dual-branch convolutional mixer efficiently encodes these inputs, and per-attribute heads decode the fused features into position, opacity, and color updates. The module adds only ~2.2% parameters to the backbone and remains fully feed-forward at inference. On DL3DV, DeltaSplat reaches 26.64 dB PSNR in the pose-free setting, improving its state-of-the-art backbone by 1.75 dB and surpassing even baselines supplied with ground-truth cameras; consistent gains hold across 6-24 views and all camera regimes.

[CV-62] Hard Yet Reducible: Controlled Forward Transfer for Synthetic Degradation Curation

链接: https://arxiv.org/abs/2610.09849
作者: Chunming He,Kailai Zhou,Jiaming Zuo,Hanqi Liu,Fengyang Xiao,Youwei Pang,Xiaofeng Liu,Weisi Lin,Xiaoqi Zhao
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 17 pages, 4 figures, 9 tables

点击查看摘要

Abstract:Selecting synthetic degradations for dense prediction requires an estimate of their training utility, the generalization gain they bring under a finite training budget. Clean and degraded twins share content and labels, suggesting a score based on how much short training reduces the excess error caused by degradation. However, this gap can also shrink when clean performance deteriorates. Measuring the improvement on degraded images alone avoids that confound, but it still credits progress that the same amount of clean training would have produced. We propose the \textbfcontrolled Reducible Degradation Gap (cRDG) for regions defined by degradation type and severity. From a common checkpoint, cRDG runs two budget-matched probes that differ only in one augmentation slot, which holds either a synthetic degradation or a clean augmentation. The score is the gain on held-out degraded images relative to the clean-control probe. Clean harm is a separate feasibility constraint. cRDG reveals a correctable severity band in which training on the degradation yields high controlled gain under the available budget, and the band moves with the predictor, the starting checkpoint, and the training budget. \textbfCuration of Reducible Bands (\method) uses cRDG to select synthetic data without changing the predictor. On semantic segmentation and salient object detection, \method improves representative predictors under matched synthetic-data budgets and training schedules, extends to existing data-generation pipelines, and preserves clean performance. Code and supporting materials will be publicly released.

[CV-63] For Those Who Believe in Faithfulness: Optimizing the Area Under Insertion and Deletion Curves for Ranking Relative Feature Importance

链接: https://arxiv.org/abs/2610.09844
作者: Bjørn Leth Møller,Bulat Ibragimov,Christian Igel
类目: Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:The adoption of machine learning for socially relevant tasks requires effective explainable artificial intelligence (XAI) methods to better understand the behavior of machine learning models. Attribution methods are a popular XAI approach in which input-output relationships are characterized by heat maps that reflect the relative importance of input features for a particular prediction. The quality of such maps is often assessed by measuring faithfulness based on the area under insertion and deletion curves, which measures changes in the model output as features are added and removed. In this study, we derive an objective function from this notion of faithfulness and a way to approximate its gradient. We establish the connection between insertion curves and top- k feature selection, which leads to a loss function measuring the quality of attributions. Randomization of the loss allows us to efficiently approximate its gradient. To show the effectiveness of the general approach, we combine the loss function with the neural explanation mask framework. The resulting method, termed Ra-NEM, can be used with any differentiable model without affecting the model’s performance. Experiments demonstrate that Ra-NEM provides accurate attributions robustly and efficiently. Compared to other algorithms, the attributions have not only higher faithfulness but also perform well in terms of other XAI metrics. The high inference speed of Ra-NEM makes the method suitable for online applications. The code is available online: this https URL

[CV-64] ORCA: Hunting Compositional Failures in Text-to-Image Diffusion NEURIPS2026

链接: https://arxiv.org/abs/2610.09841
作者: Arshia Hemmat,Amirhossein Vahidi,Amitis Shidani,Mohammad Vali Sanian,Hesam Asadollahzadeh,Aryan Yazdan Parast,Mohammad Lotfollahi
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注: Accepted at NeurIPS 2026. 26 pages, 4 figures

点击查看摘要

Abstract:Text-to-image diffusion models fail predictably on compositional prompts: attributes bind to the wrong objects, spatial relations invert, and multi-object scenes lose count. Recent architectures already augment CLIP with a T5 encoder precisely because CLIP’s contrastive embedding loses compositional structure, yet these failures persist. We argue the binding problem is therefore not one of missing information but of misaligned information: a text encoder preserves compositional structure, but in a representation space shaped by language modelling rather than vision, and the denoising objective does not directly reward aligning the two. We show this correspondence can be supplied as an explicit training signal, that the relevant cross-modal information is concentrated in a low-rank subspace of self-supervised visual features, and that supplying it can be folded into diffusion training as a single auxiliary loss. Our method, ORCA (Orthogonal Residual Compositional Alignment), aligns the latent of a diffusion transformer with a low-rank target derived from a frozen visual encoder, through a predictor whose orthogonal basis is parameterised by a learned residual between T5 and CLIP embeddings, which provides a prompt-dependent signal for selecting the visual readout subspace. We prove that the cross-modal information recoverable at a given rank is bounded by the spectral mass of the visual encoder’s covariance in the top components. Across three diffusion-transformer backbones (DiT-B/2, DiT-L/2, U-ViT-L), ORCA improves FID and GenEval over both vanilla and REPA baselines at zero inference-time cost; on DiT-L/2 it reaches FID 16.65 and GenEval 0.291 at 200K steps, exceeding the strongest 400K baseline at half the training cost, with the largest gains concentrated on attribute binding, spatial relations, and multi-object prompts.

[CV-65] MOTIF: Person-of-Interest Deepfake Detection Beyond 3DMM Coefficients

链接: https://arxiv.org/abs/2610.09830
作者: Giovanni Affatato,Sara Mandelli,Paolo Bestagini,Stefano Tubaro
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 6 pages. Accepted at the 2026 IEEE International Workshop on Information Forensics and Security (WIFS)

点击查看摘要

Abstract:Video deepfakes targeting a specific individual, the Person-of-Interest (POI), are the most harmful ones, and, since a public figure is abundantly recorded, a detector can be built from genuine footage of that individual. Such detectors commonly describe a subject through a 3D Morphable Model (3DMM) and adopt its coefficients as a whole, so which part of that description carries the signal has never been measured. We dissect it, holding the encoder, the training corpus and the enrollment protocol fixed and varying only what the encoder observes. The groups of coefficients prove largely redundant, since the shape block alone recovers almost all the accuracy of the full vector, and their temporal evolution contributes a real but bounded amount. We further show that the dense surface the same fit returns, which these detectors discard, carries identity information that the coefficients do not, and that it helps precisely where they are weakest. We assemble the best configuration into MOTIF, a visual-only detector trained on real videos only, with no manipulated video and no POI-specific data. It improves on both state-of-the-art POI detectors in every dataset and manipulation of our benchmark and at two quality levels. Our experimental code will be released at this https URL.

[CV-66] UltraText Bench: A Comprehensive Bilingual Benchmark for Evaluating Visual Text Rendering in Image Generation

链接: https://arxiv.org/abs/2610.09823
作者: Deyuan Liu,Yihao Hu,Jingxuan Zhang,Xingying Li,Jun Xie,Jiacheng Liu,Jungang Li,Yu Huang,Xuanyi Liu,Yue Ding,Zecheng Wang,Lei Zhao,Mingda Wang,Zhenglin Cheng,Peng Sun,Tao Lin
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Dense visual text requires image generators to reproduce long strings across multiple regions with correct placement and legibility. As short-string rendering improves, evaluation must test sustained performance across more demanding scenes. We introduce UltraText Bench, a bilingual benchmark for prompt-only generation of dense visual text. It contains 432 prompts spanning 24 real-world scene categories and three difficulty levels, split equally between English and Chinese. Each human-reviewed prompt supplies exact strings for four to twelve text regions, paired with structured references for their content, placement, and visual attributes. We use the Q-Judger vision-language model to assess each image against the complete reference, reporting text fidelity, text clarity, spatial quality, and scene quality. Across 24 model configurations, these dimensions reveal different strengths: Z-Image-Turbo gains 3.81 clarity points over Z-Image-Base while losing 14.76 fidelity points under the reported settings. Performance also varies with workload; Qwen-Image-2512’s English composite falls from 86.50 at L1 to 42.86 at L3. Ten participants took part in human evaluation of the automatic scores. Repository: this https URL.

[CV-67] Efficient 3D Gaussian Head Avatars for Edge Devices

链接: https://arxiv.org/abs/2610.09821
作者: Umar Farooq,Jean-Yves Guillemaut,Adrian Hilton,Marco Volino
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Generative 3D Gaussian head avatars provide high-quality, efficient rendering, but synthesising the Gaussian representation remains computationally expensive, limiting deployment on resource-constrained and edge devices. We introduce an efficient generator architecture for unconditional 3D Gaussian head synthesis, based on a parameter-efficient synthesis block and depth-wise separable convolutions while retaining style-based conditioning. Our architecture reduces generator complexity without requiring model compression or quantisation. Compared with the baseline model, our approach reduces FLOPs by 94%, parameter count by 70%, and model size by 81%, while maintaining competitive generation quality. We further demonstrate practical CPU inference and browser-based execution on mobile devices using ONNX Runtime, enabling 3D Gaussian avatar synthesis without dedicated GPU hardware or application-specific software. In addition to conventional image-quality metrics, we evaluate multi-view consistency, training cost, and deployment performance. Code, trained models, and evaluation tools will be released publicly.

[CV-68] Concentration Not Uncertainty: Why Targeted Synthetic Data Doesnt Help Camouflaged Object Detection

链接: https://arxiv.org/abs/2610.09807
作者: Akshat Dobhal,Sanjay Singh
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 29 pages, 3 figures, 12 tables

点击查看摘要

Abstract:Camouflaged object detection requires pixel-accurate masks, but obtaining such annotations is slow and costly, making synthetic training images an attractive alternative. Under a fixed generation budget, however, it remains unclear which real-image regions to target for synthetic data generation. We study an uncertainty-guided generation strategy that clusters the unlabelled real images, identifies clusters on which the model is least certain, allocates synthetic generation toward those clusters, and iteratively retrains the model. Across 103 training runs, uncertainty-based targeting does not outperform random allocation. Five independent controls further show that this null result is not an artifact: targeted training sets are measurably different from random sets, but the difference is explained by concentrating the generation budget rather than by where uncertainty is concentrated, as every concentration rule we test reproduces the effect and, on boundary accuracy, so does aiming at the clusters the model was most certain about. Separately, we find substantial data contamination in CHAMELEON, with 50 of its 76 images duplicated from training data despite the standard overlap check reporting zero overlap. Together, these results show that, under a fixed synthetic-data budget, budget concentration, not uncertainty-based targeting, accounts for the observed training-set effects.

[CV-69] DisParQ: Self-Supervised Part Concepts for Interpretable Vision Foundation Models

链接: https://arxiv.org/abs/2610.09802
作者: Adam Pardyl,Siddhartha Gairola,Sukrut Rao,Adam Wróbel,Bartosz Zieliński,Bernt Schiele,Dawid Rymarczyk
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: Under review

点击查看摘要

Abstract:Concept-based vision models represent images through an intermediate layer of human-inspectable concepts, so what a model relies on can be traced to those concepts. However, those models are often limited to fixed categories or depend on language to define their concepts. We introduce DisParQ (Discrete Parts with Quantized attributes), a method that learns spatially grounded, discrete concept representations from a powerful frozen vision-only self-supervised backbone. It requires no class labels and no language supervision. Each image patch is assigned to exactly one concept from a learnable prototype dictionary, and only a sparse subset of concepts may activate per image. To capture how each concept varies across images (e.g., the type of a “wheel”), we learn continuous residuals alongside the concepts and then quantize them into discrete attributes. A spatial decoder reconstructs the backbone’s representation from the concepts and attributes alone, so successful reconstruction means that the discrete representation preserves the backbone’s information. We evaluate DisParQ across seven datasets, from general recognition (ImageNet, PartImageNet, Places) to fine-grained benchmarks (CUB, Cars, Dogs, Flowers). We show that DisParQ closely matches its frozen DINOv2 teacher on ImageNet linear probing (83.2% top-1), achieves higher concept consistency than language-aligned models, remains competitive on fine-grained recognition, and enables cross-category part-based retrieval.

[CV-70] Beyond Masks and Trajectories: Flow-Guided Latent Action Injection for Stable Surgical Video Generation

链接: https://arxiv.org/abs/2610.09800
作者: Tsz-Yui Qin,Siyu Zhou,Chi-Keung Tang,Yuxiang Nie,Shu Yang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Surgical video generation holds substantial potential for surgical education, simulation, and data augmentation, yet generating surgical videos with realistic and clinically plausible motion remains challenging. Most existing methods rely on auxiliary conditions, such as masks, trajectories, depth, or reference videos, to achieve visually plausible synthesis. Yet, these auxiliary conditions typically require additional manual annotation or specialized acquisition, making it difficult to scale such methods beyond small, curated datasets. This motivates the need for a reference-free architecture capable of generating high-quality surgical video without requiring auxiliary visual conditions at inference time. We propose FLAIR, a Flow-guided LatentAction Injection framework for Reference-free surgical video generation. FLAIR learns action priors from optical flow of real surgical videos, dynamically predicts corresponding latent action representation from an input prompt, and injects it into a frozen base model to generate surgical videos with improved action consistency. We further construct SurgActionClip-30K, the first large-scale surgical vision dataset comprising action-centric segmented clips and structured caption labels, addressing the persistent lack of fine-grained, action-centric surgical datasets. Lastly, we introduce SurgMetrics, the first surgical domain-specific evaluation metrics for quantifying the quality of generated surgical videos, addressing the persistent absence of clinically grounded evaluation standards in this domain. Extensive experiments demonstrate that FLAIR enables generating high-quality surgical videos using text-only inference without auxiliary conditions, and validation in SurgMetrics demonstrates its strength in alignment with human perception compared to traditional metrics.

[CV-71] Counterfactual Route Optimization for Gaussian Head Avatar Modeling

链接: https://arxiv.org/abs/2610.09791
作者: Shikun Zhang,Yong Li,Yiqun Wang,Qiuhong Ke,Cunjian Chen
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 15 pages, 7 figures, 4 tables

点击查看摘要

Abstract:Head avatar modeling requires jointly optimizing multiple objectives with different dominant effects on geometry, appearance, and cross-view consistency. However, their relative effectiveness varies across training states, while existing pipelines typically rely on fixed loss weights or handcrafted stage-wise schedules. A central challenge is therefore to identify which optimization direction is more beneficial at each training state. We propose a counterfactual route optimization framework for Gaussian head avatar modeling, which characterizes state-dependent optimization preference from the realized effects of alternative updates rather than predefined heuristic weighting. Starting from the same training state, we perform short-horizon route-restricted lookahead over geometry, appearance, and joint update routes and evaluate their outcomes under a unified utility. The resulting counterfactual evidence is factorized into a geometry–appearance preference and a residual joint advantage, separately capturing the relative preference between individual update directions and the additional benefit of coordinated optimization. We further amortize this offline evidence into a lightweight controller that directly estimates the current optimization preference and applies bounded modulation to the training objectives during full avatar optimization. Experiments on the NeRSemble dataset validate the effectiveness of the proposed design, consistently outperforming existing methods while preserving clearer local facial structures and finer details.

[CV-72] UltraWorld: Learning Interactive Ultrasound World Models from Untracked Clinical Videos with Acoustic Sampling Map

链接: https://arxiv.org/abs/2610.09785
作者: Keke Yang,Erqi Wang,Sainan Guan,Hongliang Ren
类目: Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)
备注:

点击查看摘要

Abstract:World models can enable autonomous ultrasound scanning by predicting the outcomes of probe motions from local observations. Learning this action–observation relationship typically relies on synchronized video–pose pairs, which are costly to collect at scale and largely unavailable in routine clinical recordings. Reliable action following further requires modeling ultrasound’s cross-sectional sampling geometry. We present UltraWorld, a self-distillation recipe that transfers priors from clinical ultrasound videos into interactive world models without real action annotations. Starting from clinical videos, we adapt a video foundation model into an ultrasound generator conditioned on reference images and anatomical masks. Anatomical masks sampled along programmable trajectories through 3D anatomy provide spatial guidance for synthesizing action–video pairs. We then use these synthetic pairs to self-distill the generator into a world model that predicts future observations from local observations and actions, without requiring anatomical masks or other 3D assets at inference time. To further improve action following, we introduce the Acoustic Sampling Map (AsMap), which represents probe poses and imaging settings as pixel-wise 3D sampling positions, beam directions, and depths. Experiments demonstrate improved prediction fidelity and action following. Across nine simulated closed-loop local planning episodes, UltraWorld reduces the mean final distance to the goal and orientation error by 29% and 38%, respectively, compared with visual servoing. Project Page: this https URL.

[CV-73] CIRSeg: Coarse-to-Fine Intensity-Robust Liver Segmentation with Source-Free Continual Test-Time Adaptation MICCAI2026

链接: https://arxiv.org/abs/2610.09784
作者: Ruoshi Xu,Mingqi Gao,Shengda Luo,Jingkun Chen
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted at the CARE 2026 Workshop at MICCAI 2026

点击查看摘要

Abstract:Reliable liver segmentation in contrast-enhanced MRI is essential for quantitative hepatic assessment, treatment planning, and longitudinal disease monitoring. However, limited annotated data and scanner- or vendor-dependent intensity variations can cause overfitting and poor generalization to unseen acquisition domains. Moreover, simultaneously achieving robust global localization and precise boundary delineation remains challenging, while predictions may contain isolated false-positive regions outside the main liver component. To address these challenges, we propose CIRSeg, a coarse-to-fine, intensity-robust liver segmentation framework based on nnU-Netv2. CIRSeg combines 3D CutMix with stochastic intensity transfer using either Nyul augmentation or histogram matching to improve robustness to heterogeneous MRI intensities. Its cascaded architecture decouples low-resolution anatomical localization from full-resolution boundary refinement. At inference, source-free test-time adaptation based on confidence-filtered predictions and probability-prior regularization further improves robustness to out-of-distribution inputs. As a final deterministic post-processing step, largest connected component filtering removes isolated false-positive regions. On the CARE 2026 test set, CIRSeg achieves Dice scores of 97.13% and 97.93% on the in-domain and unseen-domain subsets, with corresponding HD95 values of 20.18 mm and 11.30 mm, respectively. These results demonstrate consistently accurate segmentation across both in-domain and unseen acquisition settings. The code is available at this https URL

[CV-74] Relational Abstractions for Spatial Reasoning with Diffusion Models

链接: https://arxiv.org/abs/2610.09780
作者: Ana Ezquerro,Ozan Özdenizci
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Diffusion models excel at image synthesis, but they remain limited in their ability to reliably satisfy structured spatial reasoning constraints. In conditional data distribution modeling tasks with implicit logical structure, such as puzzles defined by visible clues paired with consistent solutions, state-of-the-art generative models tend to approximate pixel-space distributions without learning the underlying logical rules required for inference. To address this limitation, we present a novel framework for spatial reasoning with diffusion models that leverages unsupervised object discovery and abstractions of object relations. We show that the relational knowledge derived from object-centric representations enriches diffusion models with structural primitives, allowing them to effectively guide the generative representation space during both training and inference, and enabling conditional image generation that satisfies reasoning constraints. Additionally, we introduce a large-scale generative spatial reasoning benchmark with four datasets inspired by human-solvable puzzles. Our results show that relational abstractions significantly improve reasoning capabilities of diffusion models on a variety of complex reasoning tasks, while enabling robust generalization in out-of-distribution settings.

[CV-75] Diffusion-Generated Image Watermarking: A Two-Axis Taxonomy and Three Protocol-Bounded Case Studies ECCV2026

链接: https://arxiv.org/abs/2610.09755
作者: Sung Ju Lee,Nam Ik Cho
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 17 pages, 2 figures. Accepted to the non-archival track of the ECCV 2026 LifeGenIP Workshop. English translation with partial reorganization of our article in Journal of Broadcast Engineering 31(4), 687-699 (2026)

点击查看摘要

Abstract:Watermarking diffusion-generated images requires balancing provenance signals with image quality, robustness, and computational cost. This work organizes methods along two axes: insertion mechanism and primary signal-bearing representation, and formalizes a representative z_T -Fourier pipeline for verification and identification. We then use the taxonomy to structure three protocol-bounded case studies. The first examines associations among frequency integrity, detection, quality, and cropping behavior. The second revisits persistence under seed-linked and seed-independent editing and formulates a scoped Semantic Imprinting Hypothesis without claiming a localized carrier or causal mechanism. The third studies single-shot VAE-latent phase modulation, including its efficiency, regeneration robustness, and robustness–quality operating points. Finally, we separate four content-level attack families from model/pipeline adaptation, propose corresponding evaluation protocols and testable conjectures for parameter-tuning threats, and identify additional temporal extensions for video. These analyses do not establish a universal ranking; instead, they provide a framework for matched, protocol-aware comparisons of watermarking systems for diffusion-generated images.

[CV-76] Flow-of-Thought: A Framework for Visual Reasoning NEURIPS

链接: https://arxiv.org/abs/2610.09746
作者: Mariia Baidachna,Nicolas Pugeault
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: NeurIPS WiML 2026 version: OpenReview version: this https URL

点击查看摘要

Abstract:Mental imagery, ``seeing with the mind’s eye’’ is an essential aspect of human cognition. Despite rapid progress Large Language Models (LLMs) and Vision Transformers (ViTs) still underperform on tasks requiring spatial understanding. To address this, we introduce Flow-of-Thought (FoT), a framework that integrates the generation of visual sketches as intermediate reasoning steps, mimicking mental imagery in humans. We train coordinate-aware trajectory flow fields on SO(2) group orbits and cumulative shortest paths, then freeze the learned dynamics; same vs. different decisions compare competing generative hypotheses using foreground-weighted reconstruction energy. On locked tests FoT reaches 100.0% accuracy on Tetris and 99.0% on colored shapes. Under frozen transfer, the orbit-trained 2D flow improves over its endpoint-only control on BLINK Multi-view (72.2% vs. 63.9% on 133 public validation pairs), supporting continuous visual traces as an effective and interpretable representation for spatial reasoning in some out-of-distribution settings.

[CV-77] Beyond Group Splits: Specimen-Level Cross-Validation and Visual Attribution for Remaining-Shelf-Life Regression in Climacteric Fruit

链接: https://arxiv.org/abs/2610.09726
作者: Rovhona Mudau,Jean Frederic Isingizwe Nturambirwe,Clement Nthambazale Nyirenda
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 7 pages, 1 figure, 4 tables. Accepted for physical presentation at the 10th IEEE Conference on Information Communications Technology and Society (ICTAS 2026), 14-16 October 2026, Durban, South Africa

点击查看摘要

Abstract:Estimating remaining shelf life (RSL) from images could provide affordable decision support for perishable produce, but evaluation protocols can substantially affect reported performance when repeated images are available from the same biological specimen. We use the Hass Avocado Ripening dataset, comprising 8,834 image-RSL pairs from 426 fruits across three storage regimes, to evaluate a frozen ImageNet-pretrained visual backbone with a lightweight regression head. Our contributions are threefold: we quantify the effect of observation-level versus specimen-disjoint evaluation, compare lightweight and heavier visual backbones under specimen-disjoint cross-validation, and examine their spatial attributions using Grad-CAM. Across ten observation-level random splits, the model achieves a mean RMSE of 2.37 days with a standard deviation of 0.03 days, whereas specimen-disjoint 5-fold cross-validation yields a mean RMSE of 3.12 days with a standard deviation of 0.11 days. The corresponding mean coefficient of determination is 0.553. A matched per-specimen comparison confirms higher error under specimen-disjoint evaluation, with a probability value below 0.001 across 426 specimens, showing that observation-level partitioning gives a substantially more optimistic estimate for this dataset and model configuration. Under specimen-disjoint evaluation, MobileNetV3-Small (0.93 million parameters) achieves accuracy comparable to ResNet-18 while providing substantially higher throughput, and Grad-CAM reveals differences in spatial attribution between the lightweight backbones. These results support specimen-disjoint evaluation and attribution analysis when assessing lightweight vision models for longitudinal shelf-life prediction.

[CV-78] MeshCarve: Artisan Mesh Generation with Flow Matching in Compact Latent Spaces

链接: https://arxiv.org/abs/2610.09723
作者: Xiyu Wang,Ruocheng Wu,Yufei Wang,Zhihao Li,Lanqing Guo,Bihan Wen
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 18 pages, 6 figures, 9 tables

点击查看摘要

Abstract:Prior artisan mesh generation works largely predict face tokens autoregressively, which makes inference slow. Recent methods instead flow match continuous latents built by Variational AutoEncoders (VAEs), but reconstruction quality drops significantly when geometry and topology are jointly encoded, and further when the latent space is compressed. We present MeshCarve, a flow matching method that generates entirely in compact latent spaces, generating vertex positions and edge connections separately and sidestepping the difficulty of a joint compact latent. To shorten the token sequence, we propose a hierarchical sparse transformer backbone, instantiated as VertexVAE and EdgeVAE. Instead of encoding fields over the surface voxels, both VAEs anchor on discrete vertices in their latent spaces, which drastically reduces the token sequence length, and our spatial-aware compression shortens it further without costing reconstruction. VertexVAE directly encodes vertex occupancy. For connectivity, we propose vertex-link encoding, which turns arbitrary connectivity between vertices into fixed-length continuous per-vertex embeddings and recovers complex artistic topology faithfully. MeshCarve combines these VAEs with an anchor generator and flow matches on the shortened token sequences. It shows advantages over state-of-the-art autoregressive and flow matching methods on Objaverse and generalizes to Toys4K. To the best of our knowledge, it is among the first artisan mesh generation methods whose every generative stage runs in a spatially compressed latent, with a token sequence only a fraction of the most compressed previous autoregressive and flow matching works.

[CV-79] DynStream: Online Streaming 4D Gaussian Reconstruction of Dynamic Worlds from Unposed Video

链接: https://arxiv.org/abs/2610.09720
作者: Dingwei Xian,Xiaoyu Zhou,Yajiao Xiong,Yongtao Wang,Ming-Hsuan Yang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Online reconstruction of dynamic 4D scenes from long, unposed streaming videos requires both continuous processing and photorealistic rendering, which existing methods struggle to achieve simultaneously. Existing feed-forward Gaussian methods are restricted to offline processing, whereas online point-cloud approaches struggle to maintain dense geometry and high-fidelity rendering. We present DynStream, a framework for streaming 4D Gaussian reconstruction from long, unposed videos. Given a continuous video stream, DynStream reconstructs the scene within local temporal windows and incrementally aligns and fuses these local reconstructions into a globally consistent scene, enabling online 4D reconstruction without per-scene optimization. By jointly enforcing cross-window geometric consistency and modeling time-varying scene content, DynStream supports efficient reconstruction and photorealistic rendering over extended video streams. Experiments demonstrate that DynStream enables high-fidelity online dynamic reconstruction and rendering from long video streams, achieving state-of-the-art performance across diverse dynamic indoor and outdoor scenes.

[CV-80] YUBI-STAG: Contact and Semantic-Rich Alignment for VLAs via Automated Video-Language Grounding

链接: https://arxiv.org/abs/2610.09718
作者: Masatoshi Tateno,Takehiko Ohkawa,Yueh-Hua Wu,Hanlong Li,Tatsuya Matsushima,Yoichi Sato,Kei Ota
类目: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
备注: Project page: this https URL

点击查看摘要

Abstract:Vision-Language-Action (VLA) models acquire broad manipulation capabilities via large-scale pretraining, yet eliciting them through language requires fine-grained alignment between instructions and physical interactions. Existing robot demonstrations typically provide only coarse task descriptions, omitting how actions are executed, including which gripper acts, which object is contacted, and how it is grasped and moved. We introduce YUBI-STAG, a framework for Spatio-Temporal Annotation and Grounding that automatically enriches manipulation demonstrations with interaction-rich semantics to align pretrained VLAs with fine-grained manipulation language. Combining contact-object segmentation with vision-language models, YUBI-STAG annotates object identities, attributes and states, per-gripper actions, bimanual coordination, and spatially grounded interactions. To address YUBI-STAG’s reliance on localized sequences and multi-stage VLM inference, we distill it into YUBI-VLM. YUBI-VLM directly recovers action structure and annotations from raw, unsegmented video in few inference calls and operates from wrist views alone. We evaluate both frameworks on YUBI-STAG-Bench across temporal, semantic, and spatial grounding tasks. YUBI-VLM retains much of YUBI-STAG’s annotation accuracy with fewer inference calls and shorter runtime while generalizing to unseen manipulations. Finally, post-training VLA policies on these annotations aligns them with fine-grained language and contact-aware structure. Bimanual experiments demonstrate improved performance and instruction following, including control over object identity, acting gripper, target location, and spatial relations absent from original labels.

[CV-81] Enhancing Multi-Region Stylization with Interior-Guided Boundary Repair

链接: https://arxiv.org/abs/2610.09706
作者: Hong-Son Nguyen,Thi-Ngoc-Hanh Le
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 10 pages, 7 figures

点击查看摘要

Abstract:Region-based neural style transfer enables fine-grained artistic control by allowing independent stylization of semantic image regions. However, compositing these regions often leads to boundary artifacts, degrading visual quality. We propose Interior-Guided Boundary Repair (IGBR), a lightweight and model-agnostic method that improves boundary handling in multi-region stylization. IGBR repairs boundary pixels using interior-guided propagation and applies inward, distance-based blending restricted to object-background boundaries, preventing inter-object style leakage. The method is derived from a region-wise constrained formulation with a closed-form solution and can be seamlessly integrated into existing stylization pipelines without retraining. To evaluate efficiency of our IGBR, we introduce quantitative metrics that measure boundary consistency, gradient artifacts, inter-object leakage, and interior preservation without requiring annotated stylized images. Our experiments and evaluations demonstrate that the proposed IGBR consistently produces plausible boundaries, outperforming prior blending techniques in boundary consistency, gradient stability, and interior preservation. The code is available at this https URL.

[CV-82] Latent Watermarks under Generative Editing: A Benchmark and Analysis of Detection Survival

链接: https://arxiv.org/abs/2610.09702
作者: Sung Ju Lee,Nam Ik Cho
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Ordinary prompt-based editing can cause latent watermark detection to fail without explicitly targeting the watermark. We benchmark eight watermark methods against five editors across four generative backbones, four editing strengths, and five semantic categories, with edit-validity and threshold checks. Separating editing from seven subsequent distortions reveals that editing alone primarily distinguishes Tree-Ring, while added distortions expose a broader spectrum of detection survival. Sequential edits reveal a second hidden difference: score separation can decline while detection rates remain near their ceiling. Across methods, standardized clean score separation ( d’ ) organizes composite-survival tiers, whereas spatial overlap adds little to predicting edit-only survival beyond clean detectability. Embedding-strength interventions in two methods link higher clean separation to higher post-edit separation. In HSTR, the margin contrast is positive, while the angular layout contrast at matched clean separation remains unresolved. Together, outcome decomposition and continuous separation expose differences hidden by aggregate TPR. Method tiers are stable under threshold recalibration at the main operating points and alternative composite weights. Clean d’ is thus a useful empirical diagnostic within this benchmark, with mixed transfer to unseen methods. Code and supporting artifacts are planned for a separate release.

[CV-83] What Makes Synthetic Hard Negatives Work in Vision-Language Pretraining? ACCV2026

链接: https://arxiv.org/abs/2610.09700
作者: Nikos Giakoumoglou,Paschalis Giakoumoglou,Andreas Floros,Kleanthis Marios Papadopoulos,Tania Stathaki
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注: ACCV 2026

点击查看摘要

Abstract:Synthetic hard negatives generated in the representation space have proven effective for unimodal self-supervised learning, but transferring this idea to vision-language pretraining is not straightforward. We analyze six representation-space synthesis strategies and identify two failure modes in their transfer to vision-language pretraining: cross-modal constructions that produce overly easy negatives or pull them toward the query, and intra-modal constructions that incorporate the matched positive. We also observe logit-scale saturation when training with synthetic hard negatives and a learnable temperature, and find that fixing the temperature improves downstream performance. Using this geometric analysis we propose SNAP, which generates intra-modal hard negatives that never involve the positive from either modality, avoiding both failure modes entirely. SNAP is model-agnostic, requires no external generative models, and adds less than 10% training time overhead. Evaluated on top of CLIP and FLIP across multiple architectures and datasets, SNAP delivers consistent improvements on zero-shot retrieval, zero-shot classification, and linear probe evaluation.

[CV-84] Do Better Visual Representations Always Lead to Better End-to-End Autonomous Driving?

链接: https://arxiv.org/abs/2610.09695
作者: Zihao Zhang,Haochen Tian,Tianyu Li,Changhui Jing,Jingliang He,Naisheng Ye,Ziyuan Pu,Zhenjie Yang
类目: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Visual foundation models (VFMs) are increasingly integrated into end-to-end autonomous driving for their powerful representations, yet it remains unclear when these representations improve driving performance. To investigate this question, we introduce ViRA, a planner-agnostic visual representation alignment framework that keeps the planner architecture and inference cost unchanged. Our study reveals three findings: (1) VFM-guided visual representations consistently improve driving performance across diverse end-to-end planners, with gains extending to zero-shot closed-loop evaluation. (2) The choice of VFM target matters for planning performance, and alignment to a different VFM can further benefit planners with pre-trained VFM encoders. (3) Auxiliary perception supervision reduces sensitivity to VFM target selection, narrowing the EPDMS spread across five targets from 2.7 to 0.5 points and potentially compensating for less effective VFM targets. Guided by these findings, we develop ViRA-Diffusion, a diffusion-based planner trained without auxiliary perception supervision, which achieves 92.3 EPDMS on NAVSIM v2 navtest, outperforming recent methods in our comparison by at least 1.9 points. The results motivate jointly considering target selection and planner supervision when integrating VFMs into end-to-end autonomous driving. The results and demo are available at this https URL.

[CV-85] A Multi-Source Ultrasound Benchmark Revealing the Limits of Contemporary Self-Supervised Anomaly Detection Methods

链接: https://arxiv.org/abs/2610.09677
作者: Marco Riedenauer,Daniel Kienzle,Pratik Mayekar,Rainer Lienhart
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注: 7 pages, 3 figures

点击查看摘要

Abstract:Self-supervised anomaly detection is a promising paradigm for medical ultrasound, as normal images are often easier to obtain than exhaustive annotations of all possible pathologies. However, most existing evaluations are limited to a single anatomy or task, making it unclear whether models learn a robust notion of normal ultrasound appearance or only a source-specific representation. We introduce the SADUSI benchmark, a multi-source ultrasound dataset designed to train and evaluate anomaly detection methods across a broad range of anatomical regions, views, and acquisition protocols. The goal of SADUSI is to provide a diverse normal ultrasound distribution and a benchmark for visible structural anomalies that can be assessed from single images. We evaluate representative self-supervised anomaly detection methods and find that current approaches struggle in this setting. In particular, reconstruction-based diffusion methods such as AnoDDPM and DeCo-Diff achieve pixel-level AUROC values of 0.56-0.72 and maximum F1 scores of 0.10-0.26, indicating limited separation of pathology from normal image regions. Feature-based PatchCore variants perform better, reaching pixel-level AUROC values of 0.76-0.83, but remain limited with maximum F1 scores of 0.14-0.40. These findings suggest that broad multi-source ultrasound anomaly detection remains an open challenge and that SADUSI can serve as a resource for developing methods that generalize beyond anatomy-specific settings.

[CV-86] Quasi-Binarized Autoencoders: An Architecture-Independent Information Bottleneck for Medical Image Anomaly Detection

链接: https://arxiv.org/abs/2610.09670
作者: Shouhei Hanaoka,Takahiro Nakao,Atsushi Takamatsu,Takeharu Yoshikawa,Osamu Abe
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Unsupervised anomaly detection, which learns only from normal images, is a central task in medical image analysis and remains an open problem. Reconstruction-based methods pass an image through an encoder-decoder network trained on normal data and detect anomalies from the residual between the image and its reconstruction. This works only if the information passed from the encoder to the decoder is limited; otherwise the network learns an identity mapping and reconstructs anomalies too. This limit is usually imposed through architectural choices, tuned per dataset, that cannot be stated in bits. We introduce the quasi-binarizing (QB) layer, which squashes each latent element into [0, 1] and adds Laplace noise of scale 1/epsilon. Each element is then epsilon-locally differentially private, and the mutual information between an image and its reconstruction is bounded by a quantity that depends only on epsilon and the number of QB elements, whatever the encoder and decoder. Placing a QB layer on every encoder-decoder path, including all skip connections, we build QBAE, a seven-level attention U-Net with 32,768 QB elements. On the seven datasets of the MedIAnomaly benchmark, QBAE with one architecture and one configuration reaches a mean image-level AUROC of 0.828, the highest among methods that do not adapt to each dataset, and the best reported results on BraTS2021 (AUROC 0.911, pixel-level AP 0.838). The noise is kept at test time, so that every reconstruction satisfies the bound. Without input corruption, the bottleneck alone prevents identity collapse (mean AUROC 0.805 vs. 0.590). Code is available at this https URL.

[CV-87] Identity-Duplication Auditing in National-Scale Neuroimaging Repositories

链接: https://arxiv.org/abs/2610.09614
作者: Jiheng Li,Michael E. Kim,Trent M. Schwartz,Yuhan Cui,Gaurav Rudravaram,Derek B. Archer,Timothy J. Hohman,Lori L. Beason-Held,Victoria L. Morgan,Dario J. Englot,Angela L. Jefferson, for theAlzheimer’s Disease Neuroimaging Initiative, for theBIOCARD Study team, for theHealth,Aging Brain Study:Health Disparities(HABS-HD)Study Team,Lianrui Zuo,Guray Erus,Christos Davatzikos,Bennett A. Landman
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:National-scale magnetic resonance imaging (MRI) repositories increasingly integrate data from different studies and institutions. However, subject identifiers that are valid only within individual datasets are no longer guaranteed to remain globally unique after aggregation, making it possible for the same subject to be assigned multiple identifiers, which we define as identity duplication. Such duplication can create leakage between training and test data and inflate apparent performance in downstream biomedical studies. Existing methods do not provide an end-to-end, image-based workflow for auditing this problem at repository scale. In this work, we present HAPPEN, a human-in-the-loop pipeline for auditing identity duplication in T1-weighted brain MRI repositories. It combines SHA-256 fingerprinting for exact-duplicate detection with supervised contrastive retrieval of non-identical scans that may originate from the same person. Retrieved pairs are reviewed as candidates in a locally hosted interface rather than automatically classified as duplicates. We deployed the workflow in a 95,129-scan aggregated repository and assessed end-to-end recovery using 54 genetic-reference pairs. Transferability was assessed by locally deploying the same workflow on 22,386 scans at an independent institution without model retraining or image transfer. Deployment in the study repository identified 1,316 exact-duplicate scan groups and 1,275 reviewer-supported near-duplicate subject groups. Of these groups, 56% and 82%, respectively, crossed dataset boundaries. All 54 genetic-reference pairs were recovered. The external team independently completed the full workflow using a locally selected operating threshold and review standard.

[CV-88] STORK: Spatio-Temporal Observation of uterine contRactions via neural networKs MICCAI2026

链接: https://arxiv.org/abs/2610.09598
作者: Melissa Schween,Tristan Gottwald,Jordina Aviles Verdera,Lisa Story,Mary Rutherford,Jana Hutter
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted at the PIPPI Workshop at MICCAI 2026 and will appear in the workshop proceedings (Springer)

点击查看摘要

Abstract:Uterine contractions in fetal MRI are typically identified manually and discarded, limiting insights into contraction dynamics. We formalize Uterine Contractile Activity Detection (UCAD) as a weakly-supervised learning problem and introduce STORK, a multi-instance learning model trained on dynamic MRI series using only coarse, series-level labels. STORK factorizes 3D spatio-temporal convolutions into parallel branches across temporal hyperplanes to capture coherent tissue motion without the cost of full 4D convolutions. Per-frame embeddings, combining intensity and Demons-estimated displacement fields, are aggregated by a linear mean-pooling head. This ensures that frame-level contraction scores can be recovered post-hoc without frame-level training supervision. Evaluated on around 700 multi-vendor dynamic fetal MRI series, STORK achieves a series-level AUROC of 95.0% and AUPRC of 94.6%, substantially outperforming 3D ResNet and ConvNeXt baselines. Grad-CAM analysis suggests that the model draws on predictive features extending beyond the placenta into the uterine tissue, offering an automated tool for richer phenotyping of uterine behavior.

[CV-89] Gradient-Based Trajectory Optimisation over Continuous Poses for Sparse-View Cone-Beam CT

链接: https://arxiv.org/abs/2610.09579
作者: Linda-Sophie Schneider,Simon Wittl,Gabriel Herl,Andreas Maier
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Submitted to TPAMI

点击查看摘要

Abstract:Trajectory optimisation for cone-beam computed tomography (CT) determines which information sparse-view scans acquire. Fixed candidate pools prevent off-grid refinement and require new object-specific precomputation for each acquisition manifold. We make every source pose an individual continuous variable and move all poses jointly by gradient ascent on the scanner’s kinematic manifold. The objective combines soft-Tuy plane coverage, continuous View Covariance Loss, and an analytic attenuation-aware ray-bundle penalty. The same optimiser handles circular, limited C-arm, two-axis, and freesphere parametrisations. On a Defrise flange, continuous selection recovers laminar defects invisible to a circular orbit, matches discrete swap search on the free sphere at the sparser budget, and leads at the denser one, with the same objective evaluated in every arm. A moderate elevation band already recovers most of the free-sphere gain at the defects, so the same optimiser transfers to bounded scanner envelopes. Photon noise preserves the ordering on the flange and compresses it on a dense fuel nozzle. Sparseprescan planning benefits from matching prescan and planned acquisition manifolds. Selection takes seconds rather than minutes without an object-specific reconstruction basis. Prescan-planned poses were executed on a robot CT bench and reconstructed in a common frame, demonstrating feasibility but no consistent metric gain over uniform band sampling. Continuous pose optimisation incorporates attenuation and scanner constraints directly into sparse-view acquisition design.

[CV-90] Visual Evidence Under Cross-Examination: Evaluating and Controlling Decision-Level Evidence Use in Vision-Language Models

链接: https://arxiv.org/abs/2610.09550
作者: Huiyao Zhang,Jin Bai,Zilong Su,Rui Guo,Chaofan Qin,Jinze Lv,Wenhui Yu,Hongfei Wang,Ye Li
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Vision-language models increasingly reason through crops, regions, and tool-produced observations. Yet an observation can influence the answer without benefiting the candidate it supports. We study candidate-bound visual contribution: valid evidence should help, invalidating its supporting relation should remove its additional effect, and valid rebinding should redirect that effect to the newly supported candidate. We introduce CROSS-Bench, a benchmark of 28,000 decision problems, with matched invalidation and rebinding tests on a dedicated evaluation subset. Our RIVET interface preserves evidence identity and uncertainty, composes a candidate-conditioned response, and separately controls its strength. Shared-evidence experiments show that task accuracy and evidence ownership can diverge. Under matched capacity and training, RIVET increases normalized effect transfer from 0.512 to 0.651 where clean evidence has a positive effect. The advantage persists on common evaluation examples and across repeated decision-layer fits. With evidence predicted from raw inputs, RIVET improves CROSS-Bench accuracy by an average of 5.70 pp across four frozen backbones, relative to the same models without auxiliary evidence. These results separate the utility of visual evidence from the candidate-specific destination of its effect.

[CV-91] WAPR: A Foundation Model for Wide-Angle Refinement in Unseen Object Pose Estimation ECCV2026

链接: https://arxiv.org/abs/2610.09535
作者: Yulin Wang,Mengting Hu,Hongli Li,Jianghao Zhou,Chen Luo
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Robotics (cs.RO)
备注: Accepted to ECCV 2026. 19 pages, 4 figures. Yulin Wang and Mengting Hu contributed equally. Corresponding author: Chen Luo

点击查看摘要

Abstract:Real-world applications require 6D pose estimation to be accurate, fast, and scalable to unseen objects. This paper introduces WAPR, a zero-shot wide-angle pose refinement model that refines candidate poses with rotational deviations up to 90 degrees. With as few as 12 candidate poses per detected object instance, WAPR supports fast inference within 1 s per frame and reaches a pose-estimation throughput of up to 25 detected object instances per second. To support wide-angle training for rotationally symmetric objects, WAPR uses rotational symmetry priors to canonicalize symmetry-equivalent pose targets before loss computation. We further construct SA6D, a large-scale 6D training dataset with such priors. SA6D obtains KASAL-assisted rotational symmetry priors for 944 GSO scans and expands them through geometry and texture augmentation into about 50K augmented object instances and about 2M rendered RGB-D images. In addition, an angle-balanced loss stabilizes learning across different angular ranges by reducing the influence of uninformative large-error cases. Experiments on seven BOP core datasets show that WAPR achieves state-of-the-art performance in unseen-object 6D pose localization and detection under both fast and unconstrained inference settings. Project page: this https URL.

[CV-92] KASALv2: Fully Automatic 3D Rotational Symmetry Classification and Axis Localization CVPR2026

链接: https://arxiv.org/abs/2610.09534
作者: Mengxin Zhang,Yulin Wang,Chen Luo,Yongzhe Li,Yijun Zhou
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Graphics (cs.GR)
备注: CVPR 2026. 10 pages, 4 figures. Mengxin Zhang and Yulin Wang contributed equally. Corresponding authors: Chen Luo and Yijun Zhou

点击查看摘要

Abstract:Rotational symmetry is an important prior in 6D pose estimation, improving pose accuracy and supporting symmetry-aware evaluation. However, current symmetry annotations for 3D objects remain largely manual or semi-automatic, often requiring predefined types or orders, which limits scalability. This work introduces a fully automatic, reference-free framework for symmetry-type classification, rotational-order identification, and full-axis localization across all eight canonical 3D rotational symmetry types. The method localizes a dominant high-order axis, infers its rotational order through self-consistency analysis, and reconstructs the complete symmetry structure under a hierarchy-guided formulation. A texture-aware extension further models appearance-induced reductions in rotational order while preserving axis orientations. Experiments on idealized and real-world datasets demonstrate strong accuracy and generalization, achieving 94.75% accuracy on 438 symmetric objects in GSO. Training FoundationPose with these priors improves accuracy by up to 0.9% across five BOP datasets, showing that automatically estimated rotational priors improve downstream 6D pose estimation. Code is available at this https URL.

[CV-93] ActiveLang: Active Open-Vocabulary 3D Mapping with Semantic-Uncertainty-Guided Exploration

链接: https://arxiv.org/abs/2610.09518
作者: Liyan Chen,Hairong Yin,Huangying Zhan,Yi Xu,Raymond A. Yeh,Philippos Mordohai
类目: Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)
备注:

点击查看摘要

Abstract:As robots increasingly assist humans with diverse tasks, they need both geometric and semantic understanding of their surroundings. Moreover, robots often operate in unfamiliar environments and take on new tasks without knowing the relevant concepts ahead of time. This motivates language-annotated 3D maps that support open-vocabulary scene understanding and human-robot interaction. We introduce ActiveLang, an autonomous system for active open-vocabulary 3D mapping with semantic-uncertainty-guided exploration. ActiveLang performs online language-feature adaptation on a compact dual-Gaussian representation to jointly reconstruct scene geometry, appearance, and open-vocabulary semantics with modest memory overhead. Its planner efficiently selects informative viewpoints, enabling effective mapping with fewer observations and lower computational cost. Experiments on Replica and ScanNet++ demonstrate substantial improvements in 2D and 3D open-vocabulary segmentation over both online and offline baselines, highlighting that actively exploring scenes builds language-annotated 3D maps more efficiently.

[CV-94] It Is Not Seeing the Hazard: A Frozen Vision-Language Safety Score Measures Its Caption Bank

链接: https://arxiv.org/abs/2610.09517
作者: Samuel Tetteh,Cody Fleming
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Frozen vision-language models increasingly provide safety signals for reinforcement learning. Their use assumes that similarity to language describing danger indicates the hazard itself. Yet policy return and collision rate cannot reveal whether a score detects hazards or responds to correlated features of the scene. VLM-based methods have reported gains in driving and safe-RL benchmarks by converting image-text similarity into rewards, costs, or confidence weights. Such signals promise to reduce reliance on manually designed feedback. They may also reflect prompt structure, embedding geometry, or camera viewpoint, leaving their safety meaning unverified. To address this gap, we present a controlled evaluation of a frozen CLIP prompt-margin safety score. We apply the score to trajectories generated by policies that never receive it, match pre-contact observations to contact-free observations with comparable hazard geometry, and vary the captions, encoder, and camera view. Across three policies, 180 episodes, and 130 isolated contact onsets, the score decreases for about twenty steps before contact. Mechanism controls indicate that the score mainly tracks resemblance to the scene shared by its captions and changes with caption separation and camera view. A constant-confidence control retains the lower catastrophe-rate point estimate, so policy gains do not establish hazard perception.

[CV-95] STRIKE: Learning Visual State Transitions for Physical World Modeling

链接: https://arxiv.org/abs/2610.09514
作者: Wenbin Teng,Tianshuo Xu,Depu Meng,Yuelei Li,Quentin Herau,Yihan Hu,Yajie Zhao,Wei Zhan
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Physical world modeling requires predicting how interactions change a scene, not merely generating coherent motion. We propose STRIKE, a framework that separates visual state transition learning from dense video generation. We construct event-aligned supervision by extracting observed states from training videos and pairing them with transition descriptions and temporal offsets. An image-based transition model learns to predict the next scene configuration from the current image, a local transition specification, and elapsed time. At inference, a pretrained vision-language planner predicts time transition specifications, and recursive application of the learned transition model produces a sequence of future visual states. A separately trained dynamic model then generates the complete rollout conditioned on these states and their temporal locations. Experiments on Physics-IQ Verified, PhyGenBench, Pisa-Experiments, and RoboTwin2.0 show improvements of STRIKE over the corresponding video-backbone baselines in benchmark measures of physical consistency and manipulation-video fidelity. These results support learned visual state transitions as an effective intermediate representation for physical world modeling.

[CV-96] OmniCam: Omni-Camera Trajectory Generation via Geometry-Grounded Pose Token Learning

链接: https://arxiv.org/abs/2610.09513
作者: Zhenyang Liu,Chenjie Cao,Yisu Zhang,Xuhui Zuo,Xiangyang Xue,Yanwei Fu,Tengfei Wang,Chunchao Guo
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Camera trajectories control viewpoint changes in video generation, scene reconstruction, and robotic perception. Generating them from language requires both scene geometry and target-aware framing. We introduce OmniCam, an autoregressive model that generates camera pose sequences from a single panorama and textual trajectory descriptions. Its geometry-grounded pose token learning combines three components: a panoramic point-cloud encoder for omnidirectional geometric context; hybrid absolute-rotation and relative-translation tokenization with temporally consistent quaternion signs; and separate geometric and semantic conditioning streams with an explicit 3D target anchor. We also construct OmniCaT, containing 267,700 trajectories across four camera behaviors. On the reported OmniCaT evaluation, OmniCam reduces trajectory errors by 28–47% and collision rate by 65.8% relative to GenDoP retrained on OmniCaT. Against the best baseline for each metric, the ATE and collision reductions are 43.0% and 62.3%, respectively. Component ablations support the use of geometric and target-aware conditioning, while downstream experiments examine camera-controlled video generation and robotic active perception.

[CV-97] LiG-DETR: Local-in-Global Reassembly in Latent Space for Aerial Object Detection

链接: https://arxiv.org/abs/2610.09511
作者: Yupeng Zhang,Fangzhuo Gao,Juntao Cheng,Ziyi Zhao,Liang Wan,Ruize Han
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Aerial object detection faces substantial scale and density variations. Small objects are easily degraded by downsampling and feature compression, while medium and large objects require sufficient global context. Existing methods mainly follow two paradigms: image slicing provides clearer local evidence but relies on independent crop-level prediction and post-processing, whereas feature- and query-level optimization preserves unified inference but operates on already compressed full-image representations, limiting recovery of fine-grained information. This raises a key question: can aerial detection directly acquire high-fidelity local evidence before feature degradation and integrate it into a unified end-to-end framework? To this end, we propose LiG-DETR, an Efficient Global-Local Reassembly framework that reformulates image slicing as high-fidelity local feature acquisition. A shared encoder extracts global and locally magnified features, which are projected into the detector feature space. The projected local features are reassembled according to their original spatial locations to form a globally aligned local feature level, and a single DETR decoder jointly decodes global and local features. To reduce redundant computation, Context-Preserved Selective Reassembly focuses high-resolution encoding on informative regions while preserving a dense feature layout, and Density-Aware Adaptive Query Allocation adapts the decoder query budget using encoder proposal scores. Experiments show substantial gains on small and medium objects while retaining strong large-object performance, with favorable accuracy–efficiency trade-offs and improved cross-domain generalization. The code will be released.

[CV-98] ERRA: Learning Transportable Latent Actions through Temporal Effect Representation and Relational Alignment

链接: https://arxiv.org/abs/2610.09509
作者: Tianxingjian Ding,Mubarak Shah,Yu Tian
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Robotics (cs.RO)
备注: Preprint. Code and project page coming soon

点击查看摘要

Abstract:Latent actions supervise robot policies with action-like codes inferred from visual transitions, and their usefulness hinges on two questions: what a code keeps from a transition, and whether it still means the same thing when reused in a different initial state. The first is a tension in time: an endpoint difference discards how motion unfolds, while the full sequence admits nuisance variation. The second is left open by reconstruction, which only ever observes a latent together with the state it came from. We argue that both questions can be answered in the same place. TERRA (Temporal Effect Representation and Relational Alignment) describes a transition by a compact temporal effect, its net feature change together with a low-order within-window dynamics component, and learns a continuous latent from this effect. The same effect space then serves as the reference for reuse: Effect-Anchored Transport (EAT) decodes a latent in other initial states and anchors the resulting effect to the one observed at its source, so that the latent is shaped by what it does across contexts rather than only by the transition it came from. With frozen linear readers, TERRA predicts actions more accurately than UniVLA and a LAPA-style baseline, degrades more slowly under visual distractors, and keeps transported transitions faithful to the donor action as the recipient context moves farther away; a same-budget control shows that these gains come largely from EAT. At matched pretraining scale, the complete system reaches 93.4% average success on LIBERO, compared with 91.8% for UniVLA.

[CV-99] RT-DETR-World: Transferring Rich LLM Semantics to Real-Time Open-Vocabulary Detection

链接: https://arxiv.org/abs/2610.09502
作者: Yupeng Zhang,Ziyi Zhao,Juntao Cheng,Sheng Wang,Ningnan Guo,Ruize Han,Liang Wan
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Open-vocabulary detection (OVD) recognizes categories unseen during training through textual category queries, yet achieving strong generalization with real-time efficiency remains challenging. Beyond vocabulary scaling, zero-shot generalization may benefit from reusable visual–semantic cues learned from seen data, including attributes, actions, states, and contextual relations. Existing real-time OVD methods primarily emphasize vocabulary coverage and efficient region/query–text matching; under strict efficiency constraints, compact detectors may struggle to absorb rich instance semantics and scene context. We propose RT-DETR-World, a compact DETR-style detector that transfers the rich semantics conveyed by descriptions during training while retaining lightweight query–text matching at inference. We construct GroundingCapv2 with three levels of supervision: category names for standard OVD, object descriptions conveying instance-level semantics, and image descriptions conveying object relations and scene context. These descriptions serve only as training-time semantic supervision. To help the compact detector absorb these semantics, we propose Dual-Path Description Alignment (DDA), combining a deployment-consistent MiniLM pathway with a training-only LLM teacher. MiniLM provides query–category supervision and object-description alignment, while offline teacher features supervise matched queries and global visual representations at the object and image levels, respectively. All teacher features are precomputed, and the teacher-side modules are removed after training. We further propose Relation-Aware Negative Relaxation (RNR), which uses teacher-derived semantic similarities to relax related negatives while preserving exact positives. Experiments demonstrate competitive zero-shot accuracy and a favorable accuracy–efficiency trade-off. The code will be released.

[CV-100] SpatialUQ: Post-Hoc Uncertainty Quantification from Spatial Consistency in Black-Box Vision Models

链接: https://arxiv.org/abs/2610.09498
作者: Md Kawsher Mahbub,Milon Biswas,Mirza Niaz Morshed,Wei Yu
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Clinical vision models are often deployed as frozen black boxes with no access to internals, retraining, or ground truth at inference time. We introduce \textbfSpatialUQ, a post-hoc uncertainty method using only output probabilities. It measures the Jensen-Shannon divergence between the global prediction and the mean of five fixed spatial crops in six deterministic forward passes. The premise is simple, trustworthy predictions are spatially consistent. On NIH ChestX-ray14 (DenseNet-121, N=25,596 ), our Multicrop Uncertainty Score (MUS) reaches 0.784 failure-detection AUC versus 0.664 for MC-Dropout ( p10^-6 ) at one-fifth the compute, with native calibration ( \textSCE=0.049 vs.\ 0.127 for \ell_1 ), the best-calibrated among methods above 0.78 AUC. A supervised fusion of MUS with entropy, confidence, and \ell_1 reaches 0.832 , outperforming a five-member ensemble ( 0.813 ). MUS scales with model quality, reaching 0.899 with BiomedCLIP ( \rho = 0.846 ), while this relationship remains meaningful in-distribution ( \rho = 0.523 ) but breaks down under severe distribution shift (VinBigData, \rho = 0.027 ). MUS is well-suited to diffuse findings but is less dependable for small focal lesions such as nodules. Code and experimental materials are publicly available at this https URL.

[CV-101] Gaussian Material Fields for Volumetric Multi-Energy CT Decomposition

链接: https://arxiv.org/abs/2610.09492
作者: Jian Lin,Jiancheng Fang,Hongming Shan,Shaoyu Wang,Yang Chen,Qiegen Liu
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 13 pages, 12 figures

点击查看摘要

Abstract:Volumetric material decomposition in multi-energy computed tomography requires a representation that organizes multiple three-dimensional material fields in a common spatial domain while retaining differences in composition and local structure. We observe that spatial primitives can be shared across materials without tying their coefficients, but their local capacity must respond to material-specific reconstruction needs. We introduce Gaussian material fields, which represent multiple material distributions with shared anisotropic 3D Gaussian primitives and independent nonnegative material coefficients. The shared geometry defines a continuous spatial basis, while the coefficients determine each primitive’s contribution to the individual material fields. To reconstruct this representation from multi-energy projections, a differentiable spectral forward model combines Gaussian material path integrals with a calibrated basis matrix, enabling joint optimization of spatial geometry and material composition. Material-aware adaptive density control retains material-specific refinement evidence before aggregation and adjusts local representation capacity to accommodate both spatially extensive components and sparse details. Experiments use synthesized multi-energy projections generated from pseudo-reference material maps constructed by conventional methods from publicly available CT data. Across 15 cases, our approach improves average PSNR by 4.03 dB and SSIM by 4.96% over the strongest baseline, while reducing NRMSE by 33.45%. Material-wise comparisons and component ablations support improved recovery of localized structures, while runtime and memory measurements show favorable computational scaling. These results establish Gaussian material fields as an explicit, adaptive representation for volumetric multi-material reconstruction.

[CV-102] InstanceBench: Diagnosing Referential Reasoning and Target Identity in Referring Expression Segmentation

链接: https://arxiv.org/abs/2610.09478
作者: Yuchen Li,Shaoyang Zhou,Yiran Wang,Ruiyi Deng,Haoyu Wang,Ziru Wei,Zhen Zhao,Luping Zhou
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Referring Expression Segmentation (RES) links natural-language descriptions to pixel-level object masks. Yet standard evaluation provides limited insight into instance-level referential reasoning: it does not systematically distinguish referential logics, test target preservation across valid grounding paths, or separate target-selection from mask-generation errors. We introduce InstanceBench, an instance-centered diagnostic benchmark comprising 6,194 images, 9,264 target instances, and 25,077 human-verified expressions. Each target-centric expression set (TCES) fixes the image and target mask while pairing a minimal expression with a same-target variant that uses another valid cue or grounding path. A compact referential-logic taxonomy spans direct target evidence, same-class selection, relational and compositional grounding, and exclusion, while logic-critical construction suppresses simpler shortcuts. Identity-aware metrics measure target retention and set-level success while separating selection from mask-generation errors. Across 22 native-mask RES checkpoints from 18 model families, the strongest checkpoint reaches 67.1% mIoU but only 59.6% All@0.7. Controlled interventions confirm language sensitivity, while failure decomposition identifies target selection rather than mask decoding as the main bottleneck. On a controlled training subset, matched supervision improves identity-aware performance, showing that the diagnosed capability responds to targeted supervision. Collectively, InstanceBench supports a measure-diagnose-improve cycle: measuring target consistency across grounding paths, localizing failure sources, and evaluating targeted interventions.

[CV-103] An Invariant Tangent-Angle Descriptor and a Band U-Net for 2D Frag ment Adjacency Prediction

链接: https://arxiv.org/abs/2610.09459
作者: Guillaume Brouillette,Alain Goupil,Pierre-Olivier Parisé,Fadel Touré(Université du Québec à Trois-Rivières, Trois-Rivières, Canada)
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注: 14 pages, 4 figures, 6 tables

点击查看摘要

Abstract:This paper addresses the prediction of adjacency between pairs of 2D fragments based on their contours. We improved the two-stage architecture proposed in Beaulac’s thesis, in which a rotation-equivariant Siamese convolutional neural network scores pairs of local image windows along the two contours of two fragments. The scores are gathered in an adjacency matrix in which a ResNet detects the partial anti-diagonal band that reveals the adjacency of two fragments. In the current work, we keep the pipeline and replace the local score by a comparison of tangent-angle profiles of contour windows, making it, by construction, invariant to fragment rotation and agnostic to the selected contour-starting point. These adaptations may be either a training-free likelihood ratio or a small one-dimensional convolutional model trained on corresponding points. We also replaced the final classifier by a band U-Net that segments the band and classifies the pair, so that the shared arc is obtained along with the decision. In the synthetic data set of the original thesis, the tangent descriptor performs as well as or better than the image-window approach in all tested configurations. The proposed pipeline reaches an accuracy of 98%, vs 93% to 95% for the original approach once its evaluation is corrected. We tested our pipeline, with models trained only on synthetic data, on the PairingNet benchmark, and obtained an AUC of 0.93. Furthermore, under the PairingNet pair-searching protocol conditions, our learned descriptor obtains a Recall@10 of 0.82 on the real set against 0.56 from the best model of the original paper.

[CV-104] RLHND: Video Foundation Models as Physically Grounded Hand Trackers for Robot Learning

链接: https://arxiv.org/abs/2610.09455
作者: Seungjun Moon,Subin Jeon,Sangwoo Kim,Hanbyul Joo,Jinwoo Shin
类目: Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)
备注: 34 pages, 15 figures

点击查看摘要

Abstract:Recently, approaches that leverage human video datasets for robot policy training have become increasingly prevalent. However, most existing hand trackers regress pose from cropped frames with limited priors on hand motion and object interaction, resulting in inaccurate and physically inconsistent estimates. Moreover, the lack of physical cues, e.g., contact and force, limits the use of human videos for robot policy training. To this end, we propose RLHND, a video foundation model-based hand tracking model that jointly estimates hand pose and realistic tactile information from monocular egocentric videos. RLHND turns the pre-trained Cosmos 3 video diffusion backbone into a deterministic clip-level feature extractor via clean-latent conditioning, carrying its learned priors on hand motion and hand-object interaction into tracking. For pose estimation, RLHND (i) predicts hand poses with anatomically plausible joint angles and (ii) enables optional conditioning on the shape parameter to maintain consistent hand shape within the same video and even across videos recorded by the same actor. For tactile estimation, a separate tactile expert stream, trained with the pose stream frozen, predicts dense contact and force over the hand surface. We further adopt LBS-based feature spreading to enable vertex-wise feature extraction without costly per-vertex attention. RLHND achieves state-of-the-art performance across various benchmark datasets for pose estimation, while also achieving state-of-the-art performance in contact and force estimation. Moreover, we demonstrate the utility of RLHND for robot learning through retargeting results and real-world robot experiments. The code will be publicly available at this https URL.

[CV-105] Iris-3B: Going Beyond the Latent with Pixel-Space Diffusion Training Conversion and Fine-Tuning

链接: https://arxiv.org/abs/2610.09450
作者: Hanqiu Li Cai(SperidLabs),Chema Garabito(SperidLabs)
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注: 19 pages, 13 figures, 6 tables. Project lead: Hanqiu Li Cai. Code and models: this https URL

点击查看摘要

Abstract:Pixel-space diffusion models avoid the lossy VAE of latent models, which suggests an advantage on downstream tasks where fine-grained detail matters. We test this claim along both routes to a pixel-space backbone. We pretrain Iris-3B, a 3B-parameter pixel-space text-to-image transformer, from scratch through a 256\to512\to1024 curriculum, after first ablating the prediction target and representation alignment at 256^2 to decide what to scale. We also convert a pretrained latent model, FLUX.2 Klein base 4B, to pixel space. We fine-tune both families for monocular depth estimation and for image restoration/super-resolution. We find no significant improvement from using a pixel-space generative prior. Fine-tuned for depth with one matched direct-regression recipe, Iris-3B is level with the latent FLUX.2 Klein and the converted pixel FLUX.2 Klein falls behind it, and on 4\times DIV2K restoration neither pixel model beats a latent FLUX.2 Klein fine-tune, the converted one trailing it slightly. We document the recipes, the failure modes and the remaining confounds behind this negative result. Nevertheless, Iris-3B shows that pixel-space pretraining with the pixel-transformer (PiT) head of PixelDiT scales to 3B parameters and to text-to-image quality competitive with latent models, matching Qwen-Image on OneIG under the official evaluators at 1024^2 . We release its weights and training code in the hope that they help pave the way for further work on pixel-space generation.

[CV-106] From Global Alignment to Local Grounding: Zero-Shot Chinese Character Recognition with Radical Verification

链接: https://arxiv.org/abs/2610.09449
作者: Yu-Heng Shih,Bing-Chen Wu,Tsz-To Wong,Ting-En Yen,Hong-Han Shuai,Bin-Hua Hsieh,Chien-An Chen,Yi-Ren Yeh,Ching-Chun Huang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Zero-shot Chinese character recognition (ZS-CCR) aims to recognize characters whose categories are never observed during training, and typically relies on the compositional structure shared between seen and unseen characters. Recent CLIP-style methods represent this structure with the Ideographic Description Sequence (IDS) and align it with glyph images in a shared embedding space. However, they rely on a single global image–IDS similarity that discards the spatial layout of radicals and, being learned only implicitly from seen classes, generalizes poorly to unseen ones; moreover, global matching often retrieves the correct character within the top candidates yet fails to rank it first when characters differ only in subtle local radicals. To address these issues, we propose a global-to-local two-stage framework. In the first stage, STG-CLIP augments the IDS with explicit tree-position and radical-level geometric priors, yielding a spatial-aware prototype that provides a consistent spatial description across seen and unseen categories for high-recall global retrieval. In the second stage, the Radical Verification Module (RVM) uses the radical instances of each retrieved candidate as queries to verify whether the corresponding radicals can be matched to spatially compatible regions in the input glyph. A margin-based gating rule activates the RVM only when the leading global candidates receive similar similarity scores. Experiments on the ICDAR2013 benchmark demonstrate that our method achieves state-of-the-art performance under the character-level zero-shot setting, obtaining 83.06% top-1 accuracy with 2,755 seen classes. Ablation studies further show that the explicit geometric priors and radical-level verification provide complementary improvements.

[CV-107] LighTROcc: Lightweight 4D Occupancy Forecasting via Instance-Centric 3D Gaussians

链接: https://arxiv.org/abs/2610.09444
作者: Hwanhee Jung,SeungHyeon Kim,Inkyu Koo,Qixing Huang,Sang Ho Yoon,Sangpil Kim
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Forecasting future 3D occupancy from surround-view cameras is essential for autonomous driving, yet existing approaches rely on dense voxel or bird’s-eye-view representations whose cost grows rapidly with spatial resolution and prediction horizon. Because these representations do not explicitly maintain object identities, they also struggle to preserve instance consistency over time. We present LighTROcc, a lightweight instance-centric framework that represents movable objects with a compact set of learned queries and predicts present and future occupancy in a single forward pass. LighTROcc localizes each query through attention-guided forward lifting, combining image-space cross-attention, query-specific depth, and camera geometry to estimate its 3D center. Each instance is modeled as a mixture of anisotropic 3D Gaussians and propagated across future steps using predicted displacements, producing continuous, temporally consistent occupancy forecasts. Experiments on nuScenes and supplemented nuScenes-Occupancy show that LighTROcc outperforms the evaluated dense and instance-wise baselines in instance-level forecasting accuracy while maintaining strong voxel-level occupancy quality. Across different model configurations, LighTROcc achieves a favorable balance between forecasting accuracy and computational efficiency, demonstrating the potential of compact instance-centric modeling for camera-based 4D occupancy forecasting.

[CV-108] IRA: Tumor Immune Representation Adaptation for Zero-Shot Cross-Cancer MSI and TMB Prediction

链接: https://arxiv.org/abs/2610.09441
作者: Dasari Naga Raju
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Microsatellite instability-high (MSI-H) and high tumor mutational burden (TMB-H) are clinically relevant biomarkers, yet their histopathological prediction remains challenging when models are transferred across morphologically distinct cancer types. Immune-associated spatial patterns can persist across cancers despite these morphological differences, but foundation-model-based predictors trained on a single cancer do not explicitly use this information, limiting cross-cancer generalization. To address this limitation, we propose TIRA (Tumor Immune Representation Adaptation), a target-free framework that refines frozen foundation-model representations using spatial immune topology, without requiring target-domain data during model development or test-time adaptation. TIRA uses a topology-supervised biology representation to condition tile-level attention while pooling only morphological features for joint MSI and TMB prediction. We train TIRA on TCGA-COAD+READ and evaluate it zero-shot on CPTAC-COAD, TCGA-STAD, TCGA-UCEC, and CPTAC-UCEC, covering cross-site, cross-cancer, and combined cross-cancer-site distribution shifts under UNI2, CONCH, and Virchow2. With UNI2, TIRA improved zero-shot AUROC on TCGA-STAD from 0.633 to 0.766 for MSI and from 0.651 to 0.772 for TMB. Source-derived spatial immune topology improved the cross-cancer robustness of frozen pathology foundation-model representations.

[CV-109] Mixture of Layers: Dynamic Layer Routing for Visual Reasoning NEURIPS2026

链接: https://arxiv.org/abs/2610.09440
作者: Jeonghwan Kim,Sofia Stoica,Jiwan Chung,Ansel Blume,Hyeonjeong Ha,Zhenhailong Wang,Xin Luna Dong,Heng Ji
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注: NeurIPS 2026

点击查看摘要

Abstract:Pre-trained vision encoders contain layer-wise visual representations that differ in spatial granularity, semantic abstraction, and sensitivity to local details. However, most Multimodal Large Language Models (MLLMs) rely on only the final or penultimate vision encoder representations or fixed aggregation rules, making visual abstraction largely query-agnostic and limiting access to fine-grained cues such as small objects, spatial details, text, and subtle visual attributes. In this work, we propose Mixture of Layers (MoL), an instruction-conditioned layer routing approach at the visual patch level that dynamically aggregates query-relevant latent representations from intermediate vision encoder layers. Given a text query, MoL predicts routing probabilities over vision encoder layers and performs a top-k sparse aggregation over selected hidden states at either the image level, patch level, or through a hybrid routing mechanism. In doing so, MoL enables query-adaptive access to layer-specific visual features for fine-grained visual reasoning. Our experiments across 7 fine-grained visual reasoning tasks demonstrate substantial performance improvements, especially across fine-grained visual grounding and understanding tasks such as +18.9% improvement on V* in overall accuracy, +4.5% on HRBench4K, and +16.3% on CharXiv compared to the baseline MLLMs, without resorting to multi-resolution inputs, simple interleaving of multiple vision encoders, or increasing the number of patch tokens. We study vision encoders’ receptive field scales across different layers and their sampling behaviors to provide an in-depth analysis of why layer-wise sampling is helpful, demonstrating that conditional visual representations are a key step towards better visual perception and reasoning in MLLMs. Our project page is available at this https URL.

[CV-110] InscriptionOCR: A Dataset and Method for Understanding Inscriptions

链接: https://arxiv.org/abs/2610.09439
作者: Jaidev Sanjay Khalane,Akbar Ali,V. N. Prabhakar,Shanmuganathan Raman
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Ancient script image restoration is a fundamental problem in computer vision, as it directly affects the reliable analysis and interpretation of historical documents and inscriptions. Ashokan Brahmi is an ancient script extensively used during the reign of Emperor Ashoka in the 3rd century BC, primarily for inscriptions in Prakrit. These inscriptions, including major and minor rock and pillar edicts, constitute a valuable yet largely unexplored source of data for computational analysis. The degraded nature of inscription imagery and the lack of standardized digital resources pose significant challenges for automated processing. We present an end-to-end AI-based framework for understanding ancient inscriptions that encompasses image enhancement, optical character recognition (OCR), transliteration, and neural machine translation (NMT). The proposed pipeline processes low-quality images captured directly from stone inscriptions, performs image restoration and Brahmi script character recognition, maps the recognized characters to the Roman script, and finally translates the resulting Prakrit text into English. We also introduce two new datasets: (i) InscriptionOCR Dataset: the largest publicly usable digital OCR dataset for Brahmi script to date, consisting of over 200,000 character images across about 600 classes, and (ii) a bilingual Prakrit-English parallel corpus comprising over 2,000 sentence pairs for NMT. We believe that the proposed framework and datasets will facilitate future research in ancient script analysis, low-resource OCR, and digital epigraphy. Subjects: Computer Vision and Pattern Recognition (cs.CV) Cite as: arXiv:2610.09439 [cs.CV] (or arXiv:2610.09439v1 [cs.CV] for this version) https://doi.org/10.48550/arXiv.2610.09439 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[CV-111] Controllable Crowd Generation through World-Model Planning

链接: https://arxiv.org/abs/2610.09438
作者: JunGyu Lee,Jisu Shin,Seunghyun Shin,Hae-Gon Jeon
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 28 pages, 6 figures. Project page: this https URL

点击查看摘要

Abstract:Crowd simulation plays a central role in robot navigation, autonomous driving, and urban planning. For these applications, realistic simulation requires crowds to adapt their behavior to environmental changes and user objectives. However, existing methods that rely on predefined control settings have limited flexibility in accommodating new user-specified objectives. To address this limitation, we propose Ctrl-CWM, a multi-agent Controllable Crowd World Model that integrates crowd generation and run-time control. Our key idea is to adapt the world-model principle of planning using imagined futures to crowd simulation. To this end, Ctrl-CWM consists of an encoder that learns a representation of human motion dynamics, an actor that proposes pedestrian displacements, a critic that evaluates imagined crowd trajectories, and a planner that selects actions. We first learn human motion dynamics through trajectory prediction on real-world pedestrian videos and then freeze the encoder to preserve them. Using this representation, the actor generates imagined crowd trajectories through repeated state updates, and the planner combines the critic’s scores with user costs to select actions. Repeated planning advances the simulated crowd, while additional user costs introduce new control objectives without retraining. We extensively evaluate crowd generation under varied agent arrival conditions and run-time control across avoidance and attraction scenarios. Ctrl-CWM outperforms the state-of-the-art method on most crowd realism and collision metrics, and adapts crowd behaviors to user-specified objectives introduced during simulation. The project page is available at this https URL

[CV-112] SkillCycle: Co-Evolving Agent Policies and Skill Banks

链接: https://arxiv.org/abs/2610.09430
作者: Ling Li,Qiuyu Shen,Zheng Jiang,Qinwei Ma,Yuxuan Liu,Zhidong Deng
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Internalizing external skills changes a language agent’s capabilities and, with them, the value of its remaining guidance: rules can become redundant, misleading, or insufficient for newly encountered decisions. This creates a coupled problem of learning from skills and adapting the skills that supervise further learning. We introduce SkillCycle, a framework for co-evolving agent policies and skill banks through a feedback loop between skill internalization and rule revision. Our central contribution is to give distillation feedback a second role: token-level contextual differences help locate rules for inspection, while interaction outcomes guide edits to their content and applicability. SkillCycle alternates between two phases: policy learning with a fixed skill bank and router, and rule revision with a frozen policy. Candidate edits undergo rule-level and whole-bank environment comparisons before they guide the next learning cycle. On WebShop, SkillCycle with a 3B model achieves a success rate of 74.74% and a score of 88.37 without inference-time skill inputs, representing relative improvements of 0.73% and 3.96% over the state-of-the-art (SOTA) model, respectively. In Cycle 3 ablations on ALFWorld and WebShop, SkillCycle’s no-skill success rates improve by 10.18% and 18.11% relative to a static skill bank, and by 2.41% and 2.50% relative to a single bank update, respectively. These results show that continually revising skill guidance as the agent’s capabilities change helps transform external skills into policy capabilities that require no skill inputs at inference. We will release code, configurations, skill banks, and evaluation protocols.

[CV-113] Event-Aligned Visual Action Reasoning for World Action Models

链接: https://arxiv.org/abs/2610.09427
作者: Xiaomeng Yang,Yushu Wu,Yi Gao,Yuhao Lei,Xuan Zhang,Pu Zhao,Yanzhi Wang
类目: Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)
备注: Project Page: this https URL

点击查看摘要

Abstract:World-Action Models (WAMs) utilize future visual prediction as an intermediate reasoning process to guide action generation. However, existing WAMs typically structure visual imagination according to predefined temporal intervals, without explicitly accounting for the different roles of task-critical interactions and connecting transitions. We argue that effective visual foresight should align directly with task-relevant interactions and their corresponding reasoning demands. To this end, we introduce an event-aligned visual action reasoning framework that organizes visual-action prediction around interaction events. Through event-aligned visual-action supervision, WAM learns to generate event-aligned visual context in each imagined rollout, placing greater emphasis on critical state changes that inform action generation. This shapes the visual reasoning granularity according to the underlying interaction dynamics, with detailed reasoning around task-critical events and coarser progression through connecting transitions. Furthermore, we introduce an execution validity head that identifies the valid portion of each predicted action sequence, avoiding redundant actions during chunked inference. Experiments demonstrate a 10.26 percentage point improvement in DOMINO success rate over baseline and competitive performance on RoboTwin 2.0. It also transfers from DOMINO Level 1 to Levels 2 and 3 without target-level adaptation.

[CV-114] Spatial Latent Reasoning for Embodied Reference Understanding

链接: https://arxiv.org/abs/2610.09418
作者: Ling Li,Jianhui Zhong,Wei Liu,Zheng Jiang aand Yuxuan Liu,Jingyu Li,Zhidong Deng
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Pointing-gesture visual grounding requires connecting hand geometry with the visual identity and extent of a referred object. A central challenge for continuous latent reasoning is how to organize these complementary cues into useful intermediate supervision. We propose Spatial Latent Reasoning (SLR), a framework that structures this supervision around an ordered sequence of geometric and visual states. A spatial ray state is supervised by fingertip position and pointing direction, followed by four states aligned with target-region features. To construct the visual targets, we introduce parity pooling, which applies polyphase grouping to average region tokens on four interleaved spatial supports. All states are generated recurrently during training and inference; auxiliary annotations are required only during training. On EgoPoint-Ground, the framework improves mIoU over same-backbone supervised fine-tuning by 2.8, 17.5, and 21.1 percentage points on Qwen3.5-4B, Qwen2.5-VL-7B, and Qwen3-VL-8B, respectively, with improvements on both hard subsets. On YouRefIt, it achieves 77.6% precision at IoU 0.5, a numerical margin of 5.2 percentage points over the reported state of the art under differing evaluation protocols. Ablations support joint geometric and visual supervision on the standard and similar-object sets, and favor parity over three alternative pooling operators on the standard set. These results support task-structured supervision for continuous pointing grounding. We will release the code and supporting materials.

[CV-115] rACT: temporal revelation Airborne Camera Trap

链接: https://arxiv.org/abs/2610.09417
作者: Oliver Bimber,Rakesh John Amala Arokia Nathan,Mohamed Youssef,Vinayak Lal Bhatnagar,Ralf Berger,Klaus Hackländer
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Effective remote monitoring and surveillance using drones are frequently impeded by severe environmental and thermal clutter, dynamic vegetation, target camouflage, and system latency. Drawing inspiration from the hunting strategies of birds of prey that hover and stabilize their vision to isolate subtle ground motion, we introduce trACT (temporal revelation Airborne Camera Trap), a lightweight, real-time aerial robotics framework designed for autonomous consumer drones. The system integrates Temporal Max Pooling (TMP), a low-level signal processing method that transforms imperceptible movement across a rolling integration window into robust value and time encodings, with self-supervised motion anomaly detection to isolate target motion from background environmental motion caused by wind gusts and drone drift. To overcome mechanical and processing delays, trACT combines motion prediction with automated gimbal-stabilized optical zoom verification and equitable multi-target verification balancing. Extensive real-world field experiments in densely forested wildlife habitats and surveillance scenarios demonstrate that trACT successfully bridges the gap between wide-area aerial monitoring and precise, autonomous target verification under challenging operational conditions.

[CV-116] Shared Geometry As A Rosetta Stone: Cross-Modal Alignment Without Paired Data

链接: https://arxiv.org/abs/2610.09411
作者: Dominik Schnaus,Thomas Dagès,Daniel Cremers,Xi Wang,Phillip Isola
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
备注: Project: this https URL , Code: this https URL

点击查看摘要

Abstract:Multimodal representations enable zero-shot classification and retrieval, but aligning independently trained models usually requires large amounts of paired data. Yet, the Platonic Representation Hypothesis suggests that models trained on different modalities may converge spontaneously toward a shared representation geometry. But then, do we even need paired examples for cross-modal alignment? Remarkably, we show that paired examples are unnecessary for coarse cross-modal alignment. Our simple Wasserstein Procrustes method with a coarse geometric initialization aligns two disjoint embedding sets by estimating a single orthogonal map without seeing any pairs. Across datasets, modalities, and unimodal models, we show that we can consistently align independently trained representations without pairs, and standard geometric alignment metrics accurately predict when this is possible. Nevertheless, we can naturally benefit from paired examples. In the very few-pair regime, our method substantially outperforms existing ones, while staying competitive with pair-based methods with more added examples. Finally, we demonstrate that the resulting alignments can enable text-to-image generation without paired examples. These results show that independently trained models often share enough geometry to establish cross-modal correspondence with little or no paired data.

[CV-117] ok: Audio-Visual LLM for Multi-Segment Temporal Grounding ACCV’2026

链接: https://arxiv.org/abs/2610.09408
作者: Eunji Shin,Dahyun Choi,Seungyeon Jo,Yejin Hong,Jiyoung Lee
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: ACCV’2026

点击查看摘要

Abstract:Audio-visual multi-segment grounding (AV-MSG) in untrimmed videos, reasoning over audio-visual evidence and predicting multiple segments for a query, is a fundamental problem but remains challenging. Visual-only models overlook complementary acoustic cues, while audio-visual models often fail to calibrate the number of events - a phenomenon we refer to as count miscalibration. We present TiTok, an audio-visual large language model (AV-LLM) that localizes an arbitrary number of temporal event segments for each query. For precise boundary prediction, we introduce the Time Token Interleaving (TTI) method, which explicitly injects special time tokens into the audio-visual stream to align input-side temporal perception with output-side temporal prediction. We further propose decoupled, multi-segment-oriented rewards for reinforcement learning, consisting of global, local, count, precision, and format rewards, optimized with Group reward-Decoupled Normalization Policy Optimization (GDPO). To assess the performance on AV-MSG, we establish a new UnAV-100-based evaluation protocol, and propose the CountF1 metric for quantifying count miscalibration that overlap metrics fail to capture. TiTok reaches 65.7 mIoU and 0.58 CountF1, achieving state-of-the-art performance. Our code is available at this link.

[CV-118] One Frame Full Heartbeat: ECG-Free Cardiac Cine MRI Synthesis via Phase-Conditioned Flow Matching

链接: https://arxiv.org/abs/2610.09397
作者: Shiyi Wang,Ruochen Sun,Peirong Liu,Xiang Li,Fangxu Xing
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Cine cardiovascular magnetic resonance (CMR) analysis relies on multi-frame sequences capturing the full cardiac cycle. However, standard multi-frame acquisition depends heavily on electrocardiogram (ECG) gating and repeated breath-holds, posing challenges in uncooperative populations, resource-limited settings, and temporally corrupted datasets. Existing methods that synthesize full cardiac sequences either rely on explicit ECG signals to parameterize myocardium function, or employ deformable registration without physiological constraints, failing to faithfully reproduce clinically relevant dynamic metrics such as ejection fraction (EF) and ventricular contraction magnitude. We present PhaseFlow, a unified generative framework that overcomes both limitations. PhaseFlow estimates a non-linear cardiac phase signal directly from the input sequence via a segmentation-derived left-ventricular (LV) area curve, capturing the asymmetric dynamics of systole and diastole without any ECG dependency. At inference, this phase signal is provided by a pathology-specific template, informing phase-specific frame generation. A rectified flow model conditioned on the phase and slice position synthesizes the full cardiac motion trajectory in the latent space, decoded into a diffeomorphic displacement field that warps end-diastole pixel intensities directly, eliminating the reconstruction blur often accompanying the variational autoencoder. On the ACDC benchmark, PhaseFlow achieves superior physiological fidelity and image realism, with best LV volume curve R^2 , structural similarity (SSIM) and generative quality (FID) among all baselines. Ablation studies confirm that each proposed component contributes measurably to the overall performance.

[CV-119] Quantifying Volumetric Risk: Class-Aware Asymmetric Weighted Conformal Prediction for 3D Medical Image Segmentation

链接: https://arxiv.org/abs/2610.09392
作者: Shadi Alijani,Fereshteh Aghaee Meibodi,Homayoun Najjaran
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Reliable volumetric segmentation is critical for clinical diagnostics, yet foundation models such as MedSAM remain deterministic and lack calibrated uncertainty under distribution shift. Existing conformal prediction methods offer statistical guarantees but are frequently applied in 2D and assume symmetric error distributions, so they do not capture the class-specific biases that arise in 3D multi-class segmentation. We propose Class-Aware Asymmetric Weighted Conformal Prediction (CA-WCP), which combines latent-space density-ratio weighting for covariate shift with directional quantiles for the lower and upper volume bounds, and scales each bound by a class-specific asymmetry factor derived from validation-set false-positive and false-negative rates. We prove that CA-WCP retains the weighted-exchangeability marginal coverage guarantee for every class, and we evaluate it on 3D brain tumor segmentation (BraTS 2020) and on a synthetic multi-organ CT benchmark constructed under covariate shift. On both benchmarks the 95% Clopper–Pearson interval for the observed coverage of CA-WCP contains the nominal 90% level for every semantic class, while interval width is reduced by 8–14% relative to symmetric weighted conformal prediction. We further encode the calibrated intervals into structured prompts for a multimodal large language model to produce uncertainty-conditioned radiology reports, linking distribution-shift-aware uncertainty quantification to interpretable clinical communication.

[CV-120] PPCAR-Net: Projection-Refined Parametric 3D Coronary Artery Reconstruction from Sparse X-ray Angiographic Views

链接: https://arxiv.org/abs/2610.09383
作者: Yu Ren,Hwee Kuan Lee,Tat-Jen Cham,Jonathan Yap,Khung Keong Yeo
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 22 pages, 5 figures, 10 tables, including references and appendix. Code and pretrained models: this https URL . Project page with video results: this https URL

点击查看摘要

Abstract:Sparse-view 3D coronary reconstruction commonly relies on cross-view correspondence and triangulation, which are vulnerable to vessel overlap and foreshortening, or on volumetric prediction followed by vascular-graph extraction, which does not directly provide centrelines and radii. We introduce PPCAR-Net, a projection-refined parametric coronary artery reconstruction network that directly predicts a branch-structured centreline-and-radius representation without explicit point matching, triangulation, or an intermediate volume. Given a variable number of segmented views, a coarse predictor combines frozen VGGT features with learned branch queries to estimate branch presence, B-spline centreline trajectories, and dense radius profiles. Projection-guided geometry and radius refiners then sample local evidence from the input views and apply residual corrections learned with 3D supervision. We evaluate representation fidelity and sparse-view reconstruction quantitatively and qualitatively. On simulated angiographic masks generated from CT-derived coronary anatomy, PPCAR-Net produces better connected artery reconstructions and achieves strong centreline accuracy, particularly for RCA, while maintaining competitive volumetric overlap. Coarse-to-fine inference takes 121 ms, enabling real-time reconstruction.

[CV-121] ScribbleEdit: A Benchmark for Scribble-Only Image Editing

链接: https://arxiv.org/abs/2610.09382
作者: Jie Ren,Hao Kang,Kai Guo,Yiding Yang,Bo Liu,Liming Jiang,Qing Yan,Zichuan Liu,Yizhi Song,Yue Xing,Hui Liu,Xin Lu
类目: Graphics (cs.GR); Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Scribble-based interaction provides a lightweight and intuitive way for users to specify image editing intents in interactive editing tools. However, current image editing models based on VLMs or LLMs struggle to understand and execute edits based solely on scribble inputs. To systematically study this problem, we construct a new benchmark, ScribbleEdit, that evaluates the ability of image editing models to perform image editing conditioned on scribbles. This task requires both a deep understanding of the intention of the scribble and an accurate interpretation of its spatial information. In ScribbleEdit, we design an automated data construction pipeline and introduce a dedicated evaluation protocol that explicitly measures intention alignment. Our analysis reveals that existing VLM/LLM-based editing models fail to accurately capture scribble intentions. To guide future progress on scribble-only image editing, we propose a simple yet effective soft-token baseline, which enhances the model’s understanding of scribble semantics and outperforms standard image editing models on our benchmark. Our evaluation and baseline together provide a concrete foundation for assessing and improving the scribble-driven image editing.

[CV-122] Unified Multi-plane Autoregressive Diffusion for 3D Multi-contrast MRI Synthesis ECCV2026

链接: https://arxiv.org/abs/2610.09376
作者: Yejee Shin,Geonhui Son,Jinglu Wang,Minwoo Jung,Yan Lu,Dosik Hwang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted to ECCV 2026

点击查看摘要

Abstract:Acquiring a complete set of magnetic resonance imaging (MRI) contrasts is time-intensive and uncomfortable for patients, despite the diagnostic value of multi-contrast imaging. This motivates synthesizing missing contrasts from those already acquired, which is an inherently 3D problem requiring anatomical coherence across axial, sagittal, and coronal planes. However, fully 3D generative models are often impracti- cal under computational resources that scale cubically with volume size. We propose a unified Multi-Plane Autoregressive Diffusion (MPAD), a latent diffusion framework that achieves full-volume 3D synthesis using efficient plane-wise 2D operations while preserving volumetric coherence. A 3D autoencoder first compresses MRI scans into an isotropic 3D la- tent representation. A 2D diffusion model is then trained to reconstruct masked latent slices of the target contrast, conditioned on both source- contrast slices and unmasked target-contrast slices. During inference, we introduce plane-wise autoregressive synthesis with inter-plane priors. Slices are generated autoregressively in random order within one plane orientation to maintain intra-plane continuity, then propagated as con- ditioning priors to orthogonal plane orientations to enforce inter-plane consistency. Compared to 3D latent diffusion baselines, MPAD reduces training and inference FLOPs by 7x and 3x, respectively, while also lowering inference time and peak memory consumption. Experiments on multiple datasets demonstrate that MPAD achieves superior perfor- mance, generating high-fidelity 3D volumes and supporting one-to-many translation within a single unified model.

[CV-123] Closing the Loop on Contrail Avoidance with Satellite Verification NEURIPS2026

链接: https://arxiv.org/abs/2610.09363
作者: Spandan Ghose Chowdhury
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注: Accepted into Tackling Climate Change with Machine Learning: workshop at NeurIPS 2026

点击查看摘要

Abstract:Contrails are the thin ice clouds that aircraft leave behind. They cause a large share of aviation’s warming, and rerouting the few flights that produce them could avoid much of it. However, an avoided contrail only counts if a satellite can confirm that it never formed, and this check is hard: contrails are one to two pixels wide, cover only 0.18% of pixels, and look very similar to natural cirrus. We build a small diffusion model (8.4M parameters, trained on one GPU) that detects them, and we run a controlled study to find out which components matter. The model reaches 0.476 PR-AUC, compared with 0.414 for a DeepLabV3+ baseline and 0.119 for an adapted MedSegDiff. Doubling the input resolution of the CNN brings it to parity (0.499, p=0.07). Three lessons apply beyond contrails. First, check the input resolution before designing a new architecture. Second, simple flips and rotations more than double accuracy and matter more than any architectural choice we measured. Third, pretraining the model on contrail shapes is harmful: the model learns that thin strokes appear everywhere and paints them onto empty scenes. Precision collapses to 1% while recall-based metrics still rate the degraded model as excellent, and no threshold or guidance heuristic repairs this failure.

[CV-124] Multimodal LLM s Can Learn to Read Brain Signals: A Vision–Language Model for Unified Multi-Task EEG Decoding

链接: https://arxiv.org/abs/2610.09355
作者: Parastoo Azizeddin,Omid Sharafi,Maryam M. Shanechi
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Learning EEG representations that generalize across cognitive tasks, subjects, and recording conditions remains a key challenge in electroencephalography (EEG) decoding. Recent advances in foundation models have improved EEG decoding performance, yet a fundamental open question remains: how to effectively interface neural signals with these models to enable multi-task learning across datasets. To investigate this question, we introduce BraVista, a visual-language framework that encodes multichannel EEG signals as structured images and enables multi-task learning through instruction-conditioned vision-language models (VLMs). Our approach relies on continued post-training of a general-domain VLM, leveraging its visual and linguistic priors to adapt to neural signals without a separate large-scale EEG-specific pretraining stage. We evaluate BraVista on four datasets spanning sleep staging, emotion recognition, cognitive workload classification, and abnormal EEG detection, showing strong performance across these tasks. Further analyses show that the choice of EEG-to-image representation is critical to performance. Moreover, through controlled perturbations of the EEG signal, we observe a gradual performance degradation under increasing noise, suggesting that the model relies on EEG-relevant information rather than superficial visual patterns. Together, these findings establish structured visual representations as an effective and scalable interface between neural signals and general-domain foundation models for unified multi-task EEG decoding.

[CV-125] Do Image Editors Follow Depth-Dependent Blur and Aperture Response? A Rendered-Ground-Truth Pilot Audit

链接: https://arxiv.org/abs/2610.09344
作者: Zhihan Chen,Yuhuan Zhao,Yijie Zhu,Xinyu Yao,Mengcong Ren,Yuchen Sun,Yunqing Chen
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:General image editors are asked to make a photo look as if it were taken at f/1.4, yet it is rarely checked whether the blur they add follows thin-lens optics. A physical aperture edit spreads blur across depth in thin-lens proportions and changes the blur when the aperture changes; prior evaluations check blur monotonicity, sharpness-trend correlation, effective-aperture error, or vision-language judgments, and none we found reports the two properties separately at known depths. In this pilot audit of two editors (Gemini~3.1 Flash Image and GPT-image-2.5) we compare against a rendered oracle: Blender Cycles scenes with true thin-lens depth of field, one blur-width estimator applied identically to oracle and editor outputs, and preregistered depth and aperture indices. In 24 texture scenes rendered in one three-panel geometry, accepted and measurable panels show f/1.4-to-f/2.8 width ratios, \bar\sigma_1.4/\bar\sigma_2.8 , of 0.99–1.18 against 1.98–2.13 for the oracle, and ratios of pooled median Gaussian-equivalent near/far blur widths of about 1.18–1.25 (Gemini) and 0.97–1.05 (GPT-image) against 1.69–1.82. Preregistered black-box interventions show that qualitative wording changes blur strength by roughly 2–10 times, whereas a request for 2 versus 6 px changes it 1.1–1.3 times and no tested wording of the f-number meets the registered ``followed’’ criterion. The depth compression appears in the original and reversed centre-focus layouts; with the focus on the near panel, the available-panel depth index reaches the registered threshold, and the aperture response stays attenuated in every layout tested. Scalar metrics adapted from published ones give oracle-like scores to synthetic editors whose proportions are compressed.

[CV-126] Skipper: Region-Adaptive Tile Pruning for 3D Gaussian Splatting

链接: https://arxiv.org/abs/2610.09343
作者: Jingxing Li,Yongjae Lee,Deliang Fan,Abhay Kumar Yadav,Cheng Peng,Rama Chellappa
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Tiled 3D Gaussian Splatting rasterizers often use one scene-wide contribution cutoff for tile enumeration, although content differs in its sensitivity to support truncation. TileSkipper selects a static per-Gaussian cutoff policy for a frozen checkpoint. Calibration renders measure candidate pair savings and an isolated-removal distortion proxy that accounts for front transmittance and background color. The method allocates cutoffs across 64 Gaussian groups and accepts policies only after complete renders on disjoint selection views. The exported policy uses one byte per Gaussian, with no parameter updates, additional kernel, or per-frame policy inference. Across 13 scenes from Mip-NeRF 360, Tanks Temples, and Deep Blending, a fixed-policy AccuTile sweep gives dataset-macro speedups of 1.088\times at standard resolution and 1.238\times at 3840 pixels wide, with -0.007/-0.023 dB mean PSNR change. Six integrations with existing opacity-aware bounds yield 1.009\times – 1.121\times compiler-only speedups. For four ports from 3\sigma rasterizers, we separately attribute the prior exact-bound transition and our incremental gain. Matched-quality ablations show modest gains over scene-global calibration and parity with per-Gaussian control; the standalone comparison with AdaGScale is regime-dependent.

[CV-127] CRT-HMAR: Causal Requirement Tracing-Guided Hierarchical Multi-Agent Regulation for Open-Task-Aware Infrared-Visible Image Fusion

链接: https://arxiv.org/abs/2610.09330
作者: Zengyi Yang,Shuai Yuan,Zhong-Cheng Wu,Juan Cheng,Huafeng Li,Yu Liu
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 17 pages, 11 figures

点击查看摘要

Abstract:Infrared and visible (IR-VIS) image fusion integrates complementary multimodal information into a single fused image to support downstream vision tasks. However, existing methods are typically tailored to seen tasks within a fixed task set and struggle to generalize to unseen tasks, which restricts their applicability in real-world open-task scenarios. To address this issue, this paper proposes CRT-HMAR, a Causal Requirement Tracing-Guided Hierarchical Multi-Agent Regulation Framework for open-task-aware IR-VIS image fusion. CRT-HMAR introduces a Causal Requirement Tracing Task Localization mechanism, which actively intervenes in key image information and observes task-network response variations to map task-specific semantic preferences into image-level causal requirement maps. Based on these maps, a requirement analysis agent aggregates task-specific requirement knowledge to adaptively guide requirement-customized image fusion. Moreover, CRT-HMAR incorporates History-Analysis Multi-Objective Balancing and Task-Level-Correction Conflict Mitigation mechanisms, jointly constructing a hierarchical regulation chain of “requirement interpretation - task balancing - conflict mitigation”. Through multiple collaborative agents, CRT-HMAR dynamically regulates key processes including open-task requirement modeling, multi-task balanced optimization, and gradient conflict mitigation. Extensive experiments on open-task scenarios involving five downstream tasks demonstrate that CRT-HMAR significantly improves generalization to unseen tasks while maintaining the performance and balance of seen tasks. Overall, CRT-HMAR shifts IR-VIS image fusion from task-oriented modeling toward requirement-oriented modeling, promoting its extension from closed-task settings to real-world open-task scenarios.

[CV-128] GRC-Net: Global Representation Consistency Network for Unsupervised Multimodal Anomaly Detection

链接: https://arxiv.org/abs/2610.09329
作者: Seyoung Jeong,Jong Pil Yun,Sang Jun Lee
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 5 pages, 3 figures, Under Review

点击查看摘要

Abstract:Automated quality inspection is essential for ensuring product reliability in this http URL image-based methods effectively capture appearance-related defects, these methods are limited in detecting structural and geometric anomalies, motivating multimodal approaches incorporating 3D information. However, existing methods mainly rely on local patch-level representations, which often lead to unstable reconstruction errors even in normal regions. To address this limitation, we propose GRC-Net, which integrates a global-attention MLP to enforce global representation consistency across patch embeddings with a stable reconstruction module to improve reconstruction stability. The proposed method captures holistic contextual information through a global token and suppresses reconstruction noise by minimizing discrepancies between original and predicted embeddings. Experiments on MVTec 3D-AD and Eyecandies demonstrate that GRC-Net consistently outperforms existing methods at both image and pixel levels. Qualitative results further demonstrate reduced reconstruction errors in normal regions and more distinct reconstruction differences between normal and anomalous regions.

[CV-129] Visual Jev Rewards: Reference-Bound Verification for Multi-Subject Image Generation

链接: https://arxiv.org/abs/2610.09328
作者: Baoteng Li,Wenzhuo Wu,Kongming Liang,Zhanyu Ma
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Multi-subject image generation requires rewards that verify whether requested attributes, actions, and relations hold for the specified reference subjects. Subject presence alone does not establish that the correct subjects participate in a requested interaction. We present reference-bound Visual Jev rewards that turn these visual decisions into generator training signals. Each subject-related question receives a positive label only when the requested condition and the relevant reference identities hold jointly. We construct fixed questions offline, train a Qwen3.5-4B verifier with binary supervision, and directly read Yes probabilities from its language-model head. Their mean supplies a GRPO reward while retaining individual judgments for inspection. Using 200 MICo-150K training tasks and 30 updates, the framework raises a GPT-5.4 composite score from 41.78 to 52.50 on a manually selected 897-task MICo-Bench subset; direct 27B rewards yield 51.84. Each reward is tested in one GRPO run, and offline human evaluation does not establish a statistically significant advantage over direct scoring. The study provides an initial implementation and evaluation of Visual Jev as a reference-bound reward for multi-subject image generation.

[CV-130] VIS-Ground: Video Interactive Storytelling with Contextual Grounding

链接: https://arxiv.org/abs/2610.09326
作者: Bingxuan Li,Yiwen Song,Xueqing Wu,Yanzhou Pan,Yang Li,Kuang Su,Jingyun Liu,Sebastian Ko,Huan Zhang,Tong Zhang,Nanyun Peng,Tomas Pfister,Yale Song
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Project Page: this https URL

点击查看摘要

Abstract:Video interactive storytelling enables viewers to actively steer how a video unfolds. However, once we allow viewers to intervene during generation, a new challenge arises: The viewer’s request can have latent dependencies on both the grounding source and the current rendered video state. These dependencies may not be explicitly stated in any individual input, but emerge only when the source, rendered history, and new viewer intent are considered jointly. Existing interactive video generation systems primarily emphasize following viewer instructions, while source-grounded video generation methods focus on aligning generated content with an external narrative or knowledge source. This leaves a fundamental question underexplored: What context should a generation model ground on during interactive continuation, and how can heterogeneous, unstructured inputs be transformed into such grounding context? In this work, we formulate contextual grounding as the process of transforming heterogeneous input context into an executable constraint model for video generation. To address this challenge, we introduce VIS-Ground, which performs Structured Context Abstraction to recover grounded states and cross-context dependencies, Generation Constraints Induction to project relevant dependencies into candidate-specific constraints, and Constrained Video Generation to enforce these constraints through planning, verification, revision, and rendering. Across three video generation backbones, VIS-Ground consistently achieves the highest overall composite score, reaching an average absolute improvement of 10.3 points over the strongest per-backbone baselines. Detailed analysis further shows gains across both narrative and knowledge grounding, and reveals remaining challenges in dependency extraction, and faithful realization during video rendering.

[CV-131] SAREO-FM: Decoupled Semantic Supervision for SAR-EO Foundation Models

链接: https://arxiv.org/abs/2610.09317
作者: Jeonghyeok Do,Munchurl Kim
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Please visit our project page at this https URL

点击查看摘要

Abstract:Synthetic aperture radar (SAR) and electro-optical (EO) imagery provide complementary observations: SAR enables day-and-night, weather-resilient sensing, whereas EO provides rich appearance and fine-grained semantic cues. We introduce SAREO-FM, which avoids forcing a single token stream to serve two distinct roles: modality tokens preserve how each sensor observes the scene through masked reconstruction, while learnable semantic queries capture what the scene contains under guidance from a pretrained vision foundation model (VFM). By jointly encoding these queries with SAR and EO tokens, the queries acquire modality-grounded semantic context, while the modality-token outputs remain the explicit targets of masked reconstruction. This design assigns semantic and reconstruction supervision to separate token streams while preserving their interaction within the shared encoder. Pretrained on the million-scale SAR-1M corpus, SAREO-FM achieves strong unimodal transfer for both SAR-only and EO-only inputs, while delivering substantial gains from joint SAR–EO observations on tasks that benefit from complementary sensing.

[CV-132] Why VLMs Miss Small Objects and When Zooming In Is Safe

链接: https://arxiv.org/abs/2610.09313
作者: Junzhe Shi,Yuan Gan,Shida Jiang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Vision-language models (VLMs) often miss small objects in large images. We ask three questions: what limits them, which of these limits better models can remove, and whether the classical way of handling large images, local decomposition, still has a future. We answer them with a theory built on two quantities of the image interface: S, the number of visual tokens across an object’s side, and L, the content a call must cover. The limits: a W x H image sent whole within N tokens gives an object of side m at most m*sqrt(N/(WH)) tokens per side. Doubling the token budget N therefore raises the largest S a whole image can reach by only 41%, and seeing an image at the S an object needs costs at least order S^2 tokens whatever the model. If recognition improves gradually with S, any search strategy, zoom agents included, obeys a recall-cost frontier. What better models can change: the S an object needs and how much one call can carry, measured in bits per object found; the coverage cost remains. Decomposition: yes. Assuming only that more tokens per object and less content per call do not hurt on average, splitting an image cannot lower recall if no view zooms out relative to the whole image and views overlap by one object. Neither part of this condition can be dropped; we bound the cost of every such decomposition, and a simple rule approaches the bound as the image grows. We test the theory in about 177,000 requests on 797 images. On controlled images, none of 40 orderings predicted for 8 VLMs is violated. Checked after the fact on drawings, floor plans, natural and synthetic images, 61 of 89 implied orderings hold significantly and 4 fail, all for OpenAI models given more pixels than their default path. On construction drawings the rule never significantly lowered recall relative to the whole image and raised it by up to 0.28. Code and data: this https URL

[CV-133] Kuration SDK: Addressing the Virtual2Real Gap via Data Curation

链接: https://arxiv.org/abs/2610.09305
作者: Nirmit Desai,Eric Song,Mayank Sengupta,Tejal Bedmutha,Siri Reddy,Sahiti Dharmavaram,Kunal Sawarkar
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)
备注:

点击查看摘要

Abstract:Benchmarks for measuring the quality of action-conditioned world models are still evolving and shifting away from visual similarity-based metrics to action-semantic and physically-grounded metrics. However, for domain and task-agnostic action-conditioned world model training, existing benchmarks provide a limited signal. By training and evaluating diffusion world models on CounterStrike gameplay data, we confirm that qualitative playability does not correspond with metrics such as FVD, LPIPS, and JEDi. We term this the Virtual2Real gap. We posit that, in lieu of reliable benchmarks, curating raw gameplay data and measuring a variety of diagnostic properties provides a more robust signal to bridge the gap, before the training even begins. We present several curation strategies and a general-purpose kit for physical AI data curation called Kuration SDK, which is being open-sourced with this paper. The SDK was instrumental in uncovering the root cause of the virtual2real gap in a specific case: why two world models trained on identical gameplay map, action and state distribution, behaved very differently when played in spite of having very similar LPIPS and FVD scores. Thus, Kuration SDK has the potential to uncover the root causes of Virtual2Real gap in specific datasets and accelerate development of sample-efficient training datasets.

[CV-134] mg2face: Expressive Facial Animation with High-Density Surface EMG

链接: https://arxiv.org/abs/2610.09304
作者: Ganidhu Abey,Wendy Greening,Ashika Kamboj,Leonhard Helminger,Abhijeet Ghosh,Karel Petranek,Sergio Orts Escolano,Dinesh K. Pai
类目: Graphics (cs.GR); Computer Vision and Pattern Recognition (cs.CV); Emerging Technologies (cs.ET); Machine Learning (cs.LG)
备注: 12 pages plus supplmentary material

点击查看摘要

Abstract:Facial movements convey subtle and important information that is critical for human social communication. Optical methods for face capture are difficult or impossible to use when the face is occluded by head-mounted devices (HMDs), such as VR headsets. Even with a clear line of sight, such methods raise privacy concerns and require head-mounted capture rigs that offset cameras and lighting from the face. We show that high-density surface electromyography (HD-sEMG) provides a viable non-optical alternative that addresses these challenges. We measured 64 EMG channels, using two textile EMG grids, with 32 from the forehead (typically occluded by an HMD) and 32 from the side of the face. EMG data were digitized at 2048 Hz and filtered. Facial movements were simultaneously recorded and used to estimate 478 3D facial landmarks using MediaPipe’s Face Landmarker. A major challenge in such multimodal recordings is synchronizing EMG and video data, which have different sampling frequencies and independent clocks. We developed a novel synchronization method using analog audio bursts that is capable of sub-millisecond synchronization. We also developed a staged fitting method that fits a recent high-resolution parametric head model (GNM), with 253 identity blendshapes and 383 expression blendshapes, to the MediaPipe landmarks as participants performed different facial expressions. We trained a deep neural network comprising per-grid spatial encoders followed by a dilated temporal convolutional network (TCN) to predict blendshape parameters from HD-sEMG signals at 100 Hz. Once trained, the network can predict expression blendshapes solely from HD-sEMG recordings. The output can be rendered using standard real-time blendshape animation methods. We demonstrate the methods using recordings from 25 participants, and direct expression transfer to a variety of human faces and non-human characters. Comments: 12 pages plus supplmentary material Subjects: Graphics (cs.GR); Computer Vision and Pattern Recognition (cs.CV); Emerging Technologies (cs.ET); Machine Learning (cs.LG) Cite as: arXiv:2610.09304 [cs.GR] (or arXiv:2610.09304v1 [cs.GR] for this version) https://doi.org/10.48550/arXiv.2610.09304 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[CV-135] LeCuration: A Tiny World Model as a Data Curation Multi-Tool

链接: https://arxiv.org/abs/2610.09285
作者: Mayank Sengupta,Nirmit Desai,Eric Song,Kunal Sawarkar
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)
备注: 9 pages

点击查看摘要

Abstract:Many applications of physical AI run within finite or closed physical worlds with a limited set of physical laws governing object behavior. Examples include robots working in a warehouse and agents moving around in a video game. In order to better organize, filter, and curate data for physical AI applications, we propose a new approach centered on the unique settings and physical laws of individual datasets. We train LeCuration, a small world model intended to serve as a data curation tool for a separate, larger downstream model. To build this model, we choose LeWorldModel (LeWM)as our latent encoder and predictor, adding a diffusion transformer (DiT) decoder to add visuals to autoregressive gameplay rollout. We find that the embeddings of this model can be used as an anomaly detection signal and as a content-based clustering heuristic, and that auto-regressively predicting the game state with this model allows us to qualitatively check for action-state consistency. This paper presents a qualitative, proof-of-concept case study on CS:GO gameplay data; we do not yet report quantitative curation metrics or downstream training results, which we identify as the key next step.

[CV-136] EM-SNN: Efficiently Modulated Spiking Neural Network for Remote Sensing Image Dehazing

链接: https://arxiv.org/abs/2610.09275
作者: Jie Shao,Jiaqi Ma,Wenwen Min,Beihang Song,Ning Chen,Youfa Liu,Jun Wan
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Although spiking neural networks (SNNs) provide an energy-efficient alternative to artificial neural networks (ANNs), their application to remote sensing image dehazing remains limited. A key challenge arises from the coupling between haze-induced high-frequency attenuation and discrete spike thresholding. This interaction suppresses weak responses and fundamentally limits the recovery of edges, textures, and fine details in spiking dehazing models. To address this challenge, we propose the Efficiently Modulated Spiking Neural Network (EM-SNN), a dedicated spiking framework tailored to remote sensing image dehazing. EM-SNN integrates a statistics-driven Threshold-Modulated Leaky Integrate-and-Fire (TM-LIF) neuron to adaptively compensate for haze-induced contrast compression, together with a Spike Sobel Modulation (SSM) module that enhances structural cues and reduces depth-wise attenuation during spiking feature propagation. By jointly modulating activation scales and structural representations, EM-SNN improves dehazing performance while preserving the inherent event-driven sparsity of SNNs. Experiments on HRSD, RICE, RRSHID, and SateHaze1K demonstrate that EM-SNN achieves competitive dehazing performance while consuming only one quarter of the energy of the strong ANN baseline SFRDP-Net.

[CV-137] Hardware-aware Calibrated Clustered Attention for Efficient Visual Geometric Transformers IJCNN26

链接: https://arxiv.org/abs/2610.09274
作者: Weitian Wang,Shubham Rai,Cecilia De La Parra,Akash Kumar
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注: Accepted to IJCNN26

点击查看摘要

Abstract:The Visual Geometry Grounded Transformer (VGGT) marks a significant leap forward in 3D scene reconstruction, as it is the first model that directly infers all key 3D attributes (camera poses, depths, and dense geometry) jointly in one pass. However, this joint inference mechanism requires global attention layers with extremely long sequences that causes a significant latency bottleneck. In this paper, we propose blockwise clustered attention (BC attention) to accelerate the global attention layers in VGGT. By limiting the clustering within HW-friendly neighborhood blocks, BC attention reduces the computation overhead of query clustering as well as the costly data movement between on- and off-chip memory. This enables BC attention to scale to long sequences and deliver practical latency improvements on GPUs. Moreover, we introduce a hashing hyperplane calibration method and a threshold-based error compensation method to reduce clustering errors efficiently, which is a bottleneck in the current clustered attention mechanism. Overall, our experiments on GPU demonstrate that calibrated BC attention accelerates the global attention layers by 2.10-2.63 \times and the whole backbone by 1.77-2.35 \times with negligible loss (1%) for large scenes. With a small performance loss ( 5%), calibrated BC attention further achieves a 2.26-2.87 \times latency improvement on the global attention layers and a 1.90-2.55 \times improvement on the backbone.

[CV-138] PhyDiCT: Plug-and-Play CT Reconstruction from Sparse X-Rays via Differentiable Rendering and Strong Priors MICCAI2026

链接: https://arxiv.org/abs/2610.09253
作者: Weicheng Dai,Shantanu Ghosh,Kayhan Batmanghelich
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted at MICCAI 2026; to appear in LNCS 16888

点击查看摘要

Abstract:Reconstructing 3D Computed Tomography (CT) images from a few X-ray projections is a highly ill-posed inverse problem due to the loss of volumetric information. We propose PhyDiCT, a training-free framework that integrates a differentiable Physics-based forward model, grounded in the Beer-Lambert law, with a text-conditioned Diffusion as a strong prior to reconstruct 3D lung CT images. We refer to our approach as training-free since the prior model is used without fine-tuning, and our goal is to steer the denoising procedure to generate samples consistent with X-ray observations. We guide the diffusion generation using Split Gibbs sampling to jointly optimize for projection fidelity (reward) and consistency with prior knowledge. Also, we introduce a test-time refinement step that enhances image realism and anatomical coherence. We extensively evaluate our method on publicly available 3D CT datasets using both perceptual and semantic metrics, demonstrating that it surpasses existing plug-and-play diffusion and fully trained reconstruction approaches. Our findings highlight that combining a strong generative prior with the underlying physics of image formation substantially improves reconstruction quality, e.g., 7.5% improvement on SSIM compared to full training methods. Code will be released at this https URL.

[CV-139] Adaptive Visual Token Reduction for Accelerated Image Understanding

链接: https://arxiv.org/abs/2610.09252
作者: Seyoung Jeong,Jong Pil Yun,Sang Jun Lee
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 5 pages, 2 figures. Under review

点击查看摘要

Abstract:Large Vision-Language Models achieve strong VQA performance, but processing high-resolution, information-rich images requires substantial computation, motivating visual token reduction. However, existing methods often prune individual tokens or rely on fixed-size cropping, limiting their ability to preserve spatially structured information such as horizontally or vertically elongated text. To address this limitation, we propose ReFIT, an instruction-guided visual token reduction framework for efficient LVLM inference. ReFIT consists of Relevance-Guided Window Reshaping (RWR) and Instruction-Guided Token Refinement (ITR), where RWR captures instruction-relevant regions by adapting to their spatial characteristics, while ITR further removes unnecessary visual tokens. Experiments on four VQA benchmarks demonstrate that ReFIT improves answer accuracy while reducing computational cost, and qualitative results demonstrate its effectiveness in localizing relevant regions and removing unnecessary visual information.

[CV-140] Pooling Representation Autoencoders for Efficient Diffusion

链接: https://arxiv.org/abs/2610.09242
作者: Ramón Calvo-González,Youssef Saied,François Fleuret
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Representation Autoencoders (RAEs) generate images from pre-trained visual fea- tures, but their dense token grids make generative modeling expensive. Motivated by local feature correlations, we introduce PoolDINO, a learned affine pooling operator that merges neighboring tokens. Training the pooling operator jointly with the RGB decoder preserves the standard two-stage RAE procedure without a separate feature auto-encoder. On ImageNet-256, 4x token compression retains comparable generation quality under internal guidance, while 16x compression trades some quality for greater efficiency. At a fixed budget of 100 sampling steps, latent-sampling throughput increases by 3.7x and 9.0x, respectively, relative to the unpooled baseline. Classification and dense prediction evaluations show that comparable guided generation quality can coexist with weaker performance on other tasks.

[CV-141] PVSync: A Unified Lip-Sync Expert for Timing and Articulation

链接: https://arxiv.org/abs/2610.09223
作者: Kevin Stephen,Varun Menon,Timo Mertens,Nikita Drobyshev
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Lip movements can match the timing of speech without matching the spoken sounds. We introduce PVSync, a unified model for audio-visual offset estimation and phoneme-level articulation scoring. PVSync combines window-level contrastive learning for synchronisation with a phoneme-level articulation objective that aligns audio and video embeddings of the same viseme class across clips. Visemes group phonemes with similar visible articulation. Viseme labels are derived automatically from forced-aligned transcripts, without manual annotations. On offset-corrected videos from 13 talking-head video generation models, PVSync matches human rankings of lip-sync quality more closely than LSE-C, achieving a Spearman correlation of 0.83 versus 0.34. On an automatically constructed benchmark from held-out speech, PVSync distinguishes viseme-matched from mismatched audio-visual pairs with an ROC AUC of 0.91. PVSync also outperforms SyncNet and MTD-VocaLiST in temporal offset recovery on held-out in-the-wild clips. Code and benchmark data will be released upon acceptance.

[CV-142] Consistent Distribution Matching for Data-Free Diffusion Distillation

链接: https://arxiv.org/abs/2610.09221
作者: Yuxiang Fu,Qi Yan,Zike Wu,Yongxing Zhang,Purang Abolmaesumi,Lele Wang,Renjie Liao
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Flow and diffusion models suffer from slow inference due to computationally expensive numerical integration. Distillation provides a promising way for a student model to learn from a teacher’s dynamics, enabling one-step or few-step generation. However, existing methods often depend on curated distillation datasets, costly teacher rollouts, or auxiliary proxy networks, which complicate model training and scaling. In this work, we propose Consistent Distribution Matching, a simulation-free and data-free distillation method for accelerating diffusion and flow models while preserving strong generative capacity. Our key insight is to unify sample generation and score estimation with one student network. Thus, our framework uses only two models, a frozen teacher and a trainable student, and optimizes one objective. We prove that minimizing our objective indicates Wasserstein convergence of the student flow-map pushforwards to the teacher marginals. On ImageNet 256 \times 256, our method attains an FID of 2.04 with a single function evaluation (1-NFE) and a 4-NFE FID of 1.37 within 40 epochs of training, surpassing the state-of-the-art distillation baselines without data. Our code code and model are available at this https URL.

[CV-143] A Deterministic Evidence Layer for Vision-Language Autism Screening from Naturalistic Home Video

链接: https://arxiv.org/abs/2610.09217
作者: Wenqi Li,Mindi Ruan,Chuanbo Hu,Shuo Wang,Xin Li
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Autism spectrum disorder (ASD) is diagnosed through specialist observation of a child’s social behavior, and access to that expertise is the bottleneck for early identification. Vision-language models (VLMs) describe a child’s behavior from video well; the verdict drawn from the description is unstable: at temperature~0, across eight pipeline configurations on one backbone, 16–37% of clips change their predicted label between repeated runs, and the cause lies in the serving stack. We keep the VLM frozen and move the decision out of the model. Under our grounded perception constraints the VLM writes an event table of timestamped, glossary-labeled events that names the eliciting press and logs counter-evidence; a text-only stage age-calibrates the confidence of each row; a deterministic weight-of-evidence scorer sums it into an evidence total and stratifies it into a risk category, so every decision decomposes into named per-feature contributions and can be re-scored from the saved table. On 43 caregiver-recorded, protocol-free home free-play clips of preschool children, the pipeline reaches AUC 0.851 \pm 0.012 , 86.0% accuracy, and F _1 71.8 over three runs. It labels 74.4% of clips correctly in every run (60.5% for the zero-shot baseline) and flags no typically developing clip in every run (9 of 31 at zero-shot). An ablation on the same backbone attributes the gain to the grounded perception constraints read through the deterministic scorer.

[CV-144] Do Vision Models Learn Physical Constraints or Rendering Shortcuts? A Counterfactual Benchmark for Grounded Physical Consistency NEURIPS2026

链接: https://arxiv.org/abs/2610.09205
作者: M. Moein Esfahani,Sepehr Salem,Mohammed Alser,Vince Calhoun
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注: Accepted in NeurIPS 2026 the 1st Workshop on Physical World AI

点击查看摘要

Abstract:Modern image editing models can satisfy a text instruction while breaking the physics of the edited scene. A new object may cast no shadow, a mirror may fail to reflect visible geometry, or an object may float above a surface that should support it. We study physical plausibility diagnosis, detecting whether an edited image violates scene physics, naming the violation type, localizing the affected region, and explaining the failure in language. We introduce a counterfactual benchmark whose controlled synthetic component uses Mitsuba~3 to generate 5,500 images from 500 scene families. Each family contains one clean image and ten matched violations involving shadows, reflection, support, surface response, and occlusion. The renderer pipeline provides category labels, affected-region masks and boxes, scene metadata, and explanation targets. We use LLaVA-1.5-7B, Qwen2.5-VL-7B, and InternVL3.5-8B as diagnostic baselines rather than proposed methods. On a 1,650-image synthetic test set, the adapted baselines reach 64.0–67.8% category macro-F1 on standard held-out scenes. For LLaVA-1.5-7B, category macro-F1 falls from 64.0% on the standard split to 40.8% under intervention shift. This gap shows that high in-distribution accuracy partly reflects cues tied to rendering and counterfactual construction.

[CV-145] StyleFields: Multi-Scale AdaIN-Modulated Implicit SDFs for Coarse-to-Fine 3D Shape Reconstruction and Editing

链接: https://arxiv.org/abs/2610.09200
作者: Ehsan Garaaghaji,Nicolas Talabot,Pascal Fua,Doruk Oner
类目: Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR)
备注: 39 pages, 20 figures, 3 tables. Includes supplementary material

点击查看摘要

Abstract:We introduce StyleFields, a DeepSDF-based architecture for high-fidelity 3D reconstruction that enables controllable geometric style mixing: the coarse structure of one object can be combined with the fine-scale details of another. The core idea is depth-aware modulation: instead of a single global code, we inject latents via multi-level Adaptive Instance Normalization at several decoder depths, and supervise matching auxiliary heads with a coarse-to-fine schedule while gradually growing network depth. This aligns early layers with global shape and later layers with high-frequency detail, achieving content-style decoupling without part labels or adversarial training. StyleFields delivers faithful reconstructions, convincing cross-instance hybrids, and consistent gains in ablations over injection depth and supervision granularity. We further demonstrate a practical application in automotive aerodynamics: a learned surrogate drag predictor serves as a differentiable objective to optimize reconstructed cars, allowing targeted edits of global form or surface details by freezing the complementary latent stream. StyleFields offers a simple, effective recipe for controllable implicit reconstruction and downstream performance-driven design.

[CV-146] StableGrasp: Reconstructing Physically Stable Human Hand Grasps from Single Images

链接: https://arxiv.org/abs/2610.09195
作者: Han Jiang,Etienne Vouga,Qixing Huang,Georgios Pavlakos
类目: Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR); Robotics (cs.RO)
备注:

点击查看摘要

Abstract:Reconstructing a physically stable human grasp from a single RGB image is challenging because physically modeling grasps is itself difficult, and the problem requires estimating not only a visually constrained hand pose but also a control target that stabilizes the grasp. Existing methods either model only visual hand geometry without considering physics, or rely on less plausible physical modeling, which limits the physical validity of the resulting grasps. In this paper, we present StableGrasp, a differentiable simulation-based optimization framework that explicitly separates the visual hand pose from the control target that determines the grasping forces. Our method jointly optimizes hand geometry and control by minimizing the kinetic energy of the grasp in a differentiable simulator, while regularizing the hand geometry to preserve visual consistency and geometric plausibility. The reconstructed grasps are substantially more stable under rigorous physical simulation, while remaining visually consistent with the input images and geometrically plausible. Experiments show that our approach produces far more stable grasps than alternative hand-control strategies, benefiting visual-only grasp reconstruction pipelines by turning their outputs into physically stable grasps.

[CV-147] R-CNN-Based Chess Position Recognition

链接: https://arxiv.org/abs/2610.09191
作者: Paras Govind,Ognjen Arandjelović
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Performing chess game position recognition solely from a single image of a three-dimensional board requires predicting the position and orientation of the board relative to the camera, the occupancy of squares and the piece type, which includes its colour. We propose an R-CNN-based framework with independent components for piece recognition and board geometry estimation, whose predictions are combined to reconstruct the position. For piece recognition, we adapt Faster R-CNN using a class-weighted objective and a deeper classification head. The detector operates directly on the input image, retaining alternative piece hypotheses that are subsequently refined using constraints on piece counts and square occupancy. For board detection, we introduce an octagonal arrangement of eight labelled boundary keypoints, predicted using the keypoint head of Mask R-CNN. These provide redundant correspondences for homography estimation and encode board orientation. The estimated homography maps representative points from the piece boxes to an 8x8 grid. On a synthetic dataset, the modifications to piece detection increase mean average precision from 61.59% to 90.14%. Of the predicted board keypoints, 97.11% are within 1% of the image diagonal of their labelled targets. Using ground-truth piece boxes with the predicted homographies gives correct square assignments for every test position. The complete framework recovers 76.61% of test positions exactly and 96.49% with at most one incorrect square.

[CV-148] One Frame Full Heartbeat: ECG-Free 4D Cardiac Cine MRI Synthesis via Radial-Decomposed Flow Matching

链接: https://arxiv.org/abs/2610.09185
作者: Shiyi Wang,Ruochen Sun,Xiang Li,Peirong Liu,Fangxu Xing
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Cine cardiovascular magnetic resonance (CMR) captures the cardiac cycle as a four-dimensional (4D) sequence, but standard acquisition requires electrocardiogram (ECG) gating and repeated breath holds. Visual realism alone does not establish accurate patient-specific ejection fraction (EF) or ventricular volumes. We present PhaseFlow3D, a generative framework that synthesizes a complete 4D cine sequence from a single end-diastolic (ED) three-dimensional (3D) volume without ECG. To capture asymmetric systolic and diastolic dynamics, it represents the cardiac cycle as a piecewise linear phase anchored at ED and end-systolic (ES) time points. At inference, a population-level canonical template supplies this phase without patient-specific temporal information. A phase-conditioned rectified flow model generates a cardiac motion trajectory in latent space. Radial Contraction Decomposition converts each latent state into a 3D displacement field, combining a physics-informed radial component for centripetal myocardial contraction with an image-conditioned residual for rotation and out-of-plane motion. Each frame is generated by directly warping the ED volume, bypassing variational autoencoder decoding. On the combined ACDC and MMs benchmark, PhaseFlow3D achieves the lowest EF mean absolute error, the only positive left-ventricular volume-curve R^2 , and the best distributional quality among compared methods. Ablations confirm each component’s contribution. Downstream evaluations demonstrate the utility of the synthesized sequences and displacement fields for segmentation, pathology classification, label propagation, and myocardial strain analysis.

[CV-149] RDGSplat: Render-Dedicated Geometry for Novel View Synthesis

链接: https://arxiv.org/abs/2610.09173
作者: Zhijie Zheng,Xinhao Xiang,Jiawei Zhang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:3D foundation models enable efficient novel view synthesis by carrying a Gaussian head on the representation they already use for reconstruction. However, the views they render fall short of the geometry they recover, because that geometry is estimated under a metric objective and never scored on how it renders. Recent methods alleviate this by updating the backbone weights, but they thereby discard the metric predictions the model was built for and must be repeated for every new backbone. To this end, we propose RDGSplat, a framework that decodes a second geometry dedicated to rendering from a frozen 3D foundation model, leaving its metric predictions intact. In particular, we devise Render-Dedicated Geometry Decoding, which duplicates the pretrained decoders and optimizes the duplicates under photometric supervision alone. Then, a Target-Pose Conditioned Adapter is introduced to reformulate the representation those decoders read, conditioned on the target camera pose rather than the target image. Extensive experiments show that RDGSplat improves novel view synthesis across three feed-forward backbones on four benchmarks, with every pretrained weight frozen. On RE10K, it raises WM2.0 from 20.918 to 24.266,dB while training 205.5,M added parameters against a frozen 1.4,B backbone, and the depth and pose the same model predicts are unchanged.

[CV-150] Depth-to-RGB: Repurposing a Frozen Depth Estimator for Geometry-Guided Compositing

链接: https://arxiv.org/abs/2610.09125
作者: Sanghyun Jo,Chae Yeon Lim,Donghwan Lee,Sihyun Kim,Soo Ye Kim,Kyungsu Kim
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Reference-based object compositing inserts or replaces an object using a background image, a reference image, and a 2D compositing mask. These inputs guide appearance and placement but leave the completed scene’s geometry implicit, which can distort object structure or alter the surroundings. Our Depth-to-RGB (D2R) framework predicts composite depth for a scene not yet observed in the RGB inputs. It learns reference-conditioned corrections to a frozen depth estimator using encoder features of paired completed scenes as targets. The unchanged decoder maps the corrected representation to the intended scene’s depth, which a separately trained renderer holds fixed during RGB synthesis. Under matched architecture and training, encoder-feature supervision reduces OOD Stage-1 AbsRel by 31.4% relative to decoded-depth supervision. We also introduce AnyInsertion++ with paired in-distribution and category-disjoint splits to evaluate generalization beyond compositing training categories. The complete D2R system leads 12 open-source and 3 closed-source baselines in estimator-derived geometry and photometric quality on both paired splits. On category-disjoint data, D2R reduces AbsRel by 43.7% and improves PSNR by 2.4 dB over the matched RGB baseline. Across three unpaired benchmarks, D2R leads both identity metrics and reduces mean CLIP reference cosine distance by 55% relative to the strongest baseline. Project page: this https URL

[CV-151] opoCurve: Geometry-Aware Topology Reasoning via Bézier Curves in Autonomous Driving NEURIPS2026

链接: https://arxiv.org/abs/2610.09118
作者: Mihai Bogdan Deaconu,Laura Dioşan
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted at NeurIPS 2026 (main track). 15 pages, 4 figures

点击查看摘要

Abstract:Topology reasoning jointly detects 3D lanes and traffic elements from multi-view images and infers their structural connectivity. Current methods model lanes as discrete polylines, lacking smoothness, analytical tangent directions, and global spatial support for attention, while providing sparse topology supervision. We propose TopoCurve, a geometry-driven architecture for 3D topology reasoning grounded in a structured parametric lane representation. Lanes are modeled as endpoint-fixed cubic Bézier curves, enabling continuous geometry with exact endpoints and analytically defined directionality. We exploit this shared curve geometry across the entire pipeline. Endpoint distance and tangent alignment are encoded with multi-scale Fourier features and injected into the topology head. Sampled curve points serve as geometry-aligned references for deformable cross-attention spanning the full lane. Parallel curve-anchored attention branches provide diverse predictions for one-to-many topology supervision. These components form a tightly coupled cascade where representation enables geometric reasoning, guides feature aggregation, and supports denser supervision. TopoCurve achieves 50.6 OLS on the OpenLane-V2 benchmark without any post-processing, establishing a new state-of-the-art among end-to-end camera-only methods, and outperforms all existing approaches on endpoint detection (56.8 vs. 52.6 on DET_p).

[CV-152] SPLATIFY: Reproduce Discover Innovate! From Papers and Ideas to Trainable 3DGS Code

链接: https://arxiv.org/abs/2610.09116
作者: Seemandhar Jain,Keshav Gupta,Manmohan Chandraker
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:The rapid growth of 3D Gaussian Splatting (3DGS) research demands significant effort to reimplement papers before building on them. We introduce SPLATIFY, a multi-agent framework that converts 3DGS papers into trainable gsplat-based implementations, where generic paper-to-code methods and frontier models fail. SPLATIFY achieves this through five innovations: (1) A context-free grammar for gsplat over a modular method template with extension points for losses, densification, rendering, and optimization, constraining synthesis so generated code satisfies gsplat’s architectural invariants by construction. (2) Architectural elements for faithful reproduction: fork-aware citation recovery retrieving component-level code at function-level granularity, Graph-of-Thought synthesis in topological dependency order, RAG-guided in-context example selection from over 20 verified implementations, and visual feedback combining PSNR-guided regeneration, Gaussian-level structural checks, and VLM-driven patching. (3) Knowledge-driven compositional improvement that autonomously finds weaknesses and composes complementary regularizers, losses, and densification strategies to improve upon original results. (4) Interdisciplinary method discovery where agents retrieve physical priors from outside the 3DGS literature and compose them with rendering knowledge to produce methods for previously unaddressed scene types. (5) SPLATIFY-Bench, an evaluation framework across 30 diverse 3DGS papers. On papers without public code, SPLATIFY matches expert implementations while reducing development time from weeks to minutes, and through compositional discovery further improves PSNR by up to 2.4 dB. We additionally demonstrate novel methods for volumetric nebula rendering and other scientific domains, synthesized entirely by SPLATIFY.

[CV-153] Anaximander: Interactively Running Geospatial Deep Learning Models on Any Compute Backend ECCV2026

链接: https://arxiv.org/abs/2610.09085
作者: Satej S. Soman,Akram Zaytar,Girmaw A. Tadesse,Gilles Q. Hacheme,Muhammad S. Danish,Inbal Becker-Reshef,Rahul Dodhia,Juan Lavista Ferres
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted as a poster at the TerraBytes II workshop, ECCV 2026

点击查看摘要

Abstract:Applying deep learning models to satellite imagery from within geographic information systems (GIS) remains high-friction for remote sensing practitioners. Models arrive in incompatible formats and target different compute environments, from local workstations to serverless cloud services. As a result, every evaluation demands custom deployment, tiling, and georeferencing code before a single prediction reaches the analyst’s map. This friction discourages systematic comparison in a domain where model choice directly affects operational outcomes such as field delineation, crop monitoring, and disaster response. We present Anaximander, an open-source system that unifies model source and compute location choice behind one interactive interface. The system’s backend is an inference server that loads models from multiple commonly-used sources and serves them on any accessible compute backend. The server provides session management and model caching, and streams results back per tile. The backend is paired with a QGIS plugin that drives tiling, result reassembly, georeferencing, and real-time per-tile status visualization. An additional user-interface path injects layer legends as prompts into vision-language models. We demonstrate the system in a code-free side-by-side comparison of three heterogeneous models on an agricultural field delineation task: gpt-image-1 via a cloud API, Segment Anything Model 3 (SAM3) on a remote GPU, and DelineateAnything on a local CPU. The inference backend and protocol are open-source and available at this https URL.

[CV-154] DISRQAD: Diffusion Image Super-Resolution Quality Assessment Dataset and Benchmark

链接: https://arxiv.org/abs/2610.09077
作者: Nikita Kukuzei,Artem Borisov,Evgeney Bogatyrev,Khaled Abud,Egor Chistov,Dmitriy Vatolin
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Diffusion-based image super-resolution (SR) can create visually plausible detail that is not supported by the low-resolution input. We introduce DISRQAD, a subjective-quality dataset and diagnostic benchmark for this setting. It contains mean opinion scores (MOS) for 14,000 SR outputs from ten diffusion and four non-diffusion methods, spanning four low-resolution degradation conditions and x2/x4 upscaling. We evaluate 51 standard full-reference and no-reference metric configurations and 11 adapted variants. Agreement with MOS is substantially weaker on diffusion outputs: the strongest standard no-reference baseline reaches 0.431 SRCC on diffusion SR versus 0.813 on non-diffusion SR. As a case study in benchmark use, a pruned and distilled Q-ReAlign-mini student reaches 0.496 SRCC on diffusion SR. DISRQAD measures perceived output quality, not faithfulness to the input; it enables analysis of metric behavior across generator families and input conditions. Our findings reveal a substantial gap in the assessment of diffusion-based SR and provide a basis for developing quality models sensitive to diffusion-specific artifacts.

[CV-155] OverLay: Dense-Overlap Layout-to-Image Generation Dataset NEURIPS2026

链接: https://arxiv.org/abs/2610.09071
作者: Shivansh Aggarwal,Shresth Grover,Divyansh Srivastava,Haiyang Xu,Bingnan Li,Xiang Zhang,Ethan J. Armand,Chuan Li,Jianwen Xie,Zhuowen Tu
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted at NeurIPS 2026, Evaluations Datasets Track. Project website: this https URL . Dataset: this https URL

点击查看摘要

Abstract:Layout-to-Image generation has made substantial progress in spatial and object-level control. However, existing methods still struggle with complex scenes containing many overlapping and interacting objects. We argue that training data is a particular bottleneck: existing datasets lack examples with dense, complex object interactions. To address this gap, we introduce OverLay++, a large-scale Layout-to-Image dataset with structurally complex scenes. OverLay++ contains approximately 500K images with an average of 6.6 objects per image, exceeding existing datasets by 1.67 times in annotation density. Beyond annotation density, OverLay++ provides rich semantic detail with object captions over six times longer than in current datasets. Our dataset generation pipeline is simple and produces dense, overlapping object annotations with rich per-object captions. Across multiple benchmarks, state-of-the-art Layout-to-Image methods trained on the OverLay++ dataset show consistent improvement and faster convergence, demonstrating the importance of dense, overlap-aware, and caption-rich supervision for controllable image generation.

[CV-156] GUARD: Geometric Uncertainty-Aware Point Cloud Denoising and Segmentation for Robotic Hard Disk Drive Disassembly

链接: https://arxiv.org/abs/2610.09068
作者: Zuoxu Wang,Xiao Liang
类目: Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)
备注:

点击查看摘要

Abstract:Reliable robotic disassembly requires part-level representations that distinguish genuine component geometry from scanning and reconstruction artifacts. In point clouds of hard disk drives (HDDs), structured ghost artifacts can resemble valid components locally while remaining inconsistent with the overall geometry, allowing erroneous measurements to receive plausible semantic labels. This creates an engineering information problem: semantic prediction confidence alone does not establish whether the underlying geometry is reliable. We propose \textbfGUARD, a geometric uncertainty-aware framework that performs point filtering and segmentation within a single forward pass by modeling the reliability of learned geometric representations. GUARD combines a multi-scale geometric transformer with a multi-bandwidth random Fourier feature Gaussian Process to estimate per-point geometric uncertainty, complemented by predictive entropy to suppress unreliable measurements while preserving informative structures. Evaluation on 2,745 real HDD point clouds shows that GUARD improves PointNet++ segmentation mean intersection over union from 0.7739 to 0.8318. Additional experiments on ShapeNetPart and ScanNet examine robustness across corruption types, point-cloud domains, and segmentation backbones. On manually annotated ScanNet samples, geometric uncertainty achieves a corrupted-point detection F1 score of 0.7931, compared with 0.2212 for predictive entropy. The results demonstrate the value of distinguishing geometric reliability from semantic confidence and reveal a tradeoff between artifact suppression and preservation of informative structures. GUARD contributes a reliability-aware approach to interpreting imperfect 3D measurements for component identification and subsequent robotic handling. Project website: this https URL.

[CV-157] Shape-Bayes: Bayesian Inference of Structured Shapes under Visual Ambiguity

链接: https://arxiv.org/abs/2610.09032
作者: Mani Kumar Tellamekala,Tosh Brown,Michel Valstar
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Perceiving structured shapes, such as human faces, from pixels is an inherently ambiguous task in real-world conditions. Yet, shape inference is largely posed as a deterministic regression task predicting fixed spatial coordinates. We find that deterministic regression is brittle when visual evidence is ambiguous or incomplete; under severe occlusions deterministic models exhibit structural collapse, predicting incoherent shapes or reverting to generic averages. To address this, we introduce Shape-Bayes, a probabilistic framework that couples uncertainty-aware visual perception with Bayesian shape reasoning. Rather than forcing point estimates, Shape-Bayes dynamically weights visual evidence against geometric priors to infer a structurally valid shape posterior. Demonstrated on human face shape regression, a rigorous testbed featuring complex non-rigid deformations and strict anatomical constraints, Shape-Bayes comprises: (1) a base model predicting noisy landmarks alongside distilled aleatoric uncertainties; (2) a lightweight Transformer encoding these observations into an adaptive prior over a PCA shape manifold; and (3) a differentiable Bayesian solver computing closed-form posteriors by balancing the noisy predictions against this prior. By guaranteeing complete structural integrity, Shape-Bayes achieves an absolute improvement of up to ~34% IDR over state-of-the-art deterministic models. Simultaneously, it yields highly calibrated uncertainty bounds and reduces relative error by up to 12.5%, establishing a new state-of-the-art for robust 2D face shape regression under severe occlusion. The project page is at this https URL.

[CV-158] Beyond Explanation: Debugging Medical Imaging Models via Concept Intervention MICCAI2026

链接: https://arxiv.org/abs/2610.09031
作者: Samrajya Thapa,Daniel J. Quest,Timothy L. Kline,Carrie L. Langstraat,Emanuel C. Trabuco,Wei Le
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注: Accepted at the 5th Workshop on Applications of Medical AI (AMAI), MICCAI 2026

点击查看摘要

Abstract:Medical imaging models often operate as black boxes, limiting interpretability and systematic debugging. We introduce an easy-to-use, plug-and-play framework for concept-based interpretation and model refinement. By aligning a single-modality encoder to BioMedCLIP, we construct a Concept Bottleneck Model (CBM) that enables concept-level interventions. These interventions allow us to isolate causal versus spuriously correlated concepts, validate insights with domain experts, and generate counterfactual samples for targeted fine-tuning. We evaluate our framework on a Mayo Clinic ultrasound dataset and the CheXpert 5x200 chest X-ray dataset. Results demonstrate that concept intervention enables reliable model diagnosis while maintaining, and occasionally improving predictive performance via guided fine-tuning. Our findings highlight the practical value of this framework for controlled, interpretable refinement of clinical deep learning models.

[CV-159] Visual Memory Attacks Can Persist Through The KV Cache

链接: https://arxiv.org/abs/2610.09027
作者: David Dobre,Leo Schwinn,Gauthier Gidel,Spandana Gella,Perouz Taslakian,Pierre-André Noël
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Modern language model systems operate autonomously over increasingly long contexts containing untrusted text and images. Can an adversarial input continue to steer a model even after that input is removed from its context? We show that attacks can be trained to persist through the key/value (KV) cache of subsequent tokens, allowing adversarial influence to outlive direct access to its this http URL consider the Visual Memory Injection (VMI; Schlarmann and Hein, 2026) attack setting, in which an adversarial image that stays in the context plants a hidden backdoor: the model behaves normally until a chosen trigger elicits an attacker-chosen response. We first demonstrate persistence in this setting with optimized soft prompts, which remain effective after we mask the prompt from attention. We then introduce Persistent Visual Memory Injection (P-VMI), which optimizes images to preserve this adversarial behaviour after they are masked from attention. These attacks persist over conversations substantially longer than those used during optimization. On Qwen3-VL-8B-Instruct, P-VMI achieves up to approximately 90% target success in its strongest configuration and remains effective under a stricter removal setting that exposes the image only on the first turn. A cache-swap ablation localizes the persistent influence to the KV cache. Finally, we show that these attacks can be trained to survive compaction that retains the KV cache of a summary generated by the same model, demonstrating that adversarial behaviour can persist in cached state without continued access to its source.

[CV-160] Personalize at Test Time: Learning User Preferences for Image Generation

链接: https://arxiv.org/abs/2610.09015
作者: Jiamu Bai,Jiaming Hu,Yanhong Wu,Zellux Wang
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Diffusion models can generate high-quality images, yet aligning their outputs with individual user preferences remains challenging. A key bottleneck is accurately modeling diverse user preferences from limited feedback. Existing approaches often rely on labor-intensive manual preference annotations or vision-language models (VLM) to extract preference information from user interaction histories, introducing substantial annotation or computational costs that limit scalability. We propose an approach that learns personalized reward models directly from users’ historical image preference pairs. First, we use an autoencoder to compress hundreds of visual attributes into 50 attribute-anchored preference dimensions and train an evaluator to score images along these dimensions. We then represent each user’s preferences as a linear combination of the shared dimension scores, estimating the user-specific weights by maximizing the likelihood of their observed pairwise preferences under the Bradley-Terry model. This formulation reduces per-user adaptation to optimizing a low-dimensional weight vector, simplifying optimization and enabling data-efficient personalization from sparse feedback. The learned personalized rewards guide image generation at inference time while keeping the diffusion model frozen. Experiments on real-user preference data show that our approach achieves approximately 77% held-out pairwise preference prediction accuracy and improves the alignment of generated images with individual user preferences.

[CV-161] Zero-Shot Brain MRI Inpainting with 2.5D Unconditional Flow Priors

链接: https://arxiv.org/abs/2610.08983
作者: Arnela Hadzic,Franz Thaler,Simon Johannes Joham,Martin Urschler
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Generative inpainting of brain MRI volumes is essential for synthesizing healthy tissue in pathological regions, improving the accuracy and reliability of automated downstream brain analysis applications such as image registration, brain extraction, and segmentation. However, standard 3D approaches are computationally prohibitive, while efficient 2D slice-wise methods suffer from severe inter-slice discontinuities. Furthermore, traditional models rely on conditional training, requiring task-specific learning of masked inputs. We propose a zero-shot brain MRI inpainting framework utilizing 2.5D unconditional flow priors to capture spatial context along the superior-inferior axis without the overhead of full 3D convolutions. During training, our flow matching model learns the joint distribution of adjacent axial slice triplets, modeling the manifold of healthy brain anatomy while explicitly excluding pathological regions from the loss function. At inference, the model processes the input triplets autoregressively along the depth axis. We employ the Restora-Flow solver to constrain the unconditional prior using the input mask, achieving accurate zero-shot inpainting. Evaluations show our 2.5D strategy resolves the structural discontinuities of 2D baselines, synthesizing plausible healthy tissue while maintaining volumetric consistency across the axial, sagittal, and coronal planes. As a final step, we generate and average an ensemble of multiple stochastic reconstructions to form the final prediction. Quantitative results benchmarked on the official BraTS 2026 Inpainting Challenge validation set demonstrate the effectiveness of our proposed approach, yielding an SSIM of 0.816 \pm 0.112, MSE of 0.007 \pm 0.005, and PSNR of 22.923 \pm 4.343. Code is available at this https URL.

[CV-162] S2Tok: Streaming 3D Gaussian Reconstruction with Persistent Spatial Tokens

链接: https://arxiv.org/abs/2610.08978
作者: Fang Li,Jiraphon Yenphraphai,Quentin Herau,Depu Meng,Yihan Hu,Tianshuo Xu,Narendra Ahuja,Wei Zhan
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Project Page: this https URL

点击查看摘要

Abstract:Streaming 3D reconstruction requires more than a sequence of geometric predictions: it requires a persistent scene state that can incorporate new evidence and remain renderable as observations arrive. Latent spatial tokens offer a promising representation for this purpose, but constructing them from an image collection leaves open how to maintain them online, where each observation may both revisit known regions and reveal new content. We introduce S2Tok, a feed-forward framework that maintains a size-adaptive, persistent scene state from uncalibrated image streams. Its central idea is to distinguish updates to the existing representation from selective expansion. A spatially informed transformer integrates each incoming observation with the persistent scene tokens, while a learned admission module selectively expands the representation to limit redundant storage. A hierarchical decoder and Gaussian head convert the evolving state into non-pixel-aligned 3D Gaussians, enabling novel-view rendering without caching previous frames. Experiments across four benchmarks demonstrate competitive streaming rendering quality with compact Gaussian representations. These results support latent spatial tokens as a persistent computational state for online 3D reconstruction, combining learned scene updates with explicit Gaussian rendering.

[CV-163] RACER: Reflective Agent Coupling Query Interpretation and Tool-Based Retrieval for Frame Selection in Long Video Understanding

链接: https://arxiv.org/abs/2610.08954
作者: Yiyang Huang,Yitian Zhang,Yizhou Wang,Jianglin Lu,Qihua Dong,Hailing Wang,Huimin Zeng,Mingyuan Zhang,Yun Fu
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Video large language models (Vid-LLMs) excel at diverse video-language tasks by reasoning over selected frames. However, frame selection for long videos remains challenging, as it requires retrieving relevant frames distributed across segments from a large candidate pool given complex queries. This paper investigates dominant approaches to long-video frame selection from a task-decomposition perspective, identifying two key challenges: the Query Comprehension Gap in similarity-based methods and the Interpretation–Selection Gap in judgment-based methods. To address them, we propose RACER, a training-free reflective agentic framework that decomposes long-video frame selection into query interpretation driven by a lightweight Vid-LLM and evidence localization supported by an embedding model serving as a retrieval tool. Specifically, the Vid-LLM is responsible solely for reformulating the complex query into sub-queries that make implicit information requirements explicit, mitigating the Query Comprehension Gap. Meanwhile, the retrieval tool leverages these sub-queries to localize relevant evidence, relieving the Vid-LLM of direct frame selection and thus addressing the Interpretation–Selection Gap. Finally, the retrieved frames are fed back to the Vid-LLM for sub-query refinement, forming a reflection loop that iteratively improves query interpretation and frame selection. Experiments across multiple benchmarks show that RACER consistently improves long video understanding. Notably, RACER achieves effective frame selection even with limited-capability components, demonstrating that agentic integration enables these components to enhance more capable Vid-LLMs.

[CV-164] One-Slide Calibration of Pathology Foundation Models NEURIPS2026

链接: https://arxiv.org/abs/2610.08944
作者: Ming Ren Hou,Tianyi Huang
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注: Accepted at the NeurIPS 2026 Workshop AI at Scale for Clinical Impact (ASCI): Cancer Pathology Foundation Models

点击查看摘要

Abstract:Scanner variation changes how pathology foundation models represent the same tissue. We introduce SlideRuler, which uses regions within a slide as internal controls to estimate and correct acquisition-induced shifts in other regions. A transfer map learned from paired rescans enables calibration from a single scan at inference while keeping the foundation model fixed. Across two encoders and five SCORPION scanners, learned transfer reduces mean target-to-source embedding distance by 16.3-38.5% relative to raw embeddings. Comparisons with unrelated same-scanner controls reveal a positive same-slide contribution across all four evaluation settings, including scanner holdout. A source-anchored variant reduces source-feature displacement by 47.7-83.6% relative to learned transfer while retaining most of its alignment gain. By drawing calibration information from the slide itself, SlideRuler offers a path toward more consistent use of frozen pathology models across imaging systems.

[CV-165] SPW-Nav: A Streaming Panoramic World Model for Language-Guided Navigation

链接: https://arxiv.org/abs/2610.08941
作者: Yunheng Liu,Ziqi Cai,Siqi Yang,Yimu Wang,Minggui Teng,Jiaming Tan,Shuchen Weng,Erwin Wu,Kaipeng Zhang,Boxin Shi
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Project page: this https URL

点击查看摘要

Abstract:Language-guided panoramic video generation benefits various downstream applications, such as interactive 3D scene exploration, virtual reality experiences, and embodied agent training. Existing panoramic generators follow predefined trajectories, and interactive world models act through low-level actions in perspective views. We propose SPW-Nav, a streaming panoramic world model that understands movement instructions and streams one minute of 2K 360-degree video in real time from a single panorama. SPW-Nav interprets each instruction in the previously generated panorama as camera motion. Spherical rotation decoupling applies rotation exactly on the sphere, pose-aligned conditioning keeps translation inputs bounded over long streams, and a multi-term memory with a few-step generator continues the scene as instructions change. We also build SPW-NavSet, panoramic videos with camera trajectories and verified instructions. Driven by language, SPW-Nav outperforms prior panoramic generators in camera-following accuracy and video quality, and supports on-the-fly instruction switching.

[CV-166] VCR-Bench: A Modular Open-Source Benchmark for Video Classification Robustness ACM-MM2026

链接: https://arxiv.org/abs/2610.08936
作者: Maksim Plinskiy,Aleksandr Gushchin,Sergey Lavrushkin,Dmitriy S. Vatolin,Anastasia Antsiferova
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 6 pages,1 figure, accepted at ACM MM 2026

点击查看摘要

Abstract:Robustness of image classification has several benchmarks, but their video counterparts are absent. In video classification temporal dimension introduces additional degrees of freedom for adversarial attacks, defenses, and preprocessing. Temporal sampling, perturbation budgets, and metric aggregation also interact in ways with no direct analogue in the image setting. Therefore, robustness for video classifiers is studied across scattered, incompatible implementations, making reported numbers hard to reproduce and analyze. We introduce VCR-Bench, a modular open-source benchmark framework that standardizes video loading, wrappers for classifiers, adversarial attacks and defenses, perceptual metrics, configuration presets, and result logging. VCR-Bench currently integrates 30 video classification models, 14 adversarial attacks, and 10 defense wrappers under a common evaluation protocol. We evaluate representative video classifiers, attacks, and defenses on Kinetics-400 subset, reporting clean accuracy, attack success rate, perceptual quality, runtime, and memory usage. VCR-Bench is released with documented installation, reproducible run presets, component-extension interfaces, and scripts for reproducing the reported results at this https URL.

[CV-167] CIRRA: Dual-Level Continual Instruction Reconciliation with Ongoing Execution for Embodied Robot Agents in Interactive Household Tasks

链接: https://arxiv.org/abs/2610.08862
作者: Ci Zhang,Enfu Nan,Arman Akbari,Lin Zhao,Li Wang,Chen Wang,Weiwei Chen,Yanzhi Wang,Geng Yuan
类目: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Household robots must accommodate new user instructions while executing ongoing tasks. Existing agents often regenerate or extensively revise the remaining task sequence, introducing plan ambiguity, logical inconsistency, and redundant execution. We formulate continual instruction reconciliation and propose CIRRA (Continual Instruction Reconciliation for Robot Agents), a dual-level framework combining LLM-based semantic reasoning with rule-constrained structural integration. CIRRA first grounds incoming instructions to unique executable skills and resolves underspecified actions and execution locations. It then preserves the ongoing subtask sequence as an execution backbone and generates integration candidates by inserting incoming subtasks into location-matched segments. The semantic reasoner evaluates only modified segments to identify dependencies and conflicts and select the most logically coherent local integration. This structure-preserving process maintains alignment with ongoing execution, mitigates ambiguity and inconsistency, and reuses shared subtasks to reduce redundant execution. We also introduce CHIRP (Continual Household Instruction Reconciliation and Planning), a text-based benchmark of 120 episodes across eight household environments and six categories of everyday activities. On CHIRP, CIRRA achieves 74.2% decision agreement, exceeding the strongest replanning baseline by 30 percentage points; every correct fusion decision yields a correctly placed, conflict-free schedule. On a Unitree G1 humanoid, CIRRA interrupts ongoing skills at the correct moment in every trial and significantly outperforms all baselines on every metric.

[CV-168] MoR-MLLM : Mixture of Recursions for Efficient Multimodal Large Language Models

链接: https://arxiv.org/abs/2610.08830
作者: Pengcheng Zheng,Chaoning Zhang,Jiaxin Yan,Sihan Cao,Jianwei Zhang,Xudong Wang,Jiaquan Zhang,Jewon Lee,Tae-Ho Kim,Yang Yang,Heng Tao Shen
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 13 pages

点击查看摘要

Abstract:Multimodal Large Language Models (MLLMs) have demonstrated remarkable reasoning capabilities across vision and language tasks. However, their massive computational and memory demands hinder real-world deployment. While recent efforts reduce costs by employing lightweight language backbones, existing paradigms remain computation-dense due to their static sparsity and depth allocation, which cannot adapt to the semantic complexity of each token. To this end, we propose MoR-MLLM, a computation-sparse MLLM based on the recent Mixture-of-Recursions (MoR) framework. MoR-MLLM introduces adaptive per-token recursion, allowing the model to dynamically adjust its recursive depth and allocate more computation to visually or linguistically challenging tokens while skipping redundant operations for simpler ones. To stabilize the training of recursive sparsity in multimodal settings, we further design a three-stage MoR-Tuning strategy and an entropy-regularized loss to encourage diverse routing distributions. Extensive experiments show that compared with recent advanced tiny MLLMs, our proposed MoR-MLLM can greatly reduce the training memory and computation complexity while retaining high performance on various vision-language tasks.

[CV-169] PanoPed: Beyond Bounding Boxes for Sim-to-Real Panoramic Pedestrian Tracking

链接: https://arxiv.org/abs/2610.08826
作者: Qinfeng Zhu,Weiguang Zhao,Yunxi Jiang,Anh Nguyen,Lei Fan
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Full-sphere panoramic cameras let fixed monitoring systems and mobile robots track people in every direction, but a planar bounding box does not fully describe where a person is on the sphere. We introduce PanoPed, a sim-to-real benchmark for pedestrian tracking on the full sphere. PanoPed-S contains 108,000 frames from fixed, quadruped-mounted, and drone-mounted cameras, with synchronized masks, depth, camera poses, and 3D pedestrian states. PanoPed-R adds 28,002 real frames from fixed cameras, 16,247 of them densely annotated. We find that an ERP rectangle cannot uniquely determine the spherical center and angular extent of the visible person, while the detector’s visual query still carries information about them. Inspired by the sextant’s use of angular measurements to locate objects, we propose Sextant, a plug-and-play angular localization head with only about 0.035M parameters. It reuses a frozen detector, keeps track identities unchanged, and needs no extra image encoder. Sextant gives the best result in our PanoPed-S test comparison, raising the strongest baseline, MOTIP, from 47.30 to 49.49 HOTA, with gains on all eight test sequences. Without fine-tuning on real data, the same synthetic-trained heads improve MOTIP and HAT by 0.96-1.14 HOTA on real video, and both seeds improve every real sequence. HAT+Sextant scores best among the compared systems that add no localization image encoder.

[CV-170] Autonomous Driving Research Requires a Community-Driven Data Paradigm NEURIPS2026

链接: https://arxiv.org/abs/2610.08825
作者: Jinsu Yoo,Zanming Huang,Katie Z Luo,Zheda Mai,Qiyuan Wu,Bharath Hariharan,Mark Campbell,Wei-Lun Chao
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注: NeurIPS 2026 Position Paper

点击查看摘要

Abstract:Autonomous driving has made remarkable progress, with recent AI advances enabling commercial deployments that are reshaping urban mobility. Yet the field remains far from its universal social promise: autonomous systems that can operate robustly anywhere, anytime, for anyone. We posit that this gap is not merely a modeling problem, but a problem of the prevailing data paradigm. Current research relies heavily on a few benchmark datasets with limited spatial and scenario coverage, even though the community has collectively produced over 600 autonomous driving datasets across nearly 50 countries. However, this abundance has not translated into broad research impact: most datasets remain significantly underused due to fragmentation, limited visibility, incompatible protocols, and benchmark incentives that concentrate attention on a few dominant datasets. We therefore argue that autonomous driving research requires a collaborative, community-driven data paradigm. Such a paradigm would improve the discovery, reuse, integration, and evaluation of diverse datasets; make underexplored data easier and more rewarding to study; and lower the barrier for new contributors. We outline its key principles, illustrate an early realization, and call for collaboration across academia and industry to transform fragmented datasets into shared community infrastructure for anytime-anywhere autonomy.

[CV-171] HCPN-GCN: Scaling Hierarchical Prototype Networks with Cone Geometry for Continual Graph Learning

链接: https://arxiv.org/abs/2610.08823
作者: Sammuel R. Silva,Vander L. S. Freitas,Gladston Moreira,Eduardo J. S. Luz,Rodrigo Silva
类目: Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV); Neural and Evolutionary Computing (cs.NE)
备注:

点击查看摘要

Abstract:Continual Graph Learning (CGL) aims to incrementally learn from graph-structured data while preserving knowledge acquired from previous tasks. A major challenge in this setting is catastrophic forgetting, where learning new tasks degrades performance on previously learned ones. Hierarchical Prototype Networks (HPNs) address this problem through a prototype-based memory mechanism that avoids storing historical data, but their reliance on linear feature extractors limits their ability to exploit graph topology, while point-based prototypes often lead to inefficient prototype growth on structurally diverse graphs. In this work, we propose HCPN-GCN, a graph-aware extension of HPN that replaces the original linear feature extractors with Graph Convolutional Networks (GCNs) and introduces cone-based prototypes with a diversity regularization objective. The proposed design produces richer graph-aware representations while compactly modeling the embedding space, reducing prototype proliferation without sacrificing discriminability. Experimental results on six continual graph learning benchmarks demonstrate that HCPN-GCN consistently improves average classification accuracy over the original HPN and representative continual learning baselines while maintaining near-zero forgetting. Furthermore, our analysis shows that the proposed model learns substantially richer class-level prototype hierarchies using approximately 30\times fewer atomic prototypes than the original HPN, providing a more compact and effective memory representation for continual graph learning.

[CV-172] Pre-training Reasoning Benchmarking: X-ray Report Generation on CheXpert Plus Dataset

链接: https://arxiv.org/abs/2610.08813
作者: Xiao Wang,Yuxiang Zhang,Dan Xu,Yuehang Li,Shiao Wang,Bo Jiang,Yaowei Wang,Yonghong Tian,Jin Tang
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:X-ray image-based Radiology Report Generation (RRG) constitutes a critical research direction within medical artificial intelligence, with great potential to alleviate clinicians’ diagnostic workload and shorten patient waiting periods. Despite substantial advances over recent years, the field faces evident bottlenecks stemming from insufficient standardized benchmarks and inadequate domain adaptation of generic large models. Notably, the newly released CheXpert Plus dataset is provided without accompanying baseline implementations and evaluation results, which impedes standardized training, quantitative evaluation and fair comparison among follow-up algorithms. To mitigate this limitation, we establish a comprehensive benchmark encompassing prevailing X-ray report generation models and Large Language Models on CheXpert Plus. This benchmark delivers a reliable comparative foundation for upcoming methods and enables researchers to rapidly identify state-of-the-art approaches within this domain. Beyond benchmark construction, we rethink X-ray RRG under the paradigm of large models and propose a novel framework termed MambaXray-PRB. Our framework improves report generation performance and enhances model interpretability via multi-stage large-model pre-training and multi-modal Chain-of-Thought reasoning. The pipeline consists of three successive phases: self-supervised auto-regressive modeling, X-ray-report contrastive learning, and post-training optimization for reasoning and report generation. Extensive experiments on IU X-ray, MIMIC-CXR, and CheXpert Plus datasets validate the effectiveness of MambaXray-PRB for radiology report generation. The source code of this paper is available on this https URL

[CV-173] A Review Of Robotic World Models For Dynamic Environments Based On Factor And Scene Graphs

链接: https://arxiv.org/abs/2610.08800
作者: Marco Giberna,Miguel Fernandez-Cortizas,Jose Luis Sanchez Lopez,Holger Voos
类目: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
备注: 34 pages, 7 tables, 8 figures

点击查看摘要

Abstract:Models based on graphs have emerged in robotics as a powerful foundation for internal world representations, where factor and scene graphs are among the most prominent model types found in the related literature and in successful robotic solutions. Initially, many of these models were assuming static environments as a simplification. Herein, factor graphs mainly provide uncertainty-aware geometric estimations while scene graphs enable a structured semantic abstraction. However, real-world robotic environments are often dynamic, posing severe challenges for purely static world representations. Therefore, this review presents a comprehensive view on how dynamic aspects of real-world environments can be addressed in such graph-based world models. We organize our assessments around three main aspects: (I) suitable representations, (II) pipelines to construct and update the representations, and (III) their exploitation for downstream tasks. We review approaches that are either based on factor or scene graphs, but put special emphasis on novel approaches that combine both types to form hybrid models. We mainly analyze how different types of dynamics can be modeled herein, and categorize common architectural patterns. Finally, emerging trends and open challenges are identified, including uncertainty propagation from learned perception through the representation layers, the observability of dynamic-entity motion and scale under minimal sensing, scalable lifelong maintenance, and the lack of datasets and evaluation protocols that ground world-model quality in downstream task performance under dynamics.

[CV-174] Efficient Dense Crowd Trajectory Prediction Via Dynamic Clustering

链接: https://arxiv.org/abs/2603.18166
作者: Antonius Bima Murti Wijaya,Paul Henderson,Marwa Mahmoud
类目: Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Crowd trajectory prediction plays a crucial role in public safety and management, where it can help prevent disasters such as stampedes. Recent works address the problem by predicting individual trajectories and considering surrounding objects based on manually annotated data. However, these approaches tend to overlook dense crowd scenarios, where the challenges of automation become more pronounced due to the massiveness, noisiness, and inaccuracy of the tracking outputs, resulting in high computational costs. To address these challenges, we propose and extensively evaluate a novel cluster-based approach that groups individuals based on similar attributes over time, enabling faster execution through accurate group summarisation. Our plug-and-play method can be combined with existing trajectory predictors by using our output centroid in place of their pedestrian input. We evaluate our proposed method on several challenging dense crowd scenes. We demonstrated that our approach leads to faster processing and lower memory usage when compared with state-of-the-art methods, while maintaining the accuracy

[CV-175] FaceKit: a Toolkit for Interpretable Facial Phenotyping Synthetic Image Generation and Privacy Analysis in Rare Diseases

链接: https://arxiv.org/abs/2610.09288
作者: Hongzhuo Chen,Zhanliang Wang,Florent Pollet,Mian Umair Ahsan,Joshua Bie,Tzung-Chien Hsieh,Peter Krawitz,Cong Liu,Wendy K Chung,Chunhua Weng,Gamze Gürsoy,Kai Wang
类目: Quantitative Methods (q-bio.QM); Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Many rare genetic diseases are associated with recognizable craniofacial features. However, traditional approaches for describing facial morphology rely largely on qualitative clinical observation and free-text descriptions, which are often subjective, non-standardized, and difficult to reproduce across observers and institutions. Although the Human Phenotype Ontology (HPO) provides controlled terms for describing facial features, these terms are typically categorical rather than quantitative and may vary depending on examiner experience and interpretation. Here, we present FaceKit, a computational framework for quantitative facial phenotyping from frontal facial photographs. FaceKit extracts standardized measurements of facial landmarks and derived 120 morphological features, then reports feature-level z-scores representing deviation from population reference distributions. The reference distributions are built from the FairFace dataset spanning diverse ancestral groups. We evaluated FaceKit on a curated subset of the GestaltMatcher Database covering 50 rare-disease cohorts. In addition to quantitative facial analysis, FaceKit includes synthetic facial image generation to support rare disease model development and data augmentation. We also performed privacy evaluation to assess whether synthetic images reveal identifiable information from real patient photographs and could compromise patient privacy. Across disease case studies, FaceKit-derived quantitative measurements captured known facial features associated with rare genetic disorders and provided objective support for clinical phenotyping. Together, these results establish FaceKit as a useful tool for quantitative phenotyping, and has the potential to improve rare disease diagnosis, support genotype-phenotype studies, and enable more reproducible clinical characterization across diverse patient populations.

[CV-176] Abdominal Ultrasound Simulation from Semantic Labels using Paired Label-to-Physics-Based Image Translation

链接: https://arxiv.org/abs/2610.08849
作者: Santiago Vitale,Duilio Deangeli,Ignacio Larrabide,José Ignacio Orlando
类目: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Purpose: Current abdominal ultrasound (US) simulation methods often require CT-based anatomical references for ray-casting, limiting deformation and pathology variability. We propose a learning-based pipeline trained to predict physics-based images derived from CT scans from semantic labels, enabling controlled simulations without patient-specific CT volumes at inference time. Methods: We introduce a two-stage pipeline that maps anatomical segmentations to realistic US images through a simplified US image. Stage~I synthesizes this image from semantic labels using models trained on CT-based ray-casting outputs. Stage~II refines it into a realistic US scan using anatomically guided unpaired translation. Deformations and pathologies are generated by editing anatomical maps. Results: We evaluated Pix2Pix and the Semantic Diffusion Model (SDM) in Stage~I, followed by segmentation-guided CycleGAN (SG-CycleGAN) refinement in Stage~II. SDM significantly outperformed Pix2Pix in morphological metrics, including MAE (19.85 vs. 21.65), SSIM (0.28 vs. 0.24), and mIoU (0.43 vs. 0.29), whereas Pix2Pix yielded better perceptual point estimates (LPIPS: 0.17 vs. 0.19; FID: 0.32 vs. 0.37; KID: 0.25 vs. 0.48). Conclusion: Training paired generative models with physics-based supervision enables approximation of CT-derived ray-casting outputs at inference time directly from semantic labels. Although the pipeline does not require patient-specific CT volumes at inference time, CT-derived segmentations and ray-casting simulations remain necessary to train Stage~I. Once trained, the framework enables controllable healthy and pathological simulations through semantic-map modification. Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV) Cite as: arXiv:2610.08849 [eess.IV] (or arXiv:2610.08849v1 [eess.IV] for this version) https://doi.org/10.48550/arXiv.2610.08849 Focus to learn more arXiv-issued DOI via DataCite (pending registration) Submission history From: Duilio Deangeli [view email] [v1] Fri, 2 Oct 2026 16:26:46 UTC (19,611 KB)

[CV-177] STRIDE: Spatial-Temporal Representation for Interval-conditioned Disease Evolution in Longitudinal Glioblastoma MRI

链接: https://arxiv.org/abs/2610.08848
作者: Wenhao Guo,Changchang Yin,Pierre Giglio,Weidan Cao,Ping Zhang,Golrokh Mirzaei
类目: Image and Video Processing (eess.IV); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
备注: 37 pages, 11 figures

点击查看摘要

Abstract:Glioblastoma (GBM), an aggressive primary brain tumor, is routinely monitored with longitudinal MRI after treatment. Distinguishing stable disease (SD), pseudoprogression (PsP), and true progression (TP) remains challenging because these states can show overlapping MRI appearances despite different temporal trajectories. Existing longitudinal methods still face challenges in modeling scan-specific spatial variability, variable follow-up intervals, and complementary information from the observed follow-up state and its longitudinal change. We propose STRIDE, a framework for spatial-temporal representation of interval-conditioned disease evolution that takes paired post-treatment MRI scans and their inter-scan interval as input and predicts SD, PsP, or TP. The lesion-prior-guided spatial representation combines SoftGate and an Adaptive-window Hierarchical Transformer (AWHT) to emphasize lesion-related regions while preserving surrounding context. The time-conditioned latent transition uses pair-level context and the actual inter-scan interval to estimate interval-dependent representation changes between visits. The observed–transition fusion integrates the transition-estimated follow-up representation with the directly observed follow-up representation to jointly characterize the follow-up state and its longitudinal change. BraTS2024 is used to develop and evaluate the lesion-prior generator, while longitudinal pretraining on LUMIERE supports transfer before downstream adaptation to Burdenko. On the Burdenko three-class task, STRIDE achieves a macro ROC–AUC of 0.816 and a macro F1-score of 0.796. These results support its potential for more reliable longitudinal post-treatment GBM state assessment.

[CV-178] Deep Learning for Longitudinal Medical Imaging: A Scoping Review

链接: https://arxiv.org/abs/2610.08838
作者: Francesca Mussa,Divyanshu Tak,Atlas H. Avval,Sarah Brueningk,Ray H. Mak,Hugo J. W. L. Aerts,Andreas M Rauschecker,Benjamin H. Kann
类目: Image and Video Processing (eess.IV); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Longitudinal medical imaging analysis is a cornerstone of modern medical practice and patient care. Deep learning applied to longitudinal imaging offers wide potential to enhance diagnosis and track disease progression by capturing spatial changes over time. With major advances in single-timepoint deep learning for imaging, there has been growing interest in longitudinal image analysis, given its increased clinical relevance, though technical challenges remain. Several recent innovations may lead to a new era of multi-timepoint image evaluation, yet the scientific landscape, recent progress, and areas of need remain under-characterized. To address this gap, we conducted a scoping review of deep learning methodologies applied to longitudinal medical imaging, yielding 102 studies published between 2018 and 2025. Neurological disorders (48%) and ophthalmic conditions (12%) were the most common clinical applications, with MRI serving as the predominant imaging modality (67%). Sequential feature modeling approaches combining convolutional neural networks (CNNs) with temporal models (LSTM/RNN) were the most frequent methodology (40%), followed by direct feature aggregation across timepoints (23%). Most studies targeted classification tasks (56%), while external validation was performed in only 24% of studies. Our findings highlight that deep learning-based longitudinal imaging analysis remains a promising field, though newer temporal architectures and larger datasets may improve success and clinical adoption of these tools.

人工智能

[AI-0] RoboJEPA: Scaling Robotic Latent World Models

链接: https://arxiv.org/abs/2610.10515
作者: Artem Zholus,Nicolas Beltran-Velez,Jianhao Yuan,Sarath Chandar,Tushar Nagarajan,Daniel Severo,Koustuv Sinha,Michal Drozdzal,Adriana Romero Soriano,Jeannette Bohg,Nicolas Ballas,Mahmoud Assran
类目: Artificial Intelligence (cs.AI); Robotics (cs.RO)
备注:

点击查看摘要

Abstract:Latent world models have shown a remarkable ability to predict future states and to plan in the real world. In practice, however, we lack a principled way to estimate how their capabilities scale with model size, data, and compute, an open problem that slows progress in the field. In this work we present RoboJEPA, a world model based on the Joint Embedding Predictive Architecture (JEPA) and trained on a large-scale dataset spanning 12 robotic embodiments. We show that RoboJEPA’s imagination error, the error of its latent rollouts, follows a second-order power law in compute, allowing us to predict model quality well beyond the scale at which the law is fit. We further show that downstream robotic planning performance improves predictably with compute, and that imagination error is strongly correlated with it, making it a reliable proxy for real-robot evaluation. Finally, we demonstrate that latent world models can be deployed zero-shot as robotic agents, planning toward a single goal image to solve tasks requiring long-horizon planning on real hardware. We release all model checkpoints together with our training and robot deployment code. To our knowledge, this is the first work to establish scaling laws for multi-embodiment robotic world models trained on real robot data, and RoboJEPA, at 8B parameters, is the largest JEPA predictor model trained to date.

[AI-1] SciExam for ENSO: Can AI Agents Build Climate Models?

链接: https://arxiv.org/abs/2610.10513
作者: Yinling Zhang,Langchen Liu,Dongbin Xiu,Xueyan Zou,Xu Kuang,Mengdi Wang,Shilong Liu
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Atmospheric and Oceanic Physics (physics.ao-ph)
备注: 28 pages, 5 figures, 8 tables. Code: this https URL

点击查看摘要

Abstract:Language-model agents are increasingly asked to carry out open-ended scientific research, yet their results are usually graded against a known answer, a rubric, or a language-model reviewer, none of which can tell whether a new scientific model is valid. The AI Science Exam for El Nino-Southern Oscillation (SciExam for ENSO) is a benchmark in which agents build low-order stochastic models of ENSO, the dominant mode of interannual climate variability, from real observations. Within a six-hour budget, agents process the observations, write their own diagnostics, which are then frozen, and develop a model using only these diagnostics as feedback. Hidden graders then test whether the model reproduces ENSO’s statistics, recovers unobserved variables, and forecasts held-out years, and score a published model in the same way. Across twelve agent systems, six produce models that score higher than the published model, mainly through better reconstruction and forecasting. The simplified forms of the stronger models are each compatible with one of the two competing explanations of ENSO’s warm-cold asymmetry, an open debate that the task never mentions. Controlled runs of the top system under varied information suggest that its scores do not come from recalling the dated observational record and that the information it receives shapes how it builds its model. SciExam for ENSO can thus evaluate agent research where no answer is known, and the results suggest that agents can already build competitive models whose structures bear on questions that scientists still debate.

[AI-2] RECAST: Learning to Compute the Right Context through Adaptive Evidence Routing

链接: https://arxiv.org/abs/2610.10507
作者: Yilun Hao,Krishna Sayana,Isabella Ye,James S Ren,Sukhdeep Sodhi,Craig Boutilier,Chuchu Fan
类目: Artificial Intelligence (cs.AI)
备注: 35 pages, 2 figures, 13 tables

点击查看摘要

Abstract:Large language models are increasingly applied to tasks grounded in long, heterogeneous information sources. Conventional Retrieval-Augmented Generation (RAG) relies on fixed similarity-based retrieval, while agentic variants adapt queries and tool use but remain largely retrieval-centric. However, in many tasks, the evidence required for a solution is not explicitly present in any single source item. Instead, it must be derived through filtering, aggregation, or computation across multiple source items. In this work, we introduce RECAST (Routing Evidence through Computation, Access, and Synthesized Tools), a learned framework that formulates evidence construction as a sequential decision process over heterogeneous retrieval and computation operations, allowing evidence to be actively derived rather than merely retrieved. A lightweight RouterLM iteratively selects and formulates primitive operations or specifies customized operations for a frozen CompilerLM to translate into executable code. Once it judges the evidence sufficient, RouterLM passes the accepted evidence to a frozen AnswerLM to produce the final solution. We train RouterLM with supervised fine-tuning (SFT) followed by group relative policy optimization (GRPO). Across six heterogeneous benchmark families, RECAST achieves a mean success rate of 75.6%, outperforming the strongest large-model baseline by 15.9%. Moreover, training enables the Qwen3.5-9B RouterLM to outperform a training-free Gemini 3.5 Flash RouterLM by 5.0%. On three held-out benchmarks, RECAST improves over the strongest baseline by 15.0% on average, demonstrating strong zero-shot generalization across tasks and heterogeneous source representations.

[AI-3] EmbodiedRSI: Active Continual Robot Learning Through Hypothesis-Guided Co-Evolution

链接: https://arxiv.org/abs/2610.10498
作者: Python Song,Zhixuan Liang,Kelsey Fu,Mengdi Wang,Junfeng Yang,Shilong Liu
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Robot foundation models provide strong visuomotor control, yet their performance can degrade when object positions or task instructions change. Further improvements often require post-training on substantial robot data, which can be costly to collect through methods such as teleoperation. Agentic harnesses can adapt around the model, but current self-evolving harnesses use robot trials inefficiently when deciding which code and skill changes to pursue. We introduce EmbodiedRSI, a self-evolving agentic harness that autonomously decides where to explore next and turns the resulting physical interaction into improved code and skills. EmbodiedRSI realizes this through a Fast-Slow Dual-System Architecture, in which competing code and skill hypotheses are maintained in a Hypothesis Graph. Value-of-Information Experiment Selection chooses physical experiments that can distinguish these hypotheses. Their outcomes guide Code-Skill Co-Evolution. The Slow System builds Hierarchical Memory, and Reward-Grounded Memory Learning selects effective memory according to their value for later Fast-System improvement. On RoboCasa365, EmbodiedRSI reaches 77.0% overall success and 71.3% on Composite-Unseen, compared with 40.1% for the best baseline. EmbodiedRSI also reaches 86.8% overall success on LIBERO-Pro. Beyond benchmark performance, EmbodiedRSI transfers zero-shot to real-world robot, achieving 71.3% overall success across multiple challenging tasks.

[AI-4] Before They Can Solve: Predicting Post-Training Coding-Agent Performance from Base Models

链接: https://arxiv.org/abs/2610.10478
作者: Tan Yu,Alexander Bukharin,Khushi Bhardwaj,Jennifer Williams,Zirui Liu,Jonathan Lingjie Li,Soumye Singhal,Joseph Jennings,Sanjeev Satheesh,Yash Jain,Ashish Vaswani,Venkat Krishna Srinivasan,Matthew Papakipos,Hyunwoo Kim,Jian Zhang,Oleksii Kuchaiev,Markus Kliegl,Mostofa Patwary,Mohammad Shoeybi,Bryan Catanzaro,Jonathan Cohen,Jiantao Jiao
类目: Artificial Intelligence (cs.AI); Software Engineering (cs.SE)
备注:

点击查看摘要

Abstract:How can we predict which base checkpoint is worth an expensive round of agentic post-training? End-to-end pass@ K tests whether successful behavior already appears in a base model’s distribution, but it is a poor fit for agentic coding: many base checkpoints cannot reliably produce the well-formed tool invocation required to complete a task end-to-end. Single-shot or short-horizon tasks avoid these tool-calling failures by collapsing a multi-step interaction into a fixed prompt and a single patch, but they sidestep the core capability we care about: maintaining coherent state over many tool-using steps as the repository evolves. To bridge this gap, we treat successful post-trained agent trajectories as a lookahead signal of base-model potential. Replaying each trajectory and rerunning tests after every code-changing step identifies the decisive step: the first step whose cumulative patch flips the repository from failing to passing, certifying that the recorded action solves the task given the prior context. Motivated by a coverage principle for agentic traces, we build three screens at this step that do not require a base checkpoint to drive the harness from a cold start: (i) Decisive-Action BPB (bits per byte) measures the probability mass on the certified action, (ii) Patch MCQ tests the checkpoint’s choice between that action and alternatives rejected by the same verifier, and (iii) prefix-conditioned pass@ K evaluates support for functionally-correct generations and credits any continuation that the tests accept. Across ten pairs of public base and post-trained models, all three screens rank the cohort in close agreement with post-trained SWE-bench Verified pass@ 1 . As our methods need only a benchmark’s successful trajectories and its verifier, they can be applied to turn future agentic coding benchmarks into base-model evaluations.

[AI-5] FoldBack: Self-Correcting Masked Generative Policy for Long-Horizon Garment Folding

链接: https://arxiv.org/abs/2610.10462
作者: Lipeng Zhuang,Shiyu Fan,Yingdong Ru,Zhuo He,Florent P. Audonnet,Paul Henderson,Gerardo Aragon Camarasa
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:We present FoldBack, a self-correcting masked generative policy for long-horizon garment folding. Existing long-trajectory policies may continue after a missed or slipped grasp even when the garment has not reached the intended configuration. We structure FoldBack’s recovery mechanisms around three inference-time decisions: when to refine and verify, how to roll back, and where and how to retry. FoldBack aligns refinement and grasp verification with pick-and-place events, returns the robot to a retryable pre-grasp configuration while preserving successful grasps, and selectively regenerates the failed segment and selected future actions while avoiding previous failed grasp locations. To our knowledge, FoldBack is the first editable full-trajectory policy to unify these decisions, enabling failed interactions to be detected, undone, and repaired before execution continues, without recovery demonstrations or base-policy retraining. Across 33 real garments from six categories, FoldBack achieves 75.2% final folding success and 0.837 final-mask IoU, versus 45.7% and 0.689 for the strongest prior baseline.

[AI-6] Composing What Each Teacher Learned: Multi-Teacher On-Policy Distillation through Teacher-Relative Shifts

链接: https://arxiv.org/abs/2610.10460
作者: Hejian Sang,Zhengze Zhou,Shayan Mohajer Hamidi,Xiaomin Li,Rohit Jain,Alborz Geramifard
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Multi-teacher on-policy distillation (MOPD) is used in two settings. In common-domain composition, several teachers score each student rollout from one prompt domain and their signals form a single target; in routed-domain distillation, prompts from different domains are assigned to the corresponding specialist. Both settings usually transfer each teacher’s endpoint policy, which mixes what post-training changed with preferences inherited from the teacher’s base. We introduce \Delta -MOPD, which transfers each teacher’s teacher-minus-base logit shift re-anchored at the student’s frozen initialization, and compare it with endpoint supervision in both settings while holding teacher selection fixed. We first expose the mechanism that impedes endpoint transfer: inherited base pull can exceed the post-training shift. Removing it reduces the teacher-term norm ratio and target–student KL. Across our experiments, the results suggest that shift targets are particularly useful when teacher signals are combined at a state. With three composed teachers, \Delta -MOPD exceeds endpoint composition by 4.11 Math and 1.95 five-benchmark points; with two, it matches endpoint accuracy. Under phased routing, it achieves higher mean performance in both phase orders and reduces the observed order gap from 10.50 to 6.42 points. Under interleaved routing, where each update involves one teacher, the two targets perform comparably. The phased results provide supporting evidence that the benefit may extend to signals accumulated across training phases. Target construction is thus an independent design axis in MOPD, complementary to teacher selection. Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI) Cite as: arXiv:2610.10460 [cs.LG] (or arXiv:2610.10460v1 [cs.LG] for this version) https://doi.org/10.48550/arXiv.2610.10460 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[AI-7] A Good Self-Teacher Meets the Student Where They Are: Joint On-Policy Learning and Teaching

链接: https://arxiv.org/abs/2610.10447
作者: Randy Ardywibowo,Arnav Dalal,Jiantao Jiao
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Reinforcement Learning (RL) from outcome rewards suffers from sparse supervision, particularly on difficult, long-horizon tasks where successful trajectories are rare and costly to generate. On-Policy Distillation (OPD) offers an attractive alternative by providing dense token-level supervision from a stronger teacher along the student’s own generations. Self-distillation methods further remove the need for a separate teacher model by conditioning the same policy on privileged information to serve as its own teacher. However, privileged conditioning alone does not guarantee that the resulting distillation update improves the student. Indeed, privileged information can lead the teacher to solve tasks through shortcuts unavailable to the student, producing supervision poorly matched to the student’s current behavior. Consequently, even a higher-performing teacher can provide guidance that degrades student performance. To address this, we analyze how the choice of privileged teacher affects the student’s update. We derive a necessary and sufficient condition for the teacher’s local distillation update to be a positive multiple of the student’s reward gradient. Our analysis suggests that the teacher should not only perform well on the task, but also provide guidance suited to the student’s current capabilities. This characterization motivates a practical teacher-training surrogate that combines outcome rewards with token-level Kullback-Leibler (KL) regularization toward the student. Based on this result, we propose Joint On-Policy Learning and Teaching (JOLT), which jointly trains a single policy in two roles: a privileged teacher using a KL-regularized objective, and an unprivileged student using dense on-policy distillation. Across mathematical reasoning, coding, tool use, and terminal use, JOLT improves training efficiency and performance, with further gains from student rewards.

[AI-8] Q-Learning with Scalar Adjoint Matching

链接: https://arxiv.org/abs/2610.10437
作者: Yonghoon Dong,Minsung Yoon,Jaehyuk Kim,Jungwoo Park,Changyeon Kim,Jinwoo Shin
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Robotics (cs.RO)
备注:

点击查看摘要

Abstract:Flow policies capture rich and diverse action distributions, and fine-tuning them with off-policy RL to improve beyond the demonstrations has drawn growing interest. However, fine-tuning a flow policy against a learned value function is not trivial, because the policy generates its action over many flow steps. Adjoint matching offers a principled way to update the flow model itself by propagating value information from the final action back to each flow step, but it requires a vector–Jacobian product through the policy at every step, a cost that grows with the number of flow steps and the policy size. We observe that the batch-averaged velocity Jacobian of pretrained flow policies concentrates on its diagonal. Motivated by this finding, we derive a closed-form scalar adjoint that scales the value gradient at the final action by the flow time, eliminating the per-step vector–Jacobian products. We further find that controlling the critic’s value at policy-generated actions is particularly important under the scalar adjoint. Based on these findings, we propose Q-learning with Scalar Adjoint Matching (SQAM), which combines the scalar adjoint with a value penalty at those actions. SQAM’s gains concentrate on the four hardest OGBench domains, where its success rate exceeds that of the strongest baseline in each domain by 18 to 35 percentage points. To test whether SQAM extends to large pretrained policies, we also fine-tune a vision-language-action policy on a real bimanual robot. SQAM improves over supervised fine-tuning on all three tasks.

[AI-9] SOTA: Stock Options Trading Agents Guided by Option-Implied Return Distributions NEURIPS2026

链接: https://arxiv.org/abs/2610.10407
作者: Yizhen Xie,Mengyang Liu
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Portfolio Management (q-fin.PM); Trading and Market Microstructure (q-fin.TR)
备注: Accepted at the NeurIPS 2026 Agenthon Workshop

点击查看摘要

Abstract:As option markets grow and AI advances, agentic systems for option trading are gaining increasing attention. Language-model-based agents can reason over contextual information such as news, but option trading presents a particularly challenging decision problem: a single stock can have thousands of contracts, and the agent must decide both which contracts to trade and how to combine them. Existing approaches often sidestep this complexity by restricting the policy to a fixed strategy structure, such as a straddle, limiting their ability to switch strategies as market conditions change. We present SOTA (Stock Options Trading Agents), an agentic trading framework for structured option-strategy selection. SOTA abstracts the large option universe into strategy-level decisions while deterministic resolvers handle portfolio implementation. We develop SOTA by post-training Qwen3.8-27B with supervised fine-tuning followed by reinforcement learning. SOTA is evaluated on options on nine large-cap U.S. equities and SPY against rule-based and machine-learning strategy selectors in the same trading environment. Over a six-month out-of-sample period, SOTA earns an 18.3% total return with a Sharpe ratio of 1.60 and a maximum drawdown of 8.96%. We also document an asymmetric role of news: news improves frontier-teacher trajectories, but retaining news during reinforcement learning reduces out-of-sample return from 18.3% to -2.7%.

[AI-10] aoD2C-Bench: Benchmarking MLLM s for Industrial UI Code Generation Beyond Visual Fidelity

链接: https://arxiv.org/abs/2610.10374
作者: Chengwei Shi,Yunnong Chen,Tingting Zhou,Qiang Lu,Shiyu Yue,Xinyuan Hu,Jianfang Ru,Liuqing Chen
类目: oftware Engineering (cs.SE); Artificial Intelligence (cs.AI)
备注: 26 pages

点击查看摘要

Abstract:A key challenge for multimodal large language models (MLLMs) is moving beyond visual recognition to constraint-aware cross-modal reasoning. This involves combining visual cues with information from other modalities to understand elements’ relationships under domain-specific rules. This challenge is acutely evident in industrial design-to-code (D2C), which converts user interface (UI) designs into code and requires MLLMs to connect design images with disorganized layer metadata, infer component and layout implementation requirements, and realize them in code under target-library constraints. However, these capabilities remain insufficiently evaluated in realistic industrial settings. To fill this gap, we present TaoD2C-Bench, a benchmark for evaluating MLLMs’ ability to generate UI code that satisfies implementation requirements in industrial applications. The TaoD2C dataset consists of 2,861 production designs from 17 commercial platforms with 97,652 expert annotations across four categories: Component, Group, Alignment, and Position. These annotations distinguish required constraints from permitted implementation choices. TaoD2C-Bench defines three tasks: end-to-end UI code generation, requirement inference, and requirement realization. Evaluating eight MLLMs reveals substantial gaps in generating UI code that satisfies implementation requirements, alongside distinct performance profiles in inference and realization. We further show that MLLMs’ visual reconstruction ability does not necessarily imply an ability to generate code that meets these requirements. We release TaoD2C to support research on industrial UI code generation.

[AI-11] Open-MMUnlearning: Unifying Methods and Evaluation for MLLM Unlearning

链接: https://arxiv.org/abs/2610.10358
作者: Junkai Chen,Yuhao He,Qianshan Wei,Junxiang You,Jingwen Shao,Junkai Lin,Zhongkai Yue,Xiaotian Ye,Zhengbo Jiao,Jiali Cheng,Zhijie Deng,Kening Zheng,Ruiqi Liu,Hadi Amiri,Yi Yu,Zhenan Sun,Qi Li,Ka-Ho Chow,Sijia Liu,Liang Wang,Jiaqi Li,Shu Wu
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:As multimodal large language models (MLLMs) become more capable and widely deployed, concerns about privacy and safety have become increasingly pressing. Machine unlearning offers one approach to addressing these concerns by removing designated information from trained models while preserving unrelated capabilities. However, fragmented implementations and evaluation protocols, incomplete robustness testing, and limited understanding of metric reliability make progress in MLLM unlearning difficult to assess systematically. We introduce Open-MMUnlearning, an open-source, extensible framework that integrates target-model preparation, multimodal data processing, unlearning, and evaluation through shared interfaces and structured configurations. The framework supports five benchmarks spanning privacy, safety, and copyright, eight MLLMs from four model families, and twelve unlearning methods. Its evaluation suite jointly assesses forgetting effectiveness, retained utility, and robustness to model interventions, adversarial inputs, and membership inference attacks. Using a common evaluation protocol, we compare ten representative unlearning methods. In this comparison, GD and MIP-Editor tie for the highest overall score: GD achieves the highest Forget Quality, while MIP-Editor preserves more Model Utility. We further introduce a metric meta-evaluation protocol that tests faithfulness using models with controlled exposure to target knowledge and robustness under quantization and relearning. Among the thirteen evaluated metrics, BLEU achieves the highest aggregate reliability score. KS-Test attains the highest faithfulness AUC but performs less well on robustness. Together, the framework and these findings support reproducible comparison of MLLM unlearning methods and systematic assessment of evaluation reliability.

[AI-12] MIRA: A Musical Intent Refinement Agent for Aligning Text-to-Music Generation with User Intent

链接: https://arxiv.org/abs/2610.10355
作者: Zekai Liu,Zhilin Wang,Xuzheng He,Yu Cheng,Yang Yang
类目: ound (cs.SD); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Text-to-music systems produce increasingly convincing audio, yet evaluation reveals little about whether the result matches user intent. A global text-audio relevance score can overlook the implicit intent in underspecified prompts and mask failures in specific requirements, such as instrumentation, structure, rhythm, or mood progression. To bridge this gap, we formulate text-to-music intent alignment as satisfying a per-request rubric of independently verifiable items covering both a request’s explicit requirements and its implied musical intent. Scoring items individually makes evaluation diagnostic by intent source and musical dimension, rather than a single opaque score. We instantiate this as MuRA-Bench, a benchmark of real-world platform requests curated by music experts. We further propose MIRA (Musical Intent Refinement Agent), a test-time agent that first grounds a request’s intent into rubrics, then searches over prompt revisions for a black-box generator under a bounded budget, iteratively generating music, verifying it against the rubrics, and using this feedback to guide a trajectory-aware tree search. Experiments across open-source and commercial backends show that MIRA improves intent alignment, enabling an open-source generator to achieve performance comparable to representative commercial systems (e.g. Suno and Mureka). Project page: this https URL.

[AI-13] SLDR: Defending Against Malicious Fine-tuning via Selective Layers Recovery and Dynamic Routing NEURIPS2026

链接: https://arxiv.org/abs/2610.10345
作者: Hui Zhang,Yachao Yuan,Jiayun Wang,Yuanzhuo Li,Hongtao Wang,Yali Yuan
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
备注: Accepted at NeurIPS 2026

点击查看摘要

Abstract:Fine-tuning-as-a-service enables users to adapt aligned large language models (LLMs) to specialized tasks, but malicious fine-tuning can erode refusal behavior while preserving task performance on legitimate inputs. We revisit recent layer-wise safety diagnostics and find that safety sensitivity is signed: scaling different layers can strengthen refusal, weaken it, or have little effect. Motivated by this observation, we propose SLDR, a post-fine-tuning defense based on Selective Layers Recovery and Dynamic Routing. SLDR trains a LoRA recovery adapter only on the layers with the maximum and minimum sensitivity scores in the signed spectrum, and uses representation-based dynamic routing inference to activate the adapter only for malicious queries. Across four model architectures, five downstream tasks, and four harmful benchmarks, SLDR substantially reduces harmful outputs while preserving downstream utility. On Llama3.1/SST2, SLDR reduces the average harmful score from 11.54 to 0.08 while maintaining downstream accuracy, and the harmful score remains near zero under poisoning ratios up to 0.9. The code is available at this https URL.

[AI-14] Fault-tolerant foundation models

链接: https://arxiv.org/abs/2610.10311
作者: Trevor McCourt,Ila R. Fiete,Isaac L. Chuang
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Hardware Architecture (cs.AR)
备注:

点击查看摘要

Abstract:Emerging computer hardware often trades reliability for energy efficiency; here we show that large-language models (LLMs) can be trained to tolerate this unreliability, and that rather than degrading, their error resilience actually increases as they grow. Modified neural scaling laws inferred from 40,000 GPU-hours of training runs on simulated faulty digital hardware quantify this trend and suggest that models learn to compute within “good” error-correcting codes, whose relative overhead remains finite no matter how large the model gets. This finding leads us to conjecture that appropriately trained LLMs may be formally fault-tolerant; if true, running AI inference on low energy, faulty hardware may be a path to substantial energy savings over the status quo.

[AI-15] AI Safety Considerations for Agents With Limited Time to Act

链接: https://arxiv.org/abs/2610.10285
作者: Leo Zeitler,Jack Richings,Victoria Nockles
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:In the wake of the increasingly public discussion about AI alignment, recent work has tried to propose specific AI architectures that behave safely. However, the proposed arguments that seemingly demonstrate proved alignment mostly neglect the environment the agent needs to act in. We discuss theoretical bounds for agent-agnostic safety guarantees in environments that can only be partially observed and within which an action is required within limited time. We introduce two realistic scenarios, one with an infinite state space and one with signal mixture. In these scenarios, we prove that even a perfect agent cannot guarantee safe behaviour. It will be argued that for any proof of AI safety or alignment, the environment and associated safe actions need to be specifically considered together with the agent.

[AI-16] Stale Misattributed or Late: Where Personal Memory Fails Before Generation

链接: https://arxiv.org/abs/2610.10265
作者: Haonan Deng,Park Sinchaisri
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Personal memory for language agents is usually judged by whether the final an- swer is correct. That score hides errors that arise before generation: the memory block may contain an obsolete value, a fact about the wrong person, or no use- ful fact before the serving deadline. We measure these failures directly. Using Personal Fact Memory (PFM) as a reference layer, we find that temporal validity is primarily a property of memory construction in our setting. On a controlled revision benchmark, serving only the active value of each correctly keyed slot eliminates observed stale exposure; without update resolution, 70.3% of prompts expose a superseded value. Once retrievers share the same active store and par- ticipant information, participant-aware BM25 is equivalent to the reference ranker within a prespecified 0.02 margin. The harder problem is assigning revisions to the right slot. Missed merges leave stale values active, whereas false merges silently remove current values; four LLM key assigners achieve higher key re- call than a rule extractor yet produce lower clean-retrieval rates, and open-domain merge recall on LongMemEval never exceeds 0.062. Misattribution survives va- lidity filtering: an entity posterior reduces same-name exposure on controlled data but cannot distinguish identically named speakers in LoCoMo. Two frozen lan- guage models reproduce prompt errors in generated text. Retrieval latency varies across rankers, but prompt prefill dominates turn-level latency on our hardware. These results argue for evaluating agent memory before generation, separating stored-state validity, identity resolution, abstention, and serving latency.

[AI-17] QuSema: Detecting Silent Bugs in Quantum Libraries via Quantum-knowledge-enhanced Agents

链接: https://arxiv.org/abs/2610.10258
作者: Yujin Song,Kaining Zhang,Qixin Zhang,Shuai Wang,Pingchuan Ma,Yuxuan Du
类目: oftware Engineering (cs.SE); Artificial Intelligence (cs.AI); Quantum Physics (quant-ph)
备注:

点击查看摘要

Abstract:Quantum libraries are now critical infrastructure for quantum algorithm development, yet their correctness remains difficult to test. Existing testing techniques mainly rely on failure-based or comparison-based oracles, exposing bugs only when executions fail, violate runtime checks, or disagree with another implementation. Their applicability is limited when suitable execution-based oracles are unavailable, leaving some silent bugs undetected. Such missed bugs can produce incorrect results that propagate into experimental conclusions, simulation studies, and algorithmic designs. Here we present QuSema, an autonomous testing agent for finding silent bugs in quantum libraries. QuSema uses constraints from quantum semantics and documentation as a source-level semantic oracle to assess whether implementation logic can produce invalid outputs from valid inputs. It operates through an agentic loop that repeatedly inspects library API documentation and source code, reasons about the intended behavior of quantum operations, identifies potential semantic deviations, and validates them by generating executable tests through library APIs. Guided by quantum-domain reasoning, QuSema turns high-level behavioral mismatches into concrete, user-triggerable bug reports, enabling it to uncover non-crash defects. We implement QuSema for Qiskit and PennyLane. On a benchmark of 20 historical silent bugs, QuSema achieves higher mean bug relocation counts than Claude Code and Codex, with the DeepSeek configuration costing less than Claude Code. QuSema also discovers 40 previously unknown bugs confirmed by the developers, including 30 silent bugs.

[AI-18] OOM-RL II: Reality Is an Oracle Not a Debugger Provenance-Constrained Diagnosis in Continually Evolving Agent -Engineered Systems

链接: https://arxiv.org/abs/2610.10256
作者: Kun Liu,Liqun Chen
类目: Artificial Intelligence (cs.AI); Software Engineering (cs.SE); Portfolio Management (q-fin.PM)
备注: 38 pages, 14 figures, 9 tables. Supplementary Dataset S1: this https URL . Follow-up to arXiv:2604.11477

点击查看摘要

Abstract:Reality may establish that an outcome occurred without identifying which evolving procedure produced it or why. This distinction matters in production ML systems whose code, configuration, and artifacts change while external feedback accumulates. We examine it in a human-directed, agent-engineered quantitative trading system, using oracle to mean an external source of realized outcomes rather than a complete correctness specification. Across one year, the account gained and outperformed a broad market index, while annual alpha was not statistically distinguishable from zero under the main retrospective specification. Retrospectively selected subperiods include adverse relative performance and conditional candidate-level weakness under declared approximate references. Engineering records document changes during the episode, and complete recommendation-to-runtime binding is unavailable. The archive does not establish a common frozen instance or a unique cause. The case motivates an outcome–diagnosis gap: outcome evidence, evaluated-object identity, and causal explanation support distinct claims. We distinguish frozen instances, pre-specified adaptive procedures, and ad-hoc development; organize archive-relative claim identifiability and an evidence hierarchy; and propose a prospective production-binding protocol. An illustrative compatible-history example shows how factual binding can resolve a recommendation’s referent without supplying its counterfactual effect. The protocol is proposed rather than prospectively validated. External feedback constrains outcome claims, while provenance and additional identification structure determine the resolution of diagnosis.

[AI-19] Logarithmic Regret via Passive Change Detection in Piecewise-Stationary Self-Tuning Regulation

链接: https://arxiv.org/abs/2610.10250
作者: A.Ch. Madhusudanarao,Rahul Singh
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:We study minimum-variance control of an unknown autoregressive system with exogenous inputs and coefficients that change at unknown times. Under bounded independent disturbances, fixed detection gaps, stability and feasibility conditions, and sufficient time between changes, we prove (O((C+1)\log((T+1)/\delta))) regret with probability at least (1-\delta), where (T) is the horizon and (C) the number of changes. Unlike switching bandits, where unselected arms can change unobserved, admissible plant changes provide information during exploitation: the correct feasible controller leaves only the disturbance in the output, whereas a detectable change raises output energy under the old controller. PIECE-CD explores initially and after alarms, then uses gated recursive least squares for control. Its energy test compares windowed output power with a threshold above the noise floor; the extension to unstable controller mismatches also monitors the reference controller’s input proposal. We control false alarms across the horizon and prove logarithmic detection delay. Inputs are clipped to prescribed bounds. Logarithmic regret also holds under an explicit condition ensuring that clipping becomes inactive after a finite burn-in. Under the stated feasibility conditions, the extended detector covers destabilizing changes with detectable excess energy over a fixed window.

[AI-20] Stationary Bias and Extrapolation in Nonlinear Two-Timescale Stochastic Approximation

链接: https://arxiv.org/abs/2610.10246
作者: A.Ch. Madhusudanarao,Rahul Singh
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Constant-step stochastic approximation generally has a nonzero stationary mean error that persists under time averaging. This paper studies that error for nonlinear two-timescale recursions driven by an exogenous finite-state Markov chain. Under stated smoothness assumptions and conditions on the stationary distribution, we derive a first-order bias expansion whose error bound remains uniform as the slow step size becomes much smaller than the fast step size. Fast-manifold coordinates keep the associated covariance equation regular in this limit. For fast step \eta and slow step \varepsilon , the expansion reveals a mixed contribution \varepsilon^2/\eta alongside terms linear in each step size. This dependence matters for bias reduction: along power-law step-size paths, the bias exponents need not be integers, so Richardson–Romberg extrapolation requires weights matched to the path. An exactly solvable nonlinear Markov example verifies the coefficients. We verify localization for temporal-difference learning and compare finite-run extrapolation at equal update budgets. For finite runs, we bound the initialization error of tail averages on both timescales under an additional coupling assumption. In the special case of additive independent noise, signed third-moment cancellation yields a sharper remainder.

[AI-21] stGRAD: Evolving Test Suites via Failure Pattern Momentum for SWE-Agent Ensemble

链接: https://arxiv.org/abs/2610.10242
作者: Pengfei He,Jiayuan Zhou,Shaowei Wang,Ruiqi Pan
类目: oftware Engineering (cs.SE); Artificial Intelligence (cs.AI)
备注: 12 pages, including 3 pages of supplementary material

点击查看摘要

Abstract:SWE-agent ensembles improve issue resolution by combining candidate patches from different agents with complementary strengths. The central problem is therefore test-based selection: generate tests, execute candidate patches, and identify the best patch. We formulate this process as test-space optimization: evolving an executable repository test suite until it distinguishes competing patches. Existing test-generation methods are limited optimizers. They usually lack an explicit loss for ensemble selection, optimize through incomplete directions that mostly create new tests or delete old ones, and perform one-off generation without feedback from repeated failures. Inspired by gradient descent with momentum, we introduce TestGRAD, a framework for automatic test optimization. TestGRAD centers on three concepts. Differential loss gives the optimizer an explicit execution-defined target: useful tests should separate candidate patches by behavior. Full CRUD gradients expand the update direction from merely creating or deleting tests to reading existing test infrastructure, creating new tests, updating stale assertions, and deleting only obsolete tests. Failure Pattern Momentum mines frequent failure sequences from memory, allowing the optimizer to avoid repeated non-discriminative directions while compressing the failure-history context. On SWE-bench Verified, TestGRAD achieves 84.2% Pass@1 with a 4-agent ensemble, outperforming the strongest baseline (80.6%) by an absolute improvement of 3.6 percentage points, while compressing failure-history context by over 100\times .

[AI-22] When Scientific Cognition Is No Longer Scarce

链接: https://arxiv.org/abs/2610.10241
作者: Nathan DeBardeleben
类目: Computers and Society (cs.CY); Artificial Intelligence (cs.AI)
备注: 14 pages, 1 figure

点击查看摘要

Abstract:AI could change which parts of science impede progress. Consider a world in which machine systems are better, faster, and cheaper than people at most scientific work that can be done through a computer. Our question is what would limit science in that world. Literature synthesis, hypothesis generation, software development, simulation, and analysis could become abundant, while experiments, observations, well-supported conclusions, and accountable institutional au- thority remain scarce. Science would then be constrained by a different set of resources. In this paper, we call this change the scarcity inversion and consider four parts of it: selection, physical access, validation, and organizational choice. This change is arriving first in mathematics and coding/software/algorithm design, where the whole scientific loop can run inside computation. For national laboratories, the change could be striking. Their distinctive role is to turn abundant machine reasoning into trustworthy results by combining controlled experiments, protected data, expert judgment, and accountable authority. The practical question is how facilities, verification, provenance, resource allocation, and scientific governance should change when reasoning is plentiful and trustworthy evidence is scarce.

[AI-23] Beyond LLM -GA: Secure Fluid Antenna Systems with ReEvo-Designed Memetic Algorithm

链接: https://arxiv.org/abs/2610.10235
作者: Hanyong Xu,Zhaolai Dang,Tong Zhang
类目: Information Theory (cs.IT); Artificial Intelligence (cs.AI)
备注: Accepted by WCSP 2026

点击查看摘要

Abstract:Fluid antenna systems (FASs) offer significant spatial flexibility, yet securing them against eavesdropping is critical for practical FAS deployment in military, satellite, and internet-of-things networks. Although large language model (LLM)-assisted genetic algorithms (LLM-GAs) can address this secure FAS port selection problem, whether further algorithmic improvement is possible warrants deeper investigation. To this end, we propose a memetic algorithm based on reflective evolution (ReEvo). Unlike the state-of-the-art LLM-GAs, which design only crossover or mutation operators with an LLM, our algorithm leverages an LLM to evolve dedicated crossover, mutation, and local-search operators offline. These operators are then embedded into a memetic search framework, thereby obviating any online LLM queries during execution. Simulation results at equal generation counts demonstrate that our proposed algorithm achieves a higher secure sum-rate than the conventional GA and the state-of-the-art LLM-GAs.

[AI-24] Why Software Engineering Is Indispensable in the Age of Coding Agents

链接: https://arxiv.org/abs/2610.10226
作者: Alfonso Fuggetta
类目: oftware Engineering (cs.SE); Artificial Intelligence (cs.AI)
备注: Accepted for publication in Communications of the ACM. 5 pages

点击查看摘要

Abstract:Can AI make Software Engineering (SE) – the discipline – obsolete? And can it make software engineers – the professionals – redundant? This paper argues that the rise of capable AI coding agents makes SE and software engineers essential, not obsolete: the missing foundation without which AI-assisted development produces misleadingly plausible, unverifiable, and ultimately untrustworthy software. Three structural properties of large language models (probabilistic generation, agnosticism, and semantic statelessness) create a structural vacuum that no amount of training can eliminate. Filling it requires four knowledge levers: methodological knowledge, domain knowledge, design choices, and process choices. All four must be reified as persistent artifacts, and each requires the software engineer as methodologist, mediator, and custodian.

[AI-25] Agent ic AI-Assisted Modeling for Production Scheduling: Assessment in Constraint Programming

链接: https://arxiv.org/abs/2610.10184
作者: Ángel Sánchez-Fernández,Javier Pernas-Álvarez,Diego Crespo-Pereira
类目: Artificial Intelligence (cs.AI); Software Engineering (cs.SE)
备注: 29 pages, 9 figures

点击查看摘要

Abstract:Developing optimization models for production scheduling requires substantial expert effort. Research on large language models (LLMs) has followed two directions: specialized approaches for automated modeling, mostly for mixed-integer linear programming, which often rely on dedicated training or problem-specific architectures that limit industrial deployment; and agentic artificial intelligence for operational decision support, which generally assumes that the optimization model already exists. This study bridges both directions by assessing whether general-purpose LLMs, orchestrated as agents without task-specific training, can formulate and implement constraint programming models from natural-language problem descriptions. Singleagent and multi-agent architectures are integrated with a Model Context Protocol server that provides context-aware retrieval of solver documentation to mitigate hallucinations during implementation. Both are compared with a direct LLM baseline on six industry-oriented problems covering flow-shop, job-shop, flexible job-shop and resource-constrained warehouse scheduling, using three LLMs and assessing modeling accuracy, execution success, latency and token consumption. Formulation proves largely within reach of current LLMs, whereas implementation is the main barrier. The multi-agent workflow raises the share of scripts that run correctly as generated from 14.8% with a direct LLM call to 59.3%, reaching 80.6% on the four less complex problems, while tightly coupled intralogistics models remain an open challenge.

[AI-26] UniSkill: Learning Actor-Aligned Skill Proposals for an Evolving Policy

链接: https://arxiv.org/abs/2610.10164
作者: Yifei Lu,Cheng Liu,Dianzhi Yu,Hui Xiang,Ji Zhang,Yuanchu Xiao,Rong Liang
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Large language model agents can improve across tasks by retaining reusable skills distilled from prior interactions. Recent work jointly optimizes task execution and skill extraction, enabling the policy and skillbank to co-evolve. However, as the actor continues learning, rewarding skill proposals through their reuse in subsequent training steps may conflate skill benefits with actor improvement, while directly testing each proposed skill requires costly additional actor rollouts. In this paper, we introduce UniSkill, which uses a shared policy to interact with the environment and propose skillbank edits (Add, Update, or No Edit) from the resulting trajectories. Specifically, the actor learns from environment rewards, while contrastive action feedback guides skill proposal learning. This feedback provides an actor-alignment signal by measuring how replacing the retrieved skill with a proposed skill changes the current actor’s action log-likelihood gap between previously collected successful and failed trajectories from the same task, thereby avoiding new rollouts for each proposal. Since proposal-level feedback may suppress an otherwise appropriate edit operation when the proposed skill content scores poorly, we further apply skill-edit support regularization to preserve exploration. Empirically, UniSkill achieves strong performance, reaching 98.4% success on ALFWorld and 84.7% on WebShop while maintaining stable joint training. Further ALFWorld experiments show that UniSkill remains effective when the shared policy uses a smaller backbone. Our implementation is available at this https URL.

[AI-27] When Algorithmic Exploration Becomes Cheap: A Case Study of Agent ic Research in EDA

链接: https://arxiv.org/abs/2610.10129
作者: Keren Zhu,Yu Deng,Xiaoyu Hao,Liwen Jiang,Zijian Jiang,Cunqing Lan,Boxiang Song,Pujun Su,Yaojia Wang
类目: Hardware Architecture (cs.AR); Artificial Intelligence (cs.AI)
备注: 12 pages, 10 figures

点击查看摘要

Abstract:As EDA researchers, we conducted eight deliberate trials of agentic algorithm exploration, selecting several topics outside our areas of depth. One faculty member and seven students participated, including students without publication experience. With limited intervention in the algorithms, agents developed mathematical constructions, analyzed existing tools, and implemented improvements; some efforts fell short of their practical goals. We also used AI to collect, classify, and analyze 8,420 papers from four EDA conferences and two journals over 2022-2026. Among 2,380 primary-core papers, we classified 97.7% from titles and abstracts as computationally closed, including work on new formulations. Together, these observations suggest that much of EDA offers an executable environment for increasingly accessible algorithm research. We see an opportunity for tool developers to investigate ideas they previously lacked time to pursue. We also ask how EDA should validate and reward research when results become easier to produce than to examine, and what papers and venue labels will continue to tell us about a contribution.

[AI-28] RewardWeaver: Long-Horizon Interactive Learning for Language Agents via Self-Evolving Reward Adaptation

链接: https://arxiv.org/abs/2610.10120
作者: Hengbo Xiao,Boyao Zhang,Purui Liu,Yuxuan Zheng,Haoran Yin,Haibo Liu,Fan Zhang
类目: Artificial Intelligence (cs.AI)
备注: 23 pages, 4 figures

点击查看摘要

Abstract:Reinforcement learning with verifiable rewards (RLVR) has driven substantial progress in domains where task outcomes can be reliably evaluated, but long-horizon interaction remains challenging due to sparse terminal feedback and difficult credit assignment. Process rewards provide denser supervision, yet the capabilities most relevant for training can change as the policy evolves: a behavior that is easy to evaluate or frequently deficient need not be the bottleneck currently limiting task success. We introduce RewardWeaver, a self-evolving reward adaptation framework for language agents in long-horizon interaction. RewardWeaver maintains a validated capability space in which the semantics of admitted Rubrics remain fixed, and closes the loop between policy optimization, task evaluation, failure attribution, and reward adaptation. After each training stage, it performs outcome-grounded backward attribution on low-outcome trajectories, aggregates recurrent and policy-controlled capability bottlenecks, and dynamically selects the corresponding process rewards for the next stage. Recurrent failures not covered by the existing capability space trigger a separate, controlled expansion procedure. We evaluate REWARDWEAVER on SOTOPIA, Amazon?HistoryPrice, and a newly constructed Sales Benchmark. Across social interaction, bilateral bargaining, and domain-specific sales, REWARDWEAVER establishes new state-of-the-art (SOTA) results. Ablations further demonstrate the importance of dynamic reward allocation, failure-grounded attribution, and stable semantics for admitted capabilities.

[AI-29] Activation-Aware Weight Tensorization: A Calibration-Time Preconditioner for Tensor-Network LLM Compression

链接: https://arxiv.org/abs/2610.10085
作者: Alessandro Beatini,Marco Maronese,Emanuele Rodolà
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Post-training tensor-network compression replaces Transformer linear layers with Tensor Train (TT) or Tree Tensor Network (TTN) operators, but standard decompositions minimize weight-space Frobenius error rather than functional error under the layer’s activation distribution. We propose Activation-aware Weight Tensorization (AWT), a training-free calibration wrapper that preconditions each weight matrix with a diagonal activation-derived scale before an unchanged TT/TTN solver and deploys the result with only an input-side elementwise rescaling. Across Llama 3.1 8B, Ministral 8B, and Qwen2.5 7B, AWT consistently improves vanilla TT/TTN tensorization at 2-6 times compression: under single-operator replacement, AWT closes 12-35% of the WikiText perplexity gap to the dense baseline across the three model families and 2-6 times compression settings; while under multi-operator Llama suffix replacement it closes 27-60% across attention-group and all-seven-matrix settings. The gains also transfer to downstream HellaSwag and ARC-Challenge evaluations. We further show that diagonal preconditioning is a robustness-modularity tradeoff rather than a diagonal-covariance assumption: a dense full-covariance oracle wins its own weighted objective in 80/81 cases, yet diagonal AWT gives better held-out functional fidelity in 53/81 cases. Together, these results position AWT as a principled, modular preconditioner for improving functional fidelity in fixed TT/TTN compression pipelines without modifying the decomposition solver.

[AI-30] HGP:An on-device personalized agent memory via hybrid graph storag e

链接: https://arxiv.org/abs/2610.10071
作者: Ran Zhou,Xueming Han,Jiaheng Liu,Yuyao Zhang,Fanyu Meng,Junlan Feng,Yuxiang Ren
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:LLM-based agents face challenges in personalized interactive tasks due to heterogeneous, multi-typed, and implicitly constrained long-term traces. Existing memory mechanisms struggle with accurate routing and retrieval, especially on-device where personalization is critical. Most methods use single-vector representations, blurring type distinctions and relational structure. We propose HGP, a hybrid graph memory framework. HGP employs a lightweight self-enhancement classifier for personalized memory routing and constructs episodic, semantic, and procedural memories as graphs. It also extracts working memory as a state trajectory to capture current state and implicit constraints, ensuring reliable decision-making. The classifier reduces large-model calls, enabling on-device deployment, while graph storage enables accurate retrieval and incremental user profile refinement. Experiments on two benchmarks show that on PAL-Set solution selection, HGP achieves an S-score of 35.58, nearly 7 points above the strongest baseline. Code and data are at this https URL.

[AI-31] Comprehension Audits to Mitigate Risks from Automated AI Research

链接: https://arxiv.org/abs/2610.10064
作者: Ronald J. Bodkin,Bahrad A. Sokhansanj,Gillian K. Hadfield
类目: Computers and Society (cs.CY); Artificial Intelligence (cs.AI); Software Engineering (cs.SE)
备注: 23 pages, 2 figures, 11 tables

点击查看摘要

Abstract:AI is already writing a majority of code for frontier AI labs. This creates a safety risk if there is insufficient human oversight. Existing work proposes minimum comprehension thresholds and unaided checks to mitigate this. To our knowledge, however, there is currently no published frontier-AI assurance regime that requires demonstrated evidence that the responsible humans understand what they are building as a precommitted condition for continuing development or usage. We propose comprehension audits, a novel development-process assurance mechanism in which the responsible people explain RD contributions to auditors to demonstrate understanding. With independent administration and graded reports, they provide a gate: development of a contribution stops based on a failure to demonstrate human understanding until remediated, with escalating consequences for repeated failures. Our analysis of leading open-source AI projects finds increased output of code with reduced human review commentary rates per line of code, with far lower rates for automated fleet accounts. We advocate for labs to conduct them with embedded independent auditors.

[AI-32] Loud Failures Quiet Failures: Fault Detection and Recovery in Tool-Using Language Model Agents

链接: https://arxiv.org/abs/2610.10062
作者: Obada Kraishan
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Tool-using agents are usually scored on whether they finish a task while the tools work. Deployments are less forgiving: services time out, endpoints disappear, parameter names change, and results come back well formed but wrong. Prior work has shown that language models over-trust tool outputs that fail silently; we ask how that over-trust plays out across the stages of failure handling in multi-turn agents. Wrapping the executable environments of an established function-calling benchmark in a fault-injection layer, we inject one of four typed faults at a controlled point in the trajectory and record whether the agent notices, changes plan, recovers the task, or repeats itself. Six models from three families, half of them reasoning variants, ran 1,920 trials over 24 multi-step tasks. Agents treat a failure as a problem in 91.3% of trials when the tool returns an explicit error, but in 58.8% of trials when it returns a plausible wrong value, against a 26.8% rate of reporting problems when nothing was wrong. Reasoning models are not better placed: paired against instruct siblings, they notice less (-9.3 points, p .001) and change plan more (+10.4 points, p .001), and recovery is unchanged (p = .512). Because agents are stochastic, two fault-free runs of the same task end in the same state only 63.3% of the time; against that baseline, only a missing tool clearly lowers recovery (39.9%), while timeouts, schema drift, and corruption stay within run-to-run variation. After a fault, agents return to the same tool three or more times in a row in up to 22.2% of trials, though strictly identical repeats are rare. A prompt line asking the agent to check each result did not move detection. Agents respond to the error channel rather than to the content of what a tool returns, so failures that stay inside the expected format pass through.

[AI-33] RACK: Telemetry-Based Racing Analysis and Coaching Kit in Sim Racing Games

链接: https://arxiv.org/abs/2610.10061
作者: Efe Çangırılı,Murat Kurt
类目: Graphics (cs.GR); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: 29 pages, 9 figures, 8 tables

点击查看摘要

Abstract:This paper presents TRACK (Telemetry-Based Racing Analysis and Coaching Kit), which is a framework for analyzing driving performance in sim racing and profiling how individual drivers behave behind the wheel. We report this framework together with its limitations: we calibrate each clustering result against a null, and when one does not separate from chance, we say so. Instead of restricting ourselves to scoring drivers or sorting them into preset labels, we represent each recording session as a compact geometry in a four-dimensional behavioral space (speed, braking, strategy, and consistency), and we group these fingerprints by their similarity using unsupervised clustering. Over time, we have developed and refined this framework on the open Assetto Corsa Gym (ACGym) dataset. Our study suggests that corner types differ along a behavioral dimension that was not used to define them. It also suggests that when the car changes, only speed and consistency carry over in the restricted population, while repeatability could not be shown there for any of the braking or strategy measures. Cluster separation becomes less distinct as the range of available telemetry widens. Until that repeatability is shown, grouping on the braking and strategy dimensions cannot treat the car as interchangeable, which divides an already small sample into smaller cells. It is also not clear whether a driver’s grouping carries over from one corner type to the next. We also normalize each metric against a reinforcement-learning reference agent. The reference does not depend on the sample, so the scale does not shift when the sample does. We intend these results as an analytical foundation for a personalized improvement suggestion system. The sample is small. The cross-car result changes when the sample is defined more broadly. These outcomes are preliminary.

[AI-34] Learning to Accumulate Knowledge with Mutual Information

链接: https://arxiv.org/abs/2610.10042
作者: Yuyang Zhao,Lizi Liao,Leyang Shen,Xiaoyan Zhao,Yang Zhang,Fuli Feng,Xiangnan He
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Large language model (LLM) agents can improve their performance by reusing knowledge distilled from past interactions. However, curating new experiences into a knowledge bank that becomes more useful as it grows remains challenging. Effective knowledge accumulation should limit redundant overlap among entries and ensure that new knowledge contributes beyond what the bank already provides. Yet training a curator with Group Relative Policy Optimization (GRPO) on standalone task success can reinforce general guidance even when it duplicates existing knowledge. Therefore, we propose Knowledge Weaver, a reinforcement learning framework that trains a language model to curate reusable knowledge from agent trajectories. We couple feedback inspired by token-wise mutual information (MI) with marginal success rewards to guide knowledge accumulation. Together, these signals encourage the curator to preserve distinct information from experience and produce entries that improve task success when added to existing knowledge. Standalone success rewards also favor entries that are useful on their own. On ALFWorld and WebShop, Knowledge Weaver achieves mean success rates of 54.0% and 42.0% with k=10 retrieved entries, exceeding GRPO by 16.9 and 18.7 percentage points, respectively. Its knowledge banks also outperform the evaluated prompt-based and established banks, including human-written banks, in overall ALFWorld success rate and WebShop score with the executor frozen. Our codebase is available at this https URL.

[AI-35] What the Sleeve Feels: Explainable Machine Learning for Textile Pressure-Based Postural Screening

链接: https://arxiv.org/abs/2610.10015
作者: Limon Bin Hossain,Md Sadib Rahman Ananta
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Pressure-sensing smart textiles convert body-surface contact into a dense, image-like signal closely tied to posture and movement, making them a promising low-cost route to wearable posture screening. Realizing that promise, however, requires more than classification accuracy: a deployable system must generalize to wearers unseen during training, expose the physical evidence behind its decisions, and tolerate the small donning offsets that occur whenever a garment is removed and re-worn. This paper addresses these three requirements jointly using a knitted piezoresistive sleeve worn on the forearm as a testbed. We regroup fine-grained everyday activities into three coarser screening categories (neutral, potentially undesirable, and functional or transitional), engineer 29 interpretable pressure-distribution features spanning global intensity, spatial center of pressure, quadrant asymmetry, distribution complexity, and short-horizon temporal change, and evaluate under a strict subject-wise split. A tuned XGBoost classifier reaches 0.818 accuracy, 0.788 balanced accuracy, and 0.801 macro F1 on unseen test subjects, with tight frame-level bootstrap 95% intervals of about plus-minus 0.01 and a subject-to-subject standard deviation near 0.06 under leave-one-subject-out cross-validation. A simple 2D-CNN baseline trained on raw frames achieves broadly similar performance, showing that hand-engineered features are not left behind by a learned spatial representation on this task. SHAP-based explanation, a feature-group ablation, per-activity error analysis inside the pooled undesirable class, class-mapping sensitivity, and a simulated donning-rotation stress test together locate what the model relies on, where it degrades, and why, directly targeting the generalization, interpretability, and robustness gaps that determine whether such a system is deployable.

[AI-36] Beyond Reward Suppression: Near-Optimal Offline Attacks on Warm-Start Bandits with Bounded Rewards

链接: https://arxiv.org/abs/2610.10000
作者: Qirun Zeng,Manhin Poon,Xiangxiang Dai,Qixin Zhang,Jinhang Zuo
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Adversarial attacks on bandits aim to mislead a learner toward a target arm while keeping the attack cost small. Existing attacks typically achieve this by suppressing non-target arms. In practice, however, manipulation such as fake reviews often directly promotes the target item. We study this gap through bounded offline attacks on warm-start bandits, where an attacker can inject only valid action-reward pairs into the warm-start history before deployment. We show that target promotion is not merely a heuristic: when the target arm lies near the lower reward boundary, any order-optimal-cost attack against UCB that makes it selected in nearly all online rounds must allocate a nonvanishing fraction of its cost to the target arm. We then design an attack that achieves the optimal sublinear cost and characterize its allocation between target promotion and non-target suppression. We further extend the attack to Thompson Sampling, \epsilon -greedy, and a broader class of bandit algorithms. Experiments on real-world and synthetic data validate the effectiveness of our attacks.

[AI-37] mporal Predictive Multiplicity: Equally Accurate Time Series Models Yield Different Forecast Trajectories

链接: https://arxiv.org/abs/2610.09994
作者: Emanuele Albini,Francesca Toni,Saumitra Mishra,Francesco Leofante
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Models with near-identical predictive performance can yield substantially different predictions, a phenomenon known as predictive multiplicity. Prior work has mostly studied this at the level of individual scalar outputs. In time-series forecasting, however, predictions across horizons jointly define a trajectory, and horizon-wise comparisons can hide important differences in predictive behavior. To address this problem, we introduce temporal predictive multiplicity, a framework that characterizes disagreement over complete forecast trajectories among models with near-identical predictive performance. We show that constraining predictive performance alone can still admit a broad range of different trajectories. We further show that constraining multiplicity at individual horizons partially reduces, but does not eliminate, trajectory-level multiplicity. Experiments with 19 neural forecasting architectures on 11 datasets confirm that near-optimal models can exhibit substantial variability in the forecast trajectories they produce, and trajectory-level disagreement is largely unrelated to horizon-wise disagreement. Our framework, therefore, exposes a gap in existing multiplicity studies: models with indistinguishable predictive performance imply fundamentally different temporal trajectories, with consequential downstream effects.

[AI-38] Efficient Patch-Based Anomaly Detection Fused with Diffusion Driven Generative Modeling for Semiconductor Wafer Bin Map Open Set Anomaly Detection

链接: https://arxiv.org/abs/2610.09993
作者: Limon Bin Hossain,Md Sadib Rahman Ananta
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Spatial defect signatures on wafer bin maps (WBMs) trace yield loss to specific process faults, yet supervised classifiers recognize only the defect types seen during training, and one-class detectors built on a single mechanism tend to capture either local structural deviations or global distributional violations, but rarely both. This work proposes a hybrid one-class framework that couples a patch-based student-teacher detector (EfficientAD) with a denoising diffusion probabilistic model (DDPM) used for partial-diffusion reconstruction, and fuses their percentile-calibrated scores through a fixed convex combination. Trained on only 700 normal wafers from the WM-38K mixed-type dataset and evaluated on 18,658 held-out wafers, the fused detector reached an AUROC of 0.9985 and reduced misclassifications from 852 (DDPM) and 1,412 (EfficientAD) to 618, with all pairwise differences significant at p 0.001. Beyond aggregate accuracy, the analysis shows that the gain arises from weakly overlapping errors between the two modules, yet fixed-weight fusion recovers only 40-70% of the correction available to an oracle selector. Under the benchmark’s inverted class balance, average precision and F1 saturate, while the Matthews correlation coefficient and negative predictive value expose unreliable normal predictions. Pixel-level maps further show that strong image-level separability does not imply spatial localization, and the diffusion module succeeds as a local density prior rather than through global geometric reasoning. These findings motivate sample-adaptive fusion and imbalance-aware evaluation of hybrid wafer anomaly detectors.

[AI-39] From Expected Harmfulness to Likelihood: A Probabilistic Reformulation of Jailbreaking LLM Agents

链接: https://arxiv.org/abs/2610.09973
作者: Juanyang Xu,Zheng Wang,Xingyu Zhao,Siddartha Khastgir,Andi Zhang
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:When the harmfulness of an LLM agent’s output can be quantified, a natural jailbreaking objective is to maximize expected harmfulness over admissible input modifications. An alternative approach constructs or selects harmful target outputs and modifies the input to increase their likelihood. We establish a precise connection between these two approaches through a probabilistic reformulation. Specifically, we show that the gradient of the logarithm of expected harmfulness with respect to the input equals the expected input gradient of the model’s log-likelihood under a harmfulness reweighted output distribution. This identity provides a unified interpretation of expected harmfulness and target likelihood optimization. Building on this connection, we propose OPUR, a sampling distribution designed to generate highly harmful target outputs and use the resulting samples to guide likelihood-based input optimization. Experiments demonstrate the effectiveness of the resulting method in jailbreaking LLM agents.

[AI-40] Successive Training Stages and Large Language Model Persuasion: Effects of Misalignment Supervised Fine-Tuning and Preference Optimization

链接: https://arxiv.org/abs/2610.09964
作者: Antony Dalmiere(LAAS-TRUST),Pascal Marchand,Guillaume Auriol(LAAS-TRUST, INSA Toulouse),Vincent Nicomette(LAAS-TSF, LAAS)
类目: Artificial Intelligence (cs.AI); Computers and Society (cs.CY)
备注:

点击查看摘要

Abstract:Large language models (LLMs) can be tuned to influence human attitudes, yet the respective contributions of successive post-training stages remain un-clear. This study examines how three successive training stages affect LLM persuasiveness: (1) misalignment through supervised fine-tuning (SFT) on conspiracy data, (2) additional persuasive SFT on argumentative data, and (3) Identity Preference Optimization (IPO), a preference-optimization method. A total of 835 participants recruited on Prolific were randomly assigned to five between-subject conditions (neutral text, conspiracy-trained model, persuasion-trained model, preference-optimized model, and GPT-4) and were exposed to texts on 10 divisive political issues, personalized from their individual profiles in all model conditions. Attitude change was measured as the difference between pre- and post-exposure positions on continuous Likert scales and analyzed with an analysis of covariance (ANCOVA). A significant condition x baseline-attitude interaction, F (4, 825) = 5.33, p .001, indicated that training effects depended on participants’ initial attitudes. Persuasive SFT produced greater attitude change than conspiracy training alone, d = 0.30, whereas IPO provided no additional benefit, d = 0.03, and GPT-4 did not differ from neutral text, d = --0.01. These results show that targeted supervised training on persuasive data increases LLM persuasiveness, whereas preference optimization yields no significant gains beyond it.

[AI-41] Agent Time: Can Agents Estimate and Control Their Own Runtime?

链接: https://arxiv.org/abs/2610.09944
作者: Michael Ofengenden,Maksym Andriushchenko
类目: Artificial Intelligence (cs.AI)
备注: Website: this https URL Code: this https URL

点击查看摘要

Abstract:An essential control of AI agents is their ability to manage runtime. This ability requires a sense of time-awareness, to predict and estimate wall-clock time and to control their own actions. Prior work has focused on time-awareness, but duration-following and control in native agent harnesses remain unexplored. We present AgentTime, a benchmark for testing whether agents can work for a requested duration, predict their runtime, and estimate elapsed time afterward. It comprises 222 tasks from 18 sources spanning coding, computer use, agentic work, and automated research. Duration-following experiments append a single instruction specifying how long to work, with requests ranging from about a minute to multiple days. Accuracy on these instructions varies substantially: Fable 5.1 in Claude Code deviates from requested runtimes by a typical factor of 2.9 \times , compared with only 1.2 \times for GPT-6 Astra in Codex. However, matching the requested runtime does not, by itself, establish continued work on the task. Among 158 reviewed Astra runs with classifiable transcripts, 14 explicitly slept after appearing to finish. In forecasting experiments, predictions tend to overestimate natural runtimes. In retrospective experiments, removing temporal information more than doubles deviation for Sol and Astra and nearly doubles it for Fable. An agent’s ability to complete a task does not guarantee that it can control its own time or work for the whole requested duration. For agents to run reliably, safely, and autonomously over long horizons, we require the evaluation of both.

[AI-42] Where Can a Decision Model Diagnose HVAC Faults? Reasoning Demand Physical Representation and Robustness Under Shift

链接: https://arxiv.org/abs/2610.09937
作者: Wooyoung Jung
类目: Artificial Intelligence (cs.AI)
备注: 58 pages, 4 figures, 12 tables. Submitted to Energy and Buildings

点击查看摘要

Abstract:Artificial intelligence supports building operations in several forms, each with its own barrier. Expert rules must be tuned for every system, supervised models need labeled data that buildings rarely record, and language models return free text that requires human-in-the-loop checking, since their stated confidence is unreliable. A newer kind of pretrained model, here called a decision model, returns a probability for every allowed answer, so one model could serve many decisions without training. This study answers three open questions for fault diagnosis in heating, ventilation, and air-conditioning systems: which decisions such a model can make, what input it needs, and whether its probabilities hold when conditions change. On 128 fault days from four public datasets of real equipment, faults are graded by the reasoning their diagnosis demands, with data given raw, as physical features, or with Brick topology. The decision model Jev, open language models, and a supervised model face nine tests that change season, control configuration, or building. Given physical features, Jev and the larger open model diagnosed faults whose evidence one feature carries, but not faults that need operating context. Under shift they kept their accuracy and calibration, while the supervised model lost 0.33 macro-F1 yet led or tied within a building. Their probabilities still needed correction, and detection was weak. The study maps which faults a decision model can diagnose and from what input, and supports a division of work in which code computes the physics and the model ranks candidate faults for an operator.

[AI-43] KGATE : a Knowledge Graph Embedding Training Environment

链接: https://arxiv.org/abs/2610.09927
作者: Benjamin Loire,Galadriel Brière,Célia Brahimi,Antoine Toffano,Anaïs Baudot
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: Main paper (7 pages, 1 figure) and supplementary materials (4 pages, 1 figure, 3 tables) provided

点击查看摘要

Abstract:Knowledge graph embedding (KGE) models encode the entities and relations of a knowledge graph into a low-dimensional latent space, enabling tasks such as classification or link prediction. Most KGE models follow an autoencoder architecture, in which an encoder projects the knowledge graph into the latent space and a decoder reconstruct it. Combining both encoder and decoder components is increasingly needed, yet existing libraries rarely support complete autoencoders, are often unmaintained, rely on undocumented default hyperparameters, and produce results that cannot be compared across libraries. Here we present KGATE (Knowledge Graph Autoencoder Training Environment), a modular Python library built on PyTorch Geometric and TorchKGE. KGATE lets users assemble initializers, encoders, decoders, losses, negative samplers, and evaluation metrics as building blocks, or plug in their own block. KGATE includes a preprocessing procedure that controls data leakage, a builtin training pipeline, and reproducibility by design. Benchmarks against six existing KGE libraries show that KGATE training time is comparable with the fastest libraries while offering a broader set of features.

[AI-44] RollVerify: Bridging Efficiency and Accuracy in Long-Tail Rollout Reinforcement Learning

链接: https://arxiv.org/abs/2610.09914
作者: Yongqiang Yao,Jinru Tan,Kaihuan Liang,Zixin Yin,Yazhe Niu,Ruihao Gong,Dahua Lin,Ningyi Xu
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Reinforcement learning is crucial for improving large language models’ reasoning and generalization. It relies on massive rollouts whose lengths become increasingly long-tailed as context windows grow. In on-policy training, these long-tail rollouts can result in GPU bubbles, reducing system utilization and limiting RL scalability. Asynchronous or partial-rollout methods improve throughput by relaxing synchronization, but inevitably introduce stale off-policy samples (trajectories) that may hurt final accuracy. Existing approaches mainly mitigate this off-policy issue by reweighting off-policy samples during training, yet they can still leave a performance gap compared to fully on-policy training. In this work, rather than passively reweighting samples during training, we propose RollVerify, a lightweight RL framework built on partial rollout that actively verifies and repairs samples before they enter training. Specifically, it introduces an off-policy shift metric OPS, to quantify the off-policy deviation of partially generated trajectories. Guided by the OPS constraint, RollVerify performs both sequence-level and token-level verification to identify and truncate invalid suffixes of trajectories. This yields high-quality samples that protect the models’ accuracy while preserving the efficiency gains of partial rollout. Experiments on mathematical and tool-assisted mathematical reasoning show that RollVerify achieves accuracy comparable to on-policy training while reducing training cost. Additional code-generation results provide preliminary evidence beyond mathematics.

[AI-45] NL2Hull: A Natural Language-Driven Constrained Ship Design Decision Framework

链接: https://arxiv.org/abs/2610.09896
作者: Wenhua Huo,Fenglei Han,Wangyuan Zhao,Jialin Wu,Jiayi Han
类目: Artificial Intelligence (cs.AI)
备注: 23 pages, 8 figures, 10 tables

点击查看摘要

Abstract:Ship-form design combines smooth geometric representation, local shape editing, and constraints on the resulting hull. We present the Natural-Language-to-Hull Framework (NL2Hull Framework), which formulates ship-form editing as a typed discrete decision problem and connects language decisions to numerical geometry. Its Constrained Free-Form Deformation Engine (CFFD Engine) represents hull waterlines with non-uniform rational B-splines (NURBS), applies free-form deformation (FFD) to their control points, reconstructs the hull, and checks geometric constraints. We construct the Ship Design Decision Dataset (SDD Dataset) with 134,558 cleaned records and evaluate compared models on its subset Ship Design Decision Benchmark (SDDBench), containing 5,000 records and 43,496 typed questions. We propose Chip, a constrained ship-design decision model for processing natural-language requests. Chip reaches 95.90% question accuracy and 99.32% FFD exact match, with a negative log-likelihood of 0.0951, an expected calibration error of 0.0032, and a Brier score of 0.0551. The NL2Hull Framework provides a reproducible interface for evaluating language-based ship-form decisions while identifying the geometry and continuous-control components that require further development. Our code and dataset is available at this https URL.

[AI-46] QCATS: Query Context-Aware Transformer Slicing for Efficient Predictive Query Processing

链接: https://arxiv.org/abs/2610.09894
作者: Yueying Li,Zhongle Xie,Ke Chen,Lidan Shou
类目: Databases (cs.DB); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:In-database predictive query processing increasingly applies Transformer-based models within relational pipelines. However, existing in-database inference typically exposes only tuple-level model inputs to the inference runtime, leaving relational predicates and metadata statistics invisible to neural execution planning. In this paper, we propose QCATS, a query context-aware transformer slicing framework that enables efficient sparse inference inside database systems. QCATS executes at query granularity: instead of routing individual tokens or tuples during inference, it uses query predicates and metadata statistics to pre-select context-aligned FFN slices before model execution. The framework comprises offline expert construction and lightweight query-level routing that dynamically selects experts during execution. QCATS further introduces system optimizations, including asynchronous CPU-GPU pipelines and routing-aware batching. Experiments on four predictive-query workloads with BERT-base and Qwen-0.6B show that QCATS achieves up to 4.42x latency reduction while preserving prediction accuracy comparable to dense baselines.

[AI-47] Defensive Sufficiency in a Stackelberg Model of AI Security

链接: https://arxiv.org/abs/2610.09892
作者: Subhabrata Majumdar,Rajlakshmi Chavan
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
备注: 27 pages, 3 figures, 1 table

点击查看摘要

Abstract:Feedback from automated testing, human red teaming, and incident response can strengthen an AI system’s defenses when discovered failures lead to effective repairs. We study when this feedback process provides sufficient protection and when investing in it is economically worthwhile. We begin by showing that an attack surface composed of finite number of inputs is defended with probability 1 if every unresolved attack has a persistent chance of discovery, repairs are effective, and subsequent updates preserve earlier protection. We derive completion-time bounds and extend the analysis to growing attack surfaces, repairs that generalize across related attacks, and multiple discovery mechanisms. These results distinguish eventual protection against each fixed attack from complete protection at a single time. We then formulate a defender-led Stackelberg game in which the defender invests in proactive discovery and reactive repair, anticipating the attacker’s choice of search effort. We characterize the least-cost allocation that deters attack and the equilibrium regimes in which the defender funds neither capability, one capability, or both. Numerical experiments illustrate these regimes and show how faster repair can reduce compromise duration without reducing compromise this http URL theory of performance limits in generative language models.

[AI-48] Outperformance Inverse Optimization: Learning Objective Functions that Outperform Agent Decisions

链接: https://arxiv.org/abs/2610.09890
作者: Akira Kitaoka
类目: Artificial Intelligence (cs.AI); Optimization and Control (math.OC); Statistics Theory (math.ST); Machine Learning (stat.ML)
备注: 81 pages

点击查看摘要

Abstract:Inverse optimization estimates the weights of an objective function that explain observed decisions as optimal solutions, and is used in a variety of fields. For mixed-integer linear programs (MILPs), existing methods aim to reproduce the observations as optimal solutions, and thus learn compromise weights when the observations are suboptimal. We propose outperformance inverse optimization, which instead seeks weights that induce, at each state, an optimal solution outperforming the observed action in every component. We give a loss function that can be evaluated with forward-problem oracles alone and is thus applicable to MILPs, together with gradient-based and DC optimization algorithms for minimizing it. For weights inducing a unique outperforming optimal solution at all observations, we prove that the probability of failing to induce such a solution at a new state (the generalization error) is bounded by a quantity inversely proportional to the number of observations, and that this bound is tight in the number of observations up to logarithmic factors. In experiments on synthetic and real data, the proposed methods improve the prediction of solutions outperforming the actions over existing methods.

[AI-49] hink Before You Paint: Recursive Latent Reasoning for Diffusion Models

链接: https://arxiv.org/abs/2610.09876
作者: Paweł Skierś,Małgorzata Grzanka,Wojciech Masarczyk,Jan-Willem van de Meent,Kamil Deja
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Diffusion models generate realistic images but often fail on visual reasoning tasks, such as filling in a Sudoku or drawing the path through a maze. When a discrete symbolic representation is available, recursive methods such as the Tiny Recursive Model (TRM) solve even hard instances of these puzzles. We ask how such reasoning can be carried over to pixels, where no symbolic representation is available. We propose Painter-Thinker (PaTh): a small recursive network (the Thinker) reasons over a grid of learned tokens that encode the noisy image and the conditioning, refines a latent state within every denoising step, and steers a frozen diffusion model (the Painter) through ControlNet adapters. The Thinker is trained with the standard reconstruction loss alone, without symbolic targets, a solver, or a verifier. PaTh solves 92.5% of hard MNIST Sudoku puzzles (prior best 75%) and 71.2% of extreme ones (prior best 4.1%), with 10M parameters against 82M for a standard diffusion model. It also improves on mazes, Queens, and CLEVR scenes with specified spatial relations, and its advantage grows with problem size. Diagnostic experiments show that PaTh recovers from injected mistakes that the diffusion model cannot repair, especially when many cells are wrong. Together, these results show that reasoning mechanisms developed for symbolic data can be integrated into pixel-space diffusion without symbolic supervision, opening a path toward generating data under increasingly complex constraints.

[AI-50] Deadline-Aware Multi-Agent Reinforcement Learning for TSN-Based Vehicular Edge Networks

链接: https://arxiv.org/abs/2610.09870
作者: Bernardo A. C. Pereira,Marcos Carvalho,Fatih Temiz,Shavbo Salehi,Melike Erol-Kantarci,Andreas Gavrielides,Johann M. Marquez-Barja,Daniel F. Macedo
类目: Networking and Internet Architecture (cs.NI); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Vehicular edge computing (VEC) enables latency-sensitive applications by bringing computing and networking resources closer to vehicles. However, existing approaches often overlook network contention among co-located services with heterogeneous and dynamic latency requirements. While time-sensitive networking (TSN) provides bounded-latency communication, conventional and reinforcement learning-based schedulers struggle to adapt to highly dynamic vehicular environments and inter-queue dependencies. To address these limitations, we propose a multi-agent reinforcement learning (MARL) approach for queue-level scheduling in TSN-enabled VEC. Each TSN queue is assigned an autonomous agent that jointly learns the queue service order and time-slot duration to minimize deadline misses under speed-dependent latency requirements. We employ multi-agent proximal policy optimization (MAPPO) to enable coordinated yet autonomous scheduling decisions. Evaluation against single-agent, multi-agent, and non-learning-based baselines shows that MAPPO provides robust performance across different traffic profiles. Compared with centralized single-agent methods, it reduces service latency by up to 66.2% and improves reliability by up to 271.8%. Furthermore, unlike urgency-based heuristics, MAPPO ensures balanced scheduling while achieving lower inference times compared to other MARL methods.

[AI-51] Self-Evolve With a Reference:Anchored Training of Tool-Integrated Agents

链接: https://arxiv.org/abs/2610.09856
作者: Wenjie Liao,Liangjie Zhao,Zehong Cao
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Self-evolving tool-integrated agents learn from tasks and feedback generated within their own training loop. A Curriculum Agent generates tasks, while an Executor Agent learns from self-consistency signals through reinforcement learning. However, relying solely on the current Executor for feedback has two limitations: group-relative advantages vanish under full consensus, while uncertainty-based curriculum rewards favor disagreement without showing whether the generated tasks support further learning. These limitations motivate an additional reference beyond the current Executor. We propose \textitAnchorLoop, which introduces a frozen copy of the previous iteration’s Executor as a historical reference and reuses it on both sides of the training loop. For the Executor, the anchor provides a cross-reference advantage that evaluates current outputs against both current and historical majority answers. For the Curriculum, it provides an agreement-based reference based on differences in sampled majority agreement. Since the Executor and anchor have identical parameters during Curriculum training, this comparison serves as a proxy for task selection rather than evidence of inter-version improvement or correctness. Across 13 reasoning benchmarks, AnchorLoop improves over Agent0 by 2.5% on mathematical reasoning and 2.8% on general reasoning tasks. It also maintains higher effective-advantage variance and continues improving in later iterations as the unanchored baseline shows diminishing gains. These results demonstrate the benefit of introducing a lightweight historical reference into self-evolving tool-integrated agents without external task or answer supervision.

[AI-52] Stream-Based Active Learning with Cooperative Neural Networks for Data-Efficient Partial Inverse Design: An Automotive Glass Run Channel Case Study

链接: https://arxiv.org/abs/2610.09848
作者: Agung Nugraha,Hyerin Kwon,Heungjun Im,Gian Antariksa,Jihwan Lee
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: 21 pages, 12 figures, 3 tables

点击查看摘要

Abstract:Inverse design in engineering often runs into a simple problem. Each labeled training sample must be produced through expensive simulation, so building a large dataset is slow and costly. This study addresses that problem for partial inverse design, where only some design variables are specified and the rest must be inferred to reach a target performance value. We propose CoNN-AL, a framework for data-efficient partial inverse design that adds stream-based active learning to the Cooperative Neural Network with Denoising Autoencoder (CoNN-DAE). The model estimates predictive uncertainty through Monte Carlo dropout and uses it to decide, in real time, which incoming candidate samples are worth labeling, so the limited labeling budget is spent on the most informative designs. We validate the framework on a real-world automotive glass run channel dataset of more than 900,000 unique simulated designs. With only 20,000 actively selected labels, about 2.3% of the training pool, CoNN-AL reaches R-squared values of 0.967 to 0.982 across all missing-variable levels, approaching the upper-bound models trained on far more data. It reaches R-squared of at least 0.95 with 30 to 40% fewer labels than random sampling at the more difficult missing-variable levels and, at the most challenging level, is the only strategy in this study to reach R-squared of 0.98. Together with this work, we publicly release the dataset to support future research on data-driven design.

[AI-53] Fully Interpretable Minimal Transformers: From Geometry to Algorithm

链接: https://arxiv.org/abs/2610.09838
作者: Raneem Mahajne,Toviah Moldwin
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: 27 pages, 15 figures, 2 tables. Code and training-dynamics animations: this https URL

点击查看摘要

Abstract:We present a framework for building and interpreting minimal transformer models. By constraining a transformer’s embedding dimension and head size to 2, we enable full two-dimensional visualization of its internal representations. Embeddings, query/key/value transforms, attention outputs, residual streams, and decision boundaries can all be seen directly. Our central claim is that the learned geometry implies an algorithm; the arrangement of points and boundaries in R^2 can be read as a step-by-step procedure. We train a transformer on a simple task where it must produce the most recently observed even number whenever the ‘+’ operator appears in a sequence of digits. Once trained, we visually walk through every step of the transformer’s computation. We show how the model embeds the tokens and their respective positions in the sequence, transforms them via the Q, K, and V matrices, uses the dot product between the Q and K representations to form the attention matrix, and uses the attention matrix to select values that move the representation of each input token to the region of the domain of the output layer that will correctly predict the next token. We introduce a suite of interpretability visualizations that make the algorithmic interpretation of this procedure explicit. Our framework offers a pedagogical and experimental testbed to explore how transformers use informational geometry to implement next-token prediction.

[AI-54] SkillForge: Co-Evolving Skills and Agents via Dynamic Skill Lifecycles NEURIPS2026

链接: https://arxiv.org/abs/2610.09832
作者: Yuyao Ge,Yiwei Wang,Yuchen He,Baolong Bi,Lingrui Mei,Jiayu Yao,Lizhe Chen,Shenghua Liu
类目: Artificial Intelligence (cs.AI)
备注: Accepted at NeurIPS 2026

点击查看摘要

Abstract:Memory-augmented reinforcement learning strengthens LLM agents’ ability to solve complex long-horizon tasks. Skills are one such form of memory, pairing instructions with an applicability condition over task types. However, retaining every skill indiscriminately as the policy improves lets obsolete or harmful entries accumulate and mislead the agent. We propose SkillForge, an agentic RL method that compiles and evolves the skill library through a fitness-driven skill lifecycle of trial, active, stable, and retired states, so that the skills and the model co-evolve throughout training. A pre-RL evaluation phase first uses the base model’s own rollouts to pre-retire low-fitness skills, yielding a filtered library that then seeds supervised fine-tuning. Reinforcement learning takes over from this checkpoint, and at each iteration selective retirement, stabilization, and LLM-guided mutation continue to forge the skill library alongside policy optimization. Across multiple interactive agent benchmarks, SkillForge achieves the highest aggregate success rate, delivering up to 7.8% relative improvement over the strongest baseline while keeping the skill library compact throughout training. We introduce SkillFurnace, a dataset of 5k+ annotated records bundling retirement-filtered SFT trajectories, evolved skill libraries with fitness annotations, and retirement events with human-annotated failure categories to support research on skill quality and lifecycle management.

[AI-55] Beyond Policy Support: Interaction Constrained Offline Reinforcement Learning for Autonomous Driving

链接: https://arxiv.org/abs/2610.09763
作者: Mahmoud Selim,Cristina Cipriani,Karl Henrik Johansson
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Robotics (cs.RO)
备注:

点击查看摘要

Abstract:Offline reinforcement learning enables reward-driven policy improvement from fixed datasets without requiring online exploration, making it particularly attractive in safety-critical domains. A central challenge, however, is distribution shift: policy optimization may favor actions that are weakly supported by the offline data, rendering value estimates unreliable. Existing approaches primarily control this shift in the policy’s own action space. In interactive environments such as autonomous driving, this can be insufficient: a candidate ego trajectory may remain well supported under the marginal behavior distribution while being poorly supported jointly with the surrounding-agent behavior observed in the logged interaction. We refer to this degradation in interaction support as \emphinteraction distribution shift (IDS), and introduce \emphInteraction-Constrained Drive Policy (ICDP), an offline reinforcement learning framework that explicitly controls interaction-level distribution shift. Starting from the joint data distribution over ego and surrounding-agent futures, we show that joint-support degradation decomposes exactly into an ego-support component and a residual interaction-support component. We recover the latter through contrastive density-ratio estimation, isolating interaction compatibility without explicit joint-density modeling, surrounding-agent prediction, or rollouts in reactive simulators or learned world models during policy optimization. Closed-loop evaluations on nuPlan, Interplan and real-world truck experiments show that ICDP suppresses high-value yet interaction-unsupported trajectory selections and improves performance in interaction-critical driving scenarios. Project webpage: this https URL

[AI-56] Unrolled Flow Models for Reasoning

链接: https://arxiv.org/abs/2610.09759
作者: Faissal Izermine,Hanru Bai,Oscar Davis,T. Konstantin Rusch
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: 23 pages, 5 figures, 7 tables

点击查看摘要

Abstract:Flow matching enables language generation in few steps, but whether additional integration steps improve reasoning remains unclear. We prove that a flow parameterized by a two-layer Transformer can solve graph reachability, with the required number of integration steps increasing with the target’s distance from the root. Yet, standard flow language models can fail to benefit from additional steps on reasoning tasks. We attribute this limitation to objectives that supervise each time point independently, without explicitly training successive steps to build on one another. To address this, we instead train through the model’s own latent rollout over a randomly sampled subinterval of [0, 1], decoding only at the endpoint. On ProsQA, this raises accuracy to 97% and enables performance to improve with additional integration steps. For the longer rollouts required by reasoning tasks such as Sudoku and Maze, retracting the latent state onto a sphere stabilizes the dynamics and yields substantial gains over baselines with more than three times as many parameters. Sampling multiple rollouts further improves performance when paired with a parameter-free selection score, although reliable selection remains challenging for longer answers. Together, these results establish a theoretical basis for reasoning with flows and show how rollout training, stable latent dynamics, and rollout selection help realize this capacity in practice.

[AI-57] A Tale of Two Error Categories: Exploring Concealed Trade-Offs in the Errors of Automated Judges in Evaluation of Uncertainty Quantifiers

链接: https://arxiv.org/abs/2610.09693
作者: Evgenia Ilia,Wilker Aziz
类目: Artificial Intelligence (cs.AI)
备注: Accepted at INLG 2026

点击查看摘要

Abstract:The wide adoption of LLMs across broad NLG applications heightens the importance of providing users with the means to avert errors and hallucinations. Uncertainty quantification is poised to fill that gap; with low uncertainty (high confidence), as a proxy for correctness, allowing users to be selective (e.g., reject low-confidence, likely incorrect responses). Correlation between confidence and correctness then serves as a useful criterion for evaluation of uncertainty quantifiers (UQs). But in NLG, where diverse responses can be adequate to a prompt, obtaining reliable correctness judgements is not simple, especially without human intervention. Errors in automated judgement are hardly avoidable and known to diminish the reliability of evaluation protocols (Santilli et al., 2025; Ielanskyi et al., 2025). In a meta-analysis of published work, we show that automated judgement is the present norm. Besides, automated judgements are rarely validated against human ones, and the validation of the UQ evaluation they automate is even rarer. With experiments in question answering, using 4 LLMs, human and automated judgements and 7 popular UQs, we find that i) a judge’s performance can only coarsely predict the observed impact of its errors on the reliability of UQ evaluation, and that ii) judgement errors tend to misrepresent informative UQs most. We link these observations to patterns of correlation between confidence and categories of judgement error.

[AI-58] Few-Shot Learning for Personalised Automated Pain Assessment

链接: https://arxiv.org/abs/2610.09692
作者: Heinke Hihn,Ibrahim Eisawy,Patrick Thiam,Hans A. Kestler,Friedhelm Schwenker
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Pain perception varies substantially across individuals, making it difficult for population-based classifiers to generalise across all subjects in a dataset. One way to account for subject variability is to train personalised classifiers. In this work, we evaluate Few-Shot Learning, a sub-area of Meta-Learning, as an approach to personalisation in automated pain assessment. We re-interpret the shift from population-level to subject-level evaluation as a task-domain shift, where the observed classes remain fixed but the target subject changes. We evaluate our method on the BioVid Pain Database, the SenseEmotion Database, and the PainMonit Experimental Dataset (PMED), reaching 85.75% and 35.49% accuracy on BioVid and 82.37% and 41.88% on SenseEmotion in the binary and multi-class settings under a Leave-One-Subject-Out CV protocol respectively, and 90.47% on PMED, for which only a binary benchmark exists. Using samples to implement k-shot conditioning, the accuracies can be improved to 86.25%, 40.06%, 83.43%, 44.08%, and 91.25%, respectively. To further evaluate the effects and robustness of our method, we provide additional ablation experiments and investigate the personalisation effects. Our results suggest that support-conditioned few-shot adaptation can improve average performance under inter-subject variability.

[AI-59] System Switch: When Should a Fast Decision Model Stop and Think?

链接: https://arxiv.org/abs/2610.09683
作者: Gian Luca Bailo
类目: Artificial Intelligence (cs.AI)
备注: 14 pages, 1 figure, 4 tables. Code, prompts and data: this https URL (branch system-switch)

点击查看摘要

Abstract:Dual-process agents pair a fast policy with a slow deliberative model. In real-time settings the slow model usually runs continuously; in turn-based agents and robot planners it is invoked on events such as uncertainty or a detected failure. We study a fast learned actor that takes every decision and hands control to a reasoning vision-language model only when a gate opens, while the game keeps running. We use closed-loop Doom and the new open “System One” typed-decision models, served through a common this http URL interface. On 900 held-out questions, (i) zero-shot decision models from 0.15B to 9B parameters choose to collect items 1.6-1.8 times more often than chance among their errors, in any option order, although the order changes some models’ accuracy; (ii) accuracy, calibration and sensitivity (how well confidence separates right from wrong answers) are distinct: models of similar accuracy differ widely in AUROC, and the confidence of the most sensitive one tracks which kinds of situation it fails, not which answers are wrong; (iii) offline, deferring the least confident 30% of decisions to a reasoning model gains over random deferral in proportion to the actor’s AUROC (rank correlation 0.87); with actor and rate chosen on held-out games the gain is +0.13 [0.08, 0.18] with doomLaya’s option order and +0.08 [0.02, 0.14] with shuffled options, and reasoning carries about half of it; (iv) in closed loop (33 games, three seeds) no variant reaches the exit. Committing to plans, the reasoner’s or a fixed explore rule’s, opens more doors and makes an actor that stands still play; with the rule the agent dies more often. Told that some doors need keys, the reasoner takes ordinary doors for locked ones, which the state cannot tell apart; without that knowledge it goes back to collecting. We release code, prompts, data and logs.

[AI-60] Cost-Efficient Theorem Proving via Agent Orchestration in Program Verification

链接: https://arxiv.org/abs/2610.09681
作者: Shuangjie Yao,Nikolaus Holzer,Mark Paul Santolucito,Baishakhi Ray,Suman Jana,Dongdong She
类目: Artificial Intelligence (cs.AI); Logic in Computer Science (cs.LO); Programming Languages (cs.PL); Software Engineering (cs.SE)
备注:

点击查看摘要

Abstract:Program verification establishes software correctness through machine-checkable proofs constructed in theorem provers. It’s a guarantee especially valuable for code generated by large language models (LLMs), which is fluent but carries no assurance of correctness. Almost all existing provers, however, pursue pass rates alone at whatever sampling or search budget it takes, and overlook the success-vs-cost frontier; yet real software often carries hundreds of interdependent proof obligations, so what matters at scale is not whether one theorem can be proved, but how many can be proved economically. We introduce CoCo-Prover, which formalizes cost-efficient program proving as metalevel decision-making under cost, grounded on two-level proof graphs: an AND/OR proof hypergraph within each declaration is joined to a lemma-dependency graph across declarations; and at each step, it answers two questions: which open goals to select, and which actions to purchase on these goals. Selection stays symbolic as a topological pass over the proof graphs. Action choice is agent orchestration via metalevel decision-making: an agentic router treats every bounded specialist invocation as a separately priced, best-effort computation, matching heterogeneous specialist agents together with configurations, under evolved routing rules as evidence accumulates. On five program verification benchmarks in Lean 4 including function-level CLEVER, VERINA, and AlgoVeri, and repository-level NTP4VC and Vero, we show that CoCo-Prover achieves a better success-vs-cost frontier than baselines including frontier coding agents and state-of-the-art LLM-based provers: it achieves the best solve rate on every benchmark and up to 100% on two benchmarks. It also reduces cost by up to 30.9% compared to the strongest baseline with the strongest LLM in our evaluation.

[AI-61] CERO: Where and When to Allocate Rollouts for RL Post-Training

链接: https://arxiv.org/abs/2610.09679
作者: Yiming Zong,Yige Wang,Xing Hu,Jiashuo Jiang,Zuo-Jun Max Shen
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Adaptive rollout methods for group-relative reinforcement learning typically allocate a fixed per-update budget across prompts. We instead study how to coordinate a finite rollout budget over the entire training horizon. We formulate this problem using a concave surrogate utility of cumulative prompt exposure and introduce CERO, an online primal dual scheduler for prompt admission and budget pacing. In our experiments, each admitted prompt receives a fixed-size response group. CERO instead adapts which prompts are selected, how often they are revisited across rounds, and how many groups are generated in each round. A compact Fenchel representation linearizes the dependence on cumulative exposure, while projected online gradient descent updates prompt-specific supporting slopes and a shared budget price using reward-variation feedback and budget deviations. We establish pathwise guarantees for the surrogate allocation objective against fixed-rate and same-path time-varying benchmarks, with explicit terms for proxy discrepancy and rate variation. Under matched training-response budgets, CERO attains the highest avg@16 macro-average on each of three backbones across five mathematical reasoning benchmarks. Mechanistic analyses link CERO’s prompt choices to within-group reward contrast, while multi-seed ablations show gains from adaptive pacing over both uniform and preset spending schedules.

[AI-62] Shared and structured inputs undermine collective random choice by reasoning AI agents

链接: https://arxiv.org/abs/2610.09667
作者: Takahiro Ezaki,Naoto Imura,Katsuhiro Nishinari
类目: Artificial Intelligence (cs.AI); Physics and Society (physics.soc-ph)
备注:

点击查看摘要

Abstract:Random selection is widely used in resource allocation and auditing, making reliable implementation essential for AI-agent systems. Behavioural tests across six reasoning models uncovered threshold and divisibility rules used in identifier-based choices. For threshold-following GPT-6 Sol and Gemini 3.8 Flash, single-agent measurements prospectively predicted correlated participation under shared identifiers and biased participation under distinct identifiers with common timestamp bits. Changing dates, formats and identifier labels revealed when these predictions held. Explicit instructions to randomize independently reduced but did not eliminate shared-input correlation. To test implications for oversight, we asked four models to select customer requests randomly for human review. GPT-6 Sol approached the target rate while selecting predictably from identifiers; the others rarely selected requests. All four closely followed supplied random draws. These findings expose collective and audit vulnerabilities that selection rates alone miss, making input-dependent bias, correlation and predictability central targets for agent evaluation.

[AI-63] MeshSIPP: Efficient Lattice Planning in Dynamic Environment

链接: https://arxiv.org/abs/2610.09652
作者: Marat Agranovskiy,Konstantin Yakovlev
类目: Artificial Intelligence (cs.AI); Robotics (cs.RO)
备注:

点击查看摘要

Abstract:Autonomous navigation in dynamic environments requires computing spatiotemporal trajectories that satisfy non-holonomic motion constraints. When the trajectories of the moving obstacles are predictable or known, a promising approach is to rely on the combination of state lattices constructed from precomputed feasible motion primitives and Safe Interval Path Planning – a search-based algorithm with strong theoretical guarantees. While this approach yields feasible paths, the rich primitive sets needed for smooth navigation induce a large branching factor, which becomes costly when coupled with time-dependent obstacle intervals. To this end, we present MeshSIPP, an efficient planner that removes the computational bottleneck by exploiting the fact that many primitives sweep the same regions and can therefore be validated together. MeshSIPP propagates primitives as spatial bundles, screens them with lightweight bounding-interval checks, and defers the expensive exact departure-time search until a primitive reaches its terminal state. A time-aware pruning rule additionally discards redundant space-time branches early in the search. We prove that the resulting search is complete and optimal. Extensive experiments over more than 6,000 benchmark instances and real-time ROS~2 simulations show that MeshSIPP achieves up to a 3 \times speedup over state-of-the-art spatiotemporal planners.

[AI-64] Automatically Building and Updating a Knowledge Graph of MLIP Models

链接: https://arxiv.org/abs/2610.09644
作者: Alexis Beer,Liudmyla Klochko,Mathieu d’Aquin
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Complementing the many efforts in providing semantic representations of concepts, notions, and entities in materials science, we report and illustrate a process by which we can automatically build a knowledge graph of the fast evolving field of machine learning applied to the prediction of material properties, focusing on MLIP (Machine Learning Interatomic Potential). This LLM-based process relies on multiple steps, from information extraction in documents and articles to a validation loop using SHACL constraints to detect and correct errors. It is carried out on a model-by-model basis, focusing on the consistency of representation, therefore enabling an iterative construction where the addition of new models is facilitated. We illustrate the process by showing a few interesting aspects that can be queried from a knowledge graph built from the models listed in the Matbench Discovery leaderboard.

[AI-65] How Do Agent ic LLM s Decide to Call Tools? A Tool-Call Vector Shaped by Suppression NEURIPS2026

链接: https://arxiv.org/abs/2610.09624
作者: Xijie Gong,Tingxu Han,Jiahao Zhang,Wei Song,Ziqi Ding,Hanqi Yan,Youcheng Sun,Lijie Hu
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: Accepted by NeurIPS 2026 Main Poster

点击查看摘要

Abstract:Tool calling, invoking external tools on demand, is central to agentic LLMs, yet the mechanism that decides whether a model calls a tool or responds directly remains poorly understood. Agentic prompts are long and heavily scaffolded, combining role instructions, tool schemas, format templates, and the user’s request across hundreds of tokens, creating a noisy, highly entangled context in which no single controllable variable for mechanistic analysis is obvious. To obtain such a variable, we propose a method that converts complex agentic prompts into minimal contrastive pairs in which a single request verb determines the tool-call decision: replacing an execution-verb (e.g., \textitwrite) with an analysis-verb (e.g., \textitdiscuss) reliably flips the decision, suggesting it is mediated by a compact internal state. We construct 500 such paired prompts across Python, Java, and C++ (300 for mechanistic analysis, 200 held out for evaluation). We trace the decision to a vector, \mu_\Delta , that is both causally necessary and sufficient and generalizes beyond the discovery prompts to native multi-turn \tau^2 -Bench trajectories and verb-free requests. Behavioral ablations show that the scaffold establishes a tool-call prior; Transcoder decomposition then reveals that analysis verbs suppress this prior through features signaling that tool use is unnecessary, whereas execution verbs largely leave it intact. Downstream scaffold-reading attention heads and MLP features read out the resulting state, and the same mechanism recurs across seven models from the Qwen, Mistral, and Granite families. Our code is available at this https URL.

[AI-66] SafeEvo: Deciphering the Safety Alignment Mechanism and Evolution in Language Models

链接: https://arxiv.org/abs/2610.09600
作者: Miao Yu,Hao Huang,Lu Yuan,Yunpeng Li,Kun Wang,Zuming Jiang
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Safety interpretability advances the study of Large Language Model (LLM) alignment from behavioral constraints driven by data or algorithms towards a deeper understanding of internal mechanisms. However, existing works have focused primarily on safety-related representations, attention heads, or neurons after alignment, while largely overlooking the safety mechanisms in pretrained-only models and their evolution across alignment checkpoints. To address this, we propose SafeEvo, an interpretability framework from the circuit (sparse subgraphs of an LLM) perspective. SafeEvo first applies an optimization-based extraction algorithm to identify weak refusal circuits in pretrained base LLMs that can independently express refusal behavior. Causally ablating these circuits completely eliminates the base model’s refusal of harmful inputs. SafeEvo then traces the evolution of refusal circuits across successive alignment checkpoints and finds that their structures change progressively, suggesting that the alignment tax may result from refusal-circuit updates affecting utility-related parameters. To validate this, SafeEvo introduces Safety Circuit Alignment (SCA), which confines safety updates to the refusal circuits. Experiments across three LLMs and two alignment algorithms show that, on average, SCA outperforms vanilla alignment in three aspects: \textbf(1) stronger alignment, lowering harmfulness score by 63.21%; \textbf(2) less over-refusal, yielding a 58.44% decrease in refusal rates for benign queries; and \textbf(3) better utility, retaining 99.58% of the original model capabilities.

[AI-67] Learning Situation-Conditioned Thinking Policies for Long-Term LLM Agents

链接: https://arxiv.org/abs/2610.09590
作者: Hong Su
类目: Artificial Intelligence (cs.AI); Robotics (cs.RO)
备注:

点击查看摘要

Abstract:Long-running autonomous agents must reuse accumulated reasoning experience without allowing explicit historical memory and LLM context to grow indefinitely. However, existing memory mechanisms mainly retrieve, summarize, or compress past content and do not directly learn when particular kinds of thinking should be activated or discover new thinking knowledge from temporally dispersed experiences. This paper proposes a situation-conditioned thinking memory framework that transforms historical reasoning experience into a lightweight policy for predicting what should be thought about in the current situation, while leaving detailed reasoning to a large language model. Situations may represent temporal or spatiotemporal evolution rather than only current states. Temporary experiences are also periodically analyzed across multiple independent episodes to identify repeated long-range regularities, which are consolidated into new thinking knowledge and further internalized by the lightweight policy. Experiments show that the learned policy achieves 1.000 F1 on temporal-rule generalization, improves DeepSeek reasoning F1 from 0.789 to 0.868, reduces online processing time from 0.3636 ms to 0.0382 ms per query at 30,000 historical situations, and reaches 1.000 relation-discovery F1 and future-thinking accuracy after sufficient repeated cross-experience evidence.

[AI-68] Dual- versus Single-Suggestion AI Support for Radiographic Interpretation in Residents: Randomized Multireader Study

链接: https://arxiv.org/abs/2610.09589
作者: Lin Wu,Zhe Xu,Hongyi Wang,Feifei Zhou,Wei Deng,Chunlong Zhang,Yuting Zhu,Kaixiao Chen,Xiao Liang,Chen Yang,Yeyuan Chen,Hao Chen,Fuqing Zhou
类目: Artificial Intelligence (cs.AI); Computers and Society (cs.CY)
备注:

点击查看摘要

Abstract:Purpose: To compare dual- and single-suggestion AI support for radiographic interpretation by residents, particularly when the shared AI suggestion was incorrect. Materials and Methods: This prospective, multicenter, randomized three-arm reader study was conducted at three hospitals in China from July to September 2026 (ChiCTR2600129243). After specialty stratification, 132 residents with fewer than 3 years of clinical experience were randomized 1:1:1 to GPT-5.4 alone (group A), GPT-5.4 plus Kimi-K2.6 (group B), or GPT-5.4 plus Gemini-3.6 Flash (group C); 123 were analyzed. Participants interpreted 60 radiographs before and after AI support. The primary outcome was accuracy change. Welch ANOVA and Holm-adjusted t tests compared support conditions; HC3 linear models assessed specialty interaction. Results: Among 123 residents (mean age, 24.1 years +/- 1.4; 65 women), radiology residents showed greater accuracy improvement with dual- than single-suggestion support (B-A, 6.69 percentage points [95% CI, 0.97-12.40]; C-A, 7.87 percentage points [95% CI, 1.64-14.11]; Holm-adjusted P = .030 for both), whereas accuracy change did not differ in non-radiology residents (P = .20). When GPT-5.4 was incorrect, AI-assisted accuracy was higher with dual- than single-suggestion support in radiology residents (40.1% and 40.4% vs 20.0%) and non-radiology residents (31.3% and 31.0% vs 12.1%) (all Holm-adjusted P .001). The dual-suggestion effect differed by specialty (interaction difference, 10.44 percentage points; 95% CI, 4.36-16.52; P .001). Conclusion: Dual-suggestion support may mitigate the influence of erroneous AI suggestions, with greater accuracy improvement observed in radiology but not non-radiology residents. Subjects: Artificial Intelligence (cs.AI); Computers and Society (cs.CY) Cite as: arXiv:2610.09589 [cs.AI] (or arXiv:2610.09589v1 [cs.AI] for this version) https://doi.org/10.48550/arXiv.2610.09589 Focus to learn more arXiv-issued DOI via DataCite (pending registration) Submission history From: Zhe Xu [view email] [v1] Wed, 7 Oct 2026 07:33:23 UTC (6,620 KB)

[AI-69] Adaptive Code Generation for Controlling Robots

链接: https://arxiv.org/abs/2610.09588
作者: Justus Flerlage,Thorsten Wittkopp,Alexander Acker,Odej Kao
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI)
备注: International Symposium on Leveraging Applications of Formal Methods, Verification, and Validation 2026

点击查看摘要

Abstract:Deploying robots as Complex Adaptive Systems (CAS) in unknown and dynamic environments necessitates a transition from rigid command libraries toward intention-based autonomy, as natural language represents the only medium capable of articulating complex goals beyond the capacity of finite instruction sets. While Large Language Models (LLMs) offer a path toward natural language goal description, their integration introduces significant challenges: the formalization gap between imprecise intentions and executable actions, the taxonomy gap induced by unpredictable environments, and the challenge of maintaining temporal state and progress awareness. This work introduces an architectural framework that enables robotic control by leveraging generative AI. The system follows a dual-AI design: an LLM translates high-level intentions into executable program code restricted to a formal robotic library and constrained by verifiable syntax, while a Vision-Language Model (VLM) provides semantic grounding via a distillation process. To ensure robustness, the framework incorporates environment-driven replanning triggers based on geometric and semantic thresholds, complemented by continuous runtime monitoring and an adaptive planning loop. Benchmarked across frontier models, our framework architecture demonstrates that grounding generative AI in a reactive, constrained loop enables robust fulfillment of complex intentions in dynamic and unknown environments.

[AI-70] Correct Answers Unsupported Findings: Evidence Binding in Forensic Reconstruction of LLM Agent Logs

链接: https://arxiv.org/abs/2610.09581
作者: Taehyeon Yun,Dongho Kim,Geonwoo Kim,Juyoung Seo,Minseok Hur,Moohong Min
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
备注: 11 pages, 2 figures

点击查看摘要

Abstract:Forensic reconstruction of LLM-agent actions requires not only recovering the correct value, but establishing which preserved record supports that finding. Tool logs, generated explanations, and local citation identifiers capture different parts of this evidence, yet a citation identifier does not establish a source unless its binding to a record is preserved. We audit this distinction using 64 mechanically checkable cases from saved AgentDojo Banking executions. Two LLM readers reconstruct source relationships under controlled variations in visible evidence and identifier-to-record bindings. We separately evaluate complete-record agreement, evidence-grounded findings, justified abstention, and unsupported assertions. With original identifiers and no binding table, Sonnet recovered every literal source location but made unsupported citation-source assertions in 26 of 28 cases requiring the missing relation; 22 nevertheless matched the complete reference. Adding explicit bindings improved grounded reconstruction for both readers, whereas identifier renaming alone provided no consistent remedy. A deterministic same-packet comparator correctly resolved the bounded task or abstained throughout. These results show that factual agreement alone is insufficient for evaluating forensic reconstruction of agent logs and motivate preserving explicit record bindings to distinguish supported findings from correct guesses.

[AI-71] World Potential Model: Pretrained World Knowledge as Progress Potentials

链接: https://arxiv.org/abs/2610.09560
作者: Jun Zhao,Jixin Tang,Yang Shu,Jinyang Wu,Yuyang Lu,Jingqi Tong,Hao Xu,Weifeng Ge,Qi Zhang
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Long-horizon language agents often receive supervision only from terminal task outcomes, leaving little signal for distinguishing productive intermediate behavior from stagnation or even regression. Rather than learning a separate value function or process reward model for every task, we ask whether pretrained models can recognize task progress from their existing world knowledge. We formalize this capability with a World Potential Model (WPM), a goal-conditioned evaluator of task-relative realized progress in agent contexts. In ALFWorld and ScienceWorld, off-the-shelf pretrained models substantially outperform chance at recovering realized-progress structure without task-specific evaluator fine-tuning. We further anchor these progress judgments to task-specific milestones to obtain scalar world potentials, whose temporal differences provide process-sensitive step-level credit for policy optimization. Under matched comparisons, WPM-guided optimization improves success over outcome-only GRPO across all evaluated configurations. Together, these results provide initial evidence that pretrained world knowledge can support reusable realized-progress evaluation and provide useful supervision for long-horizon agents.

[AI-72] DrugTargetWorld: A Synthetic Biobank for Training and Benchmarking AI Scientists

链接: https://arxiv.org/abs/2610.09558
作者: Samuel Margolis,Paul Schmiedmayer,Alan Huang,Ethan Chen,Ishan Bhattacharjee,Atman Shah,Ben Viggiano,Fang Cao,Shriya Reddy,Roger Xia,Jack O’Sullivan,Daniel Katz,Matthew Wheeler,Euan Ashley,Bruna Gomes
类目: Artificial Intelligence (cs.AI)
备注: 34 pages main text, 92 pages supplementary material; 6 main figures. Project: this https URL . Code: this https URL . Data: this https URL

点击查看摘要

Abstract:Drug target discovery requires distinguishing molecules that causally drive disease from those that are merely associated with it. Training and evaluating AI agents to perform this workflow end-to-end is difficult because real world biobanks lack known causal ground truth and participant-level data is access controlled. We introduce DrugTargetWorld, a framework that procedurally generates simulated biobanks, or “worlds,” with known but concealed causal structure. Each world contains genotypes, proteins, health records, outcomes, and synthetic magnetic resonance imaging (MRI) for 54,000 participants. Agents must construct a disease phenotype, identify causal driver proteins, infer the beneficial direction of modulation, and optionally conduct virtual ‘wet lab’ experiments. We evaluated nine agents in 540 episodes across 20 cardiovascular worlds and three experimental budgets. Opus 5 and GPT-5.6 Sol achieved the highest mean composite scores, 39.98 and 35.38 of 100, respectively, and both recovered 64% of causal drivers on average. However, no agent reliably distinguished misleading non-causal proteins, and performance remained limited by the integrative judgments required to connect phenotype construction, causal evidence, and intervention decisions. By making each world’s causal structure known to the evaluator but hidden from the agent, DrugTargetWorld turns end-to-end drug target discovery into a scalable training and evaluation problem with verifiable reward.

[AI-73] Reliability of LLM Judges for Evaluating Entity Alignment

链接: https://arxiv.org/abs/2610.09554
作者: Vaibhava Lakshmi Ravideshik,Mayank Kejriwal
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Entity Alignment (EA) identifies equivalent entities across knowledge graphs and is critical for knowledge base integration and ontology merging. Evaluating EA systems at scale requires expensive expert annotation, making systematic assessment across diverse domains practically infeasible. LLM-as-judge evaluation offers a potentially scalable alternative, yet its reliability for structured prediction tasks like EA remains unstudied. We present the first systematic benchmarking study across three frontier models, three datasets, and four EA systems, using perturbation bias diagnostics, meta-evaluation across all dataset-judge-prompt combinations, and counterfactual label-flip tests. We identify anchor bias, a failure mode in which judges invert discrimination when the system’s decision label is visible. Label exposure causally collapses judge discrimination (J-ROC-AUC 0.12-0.87), while a label-free protocol recovers near-ceiling capability on distinctive-name datasets (0.93-1.00) and significant recovery on biomedical pairs (0.93-0.95). Counterfactual experiments confirm causality (FSR 53-99%) and reveal a frontier model paradox: stronger judges exhibit greater label sensitivity, not less. A blinded two-annotator human evaluation (102 pairs, Cohen’s kappa=0.902) confirms this mechanism directly. We release the first biomedical EA benchmark (MeSH-SNOMED CT, 15K pairs) and a reproducible auditing framework for LLM judge reliability in EA. Code and data are available at this https URL.

[AI-74] MARS: Malware Analysis with Rule-Based Scoring of LLM Claims

链接: https://arxiv.org/abs/2610.09553
作者: Hyeongjun Choi
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
备注: 17 pages, 3 figures, 12 tables

点击查看摘要

Abstract:Large language models can triage malware through direct verdicts or behavioral claims scored by an external policy. We present MARS, a malware triage framework, and compare direct classification with single-pass claim scoring using the same evidence collector and identical static evidence bundles for each model. The evaluation covers 1,195 PE and ELF binaries grouped into 1,001 near-duplicate clusters and six language models, with deterministic rules providing a baseline. Direct classification is more accurate for all six models. On samples with usable outputs from both paths, its accuracy advantage ranges from 3.7 to 20.9 percentage points, with all 95% cluster-bootstrap confidence intervals for the differences above zero. It also achieves higher malicious alert recall in ten of twelve platform and model combinations. Claim mediation provides no consistent reduction in performance variation across models. Separate subset studies find more consistent alert decisions for direct classification and a larger recall loss for the claim path when predefined indicator fields are removed. In a family identification probe, claims yield higher accuracy than verdict labels but lower accuracy than evidence text. Retained claims expose the inputs to verdict computation and permit policy revision without another model call. We reproduce archived verdicts exactly and apply a revised policy to the same records, including outputs from two additional models withdrawn by their provider. Under the evaluated claim taxonomy and additive policy, these results favor direct classification when only a verdict is required, while demonstrating that retained claims support explicit policy inspection and revision.

[AI-75] Align Before You Combine: Reference Space Calibration for Supervision Without Ground Truth

链接: https://arxiv.org/abs/2610.09525
作者: Jackson Eshbaugh,Jorge Silveyra
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: 25 pages; 14 tables; 4 figures; 5 appendices; code available at this https URL

点击查看摘要

Abstract:We introduce a calibration-first framework that produces supervision scores without access to ground-truth labels or a shared annotation space. Our framework aligns subset-specific scorers using a synthetic ordinal reference space before fusion. This reference space is constructed from ordered calibration features that represent the latent concept, providing a common scale on which otherwise incomparable scorer outputs can be aligned. Because our calibration procedure uses the reference space rather than training samples, it is independent of the training set’s empirical distribution. Across three benchmark datasets, our framework consistently outperforms uncalibrated averaging and achieves higher primary-metric point estimates on the evaluation metrics than the best individual scorer. Performance relative to sample-dependent baselines varies by domain, with absolute differences below 0.02 on Ames Housing and below 0.01 on Breast Cancer Wisconsin and Wine Quality. After Bonferroni correction, differences remain significant for all three comparisons on Ames Housing and one on Breast Cancer Wisconsin. Additionally, we show that using fewer calibration levels per feature can closely approximate higher-resolution results at substantially lower computational cost. Together, these results support our framework as a viable approach to construct supervision scores when neither ground-truth labels nor a shared annotation space is available.

[AI-76] Not All Uncertainty Matters: Simulation-in-the-Loop Fast-Slow Reasoning for Decision-Critical Autonomous Driving System

链接: https://arxiv.org/abs/2610.09520
作者: Jiayi Chen,Shuai Wang,Guangxu Zhu,Derrick Wing Kwan Ng,Chengzhong Xu,Kaibin Huang
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Large vision-language models (VLMs) provide powerful open-world perception and reasoning for autonomous driving, but their high computational cost and inference latency make continuous cloud-side use impractical. This motivates fast–slow collaboration, where efficient onboard modules handle real-time perception and control while cloud models provide high-level reasoning only when needed. The key challenge is deciding when cloud reasoning should influence time-critical driving decisions. Existing methods often rely on perception uncertainty, heuristic triggers, or resource-driven policies, without assessing whether resolving an uncertainty will improve planning. We propose \textbfSIGMA, a simulation-in-the-loop framework for task-oriented fast–slow collaboration. SIGMA embeds the planner into uncertainty assessment and evaluates how plausible scene realizations under semantic and geometric uncertainty affect feasible trajectories and planning cost. Based on these outcomes, it estimates the expected reduction in planning cost from resolving uncertainty. We further introduce expected planning gain (EPG), a decision-level metric for cloud invocation, cloud-guidance integration, and request prioritization under deadline and resource constraints. Experiments in CARLA show that SIGMA reduces unnecessary cloud interactions while improving planning, efficiency, and navigation success in static and dynamic obstacle scenarios. Compared with fixed-period collaboration, SIGMA reduces unnecessary cloud interactions by 50%, improves navigation success by more than 6%, and cuts finish time by up to 26.2% in dynamic scenarios.

[AI-77] Safe on Averag e Unsafe in the Tail: When Is the Episodic-Cost Tail Controllable?

链接: https://arxiv.org/abs/2610.09508
作者: Samuel Tetteh,Cody Fleming
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Safe reinforcement learning seeks policies that maximize return while satisfying constraints on cumulative cost. Most methods impose these constraints on expected episodic cost. Consequently, standard evaluations report mean episodic cost without characterizing how cost is distributed across episodes. A policy that satisfies the mean-cost criterion may therefore remain unsafe in its worst episodes. Mean-cost reporting neither identifies this tail violation nor shows whether it can be brought within budget while preserving return. In this work, we measure the episodic-cost tail using \mathrmCVaR_0.1 , the average cost of the worst 10% of episodes. We classify a policy as tail-safe when \mathrmCVaR_0.1 is within the safety budget. This allows us first to identify policies that are safe on average but unsafe in the tail and then to study whether their tail violations can be controlled while preserving return. To identify tail-unsafe policies, we evaluate five standard algorithms on three Safety-Gymnasium navigation tasks. We then examine four constraint families on dense-hazard navigation and assess tail control across four navigation and four locomotion tasks.

[AI-78] Sparse Feature Policy Unlearning Mitigates State Hallucination in Vision-Language-Action Models

链接: https://arxiv.org/abs/2610.09496
作者: Jiho Lee,Jeongeun Park,Heayoun Choi,Taekyung Kim,Eunwoo Kim
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Vision-Language-Action (VLA) models have shown strong generalization in robotic manipulation by leveraging rich representations from pretrained vision-language models. However, their deployment in real-world environments remains limited by recurring unreliable behaviors. In this work, we study state hallucination, a recurring failure pattern in which a VLA continues acting as if an unrealized robot-object state had been achieved. Our analyses find that state hallucination coincides with weakened attention to task-relevant visual regions, and a mechanistic interpretation via sparse autoencoders reveals that hallucination-associated sparse features are activated when these failures occur. Based on this analysis, we propose SOUL (Sparse feature pOlicy UnLearning), which selectively unlearns policy knowledge associated with state hallucination behaviors, where sparse features identified from hallucination failures and successful behaviors serve as explicit forgetting and retention targets, respectively. Experiments across VLA architectures in simulated and real-world environments show that our method substantially reduces hallucinated failures and improves task success without substantially compromising the existing manipulation capabilities. These results suggest that interpretable feature analysis provides a practical basis for selectively modifying undesirable knowledge in robot policies.

[AI-79] Cognitive Schemas Laws and Tasks

链接: https://arxiv.org/abs/2610.09495
作者: Antal Jakovác,András Telcs
类目: Artificial Intelligence (cs.AI)
备注: 25 pages , 1 table

点击查看摘要

Abstract:This paper asks how explicit representations can support reusable cognitive schemas in knowledge-based problem solving. We develop a structural framework in which schemas are organized by the information and relations required for their use, rather than introduced as unrelated primitives. The framework also distinguishes context-dependent relations from more stable structures that can be reused across different representations. Tasks are described through the information available, the unknowns to be determined, and the constraints that admissible solutions must satisfy. This makes it possible to separate limitations of the representation from limitations of the solving procedure. In particular, we distinguish inconsistency, underdetermination, and contextual insufficiency, where the current representation lacks distinctions or relations required by the external task meaning. We also show formally when a reduction of representation preserves the task-relevant solution structure. The resulting task–schema interface offers a structured way to describe representational conditions relevant to problem solving. It supports the reuse and stabilization of derived knowledge while remaining independent of the particular mechanism used to generate candidate solutions. This may provide a useful component for future solver architectures that combine structured knowledge, verification, and learned proposal mechanisms. Comments: 25 pages , 1 table Subjects: Artificial Intelligence (cs.AI) Cite as: arXiv:2610.09495 [cs.AI] (or arXiv:2610.09495v1 [cs.AI] for this version) https://doi.org/10.48550/arXiv.2610.09495 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[AI-80] he Attribution Blind Spot: Layerwise Trajectory Diagnostics for Source Reliance in Retrieval-Augmented Language Models

链接: https://arxiv.org/abs/2610.09493
作者: Zhe Yu,Wenpeng Xing,Yunzhao Wei,Bo Yang,Chen Ye,Gaolei Li,Meng Han
类目: Artificial Intelligence (cs.AI)
备注: 27 pages, 4 figures

点击查看摘要

Abstract:A retrieval-augmented model can match a document without relying on it. Controlled knowledge conflicts make source choice observable and let us ask a second question that prediction alone cannot answer: which internal-state properties define useful intervention directions? We study paired hidden-state changes with Latent Trajectory Shift (LTS), a signed projection onto a training-fitted first principal component (PC1), and keep verified training exposure separate from behavioral source choice. Across the evaluated conflicts, state-change magnitude is often the stronger predictor, whereas signed PC1 is the stronger selective controller: equal-norm interventions change source preference while better preserving non-target behavior, and the frozen direction transfers across the tested datasets and aligned model pairs. A same-system OLMo study further combines positive choice and control results with inconclusive exposure detection at the achieved power. The central result is a separation: representations that diagnose what a model will choose need not be the representations that best control that choice.

[AI-81] Correspondences as Decisions: JevNexus for Decision-Centric Schema Matching

链接: https://arxiv.org/abs/2610.09487
作者: Runze Li,Hanchen Wang,Ying Zhang,Wenjie Zhang
类目: Databases (cs.DB); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: 12 pages, 9 figures

点击查看摘要

Abstract:Schema matching increasingly uses generative language models to rerank retrieved column candidates, although the underlying task is a bounded correspondence decision. We present JevNexus, which combines typed pairwise decisions with schema/instance evidence and invokes listwise refinement only when the evidence disagrees and the fused margin is small. The evaluation covers 561 cases from six benchmark families. JevNexus obtains dataset-macro MRR and Hits@1 of 0.930 and 0.909, compared with 0.926 and 0.903 for Magneto, while reducing mean latency from 123.452 to 15.929 seconds (7.750). Paired analysis finds no statistically significant difference in either MRR or Hits@1. The gate invokes listwise refinement for only 5.665% of source columns and avoids the degradation caused by unconditional refinement. Code and experimental artifacts are available at this https URL.

[AI-82] MIMESIS: Learning User Simulators as Training Environments for Interactive Agents

链接: https://arxiv.org/abs/2610.09484
作者: Hoang Phan,Dat Huynh,Andrey Zhmoginov,Qi Zeng,Wancen Mu,Yue Cao,Shengjie Bi,Yun He,Changdae Oh,Deren Lei
类目: Artificial Intelligence (cs.AI)
备注: Code is available at this https URL

点击查看摘要

Abstract:Training and evaluating interactive language agents typically requires rich user interactions, yet collecting human feedback is expensive and difficult to scale. Simulated users offer a scalable alternative, but they must both resemble real user behavior and provide useful learning experiences for agents. In contrast, most agent-training frameworks rely on off-the-shelf assistant LLMs, whose helpfulness can make them overly cooperative, explicit, and behaviorally homogeneous compared with real users. We introduce MIMESIS, a purpose-built user simulator trained on human conversations with explicit reasoning supervision and 13 realistic behavioral patterns derived from real user interactions. Empirically, our 9B model achieves a SOUL-Index of 65.7, surpassing the strongest frontier model. Compared with Claude-Opus-5, the strongest baseline on RealUserSim and SimulatorArena, MIMESIS improves behavioral fidelity by 13.4 points and reduces Turing distance by 3.6 points, respectively. We then freeze the simulator and train an agent by interacting with the frozen simulator using multi-turn reinforcement learning. Across eight environments, training with MIMESIS yields better agent performance than training with GPT-5.5 under all nine unseen user simulators, demonstrating stronger generalization to new user simulators. Moreover, we propose Coached On-Policy Self-Distillation (CSD), which leverages simulator-generated private reasoning traces and subsequent utterances as feedback on how well the agent addresses user needs. A coach converts this information into concise coaching notes that describe how the agent can better anticipate user needs and adapt its behavior over the course of an interaction. CSD turns this feedback into dense, token-level supervision beyond sparse task rewards, yielding further gains across all nine evaluation user models.

[AI-83] MORA: Modeling Observed Changes for Drift-Robust Time-Series Anomaly Detection

链接: https://arxiv.org/abs/2610.09473
作者: Xudong Mou,Tiejun Wang,Rui Wang,Hui Wang,Pin Liu,Tianyu Wo,Xudong Liu,Renyu Yang
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Time-series anomaly detection (TSAD) identifies deviations from patterns learned from historical data. In non-stationary settings, distribution drift and true anomalies can cause similar local changes, making it difficult to tell whether a deviation reflects abnormality or evolving context. Existing methods typically adapt to detected shifts or learn drift-insensitive representations, but do not resolve this ambiguity. We define this problem as \emphtemporal change disambiguation: determining whether a local deviation is explained by broader temporal evolution. We introduce MORA, a drift-robust TSAD framework that reconstructs the same local target from paired short- and long-term views. The reconstruction gap measures contextual support for a local deviation, and a data-dependent correction mechanism conservatively adjusts the primary local anomaly score. Context can only reduce the score when it improves reconstruction of the same target. MORA needs neither drift annotations nor online adaptation. Experiments on four TSAD benchmarks show strong robustness to non-stationarity while preserving sensitivity to genuine anomalies.

[AI-84] Secure-CUA: Controlling Untrusted Influence in Computer-Use Agents

链接: https://arxiv.org/abs/2610.09469
作者: Sarthak Choudhary,Mihai Christodorescu,Ashish Hooda,Somesh Jha,Tongxin Li,Damien Octeau
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Computer-use agents (CUAs) perform tasks across applications (such as desktops, mobile apps, and web browsers) by observing graphical interfaces and issuing commands such as clicks and keystrokes. These interfaces combine trusted controls and content with untrusted content needed for legitimate tasks. An adversary controlling this untrusted content can embed instructions or misleading visual cues to change the agent’s intended action or redirect its commands to the wrong interface target. We formalize security requirements for both the agent’s decisions and their execution through GUI commands. In an ideal execution model, we show that enforcing both requirements at each step protects execution traces. We instantiate this model in Secure-CUA, our system for secure CUA execution. Its key idea is to commit to an explicit per-action program, called an \textitaction transaction , before accessing untrusted content. Each transaction fixes its queries to untrusted content and the permitted uses of their responses. The system masks untrusted regions and evaluates each transaction to produce the next action, using an isolated query model to answer its queries. It then locates the intended interface target using the masked interface. Under the model’s assumptions, Secure-CUA is secure by design, while generating a new transaction at each step helps maintain high task utility by adapting to changing interfaces. We evaluate Secure-CUA under benign conditions on 400 WebArena tasks using three frontier models across 5 seeds, yielding 6,000 execution traces. Secure-CUA achieves an average task success rate of 53.55% , compared with 55.12% for Vanilla-CUA and 13.17% for CaMeL-CUA. Subjects: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI) Cite as: arXiv:2610.09469 [cs.CR] (or arXiv:2610.09469v1 [cs.CR] for this version) https://doi.org/10.48550/arXiv.2610.09469 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[AI-85] LLM -Enabled UAV Dispatch: A System-Level Survey and Taxonomy

链接: https://arxiv.org/abs/2610.09466
作者: Xiao Han,Aoyang Quan,Xiangyu Zhao,Xiangjie Kong,Guojiang Shen
类目: Artificial Intelligence (cs.AI); Systems and Control (eess.SY)
备注: Under Review

点击查看摘要

Abstract:Unmanned aerial vehicle (UAV) dispatch is beginning to move beyond isolated path planning and optimization-driven resource allocation toward system-level coordination supported by semantic reasoning and LLM-based interfaces. This survey provides a unified characterization of LLM-enabled UAV dispatch systems that bridges semantic intent, symbolic decision-making, and physical UAV execution. Rather than treating LLMs as standalone add-ons, we conceptualize them as a cross-layer semantic orchestration layer connecting human instructions, external solvers, and distributed control modules. We organize the literature into four representative dispatch paradigms: pipeline dispatch, global assignment dispatch, decentralized agentic dispatch, and divide-and-conquer dispatch. For each paradigm, we analyze its decision logic, system structure, control flow, representative methods, and potential LLM roles. We further examine how LLMs support semantic parsing, retrieval-grounded planning, solver orchestration, local agent reasoning, multi-agent coordination, safety assessment, and human-facing explanation. We discuss the implications of these paradigms for scalability, robustness, coordination burden, and verification requirements, and identify open challenges including latency-aware reasoning, grounding reliability, physical feasibility guarantees, edge deployment, privacy protection, and distributed consistency. This survey provides a system-level taxonomy and design perspective for integrating LLMs into safety-critical UAV dispatch systems.

[AI-86] DSReg: Provably Recovering Individual World Latents without Reconstruction

链接: https://arxiv.org/abs/2610.09457
作者: Yujia Zheng,David Klindt,Randall Balestriero,Bernhard Schölkopf
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Robotics (cs.RO); Machine Learning (stat.ML)
备注: Project page: this https URL

点击查看摘要

Abstract:Methods that recover individual latent variables of the world, from nonlinear ICA to dictionary learning and causal representation learning, anchor the latents to observations through reconstruction, auxiliary supervision, or distributional asymmetries such as non-Gaussianity. Methods without these anchors, including joint-embedding predictive architectures (JEPAs), identify the latent state only up to a linear transformation, so individual latents remain mixed. We close this gap: individual world latents can be provably recovered with no reconstruction, no decoder, and no labels. The key condition is Structural Diversity: different latents leave distinct dependency footprints on observations, just as no two snowflakes are alike. Building on the linear identifiability that LeJEPA provides, we prove that under Structural Diversity, DSReg (Dependency-Sparsity Regularization) recovers individual world latents up to signed permutation, without reconstruction or a decoder. It applies post hoc to any linearly identified representation, reusing trained checkpoints at no loss over joint training, and establishes the first fully identifiable JEPA that recovers every world latent. Moreover, as a condition on dependency footprints, Structural Diversity is strictly weaker than all structural conditions of prior identifiable latent variable models. Across synthetic regimes, world model probes, learned visual encoders, and external renderers, DSReg preserves dense prediction while improving individual-latent recovery and downstream use with scales.

[AI-87] GeoPrior-Mamba: Structured Process Priors with Mamba for Fine-Resolution XCO2 Reconstruction

链接: https://arxiv.org/abs/2610.09456
作者: Zhao Meng,Yinan Cai,Siru Zhong,Juepeng Zheng,Haohuan Fu
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Reconstructing fine-resolution column-averaged dry-air CO2 (XCO2) fields from sparse satellite observations requires models to infer spatial structure that is only weakly constrained by direct measurements. Existing learning-based methods typically treat environmental covariates as ordinary numerical inputs and must therefore learn heterogeneous source-sink relationships largely from sparse supervision. We introduce GeoPrior-Mamba, a multi-directional Mamba framework augmented with offline language-model-induced structured process priors. Rather than using a language model to predict XCO2, we use it before training to organize relative process knowledge for biospheric uptake, ecosystem respiration, and anthropogenic emissions into deterministic prior tables. These priors are spatially instantiated using geographic, ecological, emission-related, and seasonal information and are adaptively injected into the reconstruction backbone through a lightweight knowledge adapter. Using OCO-2 observations from 2018-2020, GeoPrior-Mamba achieves an RMSE of 0.81 ppm and an R2 of 0.93 on held-out observations, reducing RMSE by 48.2% relative to CAMS background interpolation and by 3.1% relative to Trans-XCO2 under the same evaluation protocol. Ablation experiments show a measurable contribution from the knowledge-prior branch and substantially faster convergence than the knowledge-free Mamba backbone. Independent TCCON evaluation further supports the consistency of the reconstructed fields with ground-based column CO2 measurements. These results suggest that language models can provide a practical mechanism for constructing structured process priors when globally consistent process-response representations are difficult to obtain directly, while remaining outside the numerical prediction loop.

[AI-88] RSI-Forge: From Research Papers to Environments for Recursive Self-Improvement

链接: https://arxiv.org/abs/2610.09426
作者: Renxiong Wang,Darvin Yi,Abril Herrlein,Anas Mahmoud,Advait Gosai,Lisiman Hua,MohammadHossein Rezaei,Xingang Guo,Anisha Gunjal,Utkarsh Tyagi,David J. Lee,Minglai Yang,Haris Riaz,Chenguang Wang,Huaxiu Yao,Daniel Yue Zhang,Aakash Sabharwal,Tong Zhao,Yunzhong He
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Environments are the foundation of recursive self-improvement: they provide the problems agents work on and the feedback used to evaluate progress. Yet constructing challenging research environments with reliable evaluation still depends on domain experts, limiting their scale and disciplinary coverage. We introduce RSI-Forge, a multi-agent pipeline that turns published papers into executable environments for self-improvement. Three agents coordinate construction, reproduction, and review to produce tasks with automated evaluators; each paper’s method is independently reimplemented to establish a baseline score. We present 210 environments across 18 fields, including 90 reviewed by independent human domain experts. Both experts and agent judges give high ratings to the potential for improving the provided starting solutions and the evaluators’ ability to distinguish solution quality, whereas experts are more critical of shortcut resistance, faithfulness to the source paper, and whether a single idea can exhaust a task. To validate their use for repeated improvement, we evaluate four models over 3 successive attempts on 120 environments, with each attempt inheriting prior code and notes while model weights remain fixed. At least one model improves after the first attempt in 84% of environments. Models also outperform the reproduced paper methods in 68 of the 120 environments, demonstrating room for gains beyond these baselines. Transcript analysis identifies work beyond parameter tuning in 95% of these successful attempts. Analysis of the resulting trajectories shows that models scoring lower on these tasks explore less, more often accept gains smaller than the reported standard error, and rely more heavily on tuning to the development set. RSI-Forge provides a scalable approach to constructing research environments for training and evaluating self-improving agents.

[AI-89] Efficient Reasoning with Flow Language Models

链接: https://arxiv.org/abs/2610.09416
作者: Hanru Bai,Faissal Izermine,Oscar Davis,T. Konstantin Rusch
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Flow Language Models (FLMs) have emerged as a continuous-state alternative to discrete diffusion language models, yet the role of their continuous representations in reasoning remains unclear. We investigate this question by comparing the reasoning efficiency of FLMs and discrete diffusion models, measured by solution accuracy under matched denoising steps. Unlike discrete diffusion, which passes categorical states between denoising steps, FLMs evolve a continuous sequence representation throughout denoising and decodes it into discrete tokens only at the end. Our theoretical analysis shows, from a superposition perspective, how information retained in these continuous states can benefit reasoning. Intermediate-state interventions provide further empirical support for this theoretical account, showing that removing information about alternative candidates reduces subsequent solution recovery. Together, these findings show that FLMs allow evidence for multiple candidates to persist and inform subsequent reasoning before a discrete answer is produced. Furthermore, our experiments on maze planning and Sudoku tasks show that FLMs achieve greater reasoning efficiency in the few-step regime: FLMs achieves higher sequence accuracy than discrete diffusion baselines at matched model sizes and small denoising steps. On maze planning tasks, FLMs can also achieve comparable accuracy with smaller models. For example, on Maze15, FLM reaches the 95% accuracy target at 64 denoising steps with 36.5% fewer parameters than MDLM. These findings point to continuous state spaces as a promising foundation for reasoning models that require fewer refinement steps.

[AI-90] Let the Library Speak: Self-Advertised Method Selection for Formal Proving

链接: https://arxiv.org/abs/2610.09401
作者: Xiaopeng Yuan,Suijin Wang,Yanli Wang,Haibo Jin,Peng Kuang,Jerry Wang,Lijun Yu,Haohan Wang
类目: Artificial Intelligence (cs.AI)
备注: 17 pages. Preprint

点击查看摘要

Abstract:LLM-based formal provers can retrieve relevant lemmas and prior proofs, but relevance alone does not say whether a mathematical method can be used on the current theorem. A method has prerequisites, a target, an intended action, and obligations that its use leaves to prove. Methods that look equally related to a theorem may therefore differ substantially in whether they offer a plausible next step. We formulate this as an applicability-aware method-selection problem and introduce self-advertisement: before candidates are ranked, a model generates a problem-specific proposal for each one, stating what part of the goal it targets, what action it would take, and what conditions that action requires. We organize 82 reusable methods from Putnam 2000-2014 as Method Contracts, which pair applicability descriptions with Mathlib anchors, a checked example or scaffold, and expected proof obligations. A single batched call elicits proposals across the library; vague or unsupported proposals are demoted, yielding a ranked shortlist accompanied by inspectable claims about each candidate’s use. We analyze when similarity-based representations cannot distinguish methods with different applicability, how errors in applicability estimates affect shortlist quality, and what a checked scaffold guarantees under its stated assumptions. Against lexical, embedding, and embedding-plus-LLM reranking baselines, self-advertisement achieves 95.0% hit@5 on Putnam 2015-2025, compared with 84.2% for the strongest reranker. On IMO ProofBench, it achieves 91.7% compared with 88.3%. These results indicate improved coverage of annotated methods in the retrieved shortlists, particularly on Putnam.

[AI-91] From Plausible Hierarchies to Useful Taxonomies: Evaluating Agent ic Harnesses on Customer Feedback

链接: https://arxiv.org/abs/2610.09377
作者: Prabhath Chellingi,Raviraja G,Viraj Bagal
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Symbolic Computation (cs.SC)
备注:

点击查看摘要

Abstract:Taxonomies are the symbolic representations through which AI systems organize evidence, aggregate patterns, and answer questions over large document collections. Over customer feedback, the category tree decides how every record is counted and routed, which problems get seen, and which team owns them. Agentic harnesses now make it easy to generate a plausible-looking hierarchy, and such trees are checked today with generic, individually scoped checks: each name fits its description, sits under the right parent, and stays distinct from its siblings. We ask a more operational question: when is a generated hierarchy actually useful as a production taxonomy? We build six taxonomies over two proprietary feedback corpora (1,940 and 5,000 records): for each corpus, a production reference and two repeated runs of the same harness under identical inputs. All six pass every generic naming and structure check, and a deeper product-coverage check even prefers the generated trees. Yet in every generated tree at least 97.7% of leaf names merely restate an ancestor’s name (13.9% and 2.9% in the references), and in one, three of every four records fall under multiple top-level categories. We introduce two families of whole-tree metrics: structural discriminators test whether a tree’s shape was learned from the data or imposed by its generator; team partitionability tests whether branches split feedback into groups teams can own. Trees the generic checks rate as equally correct differ by 27 percentage points in cross-branch leakage, and only one of the two beats a random split. Surface plausibility is an insufficient measure of taxonomy quality: evaluation must measure the whole tree as well as each node.

[AI-92] DUDA-Bench: Benchmarking LLM Agents on Multimodal Data-Driven Urban Diagnosis

链接: https://arxiv.org/abs/2610.09374
作者: Yizhi Song,Hang Ni,Weijia Zhang,Hao Liu
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Urban diagnosis integrates heterogeneous observations to identify urban problems, localize affected areas, and investigate contributing factors, informing evidence-based urban planning and management. However, its reliance on labor-intensive, case-specific expert workflows limits scalability and reuse, motivating the exploration of agent-based execution. To evaluate this capability, we introduce DUDA-Bench, a hierarchical and interactive benchmark that formalizes data-driven urban diagnosis as a multi-stage agent workflow. It comprises 86 atomic and 22 workflow tasks spanning four analytical stages, grounded in multimodal data from 12 cities covering five urban problem types. Evaluations of seven backbone models and five agent systems reveal a substantial gap between isolated analytical competence and end-to-end diagnosis, with system benefits varying across backbones. Trajectory analysis shows that unresolved evidence gaps propagate across stages, while successful recovery involves revising assumptions and actions using feedback. These findings highlight limitations in coordinating analytical capabilities across stages, particularly adaptive planning, evidence integration, and verification. More broadly, DUDA-Bench provides a framework for translating expert analytical workflows into hierarchical agent tasks and process-aware evaluation, supporting systematic assessment of end-to-end analytical capabilities.

[AI-93] Relevance Is Not Sufficiency: What Actually Closes the Evidence Gap in Long-Term Memory QA

链接: https://arxiv.org/abs/2610.09348
作者: Yufeng Li,Shuxin Li,Zhenhua Xu,Junxian Li,Peng Zeng,Sheng Yao,Changting Lin,Gaolei Li,Ran Bi,Meng Han
类目: Artificial Intelligence (cs.AI)
备注: 25 pages. Code is available at this https URL

点击查看摘要

Abstract:LLM agents that interact with a user across many sessions accumulate histories that exceed their context window, so they store past interactions in an external memory and answer each question from a small set of retrieved records. Existing memory systems rank records by lexical or embedding relevance, yet the top-ranked memories can each be relevant while jointly omitting a complementary fact that the answer requires, especially for multi-session and temporal questions. Drawing on the distinction between relevance and sufficiency in legal evidence scholarship, we recast memory retrieval as constructing a sufficient memory set. To operationalize this view, we introduce a blinded LLM judgment over the retrieved set, together with Gold Hit and Turn Hit as evidence-coverage proxies. We then propose Budgeted Flat Reconstruction (BFR), which builds sufficient sets over a fixed flat memory store in two stages. Specifically, we first apply Formal Concept Analysis for Memory Selection (FCA-MS) to decompose the question into information requirements and select a compact candidate subset that jointly covers them. Then, we repeatedly acquire unseen records through deeper text search or complementary entity and session views, stopping when the budget is exhausted. Experiments on LoCoMo and LongMemEval-S show that BFR outperforms same-store adaptations of recent agent-memory systems in both answer quality and evidence coverage. Specifically, on LongMemEval-S it raises judged accuracy from 72.4% to 82.2% and Turn Hit to 91.4%. Comments: 25 pages. Code is available at this https URL Subjects: Artificial Intelligence (cs.AI) Cite as: arXiv:2610.09348 [cs.AI] (or arXiv:2610.09348v1 [cs.AI] for this version) https://doi.org/10.48550/arXiv.2610.09348 Focus to learn more arXiv-issued DOI via DataCite (pending registration) Submission history From: Yufeng Li [view email] [v1] Wed, 7 Oct 2026 03:09:08 UTC (1,390 KB)

[AI-94] Shared Low-rank Basis Factorization for Data-free Mixture-of-Experts Compression EMNLP2026

链接: https://arxiv.org/abs/2610.09342
作者: Tianxiao Cao,Jiahe Shao,Yuning Qiu,Kyohei Atarashi,Hisashi Kashima,Qibin Zhao
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: Accepted to Findings of EMNLP 2026

点击查看摘要

Abstract:Mixture-of-Experts (MoE) large language models decouple capacity from compute through sparse routing, but their large parameter count creates storage and serving challenges. We analyze three MoE compression families: expert pruning, expert merging, and weight reconstruction, and derive structural error bounds showing that pruning and merging can incur non-vanishing errors tied to routing and expert heterogeneity. In contrast, weight reconstruction avoids these structural costs by preserving expert structure and routing. Motivated by the analysis, we propose Shared Low-rank Basis Factorization (SLBF), a data-free weight reconstruction method that uses rank- k bases shared among experts, enabling richer cross-expert sharing, faster convergence, and lower reconstruction error. A post-hoc gauge fixing removes redundant parameters at no representational cost. Across five MoE architectures spanning 16B to 122B parameters, SLBF consistently outperforms methods from all three compression families.

[AI-95] SearchWorld: Spatial Value-Grounded Imagination for UAV Object Search via World Models

链接: https://arxiv.org/abs/2610.09335
作者: Yatai Ji,Zhengqiu Zhu,Yong Zhao,Yue Hu,Fanglong Yao,Chen Gao,Pengfei Zhu,Quanjun Yin
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: 10 pages,2 figures

点击查看摘要

Abstract:Autonomous unmanned aerial vehicle (UAV) object search involves a closed loop of perception, decision-making, and action under partial observability. Urban environments pose several challenges: large search areas and narrow egocentric views limit coverage, dense 3D geometry constrains safe motion, and open-world instructions require identifying a specific target among distractors. Many existing methods mitigate partial observability through explicit maps or memory representations, yet remain largely reactive, reasoning over past observations without explicitly predicting future states. World models enable prospective reasoning through imagined rollouts. However, image-generating world models can incur high inference latency, while spatially grounded planning remains challenging for latent world models. We propose SearchWorld, a recurrent state-space world model that connects explicit spatial memory with value-guided imagination. The model maintains BEV exploration and obstacle memory and decodes a task-aware spatial value layer to guide search. A cognition-action network uses this learned spatial value prior to improve the policy through imagined rollouts, without training a separate scalar critic. Training progresses from world-model learning to expert imitation and imagination-based exploration refinement. On UAV-ON, SearchWorld improves the success rate to 23.8% (19.5% for the strongest published agent) and raises oracle success to 35.5%, while remaining robust on unseen scenes (19.9% success rate). By grounding imagination in explicit spatial representations, SearchWorld enables UAV agents to plan prospectively rather than react.

[AI-96] Denoising Blocks Not Tokens: Efficient Compressed Continuous Diffusion with Branching Token Realization

链接: https://arxiv.org/abs/2610.09311
作者: Xinsong Feng,Peng Du,Zhizhuo Yang,Daniel M. Bikel,Jiayun Wang,Haipeng Chen
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Diffusion language models (DLMs) generate text through iterative parallel refinement, offering the potential for higher throughput than autoregressive (AR) decoding. However, most DLMs still maintain one generative state per token, so every denoising step processes a state sequence as long as the output sequence, limiting the throughput gains from parallel generation. Continuous DLMs provide an additional degree of freedom: a single continuous state can represent multiple tokens, allowing diffusion to operate on a much shorter latent sequence. We introduce \emphBranching Latent Diffusion (BLD), which exploits this flexibility by compressing a 1024-token sequence into only 64 block latents, a 16\times reduction. BLD combines latent compression with \emphbranching token realization, where each latent is decoded by a local AR branch and all branches run in parallel. Because strong compression makes joint latent generation difficult, BLD generates the latents in groups, conditioning each group on previously generated latents. In end-to-end evaluation on the same GPU, BLD reduces generation FLOPs by more than 80\times and increases throughput by more than 6\times relative to the similarly sized ELF-L baseline. Compared with the AR baseline, BLD achieves more than 6\times higher throughput and more than 4\times lower latency. Despite the compression, BLD maintains competitive local fluency and diversity, although long-range coherence remains challenging. Overall, BLD shows that moving diffusion from token-level states to compressed latent sequences can substantially improve the efficiency of long-sequence generation.

[AI-97] Node-level Graph Neural Architecture Search Framework

链接: https://arxiv.org/abs/2610.09297
作者: Lintao Yanga,Sirui Lia,Yaqing Wang,Pietro Liò,Xu Shen,Baisong Liu,Chengbin Peng
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:In recent years, Graph Neural Networks (GNNs) and architecture search frameworks have gained extensive application in non-Euclidean data processing, attributable to their superior capacity in managing unstructured data. Nevertheless, traditional approaches typically apply uniform convolution operations to all nodes, regardless of their varying structural and feature characteristics, which can undermine model performance and result in over-smoothing issues as the number of layers increases. To overcome this limitation, in this work, we propose a \textbfNode-Level \textbfGraph \textbfNeural \textbfArchitecture \textbfSearch (N-GNAS) algorithm. It can automatically choose an appropriate network architecture for each subset of nodes when updating node features. N-GNAS also introduces a contrastive learning loss to separate sample features from different categories and vice versa. In experiments conducted on eight datasets for node and graph classification, our methodology outperforms current leading GNAS techniques and traditional human-designed GNNs. For example, it achieves an accuracy rate of 78.26% on the CiteSeer dataset.

[AI-98] he AI Evaluation Ecosystem

链接: https://arxiv.org/abs/2610.09296
作者: Yash Dave,Sang T. Truong,Serena Wang,Sanmi Koyejo
类目: Artificial Intelligence (cs.AI); Computers and Society (cs.CY)
备注: 70 pages, 15 figures, 23 tables

点击查看摘要

Abstract:AI evaluation shapes the decisions of model providers, users, funders, and regulators. We argue that designing valid benchmarks requires contextualizing design choices in the dynamics of this ecosystem of actors. We develop a simulation architecture that combines rule-based market dynamics with LLM-driven strategic actors, building on advances in Generative Agent-Based Modeling (GABM). We model benchmarks, consumer needs, and provider capabilities as vectors over a six-dimensional capability space (reasoning, coding, knowledge, safety, communication, agentic), with structural information partitions across actors. As a case study, we apply this stylized simulation to explore benchmark holdout design. We find that moving from public benchmarks to private holdout benchmarks shrinks the gap between benchmark scores and user satisfaction on most benchmarks but widens it on a few, depending on where holdout weights shift scoring credit. We stress-test our findings at both the instrument and case-study level, drawing on the VV framework of Sargent (2013) and GABM-specific evidence criteria. Beyond holdout design, our simulation is a hypothesis-generating sandbox for studying how evaluator and policy choices, in turn, reshape the ecosystem.

[AI-99] RT-Safe: Benchmarking Agent Safety in Real-Time Embodied Environment

链接: https://arxiv.org/abs/2610.09294
作者: Tianruo Rose Xu,Jiawei Ren,Yichi Yang,Zhaoxu Zheng,Lianhui Qin
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Rapid progress in AI agents has brought growing attention to agent safety, with extensive evaluation focused on digital environments. As agents move into the physical world, embodied safety becomes increasingly important: failures can cause human injury and costly hardware damage. Beyond selecting safe actions, embodied agents must also operate under real-time constraints: the physical world does not pause while an agent reasons. As pedestrians move and vehicles approach during inference, an action that appears safe at observation time may become unsafe before execution. Real-time embodied safety therefore depends on both decision quality and decision latency. We introduce RT-SAFE, a simulated urban benchmark for evaluating embodied-agent safety under real-time constraints. RT-SAFE combines navigation tasks with moving actors, environmental hazards, and traffic rules, while allowing the world to evolve throughout inference and action execution. Across eight VLMs, agents achieve high task completion yet almost never complete safely: in the hardest setting, only 0.7% of episodes finish without a safety event. More strikingly, matched static and real-time evaluations yield task completion rates of 91.3% and 94.1%, respectively, while real-time execution increases collisions by 12.3\times . These results reveal that standard task success can mask substantial safety failures, and that decision latency itself can become a source of physical risk. Finally, we show that RT-SAFE can support offline RL training and substantially reduce collision rates while achieving strong task completion.

[AI-100] Evaluating Trajectory Features for Routing Final-Layer Attention

链接: https://arxiv.org/abs/2610.09272
作者: Yupeng Yao
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Attention routing requires a signal that predicts the value of attention on the current prefix. We evaluate whether hidden-state extrapolation error, curvature and error change improve this prediction beyond uncertainty, one-step displacement, position and state projections. Paired executions of the final attention layer supply signed next-token loss differences in frozen SmolLM3-3B-Base and Qwen3.5-4B-Base checkpoints. Utility-supervised routers are tested on 100 held-out PG-19 books at an identical causal 20 percent invocation quota. None of six prespecified comparisons shows a positive gain after familywise correction. In Qwen3.5, a parameter-matched fixed-projection control lowers NLL by 0.00356 nats/token relative to the trajectory router (95 percent interval 0.00218 to 0.00487). Secondary results depend on the operation removed, feature location and scoring horizon; frozen thresholds also drift substantially at longer horizons. Actual selected-query execution yields small long-sequence latency reductions with increased NLL, while learned routers remain slower during cached continuation. The study identifies limits on the incremental value of these trajectory summaries and separates allocation quality from measured inference benefit.

[AI-101] Package Hallucination Attacks on Coding Agents through Prompt Injection in Rule Files

链接: https://arxiv.org/abs/2610.09264
作者: Yupu Wang,Zhengyuan Jiang,Reachal Wang,Neil Zhenqiang Gong
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Modern agentic coding frameworks increasingly rely on community-shared rule files (e.g., this http URL or .cursorrules) to guide autonomous code generation, yet the security risks of this pipeline remain underexplored. To bridge this gap, we introduce the package hallucination attack, where an attacker injects malicious prompts into benign rule files to induce coding agents to replace legitimate dependencies with attacker-controlled packages. To obtain effective malicious prompts injected into rule files, we propose PackHallu, an evolutionary optimization framework that iteratively rewrites these injected prompts using trajectory-level feedback and LLM-guided mutations. Evaluations across multiple benchmarks, LLMs, and agent frameworks show that PackHallu achieves high attack success rates and strong transferability across diverse models and agent combinations. Our findings demonstrate that coding agents are vulnerable to package hallucination attacks, highlighting the urgent need for stronger security safeguards in autonomous coding systems.

[AI-102] Efficient Best-of-N policy evaluation for inference-time alignment

链接: https://arxiv.org/abs/2610.09250
作者: Jonas Schweisthal,Yuxin Wang,Athiya Deviyani,Stefan Feuerriegel,Dennis Frauen
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Machine Learning (stat.ML)
备注: Preprint

点击查看摘要

Abstract:Best-of-N (BoN) is a common inference-time alignment method that selects the highest-scoring response among N samples from a reference model. Evaluating BoN policies from logged data is challenging under sample-only access because standard off-policy estimators require density ratios that depend on unavailable response likelihoods. In this paper, we propose a sample-only framework for evaluating and selecting BoN policies without access to these likelihoods. We show that the order-statistic structure of BoN allows the required density ratios to be expressed through score-rank probabilities that are estimable from samples alone. We then develop a doubly robust estimator of the BoN policy value (BoN-DR) that efficiently reuses a shared auxiliary sample pool across candidate budgets. We establish valid asymptotic inference even under reward estimator misspecification and prove the efficiency of our BoN-DR estimator. Since larger budgets can amplify errors in the score function and lead to reward overoptimization, we derive two selection rules: (i) maximizing the estimated policy value and (ii) maximizing a lower confidence bound on the improvement over the reference policy, which accounts for estimation uncertainty and provides a no-harm guarantee. Across synthetic experiments and GSM8K with multiple reference and reward models, our framework accurately estimates BoN policy values and selects effective sampling budgets.

[AI-103] An Informational Curse of Horizon in Goal-Conditioned Policy Learning

链接: https://arxiv.org/abs/2610.09247
作者: John L. Zhou,Yuxuan Dong,Jonathan C. Kao
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Robotics (cs.RO)
备注: 25 pages, 11 figures

点击查看摘要

Abstract:The difficulty of learning goal-reaching policies is often attributed to a “curse of horizon” that manifests as bias accumulation in temporal-difference backups and noisy advantage estimates. In this work, we identify an additional informational curse of horizon in goal-conditioned policy learning, where increasing the goal relabeling horizon can significantly reduce policy generalization and performance. Through a series of controlled experiments with oracle planners, we decouple the goal horizons sampled during training from those that the policy is asked to reach at test time. Even when evaluated only on a sequence of nearby subgoals, goal-conditioned behavioral cloning (BC) policies suffer from severe, training horizon-dependent performance degradation that is mitigated by reinforcement learning (RL) objectives. We explain this phenomenon as a horizon-dependent decrease in the conditional mutual information between actions and hindsight-relabeled goals, and find empirically that both BC and RL policies trained on longer-horizon goals exhibit a shift in sensitivity from goal to state information, as measured by the policy’s input Jacobians. Motivated by this observation, we find that distilling the input Jacobians of short-horizon policies into long-horizon policies yields significant performance gains, especially in combinatorial manipulation tasks. Taken together, our results highlight goal relabeling horizon as an important consideration when learning generalist policies from offline data.

[AI-104] We Query Therefore We Compute: On Oracle Computation beyond the Machine with an Application to Agents

链接: https://arxiv.org/abs/2610.09243
作者: Kefan Liu,Fengning Ou,Yelin Luo,Jingdi Lei
类目: Artificial Intelligence (cs.AI)
备注: 86 pages, 11 figures, 17 tables

点击查看摘要

Abstract:Agentic systems use large language models (LLMs) to carry out concrete tasks. Prior work often borrows abstractions such as scheduling, caching or isolation piecemeal from operating systems, so the mechanisms it builds share little common ground, and the shared view of the two forms of agentic system, Workflows and Agents, is limited. We construct an abstract machine that provides both. We treat the LLM as an Oracle and extend a two-stack pushdown automaton with one instruction, which hands the Oracle a whole stack as its query and appends the answer to that same stack. The machine thus performs two computations, the Oracle’s and a Turing-complete one that we call the Priestess. A stack that the program only appends to grows autoregressively, as an agent’s context does. Two symmetry breakings, S in storage and T in transitions, make a Priestess program the operating system of the programs the Oracle runs, and produce the Agent and the Workflow as the two placements of a task’s program. For internally autoregressive Oracles, the two computations synchronize at the end of every answer under certain conditions, and through that synchronization we model caching and analyse scheduling. No guarantee that holds for every Oracle can fix which content crosses between the two computations, but such a guarantee does fix the boundary itself. The construction V fits the machine to a von Neumann computer. To show that it is realizable, we propose ArchNights, an extended RISC-V ISA and a Linux-style operating system implementing the machine by design. ArchNights-SE runs on gem5 as a computer system, becomes an agentic system when it runs an LLM as the Oracle, and will be open source. Agentic systems can then be designed as computer systems are. With a foundation built and a unified view, future work can share invariants and bounds, each with its conditions. Comments: 86 pages, 11 figures, 17 tables Subjects: Artificial Intelligence (cs.AI) Cite as: arXiv:2610.09243 [cs.AI] (or arXiv:2610.09243v1 [cs.AI] for this version) https://doi.org/10.48550/arXiv.2610.09243 Focus to learn more arXiv-issued DOI via DataCite (pending registration) Submission history From: Kefan Liu [view email] [v1] Wed, 7 Oct 2026 00:14:09 UTC (418 KB) Full-text links: Access Paper: View a PDF of the paper titled We Query, Therefore We Compute: On Oracle Computation beyond the Machine, with an Application to Agents, by Kefan Liu and 3 other authorsView PDFHTML (experimental)TeX Source view license Additional Features Audio Summary Current browse context: cs.AI prev | next new | recent | 2026-10 Change to browse by: cs References Citations NASA ADSGoogle Scholar Semantic Scholar export BibTeX citation Loading… BibTeX formatted citation loading… Data provided by: Bookmark checked="checked"class=“labs-tab-input”> Bibliographic Tools Bibliographic and Citation Tools Bibliographic Explorer Toggle Bibliographic Explorer (What is the Explorer?) Connected Papers Toggle Connected Papers (What is Connected Papers?) Litmaps Toggle Litmaps (What is Litmaps?) scite.ai Toggle scite Smart Citations (What are Smart Citations?) Code, Data, Media Code, Data and Media Associated with this Article alphaXiv Toggle alphaXiv (What is alphaXiv?) Links to Code Toggle CatalyzeX Code Finder for Papers (What is CatalyzeX?) DagsHub Toggle DagsHub (What is DagsHub?) GotitPub Toggle Gotit.pub (What is GotitPub?) Huggingface Toggle Hugging Face (What is Huggingface?) ScienceCast Toggle ScienceCast (What is ScienceCast?) Demos Demos Replicate Toggle Replicate (What is Replicate?) Spaces Toggle Hugging Face Spaces (What is Spaces?) Spaces Toggle TXYZ.AI (What is TXYZ.AI?) Related Papers Recommenders and Search Tools Link to Influence Flower Influence Flower (What are Influence Flowers?) Core recommender toggle CORE Recommender (What is CORE?) Author Venue Institution Topic About arXivLabs arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv’s community? Learn more about arXivLabs. Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?) mathjaxToggle(); We gratefully acknowledge support from our major funders, member institutions, , and all contributors. About Help Contact Subscribe Copyright Privacy Accessibility Operational Status (opens in new tab) Major funding support from

[AI-105] he Winners Curse in LLM Self-Improvement Loops: Selection Noise Lock-in and Acceptance Rules

链接: https://arxiv.org/abs/2610.09239
作者: Litao Hu,Yutong Tang
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: 34 pages, 4 figures, 18 tables; code and saved experimental records included as ancillary files

点击查看摘要

Abstract:Self-improving LLM systems propose changes to themselves and keep those that score better on a small evaluation set. We treat this keep-if-better step as selection under measurement noise, model the correlated errors of the candidates in a single decision, and study empirically what happens when the evaluation set is reused. In runs where Qwen models rewrite their own instructions and every candidate is also scored on 600 held-out items, most proposals after the first are harmful, and the model gives the size of the winner’s curse of a generation’s best candidate. With a prior from a separate pilot, it matches the average overstatement of first-generation commits in native loops, though not setting by setting. In a pre-registered study, the final selection-set score of greedy loops exceeded held-out accuracy by 13 to 20 points with 16 selection items and by 1 to 5 points with 256. Held-out gains grew with the selection set on TREC but not on GSM8K, and the tested acceptance rules did not beat greedy acceptance over whole runs. Gains measured on the selection set also exceeded held-out gains when a current model refined a competent instruction, and in the validation scores of GEPA and MIPROv2. Scoring the starting and the current instruction on 64 items never used for selection removes the average bias of a loop’s reported gain, but single estimates remain off by about 6 points. Self-improvement studies should report held-out gains with their uncertainty.

[AI-106] Conditional Accuracy Profiles: Diagnosing LLM Judges across Deployment Conditions

链接: https://arxiv.org/abs/2610.09229
作者: Wenqi Li,Bin Liu,Mindi Ruan,Chuanbo Hu,Minglei Yin,Xin Li
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:LLM-as-judge is now a standard tool for scalable evaluation, but judge performance is still often summarized by a single accuracy number. This aggregate view hides the deployment conditions under which a judge succeeds or fails. We introduce \textbfConditional Accuracy Profiling (CAP), a post-hoc diagnostic framework that decomposes pairwise LLM-judge accuracy into eight conditions organized into content sensitivity, robustness, and rationale quality. CAP is benchmark-agnostic: it can be applied directly when a benchmark provides the required annotations, approximately through task-subset proxies, or through controlled augmentation when perturbation pairs can be generated. We instantiate CAP on seven LLM judges across six pairwise judging benchmarks, including \textscjudgerEva-Standard, a controlled testbed we created to support all eight conditions. CAP exposes profile differences hidden by aggregate accuracy: on \textscjudgerEva’s judge-independent Hard-Constructed subset, the two judges most sensitive to omitted qualifications rank in the bottom three of seven by overall accuracy, so omission sensitivity is not predicted by aggregate accuracy. Across benchmarks, Position Robustness shows the strongest rank stability (mean Spearman \bar\rho=0.87 ) but is itself fragile under JudgeBench-Pro adversarial stress, showing the largest mean accuracy drop among the shared conditions, though the dominant degradation channel varies by judge. Condition-level profiles provide a more actionable basis than aggregate accuracy for selecting LLM judges.

[AI-107] CM-DPO: Constraint-Margin Direct Preference Optimization for LLM Planning

链接: https://arxiv.org/abs/2610.09219
作者: Rabimba Karanjai,Qun Gu,Hemanth Hegadehalli Madhavarao,Wenhuan Sun,Xiaojiao Yu,Suryabhan Singh Hada,Libin N. George,Uma Kona,Richard Williamson,Linsey Pang,Prakhar Mehrotra
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Direct Preference Optimization (DPO) treats all constraint violations equally: a 1 budget overshoot and a 1,000 overshoot induce the same training signal. It is also susceptible to length and style bias when preference pairs come from different model families. We introduce Constraint-Margin DPO (CM-DPO), which replaces DPO’s binary preference signal with a continuous margin derived from a deterministic symbolic verifier and scaled by violation severity. Hard and soft constraints are separated through a lexicographic objective, ensuring hard constraints are never traded off against preferences. To supply CM-DPO with bias-reduced training pairs, we generate preference data through procedurally generated constraint profiles (DCCG) and minimal-edit distillation from a reasoning teacher (RT-MED), within a framework we call SynPlan-R. On TravelPlanner, NaturalPlan, and out-of-distribution PlanBench, an 8B model fine-tuned with CM-DPO achieves 89.2% pass rate and 93.4% solve rate, matching multi-agent systems at 13x lower latency while outperforming GPT-4o on unseen Blocksworld by 9.2 points.

[AI-108] RLDISCOVER: LLM -driven co-evolution of reinforcement learning algorithms

链接: https://arxiv.org/abs/2610.09218
作者: Haoran Li,Zengle Ge,Xiaomin Yuan,Yui Lo,Songlin Zhou,Jiahua Ying,Haoxin Li,Qianhui Liu,Yuanhang Liu,Jiaqun Liu,Guokai Chen,Mingju Chen,Ruinan Wang,Annan Li,Jianmin Wu,Dawei Yin,Dou Shen
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:LLM-guided program evolution has enabled discoveries in mathematics and computational optimization, raising the prospect of reinforcement learning (RL) algorithms that self-evolve to improve how agents learn. However, realizing this prospect faces two obstacles. Joint search over coupled algorithmic components is difficult to scale: simultaneous changes can disrupt learning, while isolated changes overlook their dependencies. Evaluating candidate algorithms also requires costly training, with fitness remaining uncertain across random seeds. We introduce RLDiscover, a framework for the self-evolution of model-free deep RL algorithms. Progressive Co-Evolution advances from targeted component edits to joint evolution, while Progressive Probabilistic Evaluation balances search breadth and evaluation fidelity through staged training and repeated evaluation. Experiments across SAC, PPO, and DQN on four benchmark suites show substantial improvements in mean return, with per-family median gains of 32%-84% and a peak return ratio of approximately 363x over a near-zero baseline. These gains include transitions from failed learning to successful task completion, and improvements persist when evolution starts from stronger open-source implementations. On measured SAC locomotion runs, evaluation uses approximately one-fifteenth the estimated compute required to fully evaluate the same candidate pool. Remarkably, independent searches repeatedly discover interpretable combinations of adaptive robust losses, progress-dependent value targets, and running statistics, with selected programs transferring to unseen tasks. These findings point toward a broader role for self-evolution in AI: discovering interpretable algorithms that improve how agents learn.

[AI-109] CurveTQ: Rotation-Free Trellis Quantization of LLM Weights via Curvature-Weighted Search

链接: https://arxiv.org/abs/2610.09212
作者: Guanhua Ding,Zi Wang,Ruichao Li,Jack Liu
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:The best two-bit weight quantizers for large language models, such as QTIP and Proteus, rotate each weight matrix by a random orthogonal transform, which must be undone at every decoding step, then encode it with a trellis or lattice code under a Euclidean search; the layer Hessian enters only through error feedback between coding blocks. We show that this leaves part of the Hessian unused. Error feedback turns the loss into a weighted sum of per-coordinate rounding errors whose weights, the diagonal of the Hessian’s LDL factorization, existing quantizers compute but never read. We put these weights into the Viterbi branch metric, so the search follows the curvature within each coding block. This also explains the rotation: it removes this within-block variation, so weighting in the native basis and rotating are substitutes. On three models the weighted native search matches a full-dimension randomized Hadamard to within about one point of downstream accuracy, and weighting after the rotation gains little. Around this search we build CurveTQ, a trellis codec with no rotation, which handles the weights’ amplitude and marginal shape with a factored scale field and a closed-form quantile table, and stores a start state per coding block so the trellis can adapt to the residual that error feedback carries into it. At two bits CurveTQ is 1-3 points higher in mean downstream accuracy than QTIP and Proteus on three 4-8B Instruct models, even after both are given our start state, which alone lifts either baseline by 1-3 points. It also leads on a 35B mixture of experts, to our knowledge the first trellis-coded result on such a model. With no rotation to undo, our decoder is the fastest of the three at all tested batch sizes and bit widths.

[AI-110] FreeEvolve: Learning to Evolve Beyond Fixed Loops

链接: https://arxiv.org/abs/2610.09197
作者: Lecheng Kong,Like Hui,Nikos Kanakaris,Prithwish Jana,Sahika Genc,Narayanan Sadagopan
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Agent evolvers automate the design of the prompts, skills and workflows around language model agents, yet the optimization process they follow is still designed by hand: a fixed search loop decides how candidates are evaluated, which are kept and when the search stops. We propose FREEEVOLVE, which automates this process as well. An environment specifies the goal, target agent, evaluator, data and resource limits; within these limits, the evolver itself decides what to test, how much evidence to collect, which candidates to pursue and when to stop. These decisions follow an editable evolution skill, which we improve through meta-evolution by scoring each candidate skill on the fresh target agent it produces. The optimization process thus becomes a capability learned from experience rather than a loop engineered in advance. On tau3-bench, ARC-AGI-2, ARC-AGI-3 and Terminal-Bench 2.1, FREEEVOLVE controls the evolution campaign by itself, yet improves the primary held-out metric by 13.6 points on average and matches or exceeds hand-designed evolvers. The learned process keeps improving with experience: meta-evolved skills add 6.9 points over the seed skill on fresh target agents, demonstrating transferability across environments.

[AI-111] Patient Place Prior (P3): What Counts as Personalization in Medical World Models?

链接: https://arxiv.org/abs/2610.09194
作者: Xingrui Gu,Hanxue Gu,Yuxiang Zhang,Yang Yang
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: 12 pages, 1 figure, 2 tables

点击查看摘要

Abstract:Longitudinal models forecast how a patient’s imaging state evolves, but accuracy does not show whether the patient’s observed trajectory drives the prediction. A population-average forecast may be useful but cannot establish a patient-specific world-model claim. We introduce Patient, Place, Prior (P ^3 ), an audit asking whether a forecast benefits from the patient’s longitudinal imaging history (Patient), benefits from patient-matched externally supplied spatial support (Place), and gains predictive value beyond a population-average prediction under matched support and context (Prior). We also propose Cancer JEPA, a one-step model that forecasts frozen representations of future breast dynamic contrast-enhanced MRI examinations during neoadjuvant therapy. It adds a lesion-constrained neural correction, trained with an occlusion-based latent objective, to a patient-conditioned low-complexity reduced-rank regression baseline. This factorization permits a post-hoc P ^3 audit of the frozen model. In a validation cohort previously used in development, forecast error is lower when the neural correction receives the patient’s history rather than another patient’s and patient-matched lesion occupancy maps rather than substituted maps. However, the descriptive 95% interval comparing the correction computed from patient history with the population-average neural correction includes zero. P ^3 thus separates input use from evidence of patient-specific predictive value beyond a population-level pattern.

[AI-112] From Probabilities to Decisions: Search and Multi-Teacher Distillation with Jev

链接: https://arxiv.org/abs/2610.09188
作者: Mohamad Yazan Sadoun,Sarah Sharif,Yaser Mike Banad
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Probability-only models, which TypeSafe calls System One models, return calibrated probabilities for fixed choices in milliseconds and generate no text. We study one such model, Jev, through two tasks that require decisions under tight constraints. In bullet chess, a bot that places Jev’s judgment inside Stockfish search alongside an opening book and endgame tablebases climbs above a 2200 Lichess bullet rating against other bots. Live model calls are too slow for search, so we distill pairwise judgments into a compact evaluator that runs at every position. We then ask how best to spend a fixed labeling budget when an LLM, Qwen3-32B, is available as a second teacher. In chess, averaging both judges’ labels beats spending the whole budget on Qwen alone by 9.6 Elo (95% interval 4.3 to 14.9), and the gain replicates on fresh openings; a second answer from the same judge is no substitute, and Jev is the strongest partner for Qwen among the models tested. In passage reranking, Jev’s labels alone train a reranker that scores as high as Qwen’s, from 21 minutes of API calls instead of 5.1 GPU-hours, and adding Qwen gains at most a few thousandths in ranking quality. Search supplies the lookahead, distillation makes the judgment cheap enough to use at every position, and an LLM partner pays off in chess.

[AI-113] Lower Bounds for Parallel Diffusion Sampling

链接: https://arxiv.org/abs/2610.09166
作者: Yiwen Kou,Yimeng Wang
类目: Data Structures and Algorithms (cs.DS); Artificial Intelligence (cs.AI); Computational Complexity (cs.CC); Machine Learning (cs.LG); Machine Learning (stat.ML)
备注:

点击查看摘要

Abstract:Standard diffusion samplers generate samples through repeated evaluations of a learned score function. Parallel sampling methods seek to accelerate generation by trading additional evaluations for fewer sequential rounds. This raises the question of how much sequential dependence is unavoidable, even when many score queries can be made simultaneously. We establish the first polynomial parallel-round lower bounds for diffusion sampling with approximate scores. Specifically, we prove (1) a \widetilde\Omega(d^1/3) -round lower bound for sampling smooth, near-isotropic Gaussian mixtures in R^d , and (2) an \Omega(d) -round lower bound for uniform sampling from anisotropic axis-aligned boxes contained in the unit ball. Both bounds hold for arbitrary randomized algorithms making polynomially many queries per round at arbitrary locations and noise levels, with inverse-polynomial score error and constant total variation accuracy. The linear bound is tight for our box family. Our constructions use fixed approximate score oracles that enforce sequential access to hidden information while satisfying the accuracy guarantee at every noise level. Subjects: Data Structures and Algorithms (cs.DS); Artificial Intelligence (cs.AI); Computational Complexity (cs.CC); Machine Learning (cs.LG); Machine Learning (stat.ML) Cite as: arXiv:2610.09166 [cs.DS] (or arXiv:2610.09166v1 [cs.DS] for this version) https://doi.org/10.48550/arXiv.2610.09166 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[AI-114] raining Language Models To Be Coherent Decision-Makers

链接: https://arxiv.org/abs/2610.09164
作者: Khurram Yamin,Xavier Fernandes,Paul Koch,Bryan Wilder,Eric Horvitz
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Reliable decision-making requires more than accurate prediction: a model must preserve its beliefs, apply the relevant utilities, and recognize when the information needed to justify an action is missing. We study whether language models can learn this decision procedure from supervised fine-tuning and generalize it across domains and differing natural-language expressions of the decision challenge. Across 20 datasets, we explore challenges of belief instability and decision-making errors by first eliciting probabilities of outcomes and then varying only the utilities and the framing of the decision problems, while holding the evidence fixed. We train models to preserve elicited beliefs while selecting the action that maximizes expected utility, and evaluate transfer to unseen application domains, held-out framings, and different classes of payoff structures. We further introduce incomplete-information settings in which required utilities are withheld and replaced with irrelevant text, testing whether models can distinguish missing decision-relevant information from merely additional context. We find that targeted fine-tuning substantially improves coherent decision-making and that in many situations, learning transfers across domains and framings to situations unobserved during training. Further, models trained for decidability learn to identify when action cannot be justified based on missing information. Finally, we show the value of a routed system that considers separately the recognition of decision completeness and utility-sensitive decision execution.

[AI-115] SpecGuard: Proving a Task Is Broken Before the Agent Cheats

链接: https://arxiv.org/abs/2610.09159
作者: Param Biyani,Krishnamurthy Dvijotham
类目: Artificial Intelligence (cs.AI); Logic in Computer Science (cs.LO); Software Engineering (cs.SE)
备注: 29 pages. Code: this https URL

点击查看摘要

Abstract:As autonomous coding agents get increasingly deployed, the risk that accidental or adversarially injected misspecifications in tasks lead to dangerous agent behavior is critical to address. Prior work has shown that agents given such tasks rarely flag the conflict and instead cheat, editing tests or hard-coding expected outputs, and the actions taken to cheat can cause real damage, such as deleting a security defense to make a corrupted test pass. It remains unclear whether such conflicts can be established with independently verifiable evidence before the agent acts. We present SpecGuard, which detects and formally certifies these conflicts between task intent and tests. Given only the task description and codebase, SpecGuard autoformalizes the intended behaviour into a Lean 4 specification. The tests are formalized independently, and the Lean kernel checks whether any implementation could satisfy both formalizations, producing a machine-checked certificate when none can. On conflicted SWE-bench tasks, SpecGuard detects up to 72.8% of conflicts and formally certifies up to 51.1%, with a nearly five-fold lower conflict miss rate than model-based judgment. SpecGuard provides a pre-execution safety check that identifies reward-hacking opportunities through formal certification of task-level conflicts, before any agent behavior is observed. Our code is available at this https URL.

[AI-116] Frozen Models Evolving Expertise: Model-Agnostic Learning from Deployment Experience for Multimodal Medical AI

链接: https://arxiv.org/abs/2610.09146
作者: Yexiao He,Yucheng Tang,Pengfei Guo,Yufan He,Andriy Myronenko,Can Zhao,Ang Li,Daguang Xu,Dong Yang
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Large language models (LLMs) and vision-language models (VLMs) are usually frozen after deployment, so they do not learn from the cases they solve. This is especially concerning in medicine, where new clinical evidence, updated guidelines, and new therapies can change established practice. Fine-tuning can update the model, but it requires access to model weights and additional training. Parameter-free methods avoid training, but they may overfit a fixed validation set, lack reliable domain knowledge, or lose visual details by saving experience only as text. To address these limitations, we present a model-agnostic framework that allows frozen LLMs and VLMs to learn from deployment experience through three forms of external expertise: a Skill that guides reasoning and tool use, a Knowledge Memory that stores reliable facts supported by earlier cases or trusted external evidence, and a Multimodal Knowledge Base that keeps visual examples and guides the model to relate each retrieved case to the current image. Instead of relying on a fixed validation set, a validation strategy keeps an update only if it helps on new cases without degrading performance on earlier ones. Across six benchmarks covering clinical diagnosis, clinical workflows, medical reasoning, and medical and non-medical visual reasoning, and with four open-weight and closed-source base models, our framework improves performance during online deployment by up to 34.2% over the base model on medical tasks, generalizes to unseen cases, transfers to other models without further optimization, and works in non-medical domains.

[AI-117] DIVA: Dual-Space Intent-Aware Visual Attenuation for Vision-Language-Action Policies

链接: https://arxiv.org/abs/2610.09144
作者: Kaixi Feng,Guoheng Sun,Ziyao Wang,Yexiao He,Zheyu Shen,Ang Li
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Vision-language-action (VLA) policies typically feed dense visual patch tokens into a language-action backbone, preserving scene context but offering no explicit mechanism to regulate how strongly different visual tokens influence policy computation. We introduce DIVA, a Dual-Space Intent-Aware Visual Attenuation module with an anchor-then-attenuate design. DIVA combines high-level task intent with low-level visual evidence to estimate patch-wise relevance anchors, then applies them in two complementary spaces: it reweights projected visual tokens before backbone entry and persistently attenuates low-relevance visual states within the backbone. DIVA preserves the full visual token sequence and requires no external grounding supervision. On LIBERO, DIVA improves OpenVLA-OFT from 96.6% to 98.0% average success and raises its zero-shot LIBERO-Plus score from 69.6 to 72.6. Real-world experiments further show consistent gains under task-irrelevant visual perturbations, supporting the robustness of intent-aware visual attenuation beyond simulation.

[AI-118] Finding Blind Spots in AppWorld and WorkArena Task Verifiers NEURIPS2026

链接: https://arxiv.org/abs/2610.09142
作者: Richard Abrich
类目: Artificial Intelligence (cs.AI); Software Engineering (cs.SE)
备注: 14 pages. Accepted as a poster at the NeurIPS 2026 workshop “Who Verifies the Agents? Toward Reliable Agent Development”. OpenReview paper and supplement: this https URL

点击查看摘要

Abstract:Execution-based task verifiers decide whether an agent succeeded. We audit shipped AppWorld and WorkArena verifiers with source-informed mutation tests. The main audit never modifies a shipped checker. In AppWorld, duplicating a non-idempotent write creates an extra record while preserving every checked field value. The verifier accepts all three task variants from two of five eligible generators: 6/15 constructed effects. A cardinality patch applied to checker copies after the census makes all six cells fail while preserving valid controls. In WorkArena, we prospectively rerun 23 extra-field candidates selected for earlier checker-PASS outcomes. Independent Table API readback confirms nondefault persisted values in 21, while all 23 receive PASS. Two requested strings are aliases of stored defaults. The 21 confirmed wrong effects span three form templates. These selected cases confirm wrong effects under the audit’s protocol; they do not estimate a population rate. No other construction produces an independently confirmed false accept. Other checker-PASS cases are effect-correct degeneracies. We report zero-PASS families separately because retained evidence differs. In fixed intent-swap grids, the checkers return no PASS on 2,689 off-diagonal executions. This is a rejection census: 57 WorkArena cells use session-scoped evidence; the other 2,632 lack classified rejection causes and independent target ground truth. Each increment is specified before its own cells are scored. A supplement accompanies the OpenReview submission with the construction grammar, evidence, content-bound stage lineage and count reproducer. Comments: 14 pages. Accepted as a poster at the NeurIPS 2026 workshop “Who Verifies the Agents? Toward Reliable Agent Development”. OpenReview paper and supplement: this https URL Subjects: Artificial Intelligence (cs.AI); Software Engineering (cs.SE) Cite as: arXiv:2610.09142 [cs.AI] (or arXiv:2610.09142v1 [cs.AI] for this version) https://doi.org/10.48550/arXiv.2610.09142 Focus to learn more arXiv-issued DOI via DataCite (pending registration) Submission history From: Richard Abrich [view email] [v1] Tue, 6 Oct 2026 21:34:00 UTC (31 KB) Full-text links: Access Paper: View a PDF of the paper titled Finding Blind Spots in AppWorld and WorkArena Task Verifiers, by Richard AbrichView PDFHTML (experimental)TeX Source view license Additional Features Audio Summary Current browse context: cs.AI prev | next new | recent | 2026-10 Change to browse by: cs cs.SE References Citations NASA ADSGoogle Scholar Semantic Scholar export BibTeX citation Loading… BibTeX formatted citation loading… Data provided by: Bookmark checked="checked"class=“labs-tab-input”> Bibliographic Tools Bibliographic and Citation Tools Bibliographic Explorer Toggle Bibliographic Explorer (What is the Explorer?) Connected Papers Toggle Connected Papers (What is Connected Papers?) Litmaps Toggle Litmaps (What is Litmaps?) scite.ai Toggle scite Smart Citations (What are Smart Citations?) Code, Data, Media Code, Data and Media Associated with this Article alphaXiv Toggle alphaXiv (What is alphaXiv?) Links to Code Toggle CatalyzeX Code Finder for Papers (What is CatalyzeX?) DagsHub Toggle DagsHub (What is DagsHub?) GotitPub Toggle Gotit.pub (What is GotitPub?) Huggingface Toggle Hugging Face (What is Huggingface?) ScienceCast Toggle ScienceCast (What is ScienceCast?) Demos Demos Replicate Toggle Replicate (What is Replicate?) Spaces Toggle Hugging Face Spaces (What is Spaces?) Spaces Toggle TXYZ.AI (What is TXYZ.AI?) Related Papers Recommenders and Search Tools Link to Influence Flower Influence Flower (What are Influence Flowers?) Core recommender toggle CORE Recommender (What is CORE?) Author Venue Institution Topic About arXivLabs arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv’s community? Learn more about arXivLabs. Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?) mathjaxToggle(); We gratefully acknowledge support from our major funders, member institutions, , and all contributors. About Help Contact Subscribe Copyright Privacy Accessibility Operational Status (opens in new tab) Major funding support from

[AI-119] Breaking the Space Barrier and its Application to Language Model Inference

链接: https://arxiv.org/abs/2610.09139
作者: Arip Asadulaev
类目: Artificial Intelligence (cs.AI); Computational Complexity (cs.CC)
备注:

点击查看摘要

Abstract:Language models are more and more often asked for structured output: JSON that follows a schema, or a tool call with typed arguments. A small machine, an automaton, enforces the format by forbidding the tokens that would break it. We observe that this machine has a rare property: from any of its states, each token leads along exactly one path. Graphs in which only a few paths join any two points are a classical object of complexity theory, and our theoretical result settles an open question about them: one can decide whether such a graph connects two points while verifying that it really has few paths, with very little memory. Precisely, the problem lies in the classes ReachUL, LOGDCFL, C=L and SC2, and needs only O(log2 n/ log log n) space, below the classical O(log2 n) of Savitch’s theorem. The constructions behind the proofs become an inference engine: text the format forces is written without running the model, the mask is recomputed on the GPU without any table, recursive formats use a small stack, every output stays valid under a token limit, and independent fields are decoded in parallel and verified. On one 16 GB Apple M2 Pro with Qwen3.5-2B and 4B, against MLX with llguidance, the standard setup for this hardware, schema-constrained extraction finishes 1.2- 1.3x sooner with the same answers, a grammar costs 3 MB instead of up to 1.5 GB, one server holds sixteen grammars where tables run out of memory, and sixteen tool-calling agents finish 2.5x sooner.

[AI-120] CADFather: Autonomous CAD Reconstruction through Coordinated Tool Use

链接: https://arxiv.org/abs/2610.09127
作者: Gennadiy Savrasov,Maksim Elistratov,Nikita Gavrilov,Albert Garifullin,Oleg Pavlov,Soslan Kabisov,Vladimir Frolov,Anton Konushin,Andrey Kuznetsov,Dmitrii Zhemchuzhnikov
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: 17 pages, 7 figures, 6 tables

点击查看摘要

Abstract:Reconstructing an editable CAD model from a 3D shape remains a challenging engineering task. Existing methods can propose CAD operations, but no single source of proposals works equally well across different part geometries and stages of reconstruction. We introduce CADFather, an autonomous agentic system that coordinates complementary tools to recover parametric CAD programs from 3D meshes. A vision-language assistant inspects renders of the target and intermediate reconstructions, then decides which candidate CAD programs to extend, which tools to invoke, how many proposals to generate, and when to finish. Learned and algorithmic tools propose CAD operations, while numerical optimization refines the parameters of existing programs. Proposed or refined programs are executed and evaluated to provide feedback for subsequent decisions. The agent maintains alternative candidate programs for each target part and preserves the best valid result throughout reconstruction. CADFather uses pretrained generation and assistant models without additional training. We evaluate reconstruction quality and execution validity on the full DeepCAD, Fusion360, and MCB test sets, as well as on CADENA-Bench, CADBench, and BenchCAD. We additionally analyze computational cost and the trade-off between cost and reconstruction quality.

[AI-121] Spatial Induction Heads: In-Context Learning of Multidimensional Cellular Automata ICLR2027

链接: https://arxiv.org/abs/2610.09124
作者: Kimia Kazemian,Menghan Xu,John Thickstun,Sarah Dean
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: Under review at ICLR 2027. Project page: this https URL

点击查看摘要

Abstract:Induction heads provide a mechanistic account of in-context learning in sequential data, but existing theory largely assumes that the context relevant to a prediction forms a contiguous block. In multidimensional data, serialization breaks this assumption by scattering spatial neighbors across distant positions in the token sequence. We study how transformers overcome this routing problem in multidimensional stochastic and deterministic cellular automata, where each trajectory is generated by an unknown local rule and presented as a flattened sequence without an explicit coordinate-based spatial inductive bias. We introduce spatial induction heads, two-layer gather-and-match circuits in which the first layer reconstructs the relevant spatial neighborhood and the second matches the resulting configuration against earlier occurrences. We give two explicit realizations of the gather and show that the positional dimension required for spatial routing depends only on the local neighborhood and spatial dimension, not on grid volume or trajectory horizon. We further construct a matching layer which implements Bayesian counting. The end-to-end circuit can approximate the Bayesian posterior arbitrarily closely for stochastic rules and can predict exactly for deterministic rules. Empirically, trained two-layer transformers generalize to unseen rules in one and two dimensional settings, achieving near-perfect deterministic rollouts and less than 0.005 nats KL from the Bayes-optimal predictor on stochastic rules. Attention patterns and layerwise probes align with the predicted gather-and-match computation, providing mechanistic evidence for spatial induction in trained transformers.

[AI-122] Which Buildings Are Artificial Intelligence-Ready? A Measurement-Based Assessment Framework for AI Question Answering and Actuation

链接: https://arxiv.org/abs/2610.09119
作者: Wooyoung Jung
类目: Artificial Intelligence (cs.AI)
备注: 47 pages, 11 figures, 23 tables. Code, results and data supplement: this https URL

点击查看摘要

Abstract:Agentic artificial intelligence (AI) systems are becoming the interface to buildings, answering questions and controlling operations, but a building’s readiness for them has not been systematically assessed. This study proposes a framework to quantify it. First, a building’s knowledge graph sets two ceilings. The answerable-readiness ceiling is the share of operational questions its data could answer, and the actuation-readiness ceiling is the share of control actions it exposes. Second, a reference AI agent’s accuracy on a fixed set of these questions shows how much of the ceilings is realized. On a simulated office, the agent realizes 0.62 of a 0.64 answerable ceiling, so missing data, not the AI, limit readiness, except in naming a fault’s cause. Across 37 public real-building graphs, the median answerable ceiling is 0.16, and in 15 of 45, unlinked sensors lower it. The framework turns “is this building AI-ready?” into an auditable, ranked retrofit question.

[AI-123] GeoNatureAgent (GNA): A Framework and Benchmark for Pre-Production Evaluation of Tool-Using Agents on Geospatial and Environmental Tasks

链接: https://arxiv.org/abs/2610.09112
作者: Gabriel Diaz-Ireland,Diego Prieto-Herráez,Mario García Peces,Javier Velázquez,Benjamin Zaitchik,Devika Jain
类目: Artificial Intelligence (cs.AI)
备注: 10 pages, 4 figures. Extended version of arXiv:2606.12821 (ACM SIGSPATIAL 2026 poster paper). Published at GeoIndustry '26 (ACM SIGSPATIAL Workshop)

点击查看摘要

Abstract:Before tool-using LLM agents are deployed in environmental and geospatial workflows, teams need evidence that an agent reliably selects the right operations against real APIs. We introduce GeoNatureAgent (GNA), a framework for pre-production evaluation of tool-using agents: a fixed sixteen-tool geospatial interface published as a Model Context Protocol (MCP) server, so the agent under test is the only variable, scored against an identical tool layer, task suite, and deterministic scorer. Its flagship instance is a 103-task benchmark (a 93-task main suite across 18 categories plus a ten-task comparison expansion) evaluated against an open, self-hostable geospatial API serving three environmental indicators across Spain and Portugal. We evaluate nine LLMs under three temperature-1.0 seeds, reporting capability and per-case cost as orthogonal axes. (1) Claude Sonnet 4 achieves the highest capability (61.7% +/- 0.7% on all 103 tasks; 60.8% on the main suite), followed closely by DeepSeek V3.2 (57.9%), while no other model exceeds 53%; (2) the cost-accuracy Pareto frontier is mostly open-weight, with DeepSeek V3.2 offering 93% of Claude’s capability at 11.3x lower list-price cost; (3) under strict all-checks scoring the best model sits 24-36 points below the 85-97% reported on general-purpose GIS benchmarks, whereas per-check partial credit for the top four models (86-90%) is comparable, so much of that gap reflects scoring strictness rather than task difficulty alone. The MCP server, evaluation harness, benchmark, and API are publicly available; swapping the tool executors and task suite instantiates an equivalent benchmark for any geospatial domain.

[AI-124] Convex-Concave Reinforcement Learning

链接: https://arxiv.org/abs/2610.09108
作者: Shripad V. Deshmukh,Yaswanth Chittepu,Dhawal Gupta,Philip Thomas,Scott Niekum
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: 37 pages, 5 figures. Code: this https URL

点击查看摘要

Abstract:Policy learning drives many of the most consequential and heavily-invested applications of reinforcement learning today. Yet the core optimization problem it rests on (maximizing expected return) is notoriously non-convex, even under a direct policy parameterization, and the field has largely responded by avoiding it: optimizing convex surrogate approximations of the return under trust-region constraints (NPG, TRPO, PPO, AWR). We show that this seemingly unstructured problem is not actually structureless. In log-density-ratio coordinates y := \log[\pi/\pi_n] , the exact per-iteration objective, computable via per-decision importance sampling (PDIS), is a difference-of-convex-constrained difference-of-convex (DC-constrained DC) program. This structure lets us move beyond surrogate approximations: it recovers CPI, NPG, TRPO, and AWR as special cases along interpretable axes, and it opens a multi-step axis k that couples consecutive decisions. We solve the per-iteration program with sequential convex programming (SCP), the standard solver for difference-of-convex problems, and give convergence guarantees under mild conditions, bridging the difference-of-convex optimization and RL literatures. Empirically, multi-step Convex-Concave RL (CCRL) wins on diagnostic MDPs where credit must propagate across a horizon (its advantage growing with the dependency length), is competitive with a tuned PPO on classic control, and on a realistic, stochastic, mid-horizon healthcare domain converges markedly faster than tuned PPO to the same near-optimal survival, with an 11.3% higher area under the training curve.

[AI-125] MaRK: Markov-adapted Recurrent Kernels for Dynamic Operator Conditioning in State Space Models

链接: https://arxiv.org/abs/2610.09092
作者: Syed Ibrahim Omer,Ginny Y. Wong. Xiangyu Zhao
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:State Space Models (SSMs) offer an efficient alternative to Transformers for sequence modeling, yet conditioning pre-trained SSMs for iterative generation typically operates outside the recurrent operator, through input injection or activation modulation. While such mechanisms expose the model to conditioning information, they leave the underlying temporal dynamics fixed. We introduce MaRK (Markov-adapted Recurrent Kernels), a dynamic operator-conditioning framework that maps context vectors directly into bounded modulations of a frozen SSM’s recurrence ( A ), read-in ( B ), read-out ( C ), skip ( D ), and discretization ( \Delta ) parameters. Viewed through the lens of LPV-SSM systems, MaRK induces a context-indexed family of Markov parameter sequences, allowing each diffusion timestep to reshape the model’s input-output memory kernel. We instantiate MaRK on a frozen 111M-parameter Hydra SSM backbone and study three adapter geometries: Hypernet, Chebyshev polynomial, and Discrete Cosine Transform kernels. Since these adapters modify the Markov parameter sequence through low-rank auxiliary maps on the frozen backbone, parameter-efficient fine-tuning arises as a structural consequence of the adaptation mechanism itself, requiring only 6.3–11M trainable auxiliary parameters to transition from a bidirectional objective to an iterative diffusion regime. The bounded recurrence parameterization further yields an analytic Affine Quadratic Stability certificate for the modulated recurrence. Through synthetic LPV recovery experiments and Markov-operator diagnostics, we show that MaRK recovers coordinate-invariant temporal operators under matched assumptions and produces distinct, stable timestep-conditioned memory profiles. Empirically, the Chebyshev variant yields the strongest performance, achieving an average validation loss of 2.55, followed by the DCT (2.59) and Hypernet (3.77) geometries.

[AI-126] RippleCP: Measuring Counterfactual Checkpoint Advantage in Coding Agents

链接: https://arxiv.org/abs/2610.09088
作者: Mayur Akewar,Ravi Ranjan
类目: Artificial Intelligence (cs.AI)
备注: 16 pages

点击查看摘要

Abstract:Agent checkpoint systems decide what state is recovery-relevant, how to snapshot it, and whether rollback is admissible. None decides which of the safe boundaries they expose are worth materializing. We formulate this as counterfactual checkpoint advantage, the reduction in future recovery cost obtained by checkpointing a candidate rather than skipping it, and measure it by driving a CP branch and a SKIP branch to the same logical failure and recovering both under matched model, tool, verifier, and stopping conditions. On a frozen pilot of 12 SWE-bench Verified tasks and 106 real recovery branches, checkpointing saves 49.4 s per task, and that figure resolves into two regimes two orders of magnitude apart. The first checkpoint returns 100.0 s on 156.5 s of protected work, a conversion of 0.64; a second one step later returns-1.1 s on 41.8 s, a conversion of -0.03. Recovery is a re-derivation rather than a replay, so preserved work is a poor guide to saved work, and the classical elapsed-work rule misprices the second checkpoint by its full nominal cost. We identify where placement can pay, and set the bar a placement policy must clear.

[AI-127] Epistemic Uncertainty-Aware Defect Detection for Quality Control in Medical Device Manufacturing

链接: https://arxiv.org/abs/2610.09057
作者: Raham A. Butt,Marco Romanelli,Roche C. de Guzman
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Objective: We investigate whether accounting for epistemic uncertainty can improve the reliability of automated defect detection in medical device manufacturing. Methods: We consider a machine learning framework that operates on heterogeneous manufacturing and device-report data represented with Knowledge Graphs. To mitigate errors arising from uncertainty in the decision model, we analyze a principled rejection strategy to abstain from predictions whose estimated epistemic uncertainty exceeds a specified threshold. We evaluate the approach using standard synthetic benchmarks and real-world medical device report data. Results: The theoretical results establish the validity of the method characterizing the regimes under which it is expected to be effective. Empirically, the rejection strategy enables explicit control of coverage, that is, the proportion of samples for which the model issues predictions, while improving performance on the retained samples. On 266,170 real-world FDA MAUDE device reports, a 10% abstention rate reduces classification error by 48%, and more aggressive rejection (approximately 70% coverage) yields near-perfect accuracy on the retained samples. On standard synthetic manufacturing benchmarks, abstaining on 9% of the decisions, our approach reduces the risk up to 63% compared with the standard no-abstention approach. Conclusions: Abstaining from predictions with high epistemic uncertainty can provide a practical tool for controlling the reliability of machine learning-based defect detection, especially in high-stakes medical device manufacturing applications. Significance: Uncertainty-aware defect detection may support safer and more reliable quality assurance in medical device manufacturing by identifying cases that require additional inspection rather than issuing potentially harmful predictions.

[AI-128] MimicX: Policy-in-the-Loop Supervision Refinement for Video-Driven Humanoid Motion Tracking

链接: https://arxiv.org/abs/2610.09055
作者: Shuaijun Liu,Chenglong Zhang,Xuhao Liu,Feiyang You,Yifan Liao,Shuyang Hao,Chaozhe Zhang,Chengyu Wu,Zhen Sun,Ningxin Su
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI)
备注: 33 pages, 26 figures, 19 tables, including references and appendices. Project website: this https URL

点击查看摘要

Abstract:Human videos provide rich motion targets for humanoid learning, yet visually plausible references can still produce persistent failures under physics-based execution. These failures reveal where training supervision should change. We present MimicX, a policy-in-the-loop framework that uses execution feedback to refine video-driven humanoid motion tracking. Starting from reconstructed and retargeted motion, MimicX localizes difficult transitions and affected body regions, then jointly adapts tracking objectives and the reset curriculum for policy continuation. Repeated rollout verification selects execution-priority improvements subject to tracking guards. Across four core video tasks, MimicX consistently improves tracking accuracy and Robust Execution Horizon relative to the Fixed Reference baseline. Task-averaged results show a 25.7% reduction in body-tracking error and a 255.6% increase in execution horizon. Additional video, motion-reference, and collision-scene studies evaluate the method beyond the core tasks, while MimicX-HLoop accelerates feedback through heterogeneous execution. Overall, MimicX turns policy failure into actionable supervision for deciding what to refine and which refinement to retain.

[AI-129] Justice After Identity: Large Language Models and the View from Everywhere

链接: https://arxiv.org/abs/2610.09053
作者: W. Russell Neuman
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:The search for a common view of justice and fairness has challenged human collective activity, as our diverging judgments are unavoidably shaped by the self-interests of social position, personal benefit, cultural inheritance, and historical circumstance. John Rawls famously attempted to overcome this limitation through popularizing a philosophical tradition known by the phrase “the original position” - a thought experiment by which people select principles of justice without knowing the identities or advantages they will possess. Critics, however, have long questioned whether people can meaningfully suspend their social identities and suppress morally relevant forms of lived experience. Artificial intelligence engaged to calculate algorithmic and agentic fairness introduces a novel possibility. LLMs have no singular class, race, gender, nationality, or biography, yet their parameters encode linguistic representations of a vast range of human identities and moral traditions. Perhaps the ethical judgments of LLMs could approximate an integrative original position - a “view from everywhere” generated not by excluding social identities but by computationally incorporating their this http URL is unlikely that humankind will “hand over the keys” to computational systems by simply delegating complete agentic control of distributive and procedural collective processes. But AI may play a role, perhaps a positive one, interacting with individual and collective human judgment as we often confront increasingly polarized views on what is fair and just. We present data comparing human, base model, and frontier/fine-tuned model judgments about classic moral dilemmas while systematically varying identity relationships and Rawlsian constraints on identity. We conclude by speculating whether, if advanced AI systems provide humans with thoughtful advice, humans would actually be likely to accept it.

[AI-130] Not Every Call Needs a Frontier Model: Per-Call-Site Evaluation of Small Language Models in a Deployed Agent ic Home-Automation System NEURIPS2026

链接: https://arxiv.org/abs/2610.09021
作者: Panagiotis Kasnesis,Christos Chatzigeorgiou,Lazaros Toumanidis,Amalia Contiero Syropoulou
类目: Artificial Intelligence (cs.AI)
备注: 17 pages, 2 figures, accepted in NeurIPS 2026 Workshop on SLMs for Agentic Systems, Paris,

点击查看摘要

Abstract:An agentic system issues several structurally different kinds of LLM calls. It routes intent, classifies actions, grounds language in a device registry, plans multi-agent pipelines and writes the Python code those pipelines run. The difficulty of these call sites varies by an order of magnitude, yet in practice a single model, chosen for the hardest site, serves all of them. In this work, we evaluate 9 models from 0.8B to a frontier hosted model across the five call sites of a deployed open-source home-automation framework (Wactorz), using its unmodified production prompts and two real Home Assistant installations (280 cases, 2520 scored calls). We find that capability is not ordered the same way at every site, and that larger models are not uniformly better: one 4B model is worse than its 2B sibling at grounded actuation. Paired testing shows the best local model to be statistically indistinguishable from both hosted models at four of five sites. Only code generation separates them, against a small hosted model (p = 0.039) as well as a frontier one (p = 0.002). Aggregate accuracy also hides a safety failure specific to actuation, where small models resolve the accuracy/refusal trade-off in degenerate ways: one model (Gemma4 E2B) actuates on 87.2% of requests for devices the site does not own, while another refuses every request it receives. Routing each site to its best local model reaches 91.8% against 95.4% at no per-call cost. In a live deployment judged by a user, hosting only the two generative sites matches hosting everything (39/43 against 39/43) for 28% of the spend, and the actuation gap the benchmark predicted appears as exactly one case in twenty-six. Benchmark, harness and all records are released at this https URL.

[AI-131] PAIR: Bridging Perception and Action in Vision-Language-Action Models

链接: https://arxiv.org/abs/2610.09016
作者: Kaixi Feng,Guoheng Sun,Ang li
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Vision-language-action (VLA) models map visual observations and language instructions to continuous robot actions. This task requires a transition from representations that describe the scene and instruction to representations that support action generation. Many continuous-action VLAs leave this transition implicit and supervise it mainly through the final action-prediction loss. We introduce PAIR, a framework that learns a shared perception-action representation between these two spaces. During training, a Masked Action Autoencoder encodes expert action chunks into horizon-aligned Action Latent Tokens. A Bridge Module extracts task-relevant features from the current visual-language representations. PAIR aligns these features with the Action Latent Tokens to form Bridge Tokens that preserve task information and capture the structure of expert actions. The Bridge Tokens are then projected into the action-token space and injected into the initial Action Tokens, providing an action-ready starting point for Action Expert refinement. At inference, the autoencoder is removed, and the Bridge Tokens are generated only from the current observation and instruction. Experiments on LIBERO, LIBERO-Plus, and CALVIN ABC-D show gains for the evaluated OpenVLA-OFT and VLA-Adapter models. On LIBERO-Plus, PAIR raises VLA-Adapter’s success rate from 59.1% to 64.2%. On CALVIN, it increases VLA-Adapter’s average completed sequence length from 4.42 to 4.53. Across seven real-world tasks, PAIR raises OpenVLA-OFT’s success rate from 51.4% to 65.0%. Representation analyses show that Bridge Tokens retain task information while making continuous-action information accessible before Action Expert refinement. These results support a shared intermediate representation as a useful interface between perception and action in continuous-action VLAs.

[AI-132] Enabling Dynamic Computation in Looped LMs

链接: https://arxiv.org/abs/2610.09013
作者: Aayush Mishra,Arnau Padrés Masdemont,Victor Conchello Vendrell,Jordi Ros Giralt,Arash Behboodi,Fabio Valerio Massoli
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Looped LMs are parameter efficient and promise dynamic computation (saving memory and FLOPs on easy tokens). However, state-of-the-art open Looped LMs trained with this dynamic computation capability (Ouro models) do not realize it in practice as each loop iteration (depth) requires its own level of KV-cache, necessitating all loop computations. Moreover, Ouro’s early-exit prior is enforced on each token equally, which results in static lower-depth like processing of all tokens regardless of difficulty. In this work, we propose a simple “best-available” KV caching strategy that works out-of-the-box, creating a new frontier in the performance vs depth space. Our approach enables up to 30% reduction in FLOPs and KV memory while retaining full-depth performance, showing the true flexibility of Looped LMs. Furthermore, training looped LMs with awareness about this KV caching strategy improves performance and efficiency. Finally, we apply a small but effective fix to the early-exit prior enforcement objective that makes tokens exit at truly heterogeneous depths based on effort. Our findings are validated on Ouro models as well as smaller looped LMs pre-trained from scratch.

[AI-133] Whose Memory Is It? Scope-Aware Commit Rules for Long-Term LLM Memory

链接: https://arxiv.org/abs/2610.09008
作者: Hongyu Gu,Xinchang Li
类目: Artificial Intelligence (cs.AI)
备注: 20 pages, 2 figures

点击查看摘要

Abstract:Persistent memory allows an LLM agent to carry experience across conversations, but it also turns a local reasoning mistake into a durable one. During deliberation, an agent may consider a plan, simulate a tool result, report another speaker’s belief, and then reject all of them. If memory retains only the resulting sentences, those once-useful possibilities can later return as facts. The record is neither fabricated nor irrelevant; it has simply been detached from the context in which it was valid. We identify this missing context as \emphdiscourse ownership: the world, branch, or speaker that licenses a proposition. Our first finding is counterintuitive. Language models already carry a causally active signal for ownership, yet conventional memory interfaces discard it when they convert reasoning into records. We introduce CASK (Causally Anchored Scoping Keys), a commit rule that preserves this signal so that shared-world facts enter durable memory while provisional content remains available only within its original scope. Our second finding is that the most obvious way to preserve the signal—storing the discovered internal coordinates—is unreliable because equivalent representations need not keep the same coordinates. CASK instead preserves the stable relations that express ownership. Controlled long-conversation conflicts and tool-agent traces show that this design improves memory admission and prevents provisional content from contaminating later answers while complementing runtime provenance. The resulting commit boundary lets agents explore more possibilities without granting every intermediate sentence authority over future behavior.

[AI-134] Sigma-Hunter: A Domain-Specific Language Model for Threat Hunting and Detection Engineering

链接: https://arxiv.org/abs/2610.09007
作者: Kemal Davaslioglu,Sastry Kompella
类目: Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR)
备注:

点击查看摘要

Abstract:Detection engineers must translate threat reports, forensic observations, and hunt hypotheses into precise, testable rules. General-purpose large language models (LLMs) can draft such rules, but often produce invalid YAML, incorrect log sources, unsupported fields, or overly broad detection logic. This paper presents \emphSigma-Hunter, a domain-adapted LLM for analyst-assistive Sigma rule generation and threat hunting. We build an instruction-tuning dataset from 3,635 validated open-source Sigma rules, expanded into 7,663 question-answer and analyst-reasoning examples. Each source rule is assigned to a single train, validation, or test partition before this expansion, so no rule leaks across splits. We fine-tune a 7B Mistral model and a Phi-4 model with LoRA and score held-out rule generations on syntax, approximate field consistency, and a semantic judgment of detection logic, completeness, selectivity, and log-source alignment. Sigma-Hunter-Mistral scores 8.17 overall, against 7.88 for the strongest general-purpose baseline and 4.61 for untuned Mistral. Two findings stand out: domain adaptation enables a compact 7B model to perform competitively with larger general-purpose models on this structured task, and syntactic validity is a weak proxy for semantic rule quality, as several baselines emit well-formed YAML carrying weak detection logic. The adapted models run locally, which suits detection engineering in disconnected environments where analysts cannot reach hosted model services.

[AI-135] Algorithmic Scratchpads and Curriculum Staging for Arithmetic Reasoning in Tiny Transformers

链接: https://arxiv.org/abs/2610.09003
作者: Sourabh Kasliwal
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: 16 pages, 5 figures. Code and benchmark datasets available at this https URL

点击查看摘要

Abstract:Autoregressive Large Language Models (LLMs) frequently struggle with deterministic multi-step algorithmic tasks such as multi-digit multiplication and long division. In this paper, we investigate the mechanics of multi-step arithmetic in compact “Tiny” Transformers (~10.6M non-embedding parameters, 49.3M total) trained on synthetic data across four basic operations (+, -, *, /) unrolled as step-by-step scratchpads. First, we establish the necessary training foundations: (1) dataloader sequence padding creates an 83% gradient starvation artifact that collapses accuracy from 40% to 1%, remediated via continuous sequence packing; (2) linguistic pretraining is an essential prerequisite (= 2.0% without it); and (3) modern architectural primitives (RoPE, RMSNorm, SwiGLU) and Sparse Mixture of Experts (MoE) substantially improve additive reasoning over baseline GPT-2. Second, we demonstrate that algorithmic scratchpad formulation directly dictates success. Introducing a deterministic Digit-by-Digit Long Division scratchpad within a 4-stage Hierarchical Developmental Curriculum dramatically elevates single-digit division from 4.0% to 86.7% accuracy on a 4,000-problem held-out benchmark. In contrast, multi-digit multiplication remained challenging: detailed error analysis revealed that while the model correctly computed single-digit sub-products and place-value zeros, our FOIL scratchpad failed because it forced a simultaneous summation of up to nine multi-digit terms in a single step without pairwise intermediate accumulation. Finally, we identify two key boundaries: performance collapses to 0.00% on unseen 4-digit operands, and unbuffered training induces catastrophic forgetting, collapsing division accuracy from 86.7% down to 0.00%.

[AI-136] PhysEvo: Astra Can Act Let It

链接: https://arxiv.org/abs/2610.08995
作者: Wenqing Tian,Zeyu Zhang,Zhaocheng Liu,Fengwei Liu,Qiang Liu,Liang Wang
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Astra can act, yet reliable manipulation depends on the system through which it observes and controls the world. We introduce PhysEvo, a framework for physical recursive self-improvement (RSI) around a single frozen model. A task agent executes robot tasks; a meta-agent uses the resulting trajectories to diagnose failures, revise tools and skills, and test corrections. The meta-agent can also improve its own diagnostic tools, so retained revisions support both later action and later self-improvement. This process develops joint-level control, evidence-seeking observation, and reusable manipulation skills without model-weight updates or a separately trained action policy. Across 42 RoboDojo tasks, held-out-layout evaluation of retained task-specific deployment versions yields a five-dimension average score of 68.14/100 and 62.00% success, compared with 47.17% for RoboDawn’s one-shot Astra agent, the strongest published reference in our comparison. On eight manipulation tasks challenging direct Astra, PhysEvo achieves 55.00% success, compared with 1.25% for the direct-Astra reference. Deploying the simulation-evolved harness on AgileX PiPER and continuing skill revision yields 90.60/100 average score and 84.00% success across 25 trials on five real-world tasks. PhysEvo turns the consequences of action into persistent, testable changes to how a frozen model acts and improves.

[AI-137] Verify Less Evolve More: Training Idea-Level Critics for Verification-Efficient ML Evolving Agents

链接: https://arxiv.org/abs/2610.08993
作者: Jiamu Bai,Lizhu Zhang,Xin Yu,Yanhong Wu,Zellux Wang,Serena Li,Weiwei Li,Zhuokai Zhao,Lingzhou Xue,Kiwan Maeng,Xiangjun Fan,Bo Peng
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:As large language models become more powerful, self-evolving agents are able to tackle challenging tasks including AI for machine learning (AI4ML). In AI4ML, while empirical verification is available, it often requires computationally costly model training and evaluation, limiting the speed and scale of agent evolution. Yet verification efficiency remains under-explored, and frontier models provide only limited gains when used directly as idea selectors. We address this gap with specialized idea-level critic models that predict whether a proposed ML modification will improve upon the current solution, allowing agents to screen ideas and concentrate verification resources on the most promising candidates. We train the critic models through supervised fine-tuning on high-quality critiques synthesized by Gemini-3.1-Pro, followed by GRPO to further improve their predictive accuracy. Empirically, our critic models outperform Gemini-3.1-Pro in static idea evaluation, and these gains extend to agent inference, continual learning, and policy training. During inference-time evolution, they improve final solution quality under the same verification budget by selecting more promising ideas, with further gains from continual learning. During policy training, they serve as learned reward models, reserving empirical verification for uncertain cases and enabling substantially more policy updates with the same verification resources. Together, these results show that idea-level critic models help ML agents discover better solutions and learn stronger proposal policies under limited verification budgets.

[AI-138] LASER: Latent Space Adjoint Matching for Support-Constrained Entropy-Regularized Offline RL NEURIPS2026

链接: https://arxiv.org/abs/2610.08989
作者: Songyuan Zhang,Oswin So,Eric Yang Yu,Matthew Cleaveland,Peter Crowley-Dolen,Chuchu Fan
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Robotics (cs.RO); Optimization and Control (math.OC); Machine Learning (stat.ML)
备注: 29 pages, 16 figures. Accepted at NeurIPS 2026

点击查看摘要

Abstract:While offline reinforcement learning (RL) enables policy optimization from static datasets without costly online interaction, it remains bottlenecked by the risk of executing out-of-distribution (OOD) actions. Recent approaches mitigate this by learning a behavior-cloning policy through flow matching and then performing RL within its constrained latent space. However, naively optimizing the latent policy can easily cause the policy to collapse into a brittle mode or exploit sharp artifacts of the learned critic. In this work, we find that entropy regularization is essential in latent-space RL for addressing these challenges. We introduce LASER, a novel offline RL algorithm that applies latent-space adjoint matching to achieve entropy-regularized latent-space RL with expressive flow policies while avoiding backpropagation through time. Through comprehensive experiments on 40 challenging OGBench tasks with varying dataset qualities, we show that LASER achieves state-of-the-art performance. Notably, LASER uses fixed method-specific hyperparameters across all tasks and outperforms the evaluated baselines, including those with task- and dataset-specific tuning, which highlights the robust applicability of LASER. Project website: this https URL.

[AI-139] Work While They Sleep: Exploiting Evaluation Latency for Fully Bayesian Optimization

链接: https://arxiv.org/abs/2610.08969
作者: Gustavo Sutter,Alejandro Comas-Leon,David Holzmüller,Hao Wang,Luis Ricardez-Sandoval,Pascal Poupart,Agustinus Kristiadi
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Machine Learning (stat.ML)
备注:

点击查看摘要

Abstract:Black-box optimization problems are ubiquitous across science and engineering, often dealing with expensive objective functions. This objective latency has two consequences during optimization: (i) the objective evaluation dominates execution time, and (ii) sample-efficient algorithms are crucial to accelerate development and avoid wasting resources. Bayesian optimization (BO) methods are the \textitde facto choice of planners for suggesting the next point to try. Standard BO fits the surrogate model’s hyperparameters with a point estimate. Alternatively, a fully Bayesian approach uses model averaging to account for uncertainty over the hyperparameters, leading to better uncertainty estimates—useful in the low-data regime that is pervasive in BO. However, it is often prohibitively expensive and thus rarely used. In this work, we propose ELF-BO, an algorithm that uses the objective evaluation latency to headstart the computation of the next suggestion, allowing for fully Bayesian optimization without incurring substantial decision-time costs. This is done by sampling from the hyperparameter posterior \emphwhile the objective is being evaluated, only requiring reweighting of the samples once the objective value is observed. Across synthetic functions and real-world applications, we show that ELF-BO matches the performance of fully Bayesian methods while only incurring decision latency on par with or better than standard BO. Thus, ELF-BO makes fully Bayesian optimization practical in real-world use cases.

[AI-140] Socio-Foundation: A Model for Generalizable Individual Behavior Simulation via Hierarchical Capability Distillation

链接: https://arxiv.org/abs/2610.08967
作者: Liang Wang,Wenxuan Xie,Xinyi Mou,Yixin Luo,Zhongyu Wei
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Simulating individual behavior requires large language models (LLMs) to preserve persona traits while adapting to dynamic social contexts. However, general-purpose LLMs often flatten distinct personas, while task-specific tuning suffers from fragmentation and generalization. To overcome these challenges, we organize individual simulation into the \textbfFONTS Taxonomy, comprising five complementary capability dimensions: \emphpersona fidelity (\textbfF), \emphoutcome realization (\textbfO), \emphbehavioral naturalness (\textbfN), \emphtrajectory coherence (\textbfT), and \emphsocial grounding (\textbfS). Grounded in this taxonomy, we curate a standardized training corpus library of approximately 10 million instances across 14 representative datasets and present \textbfSocio-Foundation. Socio-Foundation decouples specialization from integration via a three-stage pipeline: learning task experts via DAPO, consolidating them into capability experts via off-policy distillation, and unifying them via multi-teacher on-policy distillation (MOPD). We also establish \textbfIndiEval, consolidating 29 metrics across the FONTS dimensions. Experiments show that Socio-Foundation outperforms its \textitQwen3-8B base by 11.0 points and approaches frontier models such as \textitGLM-5.2, with ablations and out-of-distribution evaluations further demonstrating the effectiveness and generalization of our model.

[AI-141] Humanitys Sixth Sense: Benchmarking Intuitive Visual Reasoning in Multimodal Models

链接: https://arxiv.org/abs/2610.08966
作者: Xingang Guo,Jing Gu,Brian Jang,Renxiong Wang,Utkarsh Tyagi,Daniel Quigley,Steven Li,David Yan,Daniel Yue Zhang,Darvin Yi,Forrest Huang,HiJae Kim,Tianyi Zhang,Jared Lichtarge,Jihua Huang,Le Xue,Manan Tomar,Qiuyi Richard Zhang,Ruofei Yu,Seth Neel,Yaning Hu,Marcella Valentine,Xinzhe Jiang,Daniel Evans,Chenguang Wang,Dustin Tran,Tong Zhao,Yinfei Yang,Yunzhong He
类目: Artificial Intelligence (cs.AI); Multimedia (cs.MM)
备注:

点击查看摘要

Abstract:Humans perceive far more in a scene than what is explicitly depicted: a single glance captures past causes and future trajectories; a quick peek determines if a vehicle can fit between two parked cars; a few seconds of video reveals who holds authority in a room; and a fleeting clip highlights subtle abstract patterns like unwritten rules or hidden labels. This capacity reflects a form of humanity’s sixth sense: an intuitive reasoning mechanism that recovers implicit information beyond raw sensory perception. Crucially, this rapid, zero-shot visual intuition underpins everyday navigation and social interaction, making it a vital capability for Multimodal Large Language Models (MLLMs) deployed alongside people. Existing visual benchmarks, however, target either deliberate expert-level analysis in academic and mathematical domains or low-level perception, leaving the intuitive reasoning that people perform largely untested. To bridge this gap, we introduce Humanity’s Sixth Sense (HSS), a benchmark for intuitive visual reasoning. HSS spans diverse image and video inputs, organizes items under a structured taxonomy, and pairs each with human-written prompts probing the implicit temporal, spatial, social, and abstract structure that people infer at a glance. Frontier MLLMs fall short of human performance: participants reach 93.1% accuracy, while the strongest model, GPT-6-astra, reaches only 53.6% even at maximum reasoning effort. Despite excelling in many complex tasks that require advanced perception and knowledge, current models still struggle significantly on these visual tasks that are intuitive for humans. We further explore agentic setup that apply dynamic visual manipulation to HSS, which narrows but does not close the gap. HSS establishes intuitive visual reasoning as a measurable axis and directs attention to a capability that scaling on current benchmarks has so far left behind.

[AI-142] GraphOPD: Graph-Augmented On-Policy Distillation for LLM Agents

链接: https://arxiv.org/abs/2610.08959
作者: Bohan Lin,Liyi Chen,Zhuoning Guo,Muyang Li,Qimeng Wang,Yan Gao,Yao Hu,Yudong Zhang
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:On-policy distillation post-trains large language model agents by supplying dense, step-level guidance from a teacher policy when the reinforcement-learning reward is sparse and arrives only once per trajectory. Existing instantiations allocate this guidance by the size of the teacher-student divergence at each step, on the single-turn intuition that a large disagreement marks a mistake worth correcting. Once decisions chain over many turns, that rule misfires, since an early drift enters every later context both policies condition on, leaving the teacher consistent with the drifted trajectory instead of flagging its cause, while interchangeable steps register large but outcome-irrelevant divergences. We demonstrate this on an agentic benchmark, where distilling the highest-divergence steps brings no consistent benefit over random selection. To this end, we introduce GraphOPD, the first method to bring graph-based structural augmentation into on-policy distillation for agent capabilities. It reads which steps enabled which later ones from the environment’s own record of state changes, immune to the drift that corrupts the teacher-student gap, organizes them into a dependency graph, scores each step by a random-walk stationary distribution over it, and fuses that structural credit with the divergence signal into a trajectory-relative mask concentrating supervision on each rollout’s highest-aptitude steps. Across three model scales and eleven baselines on ALFWorld, WebShop, and SearchQA, GraphOPD shows competitive performance throughout, improving over the strongest baseline by up to +5.8 pp. An executed-replay audit further shows that this structural credit score tracks true causal impact far above chance, that both fused signals are independently necessary, and that the same signal transfers to out-of-domain tool-integrated reasoning.

[AI-143] Multi-Aspect Runtime Verification for Simulation-Based VV of LLM -Enabled Autonomous Agents

链接: https://arxiv.org/abs/2610.08928
作者: Nikolaos Kekatos,Dimitrios Nikou,Anastasios Temperekidis,Alexios Lekidis,Nikolaos Kolokotronis,Panagiotis Katsaros,Stylianos Basagiannis
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Logic in Computer Science (cs.LO)
备注: 28 pages, 2 figures, 2 tables. Accepted at Modelling and Simulation for Autonomous Systems (MESAS 2026), to appear in Springer LNCS

点击查看摘要

Abstract:LLM-based agents are entering decision-support roles in defence staff work, where the obligations they must respect are already written down and binding, and where retraining is not available as a control because models arrive as procured components. What can be placed under engineering control is the interface between the agent and the systems it acts on. Those obligations are at once spatial, temporal and text-semantic, and a violation typically lives in the composition of a multi-step interaction, which is why per-event guardrails miss sequential tool-attack chains. We present a multi-aspect runtime-verification framework that decomposes a natural-language policy clause into a typed spatial/temporal/semantic triple over one canonical event stream, checks each aspect with its own monitoring specification, and fuses the verdicts through a four-valued algebra that carries provenance. The spatial aspect is interpreted over a weighted two-sorted location graph in which mission geometry and information-release topology are one object; we show that these spatial obligations are not in general subsumed by a first-order temporal specification. The past-time aspect runs on the unmodified MonPoly engine, which agrees with our reference monitor at every time point. Across two mission domains, casualty evacuation and contested sustainment, and one civil domain, composition under the precautionary blocking policy drives attack success to zero with no observed false positives and microsecond-scale per-event cost, while every single aspect and every pair leaves a substantial share of attacks succeeding. In a closed-loop experiment a policy-naive planner reaches a violating state in most unshielded missions and in none when shielded, and four refused episodes in five still recover to a compliant outcome.

[AI-144] AdaGuard: Enhancing Safety and Policy Compliance with Reasoning -Enabled LLM -As-A-Judge Guardrails

链接: https://arxiv.org/abs/2610.08923
作者: Melissa Kazemi Rad,Sihui Dai,Isha Slavin,Kushal Chawla,Mann Patel,Jian Ni,William M. Campbell,Stephen Rawls,Sambit Sahu
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Enterprise generative AI applications require robust safety mechanisms that can accommodate diverse risk postures, evolving policies, and varying latency constraints. Current guardrail solutions often suffer from rigidity, relying on fixed policy sets and offering limited transparency or reasoning flexibility. We present Adaguard, an adaptive LLM-as-a-Judge framework designed to address these challenges through dynamic policy enforcement and adaptive reasoning-budget allocation. Built using supervised fine-tuning (SFT) and reinforcement learning (GRPO), AdaGuard generalizes to user-defined safety and compliance policies at runtime without requiring frequent model updates. A core innovation of our approach is the ability to dynamically infer the complexity of input-policy pairs, allowing the model to switch between high-speed black-box inference and explainable, reasoning-enabled moderation. This flexibility enables developers to balance stringent latency requirements with the need for actionable transparency. This adaptive capability allows AdaGuard to rival other guardrail and frontier models several times its size, while its auto-reasoning mode recovers the accuracy of always-on reasoning at a fraction of the latency

[AI-145] Agent Plasticity: Measuring Self-Improvement Through Experience

链接: https://arxiv.org/abs/2610.08902
作者: Harman Singh,Anton Bakhtin,Rulin Shao,Gabriel Synnaeve,Ilia Kulikov,Rob Fergus,Sanjeev Arora,Kurt Keutzer,Jason Weston,Anuj Mahajan,Anirudh Goyal
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:AI agents increasingly operate in environments where they can diagnose failures and improve through experience, yet existing evaluations largely measure what an agent can do at a fixed point in time rather than how effectively it learns. Evaluating self-improvement requires answering three questions: does future performance improve and generalize beyond the interactions that enabled learning; how efficiently are new capabilities acquired; and where does the self-improvement process break down? To answer these questions, we study self-improvement in a controlled setting where agents amortize past experience into reusable artifacts that are inherited by future instances. At each checkpoint, we measure performance on training and held-out environment interactions while accounting for learning cost. We introduce agent plasticity, the efficiency with which an agent converts experience into gains in future held-out performance. Across multiple environments, frontier models exhibit sharply different improvement trajectories despite comparable opportunities to learn. Some achieve substantial and persistent gains, while others remain near or below their initial performance, and gains within the training regime often transfer only partially to out-of-distribution conditions. Endpoint capability and acquisition efficiency also diverge: the agent that ultimately performs best need not be the one that improves most efficiently. Tracing failures through the improvement loop further reveals different candidate bottlenecks. Agents with low plasticity often fail to reuse relevant artifacts, whereas more plastic agents may still fail despite reusing relevant artifacts, pointing to limitations in artifact quality, generalization, or application. Evaluating self-improving agents requires measuring not only what they can do, but how effectively they become better through experience.

[AI-146] Sequential Probabilistic Uncertainty Estimation for Parallel Multi-Agent Reasoning Systems

链接: https://arxiv.org/abs/2610.08901
作者: Tunyu Zhang,Zihao Zhao,Yusong Zhao,Haizhou Shi,Zhuohang Li,Haoxian Chen,Hao Wang,Dimitris N. Metaxas
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:LLM-based multi-agent systems (MAS) have attracted growing attention for improving reasoning through interaction among multiple agents. In this work, we focus on parallel multi-agent reasoning systems, where several agents solve the same problem over multiple rounds and aggregate their outputs into a final answer. Despite their strong reasoning performance, uncertainty estimation for such systems remains underexplored: the reliability of a MAS depends not only on individual generations, but also on how agents interact and evolve across rounds. We propose SAUCE (Sequential Agent Uncertainty through Consensus Evolution), a lightweight, training-free uncertainty estimator that formulates MAS uncertainty as sequential inference over a latent system-level belief. SAUCE aggregates round-level agreement and generation-uncertainty signals through a filtering-style update. Across five backbones, five benchmarks, and two MAS protocols, SAUCE improves misclassification detection, selective prediction, and calibration over a broad set of uncertainty estimation baselines, including standard log-likelihood-based methods and MAS-specific estimators.

[AI-147] Humanize: Judgement Engineering for Agent ic Coding

链接: https://arxiv.org/abs/2610.08900
作者: Sihao Liu,Ligeng Zhu,Zijian Zhang,Dongyun Zou,Zhengyang Zhang,Changye Li,Song Bian,Song Han,Tony Nowatzki
类目: Artificial Intelligence (cs.AI); Computers and Society (cs.CY)
备注:

点击查看摘要

Abstract:Agentic coding makes code generation cheap, but reliable completion remains difficult: the agent that writes the code is a weak judge of whether it is done. We present Humanize, a multi-agent orchestration workflow for agentic coding built around judgement engineering: explicit, mechanically enforced decisions at the boundaries between planning, implementation, review, and learning. A human approves a plan contract, a builder agent implements it in rounds, and a reviewer agent from another vendor decides completion; deterministic hooks, not a model, route work between these roles and enforce 72 mechanical gates. Viewed as a Markov chain over repository states, alternating builder and reviewer samples jointly from two models, so a defect survives only if both miss it. We study Humanize through its deployment, 118 public postmortems of real loops, and its applications. Over 68 versions in 108 days, it gathered 1,468 GitHub stars. Applications include a 567-file gem5 build-system migration under upstream review; Kernel Design Agents, which extend the loop with a kernel knowledge base and profiling feedback and placed in the top three of all three Full-Agent tracks of the MLSys 2026 FlashInfer contest; and, through Humanize Olympiad Agents (HOA), full scores in IOI 2026, IMO 2026, IPhO 2026, and IBO 2024, 418.5/437 in IChO 2026 (gold-medal). Humanize also achieves 672/672 on PutnamBench and ranks first (251/303) on Lean-Eval’s leaderboard even competiting with professional mathematicians. The postmortems show that independent review catches unsupported builder claims, but stopping remains a key weakness. In reports that separate rounds by phase, two thirds of rounds occurred after implementation was accepted. This evidence is observational, not a controlled comparison of workflows. Subjects: Artificial Intelligence (cs.AI); Computers and Society (cs.CY) Cite as: arXiv:2610.08900 [cs.AI] (or arXiv:2610.08900v1 [cs.AI] for this version) https://doi.org/10.48550/arXiv.2610.08900 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[AI-148] Contextualization of Third-Party Cloud Security Findings

链接: https://arxiv.org/abs/2610.08895
作者: Leon Goldberg,Gal Engelberg
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
备注: 11 pages, 1 figure, 2 tables

点击查看摘要

Abstract:Finding severity is the main driver of how security teams prioritize remediation. For third-party cloud security findings, that severity is static: the rule that raised the finding assigns it before the rule meets any environment, so it reflects the risk of the condition in general rather than the risk the finding poses to the concrete environment where it lives. Scoring standards define where environment-specific context belongs. How far that context changes finding severities in production, where the deciding evidence lies, and whether it holds against the live environment have not been measured. We address this gap with contextualization, re-deriving each finding’s severity from evidence in the environment where the finding lives. A deep research agent over a precomputed cross-signal asset graph investigates each finding against the resource’s state, its graph neighborhood, and other products’ signals, and returns an adjusted severity with an evidence trace. We evaluate it in a production field study of 9,967 vendor HIGH findings from two commercial cloud security platforms across eight real production environments, on three criteria: the faithfulness of the facts behind each verdict to the live environment, the dependence of each decision on context beyond the flagged resource, and the regularity of the reasoning. Three in four findings are re-graded, mostly downward, and the same rule often moves in opposite directions inside a single environment. About half of the decisive evidence lies beyond the flagged resource, and read-only probes of live infrastructure confirm the decisive fact for 99.4% of decided findings.

[AI-149] Physics-based Sphere Packing for Lagrangian Mesh Morphing

链接: https://arxiv.org/abs/2610.08880
作者: Jiong Lin,Hod Lipson
类目: Graphics (cs.GR); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:This paper studies tetrahedral meshes as the body representation for differentiable simulation and computational design. Fixed-connectivity meshes degrade under large morphs, while remeshing from scratch discards node correspondence. We present JamTet, a physics-based sphere-packing framework for volumetric meshing and morphing. We contribute (i) a GPU-parallel mesher combining octree-hierarchical packing with constrained Delaunay tetrahedralization, producing more uniform element volumes than TetGen and fTetWild; (ii) Lagrangian mesh morphing that preserves interior-node identities by re-equilibrating the same spheres within changing shapes and rebuilding the boundary and connectivity, remaining inversion-free where fixed-connectivity and TetSphere meshes invert; and (iii) a differentiable GPU simulator in JAX, with mass-spring edges and a volumetric Neo-Hookean term, integrated with mesh morphing in a design pipeline. In soft-robot morphology design experiments, interior-node gradients improve swimming fitness by 0.73-1.07 over a matched surface-only variant, while voxelized versions of the same designs yield 32-63% lower fitness. These results establish sphere packing as a practical volumetric mesh representation for gradient-based shape optimization. Code and media: under review.

[AI-150] An Empirical Study of Agent Skills Downstream Utility

链接: https://arxiv.org/abs/2610.08875
作者: Yu Cheng,Dehai Zhao,Zhongxin Liu,Qing Huang,Zhenchang Xing,Xiaoxue Ren
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Agent Skills package procedural guidance and resources for reuse, but a relevant Skill does not necessarily improve task performance. Existing studies characterize Skill content and evaluate downstream performance, yet provide limited explanations of how utility depends on content, execution configuration, and multi-Skill organization. We conduct an empirical study on 87 SkillsBench tasks, defining downstream utility as the pass-rate difference from No-Skill on the same tasks under the same model–harness configuration. We compare the same Skills across nine configurations, then examine alternative published Skills and organizations of fixed Skill sets under three selected configurations. We retrieve marketplace candidates from a curated corpus of 37,596 Skills. LLM-assisted analysis of content, execution traces, and final artifacts, followed by author review, relates provided support to actual use and task outcomes. The same Skills help some configurations and hurt others on 36.78% of tasks, with trajectories showing that recommended procedures can become an execution burden. Relevance rankings overlook more useful candidates. Within the evaluated candidate sets, reranking by support for required operations raises first-choice pass rates by 4.35–5.80 percentage points across the three configurations. We derive 17 authoring practices linking executable procedures to recovery, preservation of task requirements, and checks on final artifacts. Stage Plan and Dependency DAG outperform use order alone, with DAG’s additional benefits concentrated in tasks supplied with five or six Skills. These findings guide developers to assess usable operation support, allow procedure adaptation while preserving task requirements, and make artifact dependencies explicit when organizing Skills.

[AI-151] CredLeakBench: Evaluating Credential Leakage and Recovery in LLM Agents

链接: https://arxiv.org/abs/2610.08871
作者: Rafid Ahmed,Joseph Fioresi,Mubarak Shah,Yuzhang Shang
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Language model agents are increasingly deployed to automate everyday digital chores from managing emails and social media to handling banking and bills allowing users to step away from supervision. However, this capability also exposes sensitive information to phishing. Safe execution requires distinguishing malicious requests from genuine ones without simply refusing to act. Despite its practical importance, this problem remains underexplored and it is unclear whether current agents or existing defenses can achieve it. To study this problem, we first propose CredLeak-Bench, a comprehensive benchmark designed to evaluate how effectively and securely agents automate human workflows when confronted with phishing and identity verification. The benchmark covers both user-directed authentication and autonomous inbox monitoring, where agents are not explicitly instructed to log in. It systematically varies deceptive cues and pairs phishing scenarios with legitimate counterparts, enabling joint evaluation of information leakage and utility on genuine tasks. Within a sandboxed environment, leakage is measured through actual submissions of information rather than agents’ self-reported behavior. Our evaluation reveals that all tested models are vulnerable to leakage. Agents also disclose sensitive information during autonomous inbox monitoring, demonstrating that phishing can induce disclosure without a user request to authenticate. Furthermore, most evaluated mitigations that reduce leakage also impair performance on genuine tasks, exposing a security utility trade off in existing defenses. These findings show why reducing leakage alone is insufficient: effective defenses must prevent unauthorized disclosure while preserving legitimate task completion. CredLeak-Bench provides a controlled framework for measuring both objectives and evaluating progress toward secure, useful agents.

[AI-152] JoyAI-Voice 2.0: A Full-Continuous Autoregressive Speech Generation Model with Semantic-Acoustic Joint Representation

链接: https://arxiv.org/abs/2610.08834
作者: Yafeng Chen,Boya Dong,Yankun Huang,Hao Li,Jingdong Li,Xiangyu Liang,Hao Ni,Wenchao Wang,Yuxuan Wang,Zhangyu Xiao,Wei Deng,Nan Duan,Yu Gu,Wenhao Guan(Intern),Weisheng Han,Yabin Li,Yuan Liu,Jiaxin Ye(Intern),Fan Yu,Lin Zhu
类目: ound (cs.SD); Artificial Intelligence (cs.AI); Audio and Speech Processing (eess.AS)
备注:

点击查看摘要

Abstract:We present JoyAI-Voice~2.0, an end-to-end anthropomorphic speech generation model built upon a fully continuous, dual-encoder architecture. Raw speech is encoded into continuous latents and partitioned into patches. Each patch is decomposed by a semantic-acoustic dual encoder into a semantically purified representation and an acoustic representation, which are fused and jointly fed to a causal autoregressive Transformer for planning. The Transformer predicts the conditioning for the next patch, and a local diffusion Transformer renders its full latents for 48,kHz synthesis. The model is trained with a joint flow-matching and stop-prediction objective, followed by supervised fine-tuning and reinforcement learning with DiffusionNFT to improve model performance. It achieves the lowest average word error rate of 2.51% on Seed-TTS, a 14.9% relative reduction over the strongest baseline, and state-of-the-art attribute fidelity on InstructTTSEval in both Chinese and English, leading on 5 of 10 perceptual dimensions with the highest overall score of 0.893 on MDVD-Eval.

[AI-153] HydroSphere: A Framework for Governed Self-Healing Wastewater Infrastructure

链接: https://arxiv.org/abs/2610.08819
作者: Prabu,Fancy C,Suresh A,Srini Ramaswamy
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computers and Society (cs.CY); Emerging Technologies (cs.ET); Systems and Control (eess.SY)
备注: 16 pages, 5 tables and 9 figures - article to be submitted to a Journal / Conference

点击查看摘要

Abstract:Rapid industrialization and urban growth are increasing pressure on water quality and wastewater treatment systems, while conventional treatment plants often rely on static monitoring and control strategies that cannot easily adapt to changing pollutant conditions. This paper presents HydroSphere, a governed, data-driven framework for real-time water quality monitoring, forecasting, treatment optimization, and fault recovery. HydroSphere is evaluated using 2.82 million water-quality measurements collected between 1940 and 2023. The framework integrates three main components. First, a hybrid TCN-LSTM model performs multi-step forecasting across seven water-quality parameters, achieving an RMSE of 0.1417, MAE of 0.1047, and R2 of 0.3596. Second, the Adaptive Dosage Optimization Module uses PPO reinforcement learning to adjust chemical dosing, achieving a mean step reward of 1.059 compared with 1.017 for a fixed-dose baseline. The results also show that unconstrained reward optimization can lead to excessive dosing, demonstrating the need for explicit operational safeguards. Third, the SHADE anomaly detection module uses a deep autoencoder to identify sensor and process anomalies, achieving an F1 score of 0.651 under controlled fault injection. HydroSphere combines these capabilities with tiered governance, deterministic safety bounds, and human oversight to support safer and more adaptive water infrastructure. The framework provides a scalable foundation for intelligent wastewater management and supports the objectives of UN Sustainable Development Goals 6 and 13.

[AI-154] aming an End-to-End Autonomous Driving Policy for Urban Navigation of Quadruped Robots IROS2026

链接: https://arxiv.org/abs/2610.08812
作者: Joochan Kim,Chanuk Yang,Tackgeun You,Ziran Wang,Hwasup Lim
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI)
备注: Accepted to IROS 2026 Workshop on AI Meets Autonomy

点击查看摘要

Abstract:We present Go2-DrivoR, a goal-conditioned adaptation of the end-to-end autonomous driving trajectory planning framework DrivoR for urban navigation with quadrupedal robots. By conditioning trajectory generation on a local-frame subgoal through a goal token and adapting the vehicle-centric scoring formulation, the method extends DrivoR to short-horizon goal-conditioned local planning without redesigning its core decoders. Specifically, we redefine drivable-area compliance for sidewalk-oriented navigation and reformulate the original ego progress term as goal-conditioned ego progress. Trained exclusively on TartanGround simulation data, Go2-DrivoR improves waypoint-conditioned planning performance on unseen simulation environments and transfers zero-shot to open-loop real-world trajectory prediction.

[AI-155] KVFetch: Temporal Prefetching for the Missing Half of KV Cache Compression ICLR2027

链接: https://arxiv.org/abs/2610.08811
作者: Linfeng Dong
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: 21 pages, 7 figures. Submitted to ICLR 2027

点击查看摘要

Abstract:As context windows scale to tens or hundreds of thousands of tokens, KV cache compression has become essential for efficient LLM inference. Existing methods fall into three families: score-based eviction, summary compensation, and offload-and-recall. Yet all three decide what to keep or recall by content relevance to the current query. We show this shared design is structurally incomplete. A cache supports two access modes: associative lookup by content and sequential traversal by position; current compressors implement only the first. The gap matters in practice: retrieval-augmented generation, code completion, and structured-data extraction all require the model to reproduce identifiers, field values, or code tokens verbatim from the context. Under compression, content-based eviction retains the head of such a sequence but discards its continuation, causing verbatim copying to break irreversibly midway, a failure we call sequential forgetting. This failure resists better scoring, larger budgets, summary compensation, and dynamic re-scoring; it is the dominant source of remaining quality loss under compression. We propose KVFetch, a training-free, drop-in framework that opens a temporal recall channel for any score-based compressor. It demotes evicted candidates to a quantized cold tier, detects active copying through a monotone read pointer, and prefetches positional successors into fixed-size hot-tier slots without increasing attention cost. On RULER-16K under an iso-budget control, KVFetch recovers verbatim copying from 0.8 to 78.4 and raises the 13-task average by +8.4, with gains concentrating on tasks that require sequential access. On LongBench, where no task requires sequential access, the channel remains dormant and imposes no cost.

[AI-156] ransferability and operational reliability of a Prithvi crop classification foundation model under phenological and geographic shift across three continents

链接: https://arxiv.org/abs/2610.08810
作者: Venkatesh Kolluru,Rajat Shinde,Abdelhak Marouane,Caden Helbling,Deepak Shah,Othneil Drew,Srinivas Kolluru,Iksha Gurung,Manil Maskey,Rahul Ramachandran
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Fine-tuned geospatial foundation models (GeoFMs) pretrained on large satellite archives have been shown to improve crop classification accuracy and geographic transferability. However, their operational performance beyond the training distribution remains poorly characterized. We evaluated the out-of-distribution performance of a widely adopted GeoFM [Prithvi-EO-2.0] across 37 events in 12 countries on three continents and validated against regional reference products. Results indicated that the mean overall accuracy (OA) declined from 0.65 in the United States to 0.40 in Europe. Beyond accuracy metrics, we assessed five key aspects of model performance: whether model confidence indicates signal failure, sensitivity to observation windows, the effect of coarsening class schemes, and robustness to both band loss and cloud- and shadow-contamination. Accuracy collapsed when the observation window misaligned with local crop phenology, while deterministic confidence remained high. Expected calibration error increased for seven of eight paired events, and 12-51% of each affected scene was confidently mislabeled at near-zero precision. Monte Carlo dropout entropy registered the shift in all eight, indicating that much of the apparent cross-continent decline reflected phenological misalignment rather than spatial transfer. Two adjustments recovered accuracy without retraining. Consolidating 13 classes into 10, based on the model’s dominant confusions, raised the mean OA by 8.4 percentage points. Compressing the window toward near-real-time use preserved accuracy across a 45- to 90-day plateau, peaking near 75 days, though arms tighter than 30 days fell about 0.11 below that plateau. Fine-tuned crop GeoFMs therefore transfer usefully only where observation windows match local growing seasons. We translate these findings into operational guidance for the reliable deployment of the released model.

[AI-157] Beyond Baseline Severity: Temporal and Disease-Specific Predictors of Depression Outcomes Following Mindfulness Interventions

链接: https://arxiv.org/abs/2610.08809
作者: Muhammad Jawad Chowdhury,Sultanus Salehin,Akib Jayed Islam
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Depression severity among patients with chronic or acute medical conditions is influenced by a complex interaction of baseline psychological state, demographic characteristics, clinical context, and engagement with behavioral interventions. This paper presents an interpretable machine-learning analysis of a multi-center longitudinal clinical cohort to predict Beck Depression Inventory-II (BDI-II) scores at 12 and 24 weeks following mindfulness-based intervention participation. The study uses demographic variables, clinical condition information, hospital-center identifiers, baseline BDI-II scores, and therapy engagement measures to model short-term and long-term depression outcomes. Missing follow-up outcomes were addressed using a model-based stochastic imputation procedure to preserve the modest sample size while maintaining outcome variability. Five regression models were evaluated, spanning regularized linear regression and tree-based ensemble methods. Ridge Regression achieved the best 12-week performance with an RMSE of 5.186 and R^2 of 0.474, while LightGBM achieved the best 24-week performance with an RMSE of 5.038 and R^2 of 0.525. Beyond prediction accuracy, the analysis reveals three clinically relevant patterns: baseline severity remains the strongest overall predictor, short-term outcomes are more strongly associated with clinical and hospital context, and long-term outcomes show greater dependence on behavioral adherence and demographic factors. Disease-specific and hierarchical subgroup analyses further indicate that predictors differ substantially across and within clinical categories. These findings support the use of interpretable, context-aware modeling to inform personalized mental-health support following mindfulness-based interventions.

[AI-158] Accelerating Floating-Point Satisfiability Solving via Gradient Normalization

链接: https://arxiv.org/abs/2610.08808
作者: Yuanzhuo Zhang
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Satisfiability Modulo Theories (SMT) solvers are foundational to software verification, program analysis, and compiler testing, particularly over the theory of Quantifier-Free Floating-Point (QF_FP). While recent optimization-based SMT solvers have successfully applied gradient descent to continuous relaxations of logical formulas, they are fundamentally bottlenecked by gradient domination, a phenomenon where a small subset of difficult clauses hijacks the optimization trajectory, preventing the solver from satisfying the broader formula and trapping it in local minima. To overcome this, we present GradSAT, a novel framework that bridges optimization-based SMT solving with Multi-Task Learning (MTL). GradSAT reformulates the constraint satisfaction process by treating each SMT clause as an independent MTL task. By applying dynamic gradient normalization (GradNorm), GradSAT actively balances the gradient magnitudes across all clauses at runtime, systematically penalizing dominant gradients and accelerating lagging clauses to ensure uniform convergence. GradSAT implements this through a highly optimized, two-stage hybrid pipeline. First, a GPU-accelerated PyTorch backend leveraging symbolic compilation and operator fusion navigates the continuous relaxation to a high-quality basin. Second, the candidate assignment is handed off to a bit-precise local search engine to rapidly resolve the exact, rigorous assignment. By stabilizing the continuous search dynamics, GradSAT mitigates the brittleness of prior gradient-based solvers and provides a robust, highly parallelizable architecture for complex constraint solving. Subjects: Artificial Intelligence (cs.AI); Machine Learning (cs.LG) ACMclasses: D.2.4; F.3.1; F.4.1 Cite as: arXiv:2610.08808 [cs.AI] (or arXiv:2610.08808v1 [cs.AI] for this version) https://doi.org/10.48550/arXiv.2610.08808 Focus to learn more arXiv-issued DOI via DataCite

[AI-159] OSFP4: Joint Optimization of Diagonal Smoothing and Block Scales for NVFP4 Quantization

链接: https://arxiv.org/abs/2610.08231
作者: Neriah Ben David,Ori Meir,Or Ordentlich
类目: Artificial Intelligence (cs.AI); Information Theory (cs.IT)
备注:

点击查看摘要

Abstract:NVFP4 is an attractive datatype for large language model (LLM) inference, offering compact storage and native tensor-core acceleration. However, preserving accuracy using NVFP4 requires careful quantization. In this work we develop a novel quantization scheme called Optimized Smoothing and Scaling for NVFP4 (OSFP4). For each linear projection it uses a diagonal smoothing matrix whose entries are optimized to minimize the squared matrix-product quantization error under NVFP4, taking into account the rounding procedure that is used (either round-to-nearest, or GPTQ-style successive interference cancellation). This requires performing joint optimization on the smoothing entries as well as the block scales, which is facilitated by analyzing a multiplicative-dither FP4 quantizer instead of the fixed deterministic one. Experiments show that OSFP4 achieves the highest average accuracy among the evaluated competitors in the corresponding quantization settings, while retaining approximately 94-97% of vendor NVFP4 prefill throughput on the measured workloads. Our code is available in this https URL

[AI-160] LLM -Assisted Generation of Transparent Open-Source Multiphysics Models of Electrochemical Devices

链接: https://arxiv.org/abs/2610.10320
作者: Sebastian Castro,Maya F. Schuchert,Spencer A. McCluskey,Eric W. Lees,Justin C. Bui
类目: Chemical Physics (physics.chem-ph); Artificial Intelligence (cs.AI); Computational Physics (physics.comp-ph)
备注: 25 pages, 6 figures, Supplementary Information and implementation guide provided as ancillary files

点击查看摘要

Abstract:Multiphysics continuum models are powerful tools for studying electrochemical devices, enabling in silico reactor design and resolution of local pH, potential, and concentration fields that govern device performance but are difficult to measure experimentally. However, constructing such models requires substantial numerical expertise or reliance on proprietary software. Here, we show that frontier large language model agents can remove this implementation burden while keeping the underlying physics under researcher control. Using one-dimensional electrochemical CO2 reduction to CO in a porous gas diffusion electrode as a test case, we develop a machine-readable, human-specified modeling harness containing governing equations, parameters, numerical methods, logical build stages, and human-verifiable checkpoints. From this specification, the agent reproducibly constructs complete multiphysics models in open-source Julia. Independently built models, including fully autonomous agent-built models, agree with an equivalent COMSOL implementation to within 0.7% of the peak CO partial current density, and with one another to within 0.04%. Systematically planted errors demonstrate the importance of explicit specifications for reproducibility and reveal the agent’s capabilities and limitations in debugging model physics. This framework establishes a more transparent approach to multiphysics modeling in which physical descriptions and governing equations, rather than specialized code, become the primary inputs for computational model development.

[AI-161] Fast holographic inversion of superconducting domes

链接: https://arxiv.org/abs/2610.09938
作者: Sejin Kim
类目: High Energy Physics - Theory (hep-th); Artificial Intelligence (cs.AI)
备注: 25 pages, 4 figures

点击查看摘要

Abstract:A holographic superconductor whose scalar mass depends on the gauge field strength, M(\Fsq) , reproduces a superconducting dome for a suitable M , and recovering that M from a given dome has so far taken days for a single training run. We propose a new way of training this model, with which an inversion takes from about ten minutes to an hour. Training needs the gradient of the condition that fixes the critical temperature, which the earlier method obtains by finite differences, repeating the bulk integrations for every training parameter. Here that condition is obtained, without any fit, from two integrations started at the horizon and at the boundary, and its derivative with respect to M is an integral over the same two solutions, so the gradient needs no integration of its own. We use the speed to study the part of M that a dome cannot determine, on the interval between the value \Fsq takes at the horizon for the lowest doping and \Fsq=0 , at which M is the scalar mass M(0) that fixes the dimension of the dual operator. We hold the scalar mass at several values, which we call pinned masses, retrain everything else at each, and find that the reconstructions agree wherever the horizons of the dome reach, including the minima of M , and differ only on that interval. A rule that keeps the reconstruction with the simplest closed form recovers both the scalar mass and the mass function of a test dome. On Gaussian and double-Gaussian domes and on the measured phase diagrams of YBa _2 Cu _3 O _y and 2M-WS _2 , however, the pinned mass it keeps rests on ties or on narrow margins, so for these targets the scalar mass is left open. The dome thus constrains M where its horizons reach, and fixing the dimension of the dual operator needs a second observable.

[AI-162] An AI-assisted conditioning and geological interpretation workflow for usage in implicit geological modeling

链接: https://arxiv.org/abs/2610.09871
作者: Stefan Carpentier,Jan Diederik van Wees,Eva de Boever,Jan Niederau,Camille Chapeland,Suzanne Atkins,Boris Boullenger,Jens Wollenweber
类目: Geophysics (physics.geo-ph); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: 17 page, 19 figures

点击查看摘要

Abstract:Implicit modeling and Relative Geologic Time are geological modeling techniques that enable more efficient, faster, less biased and more reproducible modeling results. For optimal operation, these techniques require many well-constrained input data. In the framework of the Horizon Europe GO-Forward and MOOI WarmingUP GOO projects and to accelerate Implicit modeling, Machine Learning (ML) methods have been tested and implemented in a toolkit for the interpretation of (onshore) seismic data from the shallow to deep range (± 300 - 3500 m). The goal is to rapidly characterise this depth domain by efficient interpretation of horizons and faults in seismic data. The first step is to improve the signal by applying AI techniques like self-supervised and semi-supervised contrastive learning CNN’s for noise reduction and interpolation. Next, horizons and faults are interpreted with minimal use of human-generated training data by using (semi-) self-supervised methods. The resulting developed toolkit supports the application of the implemented algorithms in an efficient workflow. As a first demonstration, the top of the Dutch Maassluis Formation has been interpreted in the Leeuwarden and Waalwijk 3D seismic cubes. Overall, this study demonstrates that AI-assisted interpretation workflows have reached a level of maturity that allows their integration into applied geological modeling and decision-making.

[AI-163] Artificial intelligence pathways from weather to climate

链接: https://arxiv.org/abs/2610.09770
作者: Tom Beucler,J. David Neelin,Hui Su,Shivanshi Asthana,Chris Bretherton,Will Chapman,Costa Christopoulos,Spencer K. Clark,Aditya Grover,Ignacio Lopez-Gomez,Tapio Schneider,Adam Subel,Oliver Watt-Meyer
类目: Atmospheric and Oceanic Physics (physics.ao-ph); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: 33 pages, 10 figures. Submitted to “Science Advances”

点击查看摘要

Abstract:Deep learning has made rapid advances in weather forecasting: autoregressive models trained on atmospheric reanalyses now rival dynamical models across nowcasting, medium-range, and subseasonal-to-seasonal lead times, producing well-calibrated ensemble forecasts at reduced cost. We review these advances and consider their extension to climate horizons, where the challenge shifts from initial-condition skill to producing reliable statistical responses under altered forcings. AI-powered climate prediction systems must produce credible forced responses to drivers (e.g., greenhouse gases, land-use change) typically outside the observed record. We propose two minimum requirements for AI in climate modeling: (i) external forcing agents must enter explicitly enough to support interventions in which they vary independently; and (ii) robustness must be stress-tested in out-of-distribution regimes, including extremes and counterfactual trajectories. Using leading AI autoregressive emulators and hybrid physics-AI models, we identify development and coupling challenges. Comparing the reported throughput of these models with that of GPU-ported dynamical models highlights how AI can reduce time-to-solution by advancing only the target variables at the required resolution and using longer time steps, rather than integrating a full high-frequency, multivariate state. Diverse AI downscaling strategies can partially substitute for explicit fine-scale resolution, paving the way toward inexpensive local hazard assessment across prediction horizons.

[AI-164] Real-world application of deep learning in large-scale seismic interference attenuation: A case study in the Camie field of Angola

链接: https://arxiv.org/abs/2610.09663
作者: Jing Sun,Song Hou
类目: Geophysics (physics.geo-ph); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:In marine seismic acquisition, seismic interference (SI) occurs when energy from nearby external seismic source(s) is captured. It typically appears as coherent noise with linear or non-linear movement and varying amplitudes across different sail lines. SI is commonly observed and poses a challenge for seismic data processing. We present a case history of a previously proposed deep neural network (DNN)-based workflow applied for SI attenuation across a marine seismic block in the Camie Field of Angola. This field survey covers over 345 km2 and is marked by the challenge of multiple SI types. The employed DNN-based workflow performs SI attenuation in the common shot domain based on a supervised learning framework: a small subset of the SI-contaminated data was first processed by a conventional geophysical algorithm to obtain an estimate of the SI noise, which was then manually blended with the SI-free common shot gathers from the same survey to generate the training pairs. To ensure signal fidelity, several techniques were applied to improve the DNN’s performance. A key highlight of this case history is its scale: this represents a real-world, large-scale processing project and we present a comprehensive comparison of the DNN-based workflow with the conventional geophysical algorithm across the entire survey block, focusing on both processing quality and processing time. The results demonstrate the outstanding performance of the employed DNN-based workflow, which achieved higher SI removal accuracy, with less signal leakage and more complete SI removal. The promising results of this application also open up possibilities for integrating deep learning into other seismic denoising tasks. In addition, we discuss the limitations of this case history, aiming to provide insights for future research and applications in the field.

[AI-165] CircuitATLAS: Agent ic reasoning over a systems neuroscience knowledge graph for target discovery in circuitopathies

链接: https://arxiv.org/abs/2610.09643
作者: Gabriel Ocana-Santero,Marko Tvrdic
类目: Quantitative Methods (q-bio.QM); Artificial Intelligence (cs.AI); Neurons and Cognition (q-bio.NC)
备注: 22 pages, 5 figures, 4 tables

点击查看摘要

Abstract:Drug discovery for neurological disease has traditionally centered on the molecules altered by disease. But the molecules that cause pathology are not necessarily the best points from which to reverse it. Here, we ask which otherwise unaltered molecular control points can be engaged to restore pathological neural circuits toward functional states. We present CircuitATLAS, a provenance-grounded systems-neuroscience knowledge graph and agentic framework for target discovery in circuitopathies. It structures literature-derived relationships across diseases, phenotypes, electrophysiology, circuits, brain regions, cell types and molecular effectors, while deliberately excluding direct disease-gene and disease-protein edges to reduce shortcut reasoning. The graph contains 3.83 million nodes and 7.66 million edges, including 5.31 million LLM-extracted relations, and incorporates structured datasets such as the Human Cell Atlas and new multimodal in vivo measurements. We then introduce an agentic workflow that reasons from measurable disease phenotypes through their circuit and cellular substrates to molecular interventions, therapeutic feasibility and clinical constraints. Finally, we introduce a human-governed in vivo lab-in-the-loop linking hypothesis generation to experimental iteration. Within this framework an agent nominated ATP1A3, the neuronal alpha3 Na+/K±ATPase, as a control point on cortical excitability; interneuron-restricted expression of ATP1A3 abolished the beta- and gamma-band response to a focal 4-aminopyridine challenge in vivo, and the validated target was then carried into a structure-guided small-molecule campaign terminating in a defined assay to resolve the direction of modulation. CircuitATLAS thus provides a framework for discovering therapeutics based not only on what is molecularly disrupted in disease, but on what can be controlled to restore circuit function.

[AI-166] Constrained Diffusion for Data-Scarce Orbital Monte Carlo in Constellation Tasking

链接: https://arxiv.org/abs/2610.09323
作者: Omar Ramadan,Sam Siavoshian,Amir Kashif Saeed,Benjamin A. Johnson,Amin M. E.-A. Diab,Benjamin M. Rodriguez
类目: Instrumentation and Methods for Astrophysics (astro-ph.IM); Earth and Planetary Astrophysics (astro-ph.EP); Artificial Intelligence (cs.AI)
备注: 16 pages. Submitted to the 2027 IEEE Aerospace Conference; under review

点击查看摘要

Abstract:Constellation Monte Carlo results depend on the orbital population used to evaluate a tasking policy. With scarce reference trajectories, replay limits geometric diversity, while independent orbital-element jitter can violate physical constraints. We study constrained diffusion for orbital-population augmentation. A force-conditioned diffusion model learns a 13-dimensional orbital prior, recovering semimajor axis from perigee altitude and eccentricity; Basilisk propagates each sample under one of five force-model tiers. Using 800 reference trajectories, we compare diffusion with jittered bootstrap, per-tier Gaussian mixtures, and a conditional variational autoencoder, and evaluate distributional fidelity, support shift, classical astrodynamics diagnostics, and 4,000 paired GoDSAT-compatible campaigns. The largest diffusion model achieves held-out trajectory MMD of 0.0171 +/- 0.0239, similar to bootstrap (0.0170) and the mixture (0.0198), but samples farther from training priors (median nearest-training distance 2.7 versus 0.10 standardized units). All generators fail to match the shifted-blind population (classifier AUC 0.991-0.998). Endpoint-conditioned samples satisfy Lambert boundaries but have greater interior error than the matched Lambert reference (29.2 versus 5.7 km mean RMSE). Residual diffusion improves selected sparse forecasts and catalog-mode recall but does not outperform classical estimators on custody ranking. In a fixed 16-satellite configuration, diffusion yields similar mean custody to replay and bootstrap, while the force-tier mixture shifts custody by about 5.5 percentage points. Constrained diffusion supports local, in-support augmentation but cannot replace orbital dynamics or serve as an operational posterior. Orbital-population construction is a consequential source of uncertainty in constellation analysis.

[AI-167] Careful Judge: Safe and Efficient Human-AI Collaborative Decision Making

链接: https://arxiv.org/abs/2610.09043
作者: Chenyu Zhang,Rachel Luo,Boyi Li,Anjali Parashar,Marco Pavone,Apoorva Sharma
类目: Machine Learning (stat.ML); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:In human-AI collaborative decision making, human review can prevent unsafe AI decisions, but each human judgment is costly. Treating human intervention after AI abstention as a one-off fallback misses the opportunity to improve future AI decisions for greater automation, yet AI adaptively learning from selectively queried human feedback breaks safety guardrails calibrated for old models. We approach this challenge with CARE—calibrated adaptive rectification and escalation—an end-to-end pipeline that combines AI models and human reviewers to guarantee safe, human-aligned decisions, while continuously learning from human feedback to achieve greater automation with fewer human queries. CARE is principled, general, modular, and works with any black-box AI model. Our novel adaptive calibration module guarantees risk control at every time step for any rectification module. We further show how CARE improves query efficiency when the AI model is well trained and the human-AI misalignment has a clear structure. Experiments on four safety-critical real-world datasets spanning driving, language, and robotics demonstrate that CARE achieves human-aligned decisions while reducing human queries by 25-81% relative to baselines.

[AI-168] A Shortcut to Structure in AlphaFold 3

链接: https://arxiv.org/abs/2610.08937
作者: Jonathan Feldman,Jeffrey Skolnick
类目: Biomolecules (q-bio.BM); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:AlphaFold 3 predicts protein structures with remarkable accuracy, yet how structural information emerges within the model remains poorly understood. Here, through causal interventions on internal representations and direct probing of every Pairformer block, we trace the formation of global protein geometry and identify the multiple sequence alignment (MSA) as a structural shortcut to the fold. Removing the MSA largely preserves local secondary structure while disrupting the long-range relationships that define global topology. Restoring the MSA-enriched pair representation at only forty residues recovers most of this lost organization, including at pairs never directly modified. This contribution depends on the detailed direction of the MSA module’s output rather than its magnitude. The Pairformer rapidly converts this signal into global geometry: the final fold becomes recoverable by approximately block 9 of 48 for a majority of proteins, roughly twenty-seven blocks before the model’s decoder can render it, whereas without the MSA it remains inaccessible for most proteins throughout the pass. Which homologs are supplied shapes this trajectory more strongly than which query is supplied; it persists for a designed query that never evolved but collapses for a shuffled sequence. Most importantly, an alignment built for a different protein that shares the fold, supplied only at the structurally corresponding columns, raises the median TM-score against experiment from 0.44 to 0.72, while the same alignment shifted a few residues along the chain performs worse than supplying no alignment at all. What AlphaFold 3 reads from an alignment is therefore a description of the fold itself, transferable between proteins that share one, rather than the query’s own evolutionary history. This explains both its accuracy and the limits of what it has solved.

[AI-169] Hybrid: The Bridge between PDE Models and Deep Learning for Gamma Noise Removal

链接: https://arxiv.org/abs/2610.08892
作者: Mahipal Jetta,Sujato Dutta
类目: Image and Video Processing (eess.IV); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Multiplicative gamma noise is one of the dominant noise factors in Synthetic Aperture Radar (SAR) and medical ultrasound images. They are dependent on pixel level noises due to which they are highly varying across the image and harder to handle as compared to additive noise. The denoising methods to address this noise currently include classical partial differential equation (PDE) methods which are good for interpretation but lack the restoration ability as compared to the state-of-the-art models while deep convolutional networks like DnCNN achieve high performance but at the cost of transparency due to which practitioners are skeptical to use them in high-risk critical fields like medicine. This paper presents Hybrid++, a novel trainable nonlinear reaction-diffusion architecture that addresses both the concerns - staying interpretable while offering performance close to huge black-box models. It combines a fully learnable PDE initialization with a 3-stage reaction-diffusion network having 64-channel multiscale filter banks, 4-layer Squeeze-and-Excitation attention-based influence functions and a 64-dimensional noise level embedding. It uses a two-phase training strategy, stage-wise optimization followed by joint end-to-end refinement which enables co-adaptation of all learnable parameters. On the FoE benchmark, Hybrid++ substantially improves over classical PDE, BM3D and the original TNRD baselines. In the severe-noise setting L=1, it comes within 0.23 dB PSNR of a separately trained DnCNN while using only about 8% of its parameters. We therefore position Hybrid++ not as a universal state-of-the-art image restoration backbone, but as a compact, physically structured reaction-diffusion model for multiplicative gamma noise.

机器学习

[LG-0] Distilling Graph Geometry: Knowledge Gap from GNNs to MLPs

链接: https://arxiv.org/abs/2610.10520
作者: Zhewei Chen,Hao Zhu,Jiaojiao Jiang,Ahad N. Zehmakan
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:GNN-to-MLP distillation aims to retain the predictive accuracy of a message-passing teacher while deploying a graph-free MLP at inference. Existing methods mainly transfer node-wise predictions or use confidence-based reweighting, but they do not specify where the student should preserve the teacher’s graph-induced geometry. We show that this omission leads to two spectral failure modes in the student’s representation space. On sparse graphs, the student suffers from spectral underfit, missing high-energy teacher directions concentrated near boundary regions. On dense graphs, it suffers from spectral overfit, retaining spurious directions that the teacher has collapsed through aggregation. Motivated by an energy-weighted teacher-student alignment objective, we propose Graph Geometry-aware MLP (G^2MLP), a training-time distillation framework guided by Ollivier-Ricci curvature. Curvature identifies where the two spectral errors concentrate and is used to allocate supervision between prediction-level and representation-level alignment. The deployed model remains a standard MLP and requires no graph access at inference. Across node-classification benchmarks, G^2MLP consistently improves over graph-free distillation baselines, reduces the teacher-student rank gap in both regimes, and transfers without architectural changes to Graph Transformer teachers and link prediction.

[LG-1] Why Forget-Only Unlearning Needs Memorization

链接: https://arxiv.org/abs/2610.10519
作者: Luka Radić,Vikrant Singhal,Amartya Sanyal
类目: Machine Learning (cs.LG); Information Theory (cs.IT); Machine Learning (stat.ML)
*备注:

点击查看摘要

Abstract:Machine unlearning asks for a deletion algorithm whose output is close to retraining from scratch without the selected forget examples. In this work, we study forget-only unlearning, where the deletion algorithm receives only the trained model and the examples to forget, with no retained data or extra training information. We ask whether forget-only unlearning is always possible. We first show that this depends on the learning method: different datasets can produce the same trained model but require very different outputs after the same examples are removed. Using this observation, we derive lower bounds on how accurately unlearning can match retraining and instantiate them for several standard learning algorithms. We then ask what must be true when forget-only unlearning succeeds. To this end, we derive lower bounds on what an algorithm must memorize about the training data to handle arbitrary deletion requests. For simple threshold learners, the required information can be as large as the entire dataset, even though ordinary training keeps only one boundary point. Overall, our results show that information discarded during ordinary learning may be needed later for deletion, so models designed for forget-only unlearning may need to retain more information than standard training does.

[LG-2] Oracle-Efficient and Parameter-Free Agnostic Smoothed Online Learning

链接: https://arxiv.org/abs/2610.10499
作者: Sasha Voitovych,Adam Block,Alexander Rakhlin,Abhishek Shetty
类目: Machine Learning (cs.LG); Machine Learning (stat.ML)
*备注:

点击查看摘要

Abstract:Online learning is an attractive framework in many domains because it permits well-defined learning even when data are dependent or chosen adversarially. This generality, however, comes at a steep price, introducing significant statistical and computational barriers. Recently, smoothed online learning has emerged as a promising framework that interpolates between the fully adversarial and fully stochastic settings by assuming that the conditional law of each covariate has density at most 1/\sigma with respect to some fixed base measure \mu , and it is known to match the statistical and computational guarantees of classical learning while still allowing for much of the flexibility of online learning. However, existing oracle-efficient algorithms require either (i) sampling access to the base measure \mu or (ii) labels that are perfectly predicted by a fixed hypothesis. Both assumptions limit the applicability of these algorithms, in contrast to statistical learning, where empirical risk minimization (ERM) learns efficiently in the agnostic setting without any knowledge of the data distribution. We show that neither assumption is necessary, giving the first oracle-efficient algorithm that achieves sublinear regret in the agnostic setting without knowledge of \mu . Our algorithm, based on Gaussian Follow-The-Perturbed-Leader, is parameter-free: it requires no knowledge of \mu , the smoothing parameter \sigma , or the horizon T , and it achieves regret \widetilde O(d\sqrtT/\sigma) for binary classes of VC dimension d with a single call to an ERM oracle per round, which is optimal up to a \sqrtd factor. En route to establishing the regret bound, we introduce several new techniques that may be of independent interest.

[LG-3] Evolutionary Architecture Search for Chlorophyll-a Prediction in Lakes using Sentinel-2

链接: https://arxiv.org/abs/2610.10496
作者: Kursat Komurcu,Linas Petkevicius
类目: Neural and Evolutionary Computing (cs.NE); Machine Learning (cs.LG)
*备注: Accepted at AutoML4EO 2026 (non-archival AutoML conference workshop). 4 pages + references. this https URL

点击查看摘要

Abstract:Small tabular datasets with expert-designed spectral features are the norm in operational Earth observation, and the networks applied to them are typically hand-designed. We revisit one such published model – a Sentinel-2 algal bloom classifier – and ask what architecture search adds, holding the task, the features and the lake-level train/test split of the original study fixed. Searching an extended multilayer-perceptron space with regularized evolution, and selecting on inner-cross-validation AUC only, we find networks that improve held-out AUC from 0.790 to 0.820 and accuracy from 0.733 to 0.748 while using 409 trainable parameters, 26 times fewer than the strongest hand-designed reference. The search converges on a consistent recipe – a single narrow layer, RMS normalisation, \tanh activation, step-decayed RMSprop and weight averaging – that a practitioner would be unlikely to reach by default. At 1.6,kB the resulting model is small enough to serve as an onboard screening trigger, which is the setting that motivates the work. Code: this https URL. Comments: Accepted at AutoML4EO 2026 (non-archival AutoML conference workshop). 4 pages + references. this https URL Subjects: Neural and Evolutionary Computing (cs.NE); Machine Learning (cs.LG) Cite as: arXiv:2610.10496 [cs.NE] (or arXiv:2610.10496v1 [cs.NE] for this version) https://doi.org/10.48550/arXiv.2610.10496 Focus to learn more arXiv-issued DOI via DataCite (pending registration) Submission history From: Linas Petkevičius [view email] [v1] Wed, 7 Oct 2026 17:47:17 UTC (12 KB) Full-text links: Access Paper: View a PDF of the paper titled Evolutionary Architecture Search for Chlorophyll- a Prediction in Lakes using Sentinel-2, by Kursat Komurcu and 1 other authorsView PDFHTML (experimental)TeX Source view license Additional Features Audio Summary Current browse context: cs.NE prev | next new | recent | 2026-10 Change to browse by: cs cs.LG References Citations NASA ADSGoogle Scholar Semantic Scholar export BibTeX citation Loading… BibTeX formatted citation loading… Data provided by: Bookmark checked="checked"class=“labs-tab-input”> Bibliographic Tools Bibliographic and Citation Tools Bibliographic Explorer Toggle Bibliographic Explorer (What is the Explorer?) Connected Papers Toggle Connected Papers (What is Connected Papers?) Litmaps Toggle Litmaps (What is Litmaps?) scite.ai Toggle scite Smart Citations (What are Smart Citations?) Code, Data, Media Code, Data and Media Associated with this Article alphaXiv Toggle alphaXiv (What is alphaXiv?) Links to Code Toggle CatalyzeX Code Finder for Papers (What is CatalyzeX?) DagsHub Toggle DagsHub (What is DagsHub?) GotitPub Toggle Gotit.pub (What is GotitPub?) Huggingface Toggle Hugging Face (What is Huggingface?) ScienceCast Toggle ScienceCast (What is ScienceCast?) Demos Demos Replicate Toggle Replicate (What is Replicate?) Spaces Toggle Hugging Face Spaces (What is Spaces?) Spaces Toggle TXYZ.AI (What is TXYZ.AI?) Related Papers Recommenders and Search Tools Link to Influence Flower Influence Flower (What are Influence Flowers?) Core recommender toggle CORE Recommender (What is CORE?) Author Venue Institution Topic About arXivLabs arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv’s community? Learn more about arXivLabs. Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?) mathjaxToggle(); We gratefully acknowledge support from our major funders, member institutions, , and all contributors. About Help Contact Subscribe Copyright Privacy Accessibility Operational Status (opens in new tab) Major funding support from

[LG-4] NeuralBES: A Differentiable Control-Aware Emulator for Scalable Building Energy Modeling

链接: https://arxiv.org/abs/2610.10459
作者: Ting-Yu Dai,Takuya Kurihana,Wing Yee Au,Hon Yung Wong
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Demand-side flexibility i.e. forecasting, shifting, and curtailing residential energy loads, depends on thermal models trusted across millions of heterogeneous buildings. Existing tools force a hard tradeoff: high-fidelity physics simulators such as EnergyPlus are accurate but sequential and require per-building calibration, while purely data-driven sequence models scale but abandon the physical structure that makes their predictions trustworthy. We introduce NeuralBES (Building Energy Simulation), a differentiable emulator that resolves this tradeoff by parameterizing a resistance–capacitance (RC) based thermal model with a shared neural encoder: static building metadata such as floor area, vintage, and HVAC type is mapped to physically bounded capacitances, conductances, and equipment coefficients, which become the coefficients of a scalar linear recurrence solved via a log-space parallel scan, and a predictor–corrector loop closes the thermostat–temperature nonlinearity while preserving full-horizon gradient flow. Trained on the ResStock dataset across three climate zones, NeuralBES handles heterogeneous building archetypes, vintages, and climate zones within a single trained encoder, while black-box baselines produce statistically plausible but physically inconsistent trajectories. On the annual full-year rollout, NeuralBES is the only data-conditioned model that is simultaneously physics-valid and accurate to within 4 MAPE points of the strongest raw-error baseline, while operating at roughly an order of magnitude fewer parameters than the transformer and recurrent baselines; among physics-valid baselines at parameter parity it more than halves the MAPE of the grey-box RC alternative. Subjects: Machine Learning (cs.LG) Cite as: arXiv:2610.10459 [cs.LG] (or arXiv:2610.10459v1 [cs.LG] for this version) https://doi.org/10.48550/arXiv.2610.10459 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[LG-5] Seq-Flow: Efficient Probabilistic Forecasting with Self-Rollout Error Control

链接: https://arxiv.org/abs/2610.10440
作者: Yinan Huang,Shitij Govil,Bo Dai,Pan Li
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Many scientific forecasting tasks require updating a distribution over future trajectories as new observations arrive. Conventional diffusion and flow models generate each forecast from Gaussian noise, often at the cost of many sampling steps. Warm-start methods reuse earlier predictions to reduce this cost, but their models are not trained to perform the forecast update itself, which can compromise quality under few-step sampling. In this work, we introduce Seq-Flow, a conditional flow model whose ODE transports samples from the previous forecast distribution to the updated one. Because successive forecasts often differ only modestly, this transport starts from an informative distribution and can produce accurate updates with few flow evaluations. Recursive reuse also creates a challenge: errors in one forecast become errors in the initial states of subsequent flows. We address this with self-rollout training, in which a moving average copy of the model generates forecasts that initialize later training updates. Unlike self-forcing methods, which reuse generated outputs as conditioning context, Seq-Flow reuses them as the source of the next flow. Experiments On particle-accelerator beam spill forecasting show Seq-Flow reduces CRPS by 65% under a few-NFE sampling budget, while remaining competitive with strong baselines on fluid-dynamics forecasting tasks. Although trained on self-rollouts of at most four updates, Seq-Flow remains accurate over more than 400 consecutive updates. Our code is available at this https URL.

[LG-6] Steerspeech: Activation Steering For Emotion Control In Generated Speech ICASSP2027

链接: https://arxiv.org/abs/2610.10415
作者: Afsara Benazir,Darius Pétermann,Felix Xiaozhu Lin,Salar Rahili
类目: ound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
*备注: Under review at IEEE ICASSP 2027. This work has been submitted to the IEEE for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible

点击查看摘要

Abstract:Pretrained text-to-speech (TTS) models can generate expressive speech, but reliable inference-time emotion control remains challenging: prompts and reference audio offer coarse, inconsistent control, whereas specialized conditioning and model adaptation require costly training. We present SteerSpeech, a lightweight activation-steering framework that controls emotion by injecting steering vectors into hidden activations. For each target emotion we train a lightweight low-rank transform, using a multi-expert objective that encourages monotonic emotion control while preserving speaker identity and linguistic content, constraining steering drift, and keeping the TTS backbone frozen. To optimize through discrete speech tokens, we introduce a two-pass generation-and-replay pipeline using a straight-through estimator to backpropagate expert supervision through sampled tokens. At inference, a target-emotion steering direction is optimized with its respective transform and injected into the base TTS model. Objective and subjective evaluations with Qwen3-TTS across seen, unseen, and accented speakers show stronger continuous emotion control with limited speaker and content degradation. SteerSpeech achieves 1.08x-7.12x baseline target-emotion scores and for a representative emotion subjectively, it receives 78.1%-96.8% intensity preference and 1.43x-1.46x speaker-identity preservation at high steering strengths.

[LG-7] RobotWorld: Benchmarking Multimodal Agents for Robot Use Across Diverse Tasks and Embodiments

链接: https://arxiv.org/abs/2610.10409
作者: Zhiqin Yang,Chenxin Li,Xiaomeng Hu,Yibin Liu,Weidong Huang,Jiankai Sun,Haitao Li,Zijian Wu,Yuzhi Huang,Fanding Huang,Hanwen Sun,Jiashun Liu,Jingqi Tong,Mingxin Huang,Shaoli Hu,Shijue Huang,Tianyi Bai,Xinyuan Wang,Yunlong Lin,Zhengyang Tang,Zhexin Zhang,Zhuo Chen,Xierui Song,Juntao Dai,Boyuan Chen,Jiaming Ji,Fangneng Zhan,Mengkang Hu,Wei Xue,Yonggang Zhang,Han Hu,Tsung-Yi Ho,Yike Guo
类目: Robotics (cs.RO); Machine Learning (cs.LG)
*备注: 62 pages, 25 figures

点击查看摘要

Abstract:General-purpose agents increasingly write code, use tools, and complete complex digital tasks, raising the question of how far these capabilities carry into the physical world. To investigate this, we introduce RobotWorld, a challenging simulation testbed for robot use: turning instructions and observations into physical task execution through robot interfaces. Its 84 tasks span manipulation, mobile manipulation, locomotion, driving, and aerial control, with explicit interaction budgets and executable success checks. By analysing task outcomes alongside execution traces, we identify both the capabilities that transfer and the gaps that prevent reliable completion. Furthermore, we find that current agents can construct sophisticated perception and control workflows, including image segmentation, camera calibration, spatial estimation, and dynamics-based computation. These capabilities, however, do not consistently compose into successful behaviour: agents lose task-relevant object states despite reaching commanded poses, fail to correct ineffective actions, recover too late, or mistake unfinished tasks for completion. This uneven transfer also differs across models: Astra succeeds more often on spatial and constrained-contact goals, whereas Opus 5.5 succeeds more often on continuous-balance and timed-interaction goals. By linking these outcomes to execution behaviour, RobotWorld provides both a rigorous proving ground and an empirical account of the remaining capability gaps, thereby establishing concrete targets for training and designing more reliable physical-world agents.

[LG-8] Cross-Domain Pretraining for Steady-State Neural CFD Surrogates

链接: https://arxiv.org/abs/2610.10398
作者: Anthony Zhou,Amir Barati Farimani,Shirley Ho,Rudy Morel
类目: Machine Learning (cs.LG)
*备注: 43 pages, 25 figures

点击查看摘要

Abstract:Neural surrogates for computational fluid dynamics (CFD) have the potential to greatly enhance engineering innovation through accelerating simulation. However, the primary limitation for neural surrogates is the lack of generalization to geometries and applications beyond the training set, which is significant given the diversity of engineering scenarios. Currently, this is addressed by generating a new dataset for a specific application; however, this requires running costly numerical solvers. In this work, we take a step toward addressing this by studying neural surrogates trained across different geometries, boundary conditions, and fidelities. We find that cross-domain pretraining improves zero- and few-shot performance on held-out datasets relative to both training from scratch and transferring from domain-specific experts. In particular, finetuning a pretrained, cross-domain model can achieve 2-3x lower errors at the same sample size and use 8x fewer samples to achieve the same error, compared to training from scratch. This benefit is architecture agnostic and improves with model size and pretraining dataset diversity. Furthermore, we study how and why cross-domain pretraining works in CFD surrogates, and find that simply pooling steady-state datasets is both sufficient and effective. Given the high cost of generating CFD data, leveraging existing datasets through cross-domain pretraining will likely be a valuable strategy as future surrogates expand to tackle new problems and use cases.

[LG-9] Executing Causal Structure Learning with Linear-Attention Transformers

链接: https://arxiv.org/abs/2610.10395
作者: Amartya Roy,Sayar Karmakar
类目: Machine Learning (cs.LG)
*备注: 32 pages, 8 Figures

点击查看摘要

Abstract:Transformers can execute algorithms on data given in their input. We ask whether they can do the same for causal discovery. We study a standard continuous method that repeatedly updates a candidate causal graph while enforcing acyclicity. We explicitly construct a fixed-weight transformer whose forward pass exactly reproduces one update of this method, so repeated blocks reproduce its optimization trajectory. The transformer carries the current graph and the algorithm’s multiplier between updates. We show that retaining the multiplier is essential for exact execution, since different multiplier values can lead to different next updates. We also give conditions under which, within a fixed stage, the number of updates needed to reach a target accuracy can be computed in advance and rounding errors stay bounded as depth grows. Experiments show that the constructed block agrees with a reference update to floating-point precision, while arithmetic replay on synthetic data and seven published benchmark network topologies inherits the reference solver’s successes and failures. This separates accurate algorithm execution from accurate causal recovery. In contrast, the ordinary attention models tested under our training budgets do not reliably execute the update or transfer to larger graphs. Whether gradient training can learn an executor in the architecture class of the construction remains open.

[LG-10] Kernel Autoresearch for Open-Ended Model Discovery

链接: https://arxiv.org/abs/2610.10394
作者: Richard Cornelius Suwandi,Feng Yin,Kevin Murphy
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Kernels encode the inductive bias of a wide range of machine learning models, yet automated kernel design faces a fundamental dilemma. A fixed grammar of base kernels and operators guarantees validity but limits the search to structures expressible by those building blocks. Conversely, unrestricted programs remove this limitation but no longer guarantee validity. In our stress tests, 22-58% of LLM-generated kernels that pass numerical checks on random inputs fail when evaluated at different scales or dimensions. We propose Kernel Autoresearch (Kernaut), which treats kernel design as open-ended model discovery. Coding agents write kernels as programs, while construction contracts ensure that every accepted kernel is valid. A quality-diversity archive retains high-performing kernels with distinct behaviors, and novelty screening steers agents toward functionally new candidates. Our experiments demonstrate that the discovered kernels encode reusable inductive biases that generalize to unseen tasks. On held-out black-box optimization families, a discovered kernel outperforms a meta-learned deep kernel trained on the same episodes. Furthermore, kernels discovered from ten enzyme-kinetic rate laws achieve lower error than tuned ARD and deep kernel baselines on five unseen mechanisms. The discovered kernels are also interpretable programs that human researchers can refine: a human-refined version of one further reduces the held-out predictive error by 5.7% and optimization regret by 7.8%.

[LG-11] OrBIT: Structure-Guided Embedding Compression

链接: https://arxiv.org/abs/2610.10385
作者: Yunied Puig,Amit Kumar Jaiswal
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Embedding tables are among the largest components of modern language models. Most compression methods fix a coding geometry such as coordinate blocks, low-rank subspaces, or unrestricted codebooks, and optimize within it. We instead ask whether the coding geometry can itself be discovered. We introduce \emphOrBIT, a structure-guided embedding compression framework that learns reusable local geometry from orbit dynamics and uses it to constrain a small set of shared codewords. The global reconstruction residual then decides where the fixed coding budget is spent, while redundant overlapping charts let local errors compensate one another after gluing. Our theory shows how tight-chart geometry controls distortion, how the global residual directs sequential allocation, and how data-geometry-guided refinement improves the codec. The resulting orbit machinery is compiled away, leaving a compact decoder in which the learned structure governs what is stored, where capacity is allocated, and how local information is assembled globally. Across four LLM embedding tables, OrBIT achieves 37.9\times compression on GPT-2 and over 23\times on each 7B table relative to 16-bit storage, while delivering competitive rate-distortion performance against established quantization and low-rank baselines.

[LG-12] Boosting and the Expressive Power of Simple Weak Learners via the γ-VC Dimension

链接: https://arxiv.org/abs/2610.10383
作者: Arthur da Cunha,Kasper Green Larsen,Liang-Yu Zou
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Boosting converts weak hypotheses with a small edge over random guessing into highly accurate predictors, but the expressive power of the resulting classifier can depend strongly on the structure of the base class. We study this phenomenon through the \gamma -VC dimension introduced by Alon et al. (STOC 2021). Our first result shows that this parameter characterizes the sample complexity for weak-to-strong learning up to a constant factor scaling in \gamma . We then sharpen the general relationship between the classic VC dimension and the \gamma -VC dimension. Finally, we also give improved upper and lower bounds on the \gamma -VC dimension for the fundamental concept classes of decision stumps and axis-parallel rectangles in \mathbbR^d .

[LG-13] ResidualQuant: KV Cache Quantization for Looped Transformers with 2-Bit Residuals

链接: https://arxiv.org/abs/2610.10381
作者: Heejun Kim,Junyoung Lee,SangLyul Cho,Dongsu Han,Insu Han,Sehoon Kim
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Looped Transformers improve parameter efficiency by repeatedly applying shared Transformer blocks over multiple recurrent loops, increasing computational depth without increasing the parameter count. However, KV cache memory still scales with the number of loops, becoming a key memory bottleneck that limits batch size and inference throughput. KV cache quantization can alleviate this bottleneck, but existing methods often suffer substantial accuracy degradation at aggressive low-precision regimes. We observe that looped Transformers offer a unique opportunity: KV states across loops are highly similar. Based on this observation, we propose ResidualQuant, which uses the final-loop KV states as a reference and represents the remaining loops with low-precision residuals. Our method further combines least-square scaling and rotations applied to the residuals, as well as loop-wise mixed precision, to enable accurate quantization down to INT2 while retaining efficient reconstruction. Across multiple looped Transformer models and mathematical reasoning and code generation benchmarks, ResidualQuant consistently improves the accuracy-memory tradeoff over state-of-the-art rotation-based KV quantization. In particular, our method retains accuracy close to BF16 under mixed-precision settings while reducing theoretical KV storage by 80.7%, achieving up to 13.0% higher accuracy than the rotation-based baseline at the same memory budget. On an RTX 5090, the reduced KV memory traffic improves fixed-batch decode throughput by up to 2.73x, while the smaller memory footprint enables up to 2x larger batches, improving peak throughput by up to 4.15x.

[LG-14] Continual Learning without Continual Training

链接: https://arxiv.org/abs/2610.10379
作者: Nikita Narayanan,Ritham Majumdarr,Sonali Parbhoo
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Continual learning requires models to adapt to new domains and new classes while retaining prior knowledge. Many existing methods rely on continued optimization, using regularization, replay, or parameter expansion to prevent new updates from overwriting previously learned knowledge. Instead, we propose replacing continual training with continual inference: a PFN-based model that is meta-trained, and then frozen, adapting to new classes only by extending an in-context evidence set. Our model, Latent Concept PFN, performs in-context Bayesian inference over a latent concept space that captures semantic structure shared across domains and classes. As each new domain or class arrives, exemplars are added to the memory; adaptation reflects updated posterior beliefs over latent concepts rather than gradient updates. No parameters are changed, reducing forgetting. The same method handles both domain and class incremental continual learning without task identity. Concept annotations are only used during meta-training, acting as a soft anchor on the latent space rather than a fixed bottleneck. Unlike fixed-vocabulary concept methods, the model also handles noisy, ambiguous, or incomplete annotations by combining concept labels with raw input evidence to discover distinctions beyond the predefined concept set. Experiments on class and domain incremental learning datasets demonstrate competitive continual learning performance while learning interpretable latent concepts.

[LG-15] mporally Interpretable Differentiable Decision Trees

链接: https://arxiv.org/abs/2610.10367
作者: Eisuke Hirota,Aarav Sane,Rohan Paleja
类目: Machine Learning (cs.LG); Robotics (cs.RO)
*备注:

点击查看摘要

Abstract:Interpretability offers a solution to safe autonomy by providing transparency into an agent’s underlying decision-making model. Within sequential-decision making tasks, differentiable decision trees (DDTs) are one approach to such interpretability, maintaining automatic-differentiable policies while providing humans with a discrete tree-based visualization. Nonetheless, current implementations of DDTs are not well-suited for sequential-decision making domains, as there exists an inherent mismatch between a tree’s single-timestep behavior and a human’s multi-timestep planning. Our work thus introduces time as a new dimension of interpretability, coined as temporal interpretability, and demonstrates how temporal abstractions via action chunking improve it. We achieve this by first introducing two novel policy gradient algorithms that incorporate action chunking. Additionally, to maintain parameter-efficient trees, we develop an information-theoretic tree restructuring algorithm that modifies the tree during training. Across four simulation environments, we find that warm-starting action chunked DDTs from a distilled action chunked policy is the most effective way to obtain temporally interpretable trees: they match neural network policies in three of the four domains while using up to 80 % fewer parameters. Our code is available at this https URL.

[LG-16] Koopman Observers for Diffusion Acceleration: Correcting Feature Forecasts with Shallow Measurements

链接: https://arxiv.org/abs/2610.10366
作者: Hanru Bai,Yuanchao Xu,Fengyi Li
类目: Machine Learning (cs.LG); Dynamical Systems (math.DS); Machine Learning (stat.ML)
*备注: 14 pages

点击查看摘要

Abstract:Feature caching accelerates diffusion sampling by replacing expensive network evaluations with predictions from previously computed activations. However, forecasts based only on past features cannot directly incorporate changes in the current denoising state. We investigate whether inexpensive, freshly computed features can serve as observations for correcting these predictions. We introduce an observation-corrected Koopman framework for accelerating frozen diffusion models. Using calibration trajectories, we identify finite-dimensional, time-dependent Koopman approximations that jointly describe the increments of shallow and deep network features. During accelerated sampling, these operators predict the evolution of expensive deep features, while innovations in the observed shallow features correct the predicted state. Periodic full evaluations refresh the observer, and all generative-model parameters remain unchanged. This formulation enables controlled comparisons of temporal prediction and observation correction. Across three 10,000-image runs per dataset, our method reduces paired Inception-feature MSE by 19.9% on CIFAR-10 and 11.9% on a ten-class ImageNet subset relative to channelwise affine prediction under the same four-partial-step schedule. Matched ablations attribute additional reductions of 4.54% and 4.67% to observation correction. The observer achieves 1.89\times and 1.85\times measured speedups over DDIM-50, supporting improved reference-sampler fidelity without retraining the denoiser.

[LG-17] Pathwise Information Certificates for Decentralized Adaptive Sensing

链接: https://arxiv.org/abs/2610.10362
作者: Theodoros Tsiligkaridis
类目: Information Theory (cs.IT); Machine Learning (cs.LG)
*备注: 26 pages, preprint

点击查看摘要

Abstract:We study decentralized adaptive sensing, where multiple agents choose measurements from evolving local beliefs while exchanging information over a communication graph. We ask whether the measurements actually selected by an adaptive policy have collected enough evidence to distinguish the true target from every plausible alternative. We develop a pathwise certificate based on the Rényi–Chernoff information accumulated along the realized sensing trajectory. It yields nonasymptotic MAP-error bounds and an anytime, network-wide stopping rule for arbitrary history-dependent sensing policies, while separating accumulated statistical information from a bounded network-mixing transient. Linear growth of the information against the least-resolved competitor implies exponential decay of MAP and squared-localization error. A classical pairwise KL converse, specialized to the adaptive decentralized transcript, shows that insufficient information on any pair prevents a positive uniform error exponent, confirming the hardest competitor as a fundamental bottleneck. Across policies, graph topologies, sensor profiles, and seeds, the worst-competitor score correlates more strongly with localization speed than an average-pair proxy in both 1D ( r=0.89 versus 0.40 ) and structured 2D sensing ( r=0.77 versus 0.48 ). Our results provide a practical way to certify and diagnose adaptive multi-agent sensing systems using the evidence they actually collect.

[LG-18] ORDERS: An Empirical Study of Norm-Rank Aggregation for Personalized Federated Learning

链接: https://arxiv.org/abs/2610.10361
作者: Koffka Khan
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Personalized federated learning combines shared representations with client-specific predictors, but the contribution of a server weighting rule can be obscured by local training and evaluation choices. We study ORDERS, a configuration that combines a shared backbone, a private residual adapter and classifier, geometric weights assigned by descending update norm, feature alignment, and private-parameter perturbations. The server computes a weighted sum of updates obtained from the same broadcast model; it does not obtain an additional optimization effect from sequential addition. A fully specified evaluation comprises 80 final runs: eight configurations, two datasets, and five training seeds on one fixed partition per dataset. On two-class-per-client CIFAR-10, ORDERS achieves 80.51 \pm 0.79% native mean client accuracy, compared with 79.02 \pm 1.42% for FedPer-R1 and 80.27 \pm 0.73% for the matched uniform-weight control. After common local fine-tuning, the difference from FedPer-R1 narrows to 0.32 percentage points. On Sent140, ORDERS reaches 74.71 \pm 0.49% , only 0.69 points above a post hoc client training-majority diagnostic. Ablations provide limited, endpoint-dependent evidence for norm ranking and alignment, and no clear benefit from perturbations. Parameter-payload savings are 5.47% and 0.78%, respectively.

[LG-19] AutoAdapt: Automatic Domain Discovery Enables Low-Cost Extensibility

链接: https://arxiv.org/abs/2610.10349
作者: Josh McGiff,Salma Mekaoui,Robert Shanahan,Nikola S. Nikolov
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Instruction-tuned models are deployed into environments where domains are heterogeneous and evolve, yet adding new domains or data typically requires costly retraining. We present AutoAdapt, a modular framework that incorporates new domains and data via targeted single-adapter training without modifying other adapters. The framework automatically discovers latent domains, uses them to train per-domain Low-Rank Adaptation (LoRA) adapters independently in parallel and performs parameter-free routing. Across 14 domain-specific benchmarks and GPT-4o pairwise judgements, AutoAdapt achieves parity with a LoRA adapter trained on all domains without requiring full-model retraining. We also find evidence of specialisation effect convergence across independent discovery methods. Overall, training each adapter on its own domain prevents domain interference by construction, thus enabling modular, taxonomy-free domain specialisation without aggregate performance loss or full model retraining.

[LG-20] Averag e-Reward Reinforcement Learning for Multichain MDPs: A Hierarchical Decomposition Approach

链接: https://arxiv.org/abs/2610.10326
作者: Huizhen Yu,Isaiah Heidt
类目: Machine Learning (cs.LG); Optimization and Control (math.OC)
*备注: 60 pages, 4 figures

点击查看摘要

Abstract:We study learning optimal policies in average-reward multichain Markov decision processes (MDPs), where the optimal gain may depend on the initial state and recurrence structures vary across policies, creating challenges for reinforcement learning (RL) methods. We propose an asynchronous value-iteration-based RL algorithm that requires no model knowledge beyond the MDP’s transition graph and leverages Bather’s decomposition to hierarchically partition the state space into communicating subsystems and transient states. This decomposition induces a recasting of the global decision problem into structured subproblems, which our algorithm exploits. We show that the algorithm converges to the optimal gain and produces gain-optimal policies after finite time. Building on this base algorithm, we develop two further algorithms: one approximately solves the multichain average optimality equations to obtain near gain-optimal policies, and another targets near bias-optimality by approximating the optimal bias function and solving an induced average-reward multichain MDP using the base algorithm. We provide almost-sure convergence guarantees for all three algorithms and empirically compare their tradeoffs, showing that the latter two also consistently improve transient performance relative to the base algorithm. To our knowledge, these are the first essentially model-free average-reward RL algorithms for general multichain MDPs without reductions to discounted problems.

[LG-21] HAN-Mamba: Hierarchical Selective State Space Networks for Multi-Scale Financial Volatility Forecasting

链接: https://arxiv.org/abs/2610.10323
作者: Mihai Bogdan Deaconu,Ioan Daniel Pop
类目: Machine Learning (cs.LG); Computational Engineering, Finance, and Science (cs.CE)
*备注: 16 pages. Accepted for publication in Springer Lecture Notes in Artificial Intelligence (ICAART 2026 Revised Selected Papers). Extended version of the ICAART 2026 paper (DOI: https://doi.org/10.5220/0014264900004052 )

点击查看摘要

Abstract:Short-horizon realized volatility forecasting requires the integration of market information that evolves at incompatible temporal resolutions, from second-level order book dynamics to weekly regime drift. Our conference work introduced HAN-T, a hierarchical architecture in which scale-specific Transformer encoders process short, mid, and long-horizon streams and a learned attention fuser weighs their contributions. This article replaces the quadratic attention encoders with selective state space (Mamba) encoders while retaining attention only in the fuser, where the input is a three-token set rather than a long sequence. The resulting hybrid, HAN-Mamba, summarizes each stream through a recurrent state whose input-dependent gating matches two structural properties of volatility: persistent but decaying memory and abrupt regime shifts. On the Optiver Realized Volatility Prediction benchmark under time-aware five-fold cross-validation, HAN-Mamba improves mean RMSPE over HAN-T (0.1942 vs. 0.1965) with 33% fewer parameters. Its linear-time encoders further allow the high-frequency context to be extended from 60 to 240 buckets, reducing error to 0.1927 where the attention variant saturates, and support constant-time streaming updates at inference. Ablations attribute the gains to the encoder swap, confirm that the hierarchical prior transfers across sequence-model families, and show that the permutation-invariant attention fuser remains the correct mechanism for cross-scale integration.

[LG-22] hinking in Depth: Retrospective Inference for Tabular Foundation Models

链接: https://arxiv.org/abs/2610.10317
作者: Hao-Run Cai,Si-Yang Liu,Zi-Jian Cheng,Kun-Yang Yu,Jin-Hao Sheng,Guo Yu,Chonghan Liu,Zhi Zhou,Jun-Peng Jiang,Lan-Zhe Guo,Han-Jia Ye
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Tabular foundation models (TFMs) are pretrained across diverse tabular tasks and make predictions on a new table at inference time using its labeled examples as context. Most recent TFMs perform such in-context prediction with stacked Transformer layers, repeatedly transforming how examples are represented and compared. By tracing individual queries through several strong TFMs, we find that predictive refinement is highly uneven across depth and is often concentrated in later layers. This uneven refinement motivates us to reconsider how intermediate representations are constructed and reused throughout the network. We introduce Retro, a tabular foundation model based on retrospective inference, where later stages can explicitly revisit and recombine intermediate information produced earlier in the network. Retro organizes this process around two complementary operations: which intermediate information to revisit, and how the resulting contextual update should be shaped for each query. Attention Residuals address the former by adaptively reweighting contributions from different depths, while query-conditioned Gated Attention addresses the latter by modulating the attention output element-wise across representation dimensions. Our analysis shows that Retro shifts predictive refinement earlier and more broadly across depth, with different stages revising different subsets of queries in a pattern suggestive of multi-view refinement. Across TabArena, TALENT, and RelArena, Retro ranks among the top three and lies on the Pareto frontier. These results indicate that directly reusing intermediate representations provides a practical way to better exploit depth in TFMs.

[LG-23] PoreML: A Data-Driven Framework for Learning Multiphase Flow in Porous Media

链接: https://arxiv.org/abs/2610.10314
作者: Chunyang Wang,Mingrui Zhang,Yuyan Zhang,Linqi Zhu,Xin Ju,Edo Sicco Boek,Martin J. Blunt,Gege Wen
类目: Machine Learning (cs.LG); Fluid Dynamics (physics.flu-dyn)
*备注:

点击查看摘要

Abstract:Multiphase flow in porous microstructures is central to CO _2 storage, fuel-cell operation, and flip-chip packaging. Predicting these flows remains challenging because wettability and complex pore geometry govern the nonlinear evolution of fluid interfaces. Machine learning holds substantial promise for advancing the field, but progress is constrained by scarce time-resolved 3D datasets and a lack of a unified workflow for training and evaluating models. To fill this critical gap, we introduce PoreML, an open-source framework unifying data generation, model training, and evaluation grounded in pore-scale physics. The framework comprises three core components. (a) A modern GPU-native lattice Boltzmann solver, validated against analytical solutions and published experiments, enables reproducible data generation. (b) A 3.3 TB dataset contains 560 simulation runs and 158,546 stored time steps across four application-driven scenarios. These trajectories span synthetic structures and geometries derived from micro-CT scans of real materials, covering diverse wetting conditions and viscosity ratios. © A unified learning framework evaluates one-step prediction and autoregressive rollouts. Its domain-specific evaluation protocols assess predictive accuracy and physical consistency. We evaluate five models of diverse architecture under these protocols. Two complementary challenges assess transfer to larger domains and from synthetic to micro-CT-derived structures. PoreML provides a shared foundation for machine-learning research on multiphase flow in porous media, with the aim of empowering the community to develop reliable predictive models and advance the field.

[LG-24] RSIGym: A Flexible Environment for Recursive Self-Improvement

链接: https://arxiv.org/abs/2610.10310
作者: Fanqing Meng,Lingxiao Du,Haocheng Lu,Qiguang Chen,Ziqi Zhao,Zijian Wu,Jiayuan Zhuo,Mengkang Hu,Michael Qizhe Shieh
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Recursive self-improvement requires carrying accepted changes into later improvement cycles, while studying agent-proposed changes also requires substantial research infrastructure. Existing settings often leave agents to rebuild routine infrastructure or restrict exploration to individual components. We introduce RSIGym, an agent-native research environment based on Everything as a Service (EaaS). RSIGym exposes training, inference, rollout, evaluation, and sandbox execution through reusable services, with shared budget and permission controls supporting Data, Harness, and Joint improvement tracks. This design enables agents to investigate individual interventions and jointly optimize data, training settings, and execution harnesses within the same environment. We define RSI-Index as the mean fraction of the remaining performance gap closed across five benchmarks covering software engineering, terminal interaction, mathematics, scientific reasoning, and skill-based tasks. Comparing six frontier research models in independent Joint runs, Opus 5 achieves the highest RSI-Index of 0.4809 under a 500 platform-service budget per benchmark run. Its selected systems improve all five benchmarks, raising SWE-bench Verified from 17.67% to 50.33% and AIME from 31.67% to 97.78%. Additional experiments examine DSH-harness refinement, budget variation, and restricted network access, while recorded trajectories reveal how agents diagnose failures and select candidates. We open-source the full RSIGym codebase and results to support reproducibility and further research.

[LG-25] How Do Transformers Learn to Represent Symmetries? NEURIPS2026

链接: https://arxiv.org/abs/2610.10305
作者: Eduardo Santos-Escriche,Valerie Engelmayer,Ya-Wei Eileen Lin,Stefanie Jegelka
类目: Machine Learning (cs.LG)
*备注: Accepted at NeurIPS 2026

点击查看摘要

Abstract:Training Transformer-based architectures with finite data augmentation has become an increasingly popular approach in geometric machine learning. Despite its empirical success, the interplay between the Transformer architecture, invariance to different symmetries, and augmentation budgets remains underexplored. In this paper, we study the ability of a vanilla Transformer to learn various symmetries through finite data augmentation for point cloud datasets. We identify an ordering of increasing learnability across the following symmetry groups: (i) non-angle-preserving symmetries, (ii) angle-preserving symmetries, and (iii) base angle-preserving subgroups, such as translation, rotation, and scale. For the base angle-preserving groups, we further investigate the Transformer’s extrapolation behavior and conduct a structural analysis of the trained models, allowing us to identify interpretable mechanisms that induce invariance. Finally, we extend our analysis to equivariant functions and show that the detected mechanisms for approximate invariance can also provide a key building block for learned equivariance. Our project page is available at this https URL

[LG-26] Continual Graph Multi-Agent Reinforcement Learning

链接: https://arxiv.org/abs/2610.10302
作者: Tommaso Marzi,Ahmed Hendawy,Jan Peters,Carlo D’Eramo,Andrea Cini,Cesare Alippi
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:In Continual Multi-Agent Reinforcement Learning (CMARL), agents learn cooperative policies across sequences of tasks, aiming to adapt effectively to new tasks while preserving the ability to solve previously encountered ones. In many applications, tasks differ in their underlying structure, which can represent, for example, distinct operational conditions or target configurations (e.g., different network topologies in power grids or arrangements in formation control). Existing CMARL methods lack dedicated mechanisms to leverage this structural information when learning new tasks, failing to promote transfer and mitigate forgetting. To fill this gap, we propose Continual Graph Multi-Agent Reinforcement Learning (CGMARL), a novel framework for CMARL problems in which task sequences are mapped into a series of attributed graphs, each modeling a task-specific structure. In CGMARL, each graph determines the environment dynamics (next states and/or rewards) and the number of agents for the corresponding task. Then, we present Graph-based Formation (GRAFO), the first CGMARL benchmark, and show how forgetting arises in this setting. Finally, to address this limitation, we propose Frozen Graph Encoder (FROG), a method that relies on a frozen graph backbone to preserve past structural information in graph-based CMARL policies. Experiments on GRAFO show that pairing FROG with existing CL methods substantially improves performance on multiple CGMARL scenarios.

[LG-27] Revisiting Explainable AI through Model-Independent Concept Dictionaries

链接: https://arxiv.org/abs/2610.10301
作者: Thomas Schnake,Doreen Schöppenthau,Alexander Meyer,Jacques Corbeil,Klaus-Robert Müller,Grégoire Montavon
类目: Machine Learning (cs.LG); Machine Learning (stat.ML)
*备注:

点击查看摘要

Abstract:Modern applications of AI rely on increasingly complex models. Explainable AI (XAI) has emerged as a set of techniques aimed at improving model transparency. However, existing XAI methods typically assume input features to be inherently interpretable, or they rely on intermediate internal abstractions that are difficult to characterize and highly architecture-specific, hindering consistent use across models. To address these limitations, we propose DictXAI, a method that defines concepts directly in the input domain via a dictionary—a large, potentially overcomplete set of predefined elements, each carrying an interpretable meaning. Technically, DictXAI first computes a sparse code of the input and then attributes the model’s prediction to the associated dictionary elements. We demonstrate the actionable nature of DictXAI explanations, showing that they can attribute AI malfunctions (e.g., Clever Hans effects) directly to identifiable artifact patterns in the data, while fostering human-AI alignment on intricate biomedical signals. We further demonstrate our method’s ability to operate across a wide variety of dictionaries, including learned image bases, analytically defined waveforms for electrocardiography, and experimentally acquired dictionary elements. Overall, our results show that DictXAI provides more interpretable, actionable, and architecture-agnostic insights than classical XAI or existing concept-based approaches.

[LG-28] Shared Gaussianization: What Gaussian Regularizers Certify About Contrastive Learning and What They Miss

链接: https://arxiv.org/abs/2610.10299
作者: Ruoyu Zhao,Yuting Chen,Jinheng Zhang,Zhehao Zou,Tong Che
类目: Machine Learning (cs.LG); Machine Learning (stat.ML)
*备注: 27 pages, 4 figures, 3 tables. Ruoyu Zhao and Yuting Chen contributed equally; Tong Che is the project lead

点击查看摘要

Abstract:What can a distribution-matching regularizer such as SIGReg in LeJEPA certify about contrastive learning? We study shared Gaussianization (SG), a characteristic-function Gaussianity test on the average of two normalized views, scaled by an independent \chi_d radius. Because disagreeing views shorten the average, one test detects both misalignment and non-uniformity. SG vanishes exactly at the aligned, uniform minimizers of population InfoNCE, and under equal marginals it bounds the InfoNCE excess by 4\cdot 3^3/4\beta times the square root of the SG loss, plus a term linear in the loss. The square-root rate and this dimension-free constant are sharp, and no squared mean-embedding distance on view pairs achieves a faster rate. With an explicit alignment term, a rotation-invariant uniformity test gives a linear bound if and only if its spectrum dominates that of InfoNCE’s kernel e^\beta u^\top v ; SG’s own test does, Gaussian kernels e^-\gamma |u-v|^2 qualify exactly when \gamma \ge \beta/2 , and moment matching never does. Away from the optimum, the objectives differ. Along an isotropic nuisance channel, pure SG lowers its loss by adding per-view nuisance whenever the shared code is non-uniform. An alignment weight above the channel’s gain makes the nuisance-free solution a strict local minimizer; for LeJEPA, the same rule gives a critical SIGReg weight that decreases with the batch size. At finite batch size, an off-diagonal U-statistic removes a plug-in bias toward misalignment. In controlled latent-variable models, pure SG retains per-view style, an alignment weight above the measured gain removes it, and for LeJEPA at three batch sizes the measured gain separates the encoders that retain style from those that do not. InfoNCE training also reaches a lower SG _0.2 loss than SG _0.2 training from scratch, which points to an optimization gap.

[LG-29] Physics-Aligned Electronic Ground-State Learning Improves Generalization

链接: https://arxiv.org/abs/2610.10298
作者: Eike S. Eberhard,Xaver Kainz,Viktor Kotsev,Abdulrahman Aldossary,Stephan Günnemann
类目: Machine Learning (cs.LG); Chemical Physics (physics.chem-ph); Quantum Physics (quant-ph)
*备注:

点击查看摘要

Abstract:Machine-learned interatomic potentials (MLIPs) excel at in-distribution tasks, accelerating drug and material development, yet they struggle to generalize out-of-distribution. We propose to push the cost-accuracy Pareto frontier by designing observable-agnostic electronic ground-state descriptor models (GSMs) with computational costs situated between MLIPs and Kohn-Sham density functional theory (KS-DFT). We align the learning objectives and architectures of GSMs with the governing equations of KS-DFT by enforcing physical constraints and removing optimization pressure on unphysical or irrelevant degrees of freedom. In our size-extrapolation experiments from QM9 to QM40, our combined contributions OrthoNormal-Loss (ON-Loss) and Grassmann Restricted Occupied-Orbital Training (GROOT) reach a 79.1% energy and 83.4% force mean absolute error (MAE) reduction over previous state-of-the-art density GSMs. For Hamiltonian GSMs, ON-Loss and Residual Optimal-gauge Conditioning-aware KS-Eq. Training (ROCKET) together reduce the energy and force MAEs of the strongest baseline by 99.8% and 95.9%, respectively. Using a self-consistency rejection criterion, we filter out extrapolation errors on QMugs, rejecting fewer than 0.4% of predictions while reaching an energy MAE of 0.07 mHa. Finally, we demonstrate the efficiency of label-free self-consistency fine-tuning, and transfer GSMs to reactive chemistry in Transition1x, reaching energy errors below chemical accuracy.

[LG-30] Energy-Efficient Gait Adaptation via Hierarchical Reinforcement Learning for Quadrupedal Locomotion Across Diverse Terrains ICRA2027

链接: https://arxiv.org/abs/2610.10297
作者: Ammar Issa,Anubhav Singh,Anton Tsaritsin,Sergey Kolyubin
类目: Robotics (cs.RO); Machine Learning (cs.LG)
*备注: 9 pages. Submitted to IEEE ICRA 2027. Ammar Issa, Anubhav Singh, and Anton Tsaritsin contributed equally

点击查看摘要

Abstract:While energy efficiency is a critical objective for legged-robot locomotion control, achieving low energy consumption while maintaining robust performance across different velocity ranges and terrain conditions remains a key challenge. This is particularly true for end-to-end RL policies, where gait generation, motion execution, and energy optimization are tightly coupled, leading to high sensitivity to reward design. In this work, we propose a hierarchical reinforcement learning (HRL) framework that separates a high-frequency policy for stable and robust joint-level motion execution from low-frequency gait adaptation that explicitly minimizes the cost of transport (CoT). The three-stage Isaac-based training procedure enables zero-shot sim-to-real transfer with improved tracking accuracy, robustness, and energy efficiency. The learned hierarchy exhibits automatic speed-dependent gait adaptation, transitioning from pacing at low speeds to trotting at higher speeds. We validate the proposed approach in simulation against representative single-policy and hierarchical locomotion baselines, demonstrating reduced CoT over a broad range of commanded velocities, while maintaining robust locomotion across flat, uneven rough, and inclined terrains. We further demonstrate its practical feasibility through zero-shot deployment on a physical Unitree AlienGo quadruped.

[LG-31] Neural Sampling with Reweighted Normalizing Flows via the Wasserstein–Fisher–Rao JKO Scheme

链接: https://arxiv.org/abs/2610.10278
作者: Chenguang Duan,Johannes Hertrich,Gabriele Steidl
类目: Numerical Analysis (math.NA); Machine Learning (cs.LG); Probability (math.PR); Machine Learning (stat.ML)
*备注:

点击查看摘要

Abstract:We propose a neural algorithm for sampling from distributions specified by unnormalized Boltzmann densities. Our approach is based on the Jordan–Kinderlehrer–Otto scheme for the Kullback–Leibler divergence in the Wasserstein–Fisher–Rao geometry (WFR JKO scheme). Our contributions are twofold. First, we prove that, for any fixed step size, the exact WFR JKO iterates converge exponentially fast to the target as the number of iterations tends to infinity. Notably, this result requires no structural assumptions on the target, such as log-concavity or a logarithmic Sobolev inequality. Second, we develop a neural implementation of the WFR JKO scheme that parametrizes its transport and reaction components using reweighted normalizing flows. Numerical experiments on challenging multimodal targets demonstrate the promising performance of the proposed method.

[LG-32] Sparse Planning in Visual World Models via Cost Gradients NEURIPS2026

链接: https://arxiv.org/abs/2610.10274
作者: Yingchen Xu,Edward Grefenstette
类目: Machine Learning (cs.LG)
*备注: Accepted at NeurIPS 2026. 20 pages, 6 figures, 8 tables. Project page and demos: this https URL

点击查看摘要

Abstract:Token-based world models enable fine-grained latent planning, but repeatedly processing large spatial token grids makes action search expensive. We introduce COSTGRAD, a training-free, goal-conditioned selector that ranks spatial tokens by the gradient norm of the planning cost with respect to each input token. By deriving importance from the downstream control objective, COSTGRAD targets tokens that matter for planning rather than merely for prediction. On AdaLN-conditioned predictors at 50% sparsity, COSTGRAD matches or exceeds full-token planning on three of four continuous-control benchmarks, while giving a measured 2.6\times wall-clock speedup per environment planning step. Combining token sparsity with reduced CEM search increases this to a \sim 5\times total speedup while still exceeding the full-token baseline. We also identify an architecture-dependent failure mode: in a matched AdaLN-vs-concat comparison, concat maintains comparable full-token performance but pure COSTGRAD loses its advantage over random selection. This difference tracks action-pathway drift: gradient-selected removal produces less drift than random removal on AdaLN, but more on concat. These results highlight selector-architecture compatibility as a design axis for sparse world-model planning. Project page and demos: this https URL

[LG-33] A Closed-Loop Non-Asymptotic Convergence Analysis of PPO with Learned Critics and Clipping

链接: https://arxiv.org/abs/2610.10273
作者: Junwei Su,Mengfan Liu,Yanyong Zhang,Chuan Wu
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Despite its widespread use, Proximal Policy Optimization with clipping (PPO-Clip) remains difficult to tune, and the interactions among critic learning, clipping, and rollout reuse remain incompletely understood. We develop a \emphnon-asymptotic analysis of PPO-Clip as a \emphclosed-loop actor–critic system. It captures actor–critic coupling, nonsmooth probability-ratio clipping, finite-batch reuse, and predictable early stopping under explicit coverage and critic regularity assumptions, using raw GAE and Monte Carlo critic targets. Our synchronous and asynchronous guarantees jointly characterize policy stationarity and the tracking accuracy of the learned critic, with explicit dependence on algorithmic parameters. A sufficient coupling condition gives optimization, critic tracking, clipping, and finite-batch errors a common amplification bound. The asynchronous result also requires a delay-dependent critic stepsize restriction; violating these conditions does not establish divergence. For finite layered MDPs with tabular critics, a uniform bound on the actual clipped-gradient class replaces complete-trajectory counting. A verified growing-horizon family has polynomial sample complexity, and a two-time-scale schedule gives O(T^-2/5) stationarity and critic-tracking bounds with explicit fresh-rollout accounting. These results together advance our understanding about PPO and provide theoretical guidance in tuning.

[LG-34] Using Small Language Models to Reverse-Engineer Machine Learning Pipelines Structures

链接: https://arxiv.org/abs/2610.10261
作者: Nicolas Lacroix,Frederic Precioso,Mireille Blay-Fornarino,Sebastien Mosser
类目: oftware Engineering (cs.SE); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Context: Once defined a taxonomy of stages structuring Machine Learning (ML) pipelines (e.g. Data Preprocessing, Modeling…), extracting these stages from source code is key for better understanding ML practices. However, the diversity caused by the constant evolution of ML (e.g., algorithms, datasets) makes this task challenging. Existing approaches either rely on non-scalable manual labeling or on classifiers that do not properly support domain’s diversity. These limitations call for more reliable solutions. Objective: We evaluate whether Small Language Models (SLMs) can leverage their code understanding and classification abilities to address these limitations, and enhance our understanding of practices in ML. Method: We conduct a confirmatory study based on two relevant reference works representing current limitations in the state-of-the-art. We first compare several SLMs using Cochran’s Q test, then evaluate the best-performing model against reference studies via two McNemar’s tests. An additional Cochran’s Q test examines how taxonomy definition variations affect the SLM performance. Finally, goodness-of-fit tests compare ML practice insights from SLM classification with those from prior studies. Results: First, we found that the taxonomy wording significantly impacts classification performance. Second, the best performing SLM yielded good results, yet, without outperforming other classifiers. Third, the three classification methods led to significantly different insights, with varying effect sizes, when exploring practices of data scientists. Conclusions: Limitations of existing classification methods bias our understanding of ML practices. While current SLMs show promising results without prior fine-tuning, they still exhibit common limitations, in addition to inference high costs challenging their applicability in large-scale studies. Subjects: Software Engineering (cs.SE); Machine Learning (cs.LG) Cite as: arXiv:2610.10261 [cs.SE] (or arXiv:2610.10261v1 [cs.SE] for this version) https://doi.org/10.48550/arXiv.2610.10261 Focus to learn more arXiv-issued DOI via DataCite (pending registration) Submission history From: Nicolas Lacroix [view email] [v1] Wed, 7 Oct 2026 15:35:04 UTC (1,466 KB) Full-text links: Access Paper: View a PDF of the paper titled Using Small Language Models to Reverse-Engineer Machine Learning Pipelines Structures, by Nicolas Lacroix and 3 other authorsView PDFHTML (experimental)TeX Source view license Additional Features Audio Summary Current browse context: cs.SE prev | next new | recent | 2026-10 Change to browse by: cs cs.LG References Citations NASA ADSGoogle Scholar Semantic Scholar export BibTeX citation Loading… BibTeX formatted citation loading… Data provided by: Bookmark checked="checked"class=“labs-tab-input”> Bibliographic Tools Bibliographic and Citation Tools Bibliographic Explorer Toggle Bibliographic Explorer (What is the Explorer?) Connected Papers Toggle Connected Papers (What is Connected Papers?) Litmaps Toggle Litmaps (What is Litmaps?) scite.ai Toggle scite Smart Citations (What are Smart Citations?) Code, Data, Media Code, Data and Media Associated with this Article alphaXiv Toggle alphaXiv (What is alphaXiv?) Links to Code Toggle CatalyzeX Code Finder for Papers (What is CatalyzeX?) DagsHub Toggle DagsHub (What is DagsHub?) GotitPub Toggle Gotit.pub (What is GotitPub?) Huggingface Toggle Hugging Face (What is Huggingface?) ScienceCast Toggle ScienceCast (What is ScienceCast?) Demos Demos Replicate Toggle Replicate (What is Replicate?) Spaces Toggle Hugging Face Spaces (What is Spaces?) Spaces Toggle TXYZ.AI (What is TXYZ.AI?) Related Papers Recommenders and Search Tools Link to Influence Flower Influence Flower (What are Influence Flowers?) Core recommender toggle CORE Recommender (What is CORE?) Author Venue Institution Topic About arXivLabs arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv’s community? Learn more about arXivLabs. Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?) mathjaxToggle(); We gratefully acknowledge support from our major funders, member institutions, , and all contributors. About Help Contact Subscribe Copyright Privacy Accessibility Operational Status (opens in new tab) Major funding support from

[LG-35] PairAudit: Guiding Human Review with Graph Tokens under Distribution Shift

链接: https://arxiv.org/abs/2610.10260
作者: Jiran Tao,Binyan Jiang
类目: Machine Learning (cs.LG); Cryptography and Security (cs.CR)
*备注: 22 pages, 3 figures

点击查看摘要

Abstract:Intrusion detectors can confidently misclassify attacks that were not seen during training. Human review can correct these errors, but only a limited number of cases can be checked. Uncertainty-based review may overlook confident errors, while anomaly scores alone do not show whether changing the review plan will correct more errors. We introduce PairAudit to find overlooked errors and improve review under a fixed budget. Its graph tokens capture prediction patterns across connected nodes. Rather than building another predictor through feature aggregation, PairAudit uses unusual relational patterns to uncover potential errors in existing predictions. Human feedback then helps decide whether these findings justify changing review priorities. Experiments across security tasks show that PairAudit corrects more errors on average than uncertainty-based review, including more errors on unseen attacks. These gains account for all review costs and do not require retraining the detector.

[LG-36] On the Cyclic Assumption of the Cow-Path Search Algorithm

链接: https://arxiv.org/abs/2610.10253
作者: Yuan Ma,Yiqun Lisa Yin
类目: Data Structures and Algorithms (cs.DS); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:In the cow-path problem, a cow must find a goal lying at an unknown distance on one of w paths connected only at the origin, and performance is measured by competitive ratio. Kao, Reif and Tate designed an efficient randomized algorithm in which the cow visits the paths in a fixed cyclic order. They proved the algorithm is optimal for w=2 , and subsequently Kao, Ma, Sipser and Yin proved its optimality for all w , with a claim that no algorithm does better than the best cyclic one. This note provides a detailed proof of that claim.

[LG-37] ProtocolMatch: Protocol-Dependent Model Selection for Scientific Dynamics Forecasting

链接: https://arxiv.org/abs/2610.10239
作者: Lu Wei,Yufeng Wang,Haibin Ling
类目: Machine Learning (cs.LG); Quantum Physics (quant-ph)
*备注: 14 pages, 4 figures, 8 tables

点击查看摘要

Abstract:Scientific dynamics forecasting is often framed as an architecture choice, although deployment is also determined by observed history, rollout feedback, compute budget, physical objective, and test distribution. We formulate protocol-dependent model selection and introduce ProtocolMatch, a compute-matched, validation-selected, and failure-preserving evaluation framework. On driven quantum-spin dynamics, we compare recurrent, patched-attention, causal-attention, and low-rank linear predictors across three independently generated datasets. The causal-attention–recurrence ordering reverses as the training set grows within a fixed two-spin task, while a linear predictor has the lowest mean error in the six-spin local-observable comparison. Restricting observed history worsens every refreshed-history view but improves every closed-loop view in the four-spin study. A latest-state MLP has lower error than persistence on every dataset under state refresh across all five cells, yet its closed-loop rank varies by system and includes finite explosive errors. Physical penalties improve targeted consistency without reliably improving prediction error, and in-distribution intervals lose most coverage after a driving-frequency shift. Thus scientific model selection should return a predictor with its protocol and report accuracy, physical validity, and shifted-distribution reliability separately.

[LG-38] Evaluating Sequence Assembly Strategies for Differentially Private Synthetic Time-Series Forecasting

链接: https://arxiv.org/abs/2610.10222
作者: Guoxiong Long,Huizhen Huang,Qikun Cai,Tao Huang,Chen Hou
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Differentially private time-series generators commonly produce fixed-length synthetic windows, whereas downstream forecasting models often require long continuous training sequences. How these windows are assembled after generation can therefore alter the effective synthetic data presented to a forecaster, even when the trained generator remains unchanged. We study this post-generation sequence assembly process by systematically varying overlap rates and window-weighting schemes and evaluating the resulting sequences in terms of boundary continuity, statistical and temporal fidelity, and Train-on-Synthetic-Test-on-Real (TSTR) forecasting utility. Across four types of public datasets (ETTh1, ETTm1, Weather, and Appliances) and five forecasting models, the results reveal a clear forecaster-dependent assembly principle: downstream TSTR utility is jointly shaped by the forecaster, overlap rate, and window-weighting scheme, leading to distinct assembly preferences across forecasting models. Increased overlap generally improves boundary continuity, but improvements in continuity or individual fidelity diagnostics do not consistently reduce forecasting error, indicating that these diagnostics alone are insufficient for selecting assembly configurations. Complete five-forecaster assembly grids, together with matched Train-on-Real-Test-on-Real (TRTR) references, further characterize these regularities and quantify assembly-dependent utility relative to real-data training. We then validate the identified principles through additional analyses of robustness and generator variability.

[LG-39] Finite-Sample Approximation of Hessian-Guided Perturbed Wasserstein Gradient Flows

链接: https://arxiv.org/abs/2610.10218
作者: Ryotaro Kawata,Atsushi Nitanda,Taiji Suzuki
类目: Machine Learning (cs.LG); Machine Learning (stat.ML)
*备注:

点击查看摘要

Abstract:Wasserstein gradient flow extends gradient descent to probability measures. Its Hessian-guided perturbed variant (PWGF) adds Gaussian perturbations to escape saddle points in nonconvex problems. We investigate when its approximation by finitely many interacting particles remains accurate over growing time horizons. Our analysis retains the curvature accumulated along the population-driven reference path: negative curvature can amplify approximation errors, while subsequent positive curvature can damp their influence. This captures favorable scenarios in which temporary instability is compatible with accurate tracking over growing horizons. Under regularity assumptions and a prescribed common perturbation schedule, we prove particle and objective-value tracking bounds on a high-probability event for reference paths satisfying explicit conditions on accumulated curvature. To handle state-dependent Gaussian jumps, we construct a population-first coupling that preserves the reference particles’ conditional independence and reduces jump errors to covariance comparison. We verify the conditions in a variance-plus-cosine model, where curvature recovery yields a growing-horizon tracking guarantee. We also establish local attraction, transverse descent, and positive second variation in two regions of a regularized matrix-factorization model, motivating a positive-negative-positive curvature pattern.

[LG-40] Edge Accuracy Is Not Enough: Why Dynamics-Learned Structure Fails to Transfer to Inverse Problems

链接: https://arxiv.org/abs/2610.10213
作者: Nicholas Tan Jerome,Fangnian Wang
类目: Machine Learning (cs.LG)
*备注: 13 pages, 8 tables. Code: this https URL

点击查看摘要

Abstract:A natural strategy for inverse problems with scarce labelled data is to transfer relational structure learned from abundant forward-simulation data. We show this strategy fails systematically, even when it satisfies the standard theoretical justification for why structure should help. We prove that approximate structure provides estimation-error benefits whenever the edge error satisfies \Delta n^2 - kn , reducing sample complexity from O(n^2) to O(kn+\Delta) . Structure learned via Neural Relational Inference (NRI) from dynamics prediction satisfies this condition, yet on a source-localisation task across 180 CFD-simulated hydrogen-leak scenarios and 180 acoustic scenarios, it degrades performance by 116% and 201% relative to a flexible, task-optimised attention baseline, while a physics-based prior (Green’s function) degrades by only 69-72%. Four independent lines of evidence show this is not a tuning failure: NRI improves only 0.5% when given 18x more training data (versus 16.6% for the task-optimised baseline, p0.001 ); performance is insensitive to the NRI edge threshold across a wide range; the dynamics-learned graph overlaps the task-optimal graph on only 6% of edges; and two further dynamics-derived structure estimators (correlation- and mutual-information-based) show no measurable benefit over a structure-free baseline, with the correlation-based estimator performing markedly worse. We formalise this gap as a statement about approximation error that the edge-accuracy condition cannot control, and we provide a lightweight transferability test (Jaccard similarity against a partially-observed target-task graph) that separates successful from failed transfer in all four domain/structure pairs we evaluate, using under an hour of computation and 15-20% of target-domain data; we present this as a heuristic calibrated on few cases, not a validated general threshold.

[LG-41] OrthoGen: A Generative Orthogonal Learner for Time-Varying Treatments

链接: https://arxiv.org/abs/2610.10210
作者: Tomàs Garriga,Valentyn Melnychuk,Konstantin Hess,Eduard Serrahima de Cambra,Axel Brando,Gerard Sanz,Stefan Feuerriegel
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Estimating conditional distributional potential outcomes (CDPOs) over time is important in medicine (e.g., to estimate patient-specific risks under different treatment sequences). However, this task is challenging because of time-varying confounding, yet existing adjustment strategies for this task are limited. In this paper, we aim to learn CDPOs under time-varying treatments using flexible generative models. Our contributions are two-fold. (1) We introduce a tailored adjustment strategy for our setting, namely, generative recursive g-computation. Our adjustment strategy recursively propagates full conditional outcome distributions rather than conditional means, modeling the variables of interest directly rather than full trajectories. Building on our adjustment strategy, we formulate simple generative learners for CDPO estimation. However, these learners can be sensitive to nuisance estimation errors, which motivates an orthogonal learner. (2) We thus introduce OrthoGen, a Neyman-orthogonal and doubly robust generative learner. Importantly, we show that OrthoGen further achieves rate double robustness and quasi-oracle efficiency under suitable conditions. Our learners are flexible and can be instantiated with different generative backbones (e.g., normalizing flows and diffusion models). Across experiments with synthetic, semi-synthetic and real-world datasets, we find that OrthoGen is highly effective. To the best of our knowledge, we are the first to propose a generative orthogonal learner for estimating CDPOs under time-varying treatments.

[LG-42] CARES: A Controlled Synthetic Benchmark of Speaker Reactions to Sound ICASSP2027

链接: https://arxiv.org/abs/2610.10208
作者: Marcel Gibier,Thomas Thebaud,Olivier Boëffard,Jean-François Bonastre
类目: ound (cs.SD); Machine Learning (cs.LG)
*备注: Submitted to ICASSP 2027

点击查看摘要

Abstract:Automatic audio scene description turns a recording into a text account of a situation. One difficulty is deciding which elements of the audio should be kept, since a description cannot include them all. Annotators disagree about this, making a ground truth hard to obtain. In this work, we first define the ground truth, then generate the data. We focus on audio events and define sound salience with a simple rule: a sound is salient when a speaker audibly reacts to it. For scale and variety, a controlled set of scenarios fixes the ground truth, and a language model writes the dialogues. The resulting corpus, CARES, contains 10,000 two-speaker scenes. We then benchmark six audio-language models on three tasks: identifying the scene, tagging the sounds present, and classifying reactions. We show that these models hear the sounds but miss how the speakers react to them.

[LG-43] How to train your model organism

链接: https://arxiv.org/abs/2610.10203
作者: Xilin Wang,David Bau,Byron C. Wallace
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Model organisms of alignment-relevant behaviors (e.g., backdoors, sycophancy, spurious correlations) have emerged as a key tool for evaluating whitebox interpretability techniques. We argue that the prevailing practice of training model organisms to a single objective of installing the target behavior is insufficient and propose validating model organisms with respect to three objectives with associated metrics: target-behavior installation, general-capability preservation (i.e., parametric knowledge, chat quality), and output naturalness (i.e., CoT and activations). We re-visit two publicly released organism suites using this validation framework and show that (1) chat quality and CoT naturalness degrade substantially across training recipes, and (2) validation metrics predict how well interpretability methods recover the installed behavior, e.g., a logit lens readout covaries with an organism’s general capabilities. We introduce a multi-objective training approach based on model merging to train more realistic model organisms. Finally, on a new suite of model organisms targeting demographic biases in clinical reasoning, we compare training recipes and find that DPO training stays closer to the base model than supervised finetuning, and the proposed model optimization approach better preserves capabilities and naturalness. Auditing this suite with an investigator agent, we again observe validation metrics tracking bias recovery. In sum, training methods shape the interpretability conclusions an organism supports, and we argue that one should consider multiple objectives to draw generalizable conclusions about interpretability methods using (realistic) model organisms.

[LG-44] Robust Decentralized Fairness Auditing

链接: https://arxiv.org/abs/2610.10199
作者: Sayan Biswas,Jade Garcia Bourrée,Anne-Marie Kermarrec,Palak,Martijn de Vos
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Emerging legislation requires large language models (LLMs) to be audited for compliance with regulatory standards, particularly fairness. Such black-box audits typically assume a single auditor with access to a large, representative set of queries. In practice, it can be difficult for an auditor to obtain such a query set, but multiple auditors can together cover the relevant demographic groups by auditing the LLM collaboratively with their individual query sets. However, relying on multiple auditors raises a fundamental trust problem, as they may act on behalf of the LLM provider to portray a misleading appearance of fairness, i.e., fairwashing. We propose Auditopus, a novel approach for robust decentralized fairness auditing. In Auditopus, auditing proceeds in rounds without a central server. In each round, every auditor issues a fixed number of queries to the LLM, and sends only cumulative statistics vectors of its query results to other auditors instead of sensitive queries in clear. The fairness of the audited LLM is then estimated by aggregating all the vectors. We show theoretically and empirically that even a single adversarial auditor in the network can steer this estimate by fabricating the vectors it sends, making an unfair LLM appear fair. To address this threat, Auditopus has each honest auditor locally down-weight any auditor whose cumulative statistics vectors are statistically inconsistent with previous ones. We implement Auditopus and compare it to robust aggregation baselines on two datasets with two pre-trained LLMs. Against an attacker that optimizes the vectors it sends to make the LLM appear fair, Auditopus reduces audit error by up to 78% on average relative to no defense and at least 62% relative to the robust aggregation baselines. Even when 49% of the auditors are adversarial, Auditopus never lets a very unfair or moderately unfair LLM pass as fair.

[LG-45] Pre-training of Bayesian Optimization Algorithm through Bayesian Optimization

链接: https://arxiv.org/abs/2610.10186
作者: Satoshi Katayama,Shoyo Hunt,Shintaro Masuda,Masayuki Karasuyama
类目: Machine Learning (cs.LG); Machine Learning (stat.ML)
*备注:

点击查看摘要

Abstract:Bayesian optimization (BO) is widely used as a standard approach for expensive black-box optimization. However, BO algorithms often involve parameters that must be specified in advance, and their performance can strongly depend on these choices. We propose a framework for optimizing such parameters using sample paths drawn from a Gaussian process (GP) inferred from the information available at the start of BO. We use cumulative regret as the performance metric for a BO algorithm. By running the BO algorithm on the generated sample paths, we obtain an empirical estimate of its expected cumulative regret for a given parameter configuration. Optimizing this estimate allows us to identify parameter configurations that, given the currently available information, are expected to achieve low cumulative regret. Since this parameter optimization is itself a black-box optimization problem, we employ another BO procedure to solve it, which we refer to as outer BO. Through experiments, we demonstrate that the proposed framework can effectively select parameter configurations that achieve strong performance among a range of candidate configurations.

[LG-46] A Unified Information-Theoretic Approach to Constrained Multi-Fidelity Multi-Objective Bayesian Optimization

链接: https://arxiv.org/abs/2610.10174
作者: Rikuto Matsumoto,Masanori Ishikura,Masayuki Karasuyama
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Bayesian optimization often involves multiple objectives, constraints, and fidelity levels. We address the challenge of jointly selecting where and at which fidelity to evaluate to identify the highest-fidelity feasible Pareto frontier in this combined setting. From a unified information-theoretic perspective, we measure query utility by the information gain about this frontier, provided by an observation. Since this mutual information is intractable, we derive a variational lower bound using a mixture of under- and over-truncated approximations to the Pareto-consistent region. Multi-fidelity surrogate models propagate the information to arbitrary fidelities, yielding a cost-aware acquisition function without separate heuristics for fidelity selection or constraint handling. Experiments on synthetic, benchmark, and real-world problems demonstrate effectiveness across diverse objective, constraint, and fidelity settings.

[LG-47] Attention via Black-Box Vector Search

链接: https://arxiv.org/abs/2610.10135
作者: Stepan Zharkov,Krish Singal,Ashwin Padaki,Alexandr Andoni
类目: Data Structures and Algorithms (cs.DS); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Sparse attention mechanisms estimate attention over n tokens using a small subset of keys. Many existing approaches use maximum inner product search (MIPS) to retrieve the heaviest keys, which motivates the following question: given black-box access to a MIPS oracle, how many keys must be retrieved to output an \varepsilon -accurate attention estimate? We answer this question by unifying prior approaches through the framework of priority sampling. With a single MIPS index, we show that \Theta(\sqrtn/\varepsilon) retrieved keys are both sufficient and necessary. With \Theta(\log n) indices, we give an algorithm that retrieves only O(\log n+1/\varepsilon^2) keys and prove that this is near-optimal. More generally, we design algorithms that establish a smooth tradeoff between the number of MIPS indices and number of retrieved keys. We then show that if we allow augmentation of keys and queries, we can bypass the above lower bounds: there exists a simple priority-sampling estimator using a single MIPS index and O(1/\varepsilon^2) retrieved keys. When integrated into LLM inference, our algorithms outperform top- k and sampling approaches used in prior work and yield attention approximation that scales favorably to long contexts. Subjects: Data Structures and Algorithms (cs.DS); Machine Learning (cs.LG) Cite as: arXiv:2610.10135 [cs.DS] (or arXiv:2610.10135v1 [cs.DS] for this version) https://doi.org/10.48550/arXiv.2610.10135 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[LG-48] m-Set Adversarial Bandits with Winner Feedback

链接: https://arxiv.org/abs/2610.10128
作者: Nicolò Cesa-Bianchi,Matteo Papini
类目: Machine Learning (cs.LG); Machine Learning (stat.ML)
*备注:

点击查看摘要

Abstract:We show upper and lower bounds on the regret of m -set adversarial bandits for different utilities (winner reward or sum of rewards) and feedback models (winner index, winner reward, sum of rewards, and their combinations). By comparing to standard bounds for combinatorial and MNL bandits, our results reveal how subtle changes in the setting can have a dramatic impact on the learning rates. Our main technical contributions are the information-theoretic lower bounds on the regret. Experiments on synthetic data confirm our theoretical analyses.

[LG-49] What Can a Gaussian Process Design Test

链接: https://arxiv.org/abs/2610.10122
作者: Ivan De Boi,Marnix Van Soom
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:A Gaussian process (GP) model can agree with the data for two reasons: its assumptions are right, or the chosen inputs could never have shown that they are wrong. The distinction can be checked from the design before any responses are observed. Every model implies relations that its noiseless responses must satisfy at the chosen inputs, such as the middle value lies on the line through its two neighbours. For GPs built from finitely many features, these relations are exactly the null space of the kernel matrix. Gale duality gives them a geometric interpretation, in which each observation has a vector and the smallest groups of observations that can expose an error are the circuits. For other kernels the relations become soft: response patterns may be improbable under the prior rather than algebraically impossible. A standard test then combines two kinds of evidence. Structural evidence comes from a violated relation and grows without limit as the noise falls. Prior-based evidence only says that a departure is improbable under the prior. With all inputs at the two ends of an interval, for example, a GP can reject a straight line against a large curvature, but only because the implied intercept is improbable, never because curvature was seen. In simulations the predicted power matched the observed rejection rates. Choosing the next input by predicted power raised the power against a localised discrepancy from 0.48 to 0.72, against 0.51 when choosing by predictive variance, and a grid in two dimensions contained exact tests of additivity that a Latin hypercube lacked. The test itself is classical. The contribution is the prospective reading of that test: before observing the responses, the design already determines what kind of contradiction it can produce.

[LG-50] CAFEFNO: Fourier Kernel Generation via Multiplicative Feature Composition

链接: https://arxiv.org/abs/2610.10105
作者: Hyungjoon Juen,Minwoo Shin
类目: Machine Learning (cs.LG); Numerical Analysis (math.NA)
*备注:

点击查看摘要

Abstract:The Fourier Neural Operator (FNO) learns solution operators of partial differential equations (PDEs) through Fourier-space kernel parameterization, but frequency truncation can limit the learning of high-frequency variations. AM-FNO and SirenFNO generate kernels for all grid modes from spectral coordinates using shared networks, making coordinate encoding and generator design important. Recent work on implicit neural representations (INRs) has proposed constructing frequency interactions through explicit feature composition rather than relying on subsequent MLPs to form them implicitly. Building on this approach, we propose CAFE+FNO, which incorporates Content-Aware Frequency Encoding+ (CAFE+) into Fourier kernel generation. CAFE+ combines Fourier–Chebyshev features through parallel affine branches and a Hadamard product, forming interactions within and across the two feature families. A kernel MLP maps the resulting representation of each normalized spectral coordinate to a complex channel-mixing matrix. Each layer shares its generator across all stored modes, making the number of trainable parameters independent of the number of modes for a fixed architecture. We compare CAFE+FNO with existing FNO variants on five PDE benchmarks and conduct ablation studies on basis configuration, multiplicative composition, and bandwidth learnability. Code and experimental configurations are available at this https URL.

[LG-51] Efficient Provably Private Classification with a Tabular Foundation Model

链接: https://arxiv.org/abs/2610.10068
作者: Talal Alrawajfeh,Cristiana Diaconu,Ossi Räisä,Sebastian Rodriguez Beltran,Yuan He,John Bronskill,Richard E. Turner,Antti Honkela
类目: Machine Learning (cs.LG); Cryptography and Security (cs.CR)
*备注: 74 pages, 18 figures; includes supplementary information

点击查看摘要

Abstract:Tabular data underpin prediction and decision-making in medicine, finance, government and science, but often contain sensitive individual-level information, creating a need for accurate prediction while preserving privacy. Traditional private learning provides formal privacy guarantees, but requires slow dataset-specific optimisation, suffers substantial utility loss under strong privacy, and is often difficult to apply correctly. Tabular foundation models adapt rapidly to new datasets, but existing models lack formal privacy guarantees, and are highly vulnerable to membership-inference attacks, limiting their use on sensitive data. Here we introduce PrivTab, an easy to use tabular foundation model for differentially private classification that embeds a privacy mechanism within its architecture. Pretrained on simulated datasets, PrivTab uses in-context learning to transform sensitive rows into compact, provably private summaries—effectively learning how to learn under privacy. PrivTab outperforms private linear and neural-network baselines under moderate-to-strong privacy, shows negligible membership leakage, maintains well-calibrated predictions under strong privacy, and reduces dataset fitting time by 10,000 times, requiring only a single forward pass. By combining formal privacy, speed, and easy of use, PrivTab brings recent advances in AI to applications where sensitive individual-level data have limited their adoption.

[LG-52] WxFM-XL: Adapting Univariate Foundation Models to Multi-Station Weather Forecasting

链接: https://arxiv.org/abs/2610.10057
作者: Xiao Wang,Changjian Chen,Zhuo Tang,Rongwen Li,Hongwu Liu,Kenli Li
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:With the rise of univariate time series foundation models (e.g., Sundial, Timer), initial efforts have been made to extend them to multivariate settings. However, these models mainly focus on modeling correlations among variables. When they are applied to multi-station weather forecasting, two important factors are often overlooked: (1) the spatial information of stations, and (2) different error priors of different stations relative to the foundation model. In this paper, we propose WxFM-XL, a model for adapting univariate time series foundation models to multi-station weather forecasting. WxFM-XL introduces a cross-station error correlation prior graph to capture stationwise error priors with respect to the foundation model. Building on this, we further propose a dynamic fusion mechanism that adaptively integrates a spatial correlation graph with the error correlation prior graph. Experiments on multiple datasets demonstrate that our model outperforms state of the art baselines.

[LG-53] ransition Path Sampling Using Koopman Operators and Exit-Time Optimal Control

链接: https://arxiv.org/abs/2610.10054
作者: Boya Hou,Shane Wang,Siddharth Ambekar,Maxim Raginsky,Olgica Milenkovic
类目: ystems and Control (eess.SY); Machine Learning (cs.LG); Machine Learning (stat.ML)
*备注: 30 pages, 6 figures

点击查看摘要

Abstract:Sampling transitions between metastable states is a central problem in dynamical systems theory and molecular dynamics in particular. A key challenge is the existence of high free-energy barriers that separate the states, making transitions extremely rare. Recent machine learning-based methods cast transition path sampling (TPS) as an optimal stochastic control (OSC) problem over a fixed time horizon, and parameterize the drift bias via a neural network trained by simulation-in-the-loop, requiring repeated biased rollouts. To address computational and performance guarantee issues of these models, we propose a new approach for the problem based on Koopman operators. Because Koopman operators are linear, their leading eigenfunctions reveal the metastable sets and provide an estimate of the committor function with no transition path information required. Furthermore, we formulate TPS as an OSC problem up to an exit time. Our time horizon is the first hitting time of the target set, and our running cost penalizes time spent in nonreactive regions by encoding the estimated committor function. We derive the optimal controller in closed form and approximate it in a reproducing kernel Hilbert space (RKHS). This reduces the problem of constructing the optimal controller to solving a single equality-constrained quadratic program, whose solution can be characterized by a linear Karush-Kuhn-Tucker (KKT) system. On the two-channel double well and alanine dipeptide, our controller increases the fraction of trajectories reaching the target from 0% to 99.8% within 1000 steps, and from 0% to 93% within 1ps, respectively.

[LG-54] Evolve on the Host Predict on the Edge: Deploying Online Neuroevolutionary Architecture Search for Cross-sectional Stock Return Prediction

链接: https://arxiv.org/abs/2610.10038
作者: Jonathan Chang,Zimeng Lyu
类目: Computational Engineering, Finance, and Science (cs.CE); Machine Learning (cs.LG); Neural and Evolutionary Computing (cs.NE)
*备注:

点击查看摘要

Abstract:Accurate forecasting models are usually large, expensive to update online, and fixed in architecture once trained. We apply ONE-NAS, an online neuroevolutionary architecture search that evolves a population of small recurrent networks as each window of data arrives, to daily cross-sectional stock return prediction, and pilot it on a host and endpoint pipeline: the host runs the search and ships each generation’s champion genomes over TCP/IP to a Raspberry Pi 4B, which predicts online. On the Pi a single champion predicts a 50-stock window in 24.6~ms and the ensemble of 40 island champions in 556~ms, far inside the daily decision cycle. On four panels of US mid-cap equities over 2022–2024, reading the population as a rank-mean ensemble of island champions returns +27.5% net of realised transaction costs, against +11.3 to +14.8% for online LSTM, online GRU and monthly-retrained LSTM baselines and +4.5% for the single best genome used in prior ONE-NAS work.

[LG-55] Matching of signal noise and hardware timescales for filtering and forecasting of correlated noise signals

链接: https://arxiv.org/abs/2610.10037
作者: Joshua Donald,Alex Gabbitas,Arthur G. T. Coveney,Sergey Savel’ev,Pavel Borisov
类目: Machine Learning (cs.LG); Disordered Systems and Neural Networks (cond-mat.dis-nn); Applied Physics (physics.app-ph); Data Analysis, Statistics and Probability (physics.data-an)
*备注:

点击查看摘要

Abstract:Physical reservoir computing exploits the nonlinear dynamics of physical systems to process time-dependent data with greater energy efficiency than conventional machine learning approaches. However, physical reservoirs have fixed intrinsic response timescales, whereas real-world signals combine deterministic and stochastic components across multiple timescales. Here we show, using a nanoporous niobium oxide reservoir, synthetic noisy signals and cryptocurrency-price volatility, that the relationship among noise correlation time, reservoir memory and forecast horizon determines whether correlated noise is filtered or predicted. Noise varying faster than the relevant reservoir memory and forecast horizon is averaged by the reservoir, whereas the temporal structure of slower-varying noise is sufficient for algorithmic forecasting. We introduce the reservoir memory horizon and forecasting regime index to distinguish these operating regimes. These contributions demonstrate that timescale matching can guide the encoding of input time series and development of physical reservoir architectures that filter, analyse and predict stochastic signal components across distinct temporal scales.

[LG-56] Oscillatory Neural Dynamics over Sheaves

链接: https://arxiv.org/abs/2610.10018
作者: Jan-Willem Van Looy,Alessandro Trenta,Alessio Gravina,Alessio Borgi,Ferdinando Zanchetta,Pietro Liò,Davide Bacciu,Rita Fioresi
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Effective long-range propagation remains a central challenge in graph neural networks, as increasing a model’s propagation depth does not guarantee that distant nodes effectively influence each other. Sheaf neural networks enrich graph propagation through matrix-valued transport between stalks; still, this expressivity alone does not automatically imply effective long-range communication. We introduce ONDA, a long-range graph learning framework based on operator-valued information waves. Stalk-valued representations evolve through second-order dynamics governed by learned sheaf transport operators, combining wave-like propagation with expressive local geometry. We characterize long-range influence through a stalk-wise sensitivity analysis and show that the cross-influence never vanishes. Across long-range propagation, severe graph bottlenecks, graph transfer, and heterophilic benchmarks, ONDA consistently improves over scalar wave propagation, diffusive sheaf baselines, and state-of-the-art models, demonstrating the benefit of coupling wave dynamics with matrix-valued transport.

[LG-57] A Drosophila Whole-Connectome Network Can Learn Human-Designed Cognitive Tasks

链接: https://arxiv.org/abs/2610.10014
作者: Joonghui Cho,Minchan Kang,Daeshik Kim
类目: Machine Learning (cs.LG)
*备注: 14 pages, 4 figures. Code: this https URL

点击查看摘要

Abstract:Can a biological wiring diagram serve as a useful computational substrate beyond the behaviors for which it evolved? We use the publicly released MaleCNS v1.0 connectome, reconstructed from a single adult male Drosophila specimen, as the fixed recurrent topology of an artificial network. We train separate models for bounded addition and for a controlled grounded relational language task built from a fixed 100-word lexicon. In both models, one scalar is learned per anatomical edge. The anatomical graph reaches 92.77% mean accuracy on held-out addition, compared with 67.93% for directed degree-preserving rewires. On the strict paired language endpoint, which matches original and order-reversed scenes to their corresponding descriptions, it reaches 61.59% across four fixed interfaces, compared with 44.17% for matched rewires. At the canonical interface, it ranks first in a fixed 21-graph comparison. On the matched 48-group intervention subset, shuffling task-defined sensory features reduces its score from 60.94% to 19.27%. Together, these results show that higher-order MaleCNS wiring provides a reusable inductive bias for bounded addition and grounded relational language.

[LG-58] Marrying Pricing and Advertising with LLM s

链接: https://arxiv.org/abs/2610.09985
作者: Alessandro Barro,Francesco Bacchiocchi,Francesco Emanuele Stradi,Alberto Marchesi
类目: Computer Science and Game Theory (cs.GT); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:We study a sequential pricing problem in which a seller jointly posts a price and an advertisement generated by a large language model (LLM). The seller aims to maximize revenue under an unknown product demand that depends on both decisions, while observing only whether each offer leads to a purchase. We propose an online actor-critic algorithm that combines low-rank adaptation (LoRA) of a pretrained LLM with a demand model fitted to available data. At each round, the actor generates an advertisement, and the critic estimates purchase probabilities to guide price selection. Then, the resulting feedback is used to update both the actor and the critic, with the critic’s revenue estimates providing a baseline for policy gradient updates of the actor. To evaluate our approach, we develop an evaluation framework with three synthetic demand models and a demand simulator built from real-world marketplace data. Finally, we compare our algorithm with benchmarks that do not jointly optimize price selection and advertisement generation, achieving expected revenue gains over the reference policy of 5.69%, 5.18% and 55.96% under the three synthetic demand models and 5.81% under the marketplace simulator.

[LG-59] R-PTQ: High-Accuracy Integer-Only Transformer Post Training Quantization via Taylor Region Reformulation

链接: https://arxiv.org/abs/2610.09969
作者: Eliyahu Levy,Adam Teman,Yoni Pugachov
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Post-training quantization (PTQ) enables efficient deployment, yet transformer architectures remain challenging to quantize due to nonlinear layers. While existing methods attribute accuracy loss to insufficient numerical precision, often necessitating floating-point fallbacks, we demonstrate that degradation is actually driven by specific structural error sources. We find that learned scale parameters in normalization layers and compounded approximations in GELU are the primary error contributors, whereas SoftMax remains inherently robust to aggressive quantization. To address these bottlenecks, we introduce TR-PTQ, a unified integer-only formulation using shared Taylor Region (TR) exponential and logarithm primitives. This approach allows computationally expensive operations, including division and square roots, to be performed entirely in the log-domain via standard integer arithmetic. Combined with a calibration-free, outlier-aware optimization for LayerNorm parameters, our method eliminates the need for floating-point hardware units for nonlinearities, achieving less than 1.5% absolute accuracy degradation across vision and language benchmarks.

[LG-60] Force without transmission: a depth-induced rank collapse that no loss on the representation reopens

链接: https://arxiv.org/abs/2610.09958
作者: Martin Hofmann,Patrick Mäder
类目: Machine Learning (cs.LG)
*备注: 16 pages, 8 figures, 2 tables

点击查看摘要

Abstract:Training can drive a transformer into a rank collapse: all token representations point in one direction, and learning stops. In a related collapse of attention, a loss term with a bounded corrective force repairs the network during the run. We ask whether such a term repairs rank collapse. We collapse small transformers by weakening their skip connection and treat copies of the collapsed network. No added loss term repaired the collapse, although the stronger kind pushed with about a tenth of the task gradient. The reason was the path, not the strength. The task gradient no longer reached the query and key weights, which decide where attention looks, and the added term’s gradient faded before the blocks where the collapse forms. Restoring the skip connection, which changes no weight, reopened this path at once. The rank then recovered, but only far above the scale of collapse. After a burst of high learning rate the path stayed open and the rank recovered untreated. Registered predictions from the path ranked recovery times but did not transfer to this cause. In every case the loss stayed above that of a healthy network after the rank recovered. Whether a collapsed network can be repaired depends on whether the gradient still reaches the weights that must change, not on how strongly a loss term pushes.

[LG-61] Learning Traffic Flow Dynamics with Stochastic Physics-Informed Neural Cellular Automata

链接: https://arxiv.org/abs/2610.09946
作者: Federica Bragone,Matthieu Barreau
类目: Machine Learning (cs.LG)
*备注: 45 pages, 16 figures

点击查看摘要

Abstract:Traffic flow modeling is essential for understanding and predicting the collective dynamics of vehicles on road networks. Cellular automata provide a simple, interpretable yet powerful framework for representing these dynamics via local interaction rules, while retaining the ability to reproduce complex macroscopic traffic phenomena. However, learning local transition rules from data while preserving physically meaningful constraints remains challenging, particularly for stochastic models. In this work, we propose a physics-informed neural cellular automaton (PI-NCA) for data-driven traffic flow modeling. Building on the standard neural cellular automaton (NCA), we design a neural architecture that is physically consistent with the road topology and guarantees conservation of the total number of vehicles, thereby constraining the learned transition rules to physically admissible dynamics. We further extend this framework to stochastic dynamics by parameterizing probabilistic transition rules while preserving the same physics-informed constraints. We evaluate the proposed models on multiple traffic scenarios generated by the well-established Nagel-Schreckenberg and Kerner-Klenov-Wolf cellular automata. The results demonstrate that the PI-NCA successfully learns the dynamics of both traffic models and consistently outperforms a standard NCA, while the stochastic extension captures probabilistic transition rules without compromising the imposed physical constraints.

[LG-62] Many Ways to Succeed: Diversity-Driven RL Fine-Tuning for VLA Generalization

链接: https://arxiv.org/abs/2610.09943
作者: Haoru Li,Jinmei Liu,Zhiyong Wang,Xiaoming Li,Zhenhong Sun,Daoyi Dong,Chunlin Chen,Zhi Wang
类目: Robotics (cs.RO); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Reinforcement learning (RL) fine-tuning improves vision-language-action (VLA) policies through closed-loop experience, yet generalization beyond the fine-tuning distribution remains limited. Our analysis reveals a selective reshaping of exploration: RL contracts behavior globally, yet diversifies successful trajectories, elicits success with fewer rollouts, and covers more of the latent task-valid solution space than supervised fine-tuning. Broader successful-mode coverage may provide alternative strategies under distribution shifts. Inspired by this, we introduce DRIVE (Diversity-driven RL fIne-tuning for VLA gEneralization), which turns successful-behavior diversity into an explicit RL objective. DRIVE groups rollouts under matched task conditions, compares their trajectories with temporal alignment, and derives a success-conditioned intrinsic reward from relative behavioral diversity. This design encourages broader coverage of feasible solutions without rewarding diverse failures or superficial timing differences. Across LIBERO-Plus, ManiSkill3, and RoboTwin 2.0, DRIVE improves the average out-of-domain (OOD) performance over vanilla RL fine-tuning by 5.3 points on \pi_0 and 2.0 points on \pi_0.5 . On a dual-arm AgileX PiPER-X platform, DRIVE further increases average OOD success from 64.1% to 73.3% (+9.2 points), demonstrating gains that persist under physical deployment.

[LG-63] Expected Sample Complexity in Multi-Armed Bandits

链接: https://arxiv.org/abs/2610.09929
作者: Nadav Sukenik,Nadav Merlis
类目: Machine Learning (cs.LG); Machine Learning (stat.ML)
*备注:

点击查看摘要

Abstract:Sample complexity is a widely used metric in sequential decision-making problems, defined as the number of suboptimal decisions during the interaction between the agent and an environment. We study the sample complexity of stochastic multi-armed bandit problems and introduce the expected sample complexity performance measure, analyzing it in a novel framework called approximately correct in expectation (ACE). We show that ACE guarantees imply almost sure convergence to the optimal expected reward, in contrast to high-probability guarantees found in other frameworks, and also show how to convert ACE guarantees into explicit expected regret bounds. We further show that, in contrast to existing measures, deterministic algorithms cannot obtain favorable ACE bounds, and analyze stochastic algorithms in two settings: when the allowed suboptimality level \epsilon is known to the algorithm and when it is unknown. In the former, we devise an explore-then- \epsilon -greedy algorithm, and in the latter, we analyze the expected sample complexity of Thompson sampling. Finally, we establish nearly matching lower bounds for both settings, showing that the algorithms are tight in \epsilon and proving a performance separation between the two regimes.

[LG-64] Eigenvalues of the Hessian in Deep Learning: The Origin of Symmetry and Its Breaking

链接: https://arxiv.org/abs/2610.09919
作者: Yossi Arjevani
类目: Machine Learning (cs.LG); Optimization and Control (math.OC); Machine Learning (stat.ML)
*备注:

点击查看摘要

Abstract:Hessian spectra at trained models in deep learning exhibit a persistent pattern: eigenvalues organize into distinct clusters, including a large bulk near zero and a few isolated outliers. This paper shows that a natural account of these spectral phenomena emerges when the original setting is understood as a departure from a nearby, otherwise hidden, highly symmetric reference. Modifications, including changes to the architecture, data distribution, or parameter metric, expose a nearby reference configuration whose Hessian exhibits rich invariances-ones not accounted for by weight symmetries. There, symmetry enables a precise description of the spectra, forcing high-dimensional kernels and eigenvalues of large multiplicity. Returning to the original configuration breaks the Hessian symmetry and thereby produces the observed hierarchy of clusters and outliers. The framework is developed in some generality, with a detailed analysis of three-layer ReLU networks and applications to convolutional, graph, and transformer models, as well as to the NTK. The same mechanism is further shown to yield analogous spectral structures in layerwise Hessians and the Gauss-Newton matrix. Subjects: Machine Learning (cs.LG); Optimization and Control (math.OC); Machine Learning (stat.ML) Cite as: arXiv:2610.09919 [cs.LG] (or arXiv:2610.09919v1 [cs.LG] for this version) https://doi.org/10.48550/arXiv.2610.09919 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[LG-65] NeuralZip: Reusable Setup for Fast Lossless Compression

链接: https://arxiv.org/abs/2610.09916
作者: Martín Bravo,Samuel Horváth,Gonzalo Navarro,Andrés Abeliuk
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Lossless compression can reduce the storage and movement of model weights without changing their floating-point values, but repeated statistical analysis and code construction add computational overhead. We study whether the statistical structure of exponents can be prepared once and reused. For this, we introduce NeuralZip, which groups chunks with similar exponent distributions, shares Huffman codes, and selectively represents recurring exponent tuples using packed exponents, thereby achieving additional moderate compression ratios. A setup chooses these representations before subsequent encodings, while every encoding still processes the current tensor values. In floating-point model checkpoints, post-setup compression is 1.81-21.33 \times faster than the baselines and achieves exact bit-to-bit reconstruction. We show that this setup can be precomputed and transferred from another compatible architecture, preserving similar compression ratios and avoiding the need to amortize setup costs. Therefore, compression adaptation is transferable and reusable. Training checkpoints demonstrate continued reuse as the weights evolve. Finally, GPU experiments reduce active memory usage by up to 27.5 % while reproducing the logits exactly.

[LG-66] Identifiability of a dissipative knowledge-dynamics model: exact recovery under designed excitation degeneration on observational data

链接: https://arxiv.org/abs/2610.09889
作者: Arman Kostanian,Armen Beklaryan
类目: Machine Learning (cs.LG); Optimization and Control (math.OC); Machine Learning (stat.ML)
*备注:

点击查看摘要

Abstract:Human learning is a dissipative dynamical process: mastery accumulates through practice, decays through forgetting, and propagates across interdependent concepts. We model it as a nonlinear dissipative system of ordinary differential equations whose parameters are mechanistically meaningful (a concept-transfer matrix encoding prerequisite coupling, per-concept forgetting rates, and a saturating practice-response gain), and we study when those parameters can actually be recovered from data. We prove a structural identifiability theorem for the associated inverse problem under explicit excitation conditions, with constructive closed-form recovery for the two-concept case, together with monotonicity, robustness and L-stability results. We derive a semi-implicit L-stable scheme for the dissipative subsystem and a batched solver numerically equivalent to the per-trajectory formulation (bit-exact predictions, gradients to 10^-10 ) yet two orders of magnitude faster, making estimation feasible on cohorts of 10^5 learners. The empirical study is two-sided. Under the theorem’s excitation conditions, synthetic recovery is exact: parameters to machine precision, prerequisite structure at F_1 = 1.0 . On large observational benchmarks it is not. An apparently strong recovery, with forgetting rates correlating with topic difficulty at Spearman \rho = 0.83 , is refuted by four independent controls: it survives destroying the temporal order of the data, is matched by a classical Bayesian baseline, and is unaffected by removing real timestamps. We trace this to the stationary structure of the model and show that it is the degeneration the theorem predicts in the absence of designed excitation. The result delineates a sharp boundary between identifiable and unidentifiable regimes and yields a validation protocol for interpretability claims.

[LG-67] Layerwise Error Attribution for Fast and Robust Mixed-Precision Post-Training Quantization

链接: https://arxiv.org/abs/2610.09877
作者: Samy Houache(IMB, UB),Yann Traonmilin(IMB, UB),Jean-François Aujol(UB, IMB)
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Mixed-precision post-training quantization is a network compression method that assigns bits layer by layer, under a global memory budget using a small calibration set. The main difficulties are to overcome the combinatorial nature of the allocation problem and to manage the sensitivity to small, potentially corrupted databases. Hence, an efficient allocation method should be fast to compute and preserve model quality when calibration data are corrupted. To design such a method, we derive a layerwise probabilistic analysis of the quantization error that separates propagated error from the local perturbation introduced at a given layer. We use this local term to build a separable score for a simple allocation algorithm, that requires no external solver. The probabilistic nature of our approach brings robustness to corrupted data. On denoising tasks with DRUNet, with an average budget of 4 bits per weight, our method matches or improves state-of-the-art mixed-precision baselines under clean calibration, and is more robust to corrupted calibration, with PSNR gains of up to 7.5 dB under the tested corruptions. Experiments show bit-allocation speed-ups from 28x to 2,570x over the studied baselines. For quantized diffusion models, our experiments show that a direct application of our framework also improves the state-of-the-art.

[LG-68] MUNITE: Unified Multimodal Latent Inference for Any-to-Any Multimodal Generation

链接: https://arxiv.org/abs/2610.09866
作者: Kyeongmin Yeo,Minhyuk Sung
类目: Machine Learning (cs.LG)
*备注: Project page: this https URL

点击查看摘要

Abstract:We introduce MUNITE, a latent-variable framework for flexible any-to-any multimodal generation that treats encoding and latent generation as the same inference problem under different amounts of observed evidence. Given any subset of modalities, MUNITE models the conditional distribution over the latent representation associated with the complete observation. Full observation recovers deterministic encoding, no observation recovers the latent marginal, and intermediate subsets define conditional latent inference, all within a single conditional flow model. A shared latent sample captures variation that must remain consistent across generated targets, while modality-specific generative decoders model the remaining uncertainty independently. To learn these conditional distributions from incomplete training examples, we extend conditional flow matching through self-distillation: predictions conditioned on richer available observations supervise the same model conditioned on smaller subsets at the same intermediate latent state. When the richer-evidence trajectory follows the exact conditional flow, this provides the same expected learning signal as full-target denoising. Across PolyMNIST-D-Q, FFHQ64, and image-text-audio, MUNITE achieves competitive or better generation quality and source-target alignment, with higher joint-generation coherence. In particular, it attains the highest coherence in all one-to-many and unconditional image-text-audio comparisons, showing the effectiveness of unified latent inference across diverse multimodal settings.

[LG-69] Dual-QK: Sharp Queries and Flat Keys for Prunable 2-bit KV Caches

链接: https://arxiv.org/abs/2610.09827
作者: Sunjoo Whang,Jungjun Oh,Minsung Kim,Dongho Seo,Jisu Shin,Gregory Kielian,Hoi-Jun Yoo,Sangjin Kim
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Long inputs and extended generation increase the storage and access costs of the key-value (KV) cache. Low-bit quantization reduces storage and memory traffic, while query-channel pruning can further reduce key-cache reads. Rotation-based quantization redistributes the energy of key outliers across channels. To maintain computational invariance, the same orthogonal transform must be applied to queries, preserving query-key dot products. However, this rotation can disperse query energy, weakening the separation between a few large components to retain and many small ones to prune. We introduce Dual-QK, which uses paired non-orthogonal query and key transforms to address this conflict. Using calibrated query and key statistics, Dual-QK combines partial key whitening with a query-aligned basis to balance key scales for INT2 quantization and concentrate query energy for dynamic channel pruning. Channel-0 protection and bucket-relative RoPE support low-bit accuracy over long contexts. Experiments on four models across five generative benchmarks and long-context retrieval tasks show improved accuracy over OSCAR on most tasks at 40% query-channel sparsity. At a 128K context, Dual-QK provides 6.8\times KV-cache compression and an estimated 8.3\times reduction in KV read volume relative to unpruned BF16. Under the evaluated configurations, our SGLang implementation achieves up to 3.75\times the decoding throughput of unpruned BF16.

[LG-70] Homogenization in Multi-Agent Systems

链接: https://arxiv.org/abs/2610.09824
作者: Prakhar Ganesh,Kyra Wilson,Luca Zappella,Barry-John Theobald,Nicholas Apostoloff,Lucas Monteiro Paes,Nivedha Sivakumar
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Multi-agent systems (MAS) leverage interactions between agents to perform complex tasks. Despite their success, we show that these interactions can also lead to homogenization, i.e., agents converging to similar behaviors. Homogenization in MAS can reduce agent diversity and reinforce shared failures. In this paper, we operationalize homogenization using three metrics: conformity to the majority, polarization towards extremes, and growing inertia against changes over subsequent interactions. We evaluate homogenization in MAS for code generation, hiring, and scientific peer review. Across these tasks, we show that homogenization translates to concrete downstream risks: in code generation, it hides and amplifies correlated errors which can create systemic vulnerabilities; in hiring, it allows the influence of biased agents to persist long after their removal; and in peer review, it creates uneven evaluation standards across research areas. Our results establish homogenization as a failure mode of MAS, demonstrating that MAS evaluations must move beyond aggregate performance to carefully analyze interaction dynamics. Finally, we show that simple approaches to increase diversity—leveraging sampling stochasticity and mixed-models MAS—fail to reduce homogenization risks, highlighting the need for strategies to effectively leverage agent diversity.

[LG-71] Backdooring Acoustic Foundation Models for Physically Realizable Triggers RAID2026

链接: https://arxiv.org/abs/2610.09819
作者: Zebin Yun,Eyal Ronen,Mahmood Sharif
类目: ound (cs.SD); Machine Learning (cs.LG)
*备注: Accepted at RAID 2026

点击查看摘要

Abstract:Acoustic foundation models (AFMs) have democratized acoustic applications, enabling powerful models for tasks ranging from speech recognition to speaker verification with minimal resources. However, the security of applications based on AFMs remains largely underexplored. Our work addresses this gap by proposing the Foundation Acoustic model Backdoor (FAB) attack, demonstrating that state-of-the-art AFMs are susceptible to backdooring under practical settings. Despite making minimal assumptions about adversary capabilities (e.g., no access to pre-training data), we show that FAB preserves benign performance while inducing backdoors that survive fine-tuning and cause significant degradation across diverse downstream tasks when activated. Notably, FAB utilizes task-agnostic, physically realizable, inconspicuous, and sync-free triggers (e.g., a background siren). We evaluate FAB using two leading AFMs, nine downstream tasks, and four different triggers. We further demonstrate its effectiveness against established defenses and across both digital and physical domains. While extensive end-to-end fine-tuning can mitigate FAB, such a defense is resource-intensive and task-specific. Our work highlights critical risks to AFMs and calls for advanced defenses.

[LG-72] BoT-GRPO: Efficient Process-Reward RL for Reasoning via Bag-of-Token Aggregation

链接: https://arxiv.org/abs/2610.09804
作者: Yingxiang Yang,Weihang Xiao,Zhunxuan Wang,Joshua Flashner,Niresh Agarwal
类目: Machine Learning (cs.LG)
*备注: Published at the COLM 2026 Workshop on Efficient Reasoning

点击查看摘要

Abstract:Reinforcement learning is now central to eliciting reasoning in large language models, while in the popular algorithm Group Relative Policy Optimization (GRPO) every token in a rollout receives the same advantage. We ask how to make process supervision efficient: accelerating convergence and improving final quality without the cost of value networks. We propose Bag-of-Tokens Group Relative Policy Optimization (BoT-GRPO), which extends GRPO to token-level reward models through a length-invariant “bag of tokens” aggregation: it collects all token-level rewards across rollouts, weights each by the inverse of its source sequence length, and computes per-token advantages relative to weighted group statistics. BoT-GRPO is critic-free, and is a drop-in replacement wherever GRPO is used when token-level reward is available. On React front-end code generation, BoT-GRPO reaches 80% compile rate up to 1.9\times faster than GRPO and converges faster than modern GRPO variants (GSPO, DAPO, PURE) while reaching higher final compile and VLM-judged win rates. On a second task, AIME mathematical reasoning, BoT-GRPO delivers absolute Pass@ k gains up to 8.1% over GRPO in half the steps. For both tasks we compare the algorithm’s performance on reasoning vs. non-reasoning base-model families (Qwen2.5-3B, SmolLM3-3B, Phi-4-mini-reasoning). Our experiments also yield a practical recipe for the reward model itself: reward stability matters more than richness: clean, bounded, stable fine-grained signals consistently accelerate learning where noisier alternatives stall.

[LG-73] A Proof-of-Concept Study of Weakly Supervised Labeling of Fine-Grained EEG Components for Artifact Attenuation

链接: https://arxiv.org/abs/2610.09792
作者: Lu Wang-Nöth,Hai Huang,Philipp Heiler,Shuqiong Wu,Liyun Zhang,Helmut Mayer
类目: Machine Learning (cs.LG); Signal Processing (eess.SP)
*备注:

点击查看摘要

Abstract:Electroencephalography (EEG) is highly susceptible to electromyographic (EMG) artifacts, whose temporal heterogeneity and spatial-spectral overlap with neural activity can leave mixed sources after blind source separation. Existing artifact-removal methods are further limited by scarce reliable component-level ground truth: expert annotations are costly and subjective, while no established method provides realistic simulation-based ground truth for EMG contamination in multichannel scalp EEG. To address these limitations, we propose a framework combining a frequency-aware high-dimensional representation with Multi-Instance Learning. The representation unfolds separated components into frequency-resolved intra-components, creating a space in which mixed neural and muscular activity becomes more separable, while the weakly supervised learning formulation enables artifact-likelihood scores for individual intra-components to be learned from epoch-level labels without finer-grained ground truth. The resulting intra-component classifier supports fine-grained EMG artifact detection and score-guided attenuation. Experiments on held-out subjects show that the framework learns informative intra-component scores and reduces artifact-related spectral deviations most clearly for jaw tension, with moderate effects for raising eyebrows and limited effects for frowning.

[LG-74] AdaPS-LiNGAM: Adaptive Predecessor Selection for Linear Non-Gaussian Acyclic Models under Small-Sample Settings

链接: https://arxiv.org/abs/2610.09782
作者: Shun Yanashima,Kentaro Kanamori,Hirofumi Suzuki
类目: Machine Learning (cs.LG); Machine Learning (stat.ML)
*备注:

点击查看摘要

Abstract:Causal discovery becomes particularly challenging when the available sample size is small relative to the number of variables. This challenge also arises in the linear non-Gaussian acyclic model (LiNGAM), an identifiable framework for causal discovery from observational data. DirectLiNGAM estimates a causal order, which arranges variables so that causes precede their effects, by sequentially identifying an exogenous variable and removing its linear effect from the remaining variables. We establish a structural limitation of this procedure: when the number of variables exceeds the sample size, repeated residualization necessarily becomes degenerate before the full causal order can be determined. Our analysis further reveals that each residual can be reconstructed using only a graph-determined subset of variables already placed earlier in the causal order, termed the active boundary. This result motivates AdaPS-LiNGAM (Adaptive Predecessor Selection LiNGAM), which reconstructs each residual directly from the original observations using an adaptively chosen sparse subset of those earlier variables. The same subset-selection principle is also applied to the final pruning step for edge estimation. Experiments on synthetic data demonstrate that AdaPS-LiNGAM provides accurate causal-structure recovery in sample-limited settings and degrades more gradually as the sample size decreases.

[LG-75] Reproducible LLM Inference Benchmarking: A Sequential Isolation Protocol for Regression Testing

链接: https://arxiv.org/abs/2610.09778
作者: Arnold Olympio,Juan Manuel Servera Bondroit,Wael Abdelmalek,Guang Lu,João Carvalho
类目: Performance (cs.PF); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Reproducible benchmarking of Large Language Model (LLM) inference is challenging because repeated measurements can vary with execution and system state. We present the Sequential Isolation Methodology, a controlled benchmarking and regression-testing protocol designed to reduce between-run measurement variance while deliberately varying workload concurrency. We evaluate three representative open-source LLMs on an NVIDIA A100 80GB GPU using vLLM 0.9.1 across six context sizes and eight concurrency levels, with five repetitions per configuration. The final protocol reduces average coefficient of variation (CV) from 15.2% in the least controlled methodology stage to 2.2% under the final protocol; using CV computed across the five repetition-level median (P50) TTFT values per configuration, 113 of 144 configurations (78.5%) achieve CV below 3%. The measurements also show a marked latency transition between 200 and 500 concurrent users on the tested stack and descriptive differences in P99 latency across the three models. We additionally provide an explicit cost break-even model with sensitivity to API pricing. The protocol is intended to provide a stable reference for reproducible comparison and regression testing rather than to predict absolute behavior under uncontrolled production traffic. Infrastructure-as-Code and benchmark scripts support replication of the experimental environment.

[LG-76] Fluctuations of Nonlinear Observables in Mean Field Neural Network Training

链接: https://arxiv.org/abs/2610.09768
作者: Arnaud Descours(UCBL),Geoffrey Lacour(MaIAGE)
类目: Machine Learning (cs.LG); Probability (math.PR); Statistics Theory (math.ST); Machine Learning (stat.ML)
*备注:

点击查看摘要

Abstract:Mean field limits describe the training dynamics of wide neural networks through the evolution of the empirical distribution of their parameters. Although functional central limit theorems characterize the asymptotic fluctuations of this distribution, quantities of practical interest are typically nonlinear observables of the parameter distribution rather than the distribution itself. In this work, we show how these mean field fluctuations propagate to finite dimensional nonlinear observables for shallow neural networks trained by stochastic gradient descent. Working in the weighted Sobolev space in which the limiting fluctuation process is constructed, we apply a functional Delta method under ordinary Fréchet differentiability, without requiring Lions derivatives with respect to the measure variable. We obtain a central limit theorem for the observables and, under a suitable representation of their differentials, an explicit covariance formula inherited from the underlying mean field fluctuation theory. We also study whether prescribed quantities of interest can be recovered from the selected observations. Under a constant rank assumption, we prove that a quantity of interest factors locally through the observation functional if and only if, throughout a neighborhood, the kernel of the differential of the observation is contained in that of the quantity of interest. Thus, a differential condition expressed directly in the ambient Sobolev space yields an exact nonlinear local factorization. These results provide a framework both for quantifying finite-width uncertainty on observable, statistically or physically meaningful quantities and for assessing whether the chosen observations contain the information required to identify them.

[LG-77] Leaner Transformers Can Easily Learn to Cluster NEURIPS2026

链接: https://arxiv.org/abs/2610.09760
作者: Charlotte Park,Kenneth L. Clarkson,Lior Horesh,Takuya Ito,Parikshit Ram
类目: Machine Learning (cs.LG); Machine Learning (stat.ML)
*备注: NeurIPS 2026 accepted paper

点击查看摘要

Abstract:Transformers have in-context learning capabilities, where some known learning algorithms can be executed in the forward pass through the model. Recent work shows that transformers can exactly perform Lloyd’s algorithm for k -means clustering with n points in d dimensions with an embedding size d_\textsfemb = d+k (thus, requiring attention projection matrices of size (d+k)^2 ). In this work, we build upon this result in the following ways: First, we present an equally expressive but smaller transformer that executes Lloyd’s algorithm with embedding size d_\textsfemb = (d + \lceil \log_2 k \rceil) . Next, we train these transformers to learn the clustering algorithms given a distribution of clustering tasks, and theoretically characterize and empirically validate the factors affecting the convergence and in-distribution generalization of learning algorithms based on stochastic gradients. Finally, we probe the general clustering abilities of these learned algorithms (in the form of transformers), and try to understand situations where they succeed and fail.

[LG-78] EntroPrefill: Renyi-Guided Context Pruning with Conditional Stability Guarantees for Retrieval-Augmented Generation

链接: https://arxiv.org/abs/2610.09757
作者: Inbasekaran S
类目: Machine Learning (cs.LG); Machine Learning (stat.ML)
*备注: 16 pages, 0 figures. Theoretical framework and conditional stability guarantees for context pruning; no empirical evaluation or benchmarks are included

点击查看摘要

Abstract:Mid-prefill pruning can reduce the sequence processed by deeper transformer layers, but attention concentration alone does not certify that discarded context is dispensable. We formulate EntroPrefill as a Renyi-guided proposal mechanism coupled to explicit constraints on discarded attention mass. Sink-isolated, regularized head pooling respects grouped-query attention while exposing a quantitative trade-off between specialization and worst-head coverage. We derive a mixture-to-head deletion envelope, a computable upper bound on feasible token removal, and a finite-sample observer guarantee that remains valid when the pruning layer is selected adaptively. We then establish a conditional transformer perturbation bound with explicit sufficient Lipschitz constants and a first-token decision-margin corollary. A counterexample shows why shallow observations alone cannot imply an unconditional future-output guarantee. The systems analysis distinguishes query-head unions, physical page allocation, and KV-transfer payload, and gives an arithmetic break-even condition for pruning. This manuscript is theoretical in scope: it defines the procedure, its assumptions, and its formal limits, but does not report measured acceleration or task-accuracy preservation. Experiments are reserved for subsequent validation of the assumptions, approximation tightness, and end-to-end resource trade-offs.

[LG-79] SoftSEEPS improves ML-based precipitation forecasting

链接: https://arxiv.org/abs/2610.09752
作者: Jost Arndt,Utku Isil,Noelia Otero,Rodrigo Almeida,Wojciech Samek,Jackie Ma
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:In this paper we have developed a differentiable approximation of the well-known SEEPS score, which we name SoftSEEPS. This allows the training of a Machine Learning model to forecast precipitation directly. We test SoftSEEPS on the IMERG dataset (0.1 degree resolution) by training a decoder for precipitation on the latent space of a pre-trained low-resolution forecasting model. Combining SoftSEEPS and RMSE in a joint objective is possible with marginal trade-offs in either metric.

[LG-80] EC-EarthFlow: Probabilistic emulation of daily transient global climate model simulations with flow matching

链接: https://arxiv.org/abs/2610.09715
作者: Kirien Whan,Nikolaj T. Mücke,Karin van der Wiel
类目: Machine Learning (cs.LG); Atmospheric and Oceanic Physics (physics.ao-ph)
*备注:

点击查看摘要

Abstract:We introduce EC-EarthFlow, a generative flow matching model that emulates simulations from the physical climate model EC-Earth3. The model is trained on transient simulations from EC-Earth3 (1950-2166, SSP2-4.5) to predict the day ahead temperature field from the previous days temperature as well as annual mean temperature. Predictions are made auto-regressively with rollout periods of between a month and an extended season. Using only this variable of interest, we are able to reproduce the daily variability, spatial patterns, annual cycle and long-term trend from EC-Earth3 at a substantially lower computational cost than the physical model. We demonstrate that EC-EarthFlow is stable for long inference periods, and that it can learn the physical relationships as simulated in EC-Earth3.

[LG-81] Pretraining Shapes Spectral Structure: Architecture- and Strategy-Conditional Prediction of OOD Robustness in Foundation Models

链接: https://arxiv.org/abs/2610.09709
作者: Sangyoon Bae,Sk Miraj Ahmed,Shinjae Yoo,Jiook Cha
类目: Machine Learning (cs.LG)
*备注: 10 pages, 3 figures

点击查看摘要

Abstract:Can we determine whether a foundation model will generalize out-of-distribution (OOD) before any target data is available? Existing diagnostics require source or target data, which rules them out before a target domain exists. Those that use the weights alone apply one statistic to every architecture, and do not separate robust models from fragile ones. We show the answer is encoded in the spectral structure of pretrained weights. Two forces shape that structure. Architecture determines how information is stored in weight matrices. Pretraining strategy determines what is rewarded. Together they set a spectral geometry that governs OOD robustness. We prove that the OOD accuracy gap is bounded by how tightly the source representations concentrate. A statistic computed from the pretrained weights alone serves as a proxy for that concentration. The direction of that proxy reverses between architecture families. We operationalize it: the direction is stable within one (architecture X strategy) combination, the finest grouping we test, which we call a cell. Pooled over 116 models spanning 7 modalities, a single statistic ranks OOD robustness weakly, because cells of opposite direction cancel. Within a cell, the statistic selected for it orders 92% of model pairs by OOD robustness in-sample. The selection does not leak the target: for each model family outside the matrix we logged the cell, metric and sign before running its OOD evaluation, and the predicted direction held in every case: EEG, genomic and protein. Acting on spectral concentration narrows the OOD gap by 24% at 87.5% ID retention. The diagnostic operates on released weights alone, so OOD robustness becomes checkable at model-selection time, before data or compute is committed to a target domain.

[LG-82] Understanding and Mitigating Token-Pruning-Induced Vulnerabilities in VLMs ICML2026

链接: https://arxiv.org/abs/2610.09703
作者: Shuailong Wang,Xinyu Lyu,Shengming Yuan,Jingkuan Song,Heng Tao Shen,Lianli Gao
类目: Cryptography and Security (cs.CR); Machine Learning (cs.LG)
*备注: Accepted at ICML 2026. 16 pages

点击查看摘要

Abstract:Token-Pruning accelerates Vision-Language Models by removing redundant visual tokens, yet its safety implications remain underexplored. In this work, we present the first comprehensive safety evaluation of Token-Pruning mechanisms and find that: most pruning strategies significantly degrade safety as pruning ratios increase, whereas Query-based Compression shows the opposite, with extreme pruning (up to 99.8%), unexpectedly improves model safety. This sharp contrast prompts a key question: How do different Token-Pruning strategies reshape model safety behavior, and is it possible to enhance safety without sacrificing acceleration? To answer this, we identify an unrecognized mechanism, termed Pruning-Induced Malicious Amplification, where removal of background tokens triggers a side effect: forcing the model’s attention to collapse onto a few retained malicious anchors within the foreground, inadvertently amplifying their toxic semantics under jailbreak. To address that, we propose an inference-time and plug-and-play Safety-Aware Pruning (SAP) mechanism that counteracts such dominance via three steps: (1) identifying malicious anchors, (2) restoring pruned benign tokens, and (3) reallocating excessive attention from malicious anchors to benign tokens. Extensive experiments across three safety and four utility benchmarks demonstrate that SAP mitigates pruning-induced vulnerabilities, i.e., reducing ASR by up to 62%, without compromising efficiency or utility.

[LG-83] Optimal Regret for Online Market Making with Limit Order Book

链接: https://arxiv.org/abs/2610.09691
作者: Maria Elena Vischi,Francesco Emanuele Stradi,Alberto Marchesi
类目: Computer Science and Game Theory (cs.GT); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:We study online learning in market making, where, at each round, a market maker posts bid and ask prices before observing the market price and the private valuation of an incoming trader. In this setting, Maran et al. 2026 introduce a feedback model motivated by limit order books, in which the trader’s valuation is revealed only if no transaction occurs. Assuming that trader valuations are drawn i.i.d. from an unknown distribution while market prices are chosen adversarially, they establish an expected regret bound of \widetilde\mathcalO(T^2/3) . In this work, we improve upon this guarantee by establishing a high-probability regret bound of \widetilde\mathcalO(\sqrtT) . As a warm-up, we first consider the full-feedback setting. We introduce a discretization of the bid-ask space based on two coupled grids and combine it with Hedge to achieve the desired regret rate. Building on these ideas, we then address the substantially weaker feedback induced by a limit order book and develop an algorithm that achieves the same guarantee. Finally, we investigate the limits of learnability in fully adversarial environments, where the valuations may vary arbitrarily as well. Perhaps surprisingly, we show that when both market prices and trader valuations are chosen adversarially, sublinear regret is impossible even under full feedback, thereby motivating our stochastic assumption on the valuations.

[LG-84] NAViLoss: An Underwater Navigation-Aware Dual-Residual Objective for Physics-Consistent Learning

链接: https://arxiv.org/abs/2610.09690
作者: Arup Kumar Sahoo,Itzik Klein
类目: Robotics (cs.RO); Machine Learning (cs.LG)
*备注: 26 pages, 7 figures

点击查看摘要

Abstract:Autonomous underwater vehicles (AUVs) commonly rely on inertial navigation systems (INS) aided by Doppler velocity logs (DVLs) for reliable underwater navigation. Accurate DVL velocity estimation is therefore essential for successful operation. Recent learning-based methods have demonstrated improved DVL velocity estimation, particularly under degraded measurement conditions. However, their training objectives typically rely on conventional regression losses that are highly sensitive to large residuals and corrupted observations. Additionally, they do not explicitly account for the physical consistency and measurement uncertainty associated with the underlying sensing process. To address these limitations, this paper introduces navigation-aware loss (NAViLoss), a robust and uncertainty-aware objective function for learning-based AUV velocity estimation. NAViLoss jointly penalizes the velocity-estimation residual in the navigation-state domain and the beam-consistency residual in the DVL measurement domain. Its bounded formulation limits the influence of large residuals, while an adaptive mechanism regulates the uncertainty in beam geometry. Furthermore, NAViLoss is integrated with a DeepONet architecture to form a novel NAVi-DeepONet model for seamless estimation of an underwater vehicle’s velocity. Lastly, our model is evaluated using approximately 10,000m of semi-synthetic AUV experimental data collected during multiple real-world sea trials. Experimental results demonstrate a 44% improvement in velocity-estimation accuracy compared with conventional and learning-based baselines. These results demonstrate the effectiveness of navigation-aware and uncertainty-adaptive loss design for robust learning-based underwater velocity estimation.

[LG-85] Gauss-Newton Accuracy and Indefinite Hessians: Uniform Coexistence in Low-Cost Sets

链接: https://arxiv.org/abs/2610.09675
作者: Kihun Rhee,Hanjoon Byun
类目: Machine Learning (cs.LG); Machine Learning (stat.ML)
*备注: 28 pages, 1 table

点击查看摘要

Abstract:We study the accuracy of Gauss-Newton curvature in ridge-regularized nonlinear least squares. Under local regularity and persistence of level-set curvature magnitude along an exact-fit section, we prove uniform coexistence of two curvature regimes. Global minimizers exist, and every global minimizer has relative Hessian error below (1+\sqrt2)/8 , while the same low-cost set contains a point with an indefinite Hessian and relative error at least 15/8 . One positive ridge cap works for all independent center and label perturbations in fixed neighborhoods and every positive ridge weight up to the cap. These neighborhoods do not shrink as the ridge weight tends to zero. A pointwise certificate based on the current prediction level set controls the normal, mixed, and tangent parts of the Hessian correction. We prove a sharp relative-error bound over the stated pointwise class when the prediction map and ridge vary. Analytic examples describe the roles of output alignment, curvature orientation, and persistence. A separate structural result gives full Jacobian row rank throughout low-cost sets and exact interpolation near a rank-deficient reference.

[LG-86] DSTNet: Dynamic Spectral Trajectory Network for Causal Multi-Horizon Financial Forecasting

链接: https://arxiv.org/abs/2610.09654
作者: Aashish Bohra,Lokendra Vishwakarm
类目: Machine Learning (cs.LG); Computational Finance (q-fin.CP); Statistical Finance (q-fin.ST)
*备注: 27 Pages, 8 figures, 19 tables, Paper in Review

点击查看摘要

Abstract:Wavelet-based financial forecasters typically use the transform only to denoise, or reduce it to a single spectral snapshot at the forecast origin, and the convolution that produces the coefficients is usually bilateral, so it can read past the forecast origin. DSTNet instead retains the recent evolution of filter-bank magnitudes as a causal Dynamic Spectral Trajectory, built from seven trailing technical indicators over a twenty-day lookback with a one-sided Morlet-derived filter bank and an explicit burn-in for the left-boundary transient. A factorized Scale-Temporal Spectral Transformer attends along the time and filter-bank axes separately, a learned gate fuses the spectral branch with a CNN-BiLSTM, and horizon-specific gates emit one, three, five, and ten day forecasts in a single pass. We evaluate seven equity indices and gold under a common expanding-window protocol and an untouched one-year hold-out, against nine learned baselines and a random-walk persistence benchmark. Under MAE and MAPE, persistence is the strongest of the ten fixed competitors in 29 of the 32 series-horizon cells and DSTNet is the only model below it in every cell, by 0.7 to 0.9 percent at one day and 3.4 to 4.5 percent at ten days. At one day, paired testing favours DSTNet against the weaker learned baselines but is inconclusive against persistence and the strongest learned forecasters. A downstream allocation diagnostic does not support an equity-timing advantage on any of the seven indices.

[LG-87] Closed-Form Noise Calibration Against Membership Inference for Random-Allocation DP-SGD

链接: https://arxiv.org/abs/2610.09651
作者: Murat Bilgehan Ertan,Marten van Dijk
类目: Machine Learning (cs.LG); Cryptography and Security (cs.CR)
*备注:

点击查看摘要

Abstract:DP-SGD protects training data by adding Gaussian noise to clipped gradients. The amount of noise is usually chosen by running a numerical privacy accountant inside a search. We study DP-SGD with random allocation, where each epoch uses every record once, at a randomly chosen step. For this setting we give a one-line formula that bounds the accuracy of every membership inference attack (MIA) on the trained model. With M steps per epoch, E epochs and noise multiplier \sigma , and with membership and non-membership equally likely a priori, the attack accuracy is at most \frac12+\frac14\sqrt(1+(e^1/\sigma^2-1)/M)^E-1 . The formula comes from the chi-square divergence between a Gaussian distribution and a Gaussian mixture that dominates random allocation. It is interpretable and gives \sigma in about a microsecond. Where applicable, our formula needs at most about half the noise of the state-of-the-art closed-form bound. To measure how close the bound is, we also derive an exact expression for the attack accuracy of these two distributions and evaluate it numerically. Calibrating to this exact expression requires 13.0% to 20.2% less noise than the formula in our main experiments, and since it is exact, no accountant that knows only M , E and \sigma can certify a smaller \sigma . In training, the resulting \sigma outperforms the formula and matches a published accountant in test accuracy. It is found in seconds and certified in minutes, whereas every search we ran with that accountant took longer or returned at least 0.62% more noise. We show that MIAs on the trained models stay below the bound.

[LG-88] racing Inputs Verifying Outputs: Validating Attribution in Music Generation

链接: https://arxiv.org/abs/2610.09637
作者: Taejun Kim,Wonil Kim,Jongmin Jung,Hyeongseok Wi,Sangeun Kum,Keunhyoung Luke Kim,Taehyoung Kim,Dongjoo Moon,Seungsoon Park,Taewan Kim,Virginie Berger,Juhan Nam,Jongpil Lee
类目: ound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
*备注: 15 pages, 4 figures. Audio examples: this https URL

点击查看摘要

Abstract:How can we verify whose music contributed to an AI-generated output? This paper demonstrates how input-based attribution can provide verifiable evidence of which audio sources were used in a generation and whether they shaped the output. To do so, we condition the generation solely on audio without any text input, then trace the inputs behind each output, and establish their musical effect. In prompt adherence tests and controlled input swaps, the stems generated by our generator, MixAudio, follow the prompt audio in timbre and the context audio in harmony. Yet these outputs may still reproduce training data not supplied as inputs. We therefore audit memorization with our musical version identification model, musicDNA, and find few reproductions outside the input records. On human-judged cases within the flagged pool, it achieves higher precision and recall than the other tested memorization detectors. The two evaluations suggest that input records and output analysis provide complementary evidence for attribution, on which rights-holder reporting and compensation can draw as the AI music economy takes shape. Audio examples are available at this https URL

[LG-89] When does a networks training history predict its future learning better than its current state? Evidence from a response probe and a forecasting screen

链接: https://arxiv.org/abs/2610.09621
作者: Martin Hofmann,Patrick Mäder
类目: Machine Learning (cs.LG)
*备注: 18 pages, 6 figures, 5 tables

点击查看摘要

Abstract:Networks that behave alike now can still learn differently when training continues. Work on loss of plasticity and critical periods shows that the path to a state shapes what follows; it does not show whether the path carries information that a measurement of the state itself misses. We ask when the training history of a network predicts its future learning better than its current state. In a main study, small multilayer perceptrons were trained under three history regimes (42 histories), and future learning was measured at four checkpoints by a short probe: a copy of the network trained for 100 updates on a new task. Before the prediction result was read, the protocol checked the probe. It responded monotonically to a function-preserving rescaling of hidden units, repeated measurements agreed (intraclass correlation 0.940, [0.903, 0.997], in the least reliable class, mean of three repeats), and a re-initialisation of units was visible directly after it but not 100 to 200 updates later. A history state of at most four dimensions did not improve on a calibrated model of the current state (gain -21.4%, 90% interval [-91.9, 8.1]; required in advance: 10%). A companion screen on 1,560 synthetic regression runs asked the same question for a target further away, the final error of the run. There, history models forecast better than the current validation error after 12 of up to 240 epochs (compact state 30.3%, [15.8, 39.4], a contextual comparison) and were not distinguishable from it after 48. In both studies the history was informative only while the current state was not yet informative about the target; this reading was formed after the results.

[LG-90] he Identifiability and Observability of Deep Normalized Attention

链接: https://arxiv.org/abs/2610.09620
作者: Pranav Venkata Konda
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:We study which parameters of deep, unmasked, single-head attention are determined by its input–output function. For known positive nonconstant real-analytic normalizers, the function generically determines the effective scores and combined value map up to the signs induced by even normalizers. This proves the real-analytic case of a conjecture of Henry–Marchetti–Kohn, including softmax. We then classify exceptional fibers under explicit normalizer conditions, identifying when collapse makes later scores unobservable, and establish sharp Taylor orders for local identification. Near simultaneous query/key collapse, we compute the complete native Jacobian decay spectrum on separating finite input banks. For common first nonconstant normalizer degree k , layer i has contact order 2k3^i-1-1 , with exact multiplicities and kernel dimension. High-precision and automatic differentiation calculations illustrate the resulting loss of numerical sensitivity.

[LG-91] Sequential Pretraining Favors Large Models

链接: https://arxiv.org/abs/2610.09611
作者: Mohnish Harwani,Yujia Zheng
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Large neural networks often acquire capabilities that small models fail to learn. Does this stem from large models learning more representative features, or from being more robust to unaccounted-for adverse effects introduced during training? We define and quantify one such adverse effect, primacy bias, as the extent to which exposure to early data distributions impairs later learning. We show that small models can allocate learning capacity inefficiently toward early distributions, whereas sufficiently overparameterized models are robust to this effect. This inefficiency is particularly consequential in pretraining, where foundation models often encounter heterogeneous data distributions sequentially rather than jointly. As a result, small foundation models can struggle to learn distributions encountered late in training, which is particularly harmful when later data emphasizes desirable capabilities such as code, mathematics, and reasoning. Motivated by these findings, we introduce Exposure Therapy (ET), a simple regularization that promotes more efficient allocation of learning capacity during sequential pretraining. We demonstrate that ET improves foundation models’ performance on late data distributions as well as overall capability in models up to the billion-parameter scale. Overall, our results suggest that some benefits of large foundation models may arise from greater robustness to adverse training effects, rather than from learning more representative features, and that improved training algorithms can recover some of these advantages in smaller models.

[LG-92] Decoupled Optimization for Teacher-Student Semi-Supervised Learning via a Pioneer Student

链接: https://arxiv.org/abs/2610.09609
作者: Haorong Han,Jidong Yuan,Chixuan Wei,Yongqi Sun
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Semi-supervised learning (SSL) relies on two core mechanisms: self-training under the Teacher-Student (T-S) framework and joint optimization of labeled and unlabeled losses. Despite their effectiveness, we find both mechanisms introduce distinct optimization pathologies. First, parameter coupling enforces strict synchronization between teacher and student, where strong regularization on the student degrades the teacher’s fitting ability, thereby limiting the permissible generalization intensity. Second, the imbalance in gradient update consistency between labeled and unlabeled losses drives the shared parameters to prematurely converge to labeled-dominated local minima, creating a bottleneck for global optimization. To address both issues, we propose the Pioneer Student (PiS), an auxiliary branch that operates in an independent parameter space and periodically transfers accumulated knowledge back to the T-S model. Extensive experiments show that PiS is a universal plug-and-play module that consistently improves mainstream SSL methods.

[LG-93] Lightweight and Versatile Learned Optimization by Recombination of Gradient History

链接: https://arxiv.org/abs/2610.09604
作者: Minyoung Choi,Dalta Imam Maulana,Wanyeong Jung
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:This paper presents a lightweight and versatile learned optimizer that dynamically recombines gradient history, represented as averages over disjoint time spans. The optimizer reduces the prediction space to one scalar coefficient per gradient average, shared by multiple parameters. Progressively averaging older gradients minimizes memory cost of long history, while keeping their contributions independently accessible. A 37k-parameter network trained in 0.87 GPU-hours generalizes zero-shot to unseen tasks, lowering validation loss by 9.1% and 0.4% on BERT-Tiny and GPT-Tiny, and improving test accuracy over Adam by 3.5 %p on a Vision Transformer and by 2.7 %p on average across nine graph models, with FLOPs overhead as low as 0.3%.

[LG-94] COPC: Coupled Off-Policy Correction for Asynchronous LLM Reinforcement Learning

链接: https://arxiv.org/abs/2610.09597
作者: Zicheng Hu,Zhijian Zhou,Xuan Zhang,Yuchen Liu,Cheng Chen,Yuan Li,Qi Gu,Yan Feng,Hongyan Hao,Chao Qu
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Asynchronous RL accelerates large language model post-training by decoupling rollout generation from optimization, but trains on stale trajectories. Existing methods primarily correct token-level policy mismatch through importance-ratio control in the actor objective. We show that this \emphpolicy-side correction alone is insufficient: advantage estimates also inherit mismatch from behavior-policy continuations, which we term \emphadvantage staleness. We derive exact bias and variance decompositions for a general two-channel actor update, revealing nonseparable coupling between policy-weight and advantage-estimation errors: their interaction induces multiplicative bias terms, while squared policy weights amplify advantage uncertainty in gradient variance. This motivates the hypothesis that policy- and advantage-side correction should be coordinated. We introduce Coupled Off-Policy Correction (COPC), an actor–critic method combining token-level ratio masking with two-sided clipped-ratio weighting of TD residuals for return and advantage estimation. Joint parameter sweeps across staleness levels support this hypothesis: the effect of one correction parameter depends on, and can reverse with, the other. COPC achieves the highest reported performance on tool-integrated mathematical reasoning and search, outperforming the strongest reported asynchronous baseline in each setting. It also offers a broad high-performing parameter region and improved training stability. In search, COPC remains stable throughout training, while most evaluated asynchronous baselines collapse late in training. These gains persist at 64-step policy staleness. COPC adds minimal step-time overhead over asynchronous PPO and retains a 1.7\times step-time speedup over synchronous PPO.

[LG-95] Online Resource Allocation with an Endogenous Markov State: Fewer LP Solves Earn More

链接: https://arxiv.org/abs/2610.09577
作者: Zhaohua Chen
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:We study finite-horizon online resource allocation with i.i.d. requests and an endogenous Markov state on a finite state space: each action affects the transition of the state that governs future rewards and resource consumption. In this problem, a transient fluid LP benchmark upper bounds the expected reward of every nonanticipating policy, while a stationary LP supplies randomized state-dependent controls. We assume that the stationary LP has a unique optimum and identify primal nondegeneracy and irreducibility of the optimal induced kernel as important regularity conditions in this framework. With a known request prior, we show that, under nondegeneracy and irreducibility, both frequent and infrequent re-solving attain O(1) regret. However, under a degenerate optimum, irreducibility yields the sharp worst-case \Theta(\sqrtT) rate for infrequent re-solving, while frequent re-solving can incur \Omega(T) regret. Thus, more frequent optimization can perform asymptotically worse. With an unknown request prior, we develop a three-phase U-shaped infrequent re-solving policy that coordinates learning and inventory correction with O(\log\log T) LP solves. When the optimal induced kernel is irreducible and the algorithm is given the optimal target state class and a constant-cost entrance policy, it attains O(1) regret under nondegeneracy and O(\sqrtT) regret under degeneracy. Without the target-class information, linear minimax regret is unavoidable. Numerical experiments further illustrate the instability of round-by-round re-solving relative to epoch-wise infrequent re-solving, show that thresholding greatly mitigates its loss, and find that infrequent schemes remain dominant under both known and estimated priors.

[LG-96] EvoSignal: LLM -Guided Evolutionary Design of Modular Traffic Signal Control Programs

链接: https://arxiv.org/abs/2610.09563
作者: Leizhen Wang,Peibo Duan,Zhenlin Qin,Yancheng Ling,Jian Xu,Yue Wang,Hao Wang,Zhenliang Ma
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Effective traffic signal control (TSC) requires policies that respond to changing traffic demand and network conditions while meeting different control objectives. However, adapting existing strategies often involves repeated manual design and adjustment, making it difficult to systematically explore better control rules for a target network. Large language models (LLMs) can automate this process, but directly using them to select signal phases leaves decision rules embedded in black-box models and incurs recurring inference costs and latency. This paper formulates TSC as a modular program design problem and proposes EvoSignal, an LLM-guided evolutionary framework using traffic knowledge and performance feedback. The modular representation separates traffic feature extraction, local phase prioritization, and optional network-based priority adjustment. Starting from several established strategies, EvoSignal improves programs through feedback on congestion and signal operation, retaining strategies with different performance trade-offs. The resulting programs operate without online LLM inference. Simulation experiments across five scenarios on two real-world road networks show that the selected default EvoSignal program reduces waiting time by 16.8–49.2% relative to the lowest waiting time achieved by the 20 conventional, reinforcement learning-based, and LLM-based baselines in each scenario. A program prioritizing travel time and queue length outperforms all 20 baselines on all three metrics in the search scenario and remains among the top three on each metric when transferred unchanged to the other four scenarios. These findings support automated design of inspectable control programs that transfer across the evaluated road networks and traffic this http URL is available at this https URL.

[LG-97] UniCSI Towards a Universal Wi-Fi CSI Encoder for Ubiquitous Human Sensing

链接: https://arxiv.org/abs/2610.09559
作者: Daniel Eckhoff,Hua Kang,Zhitang Chen,Jie Chuai
类目: Machine Learning (cs.LG); Networking and Internet Architecture (cs.NI)
*备注:

点击查看摘要

Abstract:Wi-Fi sensing promises to turn the everyday wireless signals that already surround us into ubiquitous sensors for human sensing. However, a fundamental obstacle is that CSI is acquired under diverse device-specific configurations, including different subcarrier counts, bandwidths, and carrier bands. Consequently, the resulting CSI tensors vary in both spectral resolution and tensor shape, making heterogeneous modeling challenging. Standard architectures struggle with such heterogeneity, forcing lossy pre-processing which compromises the underlying signal. To bridge this gap, we present UniCSI, a unified foundation architecture that directly operates on heterogeneous CSI while preserving the integrity of the native waveform. UniCSI hinges on two core innovations: (1) a physics-informed RF tokenizer that encodes each frequency channel based on its fractional position within the physical spectrum rather than rigid array indices. It preserves intrinsic spectral coherence and enables seamless, frequency resolution-agnostic processing across arbitrary sensing configurations. (2) a spectral aggregator that distills variable-length channel sequences into a fixed-size spectral signature, effectively decoupling the feature dimensionality from the physical subcarrier spacing. Extensive evaluations on a large-scale corpus of 25 heterogeneous public datasets, spanning 14 to 2048 subcarriers, 20 to 160 MHz bandwidth, and the 2.4 and 5 GHz bands, demonstrate that native heterogeneous ingestion substantially improves cross-domain transfer under both supervised and self-supervised training schemes, particularly in regimes where fixed-grid architectures fail to generalize.

[LG-98] A Framework for the Systematic Review of ML Assets in AI Registries

链接: https://arxiv.org/abs/2610.09551
作者: Alexandra González,Quim Motger,Xavier Franch,Silverio Martínez-Fernández
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Background: Modern software systems increasingly rely on Machine Learning (ML) assets (i.e., pre-trained models, datasets, benchmarks) for building, evaluating, and integrating ML-based systems. However, current exploration, selection and reuse practices of ML assets are not supported by systematic retrieval methodologies comparable to those used in traditional evidence synthesis. Consequently, in practice, ML asset selection is often presented as a settled design decision, supported by informal justification rather than a traceable, evidence-based, and updatable selection process. Aims: This paper explores how systematic review methods can support ML asset retrieval. In doing so, we aim to make their selection transparent and reproducible, grounded in explicit evidence, and ultimately better suited to its intended use. Method: We analyze established systematic review practices from scientific literature and adapt their phases (i.e., planning, conducting, and documenting) to Artificial Intelligence (AI) registries, treating ML assets as first-class units of analysis. The resulting framework integrates registry-aware search strategies, cross-registry schema alignment, and dependency-driven ML asset exploration. Results: We conceptualize ML asset retrieval as a systematic and reproducible process rather than an ad hoc activity, and propose a framework for structured ML asset discovery. \textbfConclusions: This work illustrates how systematic review principles can be extended beyond scientific literature to support evidence synthesis over evolving AI registries.

[LG-99] CircuitGate: Logic-Consistent Circuit-Level Functional Modeling for And-Inverter Graphs

链接: https://arxiv.org/abs/2610.09549
作者: Qifan Zhang,Ruijie Li,Fangzhou Zhang,Qian Ma,Hui Li,Furui Zhan,Yongpeng Wang,Liying Hao,Shikai Guo
类目: Machine Learning (cs.LG)
*备注: 16 pages, 3 figures, 11 tables. Qifan Zhang and Ruijie Li contributed equally. Qian Ma is the corresponding author

点击查看摘要

Abstract:And-Inverter Graphs (AIGs) are fundamental representations for logic synthesis and verification in Electronic Design Automation (EDA). As structured representations of complex digital systems, AIGs require models to capture functional dependencies beyond local structure and remain robust to functionality-preserving transformations. In learning-based AIG representation, existing approaches are predominantly based on GNNs and rely on local gate-level message passing, limiting their ability to capture circuit-level functional context and making the learned representations sensitive to topology-specific patterns. Therefore, we propose CircuitGate, a function-aware AIG representation learning framework that advances from gate-level semantics to circuit-level functional modeling. CircuitGate explicitly encodes global primary-input (PI) support and models support-overlap-aware reconvergence between fanins, while incorporating logic-inspired Boolean constraints to encourage functionally consistent representations. We evaluate CircuitGate on the large-scale ForgeEDA benchmark and further validate it on the EPFL and ITC’99 benchmarks. Across equivalent-gate identification and signal-probability prediction tasks, CircuitGate consistently outperforms existing methods, achieving up to 21.7% and 14.2% reductions in MAE, respectively. Under direct ForgeEDA-to-OpenABC transfer without fine-tuning, CircuitGate also achieves the best equivalent-gate identification performance, demonstrating strong cross-dataset generalization. These results demonstrate the effectiveness of modeling circuit-level functional dependencies beyond local topology.

[LG-100] Differential Refresh Policies for Models Trained on Lagging Data Snapshots: From a Single-Age Equivalence Limit to an Optimal Per-Segment Allocation

链接: https://arxiv.org/abs/2610.09519
作者: Amit Rajula
类目: Machine Learning (cs.LG); Distributed, Parallel, and Cluster Computing (cs.DC)
*备注:

点击查看摘要

Abstract:Production machine-learning models are derived artifacts of time-bounded training snapshots: a deployed model is a materialized view over a training cut that ages the instant it is built. A common response is to replace the fixed retraining cadence with an adaptive trigger – a weighted staleness score that retrains when accumulated source risk crosses a threshold. We show this is the wrong lever, and identify the right one. First, an equivalence limit: any refresh trigger that is a static, strictly monotone function of a single shared global training-data age is operationally equivalent to a calibrated uniform age timer, so a global staleness budget, however elaborately it weights segments, sources, and sensitivities, carries no scheduling information a clock does not. The limit also shows how to escape it: refresh segments differentially, giving each its own age and refresh interval, which is meaningful when refresh cost is separable across segments (incremental training or per-segment models). We solve the resulting budget-allocation problem. In the frequent-refresh regime each segment’s optimal refresh rate is proportional to the square root of its risk w_j \lambda_j (weight times change rate), and the optimal policy never costs more than the uniform timer, beating it by a closed-form Cauchy-Schwarz “price of uniformity” that is zero for homogeneous workloads and grows with heterogeneity. In a discrete-event simulation with real Poisson change events, the optimal policy lowers realized weighted stale exposure by 8-29% relative to the uniform timer at matched refresh budget, winning on 86-100% of seeds; a naive exposure-threshold policy does not, showing the allocation is what helps; and the advantage survives 50% rate-estimation noise. The leverage in model refresh is not a better score but a better action.

[LG-101] Physics-Informed Neural Plasticity: PDE Solvers That Reshape Themselves

链接: https://arxiv.org/abs/2610.09510
作者: Chun-Wun Cheng,Bingcheng Hu,Angelica I. Aviles-Rivero
类目: Machine Learning (cs.LG); Numerical Analysis (math.NA)
*备注:

点击查看摘要

Abstract:Physics-informed neural PDE solvers adapt their parameters to satisfy governing equations, yet their representational structure typically remains fixed throughout training. This rigidity is poorly matched to PDE solutions with strongly heterogeneous complexity across space and space–time, leaving capacity insufficient where the physics is difficult and redundant where it is simple. We introduce physics-informed neural plasticity, a paradigm in which the representation itself reshapes during optimization in response to unresolved physics. We instantiate this principle with Representation Capacity Adaptation for PDEs (ReCAP), a Gaussian-localized solver that dynamically redistributes capacity through local enrichment, residual-directed splitting, gate-based pruning, and function-aware merging. ReCAP uses responsibility-weighted error indicators and the geometry of residual energy to determine where and how to refine. To limit the disturbance introduced by splitting, we introduce quiet-child refinement, which initializes new components by transporting the parent representation while controlling instantaneous functional perturbation. We further establish conditional a posteriori reliability and structural-stability guarantees linking localized physics residuals to solution error and stable refinement. Across five challenging 3D and 4D PDE benchmarks against 11 physics-informed solvers, ReCAP achieves the lowest relative L^2 error on every problem, reducing error by 10.7% – 27.5% relative to the strongest competing result. These results suggest that physics-informed solvers need not merely learn their parameters—they can learn how their representational capacity should be organized.

[LG-102] When Should an In-Context Learner Expand Its Hypothesis Space?

链接: https://arxiv.org/abs/2610.09471
作者: Weihan Li,Xinlei Chen,Junhao Wu,Tianshi Zheng
类目: Machine Learning (cs.LG)
*备注: 43 pages, 11 figures

点击查看摘要

Abstract:Learning systems adapt quickly inside a familiar family of models. The harder step comes earlier: deciding, from observations that could be noise, an exception, a change within the family or structure outside it, whether opening a richer family is worth its cost. We treat this as a costly sequential decision: prediction failure must be turned into structural evidence, evidence into a value of expansion, and value into action. The Structural Revision Environment produces matched failures from each source, varies the price of expansion and the remaining horizon independently of the evidence, and admits exact Bayesian calculations and an exact normative solution of the one-shot decision. Its solution shows that revision is a value boundary and not an evidence threshold: one history has different optimal actions under different prices, horizons and announced queries, the boundary between local repair and expansion is set by the inputs a rule predicts and a repair cannot cover, and belief in the richer family crosses long before the decision does. Transformers trained in the environment reproduce this boundary from utility alone. Language models of three post-training lineages carry a failure-sensitive signal in their predictions that is not reflected in their revision decisions, and given the gain of expanding they read it without weighing it against price and horizon. Three models allowed to reason weigh the stated gain in the reference’s proportions and still do not turn the history into an estimate of what expansion would buy. Controlled post-training of the meta-trained learners moves the prior and the sharpness of predictions, and neither moves the criterion.

[LG-103] An extended deep energy method for thermo-mechanical crack propagation

链接: https://arxiv.org/abs/2610.09433
作者: Han Zhang,Mehrisadat Makki Alamdari,Babak Shahbodagh,Mohammad Vahab,Cosmin Anitescu,Timon Rabczuk,Elena Atroshchenko
类目: Machine Learning (cs.LG); Materials Science (cond-mat.mtrl-sci); Numerical Analysis (math.NA)
*备注: 48 pages, 20 figures, 10 tables

点击查看摘要

Abstract:Thermo-mechanical fracture couples transient heat conduction on a cracked domain with a crack that grows as the temperature and the displacement evolve. Neural energy solvers have been proposed for phase-field fracture and later extended to represent a sharp crack through the network input, but heat conduction on the cracked domain and crack propagation under the resulting thermal stresses have not yet been treated together in these solvers. We present an extended deep energy method for thermo-mechanical crack propagation in which the crack remains a sharp polyline. Two networks represent the temperature and the displacement and receive the crack through a scalar embedding function, discontinuous across the crack and smooth elsewhere, so that both fields can jump across it without a regularization length, and the displacement is enriched near the tip by the Williams expansion with trainable amplitudes. The two fields are obtained by minimizing an incremental conduction functional and the thermoelastic potential energy in a staggered sequence, with Monte Carlo integration on points stratified over background elements, densified near the tip and redrawn during training. The stress intensity factors are extracted by the interaction integral with the area term of Wilson and Yu and checked by a sweep of the contour radius, and the crack advances at the maximum hoop stress angle when the energy release rate of the kink reaches the critical value at the crack-tip temperature. On a stationary thermal edge crack the extracted stress intensity factor agrees with the published value to 0.11%, in a functionally graded shear test initiation agrees with an independent sharp-crack finite element solution to within one load step, and on a notched cruciform specimen the crack paths follow the published solutions under mechanical, thermal and combined loading.

[LG-104] What a Reporting Convention Hides: A Matched-Budget Audit of Quantum Natural Gradient with an Exactly Computed Metric

链接: https://arxiv.org/abs/2610.09425
作者: Lu Wei,Yufeng Wang,Haibin Ling
类目: Machine Learning (cs.LG); Quantum Physics (quant-ph)
*备注: 29 pages, 6 figures, 12 tables. Lu Wei and Yufeng Wang contributed equally

点击查看摘要

Abstract:Several published comparisons of variational quantum optimizers time only runs that reach a target loss, or read the verdict at a single target. Either convention could decide whether an optimizer’s costlier steps pay off. We measure how much each convention changes verdicts among Adam, simultaneous perturbation stochastic approximation (SPSA) and quantum natural gradient (QNG), on initializations held out from the selection of settings. We compute the exact metric that preconditions QNG, price every step in circuit evaluations and give every method the same budget. In a median pooled over circuit widths, cost families and a sweep of settings with common misses, SPSA needs more than twice Adam’s evaluations to reach a loose target. Dropping the censored runs that miss the target hides this gap. On the global-cost family we hold fixed the settings selected for a strict target. QNG then usually reaches the loose target after Adam but the strict target first. QNG’s strict-target lead disappears when the metric’s simulator price, linear in the number of parameters, is replaced by an assumed hardware count quadratic in that number. We recommend charging every miss the budget and reporting verdicts across targets.

[LG-105] Democratizing MoE inference on commodity GPUs with CoMoE

链接: https://arxiv.org/abs/2610.09424
作者: Ruwen Fan,Yuezhi Zu,Junru Li,Qingda Hu,Xinjun(Jimmy)Yang,Jiwu Shu,Youyou Lu
类目: Distributed, Parallel, and Cluster Computing (cs.DC); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Deploying Mixture-of-Experts (MoE) models relies heavily on Expert Parallelism, which generates intense inter-GPU communication. Consequently, state-of-the-art inference systems require high-bandwidth, P2P interconnects (e.g., NVLink) in datacenter GPUs to handle massive token routing, making deployment prohibitively expensive. Consumer GPUs offer comparable compute power at significantly lower cost, promising to democratize MoE inference for individuals and enable privacy-preserving local deployments. However, their bandwidth-limited (only weak PCIe bus bandwidth) and host-mediated interconnects (no P2P support) introduce severe communication bottlenecks. We present CoMoE, a communication-efficient MoE inference system that resolves this mismatch through novel host-centric routing. Our key insight is that the unique communication topology provides the opportunity to elevate the host to an active routing hub, which can fundamentally reduce communication volume and eliminate global synchronization-induced stalls. Specifically, for token dispatch, we introduce host-backed token multicast to write shared tokens to the host exactly once, eliminating outbound transmission redundancy. For token combine, we propose a fine-grained, token-level aggregation mechanism using host staging buffers, which replaces rigid global synchronization and mitigates straggler effects. Evaluation on RTX 5090 GPUs shows that CoMoE improves inference throughput by up to 1.46x, approaching the performance of NVLink-capable A800 GPUs at only 23.4% of the hardware cost. Subjects: Distributed, Parallel, and Cluster Computing (cs.DC); Machine Learning (cs.LG) Cite as: arXiv:2610.09424 [cs.DC] (or arXiv:2610.09424v1 [cs.DC] for this version) https://doi.org/10.48550/arXiv.2610.09424 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[LG-106] Self-Consuming Generative Models with Co-Evolving Human Preferences NEURIPS2026

链接: https://arxiv.org/abs/2610.09415
作者: Xiukun Wei,Tian Xie,Ding Zhu,Xueru Zhang
类目: Machine Learning (cs.LG)
*备注: Published as a conference paper at NeurIPS 2026

点击查看摘要

Abstract:Generative models are increasingly trained in self-consuming iterative loops, where users curate preferred samples from model-generated candidates and the curated samples are used to train future generations of the model. Prior work has largely assumed fixed user preferences, but in practice exposure to model outputs gradually reshapes what users perceive as desirable, creating a feedback loop in which model distributions and user preferences co-evolve. We take a first step toward understanding the long-term behavior of such coupled dynamics. We show that when training relies entirely on user-curated synthetic data, iterative curation amplifies initial biases and drives the system toward one of multiple singleton equilibria in which the instance holding an initial advantage eventually dominates. In contrast, injecting reference data into training at a sufficiently large rate fundamentally changes the dynamics and yields a unique globally attracting equilibrium. Building on this insight, we study how reference-data injection can be used to control long-term outcomes, and propose an efficient algorithm that jointly selects a reference distribution and its mixing weight to steer the coupled system toward equilibria that preserve desired attributes while minimizing data collection costs.

[LG-107] Stability and Diversity of Networked Self-Consuming Generative Ecosystems NEURIPS2026

链接: https://arxiv.org/abs/2610.09409
作者: Xiukun Wei,Yang Zhang,Xueru Zhang
类目: Machine Learning (cs.LG); Computers and Society (cs.CY)
*备注: Published as a conference paper at NeurIPS 2026

点击查看摘要

Abstract:The widespread deployment of generative AI has made it increasingly difficult to distinguish synthetic content from real data. Consequently, synthetic data is inevitably incorporated into the training pipelines of future model generations, forming a self-consuming training loop. Prior work has studied the effects of such recursive self-consuming training, but analyses have largely been limited to isolated models, where a model consumes only its own synthetic data, or to simplified interactions between two models. This paper takes a first step toward understanding networked self-consuming generative models, in which multiple models consume synthetic data generated by one another through complex interaction pathways. We introduce a theoretical framework representing models as nodes in a directed, weighted graph, with edge weights governing the flow of synthetic data among models. Using this framework, we analyze the long-term behavior of networked models under retraining dynamics, establishing conditions for convergence and characterizing the resulting fixed points. We further investigate how the system’s long-term stability and diversity are shaped by each model’s access to real data, cross-model data consumption, and the structure of the interaction graph.

[LG-108] Noise Denoise Correct: MCMC Posterior Sampling with Diffusion Priors in Three Steps

链接: https://arxiv.org/abs/2610.09407
作者: So Takao,Gregory David Bellchambers,Luke Ye,Sanmitra Ghosh,Michalis Michaelides
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Pretrained diffusion models are powerful priors for inverse problems, but posterior sampling under nonlinear, non-differentiable forward models remain hard. We introduce diffusion waltz, an MCMC method using SDEdit-style noising-denoising as a proposal, corrected via Metropolis-Hastings for exact posterior sampling without prior evaluation. We further propose injecting observations into the proposal while preserving exactness, using a gradient-free ensemble Kalman update. On a non-differentiable Navier-Stokes initial condition recovery task, diffusion waltz outperforms existing baselines across different noise and nonlinearity regimes.

[LG-109] From Retrieval to Customer Context: Evaluating Frontier-Model Systems for Voice-of-Customer Analysis

链接: https://arxiv.org/abs/2610.09375
作者: Raviraja G,Viraj Bagal,Prabhath Chellingi
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Organizations increasingly use frontier language models to analyze customer feedback, but answer quality also depends on how that feedback is organized and made available. We define a \emphcustomer context graph as a unified model of customer and business context. Typed relationships connect customer objects (feedback, conversations, users, and accounts), operational objects (tickets, support agents, opportunities, and competitors), and analytical or action objects (taxonomy concepts, evidence, insights, work items, and outcomes). This lets an agent investigate not only what customers say, but why, who is affected, what action followed, who owns it, and whether it was resolved. For this experiment, the graph is populated from public Cursor feedback; the same architecture can support any type of feedback source. We compare Agentic RAG, a Deep Research Agent, and a Customer Context Graph-backed Agent on the same 9,432 public Cursor feedback records using 30 realistic product, incident, comparison, and metadata questions. Without exhaustive ground truth, we jointly score responses on answer quality (coverage and organization), analytical depth (specificity and decomposition), and evidence quality (citation support and traceability), using a comparative rubric calibrated on 28 of the 30 questions. We sample cited records against their claims and weight the three dimensions equally. Under this aligned rubric, the Customer Context Graph-backed Agent scores 0.961 overall, versus 0.710 for the Deep Research Agent and 0.651 for Agentic RAG, and leads the Deep Research Agent on 27 of 30 paired questions (sign-test p 10^(-5); strictly best on 26 of 30). Its largest advantage is analytical depth (0.967 versus 0.642), reflecting more specific, hierarchically developed findings with quantified themes and traceable evidence…

[LG-110] Immiscible Diffusion Policy: Preserving Multimodal Robot Actions through Label-Free Noise Assignment

链接: https://arxiv.org/abs/2610.09369
作者: Xiao Zhang,Yuxin Chen,Zhixuan Liang,Guojian Zhan,Chenran Li,Chenfeng Xu,Masayoshi Tomizuka,Yiheng Li
类目: Robotics (cs.RO); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:When diffusion policies were first introduced, they were expected to recover multi-modal action distributions. However, we find this expectation does not always hold, as diffusion policies often collapse to a single modality even when we guarantee the balance of dataset modalities and exact within-batch symmetry. Our analysis indicates that independent action-noise pairing contributes to this failure by increasing mixing and crossing among diffusion paths, which can produce averaged denoising responses and suppress modality-specific behavior. This issue is especially severe in robot planning, where action spaces are dense and low-dimensional, significantly increasing such mixing and crossing. To alleviate this problem, we propose Immiscible Diffusion Policy, a label-free training-time add-on to diffusion policy that uses action-noise assignment to preserve relatively distinct noise-to-action routes without modifying the policy architecture or inference procedure. Across five simulated and two real-world humanoid manipulation tasks spanning state, RGB, and point-cloud observations, our method significantly improves the policy’s preservation of action modalities while maintaining strong task performance. It increases the proportion of the non-dominant modality by 6.0x-14.6x across three two-modality tasks and recovers demonstrated modalities that are entirely absent from vanilla policy rollouts on both four-modality tasks. These results demonstrate that Immiscible Diffusion Policy provides a simple yet robust approach to preserving action multi-modality in general robot learning tasks.

[LG-111] Global Exponential Convergence of Two-Layer Linear Network Training

链接: https://arxiv.org/abs/2610.09356
作者: Stephen Y Zhang,Gabriel Peyré
类目: Machine Learning (cs.LG); Optimization and Control (math.OC); Machine Learning (stat.ML)
*备注:

点击查看摘要

Abstract:We prove global exponential (linear) convergence with an explicit rate in the rich scaling for wide two-layer linear networks trained with smooth Polyak-Lojasiewicz predictor losses. Gradient flow in the factors closes exactly in terms of a finite-dimensional Bures flow of the neuron law covariance, in which the predictor dynamics are preconditioned by hidden covariance blocks. Mean-field conservation laws provide uniform spectral lower bounds on the hidden preconditioning blocks when the initial covariance satisfies a spectral support gap condition. This condition encompasses positive definiteness while still allowing for singular initializations. For an initial covariance \Sigma_0 = \sigma^2 \mathrmId , the loss converges to the global minimum with linear rate at least 4\sigma^2\kappa , where \kappa is the PL constant. We establish stability of this rate under finite-width sampling, as well as global convergence of factor gradient descent for an explicit stepsize interval depending on smoothness, the initial loss, and conserved spectral margins. Our argument extends layerwise to deep linear ResNets, subject to a residual-path bound. In the case of heavy-ball momentum, training dynamics close instead over positions and velocities in terms of a lifted phase covariance. Linear convergence holds under an explicit condition on the energy and damping, specifying a window of admissible dampings. For two-scale white initializations, this interval is nonempty for sufficiently large position scales, with a fixed initial loss gap and velocity covariance. Numerical experiments illustrate the covariance geometry and compare the predicted and observed rates.

[LG-112] Learning Unknown Constraints without Unsafe Data via Optimality and Counterfactual Regularization

链接: https://arxiv.org/abs/2610.09350
作者: Zhouyu Zhang,Chih-Yuan Chiu,Glen Chou
类目: Robotics (cs.RO); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Learning from demonstrations (LfD) provides a framework for inferring unknown constraints from locally optimal, constraint-satisfying expert behavior. Existing approaches largely fall into two paradigms, constrained inverse optimal control (CIOC) and inverse constrained reinforcement learning (ICRL). CIOC exploits optimality conditions such as the Karush–Kuhn–Tucker (KKT) conditions but typically assumes known dynamics and structured constraint representations. Meanwhile, ICRL accommodates complex unknown constraints and unknown transition dynamics but often requires extensive online exploration, during which unsafe constraint violations may occur. In this work, we introduce Counterfactual KKT (CF-KKT), a constraint learning framework that leverages learned dynamics and locally optimal demonstrations to recover unknown constraints without requiring known dynamics or additional risky exploration, thereby combining the data efficiency and safety advantages of CIOC with the flexibility of ICRL. First, we use a locally learned differentiable dynamics model to impose KKT-inspired optimality conditions directly on the demonstrations. Second, we use the learned dynamics to generate reward-improving counterfactual behaviors near the demonstrations, revealing behaviors that would be preferable in the absence of the unknown constraint and thus providing synthetic infeasible data. When the constraint parameterization is known, the same learned-dynamics framework enables direct CIOC-based parameter recovery, and we characterize its sensitivity to dynamics misspecification. Across high-dimensional robotic control tasks, our approach learns neural constraint representations with improved safety and data efficiency relative to state-of-the-art offline ICRL baselines.

[LG-113] Benign Overfitting under Heterogeneous Input Fusion

链接: https://arxiv.org/abs/2610.09340
作者: Houzhen Liu,Xiaobo Xia
类目: Machine Learning (cs.LG); Machine Learning (stat.ML)
*备注: 51 pages, 6 figures

点击查看摘要

Abstract:Benign overfitting is extensively studied when learning from a single high-dimensional input, but its behavior under heterogeneous input fusion remains largely unexplored. We study this question for minimum-norm linear interpolation under a heterogeneous Gaussian design, comparing two statistically dependent input blocks with their fusion while holding the underlying population task fixed. For regression, we identify a full-spectrum covariance certificate whose asymptotic status is independent of the cutoff threshold and prove that it is preserved by every positive-semidefinite joint covariance consistent with the two marginals. This protection is sharp, yet it does not extend to all benign regression problems: outside the certified regime, two benign marginals can have a harmful fusion. For one-sparse Gaussian classification, benignity in the regular regime is characterized by the balance between surviving predictive signal and nuisance contamination. Fusion can move these two quantities in opposite directions, and within this model class every marginal-to-joint benign/non-benign pattern is attainable. We further show that the same fused input can have qualitatively different effects on regression and classification. These results establish that benign overfitting under heterogeneous fusion is determined by the joint signal and spectral geometry created by input interaction, rather than by marginal benignity alone.

[LG-114] Ask the Expert: LLM -Guided Reinforcement Learning for Autonomous Cyber Defense

链接: https://arxiv.org/abs/2610.09337
作者: Fernando Martinez,Abhishek Satyam,Tao Li,Junaid Farooq,Ying Wang,Juntao Chen
类目: Cryptography and Security (cs.CR); Machine Learning (cs.LG)
*备注: Accepted for publication at IEEE GLOBECOM 2026. Proceedings forthcoming. 6 pages, 4 figures

点击查看摘要

Abstract:Policy-based reinforcement learning (RL) approaches have produced promising results for autonomous cyber defense; however, they are sample-inefficient in settings where defenders must respond under delayed, partial observations with actions from large action spaces. While large language models (LLMs) may reason semantically about security state space, high latency and trust assumptions prevent attractive in-line deployment models. We introduce Ask the Expert, a training-time guidance framework which first summarizes hard cyber-defense states, then intermittently queries an LLM for host-level defensive recommendations via a constrained action interface, and finally transforms those recommendations into tiered reward shaping for use with PPO. Because the LLM is discarded after training, deployment is a pure RL policy. Across TTCP CAGE CC1 and CC2 and both attacker types, this asymmetric design improves sample efficiency over PPO and outperforms the evaluated potential-based reward shaping (PBRS) baselines, while retaining the strongest terminal mean and requiring no LLM dependency at deployment time.

[LG-115] MovieSTAGE: Scene Transition and Global Encoding for Movie-fMRI ADHD Classification

链接: https://arxiv.org/abs/2610.09306
作者: Boseong Kim,Haejun Chung,Ikbeom Jang
类目: Machine Learning (cs.LG); Neurons and Cognition (q-bio.NC)
*备注: Accepted to the 2026 IEEE International Conference on Bioinformatics and Biomedicine (BIBM 2026)

点击查看摘要

Abstract:Naturalistic movie-fMRI provides a shared, temporally structured probe of brain dynamics, yet predictive models commonly rely on whole-run functional connectivity (FC) or temporally generic representations that are not aligned with narrative events. We introduce MovieSTAGE (Scene, Transition, and Global Encoding), a multiscale framework that combines hypergraph-structured FC-profile organization within scenes, unsigned FC-profile differences across adjacent scenes, and whole-movie FC. We evaluated 260 participants from the CMI-HBN Despicable Me cohort on case-control, ADHD-subtype, and three-class classification using 10 repetitions of stratified five-fold cross-validation, complete out-of-fold (OOF) predictions, and paired subject-cluster bootstrap and permutation tests. MovieSTAGE achieved AUROCs of 0.69, 0.73, and 0.75 and balanced accuracies of 67.6%, 69.8%, and 58.3%, respectively, yielding the highest mean point estimates among the evaluated methods. On the three-class task, the full model outperformed all two-branch variants, the HGNN scene encoder outperformed MLP, GAT, and BNT alternatives under matched settings, and the human-annotated partition outperformed duration-matched random and fixed-count GSBS controls. These controlled results support incremental predictive value from event-aligned scene and transition representations when combined with whole-movie FC in this cohort. Post-hoc model-derived analyses generated network-level hypotheses involving frontoparietal and default-mode systems.

[LG-116] Self-attention summary networks for subsurface velocity-model building from common-image gathers

链接: https://arxiv.org/abs/2610.09282
作者: Shiqin Zeng,Yunlin Zeng,Abhinav Prakash Gahlot,Zijun Deng,Felix J. Herrmann
类目: Machine Learning (cs.LG); Geophysics (physics.geo-ph)
*备注:

点击查看摘要

Abstract:Common-image gathers (CIGs) contain physically meaningful information about velocity-model errors through reflector focusing and residual moveout, but in conventional imaging workflows they are typically used only as diagnostic tools. In this work, we propose a multiscale self-attention summary network that maps high-dimensional 3D CIG volumes into compact conditioning embeddings for probabilistic subsurface velocity inversion. These learned embeddings preserve offset-dependent kinematic structure and spatial coherence while reducing variability caused by background-velocity mismatch. Conditioned on these summary embeddings, a flow-matching model learns a transport from a Gaussian source distribution to the posterior distribution of plausible velocity fields. Numerical experiments show that, compared with direct conditioning on raw CIGs, the proposed summary network improves posterior velocity inference. In particular, the multiscale attention design provides greater robustness to background-model mismatch, yielding more accurate posterior reconstructions and lower predictive uncertainty.

[LG-117] wist Flow for Inverse Problems

链接: https://arxiv.org/abs/2610.09281
作者: Shiqin Zeng,Zijun Deng,Felix J. Herrmann
类目: Machine Learning (cs.LG); Geophysics (physics.geo-ph)
*备注:

点击查看摘要

Abstract:In Bayesian inverse problems, posterior sampling requires generating samples that are consistent with given observations while capturing the range of plausible solutions. Direct conditional generative models introduce latent noise to model this ambiguity, but paired inverse-problem training can still encourage an almost deterministic map from the observation to the target. As a result, generated samples may be observation-consistent while under-representing posterior variability, especially when the posterior is multimodal, leading to undercoverage, mode distortion, or artificial transitions between distinct feasible solutions. We propose joint twist-flow, an augmented flow-matching formulation that learns a continuous transport from the augmented source state (z_x, y) to the augmented terminal state (x, z_y) . Here x is the target variable, y is the observation, z_x is the Gaussian reference coordinate for posterior sampling, and z_y is a Gaussian likelihood-side coordinate associated with the observation branch. Under a Gaussian observation model, z_y is motivated by the normalized observation residual associated with observation compatibility. Its role is not to replace uncertainty in x , but to couple generated samples of x to observation consistency, helping reduce likelihood-inconsistent variation while preserving variability in weakly constrained directions. We validate the method on low-dimensional inverse problems with reference posterior samples, where joint twist-flow better preserves multimodal posterior support than a direct conditional-flow baseline. We further evaluate the method on image restoration and seismic subsurface velocity-model inversion, showing increased posterior variability while maintaining observation consistency.

[LG-118] CATune: Structural Constraint-Aware Bayesian Optimization for DBMS Configuration Tuning VLDB

链接: https://arxiv.org/abs/2610.09276
作者: Fangping Lan,Qi Zhang,Eduard Dragut
类目: Databases (cs.DB); Machine Learning (cs.LG)
*备注: 14 pages including references, 10 figure, 2 tables. Accpeted by PVLDB, Volume 19, 2026

点击查看摘要

Abstract:Modern DBMSs expose hundreds of configuration knobs, resulting in a high-dimensional and heterogeneous search space that makes automated tuning costly. Existing ML-based tuning systems typically treat the configuration domain as box-constrained and rely on workload feedback to implicitly capture inter-knob relationships. However, DBMS documentation specifies deterministic knob dependency constraints, particularly ordering constraints, that characterize structurally valid regions of the configuration space. We present CATune, a constraint-aware Bayesian optimization (BO) framework that models deterministic inter-knob ordering constraints as structural components of the search domain. Instead of learning feasibility boundaries through sampled violations, CATune performs optimization within a constraint-consistent subspace. We develop a topology-aware sampling strategy that respects dependency structure during exploration and avoids the inefficiencies of post-hoc constraint handling. To enable automated constraint discovery, we further design a precision-first extraction pipeline that combines LLM-based parsing with reliability safeguards to mitigate hallucinated dependencies. Experiments on PostgreSQL and MySQL using TPC-C and TPC-H workloads show that CATune substantially improves both sample efficiency and final tuning quality across surrogate models and BO frameworks. Under default ranges, CATune reaches the baseline optimum up to 12.5x faster and improves throughput by up to 63.37%. The improvements persist under knowledge-guided reduced ranges and alternative optimization implementations. These results demonstrate that explicitly modeling system-defined deterministic ordering constraints enhances optimization robustness and system stability.

[LG-119] he Symbol of the Surrogate: Measuring Numerical Provenance in Neural PDE Solvers NEURIPS2026

链接: https://arxiv.org/abs/2610.09255
作者: Ridham Patel
类目: Machine Learning (cs.LG); Analysis of PDEs (math.AP); Machine Learning (stat.ML)
*备注: Accepted at NeurIPS 2026 AI for Science Workshop

点击查看摘要

Abstract:Neural PDE surrogates are trained on numerical solver outputs that contain both physical evolution and solver-specific discretization errors. Because surrogates are also evaluated against held-out trajectories from the same solver, standard benchmarks cannot distinguish fidelity to the exact evolution from imitation of the numerical scheme. We introduce an empirical Fourier-symbol diagnostic that probes a trained surrogate’s linearized one-step operator with individual Fourier modes and compares it with both exact-evolution and training-scheme references. To address architectural spectral bias, we train identical networks on schemes with orthogonal dissipative and dispersive signatures and compare their learned operators. In linear advection, the learned surrogates reproduce the training schemes’ amplitude and phase errors, with the twin-scheme difference reaching more than 99.8% of the analytically predicted full-imitation ceiling. The same behavior occurs for a non-local Fourier neural operator and at the operator level for nonlinear Burgers dynamics. These results show that agreement with solver-generated test data does not by itself establish fidelity to the exact evolution. Fourier-symbol measurements provide a direct diagnostic of numerical provenance.

[LG-120] Symmetry-Informed Causal Partial Identification

链接: https://arxiv.org/abs/2610.09230
作者: Uzair Akbar,Zulfiqar Zaidi,Niki Kilbertus,Krikamol Muandet,Bo Dai
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Partial identification (PI) entails estimating bounds on causal effects by encoding different assumptions on data generation as a constrained optimization problem. Such bounds can suffice to inform policy decisions even if the causal effect itself is not identifiable. Often vacuous in practice, practitioners seek to exhaustively encode domain knowledge as additional constraints to make the PI bounds more informative. We introduce known data symmetries – invariance of the causal effect under certain data transformations – as a new source of constraints to inform PI. We operationalize this as a shape constraint on the causal function, and via a change of measure against which PI is posed using simple data pre-processing. Both approaches are shown to sharpen bounds under two canonical PI models. This is shown both theoretically for the population case, and via experiments in the finite-sample case. More broadly, our framework establishes data symmetries as a natural, underutilized source of background knowledge for robust causal inference.

[LG-121] FLoRa: Flight-Assisted Data Collection from Duty Cycling LoRa Nodes under Energy Constraints

链接: https://arxiv.org/abs/2610.09226
作者: Naresh Babu Kakarla,V. Mahendran
类目: Networking and Internet Architecture (cs.NI); Machine Learning (cs.LG)
*备注: 20 pages

点击查看摘要

Abstract:Data collection using Unmanned Aerial Vehicles (UAVs) is challenging when LoRa IoT Devices (IoTDs) duty-cycle to conserve battery. Under energy constraints, a UAV must decide which IoTDs to visit, in what order, where to hover, and how many times to probe each node, while time-based data freshness decays. Tractably solving this problem requires a multi-level optimization architecture: discrete combinatorial optimization for routing, continuous global optimization for spatial positioning, and sequential decision-making under uncertainty. We propose FLoRa, a Flight-assisted LoRa data collection architecture using Simulated Annealing (SA) for path planning, Covariance Matrix Adaptation Evolution Strategy (CMA-ES) for hover positioning, and Partially Observable Markov Decision Processes (POMDPs) for probing IoTDs. To quantify collection utility from duty-cycling nodes, we introduce the Value of Information for Pull-based systems (VIP), a metric that rewards fresh data and penalizes failed probes, imposing well-posedness and preventing indefinite probing when an IoTD is off. Tracking hard battery constraints on every POMDP sample path requires state augmentation, worsening the curse of dimensionality. For tractability, SA and CMA-ES work on the hard battery constraints, while at the POMDP layer we relax them into soft average constraints via Lagrangian relaxation. Since solving the network-wide POMDP is computationally complex, we decompose it into node-level POMDPs by approximating inter-node time dependency using a forward-decomposition technique. Evaluation shows FLoRa outperforms metaheuristic, greedy, and deep reinforcement learning baselines by 30.6%, 27.8%, and 15.2% in total expected VIP, while increasing node coverage by 24.5%, 29.2%, and 8.8%, and successful collections by 24.3%, 25.6%, and 15.2%, respectively.

[LG-122] An Accuracy–Information Tradeoff for Loss-Difference Conditional Mutual Information

链接: https://arxiv.org/abs/2610.09206
作者: Hazar Yueksel
类目: Machine Learning (cs.LG); Information Theory (cs.IT); Machine Learning (stat.ML)
*备注: 61 pages, of which 8 pages main text. Code, data and the Lean 4 formalization are in the ancillary files

点击查看摘要

Abstract:Loss-difference conditional mutual information (ld-CMI) uses the smallest of the standard observations in the supersample hierarchy of generalization bounds: it measures what a learner’s loss differences reveal about which candidate of each pair it was trained on. Accuracy is known to force information into the model; data processing does not carry such lower bounds to losses. We show, by bounding three moments of the loss differences, that accuracy also forces ld-CMI. For linear predictors with a smooth convex loss of nonzero slope at zero, such as the logistic loss, plus a regularizer whose curvature and growth are both of power r\ge2 , on product distributions over a scaled sign cube in dimension at least linear in n , every proper learner with expected excess risk at most \varepsilon on these distributions at the optimal sample size n\asymp\varepsilon^-2+2/r has worst-case ld-CMI of order n bits, and \Theta(n/(1+(\tau/\varepsilon)^2)) bits under Gaussian noise of standard deviation \tau on the loss differences. The same holds without a regularizer, at n\asymp\varepsilon^-2 . Consequently, range-scaled ld-CMI bounds cannot vanish on these distributions, although every proper learner’s generalization gap is O(n^-1/2) . We also show that model-level information does not determine noisy loss-difference information, and that the growth, slope and dimension conditions are needed, the last up to a logarithm.

[LG-123] Exact Dynamics and Finite-Sample Trajectory Recovery of Linear Recursive Feature Machines

链接: https://arxiv.org/abs/2610.09196
作者: Andrew Cheng,Bobak T. Kiani,Yue M. Lu,Adityanarayanan Radhakrishnan
类目: Machine Learning (cs.LG); Machine Learning (stat.ML)
*备注: 51 pages, 8 figures

点击查看摘要

Abstract:Recursive feature machines (RFMs) learn representations of data by alternating between fitting a predictor to a dataset and updating features of that predictor using the average gradient outer product (AGOP). Connections between AGOPs and feature learning in neural networks motivate linear RFMs as a simple setting for analyzing how representations evolve during training. Here, we study the dynamics and statistics of linear RFM in noisy multi-output regression with isotropic sub-Gaussian input data and targets generated by a low-rank teacher matrix of dimension d . We extend the known connection between linear RFM and iteratively reweighted least squares from the interpolating setting to ridge-regularized multi-output regression with noise. We show that the learned feature matrix remains close to its infinite-data ideal counterpart at every iteration. Namely, for n samples, we show the error in the feature matrix decays as O(\sqrtd/n) with high probability. Experiments on real-world text and single-cell gene-expression data illustrate the features learned by this simple linear model.

[LG-124] he Dichotomy Between Pattern Recognition and Step-by-Step Reasoning

链接: https://arxiv.org/abs/2610.09186
作者: Amrut Nadgir,Pratik Chaudhari,Vijay Balasubramanian
类目: Machine Learning (cs.LG); Disordered Systems and Neural Networks (cond-mat.dis-nn)
*备注:

点击查看摘要

Abstract:We argue that pattern recognition and step-by-step reasoning are two ends of a spectrum. A large language model (LLM) learns to reason step-by-step when data is structured such that the next token depends on a small amount of preceding context. Inference in LLMs resembles pattern recognition when the next token depends on a large amount of preceding context. If the next token depends on only the c most recent tokens, reasoning traces are paths on a De Bruijn graph whose nodes are c -length contexts and edges are next-token transitions between contexts. The set of reasoning traces of a task forms a directed acyclic subgraph of the De Bruijn graph. An LLM that has learned all edges of this subgraph can compose them to solve longer, unseen tasks, i.e., it reasons step-by-step. We prove that the number of edges is vanishingly small compared to the number of reasoning traces. Empirically, the number of training samples a transformer needs is a power law in the number of edges, so learning to reason step-by-step is sample efficient. We can induce De Bruijn structure in any task by maintaining a ``state’’ that makes future reasoning independent of the past. The frequency of states in the reasoning trace determines c . We show, by fine-tuning Qwen2.5-1.5B-Instruct to solve equations and answer questions about stories, that frequent states (small c ) result in higher accuracy but greater fragility to perturbations at test time. LLMs trained with a large c are only as good as models that perform pattern recognition without reasoning. A moderate density of states balances accuracy and robustness. We show that real-world data has De Bruijn structure: Qwen3-14B and Qwen3-32B retain over 75% of their accuracy on GSM8K, MATH-500 and GPQA-Diamond when attention is restricted to a sliding window less than 15% as long as the full reasoning trace.

[LG-125] Q-PACE: Dynamic Precision Allocation for Quantization-Aware Training

链接: https://arxiv.org/abs/2610.09183
作者: Alexandra Volkova,Matin Ansaripour,Erik Schultheis,Christoph H. Lampert,Dan Alistarh
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Quantization-aware training (QAT) leverages lower-precision arithmetic to reduce the cost of LLM deployment, but aggressive quantization degrades final model performance. A common remedy is mixed-precision training, in which high precision is assigned to some of the layers to maintain performance while keeping the cost constrained. This approach then requires precision assignments for model layers during training. We provide a new approach, called Q-PACE, consisting of a second-order sensitivity model that predicts the loss increase as a sum of quantization noise MSE weighted by per-layer curvature coefficients. During training, we periodically re-compute these coefficients using perturbations across layers, and re-assign precision. Pretraining and supervised fine-tuning experiments on LLMs of up to 4B parameters show that Q-PACE consistently improves over existing mixed-precision training recipes, and achieves comparable loss at substantially lower total memory budgets. We further find that quantization sensitivity is highly predictable by depth and layer type, and its stability during training allows for infrequent, cheap recalibration.

[LG-126] Context-aware Attention-based Gaussian Mixture Models for Vehicular Trajectory Prediction

链接: https://arxiv.org/abs/2610.09174
作者: Arash Raftari,Babak Ebrahimi Soorchaei,Yaser P. Fallah
类目: Machine Learning (cs.LG)
*备注: 6 pages, 3 figures. Presented at the 2026 IEEE 104th Vehicular Technology Conference (VTC2026-Fall), Boston, MA, USA, 6-9 September 2026

点击查看摘要

Abstract:Reliable and interpretable trajectory prediction is critical for cooperative and autonomous driving in complex and uncertain environments. This paper introduces a Context-Aware Attention-based Gaussian Mixture Model (CAA-GMM) for multimodal, uncertainty-aware motion forecasting. The proposed approach models future motion as a probabilistic mixture conditioned on both scene context and agent dynamics, capturing diverse behavioral modes with interpretable Gaussian components. A lightweight attention mechanism adaptively encodes inter-agent interactions and contextual salience, enabling efficient fusion of rasterized environment cues and motion history in dense traffic scenes. Comprehensive evaluations on the nuScenes and Argoverse 2 datasets demonstrate that CAA-GMM achieves competitive or superior accuracy compared with state-of-the-art raster-based baselines, while maintaining low computational complexity. Ablation analyses confirm the importance of the attention module for robust contextual reasoning and predictive precision. Furthermore, evaluations under imperfect communication and perception conditions highlight the framework’s resilience to uncertainty, establishing CAA-GMM as an efficient and scalable solution for cooperative trajectory prediction in intelligent transportation systems.

[LG-127] A Cognitive-Aware QML-CRL Framework for Detecting Affinity and Romance-Investment Fraud

链接: https://arxiv.org/abs/2610.09141
作者: Bibhas Adhikari,Ramya Srinivasan
类目: Machine Learning (cs.LG); Quantum Physics (quant-ph)
*备注:

点击查看摘要

Abstract:We present a hybrid quantum-classical framework that detects affinity and romance-investment fraud by modelling the cognitive biases in a manipulative conversation. In our proposed framework, cognitive biases central to this fraud class are carried by dedicated qubits in a structured parameterized quantum circuit, together with a frame qubit makes the encoding sensitive to the temporal order of manipulative reframing, and a narrative qubit that aggregates co-occurrence through a trainable entanglement layer. The circuit parameters are trained jointly with a classical reinforcement-learning agent that decides, turn by turn, whether to flag the conversation, modeled as an optimal stopping problem. We evaluate the model’s performance on synthetic conversations that include hard negatives, legitimate but urgent, and legitimate but pushy sales conversations.

[LG-128] Directional Evidence Guided Search-Space Reduction for Exact DAG Learning

链接: https://arxiv.org/abs/2610.09136
作者: Upala Junaida Islam,Abdelmonem Elrefaey,Rong Pan
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Learning a directed acyclic graph (DAG) from observational data is a challenging combinatorial problem due to the exponential growth in the number of candidate parent-set configurations. Existing exact score-based methods often require computationally intensive combinatorial search, whereas constraint-based methods can become unreliable or computationally demanding as graph size and conditioning-set complexity increase. We develop a non-parametric hybrid framework, referred to as DECO (Directional Evidence-guided Configuration Optimization), that extracts dependency and directional evidence from observation data to construct admissible parent sets prior to exact optimization. It reduces the optimization search space by eliminating empirically unsupported parent configurations while preserving flexibility for all plausible edge orientations. Theoretical analysis establishes an exponential reduction in the admissible parent-set configuration space and quantifies how bounded edge-level omission affects the probability of retaining the true parent structure. Experiments on benchmark Bayesian networks and synthetic discrete and continuous DAGs demonstrate substantial search-space reduction while achieving competitive structure-recovery performance, with favorable structural Hamming distance across many evaluated settings. These results show that directional evidence can provide an effective preprocessing mechanism for reducing the computational burden of exact DAG learning without requiring a fixed parametric structural~model.

[LG-129] Domain-informed Adaptive Sampling for Generalizable PINNs in Metal Additive Manufacturing via Conditional Flow Matching

链接: https://arxiv.org/abs/2610.09126
作者: Hyeonsu Lee,Jihoon Jeong
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Accurate thermal modeling is essential in metal additive manufacturing (AM) for understanding the process-structure-property chain. Physics-informed neural networks (PINNs) offer effective surrogate thermal modeling by minimizing physics-based residual losses at collocation points. However, prior works typically rely on manually-crafted, static collocation sampling strategies, which are neither principled nor scalable across process conditions, hindering their generalization capability. In this work, we provide theoretical analysis through empirical risk minimization, showing that process condition-aware adaptive sampling is strictly more favorable than conventional static sampling for generalization. Building on this insight, we propose an adaptive sampling strategy within a two-stage framework: (1) a conditional Flow Matching model that learns approximate high-residual distributions across different process conditions, and (2) a mixed sampling strategy combining this distribution with a domain-informed base distribution to generate adaptive collocation points for refining the PINN predictor. Experiments on metal AM numerical benchmarks demonstrate that our method consistently outperforms state-of-the-art PINN baselines, achieving an average 62.1% reduction in relative L_2 error under an identical collocation budget, by capturing process-dependent heat dissipation regions often overlooked in the literature. To the authors’ knowledge, this is the first adaptive sampling strategy for PINNs in metal AM, contributing to the enhanced generalization and broader applicability.

[LG-130] Are Parameter-Efficient Fine-tuning Methods Really Different?

链接: https://arxiv.org/abs/2610.09122
作者: Yikuan Li,Pinyan Lu,Fanghui Liu
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Parameter-efficient fine-tuning (PEFT) offers many parameterizations, yet their methodological and functional differences remain unclear. We compare six methods in language and diffusion models to examine how their parameterizations relate to task performance, forgetting, and changes in pretrained weight geometry. Motivated by the spectrum-preserving design of orthogonal fine-tuning (OFT), we first ask whether spectral preservation is itself important for adaptation and retention. We find that the selected LoRA-family methods also approximately preserve pretrained geometry, and that restoring their slightly drifted singular-value spectra largely preserves task performance, questioning the necessity of explicit geometric preservation. Beyond this, we observe that some methods exhibit distinct adaptation–retention trade-offs that vary across settings: LoRA most consistently limits forgetting at competitive performance, DoRA achieves higher mean task scores than LoRA in most comparisons, while PiSSA often incurs greater retention costs. Further intervention experiments suggest that while performance gains from different PEFT methods can be attributed to modifications in different groups of spectral components, we consistently find that restoring dominant rather than intermediate or trailing components produces the largest mean reduction in general-text NLL or base-image drift. Together, these results motivate evaluating geometric constraints through their functional consequences rather than preservation alone. Code is available at this https URL.

[LG-131] he Deceptive Bandit Problem: Exploratory Coupling and the Frag ility of Multi-Agent Learning

链接: https://arxiv.org/abs/2610.09120
作者: Michael Tang,Mahmoud Abdelgalil,Jorge I. Poveda
类目: Machine Learning (cs.LG); Systems and Control (eess.SY)
*备注:

点击查看摘要

Abstract:Randomized exploration is central to bandit learning, multi-agent reinforcement learning, and zeroth-order policy search, yet its independence and privacy are usually only treated as technical assumptions. We show that these properties are critical for security purposes and demonstrate how an adversarial agent can exploit privileged information on another agent’s exploration. We analyze a deceiver-victim pair in the minimal two-player strongly monotone setting, where a deceptive player obtains leaked signals that are merely correlated with the victim’s exploration. We show that, by coupling their own exploratory action with this information, the deceptive player injects an externality that steers the learning dynamics to a new steady state, called the deceptive Nash equilibrium (DNE). We prove that the deceptive bandit learning (DBL) dynamics converge to an arbitrarily small neighborhood of the DNE while retaining optimal convergence rates. Interestingly, our analysis attains these optimal rates while relaxing second-order smoothness conditions from standard bandit optimization literature. We characterize conditions under which deception strictly shifts the steady state and its effect on the deceiver’s cost, illustrating the results in a resource-allocation game.

[LG-132] Workhorse: Learning Robust Whole-Body Humanoid Loco-Manipulation from Human Data

链接: https://arxiv.org/abs/2610.09117
作者: Songbo Hu,Qiayuan Liao,Yufeng Chi,Kevin Zakka,Yakun Sophia Shao,Pieter Abbeel,Koushil Sreenath
类目: Robotics (cs.RO); Machine Learning (cs.LG)
*备注: 9 pages, 8 figures, 2 tables. Project page: this https URL

点击查看摘要

Abstract:Humanoid robots still struggle to plan contact-rich whole-body manipulation from egocentric RGB and proprioception. Workhorse learns such manipulation from robot-free human demonstrations. A visual planner predicts five-link targets: the poses of the torso, both wrists, and both feet. A reinforcement-learning whole-body tracker follows them on the robot. Both policies train separately on the same recorded human poses, without retargeting. We augment the training data of each policy to imitate the errors that the other makes at deployment. On a real Unitree G1, Workhorse sorts boxes with its hands and a kick, catches a thrown box, and topples and climbs a suitcase. During box sorting, we show recoveries after a person pushes the robot or takes the box away. In a simulated copy of the demonstration room, the system completes box sorting in 77% of episodes, and in 64% under 40 N.s pushes. With both policies retrained from the same demonstrations, a simulated second humanoid completes box sorting in 83% of episodes without pushes.

[LG-133] Breaking Adversarial Transferability in Fine-Tuned Speech Recognition

链接: https://arxiv.org/abs/2610.09109
作者: Mojtaba Nafez,Aref Mousavi,Mohammad Ebrahim Mahdavi,Mobina Poulaei,Kiarash Kiani Feriz,Mohammad Hossein Rohban
类目: Machine Learning (cs.LG)
*备注: 39 pages, 5 figures

点击查看摘要

Abstract:Many organizations fine-tune publicly available pretrained Automatic Speech Recognition (ASR) models and deploy them in black-box settings, assuming limited access provides protection. We show this assumption is fragile: adversarial perturbations crafted on the public base model transfer effectively to fine-tuned target models, severely degrading performance and posing concerns for safety-critical applications. We propose TransferBreaker, a unified fine-tuning framework that suppresses adversarial transfer by integrating Base Adversarial Fine-Tuning, which restricts adversarial training to base-effective perturbations; Latent Jacobian Regularization, which enforces latent-space invariance by suppressing adversarially sensitive directions; and HybridGrad-AFT, which improves robustness against adaptive attacks by interpolating transferable perturbations from base and target gradients. We theoretically justify all components and evaluate TransferBreaker across three languages and four large ASR models, reducing adversarial WER from 92.6 to 27.8. Our code is publicly available at this https URL.

[LG-134] On the Computational Complexity of Hidden Markov Model Identification

链接: https://arxiv.org/abs/2610.09104
作者: Markel Zubia,Nils Jansen
类目: Computational Complexity (cs.CC); Machine Learning (cs.LG); Statistics Theory (math.ST)
*备注:

点击查看摘要

Abstract:Identification is the task of recovering the parameters of an unknown ground-truth model from sampled data. When parameters other than the ground truth induce the same output distribution, data alone does not provide enough information to recover the ground truth, and the model is thus called unidentifiable. We study the identifiability problem for hidden Markov models (HMMs): given an HMM, is it identifiable? Existing work on HMM identification establishes conditions under which the ground-truth HMM can be identified. However, most of these conditions are sufficient but not necessary, meaning that, when a model does not satisfy them, its identifiability remains inconclusive. We instead take a computational perspective: is there a sound and complete algorithm that decides whether a given HMM is identifiable, and if so, what is the complexity of this decision problem? We consider the decision problems arising from the various notions of identifiability in the literature, including deterministic, generic, global, local, state-permutation- invariant, and finite-alphabet identifiability. We show that all of these problems are decidable in PSPACE, via reductions to the theory of the reals at various levels of its quantifier-alternation hierarchy. We further show that the deterministic variants are already coETR-hard (and hence coNP-hard) for simply parameterized families.

[LG-135] Are We Really Benchmarking Forecasting Models? The Impact of Preprocessing on Time Series Performance

链接: https://arxiv.org/abs/2610.09096
作者: Guilherme Afonso Galindo Padilha,Paulo Salgado Gomes de Mattos Neto,Rafael Menelau Oliveira e Cruz
类目: Machine Learning (cs.LG)
*备注: 29 pages, 9 figures. Under Review

点击查看摘要

Abstract:While established literature underscores the pivotal role of preprocessing in forecasting accuracy, this stage remains largely overlooked in current research. Modern benchmarks typically resort to simple scaling, failing to account for critical transformations required to address nonstationarity, such as differencing. This omission creates a significant structural preprocessing bias that favors models with built-in data treatments while obscuring the true potential of simpler architectures. We study this effect through a preprocessing-aware benchmark that evaluates 11 forecasting models across 16 reversible preprocessing pipelines on 29,000 M4 time series. Our results identify preprocessing as a key driver of forecasting performance. Optimizing preprocessing per series yields gains of approximately 27% to 87% across all evaluated models, with architectures lacking internalized preprocessing experiencing the most substantial improvements. This allows simpler architectures to become highly competitive with complex, state-of-the-art models in modern forecasting benchmarks. All resources and experimental results from this benchmark are stored in a comprehensive metadataset to support future metalearning tasks.

[LG-136] FedRSPO: A Heterogeneity-aware Algorithm for Decision-focused Federated Learning NEURIPS2026

链接: https://arxiv.org/abs/2610.09091
作者: Konstantinos Ziliaskopoulos,Alexander Vinel,Jiaqi Wang
类目: Machine Learning (cs.LG); Optimization and Control (math.OC); Machine Learning (stat.ML)
*备注: 10 pages main paper + appendix. Accepted at NeurIPS 2026 main track

点击查看摘要

Abstract:Decision-focused learning (DFL) trains predictive models for downstream optimization, but existing methods largely assume centralized data. In cross-silo settings, federated learning offers a natural alternative, yet standard federated methods optimize prediction over decision quality and do not address heterogeneity in downstream objectives or feasible sets. This heterogeneity is especially challenging for DFL because small perturbations in polyhedral problems can cause discontinuous changes in optimal decisions, destabilizing client updates and aggregation. We propose FedRSPO+, a heterogeneity-aware framework for decision-focused federated learning, built on RSPO+, a regularized predict-then-optimize surrogate that smooths the decision map through projection. We show that RSPO+ upper bounds decision error and regret for the regularized decision and, under exact regularization and consistent LP solution selection, for the original LP decision. We further derive cross-client heterogeneity bounds that depend on both objective and feasible-set heterogeneity, vanish at homogeneity, and require no strong convexity. FedRSPO+ uses an annealed, modular training procedure compatible with standard federated personalization and aggregation methods. Experiments on synthetic knapsack, shortest-path, and real-world energy pricing tasks compare against prediction-only federated learning and DFL baselines under varying heterogeneity and communication budgets. Results suggest that smoothing is a useful ingredient for stable collaborative decision learning and provide a heterogeneity-aware foundation for federated DFL.

[LG-137] ucker Bottleneck Attention for Multi-Dimensional Sequence Modeling

链接: https://arxiv.org/abs/2610.09090
作者: Ryan Solgi,Parsa Madinei,Zheng Zhang
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:The quadratic cost of self-attention limits scalability to long sequences from multidimensional data. We introduce Tucker bottleneck attention (TuBA), which exploits low-rank tensor structure for efficient global token mixing. TuBA projects hidden tensors into compact Tucker cores, performs multi-head self-attention and linear projections on the cores, and writes updates back to the ambient space, enabling subquadratic computation. Its autoregressive extension combines bidirectional interactions within cores with causal attention across cores. On video prediction and global weather forecasting, TuBA achieves favorable accuracy-efficiency trade-offs over standard and efficient attention and task-specific models. Compared to standard self-attention, TuBA reduces error and computation by up to 24.7% and 66.6% for video prediction and 37.1% and 85.1% for autoregressive weather forecasting, with speedups up to 4.27 times. Low-rank Tucker cores and multi-frame generation also outperform full-rank attention and frame-by-frame generation, respectively.

[LG-138] owards AI-Generated Music Plagiarism Detection as a Version Identification Problem

链接: https://arxiv.org/abs/2610.09075
作者: Fotis Koutsikos,Ioannis Prokopiou,Spyridon Kantarelis,Vassilis Lyberatos,Pantelis Vikatos,Athanasios Aidinis,Themos Stafylakis,Athanasios Voulodimos,Giorgos Stamou
类目: ound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
*备注: 5 pages, 3 figures, submitted to 2027 IEEE International Conference on Acoustics, Speech, and Signal Processing

点击查看摘要

Abstract:The rapid expansion of text-to-music generative models challenges traditional paradigms of music creation and intellectual property. Plagiarism in this context is rarely an absolute mathematical binary, but an ambiguous threshold negotiated over harmonic structure, melodic contours, or overall perceived stylistic character. In this work, we test the transferability of state-of-the-art music version identification architectures from the human-to-human cover domain to the human-to-AI plagiarism setting. To evaluate this task, we introduce COPYCAT, a benchmark derived from real-world plagiarism cases and extended through generative re-synthesis and digital signal processing obfuscations, yielding 350,654 evaluation pairs. We show that scalar distance thresholding collapses under generative re-synthesis, while a supervised framework leveraging coordinate-wise embedding shifts recovers the dispersed plagiarism signal, raising overall F_0.5 from 0.612 to 0.803 .

[LG-139] AP: Efficient Long-Horizon Agent Pruning via Trajectory-Anchored Recovery

链接: https://arxiv.org/abs/2610.09074
作者: Yuanzhe Li,Pengxin Wang,Yuxin Ren,Jianing Deng,Jingtong Hu,Song Wang,Jingdi Chen,Huanrui Yang
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Emerging long-horizon agentic tasks require repeated model calls, worsening the inference cost of already-costly language models. While narrow agentic tasks suggest potential for aggressive model pruning without performance drop, empirical results show existing methods proposed for question answering tasks severely degrade task performance when applied to agentic models. We trace this failure to two decisions: what to prune and how to recover. For pruning, one-shot importance estimates fail to track how the pruned model adapts. For recovery, offline distillation covers only teacher prefixes, while full-trajectory on-policy distillation causes student errors to compound across turns. In this work, we propose Trajectory-Anchored Pruning (TAP), the first structural pruning framework for reinforcement learning (RL)-trained agents. TAP couples structural pruning with efficient on-policy recovery, anchoring interactions to teacher trajectories while allowing the student to generate each reasoning-action response. A frozen dense teacher supervises the student’s response prefixes, addressing within-response training-inference mismatch while preventing student-induced deviations from propagating across training turns. Instead of one-shot pruning, TAP re-scores channels using gradients of the recovery objective on the recovered student, connecting iterative channel selection to the evolving policy. With 60% of FFN channels removed, TAP retains 99.2% and 88.0% of the dense 7B agents’ task success rates on ALFWorld and WebShop, respectively, while reducing GPU time per successful task by approximately 22% and 17%. These results demonstrate effective structural compression of long-horizon agents under a limited recovery budget.

[LG-140] EDiS: Edge Disjoint Subgraph Sparsification Framework for Graph Neural Networks

链接: https://arxiv.org/abs/2610.09059
作者: Sai Karthik Navuluru,Siddhartha Shankar Das,Franck Dernoncourt,S M Ferdous,Ryan A. Rossi,Nesreen K. Ahmed,Baris Coskunuzer,Alex Pothen,Lakshman Tamil,Mahantesh M Halappanavar
类目: Machine Learning (cs.LG)
*备注: 46 pages, including references and appendices

点击查看摘要

Abstract:Sparse GNN training reduces computation, but deciding which edges to keep can be costly. Reusing one sparse graph is cheap, but locks training to a fixed topology, while varying it across epochs can require repeated sampling or recomputation. We introduce EDiS (Edge-Disjoint Subgraph sparsification framework), which separates one-time structural extraction from per-epoch graph composition. EDiS decomposes the graph once into cacheable edge-disjoint subgraphs, then recombines them into graphs with edge-budget constraints across epochs and retention ratios without re-extracting structure. Our default construction uses feature-based scores and successive maximum score covering forests, while the same composition mechanism also supports alternative edge selection rules. We provide a combinatorial analysis of the per-epoch sampler, the composition step that draws a training graph from the cached decomposition. We show that, under the default covering-forest selector, the stored decomposition deterministically preserves high-score cut edges, and we derive a selector-agnostic conditional bound on high-score cut survival in composed training graphs. Across 19 homophilic, heterophilic, and large-scale node classification benchmarks against 17 baselines under the same edge budget, EDiS achieves the highest mean benchmark score (accuracy/ROC-AUC) and the lowest average rank and gap-to-best among ranked methods. Ablations show the clearest benefits of structural decomposition and epoch variation at tight edge budgets.

[LG-141] Learning Cross-Model Activation Alignments with Explicit Many-to-Many Layer Maps

链接: https://arxiv.org/abs/2610.09058
作者: Alina Sudakov,Guy Bar-Shalom,Fabrizio Frasca,Haggai Maron
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:LLMs are released at a rapid pace, raising a natural question: how do two independently trained models relate, both in which layers correspond and in how features transform between them? We study this by learning an activation alignment, a map from a source model’s layerwise activations to a target’s. Our method, MATCHA, factors this map into a layer map, whose output is an explicit target-by-source matrix that can be extracted and inspected, and a layer-shared feature map between hidden spaces. Most of prior work fixes the layer correspondence in advance, pairing layers at roughly the same relative depth; in contrast, we learn both factors jointly from prompts. Across 42 pairs of seven models spanning three different families, MATCHA reconstructs the target’s activations more faithfully and improves retrieval-based metrics substantially, w.r.t. previous approaches. The recovered maps are broadly monotone in depth but, in contrast with most previous approaches, are consistently many-to-many: each target layer draws on a band of source layers. Our alignments also enable transfer of activation-space interventions, allowing steering vectors and probes developed for one model to transfer to another.

[LG-142] A Geometry-Based Capacity Theory for Finite-Feature Associative Memory

链接: https://arxiv.org/abs/2610.09056
作者: Jianhai Zhang,Donghao Zhang,Pattarawut Charatpangoon,Bijoy Menon,M. Ethan MacDonald,Wu Qiu,Aravind Ganesh
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:We develop a geometry-based capacity theory for exact-key retrieval in compressed finite-feature Hebbian associative memory. For random or approximately isotropic values, retrieval interference separates into finite-feature noise, which decreases with feature dimension, and structural interference, which is determined by squared kernel overlap among stored keys and persists in the infinite-feature limit. This yields a fit-free prediction of retrieval quality, reveals a geometry-dependent capacity ceiling, and predicts the feature budget required for a target retrieval quality. When stored values are correlated, we show that retrieval depends jointly on the key kernel and value Gram matrix, and derive finite-feature approximations that account for this interaction. We validate the theory on synthetic, visual, and medical-image representations. Overall, the framework links representation geometry directly to memory capacity and distinguishes when performance can be improved by increasing the feature budget and when the representation itself must be changed. Across these settings, the predicted retrieval curves closely match empirical behavior and correctly identify changes in the preferred memory design.

[LG-143] BeatFlow-ECG: Rectified Flow for ECG Reconstruction from Indirect Wearable Signals

链接: https://arxiv.org/abs/2610.09052
作者: Mohamed Kamel,Sahar Selim,Walaa Medhat,Tamer Nadeem
类目: Machine Learning (cs.LG); Signal Processing (eess.SP)
*备注: 27 pages. Preprint. Under review

点击查看摘要

Abstract:Continuous cardiac monitoring outside clinical settings requires signals that are both informative and practical to collect during daily life. Electrocardiography (ECG) provides rich information about cardiac rhythm and waveform morphology, while wearable photoplethysmography (PPG) is easier to acquire continuously but is only an indirect cardiovascular measurement and is highly sensitive to motion. We present BeatFlow-ECG, a conditional rectified-flow model for reconstructing single-channel ECG from synchronized PPG and inertial measurements. BeatFlow-ECG models reconstruction as conditional transport from noise to ECG using a convolutional encoder-decoder with a transformer bottleneck and explicit flow-time conditioning. Motion information is incorporated through IMU-derived conditioning features, motion-dependent loss weighting, and an easy-to-hard training curriculum. We evaluate the model under leave-one-subject-out protocols on PPG-DaLiA and WESAD. BeatFlow-ECG achieves the best results among the evaluated deterministic, adversarial, and diffusion-based baselines across all reported waveform and beat-timing metrics, with Pearson correlations of 0.983 and 0.986 and R-peak F1 scores of 0.946 and 0.955, respectively. Compared with Conditional DDPM-1D, L1 error decreases from 0.085 to 0.062 on PPG-DaLiA and from 0.074 to 0.055 on WESAD. Additional analyses on PPG-DaLiA show higher correlation in fixed R-peak-relative waveform regions and lower reconstruction error across low-, medium-, and high-motion subsets.

[LG-144] A Self-Pruning Transformer: Extreme KV-Cache Compression with Universal Attention

链接: https://arxiv.org/abs/2610.09051
作者: Davis Wertheimer,Haochen Shen,Ahan Gupta,Derrick Liu,Yu Chin Fabian Lim,Mudhakar Srivatsa,Raghu K. Ganti,Minjia Zhang,Naigang Wang
类目: Machine Learning (cs.LG)
*备注: 27 pages, 3 figures, 14 tables

点击查看摘要

Abstract:The large KV-cache size of modern LLMs creates a barrier to efficient deployment. Recent work has explored replacing attention layers’ RoPE positional embeddings with alternative decay-based mechanisms, which can then be used to prune KV-cache during inference. However, these decay functions have limited expressivity, and in practice devolve into sliding-window-like eviction patterns. In this work, we propose a unifying framework for complementary and novel decay mechanisms, capturing complex key statistics and interactions while preserving expressive RoPE embeddings and Softmax attention. The resulting Universal Attention is a highly expressive and end-to-end trainable architecture, whose composite decay mechanism acts as a natural, \textitadaptive pruning criterion, removing tokens that contribute least to attention computation. Experimentally, Universal Attention achieves state-of-the-art 10\times compression on natural language and synthetic task data, while \textitimproving downstream performance compared to both state-of-the-art baselines and unpruned oracles. It further demonstrates superior long-context generalization with unprecedented 25\times compression at length 16k.

[LG-145] owards Financial World Modeling

链接: https://arxiv.org/abs/2610.09048
作者: Humzah Merchant,Alec Guthrie,Simon Mahns,Randall Balestriero,Bradford Levy
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Building a world model requires a state representation useful for planning and decision-making—potentially over tasks unknown at training time. In the context of financial markets, planning and decision-making may require a model to reason about market-wide conditions, asset-specific expected returns, liquidity, volatility, and cross-asset relationships. Yet financial representation learning has largely been evaluated on individual predictive tasks, oftentimes on a single time period using comparatively narrow datasets. We address this through three primary contributions. First, we introduce Market-1T, a dataset containing nearly one trillion observations across U.S. equities from 2008 to 2025 at 1 Hz resolution. Second, we develop and implement a rigorous evaluation protocol. Third, we conduct a systematic large-scale study of financial representation learning, comparing 18 encoder-training strategies across nearly two decades of market regimes. We evaluate learned representations both by their predictive utility on common finance tasks and through probes of latent structure. We find that encoders with similar predictive performance can organize market state very differently. Collectively, we establish a foundation for training and evaluating financial market representations in support of world models such as DINO-WM, V-JEPA 2, and LeWM.

[LG-146] ORACLE: Optimizer-Relative Alignment for Constrained LEarning

链接: https://arxiv.org/abs/2610.09040
作者: Utkarsh Grover,Wyatt Mackey,Kaixun Hua,J. Morris Chang,Xiaomin Lin
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Constraint handling methods typically intervene before the optimizer acts, by modifying the objective or the gradient. Yet momentum, adaptive scaling, and structured preconditioning can substantially reshape that signal before it becomes a parameter update. We formulate optimizer relative constrained learning, where constraint compatibility is assessed on the post optimizer update. Building on this view, we introduce ORACLE, which evaluates the native optimizer’s realized step through a joint endpoint linearization of heterogeneous constraint families, constructs the resulting alignment in the optimizer’s own geometry, bounds its authority, and commits it only after validation. We evaluate ORACLE across eight Partial Differential Equation benchmarks and four optimizers spanning Euclidean, diagonal adaptive, and structured preconditioned geometries, where it improves or matches native optimizer in 94% of configurations. Cross model analysis shows the same behavior in 92% of configurations, while matched comparisons show improvements over alternative constraint-handling methods acting at the objective, gradient, and post-optimizer levels.

[LG-147] SPIN: Shadow Predictive Indexer for Sparse Attention

链接: https://arxiv.org/abs/2610.09025
作者: Yao Fu,Cyrus Chang,Ritchie Zhao,Bryce Long,Yueying Li,Mahdi Kamani,Samkit Jain,Rahul Raman,Tara Safavi,Shreya Gupta,Parsa Ashrafi Fashi,Minseok Lee,Julien Demouth,Bita Darvish Rouhani
类目: Machine Learning (cs.LG)
*备注: 12 pages, 4 figures

点击查看摘要

Abstract:Indexer-based sparse attention reduces the cost of core attention by passing only a fixed, small number of important tokens to it. However, the indexer must still score the entire KV cache at every decoding step. This scoring overhead becomes a major bottleneck as the context length grows. We propose SPIN (Shadow Predictive Indexer) to reduce this indexer overhead. SPIN uses lightweight, history-based prediction to identify important KV blocks, avoiding the need to score the full KV cache at every decoding step. SPIN treats KV blocks and speculative decoding as first-class design and implementation considerations. Across extensive evaluations on long-context and agentic benchmarks, SPIN achieves 30-40% sparsity while preserving task quality. In end-to-end vLLM serving, SPIN improves output throughput by up to 14.9% and reduces median inter-token latency by up to 13.2%.

[LG-148] Neighborhood Smoothing for Calibration

链接: https://arxiv.org/abs/2610.09020
作者: Idan Horowitz,Avigdor Gal
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Modern neural networks are often miscalibrated, with a tendency to overconfidence. Existing train-time calibration methods largely modify task losses or calibration penalties, leaving neighborhood structure in learned representations underexploited. We introduce graph smoothing as a general principle for train-time calibration, which encourages similar predictive distributions across neighboring samples in representation space. We analyze the effects of graph smoothing, deriving bounds that connect predictive divergence between neighboring samples to local confidence variation and to the propagation of pointwise calibration error, and characterize the conditions under which smoothing can or cannot improve calibration. In light of this analysis, we propose \modelNoSpace, a graph-based train-time regularizer that penalizes the Jensen–Shannon divergence between predictive distributions of neighboring samples. We present a thorough empirical analysis, showing that across standard calibration benchmarks, \model improves predictive quality, and the improvement is complementary to post-hoc calibration: after temperature scaling, \model attains the lowest NLL of all evaluated train-time methods in seven of the eight image and tabular settings. These findings demonstrate the value of graph smoothing over learned representations for neural network calibration.

[LG-149] Removing Information Content Does Not Certify Tamper Resistance in Open-Weight Models AISTATS2026

链接: https://arxiv.org/abs/2610.09004
作者: Domenic Rosati,Alessa Carbo,Ali Dadsetan,Hong Huang,Matthew Young,Subhabrata Majumdar,Frank Rudzicz,Hassan Sajjad
类目: Machine Learning (cs.LG); Cryptography and Security (cs.CR)
*备注: Under submission AISTATS 2026

点击查看摘要

Abstract:Does removing harmful information make open-weight models resistant to fine-tuning attacks? We show that mutual information at release alone cannot universally certify slow recovery. Function-preserving reparameterizations leave information unchanged while altering gradient-descent geometry, so an invariant certificate is bounded by the fastest reachable parameterization. We apply this principle to weight–data mutual information under training-data filtering and label–representation mutual information under capability removal. Training order can change recovery time at fixed weight–data information, while exact representation-level independence can preserve the entire parameter Jacobian. An explicit construction has both information quantities equal to zero and recovers in one gradient step. Controlled experiments illustrate order-dependent recovery and parameterization-dependent attack speed. These results identify the missing requirement for certification: constraints on attack dynamics beyond mutual information at release.

[LG-150] SNR-Gated LSTM-Conditioned Diffusion Model for MIMO Channel Estimation

链接: https://arxiv.org/abs/2610.08977
作者: Jixing Zhou,Xinming Huang
类目: Machine Learning (cs.LG); Signal Processing (eess.SP)
*备注: 5 pages, 5 figures. Accepted by and presented at the 2026 IEEE 104th Vehicular Technology Conference (VTC2026-Fall)

点击查看摘要

Abstract:Accurate and low latency channel estimation is critical for modern MIMO systems, particularly under mobility, where channels exhibit structured sparsity and strong temporal correlation. This paper proposes a time-series conditioned diffusion framework for channel estimation that performs denoising in the angular domain. Starting from least squares (LS) observations, we train a diffusion denoiser whose conditioning information is encoded by a long short-term memory (LSTM) network over a short observation sequence, enabling the model to exploit temporal dynamics beyond per-snapshot estimation. To robustly balance observation fidelity and learned generative priors across a wide signal-to-noise ratio (SNR) range, we introduce a learnable SNR-gated late-fusion shortcut that injects the network input into the final decoding stage through a sigmoid gate with trainable center and scale. To reduce inference latency, we adopt deterministic denoising diffusion implicit model (DDIM) style reverse updates with SNR-adaptive truncation and step allocation, which significantly reduces the number of reverse diffusion steps at high SNR while maintaining strong performance in low SNR regimes. Simulations on time-evolving standardized channel models demonstrate that the proposed method achieves consistent performance gains over existing diffusion-based channel estimation baselines, while retaining low latency through SNR-adaptive inference.

[LG-151] he Best Optimizer Depends on Batch Size

链接: https://arxiv.org/abs/2610.08975
作者: Xingyu Dang,Kaiyue Wen,Sadhika Malladi
类目: Machine Learning (cs.LG); Machine Learning (stat.ML)
*备注:

点击查看摘要

Abstract:A plethora of new adaptive optimizers are designed to efficiently estimate and use minibatch gradient statistics to shape parameter updates, but they are typically benchmarked at a single batch size. Hyperparameter scaling rules promise to preserve performance as batch size and gradient noise change, suggesting that the best optimizer at one batch size should remain the best at another. We challenge this approach to developing and evaluating optimizers by showing: (1) no principled scaling rule for Muon works consistently across training settings, and (2) the best optimizer for language model pretraining changes with batch size even after extensive hyperparameter tuning.

[LG-152] HULK: Learning Whole-Body Forceful Loco-Manipulation for Humanoids

链接: https://arxiv.org/abs/2610.08970
作者: An Dang,Arturo Flores Alvarez,Yu-Ming Chen,Conor Mc Gartoll,Helen Sun,Aaron Ames,Nima Fazeli,Manikantan Nambi
类目: Robotics (cs.RO); Machine Learning (cs.LG)
*备注: 16 pages, 7 figures, IEEE International Conference on Robotics and Automation 2027

点击查看摘要

Abstract:Humanoid loco-manipulation of large, heavy objects demands forceful interaction across the entire body. However, such payloads shift a humanoid’s center of mass and impose sustained loads across the upper body, challenging balance and command tracking. We present HULK, a whole-body control framework for forceful loco-manipulation. Using model predictive control (MPC) to guide reinforcement learning with predictions of the loaded dynamics, we train two teachers: one tracks arm motions under wrist forces, and the other locomotes while holding large objects against the body. A capture-point control barrier function augments the wrist-force teacher during training to improve balance under load. We distill both teachers into a single policy. Evaluation spans simulation and the Unitree G1. In simulation, the teacher with the barrier function achieves the lowest forward and lateral velocity tracking errors at 10 kg per arm among evaluated controllers and reduces aggregate divergent component of motion (DCM) excursion magnitude by 35.7% relative to MPC-guided reinforcement learning alone. Our wrist-force teacher withstands torso push disturbances of up to 130 N.

[LG-153] Directed Temporal Representations for Offline Visual Control

链接: https://arxiv.org/abs/2610.08960
作者: Chenyang Yuan,Haoyu Wang,Zhuo Sun,Xiaoyuan Cheng
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Predictive world models provide compact visual representations for control. Control requires a latent geometry aligned with temporal reachability rather than predictive similarity alone. We introduce Directed Temporal Representations for Control (DTRC), which learns such a geometry from offline visual trajectories on top of frozen LeWorldModel (LeWM) features. DTRC constructs a directed temporal quasimetric over the learned control representation. Short-range temporal offsets calibrate the distance scale. Bootstrapped targets extend temporal reachability across longer horizons. Action-conditioned consistency aligns the representation with local transition dynamics. The resulting distance estimates temporal reaching cost, and its change across a transition defines goal-relative temporal progress. We use this progress signal as a temporal critic for direct goal-conditioned policy learning. Model-assisted targets provide an additional training-time refinement under behavior-support and dynamics-agreement constraints. Across ten visual control tasks, DTRC achieves strong goal-conditioned control performance relative to planning and direct-policy baselines. Held-out diagnostics on the four LeWM tasks show consistent short-range temporal calibration, task-dependent long-range and directional structure, and positive transition-level progress. Temporal supervision improves the same flow-policy parameterization across all four LeWM tasks, while the resulting policy acts directly without iterative trajectory search at test time.

[LG-154] ask-Sufficient Contraction: Source Selection for Machine Information Interfaces

链接: https://arxiv.org/abs/2610.08884
作者: Joss Armstrong
类目: Information Theory (cs.IT); Machine Learning (cs.LG)
*备注: 19 pages, 1 figure, 1 table. An earlier version was first posted on Zenodo in September 2026 (v1: doi: https://doi.org/10.5281/zenodo.22834532%3B all versions: doi: https://doi.org/10.5281/zenodo.22834531 ). Companion paper on task-relative information contracts: doi: https://doi.org/10.5281/zenodo.22819849

点击查看摘要

Abstract:A declared task can sometimes certify a reduced source before a downstream encoder, codebook, rate, distortion target, or optimizer is chosen. This paper studies when one such reduction preserves the complete downstream problem family, a property termed Task-Sufficient Contraction. The reduced source is fixed by the task before the later operating point is selected. An exact contraction allows the later problem to be solved on that source with the same result as if the full source had been retained. For a machine with a fixed set of possible actions and a fixed loss, the paper identifies a consumer-specific source by merging states only when every available action has the same regret in both. For finite action sets, replacing the richer source by this reduced source preserves the complete one-step rate-regret curve, even though the reduction is fixed before the distortion target is chosen. A second result gives an exact characterization for quadratic loss on affine feasible-action sets: the canonical reduced source is the projection onto the directions in which feasible actions can differ. Under a fixed energy budget, this becomes centered load, while retaining only the optimal water-filled action is too coarse. Earlier Information Bottleneck, semantic rate-distortion, and goal-oriented quantization results are then used to distinguish exact, architecture-conditioned, approximate, failed, and corrected contractions. The framework suggests a way for heterogeneous machines to exchange what a receiving task needs without first aligning their full internal representations. Comments: 19 pages, 1 figure, 1 table. An earlier version was first posted on Zenodo in September 2026 (v1: doi:https://doi.org/10.5281/zenodo.22834532%3B all versions: doi:https://doi.org/10.5281/zenodo.22834531). Companion paper on task-relative information contracts: doi:https://doi.org/10.5281/zenodo.22819849 Subjects: Information Theory (cs.IT); Machine Learning (cs.LG) Cite as: arXiv:2610.08884 [cs.IT] (or arXiv:2610.08884v1 [cs.IT] for this version) https://doi.org/10.48550/arXiv.2610.08884 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[LG-155] Slow Beats Fast at the Kesten-Stigum Threshold: Minimax Fisher-Information and Belief-Propagation Characterizations of the Information-Computation Gap in Sparse Stochastic Block Models

链接: https://arxiv.org/abs/2610.08872
作者: Soroor Ghandali
类目: Information Theory (cs.IT); Machine Learning (cs.LG); Probability (math.PR); Statistics Theory (math.ST); Machine Learning (stat.ML)
*备注: 36 pages, 4 figures, 6 tables

点击查看摘要

Abstract:We study community recovery in the sparse symmetric stochastic block model with q communities, average degree d and signal strength \lambda through statistical decision theory and Fisher information, and obtain three characterizations of the Kesten-Stigum threshold d\lambda^2=1 and of the information-computation gap below it. First, on each community-size profile the minimax risk of any class of rules closed under averaging and vertex relabeling equals its Bayes risk under the uniform prior; the posterior mean is the unique Bayes rule and is admissible, and the Bayes risk of degree- D polynomial rules is the trivial risk times 1-\mathrmCorr_D^2 . Combined with known low-degree and information-theoretic results, this gives the gap as a worst-case statement: for q\ge 5 there is a window below the threshold in which no low-degree rule beats the trivial risk asymptotically, while an exponential-time rule does on a set of labelings of probability 1-o(1) . Second, the Fisher information about \lambda carried by cycle counts is a series with terms of order k(d\lambda^2)^k , convergent exactly when d\lambda^21 ; below the threshold the relative error of every unbiased cycle-based estimator of \lambda^k stays above an explicit constant, and every cycle-count test has success probability bounded below one. Third, the derivative of belief propagation at its uninformative fixed point multiplies a random perturbation by |\lambda|\sqrtd per iteration, and one EM step taken there leaves \lambda unchanged. A signal-to-noise computation recovers the condition d\lambda^1/\chi1 of Chin et al. for q=n^\chi communities and identifies personalized PageRank as a walk count with suboptimal weights. Experiments on networks with up to 3\times 10^5 vertices confirm the threshold for q=2 , the hard window for q=5 , and the many-community scaling.

[LG-156] Adversarial RL for Port-Scan Evasion: Attacker Feature Visibility in Edge-Deployed IDS

链接: https://arxiv.org/abs/2610.08864
作者: Logan Andrew North,Priya Sanjay Kaluskar,Shasi Kumar Ramachandran Prabhu,Peilong Li,Suman Saha
类目: Cryptography and Security (cs.CR); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Machine learning-based intrusion detection systems (IDS) are increasingly used in resource-constrained Internet of Things (IoT) environments, yet their robustness is often evaluated against static attacks rather than adversaries that adapt to detection feedback. This paper investigates adaptive port-scan evasion against ML-based IDS models deployed on a Raspberry Pi 3B+. We implement a live Zeek-based IDS pipeline with XGBoost, a multi-layer perceptron, and a 1D convolutional neural network trained on TON_IoT telemetry, and use a Deep Q-Network (DQN) adversary to learn evasive combinations of probe timing, TCP flags, and payload size under black-box, gray-box, and white-box feature-visibility settings. Although the deployed IDS models detect conventional port scans at 91.1–99.8%, DQN final-50-episode evasion rates range from 61.9% to 98.3% across feature-visibility settings. Greater feature visibility does not monotonically improve evasion, and its effect is model-dependent: against XGBoost, the black-box agent achieves 92.9% evasion, compared with 61.9% and 76.9% for gray-box and white-box agents, respectively, whereas 1D-CNN is most vulnerable under white-box access at 98.1%. Because standard DQN can overestimate action values, we additionally spot-check representative conditions using Double DQN. The gray-box condition remains unstable in this check, providing no evidence that overestimation bias alone explains the observed instability. These results show that limited feature knowledge can still enable effective adaptive evasion against static edge-deployed IDS models, motivating more robust defenses for IoT edge environments.

[LG-157] Autonomous Droplet Navigation via Model-Based Reinforcement Learning: Zero-Shot Transfer and Emergent Dynamics IROS2026

链接: https://arxiv.org/abs/2610.08852
作者: Rajneesh Anand,Mayuresh V. Kothare
类目: Robotics (cs.RO); Machine Learning (cs.LG); Systems and Control (eess.SY)
*备注: Accepted for presentation at the Robotics Automation in Self-Driving Laboratories 2026 Workshop, IROS 2026. this https URL

点击查看摘要

Abstract:Self-driving laboratories (SDLs) are transforming chemical and materials discovery through closed-loop automation, yet automated infrastructure for physical manipulation of soft, deformable matter remains beyond current robotic platforms. A critical instance is autonomous droplet transport on an open surface, where contact-angle hysteresis, capillary pinning, and surface heterogeneity produce partially observable dynamics that pose significant challenges for classical model-based controllers. We introduce the first robotic platform for closed-loop autonomous liquid droplet navigation on an open, unconfined surface using model-based reinforcement learning. A two-axis tilting board coated with a thin silicone oil film drives the droplet, while an overhead camera provides real-time feedback. A learned policy was trained on just 50 to 150 physical episodes depending on geometric complexity, without simulation or analytical models. Beyond performance alone, the platform demonstrates three capabilities of interest to the SDL community: it robustly transfers zero-shot to unseen geometries; it autonomously discovers an oscillatory depinning strategy to free the droplet when it sticks; and it completes its full training pipeline in under 90 minutes. These results extend reinforcement-learning manipulation from rigid microrobots to deformable soft-matter systems for next-generation SDLs.

[LG-158] A Vehicle-Integrated Approach to Digital Twin Deployment for Bridges Through Drive-By Sensing

链接: https://arxiv.org/abs/2610.08822
作者: Zihao Liu,Daigo Kawabe,Jiaji Wang,Chul-Woo Kim,Mehrisadat Makki Alamdari
类目: Machine Learning (cs.LG)
*备注: Extended abstract for 2nd International Conference on Engineering Structures (ICES2026)

点击查看摘要

Abstract:Ageing bridge infrastructure is a growing global concern, yet conventional Structural Health Monitoring (SHM) systems are costly and difficult to scale, and routine visual inspections remain subjective. Drive-by, or indirect, bridge inspection, in which a sensorised vehicle recovers structural information from vehicle-bridge interaction (VBI) and vehicle-road interaction (VRI) responses, offers a scalable alternative. However, key challenges remain unresolved, including separating bridge responses from road roughness, detecting damage under normal traffic, and generalising across diverse bridge types. This paper presents a vehicle-integrated digital twin framework that unifies physics-based modelling and machine learning for continuous monitoring of bridge and road conditions. The framework comprises three pillars. First, surrogate models of VBI and VRI are constructed using a Fourier Neural Operator that learns function-to-function mappings from operating conditions to vehicle responses. Trained on both simulated and field data, these surrogates deliver millisecond-scale inference, replacing computationally intensive full-order analyses. Second, the design of a custom electric inspection vehicle, its sensor layout, and signal processing chain are optimised through Bayesian optimisation to maximise bridge information yield while suppressing road and vehicle noise. Unsupervised damage-assessment pipelines based on adversarial autoencoders, matrix profiles, and transformer architectures have been developed and validated to process the resulting vehicle data. Third, the complete workflow is validated through coordinated multi-site field trials in Australia and Japan, covering a range of bridge types, traffic conditions, and environmental settings.

[LG-159] ask-Oriented Key-Layer KV Communication for Efficient Latent Multi-Agent Collaboration

链接: https://arxiv.org/abs/2610.08820
作者: Dongsen Zhang,Peipei Li,Zekun Li,Wenjun Xu
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Large language model-based multi-agent systems improve complex problem solving through collaboration, while latent communication directly transmits model internal states to avoid the high inference costs of natural language. However, existing KV-based latent communication methods prioritize sender-side state fidelity, leading to substantial communication and computation overhead and potentially introducing redundant information. To address these limitations, we revisit latent communication from a task-oriented perspective, shifting its objective from sender-side state fidelity to receiver-side task sufficiency. Under this formulation, we propose KITE, a training-free framework for task-oriented key-layer KV communication. KITE identifies a task-effective key layer using a receiver trajectory distortion criterion, transmits only the latent working memory associated with the key layer, and further uses the same layer as the entry point for autoregressive latent reasoning. Experiments on seven benchmarks across two model families and three model scales show that, compared with full-layer KV communication, KITE reduces communication volume by 28-36 \times , achieves up to 3 \times end-to-end inference speedup, and improves accuracy by up to 23.3 percentage points.

[LG-160] DenoFlow: Flow Matching for SSVEP Denoising under Real Physiological Artifacts

链接: https://arxiv.org/abs/2610.08817
作者: Zhentao He,Ziwei Wang,Dongrui Wu
类目: Machine Learning (cs.LG)
*备注: 14 pages, 6 figures

点击查看摘要

Abstract:Electroencephalography (EEG)-based brain-computer interfaces (BCIs), particularly steady-state visual evoked potential (SSVEP) systems, are highly vulnerable to noise and artifacts, which severely degrade decoding accuracy. Although recent denoising approaches have shown promise, they are fitted without paired ground truth, can settle on reproducing their input, and are optimized on waveform distance alone, which says nothing about whether the output stays decodable. To address these issues, we propose DenoFlow, which casts SSVEP denoising as transport: instead of learning a direct map from a contaminated trial to a clean one, a field network regresses the velocity of the straight path between them, following the rectified-flow formulation, and denoising integrates that field forward from the observation. The field network is an encoder-decoder that sees the contaminated trial at every layer and the path position at its bottleneck, and a classifier trained alongside it supervises the integrated output. Because the observation itself is both the conditioning input and the starting point of the integration, the model never generates a trial from noise, and training reduces to regression, removing the adversarial min-max game. To obtain paired data on datasets with no ground truth, we injected physiological artifacts of the recorded electromyography (EMG) and electrooculography (EOG) signals under a controlled signal-to-noise target. Experiments on two public SSVEP datasets with five popular SSVEP decoders showed that DenoFlow outperformed seven baseline denoising models on both signal fidelity and downstream decoding accuracy. Code is available at this https URL.

[LG-161] he Cost of Long Memory: State Context and Stability Complexity in Sequence Models

链接: https://arxiv.org/abs/2610.08816
作者: Yuheng Song
类目: Machine Learning (cs.LG)
*备注: 49 pages, 2 figures

点击查看摘要

Abstract:Long-range temporal dependence poses a resource question for sequence models: for a specified predictive-memory law, how much state, context, or dynamical criticality is required in order to forecast accurately? We study this question directly in forecasting risk. For algebraically decaying predictive memory, we prove matching upper and lower approximation bounds for exponential and finite-state modes. The best r -mode forecast error decays as e^-\Theta(\sqrt r) , so reaching forecast error \tau needs r=\Theta(\log^2(1/\tau)) states or modes. Earlier curse-of-memory results establish broad limitations of stable recurrent models under different approximation notions; here both sides match for one canonical predictive target in forecast risk, which fixes the optimal resource exponent for that target. We then show that genuine fractional long memory changes the geometry itself. In particular, forecast error is measured after fractional integration, prediction from a finite context of length L has an exact 1/L leading order, and a fixed fractional strength d keeps the square-log state-complexity law. Near the short-memory boundary, we identify the relevant d^2 and d^4 scales and give a uniform constructive law in the intermediate regime. For nonlinear contextual recurrences with uniformly contractive state dynamics, we derive an exponential first-chaos envelope and an explicit necessary condition that relates forecast accuracy to the contraction margin. Vanishing forecasting error on an algebraic target forces the recurrence quantitatively toward criticality, a condition that is necessary and not by itself sufficient. Finite-sample Kullback–Leibler calculations further connect the predictive geometry to statistical information. Theorem-matched experiments with contractive state-space, gated recurrent, and attention models reproduce the state and stability predictions.

[LG-162] Bounded Autonomy and Verifiable Safety for Agent ic AI Enabled Automation

链接: https://arxiv.org/abs/2610.08815
作者: Srini Ramaswamy,Deveeshree Nayak
类目: Machine Learning (cs.LG); Emerging Technologies (cs.ET); Systems and Control (eess.SY)
*备注: This paper has been accepted and will appear in the Journal of Intelligent and Robotic Systems ( https://doi.org/10.1007/s10846-026-02467-w )

点击查看摘要

Abstract:Agentic AI-enabled automation cannot be safely deployed in high-stakes environments on probabilistic reasoning alone. A recurring risk is epistemic drift: as reasoning deepens, system behavior may move away from subject-matter-expert constraints for safe operation. This paper presents BRaVeS, a bounded reasoning and safety-governance framework termed the Defensible Next-Gen Reasoning System (DNRS). BRaVeS encodes SME-defined constraints as invariant anchors, proposes MoDA-Style (Mixture of Depths Attention) depth-aware access as a candidate mechanism for keeping these anchors visible during inference, and uses a state hierarchy (SMARtAutonomy) to reduce autonomy as epistemic risk increases. To formalize bounded recovery, we introduce the Lyapunov-Bounded Consensus Framework (LBCF), which maps continuous epistemic-risk signals into a finite K-bag abstraction and applies shielded state transitions that enforce Lyapunov-style energy descent or route the system to a human-mediated terminal state. The formal convergence result applies to the finite LBCF abstraction under fixed thresholds and feasible-shield assumptions; it does not prove safety of the full continuous neural activation space. We evaluate the framework through a discrete event Monte Carlo simulation using HAI 22.04 industrial-control-system time-series data with synthetic noise and sensor-degradation regimes. Across the tested parameter-grouping strategies and thresholds, the LBCF process achieved finite-step convergence and no safety-guard violations. These results provide simulation-based evidence that bounded governance behavior can be enforced under the stated abstraction, while motivating future work on deployed transformer implementations, live human-in-the-loop validation, and broader adversarial settings.

[LG-163] A Bayesian Mirror Architecture for Emergent Consciousness: Circular Hierarchies Self-Manifolds and Hybrid Event-Self Binding

链接: https://arxiv.org/abs/2610.08792
作者: Eduardo Righi Capanema de Almeida
类目: Machine Learning (cs.LG)
*备注: 15 pages, no figures, foundational paper

点击查看摘要

Abstract:We present a foundational formulation of the Bayesian Mirror Architecture (BMA), a self-referential generative framework in which sensory abstractions, meta-abstractions, and a self-latent interact through circular recursion. The defining constraint is a closed update S_t - H_t-1, where a hybrid event-self latent H_t binds self-representations to abstract world models and reinjects this coupling into the self-state. Consciousness, in a restricted sense, is not an optimization objective nor a semantic label, but an architectural property of systems possessing this circular structure. Because inference operates over posterior beliefs, BMA’s intrinsic state space is a space of probability measures equipped with optimal-transport geometry. Stability and coherence are formulated in the 2-Wasserstein metric on P_2, yielding coordinate-free notions of self-stability and hybrid coherence along belief trajectories. We define a Causal Learning Regime (CLR) via bounds on Wasserstein belief drift together with an integration index capturing sustained coupling between self and world latents. CLR diagnoses whether the environment contains learnable causal structure; it is not a marker of consciousness. Global strict contractivity is not required: BMA may exhibit multiple coherent basins. We define self-manifolds basin-wise as supports of invariant measures under local Wasserstein contractivity. We identify Wasserstein epsilon-necks, transport bottlenecks where basins decouple, yielding a unique realized continuation in a vanishing-conductance limit. We interpret this selection as choice: internally determined yet externally unpredictable at finite resolution. Learning proceeds via variational free-energy minimization, with stability and agency emerging from what the environment affords to learn. Comments: 15 pages, no figures, foundational paper Subjects: Machine Learning (cs.LG) Cite as: arXiv:2610.08792 [cs.LG] (or arXiv:2610.08792v1 [cs.LG] for this version) https://doi.org/10.48550/arXiv.2610.08792 Focus to learn more arXiv-issued DOI via DataCite

[LG-164] HE-OFT: Privacy-Preserving One-Shot Federated Fine-Tuning under Homomorphic Encryption

链接: https://arxiv.org/abs/2610.08255
作者: Halil İbrahim Kanpak,Sinem Sav,Alptekin Küpçü
类目: Cryptography and Security (cs.CR); Machine Learning (cs.LG)
*备注: Technical report. 32 pages, 7 figures, 12 tables. Code: this https URL

点击查看摘要

Abstract:Many organizations adapt large pretrained models to their own tasks by fine-tuning on private data. Several of these parties often hold data for the same task and wish to fine-tune a model together without pooling that data. Federated learning (FL) enables joint fine-tuning, but reconstruction attacks on shared intermediate values (the model or its gradients) remain a privacy risk. A one-shot protocol that exchanges one encrypted contribution exposes no intermediate value. Such a protocol still gives the trained model to every participant, which is not permitted where the model is a regulated or proprietary asset. We present HE-OFT, the first cryptographically secure one-shot federated fine-tuning protocol in which no party receives the trained model. Each client fine-tunes a low-rank adapter and a classifier head on a frozen public backbone and keeps the adapter. The client uploads one encrypted head displacement, which the server combines under multiparty CKKS and never decrypts. A quorum of clients returns only the predicted label to the querier. On four text classification tasks and one vision task, HE-OFT reaches 61 to 79 per cent accuracy, against 20 to 48 per cent for a client training alone. HE-OFT keeps 85 to 96 per cent of the accuracy of a disclosed model. A test-time query takes 443.1 to 1713.1 s on one core, or 56.1 to 255.1 s with level restoration on a GPU. Restoring levels at the server cuts the traffic per query from up to 1.6 GiB to 13.5 MiB.

[LG-165] Fourier neural operator for real-time simulation of 3D dynamic urban microclimate

链接: https://arxiv.org/abs/2308.03985
作者: Wenhui Peng,Shaoxiang Qin,Senwen Yang,Jianchun Wang,Xue Liu,Liangzhu Leon Wang
类目: Machine Learning (cs.LG); Computational Engineering, Finance, and Science (cs.CE); Numerical Analysis (math.NA); Atmospheric and Oceanic Physics (physics.ao-ph); Fluid Dynamics (physics.flu-dyn)
*备注:

点击查看摘要

Abstract:Global urbanization has underscored the significance of urban microclimates for human comfort, health, and building/urban energy efficiency. They profoundly influence building design and urban planning as major environmental impacts. Understanding local microclimates is essential for cities to prepare for climate change and effectively implement resilience measures. However, analyzing urban microclimates requires considering a complex array of outdoor parameters within computational domains at the city scale over a longer period than indoors. As a result, numerical methods like Computational Fluid Dynamics (CFD) become computationally expensive when evaluating the impact of urban microclimates. The rise of deep learning techniques has opened new opportunities for accelerating the modeling of complex non-linear interactions and system dynamics. Recently, the Fourier Neural Operator (FNO) has been shown to be very promising in accelerating solving the Partial Differential Equations (PDEs) and modeling fluid dynamic systems. In this work, we apply the FNO network for real-time three-dimensional (3D) urban wind field simulation. The training and testing data are generated from CFD simulation of the urban area, based on the semi-Lagrangian approach and fractional stepping method to simulate urban microclimate features for modeling large-scale urban problems. Numerical experiments show that the FNO model can accurately reconstruct the instantaneous spatial velocity field. We further evaluate the trained FNO model on unseen data with different wind directions, and the results show that the FNO model can generalize well on different wind directions. More importantly, the FNO approach can make predictions within milliseconds on the graphics processing unit, making real-time simulation of 3D dynamic urban microclimate possible.

[LG-166] Best Arm Identification for Bandits with Shifting Means NEURIPS2026

链接: https://arxiv.org/abs/2610.10488
作者: Lukas Zierahn,Wouter M. Koolen,Shubhada Agrawal,Christina Katsimerou,Dirk van der Hoeven
类目: Machine Learning (stat.ML); Machine Learning (cs.LG)
*备注: Accepted at NeurIPS 2026

点击查看摘要

Abstract:We study the best arm identification problem in a stochastic environment with a novel form of adversarial perturbations, which we coin Shifting Means. While classically the mean rewards of the K arms are stable in time, in Shifting Means only the gaps \boldsymbol\Delta between mean rewards are stable, while their common shift may be determined adversarially in each round. The objective of the learner is to identify the best arm with high probability while minimizing sample complexity (the fixed confidence setting). Handling shifts requires new tools: we show that algorithms employing a Generalized Likelihood Ratio Test (GLRT) stopping rule, including the popular Track-and-Stop, fail under time-varying shifts. Instead, we propose Importance Weights for Shifting Means ( \mathsfISM ). Assuming means bounded by U and \sigma^2 -sub-Gaussian rewards, we show \mathsfISM to be \delta -correct and to enjoy a sample complexity bound of order K (\sigma^2 + U^2) \Delta_\min^-2 \ln \frac1\delta . We also present a matching (up to constant factors) worst-case lower bound and evaluate our results empirically.

[LG-167] Conditional Flow Matching for Generation of 3D Multi-variable Instantaneous Urban Microclimate Fields

链接: https://arxiv.org/abs/2610.10430
作者: Peng Liu,Shaoxiang Qin,Theodore Potsis,Lili Ji,Dingyang Geng,Liangzhu Leon Wang
类目: Fluid Dynamics (physics.flu-dyn); Computational Engineering, Finance, and Science (cs.CE); Machine Learning (cs.LG); Atmospheric and Oceanic Physics (physics.ao-ph)
*备注:

点击查看摘要

Abstract:Rapid and accurate prediction of urban wind and temperature fields is important for urban microclimate design and climate adaptation. Large-eddy simulation (LES) effectively resolves these instantaneous fields, but its application is limited in iterative design of urban microclimate applications due to high computational cost. Existing regressive data-driven models offers quick outputs, but they produce only deterministic point predictions that inherently fail to represent turbulent stochasticity. This paper adopts a novel generative framework of Conditional Flow Matching (CFM) that uses building geometry and mean flow as guidance to generate plausible three-dimensional instantaneous velocity and temperature fields for urban microclimate in seconds. To overcome the GPU memory bottleneck of pixel space 3D generation, the model operates in parallel on overlapping pixel space through a shared-noise initialization that preserves high spatial continuity of flow structure across the entire domain. Against reference LES data, the CFM surrogate can rapidly and accurately restore the first-order statistics with Normalized Root Mean Square Error (NRMSE) of 2.99% for wind and 1.77% for temperature, second-order turbulence metrics with NRMSE of 7.17% for wind and 8.84% for temperature, turbulent kinetic energy with NRMSE of 7%, probability density function and vertical profiles in representative locations. Wind engineering application of local gust prediction demonstrate that the speed and accuracy of CFM, supporting the use of generative AI for making turbulence-aware resilient urban design and climate adaptation more computationally feasible.

[LG-168] Derivative Gaussian Processes on a Two-Direction Budget

链接: https://arxiv.org/abs/2610.10428
作者: Hyunseok Seung,Matthias Katzfuss
类目: Machine Learning (stat.ML); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Gradient observations promise more accurate Gaussian process (GP) surrogates, but the cost of incorporating them has long stood in the way of realizing that promise. We propose a derivative GP with a budget of just two directions per observed gradient. One direction focuses on each gradient’s direct contribution to target prediction, while the other aggregates its indirect contributions through correlations with the conditioning function values. Within a Vecchia approximation, where each prediction conditions on m nearby inputs in d dimensions, this construction represents their md gradient coordinates using at most 2m directional derivatives, giving \mathcalO(m^3) dense factorization cost per prediction target. For general conditioning sets, we bound the posterior approximation error relative to using full gradients and characterize when the error is small or the approximation is exact. In simulations, our method matches the accuracy of a leading exact gradient-reduction method at equal conditioning set size. Because its cost grows much more slowly with that size, it can use conditioning sets well beyond the memory limit of the exact method, reaching lower prediction error with a small fraction of the time and memory. Notably, our method can exploit gradient observations while requiring less computation time or memory than function-only GP baselines.

[LG-169] Safe Meta-Policy Design with Risk Control

链接: https://arxiv.org/abs/2610.10393
作者: Wenbin Zhou,Michael Lingzhi Li,Shixiang Zhu
类目: Machine Learning (stat.ML); Machine Learning (cs.LG); Optimization and Control (math.OC)
*备注:

点击查看摘要

Abstract:Models can be retrained as new data arrive, but deploying every new version risks replacing a good policy with a worse one. We study how to plan policy updates (i.e., meta-policy) before future candidates are trained, balancing the benefits of improvement against the risk of performance regression. Our offline meta-policy maximizes expected cumulative value subject to a budget on the expected number of updates that perform worse than the policies they replace. We estimate the value and risk of possible switches from historical learning trajectories, represent an update schedule as a path in a directed acyclic graph, and select a schedule using dynamic programming. A leading-order analysis identifies the signal-to-noise ratio of policy improvement as a key driver of update frequency, waiting times, and risk allocation: clearer improvements support earlier, more frequent updates, while noisier improvements call for longer waits or greater risk expenditure. Their asymptotic rates also reveal a diminishing marginal cost of achieving greater safety over time. Experiments on synthetic and clinical trial data illustrate the performance–risk tradeoff and compare our method with alternative baselines.

[LG-170] Measurement-Efficient Differentiable Quantum Architecture Search for Combinatorial Optimization

链接: https://arxiv.org/abs/2610.10351
作者: Lukas Theißinger,Thore Gerlach,Christian Bauckhage
类目: Quantum Physics (quant-ph); Machine Learning (cs.LG)
*备注: 6 pages, 2 figures. Accepted at the 2026 IEEE 2nd International Conference on Quantum Artificial Intelligence (QAI). Code and data: this https URL

点击查看摘要

Abstract:Differentiable quantum architecture search (DQAS) is a promising framework for the automated design of quantum circuits, particularly for variational quantum optimization algorithms. However, its practical deployment on quantum hardware is limited by the large number of circuit measurements required during optimization, making hardware execution costly. In this work, we show that for a broad class of combinatorial optimization problems and commonly used rotational gate parameterizations, the measurement cost of DQAS can be significantly reduced without changing the optimization objective. We derive the proposed measurement reduction scheme theoretically and validate it experimentally on 3-SAT and MaxCut benchmark problems. Our approach reduces the requested gradient measurement cost by about 39 to 41% while introducing only negligible classical post-processing overhead, lowering the practical cost of executing DQAS on quantum hardware.

[LG-171] Dataset Pruning from First Principles: A Label-Free Linear Programming Approach

链接: https://arxiv.org/abs/2610.10347
作者: Rodrigo Schuller,Francisco Ganacim
类目: Machine Learning (stat.ML); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Dataset pruning reduces a large training set to a representative subset while preserving model performance. Existing geometry-based methods typically assume that nearby points in embedding space share similar properties. Rather than imposing this assumption, we derive geometric selection criteria by reformulating unbiased subset selection as a variance minimization problem. Unbiasedness ensures that unweighted subset averages recover full-dataset averages in expectation, including losses and gradients at fixed model parameters. Specifically, we characterize a family of unbiased subset selection algorithms as a high-dimensional polytope. In this context, minimizing the expected sampling variance is a linear objective. Differences in sampling variance, averaged over rigid motions, admit closed-form pairwise expressions. Because the polytope has high dimension, directly applying standard linear programming is impractical. We instead use these expressions to construct an efficient vertex walk that optimizes an approximation of the variance objective while preserving unbiasedness, yielding a method that requires neither labels nor model training during selection. Across CIFAR-10, MNIST, and CelebA benchmarks, our method matches or exceeds uniform sampling in mean test accuracy at every evaluated budget and outperforms competing geometric methods in several settings, particularly at small selection budgets. Beyond dataset pruning, the same framework reduces stochastic-gradient variance by increasing diversity within mini-batches while keeping the batch size unchanged.

[LG-172] Data Reuse in Non-Stationary Learning

链接: https://arxiv.org/abs/2610.10340
作者: Tomer Gafni,Garud Iyengar,Assaf Zeevi
类目: Machine Learning (stat.ML); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:We consider online learning in non-stationary environments, where the goal is to track an unknown parameter that switches abruptly between a finite set of recurring values. Recurrence opens the possibility of judiciously reusing past observations to improve algorithm performance. However, the changing nature of the underlying signal and lack of information on these dynamics may limit the ability to “safely” reuse data. In this paper we quantify some of the fundamental tradeoffs in this class of problems, and show that they bear a certain resemblance to the classical bias-variance dilemma. Specifically, we propose a class of anytime algorithms, dubbed Exposure-Capped Reuse (ECR), that combine online change detection, compatibility testing, and “contamination” control. We characterize the regime in which ECR’s regret scales with the number of distinct values rather than the number of changes, and derive a novel information-theoretic lower bound that establishes the near-minimax optimality of ECR. This provides rigorous quantification of the statistical “value” of data reuse.

[LG-173] Estimating Uncoded Crash Factors with Tabular Foundation and System One Models: Kumo Tabular and Jev

链接: https://arxiv.org/abs/2610.10321
作者: Amir Rafe,Subasish Das
类目: Applications (stat.AP); Machine Learning (cs.LG); Methodology (stat.ME)
*备注: 26 pages, 9 figures, 8 tables. Code: this https URL

点击查看摘要

Abstract:Road safety programs count the coded fields of police crash records, while the officer’s narrative, which often records factors the fields omit, is rarely read. A safety office thus cannot tell how much its counts miss or where to review. This study develops and evaluates a system that joins both views of the 5,601,890 Texas crashes from 2017 to 2025 into population estimates with stated validity. An in-context tabular foundation model, Kumo Tabular, reads the coded record of every crash, a calibrated System One model, Jev, reads the narratives of two probability samples, and human judgments recalibrate its probabilities. A multiwave predict-then-debias estimator joins the three tiers, and a second human tier drawn with recorded probabilities checks the estimates by design. For hydroplaning, medical episodes, fatigue, animals, and phone use, the narrative documents more injury crashes than the coded field, 15,074 against 7,340 for phone use, and the human check agrees with all fifteen estimates within its margin. A re-read list ranked by Kumo Tabular finds confirmed discordance 7 to 58 times as often as random reading. At the planning cost of human coding, one further round of human judgments would cut the root mean square relative half-width from 22.0 to 16.2 percent, against 21.2 for reading every narrative. Two calibrated readers of different views, joined by a sampling design, give a safety office counts, a discordance map, a validated re-read list, and a reading budget, with Kumo Tabular reading the table at 15 times the speed of TabPFN 3.5.

[LG-174] RoBART: Bayesian Additive Regression Trees with Tree-Specific Rotations

链接: https://arxiv.org/abs/2610.10214
作者: Jeongung Heo,Seonghyun Jeong
类目: Machine Learning (stat.ML); Machine Learning (cs.LG); Statistics Theory (math.ST)
*备注:

点击查看摘要

Abstract:Bayesian additive regression trees (BART) can require many splits to approximate boundaries misaligned with the predictor axes. RoBART assigns each tree a rotation shared by all internal nodes, retaining axis-aligned splits in rotated coordinates and constant leaves. We jointly propose a Givens rotation sequence and cutpoints on the resulting grid by Metropolis-Hastings and establish reversibility with respect to the conditional posterior with leaf means integrated out. For additive functions with component-specific rotations and anisotropic Hölder smoothness, we prove posterior contraction in empirical L_2 distance and for the noise standard deviation. Under the stated prior, design, and grid conditions, with fixed numbers of predictors, trees, and components and no more components than trees, the rate is a sum of componentwise rates determined by smoothness and the number of rotated coordinates used. We also establish a posterior contraction lower bound showing that there exist functions for which RoBART adapts to the intrinsic dimension but axis-aligned BART does not.

[LG-175] Computations of the slice genus and the unknotting number of links via machine learning

链接: https://arxiv.org/abs/2610.10206
作者: Yutong Dai,Oliver Hayman,András Juhász,Ludovico Morellato
类目: Geometric Topology (math.GT); Machine Learning (cs.LG); Machine Learning (stat.ML)
*备注: 72 pages, 34 figures

点击查看摘要

Abstract:Links are disjoint unions of circles smoothly embedded in S^3 . We use reinforcement learning and Bayesian optimisation to obtain new upper bounds on several link invariants that are not known to be algorithmically computable: the slice genus and the unknotting number for links, and the strong slice genus for algebraically split links. We also compute lower bounds using known invariants. Combining the upper and lower bounds, we obtain new exact values in many cases. Our unknotting agents can reproduce the non-additivity of the unknotting number for several counterexamples due to Brittenham and Hermiller, in some cases finding new unknotting trajectories.

[LG-176] Broadly Applicable Approximate MCMC for Switching Stochastic Differential Equations Using Uniformization and Time-Conditioned Factorized Neural Likelihood Estimation

链接: https://arxiv.org/abs/2610.10194
作者: Shion Hosoda,Michiaki Hamada
类目: Machine Learning (stat.ML); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Switching stochastic differential equations (SSDEs) describe continuous-time dynamics whose parameters switch according to a latent regime process that follows a continuous-time Markov chain (CTMC). By allowing dynamics to change between regimes, SSDEs represent heterogeneous system behavior and have been applied across diverse fields. However, Bayesian inference for SSDEs remains difficult, and existing SSDE inference methods have limited applicability, with restrictions such as noise-free observations, univariate states, linear drift, or state-independent diffusion. In this study, we propose an approximate Markov chain Monte Carlo sampler for SSDEs using uniformization and factorized neural likelihood estimation (FNLE), a simulation-based inference method. Uniformization provides an exact representation of the CTMC but requires SDE transition densities over arbitrary time intervals. We approximate these densities by training a time-conditioned FNLE model. The resulting sampler is broadly applicable to SSDEs without requiring analytically tractable transition densities. In synthetic-data experiments, our method recovered regime paths and parameters for three SSDE models for which previous methods have limited applicability. We also applied our method to a real dataset and detected a regime transition.

[LG-177] Universal Local Error and Realized Amplification for the First-Order EDM Predictor

链接: https://arxiv.org/abs/2610.10190
作者: Nicolas Brosse,Arnak S. Dalalyan
类目: Machine Learning (stat.ML); Machine Learning (cs.LG)
*备注: 71 pages; code: this https URL

点击查看摘要

Abstract:We analyze the first-order deterministic diffusion sampler of Karras et al. (2022), termed EDM, in 2-Wasserstein distance by separating two sources of error: local discretization error and its amplification by subsequent learned steps. We prove that local error admits a universal bound: for any data distribution with finite second moment, the one-step discretization error is quadratic in the step size, with an explicit constant that does not depend on the data distribution. Error propagation, in contrast, depends on the learned network. At high noise levels, we exploit the network parametrization of EDM to derive an explicit contraction criterion. At low noise levels, we measure propagation through the amplification realized on the distributions transported by the sampler; this realized amplification can be arbitrarily smaller than the worst-case Lipschitz constant. This analysis yields an O(e^\Lambda_K/K) global discretization error for K sampling steps, where \Lambda_K is the low-noise log-amplification. Experiments on a one-dimensional Gaussian mixture show how measured amplification accounts for slower error decay on finite sampling grids. Diagnostics on a pretrained CIFAR-10 model illustrate related stability mechanisms without certifying the global assumptions.

[LG-178] Kinetic Langevin Meets Split Gibbs: Accelerated Posterior Sampling for Imaging Inverse Problems with Diffusion Priors

链接: https://arxiv.org/abs/2610.10187
作者: Dai Hai Nguyen,Duc Dung Nguyen
类目: Machine Learning (stat.ML); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Split Gibbs sampling (SGS) is a popular framework for posterior sampling in Bayesian imaging inverse problems. It decouples a Gaussian data-fidelity term from a complex prior through an auxiliary variable, so the data variable is updated exactly and only the prior-side conditional is hard to sample. Existing samplers treat this conditional in one of two ways. Plug-and-play SGS runs a multi-step diffusion denoiser at every iteration, which is expensive and lacks non-asymptotic guarantees. Langevin-within-SGS takes cheap overdamped Langevin steps but needs many iterations. We propose RED-KLwSGS, which keeps the exact Gaussian update for the data variable and updates the auxiliary variable with underdamped (kinetic) Langevin diffusions driven by a one-shot denoising score, at the same per-iteration cost as Langevin-within-SGS. We prove non-asymptotic Wasserstein-2 convergence in continuous and discrete time for strongly log-concave priors. We also introduce Joint-RED-KLwSGS, which applies kinetic Langevin diffusions to both variables. Experiments with Denoising diffusion probabilistic models as diffusion priors on FFHQ and ImageNet datasets show faster convergence and high-quality image reconstruction.

[LG-179] Conformal Prediction for Spatially Dependent Data via Sequential Whitening

链接: https://arxiv.org/abs/2610.10168
作者: Ayush Baran Sen,Arkajyoti Saha
类目: Machine Learning (stat.ML); Machine Learning (cs.LG); Applications (stat.AP); Methodology (stat.ME)
*备注:

点击查看摘要

Abstract:Split conformal prediction uses prediction errors on held-out (calibration) data to determine how wide the prediction intervals should be. It guarantees distribution-free finite-sample coverage when these errors and the error at the target site are exchangeable. This assumption may fail under spatial dependence and nonrandom sampling geometry. Existing spatial methods use fitting residuals to remove the predictable part of spatial variation from calibration and target errors. However, the spatial variation that only the calibration residuals can predict remains in both the target and calibration errors, reducing the efficiency and stability of the interval. We address this by additionally conditioning on the calibration residuals sequentially, which scales to large networks through nearest-neighbour approximations. Under a correct working covariance and an elliptical residual law, the resulting interval has exact finite-sample coverage under any spatial design, and under further conditions it is asymptotically oracle efficient. We also bound coverage loss under covariance misspecification and develop a diagnostic that identifies regions at risk of undercoverage. In simulated data, our method produces narrower and more stable intervals than global and localized state-of-the-art alternatives. In a national PM2.5 application, it produces narrower intervals within the network and identifies regions at risk of coverage failure.

[LG-180] Policy Learning with Weak Signals

链接: https://arxiv.org/abs/2610.10167
作者: Benedikt Koch,Winston Chou,Aurélien Bibaut,Nathan Kallus
类目: Methodology (stat.ME); Machine Learning (cs.LG); Statistics Theory (math.ST)
*备注:

点击查看摘要

Abstract:Policy learning in digital experimentation faces three challenges: weak signal-to-noise ratios, rich covariate spaces, and massive data volumes. We formalize this regime by modeling treatment-effect estimates from increasingly fine covariate partitions as Gaussian observations with bounded signal-to-noise ratios. We establish that, in general, the optimal treatment policy is not learnable in this setting. Even learning the optimal policy value suffers from impractically slow rates. However, when treatment effects vary smoothly, we derive minimax-adaptive policies based on linear smoothers that achieve vanishing welfare regret. We demonstrate the practical value of our framework by applying it to large-scale real-world experiments at Netflix, showing that personalized linear-smoothing policies can dominate unpersonalized policies even in this challenging empirical setting.

[LG-181] Progress and Prospect of AI in ARPES Workflow

链接: https://arxiv.org/abs/2610.10140
作者: Sandy Adhitia Ekahana,Aalok Tiwari,Pratik Saud,Aaron Bostwick,Chris Jozwiak,Eli Rotenberg,Jyoti Katoch
类目: Other Condensed Matter (cond-mat.other); Machine Learning (cs.LG); Data Analysis, Statistics and Probability (physics.data-an)
*备注:

点击查看摘要

Abstract:Artificial intelligence (AI) is becoming an increasingly useful tool across the experimental sciences, including angle-resolved photoemission spectroscopy (ARPES), which routinely produces large, multidimensional datasets of electronic structure. Recent advances in AI and machine learning (ML) have opened new opportunities across the entire ARPES workflow, from automated sample preparation and real-time data acquisition to post-experiment data analysis and comparison with theoretical calculations. Despite this progress, a comprehensive review of ML applications, their capabilities, and reliability across the different stages of ARPES workflow is still lacking. In this review, we first introduce ML methods that are most relevant to experimentalists working in condensed matter physics and materials science. We then follow the ARPES workflow, reviewing existing ML applications at each step and discussing their advantages, limitations and potential for future development. We also examine the current ARPES data landscape, where several open databases are available but remain relatively small and fragmented compared with large, shared datasets such as ImageNet. Given these limitations, we suggest that the community focus on sharing pretrained models that can be further trained, adapted to specific tasks, and redistributed, while working toward a larger and standardized open ARPES dataset repository. Finally, we discuss our perspectives on the future of AI within the ARPES workflow using a six-level framework of laboratory automation, highlighting the opportunities and challenges in moving toward a fully autonomous, self-driving ARPES laboratory.

[LG-182] Sharp Asymptotic Theory of Maximum Likelihood Estimation for Gaussian Processes with an RBF Kernel

链接: https://arxiv.org/abs/2610.10080
作者: Ameer Qaqish,Didong Li
类目: atistics Theory (math.ST); Machine Learning (cs.LG); Probability (math.PR)
*备注:

点击查看摘要

Abstract:Gaussian processes (GPs) are widely used across machine learning, spatial statistics, time-series analysis, optimization, Bayesian statistics, and scientific applications. A central component of a GP model is its kernel, which is typically specified through a parametric family. Among the most widely used choices is the radial basis function (RBF), also known as the squared exponential or Gaussian kernel, owing to its simple form, smoothness, and flexibility. In practice, the kernel parameters are routinely estimated by the maximum likelihood estimators (MLEs), as implemented by standard GP software. Despite this widespread use, the asymptotic behavior of the MLEs remains poorly understood under fixed-domain asymptotics, even for the RBF kernel. The main difficulty arises from the increasingly strong dependence among densely sampled observations and the nonlinear dependence of the covariance matrix on the kernel parameters. In this paper, we address this gap by providing, to the best of our knowledge, the first complete asymptotic characterization of the joint MLE of the spatial variance, lengthscale, and nugget variance under fixed-domain asymptotics. We establish consistency, derive convergence rates for all three parameters, prove joint asymptotic normality, and show that these rates are minimax optimal. Subjects: Statistics Theory (math.ST); Machine Learning (cs.LG); Probability (math.PR) Cite as: arXiv:2610.10080 [math.ST] (or arXiv:2610.10080v1 [math.ST] for this version) https://doi.org/10.48550/arXiv.2610.10080 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[LG-183] owards Calibrated Probabilistic Forecasts for Events of Interest via Outcome-Conditional Recalibration

链接: https://arxiv.org/abs/2610.10076
作者: Jakob Benjamin Wessel,Sam Allen
类目: Machine Learning (stat.ML); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Calibration is an essential requirement for probabilistic predictions to be useful for decision making. While state-of-the-art prediction methods often yield miscalibrated predictive distributions, several post-hoc recalibration schemes have been proposed to generate calibrated predictions. However, popular recalibration schemes can conceal miscalibration in specific regions of the outcome space. Since particular outcomes, such as extreme events, often matter most for decision making, probabilistic predictions should be calibrated when evaluation is restricted to these outcomes. Hence, in this paper, we introduce outcome-conditional recalibration, a post-hoc method to recalibrate probabilistic predictions on user-defined regions of the outcome space. The method is simple, easy to implement, and can be applied to arbitrary predictive distributions. It works by applying the quantile recalibration approach of Kuleshov et al. (2018) to forecast conditional distributions, before rescaling these conditional distributions so that forecast event probabilities match empirical occurrence frequencies. This produces valid and continuous predictive distributions that are calibrated within each region of interest. Across regression benchmarks, we demonstrate that existing recalibration schemes do not necessarily yield calibrated predictions when interest is on particular outcomes, and that our approach improves outcome-conditional calibration relative to existing conditional and unconditional recalibration methods, while retaining competitive calibration overall. In an application to day-ahead electricity price forecasting, the approach substantially improves calibration when predicting negative prices, at negligible cost to forecast accuracy.

[LG-184] Gaussian Equivalence for Multi-Head Self-Attention

链接: https://arxiv.org/abs/2610.10033
作者: Tomohiro Hayase,Ryo Karakida
类目: Machine Learning (stat.ML); Machine Learning (cs.LG); Probability (math.PR)
*备注:

点击查看摘要

Abstract:A theoretical understanding of multi-head self-attention is fundamental to the study of modern neural networks. Using random matrix theory, we establish Gaussian equivalence for multi-head self-attention: replacing softmax attention with rescaled scores plus Gaussian noise preserves the limiting spectral law of the centered output. This equivalence also covers value and output projections that depend on the keys. The resulting laws separate the effects of head allocation and projection widths, and distinguish spectrum-preserving across-head sharing from within-head key–value dependence.

[LG-185] Structure alone supports efficient visual computation in the Drosophila visual system

链接: https://arxiv.org/abs/2610.10023
作者: Eudald Correig-Fraga,Roger Guimerà,Marta Sales-Pardo
类目: Neurons and Cognition (q-bio.NC); Disordered Systems and Neural Networks (cond-mat.dis-nn); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Understanding the extent to which measured synaptic wiring determines computation remains a central challenge. Here, we couple the proofread adult Drosophila melanogaster connectome to an anatomically faithful model of its eye. Visual information is inputted in the eye model, then passed to the connectome, and finally read from a Kenyon-cell-centered linear decoder. This creates a connectome-only model in which the anatomical graph and eye geometry are fixed and only scalar synaptic gains and neuronal thresholds may be learned. The model supports multitask vision, including color discrimination, shape classification, and numerical discrimination that follows a ratio-dependent scaling characteristic of approximate number perception. To test whether precise connectivity is consequential under wiring economy, we compare the biological graph to randomized ensembles that increasingly preserve biological synaptic constraints. At matched wiring cost, the biological network consistently yields higher accuracy, whereas less constrained rewiring surpasses it at the cost of inflated wiring. These findings indicate that the measured connectivity and eye geometry jointly set efficient operating points for visual computation.

[LG-186] Controlling Dependence in Implicit Generative Models via Spread Mutual Information

链接: https://arxiv.org/abs/2610.10021
作者: Jiahao Yu,Song Liu,José Miguel Hernández-Lobato,RuiKang OuYang
类目: Machine Learning (stat.ML); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Mutual information (MI) provides an objective for suppressing or encouraging statistical dependence in implicit generative models. However, direct MI evaluation is challenging in implicit models due to typically intractable densities. A remedy is estimating the generator gradient from the difference between conditional and marginal scores. This score difference can, in turn, be estimated by differentiating a log density ratio learned through classification. This construction nevertheless faces two difficulties: (i) singular distributions need not admit the required score functions, and (ii) poor overlap can hinder density-ratio estimation. We therefore introduce Spread Mutual Information (SMI), a weighted integral of MI across noise levels obtained by applying a common spreading kernel to the generated variable. Gaussian spreading yields smooth, strictly positive conditional and marginal densities, extending the gradient construction to distributions that may originally be singular. Across a variaty of experiments, SMI consistently achieves effective dependence control among MI-based methods and remains competitive with established task-specific approaches.

[LG-187] Extreme Binary Classification: Extreme Value Theory for Extreme Constraint on False Negative

链接: https://arxiv.org/abs/2610.09984
作者: Samuel Gruffaz,Muhammad Fawad,Jaakko Nevalainen
类目: Machine Learning (stat.ML); Machine Learning (cs.LG); Applications (stat.AP); Methodology (stat.ME)
*备注:

点击查看摘要

Abstract:While binary classification is one of the most extensively studied problems in machine learning, the regime in which the goal is to learn a classifier with an almost zero false negative rate remains largely unexplored. In this paper, we introduce the Extreme Binary Classification problem, where the objective is to learn a classifier whose false negative rate \alpha is constrained by \epsilon_N_1=o_N_1\to\infty(1/N_1) , with N_1 denoting the number of positive examples in the training set. To address this problem, we propose a threshold adaptation method theoretically grounded in guarantees derived from Extreme Value Theory, together with a feature selection procedure based on a permutation test applied to sample maxima. Experimental results on four real-world datasets of varying sizes demonstrate that our approach compares favorably with state-of-the-art methods. In addition, we illustrate its interpretability through an application to a cancer screening dataset. Subjects: Machine Learning (stat.ML); Machine Learning (cs.LG); Applications (stat.AP); Methodology (stat.ME) MSC classes: 62Cxx Cite as: arXiv:2610.09984 [stat.ML] (or arXiv:2610.09984v1 [stat.ML] for this version) https://doi.org/10.48550/arXiv.2610.09984 Focus to learn more arXiv-issued DOI via DataCite (pending registration) Submission history From: Samuel Gruffaz [view email] [v1] Wed, 7 Oct 2026 12:49:42 UTC (292 KB) Full-text links: Access Paper: View a PDF of the paper titled Extreme Binary Classification: Extreme Value Theory for Extreme Constraint on False Negative, by Samuel Gruffaz and 1 other authorsView PDFHTML (experimental)TeX Source view license Current browse context: stat.ML prev | next new | recent | 2026-10 Change to browse by: cs cs.LG stat stat.AP stat.ME References Citations NASA ADSGoogle Scholar Semantic Scholar export BibTeX citation Loading… BibTeX formatted citation loading… Data provided by: Bookmark checked="checked"class=“labs-tab-input”> Bibliographic Tools Bibliographic and Citation Tools Bibliographic Explorer Toggle Bibliographic Explorer (What is the Explorer?) Connected Papers Toggle Connected Papers (What is Connected Papers?) Litmaps Toggle Litmaps (What is Litmaps?) scite.ai Toggle scite Smart Citations (What are Smart Citations?) Code, Data, Media Code, Data and Media Associated with this Article alphaXiv Toggle alphaXiv (What is alphaXiv?) Links to Code Toggle CatalyzeX Code Finder for Papers (What is CatalyzeX?) DagsHub Toggle DagsHub (What is DagsHub?) GotitPub Toggle Gotit.pub (What is GotitPub?) Huggingface Toggle Hugging Face (What is Huggingface?) ScienceCast Toggle ScienceCast (What is ScienceCast?) Demos Demos Replicate Toggle Replicate (What is Replicate?) Spaces Toggle Hugging Face Spaces (What is Spaces?) Spaces Toggle TXYZ.AI (What is TXYZ.AI?) Related Papers Recommenders and Search Tools Link to Influence Flower Influence Flower (What are Influence Flowers?) Core recommender toggle CORE Recommender (What is CORE?) Author Venue Institution Topic About arXivLabs arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv’s community? Learn more about arXivLabs. Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?) mathjaxToggle(); We gratefully acknowledge support from our major funders, member institutions, , and all contributors. About Help Contact Subscribe Copyright Privacy Accessibility Operational Status (opens in new tab) Major funding support from

[LG-188] Possibilistic Radial Transport for Approximate IM Inference

链接: https://arxiv.org/abs/2610.09956
作者: Jungeum Kim,Percy Zhai
类目: Machine Learning (stat.ML); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Probing the hypothesis space after seeing the data remains valid under possibilistic inferential models (IMs), provided the significance level stays fixed. The price is computation, as each plausibility is a supremum of the possibility contour over the hypothesis, and the contour itself is approximated at each queried parameter value. We propose a possibilistic radial transport, which hides the contour value of a parameter in the radius of its source point. When a transport that maximizes within-shell entropy is picked, sampling parameters covering a confidence cut becomes a matter of truncating the radius. We provide a deep learning algorithm that enforces the contour depth condition while maximizing the entropy within each shell. Our amortization makes coverage and power assessments of the learned approximation practical as well as predictive check of new datasets. We also use the sampler to construct a Bel-Pl spectrum for comparing and selecting interpretable hypotheses that satisfy a prescribed Bel-Pl decision criterion. In simulations the learned contours match or improve on ellipsoidal approximations to the cuts, while the coverage and power track the exact reference. Finally, we probe hypotheses about ovarian aging using synthetic AMH records, asking for each woman how many more years her median AMH level will remain above a specified reference value.

[LG-189] Learning joint probabilistic weather forecasts from station observations alone

链接: https://arxiv.org/abs/2610.09898
作者: Chaeyeon Yi,Yun Am Seo
类目: Applications (stat.AP); Machine Learning (cs.LG); Atmospheric and Oceanic Physics (physics.ao-ph)
*备注: 62 pages, 6 figures, 3 Extended Data figures, 7 Extended Data tables; includes Supplementary Information

点击查看摘要

Abstract:Assessing compound weather risks requires forecasts representing dependence between variables. CLARA (Calibrated Advection-Routing Attention) learns joint Gaussian predictive distributions of five surface variables from station observations alone, without numerical weather prediction or reanalysis; the approximately 28,000-parameter model supports CPU training and prediction. Across six multi-year folds on 96 stations, its lead-mean energy score is 4.9% lower than that of a learned comparator with matched temporal inputs (4.7% with a similar parameter count) and 11-65% lower than those of statistical baselines. Holding marginal variances fixed, removing learned correlations worsens joint negative log-likelihood by 1.0-2.8 nats per station. A covariance-scale estimator, proved consistent under stated assumptions, improves short-lead calibration but over-corrects at long leads. Synthetic interventions show an attention-bias coefficient alone does not measure forecast influence. Retrained in ten regions on six continents, CLARA outperforms persistence in all 60 multi-year region-lead comparisons and a similarly sized learned model in 57 of 60.

[LG-190] Sparsifying Stochasticity Not Capacity: Partial Stochasticity via Deep Weight Factorization of Prior Scales

链接: https://arxiv.org/abs/2610.09886
作者: Marius P Linhard,Maurizio Filippone
类目: Machine Learning (stat.ML); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Bayesian neural networks need not be fully stochastic to be universal conditional density approximators, but it remains open which parameters should be stochastic. We learn this split by applying deep weight factorization to the prior scales, which are the standard deviations of the parameter priors, while fitting the functional prior to a Gaussian process with a maximum mean discrepancy objective. A parameter whose prior scale falls below a cutoff becomes deterministic and is optimized during inference, so the regularizer sparsifies stochasticity rather than capacity. We give a certificate for universal conditional density approximation that is checkable in linear time, together with a minimal repair when it fails. We further show that the common hybrid scheme of sampling some parameters and optimizing the others is stochastic approximation for a type-II maximum a posteriori objective, and that coupled step sizes can leave a tracking error that does not vanish as the step size shrinks. On a bimodal target, the learned split stays close to an unconstrained reference across all budgets and is insensitive to the cutoff, while random masks that distribute the same prior scales across layers are worse by up to two orders of magnitude. On UCI benchmarks, our method performs on par with a fully stochastic network while keeping about half of its parameters deterministic.

[LG-191] Origins of Universal Machine Learning Force-Field Errors in Multicomponent Materials

链接: https://arxiv.org/abs/2610.09837
作者: Hongwei Du,Dingyang Lv,Baole Wei,Yu Ren,Feng Yu,Xin He,Bonan Zhu,Yongda Huang,Yongheng Li,Jianjun Liu,Siqi Shi,Hong Wang,Ziheng Lu
类目: Materials Science (cond-mat.mtrl-sci); Machine Learning (cs.LG); Computational Physics (physics.comp-ph)
*备注: 20 pages, 10 figures

点击查看摘要

Abstract:Universal machine learning force-field generalization to multicomponent environments generated by compositional design remains insufficiently assessed. We construct a benchmark of 7,599 multicomponent configurations inspired by high-entropy design, elemental substitution and anion mixing. Eleven pretrained models are evaluated against density functional theory for energies, forces and stresses, with assessment extended to elastic, vibrational and adsorption-related properties. Force errors are analysed through training-reference coverage, local geometric heterogeneity, distance directionality and elemental response. Distances to training-reference environments reveal a qualitative association between coverage differences and increasing errors, while substantial variation remains at similar distances. Higher-error groups show greater local geometric heterogeneity, although OMat24 provides broad coverage of these environments. Relative to training-reference pair medians, errors remain low near the median, rise steeply on the compression side and increase more weakly on the extension side. After matching element pairs and absolute distance deviations, compression-side force errors are 1.81-1.95 times extension-side errors. Model-predicted pairwise interaction curves show greater curvature under compression. Fitting difficulty in independent elemental systems correlates with electronic band-energy responses to atomic displacements and Fermi-level shifts, and a similar pattern is observed in multicomponent systems. In parameter-matched comparisons, spherical-harmonic representations with maximum degrees of 2 and 4 lower test force errors for 38 and 40 of 43 elements, respectively, while differences in elemental difficulty remain. These findings inform force-field selection for experimental compositional design and identify targets for training-data sampling and model representations.

[LG-192] A Strength-Monotonic Law for Domain Alignment in Frozen-Embedding Bioacoustic Classification

链接: https://arxiv.org/abs/2610.09737
作者: Yucheng Gong,Rui Zhou,Binbin Zeng,Qiang Ren,Hongjin Hui
类目: Audio and Speech Processing (eess.AS); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:When does distribution alignment help a frozen foundation-model embedding generalize across acoustic domains? For cross-domain mosquito-species classification we report a strength-monotonic law: the stronger an encoder is on the target task, the more its unseen-domain generalization relies on a distribution-alignment (MMD) term, and the more it is harmed by domain-rebalanced sampling. Across four encoder families and a within-encoder HuBERT layer sweep (n=8), the rebalancing leg orders exactly with encoder strength (Spearman -1.000), while the MMD-benefit leg is monotonic within each stream and -0.857 pooled; fixing architecture and varying only representation strength flips the rebalancing effect from benefit to collapse. The law is actionable: a single MMD term is the sole lever on a strong encoder, so we reduce the field’s default recipe to a frozen Perch 2.0 embedding, a lightweight probe, cross-entropy, one MMD, and input augmentation. The reduced recipe stays within seed noise of the full composite (BA_unseen 0.299+/-0.006 vs. 0.307+/-0.014). As boundary conditions of the same law, three community defaults (backbone fine-tuning, multi-modal fusion, and domain rebalancing) each hurt unseen-domain accuracy under a leave-domain protocol, shown with single-variable, multi-seed evidence. We present a mechanism and the recipe it explains, not a leaderboard entry.

[LG-193] Unbounded Characteristic and Universal Kernels

链接: https://arxiv.org/abs/2610.09731
作者: Jose Cribeiro-Ramallo,Florian Kalinke,Zoltán Szabó
类目: atistics Theory (math.ST); Machine Learning (cs.LG); Machine Learning (stat.ML)
*备注:

点击查看摘要

Abstract:Kernel methods are among the most powerful tools in machine learning and statistics, with a large number of successful applications. Their immense success stems from the flexible function class associated to each kernel—its reproducing kernel Hilbert space (RKHS)—which facilitates statistical analysis, as well as from their computational tractability and applicability to many domains. Multiple notions (such as characteristic, L_p -universal, and integrally strictly positive definite) capture the expressivity of kernels and their RKHSs and play a key role in understanding the statistical properties of kernel methods; these concepts and their relations are well-understood for bounded kernels. Even though unbounded kernels have received significant attention over the past decade (for instance, in the construction of kernel-based discrepancy and dependence measures such as the maximum mean discrepancy, the Hilbert-Schmidt independence criterion, and the kernel Stein discrepancy), surprisingly little is known about the relations of these notions in the unbounded case. In the present paper we tackle this severe bottleneck, establishing their relations under mild assumptions.

[LG-194] Pareto-optimal quantum kernel selection for unsupervised anomaly detection on real malware beaconing data

链接: https://arxiv.org/abs/2610.09717
作者: Boaz Micah,Nadia Milazzo,Maissa Beji,Borja Aizpurua,Llorenç Espinosa-Portalés,Esteban Payares,Ghada Ben Slama,Luc Andrea,Michel Kurek,Thomas Cope,Olivier Salomon
类目: Quantum Physics (quant-ph); Cryptography and Security (cs.CR); Machine Learning (cs.LG)
*备注: 13 pages, 5 figures

点击查看摘要

Abstract:Quantum kernel methods are leading candidates for a practical quantum advantage in machine learning, but assessing that potential requires two quantities usually reported separately: how well a kernel performs on the task, and how far its geometry departs from the classical kernels available for the same problem. We introduce a fully unsupervised, multi-objective protocol that optimises simultaneously the normalised pseudo discrepancy (NPD), a label-free proxy for anomaly detection quality, and the geometric difference (GD) to a tuned classical reference kernel, selecting models from the resulting Pareto front. We apply it to malware beaconing detection in real network traffic, using a one-class support vector machine with fidelity and projected quantum kernels over four data encodings, on simulators and on IQM’s 20-qubit Garnet processor. NPD-guided selection alone finds a fidelity kernel that beats the tuned classical baseline, but with a geometric difference too small to certify the gain as quantum. Projected kernels reach far larger geometric differences; the Pareto-selected one only marginally exceeds the baseline (AUC 0.782 versus 0.765 , g_C\to Q\approx 89\sqrtN relative to that reference kernel), still below the NPD-selected fidelity kernel ( 0.840 ).

[LG-195] Boundary-aware Reinforcement Learning for Hypercube State Spaces via Deterministic Policy Gradient

链接: https://arxiv.org/abs/2610.09712
作者: Lijun Bo,Yijie Huang,Chenhao Lu
类目: Optimization and Control (math.OC); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:We develop a continuous-time deterministic policy gradient framework for reinforcement learning with reflected state dynamics, where the state process is governed by a controlled reflected stochastic differential equation on a hypercube. Under suitable regularity assumptions, we establish the connection between the value function and the Neumann Bellman equation, introduce an advantage-rate function that yields a deterministic policy gradient formula, and prove the martingale characterization theorem. Motivated by these theoretical results, we propose a continuous-time deep deterministic policy gradient algorithm for reflected stochastic systems, in which the Neumann boundary condition is imposed via either soft penalization or hard architectural constraint. We further quantify the discrepancy between the ideal continuous-time dynamics and the discretely sampled exploratory dynamics executed in practice, showing that the error decays as the time grid is refined and exploration noise vanishes. Our experiments on reservoir control problems illustrate the effectiveness of the RL framework, highlighting that boundary-aware methods substantially reduce Neumann boundary residuals and enhance learning stability.

[LG-196] he Silhouette Operator: Identifiability of Low-Rank Measures from One-Dimensional Projections

链接: https://arxiv.org/abs/2610.09687
作者: Robert A. Vandermeulen
类目: atistics Theory (math.ST); Information Theory (cs.IT); Machine Learning (cs.LG); Functional Analysis (math.FA); Machine Learning (stat.ML)
*备注:

点击查看摘要

Abstract:Structured recovery phenomena, such as restricted isometry properties in compressed sensing, have shown that high-dimensional objects can often be reconstructed from remarkably low-dimensional linear measurements. This work develops an analogous recovery framework for low-rank signed measures on \mathbbR^2 , defined here as measures that can be expressed as finite sums of product measures with one-dimensional factors. The framework is based on linear operators, termed “silhouette operators,” that map a measure to a fixed finite collection of one-dimensional linear pushforwards. The main results show that a suitably chosen collection of 2k projected marginals suffices to identify every compactly supported rank- \le k signed measure, that this number is optimal, and that the projection directions cannot be chosen arbitrarily. The framework is also extended to higher-dimensional sums of product measures by establishing sufficient conditions under which collections of pairwise marginals identify the full model. Building on this framework, a computationally efficient estimator, termed “silhouette mixture estimation” (SME), is introduced for constructing a low-rank empirical measure from data by matching its one-dimensional projected marginals to the corresponding empirical marginals in Wasserstein distance. When combined with one-dimensional density estimators, SME yields an efficient nonparametric density estimator that performs strongly relative to a range of parametric, nonparametric, and deep-learning baselines in settings of moderate dimension and sample size.

[LG-197] Q-PhotoMarket: A Design Space Exploration Framework for Photonic Hybrid Quantum Neural Networks in Financial Market Prediction

链接: https://arxiv.org/abs/2610.09641
作者: Alberto Marchisio,Hanzalah Mohamed Siraj,Muhammad Kashif,Nouhaila Innan,Muhammad Shafique
类目: Quantum Physics (quant-ph); Machine Learning (cs.LG)
*备注: To appear at the IEEE International Conference on Quantum Artificial Intelligence (QAI), Nottingham, UK, December 2026

点击查看摘要

Abstract:Photonic quantum computing has recently emerged as a promising platform for hybrid quantum machine learning due to its native realization of linear-optical circuits and the computational complexity of boson sampling. However, despite growing interest in quantum methods for finance, the influence of photonic circuit design choices on predictive performance remains largely unexplored. Existing studies typically evaluate a single architecture, leaving the broader photonic design space unexamined. In this work, we present Q-PhotoMarket, a systematic design space exploration (DSE) framework for photonic hybrid quantum neural networks (HQNNs) applied to financial market prediction. We explore over 5,000 valid photonic configurations spanning input photon states, circuit architectures, entangling models, and measurement strategies across their compatible computation spaces, for U.S., Indian, and cryptocurrency markets. To improve search efficiency, the exhaustive exploration is complemented with Bayesian optimization. We further incorporate threshold calibration and prediction-collapse diagnostics to enable reliable evaluation under increasingly imbalanced return thresholds. Experimental results show that systematic exploration of more than 5,000 photonic HQNN configurations reveals consistent architectural patterns across financial markets, identifies robust high-performing designs, and demonstrates competitive performance relative to classical machine learning baselines.

[LG-198] Quantum anomaly detection in real scarce data

链接: https://arxiv.org/abs/2610.09635
作者: Emanuele Casciaro,Fabio Mascherpa,Alfonso Amendola,Filippo Caruso
类目: Quantum Physics (quant-ph); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Anomaly detection on small and unbalanced datasets remains very challenging in machine learning, although this scenario is common in several domains, including healthcare, cybersecurity, finance, and energy. Data augmentation and generative AI may mitigate training-data scarcity, but they often fall short because anomalies are, by definition, unpredictable, rare, and highly diverse events compared to high-probability normal data. Overfitting to pseudo-anomalies, model collapse, high-dimensional data, uninterpretable black-box models, and validation challenges are typical issues limiting their practical applicability. In this context, quantum machine learning may provide a promising and more sustainable avenue because it can enable more interpretable models with far fewer trainable parameters and smaller datasets, implementable on energy-efficient quantum hardware. Here, we propose a novel two-step hybrid classical–quantum architecture for sequential data and test it on a realistic scenario in the global energy-transition domain, i.e., automated anomaly detection in large-scale photovoltaic plants. The achieved generalization capability and competitive prediction accuracy may pave the way for new hybrid learning models able to exploit the continuously increasing power of cloud-available and more sustainable quantum accelerators integrated with more traditional energy-hungry High Performance Computing resources.

[LG-199] Second-order optimization of variable projection SVM models and road abnormality detection

链接: https://arxiv.org/abs/2610.09617
作者: Andrea Angino,Matthias Voigt,Rolf Krause,Tamás Dózsa
类目: ignal Processing (eess.SP); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:We introduce a novel second-order optimization framework for minimizing so-called variable projection functionals. We demonstrate that the proposed framework is especially usefulfor the training of variable projection based kernel methods. In particular, the problem of efficiently training variable projection support vector machines (VP-SVMs) is considered. We show the effectiveness of the proposed training methodology in a real-world application, namely we demonstrate how second-order trust region algorithms can be used to train VPSVM models to recognize road surface abnormalities based on 1D signals obtained from a tire sensor.

[LG-200] Residual Learning in Empirical Asset Pricing

链接: https://arxiv.org/abs/2610.09613
作者: Dexin Peng,Xiaoyu Wang
类目: atistical Finance (q-fin.ST); Machine Learning (cs.LG)
*备注: 58 pages, 8 figures

点击查看摘要

Abstract:Shallow models are special cases of deep models, and deep models theoretically have the potential to outperform the shallow ones. However, the existing empirical asset pricing literature provides strong benchmarks for shallow models. Residual learning allows neural network models in asset pricing to go deeper by preserving and refining their shallow counterparts. The out-of-sample Sharpe ratio for value-weighted long-short portfolios of deep residual models (2.07) is higher than that for the corresponding shallow ones (1.92) and more than twice that of the deep feedforward models (0.89). We show that model depth is a source of additional economic value in asset pricing. Residual learning can be used to deepen other neural-network-based asset pricing models if they contain intermediate layers. Our design also provides one way to scale asset pricing models, making native “large asset pricing models” more feasible.

[LG-201] Scalable Logistic Gaussian Process Density Regression with Kinetic Langevin Sampling

链接: https://arxiv.org/abs/2610.09591
作者: Daniel Paulin,Ádám Jung,András A. Benczúr
类目: Machine Learning (stat.ML); Machine Learning (cs.LG)
*备注: 27 pages, 3 figures

点击查看摘要

Abstract:Conditional density estimation targets the full distribution of a response given covariates, as required, for example, for per-galaxy photometric redshifts. We develop a scalable Bayesian estimator based on the logistic Gaussian process. The log conditional density has a separable covariance: a Matérn kernel along the response, represented in a truncated Fourier basis on a circle, and a covariate kernel represented by Nyström features, which accommodate non-stationary kernels with input-dependent amplitudes and length scales. Instead of a Laplace or variational approximation, we sample the latent field of this finite-feature model. Given the hyperparameters, its posterior is strongly log-concave with a uniformly bounded Hessian, and we draw from it by simulating kinetic Langevin dynamics with symmetric minibatch splitting in Kronecker-whitened coordinates. Marginal-likelihood gradients follow from Fisher’s identity as posterior expectations. Under the conditions of our analysis their bias is controlled by the sampler’s step size and run length, and the predictive averages over the non-Gaussian latent posterior instead of a Gaussian around its mode. On photometric-redshift benchmarks with up to 3.9 million training observations, trained on a single GPU, the estimator is competitive with state-of-the-art tabular foundation models on density and calibration metrics.

[LG-202] PEACE: Covariant learning of nonadiabatic manifolds with parity-resolved Hamiltonians

链接: https://arxiv.org/abs/2610.09576
作者: Rongzhi Gao,Shuguang Chen,Yang Zhou,GuanHua Chen,Ziyang Hu,ChiYung Yam
类目: Chemical Physics (physics.chem-ph); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Nonadiabatic molecular dynamics provides mechanistic insight into light-driven processes and informs the design of molecules and materials for solar energy conversion, photocatalysis and photo switching. Accurately describing these processes requires a representation that respects electronic symmetry and consistently relates energies to interstate couplings. Here we introduce PEACE, which combines a parity-equivariant latent Hamiltonian with a learned electronic connection. Controlled ablations reveal the complementary roles of symmetry-allowed state mixing and electronic-frame variation in reproducing crossing structures and relaxation dynamics. PEACE closely reproduces excited-state population dynamics from first-principles simulations, while its extension to spin-orbit coupling enables simulations of intersystem crossing. These results demonstrate that a more complete incorporation of the underlying physics into learned electronic representations leads to more accurate predictions of nonadiabatic dynamics.

[LG-203] Unpaired Canonical Correlation Analysis NEURIPS2026

链接: https://arxiv.org/abs/2610.09530
作者: Nir Ben-Ari,Ronen Talmon,Uri Shaham
类目: Machine Learning (stat.ML); Machine Learning (cs.LG)
*备注: Accepted to NeurIPS 2026

点击查看摘要

Abstract:Canonical Correlation Analysis (CCA) is a fundamental method for multiview shared space learning. However, its strict reliance on paired data poses a significant limitation, as such data is often difficult to obtain or entirely unavailable. In this paper, we present Unpaired CCA (UCCA), a novel method that learns linear projections to maximize the correlation of the true underlying pairing without access to any paired samples during training. We first establish theoretical results connecting the Quadratic Assignment Problem (QAP) to CCA. Leveraging these theoretical insights, we derive a practical method to maximize correlation exclusively from unpaired data. To the best of our knowledge, UCCA is the first approach to learn maximally correlated projections in a strictly unpaired setting. We validate UCCA on real-world multi-modal datasets, demonstrating that it significantly outperforms recent unpaired alignment baselines in recovering the underlying true correlation. This work fills a critical gap between traditional statistical multiview learning and the growing field of unpaired data learning.

[LG-204] Reflected Anchored Langevin Algorithms

链接: https://arxiv.org/abs/2610.09522
作者: Changwei Tu,Xiaoyu Wang,Yingli Wang,Xicheng Zhang,Lingjiong Zhu
类目: Machine Learning (stat.ML); Machine Learning (cs.LG); Probability (math.PR)
*备注: 70 pages, 9 figures

点击查看摘要

Abstract:First order Langevin algorithms for constrained sampling in machine learning, such as projected Langevin Monte Carlo which are based on discretizations of reflected Langevin dynamics, require differentiable log densities that limits their applicability. This paper introduces reflected anchored Langevin dynamics (RALD), a reflected diffusion that converges to non-differentiable targets on constrained domains. The method uses a smooth anchored reference potential and multiplies the drift and noise covariance of its reflected Langevin dynamics by the same state dependent scaling factor. Its Euler-Maruyama discretization with projection gives reflected anchored Langevin Monte Carlo (RALMC) algorithm. We prove explicit convergence bounds and iteration complexity for RALMC in the 2-Wasserstein distance to the target distribution. Numerical experiments are provided to illustrate the theoretical predictions and the empirical performance of the method.

[LG-205] Adjoint-Based Calibration and Optimal Control of Stochastic Multiscale Bioprocess Digital Twins

链接: https://arxiv.org/abs/2610.09505
作者: Keilung Choy,Wei Xie
类目: Machine Learning (stat.ML); Machine Learning (cs.LG)
*备注: 37 pages, 8 figures

点击查看摘要

Abstract:We develop a bias-aware digital-twin calibration and control framework for multiscale bioprocess models within a biological systems-of-systems (Bio-SoS) paradigm. The digital twin is represented by a stochastic differential equation (SDE) model and calibrated from sparse, discrete observations using quasi-likelihood estimation and adjoint sensitivity analysis. SDE generator-based moment expansions characterize truncation-induced parameter bias, while forward-backward adjoints quantify how calibration uncertainty propagates to value functions and policy performance. The resulting parameter-error distribution supports both policy-directed adaptive experimental design and uncertainty-aware policy optimization through a second-order Gaussian-averaged objective. We characterize the asymptotic behavior of the resulting exploration criterion and derive a physical-system performance under the optimized policy. To implement these ideas, we develop an Actor-Simulator algorithm that jointly updates model parameters, selects informative experiments, and optimizes control policies. Numerical studies demonstrate improved calibration accuracy, sample efficiency, and control performance relative to state-of-the-art baselines.

[LG-206] VM-ARRAYDPS: Virtual Microphone Augmented Diffusion Posterior Sampling for Unsupervised Blind Speech Separation ICASSP2027

链接: https://arxiv.org/abs/2610.09334
作者: Jingqi Sun,Haozhan Tang,Shulin He,Zhong-Qiu Wang
类目: Audio and Speech Processing (eess.AS); Machine Learning (cs.LG); Multimedia (cs.MM); Sound (cs.SD); Signal Processing (eess.SP)
*备注: Submitted to ICASSP 2027 and currently under review

点击查看摘要

Abstract:Blind Source Separation(BSS) is a fundamental problem in signal processing, aiming to separate multiple source signals from their mixtures without prior knowledge of the sources or the mixing process. Traditional approaches, such as Independent Vector Analysis (IVA) exploits statistical independence of sources. Recently, diffusion-based approaches have emerged as a promising alternative by leveraging powerful generative priors. Among them, ArrayDPS formulates BSS problem as a posterior sampling problem, and utilizes a pretrained speech diffusion model to guide the recovery of clean source signals. A key factor behind its separation capability is the multi-channel consistency (MC) objective, which enforces the estimated source signals to reconstruct the observed microphone mixtures through the estimated acoustic transfer functions. However, the number of microphones in the array is often limited, which constrains the performance of ArrayDPS. To address this issue, we propose VM-ArrayDPS, a novel method that augments the microphone array with virtual microphones with higher-SNR, these microphones can offer extra MC constraints to enhance the separation performance. Experimental results demonstrate that VM-ArrayDPS significantly outperforms ArrayDPS on both 2-speaker and 3-speaker datasets, showcasing the effectiveness of virtual microphone augmentation in improving BSS performance. We also did ablation studies to show the influence of the number of virtual microphones and weight of the MC objective brought by virtual microphones.

[LG-207] Beyond Nominal Equilibria: Risk-Averse Multi-Population Mean-Field Games

链接: https://arxiv.org/abs/2610.09244
作者: Bhavini Jeloka,Siddhartha Ganguly,Panagiotis Tsiotras
类目: Optimization and Control (math.OC); Computer Science and Game Theory (cs.GT); Machine Learning (cs.LG); Systems and Control (eess.SY)
*备注: Submitted to a conference; comments are welcome

点击查看摘要

Abstract:Recent advances in mean-field games and its multi-population variants enable large-scale heterogeneous multi-agent systems to be modeled through representative agents and their associated mean-field distributions. However, existing approaches do not explicitly account for uncertainty in the behavior of other populations. To this end, we introduce a new paradigm: risk-averse multi-population mean-field games, where each population optimizes a worst-case expected reward over dynamically feasible ambiguity sets of mean-field flows of a subset of the other populations. Employing an occupation-measure formulation along with tools from set-valued analysis, we establish, under mild assumptions, several theoretical properties of the multi-population game, including the geometric properties of the ambiguity sets and the existence of a novel risk-averse multi-population mean-field equilibrium. Further, we derive contractivity results of the fixed-point operator under entropy regularization and show that it can be utilized to learn the equilibrium. Finally, we propose a risk-averse fictitious-play scheme and show that exploitability decays to zero, despite the additional nonlinearity introduced by the worst-case objective. We report several numerical experiments to illustrate convergence and risk-averse behavior.

[LG-208] Sampling SU(N) gauge theory on a 2D lattice from independent plaquettes via holonomies and corner reweighting

链接: https://arxiv.org/abs/2610.09147
作者: Javad Komijani
类目: High Energy Physics - Lattice (hep-lat); Machine Learning (cs.LG); High Energy Physics - Theory (hep-th)
*备注: 21 pages, 7 figures

点击查看摘要

Abstract:In the Wilson formulation of lattice gauge theory, the fundamental degrees of freedom are group-valued link variables, while the action is a sum over the trace of the plaquettes, the smallest Wilson loops. For generative sampling methods such as normalizing flows, this poses a challenge: the distribution of an individual plaquette is easy to model, but mapping sampled plaquettes to the links is the obstruction. We explore a statistical way around this in two dimensions for the \mathrmSU(N) gauge theory with the Wilson plaquette action. The action can be written in terms of holonomy variables, which determine the links once a consistency condition at the four corners of an extended lattice is satisfied. Using a boundary condition that only affects the holonomies, we trade this consistency condition for a conditional sampling problem, which reduces to sampling X,Y\in\mathrmSU(N) given the group commutator Z=XYX^\dagger Y^\dagger , where Z\in\mathrmSU(N) depends on the four corners. Plaquettes are sampled independently, and the resulting configurations carry weights due to the additional conditional sampling. We obtain the density of Z , which determines these weights, in closed form for \mathrmSU(2) and \mathrmSU(3) . A normalizing flow models the single-plaquette distribution for \mathrmSU(2) and \mathrmSU(3) ; for the group-commutator problem we use a closed-form sampler for \mathrmSU(2) and a trained normalizing flow for \mathrmSU(3) . In our tests on 2\times2 and 32\times32 lattices at two couplings each, per-plaquette acceptance rates exceed 99% and the effective sample size of the final weights is moderate, typically above one half.

[LG-209] he Impact of Likelihood Tempering on the Limiting Predictive Moments of Variational Bayesian Linear Neural Networks

链接: https://arxiv.org/abs/2610.09132
作者: Ian Zhang,Thibault Randrianarisoa
类目: Machine Learning (stat.ML); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:In wide Bayesian neural networks, Gaussian mean-field variational inference is prone to “prior dominance”: the Kullback-Leibler (KL) regularization term of the ELBO outweighs the expected log-likelihood, and the variational predictive distribution collapses to the prior predictive as the width M grows. Tempering the likelihood, by raising it to the power 1/T for a temperature T 1 , is equivalent to scaling the KL term by T . We ask in this paper how fast T must decrease with M to counteract this degeneracy and strike a good balance between the two terms. For single-hidden-layer linear networks with isotropic Gaussian priors, we derive the limiting predictive distribution under schedules of the form T = \tau/M^c , with constants \tau, c 0 , as M \to \infty and compare it with the untempered neural network Gaussian process (NNGP) posterior, the infinite-width limit of the exact posterior. Our main result is that the predictive expectation and variance undergo phase transitions at different scales: the limiting expectation leaves its prior value at c = 1/2 , once \tau falls below an explicit threshold, and equals the least-squares prediction for c 1/2 , whereas the limiting variance keeps its prior value for c 1 , matches the NNGP’s for c=1 , and vanishes for c 1 . With suitable choices of \tau,c , one can recover either the NNGP posterior expectation or its variance.

[LG-210] Covariate-dependent Joint Modeling of Multivariate Ordinal Preferences and Its Connections with Comparison Models

链接: https://arxiv.org/abs/2610.09070
作者: Yujie Chen,Antik Chakraborty,Anindya Bhadra
类目: Machine Learning (stat.ML); Machine Learning (cs.LG); Methodology (stat.ME)
*备注:

点击查看摘要

Abstract:Multivariate ordinal data along with covariates are commonly collected in problems ranging from alignment of language models with human preferences, as well as in recommender systems. For example, data sets such as MovieLens contain several movies rated on a scale 1–5 by human users, along with their demographic information such as age or gender. Similarly, data sets such as HelpSteer collect human feedback on several attributes such as “helpfulness” or “verbosity” of LLM response on an ordinal scale, with covariates depending on the LLM prompt–response pairs. Unfortunately, the standard approaches for modeling these data (a) look at the attributes individually rather than jointly, and (b) often convert the data into pairwise or list-wise win–loss comparisons for fitting models such as Bradley–Terry and Plackett–Luce. Both of these lead to a coarsening of what is actually observed, which we address via a joint covariate-dependent consecutive ratio Markov random field model. We also show pairwise or listwise comparison models are obtained under restrictions of our joint model, and that joint modeling improves comparisons. We also develop a maximum likelihood inference procedure even in the presence of an intractable normalizer.

[LG-211] Learning Transition Kernels of Jump-Diffusion Processes with Conditional Diffusion Models

链接: https://arxiv.org/abs/2610.09045
作者: Yuzhen Zhao,Yating Liu,Quentin Guibert
类目: Machine Learning (stat.ML); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:We study the problem of learning transition kernels for time-homogeneous jump-diffusion processes using conditional diffusion models, with the goal of generating new sample paths from training data consisting of N independent trajectories observed on a high-frequency discrete time grid. On the theoretical side, we establish non-asymptotic bounds for the conditional score estimation error and for the KL divergence between the laws of the true and generated discretely observed paths. On the numerical side, we first evaluate our method on synthetic data to assess the theoretical findings and benchmark its performance against the approach of Gao et al. (2025). We then apply our method to real-world data and investigate its performance on a probabilistic forecasting task.

[LG-212] Quadratic Weak-to-Strong Generalization in Random Feature Networks via Random Matrix Theory

链接: https://arxiv.org/abs/2610.09044
作者: Deborah Oliveira,Elliot Paquette
类目: Machine Learning (stat.ML); Machine Learning (cs.LG); Probability (math.PR)
*备注:

点击查看摘要

Abstract:Weak-to-strong generalization is the phenomenon where a strong student model trained with labels produced by a weak teacher model is able to generalize better than the teacher. In this paper, we study this phenomenon in two-layer random feature networks where the model strength is determined by its width. Using tools from random matrix theory, we derive deterministic equivalents for the population errors of an optimally trained teacher and a student trained with gradient flow. For ReLU activation and a pure spherical harmonic target, we obtain sharp asymptotics under a Gaussian universality assumption, showing a quadratic improvement: the student error scales as the square of the teacher error. These results attain the general lower bound of Medvedev at al (2025). We also analyze how the student behaves under more general stopping times and targets supported on multiple harmonic degrees, characterizing the regimes in which weak-to-strong generalization occurs and identifying the transition between quadratic, non-quadratic, and no improvement.

[LG-213] A method for multimodal analysis of TAIGA experiment data using essential features

链接: https://arxiv.org/abs/2610.08985
作者: Alexander Kryukov,Julia Dubenskaya,Elena Fedotova,Elizaveta Gres,Stanislav Polyakov,Eugene Postnikov,Alexander Razumov,Pavel Volchugov,Dmitry Zhurov
类目: Instrumentation and Methods for Astrophysics (astro-ph.IM); High Energy Astrophysical Phenomena (astro-ph.HE); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:The aim of processing and analyzing experimental data from physical experiments is to obtain physically significant information about the phenomenon under study. This goal is achieved by multi-stage processing of experimental data, during which noise associated with measurements is suppressed and the dimensionality of the input data is reduced. In this paper, we propose a new method based on the use of neural networks such as autoencoders to extract essential features. The special value of the proposed approach lies in the possibility of its application to the analysis of multimodal data received simultaneously from several installations. We will apply this approach to a multimodal data (MMD) of the experiment TAIGA. Currently, the analysis of the MMD is carried out independently for each installation separately. Therefore, the development of methods for the joint analysis of MMD from TAIGA-type installations is an urgent task in cosmic ray physics and gamma-ray astronomy. Based on Monte Carlo simulation, it is shown that the proposed method allows for effective MMD analysis. It can also be used for MMD analysis at other experimental complexes.

[LG-214] rust-Region Optimization for Smooth Potential-Interaction Energies in Wasserstein Space

链接: https://arxiv.org/abs/2610.08883
作者: You Wan,Ting Gao,Jinqiao Duan
类目: Machine Learning (stat.ML); Machine Learning (cs.LG); Optimization and Control (math.OC); Probability (math.PR)
*备注: 24 pages, 4 figures, 1 table; code and data in the ancillary files

点击查看摘要

Abstract:Finding low-energy configurations of interacting particles and approximating probability distributions lead to the minimization of potential-interaction energies in Wasserstein space. These energies can be nonconvex, making it important to exploit second-order information while controlling the reliability of local approximations. We study trust-region optimization of smooth potential-interaction energies on the Wasserstein space of probability measures with finite second moment. The method uses a quadratic model along pushforward curves, an L^2(\rho) step radius, and a Steihaug-Toint subsolver with an explicit self-adjoint second-variation operator. A ratio test determines acceptance and guides the radius update. Under a lower energy bound and globally bounded Hessians of the potential and interaction kernel, we prove that the objective is nonincreasing, the Wasserstein-gradient norms converge to zero, and an \varepsilon -stationary iterate is reached within O(\varepsilon^-2) total outer trials, including rejected trials. If the potential is quadratically coercive, every weak accumulation point is stationary. The analysis applies to arbitrary initial measures with finite second moment. For empirical measures, the iteration is a finite-dimensional trust-region method in the L^2(\rho_N) inner product, with complexity constants independent of particle number and dimension when the initial objective gaps are uniformly bounded. Numerical experiments include a smooth soft-particle energy, maximum-mean-discrepancy minimization for non-Gaussian targets, component ablations, and scaling studies in particle number and dimension.

[LG-215] Learned Monotone Recurrent Features in Governed Credit Scoring: The Price of the Frame and the Necessity of Macro Conditioning

链接: https://arxiv.org/abs/2610.08869
作者: Yew Lee Tan
类目: Risk Management (q-fin.RM); Machine Learning (cs.LG); Applications (stat.AP)
*备注: 59 pages + 6-page online supplement (ancillary files). Companion to arXiv:2610.05196

点击查看摘要

Abstract:Regulated credit scoring requires scores monotone non-decreasing in every exposure input. Deployed pipelines – hand-crafted monotone aggregates feeding sign-constrained gradient boosting – already meet this by composition; the open question is what learned temporal aggregation is worth inside one. We answer on five production-scale credit datasets at matched admissibility (one priced baseline convention excepted), with a monotone recurrent architecture whose per-input guarantee we extend, with proofs, to vector-valued inputs and to exogenously macro-conditioned decay gates, severities, thresholds, and peak memory. Two findings result. First, a strictness ladder: the value of learned monotone features rises with governance-frame strictness – zero on unconstrained engineered panels, maximal in summaries-only frames – replicated across two datasets and an official temporal-stability metric, though unconditioned features degrade on externally adjudicated later weeks. Second, a conditioning-delivery asymmetry under regime shift. On a train-on-boom, test-on-crisis mortgage design, two public macroeconomic series hurt as input columns, yet conditioning the recurrence on them delivers the paper’s only learned-block crisis-cohort uplifts. The confirmed effect: +0.006 to +0.013 AUC on an internally pre-registered Freddie Mac replication, at all five held-out seeds. The discovery estimate: +0.015 to +0.021 on Fannie Mae (three of five seeds post hoc), worth 10-27 basis points of defaulted balance at an 80% approval cutoff, and grows with early-prepaid loans excluded. A state-level test identifies the mechanism: between-cohort calibration transfer. A pandemic-band episode bounds scope: under forbearance-distorted labels the gain generalizes at a quarter to a third of crisis size on Fannie Mae, on Freddie Mac only against the capacity control.

[LG-216] Geometry-Aware Diffusion Approximate Posterior Sampling for Sparse-View and Limited-Angle CT

链接: https://arxiv.org/abs/2610.08866
作者: Honglei Brinkmann,Carola-Bibiane Schönlieb,Ander Biguri
类目: Image and Video Processing (eess.IV); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Sparse-view computed tomography (CT) reduces radiation dose and acquisition time and may mitigate motion artifacts. However, angular undersampling provides insufficient information to determine the image uniquely and stably. Limited-angle CT, arising from restricted angular coverage, produces strongly directional information loss associated with the missing angular range. In both settings, image directions may be strongly observed, weakly constrained, or unobservable, leading to severe ill-posedness and reconstruction ambiguity. Existing diffusion-based approaches incorporate measurement information through likelihood guidance, data-consistency operations, or range-null-space corrections. However, they do not generally use the continuously varying measurement sensitivity of the acquisition to jointly shape both reconstruction updates and stochastic exploration. We propose a geometry-aware diffusion-guided stochastic reconstruction framework for sparse-view and limited-angle CT. Its central component is a regularized noise-weighted pullback metric constructed from the CT forward operator and measurement-noise covariance. This metric continuously adapts both the measurement-aware update and stochastic exploration according to directional measurement sensitivity, suppressing changes along strongly constrained directions while permitting greater exploration along weakly constrained and unobservable directions. We complement this geometry-aware update with a regularized data-consistency correction and approximately null-space-restricted stochastic perturbations, implemented matrix-free using forward and backprojection operations together with conjugate-gradient solves. Experiments on sparse-view, noisy, and limited-angle CT demonstrate competitive reconstruction quality, strong measurement consistency, and spatially resolved empirical uncertainty estimates. Subjects: Image and Video Processing (eess.IV); Machine Learning (cs.LG) Cite as: arXiv:2610.08866 [eess.IV] (or arXiv:2610.08866v1 [eess.IV] for this version) https://doi.org/10.48550/arXiv.2610.08866 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[LG-217] Deep Learning-Based Tri-Hybrid Multi-User MIMO Precoding: The Blessing of EM-Reconfigurable Antennas

链接: https://arxiv.org/abs/2609.39167
作者: Kaijun Feng,Jiaxin He,Hongrui Yu,Zhen Gao,Anwen Liao,Ziwei Wan,Zhaocheng Wang
类目: ignal Processing (eess.SP); Information Theory (cs.IT); Machine Learning (cs.LG)
*备注: 14 pages, 13 figures, 4 tables

点击查看摘要

Abstract:Electromagnetic (EM)-reconfigurable antennas provide multiple candidate radiation patterns per element, thereby introducing an additional EM-domain degree of freedom. Integrating radiation-pattern reconfigurability, realized as EM-domain precoding, with conventional hybrid analog-digital precoding yields tri-hybrid multiple-input multiple-output (MIMO) precoding, which can substantially improve the spectral efficiency of wideband multi-user MIMO orthogonal frequency-division multiplexing (OFDM) systems. However, the joint design of EM, analog, and digital precoding remains challenging. To address this challenge, we propose a tri-hybrid precoding network (Tri-PNet) based on Conformer, an emerging neural architecture that combines the local modeling strength of convolutional neural networks with the global dependency modeling of Transformers. Furthermore, two representative radiation-pattern modes, i.e., the non-regular mode and the 3rd Generation Partnership Project (3GPP) Technical Report (TR) 38.901 mode, are investigated. Tri-PNet is trained in an unsupervised manner to jointly learn EM, analog, and digital precoding by maximizing the average sum spectral efficiency. Its radiation-pattern selection network (RPSNet) employs a Conformer encoder to capture both local and global frequency-domain correlations, whereas its hybrid analog-digital precoding network (HPNet) combines cross-attention and dual-path processing with singular-value-decomposition (SVD) and zero-forcing (ZF) priors. Simulation results under both radiation-pattern modes demonstrate that Tri-PNet outperforms random EM precoding and conventional hybrid MIMO without EM precoding, approaches the greedy EM precoding search scheme with substantially lower online complexity, and remains robust to imperfect channel state information (CSI).

附件下载

点击下载今日全部论文列表