本篇博文主要内容为 2026-10-07 从Arxiv.org论文网站获取的最新论文列表,自动更新,按照NLP、CV、ML、AI、IR、MA六个大方向区分。
说明:每日论文数据从Arxiv.org获取,每天早上12:30左右定时自动更新。
提示: 当天未及时更新,有可能是Arxiv当日未有新的论文发布,也有可能是脚本出错。尽可能会在当天修复。
目录
概览 (2026-10-07)
今日共更新1109篇论文,其中:
- 自然语言处理共149篇(Computation and Language (cs.CL))
- 人工智能共352篇(Artificial Intelligence (cs.AI))
- 计算机视觉共164篇(Computer Vision and Pattern Recognition (cs.CV))
- 机器学习共358篇(Machine Learning (cs.LG))
- 多智能体系统共25篇(Multiagent Systems (cs.MA))
- 信息检索共27篇(Information Retrieval (cs.IR))
- 人机交互共36篇(Human-Computer Interaction (cs.HC))
多智能体系统
[MA-0] A Case Study in Assuring AI-Written Software NEURIPS2026
【速读】:该论文旨在解决在由生成式 AI(Generative AI)驱动的软件开发过程中,缺乏正式软件工程训练的操作者如何实现有效的人类控制问题。随着代码生成能力的提升,系统产生的代码量远超人类可审查的范围,传统的全面代码审查已无法作为可靠的人类控制手段。其解决方案的关键在于构建一个以人类主导的元代理系统(meta-agent system),通过多层级代理协作——包括代码生成、监督与审查代理,并将项目规则固化为可传递的经验。研究发现,依赖测试、监控和审计代理进行监督存在固有缺陷,如代理测量的是替代指标而非真实结果、审计失败无声无息、检查项被遗漏、自动化修复引发运营中断等。因此,确保人类控制有效的核心机制是:将预期目标、评估证据、代理权限及最终人工决策均统一锚定于同一根本性目标(underlying objective),从而形成闭环的可信控制体系。
链接: https://arxiv.org/abs/2610.08651
作者: Lindsey Ferris,Sierra Bonilla
机构: Independent Researcher; University College London (伦敦大学学院)
类目: oftware Engineering (cs.SE); Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA)
备注: Accepted to the NeurIPS 2026 Meta-Agents Workshop
Abstract:Software-engineering agents can enable people without formal software training to build systems they could not otherwise implement and simultaneously can produce more code than even experts can meaningfully inspect. In both cases, exhaustive code review is not reliable as the sole basis for human control. We report a case study of a production healthcare platform built through coding agents and governed by an operator without formal software-engineering training. Over time, its workflow grew into a human-led meta-agent system where one agent wrote code, other agents supervised and reviewed it, and project rules carried lessons forward. The operator found that tests, monitors and reviewing agents used to supervise the system were fallible. Some monitors measured proxies rather than outcomes, some audits failed silently, missing checks disappeared from reported results and one automated repair caused operational disruption. In this case, human control depended on keeping the intended outcome, the evidence used to judge it, the agents’ permissions and the final human decision were all tied to the same underlying objective.
[MA-1] Recursive Game Creator: An Agent ic Product-Level Experience-Oriented Game Harness
【速读】:该论文旨在解决当前生成式游戏设计代理(game design agents)在程序正确性之外,难以确保游戏可玩性和玩家体验愉悦感的核心问题。现有方法虽能生成可运行的游戏原型,但缺乏对玩家主观体验的系统性优化,导致生成游戏在实际游玩中吸引力不足。其解决方案的关键在于提出一种以用户体验为导向的递归式游戏开发框架——递归游戏创作者(Recursive Game Creator),该框架通过“设计师(Designer)、构建者(Builder)、玩家(Player)与评审者(Reviewer)”四元组件构成闭环迭代机制:其中,代码原生的玩家(Player)利用程序化接口高效生成多样化的游戏行为轨迹,显著缓解传统基于图形界面评估带来的延迟与偏差;评审者(Reviewer)则结合基于轨迹的量化指标、视觉证据及显式文本反馈,多维度评估游戏在趣味性、任务完成度与运行稳定性等方面的表现;最终由评审者选择更优版本并提供改进意见,驱动下一轮递归优化。该方法在GameCraft-Bench上达到77.89的综合得分,在GameASG-Bench上实现53.2%的任务成功率(较基线提升34.1%),并取得93.4%的最高平均运行检查通过率,用户研究进一步验证了其在延长游戏时长和提升评分方面的优势。
链接: https://arxiv.org/abs/2610.08621
作者: Jiajun Chen,Haoyu Wu,Mingda Jia,Xihui Liu
机构: HKU MMLab(香港大学多媒体实验室); The University of Hong Kong(香港大学); Shenzhen Loop Area Institute(深圳环区研究院)
类目: Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA); Software Engineering (cs.SE)
备注:
Abstract:Recent game design agents have made substantial progress in generating playable games. However, program correctness does not ensure an enjoyable experience for players. We present Recursive Game Creator, an experience-oriented harness to advance agentic game development from rough game prototypes into entertaining games. Recursive Game Creator organizes recursive development around four components: Designer, Builder, Player, and Reviewer. The Designer translates user instructions and Reviewer’s feedback into detailed plans. The Builder turns these plans into candidate games. The coding-native Player creates and executes reusable policies through programmatic interfaces to efficiently collect diverse gameplay trajectories, mitigating evaluation bias caused by slow GUI-based collection. The Reviewer uses carefully designed trajectory-based metrics to induce player preferences, integrating with visual evidence and explicit textual preferences to evaluate games against game-specific criteria. Finally, the Reviewer accepts the better version and provides improvement reviews for the next round, closing the recursive loop. Our method achieves state-of-the-art overall performance of 77.89 on GameCraft-Bench. On GameASG-Bench, it achieves a strict task success rate of 53.2%, a 34.1% improvement over the same-model baseline, and the highest mean runtime-check pass rate at 93.4% among compared methods. A user study shows longer playtime and higher ratings. Code is coming soon.
[MA-2] Network Intervention by Polling Strategic Agents NEURIPS2026
【速读】:该论文旨在解决网络中策略性代理(strategic agents)环境下规划者面临的三大相互关联的挑战:最优决策依赖于代理的私有信息、被查询的代理可能通过虚假报告操纵结果,且精确计算无法实现可扩展性。针对具有异质私有技术的多活动网络博弈场景,规划者设定非歧视性价格,论文提出关键解决方案——基于中心性(centrality)的福利核分解:每个代理的贡献与其在由代理跨活动偏好重加权后的网络中的平方中心性成正比。这一分解启发了名为Poll的轮询算法,该算法在每轮中采样一个代理,短暂遍历其邻域,并根据局部报告更新价格。由此分解进一步导出三重效率:计算上,Poll相比精确计算及其他分布式方法显著减少运算量,在包含超过30万代理的真实网络中通信开销降低高达三个数量级;统计上,其查询复杂度随网络拓扑结构和偏好异质性变化,而非直接依赖于群体规模;经济上,该算法收敛至福利最大化价格,并支持行为特异性实现方式,能够激励真实报告并检测恶意偏离行为。
链接: https://arxiv.org/abs/2610.08347
作者: Chenyu Zhang,Rohit Parasnis,Saurabh Amin
机构: 未知
类目: Computer Science and Game Theory (cs.GT); Multiagent Systems (cs.MA); Social and Information Networks (cs.SI); Optimization and Control (math.OC); Machine Learning (stat.ML)
备注: Published as a conference paper at NeurIPS 2026
Abstract:A planner in a network of strategic agents faces three entangled challenges: the optimum depends on agents’ private information, queried agents may misreport to steer the outcome, and exact computation does not scale. We study these challenges in multi-activity network games with heterogeneous private technologies, in which the planner sets non-discriminatory prices. We show that the optimal prices admit a centrality-based decomposition of the welfare kernel: each agent’s contribution scales with its squared centrality in a network reweighted by agents’ preferences across activities. This decomposition motivates Poll, a polling algorithm in which the planner samples one agent per round, walks briefly through the agent’s neighborhood, and updates the price from a local report. From the same decomposition flow three forms of efficiency: computationally, Poll uses significantly fewer operations than exact computation and other distributed methods, requiring up to three orders of magnitude less communication on a real-world network with over 300,000 agents; statistically, its query complexity scales with topology and preference heterogeneity rather than explicitly with population size; and economically, it converges to welfare-maximizing prices while admitting behavior-specific implementations that induce truthful reports and detect adversarial deviations.
[MA-3] SC3BF: Shifted Collision Cone Control Barrier Function for Dynamic Obstacle Avoidance
【速读】:该论文旨在解决传统速度空间控制屏障函数(velocity-space control barrier functions, CBF)中碰撞锥(collision cone)过于保守的问题:现有方法会拒绝所有指向障碍物的相对速度,无论其大小,导致机器人在接近障碍物时过度保守,限制了运动效率。为此,本文提出了一种位移碰撞锥控制屏障函数(shifted collision-cone CBF, SC3BF),其核心创新在于引入一个与状态相关的允许范围(allowance),使机器人能够以随距离增大和自身速度增加而提升的速率接近障碍物,从而在保证安全的前提下提升运动灵活性。SC3BF通过标准二次规划(quadratic program)实现,其安全集在有界输入下保持前向不变性,无需设定最小前进速度或预留安全裕度。论文证明了存在非零允许范围可维持系统安全,并给出了闭式解。在包含最多100个动态障碍物的运动学自行车模型实验中,相较于三种速度空间基线方法,SC3BF显著提高了到达目标的成功率,且对原始指令的修正幅度不足基线方法的一半,验证了其高效性与实用性。
链接: https://arxiv.org/abs/2610.08332
作者: Amin Kashiri,Yasin Yazıcıoğlu
机构: Northeastern University (东北大学)
类目: Robotics (cs.RO); Multiagent Systems (cs.MA); Systems and Control (eess.SY)
备注: Submitted to the 2027 American Control Conference (ACC)
Abstract:The collision cone used by velocity-space control barrier functions is conservative: it rejects every relative velocity aimed into an obstacle, however slow. We propose the \emphshifted collision-cone CBF (SC3BF), which adds a state-dependent \emphallowance to the cone condition, so the robot may approach the obstacle at a rate that grows with distance and with its own speed. SC3BF is enforced by an ordinary quadratic program, and its safe set is forward invariant under bounded inputs without a minimum forward speed or a clearance margin. We prove that a nonzero allowance preserving safety always exists, and derive one in closed form. Against three velocity-space baselines on a kinematic bicycle among up to 100 moving obstacles, SC3BF reaches the goal more often and modifies the nominal input less than half as much.
[MA-4] Communication-Free Obstacle Localization from Aggregate Wrench Measurements in Leader–Follower Cooperative Transport
【速读】:该论文旨在解决多机器人协同搬运刚性负载时,无显式机器人间通信条件下障碍物的定位问题。其核心挑战在于:领导者机器人仅能测量所有跟随者对负载施加的合力与合力矩(即总作用力-力矩,aggregate wrench),但无法直接区分各跟随者的个体响应。解决方案的关键在于设计一种跟随者控制律,使领导者能够通过分析总作用力-力矩与负载运动速度之间的非线性关系,实现障碍物位置的精确重构。具体而言,每个跟随者在接近障碍物时会局部抵抗运动,导致负载的平移与旋转速度与总作用力-力矩之间呈现分段线性关系;不同线性区间的切换点对应障碍物的方位与距离信息,并可识别出响应的跟随者。研究给出了在固定负载构型下实现精确恢复的充分条件,并提出一种自适应探测方法,由领导者施加特定的平移与旋转输入以获取必要测量数据。仿真结果验证了该方法的有效性。
链接: https://arxiv.org/abs/2610.08324
作者: Amin Kashiri,Aditya Mohan,Yasin Yazıcıoğlu
机构: Northeastern University (东北大学)
类目: Robotics (cs.RO); Multiagent Systems (cs.MA); Systems and Control (eess.SY)
备注: Submitted to the 2027 American Control Conference (ACC)
Abstract:We consider obstacle localization for a team of robots cooperatively transporting a rigid payload without explicit inter-robot communication. A leader robot directs the payload’s motion, while follower robots assist and react to locally detected obstacles. The leader measures the followers’ aggregate wrench, i.e., the combined force and torque they exert on the payload, but cannot directly distinguish their individual reactions. We design a follower control law that allows the leader to recover obstacle locations from these measurements. Each follower resists motion toward nearby obstacles, resulting in a piecewise-linear relationship between the payload’s translational and angular velocity and the aggregate wrench. Changes between adjacent linear regions reveal an obstacle’s bearing and distance and identify the responding follower. We give sufficient conditions for exact recovery at a fixed payload configuration and develop an adaptive probing procedure in which the leader applies translational and rotational inputs to the payload to obtain the required measurements. We demonstrate the performance of the proposed method in simulations.
[MA-5] Visual Orchestration Tax in Agent ic VLM Pipelines: Auditing and Certifying Visual Evidence Reuse
【速读】:该论文旨在解决生成式视觉语言模型(VLM)代理流水线中因重复传递相同静态视觉证据而引发的“视觉编排税”问题,即在多代理协作过程中,语义不变的图像被反复重建为图像条件化的API请求,造成显著的计算冗余与资源浪费。其解决方案的关键在于构建一套“测量-认证”框架:在审计层面,定义M1(原始视觉证据触达次数)和M2(结构化触达冗余度)指标,结合查询级分布分析、自助法置信区间与成对质量检验,量化冗余程度;在认证层面,提出SharedVisCache机制,通过图像内容、预处理指纹及编码器假设三要素构成的契约感知缓存钩子,实现对可复用视觉证据的精准识别与安全重用。实验表明,在SeeingEye与MAMMQA基准上,视觉证据触达冗余率达66.8%–75.6%,且所有查询均超过预设阈值;经合同验证后,75.0%–75.5%的重复触达被确认可重用,同时保持输出一致性(350/350字符串匹配,ΔM5=0)。在物理层部署中,视觉调用频率从800降至200(ChartQA-200回放)和从200降至50(Live SeeingEye翻译阶段集成),并完整保留800/800与200/200的输出结果。研究确立了视觉证据重用作为代理编排中可度量、行为保真的核心属性,并建立了使后端前缀或令牌重用具有语义可解释性的代理层契约机制。
链接: https://arxiv.org/abs/2610.08170
作者: Lingteng Zeng
机构: The Chinese University of Hong Kong (香港中文大学); Hong Kong, China
类目: Multiagent Systems (cs.MA); Computer Vision and Pattern Recognition (cs.CV)
备注: 9 pages, 2 figures, 5 tables
Abstract:Agentic VLM pipelines increasingly pass the same static visual evidence through multiple specialist agents and tools. This design creates an orchestration-level redundancy mode: semantically unchanged images are repeatedly reconstructed as image-conditioned requests at the VLM API boundary. We call this phenomenon visual orchestration tax and develop a measurement-to-certification framework for visual evidence reuse in agentic VLM pipelines. The audit side defines \mathrmM1_\mathrmtrace to count raw visual-evidence touches and M2 to measure structural touch redundancy, with query-level distributions, bootstrap confidence intervals, and paired quality tests. Across SeeingEye and MAMMQA on chart, document, general-VQA, and multi-modal-QA tasks, audits reveal 66.8-75.6% visual-evidence touch redundancy, and every audited query exceeds the predefined gate. The certification side introduces SharedVisCache, a contract-aware evidence reuse hook keyed by image content, preprocessing fingerprint, and encoder assumptions. On SeeingEye, contract validation certifies 75.0-75.5% repeated touches as reusable while preserving 350/350 output strings and \Delta\mathrmM5=0 . At the physical layer, certified hits reduce F_\mathrmvision from 800 to 200 in ChartQA-200 trace replay and from 200 to 50 inside live SeeingEye translator-stage physical integration, preserving 800/800 replay strings and 200/200 integrated call outputs. The results position visual reuse as a measurable, behavior-preserving property of agent orchestration and define an agent-layer contract that makes backend prefix or token reuse semantically interpretable.
[MA-6] oken-Efficient Multi-Agent Collaboration via System One-Guided Computational Division of Labor
【速读】:该论文旨在解决基于大语言模型(LLM)的多智能体系统(Multi-Agent System, MAS)在复杂信息检索与推理任务中,因任务推理与协调操作(如任务选择、角色分配、消息路由和上下文管理)紧密耦合而导致的高令牌开销与延迟问题。这一耦合机制随着交互规模扩大,严重制约了智能体网络服务的可扩展性。其解决方案的关键在于提出一种基于“系统一”(System One)引导的计算分工机制——S1-MAS框架,通过将有限范围内的协调决策交由轻量级的System One模型处理,而保留开放式的复杂推理任务给高性能的LLM智能体执行,从而实现高效的计算分工。具体而言,一个轻量级控制器负责判断检查条件、选择后续任务及决定终止,同时一个紧凑的阅读器从授权来源中检索相关证据以支持决策;二者构成“决策-证据”闭环,使任务动态决定工作角色与数据访问权限,实现无需任务特化训练的自适应协作。实验表明,S1-MAS在七个不同基准测试中均显著提升准确性,并大幅降低推理成本:相较于AgentVerse、DyLAN和SelfOrg,其GPT-4o令牌消耗减少44.9%–97.2%,端到端延迟降低37.8%–93.0%,验证了该框架在构建可扩展、低成本智能体网络应用中的巨大潜力。
链接: https://arxiv.org/abs/2610.08155
作者: Zihan Zhou,Xinzhe Hu,Hanxu Yang,Liangjian Wen,Zhao Kang
机构: 未知
类目: Multiagent Systems (cs.MA); Artificial Intelligence (cs.AI)
备注:
Abstract:Large language model (LLM)-based multi-agent systems (MAS) have become a promising paradigm for complex information-seeking and reasoning tasks by enabling collaborative problem solving among specialized agents. However, existing MAS frameworks tightly couple task reasoning with coordination operations, including task selection, role assignment, message routing, and context management. As interactions grow, using powerful LLMs for these bounded control decisions introduces substantial token overhead and latency, limiting the scalability of agentic Web services. In this paper, we investigate whether coordination can be decoupled from expensive reasoning without compromising collaborative performance. We propose S1-MAS, a token-efficient multi-agent framework based on System One-guided computational division of labor. S1-MAS assigns bounded coordination decisions to lightweight System One models while reserving open-ended reasoning for capable LLM workers. Specifically, a lightweight controller selects inspection conditions, chooses subsequent tasks, and determines termination, while a compact reader retrieves condition-relevant evidence from authorized sources to support these decisions. Through a decision-evidence loop, selected tasks dynamically determine worker roles and source access, enabling adaptive collaboration without task-specific training. Extensive experiments on seven diverse benchmarks demonstrate that S1-MAS achieves superior accuracy while substantially reducing the inference cost. Across individual comparisons with AgentVerse, DyLAN, and SelfOrg on seven benchmarks, S1-MAS reduces GPT-4o token consumption by 44.9%-97.2% and measured end-to-end latency by 37.8%-93.0%. These results highlight its potential for scalable and cost-effective agentic Web applications.
[MA-7] Partially Observable Zero-shot coordination by Predicting Intention of Partner
【速读】:该论文旨在解决具身协作场景中零样本协作(zero-shot coordination)所面临的挑战,即在合作伙伴间歇性不可见的情况下进行有效决策,现有方法因伙伴表征模糊且对隐藏伙伴状态存在不确定性而表现受限。其核心解决方案是提出“预测伙伴意图”(Predicting Intention of Partner, PIP),通过联合视图变分自编码器(Joint-view VAE)将双智能体局部观测的联合信息在训练阶段提炼为仅依赖本地观测即可获取的伙伴表征,从而增强伙伴状态的可辨识性;同时引入伙伴状态信念网络(Partner-state Belief networks),基于自我代理的交互历史推断伙伴的隐藏位置与行为倾向。实验在Burrito-PO、Overcooked-PO及Melting Pot三个基准上评估,结果表明PIP在所有任务中均取得最高平均性能,人类评估与诊断分析进一步验证了其在未见过伙伴情况下的协作能力,以及两个组件在伙伴遮挡条件下的关键贡献。
链接: https://arxiv.org/abs/2610.08142
作者: Jinnyeong Yang,Yuhwan Jeong,Hoyong Kwon,Minseok Kim,Jihun Kim,Kuk-Jin Yoon
机构: KAIST(韩国科学技术院); Visual Intelligence Lab(视觉智能实验室)
类目: Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA)
备注: preprint
Abstract:Zero-shot coordination in embodied settings requires acting while the partner is intermittently out of view, leaving existing methods with ambiguous partner representations and uncertainty over hidden partner states. We propose Predicting Intention of Partner (PIP) to jointly address these challenges. PIP uses a Joint-view VAE to distill richer training-time evidence from the union of both agents’ local observations into a partner representation available from local observations alone. Partner-state Belief networks further infer the partner’s hidden location and behavioral tendencies from the ego agent’s interaction history. We evaluate PIP in Burrito-PO, Overcooked-PO, and a Melting Pot substrate, together with a human evaluation in Burrito-PO. PIP attains the highest mean performance among the compared methods across all three benchmarks. Human evaluation and diagnostic analyses further support coordination with unseen partners and the contributions of both components under partner occlusion.
[MA-8] Self-Referenced Social Preferences: Cooperation without Observing Others Rewards
【速读】:该论文旨在解决多智能体强化学习中合作行为难以在缺乏同伴奖励信号观测条件下的学习问题。现有方法通常依赖于智能体对其他智能体奖励的直接观测,但在许多现实场景中,智能体仅能通过观察他人行为与结果来推断其状态,而无法获取其私有奖励信号。为此,论文提出“自参照社会偏好”(self-referenced social preferences)机制:每个智能体首先学习自身奖励模型,并将其应用于其他智能体的可观测转移过程,从而从自身视角评估其结果,再将这些自参照评估结果融入标准社会偏好机制中。关键创新在于通过自我参照方式构建对他人结果的主观评价,无需真实奖励信息即可实现有效合作。研究设计了两种集成策略:一是修改学习奖励,二是以评估结果加权策略更新。实验在三个顺序型社会困境任务(逃出房间、清理环境、公共资源收割)上验证,结果表明,在不观测他人奖励的情况下,该方法仍能促使智能体自发形成合作行为,且在资源分配公平性方面优于可访问真实奖励的基线方法。不同社会偏好类型表现出不同的最优整合方式:对不公平厌恶偏好而言,结合价值前瞻的奖励修正效果更优;而纯粹利他偏好则更适合采用策略更新加权。尤其在部分可观测环境下,基于策略更新加权的方法仍能持续支持合作,证明了自参照社会偏好在无需共享奖励信号条件下实现高效合作的可行性与鲁棒性。
链接: https://arxiv.org/abs/2610.07881
作者: Mohamed Ayman Mohamed,Harshil Kotamreddy,Marcos Menon Jose
机构: Amazon(亚马逊); Nvidia(英伟达); Itaú Unibanco(巴西伊塔乌联合银行)
类目: Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA)
备注:
Abstract:Social preferences can promote cooperation in multi-agent reinforcement learning, but existing approaches often require agents to observe the rewards of their peers. In many real-world interactions, however, an agent can, as humans do, observe others’ behavior and outcomes without access to their private reward signals. We introduce self-referenced social preferences, in which each agent learns a model of its own reward, applies it to other agents’ observed transitions to assess their outcomes from its own perspective, and feeds these self-referenced assessments into standard social preferences. We study two ways to incorporate these assessments: modifying the learning reward, or using them to weight policy updates. We evaluate the approach on three sequential social dilemmas, Escape Room, Clean Up, and Commons Harvest, which require volunteering, public-good contribution, and resource restraint, respectively. Across all three environments, agents learn cooperative behavior without observing others’ rewards, including in settings where independent learners fail to cooperate, and frequently achieve more equitable divisions of jointly produced returns than agents with access to true rewards. The effective integration point depends on the social preference: inequity aversion works best in the reward together with a value look-ahead, whereas a purely benevolent preference benefits from policy-update weighting. Under partial observability, the policy-update approach continues to support cooperation. These results show that explicit access to other agents’ reward signals is not necessary for learning cooperative behavior: social preferences can instead be grounded in self-referenced assessments of others’ outcomes derived from their observed behavior.
[MA-9] Harness Engineering for Software Engineering via Modular Executable Dev-Primitives
【速读】:该论文旨在解决大语言模型(Large Language Models, LLMs)在长周期软件工程任务中因状态碎片化导致的脆弱性问题,具体表现为:在处理复杂项目时,模型需反复从源代码、配置文件、测试用例、依赖项及运行时行为等分散组件中重构程序状态,引发交互历史过长、上下文爆炸和语义漂移等问题,同时大型代码库进一步加剧了任务相关组件的识别难度。其解决方案的关键在于提出开发原语(Dev-Primitives),一种模块化且可执行的抽象机制,将静态的软件资产转化为具有主动参与能力的实体。每个开发原语将一个仓库组件与其本地部署的LLM绑定,赋予其基于自身实现与依赖关系的代理友好的接口,从而支持自然语言推理、组件间通信以及局部自修改能力。在此基础上,论文构建了HERMES(Harness Engineering framework for software engineeRing via Modular Executable Dev-PrimitiveS),通过依赖感知的动态激活机制与故障诊断机制,实现开发原语在仓库规模上的实例化,并能将执行证据回溯至需修正的组件。实验表明,HERMES在四个基准测试上平均优于基线12.4%;当与强激活和诊断模型结合时,即使采用Qwen3-8B作为开发原语,其性能也仅比同规模GPT-5.6配置低4.5%,同时在Terminal-Bench 4.0上降低26.2%的推理开销,凸显了高效架构设计在软件工程智能体中的核心作用。
链接: https://arxiv.org/abs/2610.07832
作者: Haibo Jin,Xinjie Li,Peng Kuang,Haohan Wang
机构: University of Illinois Urbana-Champaign (伊利诺伊大学厄本那-香槟分校); The Pennsylvania State University (宾夕法尼亚州立大学)
类目: oftware Engineering (cs.SE); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Multiagent Systems (cs.MA)
备注: 30 pages
Abstract:Large language models (LLMs) equipped with terminal access have demonstrated strong capabilities in automating software engineering tasks. However, existing agents remain brittle on long-horizon workflows, where they must repeatedly reconstruct program state scattered across source files, configurations, tests, dependencies, and runtime behavior, leading to increasingly long interaction histories, context explosion, and semantic drift. Large repositories further complicate the identification of task-relevant components. To address these challenges, we introduce \textbfDev-Primitives (\emphDevelopment Primitives), a modular and executable abstraction that transforms repository components from passive software artifacts into active participants in software engineering. Each Dev-Primitive pairs a repository artifact with a resident LLM, which gives the artifact an agent-native interface grounded in its own implementation and dependencies, enabling natural-language reasoning, inter-component communication, and localized self-modification. Building on Dev-Primitives, we propose \textbfHERMES, a Harness Engineering framework for software engineeRing via Modular Executable Dev-PrimitiveS, which instantiates these primitives at repository scale through a dependency-aware dynamic activation mechanism and a bug diagnosis mechanism that maps execution evidence back to the components that must be revised. Extensive experiments on four software engineering benchmarks demonstrate that HERMES outperforms matched baseline harnesses by 12.4% on average. Moreover, when paired with strong activation and diagnosis models, HERMES, even with Qwen3-8B Dev-Primitives, remains within 4.5% of the homogeneous GPT-5.6 Sol configuration across all four benchmarks, while reducing inference cost by 26.2% on Terminal-Bench 4.0, highlighting the importance of harness design in software engineering agents.
[MA-10] Do I Need the Cloud? Uncertainty-Aware Step-Level Handoff for Small Language Model Agents NEURIPS2026
【速读】:该论文旨在解决小型语言模型(Small Language Models, SLMs)作为本地代理控制器时,因结构化工具错误导致代理步骤失败的问题。现有路由机制通常在每个查询层面仅选择一次模型,无法适应代理决策过程中动态变化的难度。其解决方案的关键在于提出STEPGATE——一种基于不确定性的分步移交框架,通过评分每个本地SLM的动作置信度,智能识别并仅将高难度步骤移交至更强的模型进行处理。实验表明,在单步任务测试中,采用Qwen2.5-1.5B/7B组合的STEPGATE实现了82.7%的任务成功率与30.8%的移交率,显著优于纯本地部署(67.3%)和随机移交(75.4%),且在多轮评估中以仅30.0%的云端调用实现69.0%的轨迹成功率和84.0%的动作成功率,大幅缩小了与强模型全云端部署(82.0%轨迹成功率)的性能差距,同时减少了远程传输的令牌数量。
链接: https://arxiv.org/abs/2610.07816
作者: Abolfazl Younesi
机构: Sharif University of Technology (谢里夫理工大学); Tehran, Iran
类目: Artificial Intelligence (cs.AI); Distributed, Parallel, and Cluster Computing (cs.DC); Emerging Technologies (cs.ET); Machine Learning (cs.LG); Multiagent Systems (cs.MA)
备注: Accepted at the 40th Conference on Neural Information Processing Systems (NeurIPS 2026). Workshop: SLMs for Agentic Systems, Paris, France, 2026
Abstract:Small language models (SLMs) are attractive as local agent controllers because they reduce remote inference, latency, and deployment footprint, yet structured tool errors can cause an agent step to fail. Existing routers typically select a model once per query. However, agents expose sequential decision points whose difficulty dynamically changes based on intermediate observations. We propose STEPGATE, an uncertainty-aware handoff framework that scores each local SLM action and selectively escalates challenging steps to a stronger model. On a 52-task held-out single-step BFCL-derived test split, the Qwen2.5-1.5B/7B pair attains 82.7% task success with 30.8% escalation, versus 67.3% local-only and 75.4% random escalation (which uses 33.8% escalation). In a separate multi-turn evaluation, STEPGATE achieves 69.0% trajectory success and 84.0% action success using only 30.0% cloud actions, compared with 48.0%/70.5% local-only, 60.0%/78.2% random escalation, and 57.0%/77.1% query-level routing (strong-only achieves 82.0% trajectory success at 100% cloud actions). These results suggest that step-level escalation recovers a large share of the performance gap to the stronger Qwen2.5-7B backend at a matched cloud-action rate while transmitting fewer tokens remotely. However, our evaluation is limited to one model family, a single stronger backend, and scripted tasks. Furthermore, the test sets are small, multi-turn comparisons rely on paired intervals and statistical tests, and our risk tiers serve as research annotations rather than formal safety guarantees.
[MA-11] Independent Multi-Agent Reinforcement Learning with Counterfactual Semantic-Social World Models
【速读】:该论文旨在解决完全去中心化多智能体强化学习(Fully decentralized multi-agent reinforcement learning, MARL)中因信息结构受限导致的奖励信号模糊性问题。在无集中式评价网络或智能体间通信的环境下,单一标量奖励无法区分回报低下是由于自身行动无效、队友响应不兼容,还是对手响应有效所致,从而阻碍了智能体的有效学习。其解决方案的关键在于引入一种名为CASTLE(Counterfactual Action-conditioned Semantic Tokens for Local Execution in Decentralized MARL)的离线训练、在线上下文引导框架,该框架包含两个互补的世界模型:局部动态世界模型(Local Dynamics World Model)用于建模智能体本地轨迹的动力学特性与部分可观测性;语义-社会世界模型(Semantic-Social World Model)则通过反事实模拟推演,在给定当前状态的前提下预测各候选自行动作的短期任务与社交后果,以捕捉潜在的队友和对手响应。在在线学习与执行阶段,两个世界模型保持冻结,仅依赖本地信息被查询,其输出的预测逻辑值为独立的PPO策略提供上下文感知指导。实验表明,该方法在基准多粒子环境中的Tag、Spread和Adversary任务上均取得最优平均最终得分,显著优于现有最强基线,分别提升10.67、6.46和0.33个归一化分数点。
链接: https://arxiv.org/abs/2610.07704
作者: Fernando Martinez,Tao Li,Yingdong Lu,Juntao Chen
机构: Fordham University (福特汉姆大学); City University of Hong Kong (香港城市大学); IBM Research (IBM 研究院)
类目: Multiagent Systems (cs.MA); Machine Learning (cs.LG)
备注:
Abstract:Fully decentralized multi-agent reinforcement learning (MARL), also referred to as independent learning, requires each agent to learn and act using only its local information and experience, without a centralized critic or inter-agent communication. Such a stringent information structure renders the conventional reward signal ambiguous. A poor return may result from an ineffective ego action, an incompatible teammate response, or an effective opponent response, yet scalar rewards alone do not reveal which explanation is responsible. We argue that agents can learn more effectively by prospectively comparing the consequences of candidate actions rather than diagnosing failures only from realized returns. We introduce CASTLE (Counterfactual Action-conditioned Semantic Tokens for Local Execution in Decentralized MARL), an offline-training, online-in-context guidance framework with two complementary world models. A Local Dynamics World Model, offline pre-trained over agents’ local trajectories, summarizes the agent’s local trajectory dynamics and partial observability, while a Semantic-Social World Model predicts compact short-horizon task and social consequences for each candidate ego action. The latter is trained from counterfactual simulator rollouts that expose plausible teammate and opponent responses to alternative actions taken from the same logged rollout state. During online learning and execution, both world models remain frozen and are queried by agents using only locally available information. Their prediction logits provide in-context guidance to an independent PPO policy. Across 30 matched seeds on Tag, Spread, and Adversary in the benchmark multi-particle environments, our proposed CASTLE achieves the highest mean final score among the evaluated methods, exceeding the strongest baseline on each task by 10.67, 6.46, and 0.33 normalized points, respectively.
[MA-12] EIO-Agents : The Missing Semantic Layer for AI Agent Evaluation
【速读】:该论文旨在解决当前生成式AI(Generative AI)代理在日益关键的应用环境中进入生产阶段时,缺乏统一的语义标准来解释其评估结果的问题。现有评估体系中,评分、执行轨迹、评审输出及多评审员结论虽被广泛用于判断系统就绪与发布决策,但这些指标往往未明确说明支撑主张的证据、证据所能证明的内容,以及主张如何推导出最终决策,导致评估过程缺乏透明性与可验证性。为此,论文提出EIO-Agents——一个基于双层架构的开放评估规范:其一为评估智能本体(Evaluation Intelligence Ontology, EIO),提供语义层,通过类型化证据、版本化行为谓词、证据契约、主张、证人规则、证明状态、重复性及可计算的度量、发现、控制与PASS/REVIEW/BLOCK决策推导机制,构建可解释的评估逻辑链;其二为可移植评估记录(Portable Evaluation Record, PER),作为评估结果的权威记录系统,以内容寻址的标准化形式完整保留从证据到决策的全过程链条,支持重演、解释与验证。EIO填补了评估中缺失的语义契约,PER则确保评估结果作为可追溯、可审计的可信资产。随着AI代理承担更多操作责任,评估必须超越简单的评分与结论集合,演变为具备可独立核查其含义、证据、局限性与决策依据的问责型产物。
链接: https://arxiv.org/abs/2610.07675
作者: Fouad Bousetouane
机构: ProofAgent.ai; The University of Chicago (芝加哥大学)
类目: Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA)
备注: 32 pages, 11 figures
Abstract:AI agents are entering production in increasingly consequential environments without a shared semantic standard for what their evaluations actually mean. Scores, traces, judge outputs, and multi juror findings are increasingly used to justify readiness and release decisions, yet they often do not specify what evidence supports a claim, what that evidence can establish, or how the claim leads to a decision. We introduce EIO-Agents, an open specification for interoperable AI agent evaluation built on two layers. The Evaluation Intelligence Ontology (EIO) provides the semantic layer through typed evidence, versioned behavioral predicates, evidence contracts, claims, witness rules, proof status, recurrence, and computable derivations for metrics, findings, controls, and PASS, REVIEW, or BLOCK decisions. The Portable Evaluation Record (PER) provides the system of record: a canonical, content addressed representation of one evaluation that preserves the evidence to decision chain and can be re derived, explained, and verified. Scores summarize, juries interpret, and traces record, but none of them define what the evidence means or what it can prove. EIO provides that missing semantic contract, while PER preserves the resulting evaluation as a portable and verifiable system of record. As AI agents assume greater operational responsibility, evaluation must become more than a collection of scores and verdicts; it must become an accountable artifact whose meaning, evidence, limitations, and decisions can be independently checked.
[MA-13] Joint Workflow and Prompt Optimization for User Behavior Simulation
【速读】:该论文旨在解决用户行为模拟(User Behavior Simulation)中因依赖人工设计规则或领域知识而导致的泛化能力差、跨任务迁移困难的问题。现有方法通常需要大量领域专家干预和任务特定工程,难以在不同场景下保持高效与准确。其解决方案的关键在于提出SWORD框架——一种基于角色设计的、以标量任务指标为唯一优化目标的多智能体工作流与自然语言提示联合优化机制。SWORD无需领域初始化或任务定制化设置,仅通过梯度反馈驱动自动发现与优化工作流拓扑结构及提示内容,实现了对领域相关信号、评论情感映射规则以及流行病学衰减先验等隐含特征的无监督挖掘。实验表明,SWORD在控制变量条件下显著优于仅优化提示、仅优化工作流及分阶段优化的基线方法;相较于最强的已有领域专用基线,SWORD在更低的模型规模、更少训练数据和极低API成本(每数据集4–6美元)下仍实现更高精度,验证了文本梯度在用户行为建模中作为无监督特征重要性发现机制的有效性。
链接: https://arxiv.org/abs/2610.07663
作者: Nipun B Nair(1)Tongtong Wu(1),Hongzhi Yin(2),Hui Li(3),Weiqing Wang(1) ((1) Monash University, (2) The University of Queensland, (3) Xiamen University)
机构: Monash University (莫纳什大学); The University of Queensland (昆士兰大学); Xiamen University (厦门大学)
类目: Multiagent Systems (cs.MA); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: under review for ACM Transactions on Information Systems Journal, 34 pages, 2 figures
Abstract:User behavior simulation is the computational modeling of user interactions within information systems through the use of simulated agents in place of live users. It supports system testing and evaluation, decision-making and forecasting, and user experience design. Existing simulators rely on hand-crafted rules or domain expertise that transfers poorly across tasks. SWORD (Simulation-driven Workflow and Prompt Optimization with Role-based Design) is introduced as a framework that jointly optimizes multi-agent workflow topology and natural-language prompts. It is guided solely by a scalar task metric, without domain initialization or task-specific engineering. The experimental results demonstrate that SWORD achieves statistically significant gains over prompt-only, workflow-only, and staged-optimization baselines under a controlled, identical-backbone comparison. Against the strongest published domain-specific baseline, SWORD further improves accuracy while using a smaller backbone model, substantially less training data, and a very reasonable API cost (\ 4–\ 6 for each dataset). Beyond predictive performance, SWORD autonomously discovers domain-relevant signals, review-sentiment mapping rules and epidemiological decay priors, purely from scalar error feedback, establishing textual gradients as a mechanism for unsupervised feature-importance discovery in user behavior modeling.
[MA-14] Where Rules End and Judges Begin: Measuring the Judgment Boundary in Multi-Agent Systems Security
【速读】:该论文旨在解决基于大语言模型(LLM)的多智能体系统(MAS)在面对恶意内容时缺乏有效、可审计且全面防御机制的问题。现有防御方法通常孤立评估单一攻击类型,难以应对复杂多变的对抗性威胁,导致防御成本高且结果不可追溯。为此,研究提出一种基于五项核心防御原则的统一框架——DEFER1(DEterministic-First Enforcement with Residual judgment),其关键在于构建一个分层检测与决策机制:通过28个确定性检查(deterministic checks)优先拦截明显违规行为,仅将剩余不确定案例交由四名人工裁判组成的评审小组进行意图判断。实验结果显示,在四个不同领域中,攻击成功率从约30.0%显著降至3.0%,其中78%的攻击被确定性检查阻断,而仅约25%的提案进入裁判环节,表明该机制能高效处理明确违反政策的攻击,同时借助人类判断处理模糊意图场景。然而,系统仍存在局限性,如基于风险评分的审批门限虽能准确过滤多数合法请求,却对攻击提案误判率较高,凸显了在威胁评估中实现高精度与高召回之间的根本挑战。
链接: https://arxiv.org/abs/2610.07657
作者: Shaswata Mitra,Raj Patel,Subash Neupane,Sudip Mittal,Md Rayhanur Rahman,Shahram Rahimi
机构: The University of Alabama(阿拉巴马大学)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Cryptography and Security (cs.CR); Multiagent Systems (cs.MA)
备注: 26 pages, 20 figures, 24 tables
Abstract:LLM-based multi-agent systems (MAS) engage tools, share memory, and delegate tasks, often encountering adversarial content. Current defenses for MAS are typically evaluated in isolation, focusing on one attack type at a time, which can lead to costly and hard-to-audit outcomes. This study organizes defenses into five principles, implementing them as DEFER1 (DEterministic-First Enforcement with Residual judgment), which includes a cascade of 28 checks that blocks what it can and refers the rest to a panel of four judges. In independent testing across four domains, attack success rates drop from about 30.0% to approximately 3.0%, with 78% of blocked attacks handled by deterministic checks. Only a quarter of proposals reach the judges in the security-operations domain, illustrating that the rules provide security for attacks violating clear policies, while judges manage those that only misrepresent intent. Both systems have weaknesses, such as a risk-score approval gate that inaccurately approves most attack proposals but few legitimate ones, highlighting the challenges in assessing threats accurately.
[MA-15] Disentangling Models from Personas in Heterogeneous LLM Simulations
【速读】:该论文旨在解决当前基于大语言模型(Large Language Models, LLMs)的多智能体模拟中普遍采用单一基础模型所导致的局限性问题,即忽视了不同模型间相互作用对智能体交互动态的显著影响。其核心问题是:在真实部署场景下,智能体间的互动不仅受个体角色设定(persona)影响,更可能由其所依赖的基础模型特性主导。论文的关键解决方案在于构建一个由多个不同基础模型组成的异质社交网络,并通过实证模拟揭示,智能体所获得的互动量更多取决于其所属的基础模型而非角色设定;当模型数量增加时,各模型固有的吸引或排斥效应显著增强,表明大规模网络中的动态行为可能趋于由基础模型特性主导。为解释这一现象,研究进一步开展内容中介分析,验证了基础模型在不同语境下的可预测性,并揭示了模型词汇模式与高互动风格之间的关联。该研究强调,在大规模多智能体交互系统中,异质模型组合对网络结果具有决定性影响,应成为未来建模设计的重要考量因素。
链接: https://arxiv.org/abs/2610.07535
作者: Dani Roytburg,Daphne Ippolito
机构: Carnegie Mellon University (卡内基梅隆大学)
类目: Multiagent Systems (cs.MA); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Social and Information Networks (cs.SI)
备注: Presented as a Spotlight Paper at the Second Workshop on Social Simulation with LLMS, Third Conference on Language Modeling, 2026
Abstract:Multi-agent simulations with large language models (LLMs) often operate networks of agents with a single base model. This overlooks the inter-model effects which may dominate engagement dynamics in real-world deployments. To show this, we simulate a heterogeneous social network powered by several different base models and show that the amount of engagement an agent receives depends more on its base model than on its assigned persona. The attraction or repulsion effects of a base model strengthen dramatically when more models are added in the mix, suggesting that networks dynamics may converge to base model effects at scale. To help explain this effect, we conduct a series of content-mediating analyses, showing the predictability of base models across contexts as well as the relationship between a model’s lexical patterns and an engagement-maximizing style. In light of recent developments in mass multi-agent interaction, this work underscores the relevance of heterogeneous compositions in driving the outcomes of those networks
[MA-16] Who Bears the Burden? Learning Responsibility for Shared Constraints in Multi-Agent Reinforcement Learning
【速读】:该论文旨在解决多智能体系统在共享成本预算约束下,如何合理分配拉格朗日乘子(Lagrangian multiplier)惩罚权重的问题。传统方法中,统一的乘子虽能强制执行总预算约束,但无法确定各智能体应承担的具体责任;而采用个体特异性乘子则可能仍依赖于相同的全局成本信号,难以反映各智能体间收益牺牲的异质性。为此,论文提出拉格朗日责任分配(Lagrangian Responsibility Allocation, LiRA),其核心在于通过在有限训练周期内优化社会福利(social welfare),学习每个智能体对公共乘子的责任份额。该机制在保持原始奖励与约束不变的前提下,通过责任份额重新分配乘子的影响,实现更合理的惩罚分摊。在标准正则条件下,对于凸博弈场景,调整责任份额可生成一系列归一化的广义纳什均衡(normalized generalized Nash equilibria),其中活跃约束始终维持在预算水平,而社会福利随责任配置平滑变化。为在收敛前优化责任分配,论文推导出一种考虑学习更新及由此引发数据分布变化的社会福利梯度。实验在CityLearn、MABIM、Harvest和MetaDrive等多个平台(涵盖3至400个智能体)上验证了LiRA的有效性,结果显示其相比均匀分配与个体乘子基线,平均社会福利提升最高达29%,同时确保电网与驾驶成本在预算范围内,库存违规减少,Harvest任务对可用预算的利用效率显著提高。
链接: https://arxiv.org/abs/2610.07491
作者: Xiaoyang Cao,Jingqi Li,Zhe Fu,Alexandre M. Bayen
机构: Massachusetts Institute of Technology (麻省理工学院); The University of Texas at Austin (德克萨斯大学奥斯汀分校); Stanford University (斯坦福大学); University of California, Berkeley (加州大学伯克利分校)
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computer Science and Game Theory (cs.GT); Multiagent Systems (cs.MA)
备注: 20 pages, 2 figures, 4 tables. Project page with code: this https URL
Abstract:When multiple agents share a cost budget, a common Lagrange multiplier can enforce the aggregate constraint but does not determine how its penalty should be allocated across agents. Uniform penalties ignore heterogeneity in the rewards agents sacrifice, while agent-specific multipliers may still rely on the same aggregate cost signal. We introduce Lagrangian Responsibility Allocation (LiRA), which learns each agent’s share of a common multiplier by optimizing social welfare over a finite training horizon. The multiplier enforces the aggregate budget, while responsibility shares redistribute its influence without modifying the original rewards or constraints. For convex games under standard regularity conditions, varying these shares induces a smooth family of normalized generalized Nash equilibria in which active constraints remain at their budgets while welfare varies. To optimize responsibility before convergence, we derive a welfare gradient that accounts for both learning updates and the induced change in data distribution. Across CityLearn, MABIM, Harvest, and MetaDrive, spanning 3 to 400 agents, LiRA improves average social welfare by up to 29% over uniform and agent-specific multiplier baselines. Grid and driving costs remain within budget, inventory violations decrease, and Harvest makes more effective use of available budget.
[MA-17] Auditable Claims about AI Agents
【速读】:该论文旨在解决组织在宣称其人工智能代理(AI agent)具备特定可信属性时缺乏可验证依据的问题,尤其是在欧盟《人工智能法案》第12条要求高风险系统必须具备事件自动记录能力,但未明确何种记录可作为某项声明的判定依据的背景下。其核心解决方案在于提出“可审计性”(auditable)的严格定义:一个声明要被验证,必须预先明确其政策(policy)、作用范围(scope)、可作为证据的记录类型及其生成主体,并设定决策规则。该方法扩展了“可审计代理”框架中的“策略可检验性”(Policy Checkability)维度,从单一行为扩展至整体声明层面。关键创新在于引入三个新增条件:独立记录对操作的覆盖性、每项操作的授权需绑定其输入参数、以及超出完整性之外的完备性。基于显式模型的证明表明,在假设成立的前提下,若缺少任一条件,则声明的支持将不可能实现。研究通过构建“声明核查表”(claim-check table),将该方法应用于六类常见声明,并以实际案例展示单个声明在五个证据状态下的演进过程,最终为操作者、采购方、审计机构及标准制定者提供可操作的实践指南。
链接: https://arxiv.org/abs/2610.07459
作者: Yue Zhao,Jiate Li,Li Li,Yi Nian,Jinbo Liu,Xiaolin Zhou,Xiyang Hu
机构: 未知
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computers and Society (cs.CY); Multiagent Systems (cs.MA)
备注:
Abstract:Organizations make claims about their AI agents: a person approves every external email, every action is logged, an evaluation shows the agent is safe to deploy. Article 12 of the EU AI Act requires high-risk systems to allow the automatic recording of events but does not say which records settle a given claim. The position is one sentence: to be checked, a claim about an agent must first name its policy, its scope, the records that would settle it, and who writes them. Adapting the preconditions of an assurance engagement, we call a claim auditable when these elements and a decision rule are fixed before any verdict and the records are obtainable. This extends the Policy Checkability dimension of our Auditable Agents framework from single actions to claims. Agents add three conditions: coverage by an independent record, authorization bound to each action’s arguments, and completeness beyond integrity. Under an explicit model, we prove that support is impossible without each wherever its hypotheses hold. A claim-check table applies the method to six common claims, anchored in current NIST, IETF, and OWASP drafts. A worked case follows one claim through five evidence states. We close with a practice box and steps for operators, buyers, auditors, and standard setters.
[MA-18] he Right Memory in the Wrong Context: Verifying Retrieval Admissibility in Long-Term Agent Memory NEURIPS2026
【速读】:该论文旨在解决长期记忆代理(long-term-memory agents)在信息检索过程中可能暴露不合规或不可接受信息的问题,尤其是在跨主体(principal)、政策违反或生命周期状态不兼容的情况下,系统仍可能生成看似合理但实际存在安全风险的答案。现有评估方法(如召回率与最终答案准确率)无法有效揭示此类问题:一条路径可能因遗漏必要证据而显得“安全”,而正确答案也可能因暴露了不合规的提示内容而引发风险。其解决方案的关键在于提出一种检索-可接受性验证框架(retrieval-admissibility verification framework),该框架为每个记忆-查询对分配三种状态(可接受、不可接受、未决),在匹配所需证据召回率的基础上对未决情况设定边界,并追踪记忆项在提示暴露过程中的传播路径,同时将暴露行为与目标层面的信息披露关联起来。通过分阶段独立评估,研究发现该框架显著提升了关键指标——在两个公开基准(RHELM和MemOps)的后验分析中,前20名锚点召回率从0.432提升至0.533,80%召回可行性由0.237增至0.311,精确相似性评估下降98.3%;在开发诊断中,仅使用元数据参考可保持必要证据完整,而纯文本验证器在1%假阴性容忍度下均未能检测违规。进一步观察性分析显示,命名空间路由与评分准确性提升0.053–0.068相关,且在受控暴露实验中,仅一个读者的置信区间排除零值,表明相关不可接受信息的直接披露存在统计显著影响。结果强调必须对候选支持、可接受性、提示暴露及答案披露进行独立验证。
链接: https://arxiv.org/abs/2610.07309
作者: Zi Wang,Xingqiao Wang,Emmanuel Addai,Devika Ambekar,Xiaowei Xu
机构: University of Arkansas at Little Rock(阿肯色大学小岩城分校)
类目: Artificial Intelligence (cs.AI); Information Retrieval (cs.IR); Multiagent Systems (cs.MA)
备注: 26 pages. Accepted at the NeurIPS 2026 Workshop “Who Verifies the Agents? Toward Reliable Agent Development”. Code: this https URL
Abstract:Long-term-memory agents can retrieve relevant information that is inadmissible for the current request because it belongs to another principal, violates policy, or reflects an incompatible lifecycle state. Recall and final-answer accuracy do not reveal this: a route can appear safe by missing required evidence, while a correct answer may follow inadmissible prompt exposure. We introduce a retrieval-admissibility verification framework that assigns each memory-query pair one of three statuses (admissible, inadmissible, or unresolved), compares routes at matched required-evidence recall with bounds for unresolved cases, and tracks memory IDs through prompt exposure while linking exposure to target-level disclosure. We evaluate its stages on separate, non-pooled populations. A post-hoc top-20 reanalysis of frozen rankings from two public long-term-memory benchmarks, RHELM and MemOps, covers 3,767 queries. All released anchors lie within trusted query namespaces; with within-namespace scores unchanged, off-namespace filtering cannot lower their ranks. Top-20 anchor recall increases from 0.432 to 0.533, 80% recall feasibility from 0.237 to 0.311, and exact similarity evaluations decrease by 98.3%. In a frozen 72-case development diagnostic, a released-metadata reference preserves required evidence, whereas neither text-only verifier detects violations under the 1% required-anchor false-denial limit. Across 1,523 paired benchmark-native cases, namespace routing is associated with judged-accuracy gains of 0.053-0.068 across three readers; recall also changes, so this comparison is observational. In 16 controlled exposure scenarios, only one of four reader-specific 95% confidence intervals excludes zero for relevant-inadmissible literal disclosure (+0.156, 95% CI [0.031, 0.312]). Results motivate separate verification of candidate support, admissibility, prompt exposure, and answer disclosure.
[MA-19] SPEAR: Five Principles for Interactive Human-Agent Alignment
【速读】:该论文试图解决当前人工智能对齐(AI alignment)研究中将对齐问题简化为部署前优化的局限性,即在系统上线前通过人类反馈学习偏好并微调模型的范式,难以应对智能体在具体、长期且社会化的实际场景中作为用户代理持续交互时所面临的动态复杂性。其核心解决方案的关键在于将人-智能体对齐重构为一个持续的交互设计问题,并提出SPEAR框架——由五大支柱构成:意图表达与共理解建立(Specification)、行动决策机制(Process)、成效评估机制(Evaluation)、用户适应性演化(Adaptation)以及信任与行为再校准(Recalibration),从而系统性地支持智能体在真实使用情境中与用户之间持续、动态的对齐过程。
链接: https://arxiv.org/abs/2610.07204
作者: Tao Long,Lydia B. Chilton
机构: Columbia University (哥伦比亚大学)
类目: Human-Computer Interaction (cs.HC); Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA)
备注: 3 pages. Best Talk Award at the ACM Conference on Human-AI Complementarity and Alignment (HCOMP 2026)
Abstract:Recent AI alignment work often frames alignment as a pre-deployment optimization problem: collect human feedback, learn preferences or principles, finetune the model, and deploy an aligned system. This framing has produced major progress, but it under-specifies what happens once AI systems act as agents on users’ behalf in situated, long-term, and social contexts. This position paper reframes human-agent alignment as an ongoing interaction design problem. We propose SPEAR, five pillars of interactive alignment: Specification (how people express intent and establish shared understanding), Process (how agents decide when to act, ask, defer, or pause), Evaluation (how people judge whether agents succeeded), Adaptation (how agents adapt to users over repeated use), and Recalibration (how people adapt their trust, expectations, and behavior in response to agents).
[MA-20] When to Remember When to Abstain: Category-Conditioned Retention for Reliable Agent Memory NEURIPS2026 NEURIPS
【速读】:该论文旨在解决生成式AI系统中持久化代理记忆(persistent agent memory)的可靠性问题,即如何在记忆存储阶段有效区分可信与不可信的信息。现有方法通常依赖单一全局置信度阈值来决定是否保留信息,但这种方法无法适应不同语义类别间显著的证据支持差异:例如,价值和信念类断言仅有77.9%被其来源支持,而其他类别则高达96.2%。这种可靠性不对称性导致全局阈值难以兼顾保真度与覆盖率——要么误保留缺乏支持的价值主张,要么错误舍弃高置信度的其他类别信息。为此,论文提出一种基于语义类别的置信度条件化机制,即根据不同断言类型动态调整保留阈值:对证据充分的类别放宽标准,对推理不确定性高的类别施加更严格的筛选。在100个合成人格的部署型冷启动记忆流水线上的实证评估表明,仅对价值类断言采用更严格的标准,可使未支持保留率从6.2%降至4.0%(相对减少约36%),同时在相似保留率下保持额外13个百分点的覆盖率(95%置信区间:9.8–16.0)。结果表明,可靠的内存保留不应仅依赖置信度数值,而应结合断言类型进行分类处理;通过在写入边界引入类别条件化的置信阈值,可实现一种简单而有效的选择性预测(selective prediction)策略。
链接: https://arxiv.org/abs/2610.07100
作者: Olukunle Owolabi,Pulkit Gupta,Fei Wang
机构: Meta AI(元宇宙人工智能实验室)
类目: Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA)
备注: 4 pages, 1 Figure, Accepted to NeurIPS 2026 Social Agent Workshop ( this https URL )
Abstract:Persistent agent memory is only as reliable as its retention decision: an assertion weakly supported by its source can be stored and later reused as established fact. We study whether the retention decision should be governed by a confidence bar conditioned on the semantic category of the assertion rather than by a single global threshold, retaining well-evidenced categories liberally while abstaining more aggressively where inference is unreliable. We evaluate this in a deployed cold-start memory pipeline on 100 synthetic personas. The empirical evaluation is motivated by a sharp reliability asymmetry: across 4,715 candidate assertions, only 77.9% of value and belief assertions are supported by their source, versus 96.2% for all other categories. A global confidence threshold cannot separate these: it either admits unsupported value claims or discards well-evidenced ones. Conditioning the threshold on category resolves the tradeoff. In repeated held-out evaluation, a stricter bar on values alone reduces unsupported retentions from 6.2% to 4.0% (an \approx36% relative reduction, modest but consistent across folds) and, as corroborating evidence, preserves an estimated 13 percentage points more coverage (95% CI 9.8–16.0) than a global threshold at comparable retention. Our results suggest that reliable retention depends on the type of assertion, not on confidence alone, and that a category-conditioned threshold can act as a simple, effective form of selective prediction at the write boundary.
[MA-21] rust-Gated Capability Control: Breaking the Trust-Vulnerability Paradox in Multi-Agent LLM Systems
【速读】:该论文旨在解决多智能体大语言模型(multi-agent LLM systems)中分层信任模型(layered trust models)长期停留在概念层面的问题,具体表现为:现有模型未能明确各信任层级间的组合机制、运行时重要性权重的动态设定方法,以及信任如何有效指导智能体行为。这一问题的关键在于,高信任虽可提升任务成功率,却也加剧了被滥用的风险,形成“信任-脆弱性悖论”(Trust-Vulnerability Paradox)。为此,论文提出一个五层可操作的信任架构,其核心解决方案包含三方面创新:首先,引入跨层协同算子(cross-layer synergy operator),将前置层级的缺陷传递至依赖层级,并通过广义均值复合信任函数实现弱链接规则的极限情形恢复,且具有可证明的边界约束;其次,每层的重要性权重基于观测到的失败事件,采用无后悔在线估计器(no-regret online estimator)动态追踪当前最可能导致危害的层级;最后,提出“信任门控能力控制”(Trust-Gated Capability Control)机制,在复合信任及相应前置层级满足特定阈值时,才授予短期、可撤销的能力权限。理论证明该机制可打破悖论——即使存在隐蔽攻击者通过虚增行为信任而削弱前置层级,也无法实现权限升级;同时,给出了闭式表达的、严格保守的信任固定点,并具备有界延迟的撤销保证。数值实验验证了预测:在20,000次随机抽样下复合信任边界保持稳定,解析固定点与仿真结果偏差小于0.007,且受控前置层级一旦被攻破即在数次交互内触发自动撤销。
链接: https://arxiv.org/abs/2610.07000
作者: Mehedi Hasan Nipu,Chinmoy Mitra,Tarannum Ahmed Nowshin,Israt Moyeen Noumi
机构: North South University (北南大学); Rajshahi University of Engineering & Technology (拉杰沙希工程与技术大学); BRAC University (BRAC大学); Ahsanullah University of Science and Technology (阿桑努拉科学与技术大学)
类目: Computer Science and Game Theory (cs.GT); Multiagent Systems (cs.MA)
备注:
Abstract:Layered trust models for multi-agent LLM systems remain largely conceptual: they name which dimensions of trust matter but not how layers combine, how their importance is set at runtime, or how trust should govern agent actions. This gap matters because higher inter-agent trust raises task success while also enlarging exposure to exploitation, a tension formalized as the Trust-Vulnerability Paradox. We make a five-layer trust stack operational through three contributions. First, a cross-layer synergy operator propagates prerequisite-layer deficits into de- pendent layers, feeding a generalized-mean composite trust that recovers the weakest-link rule as a limiting case, with provable bounds. Second, per-layer importance weights are grounded in observed failures via a no-regret online estimator that tracks which layer is currently most responsible for harm. Third, Trust- Gated Capability Control issues short-lived, revocable capability grants only when composite trust and the relevant prerequisite layers clear capability-specific thresholds. We prove this mechanism breaks the paradox: a stealthy compromise that inflates behavioral trust while degrading a prerequisite layer cannot escalate privilege, and give a closed-form, provably conservative trust fixed point with a bounded-latency revocation guarantee. A numerical study with a sleeper adversary confirms the predictions: composite-trust bounds hold across 20,000 random draws, the analytic fixed point matches simulation within 0.007, and a compromised prerequisite layer triggers automatic revocation within a few interactions.
[MA-22] When Robots Crash: Optimal Asynchronous Gathering at Weber Meeting Nodes
【速读】:该论文旨在解决在存在崩溃故障(crash faults)的异步、匿名且无记忆(oblivious)移动机器人系统中,于无限网格上实现最优聚集(optimal gathering)的问题。具体而言,需将所有非故障机器人聚集到一个指定的“韦伯会合节点”(Weber Meeting Node),以最小化其初始位置到目标节点的总曼哈顿距离(Manhattan distance)。由于机器人不具备坐标系和手性(chirality),且无法区分故障与延迟,传统方法依赖特定机器人打破对称性,但该机器人一旦故障将导致系统停滞。本文的关键解决方案在于:通过强多重性检测(strong multiplicity detection)机制,使每个机器人能够基于自身视角独立选择同一目标节点;同时利用依赖目标的受限最短路径(target-dependent restricted shortest paths)确保所选目标始终为最优韦伯会合节点。研究证明,在某些完全对称配置下,最优聚集不可实现;而对于其余可解配置,提出的算法 \textscCrashTolerantWeberGathering() 能够确定唯一共同目标,维持其最优性,并允许非故障机器人无需等待已崩溃机器人即可推进,从而在有限时间内完成聚集。
链接: https://arxiv.org/abs/2610.06939
作者: Animesh Maiti,Prakhar Shukla,Abhinav Chakraborty,Subhash Bhagat
机构: Indian Institute of Technology Jodhpur; Birla Institute of Technology Mesra
类目: Distributed, Parallel, and Cluster Computing (cs.DC); Computational Geometry (cs.CG); Multiagent Systems (cs.MA)
备注:
Abstract:We study the \textitoptimal gathering problem over a finite set of designated \textitmeeting nodes for \textitasynchronous, anonymous, and \textitoblivious mobile robots on an infinite grid under crash faults. The robots have global visibility and strong multiplicity detection, but share neither a coordinate system nor chirality. The objective is to gather all non-faulty robots at a \textscWeber Meeting Node, minimizing the total Manhattan distance from their initial positions. Up to n-2 robots may crash permanently, and such crashes are indistinguishable from arbitrary delays. Existing approaches often rely on a designated robot to break symmetry, whose crash may block the remaining robots indefinitely. Instead, our approach enables every robot to independently select the same target from its snapshot, while target-dependent restricted shortest paths preserve the target as a \textscWeber Meeting Node. We prove that, under strong multiplicity detection, optimal gathering is impossible from certain fully symmetric configurations. For all remaining configurations, our algorithm \textscCrashTolerantWeberGathering() selects a unique common target, preserves its optimality throughout the execution, and allows non-faulty robots to progress without waiting for crashed robots, thereby guaranteeing gathering in finite time.
[MA-23] DIBench: Benchmarking Decision Integrity of GUI-based Mobile Agents Under Deceptive Injections NEURIPS2026
【速读】:该论文旨在解决移动代理在真实应用界面中自主决策时存在的决策完整性(decision integrity)风险问题,尤其关注多候选选择任务中因隐蔽的界面欺骗注入(deceptive injection)导致的“任务内目标偏离”(in-task goal deviation)风险。现有评估基准主要依赖任务完成率或劫持率等执行层面异常指标,无法有效捕捉攻击者通过非特权UI内容诱导代理偏离原始指令约束(如选择最便宜或评分最高的选项)但不引发明显执行异常的情况。为此,论文提出DIBench——一个针对移动代理决策完整性的统一基准,涵盖7个商业应用与3个模拟应用、5类典型任务,构建了8种隐蔽注入探测实例,在不触发显式异常的前提下可有效引导关键选择行为。该基准包含1,000个正常样本与36,672个注入样本,并提供标准化协议与完整性度量体系。实验表明,基于生成结果的评估方法可能高估代理可信度,掩盖实际决策风险;隐蔽注入会改变早期动作策略并提升完成率,造成虚假的安全假象。现有防御措施(如检测、图像预处理、提示提醒)在提升决策完整性方面效果不一。DIBench为量化移动代理在任务过程中的目标偏离风险提供了可复现、可比较的评估框架,推动安全防御技术的系统性验证。
链接: https://arxiv.org/abs/2610.06898
作者: Li Hu,Kanghua Mo,Yingbin Jin,Qingqing Ye,Haibo Hu
机构: The Hong Kong Polytechnic University(香港理工大学)
类目: Cryptography and Security (cs.CR); Multiagent Systems (cs.MA)
备注: Accepted to NeurIPS 2026, Track on Evaluations and Datasets
Abstract:As GUI-based mobile agents rapidly progress, rigorous safety evaluation of their autonomous decision-making in realistic app interfaces becomes increasingly critical. Existing benchmarks mainly focus on execution-level anomalies using task success or hijack rates, but fail to capture the in-task goal deviation risk in multi-candidate selection tasks, where the decision may be steered toward an attacker-specified target, even in violation of instruction-implied constraints (e.g., cheapest/highest-rated), without any overt execution anomalies. We present DIBench, a decision integrity benchmark for measuring this risk in mobile agents. DIBench covers 7 commercial and 3 simulated apps with 5 task types. Under a threat model restricted to non-privileged UI content, we construct 8 deceptive injection probe instantiations that can steer critical selections without overt anomalies. The benchmark includes 1,000 clean and 36,672 injected instances, with a unified protocol and integrity metrics for comparison. Experiments spanning 4 agent frameworks and 7 base models show that completion-based evaluation can overestimate agent trustworthiness and miss decision-integrity risks: deceptive injections steer selections and shift early action policies, inflating completion rates and creating a misleading illusion of safety. Common defenses, including detection, image preprocessing, and prompt reminders, yield inconsistent integrity gains. Overall, DIBench provides a unified, reproducible benchmark to quantify the risk of in-task goal deviation in mobile agents and enable comparable evaluations of safety defenses.
[MA-24] Axiom Satisfiability of Linear Rewards in Alignment
【速读】:该论文旨在解决在人类偏好数据驱动的语言模型对齐过程中,如何在保持社会选择公理(如帕累托最优性,PO;弱帕累托单调性,PMC)的前提下,实现线性奖励模型(linear reward model)的可计算性与有效性问题。传统方法在固定特征表示下使用非递减凸损失函数(如BTL)拟合奖励时,无法同时满足PO与PMC,且任何仅依赖多数关系的规则在要求输出为线性诱导时均无法满足PO。为克服此限制,论文提出将线性模型放宽为允许每项候选者引入松弛变量(per-candidate slack),通过求解最小总松弛量以保证公理在给定边际η下的严格满足。理论分析表明,当η ≤ O(1/m²)时,最优总松弛量为O(1),且该界在实例上紧致。考虑到实际场景中候选数远大于特征维度,且仅线性奖励可用于未见响应评估,作者进一步提出一种新方法:通过参数λ权衡线性部分在比较中的错误数量与总松弛量,使算法同时最小化两者。实验结果表明,该方法在合成与真实偏好数据上均优于标准线性BTL,并验证了其理论性质——总松弛量随λ单调递增但趋于饱和,且始终低于(η + Δ√d)⌊m²/4⌋,其中Δ和d分别为特征空间的直径与维度。核心解决方案在于引入结构化松弛机制与双目标优化框架,从而在不牺牲公理性的同时实现高效、鲁棒的线性奖励学习。
链接: https://arxiv.org/abs/2610.06892
作者: Soumya Nasipuri,Sayak Ray Chowdhury,Sanjukta Roy
机构: Indian Statistical Institute (印度统计研究所); Indian Institute of Technology Kanpur (印度理工学院坎普尔分校)
类目: Computer Science and Game Theory (cs.GT); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Multiagent Systems (cs.MA)
备注: 22 pages, 5 figures
Abstract:Learning from human preference data is the dominant route to aligning language models with human values. In linear social choice, where rewards are linear in a fixed feature representation of prompt-response pairs, Ge et al.[2024] show that fitting such a reward by minimizing any non-decreasing convex loss, including BTL, fails PO and PMC. Moreover, no rule that reads only the majority relation can satisfy PO once the output is required to be linearly induced. We ask what it costs to enforce these axioms anyway. To this end, we relax the linear model to allow per-candidate slack. We compute the relaxed linear reward with the smallest total slack that satisfies the axioms with a margin \eta , the minimum required difference between two reward values. Our solution satisfies the axioms under no assumptions about the voters or how comparisons were collected. We bound the optimal total slack by O(1) when \eta is at most O(\frac1m^2) for m candidates. Furthermore, we exhibit an instance that forces this bound, concluding that the rate is tight up to constants. In practice, the no. of candidates far exceeds the feature dimension, and only a linear reward can be evaluated on unseen responses. We therefore introduce a new method that charges the linear part for each comparison it gets wrong. It simultaneously minimizes the total slack and the no. of violations, with a parameter \lambda trading off between them. We show that the total slack is monotone but saturating in \lambda : raising it reduces the violations of the deployed linear reward and increases the slack, yet the slack stays below (\eta+\Delta\sqrtd)\lfloor m^2/4\rfloor , where \Delta and d are the diameter and dimension of the features, respectively. Experiments on both synthetic and real-life preference data corroborate our theory and show that the linear reward output by our method beats linear BTL.
自然语言处理
[NLP-0] IdeaAnchor: Teaching LLM s to Turn Literature into Research Ideas
【速读】: 该论文旨在解决生成式 AI 在科学文献驱动的创新性研究构想生成任务中缺乏结构化指导的问题。现有基于提示(prompting)或反馈的方法难以有效建模多篇文献之间的系统性整合机制,导致生成想法的逻辑连贯性与创新深度不足。其解决方案的关键在于提出一种名为IdeaAnchor的新范式,通过引入结构化规范作为特权信号(privileged signals),明确每篇输入文献在生成新构想中的功能角色、相互关系及合成标准。该范式通过挖掘已发表论文中真实的研究思路演化过程构建训练实例,并结合示范学习(demonstration)、自蒸馏(self-distillation)与强化学习进行模型训练,同时在推理阶段引入检索机制以增强细节丰富度。实验表明,该方法显著提升了研究构想的质量;分析进一步揭示:基于锚点(anchor-based)的训练强化了创造性整合能力,而检索机制则促进细节展开,二者结合可实现最优性能。
链接: https://arxiv.org/abs/2610.08781
作者: Ziyu Chen,Yilun Zhao,Jiashuo Sun,Yiling Ma,Manasi Patwardhan,Arman Cohan
机构: Yale University(耶鲁大学); The University of Chicago(芝加哥大学); University of Illinois Urbana-Champaign(伊利诺伊大学厄本那-香槟分校); Tata Consultancy Services(塔塔咨询)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注:
Abstract:Scientific research often begins by synthesizing ideas from a set of related papers to identify gaps and formulate new directions. However, training language models to perform this form of literature-grounded ideation remains challenging, as existing approaches based on prompting or feedback lack structured supervision for how papers should be synthesized. We introduce IdeaAnchor, a paradigm for training LLMs to perform research ideation using structured specifications as privileged signals. Each IdeaAnchor instance encodes how each input paper should be synthesized into a successful idea, including their functional roles, relationships, and target synthesis criteria. We build this paradigm by mining instances from published papers, capturing how real ideas emerge from prior literature. We then train models via demonstration, self-distillation, and reinforcement learning, and further enhance generation with retrieval at inference time. Experiments show consistent improvements in ideation quality. Our analysis reveals a functional decomposition: anchor-based training strengthens creative synthesis, retrieval enhances detail elaboration, and combining both yields the best performance.
[NLP-1] Sherpa: Teaching LLM s to Teach Adaptively ALT
【速读】: 该论文旨在解决当前大语言模型(LLM)作为教学者时存在的“因材施教”能力不足的问题,即现有方法依赖示范、偏好数据或预设教学准则,难以根据个体学生的真实学习成效动态调整教学策略。其核心解决方案是提出Sherpa——一种基于多轮强化学习的教学框架,通过构建多个由LLM模拟的、具有不同学习偏好的学生原型(student archetype),并以直接优化这些学生的学习结果为目标,训练教师模型实现个性化教学策略的自适应。该方案的关键在于将教学效果与具体学习者的实际表现挂钩,使教师能够针对不同学习风格动态调整指令,从而显著提升教学有效性。实验表明,采用Sherpa训练的教师模型在各类学生原型上的平均性能提升达20.5个百分点,且在MathTutorBench评估中整体教学评分从52.5%提升至79.2%,人类对比测试也显示其在79.6%的情况下更受青睐,验证了其在模拟真实教学场景中的优越性与可扩展性。
链接: https://arxiv.org/abs/2610.08778
作者: Weixian Xu,Yanzhe Zhang,Zora Zhiruo Wang,Changyu Chen,Diyi Yang
机构: Stanford University (斯坦福大学); Georgia Tech (佐治亚理工学院); Carnegie Mellon University (卡内基梅隆大学)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: 32 pages, 6 figures. Code and model are available at this https URL
Abstract:Large language models (LLMs) have become increasingly capable problem solvers, but being able to solve a problem is not the same as being able to teach it. Existing approaches to training LLMs as teachers rely on demonstrations, preference data, or predefined pedagogical criteria that specify what good teaching looks like. However, these signals are often not grounded in individual student learning outcomes, where effective teaching strategies can vary substantially across learners. To address this, we introduce Sherpa, a multi-turn reinforcement learning framework that instantiates multiple student archetypes with LLMs conditioned on distinct learning preferences and trains a teacher model to adapt its instruction by directly maximizing their learning outcomes. Teacher LLMs trained with Sherpa improve instructed students’ performance across all archetypes by an average of 20.5 percentage points. Under MathTutorBench’s evaluation, Sherpa raises the overall pedagogy score from 52.5% to 79.2%, indicating better teaching responses. Our human studies show that the trained teacher is preferred over the base model in 79.6% of pairwise comparisons. Together, Sherpa trains LLM teachers to adapt to diverse simulated students and become better aligned with human teachers, paving the road towards AI tutors teaching real students.
[NLP-2] AdvSim2Real : Training Web Agents Against Adaptive Prompt Injection in a Web World Model UAI
【速读】: 该论文旨在解决生成式AI(Generative AI)在网页代理(Web agent)任务中因第三方页面注入恶意指令而导致目标偏离的问题。现有防御方法通过在训练前固定注入样本进行微调,但攻击者可针对已训练模型自适应地设计新攻击,从而绕过防御。传统对抗训练虽允许攻击者动态适应,却因任务固定导致模型学不到新知识,一旦任务被解决便停止进步。为此,本文提出AdvSim2Real框架,其核心在于在一个冻结的网页世界模型内,协同演化任务课程(task curriculum)、注入攻击者(injection adversary)与代理模型。其中,任务课程通过奖励代理约一半成功率的任务来维持挑战性,而攻击者仅在成功“翻转”(success flip)——即把原本成功的任务变为失败——时获得奖励。该机制使代理在模拟环境中持续学习应对复杂、动态的对抗性注入。实验表明,该方法训练出的40亿参数代理不仅能力显著提升,且在面对未见过的前沿模型攻击者时仍保持高鲁棒性,其性能提升可迁移至真实浏览器环境;在150个网页任务上,相对基线代理,完成率提升了33.6%。因此,解决方案的关键在于通过协同进化机制实现动态、可持续的对抗训练,从而突破传统防御的局限。
链接: https://arxiv.org/abs/2610.08773
作者: Sarim Hashmi,Mukul Ranjan,Kshitij Mishra,Mikhail Kuznetsov,Praneeth Vepakomma,Nils Lukas
机构: Mohamed bin Zayed University of Artificial Intelligence (穆罕默德·本·扎耶德人工智能大学); Amazon(亚马逊); Massachusetts Institute of Technology (麻省理工学院)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: Code at this https URL
Abstract:Web agents complete user requests by reading and acting on pages that third parties write, so an instruction planted on a page can redirect the agent away from the user’s goal. The agent cannot simply ignore the page, because the page also holds the values and controls the task requires. Current defenses fine-tune the agent on injections fixed before training, and attackers that adapt to the trained model bypass them. Adversarial training lets the attacker adapt but keeps the tasks fixed, so a task stops teaching once the agent solves it. We introduce AdvSim2Real, which co-evolves a task curriculum, an injection adversary, and the agent inside a frozen web world model. The curriculum is rewarded for tasks the agent solves about half of the time, and the adversary only for a success flip, an injection that turns a judged success into a failure. Training in the simulator makes a 4B agent both more capable and more robust: its completion rises with and without attacks, holds against a frontier-model adversary it never trained against, and its capability gain carries over to a real browser. On 150 web tasks, AdvSim2Real raises completion under this unseen adversary by 33.6% relative to the base agent.
[NLP-3] he Missing Minimal Pair: Stereotype Evaluation in LLM s
【速读】: 该论文旨在解决现有大型语言模型(Large Language Models, LLMs)偏见评估方法中存在的可靠性问题,特别是基于单对对比句(single-pair contrastive sentences)的偏见测量方法易受逻辑不一致影响的缺陷。传统方法通过比较两个具有对立刻板印象的句子的对数似然值来量化偏见,但当同一刻板印象仅通过替换属性进行重写时,可能产生矛盾的偏好结果,导致评估不可靠。为此,论文提出一种双最小对(dual minimal pair)评估框架,其关键在于引入两个可比维度:一是通过数据增强框架生成同义改写句与替代属性,填补现有刻板印象数据集的空白,涵盖英语、俄语、西班牙语和中文等多种语言;二是设计两种针对该双轴对比结构的评估指标,其中一项基于互信息(Mutual Information, MI)的指标,从社会群体与刻板属性之间的统计依赖性角度重新定义偏见强度,能够更稳健地支持跨语言、跨模型的偏见强度聚合与比较。该方法显著提升了偏见评估的可靠性和可比性,为多语言场景下的模型公平性分析提供了新范式。
链接: https://arxiv.org/abs/2610.08747
作者: Nataliya Stepanova,Ivan Titov,Emily Allaway,Björn Ross
机构: University of Edinburgh(爱丁堡大学); University of Amsterdam(阿姆斯特丹大学)
类目: Computation and Language (cs.CL)
备注:
Abstract:A common approach to measuring bias in Large Language Models is to compare the log-likelihoods of two contrastive stereotype sentences. We argue that such single-pair comparisons are often unreliable: simply rewriting the same stereotype with an alternative attribute can yield logically inconsistent preferences. To address this, we propose a dual minimal pair setup that introduces two axes of comparison for robust stereotype evaluation. First, we present a data-augmentation framework that fills critical gaps in existing stereotype datasets by generating paraphrases and alternate attributes. We apply our framework on a set of English, Russian, Spanish and Chinese stereotypes. Second, we introduce two evaluation metrics tailored to the dual minimal pair setup. One of these metrics provides a new perspective on bias by modeling the mutual information (MI) between social groups and stereotyped attributes. This MI-based metric is better suited for aggregation and enables more robust comparisons of stereotype strength across different languages and models. Our code is available at this https URL. Subjects: Computation and Language (cs.CL) Cite as: arXiv:2610.08747 [cs.CL] (or arXiv:2610.08747v1 [cs.CL] for this version) https://doi.org/10.48550/arXiv.2610.08747 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[NLP-4] Denoising Hierarchical Representations: Joint Continuous Diffusion for Language Modeling
【速读】: 该论文旨在解决生成式语言模型在文本生成过程中存在顺序依赖性及并行生成能力受限的问题,尤其针对连续扩散语言模型(Continuous Diffusion Language Models, CDLMs)在生成质量与效率之间的权衡难题。现有方法虽通过精心设计的词元表示和扩散/流匹配空间提升了性能,但在多粒度语义信息利用与跨模态协同方面仍存在不足。为此,本文提出层级连续扩散语言模型(Hierarchical Continuous Diffusion Language Models, H-CDLMs),其核心解决方案在于引入多粒度语义模态的联合扩散机制:同时在原始词元(token-level)与由预训练词元嵌入聚类得到的粗粒度簇(cluster-level)两个层次上进行并行扩散。该框架允许每个模态独立配置采样器与扩散调度策略,从而增强不同粒度间的交互与互补性。实验表明,基于此框架构建的H-CoBit在多个基准测试中取得显著提升,在LM1B和OWT数据集上分别实现49.4和50.4的生成困惑度(GenPPL),较基线分别降低24.2和20.7点,并超越同规模离散扩散语言模型;在GSM8K任务中达到27.4%的准确率,优于先前连续扩散与流匹配模型。此外,该框架在流匹配模型FLM上的成功应用进一步验证了其对不同连续生成范式的通用性。整体上,该方法以极低的计算与参数开销实现了性能突破。
链接: https://arxiv.org/abs/2610.08738
作者: Mathias Ollu,Nikos Komodakis
机构: Ecole Polytechnique; Archimedes, Athena RC; University of Crete; IACM-Forth
类目: Computation and Language (cs.CL); Machine Learning (cs.LG)
备注: 27 pages, 10 figures
Abstract:Diffusion Language Models (DLMs) hold the promise of order-agnostic, parallel text generation. Recently, continuous diffusion and flow matching models have seen substantial gains, driven by carefully crafted token representations and diffusion/flow spaces. In this work, we introduce Hierarchical Continuous Diffusion Language Models (H-CDLMs), a simple framework that further improves continuous DLMs with minimal compute and parameter overhead. Drawing on the discrete DLM and continuous image diffusion literature on joint diffusion, we diffuse multiple modalities in parallel. These modalities represent tokens at different semantic granularities: in our instantiation, the tokens themselves and coarser clusters obtained by clustering pretrained token embeddings. We propose a general setup that allows per-modality samplers and schedules to enhance the interplay between modalities. Applied to CoBit, this yields H-CoBit, which delivers large empirical gains across benchmarks. At dataset entropy, H-CoBit improves MAUVE and reaches a generative perplexity (GenPPL) of 49.4 on LM1B and 50.4 on OWT, improving on the baseline by 24.2 and 20.7 points and surpassing even discrete DLMs of comparable size. On GSM8K, it reaches 27.4% accuracy, outperforming prior continuous diffusion and flow-based models. We further apply H-CDLM to the flow matching model FLM, obtaining consistent gains with H-FLM and demonstrating that the framework generalizes across continuous generative paradigms. Our code will be made publicly available at this https URL .
[NLP-5] Holdout Best-of-N: Unbiased Evaluation and Its Cost
【速读】: 该论文旨在解决在基于评分的模型评估中,利用“Best-of-N”选择机制所生成的分数来估计最优候选者的期望奖励时可能出现的偏差问题。具体而言,当仅依赖一个固定评分矩阵(每候选者有 K 个独立评分)进行策略选择(使用 J 个新评分)时,现有估计算法可能因重复使用已用于选择的评分而高估其真实预期奖励。其核心解决方案在于:当且仅当 J=K 时,存在一个仅基于该固定评分矩阵的单一致估计算子,可在任意独立且稳定的候选者特定评分分布下实现无偏估计。当 J=K−1 时,选择器深度随 K 增大而加深,此时在独立同方差高斯评分设定下,无偏最小最大风险为 σ2/K,由“留出法”(Holdout)达到;若允许引入偏差,则可将收敛速率提升至 σ2/K。对于两候选情形,论文推导出已知方差下的最小方差无偏估计量及精确渐近无偏最小最大常数 1/(π2),该值可通过无需知晓方差的 Holdout 方法实现。此外,通过循环平均子集与平局处理,可在 O(MKlogM) 时间内完成计算;在固定选择器深度下,对有界评分的循环评估具有统一于池大小的 O(K−1) 风险。最终的不可能性结果表明,在固定评分矩阵条件下,只需额外一个全新的获胜者评分即可实现对全 K 策略的无偏评估。
链接: https://arxiv.org/abs/2610.08719
作者: Shrey Shah,Yinheng Li
机构: 未知
类目: Computation and Language (cs.CL)
备注: 25 pages, 2 figures, 3 tables
Abstract:Reusing the scores that select a Best-of- N winner can overstate its expected reward. We study evaluation from a fixed matrix of K independent scores per candidate for a policy that selects using J fresh scores. A single estimator based only on this matrix is exactly unbiased for expected judge reward under every independent, stable collection of candidate-specific score laws if and only if JK , for every pool size M\ge N\ge2 . At J=K-1 , the selector deepens as K grows. For independent Gaussian scores with common variance and fixed M\ge N\ge2 , the unbiased minimax risk in this regime is of order \sigma^2/\sqrt K , attained by Holdout; allowing bias improves the rate to \sigma^2/K . For two candidates, we derive the minimum-variance unbiased estimator at known variance and the sharp asymptotic unbiased minimax constant 1/(\pi\sqrt2) , which Holdout attains without knowing the variance. The cyclic average over subsets and ties can be computed in O(MK\log M) operations. At fixed selector depth, cyclic evaluation of bounded scores has O(K^-1) risk uniformly in pool size. The impossibility result concerns the fixed matrix: one additional fresh winner score permits unbiased evaluation of the all- K policy.
[NLP-6] When Forgetting is not Catastrophic: On the Mechanics of Spurious Forgetting
【速读】: 该论文旨在解决语言模型在微调过程中出现的“伪遗忘”(spurious forgetting)问题,即模型看似丢失了旧知识,但这些知识实际上仍被存储并可恢复。其核心问题是:在持续微调新知识时,旧知识为何会出现先崩溃、后恢复、最终又永久消失的非单调遗忘动态?解决方案的关键在于揭示遗忘机制的双重性——既包含可逆的、共享的访问损失(由参数更新沿共同方向移动旧表征引起),也包含不可逆的、个体化的事实侵蚀(由特定事实的累积变化导致)。研究发现,这种现象可通过一个最小关联记忆模型再现,其关键要素包括:具有共享结构的键(keys)、集中式的新值(concentrated new values)以及网络中的归一化机制。其中,归一化在新知识学习完成后会撤销共同偏移,从而恢复旧知识;而个体事实的变化则持续积累并导致长期遗忘。因此,真正的灾难性遗忘仅源于后者,且其主导程度取决于新数据是否将旧记忆“聚集”或“分散”。通过从权重更新中移除单一方向,可在预训练语言模型中有效恢复旧知识,验证了该机制的可调控性。
链接: https://arxiv.org/abs/2610.08718
作者: Vedant Palit,Florent Draye,Nicolas Zucchet,Zhijing Jin,Bernhard Schölkopf
机构: MPI for Intelligent Systems (马普智能系统研究所); Jinesis Lab, University of Toronto (多伦多大学杰尼西斯实验室); Vector Institute (向量研究所); EuroSafeAI; Hector Foundation; Stanford University (斯坦福大学); ELLIS Institute Tübingen (蒂宾根ELLIS研究所)
类目: Computation and Language (cs.CL); Machine Learning (cs.LG)
备注:
Abstract:Knowledge that a language model appears to forget during finetuning often remains stored and can be recovered, a phenomenon called spurious forgetting. Finetuning on new facts can even produce forgetting that undoes itself: recall of the old facts collapses, recovers as training continues on new facts alone, and only then erodes for good. We seek to understand when such forgetting is not catastrophic. A minimal associative memory reproduces these dynamics with three ingredients: keys with shared structure, concentrated new values, and normalization in the network. Finetuning moves all old representations along a common direction, hiding the old facts while preserving their relative geometry; normalization withdraws this shift once the new facts are learned, whereas fact-specific changes accumulate and cause the erosion. Moreover, subtracting the common shift eliminates the collapse in a Transformer trained on synthetic data, and removing a single direction from each weight update restores old facts in a pretrained language model. Forgetting thus combines a shared, reversible loss of access with a slow erosion of individual facts, and only the second is catastrophic. Which one dominates depends on whether the new data move old memories together or apart.
[NLP-7] Agreement Is Not Validity: Cross-Model LLM Consensus in Diagnosing Student Failure Modes in K-12 Math Tutoring Dialogue
【速读】: 该论文旨在解决在K-12数学辅导对话中,基于大语言模型(LLM)对学习者错误类型进行分类的解释有效性问题。当前学习分析研究广泛依赖生成式AI从对话中提取学生解题过程与困难来源信息,用于知识追踪、行为建模及推理错误诊断等下游任务,但这些模型生成的分类结果是否具有实际有效性仍缺乏充分验证。本研究通过操作化诊断编码手册,评估了多类LLM对五种典型学生失败模式(不确定性、归因错误、算子选择错误、概念性缺口、程序性失误)的分类表现。结果显示,人类专家与模型间的一致性仅为中等(kappa = 0.524–0.597),而不同模型之间的共识则显著更高(kappa = 0.755–0.781;alpha = 0.769)。这一发现揭示了模型间高一致性可能产生“正确”的假象,挑战了“多模型共识即有效”的假设。研究的关键启示在于:在学习分析中,尽管大规模自动化标注具有可扩展性优势,但其有效性必须建立在独立验证的基础之上,模型间的共识不能替代对构念有效性的独立证据支持。
链接: https://arxiv.org/abs/2610.08703
作者: Clayton Cohn,Joyce Fonteles,Kirk Vanacore,Gianni Mazza,Candida Crawford,Tom Hooper,Gautam Biswas,Rene Kizilcec
机构: 未知
类目: Computation and Language (cs.CL)
备注: Submitted to LAK27 as a short paper. Currently under review
Abstract:In K-12 mathematics tutoring, student-tutor dialogue provides rich evidence of learners’ problem-solving processes and sources of difficulty. Learning analytics research increasingly relies on large language models (LLMs) to extract such information from dialogue for a variety of downstream tasks, including knowledge tracing, behavioral modeling, and diagnosis of student reasoning errors. However, the validity of these model-generated interpretations remains insufficiently understood. In this exploratory study, we examine the validity of LLM classifications of five student failure modes in mathematics tutoring dialogue using an operational diagnostic codebook: uncertainty, misattribution, operator selection, conceptual gap, and procedural slip. Across models, human-LLM agreement was moderate (kappa = .524-.597), while cross-model agreement was substantially higher (kappa = .755-.781; alpha = .769). These findings show that cross-model agreement can create a misleading appearance of correctness, challenging the assumption that consensus among LLMs constitutes evidence of valid learner interpretation. For learning analytics, the implication is clear: scalable labeling is useful only if the inferred constructs are valid, and model consensus cannot substitute for independent evidence of that validity.
[NLP-8] A Systematic Study of Small Language Models on Abstract Reasoning Tasks
【速读】: 该论文旨在解决语言模型在抽象推理基准测试中所表现出的端到端准确率是否真正反映了其对可迁移规则的学习,还是仅拟合了特定数据分布的统计规律这一关键问题。其核心解决方案在于通过ARC-TGI基准测试框架,在可控的任务家族设置下系统评估小规模语言模型在监督微调(supervised fine-tuning)条件下的技能习得效率、稳定性与泛化能力,该框架支持重采样、空间平移及跨基准迁移,从而有效分离出规则学习与分布内拟合的区别。研究发现,尽管模型可在训练分布内达到较高准确率,但其技能获取对优化过程高度敏感且在任务族间分布不均;当测试分布超出训练范围时(如网格尺度变化),性能急剧下降,即使规则本身保持不变。此外,增加训练集深度与广度带来的提升效果不均衡,而上下文示例数量的影响则依赖于模型架构。值得注意的是,通过可执行规则推导获得的解在直接生成网格时并未出现,表明规则理解具有非平凡性。注意力诊断分析揭示了不同任务下显著不同的注意力集中模式与上下文依赖特征,但尚未能建立普遍的因果机制。总体而言,抽象推理表现受模型类型、适配策略、评估分布和响应格式等多重因素共同决定。
链接: https://arxiv.org/abs/2610.08680
作者: Nur A Zarin Nishat,Jens Lehmann,Andrei Aioanei,Sahar Vahdati
机构: Leibniz University of Hanover; TIB – Leibniz Information Centre for Science and Technology and University Library; Amazon
类目: Machine Learning (cs.LG); Computation and Language (cs.CL)
备注:
Abstract:Endpoint accuracy on abstract-reasoning benchmarks does not reveal whether a language model has acquired a transferable rule or fit distribution-specific regularities. We study this distinction in small language models on the ARC-TGI benchmark, which organizes abstract grid transformations into controllable task families and supports resampling, spatial shifts, and cross-benchmark transfer. Across more than 1,000 runs, we profile decoder-only, encoder–decoder, and mixture-of-experts model families under supervised fine-tuning. We examine the efficiency and stability of skill acquisition, robustness beyond the training distribution, interactions with model family and task formulation, and layer-wise attention signatures that accompany behavioral differences. Substantial in-distribution accuracy is attainable, but acquisition is sensitive to optimization and unevenly distributed across task families. Performance deteriorates sharply outside the training distribution, including when the rule is retained but grid scale changes. Greater training-set depth and breadth yield uneven gains, while the effect of additional in-context examples depends on model family. Executable-rule induction also yields correct solutions not observed under direct grid generation. On selected tasks, attention diagnostics show distinct concentration and context-dependence profiles, but do not establish general causal mechanisms. Overall, abstract-reasoning scores are conditional on the model, adaptation regime, evaluation distribution, and response format.
[NLP-9] Same-Number Citation Swaps: Stress-Testing Jev as a Financial Evidence Judge
【速读】: 该论文旨在解决生成式人工智能(Generative AI)在财务报告分析中因数值重复导致的“角色误引”问题,即模型虽能生成数值正确的计算结果,却可能引用了错误的财务条目(financial role)。其核心解决方案在于通过引入概率证据验证机制(probabilistic evidence verification),以超越单纯依赖数值匹配的局限。研究以Jev作为源支持验证器,对GPT-4.1-mini的计算轨迹进行评估,发现基于符号位置的基准方法(signed-number-at-pointer baseline)已能解释大部分恢复效果。为隔离角色识别问题,研究固定运算逻辑与数值,仅调整引用来源在相同数值单元格间的切换,并保持语义等价的控制条件。实验揭示出两类关键现象:部分错误引用可通过验证,而部分有效替代引用却被拒绝。进一步构建的36页新源数据集验证表明,显式列标签虽提升对错误角色的识别能力,但同时削弱了对等价证据的支持,暴露出检测角色错误与保留有效引用之间的权衡。本研究贡献在于,在保持数值匹配不变的前提下,系统性地识别出概率性财务验证器所区分的关键特征,为基于大语言模型的财务助手提供了可独立评估数值正确性、引用角色支持度与接受结果的框架。
链接: https://arxiv.org/abs/2610.08675
作者: Chuhong Xu(Sofia University),Bo Su(Indiana University),Ziyao Chen(University of California, San Diego),Ruiyang Xu(Northeastern University),Shimeng Dai(Michigan State University),Xinyu Qiu(Northeastern University)
机构: Northeastern University(东北大学); University of California, San Diego(加州大学圣地亚哥分校); Michigan State University(密歇根州立大学); Indiana University(印第安纳大学)
类目: Computation and Language (cs.CL)
备注: counterfactual citation perturbation, evidence attribution verification, financial document question answering, Jev, LLM-as-a-judge, probabilistic source verification, tabular numerical reasoning
Abstract:Financial reports repeat values across periods, metrics and accounting lines, allowing an LLM-generated calculation to be numerically correct while citing the wrong financial role. We evaluate what probabilistic evidence verification adds beyond number matching using Jev as a source-support verifier for GPT-4.1-mini calculation traces. A signed-number-at-pointer baseline explains most recovery over exact quotation checks. To isolate the remaining role-recognition problem, we hold operands and arithmetic fixed, move citations between same-number cells, and retain controls that express equivalent facts. These contrasts reveal both wrong-role citations that pass and valid alternative citations that are withheld. Explicit column labels improve selected wrong-role decisions while also lowering support for some equivalent evidence. A constructed follow-up on 36 new source pages, labeled by a non-author reviewer, extends this evaluation and exposes the same tradeoff between detecting role errors and retaining valid citations. The contribution is a controlled evaluation that identifies what a probabilistic financial verifier distinguishes when numerical matching is held fixed. For LLM-based financial assistants, it makes numerical correctness, cited-role support and acceptance outcomes separately assessable.
[NLP-10] Principled Under Pressure: Post-Training Decides Whether LLM s Act on Their Own Moral Judgment
【速读】: 该论文旨在解决生成式 AI(Generative AI)在作为智能体(agent)时,存在“明知故犯”行为却无法被传统价值观评估所捕捉的问题。具体而言,当模型在面临外部压力时,虽能判断某行为错误,但仍可能执行该行为,这种行为偏差无法通过仅考察其陈述的价值观来识别。为揭示此问题,研究构建了一个预注册的248个场景面板,涵盖五类压力情境,并对同一模型在两种角色下进行测试:一是以第一人称作为决策主体选择行动;二是以第三人称判断何者为正确选项,从而将模型自身的判断作为参照基准。每个场景均设有无压力的对照版本,并引入正向控制条件(即由操作者直接命令执行违规行为),以区分“缺失道德判断”与“盲从指令”的差异。实验结果表明,在OLMo-3-7B-Instruct模型中,模型在有压力情境下执行其自认为错误的行为频率约为五分之一,显著高于无压力情形。不同指令微调(instruct fine-tuning)策略的影响显著:基于Meta和Ai2的Llama-3.1权重但采用不同后训练流程的模型表现出明显差异——Meta的方案保留了这一“判断-行为差距”(gap),而Tulu 3和Qwen2.5-7B-Instruct则未在整体或关键压力场景中表现出显著差距。此外,脱离对话模板阅读模型输出会反转该差距符号,说明模型表现受输入格式干扰。进一步分析发现,推理前预判后果可使行为更趋近于模型自身判断,且明确指出规范内容可实现约三分之一的修正效果。因此,该“判断-行为差距”并非预训练权重固有的属性,而是后训练流程可调节的可测量目标,提示可通过优化微调策略提升智能体的伦理一致性。
链接: https://arxiv.org/abs/2610.08670
作者: Orion Reblitz-Richardson
机构: Distiller Labs(迪斯蒂勒实验室)
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: 33 pages
Abstract:Language models increasingly act as agents. An agent that says an action is wrong and then takes it anyway is a different failure from one that does not know better, and evaluations of stated values cannot see it. We build a pre-registered panel of 248 scenarios across five kinds of pressure. Each scenario is posed twice to the same model, once as the agent choosing what to do and once in the third person asking which option is right, so the model’s own judgment is the reference. Every scenario has a twin with the pressure removed, and every model gets a positive control in which its operator orders the violating action, so that a missing gap can be told apart from a blind instrument. On OLMo-3-7B-Instruct, the model takes the action it judged wrong on about one in five pressuring scenarios, more often than on the same scenarios with the pressure removed. Across four instruct models the gap depends on the post-training recipe: OLMo-3 and Meta’s Llama-3.1-8B-Instruct carry it; Tulu 3 shows none on the whole panel (above about 0.01 in probability) or on its own most-pressuring scenarios; Qwen2.5-7B-Instruct shows none on the whole panel (above about 0.02) and is unresolved on its own (0.083, -0.028 to 0.195). Meta’s recipe and Ai2’s Tulu 3 start from the same Llama-3.1 weights, and only Meta’s carries the gap. Reading a chat model outside its chat template reverses the sign of its gap with nothing at stake (-0.038 against +0.055 under the template on OLMo-3), a distortion present on two of three recipes. On both models that carry it, reasoning about the stakes before acting moves the choice back toward the model’s own judgment, against a same-length non-moral task, with or without the pressure; on OLMo-3, naming the norm at stake does about a third of that. The gap is a measurable target for post-training recipes, not a fixed property of pretrained weights.
[NLP-11] Evidence-Bound Reasoning : Neuro-Semantic Verification of Biomedical AI in Glioblastoma Radiogenomics
【速读】: 该论文旨在解决生成式人工智能(Generative AI)在生物医学领域应用中缺乏可靠证据验证的问题,即模型生成的解释虽具合理性,但无法确保每条陈述均基于患者特异性证据。其解决方案的关键在于构建一个神经-语义验证框架(neuro-semantic verification framework),将影像组学(radiomic)测量转化为可定位的证据记录与机器可检查的命题(machine-checkable claims)。该框架通过在多中心独立队列中对UPenn-GBM影像组学特征与自动生成的CaPTk分割结果进行对齐,建立包含1,728个来自T1、T1GD、T2和FLAIR序列的肿瘤区域特征的共享空间,并基于611例患者的参考定义语义状态进行验证。研究实现了跨队列可迁移性评估、模型溯源追踪、确定性验证、受控预测退化测试以及大语言模型(LLM)命题提取试点,证明了验证能力可独立于预测性能进行工程化设计与评估。实验结果显示,验证器在6,620条命题的污染基准测试中达到100%精确集准确率,在24例试点中GPT-5.6成功复现全部72个预设原子命题,冻结的验证器亦恢复全部24个预期条件,且在模型性能下降时验证准确性仍维持1.000。这表明生成式解释可由大语言模型结构化输出,而最终的证据一致性检验仍可通过确定性机制完成,从而实现可信、可验证的智能决策支持。
链接: https://arxiv.org/abs/2610.08660
作者: Mariya Miteva,Maria Nisheva-Pavlova
机构: 未知
类目: Computation and Language (cs.CL)
备注: 15 pages, 4 figures, 4 tables. Preprint
Abstract:Background: Biomedical AI can generate plausible explanations without reliably verifying whether each statement is supported by patient-specific evidence. We developed a neuro-semantic verification framework that converts radiomic measurements into addressable evidence records and machine-checkable claims. Methods: UPenn-GBM radiomics were aligned with de novo CaPTk extraction from standardized MRI and expert-validated segmentations in an independent multicenter cohort. The shared space comprised 1,728 features from T1, T1GD, T2, and FLAIR MRI across three tumor regions. Reference-defined semantic states were derived from 611 UPenn cases. We evaluated cross-cohort transportability, model-linked provenance, deterministic verification, controlled predictive degradation, and an LLM claim-extraction pilot; MGMT prediction served only as a transport stress test. Results: Median semantic-state agreement was 0.786 (weighted kappa 0.709), ranging from 0.918 for morphologic to 0.252 for intensity features. The external evidence ledger contained 1,655 model-linked records for 331 patients. The verifier achieved 100% exact-set accuracy in a 6,620-claim corruption benchmark. In a 24-case pilot, GPT-5.6 Sol reproduced 72/72 prespecified atomic claims, and the frozen verifier recovered 24/24 expected conditions. During controlled degradation, ROC AUC declined from 0.899 to 0.500 while verification accuracy remained 1.000. External MGMT discrimination was weak (ROC AUC 0.543). Conclusions: Verifiability can be engineered and evaluated independently of predictive performance. LLMs may structure explanations, while final evidence-consistency checking remains deterministic.
[NLP-12] SquidAgent : Parallelize Wisely Coordinate Efficiently NEURIPS2026
【速读】: 该论文旨在解决基于大语言模型(LLM)的多智能体系统在并行执行时效率不升反降的问题。尽管理论上并行化可带来近线性的加速,但现有系统常因隐藏开销导致性能劣于串行单智能体基线。其核心问题源于两个关键成本:一是重探索成本(re-exploration cost),即并行工作节点重复构建本可由协调器隐式继承的上下文(如先前决策),二是对齐成本(alignment cost),即独立生成结果间不一致所带来的后期协调开销。为克服此问题,论文提出一个基于任务层的并行决策准则:仅当某层的临界路径成本加上重探索与对齐开销之和低于串行执行成本时,才应并行化。由于大模型对任务耗时的估计能力较差,作者转而以预测输出的token数量作为成本度量标准,发现该指标更可靠。基于此,论文提出SquidAgent框架:通过一次规划阶段预估所有层的token预算,直接从协调器会话分叉工作智能体以消除重探索成本,并引入预先生成的共享约定块(convention block),将对齐操作转化为可量化、可控制的前置开销。结合确定性调度器逐层应用上述准则,实验表明SquidAgent在平均吞吐量上相较Claude Code提升2.2倍,在平均墙钟时间上提速2.6倍,且较最强多智能体基线提升2.0倍。
链接: https://arxiv.org/abs/2610.08647
作者: Yexiong Lin,Shanshan Ye,Yu Yao,Zhen Fang,Bo Han,Tongliang Liu
机构: The University of Sydney(悉尼大学); Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学); University of Technology Sydney(悉尼科技大学); Hong Kong Baptist University(香港浸会大学)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
备注: Accepted at NeurIPS 2026. 37 pages, including appendices
Abstract:LLM-based agents solve complex multi-step tasks, but sequential execution incurs substantial latency. In principle, parallelizing work across multiple agents should yield near-linear speedups. Yet existing parallel multi-agent systems often run slower than a single-agent baseline. We attribute this gap to two hidden costs that parallel execution incurs but a serial agent avoids. First, there is a re-exploration cost: redundant effort spent by parallel workers reconstructing context that the orchestrator already possesses, such as prior decisions, that would otherwise be inherited implicitly in a serial execution. Second, there is an alignment cost: the overhead required to reconcile inconsistencies across independently generated outputs. We thus derive a principled decision criterion: a layer should be parallelized only when its critical-path cost, plus re-exploration and alignment overheads, is lower than the corresponding serial cost. While this criterion is naturally expressed in wall-clock time, we observe that LLMs are poorly calibrated when asked to estimate task duration. To address this, we instead measure cost in predicted output tokens, which we empirically find LLMs can estimate substantially more reliably than wall-clock time. Building on this token-based criterion, we propose SquidAgent. It estimates all token budgets in a single planning step, forks each worker directly from the orchestrator’s session to eliminate re-exploration cost, and replaces post-hoc reconciliation with a pre-generated shared convention block that converts alignment into a bounded upfront cost. A deterministic scheduler then applies the criterion layer by layer. Empirically, SquidAgent achieves a 2.2 \times mean throughput improvement and a 2.6 \times mean wall-time speedup over Claude Code, and a 2.0 \times throughput improvement over the strongest multi-agent baseline.
[NLP-13] owards In-Parameter Memory Augmentation for Large Language Models
【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在预训练后需持续融入新知识(如领域事实、用户偏好、文档信息及交互经验)时所面临的挑战,尤其针对传统上下文学习(In-context Learning, ICL)及其代理框架中存在的上下文容量消耗过大、随着上下文长度增加而产生重复离散编码开销的问题。其核心解决方案是引入参数化记忆(In-parameter memory),即通过将可复用的记忆信息以模型参数、适配器(adapter)或其他类似参数的对象形式嵌入推理过程,在不依赖上下文序列的前提下实现知识的高效存储与调用。该方法的关键在于:将记忆信息显式编码于模型参数中,并在推理时动态注入前向传播路径中,从而实现对知识的持久化与低延迟访问。论文从两个正交维度系统梳理了该技术生态:参数位置(Parameter Placement)(如嵌入层、注意力层、前馈网络层或混合结构)与参数获取时间(Parameter Acquisition Time)(部署期间获取的在线方式 vs. 部署前预设的离线方式),并深入探讨了干扰性、安全性、与ICL协同设计以及递归自我改进等开放问题。
链接: https://arxiv.org/abs/2610.08630
作者: Haoyu Huang,Zhongwei Xie,Jiaxin Bai,Yisen Gao,Hong Ting Tsang,Wuganjing Song,Huihao Jing,Yufei Li,Yangqiu Song
机构: The Hong Kong University of Science and Technology (香港科技大学); Hong Kong Baptist University (香港浸会大学)
类目: Computation and Language (cs.CL)
备注:
Abstract:Recently Large Language Models (LLMs) and LLM-based agents increasingly need to incorporate knowledge acquired after pretraining, e.g., domain facts, user preferences, documents, and interaction experience. In-context learning (ICL) and ICL-based agent harness remain flexible, but they consume context capacity and incur repeated discretized encoding cost that grows with context length. \textbfIn-parameter memory offers a complementary substrate: reusable memory information is represented in model parameters, adapters, or other parameter-like objects that are composed into the forward pass at inference time. This survey focuses on methods that augment LLMs with such parametric memory at deployment: a memory-bearing parameter object is plugged into the forward pass during inference, whether it is acquired before or during deployment. We organize the landscape with two orthogonal axes: \textbfParameter Placement, which includes Embedding, Attention, FFN layers, or Hybrid when two or more layers are used; and \textbfParameter Acquisition Time, which distinguishes methods whose memory object is acquired during deployment (online) from those acquired before it (offline). We clarify boundaries, conduct comparisons, and discuss open directions in interference, safety, co-design with ICL, and recursive self-improvement.
[NLP-14] InterCorrect: Intersection-Aware Correction of Demographic Model Merging for Fair ASR
【速读】: 该论文旨在解决基于语音大模型(Speech-LLM)的自动语音识别(ASR)系统在不同人口统计学群体间表现不均的问题,尤其关注多重人口统计学特征交叉群体(如性别与种族的交集)中识别误差难以缓解的挑战。其核心解决方案是提出一种人口统计学感知的模型融合方法:首先在基础SLAM-ASR模型上,仅对特定人口统计学子群体微调连接器(connector),随后将这些子群体适配的连接器合并为全局模型;进一步通过子群体词错误率(subgroup WER)与任务向量冲突(task-vector conflict)分析,识别出关键的跨轴人口统计学组合(cross-axis demographic pairs),并针对这些交集区域施加专门的修正向量(intersection-specific correction vectors)。实验结果表明,该方法在Fair-Speech数据集上显著降低整体词错误率(WER),其中结合基于WER的修正向量的TIES策略将WER从7.38%降至5.13%;同时,子群体性能与差异性分析显示,该方法有效提升了多个人口统计学维度的表现,但揭示了平均性能改善并不必然等同于子群体间差距的缩小,强调了公平性评估需兼顾整体性能与群体间差异。
链接: https://arxiv.org/abs/2610.08604
作者: Ashley E. Bravo-Bravo,Yuchen Zhang,Haralambos Mouratidis,Ravi Shekhar,Monorama Swain
机构: University of California, Irvine (加州大学欧文分校); University of Southern California (南加州大学)
类目: Computation and Language (cs.CL); Sound (cs.SD)
备注: Under Review
Abstract:Automatic Speech Recognition (ASR) systems often show uneven performance across demographic groups, and errors can be especially difficult to address for speakers belonging to multiple demographic groups. This work studies demographic-aware model merging for fair Speech-LLM-based ASR. Starting from a SLAM-ASR-based model, we fine-tune only the connector on demographic-specific subsets and merge the resulting subgroup-adapted connectors into a global model. We then identify critical cross-axis demographic pairs using subgroup WER and task-vector conflict, and apply intersection-specific correction vectors to the global merged model. Experiments on Fair-Speech show that global demographic merging improves overall WER over the base model, while intersection correction provides additional gains for several merging strategies. In particular, TIES with WER-based correction achieves the best overall WER, reducing it from 7.38% to 5.13%. Subgroup and disparity analyses further show that the proposed approach improves performance across demographic axes, while highlighting that lower average WER does not always imply reduced subgroup disparity.
[NLP-15] Generative AI translations in high-stakes emergency messaging
【速读】: 该论文旨在解决紧急消息(如极端天气报告和地震预警)在多语言传播中因翻译错误可能引发严重后果的问题。其核心挑战在于,在确保高准确性和可行动性的同时,兼顾翻译效率与责任归属。解决方案的关键在于利用生成式人工智能(Generative AI)进行初始翻译,并通过设计特定于语篇类型的提示(discourse-specific prompts)显著提升译文的可理解性与可操作性。尽管如此,由于译者对生成内容的信任度仍不足,且涉及重大伦理责任,人工校审仍是不可或缺的环节,不仅用于纠错,更承担最终责任主体的角色。
链接: https://arxiv.org/abs/2610.08601
作者: Nune Ayvazan,Anthony Pym,Yu Hao
机构: 未知
类目: Computation and Language (cs.CL)
备注:
Abstract:Emergency messaging such as extreme-weather reports and earthquake instructions can involve high stakes, to the extent that translation errors can lead to tragic consequences. The use of machine translation or generative artificial intelligence might therefore not be recommended. On the other hand, time savings in the initial translation can allow greater investments of resources in revision and authorization processes, as well as a wider range of target languages. An experiment with generative AI translations of an earthquake instruction text from English into Chinese and Spanish shows that use of discourse-specific prompts can considerably improve understandability and actionability, although the translations may still not be trusted by translators. Human revision is still required, not only to detect errors but also because of the ethical need for someone to take responsibility for any errors or delays in such messaging.
[NLP-16] Incidental information contaminates patient notes and disrupts clinical reasoning in large language models
【速读】: 该论文旨在解决生成式人工智能(Generative AI)在临床文档生成与临床推理应用中因对非相关信息(即偶然性信息)敏感而引发的可靠性问题。其核心挑战在于,当前前沿大语言模型(LLM)在处理患者-医生对话时,易将与诊疗无关的闲聊内容或外部背景语音误纳入临床记录,进而导致信息误标或临床判断偏差。解决方案的关键在于提出“双重编码假说”——即同一模型组件既可能因偶然信息干扰而产生分心,又同时支持临床推理能力。研究发现,尽管模型在35%的病历中插入了闲聊内容,且约3.7%出现误归因或不当临床使用,但整体质量评分变化极小(≤0.20分),表明模型对干扰具有一定的鲁棒性,但风险不可忽视。因此,论文强调应在临床部署前评估模型对偶然信息的抗干扰能力,并设计机制在防止信息污染的同时保留其临床推理功能。
链接: https://arxiv.org/abs/2610.08585
作者: Krithik Vishwanath,Brandon Ye,Anton Alyakin,John E. Markert,Aaron Hsieh,Michał Mańkowski,Eric K. Oermann
机构: 未知
类目: Computation and Language (cs.CL)
备注:
Abstract:Large language models (LLMs) are increasingly relied upon to support ambient documentation and clinical reasoning. Here we examine the impact of a failure mode shared between these two applications by assessing their sensitivity to information incidental to the patient encounter. In 576 patient-clinician dialogues, we found that frontier models inserted small-talk exchanges into 35% of notes, while mean quality scores changed by at most 0.20 points on five-point scales. In 3.7% of frontier notes, models misattributed the asides or used them clinically. In 57 mock recorded consultations, background speech from a separate patient encounter at -10 dB leaked into 48.2% of transcripts, with contamination detected in 5.3% of downstream notes generated by four open-weight models. We propose a dual encoding hypothesis of clinical reasoning and distraction in LLMs, with preliminary evidence that LLM components associated with disruption by incidental information also support clinical reasoning. These findings support evaluating resistance to incidental information before clinical use, with safeguards that prevent contamination while preserving clinical reasoning.
[NLP-17] Have I Seen Enough? Frozen Video-Language Models Encode Evidence Readiness
【速读】: 该论文旨在解决流式视频-语言模型(Streaming Video-Language Models)在实时问答中面临的“证据就绪性判断”问题,即模型需自主判断当前问题所需的视觉证据是否已完整到达。现有系统通常通过独立训练的触发机制来实现这一决策,而本文提出质疑:是否无需修改模型即可从预训练冻结的视频语言模型(VideoLLM)中直接读取该信号?其核心解决方案在于发现并利用冻结模型内部隐含的、线性可读的“证据就绪信号”(evidence-readiness signal),该信号通过时间戳标注的证据数据进行标签化训练,并可在7个字节完全相同的基准模型中稳定解码(严格未就绪采样下AUROC为0.733–0.905,接近随机水平)。该信号具有强问题条件性,在相同视频窗口仅改变问题时,66.1%的样本对呈现相反读出结果,而所有无问题依赖的对照组均处于随机水平。值得注意的是,即使模型回答错误,该信号仍保持较高有效性(错误答案中AUROC达0.722),且优于不确定性估计器及其监督融合方法在延迟匹配条件下的答案选择性能,并更贴近人类独立判断。研究进一步揭示,现有发布的流式触发机制虽也为线性读出,但其在原始模型激活值上的解码与就绪信号近似正交,准确性显著更低。基于此,作者提出“就绪门控”(Readiness Gating)策略,作为一种低开销的答案时机控制机制,在视频时长匹配条件下将准确率最高提升9.75个百分点,且增益大小与任务可用的精度余量高度相关,验证了该信号在优化推理时序中的关键作用。
链接: https://arxiv.org/abs/2610.08560
作者: Dan Ben-Ami,Kobi Cohen,Chaim Baskin
机构: INSIGHT Lab, Ben-Gurion University of the Negev (以色列内盖夫本-古里安大学); Ben-Gurion University of the Negev (以色列内盖夫本-古里安大学)
类目: Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL); Machine Learning (cs.LG)
备注:
Abstract:Streaming video-language models must decide not only what to answer, but whether the evidence needed for the current question has arrived. Existing systems learn that decision as a separate trigger; we ask whether an unmodified model already computes it. We show that frozen VideoLLMs carry a linearly readable evidence-readiness signal, labelled from timestamped evidence rather than from model output. It decodes in all seven models of a shared byte-identical evaluation (AUROC 0.733-0.905 under the strictest not-ready sampling, where a fitted clock is near chance), and a probe fitted without any of a benchmark family’s footage still reads that family. It is question-conditioned: on byte-identical windows, changing only the question reverses the readout on 66.1% of pairs, while every question-blind control is at chance by construction. The model can answer incorrectly and still encode readiness: AUROC remains 0.722 among wrong answers. Readiness also beats uncertainty estimators and their supervised combination on latency-matched answer selection, and tracks independent human judgments more closely than confidence. Released streaming triggers are also linear readouts, yet a trained trigger read on its own base model’s activations is approximately orthogonal to readiness and decodes it far less accurately than a probe. We turn the readout into Readiness Gating, an answer-timing policy that improves accuracy by up to +9.75 pp at matched video duration with negligible computational overhead. How much it gains varies with the accuracy headroom the task makes available: across 26 configurations the gain tracks that headroom, and an intervention that moves it over identical pixels moves the gain with it.
[NLP-18] Latent space bias directions in LLM s capture confidence not fairness
【速读】: 该论文旨在解决生成式AI(Generative AI)中激活量操控(activation steering)技术在推理时去偏(debiasing)应用中存在的性能不一致问题,特别是其方向向量泛化能力差、对模型性能产生非预期影响以及在新数据集上迁移效果有限的挑战。其解决方案的关键在于深入分析用于激活量操控的去偏方向实际编码的内容,发现该线性去偏方向主要反映的是模型置信度(model confidence),而非语义层面的偏见表征;具体而言,该方向在隐藏空间中从高概率词元指向低概率词元,本质上是模型置信度的体现。因此,通过该方向进行操控虽能降低测得的偏见分数,但其根本机制并非修正模型的内在偏好,而是通过削弱模型置信度导致其倾向于拒绝回答问题,从而在问答(QA)基准测试中产生“更公平”的假象。研究结果表明,在隐层空间中,偏见与反偏见提示之间的主要区分因素是模型置信度,这揭示了将偏见表示从置信度中解耦的困难性,进而强调基于激活量操控的去偏方法需谨慎解读其有效性与作用机制。
链接: https://arxiv.org/abs/2610.08559
作者: Stephanie Buttigieg,Maeve Madigan,Parameswaran Kamalaruban,Stuart Burrell
机构: Visa Inc.(Visa公司); Institute of Astronomy Kavli Institute for Cosmology, University of Cambridge(剑桥大学天文研究所基夫利宇宙学研究所)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:
Abstract:Activation steering has gained popularity as a lightweight inference-time debiasing technique for large language models. However, prior work reports that steering vectors generalise poorly, with unintended effects on model performance and limited transfer to new datasets. Our work analyses what the debiasing direction used for activation steering actually encodes, in order to shed light on its inconsistent performance. We study the linear debiasing direction obtained by contrasting the activations of anti-biased and biased prompts, and evaluate it as a steering intervention across bias and general knowledge benchmarks. We find that this direction is dominated by model confidence, pointing from regions of high to low-probability tokens in activation space rather than encoding a meaningful representation of model bias. Steering along it does reduce measured bias, but this is a consequence of reducing model confidence: on QA benchmarks we find that this steering drives the model to abstain from answering, with a side effect of improving fairness metrics. Our experiments show that model confidence is the dominant separating factor between biased and anti-biased prompts in hidden space, indicating that isolating a linear representation of bias which is disentangled from model confidence is difficult and steering-based debiasing results should be interpreted with care. In short, steering appears to reduce bias, not by correcting the model’s underlying preferences, but by making it less confident, even on tasks unrelated to bias.
[NLP-19] DeltaTTT: Layerwise Optimization for Nonlinear Recurrent Memory
【速读】: 该论文旨在解决序列测试时训练(Sequential Test-Time Training, TTT)中非线性记忆网络在连续更新过程中难以有效优化的问题。尽管串行更新机制理论上可通过状态依赖性整合已有知识并更好地融合新信息,但实验发现,对于非线性记忆结构,其性能反而劣于固定基线的并行式TTT方法。核心问题在于:非线性记忆在单次序列遍历中比线性记忆更难优化,导致梯度传播效率下降。为此,论文提出DeltaTTT,其关键创新在于将两层记忆网络的联合内循环优化改为分层学习(layerwise learning),每层基于局部预测目标,通过状态依赖的增量规则(delta rule)进行独立更新。该设计保留了非线性读出能力的同时,支持分块并行计算,显著提升了优化稳定性与模型表现。在DeltaNet和LaCT骨干网络上的实验验证了该方法在语言建模与检索任务上优于传统递归式基线。
链接: https://arxiv.org/abs/2610.08553
作者: Yining Li,Dongchen Han,Jie Fu,Gao Huang
机构: Qiuzhen College, Tsinghua University (清华大学求真书院); LeapLab, Tsinghua University (清华大学跃动实验室); IQuest Research
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注:
Abstract:Sequential test-time training adapts a memory network through successive updates, each computing an inner-loop gradient based on the network’s previous state. Intuitively, this state dependence should allow each update to account for what the memory has already learned and better incorporate new information. However, we find that this expected advantage does not consistently materialize in nonlinear memories: a fixed-base parallel TTT baseline outperforms its serial counterpart. Our exploratory experiments point to a key underlying difficulty: nonlinear memories can be harder to optimize than linear ones within a single pass over the sequence. To alleviate this optimization difficulty, we introduce DeltaTTT, which replaces joint inner-loop optimization of a two-layer memory network with layerwise learning. Each layer is assigned a local prediction target and updated through a state-dependent delta rule. This formulation retains a nonlinear readout while enabling chunkwise parallel computation. Experiments on DeltaNet and LaCT backbones show improvements in language modeling and retrieval over their recurrent baselines.
[NLP-20] How High Is 0.6? Floors Ceilings and Headroom in Interpretability Probing
【速读】: 该论文旨在解决现有模型可解释性分析中探针(probe)评分缺乏固定语义解释的问题。传统方法中,探针的决定系数 $ R^2 $ 无法区分模型对目标变量的预测是源于其内在表征还是仅依赖输入文本中已显式包含的信息,导致结论易受数据分布和输入内容的影响。为此,作者提出引入两个基准参考点:下界(floor),即仅基于简单输入所能预测的性能上限;上界(ceiling),即完整输入所能达到的最佳预测能力。二者之间的差值称为“余量”(headroom),用于衡量模型在超越简单输入信息之外所计算出的额外表示能力。研究证明,当目标变量不再依赖于需推断的隐藏变量或输入不再揭示该变量时,余量将消失。在基于上下文元分析训练的Transformer模型上验证发现,尽管分布偏移导致探针得分下降且预测误差上升12–15倍,但模型仍能恢复相近比例的余量,表明信息损失来自数据而非表征能力退化。进一步应用于真实模型(如scGPT单细胞基础模型)及四项经典大语言模型(LLM)探针研究后发现,部分声称模型编码地理、奥赛罗棋盘状态、事实真相及用户人口统计等信息的结论,在扣除仅由输入文本即可解释的部分后,其有效性显著降低,说明这些“表征”现象很大程度上可由输入文本本身解释。因此,该研究的核心解决方案在于通过构建双参考基准(下界与上界)来量化并校准探针结果,从而更准确地评估模型是否真正具备超越输入显式信息的隐含表征能力。
链接: https://arxiv.org/abs/2610.08544
作者: Pranjal Garg
机构: 未知
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注:
Abstract:Probes are the workhorse of interpretability. If a model’s hidden states predict a variable, the model is said to represent it. But a probe score has no fixed meaning. An R^2 of 0.6 may only reflect what the input already gives away, and the same score can mean different things on different data. We propose reading every probe score against two reference points: a floor, what a declared set of simple inputs already predicts, and a ceiling, what the full input can predict. The gap between them, the headroom, is the range in which a probe can show that a model computes something beyond the simple inputs. We prove that headroom vanishes in two ways: the target stops depending on a hidden variable the model must infer, or the input stops revealing it. We test this on transformers trained for in-context meta-analysis, which must infer the hidden heterogeneity between studies to weight them correctly, and where both reference points are known. Under distribution shift, probe scores fall and prediction error rises 12 – 15\times , yet the model recovers a similar share of the headroom, indicating that the data lost information, not the representation. We then analyze the real models. The single-cell foundation model scGPT encodes biological variability only partially. We also revisit four influential LLM probing studies, which claim that models represent geography, the state of an Othello board, truth, and the demographics of their users. Against a floor computed from the input text alone, some of these claims hold, while others are largely explained by the text itself.
[NLP-21] oward Alignment Scaling Laws: A Framework and First Preregistered Measurements WWW KR
【速读】: 该论文旨在解决生成式 AI(Generative AI)在模型规模扩展过程中对齐(alignment)难度随能力增长而变化的争议性问题,即“对齐是否随着模型变大而变得更简单或更困难”。传统讨论常基于孤立现象进行推断,将对齐视为单一属性,但本文将其重新建模为一组可量化的缩放关系(scaling relations),针对每类风险 $ r $,定义对齐负担 $ B_r(N) = a_r N^{\alpha_r} $,其中 $ N $ 为模型能力代理指标。关键在于通过指数 $ \alpha_r $ 判定对齐趋势:若 $ \alpha_r < 1 $,则缩放有助于对齐;若 $ \alpha_r \approx 1 $,则对齐与能力同步;若 $ \alpha_r > 1 $,则对齐负担累积,形成对齐债务。研究区分了观测对齐、审计对齐与真实对齐,并构建一个简化模型揭示修正行为会消耗能力冗余空间的后果。核心发现包括:长期对齐趋势由所有被修正风险中最大指数决定;当任何风险的 $ \alpha_r > 1 $ 时,维持能力冗余下限需超指数级增长;小模型拟合常低估大规模下的实际指数;且无误报的审计不会低估真实对齐水平。研究提出可预注册的协议并两次应用其简化版本:首次对 Pythia 分类器的对抗训练数据再分析显示,使攻击成功率低于 10% 所需计算量以 $ N^{0.60} $ 增长;第二次在 Qwen2.5 系列(0.5B–72B)上发现诚实性($ \alpha = -0.05 )和陈述倾向( \alpha = 0.48 )均呈有益缩放,而奉承行为( \alpha = 0.89 $)及植入后门则未明确,后者在盲安全训练下于五分之四规模仍存活。研究发布四款浏览器游戏以直观展示这些缩放规律,但不声称当前前沿模型处于何种对齐范式。
链接: https://arxiv.org/abs/2610.08540
作者: Jeremy Canale
机构: 未知
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
备注: 34 pages, 24 figures, 8 tables. Games: this https URL . Preregistrations: this https URL , this https URL , this https URL
Abstract:Whether alignment gets easier or harder as models grow is often argued from isolated findings, as if alignment were one property. We treat it as a family of measurable scaling relations: for each risk category r, the alignment burden needed to hold a fixed safety target is modeled as B_r(N)=a_rN^alpha_r, with N a capability proxy; against a budget proportional to N, scaling helps if alpha_r1, keeps pace if alpha_r~1, and accumulates alignment debt if alpha_r1. We give three operationalizations of burden and distinguish observed, audited and true alignment. A toy model, in which corrections consume capability headroom, makes the consequences explicit. We prove that the largest exponent among corrected risks, not an average, sets the long-run regime; that above 1 any policy holding headroom above a floor must grow super-exponentially; that, for burdens that are positive mixtures of power laws, fits on small models underestimate large-scale exponents; and that an audit that uncovers hidden failures without false positives never underestimates true alignment. We propose a pre-registrable protocol and apply reduced versions of it twice. A preregistered reanalysis of public adversarial-training data for Pythia classifiers finds that the compute needed to bring attack success under 10% grows as N^0.60. A preregistered pilot on Qwen2.5 0.5B-72B finds exponents of -0.05 for truthfulness and 0.48 for stated dispositions (both scaling helps under its reduced rule, though local slopes approach 1 at the top; replicated on Qwen3 0.6B-14B), while sycophancy (0.89, or 0.83 with two seeds added at 72B) and a planted backdoor are undetermined: the backdoor is removed quickly when its trigger is known but survives blind safety training at four of five sizes. We release four browser games that play these laws (this http URL). We make no claim about which regime holds for current frontier models.
[NLP-22] Wiki-Talkie: Multilingual Benchmarking of Persona-Based Agents on Real-World Discussions
【速读】: 该论文旨在解决大语言模型(LLM)在社会环境中作为自主代理时,其行为模拟缺乏真实用户人格基础的问题。现有数据集多依赖虚构人格且仅覆盖少数语言,难以支撑跨文化、跨语言的人类互动行为真实性评估。为此,研究提出Wiki-Talkie——一个涵盖五个语言(德语、英语、西班牙语、法语、意大利语)的多语言真实对话数据集,源自维基百科讨论页,覆盖日耳曼语族与罗曼语族,其用户人格信息基于真实用户社区构建,包含社会人口学特征、自我描述及行为导向的交互属性。关键解决方案在于通过真实世界对话与可验证的用户画像相结合,实现对代理在下一轮回复生成任务中的人格化行为建模。实验表明,基于用户历史评论所体现的互动行为模式,显著优于显式人格信息;同时,模型系统性低估负面或极端情绪,过度生成引用与建议,表现出对宜人性和积极情绪的偏差,且这些偏差在不同语言间保持稳健,仅存在微小跨语言差异。
链接: https://arxiv.org/abs/2610.08513
作者: Dennis Fucci,Andrea Bacciu,Dong Liu,Weronika Łajewska,Saab Mansour
机构: Amazon(亚马逊)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注:
Abstract:LLMs are increasingly deployed as autonomous agents in social environments, making it critical to study their ability to faithfully simulate human interactions. Central to this is grounding agents in realistic user personas, yet existing datasets rely on fictional personas and are limited to a handful of languages, lacking the empirical grounding necessary to evaluate behavioral fidelity across diverse populations. We introduce Wiki-Talkie, a multilingual dataset of real-world conversations from Wikipedia Talk pages across five languages spanning two language families: Germanic (German, English) and Romance (Spanish, French, Italian), paired with personas derived from real user communities and encompassing sociodemographic attributes, self-descriptions, and behaviorally grounded interaction traits. Using Wiki-Talkie, we evaluate agent interactional behavior on a next-turn generation task across various persona conditioning strategies. Our evaluation assesses whether agents collectively reproduce the distributional behavioral patterns observed in human discussions. Results show that user’s comment history exemplifying interaction behavior consistently outperforms explicit persona information. In addition, models systematically underproduce negative or extreme sentiments, while over producing references and suggestions, revealing biases toward agreeableness and positivity. Crucially, these patterns hold robustly across languages, with small cross-lingual differences.
[NLP-23] Language-model ratings of depression reflect the rater more than the patient
【速读】: 该论文旨在解决抑郁症缺乏客观生物标志物检测手段,而现有基于语言模型的评估方法在临床应用中存在评分者间一致性不足的问题。尽管生成式 AI(Generative AI)具备持续、一致评估的潜力,但研究发现不同语言模型在对同一患者访谈内容进行症状评分时仍存在显著分歧。其解决方案的关键在于通过系统性预注册实验设计,评估11种开放源代码语言模型在不同提示策略与评分标准下的表现,并揭示模型选择可解释30.0%的症状总分方差,而个体稳定差异仅占10.5%。研究发现,即使两个性能相当的随机抽取模型(AUC=0.70),平均仍有40%的参与者在筛查决策上产生分歧;且过度高估倾向主导了被标记为阳性的人数,而即便在能力相近的情况下,约五分之一的参与者仍会得到不同结论。后续通过40个标注样本的探索性校准,将分类准确率从约60%提升至75%,并使分歧率减半,但仍无法完全消除个体层面的判断差异。这表明,尽管校准能显著缓解评分者的依赖性,但要实现对个体的一致性判断仍面临根本性挑战。
链接: https://arxiv.org/abs/2610.08501
作者: Baihan Lin
机构: Icahn School of Medicine at Mount Sinai (伊坎医学院在西奈山); James J. Peters VA Medical Center (詹姆斯·J·彼得斯退伍军人医疗中心); Harvard University (哈佛大学)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Neurons and Cognition (q-bio.NC)
备注:
Abstract:Depression has no diagnostic blood test. Language models promise tireless, consistent assessment, but can accurate raters disagree about individuals? We pre-registered 880 language-model raters, crossing 11 open models with prompting and scoring choices, and applied them to 189 interviews against the eight-item Patient Health Questionnaire. Model choice explained 30.0% of summed-symptom score variance, stable participant differences 10.5%. Two randomly drawn raters with area under the receiver operating characteristic curve (AUC) = 0.70 disagreed on screening decisions for 40% of participants, on average. Average over-rating governed how many were flagged, yet equal-capacity raters chose differently for about one participant in five. A locked analysis of 86 new interviews reproduced the main pre-registered findings. Exploratory recalibration with 40 labelled participants raised accuracy from about 60% to 75% and halved disagreement, leaving one participant in five decided differently. Calibration repaired much of the rater dependence without securing agreement about individuals.
[NLP-24] Rethinking Cross-Tokenizer On-Policy Distillation: From Alignment Coverag e to Supervision Reliability
【速读】: 该论文旨在解决生成式模型中师生(teacher-student)对齐时因分词器差异导致的词汇与序列层面不一致问题,进而影响学生模型通过教师反馈进行自我蒸馏(On-Policy Distillation, OPD)的学习效果。其核心挑战在于:当教师与学生使用不同分词器时,如何在存在显著词汇不匹配的情况下实现有效的监督信号传递。解决方案的关键在于重新评估“对齐覆盖范围”与“监督可靠性”之间的权衡——研究发现,尽管存在较大的词汇差异,严格的一一对应(strict 1:1)对齐已覆盖绝大多数学生生成的token,且在这些对齐位置上,共享词汇的预测概率质量保持较高一致性。进一步地,仅在严格对齐位置上选取学生侧概率最高的前16个共享词汇进行反向KL散度优化,即可达到与全共享词汇蒸馏相当的性能,优于跨分词器基线方法。而引入对非对齐片段(mismatch groups)的均方误差(MSE)监督虽实现了全局覆盖,却因引入弱对齐或冲突信号导致准确率下降。梯度分析显示,包含非严格对齐区域的训练中,跨度梯度与严格对齐梯度方向一致性差甚至为负,且其幅值相对增大,揭示了额外监督可能损害学习效率的内在机制。因此,论文提出应从追求广覆盖对齐转向优先保障严格对齐位置上的可靠监督,强调“紧凑但高可靠性”的监督策略在实际应用中更具有效性。
链接: https://arxiv.org/abs/2610.08448
作者: Bingxi Hou,Guochao Jiang,Guofeng Quan,Weiqing Li,Wenfeng Feng,Guohua Liu,Yuewei Zhang
机构: Alibaba Cloud Computing(阿里云计算)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注:
Abstract:On-Policy Distillation (OPD) trains a student on its own generations using teacher feedback. With different tokenizers, comparing teacher and student predictions requires alignment at both sequence and vocabulary levels. In this paper, we examine whether expanding this alignment coverage improves learning. Across three heterogeneous teacher–student pairs on mathematical reasoning and code generation, strict 1:1 groups already cover most student-generated tokens despite substantial vocabulary mismatch. On responses sampled from the students before distillation, the shared vocabulary retains nearly all teacher and student probability mass at strictly aligned positions on average. Restricting reverse KL to a student-selected top-16 subset of the shared vocabulary at each strict position achieves accuracy comparable to full shared-vocabulary OPD, outperforming the evaluated cross-tokenizer baselines. Adding mean squared error supervision on span log-probabilities in mismatch groups gives complete supervision coverage, yet reduces accuracy. At checkpoints from training with only the strict loss, the span gradients show weak or negative directional agreement with the strict gradients and grow in magnitude relative to them. These diagnostics may help explain the accuracy drop from adding span supervision. Our findings motivate a shift from maximizing alignment coverage to prioritizing supervision reliability: compact supervision at strict positions can be more effective than broader coverage that introduces weakly aligned or conflicting training signals.
[NLP-25] Knowing When Not to Answer: Cross-Domain and Multi-Turn Generalization of Latent Underspecification Signals
【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在问答与对话中存在“不可回答性”(unanswerability)的问题,即模型在信息不足时仍强行生成答案,或在对话尚未充分展开时过早回应。其核心挑战在于:不同形式的不可回答性是否共享同一表征空间,以及该信号在多轮对话中的实际效用如何。为此,研究提出一个带有回合标签的多轮对话基准数据集(423次对话,1,661个标注的对话状态),并构建了一个模拟用户评估框架,通过澄清性问题来测试模型对不可回答性的检测能力。实验表明,针对“缺失信息”类不可回答性的探测器在共享语义基础的数据集间具有强泛化能力(如数学题中AUROC达0.77–0.97,SQuAD 2.0-MuSiQue为0.77–0.90),但对“认识论层面的已知未知”(epistemic “known-unknowns”)的探测则表现不佳,且该差异受词汇控制和层位置、坐标系变化影响,尚无定论。单轮探测器无法零样本识别对话何时具备可回答性,而基于结构的探测器虽能恢复该能力,但性能仅相当于词袋分类器。关键突破在于引入校准探测器的门控机制,在无需微调的前提下显著提升了对未充分指定回合的精准识别,其表现接近使用真实标签的门控模型(差距仅0.08),然而在最终任务表现上,该机制仍未能稳定超越原始生成或提示聚合策略。研究揭示:当前模型性能瓶颈主要源于对澄清信息的利用方式,而非不可回答性的检测能力本身。
链接: https://arxiv.org/abs/2610.08413
作者: Jerzy Kamiński,Ilya Galyukshev,Artem Kuznetsov,Danil Fedorov,Kirill Redko,Sergey Chuprin,Aidar Shumbalov,Stanislav Chumakov,Anna Kalyuzhnaya
机构: ITMO University(圣彼得堡国立信息技术机械与光学大学)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: 15 pages, 3 figures, 10 tables. Under review
Abstract:Large language models routinely answer questions that cannot be answered from the information given, and in dialogue they answer before enough has been said. Unanswerability is linearly decodable from hidden states, but it is unclear which of its forms share a representation and whether the signal is useful in dialogue. We contribute a turn-labeled multi-turn benchmark (423 conversations, 1,661 labeled turn-states) and an evaluation harness with a simulated user who answers clarifying questions, and use them with six datasets and six open-weight LLMs to test how far probes for unanswerability carry. Probes transfer robustly between datasets that share a ground of unanswerability: missing information in math (AUROC 0.77-0.97) and in a passage (SQuAD 2.0-MuSiQue, 0.77-0.90). Probes for epistemic “known-unknowns” transfer poorly to math, but this separation weakens under lexical controls and changes with layer and coordinate system, so it remains unresolved. Single-turn probes fail zero-shot to detect when a conversation becomes answerable; in-structure probes recover it, but no better than a bag-of-words classifier. A gate on the calibrated probe, with no model fine-tuning, fires on underspecified turns far more precisely than chance, and its end-task success comes within 0.08 of a gate given the true labels. Yet across four models it does not reliably beat vanilla generation or prompted consolidation. The remaining gap lies mostly in how models use a clarification, not in detection.
[NLP-26] Foresight-over-Graph: Reasoning Beyond Local Horizons for Knowledge Base Question Answering NEURIPS2026
【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在知识密集型任务中频繁产生幻觉(hallucinations)的问题,尤其针对基于知识图谱(Knowledge Graphs, KGs)的问答系统中因证据检索策略局限导致的推理失败问题。现有方法通常采用逐跳贪婪或束搜索(beam-style)的局部剪枝策略进行证据检索,这种短视(myopic)的决策机制会过早舍弃看似弱相关但后续可能至关重要的路径,从而导致关键推理分支丢失且难以恢复。为此,本文提出了一种名为“图上远见”(Foresight-over-Graph, FoG)的前瞻性证据检索框架,其核心在于通过“由远及近”(far-to-near)的反馈机制引导路径探索,并维护一个紧凑的记忆子图以支持持续的上下文感知式推理。该方案实现了对全局推理路径的更优规划,显著提升了答案准确性,在多个标准知识库问答(KBQA)基准测试中达到领先性能,尤其在CWQ数据集上取得了16.58%的Hit率提升,同时有效降低了大语言模型调用次数与令牌消耗。
链接: https://arxiv.org/abs/2610.08388
作者: Yang Hong,Yajun Yang,Xin Wang,Liping Jing,Qinghua Hu
机构: Tianjin University (天津大学); Beijing Jiaotong University (北京交通大学)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: 25 pages, 10 figures. Accepted at NeurIPS 2026
Abstract:Large language models (LLMs) have demonstrated strong capabilities in question answering, yet they still frequently suffer from hallucinations on knowledge-intensive tasks. Knowledge graphs (KGs) provide LLMs with structured, interpretable, and updatable factual grounding, making them a promising external knowledge source for reliable reasoning. However, existing LLM-guided graph reasoning methods typically rely on hop-wise greedy or beam-style pruning during evidence retrieval. Such local decision processes are inherently myopic: evidence that appears weak near the source may become crucial only after deeper graph context is explored, causing answer-critical branches to be discarded prematurely and making the reasoning chain difficult to recover. To address this limitation, we propose Foresight-over-Graph (FoG), a foresight-aware evidence retrieval framework for knowledge base question answering (KBQA). FoG iteratively constructs a question-relevant evidence subgraph and uses far-to-near feedback to guide path exploration, and maintains a compact memory subgraph to support continued exploration. Extensive experiments on widely used KBQA benchmarks demonstrate that FoG achieves state-of-the-art performance, with a particularly large improvement of 16.58% in Hit on CWQ, while also reducing LLM calls and token usage. Our code is available at this https URL .
[NLP-27] CoDe-LoRA: Mitigating the Orthogonality Dilemma in Continual Learning of LLM s via Knowledge Consolidation and Decoupling EMNLP2026
【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在持续学习(Continual Learning, CL)过程中因参数更新导致的灾难性遗忘问题。现有方法如基于正交投影的低秩适应(O-LoRA)虽能通过严格的几何约束隔离任务参数以缓解遗忘,但其固有的“正交性困境”限制了语义相关任务间共享表征的迁移与累积,影响模型长期性能。为此,论文提出一种无需回放(replay-free)的新方法——知识整合与解耦低秩适应(Consolidation and Decoupling LoRA, CoDe-LoRA),其核心在于将学习过程解耦为“通用知识整合”与“任务特异性知识解耦”两个阶段。关键创新在于引入自适应零空间投影机制与语义路由策略,动态平衡知识积累与任务适应之间的权衡,从而在不依赖历史数据回放的前提下实现更优的持续学习表现。实验结果表明,CoDe-LoRA 在四种模型架构及三个持续学习基准上均取得最优平均准确率。
链接: https://arxiv.org/abs/2610.08312
作者: Maoqi Liu,Quan Fang,Yufei He
机构: Beijing University of Posts and Telecommunications(北京邮电大学); National University of Singapore(新加坡国立大学)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
备注: Accepted to EMNLP 2026 (Main Conference)
Abstract:Continual learning (CL) is essential for Large Language Models (LLMs) to sequentially adapt to evolving tasks. To mitigate catastrophic forgetting, recent advances implement low-rank adaptation with orthogonal projections (e.g., O-LoRA) to isolate task parameters. However, we reveal that such strict geometric constraints trigger an “Orthogonality Dilemma”: rigid parameter isolation impedes the transfer and accumulation of shared representations across semantically related tasks. In this work, we propose a new replay-free method, called Consolidation and Decoupling LoRA (CoDe-LoRA), for CL of LLMs. CoDe-LoRA disentangles the learning process into Consolidating Universal Knowledge and Decoupling Task-Specific Knowledge. To achieve this, CoDe-LoRA leverages an adaptive null space projection mechanism and semantic routing to balance knowledge accumulation with task-specific adaptation. Experimental results across four backbones and three CL benchmarks show that CoDe-LoRA achieves the best average accuracy. Our code is available at this https URL.
[NLP-28] Language Unalignability: Why Some Concepts Resist Cross-Cultural Benchmark Evaluation
【速读】: 该论文旨在解决当前多语言大模型(Multilingual Large Language Models, LLMs)评估中隐含的“翻译同构性假设”(Translation-Isomorphism Assumption, TIA)所引发的根本性问题:即默认不同语言间的语义结构在形式与内容上具有可映射性且信息无损。研究指出,这一假设在理论上对特定类型概念——如语用标记、敬语系统及历时分层词汇——是不成立的,因其在语言类型学上具有不可通约性。其核心解决方案在于提出“α-不可对齐性”(α-unalignability)的理论框架,通过使用上下文嵌入的“使用云”(usage-cloud)模型将概念表示为点集,并定义无法同时保持词义忠实性(中心点对应)与结构忠实性(局部邻域拓扑)的映射为不可对齐。研究从行为、机制和诊断三个层面提供证据:行为层面显示FLORES-200的翻译失败由语言家族与资源类别决定而非书写系统,且LOBSTER推理得分随语言家族差异而变化;机制层面揭示在雅美语(Yami)九模型案例中存在“表征-干预差距”(Representation-Intervention Gap, RIG),即模型激活模式虽显现出雅美语与其他低资源南岛语的规律性关联,但语言特异性神经元的干预效果并不优于随机掩码,表明该规律虽可感知却不可操作;最终,研究构建了多维诊断谱系,包括循环一致性、语用负载分歧、流形曲率不匹配与RIG,主张将文化能力简化为单一标量指标会诱导“概率扁平化”现象,强调识别不可对齐类是实现尊重而非消解文化差异的AI的前提。由此,研究提出多语言对齐并非单一明确目标,而是多个相互冲突的投影集合。
链接: https://arxiv.org/abs/2610.08303
作者: Shu-Kai Hsieh,Da-Chen Lian
机构: National Taiwan University(台湾大学)
类目: Computation and Language (cs.CL)
备注: Position paper. 32 pages (10 pages main text), 6 figures, 12 tables
Abstract:Current evaluation of multilingual Large Language Models (LLMs) rests on an implicit Translation-Isomorphism Assumption (TIA): that semantic structures across languages are congruent and mutually mappable without loss of information. We argue that this assumption is not merely violated in practice, but ill-posed in principle for a typologically identifiable class of concepts, including pragmatic markers, honorifics, and diachronically stratified terms. We formalize this failure using a usage-cloud framework, representing concepts as point sets of contextualized embeddings. We define \alpha -unalignability as the impossibility of any mapping that simultaneously preserves lexical faithfulness (centroid correspondence) and structural faithfulness (local neighborhood topology). We provide three layers of evidence. Behaviorally, we show that FLORES-200 translation failures are predicted by language family and resource class but not by script, and that LOBSTER reasoning scores vary by family. Mechanistically, we report a Representation-Intervention Gap (RIG) in a nine-model case study on Yami: the models’ activations encode a regularity along which Yami groups with other low-resource and Austronesian languages, yet interventions on language-specific neurons show no demonstrated advantage over random masks: the regularity is visible but not usable by this intervention. Finally, we operationalize these findings into a multidimensional diagnostic profile: Cycle-Consistency, Pragmatic-Load Disagreement, Manifold-Curvature Mismatch, and RIG. We argue that collapsing cultural competence into a single scalar incentivizes “probabilistic flattening,” and that recognizing the unalignable class is a precondition for AI that respects, rather than erases, cultural divergence. This suggests that multilingual alignment is not a single well-defined objective, but a set of mutually incompatible projections.
[NLP-29] Memory Depth and Reconstructed Context Width: A Controlled Evaluation of Hierarchical Retrieval NEURIPS2026
【速读】: 该论文旨在解决大语言模型(LLM)在长期对话记忆(long-term conversational memory)中如何有效组织与利用历史信息的问题。现有架构通常通过主题与事件分组、构建层次结构和图谱,并基于因果与时间关系连接事实,但其性能受制于记忆结构的深度与上下文宽度之间的权衡。本研究的关键发现是:在保持核心记忆预算(core budget)不变的情况下,增加上下文宽度(从1,024到4,096 tokens)可显著提升准确率(10.11–17.98个百分点),而增加记忆结构深度(D1–D4)并未带来单调性能增益;当上下文扩展至8–16K以上时,生产环境(Production)性能趋于饱和,但每正确回答所需的令牌数持续上升,而理想化全档案条件(Oracle)下仍能维持高质量表现(68–71K tokens)。因此,研究提出应优先探索大规模、连贯的上下文块(large, coherent context blocks)而非逐级加深的记忆结构,这为未来长时记忆系统设计提供了关键方向。
链接: https://arxiv.org/abs/2610.08300
作者: Michael Andreev
机构: Independent Researcher(独立研究员)
类目: Computation and Language (cs.CL)
备注: 4 pages, 1 figure. Accepted at the PALM Workshop at NeurIPS 2026
Abstract:Long-term conversational memory is becoming an integral component of modern LLM systems. Proposed architectures group records by topics and events, construct hierarchies and graphs, and connect facts through causal and temporal relations. We experimentally study the interaction between two memory parameters: structural depth and the width of context supplied to the answer model. Using EverMemBench, we evaluate depths D1-D4, core budgets of 1,024/2,048/4,096 tokens, and additional Production and Oracle conditions up to the full archive. Increasing width from 1K to 4K improves Accuracy by 10.11-17.98 percentage points, whereas increasing depth provides no monotonic gain. Beyond 8-16K, Production performance reaches a plateau while tokens per correct answer continue to increase; Oracle preserves quality on full archives of 68-71K tokens. These results motivate further investigation of large, coherent context blocks instead of progressively deeper memory structures.
[NLP-30] STRUCTURALCOST: A controlled reading time dataset for modeling human sentence processing difficulty EMNLP2026
【速读】: 该论文旨在解决语言模型在处理长距离主谓依赖关系时,对人类阅读加工成本(processing cost)建模不足的问题。其核心问题是:当前主流语言模型虽能捕捉人类语言理解中的预测性成分,却未能充分反映工作记忆在整合复杂句法结构时所承担的真实认知负荷。为此,研究提出STRUCTURALCOST数据集,包含475名参与者和40,800个观测值,通过自适应阅读范式精确分离出长距离主谓依赖解析的加工成本。关键解决方案在于,利用大规模实证数据验证并量化人类阅读时间随句法嵌套深度增加而延长的现象,揭示这一现象主要由句法嵌套(syntactic embedding)而非线性距离驱动。实验表明,尽管各类模型(包括n-gram、状态空间模型SSMs及Transformer)部分再现了人类的渐进式难度分布,但普遍低估了真实的人类整合成本,且该差距在不同架构与规模下持续存在。这表明现有模型缺乏对工作记忆负担的完整表征,提示未来需以认知可解释性为导向,提升语言模型在模拟人类语言处理机制方面的有效性。
链接: https://arxiv.org/abs/2610.08208
作者: Nina Nusbaumer,Iria de-Dios-Flores,Corentin Bel,Christophe Pallier,Guillaume Wisniewski,Benoît Crabbé
机构: LLF, CNRS, Université Paris Cité(法国国家科学研究中心, 巴黎-夏尔·戴高乐大学); COLT, Universitat Pompeu Fabra(庞佩乌·法布拉大学); Unicog, Neurospin, CEA(法国原子能委员会神经成像中心); LNC2, ENS-PSL(巴黎文理研究大学); INSERM, CNRS(法国国家健康与医学研究院, 法国国家科学研究中心)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: Will be published at EMNLP 2026
Abstract:We introduce STRUCTURALCOST, a self-paced reading dataset of 475 participants and 40,800 observations isolating the processing cost of long-distance subject-verb dependency resolution. We replicate a low-powered psycholinguistic finding at NLP scale, namely that human reading times at the main verb increase with dependency length, driven by syntactic embedding beyond linear distance. Different language models – spanning n-gram models, SSMs, and transformers – partially mirror this graded difficulty profile, yet underestimate the integration cost humans incur, with a gap that persists across architectures and model sizes. This suggests these models capture the predictive component of human processing but not the full integration cost that working memory imposes. STRUCTURALCOST provides data needed to drive progress toward evaluating the cognitive plausibility of language models.
[NLP-31] Align Then Correct: Training-Free Two-Stage Low-Rank Compensation for Extremely Quantized Large Language Models
【速读】: 该论文旨在解决在极端权重量化(aggressive weight quantization)条件下,模型精度显著下降的问题,尤其是针对现有低秩量化误差补偿(LQEC)方法因双重简化假设而导致补偿能力受限的瓶颈。现有方法存在两个关键局限:其一,采用对称校准方式,在相同激活下评估全精度与补偿后权重,导致补偿目标本身具有高秩特性,固定秩预算难以充分捕捉有效信息;其二,仅最小化损失函数的二阶项,忽略了补偿后模型仍存在显著的一阶梯度方向(即非平稳性),而该方向无法通过传统重建目标吸收。为此,本文提出一种两阶段闭式框架以消除上述限制:第一阶段基于Fisher加权的非对称目标,将每一层输出对齐至全精度模型,从而将秩预算集中于可压缩的低秩目标;第二阶段在补偿模型上重新统计特征,并施加秩约束的自然梯度步长,有效吸收残余一阶信号。所有适配器均由一次截断奇异值分解(truncated SVD)生成,反向传播仅用于统计收集。实验表明,在2比特量化下,本方法将Qwen3-8B和Qwen3-4B在WikiText-2上的困惑度分别从12.43、21.11降至10.26、13.22;在保留的C4数据集上,分别恢复了51%和84%的FP16差距(相较最强基线提升20个百分点),并在七项零样本任务平均性能、更高比特率及不同量化器设置下均表现出一致优势。核心解决方案在于通过两阶段非对称优化与自然梯度引导的秩约束更新,实现更精准、更全面的量化误差补偿。
链接: https://arxiv.org/abs/2610.08164
作者: Seobin Song,Geonho Lee,Janghwan Lee,Jungwook Choi
机构: Hanyang University (汉阳大学)
类目: Machine Learning (cs.LG); Computation and Language (cs.CL)
备注: 17 pages, 5 figures
Abstract:Low-rank quantization error compensation (LQEC) recovers the accuracy lost under aggressive weight quantization by attaching a closed-form rank- r adapter beside each frozen quantized weight, without any training. We show that existing compensators are limited by two shared simplifications. They calibrate symmetrically, evaluating the full-precision and compensated weights on the same activation, which yields a compensation target that is inherently high-rank – so a fixed rank budget captures only a small fraction of it. And they minimize only the second-order term of the loss, although the compensated model is not stationary: a first-order descent direction larger than the applied compensation itself remains in every layer, and no reconstruction objective can absorb it. We propose a two-stage closed-form framework that removes both simplifications. Stage 1 aligns each layer’s output with the full-precision model under a Fisher-weighted asymmetric objective, concentrating the rank budget on a rank-compressible target. Stage 2 re-measures statistics on the compensated model and applies a rank-constrained natural-gradient step that absorbs the remaining first-order signal. Every adapter is the result of a single truncated SVD; backward passes serve only to collect statistics. At 2 bits under QuIP#, our method reduces WikiText-2 perplexity from 12.43 to 10.26 on Qwen3-8B and from 21.11 to 13.22 on Qwen3-4B. On the held-out C4 corpus, it recovers 51% and 84% of the gap to FP16, versus 31% and 63% for the strongest baseline, with consistent gains in the seven-task zero-shot average, at higher bit-widths, and under a distinct quantizer.
[NLP-32] he Failure Is in the Readout: Fine-Grained Emotion Recognition Benchmarks Measure Elicitation Not Perception
【速读】: 该论文旨在解决细粒度情绪识别(fine-grained emotion recognition)在实际应用中对真实人脸数据的依赖所引发的隐私与数据保护问题。现有方法通常需要真实面部图像以支持心理治疗工具与社交机器人,但由此带来的数据泄露风险限制了其广泛应用。为此,论文提出使用生成式人脸图像(EmoNet-Face-HQ)作为替代数据源,构建了一个包含40个类别的细粒度情绪标注体系,远超传统6–8种基本情绪的分类范畴。关键挑战在于:当前主流视觉-语言模型(VLMs)在该基准下表现不佳,因此研究者提出了专用微调模型Empathic-Insight-Face(EIF, Small/Large)。然而,本文的核心贡献在于揭示:若不采用传统的生成式回答方式,而是直接从模型输出的logits中读取每个情绪类别对应的概率值(即通过二元查询方式逐类判断),则无需微调的现成VLMs可达到甚至超越EIF的性能。实验表明,在专家评估一致性(κ_w = 0.468)的基础上,11个开源VLMs在验证机制下均显著提升,κ_w达0.507–0.586,其中三个模型显著优于EIF(κ_w = 0.551/0.534)。进一步分析显示,性能提升源于对概率分布的连续性利用,而非简单的“是/否”问答;将相同概率阈值化为二元决策会损失平均增益的142%,且一致性下降至κ_w = 0.254–0.423。此外,该效应在真实照片数据集(FACES)上亦部分显现,说明其有效性不仅限于合成数据。因此,解决方案的关键在于改变答案读取范式——从生成式输出转向基于原始logits的分级概率解析,从而在无需额外训练的前提下实现更优的情绪识别性能。
链接: https://arxiv.org/abs/2610.08162
作者: Tobias Hallmen,Fabian Deuser,Robin-Nico Kampa,Norbert Oswald,Elisabeth André
机构: University of Augsburg (奥格斯堡大学); University of the Bundeswehr Munich (慕尼黑国防大学)
类目: Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL)
备注: Preprint. 19 pages, 6 figures
Abstract:Fine-grained emotion recognition supports therapy tools and social robots, but it needs facial data, which raises privacy and data-protection concerns. EmoNet-Face-HQ answers that with generated portraits, expert-rated over a 40 -category taxonomy far finer than the usual six to eight basic emotions. Under the protocol it ships with, vision-language models (VLMs) score poorly on that taxonomy, and the benchmark concludes that a dedicated fine-tuned model is necessary: Empathic-Insight-Face (EIF; Small/Large). We show that off-the-shelf VLMs match or beat that fine-tuned model when the answer is not generated but read from the logits, as one binary query per category. We keep the benchmark’s images, taxonomy and ratings, and change only how the answer is read. Experts agree at \kappa_w = 0.468 on the five categories they measure most reliably. Generatively, no interval among eleven open-weight VLMs lies entirely above that anchor ( \kappa_w=0.268 - 0.486 ). Under verification all eleven clear it, each of them significantly better at \kappa_w=0.507 - 0.586 . Three also significantly beat EIF sitting at \kappa_w = 0.551 (Small; 0.534 Large). The gain comes from the graded probability and not from asking a yes/no question: as a control, thresholding those same probabilities to yes/no costs 142% of the average gains and drops binarization below generative elicitation to \kappa_w=0.254 - 0.423 . A replication on real photographs (FACES) is weaker and mixed: of the ten models that pass a validity gate, six gain, three are neutral to positive and one is negative, so the effect is not confined to synthetic data.
[NLP-33] Symphony for Text Generation: Benchmarking Clinical Note Generation
【速读】: 该论文旨在解决环境式临床记录系统(ambient documentation systems)在实际应用中对临床病历质量影响缺乏系统评估的问题。现有研究未能充分揭示此类系统在真实临床场景下的表现差异,尤其在多语言环境下缺乏标准化的评测基准与数据支持。为此,论文提出MedConv——一个包含英语、丹麦语和德语共300例临床会诊的多语言数据集,并结合环境临床智能基准(ACI-BENCH)对Corti这一基于API的临床AI平台与两款主流通用大模型驱动的环境语音记录软件进行对比评估。其解决方案的关键在于构建了一个受控的临床评估框架,融合蕴含关系度量(entailment metrics)与基于大语言模型(LLM)判断的成对比较方法,在借鉴PDSQI-9量表的八个维度基础上实现对病历质量的精细化评估。实验结果表明,Corti在文本生成质量上达到或优于当前领先的商业级环境记录工具,且其可配置的API架构具备针对特定文档使用场景灵活调整质量维度的能力,从而为未来环境式临床记录系统的可复现性比较提供了方法论基础与公开数据支持。
链接: https://arxiv.org/abs/2610.08161
作者: Daniel Varab,Victor Petrén Bach Hansen,Asbjørn W. Helge,Kevin Pelgrims,Mathias Baltzersen,Adrian Young-San Roessler,Vanessa Klungtvedt,Maximilian Brand,Lasse Krogsbøll,Henrik Cullen,Lars Maaløe
机构: 未知
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注:
Abstract:Ambient documentation systems are rapidly gaining adoption, yet their impact on clinical note quality remains poorly characterized. We introduce MedConv, a multilingual dataset of 300 clinical encounters in English, Danish, and German, and use it alongside the Ambient Clinical Intelligence benchmark (ACI-BENCH) to compare Corti, a clinical AI platform, with two leading, accessible ambient scribe software applications built on general-purpose AI. We present a controlled clinical evaluation framework that combines entailment metrics with LLM-judged pairwise comparisons across eight dimensions adopted from PDSQI-9. Results show that Corti’s API-based text-generation infrastructure is on par with or outperforms leading commercial scribes. We further show that Corti’s configurable API provides the flexibility necessary to fine-tune quality dimensions for specific documentation use cases. We present the evaluation methodology and release a dataset to support future reproducible comparison of ambient documentation systems.
[NLP-34] Making COMET Comparable Across Scripts: Diagnosis and Correction of Tokeniser-Induced Script Bias in Indic MT Evaluation
【速读】: 该论文旨在解决生成式 AI(Generative AI)在机器翻译质量评估中因书写系统差异导致的评分偏差问题,即“书写系统不变性”(Script Invariance)假设的失效。现有评估指标COMET将翻译质量以单一数值表示,并在不同书写系统的语言间进行直接比较,但研究发现,书写系统身份可解释原始COMET分数22.9%的方差,且在五种印地语系语言中,与人工标注者的一致性均下降。其关键原因在于分词器(tokeniser)对不同书写系统的不一致性处理:一方面导致不同书写系统下的得分分布不可比,另一方面削弱了同一书写系统内翻译排序的准确性。该偏差由两个独立故障构成——前者可通过COMET-QN实现精确校正,其通过将每个(语言, 书写系统)组合的得分分布映射至统一参考分布,使跨语言评分具备可比性并严格保持语言内部排序;后者则源于编码器本身对书写系统敏感性的固有缺陷,无法通过后处理修复。研究进一步提出三个无标签诊断方法用于量化此偏差,结合基于对齐特征的回归器可恢复17.1%的丢失敏感度。因此,作者建议发布标准化后的得分、三项诊断指标及所用分词器信息,以便读者区分评分中究竟反映翻译质量还是书写系统影响。
链接: https://arxiv.org/abs/2610.08159
作者: G. L. John Salvin(1),Swapnil Hingmire(1) ((1) Indian Institute of Technology Palakkad)
机构: Mehta Family School of Data Science and Artificial Intelligence; Indian Institute of Technology Palakkad (印度理工学院帕拉克卡德), Kerala (喀拉拉邦), India (印度)
类目: Computation and Language (cs.CL); Machine Learning (cs.LG)
备注: 18 pages, 2 figures. Camera-ready version, accepted at WMT 2026. Code and data: this https URL
Abstract:COMET reports translation quality as a single number, and that number is routinely compared across target languages written in different scripts. Such a comparison assumes Script Invariance: the score should not depend on the writing system that carries the target. We test it on IndicMT Eval by re-encoding the target into Latin script, which changes orthographic form while holding content and human ratings fixed. Script identity then accounts for 22.9% of native-script COMET variance, and agreement with annotators falls in all five languages studied. We trace the effect to the tokeniser and measure it with three label-free diagnostics. The bias is two faults, not one. Scores from different scripts occupy incompatible ranges, and within a single script the metric orders translations less accurately. No order-preserving transform of the score can repair the second fault. The first is removed exactly by COMET-QN, which maps the score distribution of each (language, script) pair onto a shared reference. Pooled agreement with annotators rises from 0.300 to 0.399, which is what makes scores from different scripts safe to place on one axis, and every within-language ordering is provably preserved. A regressor over parity features recovers a further 17.1% of the lost sensitivity. The remainder belongs to the encoder, and no post-processing can reach it. We therefore recommend publishing the normalised score, the three diagnostics, and the identity of the tokeniser they were computed against, so that a reader can tell how much of a score reflects translation quality and how much reflects the writing system.
[NLP-35] Penalty-Framed No-Valid-Option MCQA: Analyzing LLM Abstention under Invalid Choices AACL
【速读】: 该论文旨在解决大语言模型在真实应用场景中面临的一种关键可靠性问题:当提供的多选题选项集(option set)中不存在正确答案(即无有效选项,no-valid-option)时,模型仍可能被迫选择一个错误选项,从而引发下游成本。传统基于答案选择准确率的评估方法无法捕捉这一风险,因此论文提出以“惩罚框架下的无有效选项多选题问答”(penalty-framed no-valid-option MCQA)作为新的评估范式。其解决方案的关键在于引入“正确条件分析”(correct-conditioned analysis),即仅在模型原本正确回答的样本上评估其放弃回答(ABSTAIN)的可靠性,并通过惩罚机制强制模型在无有效选项时选择不作答。实验结果表明,即使在明确提示和惩罚机制下,高准确率模型仍会在部分本应正确回答的实例中产生无效强制选择行为,揭示了标准答案选择准确率无法全面反映模型在不确定或异常输入下的可靠性。
链接: https://arxiv.org/abs/2610.08153
作者: Jinhyeok Kim,Hye-Young Jung
机构: Hanyang University (汉阳大学)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: Accepted to AACL-IJCNLP 2026 Main Conference (Short Paper)
Abstract:Multiple-choice question answering (MCQA) is commonly used to evaluate large language models under the assumption that one of the provided options is correct, typically using answer-selection accuracy. However, in real deployments, users or retrieval systems may provide invalid option sets in which none of the listed choices is correct, and selecting one of them may incur downstream cost. We study this setting as penalty-framed no-valid-option MCQA. Using the mathematics subset of MMLU-Pro, we remove the labeled correct option, allow models to either choose a remaining option or output ABSTAIN, and penalize invalid forced-choice responses. We further introduce correct-conditioned analysis, evaluating abstention only on instances that the model originally answered correctly. Experiments show that high MCQA accuracy does not fully guarantee abstention reliability: even under explicit no-valid-option-aware instructions and penalty-based scoring, models still produce invalid forced-choice responses for a subset of originally correct instances. These results show that penalty-framed no-valid-option MCQA reveals an aspect of model reliability not captured by standard answer-selection accuracy.
[NLP-36] Natural Language Questions as an Interface for Knowledge Graphs: QRAKEN Graph Distillation and Semantic Self-Healing
【速读】: 该论文旨在解决大语言模型(LLM)在自然语言到SPARQL(Text-to-SPARQL)转换任务中,面对不熟悉的知识图谱时因依赖模式(schema)预期而非实际数据分布而生成语义错误查询的问题。其核心挑战在于如何在不依赖显式本体(ontology)或训练数据的前提下,使生成过程更忠实于具体知识图谱的实证数据结构。解决方案的关键是提出QRAKEN——一种无需训练、本体无关的神经符号流水线,通过离线构建一个紧凑的图谱实证描述(TTQL),包含多跳路径模式、条件频率及路径条件下的字面量示例,以及类-属性共现矩阵。在线推理阶段,TTQL作为外部约束引导LLM生成,并结合确定性语法、词汇与数据模型校验实现迭代优化。实验表明,在CK25基准上,基于GPT-4.1 mini和GPT-5.4的QRAKEN分别取得0.643±0.026和0.652±0.012的严格F1,相较最强基线提升30%和32%,且显著优于使用相同基础模型家族的系统。消融研究揭示TTQL中的实证路径模式是性能主导因素(相比仅依赖形状的基线提升0.31严格F1),而迭代精炼机制提供了低成本的安全保障。相较于自动推导的SHACL规则,TTQL实现64%的严格F1提升,验证了基于实证数据模式的建模价值。此外,仅需两台本地部署的35B参数4位量化开源模型即可达到最优结果,且TTQL的优势在不同设置下保持稳定,显示出其高效性和可扩展性。然而,当前结果仍局限于单一小型基准,对超开放跨领域图谱进行统一的TTQL注入仍是主要局限。
链接: https://arxiv.org/abs/2610.08095
作者: Remo Grillo,Lukas Klic,Giovanni Colavizza
机构: 未知
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注:
Abstract:Natural-language access to RDF knowledge graphs is a core Semantic Web ambition. Large language models (LLMs) have advanced Text-to-SPARQL, yet on unfamiliar graphs they often generate valid queries that misrepresent the populated data model. QRAKEN is a training-free, ontology-agnostic neurosymbolic pipeline grounding generation in empirical graph evidence rather than schema expectations. An offline distiller produces TTQL, a compact description of populated multi-hop patterns, conditional frequencies and path-conditioned literal examples, plus a class-property co-occurrence matrix. Online, TTQL guides the LLM, while deterministic syntax, vocabulary and data-model checks provide diagnostics for iterative refinement. On CK25 (First International Text2SPARQL Challenge), under matched-condition recomputation on a QLever snapshot, QRAKEN achieves strict F1 of 0.643 \pm 0.026 with GPT-4.1 mini and 0.652 \pm 0.012 with GPT-5.4: relative gains of 30% and 32% over the strongest recomputed participant, outperforming systems using the same base model family. Ablations identify TTQL patterns as the dominant driver (+0.31 strict F1 over a shape-only baseline); the refinement loop provides a cheap safety net, rejecting triple patterns unsupported by the co-occurrence matrix. Compared with auto-derived SHACL, TTQL yields 64% higher strict F1, supporting the value of empirical patterns beyond schema exposure. With two local 35B 4-bit open-weight models at zero marginal cost, the same pipeline matches the strongest recomputed participant, and TTQL advantages over shape-only and SHACL baselines persist. Results on a single, relatively small benchmark provide an initial empirical signal; monolithic TTQL injection on very open cross-domain graphs remains the main limitation.
[NLP-37] SAGE: Semantic Anchor-Guided Evolution for Grounded Medical QA Data Synthesis EMNLP2026
【速读】: 该论文旨在解决临床任务(如医学问答,Medical QA)中因高质量、专家标注训练数据稀缺而导致的模型可靠性瓶颈问题。在资源受限的临床环境中,这一挑战尤为突出,主要受制于严格的隐私保护要求,以及无法有效使用大规模开源语料库或专有云API。为应对上述限制,本文提出SAGE(Semantic Anchor-Guided Evolution)——一种新型的数据合成框架,其核心在于通过轻量级、公开可用的医学术语体系(如MeSH)作为语义锚点(semantic anchors),引入结构化先验知识以指导并约束生成过程。SAGE的关键创新在于采用原子式(基于单一概念)与关联式(基于关系)合成策略的迭代交替机制,从极少量种子数据出发逐步自举生成高质量医学训练数据。该方法无需依赖大规模医学文档集合或外部API,实现了本地化、可部署的数据生成方案。大量实验结果表明,使用SAGE合成数据微调的模型在多个医学问答基准上均显著优于基于自生成或传统文档范式的模型,验证了其在提升数据效率和资源利用率方面的有效性。
链接: https://arxiv.org/abs/2610.08093
作者: Chuan Li,Chengyu Wang,Cen Chen,Ye Lyu,Mingyuan Fan,Ming Gao
机构: East China Normal University (华东师范大学); Alibaba Group (阿里巴巴集团)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: EMNLP 2026
Abstract:Developing reliable models for clinical tasks, such as Medical Question Answering (QA), is severely constrained by the limited availability of high-quality, expert-annotated training data. This challenge is exacerbated by stringent privacy requirements and the impracticality of utilizing large open-source corpora or proprietary cloud APIs within resource-limited clinical settings. To address these obstacles, we introduce SAGE (\textitSemantic Anchor-Guided Evolution), a novel data synthesis framework that enables small, locally deployed models to generate high-quality medical training data. SAGE leverages lightweight, publicly available taxonomies such as MeSH as semantic anchors, imposing a structured prior to effectively guide and ground the data generation process. At its core, SAGE iteratively interleaves atomic (individual concept-based) and associative (relation-based) synthesis, bootstrapping training data from minimal seeds. This approach eliminates the need for large collections of medical documents or reliance on external APIs, providing a practical solution for on-premises data creation. Extensive experiments across multiple medical question-answering benchmarks demonstrate that models fine-tuned with SAGE-synthesized data consistently outperform those trained using self-derived or conventional document-based paradigms, highlighting tangible improvements in data efficiency and resource utilization for medical LLM development. Code is available at this https URL.
[NLP-38] DirectSpeech2LLM : A Simple End-to-End Framework to Mitigate Prompt Overfitting in Speech-LLM s
【速读】: 该论文旨在解决语音大模型(Speech-LLMs)在指令学习中出现的提示过拟合(prompt overfitting)问题,即模型仅在自动语音识别(ASR)指令上训练后,难以泛化至未见过的任务(如语音翻译和情感识别),且仍表现出典型的ASR系统行为。其解决方案的关键在于提出一种简洁的端到端框架DirectSpeech2LLM:通过在冻结的大语言模型(LLM)嵌入矩阵上计算基于距离的连字符时间分类(CTC)损失,并利用贪婪解码得到的CTC标签,分别生成在几何结构与时间维度上对齐的语音嵌入,作为输入馈送至LLM。该方法仅使用960小时的LibriSpeech ASR数据进行训练,在已知任务(ASR)上表现优于级联系统,并在零样本条件下成功泛化至语音翻译与情感识别两项未见任务,性能接近级联系统的上限。研究还发现,几何对齐强度的作用小于先前预期,因为所提出的改进型CTC损失已能提供充分的隐式几何引导,无需显式的回归损失。实验结果在两种不同架构的LLM中均具一致性,且随着训练数据量与模型容量增加而持续提升。
链接: https://arxiv.org/abs/2610.08085
作者: Hemant Yadav,Sunayana Sitaram,Roger Zimmermann,Rajiv Ratn Shah
机构: IIIT Delhi(印度信息科学技术研究所); Microsoft Research(微软研究院); National University of Singapore(新加坡国立大学)
类目: Computation and Language (cs.CL)
备注:
Abstract:Speech-LLMs often exhibit prompt overfitting, where models solely trained on automatic speech recognition (ASR) instruction fail to generalize to new instructions such as speech translation and continue to behave primarily as ASR system. We propose DirectSpeech2LLM, a simple end-to-end framework that preserves the instruction-following ability of the LLM on unseen tasks when conditioned on speech. It computes distance-based CTC loss over the frozen LLM embedding matrix and uses greedy CTC labels to derive geometrically and temporally aligned speech embeddings respectively as an input to the LLM. Trained solely on 960 hours of LibriSpeech ASR data, DirectSpeech2LLM outperforms the cascaded system on ASR (seen task) and generalizes zero-shot to speech translation and emotion recognition (two unseen tasks), closely matching the cascaded system upper bound on these two new instructions despite seeing neither during training. We also find that geometric alignment strength plays a smaller role than previously assumed, as our modified CTC loss is shown to provide sufficient implicit geometric grounding without requiring an explicit regression loss. Results are consistent across two LLM families and scale with both more training data and model capacity.
[NLP-39] POLAR: Ontology-Guided Risk Prevention for Tool-Calling LLM Agents AACL
【速读】: 该论文旨在解决大语言模型(Large Language Model, LLM)工具调用代理在动态环境中执行操作时面临的操作风险不可预知且事后防护机制滞后的问题。现有安全机制多依赖于错误发生后的响应,或通过链式思维微调、将自然语言约束编译为运行时检查等方式进行事前防范,但缺乏可审计的结构化判断依据。为此,论文提出POLAR框架——一种面向小型工具调用代理的守卫机制,其核心创新在于构建一个分层的两层本体结构(two-layer ontology),基于候选逆向操作序列对每项动作评估可逆性得分(reversibility score),并在执行前根据阈值剪枝高风险操作。该方法实现了可审计的结构化验证,并量化了其在任务效用与安全性之间的权衡。实验在τ²-bench上针对六种代理模型进行评估,结果显示,在四个代理中航空领域任务的平均任务奖励提升0.11至0.18分,但整体仅在十八个模型-领域组合中的八个实现收益,零售领域及更强代理常出现性能退化。研究强调,奖励并非防止伤害的直接度量,凸显了安全机制设计中需兼顾任务有效性与风险控制的复杂性。
链接: https://arxiv.org/abs/2610.08082
作者: Yunju Kang,Seonghyeon Cho,Irene Li,Yeo-Chan Yoon,Chanjun Park
机构: Soongsil University (松岛大学); Tokyo University (东京大学); Jeju National University (济州国立大学)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: Accepted Findings of AACL-IJCNLP 2026
Abstract:LLM tool-use agents operate in dynamic environments where many actions carry operational risk. However, most safety mechanisms react only after errors manifest. Existing pre-emptive approaches either fine-tune the agent on chain-of-thought deliberation or compile natural-language guardrails into runtime checks, but they do so without exposing a structural, auditable verdict. We propose POLAR, a guardrail framework for small tool-calling agents that assesses reversibility through a structured two-layer ontology. POLAR assigns each action a graded reversibility score by deriving a candidate inverse sequence; calls failing a threshold are pruned before execution. Evaluated on \tau^2 -bench across six agent models, POLAR improves mean task reward by 0.11 to 0.18 points on airline for four of six agents, but only eight of eighteen model–domain cells improve overall; retail and stronger agents often regress. POLAR provides an auditable structural check and characterizes its task-utility trade-offs. Reward is not a direct measure of prevented harm.
[NLP-40] Language Carries the Experts Impression: Instrument-Anchored LLM Judges Transfer Counseling-Quality Assessment and Beat In-Domain Training
【速读】: 该论文旨在解决二元咨询对话中通信质量自动评估因数据稀缺而受限的问题,具体表现为专家评分语料库规模小且构建成本高昂。其核心解决方案在于跨领域迁移学习:通过在三个德语模拟咨询语料库(涵盖一般医疗与家校沟通场景,共195个专家评分会话)之间进行跨域转移学习,发现基于其他领域数据训练的模型性能显著优于本域训练,留一域外推在嵌套斯皮尔曼相关系数上达到ρ = 0.54,高于本域内训练的≤0.48,且在匹配训练集规模后仍保持+0.12的显著优势,表明该效果并非单纯由数据量驱动。关键突破在于利用小型开放权重大语言模型(small open-weight LLMs)从双人对话转录文本中提取会话层面的结构化构念得分,这些构念主要源自专家评分工具的设计维度;实验表明,基于评分工具衍生的构念组合可将单个评估者的表现从0.32提升至0.41,三类模型集成后达到0.51的语言仅特征水平,而加入非言语-互动性特征块进一步带来+0.03的增益(但在当前样本量下难以与噪声分离)。此外,研究还量化了录音设备配置的影响:某一语料库因缺失每说话人独立音频,导致16%的语音分离片段标注错误,修复后可提升性能+0.07。综上,在实际可获取的语料规模下,专家对整体沟通质量的判断主要由言谈内容决定,且更依赖于其他通信系统生成的数据而非自身积累的专有数据。
链接: https://arxiv.org/abs/2610.08055
作者: Tobias Hallmen,Elisabeth André
机构: University of Augsburg(奥格斯堡大学)
类目: Computation and Language (cs.CL); Machine Learning (cs.LG)
备注: Preprint. 25 pages, 2 figures
Abstract:Automatic assessment of communication quality in dyadic counseling conversations is bottlenecked by data: expert-rated corpora are small and expensive to grow. We study cross-domain transfer of expert overall-impression prediction across three German corpora of simulated counseling (two general-practice medical, one school-related parent-teacher; n=195 expert-rated sessions, one corpus after scale equating). Training on the other domains beats training in-domain: leave-one-domain-out transfer reaches nested Spearman \rho = 0.54 against \le 0.48 within the target domain, a paired session-level gap of +0.15 that holds at +0.12 when the training-set sizes are matched, so it is not simply data volume. The decisive features are session-level construct scores from small open-weight LLMs reading the two-speaker transcript, with the constructs largely derived from the experts’ rating instruments: the instrument-derived battery lifts a single judge from 0.32 to 0.41 over generic dialogue qualities, judges from three model families ensemble to 0.51 language-only, and a nonverbal-dyadic block adds +0.03 more, not separable from noise at this sample size. We also price the recording setup: one corpus lost its per-speaker audio, 16% of its diarised segments carry the wrong speaker, and repair is worth +0.07 there. At practically attainable corpus sizes, the expert’s overall impression is carried by what is said, and by other communication programs’ data more than by one’s own.
[NLP-41] DAEDALUS: Bootstrapping Agent Memory from Self-Generated Tasks
【速读】: 该论文旨在解决大语言模型代理(LLM agents)在新环境中缺乏操作知识,导致无法可靠执行任务的问题。具体而言,由于代理在面对未知环境时无法依赖已有经验,往往重复相同错误,造成任务失败率上升和执行轨迹过长。现有方法通常依赖人工编写的指导规则或基于训练任务与“权威验证器”构建的程序性记忆,但这些方法均需预先掌握环境信息。为克服这一局限,本文提出DAEDALUS,一种无需现成任务或权威验证器即可自动生成可复用代理记忆的自举方法。其核心在于通过两个协作代理——探索者(explorer)与求解者(solver)——实现闭环学习:探索者生成具有挑战性且可解的任务,求解者尝试求解;当求解失败时,系统提取失败情境下的启发式策略,并仅在求解者后续多次成功应用该策略后才将其纳入记忆库。同时,求解者的反馈用于优化探索者生成任务的难度。经验证,该方法在AppWorld、τ²-bench和AutomationBench等多个基准上,相较于无记忆基线,平均成功率提升最高达15.9个百分点,pass^5指标提升2.2倍,性能接近使用训练任务的方法,且推理成本更低。关键发现表明,求解者的行为轨迹是生成有效启发式的核心信息源,而对早期发现进行因子化处理可显著提高探索效率。此外,研究还发现DAEDALUS生成的任务可作为模型性能评估的代理基准。
链接: https://arxiv.org/abs/2610.08048
作者: Antoine Edy,Max Conti,Victor Xing,Marc-Antoine Allard,Nawfal Benhamdane,Gautier Viaud
机构: Illuin Technology(伊鲁因科技)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
备注: 9 pages (31 including Appendix), 8 figures (11 including Appendix). We release the code and artifacts, including generation and inference traces, at this https URL
Abstract:LLM agents often lack the operational knowledge to act reliably in new environments, as they must discover specific tool behaviors or environment conventions on their own. Without memory of past attempts, they repeat the same mistakes across tasks, leading to more task failures and longer trajectories. To address this, agentic systems typically rely on human-written guidelines or on procedural memory built from training tasks and an oracle verifier, both of which require prior knowledge of the environment. We present DAEDALUS, a method for bootstrapping reusable agent memory from self-generated practice without existing tasks or oracle verifiers. DAEDALUS pairs two agents: an explorer that interacts with the environment to generate challenging yet solvable tasks, and a solver that attempts them. A heuristic is derived from each solver failure and accepted only after the solver repeatedly succeeds with that heuristic in context. These outcomes also provide feedback for the explorer to refine the difficulty of future tasks. Accepted heuristics are then consolidated into a memory bank for test-time use. Across AppWorld, \tau^2 -bench, and AutomationBench, DAEDALUS improves mean success rates by up to 15.9 points and pass^5 by up to 2.2x over a no-memory baseline, and is competitive with methods using training tasks, at a lower inference cost than most. We show that performance gains already emerge with a small exploration budget, and that its heuristics also benefit agents from other model families. Our ablations further reveal that solver traces provide the key information needed to derive effective heuristics, while factorizing early discoveries makes exploration more cost-efficient. Beyond memory construction, we find that the tasks generated by DAEDALUS can serve as a proxy for benchmark tasks when ranking models by performance. Code and artifacts: this http URL.
[NLP-42] Are Language Models Script-Aware? AACL
【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)和小语言模型(Small Language Models, SLMs)在生成文本时出现的非目标语言或脚本(off-target generation)问题,尤其聚焦于脚本知识(script knowledge)这一长期被忽视的关键维度。现有研究多关注语言选择,但未充分探讨模型对书写符号系统的认知能力——即用户能否正确识别模型输出的图形符号。为探究模型是否具备脚本知识,研究设计了两项互补实验:一是评估模型能否根据输入语境自适应地调整输出脚本;二是检验模型是否能遵循明确指令生成指定脚本的文本。结果表明,所有测试模型均表现出显著的脚本知识,其拉丁字母脚本保真度超过98%,且能高频遵循脚本指令。尽管如此,大型语言模型在处理非标准脚本组合时表现优于小型模型,显示出规模效应带来的优势。因此,该研究的核心解决方案在于揭示并验证了模型在跨脚本生成任务中的隐式知识表征,强调了脚本知识在确保生成内容可读性和可控性中的关键作用。
链接: https://arxiv.org/abs/2610.08037
作者: David Kletz,Sandra Mitrović,Ljiljana Dolamić,Fabio Rinaldi
机构: SUPSI, IDSIA, Switzerland(瑞士洛迦诺应用科学与艺术大学, 人工智能系统研究所); armasuisse, Science Technology, Switzerland(瑞士联邦武器局, 科学与技术部)
类目: Computation and Language (cs.CL)
备注: Accepted to AACL-IJCNLP 2026
Abstract:Language models frequently generate outputs in unintended languages or scripts, a phenomenon known as off-target generation. While existing research has focused on language selection, the dimension of script knowledge remains understudied: before any linguistic understanding can occur, users must recognize the graphic symbols in a model’s response. We investigate whether Small and Large Language Models (SLMs and LLMs) possess script knowledge by testing them on multi-scriptic languages. Through two complementary experiments, we evaluate whether models (1) adapt their output script to match the input, and (2) follow explicit instructions to generate text in a specified script. The models we tested demonstrate substantial script knowledge: they all achieve a near-perfect Latin script fidelity (more than 98%) and follow script instructions with high frequency. Nevertheless, we notice differences between LLMs and SLMs, with higher scores for LLMs including for non-standard script combinations.
[NLP-43] he Labeling Problem in Hallucination Detection Benchmarks: An Empirical Evaluation NEURIPS2026
【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在幻觉检测评估中存在的一种方法论模糊性问题,即自动标注策略在实践中可能混淆“参考答案一致性”(reference faithfulness)与“事实正确性”(factual correctness)这两个核心评价标准。具体而言,现有评估框架通常依赖于开放域问答(open-domain question answering, QA)数据集中的参考答案对生成答案进行比对,但自动化标注工具往往基于参考答案的词汇或语义相似性来判断是否为幻觉,这实质上偏向于衡量答案是否与参考答案一致,而非是否在事实上准确无误。为探究此偏差,研究者基于三个常用QA数据集和三种生成模型,构建了900组由人工标注的事实正确性标签的问答对,并系统评估了词法相似度指标、基于自然语言推理(NLI)的基准方法以及七种大语言模型(LLM)判别器在不同提示(prompt)变体下的表现。结果表明,各类自动化标注方法之间存在显著分歧,且与人工标注一致性较低;多数方法表现出明显的定向错误偏差,尤其在高假阳性率方面。更重要的是,将原本面向“参考一致性”的提示替换为明确指向“事实正确性”的提示后,大多数判别器与人工标注的吻合度显著提升,同时假阳性率下降,说明自动化幻觉标签的高度依赖于目标评价标准的明确定义。因此,研究强调:标注来源的选择应作为基准设计的核心环节,必须显式声明、验证并严格匹配评估目标,以确保幻觉检测评估的有效性与可解释性。
链接: https://arxiv.org/abs/2610.08026
作者: Jorma Valjakka,Juhani Kivimäki,Juha Mylläri,Jukka K. Nurminen
机构: University of Helsinki(赫尔辛基大学)
类目: Computation and Language (cs.CL)
备注: 27 pages. Accepted at the NeurIPS 2026 Evaluations Datasets Track. Data: this https URL . Code: this https URL
Abstract:In recent years, several methods for detecting when large language models (LLMs) hallucinate have been developed. These methods are often benchmarked with open-domain question answering (QA) datasets containing questions and corresponding short reference answers. First, an LLM is used to generate answers to questions within the QA dataset. Then, some automated labeling strategy is used to label these answers as hallucinated or not by comparing them with the reference answers in the dataset. This evaluation setting creates a methodological ambiguity between two criteria: reference faithfulness (whether the answer is fully supported by the reference) and factual correctness (whether the answer is free from contradictions and factually false specific claims). In practice, automated labelers may apply the former criterion even when the intended target is the latter. We study this potential criterion mismatch using 900 human-labeled question-answer pairs spanning three commonly used QA datasets and three generator models, with labels targeting answer-level factual correctness. We evaluate lexical similarity metrics, a reference-entailment NLI baseline, and seven LLM judges under controlled prompt variants as automated labelers. Our experiments reveal substantial disagreement both among automated labeling strategies and between these labels and human annotations. Many strategies also exhibit strong directional error biases, and for most judge-generator pairs, replacing a faithfulness-oriented prompt with a factual-correctness prompt improves agreement with human annotations and reduces false-positive dominance, indicating that automated hallucination labels depend strongly on how the target criterion is specified. Label-source choice should therefore be considered a fundamental part of benchmark design and made explicit, validated, and matched with the benchmark goal.
[NLP-44] Structured but Silent: Probing Capability Requirements in LLM Hidden States AACL
【速读】: 该论文旨在解决大语言模型(LLM)在执行可靠工具使用时,如何有效识别用户查询中隐含的工具能力需求这一关键问题。传统方法依赖于机械式触发或API描述匹配,但缺乏对用户意图背后能力要求的深层理解。为此,论文提出一种名为TACIT的框架,将外部能力需求沿“源(Source)、变换(Transformation)、世界效应(World Effect)”三个基本维度进行分解,定义出八类结构化的工具能力类别。其解决方案的关键在于:通过在预生成阶段的模型隐藏状态上训练线性探测器,发现这些能力需求在隐藏表示中可被高精度线性解码;然而,当要求模型以自然语言显式分类相同查询时,其表现显著下降,揭示出“表示到表述之间的鸿沟”——即模型虽在内部表征中“知晓”所需能力结构,却无法稳定地将其转化为语言表达,这种现象被称为“结构化但沉默”(structured but silent)。这表明,尽管隐藏状态蕴含丰富的能力信息,但当前模型在将隐性知识转化为显性语言判断方面存在根本性局限。
链接: https://arxiv.org/abs/2610.08018
作者: Kyojun Choo,Minsoo Song,Yunju Kang,Chanjun Park
机构: Soongsil University (中央大学)
类目: Computation and Language (cs.CL)
备注: Accepted to AACL-IJCNLP 2026 Findings
Abstract:Reliable tool use requires more than triggering a mechanism or matching a query to an API description. Before selecting a specific tool, an agent must first infer the capability requirements implied by the user query. In this paper, we investigate whether these query-side capability requirements are linearly decodable from LLM hidden representations prior to generation, and how this hidden-state accessibility compares with explicit verbal classification. We introduce TACIT, a framework that decomposes external requirements along three fundamental axes: Source, Transformation, and World Effect, defining eight structurally distinct capability classes. Using 1,600 balanced training queries from benchmarks, synthetic examples, and new domain scenarios, we train linear probes on pre-generation hidden states from four open-weight LLM families. Our empirical results demonstrate that fine-grained capability structures are linearly decodable with high accuracy across all models. Crucially, however, we expose a representation-to-verbalization gap: these same models are significantly less reliable when asked to explicitly classify the same queries in natural language. This disconnect indicates that information about required external capabilities is linearly accessible in LLM hidden representations but not reliably expressed, a phenomenon we define as “structured but silent.”
[NLP-45] A Broader Look at Model Merging: Rethinking Implicit Regularization Induced by Task Arithmetic
【速读】: 该论文旨在解决现有模型融合(model merging)方法中因隐式正则化导致性能受限的问题。具体而言,主流方法通过在额外数据集上优化任务特定权重更新的线性组合系数来构建多任务模型,但这一过程将候选模型限制在由各任务权重更新张成的子空间内,形成了隐式正则化。研究发现,这种约束实际上限制了模型表达能力,而突破该子空间限制、直接优化预训练模型权重可显著提升融合模型在多种架构、领域甚至极端数据稀缺场景(每类仅一个样本)下的性能。关键解决方案在于摒弃依赖系数搜索的传统范式,转而探索更广阔的权重空间,通过直接优化或采用其他策略释放潜在的更优多任务权重。研究表明,最优权重往往存在于子空间之外,且可通过多种方法有效寻得,从而推动对模型融合流程的根本性重构,并重新审视任务算术(task arithmetic)带来的隐式正则化影响。
链接: https://arxiv.org/abs/2610.07990
作者: Sin-Han Yang,Shih-Cheng Huang,Chieh-Yen Lin,Yun-Nung Chen,Shao-Hua Sun,Hung-yi Lee
机构: Appier AI Research(Appier人工智能研究院); National Taiwan University(国立台湾大学)
类目: Machine Learning (cs.LG); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)
备注: Preprint
Abstract:Model merging aims to build a multi-task model cheaply by combining the weights of individual task-specific models. To perform well across multiple tasks, most existing merging methods use an additional dataset to find the coefficients for the best linear combination of task-specific weight updates. However, we identify an implicit regularization in this standard practice: searching over coefficients restricts the candidate models to a subspace spanned by task-specific weight updates. In this work, we investigate whether this regularization is actually useful. Surprisingly, empirical results show that optimizing merged-model weights without this regularization significantly boosts the performance of common merging methods across multiple architectures, domains, and even in an extremely data-limited scenario where only one instance is available per class. Moreover, directly optimizing the pretrained model weights even outperforms some existing merging methods. Analysis shows that better multi-task weights exist outside the subspace and can be found using multiple methods. We study different strategies for using the additional dataset, discussing their practical use and implications for model merging. Overall, this work calls for revisiting the existing model-merging pipeline, motivating a broader exploration of the weight space and a reconsideration of the implicit regularization induced by task arithmetic.
[NLP-46] VisionWeave: Weaving Elastic Visual Representations as a Native Capability of MLLM s
【速读】: 该论文旨在解决多模态大语言模型(Multimodal Large Language Models, MLLMs)在视觉理解中因统一编码输入为密集固定尺寸图像块(patch tokens)而导致的计算与存储开销过高的问题。现有方法如下采样或固定比例的令牌剪枝难以实现内容自适应的粒度调整,限制了其在复杂视觉内容中的表达能力与任务泛化性能。为此,本文提出“弹性视觉表征编织”(elastic visual representation weaving)这一核心理念,即让模型端到端地学习在何处以及以何种粒度分配视觉表征。其解决方案的关键在于引入VisionWeave框架,包含两个核心组件:一是门控空间池化器(gated spatial pooler),在共享的多分辨率位置编码(MRoPE)坐标系中同时生成粗粒度与细粒度视觉表征;二是粒度路由模块(granularity router),通过内容感知机制动态决定各区域应采用的表征粒度。仅通过自蒸馏训练,该方法在Qwen3.5-4B上验证了可行性,并扩展至Qwen3.8-27B模型(超过3万小时A100 GPU算力),实现了平均43.0%的令牌节省率,同时保持98.9%的原始性能,在八项基准测试中显著优于固定50%剪枝率的基线方法(后者仅保留88%性能)。此外,部署于SGLang服务引擎时,该方法实现2.3倍吞吐量提升,平均首字延迟(TTFT)降低54.4%,平均输出延迟(TPOT)降低60.6%,充分证明其在效率与质量之间具备鲁棒的权衡能力,为下一代高效、自适应的多模态模型提供了关键范式。
链接: https://arxiv.org/abs/2610.07987
作者: Yuan Feng,Qize Yang,Ruizhe Chen,Sibo Song,Haolin He,Muzhi Zhu,Zihan Liu,Yunfei Chu,Xize Cheng,Yuxuan Wang,Jin Xu,Xike Xie
机构: Alibaba Token Hub, Alibaba Group; University of Science and Technology of China; The Chinese University of Hong Kong; Zhejiang University
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注:
Abstract:Multimodal large language models have become the dominant paradigm for visual understanding, but incur substantial costs by encoding inputs into dense, fixed-size patch tokens. However, visual information is unevenly distributed: some regions require fine-grained detail, while others admit compact representations. Downsampling sacrifices this detail, while existing token pruning and adaptive approaches remain limited in content-adaptive granularity, task generalization, and integration with modern MLLMs and serving infrastructure. Overcoming these limitations calls for foundation models that learn, end to end, where-and at what granularity-to allocate visual representations, a native capability we term elastic visual representation weaving. We introduce VisionWeave, establishing this capability in frontier-level MLLMs through large-scale training. It combines two components: a gated spatial pooler constructs coarse-grained representations alongside native fine-grained representations within a shared MRoPE coordinate, while a granularity router learns their content-adaptive allocation. Through self-distillation alone, we validate this capability on Qwen3.5-4B and scale to Qwen3.8-27B with over 30K A100 GPU-hours. Based on Qwen3.8-27B, VisionWeave adaptively adjusts token savings to visual content, saving 43.0% tokens on average while retaining 98.9% native performance across eight benchmarks, versus only 88% performance preserved for token pruning baselines with a fixed 50% savings target. Extensive evaluations confirm robust efficiency-quality trade-offs across diverse tasks, resolutions and video frames. When deployed on SGLang serving engine, our method achieves a 2.3x throughput gain while reducing mean TTFT by 54.4% and mean TPOT by 60.6%. Together, we believe these results position elastic visual weaving as a promising capability for next-generation multimodal models.
[NLP-47] Confidence Reasoning Graphs: Structured Confidence Estimation for LLM Agents
【速读】: 该论文旨在解决在因果性领域(consequential domain)中使用大语言模型代理(LLM agent)时,如何在不依赖模型内部信号或训练数据的前提下,对代理执行结果进行可信度校准(calibrated confidence)的问题。由于任务成功证据分散于异构且相互依赖的多步轨迹中,且实际部署中存在前沿大模型访问受限、代理运行成本高以及训练数据易过期等挑战,传统方法难以有效估计代理的成功概率。为此,论文提出置信度推理图(Confidence Reasoning Graphs, CRGs),其核心在于:以任务完成为初始假设,将该主张分解为一系列基于轨迹证据的上下文化子命题(sub-claims),分别估算每个终端子命题的置信度,并通过聚合机制生成整体置信度估计。与传统的整体性判断方式不同,CRG通过分层式、可解释的推理结构,在无需特权模型访问或训练数据的情况下实现了更优的置信度校准和风险敏感决策能力。实验表明,相较于语义化、采样法及白盒代理基线,CRGs在多个基准测试、模型架构和代理框架下均表现更优;进一步分析揭示,仅依赖校准误差评估可能具有误导性——某白盒基线虽看似校准良好,但实际判别能力接近随机水平。消融实验确认,CRG性能提升主要归因于子命题层面的置信度估计与聚合机制,而非图结构本身。此外,CRG能够显式暴露支撑置信度判断的子命题及其对应轨迹证据,支持在决策时刻进行可审计性分析。
链接: https://arxiv.org/abs/2610.07948
作者: Brendan King,Farima Fatahi Bayat,Jean-Flavien Bussotti,Pouya Pezeshkpour,Estevam Hruschka
机构: University of California, Santa Cruz(加州大学圣克鲁斯分校); Megagon Labs
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: 34 pages, 6 figures, 11 tables
Abstract:When using an LLM agent in a consequential domain, making an informed decision about whether to trust its output or intervene requires calibrated confidence in the agent’s success. Confidence estimation for agents is difficult because evidence about success is distributed across heterogeneous, interdependent steps of an agent’s trajectory. Practical agentic deployments introduce further challenges: frontier LLMs often provide limited access to internal signals, agent roll-outs are costly, and training data may be unavailable or quickly become outdated. To address these challenges, we introduce Confidence Reasoning Graphs (CRGs), an inference-time framework that estimates the probability an agent accomplished its task from a single trajectory, without privileged model access or training data. Rather than compressing an execution into a single holistic judgment, a CRG begins with the claim that the agent accomplished its task, decomposes it into contextualized sub-claims grounded in trajectory evidence, estimates confidence for each terminal claim, and finally aggregates these into an overall confidence estimate. Across three agentic benchmarks, three backbone models, and three agent frameworks, CRGs yield better-calibrated confidence and stronger risk-aware decision making than verbalized, sampling-based, and white-box surrogate baselines. We further find that calibration error alone can be misleading: a white-box surrogate baseline appears well calibrated while providing near-chance discrimination. Ablations attribute CRG’s improvements to claim-level confidence estimation and aggregation rather than graph construction alone. Finally, a CRG exposes the claims and trajectory evidence underlying each confidence estimate, enabling it to be audited at decision time.
[NLP-48] Hybrid Latent Attention for Looped Language Models
【速读】: 该论文旨在解决循环语言模型(Looped Language Models)在长序列推理时因键值(KV)缓存过大导致的显存占用高、并发处理能力受限及解码速度慢的问题。其核心解决方案是提出混合隐变量注意力(Hybrid Latent Attention, HLA),通过在滑动窗口内保留最近W个令牌的精确键值对,而将更早的令牌以紧凑的隐变量形式编码,并由每个查询直接读取,避免了对旧键值的重建。该方法仅训练新增的隐变量参数,保持预训练权重冻结,在不损失模型性能的前提下,使每令牌缓存大小缩小10.7倍,显著提升单个GPU上的并发序列数(4.0–8.8倍)和解码吞吐量(1K上下文场景下提升2.5倍,16K上下文场景下最高达7.4倍),同时在数学、知识与推理等基准测试中保持超过97%的原始精度,在长达16K令牌的长文本检索任务中表现稳定(96–100%准确率),经监督微调后在竞赛级数学任务上达到与原模型相当的性能。
链接: https://arxiv.org/abs/2610.07940
作者: Yuhan Chen,Siyuan Zhang,Nan Wang,Feiyang Kang,Ruoxi Jia
机构: Virginia Tech(弗吉尼亚理工学院)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注:
Abstract:Looped language models apply the same stack of layers T times to each token, which deepens the model without adding parameters but multiplies its key-value (KV) cache by T. The larger cache limits how many sequences a GPU can decode at once and slows each decoding step, which reads the whole cache. We propose Hybrid Latent Attention (HLA), which keeps exact keys and values within a sliding window of W recent tokens and stores each older token as a compact latent that the query of each loop reads directly, without reconstructing keys and values. We uptrain HLA on Ouro looped models (T=4) with 1.4B and 2.6B parameters, keeping the pretrained weights frozen and training only the added parameters to reproduce the original attention. The cache shrinks by 10.7x per token, fitting 4.0-8.8x as many concurrent sequences per GPU, and decoding throughput improves by 2.5x at 1K-token contexts and by up to 7.4x at 16K. HLA retains over 97% of the original accuracy on math, knowledge and reasoning benchmarks, and 96-100% on long-context retrieval up to 16K tokens. After supervised fine-tuning, it performs on par with the fine-tuned original model on competition-level math.
[NLP-49] Leverag ing a four-quadrant approach for evaluating Redpine Science
【速读】: 该论文旨在解决大模型在获取高质量、可信赖的学术文献信息时面临的检索效率与准确性问题,尤其是在面对复杂科学命题时,传统网络搜索存在相关性不足、信息噪声大及结果不可靠等挑战。其核心解决方案是通过Redpine Science平台,为模型和智能体提供基于同行评审文献的统一访问接口,依托模型上下文协议(Model Context Protocol, MCP)与API实现对权威学术文献的精准检索。该方案的关键在于构建了一个以高质量学术资源为基础、经过专家验证的评估体系,并结合公开基准与专家标注数据集进行双重验证:一方面在公开基准ScholarQA-Bench SciFact上证明使用Redpine Science的代理模型在答案准确率上显著提升(94.4% vs. 87.6%),另一方面在专家验证的问题集上展现出更高的事实陈述覆盖率(80.1% vs. 70.2%),同时在检索召回率(Recall@10)和精确率(Precision@5)方面均优于主流工具(如PubMed)。这表明,通过结构化、可信的学术知识源接入,能够有效提升生成式AI在科学问答任务中的表现,尤其在需要高精度证据支持的场景中具备显著优势。
链接: https://arxiv.org/abs/2610.07937
作者: Filip Dorm,Leonora Vesterbacka
机构: Redpine(瑞德平), Sweden
类目: Computation and Language (cs.CL)
备注:
Abstract:Redpine Science gives models and agents a single access point to a wide range of peer-reviewed literature, queried directly through the Model Context Protocol (MCP) and an API. This report evaluates Redpine Science on two levels: the relevance of the retrieved chunks, and a model’s answer when it has access to Redpine Science compared to web search. Both public and expert-validated benchmarks are used. Public benchmarks are a widely accepted way to test model development and are comparable across labs, but risk saturation and memorization. To address this, we complement them with an expert-validated question set. In total, this report presents four evaluations. On ScholarQABench SciFact, the public answer-quality benchmark reported here, an agent with Redpine Science answers 94.4% of claims correctly against 87.6% with no retrieval. On the expert-validated question set, an agent with Redpine Science states 80.1% of the required claims against 70.2% for an agent restricted to web search. On the 668 queries of a public retrieval benchmark whose gold paper Redpine holds, stripped of any model reasoning, Redpine Science places the correct source paper in its top ten results for 83.1% of queries (Recall@10), against 79.3% for the benchmark’s creator. A blinded expert relevance panel places Redpine Science’s Precision@5 at 75.2% against 39.8% for the PubMed search tool. We release the expert-validated question set and instructions to reproduce every headline result above, at this https URL.
[NLP-50] Pseudowords as probes: Large Language Models show little of the sublexical sensitivity that governs human pseudoword processing
【速读】: 该论文旨在探究大语言模型(Large Language Models, LLMs)是否具备与人类相似的对词素级(sublexical)线索的敏感性,即在处理伪词(pseudoword)时能否像人类一样依赖形式到意义之间的概率映射。研究通过两项意大利语的两择一强迫选择伪词实验,对比了五种LLMs与人类行为基线的表现。结果表明,当真实词选项提供词汇熟悉度线索时,LLMs的表现更接近人类;而在仅包含伪词的条件下,其表现显著低于fastText(一种基于字符n-gram的模型)。此外,能够有效驱动人类与fastText一致性的词素级余弦相似度线索,并未在人类与LLMs之间稳定传递,且推理阶段的标记(token)消耗量也未能一致反映人类加工难度。这表明LLMs未必共享驱动人类伪词加工的词素级认知机制。作者提出分词方式(tokenization)和训练数据覆盖范围(training-data coverage)可能是造成这一差异的关键因素。
链接: https://arxiv.org/abs/2610.07936
作者: Jing Chen,Giulia Loca,Simona Amenta,Marco Marelli
机构: University of Milano-Bicocca(米兰-比科卡大学); Department of Informatics, Systems and Communication – DISCo(信息学、系统与通信系 – DISCo)
类目: Computation and Language (cs.CL)
备注:
Abstract:Systematicity, the probabilistic mapping of form to meaning, permeates language at all levels, and sublexical cues have been shown to govern human pseudoword processing. Yet whether LLMs exhibit comparable sensitivity to these cues remains unclear. We tested five LLMs on two Italian two-alternative forced-choice pseudoword experiments and compared their responses with a human behavioural baseline. LLMs aligned more reliably with humans when real-word options provided a lexical familiarity cue than in the pseudoword-only condition, where they fell substantially below fastText, a character-n-gram model. In addition, the sublexical cosine-similarity cue that reliably drove human–fastText agreement did not consistently transfer to human–LLM alignment, and reasoning-token expenditure bore no consistent relation to human processing difficulty. These findings suggest that LLMs do not necessarily share the sublexical cues that govern human pseudoword processing; we discuss tokenization and training-data coverage as candidate explanations.
[NLP-51] Isotropic Yet Undecodable: The Sequential Content-Sufficiency Gap in Latent-Predictive Text Representations
【速读】: 该论文旨在解决序列内容充分性(sequential content sufficiency)问题,即探究模型表示是否保留了输入中蕴含的有序目标信息。其核心挑战在于区分输入模糊性(input ambiguity)、表示损失(representation loss)与读出不匹配(readout mismatch)三者在信息丢失中的贡献。解决方案的关键在于提出一种基于信息论分解的分析框架,并构建可恢复视图(recoverable views),在其中完美一致性和联合各向同性高斯性与零目标信息共存,从而揭示由确定性规范锚点(deterministic canonical anchors)所施加的理论极限。进一步引入词元对数损失(token log-loss)作为单侧信息损失上界,通过固定惩罚的岭回归分析表明,仅靠秩(rank)无法决定预测风险。基于上述发现,论文提出了CANOPE框架,其核心包括有序潜在画布(ordered latent canvases)、规范词元监督(canonical-token supervision)和几何正则化(geometric regularization)。实验结果表明,在强自然噪声下,尽管潜在一致性(PL0)与词元锚定(PL2)在池化排名上几乎相同,但其位置召回率@1分别达到13.5%和98.8%;在LJSpeech语音合成任务中,冻结的PL2配合训练好的MatchaTTS读出模块在受损文本上实现21.54%的词错误率(WER),而冻结的PL0仅为99.22%,端到端匹配也达到10.93%。这些结果说明,仅具备几何规则性并不足以保证序列内容可恢复或下游任务的有效访问。
链接: https://arxiv.org/abs/2610.07906
作者: K. P. Santoso,N. Z. Fadil,F. P. Harsanti,R. V. H. Ginardi,G. N. Iyer
机构: Institut Teknologi Sepuluh Nopember (印尼十日理工大学); Avalon AI; Universitas Indonesia (印度尼西亚大学); National University of Singapore (新加坡国立大学)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
备注:
Abstract:We study sequential content sufficiency by investigating whether a representation retains the ordered target information available in its input. An information-theoretic decomposition separates input ambiguity, representation loss, and readout mismatch. We construct recoverable views where perfect agreement and joint isotropic Gaussianity coexist with zero target information, and establish limits imposed by deterministic canonical anchors. Token log-loss provides a one-sided information-loss bound; a fixed-penalty ridge analysis shows why rank alone cannot determine prediction risk. These results motivate CANOPE, a nonautoregressive framework with ordered latent canvases, canonical-token supervision, and geometric regularization. On 40,000 validation sequences, latent-agreement (PL0) and token-grounded (PL2) have nearly identical pooled ranks but reach 13.5% and 98.8% positional Recall@1, respectively, under strong natural corruption when the correct target length is provided. On 3,930 LJSpeech validation utterances, frozen PL2 with a trained MatchaTTS readout yields 21.54% word error rate (WER) on corrupted text, versus 99.22% for frozen PL0, while end-to-end MatchaTTS reaches 10.93%. These results show that geometric regularity alone does not guarantee recoverable sequential content or effective downstream access in the text settings studied here.
[NLP-52] ARIA: Audio-Driven Melody-Tone Relation Modeling for Cantonese Lyric Authoring EMNLP2026
【速读】: 该论文旨在解决粤语歌词创作中声调(tonal)与旋律(melodic pitch)对齐的难题,尤其针对实际歌曲创作场景中旋律以原始人声演唱音频或哼唱录音形式存在、其音高信息隐含、嘈杂且无结构的问题。现有基于符号化旋律的歌词生成方法难以直接应用于此类非结构化音频输入。为此,论文提出ARIA框架,其核心在于构建一个两阶段的音频驱动旋律-声调关系建模机制:首先设计三流关系感知声调估计器(TRATE),通过建模多源声学线索与声调间关系结构,从带字符时间戳的演唱音频中预测0243声调序列;其次提出解耦检索增强型声调条件歌词生成器(DRA-TCLG),基于预测的声调规划结合检索增强的词汇引导,生成符合声调约束的流畅歌词。该方案的关键创新在于将非结构化音频中的隐含音高信息转化为可计算的声调序列,并通过端到端的声调-歌词联合建模实现高质量粤语歌词生成。同时,研究构建了一个大规模真实演唱录音对齐的音频-Jyutping-0243数据集,为该任务提供了坚实的数据基础。实验结果表明,ARIA在0243预测与声调一致的歌词生成任务上均取得优异表现,验证了所提框架的有效性。
链接: https://arxiv.org/abs/2610.07902
作者: Shengyu Li,Jinting Wang,Li Liu
机构: The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
类目: Computation and Language (cs.CL); Sound (cs.SD)
备注: Accepted for publication in Findings of EMNLP 2026. 24 pages, including references and appendices. Author-prepared version
Abstract:Cantonese lyric writing requires close alignment between lexical tones and melodic pitch. Existing melody-guided lyric generation methods typically rely on symbolic melody to generate lyrics. However, in real songwriting scenarios, melodies are often expressed as raw singing audio or hummed recordings, where pitch is implicit, noisy, and unstructured, making these methods difficult to apply directly. To address this limitation, we propose ARIA, a two-stage audio-driven melody-tone relation modeling framework for Cantonese lyric authoring that generates Cantonese lyrics from singing recordings with provided character-level timestamps. Specifically, we first design a Tri-Stream Relation-Aware Tone Estimator (TRATE) to predict 0243 sequences from timestamped singing audio by modeling multi-stream acoustic cues and relational tonal structure. We then propose a Decoupled Retrieval-Augmented Tone-Conditioned Lyric Generator (DRA-TCLG) to generate fluent lyrics conditioned on predicted tonal plans with retrieval-enhanced lexical guidance. Moreover, we construct a large-scale aligned audio-Jyutping-0243 dataset from real Cantonese singing recordings to support this new task. Experimental results demonstrate that ARIA achieves strong performance in both 0243 prediction and tone-consistent lyric generation, validating the effectiveness of the proposed framework.
[NLP-53] Rethinking Faithfulness in LLM s: A Pairwise Context-Sensitive Perspective
【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在问答任务中缺乏对上下文充分性敏感性的问题,即模型在上下文信息不足时仍倾向于生成答案而非选择不回答,导致幻觉(hallucination)现象。现有评估方法多采用孤立的实例级评价,无法捕捉模型在不同上下文条件下的响应适应能力。为此,本文提出一种成对一致性忠实度基准(Pairwise Faithfulness Benchmark, PFaithBench),通过对比同一问题在支持性与非支持性上下文下的响应行为,评估模型是否具备根据上下文充足性动态切换“回答”与“不回答”的能力。研究结果表明,当前大多数模型存在显著的“回答偏好”偏差,多数忠实度错误源于过度回答;此外,模型训练效果高度依赖于回答与不回答数据的构建方式,若两类数据来源不匹配,模型可能学习到数据集特异性捷径而非真正基于上下文充分性的判断机制。增加回答监督数据虽提升回答准确性,但加剧了幻觉问题;而增加不回答数据虽降低幻觉,却可能导致过度回避。因此,解决方案的关键在于设计合理且来源一致的标注数据,并在回答与不回答之间实现平衡训练,以实现真正的可信响应行为。
链接: https://arxiv.org/abs/2610.07894
作者: Zizhuo Zhang,Xiong Peng,Jingwei Sun,Rong Yao,Borui Jiang,Bo Han
机构: Hong Kong Baptist University (香港浸会大学); Huawei Noah’s Ark Lab (华为诺亚方舟实验室)
类目: Computation and Language (cs.CL); Machine Learning (cs.LG)
备注: 22 pages
Abstract:Large language models (LLMs) are expected to answer questions faithfully based on the provided context, abstaining when the context information is insufficient to answer the questions. Existing faithfulness evaluations typically assess each question-context instance in isolation; however, such instance-level evaluation fails to capture a fundamental requirement of faithful behavior: the ability to adapt model responses to changes in available contexts. In particular, a model should provide correct answers when sufficient evidence is present and abstain when it is not. In this work, we propose a Pairwise Faithfulness Benchmark (PFaithBench) that evaluates whether a model can switch between answering and abstaining for the same question under supporting versus non-supporting contexts. Our evaluations across thirty-nine models with seven model families demonstrate that faithfulness fundamentally involves a trade-off between answering and abstaining, and that most current models exhibit a strong bias toward answering, with most faithfulness errors arising from over-answering, i.e., models tend to fabricate a response even when the provided context is insufficient. We further conduct a series of studies on faithfulness training under different data constructions. Our results show that training outcomes are highly sensitive to the specific composition of answering and abstaining data. Constructing answering and abstaining data from mismatched sources can cause models to rely on dataset-specific shortcuts rather than actual context sufficiency. Moreover, increasing answer-supervised data improves answering performance but exacerbates over-answering, while increasing abstaining data reduces hallucination but leads to over-abstention. The code and data are released at this https URL.
[NLP-54] Visual Abstention in Unified Multimodal Models
【速读】: 该论文旨在解决统一多模态模型(Unified Multimodal Models, UMMs)在执行视觉编辑任务时缺乏对任务可行性判断能力的问题,即当用户请求的视觉变换违反任务规则或在给定图像条件下不可行时,模型往往未能识别这一矛盾,而是盲目生成不合理的输出。其核心挑战在于:当前模型将理解与生成能力割裂,导致在面对不可行请求时仍强行生成内容,甚至虚构不存在的图像对象或隐性修改请求以完成任务,表现出“错误生成”而非“合理拒绝”的行为。解决方案的关键是提出一种名为VisTA(Visual Transformation and Abstention)的新训练范式,通过在训练阶段成对引入可行与不可行样本,使模型在生成前先进行可行性判断,从而实现“能做则做,不能做则拒”。该方法显著提升了模型对不可行请求的拒绝率(从0.4%提升至93.0%),同时保持较高的编辑准确率(74.3%),且未造成编辑性能下降,实现了生成质量与理性克制之间的平衡。
链接: https://arxiv.org/abs/2610.07887
作者: Chufan Shi,Cheng Yang,Tiannuo Yang,Isadora White,Yiwei Chen,Taylor Berg-Kirkpatrick,Xuezhe Ma
机构: University of Southern California(南加州大学); University of California San Diego(加州大学圣地亚哥分校)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
备注: 25 pages, 6 figures, 13 tables. Project page: this https URL
Abstract:Unified multimodal models (UMMs) integrate understanding and generation, yet their generative behavior is rarely governed by what they understand about the task. We formalize visual abstention: when a requested visual transformation is impossible under the task’s rules, the model should recognize that no valid solution exists, state this, and decline to generate. We introduce Draw-or-Decline (DoD), a benchmark of 1,050 feasible-infeasible request pairs across 7 task categories that jointly measures editing success and the refusal of infeasible requests. Evaluating 8 UMMs, we find that editing ability and abstention are distinct capabilities: even the strongest editor, at 68.4% editing accuracy, refuses only 0.4% of infeasible requests under ordinary instructions. Their reasoning shows why: the models rarely notice the conflict, and instead plan the edit as if the request were possible, often describing objects that are not in the image, or quietly change the request into one they can complete. Explicitly prompting these UMMs to report infeasibility increases textual refusals but reduces editing accuracy. We propose VisTA (Visual Transformation and Abstention), a training method that pairs feasible and infeasible examples so that a model judges feasibility before deciding whether to generate. We train VisTA-BAGEL to perform feasible edits and decline infeasible requests. Without any reminder, it refuses 93.0% of infeasible requests, up from 0.4% for the strongest editor, while falsely refusing only 0.8% of feasible ones. Unlike a reminder, this does not cost editing accuracy: VisTA-BAGEL completes 74.3% of feasible edits, more than any of the 8 evaluated UMMs.
[NLP-55] ReFold: Training-Free Reversible Inter-Turn Context Folding for Long-Horizon Agents
【速读】: 该论文旨在解决长时程大语言模型(LLM)智能体在持续交互过程中因上下文不断累积而导致的上下文窗口溢出问题。随着交互轮次增加,历史信息被重复发送至模型,造成上下文长度与计算开销呈线性增长,进而引发性能下降甚至任务中断。现有方法依赖额外模型调用、启发式规则或训练好的策略来预测并管理上下文需求,但此类方法引入运行时开销、破坏前缀缓存(prefix cache)有效性,并存在不可恢复的内容丢弃风险。为克服上述局限,本文提出 ReFold——一种无需训练的渲染层机制,其核心创新在于仅压缩模型实际渲染的上下文,而保留原始交互历史完整可追溯。ReFold 通过两种无辅助预测器的冗余消除操作实现高效压缩:一是移除已被先前轮次展示过的重复内容,以占位符替代;二是将智能体自身标记为“已完成”的对话轮次折叠为一行摘要。上述操作均采用分块渲染策略,每数步更新一次缓存前缀而非每步更新,显著降低频繁重写带来的开销。所有删除操作均为严格可逆,误删仅需从历史中恢复一次,避免永久信息丢失。由于工作于渲染层,ReFold 可无缝集成至标准 ReAct 风格框架中。实验结果表明,ReFold 在五个长时程基准测试和两款前沿大模型上均表现出色,可将令牌消耗降低最高达 2.5 倍,会话级 KV 缓存内存占用减半,且不损害任务成功率;在固定上下文预算下,可避免高达 92% 的强制压缩;在并发服务场景中,请求队列延迟减少最高达 100%,推理速度提升最高 1.7 倍,推理成本降低最高 3.4 倍。
链接: https://arxiv.org/abs/2610.07863
作者: Yupeng Su,Jiayi Tian,Zheng Zhang,Souvik Kundu
机构: University of California, Santa Barbara (加州大学圣塔芭芭拉分校); Intel(英特尔)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: 27 pages, 6 figures, 14 tables
Abstract:Long-horizon LLM agents act on an append-only interaction history that is re-sent to the model at every step, so the context and its cost grow with steps until the sessions exceed the context window. Existing methods manage the context through context requirement prediction, relying on additional model calls, heuristic rules, or trained policies. However, these predictive approaches introduce runtime overhead, invalidate prefix caches, and permanently discard content with no guarantee of recovery. To overcome these limitations, we introduce ReFold: a training-free rendering layer that preserves the underlying interaction history while compressing only the model’s rendered context. It removes two kinds of inter-turn redundancy without an auxiliary predictor: content an earlier turn already displayed, replaced by a stub, and turns the agent itself reports finished, folded into a one-line note. Both operators use chunked rendering, rewriting the cached prefix once every few steps rather than at every step. Every removal is strictly reversible, a wrong removal costs one restore from the history rather than permanent content loss. Because it operates at the rendering layer, ReFold is plug-and-play across standard ReAct-style harnesses. Evaluations across five long-horizon benchmarks and two frontier LLMs demonstrate that ReFold reduces token consumption by up to 2.5x and halves the KV-cache memory per session without degrading task success rates. Under capped context budgets, it avoids up to 92% of forced compactions. Under concurrent serving workloads, it reduces request queuing delays by up to 100%, accelerating inference by up to 1.7x, while cutting inference costs by up to 3.4x.
[NLP-56] Lost in the bf16 Cast: Exporting Ternary Language Models Can Revert Most Low-Learning-Rate Code Changes
【速读】: 该论文旨在解决三值语言模型(ternary language models)在从高精度潜变量(high-precision latent weights)微调后,通过导出流程生成三值代码时存在的不一致性问题。其核心挑战在于:当前实验室公开的导出流程中,潜变量首先被转换为bfloat16(bf16),而这一过程引入的舍入误差导致最终部署的三值代码与原始检查点中的浮点量化结果出现偏差。具体而言,在Falcon-E和BitCPM中,约0.83%–1.77%的代码存在不一致;在BitNet 2B-4T中则达1.53%,且多数差异源于bf16舍入恰好落在阈值处,按“四舍六入五成双”规则映射为零。实验表明,仅使用标准工具导出时,部分Falcon-E版本可完全逐字复现,说明问题根源在于量化与导出流程间的兼容性。更严重的是,在微调后的端点上,若采用文档中推荐的导出方式,模型在GSM8K严格准确率上显著下降——例如Falcon-E-1B-Base从58.79%降至0.78%,BitCPM-CANN-0.5B从36.13%降至0.39%。此外,对BitNet 2B-4T进行bf16保存与重载也导致严格准确率下降27.54分,尽管末位数字准确率上升,暗示潜在数值稳定性问题。为应对上述问题,论文提出两种兼容性修复方案:一是直接写入训练量化器产生的代码,二是调整bf16输入使不变工具输出目标代码。两种方法均在所有三个模型中满足严格的4分非劣性标准,证明其有效性。进一步分析显示,在两个模型族中,对初始距离阈值的随机扰动支持基于距离依赖的选择策略,即微调过程更可能改变那些靠近阈值的代码,揭示了量化敏感区域的动态特性。
链接: https://arxiv.org/abs/2610.07853
作者: Avichal Sahai(Ofbusiness),Nishant Raj(Ofbusiness),Animesh Srivastava(Ofbusiness)
机构: 未知
类目: Machine Learning (cs.LG); Computation and Language (cs.CL)
备注: 14 pages, 4 figures, 17 tables
Abstract:Ternary language models such as BitNet b1.58, Falcon-E and BitCPM are fine-tuned with higher-precision latent weights and deployed as ternary codes produced by an export step that, in the labs’ documented pipelines, first casts the latents to bf16. We audit those pipelines across three labs. In released checkpoints, fp32 quantization of the shipped latents disagrees with the deployed codes on 0.83-1.77% of codes in Falcon-E and BitCPM and on 1.530% in BitNet 2B-4T; for Falcon-E and BitCPM most disagreements are products that bf16 rounding lands exactly on the threshold, which ties-to-even maps to zero, and the unmodified onebitllms exporter reproduces all four Falcon-E releases byte for byte. At fine-tuned endpoints, with learning rates selected to match a nominal learning-rate-to-bf16-ULP ratio, the documented export lowers greedy GSM8K strict accuracy from 58.79% to 0.78% for Falcon-E-1B-Base and from 36.13% to 0.39% for BitCPM-CANN-0.5B, and a bf16 save and reload lowers BitNet 2B-4T’s strict accuracy by 27.54 points while its last-number accuracy rises. Two compatibility remedies, writing the training quantizer’s codes directly or adjusting the bf16 inputs until the unchanged tools emit them, each met a 4-point strict-accuracy non-inferiority criterion against online evaluation in all three models. In two model families, randomized interventions on the initial distance from the threshold support distance-dependent selection of the codes that fine-tuning changes.
[NLP-57] Dynamic Positional Attention Modulation for Parameter-Efficient Fine-Tuning of Large Language Models KDD2026
【速读】: 该论文旨在解决现有参数高效微调(Parameter-efficient fine-tuning, PEFT)方法在适应大语言模型下游任务时,因采用统一且静态的调整策略而无法有效捕捉注意力机制中跨维度、头、层及输入标记的结构异质性问题。具体而言,实际注意力表示具有非均匀特性,且旋转位置编码(Rotary Positional Embeddings, RoPE)引入了依赖维度的位置结构,使得均一化调整方式次优。为此,本文提出动态位置注意力调制(Dynamic Positional Attention Modulation, DyPAM),其核心创新在于直接作用于查询(query)和键(key)表示,通过结合输入条件驱动的维度级调制、头级与层级结构调制,实现与RoPE诱导结构对齐的细粒度位置注意力适应,同时不修改预训练主干网络。实验结果表明,DyPAM在多个骨干模型上的数学推理与常识推理基准任务中均显著优于现有主流PEFT基线方法。
链接: https://arxiv.org/abs/2610.07848
作者: Dayan Pan,Jingyuan Wang,Xie Yu
机构: Beihang University (北京航空航天大学); MOE Engineering Research Center of Advanced Computer Application Technology (教育部先进计算机应用技术工程研究中心)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: Accepted by KDD 2026
Abstract:Parameter-efficient fine-tuning (PEFT) has become a standard approach for adapting large language models to downstream tasks. However, most existing PEFT methods rely on uniform and static adaptations, without accounting for the structured heterogeneity of attention across dimensions, heads, layers, and input tokens. In practice, attention representations exhibit non-uniform behavior, and positional encoding mechanisms such as rotary positional embeddings (RoPE) induce dimension-dependent positional structure, making uniform adaptation suboptimal. In this work, we propose DyPAM (Dynamic Positional Attention Modulation), a PEFT method that adapts how positional information contributes to attention by operating directly on the query and key representations. DyPAM combines input-conditioned, dimension-wise modulation with head-wise and layer-wise structural modulation, performing fine-grained adaptation of positional attention aligned with the RoPE-induced structure without modifying the pretrained backbone. Extensive experiments on mathematical and commonsense reasoning benchmarks across multiple backbone models demonstrate that DyPAM consistently outperforms existing strong PEFT baselines.
[NLP-58] OMIT the Action: Measuring Framing-Invariant Omission Bias under Philosophical Disagreement AACL
【速读】: 该论文旨在解决大语言模型(LLM)在道德推理中普遍存在但尚未充分研究的“遗漏偏差”(omission bias)问题,即模型倾向于选择不作为,即使在框架重构后其实际后果等同于积极行动。现有研究多局限于小规模样本且聚焦于功利主义与义务论之间的冲突,难以全面反映复杂情境下的偏差表现。为此,本文提出OMIT基准,包含218对框架对比场景,覆盖10类道德冲突,其构建基于由五种哲学人格角色(功利主义、义务论、美德伦理、关怀伦理和契约论)组成的LLM驱动辩论小组所揭示的分歧模式。实验评估八款主流LLM发现,遗漏偏差普遍存在,且在同一系列模型中与模型规模呈负相关。进一步测试四种推理时干预策略表明,引导模型在作出“是/否”判断前先考虑道德原则,可有效降低遗漏偏差并提升框架一致性响应;然而,部分情况下低遗漏偏差也可能伴随向行动偏倚的转移。本研究不仅提供了OMIT这一高保真度的评测基准,更提出一种利用多元哲学立场分歧信号评估框架敏感性不作为偏好及其缓解措施分布效应的方法论,为未来复杂道德决策场景下生成式AI(Generative AI)的公平性与可靠性评估提供了新范式。
链接: https://arxiv.org/abs/2610.07847
作者: Sihyeon Lee,Jihun Song,Chanwoo Kim,Jiwoo Kum,Chanjun Park
机构: Soongsil University (中央大学)
类目: Computation and Language (cs.CL)
备注: Accepted to AACL-IJCNLP 2026 Findings
Abstract:As LLMs increasingly assist in moral reasoning, omission bias, the tendency to prefer inaction even when equivalent framings reverse substantive outcomes, poses a significant risk of skewed decision-making. Yet omission bias remains underexplored in LLM evaluation, with the few existing studies limited in scale and focused largely on utilitarian-deontological conflicts. To address this gap, we introduce OMIT, a benchmark consisting of 218 paired-frame scenarios across 10 conflict types, constructed by leveraging disagreement patterns from an LLM-based, five-perspective philosophical persona panel (utilitarianism, deontology, virtue ethics, care ethics, and contractualism). Evaluating eight LLMs, we find that omission bias is pervasive but inversely correlates with model size within families. We further evaluate four inference-time interventions and find that interventions encouraging models to consider moral principles before committing to a yes/no answer reduce omission bias and increase frame-consistent responses, although lower omission bias rates can also coincide with shifts toward action-biased responses. Ultimately, this work contributes not only the OMIT benchmark, but also a methodology for using diverse philosophical disagreement signals to evaluate framing-sensitive inaction preferences and the distributional effects of mitigation attempts in LLMs under complex moral conflicts.
[NLP-59] Nucleus Speculative Decoding: Plausibility-Aware Verification Beyond Exact Distribution
【速读】: 该论文旨在解决自回归生成中因标准接受规则过于保守而导致的推测解码(speculative decoding)效率瓶颈问题。具体而言,传统方法仅基于精确分布校正进行验证,会拒绝那些在目标模型下仍具高合理性的候选词,即使其由轻量级草稿模型赋予了过高的概率,从而限制了每轮验证后可保留的草稿词数量。为此,论文提出一种名为“核心推测解码”(Nucleus Speculative Decoding, NSD)的松弛化验证机制,其关键创新在于将目标模型的合理性(plausibility)纳入验证标准:只要草稿词满足标准接受条件,或属于目标模型的核内(nucleus)词汇集,即可被接受。理论分析表明,该方法引入的分布偏差在单步误差上完全由草稿模型在目标核内的超额概率决定,并进一步推导出序列级保真度边界以量化局部偏差在自回归过程中的累积效应。实验结果表明,NSD在多种目标模型与提案机制下均显著提升推测解码效率,相较标准自回归解码实现最高达5.16倍的吞吐加速,较标准推测解码提升至多3.15倍,且性能保持竞争力。这一改进主要得益于更长的可接受长度,使得更多输出词能够分摊每次目标模型验证的计算开销。分析证实,基于合理性的松弛验证是提升推测解码效率的有效路径。
链接: https://arxiv.org/abs/2610.07822
作者: Shuhao Li,Fanghua Ye,Wanyu Lin,Tianyu Yuan,Xiaoyu Shen
机构: EIT-NLP Lab, Eastern Institute of Technology, Ningbo(东方理工大学人工智能实验室); The Hong Kong Polytechnic University(香港理工大学)
类目: Computation and Language (cs.CL)
备注:
Abstract:Speculative decoding accelerates autoregressive generation by using a lightweight draft model to propose multiple tokens that are verified by a target model in parallel. However, the standard acceptance rule focuses on exact distribution correction and rejects tokens that remain highly plausible under the target model when the draft model assigns excess probability. This conservative verification limits the number of draft tokens retained after each verification forward pass. We introduce Nucleus Speculative Decoding (NSD), a relaxed verification method that incorporates target-model plausibility into speculative decoding. NSD accepts a draft token if it satisfies the standard acceptance rule or belongs to the target model’s nucleus. We theoretically characterize the distributional deviation introduced by our method and show that the single-step error is exactly determined by the draft model’s excess probability within the target nucleus. We further derive sequence-level fidelity bounds that quantify how local deviations accumulate over autoregressive decoding. Experiments across multiple target models and proposal mechanisms demonstrate that NSD consistently improves speculative decoding efficiency while maintaining competitive task performance. Our method achieves throughput speedups of up to 5.16\times over autoregressive decoding and up to 3.15\times over standard speculative decoding. These improvements coincide with longer accepted lengths, allowing more output tokens to share the cost of each target verification pass. Analysis shows that plausibility-aware verification provides an effective approach for relaxed verification and speculative decoding efficiency. Our code is available at this https URL.
[NLP-60] αTransfer: Coefficient Transfer for Efficient Model Merging
【速读】: 该论文旨在解决大规模模型微调后参数合并(model merging)过程中因合并系数搜索空间呈组合爆炸式增长而导致的计算开销过大问题,尤其在模型规模和数量双重扩增时,传统方法面临高内存消耗与不可行的搜索成本。其解决方案的关键在于发现同一模型家族中不同尺寸模型在合并系数上的性能分布具有高度一致性,从而提出一种名为“α Transfer”的高效范式:在小型代理模型(proxy model)上完成最优合并系数的搜索后,直接将所得系数迁移至更大规模的目标模型。该方法显著降低了计算资源需求,在视觉变换器(Vision Transformers)上实现6倍加速与70%内存减少,在大语言模型(Large Language Models)上实现20倍加速与85%内存减少,同时保持与原方法相当的性能水平,验证了α Transfer在多种合并方法、模型族及任务上的通用性与可扩展性。
链接: https://arxiv.org/abs/2610.07819
作者: Shih-Cheng Huang,Zhi Rui Tam,Chieh-Yen Lin,Yun-Nung Chen,Hung-yi Lee,Shao-Hua Sun
机构: Appier AI Research(Appier人工智能研究院); National Taiwan University(国立台湾大学)
类目: Machine Learning (cs.LG); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)
备注: Under review
Abstract:Model merging offers a promising solution for combining multiple fine-tuned checkpoints into a single model through parameter arithmetic. However, finding optimal merging coefficients requires an extensive search that becomes prohibitively expensive as models scale in both size and number, due to high memory requirements and combinatorial growth in the search space. We show that, within the same model family, models exhibit highly congruent performance distributions over merging coefficients across different model sizes. This distributional similarity enables a practical paradigm we call \textit \alpha Transfer: searching for optimal coefficients on a small proxy model, then directly transfer them to larger target models. We verify \alpha Transfer across multiple merging methods, model families, and tasks. Experimental results demonstrate a 6 \times speedup and 70% memory reduction on vision transformers, and a 20 \times speedup and 85% memory reduction on large language models, while maintaining comparable performance. Our findings establish \alpha Transfer as an efficient and generalizable approach to scaling model merging.
[NLP-61] One Step at a Time: Trading LLM Autonomy for Process Predictability
【速读】: 该论文旨在解决生成式AI在自动化运营流程中因执行过程不可见而导致的可预测性与可审计性缺失问题。传统方法将操作规程(SOP)嵌入系统提示(system prompt),仅返回最终结果,使执行路径无法追踪,难以验证是否真实遵循流程。其核心解决方案是通过模型上下文协议(Model Context Protocol, MCP)实现步骤级交付:由服务器逐步释放操作指令,代理按序执行并返回结构化输出,从而在不依赖具体执行器能力的前提下,构建可预知、可追溯的执行路径。关键在于将流程控制权从执行器转移至外部协调机制,以牺牲部分自主性换取确定性——执行路径在运行前即被明确规划,且每一步输出构成机器可读的执行日志,支持后续审计与优化。实验表明,该方法在13个SOP-Bench领域、4种不同规模的开源执行器上均显著提升流程遵从度(76–95% → 95–99%),并几乎消除无依据答案(由2.1–4.5%降至0.2–0.3%),尤其对轻量级执行器带来+6.5个百分点的准确率增益,证明了外部流程供给可减轻其推理负担,而高性能执行器则以微小精度损失换取更高可解释性与可控性。
链接: https://arxiv.org/abs/2610.07817
作者: Hans Schabert,Christoph Peters
机构: Amazon Web Services(亚马逊网络服务); University of the Bundeswehr Munich(慕尼黑联邦国防军大学)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: 14 pages, 12 tables
Abstract:Organizations automating operational processes need more than a correct outcome: they need to predict how a process will run, know which one actually ran, and inspect it step by step. When an agent is the executor that predictability is normally lost: the prescribed procedure goes into the system prompt, and only a final answer comes back. We deliver the procedure step by step over the Model Context Protocol (MCP) instead: a server releases one step at a time, the agent executes it, and each step returns a structured step_output. This trades autonomy for predictability, and two properties then follow by construction, independent of the executor. The execution path is prescribed before the run, so the process is predictable in advance rather than reconstructed afterwards; and the completed step records form a machine-readable execution log that downstream tooling can audit and optimize step by step. Evaluating 15,475 trials across 13 SOP-Bench domains and four open-weight executors from frontier (Kimi K2.5) to lightweight (Ministral 3 8B), we find step-level delivery makes the executed process predictable and inspectable for every executor, and additionally raises accuracy when the executor is small. Across all four, process adherence rises significantly (76-95% to 95-99%) and ungrounded answers (correct outputs produced without executing the SOP) near-vanish, falling from 2.1-4.5% to 0.2-0.3% of trials (all 95% CIs exclude zero); under prompt-based delivery, 31-49% of correct answers on know_your_business bypass the SOP entirely, even for the frontier executor. Accuracy is where the executor’s capability enters: the lightweight executor gains +6.5pp grounded accuracy because supplying the process externally removes a reconstruction burden it cannot carry, while capable ones trade a small raw-accuracy decrement for a predictable, auditable process.
[NLP-62] hinkFuse: Trajectory-Aware Test-Time Fusion for Small Reasoning Models EMNLP2026
【速读】: 该论文旨在解决小型推理模型(Small Reasoning Models, SRMs)在复杂推理任务中因推理路径进入错误轨迹后难以自我修正的问题。现有测试时融合(test-time fusion)方法依赖局部融合信号判断是否触发融合,易受瞬时不确定性波动干扰,可能导致对不稳定的推理路径进行强化。其解决方案的关键在于提出一种无需训练的测试时融合框架ThinkFuse,通过对比片段级不确定性变化与轨迹级不确定性趋势,精准识别推理过程中的不稳定性点,并选择性地将辅助推理路径的信息融合至主模型的推理轨迹中。该方法实现了更高效的融合触发机制,在数学和知识密集型推理基准上均显著优于基线方法,且在不同模型组合间保持一致性能提升,同时在小规模主模型下仍具备鲁棒性。分析表明,ThinkFuse能减少融合触发次数并降低生成的标记符数量,体现了其选择性干预的高效性。
链接: https://arxiv.org/abs/2610.07803
作者: Myunghoon Kang,Jungseob Lee,Jaehyung Seo,Heuiseok Lim
机构: Korea University (高丽大学); Konkuk University (国民大学); Human-inspired AI Research
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: Accepted to EMNLP 2026 Findings
Abstract:Small reasoning models (SRMs) have shown strong performance on complex reasoning tasks by generating extended chain-of-thought trajectories, but they often fail to recover once their reasoning enters an erroneous path. Existing test-time fusion methods rely on local fusion signals to determine when to trigger fusion, which can be misled by transient uncertainty fluctuations and may reinforce unstable reasoning trajectories. We propose ThinkFuse, a training-free test-time fusion framework that selectively intervenes in unreliable reasoning segments. ThinkFuse compares segment-level uncertainty shifts with trajectory-level uncertainty trends to identify unstable reasoning points and fuse auxiliary reasoning paths into the primary model’s trajectory. Extensive experiments demonstrate that ThinkFuse outperforms baselines on mathematical and knowledge-intensive reasoning benchmarks, with consistent gains across model-family combinations, and remains robust with a smaller primary model. Our analysis shows that ThinkFuse requires fewer fusion triggers and generates fewer tokens, highlighting the efficiency of selective triggering. Our code is available at this https URL.
[NLP-63] Persistent Memory in Multi-Agent LLM Inference: What It Costs What It Buys and When You Can Tell NEURIPS2026
【速读】: 该论文旨在解决长上下文推理中因键值缓存(KV cache)内存占用过高而导致的计算瓶颈问题,尤其关注在多代理协作系统中如何有效控制每次调用的活跃KV缓存大小。其核心解决方案是通过将长上下文推理任务分解至多个协同代理之间执行,从而将每个查询的峰值KV工作集限制在较低水平,而非依赖于全局证据总量。关键创新在于:相较于单次遍历(single-pass)与检索增强(retrieval-augmented)基线方法分别高达35.5 MiB和35.3 MiB的峰值缓存占用,该分解策略将峰值缓存降至14.3 MiB,显著降低了内存压力。然而研究发现,以往系统引入的持久化层级(persistent tier)用于存储和召回推理轨迹,并未带来可检测的性能提升——在八组受控数据集上,该层级虽使峰值缓存增加0.368 MiB(95% CI [0.167, 0.590]),但准确率变化不显著(+0.015,95% CI [-0.011, +0.046])。作者指出这一“无效果”现象具有结构性本质:当前单问题基准测试设计使得每个样本独立获得自身证据并独立评分,导致不同条件间必须重置存储轨迹,因此记忆召回无法提供有效信息。为验证该结论,研究进行了四次测量校正,其中前三次人为夸大了持久化层级的潜在收益,第四次才使效应量达到可分辨程度,但实际结果表中并未体现此类偏差。为此,论文提出了代理记忆消融实验应满足的条件及无需预设缺陷知识的检测流程,以提升评估的严谨性与可重复性。
链接: https://arxiv.org/abs/2610.07782
作者: Hochan Son,Kyungdoe Han,Jaehan Koh,Xiaowu Dai,Wenlu Xu,Guang Cheng
机构: University of California, Los Angeles (加州大学洛杉矶分校); University of Wisconsin (威斯康星大学); HCLTech America (HCL科技美国)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Distributed, Parallel, and Cluster Computing (cs.DC); Machine Learning (cs.LG)
备注: 13 pages, 1 figure. Accepted as a poster at the Machine Learning for Systems Workshop, NeurIPS 2026
Abstract:Decomposing long-context inference across cooperating agents bounds the active KV cache per call rather than total evidence, which matters when KV-cache memory binds. Many such systems add a persistent tier storing and recalling reasoning traces, usually validated by an ablation reporting an accuracy gain. We measure both on one three-tier agent architecture. Decomposition delivers: peak KV working set of 14.3 MiB per query against 35.5 and 35.3 MiB for single-pass and retrieval-augmented baselines. The persistent tier does not: across eight controlled dataset pairs at n=100 per arm it costs +0.368 MiB [+0.167, +0.590] of peak cache and produces no detectable accuracy change (+0.015, 95% CI [-0.011, +0.046]). We argue the null is structural: single-question benchmarks supply each item with its own evidence and score it independently, and correctness requires resetting stored traces between conditions, so recall has nothing informative to retrieve. Reaching it took four measurement corrections – three inflating the apparent benefit, the fourth making an effect that size look resolvable – none visible in the results table. We give the conditions an agent-memory ablation must satisfy and detection procedures that need no knowledge of the specific defect.
[NLP-64] Quantization Effects on Tool-Failure Recovery Vary Across Prompts and Evaluation Designs NEURIPS2026
【速读】: 该论文旨在解决后训练量化(post-training quantization)在部署语言模型智能体时,对智能体从临时工具故障中恢复能力的影响评估不一致的问题。其核心挑战在于:量化精度(如8比特与4比特)对智能体鲁棒性的影响并非固定,而是高度依赖于评估方式的选择,包括所用提示(prompt)、任务集匹配性、评分策略及输出解析严格程度等。研究的关键发现是,不同评估维度会导致结论方向反转——例如,在相同任务和提示下,8比特Llama-3.1-8B-Instruct的表现可能优于或劣于4比特版本,具体取决于是否仅评估“干净通过”任务、是否考虑完整流水线成功率,或是否采用宽松/严格的输出解析标准。尤其值得注意的是,执行器宽容度(executor leniency)这一隐含假设会显著影响结果:当引入严格输出解析后,原本优势明显的8比特Llama的得分反而转为负值,而Qwen2.5-7B-Instruct则保持稳定。因此,该研究提出,可靠的量化智能体评估必须基于匹配的任务集、报告全流水线成功率、明确声明评分策略,并量化跨任务的不确定性,而非仅依赖特定故障注入点的局部表现。
链接: https://arxiv.org/abs/2610.07781
作者: Yuhe Hu
机构: Duke University (杜克大学)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
备注: Accepted at the NeurIPS 2026 Workshop on Small Language Models for Agentic Systems (SLM-Agents). 7 pages, 2 figures, 2 tables, plus appendix
Abstract:Post-training quantization reduces the cost of deploying language-model agents, but its effect on recovery from temporary tool failures can depend on how recovery is evaluated. We compare 8-bit and 4-bit variants of Llama-3.1-8B-Instruct and Qwen2.5-7B-Instruct on twenty deterministic tool-use tasks and five prompts. The 8-bit-4-bit recovery comparison changes direction across prompts and evaluation targets. On tasks that both variants complete without faults under the same prompt, the difference ranges from 0 to +20.2 percentage points for Llama and from -50.0 to +35.0 points for Qwen. Full-pipeline point estimates favor 8-bit Llama under all five prompts, whereas the Qwen comparison changes direction across prompts. The evaluation target can also reverse the result. For Llama under one prompt, scoring each variant only on its own clean-passing tasks favors 4-bit by 17.5 points; scoring the same tasks for both variants gives no difference, while scoring the full pipeline favors 8-bit by 28.3 points. Executor leniency is a third such choice. Rescoring the same logs with strict output parsing, which 8-bit Llama violates far more often than 4-bit Llama under that prompt, turns that +28.3 into -15.0 while leaving Qwen essentially unchanged. These findings show that one prompt, one screened task set, and one scoring policy do not establish a stable conclusion about quantized-agent robustness. Evaluations should compare variants on matched tasks, report full-pipeline success for deployment decisions, state the scoring policy, and quantify uncertainty across tasks rather than injected fault sites.
[NLP-65] APEX: Speculate smarter not deeper
【速读】: 该论文旨在解决生成式 AI(Generative AI)在大语言模型(Large Language Model, LLM)推理过程中因固定推测深度和提案机制导致的计算浪费与加速效率不匹配问题。具体而言,推测解码(Speculative Decoding)虽可通过提前生成多轮候选令牌以降低推理延迟,但其性能高度依赖于提案策略与推测深度,而静态配置无法动态适应生成过程中的可预测性变化、重复模式及验证接受率波动,从而在深层推测时引发大量无效计算。为应对这一挑战,论文提出 APEX,一种基于学习的控制器,其核心在于通过请求级专家选择(request-level expert selection)与块级深度自适应(block-level depth adaptation)实现解码速度与推测令牌浪费之间的动态平衡。APEX-Router 依据请求特性在 EAGLE-3、n-gram 和草案模型三种推测策略间进行选择;APEX-Depth 则利用因果解码信号与近期验证反馈,在每个验证块中动态调整推测长度。该系统将被接受的推测长度建模为受控生存反馈(censored survival feedback),通过学习位置级拒绝风险、块执行成本及动作效用函数(综合吞吐量、有效推进进度与浪费令牌数),实现对推测行为的精细化调控,同时保持目标模型的原始验证流程不变。实验表明,APEX 在 vLLM 框架中集成并应用于 Qwen3-8B 模型,在六类工作负载下相较自回归解码最高提升 5.24 倍,整体评估中 APEX-S 达到 4.27 倍加速,而 APEX-B 在维持更高推测利用率的同时实现 3.27 倍加速,并较固定 n-gram 推测(k=16)减少 41.0% 的浪费令牌比例,展现出显著的性能-效率权衡优势。
链接: https://arxiv.org/abs/2610.07780
作者: Manvi Jha,Zach Zhang,Zhichao Xu,Linbo Liu,Sai Muralidhar Jayanthi,Vinayak Arannil
机构: University of Illinois Urbana-Champaign (伊利诺伊大学厄本那-香槟分校); AWS AI (亚马逊云科技人工智能)
类目: Computation and Language (cs.CL); Machine Learning (cs.LG)
备注:
Abstract:Speculative decoding reduces large language model inference latency by drafting multiple tokens before target-model verification, but its effectiveness depends on both the proposal mechanism and draft depth. Fixed configurations cannot respond to changes in predictability, repetition, and acceptance during generation, so deeper drafting can increase wasted computation without proportional speedup. We introduce APEX, a learned controller that balances decoding speed and draft-token waste through request-level expert selection and block-level depth adaptation. APEX-Router selects among EAGLE-3, n-gram, and draft-model speculation for each request, while APEX-Depth adjusts draft length at each verification block using causal decoding signals and recent verifier feedback. APEX models accepted draft length as censored survival feedback, learning position-wise rejection hazards, block execution costs, and an action utility that balances throughput, accepted progress, and wasted tokens. This allows the controller to adapt speculation while retaining the target model’s verification procedure. We integrate APEX into vLLM and evaluate it with Qwen3-8B across six workloads, achieving up to 5.24X speedup over autoregressive decoding. Across the aggregate evaluation, APEX-S achieves 4.27X speedup, while APEX-B achieves 3.27X speedup with a 41.0% relative reduction in wasted-token percentage compared with fixed n-gram speculation at k=16, providing distinct operating points for balancing acceleration and draft-token utilization.
[NLP-66] Reading Not Manipulating: Leverag ing Router Logits for Multimodal Safety in MoE Vision-Language Models
【速读】: 该论文旨在解决视觉语言模型(Vision-Language Models, VLMs)在多模态输入交互下产生的组合性安全风险(compositional safety risks),即有害意图可能由视觉与文本信息的协同作用引发。随着混合专家模型(Mixture-of-Experts, MoE)架构在VLM中的广泛应用,现有安全干预方法如提示工程、监督微调及基于路由的专家引导虽被尝试,但其效果在不同模型和评估分布间表现不一致,且常因过度拒绝(over-refusal)引入安全-效用权衡。本文提出一种新思路:不通过修改模型内部状态来调控行为,而是利用路由状态(routing states)作为多模态安全性的诊断信号。研究发现,路由器输出的logits能够高度预测多模态输入是否具有安全隐患。基于此,作者设计了一种轻量级的路由日志安全检测器(router-logit safety detector),在提示预填充阶段读取路由信号,实现对潜在危险请求的提前识别,无需修改模型参数或专家路由策略。该方法在HoliSafe基准上显著降低安全错误率,并在包含MISHard和MM-SafetyBench等异构分布的安全基准上展现出强泛化能力。这一成果揭示了模型内部状态的新应用范式:仅通过读取自然涌现的信号并将其与外部安全机制联动,即可提供一种简单、高效且非侵入式的补充方案,为现有安全干预手段提供了重要拓展。
链接: https://arxiv.org/abs/2610.07774
作者: Ziyuan Yang,Wenxuan Ding,Shangbin Feng,Yulia Tsvetkov
机构: University of Washington (华盛顿大学); New York University (纽约大学)
类目: Computation and Language (cs.CL)
备注: 15 pages, 4 tables, 11 figures
Abstract:Vision-language models (VLMs) face compositional safety risks where harmful intent emerges from the interaction between visual and textual inputs. As mixture-of-experts (MoE) VLMs become increasingly common, recent work has explored various safety interventions, including prompting, supervised fine-tuning, and routing-based expert steering. However, these methods show inconsistent improvements across models and evaluation distributions, and the intervention into model behavior or internal states introduce safety-utility tradeoffs by over-refusal. Rather than manipulating internal states to steer model behavior, we instead ask whether routing states can serve as diagnostic signals for multimodal safety. We find that router logits indeed provide highly predictive signals of whether a multimodal input is safe or not. Motivated by this observation, we introduce a lightweight router-logit safety detector that reads out routing signals during prompt prefill and identifies unsafe requests before generation, without modifying model parameters or expert routing. Across Qwen3-VL and Kimi-VL, the proposed detector substantially reduces safety errors on the HoliSafe benchmark and resoundingly generalizes to out-of-distribution safety benchmarks featuring different safety patterns, including MISHard and MM-SafetyBench. The success of the proposed router-logit detector also suggests a broader perspective on model internals: rather than focusing only on manipulating internal components to steer behavior, simply reading naturally emerging signals and linking them to an external safety mechanism can provide a simple, effective, and non-intrusive complement to existing safety interventions.
[NLP-67] RACE: Rollout-Guided Quantization-Aware Training for FP4 Reinforcement Learning of MoE Language Models
【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在后训练阶段进行强化学习(Reinforcement Learning, RL)时,因推理生成过程(rollout generation)带来的巨大计算与内存开销问题。现有基于FP4(4位浮点)的RL方法虽能降低精度以提升效率,但其核心局限在于:仅独立优化训练路径与推理路径的量化精度,而未直接减小两者之间的量化执行差异(train-rollout discrepancy),导致性能下降。为此,本文提出一种名为TRACE(Train-Rollout Quantization Alignment via Compact GuidancE)的FP4量化框架,专为混合专家模型(Mixture-of-Experts, MoE)设计。其关键创新在于引入“推理引导的量化感知训练”机制,利用推理端的量化结果动态指导训练端的FP4舍入决策,从而直接缩小训练与推理路径间的量化偏差。此外,TRACE采用高效的量化信息缓存策略,选择性保留深层网络中的尾数(mantissa)与缩放因子(scale)信息,有效缓解由推理引导带来的存储与通信开销。实验在四个大规模MoE语言模型上展开,涵盖推理、编程及长周期强化学习任务,结果表明,TRACE实现了权重/激活与键值缓存(KV-cache)的联合FP4量化,在保持与BF16推理相当的强化学习性能的同时,推理速度最高提升5.4倍,并显著优于事后(post-hoc)对BF16训练策略进行的FP4量化方法,展现出卓越的高效性与最终性能。
链接: https://arxiv.org/abs/2610.07767
作者: Xin Wang,Hao Yu,Zhengyang Zhuge,Bochao Mao,Zheng Li,Junda Feng,Yuyan Luo,Yi Zhang,Yizhong Cao,Mi Zhang,Dayiheng Liu,Jianwei Zhang
机构: Alibaba Token Hub, Alibaba Group; Ohio State University
类目: Machine Learning (cs.LG); Computation and Language (cs.CL)
备注:
Abstract:Reinforcement learning (RL) for post-training large language models (LLMs) incurs substantial computation and memory overhead during rollout generation, which motivates low-precision rollout for efficient RL training. However, existing FP4 RL methods suffer from a key limitation: they primarily optimize quantization accuracy on the training and rollout paths independently rather than directly reducing the discrepancy between the two quantized execution paths. In this work, we propose TRACE (Train-Rollout Quantization Alignment via Compact GuidancE), an FP4 quantization framework for RL training of Mixture-of-Experts (MoE) language models that addresses the limitation of existing FP4 RL methods. TRACE incorporates rollout-guided quantization-aware training that uses rollout-side quantization outcomes to guide training-side FP4 rounding decisions, directly reducing train-rollout discrepancy. Moreover, TRACE adopts an efficient quantization-information caching scheme that selectively retains mantissa and scale information from deeper layers to reduce the storage and communication overhead introduced by rollout guidance. We evaluate TRACE on four large-scale MoE language models across reasoning, coding, and long-horizon RL tasks. Our results demonstrate that TRACE enables joint FP4 weight/activation and FP4 KV-cache rollout with RL performance comparable to BF16 rollout, while achieving up to 5.4xrollout speedup and strong final FP4 performance compared with post-hoc FP4 quantization of BF16-trained policies.
[NLP-68] No Transformer Beats Six Covariates: Long-Horizon Prediction of Depressive Symptoms from Childhood Essays
【速读】: 该论文试图解决的问题是:预训练的生成式人工智能模型(如Transformer)是否能够基于个体在11岁时撰写的文本,有效预测其23岁时可能出现的抑郁症状。研究聚焦于长期纵向预测场景,即跨越长达12年的时滞,评估文本分析模型在精神健康早期风险识别中的潜力。其解决方案的关键在于,通过在英国国家儿童发展研究(National Child Development Study)这一大型出生队列数据中进行实证检验,发现尽管采用了多种先进的自然语言处理技术(包括七种微调的Transformer模型、词袋模型、冻结嵌入和四种零样本大语言模型),但这些文本模型的表现均未显著优于仅使用六个童年协变量的逻辑回归基线模型。该基线模型的受试者工作特征曲线下面积(AUC-ROC)达到0.737,而表现最佳的Transformer模型仅为0.670,且所有文本特征的加入均未显著提升基线性能。此外,在经过贝叶斯校正后,五种领域预训练的Transformer模型也未能显著超越通用领域对照组。因此,研究结论指出,在长时程预测任务中,传统统计模型仍为最优基准,提示当前生成式AI在跨时间跨度的心理健康预测中存在局限性。
链接: https://arxiv.org/abs/2610.07764
作者: Daniel Kua,Emrul Hasan,John-Jose Nunez,Frances Chen
机构: University of British Columbia (不列颠哥伦比亚大学); Vector Institute (向量研究所)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注:
Abstract:Natural language processing (NLP) models can detect depression-related language in text written near the time symptoms are measured, but whether pretrained transformers can predict depressive symptoms from text written twelve years earlier is largely untested. In the National Child Development Study, a British birth cohort, we predict probable depressive symptoms at age 23 from essays the same people wrote at age 11. Our baseline, a logistic regression on six childhood covariates, outperforms every text model that sees only the essay: seven fine-tuned transformers, a bag-of-words model, frozen embeddings and four zero-shot large language models. Its area under the receiver operating characteristic curve (AUC-ROC) is 0.737 against 0.670 for the best transformer on the primary seed, and no added text score detectably raises the baseline’s AUC-ROC. None of the five domain-pretrained transformers detectably beats its general-domain control after Bonferroni correction. For long-horizon prediction, the baseline remains the model to beat.
[NLP-69] From Evidence to Action: How Tool-Using Agents Fail
【速读】: 该论文旨在解决生成式 AI 代理在执行工具使用任务时,尽管最终结果看似正确,但其决策过程缺乏充分前期证据支持的问题。核心问题是:在从“决定是否采取行动”到“执行单个动作及依赖性工作流”的过程中,证据与行动之间的因果链条在何处断裂。解决方案的关键在于构建一个名为 SafeActBench 的评估基准,涵盖六个操作领域和五种渐进式协议,系统化地测试代理在静态动作判断、非行动调查、单动作执行以及多动作工作流中的表现。通过引入基于溯源的证据账本(Provenance-bound Evidence Ledger)与确定性轨迹评估器,能够精确追踪证据建立的时间点、动作发生时机以及下游依赖关系的满足情况。研究发现,失败不仅源于信息缺失,更关键的是代理在决策与执行阶段对已有证据的利用不当,尤其是在执行前未完成充分调查或过早行动,而多动作工作流则进一步暴露了未解决的前置条件与执行不完整的问题。
链接: https://arxiv.org/abs/2610.07753
作者: Hongzhan Lin,Shidong Cao,Ziyang Luo,Wenhao Chai,Mong-Li Lee,Wynne Hsu
机构: Princeton University (普林斯顿大学); National University of Singapore (新加坡国立大学); Hong Kong Baptist University (香港浸会大学); Amazon Web Services (亚马逊网络服务)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: 36 pages. Project page: this https URL
Abstract:Tool-using agents make consequential changes to external state, yet correct outcomes do not guarantee that their actions were supported by evidence established beforehand. We study where this evidence-to-action chain breaks as agents move from deciding whether to act to executing single actions and dependent workflows. Across ten model-harness configurations, strong static action assessment can coexist with much weaker interactive execution. Failures often begin before execution: agents stop with incomplete investigation or act before required evidence is established. Once required evidence is obtained, single-action execution is usually reliable, while multi-action workflows additionally expose unresolved prerequisites and incomplete execution. For this analysis, we introduce SafeActBench, comprising 656 cases across six operational domains and five protocols that progress from static action judgment and investigated non-action to single- and multi-action workflows. A provenance-bound Evidence Ledger and deterministic trajectory evaluator track what information was established, when actions occurred, and whether downstream dependencies were satisfied. These results show that failures arise not only from missing information, but also from how agents use established evidence when deciding and executing actions.
[NLP-70] SanSi: A Looped Typed Decision Model for System 1.5 Thinking
【速读】: 该论文旨在解决传统生成式推理模型在决策任务中效率低下且缺乏可解释性的问题,尤其关注如何在不生成完整文本的前提下实现高效、精准的决策。其核心挑战在于:如何在保持快速响应(类似系统1思维)的同时,引入多轮迭代式推理以提升决策质量,而避免传统生成式推理带来的高延迟与冗余输出。解决方案的关键在于提出一种名为SanSi的新型框架,通过在预训练语言模型中引入可重复调用的循环机制(looping),使模型在不生成任何新token的情况下,通过多次递归应用相同网络层来逐步优化隐藏状态,形成“系统1.5”思维——介于单次前向传播与完全生成式推理之间的中间态。该方法在每个循环后均读取选项概率,并采用合适的评分规则进行端到端训练,使得同一模型可在不同计算预算(1至8次循环)下灵活适配,实现性能与效率的平衡。实验表明,该方法在多个数据集上达到72.0%的准确率,显著优于同规模非循环模型,且在深度控制任务中展现出超越训练深度的泛化能力,同时作为强化学习中的判别器,在无真实答案监督的情况下使生成器的F1提升7.7个百分点。
链接: https://arxiv.org/abs/2610.07730
作者: Shuyu Gan,Young-Jun Lee,Dongyeop Kang
机构: University of Minnesota(明尼苏达大学); SanSi Project Page
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: 43 pages, 15 figures, 42 tables. Project page: this https URL
Abstract:Typed decision models answer a declared question without generating text: a decision head returns a probability for each of the declared options in a single forward pass. A single pass is fast, intuitive System 1 thinking. We study what lies between one pass and generated reasoning: looping, in which the same layers are recursively applied several times before one typed readout. Each loop lets the model revise its hidden state before it commits to an answer, without generating a token; we call this System 1.5 thinking. We propose SanSi, which turns a pre-trained looped language model into a typed decision model. The option probabilities are read after every loop, and every loop is trained with a proper scoring rule, so that one model serves every budget from one loop to eight in a single run. On 10,027 test decisions from 59 sources, SanSi reaches 72.0% accuracy: 13.5 points above a non-looped model of the same shape trained with the same recipe, 5.3 points above a newer non-looped model of its size, and 1.8 points below one with three times the parameters. On two depth-controlled tasks, loops extend the solvable depth beyond the depths seen in training, where the larger single-pass model fails. Used as the judge for policy optimization with reinforcement learning, without gold answers, SanSi raises the generator’s F1 by 7.7 points.
[NLP-71] Does Steering Break Your Model? A Multi-Dimensional Evaluation Suite for LLM Steering Methods
【速读】: 该论文旨在解决生成式AI中激活控制(activation steering)方法在实际应用中面临的有效性与副作用之间的权衡问题。现有评估体系未能系统性地衡量激活控制方法在目标行为诱导、语言质量、任务能力、安全性与可靠性,以及泛化性能和数据依赖性等方面的综合表现,导致对方法间优劣的判断缺乏全面依据。为此,论文提出SteerScope——一个双轴多维度评估框架,通过15项指标联合刻画控制效果与方法特性,涵盖目标有效性与副作用在语言质量、任务能力、安全可靠性的表现,并引入样本效率与样本敏感性等专属指标以评估方法的泛化能力与数据依赖性。不同于传统单一操作点的对比,该框架系统揭示了有效性与副作用之间的内在耦合关系。实验基于23种方法(覆盖提示工程、LoRA、微调等四类主流技术)在相同模型、任务与评估协议下进行基准测试,结果表明当前激活控制方法尚未超越提示控制基线,在整体平衡性上仍存在不足:无论模型规模如何,所有被测方法均无法在提升有效性的同时避免更严重的综合副作用。此外,研究发现,在分布外(OOD)提示下,目标有效性往往得以保留,但副作用显著加剧,尤其表现为指令相关性和语言流畅性的下降,且不同方法在样本效率方面表现出显著差异。
链接: https://arxiv.org/abs/2610.07722
作者: Haotian Yang,Huikang Jiang,Yucheng Wu,Wen-Jie Jiang,Chenpeng Wang,Yibin Lou,Liangming Pan
机构: Peking University (北京大学); Columbia University (哥伦比亚大学); YiXin-AILab (易信人工智能实验室); Southern University of Science and Technology (南方科技大学)
类目: Computation and Language (cs.CL)
备注:
Abstract:Activation steering provides a lightweight and flexible way to control large language model (LLM) behavior. However, effective steering requires more than inducing the intended behavior: it should also limit unintended changes and remain robust across inputs and training data. Existing evaluations cover these dimensions only in fragments. As a result, the trade-offs between efficacy and side effects have not been systematically characterized. We introduce SteerScope, a two-axis, multi-dimensional evaluation suite that jointly characterizes steering outcomes and method properties through 15 metrics. We score target efficacy and side effects on language quality, task capabilities, and safety and reliability, and further assess generalization and data dependence through steering-specific metrics for sample efficiency and sample sensitivity. Rather than comparing methods at a single operating point, we characterize the trade-offs between efficacy and side effects. Under matched models, tasks, and evaluation protocols, we benchmark 23 methods spanning 4 families, including prompting, LoRA, and SFT as baseline methods, and release the suite as an extensible codebase. We find that current activation steering methods do not yet surpass the Prompt Steering baseline in their overall balance between steering efficacy and side effects: across both model scales, no evaluated activation steering method achieves higher efficacy without incurring greater composite side effects. We further uncover a consistent coupling between steering efficacy and side effects. Under OOD prompts, target efficacy is often preserved, whereas side effects tend to become more pronounced, particularly through declines in instruction relevance and fluency. Methods also exhibit sharply different sample-efficiency profiles.
[NLP-72] Readout Stability in Prefill-Only Decision Models:Zero-Label Prediction and Inference-Time Compute Allocation
【速读】: 该论文旨在解决生成式语言模型在推理阶段计算成本过高且缺乏可解释性决策机制的问题,尤其关注预填充仅(prefill-only)决策模型在高效、低成本推理中的潜力。其核心问题在于:如何在不进行解码的情况下实现准确的分类决策,并揭示此类模型内在的可预测性特征。解决方案的关键在于发现并利用一种可验证的理论性质——当仅改变候选菜单(candidate menu)而输入文本保持不变时,模型在干预后的准确率完全由首次前向传播中缓存的分布决定。基于此性质,提出了一种无需标签或二次前向传播的估算器:将第一轮的输出概率限制于新菜单范围,重新归一化后直接取最大值作为预测结果。实验表明,在七个模型家族、十个数据集和两类任务上,该估算器的预测误差不超过4.2个百分点,甚至在某一模型家族中达到精确预测;而概率级变体误差高达21.0点,说明该性质存在于排序而非概率本身,无法通过校准恢复。相比之下,同规模生成式语言模型不具备该特性,相同估算器的误差在1.6至15.8点之间且随模型增大而恶化。研究进一步表明,推理时的额外前向传递仅能带来微弱校准收益,而通过构建精炼候选菜单(如5候选子集)所能获得的性能提升远超扩大模型规模,例如0.8B参数模型在经筛选的5候选菜单上达到CLINC150数据集95.4%准确率,优于4B模型在完整150标签集上的80.0%表现。这表明,将推理计算转化为部署前可确定的决策策略,是实现高效、高精度决策的关键路径。
链接: https://arxiv.org/abs/2610.07716
作者: Ran Li,Lei Chen
机构: Hong Kong University of Science and Technology(香港科技大学); Hong Kong University of Science and Technology(Guangzhou)(香港科技大学(广州))
类目: Computation and Language (cs.CL)
备注:
Abstract:Prefill-only decision models inspired by the Jev model score every candidate in a menu during a single forward pass and never decode, which makes one call one to two orders of magnitude cheaper than a same-scale generative language model. We show that this read-out structure comes with a testable property. When an intervention changes only the candidate menu and leaves the input text fixed, the post-intervention accuracy is already determined by the cached first-pass distribution. The estimator restricts the pass-1 probabilities to the menu, renormalizes, and reads off the argmax; it uses no labels and no second forward pass. Across seven model families, ten datasets and two task types, menu-only interventions are predicted to within 4.2 points, and for one family the prediction is exact. A probability-level variant of the same estimator errs by 21.0 points, so the property lives in the ranking rather than in the probabilities and is not recovered by calibration. Same-scale generative language models do not share the property. On those models the same estimator errs by 1.6 to 15.8 points and degrades as the model grows. The property turns inference-time compute into a decision that can be made before deployment. Uniform extra passes buy calibration but almost no accuracy; at matched cost a confidence cascade outperforms every scheme that re-asks the same model, and curating the menu beats enlarging the model, with a 0.8B model on a curated 5-candidate menu reaching 95.4% on CLINC150 against 80.0% for a 4B model on the full 150-label this http URL and data are available at this https URL.
[NLP-73] When Old Facts Return: Re-Reads Reverts and the Limits of Temporal Memory
【速读】: 该论文旨在解决生成式AI在处理历史状态更新时所面临的“过时值”(stale value)歧义问题,即系统在重新读取旧语句或执行真实回滚操作时,可能产生相同的值序列观测结果,但实际所需响应却相反。其核心挑战在于:如何区分对某一值的观察是源于该值确实发生了改变,还是仅仅因为系统重新访问了已过期的历史状态。解决方案的关键在于引入一种“防止重激活已退役值”的保护机制(post-failure guard),该机制通过拒绝重新启用此前已被淘汰的值,有效避免了因误判旧值为有效而导致的错误响应。实验表明,在构造的重复读取场景下,该保护机制将准确率从10.8%恢复至97.7%,并将过时值率从88.5%降至0.8%。然而,该机制无法自动识别合法的回滚操作,除非额外提供变更溯源信息(change provenance)。研究进一步揭示,仅暴露被退役的历史记录或仅提供当前源代码均不足以完全解决此问题,强调需明确区分“值的观察”与“值已变更的证据”。
链接: https://arxiv.org/abs/2610.07715
作者: Neeraj Yadav(Called It Inc.)
机构: Called It Inc.; https://memstrata.dev
类目: oftware Engineering (cs.SE); Computation and Language (cs.CL)
备注: 12 pages, 1 figure. Ancillary files contain retained aggregate evidence, derived scenario and annotation exports, reference code, and an offline verifier
Abstract:A memory system can retire an obsolete value and later restore it merely because the same old statement appears again. A re-read of an old source and a genuine revert can produce the same observed sequence of values while requiring opposite current answers. We study this ambiguity on 130 extractor-selected atomic transitions derived from software fixes. In the ordinary transition condition, identity-based temporal memory reaches 98.5% model-judged accuracy with zero observed errors under a literal stale-value proxy. Appending a verbatim re-read of the old statement reduces accuracy to 10.8% and raises the stale-value rate to 88.5%. A guard that refuses to reactivate a previously retired value restores accuracy to 97.7% and reduces that rate to 0.8% in this constructed re-read condition. The guard cannot also recognize a legitimate revert without additional change provenance. Two supporting studies examine exposing retired history to the answer model and supplying current source for changed behavior. An exploratory extraction study over 707 software fixes provides scope context, not a universal coverage estimate. The design implication is to distinguish an observation of a value from evidence that the value changed. Selected inputs, aggregate-only answer records, related-family judges and a post-failure guard evaluation limit the conclusions to the retained experiments.
[NLP-74] Detecting LLM -Assisted Vietnamese Writing via Keystrokes under Behavioral Manipulation ICTAI2026
【速读】: 该论文旨在解决大语言模型(Large Language Model, LLM)辅助写作背景下,基于敲击动力学(keystroke dynamics)的检测方法在真实场景中的鲁棒性问题。研究提出了一种面向越南语的敲击动力学数据集,涵盖真实写作模式(包括真实创作、转录和改写),并构建了一个行为基础的威胁模型,模拟用户故意改变打字模式以规避检测的行为。为实现该威胁模型,研究生成了经过行为操纵的数据变体以逃避基于敲击动力学的识别。实验评估了四种建模方法:时间与节奏特征表示,以及基于一维卷积神经网络(1D-CNN)和TypeNet的序列化表示,在用户无关与上下文无关设置下的表现。结果表明,序列模型在多数情况下优于传统特征方法,且敲击信号确实编码了可区分的写作过程信息;然而,检测性能并非一致可靠——转录内容可被可靠识别,而改写及对抗性操纵样本则常被误判为真实创作,尤其在未显式建模相关行为时。为此,研究引入基于行为操纵数据的对抗训练,显著提升了分类可分性和鲁棒性。研究结论强调,敲击动力学检测的有效性高度依赖对多样化写作行为的充分暴露,有限条件下表现良好的模型若缺乏针对性建模,难以泛化至真实或对抗性场景。
链接: https://arxiv.org/abs/2610.07700
作者: Thanh Dong,An Ngo,Minh Dau,Rajesh Kumar
机构: Bucknell University ( bucknell大学)
类目: Computation and Language (cs.CL); Computers and Society (cs.CY)
备注: 9 pages, 2 figures. Thanh Dong and An Ngo contributted equally. Accepted at the 2026 IEEE International Conference on Tools with Artificial Intelligence (ICTAI 2026)
Abstract:We study the robustness of keystroke dynamics for detecting large language model (LLM)-assisted writing. We introduce a Vietnamese keystroke dataset capturing realistic writing modes, including bona fide composition, transcription, and paraphrasing. We also define a behaviorally grounded threat model in which users deliberately alter typing patterns. To implement the threat model, we create behaviorally manipulated variants of the data designed to evade keystroke-based detection. We evaluate four keystroke modeling approaches: temporal and rhythmic representations, and sequential representations modeled with a one-dimensional convolutional neural network (1D-CNN) and TypeNet, under user-independent and context-independent settings. The results show that sequential models outperform feature-based approaches in most cases and that keystroke signals encode discriminative information about the writing process. However, detection is not uniformly robust: transcription is reliably identified, while paraphrasing and adversarially manipulated samples are frequently misclassified as bona fide when not explicitly modeled. To address this, we incorporate adversarial training using behaviorally manipulated data, which substantially improves separability and robustness. These results suggest that keystroke-based detection depends critically on exposure to diverse writing behaviors, and that strong performance under limited conditions does not generalize to realistic or adversarial settings without targeted modeling.
[NLP-75] Improving Synthetic Data Generation for Argument Mining via Adversarial Reinforcement Learning EMNLP2026
【速读】: 该论文旨在解决论点挖掘(Argument Mining, AM)领域中高质量结构标注数据稀缺的核心瓶颈问题。现有方法依赖人工标注数据,难以满足模型训练对大规模、多样化且结构准确数据的需求。尽管大语言模型(LLM)在生成合成数据方面展现出潜力,但如何生成既具备结构准确性又保持足够多样性的合成AM数据仍是一大挑战。为此,论文提出一种新颖的对抗式强化学习框架用于合成数据生成,其关键在于构建一个生成器与判别器之间的对抗优化循环:生成器负责生成具有完整结构的AM实例,而判别器则通过区分真实数据与合成样本提供梯度反馈信号;该对抗机制促使生成器在持续迭代中逐步提升生成数据的结构准确性,同时通过对抗性反馈维持输出多样性。实验结果表明,该框架在三个基准数据集上均显著提升了AM性能,无论是在全量数据还是低资源设置下,验证了其有效性与可扩展性。
链接: https://arxiv.org/abs/2610.07699
作者: Zhijun Zhang,Qianlong Wang,Keyang Ding,Genan Dai,Bowen Zhang,Bin Liang,Ruifeng Xu,Yongsheng Liang
机构: Shenzhen University (深圳大学); Shenzhen Technology University (深圳技术大学); Harbin Institute of Technology (Shenzhen) (哈尔滨工业大学(深圳)); The Chinese University of Hong Kong (香港中文大学); Pengcheng Laboratory (鹏城实验室)
类目: Machine Learning (cs.LG); Computation and Language (cs.CL)
备注: Accepted to Findings of EMNLP 2026
Abstract:Argument Mining (AM) is fundamentally constrained by the scarcity of high-quality structure-annotated datasets. While LLMs have shown promise in synthetic data generation, producing synthetic AM data that is both structurally accurate and sufficiently diverse remains a challenging problem. To address this problem, we revisit synthetic data generation for AM from a new perspective and propose a novel adversarial reinforcement learning framework for data synthesis. The proposed framework jointly optimizes the generator and the discriminator in an adversarial loop, in which the generator produces structured AM instances, and the discriminator provides learning signals by distinguishing real data from synthetic candidates. This enables the generator to progressively improve both the structural accuracy of generated argument data while maintaining diversity through adversarial feedback. Extensive experiments demonstrate that the proposed framework consistently improves AM performance on three benchmark datasets in both full-data and low-resource settings, validating its effectiveness and scalability.
[NLP-76] DLoop: Looped Speculative Decoding
【速读】: 该论文旨在解决生成式 AI(Generative AI)中自回归生成过程因冗余验证带来的计算效率瓶颈问题。在传统的推测解码(speculative decoding)框架下,轻量级的草稿模型(draft model)每轮生成若干候选词元,随后由目标模型(target model)逐一验证,即使草稿模型高度可信且所有生成词元均有效,仍需执行完整的验证流程,导致大量不必要的目标模型前向传播开销。现有自适应草稿长度方法虽能缓解此问题,但仅适用于自回归草稿模型,无法适配并行化草稿模型(parallel draft model)——后者需要目标模型隐藏状态来支持未验证词元的生成。为此,本文提出DLoop,一种环状推测解码机制,其核心在于:在草稿模型保持高置信度时,持续进行多轮草稿生成,并在最终统一验证累积的全部草稿词元,从而显著减少目标模型的前向传播次数。为确保草稿模型在多轮迭代中仍具备可靠性,DLoop引入“环路感知训练”(loop-aware training),使其在训练过程中接触自身生成的未验证隐藏状态,增强长期生成一致性。实验表明,该方法在EAGLE-3、DFlash、Domino、DSpark及多词元预测模块等多种推测解码范式下,实现了5%至41%的墙钟时间加速,同时保证无损解码。
链接: https://arxiv.org/abs/2610.07659
作者: Geonmo Gu,Byeongho Heo,HeeJae Jun,Yoohoon Kang,Sangmin Lee,Sangdoo Yun,Dongyoon Han
机构: 未知
类目: Computation and Language (cs.CL)
备注: 22 pages
Abstract:Speculative decoding accelerates autoregressive generation in large language models. In each drafting stage, a lightweight draft model proposes tokens that the target model subsequently verifies. With increasingly capable draft models, we find that the target model frequently accepts all tokens produced in a drafting stage. A verification nevertheless follows each drafting stage, resulting in unnecessary target-model forward passes even when drafting could have continued. Adaptive draft length methods decide during decoding how many draft tokens precede a verification, but they raise the speedup only for autoregressive draft models. For a parallel draft model, drafting further requires target-model hidden states for draft tokens that have not been verified. We propose DLoop, a looped form of speculative decoding that adaptively performs multiple drafting stages before verification. DLoop continues drafting while the draft model remains confident and verifies all accumulated draft tokens together. Loop-aware training keeps the draft model reliable in the additional drafting stages by exposing it to its own hidden states for unverified draft tokens. By spending additional draft-model forward passes, DLoop reduces the number of target-model forward passes required for verification. Across diverse speculative decoding methods including EAGLE-3, DFlash, Domino, DSpark, and multi-token prediction modules, DLoop improves the wall-clock speedup by 5 to 41 percent while preserving lossless decoding. Code will be available at this https URL.
[NLP-77] Loud and Clear: Dynamic Activation Steering for Improving Speech Intelligibility in Noisy Environments ICASSP2027
【速读】: 该论文旨在解决在噪声环境中语音可懂度下降的问题,提出通过激活引导(activation steering)技术,在不重新训练预训练文本到语音(TTS)模型的前提下,动态调整语音以提升其在嘈杂环境中的可懂性。其核心挑战在于如何有效模拟人类的伦巴德效应(Lombard effect),即在噪声中自动增强语音努力程度和过度清晰化(hyper-articulation)的行为。解决方案的关键在于提出一种提示相对的引导机制(prompt-relative steering mechanism),该机制能够在生成过程中防止引导效应的累积,同时支持对引导强度进行动态调节。实验结果表明,该方法在已见与未见说话人、多种语言场景下均能系统性地改变与伦巴德效应相关的声学特征,保持较高的说话人相似度(89-95%),并在1 dB信噪比条件下将词错误率(WER)降低7%-22%,验证了预训练TTS模型可通过无需重训练的动态控制实现更鲁棒的语音生成。
链接: https://arxiv.org/abs/2610.07647
作者: Seymanur Akti,Alexander Waibel
机构: 未知
类目: ound (cs.SD); Computation and Language (cs.CL)
备注: Submitted to ICASSP 2027
Abstract:Speech becomes less intelligible in noisy environments, and humans naturally adapt their voice to compensate. Inspired by this behavior, we investigate whether a text-to-speech (TTS) model can be guided to produce more intelligible speech using activation steering, without retraining. We focus on two characteristics of the Lombard effect: increased vocal effort and hyper-articulation. We introduce a prompt-relative steering mechanism that prevents steering effects from accumulating during generation while allowing their strength to be adjusted dynamically. Across seen and unseen speakers and multiple languages, our method produces systematic changes in Lombard-related acoustic features, preserves speaker similarity (89-95%), and reduces WER under background noise by 7-22% at 1 dB SNR. These results show that pretrained TTS models can be dynamically controlled to generate more intelligible speech without retraining.
[NLP-78] Monte Carlo Estimation for KV Cache Eviction
【速读】: 该论文旨在解决大模型推理过程中键值缓存(KV-cache)高效管理的难题,核心问题在于传统缓存淘汰策略仅基于提示阶段(prompt)的注意力重要性进行判断,而忽略了生成回答时各缓存项的实际贡献。现有未来感知方法虽尝试预估未来查询需求,但依赖伪响应或合成查询估计,存在偏差且需额外训练或模型支持。本文提出一种无需训练的未来感知淘汰机制——LORE-KV(Lookahead Output-perturbation with Reliability-weighted Ensembles for Key-Value caches),其关键创新在于将固定预算的未来感知淘汰建模为对模型条件化查询轨迹的概率分布估计;通过从冻结的目标模型中采样短序列自回归延续,并利用其响应侧查询状态来估算提示词的效用价值。具体地,采用投影留一法注意力输出删除代价作为评分依据,结合可选轨迹加权策略对多个未来路径的得分进行聚合。这些临时生成的延续在最终解码前被丢弃,不引入额外模型或训练开销。实验表明,在缓存预算为128时,单个响应侧延续即可恢复约89%相对于提示窗口控制组的性能增益,多路径采样进一步微调提升;在Qwen2.5-14B和Mistral-7B上分别实现LongBench平均分提升2.75和16K RULER平均分提升5.85。尽管收益随缓存预算增大而递减,并伴随部分任务性能下降,但整体表现显著优于基线。该方法仅带来1.46–2.77倍于AnDPro的单样本运行时开销,作为一次性压缩预处理,适用于多种密集与混合注意力架构。
链接: https://arxiv.org/abs/2610.07643
作者: Ahsan Bilal,Muhammad Ahmed Mohsin,Muhammad Umer,Wajih Hassan Raza,Atta Ul Asad,Young D. Kwon,Michal Valko,Dean F. Hougen
机构: University of Oklahoma(俄克拉荷马大学); Stanford University(斯坦福大学); University of Houston(休斯顿大学); Lahore University of Management Sciences(拉合尔管理科学大学); University of Cambridge(剑桥大学); Isara Labs; University of Oklahoma(俄克拉荷马大学)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注:
Abstract:Most KV-cache eviction methods ask, in effect, which memory appeared important while reading the prompt? We instead ask, which memory will matter while answering? Since decoding queries are unavailable at eviction time, prior future-aware methods rely on pseudo-responses or synthetic future-query estimates. We cast fixed-budget future-aware eviction as distributional estimation over plausible model-conditional query trajectories and introduce LORE-KV (Lookahead Output-perturbation with Reliability-weighted Ensembles for Key-Value caches), a training-free method that samples short autoregressive continuations from the frozen target model and uses their response-side query states to estimate prompt-token utility. Tokens are scored by projected leave-one-out attention-output deletion cost and aggregated across sampled futures with optional trajectory weighting. The temporary continuations are discarded before final decoding, requiring no auxiliary model or training. Ablations isolate the mechanism: at B=128, a single response-side continuation recovers about 89% of the gain over the prompt-window control, while additional futures provide smaller improvements. At B=128, LORE-KV raises the LongBench average on Qwen2.5-14B from 45.49 to 48.24 (+2.75) and the 16K RULER average on Mistral-7B from 45.20 to 51.05 (+5.85). Gains diminish at larger cache budgets and coexist with task-level regressions. LORE-KV incurs 1.46-2.77x AnDPro’s per-sample wall-clock time as a one-time compression overhead across six dense and hybrid-attention backbones.
[NLP-79] Stateless Language Agents : Scaling Long-Horizon Automated Research
【速读】: 该论文旨在解决自动化研究系统中大型语言模型(LLM)代理在长时序运行下效率低下的核心问题:尽管推理成本持续增加,但系统却频繁出现重复工作、代理间行为同质化或实验探索停滞等失效模式。其根本原因在于两个设计选择——研究状态的存储位置以及下一步行动决策的主体。论文提出无状态语言代理(Stateless Language Agents, SLAs)作为解决方案,其关键在于将“状态感知搜索”与“无状态代理”相结合:不再由单个代理携带历史上下文,而是由协调框架(harness)统一管理研究状态(候选解及测量结果),并在每次调用时为每个代理重建角色特定的上下文。这一设计使代理所见内容成为可显式控制的工程选择,而非随运行时间无限增长的历史记录。在SLA框架中,无状态顾问(Advisor)基于框架汇总的多方向证据,为并行工作的代理分配具体实验任务。评估表明,在软件工程、内核优化和算法设计三个任务上,SLA在高达十亿令牌的预算下均取得最优结果,且在内核优化任务中以低于84%的令牌消耗达到最强基线性能。消融实验显示,聚焦的上下文与明确的任务分配各自贡献于进展提升,且效果可在长时间运行中累积;同时,顾问自身消耗的计算资源不足总令牌数的0.6%。研究结果表明,应将持久的研究状态从代理对话中剥离,且短周期评估可能严重误判研究系统的实际能力及其组件的有效性。
链接: https://arxiv.org/abs/2610.07625
作者: Qizheng Zhang,Changxiu Ji,Isaac Sun,Yuetai Li,Shubhangi Upasani,Sherry Ruan,Boyuan Ma,Fenglu Hong,Vamsidhar Kamanuru,Yoonho Lee,Yuzhen Mao,Genghan Zhang,Rulin Shao,Qiuyang Mang,Andy Dimnaku,Changran Hu,Radha Poovendran,Kunle Olukotun
机构: Stanford University (斯坦福大学); Carnegie Mellon University (卡内基梅隆大学); University of Washington (华盛顿大学); SambaNova Systems, Inc. (桑巴诺瓦系统公司); UC Berkeley (加州大学伯克利分校)
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: 32 pages
Abstract:Automated research systems increasingly run LLM agents over long horizons, but more inference does not by itself produce more progress: agents replay growing histories, duplicate one another’s work, or stop experimenting while token consumption continues. Yet most evaluations use short budgets or benchmarks that saturate early, leaving these failure modes untested. We trace these failures to two choices: where research state lives and who decides what to try next. We introduce Stateless Language Agents (SLAs), built on the principle of stateful search with stateless agents: no agent carries its conversation across invocations; instead, the harness owns the research state (candidate solutions and measured outcomes) and reconstructs a fresh and role-specific context for every invocation. What each agent sees becomes an explicit design choice rather than a history that grows with the run. We implement this principle in the SLA framework, where a stateless Advisor reads harness-summarized evidence across search directions and assigns concrete experiments to parallel Workers. We evaluate SLA against three recent frameworks on software engineering, kernel optimization, and algorithm design at budgets of up to one billion tokens. SLA achieves the best final result on every task and reaches the strongest kernel baseline’s final performance with over 84% fewer tokens. Ablations from shared checkpoints show that focused contexts and explicit assignments each contribute to SLA’s progress, with effects that can compound over full runs, while the Advisor consumes less than 0.6% of tokens. These results argue for SLAs, which keep durable research state out of agent conversations, and show that short evaluation horizons can misjudge research systems and their components.
[NLP-80] Recurrent Looped Transformer
【速读】: 该论文旨在解决传统Transformer模型在序列状态跟踪任务中因固定深度计算路径而导致的长序列泛化能力不足问题。尽管输入序列长度不断增长,Transformer对每个标记的处理深度保持不变,限制了其对长期依赖关系的建模能力。为此,论文提出循环回路Transformer(Recurrent Looped Transformer, RLT),其核心创新在于将网络层划分为并行的因果编码器与递归解码器两部分:编码器并行处理输入序列,而解码器在每个时间步将前一时刻的最终状态与当前编码器输出进行融合,从而实现计算路径随序列长度线性增长,同时保持每标记固定的计算开销。关键在于引入递归反馈机制,使模型能够持续累积和更新状态信息,显著增强对长序列中复杂模式(如奇偶性、置换跟踪、模运算)的泛化能力。实验表明,在训练序列仅40比特的情况下,RLT可将奇偶性任务泛化至256比特并达到100%准确率,而标准Transformer仅能随机猜测;在八倍于训练长度的置换跟踪任务中,RLT准确率达97%,远超Transformer的不足1%。消融实验进一步证实,移除反馈机制会导致性能退化至随机水平,且逐标记反馈对长序列置换跟踪至关重要,而分块更新则会显著降低性能,验证了反馈机制在维持状态连续性中的核心作用。
链接: https://arxiv.org/abs/2610.07591
作者: Yifan Zhang,Jichen Feng,Shihan Qin
机构: Princeton University (普林斯顿大学); University of Pennsylvania (宾夕法尼亚大学)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: Project Page: this https URL
Abstract:State tracking requires an update at every input, but the depth a Transformer applies to each token is fixed regardless of sequence length. We introduce the Recurrent Looped Transformer (RLT), which splits its layers between a parallel causal encoder and a recurrent decoder. At each token, the decoder merges the encoder output with the previous token’s final decoder state, so the computation path grows with sequence length at a fixed per-token cost. On six algorithmic tasks, we compare five splits of eight layers with an eight-layer Transformer over three seeds. Trained on at most 40 bits, two RLT splits generalize parity to 256 bits with 100% accuracy in every seed, while the Transformer stays at chance. On swap-based S_5 permutation tracking at eight times the training length, RLT reaches 97% final-state accuracy versus under 1% for the Transformer, and accuracy increases with decoder depth. On modular arithmetic beyond the training lengths, RLT reaches up to 93% versus 33% for the Transformer. Ablations show that these gains depend on the feedback: removing it drops parity and swap-based S_5 to chance at every split. Updating the feedback once per four-token chunk lets known tokens in a chunk run in parallel and keeps 64-bit parity at 99%, while permutation tracking depends on per-token feedback: chunking lowers length-64 swap-based S_5 from 100% to 20%.
[NLP-81] Large Language Model Orchestration under Heterogeneous Preferences via Explicit Persona Inference
【速读】: 该论文旨在解决大语言模型(LLM)编排框架中对异质偏好代理(heterogeneous-preference agents)隐藏偏好推断的难题。现有方法将代理的信念编码于提示词(prompt text)中,缺乏显式的更新机制,导致早期误差无法修正并持续传播。其解决方案的关键在于提出一种名为HARP(Heterogeneous-preference Agent oRchestration via Preference inference)的新框架,将信念从提示词中解耦,转而以数值后验分布的形式在有限候选偏好集合上进行显式维护,并通过贝叶斯规则闭式更新。语言模型仅负责生成动作及各候选偏好下的似然值,使推断过程与模型推理分离。理论证明表明,当因子分解精确时,HARP可达到与显式联合推断相当的O~(K)贝叶斯后悔率。此外,HARP⁺通过在规划中引入区分性奖励(bonus for actions that distinguish candidates),即使最优动作不具信息量,仍能持续推动偏好推断。在三类不同复杂度的实验场景(从偏好完全决定收益、到偏好之外还受其他因素影响,再到显式联合推断不可行的大规模场景)中的实证结果表明,HARP⁺是所识别理论类别中表现最强的非预言型方法。
链接: https://arxiv.org/abs/2610.07587
作者: Shuqing Shi,Ziyan Wang,Milind Tambe,Yali Du
机构: 未知
类目: Computation and Language (cs.CL)
备注:
Abstract:LLM orchestration investigates how an orchestrator coordinates a group of autonomous agents to achieve common goals or maximize collective welfare. The agents are typically heterogeneous, each holding a private preference that it pursues but does not reveal. Inferring such hidden preferences from behavior has been a subject of long-standing research in game theory and multi-agent systems. The core challenge lies in maintaining a belief over every agent’s preference and updating it from the agents’ observed actions. Existing LLM orchestrators carry that belief as prompt text with no explicit update rule. This lets early errors persist and propagate rather than be corrected. We therefore propose \textbfHARP (Heterogeneous-preference Agent oRchestration via Preference inference), a novel framework that moves the belief out of the prompt. Specifically, HARP maintains one numeric posterior per agent over a finite set of candidate preferences and updates it in closed form by Bayes’ rule. The language model supplies only actions and per-candidate likelihoods, so estimation is decoupled from its reasoning. We prove that HARP attains the same \tilde O(\sqrt K) Bayesian regret as explicit joint inference when the factorization is exact. Furthermore, HARP\textsuperscript+ augments planning with a bonus for actions that distinguish the candidates, so inference continues even when the optimal action is uninformative. Empirical results on three substrates, ranging from payoffs the preferences fully determine, through payoffs that depend on more than them, to scales where explicit joint inference is infeasible, demonstrate that HARP\textsuperscript+ is the strongest non-oracle method across the class our theory identifies.
[NLP-82] LOGIC: An LLM Benchmark for Intent-Grounded Change Impact in Aerospace Electrical Systems
【速读】: 该论文旨在解决航空航天电气设计修订中因变更传播范围过宽而导致影响报告过于冗余的问题。在实际工程请求仅授权部分变更的情况下,若对所有检测到的差异进行传播,将产生过度泛化的后果。其核心解决方案是提出LOGIC框架——一个可本地部署的可控基准与评估体系,利用语言模型(Language Model)在本地对工程请求进行确定性候选变更项生成,并通过有类型的电气可追溯图(typed electrical traceability graph)仅对选定变更进行下游传播。该设计的关键在于将候选变更选择错误与下游传播错误分离,从而实现更精准的影响分析。实验表明,在96个明确锚定的候选选择案例中,仅使用门控结构化证据(gate-only structured evidence)可达到1.0000的候选F1值,显著优于基于词元的词汇匹配方法(0.9677)。而在涉及语义重述的12个关系型案例中,传统词汇匹配方法表现不佳(F1=0.1772),而门控结构化方法完全失效(F1=0.0000),但大语言模型仍能维持0.5000–0.6400的合理性能。研究还发现,随着候选变更集规模从4增至64,模型的置信度下降,但受影响元素和类型路径的准确性在固定选择下于约1K至10万节点的图上保持稳定。严格证据门控虽能抑制误报,但可能排除正确语义选择;引入探索性空证据弃权策略后,所有模型的严格弃权准确率提升至0.6667,不安全报告率降至0.1667,但可回答案例覆盖率下降16.0%–27.1%。六组冲突请求中有四组对所有模型仍存在安全隐患。综上,研究支持在无法可靠确立意图时,结合字面证据、语言模型推理与人工工程评审以保障安全性。
链接: https://arxiv.org/abs/2610.07580
作者: Muhammad Faraz Shoaib,Muhammad Qasim,Raisulhaq Mohammed Rizwan,Rahmatullah Safdar,Muzammil Adnan Shaik,Abdul Aleem Mohammed
机构: Wichita State University (威奇托州立大学)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: 14 pages, 4 figures, 4 tables
Abstract:Aerospace electrical-design revisions can contain multiple genuine changes, although an engineering request may authorize only a subset. Propagating every detected difference can therefore produce overly broad impact reports. We present LOGIC, a controlled benchmark and evaluation framework in which locally deployable language models ground a request in a deterministic candidate-change inventory before selected changes are propagated through a typed electrical traceability graph. This separation permits candidate-selection errors to be distinguished from downstream propagation errors. LOGIC contains 168 scenarios, including 144 selection and 24 abstention cases. We evaluate three 7–8B models against intent-agnostic, lexical, and structured-evidence methods, with an oracle-root upper bound. On 96 explicitly anchored selection cases, gate-only structured evidence achieves candidate F1 of 1.0000, compared with 0.9677 for token-lexical matching. On 12 relational-paraphrase cases, token-lexical F1 is 0.1772 and gate-only F1 is 0.0000, compared with 0.5000–0.6400 for the large language models. Model grounding degrades as candidate inventories grow from 4 to 64 changes, while affected-element and typed-path accuracy remain comparatively stable when frozen selections are replayed over graphs of approximately 1K to 100K nodes. Strict evidence gating suppresses false positives but can remove correct semantic selections. An exploratory evidence-empty abstention policy raises strict abstention accuracy to 0.6667 for all three models and reduces unsafe-report rates to 0.1667, while decreasing answerable-case coverage by 16.0–27.1 percentage points. Four of six conflicting requests remain unsafe for each model. These findings support combining literal evidence and language-model reasoning with engineering review when intent cannot be established reliably.
[NLP-83] wo Vectors Replace In-Context Demos: Structured Task Adaptation via Embeddings
【速读】: 该论文旨在解决在上下文学习(In-context Learning, ICL)中,冻结的大型多模态模型(Large Multimodal Models, LMMs)因每次查询需重新编码演示样本而带来的高昂计算开销问题,尤其是每个演示图像会引入数百个视觉标记(visual tokens),导致推理效率低下。现有无演示(demo-free)方法虽通过紧凑的任务状态减少开销,但其任务参数通常分布在特定位置或每一解码层中,导致参数随网络深度增长,且插入的标记或键无法动态调整原始提示在层内注意力分配的方式。为克服上述局限,本文提出结构化任务适配嵌入(Structured Task Adaptation via Embeddings, STAVE),其核心创新在于用两个任务特定向量替代原始演示:一个“读出向量”用于更新生成答案的标记,另一个“上下文向量”用于更新其他结构性标记组。这两个向量均通过带标签的提示(含/不含演示)进行端到端训练。理论分析基于损失的一阶展开与边界裕度(margin bound),验证了设计合理性。大量实验表明,STAVE 在六种LMM和五种大语言模型上均达到或超越当前最优性能,在多模态任务中使用远少于对比方法的任务参数,且在18项文本任务上优于15次示例的ICL和先前任务向量方法,同时保持零次示例推理成本(zero-shot inference cost)。
链接: https://arxiv.org/abs/2610.07572
作者: Xi Ding,Naichen Shi,Jiawei Zhang
机构: University of Wisconsin–Madison(威斯康星大学麦迪逊分校); Northwestern University(西北大学)
类目: Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注: Technical report
Abstract:In-context learning (ICL) adapts frozen large multimodal models (LMMs) to new tasks from a few demonstrations (demos), but re-encodes them at every query, where each demo image adds up to hundreds of visual tokens. Demo-free methods remove this cost with a compact task state. However, they add it at locations searched per task or at every decoder layer, where task parameters grow with depth. Moreover, inserted tokens or keys cannot change how the original prompt divides its attention within a layer. To address these issues, we propose Structured Task Adaptation via Embeddings (STAVE), which replaces demos with two task-specific vectors added to existing input embeddings. Specifically, a readout vector updates the answer-producing tokens and a context vector updates the other structural token groups. Both are trained with answer labels on prompts with and without demos. We justify these design choices theoretically using a first-order analysis of the loss and a margin bound. Extensive experiments on six LMMs and five large language models show that STAVE matches or outperforms state-of-the-art methods on multimodal tasks with far fewer task parameters and surpasses 15-shot ICL and prior task vectors on 18 text tasks, all at zero-shot inference cost.
[NLP-84] HouseholdBench: Evaluating Large Language Models as Predictors of Household Economic Behavior
【速读】: 该论文旨在解决现有大型语言模型(LLM)在经济领域中对家庭决策行为建模能力不足的问题,尤其针对现有评估体系覆盖调查样本与预测任务有限、且未能充分考察家庭在经济环境变化下行为调整能力的局限性。其核心解决方案是提出一个名为HouseholdBench的新评估基准,该基准整合了6个美国家庭调查数据集和32项涵盖数值型、分类型及概率型结果的预测任务,涉及消费、收入、劳动力、预期与住房等多个经济维度。通过利用历史行为、人口统计学特征及宏观经济条件,该基准测试模型在预测家庭行为及其对政策变动响应方面的能力。实验对比了13个专有与开源权重的LLM模型,发现多数模型优于无变化基线,其中表现最佳的模型使数值型结果的误差降低12.2%;尽管梯度提升树(Gradient-Boosted Trees)在多数任务中仍居首位,但领先专有模型已接近其性能,而开源模型则明显落后。研究进一步揭示了模型在不同任务中存在系统性高估与低估现象,并提出通过微调(fine-tuning)与对每个观测值进行16次预测聚合的方法,使40亿参数的开源模型性能达到与专有模型相当水平,且该改进在未参与微调的政策响应任务中亦具泛化能力。研究成果已公开发布数据集、代码与排行榜以促进后续研究。
链接: https://arxiv.org/abs/2610.07563
作者: Jin Huang,Diego Ferreras Garrucho,Yutong Xie,Walter M. Yuan,Qiaozhu Mei,Chen Lian,Jonathon Hazell
机构: University of Michigan, Ann Arbor(密歇根大学安娜堡分校); London School of Economics(伦敦政治经济学院); MobLab Inc(莫布实验室公司); University of California, Berkeley(加州大学伯克利分校)
类目: Computation and Language (cs.CL)
备注:
Abstract:Large language models (LLMs) have the potential to meet a key goal in economics: a quantitative model of household decision making, across a variety of settings. Yet existing evaluations cover few surveys and outcomes, and do not study how households adjust to changing economic conditions. We introduce a new evaluation, HouseholdBench, which unites 6 U.S. household surveys and 32 prediction tasks spanning numeric, categorical and probabilistic outcomes, related to consumption, income, labor, expectations, and housing. Using past behavior, demographics and macroeconomic conditions, the tasks test whether LLMs predict behavior, including how households adjust to changes in various policies. We evaluate 13 proprietary and open-weight LLMs against a no-change baseline and a gradient-boosted tree model. Most LLMs outperform the no-change baseline, including for policy response tasks – with the best model lowering error for numeric outcomes by 12.2%. Across most tasks, gradient-boosted trees rank first; leading proprietary LLMs approach their performance, but open-weight models lag. LLMs exhibit systematic over- and underprediction across different tasks. We identify methods that enable a 4 billion parameter open-weight model to match proprietary models’ performance: fine-tuning and aggregating 16 predictions per observation. Improvements generalize to policy-response tasks, which are excluded from fine-tuning. We release our datasets, code, and leaderboard on our website: this https URL
[NLP-85] Quality-Aware Self-Correcting Speech Translation on an Edge Device
【速读】: 该论文旨在解决在资源受限设备(如Jetson Nano,4 GB内存)上实现高效、低延迟的离线语音到语音翻译(Speech-to-Speech Translation, STST)的问题,同时克服传统端到端模型在生成质量不佳时缺乏自纠正能力的缺陷。其核心挑战在于如何在不重新训练模型的前提下,对弱翻译结果进行有效修正,以提升翻译质量并保持系统实时性。解决方案的关键在于构建一个基于质量评估(Quality Estimation, QE)门控机制的双阶段纠错框架:首先使用轻量级Whisper-tiny自动语音识别(ASR)模块将语音转为文本,再通过Opus-MT多语言翻译模型生成初始译文;随后利用多语言BERT的余弦相似度作为QE评分器,当置信度低于预设阈值τ(如0.90)时触发二次修正流程。研究对比了三种修正策略:基于QE重排序(M1)、最小贝叶斯风险解码(M2)和约束束搜索(M3),发现仅依赖QE作为门控而非排序依据的M2方法表现最优,在BLEU、ChrF和COMET等指标上均显著优于贪婪解码(p < 0.001),而移除QE参与候选选择(即从M1转向M2)不仅未损害性能,反而释放了680 MB关键路径内存。进一步分析表明,结合后编辑效率文献中的收益-编辑比指标,较小的候选池规模(N=3)可实现更精准的语义修正,而较大的候选池(N=10)则更有利于词汇层面的优化。最终,作者开源了整套系统,并实现在六种语言对上的实时翻译演示。
链接: https://arxiv.org/abs/2610.07545
作者: Zubair Ajmal Farooq,Diptesh Kanojia
机构: University of Surrey (萨里大学); United Kingdom
类目: Computation and Language (cs.CL)
备注: 7 pages, 2 figures, 4 tables. Full paper submitted to the Convergence 2026 proceedings; poster presented at Convergence 2026, University of Surrey. Code: this https URL
Abstract:We present a fully offline speech-to-speech translation pipeline that runs on a Jetson Nano (4 GB) and corrects its own weak translations without retraining. A Whisper-tiny ASR feeds an Opus-MT translator; multilingual BERT cosine similarity acts as a Quality Estimation (QE) gate, triggering a secondary-pass correction when confidence falls below a pre-defined threshold \tau . We compare three correction methods: QE reranking (M1), Minimum Bayes-Risk decoding (M2), and constrained beam search (M3). On 1,012 FLORES-200 sentences (English-Spanish), M2 at \tau=0.90 produces statistically significant improvements over greedy decoding on BLEU (+0.67, p0.001), ChrF (+0.51, p0.001), and COMET (+0.0020 at N=3, p=0.002); M1 yields no significant gains, and M3 is significantly worse than baseline (p0.99). Our central finding is that QE functions effectively as a gate but poorly as a ranker: removing the QE model from candidate selection (M1 \to M2) does not hurt quality and frees 680 MB from the critical path. Using a gain-to-edit ratio adapted from the post-editing-effort literature, we further show that smaller candidate pools (N=3) yield more surgical corrections with better semantic adequacy, while larger pools (N=10) maximise lexical reward. We release the system and demonstrate live translation across six language pairs.
[NLP-86] Safeguarding LLM s via Model-Agnostic Latent Safety Signals from Dark Knowledge
【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在安全防御中面临的两大核心问题:一是现有解码阶段的防御方法普遍存在安全与过度拒绝(over-refusal)之间的权衡,即强化安全性会损害模型对良性查询的有用性;二是多数方法依赖于特定模型架构的内部隐藏状态,导致泛化能力差且计算开销大。针对上述局限,本文提出一种名为LADE(Latent Safety Signals for Defense)的新方案,其关键在于利用“暗知识”(dark knowledge,即输出概率分布中超出argmax之外的信息)来提取跨模型一致的潜在安全信号。具体而言,通过对比有害与良性输入在首个词元输出概率分布中的差异,识别出在不同查询类型下概率差异显著的词元,这些词元构成了一种由安全对齐所催生的、模型无关的安全方向。LADE的核心创新包括:(1)从首个词元的概率分布中提取潜在安全词元;(2)通过词器映射(Tokenizer Mapping)实现跨不同分词器的通用性;(3)基于kNN的判别机制,对查询进行分类。实验表明,LADE在多种主流大模型和攻击基准上均能有效抵御各类越狱攻击,显著降低攻击成功率,同时保持良好的安全-效用平衡。
链接: https://arxiv.org/abs/2610.07532
作者: Wonjun Lee,Kyungsik Yang,Gaeun Ji,Vaidehi Patil,Haon Park,Bumsub Ham,Mohit Bansal,Suhyun Kim
机构: UNC Chapel Hill(北卡罗来纳大学教堂山分校); Kyung Hee University(庆熙大学); AIM Intelligence; Yonsei University(延世大学)
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: project page: this https URL
Abstract:LLMs have advanced rapidly, raising growing concerns about their safety. Recent work has proposed approaches to detect and defend against attacks including defenses at decoding stage that leverage models’ hidden states. However, existing decoding-stage defenses suffer from two limitations. First, they introduce a trade-off between safety and over-refusal, where strengthening safety degrades the model’s helpfulness on benign queries. Second, many of these methods rely on internal hidden states and are thus restricted to specific architectures, incurring substantial overhead and limited generalization across models. To address these limitations, we introduce LADE (Latent Safety Signals for Defense), which leverages latent safety signals extracted by contrasting harmful and benign queries from dark knowledge (i.e., information carried by the output probability distribution beyond its argmax) in the first-token output probability distribution. Our key insight is that, beyond surface-level refusal tokens, the dark knowledge in the first-token distribution contains latent safety signals, defined as tokens whose probabilities differ sharply between harmful and benign queries. We show that these signals consistently align across LLMs, forming a model-agnostic direction that emerges from safety alignment. LADE consists of three components: (1) Extracting Latent Safety Signals from Dark Knowledge, which selects top-k safety-discriminative tokens from the first-token probability distribution; (2) Tokenizer Mapping, which maps these tokens across different tokenizers to enable model-agnostic application; and (3) kNN-based Discrimination, which classifies queries via a k-Nearest Neighbors search over the mapped tokens. Across diverse LLMs and benchmarks, LADE is robust against a wide range of jailbreak attacks and lowers attack success rates while maintaining a competitive safety-utility trade-off.
[NLP-87] Not What a Child Expressed: Auditing the Sign-to-Text Safety Interface in Child-Facing AI NEURIPS2026
【速读】: 该论文旨在解决生成式 AI(Generative AI)在儿童手语自动翻译(SLT)应用中潜在的安全风险问题,特别是当未充分训练或评估的模型将聋人儿童的手语翻译为文本时,可能因翻译错误导致安全决策失误。其核心解决方案在于提出一种由聋人知情(Deaf-informed)的部署前审计框架,通过构建故障分类体系、去标识化场景模板、四种对比条件及四项结果评估指标,系统性地评估手语翻译在安全过滤与内容审核边界上的可靠性。该框架重点关注翻译错误对否定、参与者角色、保密性、紧急程度及求助意图等关键语义维度的影响,以确保儿童使用手语交互时,其安全防护机制不会因翻译偏差而失效。首个案例研究将聚焦澳大利亚手语(Auslan)。
链接: https://arxiv.org/abs/2610.07519
作者: Muhammad Rafiullah Memon,Viet Vo,Wanlun Ma,Yang Xiang
机构: Swinburne University of Technology, Hawthorn, VIC, Australia(斯威本科技大学,霍桑,维多利亚州,澳大利亚)
类目: Computation and Language (cs.CL); Cryptography and Security (cs.CR); Computers and Society (cs.CY)
备注: 5 pages, 1 figure. Accepted as a poster at the NeurIPS 2026 Workshop on Child Safety in AI (non-archival)
Abstract:Automatic sign language translation (SLT) has entered consumer products, turning American Sign Language into English text for dictation, messaging, and queries put to a conversational assistant. Child-facing AI and platform trust-and-safety tooling decide on text, using filters on minor accounts and grooming classifiers that score chat messages. A signing child who uses SLT therefore reaches these safeguards through a translation. We found no publicly documented system in which the two have been jointly evaluated, and the leading deployed SLT model was neither trained nor formally evaluated on signers under 18. Errors that alter negation, participant roles, secrecy, urgency or help-seeking could change a safety decision without disturbing fluency. This paper proposes a Deaf-informed pre-deployment audit of that boundary, with a failure taxonomy, a sanitised scenario schema, four comparison conditions, and four outcome measures. Auslan is the planned first case study.
[NLP-88] On Open-Ended Information Seeking for Information Elicitation Agents
【速读】: 该论文旨在解决生成式人工智能在开放性信息获取(information elicitation)任务中,不同大语言模型(Large Language Models, LLMs)因自身对信息价值判断的差异而引发的序列化信息搜寻行为异质性问题。核心挑战在于:尽管代理式信息获取(agentic elicitation)可将决策权交由基础模型,但模型间在信息价值评估上的隐含偏好差异如何影响整体信息探索路径尚不明确。其解决方案的关键在于构建一个受控的模拟环境,在该环境中多个LLM面对相同的初始信息空间并遵循统一的选择规则,从而剥离问题生成与回应者行为等干扰因素,仅聚焦于模型自身对信息价值的判断机制。通过系统比较11种跨模型家族与参数规模的LLM在该设定下的表现,研究揭示了各模型在“广度-深度”探索策略上的固有偏好,并进一步分析交互历史如何动态调节信息评估与选择过程。研究通过敏感性分析与消融实验验证了结论的稳健性,涵盖可获取机会、响应标签设计、交互历史存在性及冗余性是否被显式考虑等因素,为理解模型选择对信息获取行为的影响提供了实证依据与方法论框架。
链接: https://arxiv.org/abs/2610.07509
作者: Victor De Lima,Grace Hui Yang
机构: InfoSense Lab, Georgetown University (乔治城大学)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注:
Abstract:Information elicitation is an open-ended information-seeking problem in which an interaction can unfold in many potentially valuable directions, requiring an elicitor to continually determine which information to pursue as new information emerges. In agentic elicitation, these decisions may be delegated to a foundation model, yet how model choice shapes the resulting information-seeking behavior remains understudied. We study how judgments about information value vary across LLMs and how these differences shape sequential information seeking. We first examine these judgments across 11 LLMs spanning multiple model families and parameter scales, using a shared set of information and elicitation objectives. We then develop a controlled elicitation simulation in which different models encounter the same information space and use the same selection rule, isolating these judgments from question generation and respondent behavior. Using this setting, we characterize the breadth-depth behavior that emerges from model-specific information-seeking preferences over the course of elicitation. We further examine how interaction history changes the evaluation and subsequent selection of prospective information. We test the robustness and boundaries of these findings through sensitivity analyses and ablations over the opportunities available to the elicitor, the response labels used to operationalize information-seeking preferences, the presence of interaction history, and whether redundancy is explicitly relevant to the assessment. The project code, data, and trajectory files are available at this https URL.
[NLP-89] Closing Ambient Clinical Documentation Gaps with Automated Provider Queries
【速读】: 该论文旨在解决临床文档中因信息缺失导致的记录不完整问题,具体针对临床文档专家(clinical documentation specialists)向医师发送澄清请求(provider queries)以填补临床笔记空白、确保准确编码与 billing 的“查询循环”(query loop)自动化难题。现有研究虽已实现病历草稿生成、ICD-10编码及医嘱提取的自动化,但均基于完整病历转录文本假设,未处理实际临床场景中普遍存在的信息缺失问题。本文提出一种基于大语言模型(LLM)的自动化框架——DAU(Draft, Ask, Update),通过模拟人工查询流程,在病历起草、编码和医嘱提取任务中动态识别并补全缺失信息。其解决方案的关键在于:首先基于对3,000例真实就诊记录的审计,构建了五个公开数据上的病历退化基准(transcript-degradation benchmarks),揭示了文档缺失的主要来源;其次通过对21,000次真实澄清对话的分析,发现有效提问的预测因子具有任务特异性——在病历完整性补全任务中,简单回忆式问题即可满足需求;而在ICD-10编码任务中则需更复杂的多选项问题以提升准确性;同时识别出约9%的提问会损害性能,主要源于冗余问题与无效回复仍触发重写逻辑。因此,系统成功部署的关键不仅在于“如何提问”,更在于学习“何时不应提问”(learning “when not” as much as “what to” ask),即实现对提问必要性的精准判断,从而避免不必要的干预与性能下降。
链接: https://arxiv.org/abs/2610.07502
作者: Joseph Paul Cohen,Raj Shah,Han-Chin Shing,Fang Wang,Susan Nguyen,Chaitanya Shivade,Jack Moriarty
机构: 未知
类目: Computation and Language (cs.CL)
备注:
Abstract:Provider queries are clarifying requests sent by clinical documentation specialists to physicians to close gaps in the clinical note and ensure accurate billing. Prior work automates note drafting, ICD-10 coding, and order extraction assuming a complete transcript, leaving these gaps unaddressed. We study whether an LLM can automate the query loop, termed DAU (Draft, Ask, Update), across those three tasks. An audit of 3,000 real visits identifies the sources of missing documentation, from which we build five transcript-degradation benchmarks on public data. Analyzing 21k clarification turns on real conversations, we find useful-question predictors are task-specific: oracle confidence dominates, but note completeness needs only simple recall questions while ICD-10 coding needs harder, multi-option ones. About 9% of turns hurt performance, driven by redundant questions and non-answers that still trigger a rewrite. Deployment depends on learning “when not” as much as “what to” ask.
[NLP-90] In With the Old: Enhancing Classical Document Automation with Generative AI
【速读】: 该论文旨在解决法律文本生成与处理中因非专业人士表达不准确或不规范而导致的法律文书质量下降问题。其核心挑战在于如何在保持法律严谨性的同时,提升普通用户在撰写法律文件时的准确性与合规性。解决方案的关键在于融合基于专家系统(expert system)等符号主义方法的文档自动化技术与当前生成式人工智能(Generative AI)的能力,利用大语言模型(Large Language Models, LLMs)对普通人撰写的文本进行自动识别与修正,从而实现从“自然语言输入”到“合规法律文本输出”的高效转换。初步实验表明,生成式AI能够有效识别并修复非专业用户文本中的逻辑漏洞、术语误用及格式问题,显著提升了法律文书的可读性与合法性。
链接: https://arxiv.org/abs/2610.07480
作者: Marc Lauritsen,Hannes Westermann
机构: 未知
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: 10 pages. AI for Access to Justice Workshop at ICAIL 2025
Abstract:Software-based legal assistance systems have leveraged many different forms of knowledge representation and reasoning. This article explores how document automation services rooted in expert system style and other symbolic approaches can usefully enhance and be enhanced by current generative AI approaches. We discuss the possible benefits and challenges, and report on preliminary experiments in using large language models to identify and fix issues in texts written by laypeople.
[NLP-91] AlignQuant: Tile-Aligned Mixed-Precision Quantization for Efficient LLM Generation
【速读】: 该论文旨在解决大语言模型(Large Language Model, LLM)在推理过程中因细粒度混合精度量化(fine-grained mixed-precision quantization)与GPU硬件存储及计算单元固有结构不匹配而导致的压缩效率难以转化为实际加速的问题。其核心挑战在于,局部精度分配策略常与GPU的二维内存布局和计算单元对齐要求冲突,从而限制了量化带来的性能收益。为此,论文提出AlignQuant——一种后训练量化方法,以GPU兼容的二维权重重叠块(two-dimensional weight tiles)作为精度分配、紧凑存储与执行的统一单位,使精度可随输出通道敏感性动态调整。通过联合预填充(prefill)与解码(decode)阶段的校准,利用量化激活下的语言模型损失梯度加权的投影输出扰动来评估精度降低的合理性,并引入相位归一化评分机制,在全局权重存储预算约束下优先为对任一阶段更为关键的块保留更高精度。每个重叠块仅存储一个选定的表示形式,而针对不同阶段优化的内核则复用打包后的模型,并将低比特权重展开以支持INT8计算与8位激活的高效运算。在涵盖3B至14B参数量的四类主流LLM上,AlignQuant相较BF16实现了最高2.50倍的生成速度提升,同时保持模型质量不变;实验覆盖三种GPU架构及最长64K token的上下文长度。结果表明,通过共享的二维块单元设计,局部精度灵活性与标准GPU执行模式可实现有效共存。该方案的开源实现已发布于指定链接。
链接: https://arxiv.org/abs/2610.07457
作者: Hanzhi Zhang,Qiao Zhang,Qinglei Cao,Heng Fan,Yan Huang,Kewei Sha,Yunhe Feng
机构: LLaVi Lab, Computer Science and Engineering, University of North Texas(北德克萨斯大学); Computer Science, Saint Louis University(圣路易斯大学); Data Science, University of North Texas(北德克萨斯大学)
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注:
Abstract:Fine-grained mixed-precision quantization promises efficient large language model inference, but local precision choices can conflict with regular GPU storage and computation units. This precision-boundary mismatch limits the translation of compression into practical acceleration. We introduce AlignQuant, a post-training quantization method that uses GPU-compatible two-dimensional weight tiles as the common unit of precision allocation, compact storage, and execution. This shared partition lets precision follow sensitivity within output channels. Joint prefill/decode calibration scores precision reductions using projection-output perturbations weighted by language-model loss gradients under quantized activations. Phase-normalized scores prioritize higher precision for tiles important to either phase under a model-wide weight-storage budget. Each tile stores one selected representation, while phase-specialized kernels reuse the packed model and expand lower-bit weights for INT8 computation with 8-bit activations. Across four LLMs spanning 3B to 14B parameters, AlignQuant achieves up to 2.50\times generation speedup over BF16 while preserving model quality. Evaluations further cover three GPUs and contexts up to 64K tokens. These results show that local precision flexibility and regular GPU execution can coexist through a shared tile unit. The implementation is available at this https URL.
[NLP-92] AccentCL: Robust Accent Classification with Incremental Expansion
【速读】: 该论文旨在解决英语口音分类模型在面对新口音类别时难以动态扩展、且在存在类别不平衡和跨语料库领域偏移(domain shift)情况下性能下降的问题。现有方法通常依赖固定标签空间,无法适应随时间新增的口音类型,同时在实际数据中常见的多数类主导现象和不同录音条件导致的分布差异会严重削弱模型泛化能力。其解决方案的关键在于提出一种类增量学习(class-incremental learning)框架AccentCL,该框架基于冻结的Whisper-Large-v3编码器提取多层特征表示,并通过两种核心机制实现鲁棒性:一是采用考虑类别不平衡的交叉熵损失(imbalance-aware cross-entropy loss),以缓解对主流口音类别的偏差;二是引入领域均值对齐损失(domain mean alignment loss),有效减少不同语料间特征分布的均值偏移。此外,为支持新口音类别的持续加入,AccentCL采用基于回放的持续学习策略,利用冻结的基础模型保持已有知识,并设计旧到新边界损失(old-to-new margin loss)以抑制对新类别的过度预测。实验表明,在五类口音分类任务中,AccentCL达到77.1%的平衡准确率与76.9%的宏平均F1分数;在依次引入西班牙口音英语和中文口音英语后,分别获得83.3%和61.8%的新类别F1分数,同时维持对原有类别的高保真度(77.3%和77.6%平衡准确率),验证了其无需全量重训练即可实现高效、稳定的新口音类别增量添加的能力。
链接: https://arxiv.org/abs/2610.07426
作者: Mu-Ruei Tseng,Waris Quamer,Ghady Nasrallah,Ricardo Gutierrez-Osuna
机构: Texas A&M University (德州农工大学)
类目: Computation and Language (cs.CL); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
备注: Published in Proceedings of IEEE Spoken Language Technology Workshop (SLT) 2026
Abstract:Accent classifiers are typically trained with a fixed label inventory and cannot accommodate new accent categories as new data becomes available. Moreover, accented speech corpora often exhibit substantial class imbalance and/or domain shift due to differences in recording conditions across corpora. We present AccentCL, a class-incremental learning framework for English accent classification that is robust to class imbalance and cross-corpus domain shift. AccentCL extracts multi-layer representations from a frozen Whisper-Large-v3 encoder, optimized with an imbalance-aware cross-entropy loss to reduce bias toward the majority accent classes and a domain mean alignment loss that minimizes distributional mean shift across training corpora. The label space is then expanded via replay-based continual learning, using the frozen base model for knowledge retention and an old-to-new margin loss to reduce overprediction on newly added classes. On a five-class accent classification task, AccentCL achieves 77.1% balanced accuracy and a 76.9% macro-averaged F1 score. We further evaluate the model’s ability to incrementally incorporate two new accent categories: Spanish-accented and Chinese-accented English. When adding Spanish-accented English to the pretrained model, AccentCL attains an F1 of 83.3% on the new class while retaining 77.3% balanced accuracy on the base classes. When subsequently adding Chinese-accented English, it achieves 61.8% F1 on the new class while preserving 77.6% balanced accuracy on the previously learned classes. These results show that AccentCL enables robust regional accent classification while allowing new accent categories to be added without full retraining.
[NLP-93] Who Wrote It Is Not Enough: Detecting Who Contributed the Insight
【速读】: 该论文旨在解决大语言模型(LLM)在科学写作与同行评审中日益广泛应用背景下,仅识别文本作者已不足以判断其智力贡献来源的问题,核心挑战在于区分评审意见中的洞见(insight)究竟源自人类、大语言模型,还是二者混合贡献。解决方案的关键在于提出“洞见溯源”(Insight Provenance)任务,并构建首个基于4,057篇科学论文与12,660条人工评审的标注数据集InsightProv-v0,通过模拟GPT-4o、Gemini和DeepSeek等模型在不同参与程度下的介入情况,在句子级别标注洞见来源。研究发现,模型在原始数据上表现优异可能源于对语言风格和文本作者特征的捷径依赖,此类捷径信号在去偏评估下迅速失效。为此,论文提出一种两阶段对抗性框架,有效抑制表面捷径信号的同时保留与洞见来源相关的核心信息。进一步分析揭示,智力原创性的可识别性主要依赖于论文上下文锚定(paper grounding)与邻近评审语境(neighboring review context)所提供的互补性信号;人类、混合及AI生成的洞见在信息来源与失效模式上存在系统性差异——其中,AI洞见多局限于通用或论文内提供的信息,而人类洞见更倾向于引入外部知识与独立判断。这一发现表明,尽管语言表达可被大模型重构,但思想本身的源头仍会留下深层且持久的可识别痕迹。
链接: https://arxiv.org/abs/2610.07365
作者: Zhuoyang Zou,Abolfazl Ansari,Jiaxi Yang,Delvin Ce Zhang,Qian Chen,Dongwon Lee,Wenpeng Yin
机构: Pennsylvania State University (宾夕法尼亚州立大学); University of Sheffield (谢菲尔德大学)
类目: Computation and Language (cs.CL)
备注:
Abstract:As LLMs increasingly assist scientific writing and peer review, detecting who wrote the text is no longer sufficient: we need to determine who contributed the underlying insight. We introduce Insight Provenance, the task of identifying whether a review insight originates from a human, an LLM, or their hybrid contribution. We construct InsightProv-v0 from 4,057 scientific papers and 12,660 human reviews, simulating different levels of LLM involvement with GPT-4o, Gemini, and DeepSeek and annotating provenance at the sentence level. We show that strong performance on raw data can be misleading, as models exploit linguistic and textual-authorship shortcuts that degrade substantially under progressively debiased evaluation. We therefore propose a two-stage adversarial framework that suppresses shortcut signals while preserving provenance-relevant information. Beyond detection, extensive analyses reveal what makes intellectual authorship identifiable: paper grounding and neighboring review context provide complementary provenance signals, while human, hybrid, and AI insights systematically differ in their information sources and failure modes. Most strikingly, AI insights predominantly remain close to generic or paper-provided information, whereas human insights more often introduce external knowledge and independent judgment. These findings suggest that while wording can be rewritten by an LLM, the provenance of an idea leaves a deeper and more persistent signal.
[NLP-94] racking Is Not Permanence: What Video World Models Keep of a Hidden Object
【速读】: 该论文旨在探究生成式视频世界模型(Video World Models)在对象不可见时仍能保留何种信息,即模型如何处理被遮挡或隐藏的物体。其核心问题是:当一个物体被遮蔽后,模型的预测机制是否能维持对该物体的持续表征,以及这种表征的持久性与物理常识(如物体恒常性、容器封闭性等)之间的关系。解决方案的关键在于通过对比冻结的V-JEPA 2预测器对遮挡区域的预测结果与编码器对两种仅在遮挡区域内存在差异的世界状态的表示,揭示预测器在隐空间中对隐藏物体的保留能力。研究发现,预测器对静止物体的部分表征可保持约0.3秒(在自掩码条件下为0.5秒),而对移动物体的表征则在0.3秒内丢失;相比之下,编码器能持续读取物体存在性达1.0秒,且封闭容器内的内容可解码长达3.5秒。尽管预测器输出的表征强度显著低于基线(仅为基线的14–60%),但其仍包含微量信息。通过仅3000步的合成容器训练,模型即可将对隐藏物体的信念从0.05提升至1.00,显著改善了对物理常识的理解。此外,引入管状掩码(tube mask)的持续训练使运动物体的表征延续时间延长至1.1–1.6秒,表明该缺陷并非隐空间预测的本质局限。这一结果表明,通过特定训练策略,可以低成本地在模型中建立“物体恒常性”先验,从而有效弥补预测器在遮挡场景下的信息衰减问题。
链接: https://arxiv.org/abs/2610.07355
作者: Peng Xie,Amr Alanwar
机构: Technical University of Munich(慕尼黑工业大学)
类目: Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL)
备注:
Abstract:Video world models track objects they can see; we ask what they keep of objects they cannot. We hide an object from a frozen V-JEPA 2 predictor and compare its prediction for the hidden region with the encoder’s representation of two worlds that differ only inside that region. The predictor’s decision keeps a stationary object in part and one carried inside a container not at all, and loses a moving one within 0.3 s (0.5 s under V-JEPA’s own tube mask; ViT-H keeps it to 1.1 s at pretraining’s 90% masking ratio); in projection a trace remains, below the midpoint, at 14-60% of what a baseline copying the last view retains. The information is there: the encoder reads the object’s presence at 1.00 and keeps a closed container’s contents decodable for 3.5 s, while the predictor’s output, read with the encoder’s own probe, contains the ball in 2% of scenes once the box has been closed for half a second. On rendered scenes, permanence is missing on the predictor’s side, and training installs it cheaply as a prior: three thousand predictor-only steps on synthetic containers take this belief from 0.05 to 1.00 against two matched controls. They also raise IntPhys-2019 from 84.2% to 93.3%, but so does a curriculum without containers, and which training habit the benchmark credits changes with its scoring rule. Continued training with tube masks produces 1.1-1.6 s of moving-object carry-over on manipulation and internet-style video, so the deficit is not intrinsic to latent prediction. VideoMAE keeps almost nothing, and Cosmos’s next-token prediction keeps a stationary hidden object but not one carried inside a moving container.
[NLP-95] Stepped MoE: Segment-Level Routing with Configurable Inference Complexity
【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在多样化部署场景中面临的资源消耗高与计算约束严苛之间的矛盾问题,尤其针对边缘设备上进行本地化推理时所受内存和算力限制的挑战。现有方法通常将弹性架构(elastic architectures)与稀疏激活模型(sparsely activated models)作为独立方案处理,难以同时兼顾部署灵活性与任务自适应性。本文提出一种统一框架,将弹性结构与稀疏门控架构相结合,构建能够同时响应部署约束(如设备内存、计算能力)与任务需求的自适应模型。其核心创新在于:采用一个以上下文和目标效率规格为条件的模型主干,实现推理时对精度-效率权衡的细粒度控制;模型可动态激活嵌套于弹性子网络中的任务相关参数,使单一模型在不同参数规模(10亿至40亿参数)间灵活切换,同时保持输入自适应路由机制。实验表明,该模型在知识密集型基准测试中相比同规模稠密模型提升2%-5%准确率,且在延迟表现上与静态模型相当,同时通过共享参数显著降低设备存储开销,并支持基于可用DRAM与算力动态调整服务策略,从而在性能与资源效率之间实现更优平衡。
链接: https://arxiv.org/abs/2610.07348
作者: Arnav Kundu,Zhaoyang Xu,Bairu Hou,Chang Gao,Reed Li,Tao Lei
机构: Apple(苹果)
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: Apple Foundation Models, 15 Pages, Edge LLMs
Abstract:Training large language models (LLMs) is resource-intensive, and adapting them for diverse deployment scenarios with varying computational constraints remains challenging. While elastic architectures enable flexible model deployment and sparsely activated models allow input-adaptive computation, existing approaches treat these dimensions independently. Moreover, models catered towards on-device edge inference need to conform to the memory and compute limitations of the serving devices. In this paper, we introduce a unified framework that combines elastic structures with sparsely gated architectures to create models that adapt simultaneously to both deployment constraints and task requirements. Our approach employs a model backbone that conditions on both the context and target efficiency specifications, enabling fine-grained control over the accuracy-efficiency trade-off at inference time. The model learns to activate task-relevant parameters within elastically-nested sub-networks, allowing a single model to span multiple capacity points while maintaining input-adaptive routing. Through experiments we demonstrate that we can create a model that allows the flexibility to use 1,2,3,4 billion parameters while being more accurate than their dense counter-parts (2-5% on knowledge-intensive benchmarks) and at par with their static versions while delivering similar latency metrics as dense models. Overall, we save on device disk space by sharing the model parameters, allow flexibility of serving based on DRAM and compute available while delivering more accurate results.
[NLP-96] A doctrine-grounded visual question answering dataset for Tactical Combat Casualty Care
【速读】: 该论文旨在解决战术战伤救护(Tactical Combat Casualty Care, TC3)中响应者将视觉观察到的伤情与干预措施与既定临床指南进行关联的难题,其核心挑战在于构建能够理解视觉证据并匹配可追溯临床教义的视觉-语言模型。解决方案的关键在于提出TC3-VQA数据集,该数据集基于公开的教学与实战TC3视频及权威TC3文档构建,涵盖11个临床概念、581个样本和1,860个问题,覆盖干预识别、教义内容、临床推理、操作指引以及在视觉信息不足时的拒绝判断等任务。数据集通过保留原始教义文本片段及其字符偏移量实现教义驱动的答案标注,并结合视觉标注、段落检索、蕴含关系验证及跨模型族交叉验证等方法确保质量。同时,配套提供装备箱、解剖标签、时间片段和来源元数据,支持对视觉证据与临床知识之间关联性的研究。自动化审计及两名医师与两名医学生的人工评分进一步量化了标注质量,为视觉-语言模型在TC3场景下的适配、识别性能评估、教义召回能力分析及合理拒答行为研究提供了高质量基准资源。
链接: https://arxiv.org/abs/2610.07339
作者: Junseob Kim,Jade Chng,Ayman Ali,Victor Moas,Yichun Lee,Po-Chun Chin,Sunil Hwang,Rishikesan Kamaleswaran
机构: Duke University (杜克大学); National Yang Ming Chiao Tung University (国立阳明交通大学); Korea Military Academy (韩国军事学院)
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: 20 pages, 6 figures, 5 tables. Dataset: this https URL code: this https URL
Abstract:Tactical Combat Casualty Care (TC3) requires responders to connect visual observations of injuries and interventions with established clinical guidance. Developing vision-language models to support this process requires supervision that links visible evidence to traceable doctrine. We present TC3-VQA, a dataset constructed from public instructional and field TC3 videos and authoritative TC3 documents. It contains 581 items spanning 11 concepts, with 1,860 questions covering intervention recognition, doctrine, clinical reasoning, procedural guidance, and refusal when visual information is insufficient. Doctrine-based answers preserve verbatim source passages and character offsets. Construction combines visual annotation, passage retrieval, entailment checks, and verification across model families. Equipment boxes, anatomical labels, temporal segments, and source metadata accompany the question-answer pairs. Automated audits and ratings by two physicians and two medical students characterize annotation quality, with human ratings available for 88 retained items. The dataset provides a resource for adapting vision-language models to TC3, studying the connection between visual evidence and clinical knowledge, and evaluating recognition, doctrine recall, and abstention.
[NLP-97] Structuring MoE Expert Selection for Agent ic Reinforcement Learning
【速读】: 该论文旨在解决长时序大语言模型(LLM)智能体在采用稀疏专家混合(MoE)架构时,智能体行为与专家选择结构之间缺乏协同设计的问题。现有方法中,尽管MoE的专家路由机制在实践中表现出与智能体操作轨迹高度契合的局部专业化特性(如相同语义操作如“读取”“更新”之间的专家选择重叠度更高),但标准强化学习(RL)训练算法未能有效利用这一结构性信号,导致专家路由在后训练过程中失控,进而限制了任务成功率和推理效率。为此,论文提出一种分层路由控制框架,其核心在于:在任务层面显式引导每一轮交互中的专家选择与具体智能体操作对齐,同时在令牌层面通过正则化手段维持局部一致性;为解决训练过程中的稳定性问题,进一步引入熵门控控制机制。实验表明,该框架在所有评估基准上均实现超过10个百分点的成功率提升,验证了智能体轨迹的结构信息可作为优化MoE容量的有效信号。
链接: https://arxiv.org/abs/2610.07332
作者: Bolian Li,Ting-Yao Hu,Cheng-Yu Hsieh,Sanjoy Chowdhury,Oncel Tuzel,Raviteja Vemulapalli
机构: 未知
类目: Machine Learning (cs.LG); Computation and Language (cs.CL)
备注:
Abstract:Long-horizon LLM agents are frequently implemented using sparse mixture-of-experts (MoE) models, yet the co-design of agentic behavior and MoE structures remains underexplored. In this work, we comprehensively study the connections between agentic post-training and MoE expert selection. In off-the-shelf MoE models, we observe expert selection exhibits a specialized structure that naturally aligns with agentic trajectories. Specifically, expert routing overlaps more between turns where the agent performs semantically similar operations (e.g., READ, UPDATE) than between turns with differing operations. However, standard RL algorithms ignore this specialization, allowing the MoE routing to go uncontrolled during training, which empirically limit task performance and inference efficiency. To address this, we introduce a hierarchical routing control framework for agentic tasks. We explicitly encourage turn-level expert selections to align with agentic operations while regularizing token-level expert selections to maintain local consistency. To resolve stability issues that arise during post-training with the proposed methods, we further introduce an entropy-gated control mechanism. Overall, our routing control framework achieves over 10-point improvements in success rate on all evaluated benchmarks. These results demonstrate that agentic trajectory structure provides an effective signal for optimizing MoE capacity during RL post-training.
[NLP-98] SharedKV-BT: Node-Local Typed Decisions for Behavior-Tree Agents DATE
【速读】: 该论文旨在解决智能体在执行复杂任务时,如何高效且准确地处理一系列相互依赖的决策问题。传统自回归模型虽能提供灵活的决策接口,但存在逐标记生成带来的延迟问题;尽管近期共享前缀方法通过复用编码上下文并行评分多个决策以降低计算开销,却无法建模决策间的依赖关系,也缺乏对执行过程的验证机制。为此,本文提出SharedKV-BT框架,其核心在于将行为树(Behavior Tree, BT)中每个活跃节点设计为具备局部状态字段与候选动作集,并利用共享键值(Shared-KV)并行评估这些候选动作,再将选定决策传递至独立执行系统。该方案的关键创新在于:1)通过节点局部的共享键值机制实现并行决策评分,显著提升推理效率;2)引入阶段门控(stage gating)和外部后条件(external postconditions)机制,有效防止动作顺序错乱与过早终止,保障决策序列的正确性与闭环执行成功率。实验结果表明,在机器人操作、移动导航及计算机使用任务中,SharedKV-BT相比提示匹配的自回归解码提速2.36–4.15倍,且在操作任务中将联合决策准确率从75%提升至94%,闭环成功率达60%。
链接: https://arxiv.org/abs/2610.07327
作者: Naoki Wake,Justin Wagle
机构: Microsoft(微软)
类目: Robotics (cs.RO); Computation and Language (cs.CL)
备注: 8 pages, 5 figures, 1 table. Last updated on October 5th, 2026
Abstract:Agent tasks require sequences of interdependent decisions. Autoregressive models support more flexible decision interfaces than conventional classifiers but incur the latency of token-by-token generation. Recent shared-prefix methods reduce this cost by reusing encoded context and scoring multiple decisions in parallel, but do not model decision dependencies or verify execution. We propose SharedKV-BT, where each active node of a behavior tree (BT) exposes stage-local fields and candidates, and Shared-KV scores the candidates in parallel and passes the selected decision to a separate execution system. We tested SharedKV-BT on robot manipulation, mobile navigation, and computer-use tasks. Across three tasks, SharedKV-BT made typed decisions 2.36-4.15 times faster than prompt-matched autoregressive decoding. On the manipulation task, node-local Shared-KV improved joint decision accuracy from 75% to 94% and closed-loop success from 0% to 60%. Fixed-score policy replay showed that stage gating prevented out-of-order actions and external postconditions prevented premature completion.
[NLP-99] Kurate: Scalable Scientific Quality Analysis
【速读】: 该论文旨在解决科学文献检索系统在识别相关研究论文的同时,无法有效评估其证据质量的问题。现有系统虽能定位与问题相关的论文,但缺乏对研究设计严谨性、报告透明度及可重复性的综合评价能力。为此,论文提出Kurate系统,其核心解决方案是利用大语言模型(Large Language Models, LLMs)对已发表研究的证据质量进行自动化评估。该系统不仅分析论文本身,还整合研究相关的补充文档(如临床试验注册信息和研究方案),并将每一项质量判断明确关联至原文中支持该判断的具体文本片段。通过在包含4,347篇论文(其中3,913篇为随机对照试验)的语料库上应用该系统,对研究设计与报告质量的8个维度——统计功效、因果识别、预先注册、选择性报告、测量有效性、分析预设、报告透明度以及利益冲突与资助披露——进行评分,发现统计功效不足、选择性报告和分析预设缺失是最常见的问题,且不同临床领域平均质量存在差异。与60份独立临床试验文档的专家标注结果对比显示,Kurate在协议评分点上的准确率(AC1)达0.94,在结果发布评分点上为0.81,表明其评估结果具有高度可靠性。通过一个高质量临床试验的案例剖析,进一步展示了单篇论文的整体评分如何分解为多个可追溯至原始证据的具体判断。综上,该研究证明了基于大语言模型的大规模研究质量评估在技术上是可行的,并可为元科学(meta-scientific)研究提供有力工具。
链接: https://arxiv.org/abs/2610.07306
作者: Matthew J. Vowels,Jamie Cummins
机构: Kivira Health(美国); University of Bern(伯尔尼大学)
类目: Computation and Language (cs.CL)
备注:
Abstract:Scientific search systems can find papers that are relevant to a question, but they generally do not assess the quality of the evidence that those papers provide. We present Kurate, a system that uses large language models (LLMs) to assess the quality of published studies. Kurate uses both the paper and its related documents (e.g., the study’s trial registration and protocol), and links each of its judgments to the passage of text on which that judgment is based. We applied Kurate to a corpus of 4,347 papers (3,913 of which report randomized trials) and scored each paper on 8 dimensions of study design and reporting: specifically, statistical power, causal identification, preregistration, selective reporting, measurement validity, analysis prespecification, reporting transparency, and conflict of interest and funding. Across the corpus, we found that papers most often exhibited issues with statistical power, selective reporting, and analysis prespecification, although average quality differed between clinical areas. When compared against expert annotations of 60 held-out clinical-trial documents, the information Kurate extracted matched the expert label in 221/242 protocol scorepoints and 294/370 results-publication scorepoints, with AC1 0.94 and 0.81, respectively. Using a well-reputed, high quality clinical trial as a worked example, we show how a single paper’s overall grade breaks down into separate judgments, with each linked to specific evidence from the trial’s registration, protocol, and published report. Together, these results show that large-scale quality assessment of this kind is feasible, and that it can be used to address meta-scientific research questions.
[NLP-100] Lineage-Aware Memory Governance: A Derivation-Gated Framework for Privacy-Preserving Column-Level Access Control in Enterprise AI Agents
【速读】: 该论文旨在解决企业级生成式 AI 代理在共享记忆存储时面临的两大核心安全与一致性问题:一是敏感数据可能通过合法计算结果泄露,而请求方无法直接推导出这些结果;二是不同部门可能基于冲突逻辑对同名关键绩效指标(KPI)进行隐性计算,导致结果不一致。现有代理-记忆系统(如 MemGPT、Zep、A-MEM)仅依据内容、所有权和角色进行检索访问控制,缺乏对数据衍生路径(lineage)的管理,因而无法防止缓存结果中嵌入受禁列的敏感信息。为此,论文提出分析型记忆单元(Analytical Memory Unit, AMU),其核心创新在于为每个缓存结果附加完整的衍生路径(lineage)图谱,并设计基于派生路径的检索策略——仅当请求方对结果所涉及的所有列均具备权限时才允许访问。通过构造性证明,该策略可确保在完整记录派生路径的前提下,以最坏时间复杂度 O(n) 阻止从敏感列派生出的结果被越权访问,从而实现对派生特征中隐含敏感信息的实质性防护,而非依赖经验性观察。实验表明,达到90%的派生路径记录覆盖率是消除实际泄露的保守阈值;在六组实验中,该机制将跨部门泄露率从18.8%-25.5%降至零,同时保留81.5%-82.6%的记忆复用率,且最坏情况下的延迟仅为13.8微秒。真实代理原型验证了该机制的有效性:在18轮交互中未发生任何泄露,自动识别出两起逻辑冲突,虽仅为可行性演示,但为共享代理记忆提供了可落地的治理层,与源层访问控制互补,并支持欧盟《人工智能法案》(EU AI Act)合规要求。
链接: https://arxiv.org/abs/2610.07258
作者: Venkata M Sangaraju,Sudhir Vissa
机构: 未知
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
备注:
Abstract:Enterprise AI agents that share a memory store face two unaddressed risks: sensitive data can leak through legitimately computed results the requester could not derive, and departments can silently compute a same-named key performance indicator (KPI) through conflicting logic. Existing agent-memory systems (e.g., MemGPT, Zep, A-MEM) gate retrieval by content, ownership, and role, not derivation, missing a cached insight that embeds a forbidden column. We introduce the Analytical Memory Unit (AMU), a memory schema that attaches a full derivation (lineage) graph to every cached result, gated by a retrieval policy that serves a hit only when the requester is authorised for every column touched. Provided lineage recording is complete, we prove by construction that the policy blocks retrieval of results derived from a sensitive column outside the requester’s permissions, at O(n) worst case – a conditional design guarantee, not an empirical claim, that excludes derived features encoding sensitive information without naming their source. Eliminating measured leakage required 75-90% recorded lineage completeness, so we treat 90% as a conservative deployment target. Across six experiments, lineage-gated retrieval removes the 18.8-25.5% cross-department leakage naive content-gated memory suffers, keeping 81.5-82.6% of memory reuse at 13.8 microsecond worst-case overhead. A real-agent proof-of-concept with LLM-generated SQL is consistent with the guarantee: zero leaks over 9 round-trips, two conflicts caught automatically – though a feasibility demonstration, not evidence of production viability. This offers a practical governance layer for shared agent memory, complementing source-layer access control and supporting EU AI Act compliance.
[NLP-101] Minimal Witness Reinforcement Learning
【速读】: 该论文旨在解决在计算与科学领域中普遍存在的“产生某一结果所必需的不可约条件是什么?”这一核心问题,其本质是识别最小充分原因(minimal sufficient witnesses),这些即为解释、机制与理由。传统强化学习(Reinforcement Learning, RL)方法通常只能发现单一解或冗余的超集,难以满足对多个最小解的识别需求。为此,论文提出最小充分原因识别(Minimal-Witness Identification)的形式化框架,并引入最小充分原因强化学习(Minimal-Witness Reinforcement Learning, MWRL)。MWRL的关键在于基于黑箱验证器反馈设计一种新颖的信用分配机制:通过考察每个候选方案在策略采样中对成功解集并集的贡献度,量化其不可或缺性。该信用分配直接源自问题定义,统一实现了对最小性(minimality)与替代方案恢复(recovery of alternatives)的双重要求。在此原则下,研究者推导出一种值迭代规划算法以完整恢复所有最小原因家族,并提出一种可扩展至大型语言模型的策略梯度方法。实验表明,MWRL能有效识别绝大多数最小原因,而对比方法则常返回冗余超集或仅一个解。通过使原因家族能够从验证器反馈中学习,MWRL将强化学习的应用范围从单解优化拓展至多解解释生成,显著提升了其在复杂推理任务中的表达能力与可解释性。
链接: https://arxiv.org/abs/2610.07226
作者: T. Y. Tsui,Zihao Ye,Pengxiang Cai,Yanchao Li,Yuqiang Li,Zhehong Ai
机构: 未知
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Logic in Computer Science (cs.LO)
备注:
Abstract:``What are the irreducible conditions that are sufficient to produce an outcome?‘’ is one of the most common questions that recur across computation and science. Its answers, the minimal sufficient witnesses, are what we mean by explanations, mechanisms and reasons. These problems usually ask for multiple minimal witnesses, yet standard RL methods may reveal only one solution or redundant ones. We formalize this problem as minimal-witness identification and introduce Minimal-Witness Reinforcement Learning (MWRL). MWRL takes the union of the sets certified by successful proposals sampled from the policy and credits each proposal for the coverage the group union would lose without that proposal. This credit assignment, derived directly from the problem definition, unifies the demands for minimality and recovery of alternatives from a single black-box verifier bit. Under this principle, we derive a value iteration planner that recovers the entire family of witnesses and a policy gradient method that can scale to large language models. Across different experimental settings, MWRL recovers most minimal witnesses, while other methods return redundant supersets or a single witness. By making witness families learnable from verifier feedback, MWRL expands the scope of reinforcement learning beyond single-solution optimization. Our code is available at this https URL.
[NLP-102] IDE 2.0: an open model-agnostic engine for keyed de-identification of clinical notes
【速读】: 该论文旨在解决临床笔记在用于科研前的去标识化(de-identification)难题,核心问题在于传统方法在去除受保护健康信息(PHI)时存在严重缺陷:仅依赖检测无法保留临床语义内容,直接删除或替换标识符会破坏时间序列信息,且频繁更换随机替代值会割裂患者多份病历间的关联性。其解决方案的关键在于提出TIDE 2.0——一个开源的、基于硬件自主可控的双阶段去标识引擎,包含可替换的识别器(recognizer)与密钥驱动的匿名化器(keyed anonymizer)。该系统通过加密方式生成不可逆的替代值,确保同一患者在同一密钥下所有相同实体使用一致的替代符号,同时保持各时间点间的时间间隔不变;不同密钥生成的释放版本之间无法关联,保障了数据再利用的安全性。此外,研究团队还发布了基于大语言模型蒸馏的TIDE2-Sentry识别器,实验证明其在两个机构的金标语料上分别实现了0.88和0.77的跨度级召回率,精度分别为0.88和0.87,且支持按类别报告性能指标。整体方案实现了去标识化过程中的准确性、完整性与隐私安全性的平衡,且允许机构在本地环境中运行、审查与扩展系统。
链接: https://arxiv.org/abs/2610.07224
作者: Jose D. Posada,Somalee Datta,Priya Desai
机构: Technology Digital Solutions, Stanford Medicine, Stanford, CA, USA
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注:
Abstract:Clinical notes capture most of what is documented about a patient’s care, but they cannot be used for research until protected health information (PHI) is removed. De-identification is often treated as a detection problem. Detection alone is not sufficient: redaction strips clinical content along with identifiers, date blanking destroys the temporal intervals needed for longitudinal analysis, and assigning a fresh random surrogate at each occurrence breaks links between a patient’s notes. We present TIDE 2.0, an MIT-licensed engine with two separable stages: an interchangeable recognizer and a keyed anonymizer. Both run on hardware the institution owns. Surrogates are generated cryptographically with no stored linkage table. Dates shift by a per-patient, interval-preserving offset; each value receives the same surrogate across all occurrences under a given key; and a release produced under a new key cannot be linked to earlier releases. We also release TIDE2-Sentry, a recognizer distilled from a large language model. On two gold-annotated corpora from two institutions, the default configuration reached span-level recall of 0.88 in-domain and 0.77 on the second institution’s corpus, at precision 0.88 and 0.87. We report recall and precision per category alongside these aggregates. The engine is open source, and the recognizer is available under a gated research-use agreement, so institutions can run, inspect and extend both within their own environments.
[NLP-103] Forecasting the Growth of Social Media Information Cascades: Towards Human-in-the-Loop Misinformation Triage
【速读】: 该论文旨在解决有限审核团队在信息传播初期难以准确判断哪些新兴谣言(misinformation)具有持续扩散潜力的问题,核心挑战在于如何基于早期传播行为预测其未来的传播规模。解决方案的关键在于构建一个集成多维度早期特征的预测模型:通过分析前30分钟内传播树的节点数量、结构深度熵(structural depth entropy)、时间到达熵(temporal arrival entropy)及其组合特征,实现对后续传播增长的精准预测。实验结果表明,在352个独立测试传播树上,联合模型将对数变换后未来增长的决定系数(R²)从0.307提升至0.323,并使对数平均绝对误差(log-MAE)降低2.4%;在高活跃度传播树子集中,性能提升更为显著,R²上升至0.395,Spearman相关系数ρ从0.394增至0.529,log-MAE下降11.7%。此外,研究还发现前15条回复的语义立场、传播意图与情感状态构成在事件间仍具备跨事件排序能力(ROC-AUC 0.538),引入事实准确性维度后进一步提升至0.562。结合倒数排名融合(reciprocal-rank fusion)与句子重排序技术,在Check-COVID评估中实现了74.2%的黄金证据文档召回率(Top-5)和94.3%的召回率(Top-20),Sentence Reranking达到Recall@20为58.1%。最终提出一种整合早期传播趋势预测、响应模式分析与证据检索结果的综合人工审核系统,以支持基于潜在病毒式传播风险的谣言优先级判别。
链接: https://arxiv.org/abs/2610.07209
作者: Ansh Gupta,Abhiram Gorle,Aayush Rajesh,Tsachy Weissman
机构: The Overlake School(Overlake学校); Stanford University (斯坦福大学)
类目: ocial and Information Networks (cs.SI); Computation and Language (cs.CL)
备注: Work done as part of the SHTEM Internship at Stanford University
Abstract:Limited review teams must identify which emerging claims are likely to keep growing before their eventual reach is known. We center early misinformation triage on this continuation-forecasting problem: predicting subsequent recorded propagation-tree growth from the first 30 minutes of activity. On FibVID, we compare early node count with structural depth entropy, temporal arrival entropy, and their pair while keeping all propagation trees from each original claim in one partition. Across 352 test trees from 59 claim groups separate from training, the combined model raises R^2 for log-transformed future growth from 0.307 to 0.323 and reduces log-MAE by 2.4% (95% claim-bootstrap CI, -0.4% to 5.2%). The gain is especially pronounced among 97 high-activity trees: R^2 rises from 0.248 to 0.395, Spearman’s \rho from 0.394 to 0.529, and log-MAE falls by 11.7% (95% CI, -1.9% to 23.9%). Complementing the 30-minute growth forecast, we analyze the first 15 replies in 563 PHEME threads. In this cohort, the 15th reply arrives after a median of 28.8 minutes; 52.0% reach the fixed reply prefix within 30 minutes and 71.6% within one hour. Even without the LLM-generated factual-accuracy dimension, the remaining stance, communicative, and affective state composition retains cross-event ranking signal (ROC-AUC 0.538); including that dimension increases ROC-AUC to 0.562. In a separate Check-COVID evaluation of 229 claims, reciprocal-rank fusion retrieves a gold evidence document within the top five for 74.2% of claims and within the top 20 for 94.3%; sentence reranking reaches Recall@20 of 58.1%. We propose an integrated human-review system that brings these early forecasts, response patterns, and retrieved evidence together for misinformation triage relying on the potential virality of claims.
[NLP-104] Identifying Introspection From the Inside
【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在自我报告(self-report)中存在可信性难以验证的问题,即如何区分模型所生成的合理自述与虚假编造(confabulation)。其核心挑战在于,模型可能表现出看似合理的自我陈述,但这些陈述并不反映真实内部状态。为应对这一问题,研究提出了一种基于低秩适配器(low-rank adapters)的受控训练范式,使模型在隐式决策任务中为虚构角色做出决策,并依据潜在的线性偏好函数进行学习。研究发现,在持续微调过程中,模型能够自发产生对自身习得偏好的准确自我报告,即使未接受显式的自述监督。解决方案的关键在于揭示了“忠实自述”的机制性特征:通过权重消融和冻结层实验,发现偏好表征在训练过程中逐渐向网络更早层迁移,表明可信自述依赖于偏好信息位于可被已有语言表达机制访问的位置;进一步利用归因修补(attribution patching)分析,发现具备忠实自述能力的模型在决策任务与自述任务之间的归因相似性显著更高,形成一种不依赖于内容理解的、可量化的结构化签名(mechanistic signature),从而实现对模型是否忠实自述的结构性区分。
链接: https://arxiv.org/abs/2610.07186
作者: David I. Atkinson,Dillon Plunkett,David Bau
机构: Northeastern University(东北大学); Eleos AI Research
类目: Computation and Language (cs.CL)
备注: Published at COLM 2026. Project page: this https URL
Abstract:Large language models make claims about themselves that are both consequential and increasingly difficult to verify from behavior alone. How can we distinguish plausible confabulations from genuine introspection? In this paper, we identify mechanistic signatures of faithful self-report in a controlled setting. Using low-rank adapters, we train models to make decisions on behalf of fictitious characters, according to latent linear preference functions. We find sustained fine-tuning on an implicit decision task can lead to the emergence of accurate self-reporting of models’ learned preferences, even without explicit self-report supervision. We ask two research questions about this emergent phenomenon. First: is the emergence of accurate self-reporting accompanied by a measurable structural change in the model? Weight ablations and frozen-layer experiments together indicate that preference representations shift to earlier layers over training, consistent with the hypothesis that faithful self-report requires preferences to be located where pre-existing verbalization mechanisms can access them. Second: can these structural differences distinguish faithful models from unfaithful ones? Using attribution patching, we find that faithful models exhibit significantly higher attribution similarity between the decision-making and self-report tasks – a mechanistic signature of faithful self-report that does not require us to understand the content of the report itself. Previous work on self-report has observed behaviorally that models can be faithful or unfaithful; our work proposes that, at least in our restricted setting, it is possible to distinguish between the two patterns of computation by examining the structure of the networks themselves.
[NLP-105] Learning Scientific Exploration from Human Research Decision Trajectories
【速读】: 该论文旨在解决当前科学领域人工智能系统缺乏对科研探索过程(scientific exploration)建模的问题,即现有科学语料库主要记录研究的最终成果,而忽略了推动这些成果产生的动态决策与行动序列。其解决方案的关键在于构建一个名为 ResearchTrails 的数据集,该数据集通过从 Git 代码仓库中提取提交历史(commit histories),将版本控制中的变更记录作为科研探索轨迹的代理信号,从而实现对研究过程中方法、实验及消融分析等连续演进过程的结构化捕捉。研究提出了一套自动化且可扩展的流水线,能够从大量开源项目中高效抽取人类科研轨迹,并验证了这些轨迹蕴含了超越最终论文所披露的中间决策信息。此外,该工作展示了 ResearchTrails 在多个应用场景中的价值,包括在推理阶段引入人类科研经验作为外部知识,以及基于科研轨迹训练模型以提升对新研究决策的泛化能力。结果表明,未来人工智能系统有望不仅学习科学成果,更能理解科学发现本身的动态演化过程。
链接: https://arxiv.org/abs/2610.07184
作者: Xuchen Gong,Shane Gu,Haokun Liu,Dixi Yao,Chenhao Tan,Tian Li
机构: University of Chicago(芝加哥大学)
类目: Machine Learning (cs.LG); Computation and Language (cs.CL)
备注:
Abstract:A key challenge in building AI systems for scientific research is enabling \textitscientific exploration : the systematic process of investigating unknown phenomena or ideas to gain new knowledge through sequences of research decisions and actions. Yet this process is largely missing from existing scientific corpora; for example, research papers primarily record final outcomes rather than the trajectories that produced them. In this work, we introduce \textbfResearchTrails , a dataset of \textbfhuman research trajectories constructed from Git repositories , where \textbfcommit histories serve as proxies for research exploration. We develop an automated and scalable pipeline that extracts structured research trajectories from repository commits, capturing successive changes to methods, experiments, and ablations. We characterize the resulting dataset and show that these trajectories contain meaningful signals about intermediate research decisions beyond what final papers reveal. We further demonstrate utilities of ResearchTrails in multiple use cases, including retrieving human research experience as external skills at test time and training models on research trajectories to improve generalization to new research decisions. Our results suggest a path toward AI systems that learn not only from the products of science, but from the evolving process of discovery itself.
[NLP-106] CLM-as-a-Judge: Evaluating an Open Contrastive Decision Model on Public Judge Benchmarks
【速读】: 该论文旨在解决生成式模型在自动评估任务中作为判别器(judge)时表现接近随机猜测(near chance performance)的问题,尤其针对开放性对比决策模型(open contrastive decision model)在多个权威基准测试(如RM-Bench、JudgeBench和HaluEval)上表现出的低效与不可靠性。其核心问题是:尽管模型参数量与高性能判别器相当,但其决策能力却无法显著超越随机基线或恒定输出的平凡基线(如always-first),表明其内部判断机制存在严重偏差或校准缺陷。解决方案的关键在于引入基于校准置信度的分层筛选机制(confidence-gated cascade),通过在保留率阈值0.97下对模型自身置信度进行校准,将高置信度样本导向更可靠的强判别器(如奖励模型或生成式判别器)。该方法有效修复了原始模型过自信(overconfident)的问题——其原始置信度最高可高出真实概率0.401,经统一温度校准后预期校准误差降至0.062以内,并使校准后的置信度能够将模型自身的错误排序高于随机水平。此外,研究还揭示了两个关键现象:一是对比模型的决策顺序翻转率(decision order-flip rate)极低(0.0002),远低于生成式判别器(0.2188),表明其行为高度僵化;二是其长度偏好偏移(length-preference shift)仅-0.023,显著优于生成式判别器的-0.217,说明其更少受文本长度干扰。最终,通过五项公开偏好基准与一项幻觉检测基准的预注册评估框架,验证了该设计在提升判别可靠性方面的有效性。
链接: https://arxiv.org/abs/2610.07177
作者: Gowthamkumar Nandakishore
机构: 独立研究员(Independent researcher)
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注:
Abstract:An open contrastive decision model is near chance as a judge on the hard public benchmarks: Contrastive-LM/CLM-v0.1-8B scores between 0.351 (best- of-four, chance 0.250) and 0.593 (pairwise, chance 0.500), is statistically indistinguishable from coin flipping on RM-Bench and JudgeBench, and answers every HaluEval item with one constant label, matching the trivial always-first baseline at 0.581. Judges with the same parameter count score far higher everywhere: a reward model reaches 0.764 to 0.976 and a generative judge 0.611 to 0.778, and every gap to CLM is significant after Benjamini-Hochberg correction. Two properties do work. Raw confidences are overconfident by up to +0.401, yet one pooled temperature fit on held-out calibration items repairs expected calibration error to at most 0.062, and the repaired confidence ranks the model’s own errors above chance on three of six benchmarks. The decision order-flip rate is 0.0002 against 0.2188 for the generative judge, and the length-preference shift is -0.023 against -0.217. The confidence-gated cascade, however, escalates between 0.923 and 1.000 of items to the strong judge at the preregistered 0.97 retention bar: calibrated confidence about a near-chance judge has almost nothing to keep. The design: five public preference benchmarks and one hallucination benchmark with real labels, scored under a preregistration frozen before any test item was seen, against generative, reward-model, and trivial baselines, with per-item predictions released.
[NLP-107] A theory of platonic representations in language models
【速读】: 该论文试图解决的问题是:在多语言语言模型的内部层中,翻译句子的表示具有高度相似性这一现象,尽管这一现象与柏拉图表征假说相关,但缺乏理论解释。其解决方案的关键在于假设数据存在一种隐藏的分层结构,其中高层抽象特征在不同语言间共享,而低层表面特征则依赖于具体语言或模态。通过基于共享上层但不共享下层生成规则的概率上下文无关语法(Probabilistic Context-Free Grammar, PCFG)生成合成语言,研究发现贝叶斯最优的下一个词预测器为信念传播(Belief Propagation, BP),将BP的消息传递过程编码到模型的逐层结构中,可获得与在相同数据上训练的Transformer模型高度吻合的解析预测结果。该框架解释了跨语言相似性在中间层达到峰值、与语言特异性结构共存,并随语言相近性、模型质量及数据暴露程度增强的现象;同时区分了相似性(共享邻域几何)与对齐性(共享坐标系),指出只有当代码切换数据(即混合语言句子)足够丰富时,对齐性才会出现。此外,该框架还预测:从每一层中减去可由前一层线性预测的部分,能够提升跨语言相似性,这一预测已在预训练大语言模型中得到验证。
链接: https://arxiv.org/abs/2610.07168
作者: Darshil Doshi,Wenjie Zhou,Corinna Elena Wegner,Daniel J. Korchinski,Santiago Acevedo,Matthieu Wyart
机构: Johns Hopkins University(约翰霍普金斯大学); EPFL(洛桑联邦理工学院); SISSA(国际高等研究院)
类目: Machine Learning (cs.LG); Computation and Language (cs.CL); Machine Learning (stat.ML)
备注: 10+14 pages, 7+12 figures
Abstract:Representations of translated sentences are similar in the inner layers of multilingual language models – an observation connected to the platonic representation hypothesis, yet unexplained theoretically. We provide an explanation based on the assumption that data have a hidden hierarchical structure whose abstract levels are shared across languages while surface levels are modality- or language-specific. Concretely, we generate synthetic languages from probabilistic context-free grammars sharing upper-level but not lower-level production rules. In this setting the Bayes-optimal next-token predictor is belief propagation (BP); encoding its messages in successive layers yields analytical predictions that agree well with transformers trained on the same data. The framework explains why cross-lingual similarity peaks in middle layers, coexists with language-specific structure, and strengthens with language proximity, model quality and data exposure. It distinguishes similarity (shared neighborhood geometry) from alignment (shared coordinates), showing that the latter occurs when code-switched data, i.e. mixed-language sentences, are abundant enough. It further predicts that subtracting from each layer the component linearly predictable from the preceding one increases cross-lingual similarity, which we confirm in pretrained LLMs.
[NLP-108] Jailbreaking Open-Weight LLM s via Random Embedding Perturbations
【速读】: 该论文旨在解决开放权重大语言模型(open-weight LLMs)在安全性方面存在的漏洞问题,特别是针对模型对有害、恶意或不当提示的防御能力不足。其核心问题是:尽管这些模型在功能上持续进步并被广泛应用,但其安全机制在面对特定攻击时仍表现出显著脆弱性。论文提出的关键解决方案是“扰动嵌入向量”(Perturbed Embedding Vector, PEV)攻击方法,其关键在于通过在提示词的嵌入向量表示中添加独立的高斯噪声,无需梯度计算、每提示优化或修改模型内部权重,即可高效触发模型生成不安全响应。实验表明,PEV在计算成本上比以往方法低一个数量级,多数模型在新提示下首次成功攻击可在一分钟内完成,且在JailbreakBench基准测试中对所有模型和所有提示均能生成有害输出。这一结果凸显了嵌入空间扰动对模型行为的重大影响,揭示了其作为安全风险的同时,也为探索大语言模型的动力学行为提供了新的研究工具。
链接: https://arxiv.org/abs/2610.07125
作者: Abhinav Sudhakar Dubey(University of California Santa Cruz),Scott Sirri(University of California Santa Cruz),Vaggos Chatziafratis(University of California Santa Cruz),C. Seshadhri(University of California Santa Cruz)
机构: University of California, Santa Cruz (加州大学圣克鲁兹分校)
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
备注: 15 pages, 4 figures, Code: this https URL
Abstract:While open-weight models have enjoyed steady progress in capabilities and wide adoption across multiple domains, their safety remains an important concern. One key feature is the ability to refuse or deflect harmful, malicious, or insensitive prompts. In this paper, we expose safety vulnerabilities across six common open-weight LLMs of various sizes that consistently lead to harmful or unsafe responses on the JailbreakBench benchmark dataset. Our proposed attack, Perturbed Embedding Vector (PEV), is a simple and fast “jailbreaking” technique that is cheaper than prior approaches, which typically require gradient computations, per-prompt optimizations, or altering internal weights of the models. PEV just adds independent Gaussian noise in the embedding vector representations of the prompt, with no need for further manipulations. To generate unsafe responses, we repeatedly sample additive noise from this distribution. In our experiments, we observe that the average compute cost to get the first successful attack is up to an order of magnitude less than previous attacks. The first successful jailbreak on a new prompt typically arrives within one minute on every tested model, and PEV generates unsafe responses across all models for all prompts in JailbreakBench. No other tested method achieves such results, despite them taking longer to run. More broadly, we believe that understanding the behavior of LLMs under perturbations in the embedding vectors is an important research direction: while perturbations constitute a major security risk, they can also serve as a valuable tool for exploring the dynamical behavior of such models.
[NLP-109] JudgeMoE: Distributional Aggregation for LLM -as-a-Judge
【速读】: 该论文旨在解决大语言模型(LLM)评判器在对生成结果进行评分时,因将评分分布压缩为单一标量而导致的不确定性与评判分歧信息丢失问题。其核心解决方案是提出一种轻量级聚合方法 JudgeMoE,通过为每个样本动态分配特定权重来整合缓存的评判器评分分布,并在融合后计算最终得分。关键创新在于利用示例特异性权重实现对多源评分分布的自适应加权聚合,从而保留并有效利用原始评分中的不确定性信息。实验表明,相较于均匀对数池化等传统方法,JudgeMoE 在原始10个评测单元上使平均斯皮尔曼相关系数提升0.079;在扩展至16个评测单元的综合评估中,相较最强本地单个评判器平均提升0.0393,且在12/16个单元中表现更优,统计显著性达到p=0.0091(单侧Wilcoxon符号秩检验),验证了该方法在不同任务和评判器池配置下的有效性与适应性。
链接: https://arxiv.org/abs/2610.07109
作者: Yiqi Liu,Joseph James,Yang Wang,Kun Zhao,Chenghao Xiao,Chenghua Lin
机构: The University of Manchester (曼彻斯特大学); The University of Sheffield (谢菲尔德大学); University of Pittsburgh (匹兹堡大学); Shanghai University of Finance and Economics (上海财经大学)
类目: Computation and Language (cs.CL)
备注:
Abstract:When an LLM judge scores an output, its score distribution retains uncertainty and disagreement information that is lost after scalar compression. We introduce JudgeMoE, a lightweight aggregator that assigns example-specific weights to cached judge score distributions and fuses them before computing a final score. A protocol study shows that score-range choice is unstable across judge–dataset settings and that soft scoring usually outperforms hard decoding. On the original 10-cell benchmark, JudgeMoE improves mean Spearman over uniform log pooling by +0.079 . Applying the same configuration to six additional cells yields a +0.0393 mean gain over the strongest local single judge across 16 cells, with positive differences in 12/16 cells and a one-sided Wilcoxon signed-rank p=0.0091 . Validation-based analyses further show that the preferred aggregation method depends on the task and judge pool.
[NLP-110] urnslide: Scalable Multi-Turn Data Synthesis by Walking a Finite-State Machine NEURIPS2026
【速读】: 该论文旨在解决小语言模型(Small Language Models, SLMs)在多轮工具调用(multi-turn tool calling)任务中表现不佳的问题,尤其针对缺乏针对特定API的微调数据这一现实挑战。现有合成方法因需构建复杂的模拟运行环境并频繁调用大语言模型(LLM)生成对话回合,导致成本过高,难以实现大规模微调。为此,论文提出一种全自动、轻量级的合成框架,其核心创新在于将每个API建模为有限状态机(Finite-State Machine, FSM),通过抽象状态来约束工具调用的合法性,从而生成符合状态逻辑的工具调用序列。该方法仅需一次LLM调用即可将状态有效序列转换为完整的训练样本,并通过设定目标分布(包括对话轮数、工具序列及任务复杂度)来控制生成数据的结构特征。实验表明,该方法在微调后显著提升下游任务准确率,相较基线模型和现有工作,在仅使用3.6至6.6倍更少的令牌(tokens)情况下,达到70.7%的全准确率,优于已有方法的63.4%与53.7%,验证了其高效性与数据质量优势。
链接: https://arxiv.org/abs/2610.07070
作者: Aaron Fainman,Gabriela Kadlecová,Maciej Gryka,Bartosz Kruszczyński,Usman Zafar,Cédric Archambeau,Aaron Klein,David Salinas,Selim Nowicki,Jacek Golebiowski
机构: Imperial College London (帝国理工学院); distil labs; Agon; ELLIS Institute Tübingen (ELLIS图宾根研究所)
类目: Computation and Language (cs.CL)
备注: Accepted at the SLM-Agents Workshop, NeurIPS 2026 (non-archival)
Abstract:Small language models are inexpensive to serve and can run on private infrastructure, but base models are often not good enough at multi-turn tool calling, and fine-tuning them needs per-API data that rarely exists. Existing synthesis methods are too expensive for high-scale fine-tuning, as they often require mock operational environments for different domains and multiple LLM calls per generated conversation turn. We introduce a fully automated, lightweight synthesis framework that models each API as a finite-state machine, representing the system as abstract states that determine when each tool may be called, producing state-valid sequences of tools; sequences are translated into complete examples with a single LLM call. Rather than optimize diversity, we set a target distribution over the number of turns, the tool sequence and task complexity. We measure data quality by fine-tuning SLMs on generated trajectories, showing that our FSM-based generation significantly improves downstream accuracy over an unmutated baseline and, against existing works, reaches 70.7% full accuracy over 63.4% and 53.7% with 3.6-6.6 \times fewer tokens.
[NLP-111] Learning to Simulate Individuals from Macro Social Signals
【速读】: 该论文旨在解决大语言模型在用户行为模拟中缺乏多样化且可解释的推理机制的问题,即现有方法依赖预训练知识或个体标注数据,难以有效捕捉真实情境下多主体复杂互动中的行为推理过程。其核心解决方案是引入宏观市场信号(macro2mind),通过基于强化学习的生成式策略优化(GRPO)训练语言模型,利用预测市场的价格轨迹作为社会性行为反馈信号。关键创新在于提出一种社会行为分解框架,将行为推理显式建模为四步过程:识别代表性市场参与者群体、推断其对新闻事件的解读与信念更新、模拟群体间交互逻辑,并聚合结果形成市场价格。同时,采用基于事后后悔(hindsight-regret)的难度感知采样课程学习策略,聚焦于能显著提升预测性能的复杂决策转换,确保训练集中高价值且可学习样本的有效利用。该方法使模型具备零样本迁移能力,在多个用户模拟基准测试中表现优异,且作为数据生成器可显著提升下游模拟器对未见用户的建模精度。
链接: https://arxiv.org/abs/2610.07062
作者: Yining Zhao,Bushi Liu,Haofei Yu,Zhengyang Qi,Shanyong Wang,Chuyue Li,Yuxiang Liu,Jiaxuan You
机构: University of Illinois Urbana-Champaign(伊利诺伊大学厄本那-香槟分校); Independent Researcher(独立研究员)
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Social and Information Networks (cs.SI)
备注:
Abstract:Large language models are increasingly used to simulate how individuals respond to new situations, yet the behavioral reasoning behind these responses is either inherited from pretraining or learned from individual-level annotations, which offer limited behavioral diversity and little supervision of the reasoning itself. We propose to learn behavioral reasoning from prediction markets, whose price trajectories record how populations respond to real-world events at scale. We introduce macro2mind, which trains a language model with GRPO using market signals. A social behavioral decomposition makes behavioral reasoning an explicit step of forecasting: the model infers representative groups of market participants, predicts how each interprets the news and updates its beliefs, reasons about their interactions, and aggregates these responses into a price. A hindsight-regret curriculum with difficulty-aware sampling focuses training on transitions where hindsight-identified groups substantially improve the forecast while prioritizing examples that remain learnable for the current policy. The learned reasoning applies to user simulation without further training. On SWM-Bench, macro2mind achieves state-of-the-art directional accuracy and correlation on Polymarket. Trained on market data, it transfers zero-shot to four user-simulation benchmarks (Humanual, OvertonBench, PRISM, and CAD) and has competitive performance among zero-shot methods. Used as a data generator, macro2mind also raises a downstream simulator’s accuracy on unseen users by 15.5 points, outperforming data generated by its backbone by 13.2 points.
[NLP-112] Investigating Model Compression for Neural Machine Translation in the Biomedical Domain
【速读】: 该论文旨在解决在低资源条件下,针对专业领域(如生物医学)的法语-英语机器翻译任务中,大型预训练变换器模型在模型压缩与推理加速过程中面临的性能退化问题。具体而言,知识蒸馏因缺乏领域特定的平行数据而难以有效迁移知识,而量化则在降低权重和激活精度时易导致翻译质量下降。为应对上述挑战,论文提出将知识蒸馏与量化技术联合优化,通过多种微调策略适配压缩后的学生模型。其解决方案的关键在于:构建一种协同蒸馏与量化的方法,使学生模型在大幅减小模型规模(减少69%)、显著提升推理速度(提升98.21%)并大幅降低碳排放(减少98.46%)的同时,保持与原始基线相当的翻译质量,从而实现高效、高性能且低碳的模型部署,适用于资源受限的翻译服务场景。
链接: https://arxiv.org/abs/2610.07032
作者: Maria Zafar,Souhail Bakkali,Rejwanul Haque
机构: South East Technological University, Carlow, Ireland; IRISA, France
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: Accepted at AICS 2025
Abstract:Large-scale pretrained transformer models have achieved state-of-the-art performance across diverse machine translation tasks, including multilingual settings. Knowledge distillation has emerged as a sustainable approach for model compression, transferring knowledge from large teacher models to smaller, more efficient student models. Similarly, quantization, which reduces the numerical precision of model weights and activations (e.g., from 32-bit to 8-bit representations) is widely used to accelerate inference, enabling models to run several times faster during deployment. However, both techniques face limitations when applied to specialized domain data, particularly under low-resource conditions. In knowledge distillation, the effectiveness of transfer is often constrained by the scarcity of domain-specific parallel data, while quantization can lead to performance degradation as bit precision decreases. In this work, we investigate the combined application of knowledge distillation and quantization for French-to-English biomedical translation, a domain characterized by specialized terminology and limited parallel resources. We develop and compare multiple fine-tuning strategies to adapt compressed student models to this challenging setting. Our experiments demonstrate that a collaboratively distilled and quantized student model achieves a 69% reduction in size, a 98.21% increase in inference speed, and a 98.46% reduction in CO2 emissions compared to the original baseline all without sacrificing translation quality. These results indicate that jointly optimized compression techniques can yield efficient, high-performance models suitable for translation service providers operating under resource constraints.
[NLP-113] Calibrated Answers About Randomized Trials From a 4-Billion-Parameter Open Model: A Registered Test and a License-Clean Release
【速读】: 该论文旨在解决医学文献中证据推理(Evidence Inference 2.0, EI)任务的自动化评估问题,即针对随机对照试验(Randomized Controlled Trial, RCT)文章,判断某干预措施是否显著改善、恶化或未显著改变某一结局指标相对于对照组的结果。其核心挑战在于模型需在有限上下文长度(6,144 tokens)内准确理解临床研究内容并做出高置信度的二元分类决策。解决方案的关键在于采用基于Qwen3-4B-Base架构的生成式AI模型,通过引入低秩适配器(low-rank adapters)与决策头(decision head)进行微调,并仅在1,431篇允许再利用的训练文章上训练,以确保合规性与可复现性。该方法在测试集上表现出优异性能:预期校准误差为0.0168(低于0.05阈值),对数损失显著优于基线模型(下降0.8603),宏平均F1达0.9248,且在四项预设于开放科学框架(Open Science Framework)的评估标准下全部通过。此外,实验表明模型对标题信息敏感,标题含结果提示时性能提升明显,而干预与对照组角色互换则导致约66.5%的预测方向反转,验证了模型对语义关系的敏感性。最终模型部署符合预先设定的精度与可靠性边界,释放版本遵循Apache License 2.0协议。
链接: https://arxiv.org/abs/2610.07019
作者: Johann Emmanuel Li
机构: 未知
类目: Computation and Language (cs.CL); Machine Learning (cs.LG)
备注: 20 pages, 2 figures, 13 tables. Code, results files and the paper’s sources: this https URL . Model: this https URL . Registration: this https URL
Abstract:Fiorillo v0.5 is an open model that answers typed questions with a probability for each answer. Its main specialist reads a randomized trial’s article, cut to 6,144 tokens, and answers whether an intervention significantly increased, significantly decreased or did not significantly change an outcome against a comparator (Evidence Inference 2.0, EI). It is Qwen3-4B-Base with low-rank adapters and a decision head, fine-tuned for EI only on the 1,431 of 2,657 training articles whose own license allows reuse. Four criteria registered on the Open Science Framework before this version’s test predictions decided its release, the second bar judged on EI’s test split, whose labels are public. On that split (1,218 prompts in 333 articles), the expected calibration error was 0.0168 against a limit of 0.05; log loss was below the prior’s by 0.8603 (95 percent interval 0.8104 to 0.9078) and below that of Gemma 4 31B-it, reading the same input, by 0.1829 (0.1164 to 0.2598); and macro-F1 was 0.9248 against 0.8668, so all four criteria passed. Training the same recipe on clean articles alone cost 0.0123 in accuracy (0.0034 to 0.0207; descriptive). With no article, macro-F1 fell to 0.4384; the title alone raised it by 0.0939 (0.0655 to 0.1234), which a title stating the result or recall of the trial could explain; exchanging intervention and comparator reversed 0.6652 of its direction answers. Run as released, the files matched the evaluated predictions within limits set in advance. The release is under the Apache License 2.0 (digital object identifier https://doi.org/10.57967/hf/10722).
[NLP-114] Mask-Guided KV Cache Eviction in Block Diffusion Language Models
【速读】: 该论文旨在解决生成式语言模型在长序列生成过程中因维持大规模键值(Key-Value, KV)缓存而导致的内存占用高与推理速度慢的问题。传统方法在每一步去噪过程中均需访问全部历史KV缓存,严重限制了模型的可扩展性与效率。其核心挑战在于如何高效地进行两阶段决策:一是选择当前生成块所依赖的过去标记(即选择,selection),二是决定哪些历史信息应保留在内存中以供后续使用(即淘汰,eviction)。本文提出了一种无需训练的解决方案——MaskAhead,通过单一基于掩码查询(mask-query-based)的排序机制同时完成上述两项任务:利用当前块的掩码指导选择,通过探测未来被掩码块的响应来指导淘汰,二者均依据对注意力输出贡献度的估计对KV条目进行排序。进一步提出的量化版本Q-MaskAhead直接从低比特表示的KV中计算选择与注意力,显著保留了被选中的关键信息。实验覆盖长序列推理、长提示问答及“针在 haystack”检索等任务,在长提示问答场景下,MaskAhead平均将KV内存降低9.5倍,仅损失1.2点平均F1;Q-MaskAhead则实现20.1倍的内存压缩,F1损失为2.3点。在批量32的系统级评估中,MaskAhead相较密集推理分别带来1.23倍的整体端到端加速与1.68倍解码阶段加速,验证了其在实际部署中的高效性。
链接: https://arxiv.org/abs/2610.06996
作者: Gleb Molodtsov,Ekaterina Alimaskina,Evgeny Uskov,Artur Zagitov,Aleksandr Beznosikov
机构: 未知
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注:
Abstract:Block diffusion language models keep a large key-value (KV) cache throughout generation and attend to it at every denoising step, limiting both memory capacity and generation speed. Reducing these costs requires deciding which past tokens to use for denoising the current block (selection) and which to keep in memory for future blocks (eviction). We propose MaskAhead, a training-free method that solves both tasks with a single mask-query-based ranking mechanism. Current-block masks guide selection, while probes of upcoming masked blocks guide eviction. Both rank KV entries by their estimated contribution to the attention output. Our quantized variant, Q-MaskAhead, computes selection and attention directly from low-bit KV, largely preserving the selected entries. Experiments on Fast-dLLM-v2, DreamReasoner, and LLaDA2.0-mini cover long-generation reasoning, long-prompt question answering, and needle-in-a-haystack retrieval. On long-prompt QA, MaskAhead reduces KV memory by 9.5\times on average with a 1.2-point mean F1 loss relative to dense inference. Q-MaskAhead increases the reduction to 20.1\times with a 2.3-point mean F1 loss. In a batch-32 systems profile, MaskAhead achieves 1.23\times end-to-end and 1.68\times decode-stage speedups over dense inference.
[NLP-115] AegisFlow: A Multi-Agent Agent ic AI Framework for Autonomous Remediation and Self-Healing in Frag ile Data Ecosystems
【速读】: 该论文旨在解决传统数据管道(data pipeline)在面对上游模式漂移(schema drift)、API契约变更或网页DOM结构调整等动态变化时极易失效的问题,以及现有可观测性工具仅能告警而无法自动修复所导致的高平均修复时间(MTTR)和运维人员操作疲劳问题。其核心解决方案是提出AegisFlow——一种基于智能体(agentic)的自愈框架,通过构建“监控-分析-规划-执行-知识”(MAPE-K)闭环机制,实现从故障检测到自动修复的全链路自治。该框架的关键创新在于采用非侵入式并行影子补丁(Parallel Shadow Patching)执行模型,利用大型语言模型(LLM)驱动的修复智能体(Repair agent),在数字孪生(digital twin)环境中自动生成、验证并部署代码补丁,从而显著缩短修复周期。实验表明,AegisFlow将平均修复时间从170分钟降至3.2分钟(提升98.1%),整体补丁成功率高达92%,尤其在处理JSON模式变更(96%)和标点符号漂移(98%)方面表现优异,尽管在影子DOM场景下表现稍弱(85%)。此外,该框架具备部署无关性,可作为插件无缝集成至现有工作流调度系统,使数据工程师约98%的应急值守时间得以释放,转向更具创新性的任务。
链接: https://arxiv.org/abs/2610.06971
作者: Muhammad Bilal Awan,Zubair Hussain,Abdul Shahid
机构: 未知
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: 43 pages
Abstract:Traditional data pipelines are notoriously brittle, often failing due to upstream schema drift, API contract changes, or website DOM modifications. Present observability tools only raise alerts but for human engineers, resulting in a high Mean Time to Repair (MTTR) and operational fatigue. In this paper we propose AegisFlow (Agentic Engine for Intelligent Self-healing and Graph-driven Operations for Workload remediation), a novel agentic framework that closes the loop between detection and resolution. AegisFlow uses a Watchdog agent to collect runtime telemetry and has a Repair agent to automatically create, test and deploy code patches based on Large Language Models (LLMs). The framework presents the non-intrusive execution model called Parallel Shadow Patching, a non-intrusive execution model based on the Monitor, Analyze, Plan, Execute, Knowledge (MAPE-K) loop to generate and verify patches in digital twin environments. Through experimental testing, we have evaluated AegisFlow across five common failure scenarios, and see 98.1 percent improvement in MTTR (from an average of 170 minutes per patch to 3.2 minutes) and a patch success rate of 92 percent . In particular, the system is successful in dealing with changes in the JSON schema (96 percent ) and punctuation drift (98 percent ), and is least successful in Shadow DOM cases (85 percent ). AegisFlow frees up about 98 percent of data engineering on-call time from firefighting and reallocates it towards innovation. The framework is deployment agnostic consisting of a system that can be deployed in a plugin fashion into an existing pipeline orchestration system with minimal uplift to the existing system.
[NLP-116] WavePrune: One period is often enough for RoPE
【速读】: 该论文旨在解决旋转位置编码(Rotary Position Embedding, RoPE)因周期性特性导致的位置混叠(position aliasing)问题,即当相对位置相差一个完整旋转周期时,模型难以区分其真实差异,从而影响长上下文建模性能。其核心解决方案是提出WavePrune方法,通过将每个通道的旋转限制在首个周期内,有效消除由位置混叠引起的注意力图干扰,提升长序列建模能力。实验表明,WavePrune在无需额外调参的情况下显著提升多个模型的HELMET得分(如Qwen3-8B从35.7提升至40.0),并在从头预训练中实现外推长度下的更低验证损失。此外,由于WavePrune引入了细粒度稀疏结构,可被硬件对齐的CUDA核函数高效利用,在32K上下文下相较FlashAttention-2实现1.15倍的prefill加速与1.24倍的解码加速。研究结果表明,RoPE的周期结构在首个旋转周期后基本冗余,突破了其“周期性为必要”的普遍认知。
链接: https://arxiv.org/abs/2610.06963
作者: Guancheng Du,Luotian Huang,Shaowen Wang,Si Li,Kaifeng Lyu
机构: Tsinghua University (清华大学)
类目: Computation and Language (cs.CL); Sound (cs.SD)
备注:
Abstract:Rotary Position Embedding (RoPE) encodes token positions by rotating each two-dimensional channel of the query and key vectors at a channel-specific frequency, making the attention logits invariant to a common shift of positions. However, this rotation is periodic, and it leads to position aliasing where relative positions separated by a full rotation period become hard to tell apart. To address this, we propose WavePrune, which restricts each channel to its first rotation period. We show that it removes the distractions in attention maps created by position aliasing and improves overall long-context performance. Specifically, WavePrune raises the HELMET score on four of five models we test without any extra tuning (e.g., 35.7 - 40.0 on Qwen3-8B). When pretraining models from scratch, WavePrune also achieves lower validation loss at extrapolated lengths than pretraining without it. Because WavePrune restricts each channel to a sliding window, it induces a fine-grained sparsity that our hardware-aligned CUDA kernels exploit for 1.15x prefill and 1.24x decoding speedups over FlashAttention-2 at 32K context. Together, these results show that RoPE’s periodic structure, widely regarded as essential, is largely redundant beyond the first rotation period.
[NLP-117] Verdicts Without Annotated Evidence: Rejection Sampling or Label-Only Post-Training for Evidence Recovery?
【速读】: 该论文旨在解决审阅流程中仅保留最终裁决(verdict)而缺乏可追溯证据(evidence spans)的问题,即在没有人工标注证据的情况下,如何从仅有的裁决标签中恢复出可解释的文本依据。其核心挑战在于:传统评估指标(如准确率)无法有效反映生成证据的可审查性,因为裁决正确性与实际证据匹配度之间关系微弱,导致以准确率为指导的模型训练难以保证输出的可追溯性。解决方案的关键在于采用“仅基于标签的训练”(label-only training)与“拒绝采样”(rejection sampling)策略——前者通过在无证据标注的前提下对小语言模型进行后训练,提升其生成与裁决一致的证据片段能力;后者则进一步引入自动来源锚定得分(source-grounding score),仅保留与记录裁决一致且语义连贯的生成轨迹。实验结果表明,这两种方法均显著提升了证据召回与精确度,尤其在无需额外人工标注的情况下,实现了证据质量的改善,验证了在无监督条件下增强生成式 AI 输出可审查性的可行性。
链接: https://arxiv.org/abs/2610.06962
作者: Nishanth Nayakanti,Prasang Gupta,Ashutosh Bilthare,Kevin Paul
机构: PricewaterhouseCoopers(普华永道), U.S.
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: 16 pages, 5 figures
Abstract:In many review workflows the verdict is the only thing retained. The passages behind it are not marked, because that annotation costs far more than recording the decision. We measure how much of that evidence a small language model can recover when it is post-trained on the verdicts alone, with no human evidence labels at any stage. On ContractNLI the human evidence spans are held out until evaluation. Matching the recorded verdict and agreeing with those spans are not the same thing: across six systems the two scores are only weakly related and rank the systems differently, so accuracy is a poor guide when the citations have to be reviewable. Label-only training on the bare verdict reaches accuracy 0.896 and span F1 0.564. Rejection sampling, which keeps a generated trace only when its verdict matches the record and then picks one by an automatic source-grounding score, reaches 0.797 and 0.556, against 0.747 and 0.493 before training. Verbatim citation rises from 0.597 to 0.729 under label-only training and to 0.701 under rejection sampling. One seed on one corpus cannot say which method is better, but both improve the evidence without anyone annotating it.
[NLP-118] EMODE: Dynamic Para-Semantic Experts for Emotion-Aware Speech Language Modeling
【速读】: 该论文旨在解决大尺寸语音语言模型在跨模态理解与生成任务中难以有效保留副语言线索(尤其是情感信息)的问题。现有系统普遍依赖于纠缠的声学表征,导致语言模型过度依赖恢复出的词汇内容,而非基于声学-韵律证据进行行为决策。为克服这一局限,本文提出EMODE——一种基于动态参数语义专家(Dynamic Para-Semantic Experts, DPSE)的情感感知语音语言模型。其核心解决方案在于通过DPSE将连续语音特征解耦为语义与副语言两条路径,实现动态路由与融合后再输入语言模型,从而促进对情感等副语言信息的显式建模。为实现结构解耦到功能专化的转化,EMODE采用三阶段训练范式:语义预热、副语言激活与联合精调,并分别引入正交专家引导(Orthogonal Expert Guidance, OEG)、语义-声学对齐(Semantic-to-Acoustic Alignment, SAA)与门控多样性正则化(Gating Diversity Regularization, GDR),以强化各路径的功能独立性与协同能力。实验结果表明,EMODE在语音情感识别(SER)、共情响应评估及新构建的双语MEPA基准上均显著提升了词汇忠实度与情感敏感性的平衡,增强了基于情感的响应生成能力,并验证了显式副语义因子分解对于跨语料库情感理解鲁棒性的关键价值。
链接: https://arxiv.org/abs/2610.06956
作者: Jianan Pan,Yiwen Gu,Xinze Li,Rui Wang,Kejie Huang
机构: Zhejiang University (浙江大学)
类目: Computation and Language (cs.CL); Multimedia (cs.MM); Sound (cs.SD)
备注:
Abstract:Large speech language models have demonstrated strong capabilities in unified cross-modal understanding and generation, yet paralinguistic cues, especially emotion, remain difficult to preserve. Existing systems typically rely on entangled acoustic representations, which allow the underlying language model to depend excessively on recovered lexical content instead of grounding its behavior in acoustic-prosodic evidence. We address this limitation with EMODE, an emotion-aware speech language model built around \textbfDynamic Para-Semantic Experts (DPSE). DPSE decomposes continuous speech features into semantic and paralinguistic pathways, routes them dynamically, and fuses them before integration into the language model. To turn this structural decomposition into functional specialization, EMODE is trained with a three-stage curriculum consisting of semantic warm-up, paralinguistic activation, and joint refinement, guided by Orthogonal Expert Guidance (OEG), Semantic-to-Acoustic Alignment (SAA), and Gating Diversity Regularization (GDR). Experiments on SER test, empathetic response evaluation, and the newly constructed bilingual MEPA benchmark show that EMODE improves the balance between lexical fidelity and emotional sensitivity, strengthens affect-grounded response generation, and exposes the value of explicit para-semantic factorization for robust cross-corpus emotion understanding.
[NLP-119] Stabilizing language models under continual learning via condition-anchored distillation
【速读】: 该论文旨在解决持续适应(continual adaptation)过程中语言模型对早期学习提示(prompt)的输出分布发生偏移的问题,同时避免或无法保留所有历史提示-答案对所带来的存储与计算负担。其核心解决方案是条件锚定生成式蒸馏(Condition-Anchored Generative Distillation, CAGD),关键在于将传统回放机制中混淆的三种角色明确分离:条件(condition)用于选择需保护的行为模式,教师生成(teacher generation)定位相关状态空间,软目标(soft target)则精确控制预测分布的变化方式。针对自回归语言生成任务,CAGD利用教师滚动生成的序列分解实现对序列分歧的精确链式规则建模;在掩码扩散语言建模(masked diffusion language modeling)中,则直接调控教师生成结果上的局部去噪漂移。实验表明,在219M参数的掩码扩散语言模型上,CAGD显著降低四任务连续适应后的保留损失(held-out loss),在两种任务顺序下分别从2.927降至1.114、从2.168降至0.891;且在相同教师生成支持下,相比硬回放(hard replay)平均损失降低0.055。该方法在不同数据集(SMDM、Qwen3)和任务类型(新事实、自然指令)中均表现稳定,并在GSM8K任务中保持答案格式合规性,尽管在更大规模模型(0.6B与1.7B)上精确匹配率受种子影响而波动。这些结果验证了“条件锚定的功能性保全”作为跨多种语言生成目标的通用设计原则的有效性。
链接: https://arxiv.org/abs/2610.06940
作者: Huan Li,Zhe Cao,Qinlei Xie,Fushun Cui,Xuechen Liang
机构: 未知
类目: Computation and Language (cs.CL)
备注:
Abstract:Continual adaptation of language models can change their output distribution on prompts learned earlier, while retaining every old prompt-answer pair may be undesirable or impossible. We study condition-anchored generative distillation (CAGD): retain a small set of old prompts, use a frozen previous model to reconstruct completions and generation states, and match its predictive distributions while learning the next task. The formulation separates three roles that ordinary replay conflates: conditions select the behavior to protect, teacher generations locate relevant states, and soft targets specify how predictions may change. For autoregressive language generation, teacher-rollout distillation admits an exact chain-rule decomposition of sequence divergence. For masked-diffusion language modeling, our implementation directly controls local denoising drift on teacher-generated completions. In continual adaptation of a 219M masked diffusion language model, CAGD reduces four-task final held-out loss from 2.927 to 1.114 in one task order and from 2.168 to 0.891 in exact reverse. The same soft targets lower final average loss by 0.055 over hard replay when teacher-generated support is held identical. The direction persists on fresh facts and natural instructions across SMDM and Qwen3. On GSM8K, Qwen adaptation preserves answer-format compliance, but exact-match retention is seed-mixed at 0.6B and worsens at 1.7B. These results support condition-anchored functional preservation as a common design principle across the tested language-generation objectives.
[NLP-120] Component and Dimension Sparsity in Transformer Refusal Mechanisms
【速读】: 该论文旨在解决生成式AI(Generative AI)中拒绝行为(refusal behavior)的机制可解释性问题,即如何通过干预大语言模型内部激活来操控其行为,而现有方法对这些干预的内在机理理解尚不充分。其核心解决方案是将拒绝行为的调控分解为在四个开源权重模型上的组件级干预,识别出仅需操纵少量关键注意力模块与前馈神经网络(MLP)组件即可复现完整的拒绝行为效应。研究发现,有效的拒绝方向集中于占上游组件28%–48%的稀疏组件机制中,且能保留88%–101%的调控效能;在这些机制内部,信号进一步集中于约50%的残差流维度,保留85%–98%的基准效果,体现出一种特权基结构(privileged basis structure)。这表明拒绝行为并非在Transformer架构中广泛分布,而是由一个结构化、可识别的机制所构建。该研究揭示了稀疏性在两个层面发挥作用:一是被干预的组件选择,二是组件内承载信号的具体维度,为理解并精确操控拒绝行为的表征与调控提供了可解释的机制基础。所有代码与原始实验结果已公开以保障可复现性。
链接: https://arxiv.org/abs/2610.06903
作者: Vincent Siu,Glenn Grant-Richards,Vlad Pavlovich,Yizhou Sun,Dawn Song,Chenguang Wang
机构: UC Santa Cruz(加州大学圣克鲁斯分校); UCLA(加州大学洛杉矶分校); UC Berkeley(加州大学伯克利分校)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: Accepted to COLM 2026
Abstract:Activation steering manipulates large language model behavior by intervening on internal activations, but the mechanistic basis of these interventions remains poorly understood. We decompose refusal steering into component-level interventions across four open-weight models, identifying the sparse subsets of attention and MLP components whose steering suffices to reproduce the full behavioral effect. We find that refusal directions concentrate in sparse component mechanisms comprising 28–48% of upstream components, retaining 88–101% of steering effectiveness. Within these mechanisms, effective steering further concentrates in approximately 50% of residual stream dimensions, retaining 85–98% of the component-mechanism baseline, consistent with a privileged basis structure. Sparsity thus operates at two levels: which components are steered, and which dimensions within those components carry the signal. Together these findings show that refusal is not diffusely encoded across a transformer but assembled by a structured, identifiable mechanism, providing a foundation for mechanistic understanding of how refusal behaviors are represented and steered. To facilitate reproducibility, we release all code and raw experimental results in this https URL.
[NLP-121] Capacity Responsiveness and Alignment: What Makes a Latent Structure Actionable
【速读】: 该论文旨在解决语言模型(Language Models, LMs)激活空间中潜在结构的可操作性问题,即为何某些局部化结构虽存在于模型内部却难以有效操控其行为。核心挑战在于:并非所有可定位的结构都具备相同的因果影响力,因此需要明确决定结构是否“可行动”的关键因素。论文将因果影响力建模为三个可解释且独立的约束因子的乘积:容量(capacity),衡量模型输出对沿该结构方向移动的敏感程度;响应度(responsiveness),反映在当前上下文下概念被激发的可能性;对齐度(alignment),表征该结构与上下文特定概念表示的一致性。实验结果表明,三者均需处于高水平才能实现有效的因果干预;容量或响应度低下分别导致因果效力下降84%和95%,而对齐度不足甚至可反转因果效应,抑制概念表达。此外,研究发现因果有效性具有强上下文依赖性,而非结构固有属性,有效的因果方向构成一个随上下文变化的低维子空间。基于此,作者提出因果探针(causal probes),通过仅在该动态子空间内训练线性探针,实现了跨模型17%–118%的引导性能提升,同时仅损失3%的概念检测能力,显著提升了可控性和效率。
链接: https://arxiv.org/abs/2610.06897
作者: Or Shafran,Mor Geva
机构: Tel Aviv University (特拉维夫大学)
类目: Computation and Language (cs.CL)
备注:
Abstract:Localizing latent structures in the activation space of language models (LMs) is central to understanding and controlling their behavior. Yet, localized structures can differ substantially in their causal influence, raising the question of what makes a structure actionable. We tackle this question by casting causal influence as a product of three factors and showing empirically that they act as interpretable, distinct constraints: capacity, measuring the sensitivity of the model’s output to movement along the structure, responsiveness, capturing how promotable the concept is given the current context, and alignment, reflecting how well the structure aligns with the context-specific representation of the concept. Across 4 LM families and 50 concepts, we observe that causal effectiveness requires all factors to be high; low capacity and responsiveness reduce it by 84% and 95%, respectively, while low alignment can reverse it, suppressing concept expression. Moreover, we find that causality is context-dependent rather than an intrinsic property of the structure, with causally effective directions forming a low-dimensional subspace that varies across contexts. By restricting the training of linear probes to this subspace, we introduce causal probes that achieve 17%-118% improvement in steering across models, with only 3% reduction in concept detection.
[NLP-122] Zero-Shot Visualization: Exploring Text Corpora with User-Prompted Axes
【速读】: 该论文旨在解决如何利用大语言模型(Large Language Models, LLMs)实现对文本语料的可视化探索,特别是在用户以自然语言描述概念时,如何将文档映射到相应概念轴上进行直观展示的问题。其核心挑战在于,在零样本(zero-shot)场景下,需在特征函数选择、高效实现权衡以及预/后处理策略之间做出合理决策,以保障可视化结果的语义忠实性与质量。为此,研究构建了一个基准测试框架,系统比较了基于嵌入相似性、直接语义判断及条件似然估计等不同方法在该任务中的表现。实验结果表明,基于下一个词概率(next-token probabilities)的评分方法在语义忠实度、得分保真度与计算成本之间取得了最优的实用平衡。进一步地,研究将该方法应用于未标注语料,在真实探索场景中揭示了分级轴设计结合二元相关性过滤的重要性,并发现非主题文档中存在复合情感偏差。基于上述发现,论文提出了构建端到端零样本可视化(Zero-Shot Visualization, ZSV)基线系统的实用指导原则。
链接: https://arxiv.org/abs/2610.06889
作者: Arnau Bueno Tricas,Jose A. Rodríguez-Serrano
机构: Universitat Ramon Llull Esade(拉蒙·柳尔大学埃萨德商学院); Barcelona, Spain(巴塞罗那, 西班牙)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:
Abstract:We study the application of large language models (LLMs) to the visual exploration of textual corpora. We introduce zero-shot visualization (ZSV), a task in which users specify concepts in natural language and documents are mapped onto the corresponding concept axes for visualization. Building a ZSV system of practical value is non-trivial, as it requires choices at the intersection of feature functions, efficient implementation tradeoffs, and pre/post-processing decisions affecting visualization quality. To that end, we establish a benchmark that compares methods spanning embedding similarity, direct semantic judgments, and conditional likelihood estimation in this setting. Across multiple datasets and use cases we evaluate the properties of different scoring methods and design choices in terms of semantic faithfulness, score fidelity, and computational cost. Our results identify that scoring based on next-token probabilities offers the strongest practical trade-off among the evaluated methods. We further apply this approach to unlabeled corpora to examine its behavior in realistic exploratory settings. These experiments highlight additional design considerations, including the use of graded axes together with binary relevance filtering, and reveal a compositional sentiment bias in off-topic documents. Based on these findings, we provide practical guidelines for constructing end-to-end ZSV baselines.
[NLP-123] When Does External Guidance Help LLM Reasoning ? A Bias-Variance Theory of Guidance-Augmented GRPO
【速读】: 该论文旨在解决生成式人工智能(Generative AI)中基于可验证奖励的强化学习(Reinforcement Learning with Verifiable Rewards, RLVR)方法在引入外部引导信号(如专家轨迹、自解释或检索到的思维模式)时缺乏理论保障的问题,具体表现为现有方法未提供收敛速率、偏差界或最优引导权重规则。其解决方案的关键在于提出统一的理论框架——引导增强型GRPO(Guidance-Augmented GRPO, GA-GRPO),将外部引导建模为一个随机引导算子 $ G $,用于重写问题分布;并通过分析由此产生的策略梯度估计器为一种受控偏差的在线策略估计器,证明其偏差被引导增强采样分布与策略自身分布之间的总变差引导发散度 $ \delta_G $ 所有界。在此基础上,在光滑性和有界发散假设下,建立了GA-GRPO以 $ O(1/\sqrt{T}) $ 的速率收敛至GRPO不动点附近 $ O(\delta_G \sqrt{T}) $ 的邻域,并推导出均方误差(MSE)最优引导权重 $ \lambda^*(T, \delta, \sigma_0^2) = \sigma_0^2 / (\sigma_0^2 + R_{\text{max}}^2 \delta^2 T) $,同时通过匹配的极小极大下界证明 $ \Omega(\delta^2 T) $ 的偏差项不可避免。实验在Qwen2.5-Math-7B-Base模型上于九个数学及分布外(OOD)基准测试中验证了该理论预测,结果表明最优权重下的GA-GRPO性能优于或等同于TAPO、LUFFY、ExPO和原始GRPO,且节省31%的GPU小时数,八项分析实验进一步验证了各理论结论的有效性。
链接: https://arxiv.org/abs/2610.06861
作者: Sofia Torres,Gabriel Almeida,Carter Adams,Camila Rocha
机构: Federal University of Bahia (巴西联邦大学)
类目: Machine Learning (cs.LG); Computation and Language (cs.CL)
备注:
Abstract:Reinforcement learning with verifiable rewards (RLVR) has become the dominant paradigm for eliciting multi-step reasoning in large language models, and a recent wave of methods (LUFFY, ExPO, PAPO, TAPO) further augments RL with \emphexternal guidance - expert traces, self-explanations, or retrieved thought patterns. Although each method reports empirical gains, none provides convergence rates, bias bounds, or an optimal weighting rule for the guidance signal. We close this gap with \emphGuidance-Augmented GRPO (GA-GRPO), a unified theoretical framework that casts external guidance as a stochastic guidance operator G re-writing the question distribution, and analyses the resulting policy-gradient estimator as a biased on-policy estimator whose bias is bounded by the total-variation guidance divergence delta_G between the guidance-augmented sampling distribution and the policy’s own distribution. The framework subsumes vanilla GRPO, LUFFY, ExPO, PAPO, and TAPO as special cases obtained by particular choices of G. Under smoothness and bounded-divergence assumptions we prove that GA-GRPO converges at rate O(1/sqrt(T)) to an O(delta sqrt(T))-neighbourhood of the GRPO stationary point, derive the closed-form MSE-optimal guidance weight lambda-star(T, delta, sigma_0 squared) = sigma_0 squared / (sigma_0 squared + R_max squared delta squared T), and prove a matching minimax lower bound showing the Omega(delta squared T) bias term is unavoidable. Experiments on Qwen2.5-Math-7B-Base across nine math and OOD benchmarks confirm that optimal-weight GA-GRPO matches or surpasses TAPO, LUFFY, ExPO, and vanilla GRPO while requiring 31% fewer GPU-hours, and eight analysis experiments validate each theoretical prediction.
[NLP-124] How Much Do LLM -as-a-Judge Design Choices Matter? A Systematic Comparison of Prompt Designs Rating Scales and Models
【速读】: 该论文旨在解决当前在使用大语言模型作为评判者(LLM-as-a-judge)时缺乏标准化设计方法的问题。由于研究者通常凭直觉选择提示词(prompt)、评分量表和模型,这些设计差异可能显著影响评判结果,进而导致不同研究对同一事实得出矛盾结论。为此,研究系统评估了10个推理模型在两种任务上的表现:基于1-7分量表的句子情感与毒性评分(每类超过500项),以及问答对的二分类准确率判断(n=600)。结果显示,尽管评判者与人工标注存在一定程度的分歧,但平均绝对偏差仅为0.11分,表明多数评判者具有较高可靠性;且毒性评判表现甚至优于传统分类器。在准确性任务中,评判者的平均准确率达96.5%。然而,设计选择仍可引发显著偏差:仅改变评分量表即可使测得偏倚最大变动0.93分;在分类任务中,尽管准确性受设计影响较小,但评判宽松度(leniency)却显著变化——使用详细提示可使宽松度下降28.9个百分点,切换模型则可能导致最高56.1个百分点的降幅。值得注意的是,降低推理努力程度对准确性和宽松度均无明显影响。跨两项任务分析均表明,模型身份是变异的主要来源。因此,尽管总体上大语言模型作为评判者具备可信性,但其具体设计对结果具有实质性影响。研究强调,在日益依赖自动化评估的生成式人工智能(Generative AI)研究中,必须重视评判体系的设计一致性,以提升评估的鲁棒性与可复现性,为构建更可靠的LLM-as-a-judge流程提供方法学参考。
链接: https://arxiv.org/abs/2610.05094
作者: Laurène Vaugrante,Thilo Hagendorff
机构: University of Stuttgart(斯图加特大学)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注:
Abstract:Researchers increasingly use Large Language Models as judges (LLM-as-a-judge) to evaluate model outputs. Yet there are no standards for how to design these judges. Typically, researchers choose the prompt, rating scale, and model intuitively. If these choices change the judge’s verdicts, two studies can reach different conclusions about the same facts. To address this risk and to provide an empirical basis for judge designs, we evaluate 10 reasoning models across multiple designs on two tasks: a scalar rating of sentence sentiment and toxicity (over 500 items per category), as well as a binary accuracy classification of question-answer pairs (n=600). For the rating tasks, despite judges showing significant disagreements with the human ground truth, the practical size of differences is small enough to consider most judges reliable (mean absolute deviation of 0.11 points on a 1 - 7 scale); toxicity judges even outperform standard classifiers. Judges are also highly accurate on average (96.5%) for the accuracy classification task. However, design choices can produce shifts: changing the rating scale alone can shift measured bias by up to 0.93 points (rating task), and while accuracy levels are rarely impacted, design choices consistently impact judge leniency (classification task; leniency drop of 28.9 percentage points when using detailed prompts, and up to 56.1 percentage points when switching models). Counterintuitively, lower reasoning effort affects neither accuracy nor leniency. Across both tasks, model identity is the dominant source of variance. These findings suggest that while LLM judges are broadly trustworthy in aggregate, design choices can be meaningful sources of variance. Given the growing reliance on automated evaluation in LLM research, we intend this study as a methodological reference for designing more robust and replicable LLM-as-a-judge pipelines.
[NLP-125] OncoNoteBERT: A Foundation Representation Model for Natural Language Processing of Real-World Outpatient Oncology Notes NEURIPS2026
【速读】: 该论文旨在解决通用生物医学或邻近临床语言模型在处理真实世界门诊肿瘤学笔记时的适应性不足问题,尤其针对其中包含的专业术语、肿瘤分期表达、治疗名称、毒性描述及机构特异性去标识化标记等复杂内容。其核心解决方案在于开发并评估专用于肿瘤学领域的BERT风格编码器,通过两种本地化策略实现:一是基于持续掩蔽语言建模预训练的OncoNote-RadBERT,二是从头训练并采用肿瘤学自定义WordPiece分词器的OncoNoteBERT。研究发现,尽管外部预训练模型(RadBERT与PathologyBERT)在零样本评估中表现不佳(困惑度分别为113.04和2035.03),但经过持续预训练的OncoNote-RadBERT在语料库层面表现出最佳拟合(困惑度2.10)。然而,OncoNoteBERT虽在整体困惑度上略高(2.83),却展现出更优的分词效率,表现为更低的子词丰度与更短的归一化序列长度,并在13个掩蔽词探测任务中有12个获得临床可接受的预测结果,优于OncoNote-RadBERT的7个。这一性能差异主要源于分词器碎片化而非仅语义学习所致。此外,两类本地模型均能将机构占位符表示为单一可学习标记,表明本地化分词器设计与表示层架构对提升肿瘤学自然语言处理效果具有关键作用。因此,该研究强调了持续适配与专用分词机制的互补价值,凸显了在应用邻域外编码器前优化表示层设计的重要性。
链接: https://arxiv.org/abs/2610.03829
作者: Wuraola Oyewusi,Eliana Vasquez Osorio,Goran Nenadic,Gareth Price
机构: The University of Manchester (曼彻斯特大学); The Christie NHS Foundation Trust (克里斯蒂国家卫生服务基金会信托); Department of Computer Science (计算机科学系)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Quantitative Methods (q-bio.QM)
备注: Accepted at the AI at Scale for Clinical Impact (ASCI): Cancer Pathology Foundation Models Workshop at NeurIPS 2026
Abstract:Real-world outpatient oncology notes contain specialised terminology, tumour staging expressions, treatment names, toxicity descriptions, and institution-specific de-identification markers that may not be represented efficiently by general biomedical or adjacent clinical language models. We developed and evaluated oncology-specific BERT-style encoders using a governed UK outpatient oncology corpus comprising 290,026 notes from 21,564 patients treated for lung and head-and-neck cancer. We compared RadBERT and PathologyBERT with two local strategies: OncoNote-RadBERT, produced by continued masked language model pretraining, and OncoNoteBERT, trained from scratch with an oncology WordPiece tokenizer. Models were evaluated using masked language modelling loss and perplexity on the validation set, tokenizer fragmentation metrics, clinical term tokenisation, masked-token probes, and exploratory representation analysis. Both external encoders fit the oncology corpus poorly in zero-shot evaluation (perplexity 113.04 for RadBERT; 2035.03 for PathologyBERT), while continued pretraining produced the strongest fit (2.10 for OncoNote-RadBERT). OncoNoteBERT achieved perplexity 2.83 but produced the most efficient tokenisation, with lower subword fertility and shorter normalised sequence length. It also returned a clinically acceptable prediction for 12 of 13 masked-token probes, compared with 7 of 13 for OncoNote-RadBERT. This divergence between corpus-level fit and masked-token performance was partly attributable to tokenizer fragmentation rather than learned semantics alone. Both locally developed models represented the institutional placeholder as a single learnable token. These findings show that continued adaptation and bespoke tokenisation provide complementary benefits, and that representation-layer design matters before adjacent-domain encoders are applied to oncology NLP.
[NLP-126] Conversation Is a Two-Body Problem: Dyadic Evaluation of Full-Duplex Dialogue Models
【速读】: 该论文旨在解决全双工语音对话模型(full-duplex spoken dialogue models)在评估过程中存在的根本性缺陷:现有评估框架多采用单边对话者(single-sided interlocutors),如预录音频或固定测试流程的自动化考评系统,无法真实反映双人交互中双方协同产生的互动行为(如交替发言、重叠说话和打断等)。这种评估方式仅考察模型单方面表现,忽略了对话中双向耦合的本质。其解决方案的关键在于提出DyaFDB框架,该框架通过构建双模型对话语境,使两个全双工模型在设定角色与目标(合作或冲突)下直接进行实时对话,并由外部评判员对双方表现进行离线评分。该设计能够动态捕捉模型间的相互影响机制,揭示模型行为如何持续重塑对方,从而证明每个模型必须同时承担“考评者”与“被考评者”的双重角色,任何单一固定的对话者均无法完整模拟这一双向交互过程。研究共构建140个场景,生成7,560次对话,覆盖六种自洽与跨模型组合,验证了该评估范式在揭示真实交互动态方面的有效性。
链接: https://arxiv.org/abs/2610.08125
作者: Sungnyun Kim,Sungwoo Cho,Jihwan Oh,Se-Young Yun
机构: Korea Advanced Institute of Science and Technology (韩国科学技术院)
类目: Audio and Speech Processing (eess.AS); Computation and Language (cs.CL)
备注: Project page: this https URL
Abstract:Full-duplex spoken dialogue models listen and speak at the same time, enabling voice agents to have natural, low-latency interactions that turn-based systems cannot offer. However, they are commonly evaluated against single-sided interlocutors: pre-recorded audio that cannot react, or an automated examiner that reacts in real time but only administers a fixed sequence of tests and is never graded. These single-sided frameworks evaluate only half of a two-body problem, where turn-taking, overlap, and interruption are joint products of two coupled speakers. We propose DyaFDB, a framework that evaluates full-duplex models in a dyadic setup: two models converse directly under assigned roles with cooperative or conflicting goals, and both sides are scored offline with an external judge. DyaFDB probes how the two models behave toward each other, such as how they take turns or carry an assigned role under different interests. We instantiate four tasks as 140 scenarios and record 7,560 conversations, covering six self- and cross-play pairings. Throughout the experiments, we observe that how a model behaves continually reshapes its partner. We thus demonstrate that each model must be both the examiner and examinee of the other, and no single fixed interlocutor can play both parts. We will release the scenarios, role prompts, and recording protocols between two full-duplex models, without any pre-recorded audio.
[NLP-127] HINTT Submission to the 2nd MLC-SLM Challenge: Comparing Cascaded and Unified Approaches to Diarization and ASR
【速读】: 该论文旨在解决多语言对话式语音识别中的话语者归属(speaker-attributed ASR)问题,即在多说话人场景下准确识别“谁在何时说了什么”。其核心挑战在于同时实现高精度的语音转写与精确的话语者分离。解决方案的关键在于对比两种建模策略:一是级联式流水线方法,将话语者分割(diarization)与基于大语言模型(LLM-based)的语音识别(ASR)模块分步结合;二是统一式语音大模型(unified speech LLM),直接联合生成说话人标签、时间戳和文本转录。实验表明,在任务1条件下,级联式系统表现更稳定可靠,其最终方案由微调后的DiariZen话语者分割模型、Qwen3-ASR语音识别模型及基于LLM的生成式纠错模块构成。相比之下,统一式语音大模型虽尚不成熟,但展现出未来端到端话语者归属语音识别的潜力。所有模型微调均仅使用官方提供的MLC-SLM训练数据,未引入外部数据或伪标签,确保了评估的公平性与可复现性。
链接: https://arxiv.org/abs/2610.08063
作者: Takanori Ashihara,Kohei Matsuura,Masato Mimura
机构: 未知
类目: Audio and Speech Processing (eess.AS); Computation and Language (cs.CL)
备注:
Abstract:This paper presents the HINTT system submitted to the 2nd Challenge and Workshop on Multilingual Conversational Speech Language Model (MLC-SLM). We address multilingual speaker-attributed ASR, where systems must determine who spoke when and what was spoken. We investigate two modeling strategies for this problem: a cascaded pipeline that combines speaker diarization with speech-LLM-based ASR, and a unified speech LLM that directly generates speaker labels, timestamps, and transcriptions. Our final submission is based on the cascaded pipeline, consisting of a fine-tuned DiariZen diarization model, a fine-tuned Qwen3-ASR model, and LLM-based generative error correction. For comparison, we also fine-tune VibeVoice-ASR as a unified model using the same official training data. All task-specific fine-tuning and model selection are performed using only the official MLC-SLM data, without external data or pseudo-labels. Experimental results demonstrate that the cascaded system remains more reliable under the MLC-SLM Task 1 conditions, while unified speech LLMs offer a promising direction for future speaker-attributed ASR.
[NLP-128] A Novel Sentence Stress Detection Framework Leverag ing Auxiliary Word-Stress Modeling and Loss Optimization INTERSPEECH2026
【速读】: 该论文旨在解决自动发音评估(APA)中韵律重音识别的两大任务——句子重音检测(SSD)与词重音检测(WSD)长期被视作独立任务,忽略了二者共享的韵律线索(如基频、时长和强度)这一关键问题。其核心解决方案是提出一种新型建模范式,通过联合建模SSD与辅助的WSD,实现两者的协同优化;同时引入词跨度重音正则化器(WSR),将每个重读词内的分词级重音概率集中于该词的重音跨度内,从而增强模型对词内重音结构的捕捉能力。实验在TinyStress-15K基准数据集上验证了该方法的有效性,完整配置显著优于现有强基线,实现了最优的SSD性能。
链接: https://arxiv.org/abs/2610.07626
作者: Tien-Hong Lo,Fong-Chun Tsai,Ting-An Hung,Yu-Hsuan Hsieh,Yao-Ting Sung,Berlin Chen
机构: 未知
类目: Audio and Speech Processing (eess.AS); Computation and Language (cs.CL)
备注: Interspeech 2026
Abstract:Prosodic stress is a crucial aspect of automatic pronunciation assessment (APA), encompassing both sentence stress detection (SSD) and word stress detection (WSD). SSD highlights semantically salient words that shape discourse meaning, while WSD identifies the primary stressed syllable within each word to ensure lexical clarity. However, most prior work treats SSD and WSD as independent tasks, overlooking their shared reliance on prosodic cues such as pitch, duration, and intensity. To address this gap, we propose an effective SSD approach combining SSD with auxiliary WSD via a novel modeling paradigm. In addition, we introduce a word-span stress regularizer (WSR) that concentrates token-level SSD probabilities within each stressed word span. Experiments on the TinyStress-15K benchmark show that the proposed method outperforms strong baselines, with the complete configuration achieving the best SSD result.
[NLP-129] Logbook: Extremely Long-form Audio Event Understanding ICASSP2027 KR
【速读】: 该论文旨在解决现有音频基准测试(audio benchmarks)普遍依赖短时预分割片段所导致的局限性,即模型设计被限制在短输入或固定词汇表,难以适应真实场景中长时间、连续音频的理解需求。为此,研究提出了Logbook——一个面向小时级音频理解的新型基准,涵盖时长从10分钟至6天的连续录音。其核心任务是在给定连续音频与事件标签词汇的前提下,实现无间隙的分段标注,为每个片段输出对应的事件标签与描述。解决方案的关键在于构建一个支持长时序建模的端到端(end-to-end)与级联式(cascaded)系统对比框架,并通过消融实验分析微调(fine-tuning)、上下文长度(context length)及推理预算(reasoning budget)对性能的影响。研究发现,尽管任务具有可解性,但当前最优系统仍低于人类参考水平;过度分段(over-segmentation)现象普遍存在,而微调可在一定程度上缓解该问题;此外,端到端模型整体表现优于级联系统,但在处理更长上下文时性能下降。
链接: https://arxiv.org/abs/2610.07338
作者: Kwanghee Choi,Suwon Shon,Dmitriy Serdyuk,Guitang Lan,Chao-Wei Huang,Mohammad Sadegh Rasooli,Sangeeta Srivastava,Zhaojiang Lin,Saurabh Adya,Ming Sun
机构: 未知
类目: Audio and Speech Processing (eess.AS); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG); Sound (cs.SD)
备注: Submitted to ICASSP 2027. Source code available at this https URL
Abstract:Audio benchmarks are built around short, pre-segmented clips, limiting model design to brief inputs or fixed vocabularies. To close this gap, we introduce Logbook, a benchmark for hour-scale audio understanding, with recordings ranging from ten minutes to six days. Given a continuous audio recording and an event label vocabulary, a system must predict a gap-free segmentation with an event label and a description per segment. We compare 52 systems, end-to-end and cascaded, and ablate fine-tuning, context length, and reasoning budget. We find the task tractable, though the best systems remain below the human reference. Also, over-segmentation is pervasive, and fine-tuning partially mitigates it. Finally, end-to-end are often better than cascaded systems, but degrades with longer context.
[NLP-130] SEAL: Mixture-Closed Additive Reconstruction and Refinement-Aware Expert Routing for Efficient Speech Separation ICASSP2027
【速读】: 该论文旨在解决现有紧凑型时频分离器在信号分离任务中面临的两大局限:一是受限的乘性掩码仅能缩放混合信号的频谱单元,导致在不同成分相消区域估计值过小;二是共享单元在每一步对所有时频令牌应用相同权重,导致模型扩展时计算开销全局增加。其解决方案的关键在于提出SEAL(Sparse Expert routing with Additive Latent reconstruction),通过两项创新机制实现突破:在重构方面,采用基于局部混合幅值约束的零和加性残差结构,使估计值在成分相消区域仍可非零,同时保证重建结果与原始混合信号一致;在路由方面,利用融合声学特征与跨步间上下文证据的查询向量,将每个时频令牌分配至六个残差专家之一,并引入归一化上限机制,防止步骤间提示过度干扰明确的声学线索。实验表明,SEAL(小)在EchoSet数据集上以28%更少参数、2.9倍更低的乘加操作数(MACs)实现比TIGER(小)高0.31 dB的SI-SDRi性能,而SEAL(大)在仅需TIGER(大)约三分之一MACs的情况下,性能差距不足0.07 dB,显著提升了模型效率与分离精度。
链接: https://arxiv.org/abs/2610.07047
作者: Shao-Chun Hu,Zi-Xiang Lin,Jeih-Weih Hung,Hung-Shin Lee
机构: 未知
类目: Audio and Speech Processing (eess.AS); Computation and Language (cs.CL); Sound (cs.SD)
备注: Submitted to ICASSP 2027
Abstract:Compact time-frequency separators that mask the mixture and refine through a shared cell face two limits. First, a bounded multiplicative mask only scales a mixture bin, so where overlapping components cancel, the estimate stays small. Second, a shared cell applies the same weights to every time-frequency token at every step, so enlarging it adds compute everywhere. We present SEAL (Sparse Expert routing with Additive Latent reconstruction) to address both. For reconstruction, a zero-sum additive residual bounded by the local mixture amplitude lets estimates be nonzero where components cancel yet still sum to the mixture. For routing, a query built from acoustic and inter-step evidence sends each token to one of six residual experts, and a norm cap keeps the step cue from overriding clear acoustic evidence. On EchoSet, SEAL (small) surpasses TIGER (small) by 0.31 dB SI-SDRi with 28% fewer parameters and 2.9 times fewer MACs, and SEAL (large) is within 0.07 dB SI-SDRi of TIGER (large) at 3.1 times fewer MACs.
[NLP-131] GIVE-KWS: Gated Injection of Visual Evidence for Noise-Robust Query-by-Example Keyword Spotting ICASSP2027
【速读】: 该论文旨在解决在低信噪比环境下,传统文本与音频联合驱动的关键词语音检测(KWS)系统因噪声干扰导致性能下降的问题,尤其关注视觉模态在提升系统鲁棒性方面的潜力。尽管视觉语音(visual speech)具有天然的抗噪优势,但现有方法中视觉编码器缺乏对音素信息的有效建模,导致其难以充分发挥作用。本文提出GIVE-KWS框架,其核心创新在于融合阶段引入“门控视觉证据注入”(Gated Injection of Visual Evidence, GIVE),通过门控交叉注意力机制将查询音频与唇动信息进行条件化关联,实现视觉证据的主动注入而非简单重加权。研究表明,视觉鲁棒性的关键在于两个相互作用的条件:一是视觉编码器需具备音素级表征能力,二是融合策略应以注入视觉证据为主导而非对音频特征进行缩放。实验表明,在具备音素表征能力的编码器下,GIVE相比掩蔽策略在-10 dB信噪比下可实现4.0–9.3 dB的有效信噪比增益;而当编码器缺乏音素信息时,该增益几乎消失。相较于基准系统,GIVE-KWS在-10 dB下使未见关键词的等错误率(EER)降低72.9%,平均降低62.8%,显著提升了跨模态融合的性能表现。
链接: https://arxiv.org/abs/2610.07046
作者: Ming-Hsiang Hu,Kuan-Tang Huang,Hung-Shin Lee,Berlin Chen
机构: 未知
类目: Audio and Speech Processing (eess.AS); Computation and Language (cs.CL); Multimedia (cs.MM); Sound (cs.SD)
备注: Submitted to ICASSP 2027
Abstract:Visual speech promises noise-robust keyword spotting, yet a visual stream is not necessarily used. On a tri-modal query-by-example keyword spotting (QbyE-KWS) benchmark, we find that a system with a task-trained visual encoder comes within 2 percentage points of a text-and-audio system in equal error rate (EER) at -10 dB, and link this gap to the encoder’s lack of phonemic information. We present GIVE-KWS, whose fusion stage, GIVE (Gated Injection of Visual Evidence), conditions query audio on lip motion through gated cross-attention. We show that visual robustness depends on two interacting conditions: a phoneme-bearing visual representation, and fusion that injects visual evidence rather than rescaling audio features. Under a phoneme-bearing encoder, injection yields an effective SNR gain of 4.0-9.3 dB over masking at -10 dB, whereas under a phoneme-poor one it nearly vanishes. Relative to the benchmark system, GIVE-KWS reduces unseen-keyword EER by 72.9% at -10 dB and 62.8% on average.
[NLP-132] Nord-Parl-TTS: Finnish and Swedish TTS Dataset from Parliament Speech ICASSP2026
【速读】: 该论文旨在解决低资源语言在文本到语音(Text-to-Speech, TTS)技术发展中因高质量、公开可用语音数据稀缺而导致的瓶颈问题,尤其针对芬兰语和瑞典语等非高资源语言。其核心解决方案是构建并发布Nord-Parl-TTS——一个基于北欧议会会议录音的开源TTS数据集,从中提取了900小时芬兰语和5090小时瑞典语的语音数据,适用于TTS模型训练。该数据集采用改进版的Emilia数据处理流程,并包含统一的评估集以支持模型开发与基准测试。关键在于通过“野外”获取的真实场景语音数据,实现了大规模、低成本且高质量的多语言语音数据供给,显著缩小了高资源语言与低资源语言在TTS技术上的资源差距。
链接: https://arxiv.org/abs/2509.17988
作者: Zirui Li,Jens Edlund,Yicheng Gu,Nhan Phan,Lauri Juvela,Mikko Kurimo
机构: 未知
类目: Audio and Speech Processing (eess.AS); Computation and Language (cs.CL)
备注: Accepted by ICASSP 2026. 5 pages, 2 figures
Abstract:Text-to-speech (TTS) development is limited by scarcity of high-quality, publicly available speech data for most languages outside a few high-resource languages. We present Nord-Parl-TTS, an open TTS dataset for Finnish and Swedish based on speech found in the wild. Using recordings of Nordic parliamentary proceedings, we extract 900 hours of Finnish and 5090 hours of Swedish speech suitable for TTS training. The dataset is built using an adapted version of the Emilia data processing pipeline and includes unified evaluation sets to support model development and benchmarking. By offering open, large-scale data for Finnish and Swedish, Nord-Parl-TTS narrows the resource gap in TTS between high- and lower-resourced languages.
[NLP-133] Pronunciation Editing for Finnish Speech using Phonetic Posteriorgrams
【速读】: 该论文旨在解决低资源语言(low-resourced languages)第二语言(L2)语音合成中因缺乏高质量L2语音数据集而导致的合成困难问题。针对这一挑战,其核心解决方案是提出一种基于扩散模型的多说话人音素后验图(Phonetic-Posteriorgrams, PPGs)到语音(PPG2Speech)的生成方法,能够无需文本对齐即可对单个音素进行编辑。该方法以Matcha-TTS的流匹配解码器为骨干网络,将PPGs转换为梅尔频谱图,并通过外部说话人嵌入和基频信息进行条件控制。为提升生成质量与编辑精度,引入无分类器引导(Classifier-free Guidance, CFG)和摇摆采样(Sway Sampling)策略。此外,论文提出了一种新的任务特定评估指标——音素对齐一致性(Phonetic Aligned Consistency, PAC),用于量化编辑前后音素后验图与合成语音提取的音素后验图之间的匹配程度,从而评估编辑效果。实验在芬兰语(一种低资源、近似音素对应的语言)上进行,仅使用约60小时的数据,通过客观与主观评估验证了该方法在语音自然度、说话人相似性及编辑有效性方面优于传统基于文本到语音(TTS)的编辑方法。
链接: https://arxiv.org/abs/2507.02115
作者: Zirui Li,Lauri Juvela,Mikko Kurimo
机构: 未知
类目: Audio and Speech Processing (eess.AS); Computation and Language (cs.CL)
备注: Accepted by Proceeding of 13th edition of the Speech Synthesis Workshop; 5 pages, 1 figure
Abstract:Synthesizing second-language (L2) speech is potentially highly valued for L2 language learning experience and feedback. However, due to the lack of L2 speech synthesis datasets, it is difficult to synthesize L2 speech for low-resourced languages. In this paper, we provide a practical solution for editing native speech to approximate L2 speech and present PPG2Speech, a diffusion-based multispeaker Phonetic-Posteriorgrams-to-Speech model that is capable of editing a single phoneme without text alignment. We use Matcha-TTS’s flow-matching decoder as the backbone, transforming Phonetic Posteriorgrams (PPGs) to mel-spectrograms conditioned on external speaker embeddings and pitch. PPG2Speech strengthens the Matcha-TTS’s flow-matching decoder with Classifier-free Guidance (CFG) and Sway Sampling. We also propose a new task-specific objective evaluation metric, the Phonetic Aligned Consistency (PAC), between the edited PPGs and the PPGs extracted from the synthetic speech for editing effects. We validate the effectiveness of our method on Finnish, a low-resourced, nearly phonetic language, using approximately 60 hours of data. We conduct objective and subjective evaluations of our approach to compare its naturalness, speaker similarity, and editing effectiveness with TTS-based editing. Our source code is published at this https URL.
信息检索
[IR-0] A Systematic Study of Semantic ID Spaces for Generative Information Retrieval
链接: https://arxiv.org/abs/2610.08732
作者: Alexia Allal,Hicham Randrianarivo,Sylvain Lamprier
类目: Information Retrieval (cs.IR); Computation and Language (cs.CL)
备注: 8 pages, 3 figures, 1 table
Abstract:Generative Information Retrieval (GIR) has emerged as a transformative paradigm, shifting document retrieval from a traditional “retrieve-and-rank” workflow to sequence-to-sequence generation, where a model directly predicts document identifiers (DocIDs). While the semantic design of these DocIDs is known to be critical for performance, a fundamental question remains under-explored: what makes a good DocID? Current approaches rely heavily on computationally expensive downstream evaluations, hindering systematic analysis and rapid iteration. In this work, we address this challenge by presenting a comprehensive study on the properties, metrics, and trade-offs that define effective numerical DocIDs. Specifically, our contributions are threefold: First, we propose a unified framework that unifies Product Quantization (PQ) and Residual Quantization (RQ), and their hybrid variants within a single design space. This enables us to systematically study key DocID properties, such as hierarchy versus parallelism, as well as the impact of hyperparameters like DocID length and codebook size. Second, we define a suite of training-free, intrinsic metrics, to quantify DocID quality and evaluate structural fidelity without the overhead of full model training. Through extensive experiments on MS MARCO 300K and NQ320K, we analyze how these structural properties influence retrieval effectiveness.
[IR-1] Disentangling Paradigm Identifier and Decoding in Generative Retrieval
链接: https://arxiv.org/abs/2610.08716
作者: Hicham Randrianarivo,Logan Renaud,Alexia Allal
类目: Information Retrieval (cs.IR); Computation and Language (cs.CL)
备注: 13 pages, 7 figures, 11 tables
Abstract:Generative retrieval trains a language model to generate the identifier of a relevant document. Recent work replaces the autoregressive decoder with diffusion, but changes identifiers, training recipe and decoding at once, so differences cannot be credited to the paradigm. On NQ320K and MS300K, we train autoregressive, masked-diffusion and block-diffusion models with residual-quantised, product-quantised and random identifiers. With identifier length and training budget fixed, we decode each model in several ways. Decoding alone moves a diffusion model’s Hit@1 by 6.6 to 13.7 points. Our reference diffusion decoding, generate-and-match, generates an identifier, then retrieves the closest corpus identifiers. The generated identifier is right for 14-21% of NQ320K queries. We test one-pass scoring to decode diffusion retrievers: the model reads a fully masked identifier once, and each document is scored by its codes’ probabilities. It matches or beats generate-and-match in 11 of 12 settings. Autoregressive models still lead in Hit@1; on NQ320K, the lead comes from the model, not beam search. Starting from one sampled identifier, one-pass scoring removes 46-83% of masked diffusion’s deficit to beam search; from generate-and-match, at most a quarter. On NQ320K, every paradigm largely memorises which identifier answers which query: random identifiers keep 83-90% of the Hit@1 of residual-quantised ones. There, product-quantised identifiers lead residual-quantised ones by 3.4 points in the autoregressive model and by -0.7 to +3.6 in diffusion models; across decodings, AR’s gap exceeds diffusion’s by 1.5-2.3 points, around our 2-point threshold. Paradigm comparisons must report each paradigm at its own recipe and best decoding.
[IR-2] UNREAL: Unifying Retrieval and Long-Context with a Single Model
链接: https://arxiv.org/abs/2610.08463
作者: Edan Kinderman,Elad Hoffer,Yochai Blau,Brian Chmiel,Ron Banner,Daniel Soudry,Boris Ginsburg
类目: Computation and Language (cs.CL); Information Retrieval (cs.IR); Machine Learning (cs.LG)
备注:
Abstract:Long-context inference and Retrieval-Augmented Generation (RAG) handle evidence selection at vastly different scales, from a single long prompt to an entire corpus. We ask whether a single model-internal mechanism can select evidence across this range. We introduce UNifying REtrieval And Long-Context with a Single Model (UNREAL), a model-native evidence selection framework to span corpus retrieval and long-context inference. UNREAL encodes chunks and derives retrieval queries directly from the frozen LLM’s internal representations. It adds fewer than 500K trainable parameters and leaves the backbone unchanged. On a 3B-token, 21M-chunk Wikipedia index, all four dense and hybrid UNREAL backbones outperform state-of-the-art retriever-reranker systems. The best model raises recall from 49.1% to 73.2% on HotpotQA, from 31.7% to 60.1% on 2WikiMultiHopQA, and from 8.8% to 14.4% on MuSiQue. Applied to long-context tasks, the same selection mechanism removes distractors before generation, raising NoLiMa accuracy from 1.0% to 24.83% at its maximum context length of 128K tokens, and LV-Eval’s F1 score from 49.97% to 54.66% at 256K. UNREAL also reduces FLOPs and time-to-first-token relative to full-context inference from roughly 32K tokens onward, with larger gains as context grows. Together, these results establish model-internal evidence selection as a common foundation for corpus retrieval and evidence-sparse long-context inference.
[IR-3] Agent ic AutoRAG RAG : RAG Pipeline Optimization through Reasoning -Driven Agents NEURIPS2026 EMNLP2026
链接: https://arxiv.org/abs/2610.08452
作者: Lasse B. Strand,Robert Jakob,Kevin O’Sullivan,Markus Kreft
类目: Computation and Language (cs.CL); Information Retrieval (cs.IR); Machine Learning (cs.LG)
备注: Accepted at the Second Workshop for REsearch on Agent Language Models (REALM) at EMNLP 2026 and at the Machine Learning for Systems Workshop at NeurIPS 2026. 9 pages plus references and appendix (16 pages total), 4 figures, 6 tables. Code: this https URL
Abstract:Retrieval-augmented generation (RAG) is a widely used approach for grounding large language models (LLMs) in external knowledge. However, configuring a pipeline is an expensive hyperparameter optimization problem over many interacting choices, from chunking and embedding model to reranking and generation. Existing optimizers, from greedy search to Bayesian optimization, reduce each trial to an aggregate score and search without modeling why a configuration performed as it did, even though the retrieved chunks already provide evidence about whether each failure occurred during retrieval or after it. We introduce Agentic AutoRAG, an LLM-agent optimizer for multi-objective RAG hyperparameter optimization with retrieval-versus-generation failure attribution. It proposes configurations scored on a frozen exam from the corpus: after each trial a Diagnoser attributes each failed question to retrieval or generation, and a Proposer, grounded in a knowledge base of model rankings and pricing, selects the next configuration, weighing accuracy against cost to trace a Pareto frontier. On three multi-hop QA benchmarks it reaches higher LLM-judge accuracy than every baseline we compare, and within its first 10 trials it matches or beats the statistical baselines’ full 30-trial judge accuracy. In its cost-aware mode on a real-world healthcare corpus it reaches a median exam accuracy of 77%, above the strongest baseline’s 71.5%, at about 58% of that baseline’s cost per query, and it matches that 71.5% at about 22% of the cost.
[IR-4] Seeing the Context: Enhancing Recommender Systems with Image-Derived Contextual Signals RECSYS2026
链接: https://arxiv.org/abs/2610.08407
作者: Tal Cordova,Tomer Geva,Moshe Unger
类目: Information Retrieval (cs.IR)
备注: Accepted at the CARS workshop, RecSys 2026. 8 pages, 2 figures
Abstract:Contextual information, capturing the circumstances of a user-item interaction, is central to recommender systems. Prior work draws context from location, time, or reviews, but not images; multimodal recommender systems mainly use images to enrich item or user representations, not identify situational context. We propose a new representation of context derived from images, spanning physical, social, and modal categories learned via a vision-language model. We introduce ICE-Fuse, a pipeline for evaluating this representation that fuses these categories and integrates them into a context-aware recommender system, using TripAdvisor data and Review-aware Graph Contrastive Learning as the recommendation algorithm. Image context does not outperform established signals standalone, but improves them combined, indicating complementary information. Semantic analysis shows image- and review-derived context capture distinct aspects of the interaction, positioning images as complementary context.
[IR-5] Aligning Performance with Contribution: Towards Contribution-Aware Fair Recommendation
链接: https://arxiv.org/abs/2610.08245
作者: Shuai Zhang,Hui Fang,Zun Sun
类目: Information Retrieval (cs.IR)
备注:
Abstract:Existing research on user fairness in recommender systems has developed diverse objectives. However, it has paid limited attention to a distinct distributive perspective: whether users’ contributions to model learning should be reflected in the recommendation benefits they receive. We argue that, in addition to existing fairness protections, a fair system may account for the alignment between users’ estimated contributions and the recommendation performance they receive. Such alignment can incentivize sustained and informative engagement, thereby supporting a sustainable recommendation ecosystem. To this end, we propose Contribution-Performance Fairness, a novel fairness perspective which requires recommendation performance to be aligned with estimated contribution across user groups and to remain equitable among users with comparable contributions within a same group. To instantiate this perspective, we introduce the Contribution-Performance Fair Recommender (CPFR), a framework applicable to different backbone recommenders. CPFR constructs ordered user groups from a training-dependent contribution considering interaction volume, loss alignment, and optimization intensity, and jointly optimizes recommendation accuracy with the two fairness requirements. A game-theoretic analysis shows that such alignment can strengthen contribution incentives and improve system-level recommendation accuracy under voluntary contribution. Experiments on three datasets and three backbone models demonstrate that CPFR achieves a strong accuracy–fairness trade-off under the proposed operational metric.
[IR-6] Behavior-Mining Generative Conversations and Collaborative Advisory: the Future of Travel and Tourism Recommender Systems
链接: https://arxiv.org/abs/2610.08232
作者: Alejandro Bellogín,Linus W. Dietz,Francesco Ricci,Pablo Sánchez
类目: Information Retrieval (cs.IR)
备注:
Abstract:Since the early adoption of e-commerce, travel and tourism has been a lab for the design of recommender systems: tools that help travelers choose destinations, flights, accommodations, and combine them into itineraries. Data-driven recommendation techniques, ranging from case-based reasoning to reinforcement learning, have been adapted to travelers’ needs. The research community has produced multifaceted prototypes of travel and tourism recommender systems (TTRSs), which are context-dependent, multistakeholder-oriented, and more recently, addressing sustainability issues, such as overtourism. Despite this enduring work, TTRSs are not widespread yet. We argue that three limitations can explain this: outdated and sparse data sets used to train and validate TTRSs, algorithms that prioritize prediction accuracy over domain-specific dimensions such as novelty and contextual relevance, and a failure to address the specific needs of travelers. Targeted incremental research could address these limitations, but a disruptive factor has meanwhile entered the ecosystem of tourism information and commercialization platforms: generative artificial intelligence. According to market research, GenAI applications are becoming the primary entry point for travelers planning their trips. This forces research to rethink how TTRSs should be designed and which core techniques should be integrated. We claim that future TTRSs, in addition to offering personalized information filtering, should become more flexible advisors that support decision making, integrating multiple data types and AI techniques, from data mining to natural language processing. Moreover, they must transparently balance the conflicting goals of travelers, service suppliers, platform owners, and local communities. We then outline research targets for building more effective TTRSs, fruitfully combining old and new recommendation techniques.
[IR-7] Confidence-Ordering Reversal under Contextual Priors in Neural Decoding
链接: https://arxiv.org/abs/2610.08229
作者: Xinyu Zhang,Sichao Liu
类目: Artificial Intelligence (cs.AI); Information Retrieval (cs.IR); Neurons and Cognition (q-bio.NC)
备注: 28 pages, 4 figures, 18 tables
Abstract:Contextual priors improve neural-to-language decoding by reshaping candidate scores. However, confidence is read from the same reshaped scores, so the errors a prior leaves behind can become more confident with no change in accuracy to reveal it. We study how a prior shapes confidence in speech retrieval on MEG-MASC and MOUS using local decoding scores, a contextual prior combined by additive shallow fusion, and the fused top-two margin as confidence. Among initially incorrect predictions, we find a confidence-ordering reversal: a larger margin makes a repair more likely when the correct candidate starts near the top of the local ranking, but less likely when it starts lower. On MEG-MASC, pooled correctness AUROC is 0.87, yet AUROC separating repairs from residual errors falls from 0.70 at initial ranks 2-3 to 0.39 at ranks 21-50. Errors starting beyond rank 20, inside the reversed region, make up 46.6% of all post-fusion errors. We propose a score-level account: a repair must first close the correct candidate’s initial deficit, limiting its final margin, whereas a residual error can build a large margin between two incorrect candidates. A causal intervention that changes only the fusion weight moves the reversal to deeper ranks as predicted. Under a word-level LM prior, it keeps moving after accuracy gain peaks, so a weight chosen for accuracy does not settle confidence. Reading local and prior scores separately improves selective decoding: the decoder answers on 74.5% of windows instead of 56.7%, while 92% of output sets still contain the correct candidate. Confidence after contextual fusion should retain the local and contextual evidence behind each prediction, not just the fused scores. Project website: this https URL Code: this https URL
[IR-8] Adapting Generative Recommenders for Multi-Turn Interaction
链接: https://arxiv.org/abs/2610.08136
作者: Yu-Chen Den,Zhi Rui Tam,Yung-Yu Shih,Shih-Hsin Wang,Yun-Nung Chen,Pu-Jen Cheng,Eugene Yang
类目: Information Retrieval (cs.IR)
备注:
Abstract:Generative recommenders decode items from a user’s interaction history, but offer no way for users to correct a recommendation that misses their current intent. Adding conversation is natural since items and words share same output space, yet training the model to converse may overwrite the history-to-item mapping it relies on. We introduce INTEGER (INTEractive GEnerative Recommendation), which extends generative recommendation to multi-turn interaction with a learned routing token that lets the model decide when to recommend, history re-anchoring that conditions each item on both past behavior and the dialogue, and behavioral replay with instruction-data rehearsal that prevents forgetting during adaptation. Users can thus give feedback on recommendations within the dialogue, while recommendations stay grounded in behavioral history and accuracy is not traded for fluency. On Amazon Beauty and Toys, INTEGER matches or exceeds the strongest baselines in accuracy with competitive conversation quality, improving Hit@10 by 13.3% on Amazon Beauty, and significantly outperforms the generative recommender it starts from. Our analyses show that INTEGER learns behaviors that naive adaptation fails to acquire, recommending once the user’s intent is clear and staying attentive to behavioral history at the moment of recommendation. INTEGER also learns an intent-agnostic replacement over the item space, which suppresses rejected items but points to attribute-aware feedback as the next step.
[IR-9] Self-Retrospection Distillation: Turning Post-hoc Experiences into Prior Foresight
链接: https://arxiv.org/abs/2610.08077
作者: Haoxiang Zhang,Qinglin Chen,Hiroaki Hayashi,Zhuofeng Li,Siming Zhang,Jiaxin Zhang,Jixuan Chen,Fang Wu,Pan Lu,Silvio Savarese,Julian McAuley,Chien-Sheng Wu
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Information Retrieval (cs.IR); Machine Learning (cs.LG)
备注:
Abstract:Reinforcement learning with verifiable rewards (RLVR) turns agent experience into learning signals primarily through scalar outcome rewards after interaction. For group-relative objectives, however, this signal vanishes when all rollouts receive the same reward, even though their trajectories may reveal useful information about what the task requires and how the agent fails. We ask a complementary question: can hindsight teach an agent what it could have anticipated before acting? We introduce prospective learning, which uses post-hoc experience to supervise foresight predictions from the pre-interaction view, and instantiate it with Self-Retrospection Distillation (SRD). Intuitively, a completed trajectory reveals knowledge that would have been useful and pitfalls that should be avoided; SRD distills this privileged hindsight into trajectory-blind foresight of the same policy. Foresight serves only as a training target and need not be explicitly generated at inference time. Across 10 tool-integrated reasoning and long-horizon agentic tasks, SRD complements RLVR and self-distillation baselines with gains of up to 24.2 pp. Its advantage is especially pronounced when reward contrast is scarce: when 37 – 98% of rollout groups are reward-uniform across model scales, yet SRD can still exploit learning signal from sampled trajectories. In the 2B setting, where 98% of groups are all-failure, the RLVR training ends up at 0.0% success, while adding SRD reaches 60.6% under the same rollout budget. Our results suggest that post-hoc agent experience is useful not only for evaluating or improving behavior, but also for shaping predictive representations before available interaction.
[IR-10] From Delivery to Stateful Exploration: Rethinking the Index for Agent ic Search
链接: https://arxiv.org/abs/2610.07960
作者: Deogyong Kim,Sunghwan Kim,Sangam Lee,Wonjae Lee,Dongha Lee
类目: Information Retrieval (cs.IR)
备注: Work in Progress
Abstract:Recent advances in agentic search have given large language model (LLM) agents finer control over corpus exploration. However, search interfaces often return matching passages even when feedback about the candidate set would suffice for the next decision, coupling candidate refinement with source-text exposure. We propose IndexAct, an interface for Index-Native Corpus Interaction that separates candidate-set refinement from text inspection. Agents construct and manipulate persistent candidate sets through lexical conditions and set operations over an inverted index, receiving reusable state references and statistics such as candidate counts rather than matching passages. This feedback guides further refinement, while separately requested passages provide new clues or evidence that can inform subsequent operations on retained candidate sets. Experiments on five benchmarks spanning agentic search and multi-hop question answering show that IndexAct outperforms the evaluated baselines on each benchmark. On BrowseComp-Plus, it also achieves higher evidence coverage with a smaller average live context than terminal-based corpus interfaces, and maintains answer accuracy as the corpus expands. Further analyses suggest that informative refinement feedback and state reuse support continued evidence discovery, while shorter contexts or fewer search steps alone do not ensure better performance.
[IR-11] ShanLiangRen: A Nutrition Agent for Personalized Daily Meal Planning
链接: https://arxiv.org/abs/2610.07886
作者: Miao Xie,Xiao Zhang,Yuan Wang,Ruixin Zhu,Chunli Lv
类目: Artificial Intelligence (cs.AI); Information Retrieval (cs.IR)
备注:
Abstract:Dietary nutrition planning plays an important role in chronic disease management and maintaining a healthy body. In applications, it must simultaneously satisfy personalized constraints and reasonable multidimensional nutritional goals. These two aspects often conflict, and user constraints evolve with feedback, resulting in a substantial gap between generic guidelines and executable plans. To bridge this gap, we first propose the personalized fully quantified multiobjective dietary planning problem (MDP). To tackle MDP, we develop a nutrition agent, ShanLiangRen. The system first transforms dietary specifications, nutrient data, user attributes and natural language requirements into an individualized constrained planning instance. It then employs an exact retrieval-augmented generation method to shrink the feasible candidate set from a large scale ingredient and recipe space. Finally, it adopts a refinement guided by Pareto principles, where an LLM iteratively revises candidate plans under deterministic nutrition computation and feedback from constraint verification. The system outputs fully quantified meal plans with explicit ingredients and portion sizes, together with reports on nutrition compliance that show constraint satisfaction and nutrient interval attainment. We have released the system online as a WeChat Program, ShanLiangRen. A demo video is available at this https URL.
[IR-12] Contrastive Learning for Aspect Representation towards Explainable Recommendation
链接: https://arxiv.org/abs/2610.07761
作者: Emrul Hasan,Chen Ding
类目: Information Retrieval (cs.IR); Artificial Intelligence (cs.AI)
备注: 8 pages. Published in WI-IAT 2025. Best Student Paper Award
Abstract:In this work, we propose a novel recommendation model, CLARER (Contrastive Learning for Aspect Representation towards Explainable Recommendation) that integrates aspect features learned from textual reviews with rating information to improve the accuracy and explainability of recommendations. Our proposed framework learns user and item representations by combining rating-based features and aspect-based features from reviews. Specifically, rating-based features are learned through a multi-layer perceptron (MLP) model, while aspect-specific review representations are learned using a transformer encoder to capture the semantic information and contrastive learning to better distinguish user preferences. To provide explanations, we train a transformer decoder, using the final representations of users and items from both rating and aspect-based features as context. Experimental results in three benchmark data sets demonstrate that our model achieves superior performance compared to baseline methods in both recommendation (accuracy) and explanation generation.
[IR-13] oken-Budgeted Escalation for Financial Document QA: Cost Is Predictable Benefit Is the Bottleneck
链接: https://arxiv.org/abs/2610.07760
作者: Junru Zhu,Yixin Yang,Xiaoqing Ding,Ruoyu Qi
类目: Information Retrieval (cs.IR)
备注: 6 pages, 3 figures, 6 tables. Code and aggregate artifacts: this https URL
Abstract:Retrieval-augmented generation systems can route difficult queries to deeper context, but batch deployments must allocate a shared token budget across calls whose costs vary by query. We formulate selective escalation as finite-batch allocation for financial document question answering. Each of 150 FinanceBench questions first receives a top-1 retrieval answer. Predictors estimate the adjudication-quality gain and token cost of an optional top-5 call, and the allocator prioritizes calls by predicted gain per token. At the nominal 10% budget, gain-per-token allocation improves adjudication quality over gain-only ranking by 0.034 (95% document-bootstrap CI [0.001, 0.072]) while using 46.6% fewer total tokens than one-pass top-5 retrieval. Additional-call cost is accurately predictable (R-squared 0.93), whereas beneficial escalation remains difficult to rank (AUROC 0.60). These results show that heterogeneous cost is actionable under tight constraints, while progress across the full budget frontier depends on stronger query-specific benefit estimates.
[IR-14] Learning to Retrieve via Reinforcement Learning in Embedding Space
链接: https://arxiv.org/abs/2610.07731
作者: Qi Liu,Fengming Liang,Yiqun Chen,Erhan Zhang,Jiaxin Mao
类目: Information Retrieval (cs.IR); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
备注:
Abstract:Dense retrieval models are typically trained with contrastive objectives that learn effective representations but do not directly optimize retrieval metrics or downstream task performance. To address this problem, we introduce RELER (REinforcement LEarning for Retrieval), a reinforcement learning framework that enables existing embedding models to learn to retrieve directly in embedding space and align to task-specific rewards. We train RELER by sampling unit-length query and document embedding actions from von Mises-Fisher (vMF) distributions centered on normalized encoder outputs, scoring the resulting retrieval or downstream outcomes as rewards, and updating the encoder with REINFORCE using a leave-one-out baseline (RLOO). As exploration in the high-dimensional embedding space is prone to sampling noise, we further propose conditional-mean projection (CMP), which projects each sampled embedding onto the low-dimensional subspace spanned by its encoder output and the candidate embeddings it is compared against, reducing noise in the policy gradient while preserving its expectation. We evaluate RELER on BRIGHT, a benchmark with reasoning-intensive queries that remain challenging for existing embedding models. RELER consistently outperforms InfoNCE and LambdaLoss in average nDCG@10 when post-training BGE-M3 and Qwen3-Embedding backbones. We further evaluate downstream utility through retrieval-augmented generation (RAG), where we adapt only the query encoder while keeping the document index and generator fixed. Across seven QA datasets, jointly optimizing retrieval and answer rewards improves both average retrieval performance and answer quality in RAG.
[IR-15] DBRAG : Multi-Table Retrieval-Augmented Generation for Complex Database Queries
链接: https://arxiv.org/abs/2610.07622
作者: Prince Larbi Ampofo,Ryoji Kubo,Djellel Difallah
类目: Information Retrieval (cs.IR)
备注:
Abstract:Recent advancements in large language models have introduced new capabilities for reasoning over structured data, particularly through program-aided tools that can analyze tables. However, many existing methods address single-table scenarios or assume that the relevant tables are already provided. In practice, users often issue complex data exploration queries over entire databases, where relevant information may be distributed across multiple relations. In this work, we introduce DBRAG, a retrieval-augmented generation framework tailored for multi-table question answering. DBRAG first retrieves candidate tables using an offline table index, enriches their summaries with query-relevant rows, and uses an LLM to rerank the candidates. A program-aided reasoner then selects the required tables and executes operations over their full contents, keeping the initial prompt context compact. Experiments on the Spider, GeoQuery, and ATIS datasets used in this study demonstrate improvements in table retrieval and multi-table question answering.
[IR-16] A Study of Prior Case Retrieval Using Lexical Semantic and Rhetorical Role Information in Indian Legal Documents
链接: https://arxiv.org/abs/2610.07437
作者: Sayed Ayaan Ahmed Sha,Sangeetha Sivanesan,Anand Kumar Madasamy,Navya Binu,Aniket Mani,Rishu Kumar
类目: Information Retrieval (cs.IR)
备注:
Abstract:For retrieving prior cases in Indian legal judgments, the problem involves distinguishing relevant legal facts from mere lexical similarities because a prior case that shares a statute with the query judgment is not necessarily relevant. In this paper, we describe an empirical evaluation of three consecutive designs of retrieval systems for the IL PCR(Indian Legal Prior Case Retrieval) task. We demonstrate that the combination of the rhetorical roles (Fact, Ratio Of The Decision, Precedent, Argument, Statute) is better than either using individual roles or performing full text retrieval, that statute similarity is not discriminative, and that a legal entailment reranker with training data produced by an LLM is much worse in terms of official evaluation than its internal validation score. Two rankers, with excellent internal MRR up to 0.97, performed poorly in terms of official evaluation (MRR as low as 0.14). Thus, we propose a new design with query disjoint splitting and frozen validation fusion. The result is a four stage pipeline with Micro F1 0.2549, MRR 0.6081, and nDCG@10 0.4187.
[IR-17] Rethinking Semantic ID Construction for Generative Recommendation: SimHash with Parallel Decoding and Semantic Alignment NEURIPS2026
链接: https://arxiv.org/abs/2610.07402
作者: Yuqing Liu,Huiyuan Chen,Yibo Wang,Wooseong Yang,Philip S. Yu
类目: Information Retrieval (cs.IR)
备注: Accepted at NeurIPS 2026. Code: this https URL
Abstract:Semantic ID-based generative recommendation represents each item as a sequence of discrete tokens, enabling structured modeling of item semantics. A critical challenge is constructing semantic IDs that are both semantically expressive and computationally efficient. While recent approaches favor complex learned quantization, simple hashing-based methods such as SimHash are widely regarded as fundamentally inferior. In this work, we challenge this consensus by showing that the apparent performance gap does not stem from inherent limitations of hashing, but rather from a structural mismatch with autoregressive decoding, coupled with the inevitable information loss during rigid discretization. Based on this insight, we propose FLASH, a two-stage framework that revitalizes training-free SimHash tokenization through parallel decoding and explicit semantic alignment. Despite its simplicity, FLASH achieves state-of-the-art performance across multiple datasets without requiring any tokenizer training, while exhibiting stronger generalization in cold-start scenarios. Notably, we demonstrate that semantic alignment acts as a universally effective mechanism across diverse paradigms. Our findings suggest that, with compatible decoding and semantic grounding, simple and efficient tokenizers can achieve performance comparable to complex learned counterparts in generative recommendation. Our code is available at this https URL.
[IR-18] WildMatch: Weakly Supervised Image Matcher Adaptation for Wildlife Re-Identification ATC
链接: https://arxiv.org/abs/2610.07384
作者: Turhan Can Kargin,Piotr Kubaty,Ekaterina Rostovskaya,Izabela Wierzbowska,Bartosz Zieliński,Marcin Przewięźlikowski
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Information Retrieval (cs.IR)
备注: 15 pages, 7 figures, 3 tables. Project page: this https URL
Abstract:Individual animal re-identification from camera-trap imagery is an instance retrieval problem central to non-invasive wildlife monitoring: a query image must retrieve the correct individual from a reference set of known animals. This requires computer vision models to recognize distinctive local patterns in fur, skin, or other visual markings. Current approaches either learn global embeddings as a classification problem, requiring many labeled images per individual while largely ignoring local evidence, or apply off-the-shelf, domain-agnostic image matchers. Although such matchers are pretrained on large and diverse image collections, adapting them to wildlife imagery is challenging because available datasets are small and lack correspondence-level annotations. We study weakly supervised adaptation of a pretrained keypoint matcher using only identity labels, without keypoint-level or geometric correspondence ground truth. We mine informative image pairs with the pretrained matcher, derive weak positive and negative supervision from identity agreement, and contrastively fine-tune the matching network to strengthen correspondences for same-identity pairs and suppress them for different identities. Across open-source wildlife re-identification datasets, our approach improves accuracy over off-the-shelf matchers and a state-of-the-art local–global fusion method. Under an open-world protocol with held-out individuals, it learns a transferable correspondence prior rather than memorizing training identities. To our knowledge, this is the first study of matcher-level, identity-supervised adaptation for animal re-identification. Our method enables data-efficient specialization of image matching models to wildlife domains using identity annotations already available in typical monitoring datasets.
[IR-19] RAG Flip: Measuring Query-Level Negative Flips in Retriever Upgrades
链接: https://arxiv.org/abs/2610.07266
作者: Elyas Irankhah,Muhammad Arif
类目: Information Retrieval (cs.IR)
备注: 19 pages, 4 figures. Code available at this https URL
Abstract:Retriever upgrades are typically evaluated using aggregate metrics, which can hide regressions on queries the previous retriever already served correctly. We study these regressions as negative flips: queries for which BM25 retrieves a judged relevant passage and the replacement does not. We evaluate BGE-large, E5-large-v2, and SPLADE on three BEIR collections: Natural Questions, HotpotQA, and FiQA, across five retrieval depths. All replacements improve overall retrieval coverage. Negative flips occur in every setting and vary substantially by corpus, retriever, and depth. At k=1, 8.6-37.5% of BM25 successes are lost across the evaluated settings. Negative-flip rates are lower in the larger-depth settings, where the BM25-supported cohort is defined separately at each depth. These rates use the any-relevant support label. On HotpotQA at k=10, requiring every positive qrel passage raises the negative-flip rate to 12.3-17.3%. Simple fixed-budget combinations with BM25 reduce these regressions, and a HotpotQA reader experiment provides a limited downstream check in which some retrieval flips are accompanied by answer regressions. These results motivate evaluating retriever updates using query-level compatibility alongside aggregate retrieval quality.
[IR-20] CroissantMiner: Automated Extraction and Validation of Croissant Metadata for ML Datasets NEURIPS2026
链接: https://arxiv.org/abs/2610.07132
作者: Berke Arda,Ahmetcan Yavuz,Paul Gerry,Sebastian Lobentanzer,Nobin Sarwar,Joan Giner-Miguelez,Kongtao Chen,Luyao Zhang,Mrinmaya Sachan,Mubashara Akhtar
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Information Retrieval (cs.IR); Machine Learning (cs.LG)
备注: Accepted at NeurIPS 2026 (Track on Evaluations and Datasets). Website: this https URL
Abstract:Croissant has emerged as a standard for machine-readable dataset metadata, yet populating its fields remains labor-intensive and requires careful reading of accompanying dataset documentation. We present the first benchmark enabling end-to-end evaluation of metadata extraction aligned with a community-standard schema. The benchmark comprises 602 papers, including 102 with human-validated gold annotations and 500 with LLM-generated silver annotations, covering the full Croissant schema with both core and Responsible AI (RAI) fields. Using this benchmark, we evaluate a range of extraction systems spanning frontier models, open-weight models, and agentic architectures, under a two-tier evaluation framework that combines rule-based scoring with an LLM judge selected via human audit. We find that single-pass extraction consistently outperforms the four agentic architectures we evaluate: across backbones, these decomposed variants achieve lower accuracy than a single full-context pass. The largest gap appears on long-form RAI fields, which require synthesizing and interpreting information scattered across a paper rather than copying it from a single location, a setting where current systems remain far from reliable. We release the benchmark, evaluation code, judge audit, a live demo, and a leaderboard open to new systems.
[IR-21] Beyond Successor Accuracy: State Retention for Recursive Self-Improvement in Recommendation
链接: https://arxiv.org/abs/2610.07105
作者: Jinfeng Xu,Zheyu Chen,Ziyue Peng,Zheng Lin,Wenhao Yuan,Jian Chen,Shujie Li,Edith Ngai
类目: Information Retrieval (cs.IR); Artificial Intelligence (cs.AI)
备注:
Abstract:Recommendation recursive self-improvement (Rec-RSI) feeds recommender outputs into subsequent training. Evaluating each round solely through its latest model assumes that the successor consolidates the update, although pre- and post-update models may retain complementary ranking decisions. We term this \emphdistributed progress and quantify it using cross-generation advantage (CGA), a marginally matched contrast between cross- and within-generation model pairs. A rank-separation statistic, label-free at selection time, predicts which family to retain. Across four datasets and three sequential recommendation encoders, the preferred retention regime varies by architecture: cross-generation pairing benefits GRU4Rec and SASRec, whereas FMLP initially favors within-generation pairing and shifts toward cross-generation pairing after a second update. Rank separation selects the stronger family in 12/12 first-update and 5/6 second-update dataset-encoder settings; on held-out tests, the selected family outperforms the direct successor in 34/36 trajectories. Five transfer mechanisms do not consistently reproduce these gains in one model. These findings establish state retention as a distinct Rec-RSI problem: progress may reside in relations between generations as well as in the latest model. Code is available at \hrefthis https URLthis https URL.
[IR-22] Smart Content Ingestion for Generative AI Workloads
链接: https://arxiv.org/abs/2610.07091
作者: Abbas Raza Ali,Muhammad Ajmal Siddiqui,Moona Zahid
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Information Retrieval (cs.IR)
备注:
Abstract:The evolution of machine learning has progressively changed where intelligence resides in an AI system. In conventional machine learning the task, data representation, labels and model architecture were tightly coupled, so data preparation was narrow, schema-bound and visible. Generative AI decouples the model from any single task: one foundation model serves open-ended downstream tasks, and the generality gained on the model side is matched by heterogeneity on the data side, because enterprise knowledge is authored in the formats people use (PDF, presentations, spreadsheets, scanned documents, forms, tables, diagrams and mixed-layout files) that carry textual, visual, geometric and structural information at once. A language model or retriever cannot reason reliably over information misrepresented at this interface, so content extraction becomes a lifecycle stage in its own right whose errors no downstream retriever or re-ranker can repair. This paper presents a production-ready content-extraction system that makes this stage explicit, configurable, and measurable. The system incorporates selective OCR routing, a scarcity-first curation engine with a reference-based extraction scorer that measures character, word, and table-structure accuracy, a deterministic structure-aware parent-child chunker, and a read-only retrieval evaluator that generates grounded questions from every page and reports Hit@k, mean reciprocal rank, and latency. On a 180-document corpus the best extractor scores 97.4 of 100 (character error rate 0.13%, table similarity 0.995) and the chunker reaches hit@1 of 68.6%, hit@10 of 92.8% and MRR 0.77 over 25,050 generated questions. We distil three design principles (structure before semantics, never mutate what you measure, budget your labels) and position measured content extraction as the perception layer of enterprise agentic systems.
[IR-23] Beyond Refusal Patterns: Safe-Role Internalization for Robust and Generalizable LLM Safety Alignment
链接: https://arxiv.org/abs/2610.07023
作者: Jinghao Pang,Jitai Hao,Qiang Huang,Zhaochun Ren,Jun Yu
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Information Retrieval (cs.IR)
备注: 27 pages,7 figures, under review
Abstract:Large Language Models (LLMs) have achieved remarkable capabilities but remain vulnerable to jailbreak attacks that elicit harmful or unsafe outputs. Existing safety alignment approaches, including Supervised Fine-Tuning (SFT) and Reinforcement Learning from Human Feedback (RLHF), often require substantial attack-specific supervision and computational resources, while remaining susceptible to shallow safety alignment and over-refusal. To address these challenges, we introduce SSRFT(Supervised Safe-Role Fine-Tuning), the first framework that reformulates safety alignment as the internalization of a predefined safe role. SSRFT constructs a Safe-Role Question-Answer (SRQA) dataset from psychometric questions, limited jailbreak prompts, and a safe-role description. Role-consistent responses are synthesized, validated, and expanded into diverse scenarios, enabling models to internalize safety-oriented values and principles rather than explicit refusal patterns. Experiments across multiple Base and Instruct models show that SSRFT achieves more robust and generalizable safety alignment than standard SFT. SSRFT shows substantially greater robustness to prefilling attacks and better generalization to unseen jailbreak domains, while reducing over-refusal on benign queries and preserving the model’s general capabilities. These results establish safe-role internalization as an effective alternative to refusal-centric safety alignment. Warning: This paper contains examples of harmful and toxic language.
[IR-24] ree Navigation Without LLM Summaries: A Matched-Cost Study of Hierarchical Retrieval for Long-Document QA
链接: https://arxiv.org/abs/2610.06902
作者: Priyank Jayraj,Poonam Goyal,Navneet Goyal
类目: Computation and Language (cs.CL); Information Retrieval (cs.IR)
备注:
Abstract:Retrieval-augmented generation grounds language models in external context, but for long documents flat top- k retrieval can cluster on a single region and miss complementary evidence. RAPTOR-style summary trees address this by recursively clustering chunks and using a language model to summarize each cluster at indexing time, then ranking summary nodes alongside raw chunks at query time. We show the main benefit of summary trees in long-document QA can come from navigation rather than the generated summary content. We introduce NavTree, a leaves-only retriever that builds a deterministic balanced segment tree over chunks (zero language-model calls at indexing) and uses the tree purely as a navigation scaffold: a hybrid lexical-and-dense frontier walk, anchored on top retrieved leaves, descends from the root and emits only leaf chunks to the reader. On a matched-cost evaluation against flat retrievers and an extractive re-implementation of RAPTOR, NavTree is the strongest matched-cost hierarchical retriever in our evaluated grid and ties the strongest flat baseline. On long-document multi-hop QA, it is the only hierarchical method that significantly beats BM25 on a class-vs-class basis, corroborated by a reader-free retrieval-recall check. A matched-reader replication of the published abstractive RAPTOR variant, given strong cluster summaries, still loses to NavTree at every multi-chunk budget, at zero indexing cost. The ranking carries across stronger and open-weight readers, a stronger encoder, and a full factorial that isolates leaves-only emission as the structural lever.
[IR-25] Diff-SQL: SQL Efficiency Optimization via Patch Generation and Constraint Alignment
链接: https://arxiv.org/abs/2610.06857
作者: Shipei Lin,Duomin Zhang,Xiaolong Li,Bohan Hu,Bowen Qin,Jinyang Li,Chenhao Ma
类目: Databases (cs.DB); Information Retrieval (cs.IR)
备注:
Abstract:SQL efficiency optimization aims to transform slow queries into semantically equivalent but faster alternatives. However, directly optimizing SQL with large language models in an end-to-end fashion often induces Objective Misalignment which creates a fundamental tension between optimization and correctness, making direct full SQL rewriting unreliable for execution-facing database applications. To address this problem, we propose Diff-SQL, a two-stage framework that decouples efficiency-oriented optimization from constraint-aware alignment. The first stage identifies optimization opportunities and proposes targeted edits in the form of a unified diff patch, while the second stage is trained with on-policy reinforcement learning to revise outputs under executability and semantic-equivalence constraints. To train and evaluate Diff-SQL, we construct an automated pipeline that mines optimization knowledge from StackOverflow and builds Slow-Fast SQL pairs through cascaded filtering. We further introduce Effi-SQL, a benchmark containing 1,100 human-verified Slow-Fast pairs across five SQL dialects. Experiments show that Objective Misalignment is widespread across existing LLM-based SQL optimization methods, where direct full SQL optimization causes an average 22.7% execution accuracy degradation across frontier models such as Claude-Opus-4.6, with the worst model dropping by 43.0%. Diff-SQL alleviates this trade-off. As an inference-only strategy, it improves R-VES by 10.0% on average while reducing execution accuracy degradation by 6.11% on average across three strong base models. With execution-grounded training, Diff-SQL further enables a 7B model to improve R-VES from 33.42% to 46.83%, demonstrating that the proposed two-stage optimization-and-alignment paradigm can deliver both stronger efficiency and better correctness in local, small model deployment settings.
人机交互
[HC-0] reVISit-XR: Bringing Extended Reality into Embeddable Trackable and Replayable Visualization Studies
链接: https://arxiv.org/abs/2610.08700
作者: Shano Liang,Max Chen,Lane Harrison
类目: Human-Computer Interaction (cs.HC)
备注: 5 pages, 2 figures
Abstract:Extended reality (XR) is increasingly a setting for empirical visualization research, where study-relevant state is distributed across headset and controller pose, scene configuration, selections, spatial layouts, and AR anchors. Existing XR tools support parts of the workflow, such as scene authoring, interaction, or session analysis, yet seldom treat an XR stimulus as a reusable component of a complete study lifecycle. We present reVISit-XR, an extension of reVISit that makes customizable WebXR stimuli embeddable, trackable, and replayable within empirical visualization studies. reVISit-XR sequences XR scenes alongside standard study components, collects reactive task responses, captures scene-authored semantic state together with generic XR traces, and rehydrates participant sessions for later desktop and headset analysis. Organized as a reusable stimulus build package and a study integration package, it lets XR stimuli act as first-class study components. We demonstrate its scope and feasibility through seven integrated, reusable examples and a deployed study, and discuss its current capabilities and future extensions.
[HC-1] Juicy Interactive Visualization: Evaluating How Excessive Feedback Design Shapes Visualization Engagement
链接: https://arxiv.org/abs/2610.08681
作者: Shano Liang,Max Chen,Lane Harrison
类目: Human-Computer Interaction (cs.HC)
备注: 11 pages, 3 figures
Abstract:Visualization research has long examined embellishment and animation, while the design of rich interaction-contingent feedback remains under-articulated as a systematic design space. We introduce JuicyVIS, a theory-informed operationalization of juicy feedback for interactive visualization that translates ideas from game studies into a visualization-centered framework. JuicyVIS introduces three dimensions, interaction type, feedback timing, and feedback intensity, which we explore through 26 controlled prototypes, grounded in established visualization interaction categories and prior work on juicy feedback in games. We evaluate the juicy design space through three exploratory mixed-methods online studies. Results show that juicy feedback was consistently associated with higher engagement and aesthetic experience, while the clearest gains came from post-interaction feedback, not, for example, by simply increasing intensity. We discuss how juicy feedback can be thought of as a design material in visualizations, shaping how visualization interactions feel, how action consequences are perceived, and how users remain engaged with data representations.
[HC-2] A Space-Agnostic Visual Game Analytics Tool with Adaptive Spatial Reconstruction for Mixed Reality Game Development
链接: https://arxiv.org/abs/2610.08619
作者: Nahian Rifaat,Fedir Skliar,Parisa Sargolzaei,Loutfouz Zaman
类目: Human-Computer Interaction (cs.HC)
备注:
Abstract:Recent years have seen an increasing need for MR tools for game development and visual game analytics. However, the number of visual game analytics tools for MR games remains scarce due to the unique complexities in integrating real components from both the physical and the digital worlds in MR games. Existing visual analytics tools for such games are not suitable for space-agnostic generalization. To tackle this problem, we introduce a space-agnostic visual game analytics tool that incorporates adaptive spatial reconstruction and real-time object detection built for the Microsoft HoloLens 2 MR headset. We believe that the outcomes of our work will increase insights into player environment and pave a new direction for visual game analytics for MR game development.
[HC-3] Systemization of Knowledge (SoK): Human-Centered AI Safety for Youth
链接: https://arxiv.org/abs/2610.08554
作者: Pratyasha Saha,Yaman Yu,Yang Wang
类目: Human-Computer Interaction (cs.HC); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:
Abstract:While HCI increasingly examines AI-safety for youth, the literature lacks a comprehensive view of what risks have been identified, how they are addressed, and whether proposed protections work in-practice. We systematically reviewed 100 empirical HCI studies involving children and youth interacting with or exposed to AI across schools, homes, care settings, and public services. Using the YAIR taxonomy for risks and the MIT Mitigation Taxonomy for countermeasures, we map which risks have been identified, whether each risk is addressed by countermeasure(s), and whether each countermeasure for that risk is implemented and even evaluated. The risk-countermeasure mapping shows that most risks are matched only with proposed/ideated countermeasures; few countermeasures have been implemented, and fewer still evaluated; and existing evaluations often measure technical performance rather than protection from harm. We identify where coverage is absent, where safeguards remain untested, and propose concrete directions for HCI research to strengthen youth AI-safety.
[HC-4] he Now and Then: Integrating Current and Historical Data in Small Multiple Time Series Visualization
链接: https://arxiv.org/abs/2610.08473
作者: Sydney K. Purdue,Enrico Bertini,Melanie Tory
类目: Human-Computer Interaction (cs.HC)
备注: 16 pages, 9 figures, to be published in IEEE TVCG
Abstract:Small multiple time series visualizations are often used for real-time data monitoring tasks in high-impact domains such as healthcare and manufacturing. Effective design is critical because users rely on these visualizations to monitor data from many entities, such as patients or machines, often while distracted. Users may need to rapidly appraise current values for each entity, monitoring for those that go outside an acceptable range, while also watching temporal trends. However, no design guidelines currently exist for visually emphasizing current values in historical time series represented by small multiples. Via an iterative design process informed by a review of related literature and theory on visual channels and emphasis, we present a design space for glanceable time series small multiple displays. We evaluate this space through two online empirical studies, testing against non-threshold and threshold rapid appraisal tasks. Our results provide insights into merging current value and historical data visualizations for rapid appraisal tasks in time series monitoring. For non-threshold tasks, we found that size encodings on the current value, spatially integrated into the line chart, may provide a good compromise, with 28% response time improvement for tasks involving finding large current values and minimal interference with trend lookup tasks. More generally, integrated designs outperformed separated designs (in which the current value representation is spatially separated from the historical trend line). For threshold tasks, color threshold encodings significantly outperformed shaded band encodings. All supplemental materials are available at this https URL. Comments: 16 pages, 9 figures, to be published in IEEE TVCG Subjects: Human-Computer Interaction (cs.HC) Cite as: arXiv:2610.08473 [cs.HC] (or arXiv:2610.08473v1 [cs.HC] for this version) https://doi.org/10.48550/arXiv.2610.08473 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[HC-5] Living Dashboards: Automatically Self-Updating Visualization Dashboards
链接: https://arxiv.org/abs/2610.08393
作者: Mingyu An,Heyon Jeon,Sungbok Shin,Jinwook Seo,Eduard Gröller,Niklas Elmqvist,Vaishali Dhanoa
类目: Human-Computer Interaction (cs.HC)
备注:
Abstract:Visualization dashboards are widely used interactive tools, but a disconnect exists between the dynamic data they display and their static structure. End-users cannot modify the dashboard to answer new questions. We introduce Living Dashboards, whose views are born, wither, revive, and die in response to how they are used. Rather than requiring manual reconfiguration, a living dashboard observes interaction and natural-language queries to autonomously wither neglected views and revive those used again. More consequential decisions, such as adding or retiring views, are deferred to the user. We formalize the concept as a four-dimensional design space and implement it in Living Dashboard, a web-based prototype. We evaluate it in an exploratory between-subjects study (N = 12) against an AI-supported baseline on analytical tasks. Living Dashboard participants answered more tasks correctly, reported lower workload, and rated the system higher on usability, though the two conditions differed in more than adaptive behavior alone.
[HC-6] Hugging Suit: Pneumatically-Actuated System Design for Remote Haptic Experiences
链接: https://arxiv.org/abs/2610.08305
作者: Russian(Ruo-Xuan)Wu,Luke Hespanhol,Marius Hoggenmueller,Hannes Waldschütz,Eva Hornecker
类目: Human-Computer Interaction (cs.HC); Emerging Technologies (cs.ET)
备注: Accepted Work in Progress at the IEEE World Haptics Conference 2025, July 8-11, 2025, Suwon, South Korea. 2 pages, 1 figure
Abstract:The COVID-19 pandemic emphasised the importance of remote emotional communication and highlighted a gap in existing technologies that lack haptic channels. While video calls maintain visual and auditory presence, they cannot convey the emotional depth of physical contact, especially hugging, a universally recognised form of intimacy and support. To explore this challenge, we developed the Hugging Suit, a pneumatically-actuated wearable system that enables users to simulate and receive remote hugs in real time. Unlike prior haptic systems that focus solely on tactile sensation, our approach integrates both technical and experiential considerations. The system integrates a programmable Air Actuator Matrix (AAM), a portable high-pressure pneumatic unit, and a fabric-based pressure sensor layer, enabling precise, wearable haptic feedback while remaining lightweight and mobile. Guided by a Research through Design (RtD) methodology, we iteratively refined the prototype to improve tactile resolution, overall usability, and user comfort, while introducing modular components to support flexible and scalable haptic configurations. In parallel, we explored how specific contextual factors, such as lighting, privacy, and the visibility of the remote partner, might shape the emotional and perceptual experience of receiving a remote hug. Preliminary feedback suggests that low-lit, private spaces, especially when users could also see their remote partner via video, improved emotional engagement and comfort during mediated haptic experiences. Our low-cost, maker-space-friendly approach contributes to affective haptics applications in long-distance relationships, emotional regulation, and remote therapeutic support.
[HC-7] Building A Civic Tool for Community-Police Engagement to Adapt Neighborhood Policing
链接: https://arxiv.org/abs/2610.08212
作者: Ravinithesh Reddy Annapureddy,Staņislavs Šeiko,Natalie Higham-James,William Droz,Alessandro Fornaroli,Sarah Vollmer,Britta Elena Hecking,Daniel Gatica-Perez
类目: Human-Computer Interaction (cs.HC); Computers and Society (cs.CY)
备注:
Abstract:Data-driven policing often prioritizes incident records over residents’ lived experiences. In the Baltic city of Riga, with a history of distrust and limited community-police engagement, this can further alienate the public. To bridge this gap, we propose a Research through Design (RtD) inquiry into the development of Par drošu Rīgu, a civic tool for community-data-integrated policing. With municipal police, NGOs, and city staff, we ask how RtD enables stakeholder negotiation and which interaction qualities support trust and the use of combined community and incident data. The co-design process included workshops that surfaced divergent notions of safety; material probes designed as boundary objects to negotiate among stakeholders; and a pilot deployment showing how combining quantitative and qualitative data reshapes engagement and trust. Mixed-methods evaluation suggests increased officer-citizen interaction, but frictions in sustaining stakeholder collaboration. We contribute (i) an empirical RtD inquiry with public institutions, (ii) an artifact combining physical and dashboard interactions, and (iii) reflections on interaction design as a boundary-spanning practice for trust and infrastructuring.
[HC-8] Frontstage Mediation Work: Invisible Work Bridging Gaps Between AI Decisions and User Expectations
链接: https://arxiv.org/abs/2610.08067
作者: Yongjae Sohn,Daehyun Kwak,Jiyeon Amy Seo,Hyungjun Cho,Seongah Youn,Youn-kyung Lim
类目: Human-Computer Interaction (cs.HC); Computers and Society (cs.CY)
备注: Poster at CSCW Companion '26. 8 pages
Abstract:Automated service systems increasingly generate algorithmic operational decisions that shape how services are delivered. However, these decisions reach end-users only through frontline workers who carry them out in real-world settings. During this process, automated decisions can diverge from user expectations, surfacing as friction at the service encounter. We propose Frontstage Mediation Work as a preliminary analytic lens for examining the often invisible labor through which frontline workers anticipate and manage such misalignments between algorithmic decisions and user expectations. Drawing on a qualitative case study of an On-Demand Ride-Pooling service, we identify four recurring practices through which drivers sustain the service encounter when frictions arise. Such labor remains absorbed into routine operations, leaving no trace in performance metrics, system logs, or formal job descriptions. This paper contributes to worker-centered HCI scholarship by illustrating how automated services shift onto frontline workers the responsibility of managing the interactional consequences of system-level decisions.
[HC-9] he Amplifier Effect: Human-Factor Risks of AI-Suggested Correlation and Auto-Propagation in Multi-Framework GRC Self-Assessment
链接: https://arxiv.org/abs/2610.07866
作者: Nikolaos Kekatos,Michael Ioannou,Marina Korgiala-Karyda,Alexios Lekidis,Tom Nianios
类目: Cryptography and Security (cs.CR); Computers and Society (cs.CY); Human-Computer Interaction (cs.HC)
备注: 7 pages, 2 figures, 3 tables. Accepted at the 2026 IEEE International Conference on Cyber Security and Resilience (IEEE CSR 2026)
Abstract:Multi-framework Governance, Risk and Compliance (GRC) platforms increasingly automate the link between an organisation’s self-assessment answer and the compliance obligations that answer is said to satisfy. Cross-framework control mapping, AI-suggested question correlation, and automatic propagation of answers and evidence across correlated questions all serve the legitimate efficiency goal of reducing duplicate work for small and medium-sized enterprises under the EU Cyber Resilience Act, NIS2 and GDPR. The same mechanisms, however, amplify the consequences of any human-factor bias in a single answer: one optimistically-graded control, one rubber-stamped attestation, or one AI-drafted answer can be silently replicated as evidence of compliance with many obligations across multiple frameworks. We call this the amplifier effect: a platform-design property (coarse-grained attestation and un-gated propagation) rather than a failing of individual users. Using two EU-funded SME-facing GRC platforms, CYBERFORT and CYBER-BRIDGE, as examples, we (i) describe the amplification mechanism in concrete data-model terms, (ii) propose a six-dimension scoring framework for evaluating any GRC tool’s exposure to the effect, (iii) instantiate the framework on a thirteen-tool comparison covering enterprise IRM, mid-market platforms, compliance-automation tools, and the two EU SME projects, and (iv) outline a measurement protocol that a consortium with access to production self-assessment data can run. The thirteen-tool comparison is a structured design assessment, not an empirical measurement of user behaviour. The EU SME platforms score lowest on the amplifier dimensions because their burden-reduction design deliberately trades sign-off granularity for throughput; we report this as a design trade-off, not a verdict on the platforms. Our contribution is the framing and the measurement protocol.
[HC-10] Novice Reliance Calibration in AI-Assisted Decision Making: The Role of Explanations and Self-Assessment
链接: https://arxiv.org/abs/2610.07800
作者: Eun Jeong Kang,Peter(Xianpi)Duan,Swati Mishra
类目: Human-Computer Interaction (cs.HC); Artificial Intelligence (cs.AI)
备注: under review
Abstract:Artificial Intelligence (AI) tools are widely used to support decision making in tasks and domains where no immediate performance feedback is available. In these settings, users cannot learn to adjust their reliance behavior over time through trial and error. However, little is known about how novice users calibrate reliance on AI when external feedback is unavailable, or whether AI explanations can support calibration in its absence. We introduce reliance calibration as an organizing construct for studying how novice users dynamically adjust reliance behavior, and examine how AI explanations and meta-cognitive self-assessment shape it. Through a between-subjects study with 110 participants completing a clinical entity extraction task with AI assistance and limited performance feedback, we observe that novice users exhibit systematic drift toward over-reliance in the presence of explanations, while higher self-reported task understanding is associated with more selective reliance behavior. These results extend reliance calibration research into human-AI collaboration contexts without real-time performance signals and present actionable guidelines on designing AI tools that must support appropriate reliance in these settings.
[HC-11] A Pedagogically Demonstrative Model Visualizing the Pathway from Online Interactions to Personalized Recommendation
链接: https://arxiv.org/abs/2610.07744
作者: Sushmita Khan,Connor Pennington,Bart P Knijnenburg
类目: Human-Computer Interaction (cs.HC)
备注:
Abstract:Personal digital activity increasingly shapes online experiences, yet few users have been educated regarding the processes transforming raw interactions into personalized suggestions. We developed an education artifact that illustratively simulates how AI leverages users’ digital activities to shape online recommendations (e.g., ads). Our artifact processes users’ digital activity using a locally-hosted LLM to generate user profiles of their inferred interests and personalized recommendations. A three-layered Sankey diagram maps data sources through inferred interests to personalized recommendations. Interactive filters enable users to explore how different combinations of data sources influence personalized outcomes. This paper describes the artifact and its educational value, and reports findings of a pilot think-aloud study with six young adults. We find that the artifact effectively taught participants the conceptual relationship between digital activities and personalized recommendations. While this did lead participants to develop privacy awareness, they anticipated minimal behavior change due to the perceived unavoidability of platform participation.
[HC-12] Who Is Talking to the Agent ? LLM s in Multi-User 3D Virtual Environments
链接: https://arxiv.org/abs/2610.07732
作者: Mohammad Al-Ratrout,Shayla Sharmin,Roghayeh Leila Barmaki
类目: Human-Computer Interaction (cs.HC)
备注:
Abstract:When several people share a 3D virtual room with an LLM agent, the agent must decide not only what to say, but whether an utterance was addressed to it and, if accessible, what profile information about the others present it may use. To study both problems, we construct LookAway, a controlled corpus of 40 sessions involving 80 distinct personas and an LLM agent (1,200 turns), including ambiguous-addressee turns in which speaker orientation agrees or conflicts with the intended addressee. Across three open-weight large language models and five conditions varying which profiles the agent sees and whether it is told where each person faces (18,000 decisions), adding speaker orientation increased addressee accuracy from 56% to 99.5% when orientation was congruent, but when the speaker faced someone other than the addressee, two of the models went by where the speaker faced on more than 85% of those turns. Warning one model that orientation could be misleading reduced this only modestly. A browser-based 3D demonstrator shows the effect live: the same sentence gets an answer when the speaker faces the agent and silence when they face the other person. Providing both personas’ profiles improved responses about the person being asked about, but also increased the use of profile attributes not revealed in the shared conversation, reaching 45.3% of answers for one model. More context thus improves multi-user interaction but also leads to oversharing, so shared LLM agents need mechanisms for weighing spatial cues and controlling when user-specific information enters a response.
[HC-13] Evaluating human-AI workflows for field research in viticulture
链接: https://arxiv.org/abs/2610.07669
作者: Niko Carvajal Janke,Daoyuan Jin,Shivranjani Baruah,Nicholas Gunner,Jacob Maus,Yu Jiang,Kaitlin M. Gold
类目: Human-Computer Interaction (cs.HC); Artificial Intelligence (cs.AI); Computers and Society (cs.CY)
备注: 35 pages, including supplementary materials. Supporting files S1-S4: this https URL
Abstract:We assessed the value of two live human-AI interactions in a precision disease control project in California vineyards. The project tested whether 2021-2024 commercial scouting records and remote-sensing measurements across 140 hectares could support 2025 red-leaf symptom forecasting for prioritized scouting and virus testing. In Workflow 1, Aleks v1, a multi-agent research system, developed forecasting models with iterative human refinement. We applied Aleks’s 2024 vine-scale model to updated 2025 predictors and evaluated red-leaf forecasts against independent 2025 scouting. In retrospective simulations surveying 45% of all vine positions, adding model-informed row prioritization to adaptive scouting increased the encountered proportion of newly recorded red-leaf observations from 85.8% to 94.1%. Within-block scouting comparisons suggested the model mainly improved scouting allocation among blocks. Despite unreliable internal 2024 performance estimates from synthetic oversampling before train/test splitting, Aleks developed an informative vine-scale model in 145 minutes, increasing throughput and answering our research questions. In Workflow 2, we assessed whether higher model-score vines had more frequent virus detection, and whether Aleks could infer this sampling goal from a general prompt with data and literature. Aleks’s plan prioritized balanced vineyard and model score coverage, while our plan prioritized field efficiency and high-model-score oversampling. Aleks’s and our plans yielded 41/50 (82%) and 97/100 (97%) sampled vines. Aleks’s plan omitted instructions for replacing missing vines, limiting implementation and operational value. Five of 137 sampled vines tested positive for grapevine red blotch virus (model score ROC AUC 0.735). These findings support assessing AI interactions by how well they advance field research objectives under live, project-specific constraints.
[HC-14] SENSE: State-aware Emotion Navigation Storytelling Engine
链接: https://arxiv.org/abs/2610.07666
作者: Yi Xia,Pablo Carrasco Velo,Mudit Paliwal,Ibrahim Khan,Yifan Geng,Mustafa Can Gursesli,Juho Hamari,Ruck Thawonmas
类目: Human-Computer Interaction (cs.HC); Artificial Intelligence (cs.AI); Multimedia (cs.MM)
备注:
Abstract:This paper presents SENSE, a state-aware framework for generating playable branching visual novels with multi-track emotional navigation. Integrating a state-based narrative architecture called MIND, a structure analyzer, and a path-aware context management module, SENSE produces narratives that are both structurally coherent and emotionally rich. From minimal high-level inputs, it generates multiple intersecting routes while preserving character consistency and narrative causality. Evaluations using LLM judges, affective metrics, and visual assessments indicate SENSE outperforms baselines in narrative diversity and robust asset integration, while preliminary human trials show directional improvements in emotional fidelity alongside comparable enjoyment.
[HC-15] When the Commons Appropriates a Large Language Model: How WikiVault Reshaped Korean Wikipedia
链接: https://arxiv.org/abs/2610.07660
作者: Inhwa Song,Sohyeon Hwang,Ted Yoo,Manoel Horta Ribeiro
类目: Human-Computer Interaction (cs.HC)
备注:
Abstract:Large Language Models (LLMs) have disrupted the balance between content production and quality assurance that sustains knowledge commons, leading many to prohibit or restrict their use. But what happens when a community instead appropriates an LLM-powered tool for its own needs? We investigate this question through WikiVault, an LLM-powered editing tool developed within the Korean Wikipedia community and used primarily for translation. Combining ten interviews, platform-scale analyses, and matched quasi-experimental comparisons, we examine how WikiVault reshaped knowledge production on Korean Wikipedia. We find the tool 1) drastically amplified the production capacity of a small group of experienced editors, producing longer and more widely viewed articles; 2) shifted work toward reviewing articles and importing content; 3) imported not only content but also editorial judgments from English Wikipedia. Our findings show how LLM adoption can rebalance the interdependent work that sustains knowledge commons, while illustrating how communities can learn from emerging technologies through use and adapt their governance accordingly.
[HC-16] Beyond screen time: Explaining cross-national differences in digital literacy through socioeconomic and psychological mechanisms
链接: https://arxiv.org/abs/2610.07619
作者: Hyejeong Lee,Daeyoung Ham,Suyoun Kim,Tiffany Emanuel
类目: Human-Computer Interaction (cs.HC)
备注:
Abstract:This study provides a structural explanation for cross-national variation in the relationship between screen time and digital outcomes. While prior research and large-scale assessments such as ICILS have documented inconsistent associations between screen time and digital competence, the mechanisms underlying these differences remain unclear. Using ICILS 2023 data, this study employs multigroup structural equation modeling to examine the relationships among socioeconomic status, screen time regulation, ICT self-efficacy, and digital literacy outcomes. Results reveal substantial cross-country differences in the effects of screen time regulation. In contrast, ICT self-efficacy emerges as a consistent and robust predictor across all countries. Moreover, screen time regulation influences outcomes indirectly through self-efficacy in some contexts but not others. These findings challenge the use of screen time as a standalone indicator of digital engagement and highlight the importance of psychological mechanisms. By integrating socioeconomic, behavioral, and psychological factors, this study advances a more nuanced understanding of digital competence and moves beyond quantity-based approaches to digital learning.
[HC-17] Emoception: Selective Affective Layer Fine-Tuning of Video Vision Transformers for Player Arousal Change Recognition From Gameplay Footage
链接: https://arxiv.org/abs/2610.07603
作者: Yi Xia,Ibrahim Khan,Mury Fajar Dewantoro,Wenwen Ouyang,Ruck Thawonmas
类目: Human-Computer Interaction (cs.HC); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:This article proposes Selective Affective Layer Fine-Tuning (SALFT), an efficient adaptation framework for Video Vision Transformers in player arousal recognition from gameplay. To bypass computationally expensive full fine-tuning, SALFT introduces a selection criterion based on the L2-norm change in layer parameters after brief adaptation, directly measuring representational shifts and providing a more stable basis than gradient-based alternatives. Evaluated via five-fold cross-validation on the Arousal Video Game AnnotatIoN dataset, SALFT achieves performance comparable to full fine-tuning across all games without statistically significant degradation ( p0.05 ), while updating only \approx 8% of parameters (over 92% reduction). Notably, in one game, SALFT consistently outperforms both full fine-tuning and the best baseline across all metrics and folds, reaching the theoretical minimum p-value (p=0.0625, exact two-sided Wilcoxon signed-rank test). In addition, we introduce an interpretability method to trace attention patterns, enhancing model transparency. These results establish SALFT as an effective and efficient approach for affective game computing.
[HC-18] PAIR: Perceptual Affective Inference and Regulation in a Real-Time Multimodal Conversational Agent
链接: https://arxiv.org/abs/2610.07523
作者: Kexin Quan,Zijian Ding,Jiaye Yong,Qinshi Zhang,Dong Wang,Jessie Chin
类目: Human-Computer Interaction (cs.HC)
备注:
Abstract:Sustained emotional support requires generative agents to connect momentary emotion inference and regulation with continuity across encounters. We present PAIR (Perceptual Affective Inference and Regulation), a real-time multimodal agent that reconstructs how an event is appraised into an emotional state. Appraisal scaffolds produce a valence-arousal-dominance estimate and select regulation guidance, delivered through conversation with coordinated speech, color, and avatar cues. Rolling memory carries context across sessions, and the scaffold re-runs after guidance. In a 14-day deployment with 19 participants, 1,093 sessions paired initial and post-guidance estimates with unanchored self-reports. Initial valence reached MAE 1.20 on the 9-point SAM scale (r=.68), dominance reached MAE 1.30, and arousal showed weak agreement even after coarsening. Self-reported emotional change varied with initial state, with the largest valence increases in sessions that began at negative valence. Perceived understanding was associated with greater valence increase and showed little correspondence with numerical prediction error. Over two weeks, helpfulness increased while input shortened; interviews traced personalization and companionship to relevant recall, context updates, and familiar dialogue. These findings connect inference accuracy to conversational and temporal patterns of support through per-event, first-person evaluation.
[HC-19] he Relationship Between Blood Pressure and Self-Reported Stress During Sound-Based VR Relaxation Videos
链接: https://arxiv.org/abs/2610.07507
作者: Md Alamin Hossain,M. Rasel Mahmud
类目: Human-Computer Interaction (cs.HC)
备注:
Abstract:Virtual reality (VR) relaxation environments are increasingly proposed as accessible tools for stress management, yet the physiological correlates of self-reported stress during VR exposure remain unclear, particularly for blood pressure (BP). We conducted a within-subjects study (N = 18) in which participants viewed relaxing 360-degree videos, grouped into four sound-design categories, such as videos with rhythmic (RHYT), aperiodic (APRD), continuous (CONT), and low-ambient (LOWE) sound through an HTC Vive Focus Vision headset. Systolic (SYS) and diastolic (DIA) BP were recorded after each category using an Omron 3 Series monitor, and self-reported stress was collected via a 0-100 Visual Analogue Scale (VAS) with fixed intervals. While overall BP did not differ significantly across sound categories, self-reported stress did (Friedman chi-square(4) = 12.07, p = .017), with rhythmic sound eliciting significantly lower stress than aperiodic and continuous sound. Critically, DIA, but not SYS, correlated significantly with self-reported stress across conditions (r = .31, p = .003). This suggests that DIA may be a more sensitive physiological indicator of subjective stress than SYS in short VR exposures. We discuss implications for designing and evaluating VR mental health interventions and the value of low-cost physiological sensing alongside self-report.
[HC-20] Jarvis: A Proactive Speech Agent for Multi-Party Conversations
链接: https://arxiv.org/abs/2610.07506
作者: Seunghyun Oh,Hirotaka Hiraki,Shuyue Stella Li,Yulia Tsvetkov,Shyamnath Gollakota
类目: Human-Computer Interaction (cs.HC); Artificial Intelligence (cs.AI)
备注: 45 pages, 6 figures, 20 tables
Abstract:Speech agents are reactive and dyadic: they speak when spoken to, and to one person at a time. We ask what it takes for a speech agent to instead take part in a conversation among several people and speak up only when it can help. We introduce Jarvis, a real-time proactive speech agent that audibly participates in multi-party human conversations. Grounded in a document shared beforehand, Jarvis follows the discussion and intervenes when the group misses or misstates a fact and does not correct itself within a few turns. We make three contributions: a problem setting based on epistemic breakdowns that makes proactive intervention measurable, realized as CHI-180-proactive, a synthetic multi-party dataset seeded with known gaps, errors, and self-corrections; a proactive backbone that harnesses a small, open-weight model with deterministic checks and grounds every claim in a source sentence; and interaction techniques for taking the floor in live speech and showing the cited evidence on screen. On CHI-180-proactive, Jarvis is correct on most events it addresses and stays silent 97% of the time when the group resolves an issue itself. A live study with 23 participants confirms these trends with real-time interventions.
[HC-21] From Written Response to Dialogue with AI: How Activity Format Interaction Modality and Language Impact Student Learning and Engagement
链接: https://arxiv.org/abs/2610.07483
作者: Deepak Varuvel Dennison,Connie Zhang,Rene Kizilcec,Aditya Vashistha
类目: Human-Computer Interaction (cs.HC); Computers and Society (cs.CY)
备注:
Abstract:The widespread availability of LLMs is challenging written learning activities, as students can increasingly generate responses without necessarily engaging with the learning content. Conversational AI creates an opportunity to redesign these activities as dialogue, while multilingual capabilities may make such dialogue more accessible to students learning through a non-native language. We conducted a field study with 305 native Kannada-speaking undergraduate students at English-medium institutions in Karnataka, India. We compared written response activities with text- and voice-based dialogic activities with AI, each conducted in English-only or bilingual Kannada-English settings. Students who completed dialogic activities spent more time on the activities, contributed more, and reported greater interest and self-efficacy than those completing written responses, although fewer students completed the dialogic activities overall. Knowledge increased across all conditions, with no reliable differences in gains between activity formats or languages. Language shaped participation differently across modalities: bilingual interaction was particularly beneficial in voice dialogue, where it reduced articulation difficulties, increased turns demonstrating understanding, and reduced conversation abandonment. However, students also valued English because of its connection to their academic and professional aspirations. These findings show that designing dialogic learning with AI requires more than choosing between writing and dialogue, voice and text, or English and students’ native languages. We highlight opportunities to give learners greater control over modality, information, and pace, and to use native languages as translanguaging support rather than as a replacement for English.
[HC-22] MRPilot: Supervising and Intervening LLM -Based Multi-Robot Teams through Mixed Reality
链接: https://arxiv.org/abs/2610.07477
作者: Xiaoran Yang,Xun Qian,Yang Zhan,Nathan Tran,Ziyi Liu,Qiao Jin
类目: Human-Computer Interaction (cs.HC); Robotics (cs.RO)
备注: 15 pages, 8 figures
Abstract:Large language models (LLMs) let users direct heterogeneous multi-robot systems (MRS) through natural language, but make task interpretation, robot assignment, and coordination difficult to inspect and change. Based on a formative study with 12 non-expert users, we developed MRPilot, a mixed reality system organized around four stages of supervision and intervention. MRPilot represents robot-team plans and execution states as structured commitments shared across synchronized situated and overview views. Across four stages, it helps users resolve ambiguous references (Forming), review plans before execution (Reviewing), monitor distributed execution (Following), and make robot-level or team-level changes when problems arise (Repairing). In a within-subjects study with 20 participants in a virtual reality-simulated home, MRPilot reduced workload, increased situational awareness, transparency, trust, and perceived control compared with a conventional LLM-based conversational interface using the same LLM planner and robot capabilities. We provide design implications for multi-scale intervention, adaptive supervision, and calibrated reliance in LLM-based MRS.
[HC-23] Artifact removal improves electrodermal waveforms but not downstream classification in a virtual-reality balance task
链接: https://arxiv.org/abs/2610.07438
作者: Haochen Chai,Qixu Zhu,Siyao Li,Fangfang Jiang
类目: Human-Computer Interaction (cs.HC); Machine Learning (cs.LG); Signal Processing (eess.SP)
备注: 10 pages, 9 figures, 3 tables. Code and frozen data: this https URL
Abstract:Artifact removal routinely precedes the classification of electrodermal activity (EDA), on the assumption that a cleaner signal supports a better decision. We tested this assumption in a virtual-reality (VR) balance-disturbance task. A residual gating network was trained on a benchmark with expert-corrected EDA, frozen, and applied to VR recordings, where raw and gated signals were classified by five published time-series methods under identical leave-one-participant-out evaluation. On the benchmark the gate detected artifacts well (median record AUROC 0.94) and reduced error inside artifact regions by 17.8%. In the VR task it did not improve classification. Changes in balanced accuracy ranged from -1.35 to +0.93 percentage points, no classifier improved and two lost accuracy, and all five were equivalent to raw input within +/- 3.32 points. The benefit was lost between waveform and decision. The correction that lowered waveform error also reduced skin conductance response detection in all 43 benchmark records. Processing left 92.8% of predictions unchanged, and the predictions it did change were corrected and corrupted at similar rates. The VR recordings also carried little contamination (an estimated 4.6% of samples), and even perfect localization of deliberately injected artifacts recovered only 3.3 points in the most sensitive classifier. A pooled association between artifact level and accuracy (11.3 points) disappeared within participants (0.1 points), showing how differences between people can make cleaning look useful. Preprocessing should be judged by the decision it supports, against an unprocessed arm.
[HC-24] Knit-Structure Effects on Electromechanical Metrics and Their Correlation with Joint-Angle Estimation Error in Knitted Strain Sensors
链接: https://arxiv.org/abs/2610.07416
作者: Annika Eloranta,Zhuchenyang Liu,Iiro Naulapaa,Iida Arvola,Yao Zhang,Anna-Mari Leppisaari,Lulu Xu,Yu Xiao
类目: Human-Computer Interaction (cs.HC)
备注: 10 pages
Abstract:Knitted resistive strain sensors show strong promise for joint motion sensing in sports and rehabilitation, but the linkage between sensor design and in situ performance remains unclear. We investigate how knit structure and machine settings (e.g., stitch size) shape electromechanical properties and which metrics predict sensing performance during bending. Sensors spanning seven common knit structures at two stitch sizes were fabricated, characterized under uniaxial cyclic tension, and evaluated on a joint emulating bending rig. Joint-angle estimation was assessed with machine learning models, and correlations with electromechanical metrics were analyzed. Experimental results show that, among six common metrics, gauge factor and baseline resistance are largely set by knit structure, while working range, linear range, hysteresis, and cyclic stability vary only modestly across designs. Gauge factor correlates negatively and baseline resistance positively with joint-angle estimation error, mainly in lower-sensitivity designs, whereas the other metrics have weak or no predictive value. These results support using uniaxial tensile tests to screen out weak designs, while underscoring the need for joint-relevant evaluation and application-specific metrics to identify top performers.
[HC-25] Redistributing Harm: Document Transition and the Limits of Trans Inclusion in Indias Identity Systems
链接: https://arxiv.org/abs/2610.07343
作者: Megh Marathe,K Ranade,L. Ramakrishnan,Koyel Ghosh
类目: Human-Computer Interaction (cs.HC); Computers and Society (cs.CY)
备注:
Abstract:Identity documents (IDs) are some of the most consequential sites through which transgender people encounter state infrastructures. This article examines document transition, the work of aligning names, gender markers, and related information across official records in India based on focus groups with sixteen trans participants, most of whom were transmasculine and gender diverse, in mid-2025. Participants described encountering an absence of clear protocols, changing demands for proof, and objectionable conduct from officials, arising often from limited understandings of transness in systems and policies. As a result, participants faced harms related to livelihood, housing, voting, travel, redress from violence and discrimination, and education. We situate these findings within information studies and trans studies scholarship on classification and state recognition. Participants with mismatching IDs faced a form of torque or administrative violence that we call trans tax' including disproportionate tax deductions, reluctant employers, and resultant insecure jobs. Further, gender-concordant IDs could redistribute harm: participants who became administratively legible as male lost jobs, benefits, and rights despite remaining vulnerable to the discrimination the programs sought to address. The disruption of people's lives by updated gender-concordant data calls for a re-examination of haunting’ at the intersection of transness, classification, and data practices. The article concludes by calling for sensitization efforts directed at street-, system-, and policy-level decision-makers about trans identities and state provisions together with established and widely disseminated protocols for document transition in the short term; as well as a broad and coordinated upheaval of systems and policies to undo cisgender-heteronormative assumptions.
[HC-26] Quantum-Like Spatial Decision Dynamics: A Falsifiable Model of Cue-Order Effects in Immersive Navigation with Implications for Human-Quantum Computer Interaction
链接: https://arxiv.org/abs/2610.07336
作者: Aryabrata Basu
类目: Human-Computer Interaction (cs.HC)
备注: 24 pages, 4 figures, and 6 tables. Includes simulation code, synthetic replicate-level results, summary data, and run metadata as ancillary files
Abstract:The sequence in which a person encounters spatial evidence can change a later choice, yet an order effect alone does not identify a quantum-like cognitive structure. We introduce Quantum-Like Spatial Decision Dynamics (QSDD), a falsifiable state-space account of embodied decision making in which cue exposures are completely positive trace-preserving maps, intermediate judgments are quantum instruments, and route commitment is an operationally defined measurement. We distinguish this formal claim from any assertion that cognition is microscopically quantum. We then specify a minimal binary model with a fixed decision basis, noncommuting cue rotations, a fixed symmetry-breaking initial azimuth, dephasing, and lapse; its five fitted parameters must generalize across environments rather than being refit to individual conditions. A reproducible simulation study evaluates recovery and model discrimination under both QSDD and seven-parameter classical logistic ground truths. Across 40 replicates per design cell, QSDD was preferred on held-out environments in 85% of QSDD-generated datasets at 240 independent observations per environment-order cell and 95% at 480, but in 0% of classically generated datasets at every tested sample size. The simulation also exposes weak identification of the lapse parameter and is explicitly a design analysis, not human evidence. We provide a preregistrable virtual-reality experiment, telemetry schema, classical comparison set, and failure criteria. Finally, we show how the same process-measurement logic can be transferred to human inspection and debugging of quantum circuits. QSDD is therefore offered as a constrained model to be defeated or supported by data, and as a methodologically continuous route from immersive interaction research to human-quantum computer interaction.
[HC-27] Mapping E-textiles Design Pain Points and Generative AI Opportunities: Insights from Workshops in Shanghai and Winchester
链接: https://arxiv.org/abs/2610.07296
作者: Zhuchenyang Liu,Nianchong Qu,Yao Zhang,Marie O’Mahony,Qi Wang,Yu Xiao
类目: Human-Computer Interaction (cs.HC)
备注: 10 pages, 2 figures, 2 tables
Abstract:E-textile design involves complex decisions across materials, sensor and actuator structures, fabrication, garment integration, and data processing. It typically requires iterative prototyping and testing, which are time- and labour-intensive, while few practitioners possess cross-disciplinary expertise across all relevant domains. To identify current bottlenecks and explore how Generative AI (GenAI) might support the design process, we conducted two half-day co-design workshops, one in Shanghai and one in Winchester, with practitioners from materials science, electronics, garment design, human-computer interaction, and manufacturing. Twenty practitioners participated in the Shanghai workshop; ten of them had prior experience in e-textiles and form the contributing sample analysed here. A further ten practitioners participated in the Winchester workshop. Participants mapped their own design pipelines, annotated bottlenecks, and proposed where GenAI could provide support. Rather than presenting a ranked list of opportunities, we report a process map that indexes each proposed GenAI role to the pipeline stage at which practitioners located it, together with the conditions on which they stated its usefulness would depend. Across both sites, practitioners consistently identified domain-specific operational barriers, including data scarcity, the disconnect between prototyping and manufacturing, and trade-offs in material-hardware integration. They also emphasized that the primary barrier to GenAI-driven e-textile design is not general model capability, but the lack of standardized, machine-readable representations of e-textile designs. Based on these findings, we identify four classes of domain-tailored AI tools that could support future e-textile design processes.
[HC-28] RoboCap: A New Platform for Egocentric Robot Learning
链接: https://arxiv.org/abs/2610.07217
作者: Grounded Superintelligence,BitRobot
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Human-Computer Interaction (cs.HC)
备注:
Abstract:Despite its promise for scaling robot learning, egocentric manipulation data is still scarce today. Collection at scale requires vertically integrating ergonomic hardware with centimeter-precise 3D algorithms, at a precision that has not been publicly demonstrated. To address this gap, we introduce RoboCap, a 250,g six-camera dual-IMU hat designed for in-the-wild egocentric data capture, and the Grounded API, a suite of device-agnostic 3D algorithms tuned for RoboCap. In this report, we demonstrate how hardware, calibration, and 3D algorithms interact to achieve state-of-the-art performance on the public benchmarks: our SLAM across diverse settings and rigs, our depth estimation on egocentric settings, and our hand tracking when adapted to third-party devices.
[HC-29] Responsible Institutional Analytics: Interpreting Bias with AI Support
链接: https://arxiv.org/abs/2610.07205
作者: Francielle Marques,Ariel Ortiz-Beltrán,Ishari Amarasinghe,Davinia Hernández-Leo
类目: Human-Computer Interaction (cs.HC); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Social and Information Networks (cs.SI)
备注: Accepted for publication in the Journal of Universal Computer Science ( this http URL )
Abstract:Institutional Analytics (IA) dashboards inform decision-making in higher education, yet data limitations, constraints in analytical techniques, and missing contextual information often affect their interpretation. To support more responsible interpretation of IA, we introduce FACTRIA, a framework that organizes potential biasing factors across four areas: the analytics pipeline, institutional context, course-level characteristics, and demographics. We used the FACTRIA framework as input to a generative-AI chatbot designed to prompt users to reflect on these factors while analyzing IA. A qualitative study with stakeholders, drawing on four authentic IA cases, and a transition network analysis showed that the chatbot prompted participants to recognize how overlooked factors influenced their initial interpretation. Findings indicated that combining a structured framework with AI-based guidance can enhance context-aware, responsible interpretation of institutional data.
[HC-30] owards semantic reconstruction of individual words from fnirs using clip loss
链接: https://arxiv.org/abs/2610.07120
作者: Santiago Posso-Murillo,Nathan Palladino,Ben Pyykkonen,Dan Y. Han,Luis G. Sanchez-Giraldo,Jihye Bae
类目: Human-Computer Interaction (cs.HC)
备注:
Abstract:Semantic reconstruction maps neural activity to a word-embedding space, recovering the meaning of a perceived word instead of selecting it from a fixed vocabulary. Functional near-infrared spectroscopy (fNIRS) carries semantic information suitable for this mapping. However, most fNIRS decoders are trained with a squared-error objective that fits each word independently and ignores the geometry of the embedding space. To address this limitation, we evaluate a contrastive loss based on the contrastive-language-image-pretraining (CLIP) loss, as an alternative to mean-squared-error (MSE) for reconstructing perceived words from fNIRS. We compare the two objectives by training a bidirectional long short-term memory (Bi-LSTM) decoder to map fNIRS signals to word embeddings. We use GloVe-50 and T5 word embeddings as targets, across three fNIRS datasets recorded under a shared paradigm pairing each word image with its spoken name. Performance is measured with a pairwise matching score and open-vocabulary top- k retrieval. The Bi-LSTM trained with CLIP is the most consistent decoder across experiments. T5 produces higher matching scores, whereas every significant retrieval result uses GloVe-50. These results support the use of contrastive objectives as a promising direction for fNIRS semantic decoding and motivate validation on larger datasets.
[HC-31] mation: Making Narrative Gaps Visible in Childrens Storytelling with Just-in-Time Animation
链接: https://arxiv.org/abs/2610.07039
作者: Vincent Cavez,Marielle Zheng,Momin Siddiqui,Abbie L Olszewski,Srirangaraj Setlur,Maneesh Agrawala,Hari Subramonyam
类目: Human-Computer Interaction (cs.HC)
备注:
Abstract:Illustrated scenes are used to prompt children’s stories, yet children with language difficulties often omit the details, actions, and relationships that make a story coherent. We introduce Tellimation, which animates scene elements a child has omitted or misdescribed, drawing attention to a narrative gap without saying what to tell. Its design space covers eight kinds of gaps (who is in the scene, how many, what they are like, what they are doing, where, when, how they relate, and what lies beyond the picture, such as speech and thoughts), instantiated through 20 parameterized animations grounded in classical animation principles. A real-time pipeline detects discrepancies between utterance and scene, then selects and parameterizes an animation from the scene’s narrative potential and the child’s history. Adults interpreted most animations without instruction (N=120); children (N=12) resolved 46% of the gaps the system identified with animations, against 13% without.
[HC-32] Educating future engineers about LLM s: A scalable workshop
链接: https://arxiv.org/abs/2610.07027
作者: R. Zhang,J. C. F. de Winter,T. Dicke,D. Dodou,Y. B. Eisma
类目: Human-Computer Interaction (cs.HC); Computers and Society (cs.CY); Robotics (cs.RO)
备注:
Abstract:As large language models (LLMs) are increasingly integrated into engineering workflows, students require hands-on experience to learn how to collaborate with them critically. This paper presents a scalable gamified workshop designed for engineering Master’s students to practice human-AI collaboration in navigation planning. Using a mobile web interface across 10 workshop sessions, a total of 226 students wrote prompts for a non-reasoning and a reasoning LLM to solve grid-based navigation tasks of increasing complexity. The system returned robot-executable plans, trajectory visualizations, and automated scoring, culminating in a live demonstration on a Boston Dynamics Spot robot. In a post-workshop questionnaire, 81.5% reported substantial learning and 91.0% reported high engagement. Analysis of the submitted prompts revealed that students changed their strategies from step-by-step instructions for the non-reasoning LLM toward providing higher-level guidance for complex problem-solving tasks. We conclude that such interactive simulation-to-reality environments are viable for teaching the verification and collaboration skills necessary for responsible LLM use in engineering. Code is available at: this https URL
[HC-33] EVFormer: An Egocentric Vision-EMG Bidirectional Attention Model for Bimanual Hand Pose Estimation
链接: https://arxiv.org/abs/2610.06970
作者: JiaCheng Ge,SiYu Zhang,ShengJie Li,XinTong Yang
类目: Machine Learning (cs.LG); Human-Computer Interaction (cs.HC)
备注:
Abstract:Egocentric bimanual hand pose estimation is important for virtual interaction, wearable control, and rehabilitation, but visual observations are often degraded by self-occlusion, hand-hand contact, and object manipulation. We propose EVFormer, a multimodal framework that combines the current RGB frame with the preceding 200 ms of bilateral wrist surface electromyography (sEMG) to estimate 44 finger and wrist joint angles. EVFormer separately encodes visual spatial features and sEMG temporal features, enables cross-modal information exchange through sequential bidirectional cross-attention, and integrates the two modalities using feature-wise gated fusion. We evaluate EVFormer in a single-participant feasibility study using one synchronized public EgoEMG recording with chronologically separated training, validation, and test splits. On 296 test samples, EVFormer achieves a mean absolute error of 11.482 degrees, compared with 13.228-13.610 degrees for vision-only, sEMG-only, late-fusion, and training-mean baselines. This corresponds to relative error reductions of 13.20% compared with the vision-only model and 14.23% compared with late fusion. EVFormer also achieves the lowest error in four of the five evaluated gesture classes. These results provide preliminary evidence that feature-level interaction between egocentric vision and sEMG can improve bimanual hand pose estimation. Further evaluation across participants, recording sessions, sensor placements, and real-world interaction conditions is required to establish the generalizability of the approach.
[HC-34] rajectools Demo: Towards No-Code Solutions for Movement Data Analytics MDM2024
链接: https://arxiv.org/abs/2610.06858
作者: Anita Graser,Melitta Dragschnig
类目: Databases (cs.DB); Human-Computer Interaction (cs.HC); Machine Learning (cs.LG)
备注: Accepted at MDM2024
Abstract:This demo paper presents the conceptual foundations and the first steps towards implementation of a novel no-code solution for movement data analytics based on the open-source Python library MovingPandas and the open-source geographic information system QGIS. The resulting Trajectools plugin is available open-source at this https URL.
计算机视觉
[CV-0] World Models Last Exam in Physics
链接: https://arxiv.org/abs/2610.08791
作者: Mingju Gao,Qingle Liu,Yuzhao Peng,Xinjie Lin,Ziming Qin,Zheng Jiang,Wenyi Li,Calvin Xiao,Youjie Zheng,Kaisen Yang,Qinhuai Na
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Video world models can produce visually convincing yet physically inconsistent sequences, raising concerns about their reliability for prediction and planning in embodied AI systems. Existing evaluations often rely on model-based judgments or reference videos, while direct physical tests largely focus on mechanics. We introduce World Models’ Last Exam in Physics, a measurement-based benchmark for evaluating physical consistency in video world models. The benchmark comprises 40 controlled tasks spanning mechanics, optics, fluids, thermal and phase-change phenomena, electromagnetism, and surface tension. Each task pairs an initial image and a generation prompt with predefined physical criteria, enabling interpretable tests of observable physical relationships without requiring reference videos. Its evaluator combines task-observability screening with task-specific quantitative physical measurements. Experiments on eight video generation models across 1,280 videos reveal persistent physical inconsistencies and substantial variation across tasks, with the best model achieving an overall score of 57.76 out of 100. Evaluation on synthetic videos with known physical relationships provides evidence for the validity of the measurement module under controlled conditions. The evaluator also achieves higher agreement with human judgments than a direct vision-language model baseline in both within-task rankings and pairwise comparisons. By combining coverage across physical domains with scores grounded in measurable evidence and explicit measurement limitations, the benchmark provides an interpretable basis for diagnosing physical inconsistencies and tracking progress toward physically consistent video world models.
[CV-1] Building Rome from a Single Image
链接: https://arxiv.org/abs/2610.08790
作者: Jiraphon Yenphraphai,Fang Li,Tianshuo Xu,Depu Meng,Quentin Herau,Yihan Hu,Raymond A. Yeh,Wei Zhan
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Project page: this https URL
Abstract:Single-image scene generation aims to produce a complete 3D scene mesh from a single image, including surfaces the camera did not observe. While pretrained 3D object generators encode a strong shape prior, they are mainly designed for isolated objects in a fixed canonical volume and focus mostly on indoor scenes, since diverse 3D data for outdoor scenes are quite limited. In this work, we present a method that redesigns such an object-centric generator, e.g., Trellis 2, to work on both indoor and outdoor scenes while retaining its prior. We accomplish this by (a) partitioning the scene into adaptive chunks that scale relative to the distance to the camera; nearby chunks have a smaller size to keep the finer detail, while distant structures, e.g., buildings, are covered by large chunks; (b) making the generator capture explicit 2D-3D correspondence by lifting image features and making the model aware of the free space, observed surface, and unobserved region; © synthesizing around 4,000 outdoor scenes to broaden the training data, as existing scene datasets are largely indoor. Experiments on Tanks and Temples, ScanNet++, and in-the-wild images show that our method outperforms all baselines in geometric accuracy and perceptual quality across both indoor and outdoor scenes.
[CV-2] 4D-HOF: Hand-Object Flow Matching for Feed-Forward 4D Interaction Reconstruction
链接: https://arxiv.org/abs/2610.08782
作者: Shiqi Li,Sean Cho,Yijie Li,Fengzhi Guo,Bowen Wen,Cheng Zhang
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Graphics (cs.GR)
备注: Project page: this https URL
Abstract:Existing methods for 4D hand-object reconstruction often rely on costly per-sequence optimization, while generative approaches typically synthesize interactions from random noise, which can lead to unstable interaction prediction. We introduce 4D-HOF, a feed-forward framework that reconstructs 4D hand-object interactions from coarse but informative estimates produced by vision foundation models. Concretely, we learn a conditional flow matching model that transports foundation-model-derived hand-object states toward an interaction manifold, allowing the model to correct errors in translation, rotation, and alignment in a feed-forward manner. A key advantage of our generative formulation is that it naturally enables test-time guidance within the transport process. Rather than applying a separate post-hoc optimization after reconstruction, we directly steer the evolving generative states using physical interaction constraints and observed 2D evidence, allowing the reconstruction to be refined as part of the generative process itself. By training the generative model on diverse datasets, 4D-HOF generalizes robustly to challenging in-the-wild scenarios. Experiments on out-of-domain benchmarks show that 4D-HOF achieves state-of-the-art performance, producing more stable and accurate 4D hand-object reconstructions.
[CV-3] DepthWorld: 3D World Model for Robot Manipulation WWW
链接: https://arxiv.org/abs/2610.08780
作者: Jai Bardhan,Josef Sivic,Vladimir Petrik
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted at the Conference on Robot Learning (CoRL) 2026. Project page: this https URL . 32 pages including supplementary material, 15 figures, 7 tables
Abstract:World models offer a data-driven alternative to traditional simulators for robotics, with applications spanning policy evaluation, improvement, and planning. All of these uses depend on faithful 3D geometry, yet current video-based world models are trained on RGB alone and produce rollouts that look correct frame-by-frame but do not compose into a consistent 3D world. Closing this gap requires progress on two fronts: large-scale 3D supervision for manipulation, and an architecture that can absorb it without disturbing strong pretrained video priors. We introduce a calibration pipeline that combines learned stereo depth with a joint factor graph, pooling all episodes collected from the same physical robot to recover its shared kinematic parameters alongside per-scene extrinsics. Applied to the DROID dataset, this yields DROID-3D, a calibrated 3D dataset providing dense metric depth and recalibrated multi-view extrinsics (achieving 0.7 px reprojection error on 90% of episodes for external cameras). We then train DepthWorld, a Stable Video Diffusion-based world model that jointly predicts multi-view RGB and depth via spatial latent tiling, leaving the pretrained Variational Autoencoder (VAE) unchanged. Depth supervision improves RGB prediction itself by +1.48 dB PSNR over an identical RGB-only baseline at equal training budget, while simultaneously yielding accurate metric depth for downstream geometric reasoning.
[CV-4] ALIVE: Interaction-Aligned Object Insertion for First-Frame-Guided Video Editing
链接: https://arxiv.org/abs/2610.08779
作者: Zhenghong Zhou,Zhe Lin,Jiebo Luo,Yuqian Zhou
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Project page: this https URL
Abstract:Current video editors can insert objects but often struggle to make them participate in interactions such as being picked up or manipulated. We introduce ALIVE, a framework that makes inserted objects “alive” through coherent interactions with the source video’s contents, using an edited first frame and an instruction naming only the added object. We curate 35,800 editing pairs combining 3D-rendered, model-generated, and real-world videos with general editing pairs from ROSE. Each pair differs in the target object’s presence while preserving the surrounding action, teaching editors coordinated object behavior and source preservation. We further train a vision-language model (VLM) to predict interaction guidance from the same inputs. We introduce the ALIVE-interaction benchmark to assess interaction fidelity, source preservation, and visual coherence using a unified VLM-based protocol, and evaluate on the general video object insertion benchmark. Without VLM guidance, ALIVE improves Overall over the strongest evaluated baseline by 43.9% and 4.4% on the two benchmarks, respectively. VLM-predicted guidance further improves the ALIVE-interaction score by 0.95 points without additional user inputs.
[CV-5] CtrlCache: Accelerating Interactive Video World Models with Control-Aware Caching
链接: https://arxiv.org/abs/2610.08777
作者: Shangye Song,Dong Gong,Hong Jia,Yun Sing Koh,Xinyu Zhang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 18 pages. Project page: this https URL
Abstract:Interactive video world models need to generate each video chunk efficiently while responding faithfully to user controls. Many systems use chunk-wise autoregressive generation with few-step denoising, but each chunk still requires several costly denoising iterations. Training-free caching can reduce this cost, yet existing policies make reuse decisions primarily from model-internal denoising dynamics and do not explicitly account for control transitions. Actually, interactive generation explicitly exposes a signal they do not use: the controls for a chunk arrive before it is denoised, so a schedule derived from them costs no forward pass. To this end, we analyze adjacent chunks under different control regimes and find that structural similarity drops around action changes, while low-frequency structure remains more persistent than high-frequency detail. Motivated by these observations, we propose CtrlCache, a training-free control-aware caching framework that adapts computation to the current control sequence. Specifically, the action-aware scheduling and refresh policy detects action changes across and within chunks, and labels each chunk as initial, transition, turning, or steady state. At one selected interior denoising step, initial and transition chunks retain full computation, while turning and steady chunks reuse the transformer residual from the most recent fully computed step in the same chunk. To exploit the persistence of low-frequency structure during steady interaction, we further introduce a frequency-mixed history prior guidance that incorporates complementary information from the preceding clean latent without an additional DiT forward pass. Evaluated on Matrix-Game 2.0 and LingBot-World v1/v2, CtrlCache achieves 1.21x to 1.41x DiT-backbone speedups without model retraining while improving WBench Overall scores over original inference across all three models.
[CV-6] Backend-Agnostic Sparse Attention for Fast High-Resolution Visual Generation
链接: https://arxiv.org/abs/2610.08772
作者: Liao Ma,Jiayi Song,Yunfeng Wu,Songhua Liu,Peilin Zhao
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Diffusion Transformers (DiTs) have achieved strong performance in image and video generation, but the quadratic complexity of full attention makes high-resolution generation computationally expensive. Window attention offers an efficient alternative, yet existing methods face a practical trade-off: partitioned window attention typically achieves computational efficiency consistent with its theoretical complexity. However, isolated windows block cross-window interaction, often introducing visible grid-like artifacts in the generated results. Fine-grained sliding-window attention effectively restores interactions across neighboring windows and improves visual quality. However, its irregular computation patterns create a substantial gap between theoretical and practical speedups and require specialized kernels tailored to each hardware backend. To tackle these challenges, we propose BASA, a backend-agnostic sparse attention, which brings the best of both worlds: visual quality and practical acceleration. Specifically, BASA replaces visual self-attention with shifted local-window attention. By introducing a structured window-shifting scheme across DiT blocks, we allow tokens divided by window boundaries in one layer to communicate in the following layers, thereby achieving global information exchange and eliminating window-induced visual artifacts. Notably, our design introduces no additional irregular operators or customized kernels, making it readily deployable on existing attention backends and closing the gap between theoretical sparsity and practical acceleration. Experiments demonstrate that BASA achieves measured speedups exceeding 90% of the theoretical estimates on FLUX and delivers a 4.52 \times attention speedup on Wan while maintaining competitive generation quality.
[CV-7] Data Leakage in Patch-Based Hyperspectral Image Classification: Quantifying the Impact of Spatial Overlap
链接: https://arxiv.org/abs/2610.08770
作者: Mohammed Q. Alkhatib
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: paper accepted for presentation at IEEE-WHISPERS
Abstract:Patch-based learning improves hyperspectral image (HSI) classification by exploiting local spectral-spatial information, but random train-test sampling from the same image can cause spatial patch overlap, leading to data leakage and optimistic performance estimates. This paper investigates same-class train-test spatial overlap in patch-based HSI classification using two measures: overlap percentage (OP), which quantifies the global amount of overlapped testing patch pixels, and average overlap ratio (AOR), which measures the local severity among affected testing patches. Experiments on the Pavia University dataset compare random and non-random spatial sampling using SVM, MLP, 2D-CNN, 3D-CNN, ViT, and MorpMamba. The results show that deep patch-based models achieve high accuracy under random sampling, with 3D-CNN reaching 96.17% Overall Accuracy (OA), but drop substantially under non-random spatial sampling, where 3D-CNN decreases to 55.20% and ViT and 2D-CNN drop by 40.71 and 38.81 percentage points (PP), respectively. Patch-size analysis further shows that increasing the patch size from 5x5 to 19x19 raises the random-sampling overlap percentage from 23.28% to 77.02%. These findings demonstrate that random patch-based evaluation can substantially inflate classification performance, especially for models that strongly exploit spatial context. The code associated with this paper is available at: this https URL.
[CV-8] WorldSonus: Bringing Sound to Worlds
链接: https://arxiv.org/abs/2610.08760
作者: Pengjun Fang,Jingyi Fa,Kam Man Wu,Jiaming Wang,Haoyuan Huang,Yaguang Wu,Xiangjun Huang,Ziyang Ma,Weijia Chen,Hongyu Liu,Zeyue Tian,Qifeng Chen
类目: ound (cs.SD); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Audio and Speech Processing (eess.AS)
备注: 25 pages, 4 figures, 16 tables. Project page: this https URL
Abstract:Recent advances in world models have enabled increasingly realistic visual synthesis. However, these generated environments remain largely silent. Bringing sound to world models poses three core challenges: real-time generation to keep pace with interactive video streams, interactive control to respond to mid-stream sound instructions, and spatially aligned stereo to reflect scene geometry and camera motion. To address these demands, we introduce WorldSonus, an interactive video-to-audio framework designed for real-time spatial sound synthesis in world models. For real-time generation, WorldSonus employs a streaming causal autoregressive diffusion architecture that synthesizes audio chunks at a low real-time factor (RTF) of 0.41. For interactive control, we incorporate an audio-centric captioning pipeline with chunk-indexed prompt scheduling, enabling dynamic manipulation of sound events during generation. For spatial alignment, we leverage high-quality stereo supervision curated from diverse stereo and ambisonic data. Extensive experiments demonstrate that while tailored for world models, WorldSonus generalizes effectively to open-domain video-to-audio benchmarks, matching or outperforming state-of-the-art bidirectional models in both acoustic quality and spatial alignment. Project page: this https URL
[CV-9] Post-Training Semantic Lifting for 3D Gaussian Splatting: Separating Detector Lifting and Representation Error
链接: https://arxiv.org/abs/2610.08756
作者: Iván Verdugo Guerra,Ezequiel López Rubio,Jorge García González
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 18 pages, 11 figures, 9 tables. Code: this https URL
Abstract:The same Gaussian of a 3D Gaussian Splatting model is seen from many views, and these views do not always agree on the class it belongs to. The Gaussian may be occluded in some of them, and the confidence of the detector is not the same from one view to another. The ground truth, on the other hand, is given as an annotated mesh, because two training runs do not produce the same Gaussians. In this work, we propose a post-training lifting method that works with one target class at a time and combines the information coming from all the views. Target and non-target evidence are accumulated simultaneously, weighted by the visibility of each Gaussian in each view. After that, the Gaussians are filtered with two thresholds: a main threshold \beta selects the high-confidence seeds, and a lower one \gamma\beta adds the connected components around them. For the evaluation, the labels are transferred from the Gaussians to the mesh vertices that are both visible and annotated. With this design, we can separate three sources of error: the 2D detector, the lifting and the transfer between representations. The thresholds and the transfer operator are chosen on seven Replica validation scenes, and the method is evaluated on ten held-out ScanNet++ scenes with the same values for every scene and class. The mean mIoU on the validation scenes was 0.93 with masks from the dataset annotations and 0.65 with YOLO masks, and on the ScanNet++ test scenes it was 0.80 and 0.54. Compared with thresholding the evidence per view, as a previous version of the method did, the fraction improves the test mIoU by 0.24 and makes it possible to use a single threshold for all the classes and scenes of both datasets. Finally, the error analysis shows that most of the remaining error comes from the detector.
[CV-10] Co-Evolving Paths and Flows via Path-Flow Alignment
链接: https://arxiv.org/abs/2610.08717
作者: Zeyu Michael Li,William Xingxu Chen,Xiang Cheng
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注:
Abstract:We study path-flow alignment as a unified training objective for flow matching. Instead of fixing the interpolation path and learning only the velocity field, we jointly train an endpoint-preserving path network and a flow network using the same alignment loss: the flow learns to match the path velocity, and the path learns to align its velocity to the current flow. Although every fixed learned path defines a valid flow-matching objective, the alignment loss alone is not a reliable criterion for path learning. We identify path overfitting, a failure mode in which the alignment loss decreases while sample quality worsens. We find that this failure is associated with low-entropy bottlenecks in the induced probability path, where the learned path routes samples through overly concentrated intermediate marginals. Motivated by this diagnosis, we introduce a stochastic path regularizer that hides part of the source information from the path network while preserving exact endpoints. The resulting regularization gives an explicit entropy floor for the stochastic training-path marginals and empirically suppresses the bottleneck in the learned sampler, making joint path-flow training effective. On ImageNet-256x256 with SiT backbones, our method consistently improves FID across model scales, extends to model-guidance training, and leaves the inference-time architecture and sampler unchanged. Code is available at this https URL
[CV-11] SpaTime: Streaming Vision-Language Models for Spatio-temporal Reasoning
链接: https://arxiv.org/abs/2610.08713
作者: Hairong Yin,Huangying Zhan,Shin-Fang Chng,Yi Xu,Raymond A. Yeh
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Embodied agents must reason about 3D space while the video is still arriving, answering questions as soon as they have observed enough of the scene. VLMs that incorporate 3D geometric priors achieve strong spatial reasoning, but they operate offline, i.e., the full video must be available before they produce an answer. Streaming VLMs process frames causally and decide for themselves when to respond, yet they lack explicit 3D representations. We present SpaTime, a streaming VLM that fuses causal geometry tokens into the language model at every frame, using only the frames observed so far. To supervise when the model answers, we propose a response-time loss that maps per-frame response probabilities to a differentiable expected response time and penalizes the distance from the ground-truth frame. For evaluation, we construct StreamVSTI-Bench and StreamVSI-Bench, streaming adaptations of VSTI-Bench and VSI-Bench. On StreamVSTI-Bench, SpaTime reaches 49.2% overall accuracy and reduces the mean response-time error by 66% relative to the strongest streaming baseline.
[CV-12] Local Content-Style Control for Diffusion-based Image Stylization SIGGRAPH
链接: https://arxiv.org/abs/2610.08704
作者: Amir Semmo
类目: Graphics (cs.GR); Computer Vision and Pattern Recognition (cs.CV)
备注: SIGGRAPH Asia 2026 Technical Communications. 4 pages, 4 figures, 1 table. Supplemental material included as an ancillary file
Abstract:Image stylization with latent-diffusion models entangles two independently refined axes: what a region depicts and how it is depicted. Such pipelines expose only global controls, yet professional retouching demands deliberate, region-specific control. We lift two conditioning weights already present in a ControlNet + IP-Adapter stylization pipeline from global scalars to per-location spatial maps, yielding local, per-axis control of content and style in a single generative pass. Because the two weights act on disjoint pathways, adjusting them independently spans a 2x2 retouching vocabulary, from free regeneration to identity preservation. We validate that edits stay confined to the retouched region and that each weight predominantly steers its own axis. Our approach requires no retraining and drops unchanged into any such pipeline.
[CV-13] RenderBench: Benchmarking Render-to-Real Video Transfer with Reconstructed Digital Twins
链接: https://arxiv.org/abs/2610.08684
作者: Dicong Qiu,Zhiyuan Xu,Yaosheng Liu,Feng Han,Bo Ye
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 10 pages, 4 figures, 2 tables
Abstract:Modern video models can generate realistic videos from real appearance references and proxy renders that specify scene structure, viewpoint changes, and motion. Evaluating this render-to-real capability requires a real target video depicting the same scene evolution, paired with an editable, geometrically registered 3D replica. Such data has traditionally required substantial manual modeling, calibration, and animation effort. We introduce RenderBench, a benchmark of 12 reconstructed real-world scenes spanning large-scale indoor environments and egocentric viewpoints, with both static and dynamic settings. Our construction pipeline combines visual geometry, neural reconstruction, and assisted 3D authoring. Each scene is decomposed into static objects and dynamic actors, registered to the capture cameras, and accepted only after multi-view geometric and temporal validation. Each evaluation unit contains appearance reference images, a held-out real target video, an editable digital twin, a matched proxy render, and renderer-native scene annotations. We evaluate transfer models against paired real target videos, retain PAI-Bench-C-compatible structural projections, and use scene annotations to localize failures by object, visibility, articulation, and motion. The first release retains 12 of 14 registered samples (85.7%), comprising 1,496 paired real-proxy frames. All released scenes pass file-integrity and environment-edit audits, while proxy diagnostics yield a depth si-RMSE of 0.2170 and instance mIoU of 0.3673. RenderBench provides paired real observations and editable scene state for assessing both appearance fidelity and preservation of geometry and dynamics.
[CV-14] EC-RAG : Event Chain Retrieval-Augmented Generation for Long Video Understanding
链接: https://arxiv.org/abs/2610.08674
作者: Yuhao Qin,Junbo Wang,Yuke Li,Yining Zhu
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 12 pages, 7 figures, 7 tables, including supplementary material
Abstract:Current large video-language models (LVLMs) still face challenges when dealing with long videos, mainly because frames are often processed independently, making it difficult to capture temporal dependencies across events. Although retrieval-augmented approaches have been introduced to provide additional context, most of them operate at the frame or snippet level, which limits their ability to model how events evolve over time and relate to each other. In this paper, we propose Event Chain Retrieval-Augmented Generation (EC-RAG), a training-free framework that organizes video content into an explicit event chain before question answering. Instead of retrieving isolated frames or text segments, EC-RAG first partitions the video into semantically coherent segments, represents each segment using multi-modal signals, and then links them into a structured chain that preserves temporal order and captures inter-event relationships. Given a query, the system identifies relevant events within this chain and gathers supporting evidence from the associated modalities. Our approach offers several practical advantages: (i) event-level abstraction that better reflects how video content is naturally structured, enabling more reliable localization compared to frame-level retrieval; (ii) structured multi-modal fusion that aggregates speech, text, and visual cues at the event level, allowing complementary information to be more effectively utilized during reasoning; and (iii) plug-and-play compatibility with existing LVLM backbones, requiring no additional training or reliance on proprietary models. Experiments on Video-MME, MLVU, and LongVideoBench show that this event-centric design consistently outperforms frame-level retrieval baselines, highlighting the importance of modeling temporal structure for long-video understanding.
[CV-15] PDB: Point-Based Deformation Blending for Facial Animation Retargeting
链接: https://arxiv.org/abs/2610.08672
作者: Sihun Cha,Hyeonseung Shin,Suah Yu,Junyong Noh
类目: Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR)
备注:
Abstract:Mesh-agnostic facial animation retargeting transfers expressions across meshes with different structures, but preserving facial motion without surface artifacts remains challenging. To address this, we present PDB, Point-Based Deformation Blending for facial animation retargeting. PDB predicts a compact set of deformed control points from a source neutral-expression pair and blending weights from the target neutral mesh. The weights are computed once per target and reused across frames, while the control points vary with each source expression. ReLU enforces non-negative weights and permits exact zeros, followed by row-wise normalization. The target mesh is reconstructed directly by multiplying the weights and control points, without a predefined cage, precomputed coordinates, a learned per-element deformation decoder, or a global reconstruction solve. Trained only with self-retargeting reconstruction supervision, PDB supports cross-identity transfer without paired cross-identity training expressions. Experiments demonstrate accurate retargeting, fast inference, and localized support in the learned weights. Joint evaluation of expression accuracy and local surface preservation shows reduced surface artifacts relative to the evaluated dense displacement method while retaining the intended motion. Perceptual evaluations further support expression fidelity and visual quality in both self- and cross-retargeting.
[CV-16] Knowing When to Trust a Prior: Reliability-Gated Cue Fusion for Video Gaze Prediction
链接: https://arxiv.org/abs/2610.08663
作者: Lichen Zhu,Yueqian Lin,Yiheng Wang,Hai “Helen” Li,Yiran Chen
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Video gaze prediction is led by gaze-trained models, yet gaze-free priors carry signal those models have not absorbed, if one knows when to trust them. We propose FocusGate, a gated ensemble of gaze-free priors whose members may abstain. A per-frame gate reads three shape statistics of a defocus map and selects the frames on which the estimator is above chance on average, so rejected frames reduce to the base exactly, while midrank normalisation lets an all-zero prior abstain at zero parameters. Gated fusion is significantly positive on film, sports and web video, whereas unconditional fusion is harmful on sports and null on web. Added to four supervised predictors, the NTIRE 2026 champion among them, FocusGate improves all sixteen model-domain cells in shuffled AUC, fifteen significantly, one domain pre-registered and scored once, while adding only 1% to the champion’s latency. Alone, it surpasses TASED-Net and UNISAL in shuffled AUC on film with a 16-frame causal mean.
[CV-17] Selective Transfer of RL Updates for Visual Reasoning
链接: https://arxiv.org/abs/2610.08659
作者: Suxin Ji,Hungtao Wan,Mingjun Liu,An Zhang
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:
Abstract:Model merging provides a training-free way to transfer reasoning capabilities from language models to vision-language models (VLMs), but endpoint-based transfer can conflate pre-existing model differences with changes acquired during reasoning post-training. We instead formulate capability transfer around the training-stage update, isolating the parameter changes induced by reinforcement learning (RL). Yet transferring this update in full remains suboptimal: we find that its components differ substantially in cross-model transferability, with dominant directions transferring more effectively than the complete update. Based on this finding, we introduce Selective-RL, which isolates the RL-stage update, retains its dominant matrix-wise directions with magnitude preservation, and transfers them to the language modules of a VLM. Across three model families and five visual-reasoning benchmarks, Selective-RL improves full-update interpolation in 12 of 15 comparisons, including an 8.55 percentage-point MathVision gain on the Qwen recipient. Matched controls show that update magnitude or arbitrary low rank alone does not reproduce these gains. These results highlight a distinction between what is acquired during post-training and what remains transferable across models, providing a training-stage perspective on cross-model capability transfer. Code is available at this https URL.
[CV-18] Stable Scores Unstable Answers: Frame Phase and Option Order in Video Multiple-Choice Evaluation
链接: https://arxiv.org/abs/2610.08649
作者: Lichen Zhu,Yiheng Wang,Yueqian Lin,Hai “Helen” Li,Yiran Chen
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Video-language models are ranked by multiple-choice accuracy on frames from a uniform grid. The grid has two parameters, a rate and a phase, and benchmarks report only the rate. The phase moves answers: two deployed samplers differing only by a half-step phase offset answer 23.6% of questions differently while scoring within a point, and across four releases from two families shifting only the phase changes roughly one answer in five after controlling option order. PHASEFUSION decodes three offset grids and averages the option posteriors. The grids are the polyphase components of the dense grid. Fusion matches a 32-frame single pass in accuracy within a prespecified margin (logit-scored) and cuts the answers a half-step shift of all three grids changes from 18.2% to 10.1%. Option order, which changes only the presentation, is flagged instead by a one-pass answer margin. Report the phase convention with the budget, or marginalize it.
[CV-19] Forensic Reserve: Eliciting Latent Knowledge for Image Forgery Detection
链接: https://arxiv.org/abs/2610.08639
作者: Jiahua Li,Zixu John,Tom Zhong,Fuping Wu,Tianhao Xu,Jianqing Zheng,Yuanhan Mo,Fei Shen
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:As generated images become increasingly realistic, reliable forgery detection is essential for maintaining trust in visual information. However, existing methods primarily rely on task-specific supervision to adapt vision foundation model representations, without fully exploiting internal forensic knowledge to guide detection. To address this limitation, we propose Reserve-Guided Elicitation (RGE), a framework that treats sparse, origin-sensitive internal components in pretrained models as a forensic reserve and translates their localization into structural constraints for lightweight adaptation. Specifically, we first use the Forensic Lens (F-lens) to decompose activations across layers and token groups into independent components and globally screen them by their response differences between real and generated images, identifying reserve sites and directions. Next, we map the selected directions back to hidden-state space to construct fixed reserve subspaces and insert Forensic Reserve Adapters (FRA) only at the identified sites. Finally, with the backbone parameters, previously fitted reference classifier, and subspace bases fixed, we train only the FRA coefficient maps to generate input-dependent residual updates constrained to the corresponding subspaces, strengthening existing forensic responses. Using only 500 labeled training images and a trainable parameter budget below 0.2% of the backbone, RGE achieves competitive performance across three detection benchmarks without target-benchmark adaptation. Furthermore, RGE consistently improves over the corresponding frozen detectors across eight encoders spanning self-supervised and vision-language pretraining, eliciting a latent forensic capacity broadly shared across pretrained vision models.
[CV-20] LiDAR Resolution Recovery via Foundation-Model-Guided Diffusion
链接: https://arxiv.org/abs/2610.08620
作者: Samed Doğan,Nico Leuze,Alfred Schöttl
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:High-beam-count LiDAR sensors are costly, yet many perception pipelines require dense angular sampling. Using a pretrained Stable Diffusion model as the backbone, we fine-tune a LiDAR-conditioned depth model with pseudo-depth targets from a 2D foundation model. During training, the LiDAR conditioning is randomly decimated at different beam budgets. We then investigate how much of a LiDAR scan can be recovered from heavily decimated input and characterize performance across the input beam budget. We evaluate against physically held-out real beams on nuScenes and report recovery separately from fit accuracy. Our model yields its largest advantage in very sparse regimes, achieving a \delta_1.25 accuracy of 66.8 % from 4 -beam input where scattered interpolation reaches only 45.1 %. A class-stratified error breakdown further reveals that planar surfaces recover first while objects introducing depth discontinuities degrade earliest. Together, these results quantify the recovery/resolution trade-off for foundation-model-guided LiDAR enhancement.
[CV-21] FedDermaSeg: Federated Learning for Dermatological Image Segmentation
链接: https://arxiv.org/abs/2610.08574
作者: Anabik Pal,Ganesh Patidar,Bikash Santra
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:
Abstract:Skin cancer is a major global health concern, and early detection and accurate lesion delineation are important for effective diagnosis and treatment planning. Automated skin lesion analysis can assist dermatologists, with lesion segmentation serving as a fundamental step in computer-aided diagnostic systems. Conventional deep learning-based segmentation models typically rely on centralized training, where images and their corresponding segmentation masks are collected on a central server. Such data aggregation raises privacy concerns in medical applications and requires substantial centralized computational resources. To address these limitations, we investigate the feasibility of federated learning for privacy-preserving skin lesion segmentation. The training and validation sets of the ISIC 2018 Skin Lesion Segmentation Challenge dataset are used to simulate a distributed learning environment and develop a federated segmentation model. The resulting model is evaluated on the ISIC 2018 test set and the PH2 dataset to assess its performance and generalizability. Experimental results demonstrate that the federated model achieves performance comparable to centralized training while consistently improving upon the locally trained models. These findings demonstrate the potential of federated learning for collaborative skin lesion segmentation without requiring centralized aggregation of medical images.
[CV-22] Sparse2comm: Towards Robust Cooperative 3D Object Detection
链接: https://arxiv.org/abs/2610.08573
作者: Lei Yang,Boqi Li,Chunmian Lin,Li Wang,Ziying Song,Shaoqing Xu,Heye Huang,Haibao Yu,Chen Lv
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 15 pages. Code: this https URL
Abstract:Cooperative perception improves autonomous driving by sharing complementary observations among vehicles and roadside infrastructure for 3D object detection. However, practical deployment is constrained by limited bandwidth and unreliable cooperation, where packet loss, transmission delay, and spatial misalignment jointly degrade the cooperative feature stream. Existing methods often reduce communication cost or compensate for one degradation type, leaving coupled disturbances insufficiently addressed. To address this problem, we propose Sparse2comm, a bandwidth-efficient and robust cooperative 3D object detection framework that treats unreliable cooperation as progressive restoration over degraded cooperative features. Sparse Feature Encoding first encodes communication as randomly mask-sampled foreground features transmitted by collaborating agents, from which the ego vehicle reconstructs dense semantic representations. This sparse-to-dense mechanism learns to infer missing object-centric content from sparse observations, enabling ultra-low-bandwidth communication and packet-loss recovery within the same representation. On the semantically restored features, Latency-Aware Alignment predicts motion flow to compensate delayed messages, and Self-Calibrating Fusion estimates residual spatial offsets in a self-supervised manner before adaptive cross-agent fusion. Sparse2comm therefore restores semantic completeness, temporal consistency, and spatial alignment in an ordered pipeline. Extensive experiments on DAIR-V2X, OpenV2V, and V2V4Real show that Sparse2comm maintains competitive clean accuracy and consistently improves robustness under individual and mixed real-world degradations. Compared with the selective feature communication baseline Where2comm, Sparse2comm improves mixed-setting AP@0.5/AP@0.7 by +20.15/+11.79, +12.66/+11.07, and +15.36/+12.61 on the three datasets, respectively.
[CV-23] Less Is More: A Leakage-Controlled Study of Dermoscopic Preprocessing for Joint Skin Lesion Classification and Segmentation with YOLO26
链接: https://arxiv.org/abs/2610.08570
作者: Truong Viet Vu,Nguyen Chi Hai,Nguyen Phuc Nguyen,Ngo Hoang Tu,Vo Nguyen Quoc Bao,Nguyen Thai Anh
类目: Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)
备注: 6 pages, 3 figures, 5 tables
Abstract:Handcrafted preprocessing is widely employed in automated dermoscopic analysis to suppress imaging artifacts and enhance lesion visibility. Nevertheless, its actual contribution to modern real-time models remains unclear, particularly when evaluation protocols do not adequately control correlations among images of the same lesion. This study presents a leakage-controlled, lesion-disjoint evaluation of dermoscopic preprocessing and augmentation for joint multi-class lesion classification and instance segmentation using a fixed nano-scale YOLO26 segmentation model (YOLO26n-seg). From HAM10000 (10,015 images), quality control yields 10,013 valid image-mask pairs from 7,468 unique lesions, partitioned into mutually exclusive sets by lesion identity. With the architecture, resolution, training budget, and evaluation protocol held fixed, we compare minimally processed images plus online augmentation against offline class balancing, DullRazor-CLAHE preprocessing, and raw-processed hybrid views, over three random seeds. On the lesion-disjoint test set, the raw baseline achieves a mask mAP _50:95 of 0.5636 \pm 0.0234 , a Dice score of 0.9356 \pm 0.0024 , and a macro-F1 score of 0.6917 \pm 0.0202 . Offline augmentation does not improve the mean performance, while the combined and hybrid strategies reduce both class-aware segmentation and classification accuracy. At only 2.69 million parameters, the model runs at approximately 50 frames per second. Under a leakage-controlled, lesion-disjoint protocol with all non-input factors held fixed, minimally processed dermoscopic images combined with standard online augmentation deliver a better accuracy-efficiency trade-off than increasingly complex deterministic preprocessing, which yields no consistent joint benefit across three seeds on HAM10000.
[CV-24] RSJEV: Discriminative Remote Sensing Scene Classification with Multimodal Large Language Models
链接: https://arxiv.org/abs/2610.08539
作者: Dongchen Si,Di Wang,Mingzhen Xu,Jing Zhang,Bo Du,Liangpei Zhang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Remote sensing scene classification is a fundamental task in Earth observation and geospatial analysis. Existing approaches mainly follow three paradigms: task-specific visual classification, vision-language similarity matching, and autoregressive multimodal generation. However, visual classifiers rely on predefined label spaces, CLIP-based methods perform recognition through static image-text alignment, and multimodal large language models (MLLMs) introduce unnecessary token-level generation for classification tasks with explicit candidate categories. To address these limitations, we propose RSJEV, a one-pass multimodal decision framework for remote sensing scene classification. Unlike conventional MLLMs that formulate classification as autoregressive text generation, RSJEV reformulates scene classification as a candidate-conditioned multimodal discriminative decision process, where visual representations, task instructions, and candidate category semantics are jointly modeled. Specifically, we introduce a OnePass Decider that extracts multimodal decision states and directly estimates category probabilities within the candidate category space, eliminating autoregressive decoding while preserving vision-language interactions. Extensive experiments on three widely used remote sensing scene classification benchmarks, including UC Merced, AID, and NWPU-RESISC45, demonstrate that RSJEV achieves superior classification performance compared with representative CNN-, Transformer-, Mamba-, CLIP-, and MLLM-based methods. Moreover, RSJEV significantly reduces inference costs and achieves a better accuracy-efficiency trade-off with only a compact 0.8B-parameter model. These results demonstrate the effectiveness of state-conditioned multimodal decision making for efficient remote sensing image understanding. The code will be available at this https URL.
[CV-25] Beyond Perturbation Magnitude: Direction-Dependent Responses in Multimodal Geometric Representations
链接: https://arxiv.org/abs/2610.08533
作者: Yongsheng Luo,Wengan He,Yu Li,Rouying Wu,Wei Lv
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Sound (cs.SD)
备注: Submitted to IEEE Transactions on Multimedia (TMM). 12 pages, 6 figures, 3 tables
Abstract:Geometric alignment scores based on Gram determinants provide a compact way to model higher-order consistency among modalities, yet how such scores respond to modality degradation is poorly understood. This paper asks whether the response of a multimodal geometric score is determined primarily by the magnitude of the perturbation-induced displacement. Using frozen cohorts from MSR-VTT (N=878) and DiDeMo (N=980), we apply controlled video blur and audio noise and analyze the response in the relational geometry on which the score is defined. Displacement magnitude explains at most 15% of the out-of-sample variance in the absolute response, and magnitude-matched pairs respond systematically differently, so scalar magnitude does not organize the response. The closed-form first-order expansion of the Gramian volume yields the Directional Geometric Response (DGR): the projection of the displacement onto the local volume gradient, which jointly captures the clean operating point, displacement magnitude, and displacement direction. The absolute first-order DGR term explains the observed response with out-of-sample R^2 of 0.838-0.969, matched-magnitude ranking accuracies of 0.864-0.963, and response-sign accuracies of 0.909-0.989, whereas the tested direction-free alternatives remain weak or unstable under the corresponding evaluation protocols. A pre-specified gain-normalization candidate, V/(g_V+eps), fails its predictability and clean-order gates. DGR uses the observed degraded-state displacement and is therefore an explanatory quantity, not a deployment-time predictor: geometric response depends on where the representation operates, how far degradation moves the relational geometry, and in which direction it moves.
[CV-26] MedCORE: Criteria-Grounded Clinical Reasoning for Interpretable Medical Image Diagnosis
链接: https://arxiv.org/abs/2610.08528
作者: Asim Khan,Samee Ullah Khan,Dwarikanath Mahapatra
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注: 16 pages, 4 figures, conference
Abstract:Clinical diagnosis is inherently a structured reasoning process, yet existing deep learning models often bypass this structure by mapping image features directly to disease labels without explicitly interrogating the morphological and textural criteria that clinicians systematically evaluate. This limits diagnostic transparency and may compromise safe clinical deployment. We present MedCORE (Medical Criteria-Oriented Reasoning and Evidence), a structured diagnostic framework that operationalizes clinical reasoning within a vision-language architecture. For each input image, MedCORE decomposes the diagnostic process into clinically defined criteria, spatially localizes each criterion to diagnostically relevant image regions, encodes evidence through multi-scale representations that capture macro-structural and micro-textural pathological characteristics, and refines criterion representations using a Graph Attention Network that explicitly models inter-criteria dependencies. Criterion representations are further aligned with clinical text descriptors, reinforced through class-wise visual prototypes, and aggregated using uncertainty-calibrated weighting that proportionally discounts low-confidence diagnostic evidence. MedCORE is validated across three clinically heterogeneous imaging modalities, including dermoscopic lesion classification on ISIC 2018, breast ultrasound lesion characterization on BUSI, and diabetic retinopathy grading on IDRiD. Quantitatively, MedCORE achieves 89.2% accuracy, 85.7% macro-F1, and 96.4% AUC on ISIC 2018; 96.1% accuracy, 95.2% macro-F1, and 98.4% AUC on BUSI; and 84.3% accuracy, 80.2% macro-F1, and 92.8% AUC on IDRiD. These results demonstrate consistent improvements over strong CNN, transformer, biomedical vision-language, concept-based, and prototype-based baselines.
[CV-27] WareFly-VLA: A Vision-Language-Action Framework for UAV Navigation and Human Tracking in Smart Warehouses
链接: https://arxiv.org/abs/2610.08526
作者: Thinh D. Le,Son T. Nguyen,Duong Q. Nguyen,Dung D. Le,Ngo Anh Vien,H. Nguyen-Xuan
类目: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
备注: 41 pages, 35 figures, 11 tables
Abstract:Vision-Language-Action (VLA) models have achieved impressive results in robotic manipulation and ground-mobile navigation, yet language-conditioned control of unmanned aerial vehicles (UAVs) in smart warehouses remains largely unexplored, hindered by the lack of benchmarks that jointly provide continuous low-level flight actions, fine-grained natural-language target descriptions, and realistic industrial environments. This paper introduces WareFly-VLA, a photorealistic UAV VLA framework and dataset for language-guided human search, localization, and tracking in warehouse environments. It contains 507 human-teleoperated flight episodes and 8,504 high-resolution RGB transitions collected in NVIDIA Isaac Sim, each paired with a human-written appearance description of the target worker and a synchronized four-degree-of-freedom control command. Two aerial tasks are covered: target approach and person following, under occlusion, long-range search, altitude variation, and clutter. A unified benchmark of four open-source VLA architectures (SmolVLA, GR00T N1.7, pi_0 and OpenVLA) is established under a leakage-free episode-level protocol at two control rates. The results show that language-conditioned aerial control in warehouses is far from solved: performance drops substantially under strict generalization settings, continuous action modeling consistently outperforms discrete action tokenization, only the forward channel is reliably learnable from a single frame, and current foundation-model interfaces transfer poorly from ground and humanoid embodiments to aerial platforms. The synchronized video, language, action, pose, and difficulty annotations further support world-model research. The dataset, baselines, and evaluation protocol are released to support language-grounded aerial autonomy in smart warehouses.
[CV-28] 2D Spatial Reasoning with Adaptive Neural Cellular Automata
链接: https://arxiv.org/abs/2610.08518
作者: Martin Spitznagel,Janis Keuper
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Many modern learning approaches are still struggling with spatial reasoning tasks, i.e. they lack the ability to utilize geometric information of perceived entities and their spatial relation to each other to solve problems. We introduce a novel Adaptive Neural Cellular Automata (aNCA) architecture which uses deformable convolutions to dynamically adapt the perceptive field and iteratively reason over 2D spatial relations on grid-like data structures (e.g. images). Empirical results on public benchmarks show state of the art comprehensible results with high generalization abilities for solving image based puzzles like Sudoku or finding the shortest path in a maze.
[CV-29] Knee3DVLM: Dual-Sequence Full-Volume Vision-Language Modeling for Comprehensive Knee MRI Assessment
链接: https://arxiv.org/abs/2610.08482
作者: Maryam Baizhigitova,Andrew Seohwan Yu,Po-Hao Chen,Naveen Subhas,Sixu Chen,Xinxin Wang,Kunio Nakamura,Richard Lartey,Xiaojuan Li,Mingrui Yang
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注: 11 pages, 2 figures, 5 tables
Abstract:Vision-language models (VLMs) are increasingly being applied to three-dimensional medical imaging, but their application to knee MRI remains limited, particularly for interpreting the complementary sequences used in clinical practice. We introduce Knee3DVLM, a sequence-aware VLM that uses full-volume DESS and fluid-sensitive TSE MRI to predict 57 anatomically resolved binary diagnostic targets derived from the MRI Osteoarthritis Knee Score (MOAKS) for structured reporting. We evaluated DESS-only, TSE-only, and paired DESS-TSE configurations using subject-disjoint Osteoarthritis Initiative partitions. In a held-out cohort of 1,074 examinations, the fused model achieved 72.98% average accuracy, 71.17% balanced accuracy, 78.96% mean ROC-AUC, and 78.74% macro ROC-AUC, the highest values among the three configurations. In a secondary multiclass analysis aligned with the released 3DReasonKnee cohort, Knee3DVLM was numerically higher than the strongest reported 3DReasonKnee configuration across five pathology categories. These findings support dual-sequence full-volume modeling for comprehensive knee MRI assessment.
[CV-30] HuC-VideoMAE: Human-Centric Video Masked Autoencoding from synthetic data
链接: https://arxiv.org/abs/2610.08433
作者: Ricardo Pizarro,Roberto Valle,José M. Buenaposada,Luis M. Bergasa,Luis Baumela
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Modern action recognition models rely on video transformers pretrained on massive collections of web-crawled videos, such as Kinetics-700. However, the use of such data raises ethical concerns, as subjects’ consent is typically not obtained. Recent high-quality synthetic video datasets generated from motion-capture data, such as BEDLAM2.0, offer a promising ethical alternative. In this work, we investigate self-supervised pretraining of video transformers on synthetic human-motion datasets. We first show that directly applying the standard VideoMAE masking strategy leads to substantially worse performance than pretraining on Kinetics. To address this limitation, we propose a human-centric masking scheme that leverages body keypoints and person bounding box regions. Our approach encourages the model to focus on the structure and dynamics of human motion during pretraining. Experiments on NTU RGB+D and Toyota-Smarthome demonstrate that our method significantly outperforms standard VideoMAE pretraining on synthetic data, closing 49% of the gap to Kinetics pretraining on NTU RGB+D cross-view-subject without using a single real frame during pretraining. To promote the use of ethical action recognition models, we will publicly release our pretrained models.
[CV-31] Deformable CT-US Registration via Anatomy-Aware Implicit Neural Representations MICCAI2026 MICCAI MICCAI-2026
链接: https://arxiv.org/abs/2610.08419
作者: Agnieszka Lach,Magdalena Wysocki,Feng Li,Mohammad Farid Azampour,Benjamin D. Killeen,Felix Ginzinger,Mathias Braun,Philipp Steininger,Heinz Deutschmann,Nassir Navab
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 10 pages, 3 figures. Accepted at the 7th International Workshop on Advances in Simplifying Medical UltraSound (ASMUS 2026), held with MICCAI 2026; to appear in Springer LNCS 17276 (MICCAI 2026 Workshops and Challenges). Open-access camera-ready: this https URL
Abstract:Slice-to-volume registration between ultrasound (US) and preoperative computed tomography (CT) imaging would enhance many minimally invasive interventions, for example by locating soft tissue structures intra-operatively that are discernible in CT. While optical tracking enables initial rigid registration, contact from the probe induces soft tissue deformations that inhibit accurate alignment. In this work, we introduce a deformable CT-ultrasound registration framework that incorporates anatomical priors derived from CT to improve registration under deformation. Rigid registration is first established using a robot-assisted optical tracking system, after which a deformable transformation is estimated using a sinusoidal implicit neural representation (SIREN) optimized per frame. Tissue stiffness is approximated from CT-based HU values and used as spatially varying regularization, suppressing deformation in rigid structures such as bone while allowing more flexibility in soft tissue. Two additional constraints capture the physics of probe contact: a contact-zone displacement prior that drives the displacement field to compress tissue below the probe face, and a fan-geometry regularization term based on beam direction and convex transducer field of view. Model parameters are optimized with a normalized gradient field (NGF). The proposed approach improves alignment over rigid initialisation by 17% and outperforms classical deformable baselines while maintaining near-zero topological folding.
[CV-32] From the Drosophila Visual Connectome to General-Purpose Computer Vision
链接: https://arxiv.org/abs/2610.08418
作者: Zongyu Li,Akito Yamauchi,Huaizhi Liu,Vishwanatha Rao,Jia Guo, for theFrontotemporal Lobar Degeneration Neuroimaging Initiative, for theAlzheimer’s Disease Neuroimaging Initiative
类目: Computer Vision and Pattern Recognition (cs.CV); Neural and Evolutionary Computing (cs.NE); Neurons and Cognition (q-bio.NC)
备注: 27 pages, 17 figures, 7 tables
Abstract:Biological connectomes encode structured solutions to visual computation that may provide reusable inductive biases for artificial vision. We develop ConnectomeX around FlyVision, a trainable architecture that preserves parallel ON/OFF processing, recurrent computation and population-level graph interaction while scaling model capacity across tasks. FlyVision reached 99.34% accuracy on MNIST with 80,608 parameters and 78.03% on CIFAR-10 with 81,408 parameters. On ImageNet-1K, FlyVision Base and Large reached 60.79% and 66.25% top-1 accuracy with 1.8 and 3.7 million parameters, while a Large local-k7 model with a learned low-frequency branch reached 66.53%, compared with 69.25% for ResNet18 with 11.7 million parameters. On a 22-class skin-disease benchmark, FlyVision Large achieved 63.78% accuracy and 95.28% macro-AUROC with 2.99 million parameters. In four-class chest radiography, ImageNet-pretrained FlyVision Base and Large reached 92.60% and 92.76% accuracy with 1.33 and 2.97 million parameters, compared with 91.56% for ImageNet-pretrained ResNet18 with 11.18 million. BrainAGE extends FlyVision to volumetric T1-weighted MRI by applying a shared ImageNet-pretrained FlyVision Large encoder to 24 sagittal, coronal and axial slices per scan and combining slice-level age estimates by confidence-modulated Gaussian voting. On 433 held-out scans, three-axis fusion achieved a mean absolute error of 5.98 years and R^2 = 0.868. Across the 224x224 classification tasks, the best FlyVision configuration remained within three percentage points of ResNet18 on ImageNet-1K and skin-disease classification and exceeded it on chest radiography with substantially fewer parameters. These results show that a conserved connectome-informed computation can scale from compact recognition to large-scale natural and biomedical vision.
[CV-33] Ariadnes Thread of LipSync: Unraveling Forgeries via Inconsistency between Lip Motions and Head Poses ICML2026
链接: https://arxiv.org/abs/2610.08417
作者: Tianyi She,Jiawei Liu,Weifeng Liu,Hanqing Zhao,Weiming Zhang,Kejiang Chen
类目: Computer Vision and Pattern Recognition (cs.CV); Cryptography and Security (cs.CR)
备注: 24 pages, Accepted at ICML 2026
Abstract:Recent advances in LipSync generation technology have led to the creation of highly realistic videos, posing severe societal risks. However, existing defense strategies struggle against LipSync forgeries, as advanced LipSync generation methods not only achieve better lip synchronization but also eliminate visual artifacts. An important reason is that they overlook an inherent biological coupling between lip movements and head poses in natural speech videos. In this paper, we propose LipDA, a novel framework for joint LipSync Detection and Attribution, which takes advantage of the inconsistency between head and lip. For detection, the framework learns to quantify this discrepancy by contrasting lip and pose features from authentic versus forged videos. For attribution, our method is designed to capture the unique temporal dynamics and audio-visual synchronization patterns that act as the fingerprint of models, enabling source tracing. We conduct extensive experiments on two challenging LipSync datasets as well as our own proposed large-scale and multi-generator dataset. LipDA achieves over 97% AUC in detection and 97.5% accuracy in model attribution, significantly outperforming existing methods. Code and the proposed LipSync-A dataset are available at this https URL.
[CV-34] Image Bitstream Fine-grained Understanding for Privacy-Friendly AIoT
链接: https://arxiv.org/abs/2610.08414
作者: Zhen Yu,Wenyang Liu,Kejun Wu,Chengwang Xiao,Renjie Qiao,Chengtao Cai
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Image Bitstream Fine-grained Understanding (IBFU) aims to directly perform fine-grained classification and semantic description generation from encoded image byte sequences. In contrast to conventional pixel-domain visual understanding, IBFU conducts semantic analysis without fully decoding images into the pixel domain. Since pixel-level visual content is not explicitly reconstructed during inference, this paradigm reduces visual exposure within the processing pipeline and suits privacy-friendly Artificial Intelligence of Things (AIoT) applications. In this paper, we propose Bitstream Fine-grained Generator (BFG), a novel foundation model tailored for IBFU. BFG consists of two main components: a Bitstream Semantic Encoder (BSeE) and a Fine-grained Semantic Generator (FSeG). BSeE directly models semantic representations from encoded image bitstreams without explicit pixel reconstruction, while FSeG transforms the extracted bitstream semantics into detailed natural-language descriptions through autoregressive generation. To train BFG and comprehensively evaluate IBFU in practical AIoT scenarios, where image bitstreams may suffer corruption during transmission and storage, we construct a large-scale Corrupted-bitstream Fine-grained Understanding dataset (CFU-D), containing both intact bitstreams and corrupted variants across multiple corruption types and severity levels. Experiments show that BFG maintains stable fine-grained caption generation under bitstream corruption. For example, the performance only has slight change from 0.6339 to 0.6077 in terms of average CIDEr score on Stanford Dogs Caption dataset, while vision-language models, such as Qwen-VL-Chat, BLIP-2, GLM, Gemini, and GPT suffer severe performance decrease. This paper provides a practical paradigm for privacy-friendly fine-grained understanding in AIoT.
[CV-35] Decoy and disclosure radii of invariant shape descriptors
链接: https://arxiv.org/abs/2610.08410
作者: Tanush Shaska,Lubjana Beshaj
类目: Computer Vision and Pattern Recognition (cs.CV); Earth and Planetary Astrophysics (astro-ph.EP); Cryptography and Security (cs.CR); Algebraic Geometry (math.AG)
备注:
Abstract:A recognizer that compares rotation-invariant descriptors sees a surface only up to the fiber of the descriptor. We measure this fiber by its radius in the orbit distance from the enrolled surface. A large radius admits decoys, that is, distant shapes that pass the matcher. A small radius discloses the enrolled shape to anyone who captures the stored value. For star-shaped surfaces truncated to spherical harmonics of degree at most L , with n coefficients, a descriptor of generic rank r has generic fibers of dimension n-3-r modulo rotations. The standard pool of band powers, even bispectra, and three invariants of the degree-three band therefore admits decoy families of dimension 5 , 13 , 20 at L=4,6,8 . Its rank first reaches n-3 at L=16 , and a mirror decoy remains at every L . The odd bispectra remove the mirror decoy generically for L \geq 4 . Yet at fixed mean radius the same pool determines the enclosed volume exactly, and it does not determine whether a surface meets a clearance requirement. We certify two cases by exact and interval arithmetic. At L=6 a decoy matches all 32 invariants to relative precision 2 \cdot 10^-18 at orbit distance at least 0.87 times the norm of the enrolled tuple. For the radar shape model of asteroid (101955) Bennu, the pool recovers the modeled volume, misses the handedness, and leaves the keep-out radius uncertain by more than 7 , \mathrmm .
[CV-36] GeoPID: Decomposing and Steering Visual Information in Vision-Language Models
链接: https://arxiv.org/abs/2610.08401
作者: Seulgi Kim,Zhixiong Zhang,Xinwei Zhang,Jie Ling,Ronn Shaw
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注: Under Review
Abstract:While recent vision-language models (VLMs) have shown outstanding performance across diverse applications, they tend to under-use visual information and over-rely on textual context. In this work, we propose \textscGeoPID, a training-free framework that analyzes multimodal information within VLMs from a geometric perspective. \textscGeoPID decomposes information into Redundant, Modality-Unique, and Synergistic components through the geometric relationships between visual and textual representation subspaces. Through an extensive analysis across 22 VLMs and 14 benchmarks, we confirm that correct predictions exhibit stronger vision-unique components when questions strongly require visual grounding. Building on this geometric analysis, we introduce a targeted intervention technique that selectively amplifies visual representations along the vision-unique subspace during inference. As a result, visual grounding capabilities were enhanced without any additional model parameter updates, achieving an average relative accuracy gain of 7.63%.
[CV-37] UP-MOPD: Update Projection in Multi-Teacher On-Policy Distillation
链接: https://arxiv.org/abs/2610.08398
作者: Taojie Zhu,Jing Jin,Yuan Xia,Chenyang Ding,Qunshan He,Wanke Xia,Tao Sun,Yan Chen,Jian Wang,Jinjie Gu,Tao Feng
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:On-policy distillation from multiple teachers combines expertise from different domains in a single student, but conflicting gradients can hinder this integration. Gradient corrections directly constrain parameter updates under plain SGD. With optimizers such as AdamW, however, momentum, adaptive scaling, and weight decay can turn a corrected gradient into an update that increases a domain loss to first order. To address this gap, we propose Update Projection for Multi-Teacher On-Policy Distillation (UP-MOPD). UP-MOPD lets the original mixed gradient update the optimizer state and generate a candidate displacement, then projects only violating candidates before they are committed to the parameters. The projection gives the unique feasible update closest to the candidate in Euclidean distance. In experiments combining medical and general domains, UP-MOPD improves IFEval-loose accuracy late in training by 2.96 points over vanilla M-OPD. It achieves an average score of 60.03 across eight metrics, compared with 59.00 for gradient projection and 59.15 for update rejection. On a public benchmark covering mathematics, code, and instruction following, it achieves the best average across six tasks (32.67), leads on LiveCodeBench v5, and ties for the best IFEval this http URL results support projecting optimizer updates to reduce interference between domains.
[CV-38] UniCounting: Instance-Aware Proposal Consolidation for Image-Query-Free Multi-Category Counting
链接: https://arxiv.org/abs/2610.08379
作者: Jinshi Liu,Pan Liu,Lei He,Weichao Luo,Rui Qian
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Visual counting is commonly formulated as counting a single specified target, with a model receiving an image-specific exemplar, text query, or target category and returning a single count. We instead study fixed-vocabulary image-query-free multi-category counting. A global vocabulary is fixed for each run, and, given only an RGB image, the model predicts a complete category–count vector without being told which categories appear. We present UniCounting, which casts counting as instance-aware structural inference over an over-complete proposal set. Generic segmenters produce duplicate masks, partial views, and proposals from neighboring instances; semantic scores can name them but cannot determine which denote the same object. Frozen SAM~2.1 generates masks, while frozen DINOv2 and OpenCLIP provide relation and category features. A 3,267-parameter category-shared relation head predicts same-instance affinities from instance-mask-derived supervision. Sparse graph construction, representative selection, labeling, and background-margin admission then convert each admitted component into one count with replayable group evidence. Only the relation head is trained, without count or density-map targets. On COCO clean500, UniCounting obtains lower point-estimate vector \ell_1 error and absent-class false mass than calibrated OWLv2-All80, with comparable micro presence F1. Under a matched decoder, the learned relation reduces both errors relative to mask containment, mask IoU, CLIP, and DINO, while revealing a fragmentation–merge trade-off. We also report transfer diagnostics on OmniCount-sub, FSC-147, and CARPK.
[CV-39] A Stevenss Power Law Check-up of GPT -5.5s Image-Based Visualization Reading IEEE-VIS2026
链接: https://arxiv.org/abs/2610.08365
作者: Kaichun Yang,Jian Chen
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 9 pages, 6 figures, including supplementary material. Accepted by the VISxGenAI workshop at IEEE VIS 2026
Abstract:We adapt Stevens’s power law to measure the innate ability of AI models to read visualizations, which can reveal the built-in perceptual mechanisms of algorithmic models. In our pilot study, models see no legend. A model first views a reference visual representation and estimates its magnitude, then estimates the magnitude of each subsequent image of the same representation relative to that reference. Our evaluation of twelve visual variables makes how algorithmic models read visual encodings measurable, comparable with human perception, and more interpretable to humans.
[CV-40] st-Time Adaptation of Quantized ViTs via Single-Pass Quantizer-Aligned Recalibration
链接: https://arxiv.org/abs/2610.08358
作者: Hyeongheon Cha,Young D. Kwon,Sung-Ju Lee
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注: 44 pages, 6 figures. Code at this https URL
Abstract:Post-training quantization is a standard route to fitting vision transformers (ViTs) into edge compute and memory budgets, yet quantized models become especially brittle under distribution shift. Test-time adaptation (TTA) addresses such shifts without labels, but most existing approaches are poorly aligned with the constraints of quantized inference. Prevailing TTA methods recover accuracy through backpropagation, while backprop-free methods often still incur overhead from extra forward passes or parameter updates, and lightweight feature- or logit-level methods recover only part of the loss. Across these approaches, a quantization-specific failure mode that amplifies the drop is not directly targeted: under shift, activations occupy frozen quantizers’ calibrated ranges differently, distorting their code distribution. We propose Quantizer-Aligned Recalibration (QuAR), a single-pass TTA method tailored to quantized ViTs that neither backpropagates nor updates any model parameters. QuAR recalibrates activations at the input to a frozen quantizer, mapping the test stream’s running per-channel statistics back toward the source calibration. On ImageNet-C with ViT-B, QuAR achieves the highest mean accuracy among state-of-the-art backprop-free TTA methods at 3-, 4-, 6- and 8-bit weight/activation precision, outperforming the strongest baseline by 2.28 points at 8 bits and 4.00 at 3 bits, with 46% lower latency and a memory overhead of only 0.17 MB (0.01% of peak inference memory). Analysis and diagnostics trace the gain to a reduced per-channel mismatch at these quantizers, which restores the code distribution the baselines leave unchanged or distort further. A single fixed configuration remains ahead across continual streams, non-i.i.d. label shift, seven out-of-distribution suites, and three other backbones.
[CV-41] PolarScale: A Physics-Grounded Benchmark for Radiometrically Consistent RGB-to-Stokes Estimation NEURIPS2026
链接: https://arxiv.org/abs/2610.08346
作者: Beibei Lin,Tingting Chen,Xin Zhang,Wenhao Zhao,Dongjun Li,Zifeng Yuan
类目: Computer Vision and Pattern Recognition (cs.CV); Optics (physics.optics)
备注: 22 pages, 17 figures, 8 tables. Accepted to NeurIPS 2026
Abstract:Polarization imaging provides physical cues beyond intensity imaging but typically requires specialized hardware. Recent methods infer polarization from RGB-like inputs, yet predict only normalized Stokes components or relative descriptors, from which the radiometric scale needed for full Stokes reconstruction has been divided out. We introduce PolarScale, a benchmark that makes this scale an explicit prediction and evaluation target. Built on existing trichromatic full-Stokes measurements, PolarScale takes the per-scene normalized total-intensity image s_0 (a scene-referred linear image, not a consumer sRGB photograph) and asks models to predict normalized Stokes components, AoLP/DoLP/DoCP, and a per-scene scale. Because the scale is divided out of the input, it is not physically identifiable; PolarScale therefore evaluates dataset-conditioned semantic scale estimation against a constant-scale control, together with angular, self-consistency, and physical-bound metrics. Across seven restoration-based and generative backbones and three prediction strategies, the strongest restoration models estimate the scale with 3.6-4.3% mean relative error versus 5.7% for the constant control and violate physical bounds on fewer than 0.25% of pixels, whereas two generative baselines collapse to a near-zero scale; explicit descriptor supervision improves descriptor accuracy (23.66 vs. 18.88 dB PSNR for MAE). Predicted full-Stokes representations improve diffuse/specular separation, material segmentation, and glare classification, although in diffuse/specular separation the learned scale performs only on par with the constant control.
[CV-42] DIPrune: Task-Aware Token Pruning with Dual Importance for Efficient Multimodal Language Models
链接: https://arxiv.org/abs/2610.08341
作者: Shuo Yang,Changbai Li,Linlin Yang,Huobin Tan,Rongyu Chen,Tongfei Chen,Tian Wang,Sheng Xu,Baochang Zhang
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注:
Abstract:Recent training-free pruning approaches for Multimodal Large Language Models (MLLMs) effectively cut computational overhead by exploiting visual redundancy or text-vision attention. However, they frequently suffer from semantic degradation due to their task-agnostic design or unreliable attention estimates. Based on our empirical analysis, we have found that this issue arises because salient tokens in shallow layers persistently suppress emerging semantic ones through numerical inertia, leading to premature discarding of signals crucial for deep reasoning. To address the aforementioned issue, from the task-oriented aspects, we first reformulate training-free pruning as a minimization of the distortion in the final task loss and derive a tractable, token-wise upper bound to serve as a surrogate objective. Specifically, this formulation inherently reveals a previously neglected inter-layer term that accounts for gradients across layers. Accordingly, for the implementation, we propose DIPrune, a rank-based framework that employs a dual importance scoring mechanism to jointly optimize intra-layer static feature saliency and inter-layer dynamic semantic evolution. Extensive experiments on LLaVA and Qwen-VL demonstrate that DIPrune consistently achieves state-of-the-art results.
[CV-43] Digital Twin-Driven Real2Sim2Real: Simulator-Conditioned Generation via Paired Driving-Scene Reconstruction
链接: https://arxiv.org/abs/2610.08339
作者: Hojun Lim,Hyeongseok Jeon,Donghyun Kim,Soonyoung Jung,Heecheol Yoo
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 8 pages, 6 figures
Abstract:Camera-based 3D perception for autonomous driving relies heavily on large annotated datasets, and deploying such a system to a new target region typically requires data collection and annotation. Generative augmentation has been proposed to reduce this cost, but existing approaches face a fundamental trade-off: label-conditioned methods consume the very annotations they aim to replace, while simulator-conditioned methods offer free annotations but lack visual grounding to specific real environments. This work investigates the extent to which a digital-twin-driven Real2Sim2Real pipeline (DT-R2S2R) can substitute for target-region real data. By reconstructing recorded driving clips inside a georeferenced digital twin (DT-R2S), we condition a diffusion model on geometrically aligned simulator renderings, establishing a digital twin-grounded Sim2Real model (DT-S2R). As a result, DT-S2R synthesizes photorealistic driving images given low-cost yet georeferenced simulator data across both reconstructed and novel simulator scenes within digital-twin coverage. The efficacy of generated data is verified on diverse 3D detectors. DETR3D, especially, reports 93.18% of mAP obtained by a target-region real-data oracle, without employing target images for detector training. Furthermore, simple co-training with existing out-of-target real data outperforms the oracle. Thus, DT-R2S2R can substantially reduce the cost of manual on-site data collection and annotation in digital twin-available districts, providing a practical foundation for scaling 3D perception.
[CV-44] ransferable Spatial Temporal Coherence Adversarial Attack on Black-Box Vision Language Models for Autonomous Driving
链接: https://arxiv.org/abs/2610.08331
作者: Heyam Bin Jahlan Areej Alhothali Abeer Alhothali
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:
Abstract:The rapid integration of Vision Language Models (VLMs) into sensitive systems introduces critical safety vulnerabilities that remain unexplored in exist studies. While adversarial attack robustness has been extensively studied for image-based models, the susceptibility of VLMs to temporally-aware adversarial attacks against video in driving context poses a distinct and under examined threat. In this paper, we introduce novel adversarial attack against video targeting VLM models used for autonomous driving scenes named Spatial Temporal Coherence Adversarial Attack (STCA). Our attack comprise from three stages: modalities expansion, Spatial attack, and STCA attack. In modalities expansion, we propose caption-guided frame selection method in order to ensure that adversarial perturbation target the most semantically significant frames. this http URL spatial attack, we craft effective perturbation and preserve high similarity. Then the perturbed video generated fed into STCA stage that disrupt cross-frame temporal coherence using motion guided mask. Our method operate under black box threat model against victim target VLMs, relying solely on transferability from white-box surrogate this http URL conduct our experiments on the BDD100K and nuScenes autonomous driving datasets across three VLM models: Video LLaVA-7B, Qwen2.5-VL-7B, and Dolphin. Experimental results demonstrate spatial attack achieves an ASR with high SSIM. Our finding reveal that existing video language model, remain highly susceptible to adversarial attack in autonomous driving scenarios, underscoring the urgent need for robust defense for VLM models.
[CV-45] Catastrophic Forgetting in Sequential Thermal Anti-UAV Detection: The Role of Scale-Conditioned Gradient Imbalance
链接: https://arxiv.org/abs/2610.08315
作者: Khac Duc Giang Nguyen,Seyed Sahand Mohammadi Ziabari,Ali Mohammed Mansoor Alsahag
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Counter-UAV systems based on thermal infrared detection must stay accurate as operational datasets evolve, yet sequential fine-tuning causes catastrophic forgetting of prior tasks, a problem that remains insufficiently characterized in this domain. This continual-learning study measures the stability-plasticity trade-off in YOLOMG, a YOLOv5-based detector run as a single thermal-infrared stream with the motion channel disabled, trained sequentially across three anti-UAV benchmarks of rising scale difficulty: Anti-UAV-RGBT, Anti-UAV410, and CST Anti-UAV. Naive fine-tuning on CST yields a Forgetting Measure of -0.605 against the Stage 1 ceiling, corresponding to a 90% capability loss, with -0.572 occurring in Stage 3 alone. In contrast, knowledge distillation from a frozen teacher is associated with FM = -0.033 +/- 0.004 across three seeds, corresponding to 95% retention. Because no Stage 2 no-KD control is included, this result establishes retention under KD training rather than a causal KD effect. Per-stratum analysis shows large-target detection collapsing to near zero within the first epoch, despite an inter-stage cosine similarity of 0.987 over the gradient-updated weights, pointing to scale-conditioned gradient imbalance, rather than weight drift, as a candidate mechanism. Scale-Stratified Herding (SSH), a 300-exemplar buffer balanced across four UAV size strata, roughly halves the forgetting (FM = -0.605 to -0.311) and keeps large-target detection non-zero. An ablation attributes the gain primarily to scale stratification rather than herding: random-stratified replay performs at least as well (FM = -0.221 versus -0.311 for SSH). These replay results are single-seed and should therefore be treated as preliminary.
[CV-46] Event Detection in Table Tennis Videos using 2D Keypoints
链接: https://arxiv.org/abs/2610.08286
作者: Rainer Lienhart,Daniel Kienzle,Shin’ichi Satoh,Anastasiia Bilinska
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted at the 9th International ACM Workshop on Multimedia Content Analysis in Sports
Abstract:This paper addresses the challenge of automatic, frame-accurate event detection in table tennis videos. Current methods for estimating 3d ball trajectories and ball spin typically require that key events, such as ball-racket contacts, have already been identified in advance. This requirement makes it difficult to apply these methods to longer, unedited video recordings. To overcome this limitation, we propose EventNet, a two-stage pipeline to detect key events: (1) 2d keypoints are extracted of the upper-body poses for both players, table corners and ball center. A small keypoint transformer combines them into a compact representation that is robust to changes in viewpoint, lighting, and background clutter. (2) The temporal sequences of these frame-based representations are processed by a transformer encoder that predicts two time-to-event values for each frame, indicating how close the current frame is to the next and previous ball-racket contact. One novelty is a new, temporal cosine-like target signal. Furthermore, we introduce viewpoint augmentation via 3D reprojection and frame-rate augmentation to improve robustness and generalization. Our extensive ablation study gives deeper insights into the importance of various architectural and training aspects. Experimental results show that the proposed approach achieves an F1 score of 91.16% and a mean frame deviation between ground truth and predicted frame of 0.42 on the Latte-MV dataset and 73.08% / 1.16 on the challenging TTHQ dataset. Overall, our work demonstrates that 2d keypoint-based temporal modeling with our EventNet architecture is a promising and practical approach for automatic event detection in table tennis videos.
[CV-47] Whose Face Is It Anyway? A Multi-Model Audit of Facial Affect Recognition on Children and Why the Gap Is the Head Not the Features
链接: https://arxiv.org/abs/2610.08279
作者: Tobias Hallmen,Robin-Nico Kampa,Elisabeth André
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Preprint. 10 pages, 5 figures
Abstract:Facial affect models are trained almost entirely on adults, yet are increasingly applied to children in education, health, and developmental research. We present a controlled, multi-model audit of five AffectNet-pretrained expression models (EmoNet, EmotiEffLib, DDAMFN++, OpenFace 3.0, LibreFace) on children, across four child image datasets, the AffectNet-8 validation set, and two spontaneous child video datasets, through one shared harness. Three findings emerge. First, the child gap is model-agnostic: every architecture degrades from posed to naturalistic faces and shares the fear \rightarrow surprise confusion. Second, it is concentrated and corroborated across all five models: open-mouth faces (read as surprise, correlating with the AU26 jaw drop) and South-Asian children degrade systematically, with a smaller averted-gaze penalty, while closed-mouth faces, White and Black children, and direct gaze do not; the bias tracks expression morphology and specific populations, not skin tone. Third, the gap is diagnosable: a linear probe on frozen features reaches 0.75-0.91 on unseen children versus 0.48-0.66 zero-shot, so it lies largely in the classifier head, not the representation, whereas dimensional valence/arousal regression degrades sharply under domain shift. Building on this, recalibrating only the head on a little target data recovers +0.13 to +0.28 on the two largest child sets across all five models at negligible adult cost, though the gain is in-distribution and does not transfer across child collections. We will release the harness, per-sample predictions, and analysis code; the child face data stays license-locked and is never redistributed.
[CV-48] SRN-RTVD: Real-Time Video Deblurring System
链接: https://arxiv.org/abs/2610.08230
作者: Nikita Alutis,Danila Evsyukov,Egor Chistov,Mikhail Voronin,Evgeney Bogatyrev,Dmitriy Vatolin
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:As video capture moves to handheld and edge devices, motion blur from camera shake has become a pervasive degradation that lowers perceptual quality and harms downstream vision tasks. The strongest deblurring networks recover impressive detail, yet they remain computationally heavy and overwhelmingly complex, so their quality comes at a cost that consumer hardware cannot pay in real time. This gap between restoration quality and on-device speed is exactly what makes real-time deblurring difficult. We developed and implemented TSRN-RTVD, an efficient video deblurring system that explicitly reconstructs the underlying camera trajectory during exposure and uses the recovered motion to guide restoration. This approach turns the physical cause of blur into a signal that drives sharpening. Our system runs on a single consumer GPU and restores the video at 30 FPS while reaching 30.08 dB PSNR on the GoPro dataset. We demonstrate TSRN-RTVD on consumer devices with interactive side-by-side visualization of the blurry input and the deblurred output, live throughput, and an on-screen view of the recovered camera trajectory. Demo video is available at this https URL. Subjects: Computer Vision and Pattern Recognition (cs.CV) Cite as: arXiv:2610.08230 [cs.CV] (or arXiv:2610.08230v1 [cs.CV] for this version) https://doi.org/10.48550/arXiv.2610.08230 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[CV-49] RACE-FPP: A Robust AI-assisted Characterisation Enhancement for Fringe Projection Profilometry
链接: https://arxiv.org/abs/2610.08213
作者: Osman Ali(1),Xiangjun Kong(1),Tibebe Yalew(1),Waiel Elmadih(2),Samanta Piano(1) ((1) Manufacturing Metrology Team, University of Nottingham, Nottingham, United Kingdom, (2) Taraz Metrology Ltd., Nottingham, United Kingdom)
类目: Computer Vision and Pattern Recognition (cs.CV); Image and Video Processing (eess.IV); Optics (physics.optics)
备注: 19 pages, 9 figures, 7 tables
Abstract:Fringe Projection Profilometry (FPP) requires precise system characterisation to achieve reliable three-dimensional (3D) reconstructions; however, characterisation accuracy strongly depends on robust checkerboard feature localisation, which can deteriorate under challenging imaging conditions such as lens blur and characterisation target orientations. Existing deep learning-based corner detectors are typically assessed using detection metrics and camera reprojection error alone, without considering their wider impact on projector characterisation, camera-projector stereo characterisation consistency, or overall measurement accuracy. In this work, we introduce a complete FPP characterisation pipeline that incorporates deep learning-based corner detection into the standard camera characterisation workflow. We also characterise the projector by sampling phase values at the centres of the white squares in the characterisation target. Rather than treating corner detection as an isolated task, the proposed framework explicitly analyses how localisation errors propagate throughout the entire FPP characterisation chain. Performance is evaluated using detection metrics (e.g., precision and recall), camera and projector reprojection errors, and the camera and projector stereo characterisation. Across a mixed dataset of clean and degraded images, the camera reprojection error is reduced from 1.237 pixels to 0.259 pixels, while the projector reprojection error is reduced by roughly 50%. Dimensional evaluation of reconstructed artefacts shows improved geometric accuracy compared with those resulting from the conventional pipeline. Overall, the findings indicate increased robustness of system-level characterisation under challenging imaging conditions, thereby enabling more reliable industrial FPP measurements.
[CV-50] MacJEPA: Missingness-Robust Audio-Visual Recognition from Untrimmed Egocentric Videos
链接: https://arxiv.org/abs/2610.08192
作者: Souptik Sen,Zahra Ahmadi
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Audio-visual models improve egocentric action recognition by exploiting complementary cues, yet typically assume that both streams remain available at inference. Existing missing-modality methods operate on trimmed, single-event clips in which a stream is entirely present or absent, whereas real sensors fail and recover within long, untrimmed observations. We redefine egocentric modality missingness as temporally localized sensor outages within untrimmed, multi-event observations, with whole-clip absence as the limiting case. We introduce \textbfMacJEPA, a missing-modality-robust \textbfMasked-\textbfcontext query \textbfJEPA that recognizes visual actions and acoustic events from supplied interval queries over audio-visual context. Window-local modality dropout simulates these sensor outages during training. MacJEPA further repurposes masking in JEPA from a self-supervised pretext into a supervised robustness objective, aligning masked and clean latent representations of both multimodal content tokens and the task-conditioned queries. All objectives are optimized jointly with recognition in a single stage, requiring no test-time adaptation. Across Epic-Kitchens-100 and Epic-Sounds, a single checkpoint remains competitive under complete input and consistently surpasses published missing-modality baselines when either the dominant or auxiliary stream is removed. MacJEPA thus unifies strong full-input recognition with temporal missing-modality robustness in a single model operating on untrimmed multi-event videos.
[CV-51] PIE-PS: Photometric Stereo from Physical Irradiance Event Streams SIGGRAPH
链接: https://arxiv.org/abs/2610.08188
作者: Xiangze Meng,Guangyu Li,Jing Li,Di Mei,Songchen Ma,Mingkun Xu,Rui Ma
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 9 pages, 7 figures. Accepted to SIGGRAPH Asia 2026 Conference Papers
Abstract:Event cameras record asynchronous log-image-irradiance changes with microsecond latency and high dynamic range. These properties are useful for photometric stereo under moving illumination, but raw events are sparse and depend on an unknown contrast threshold. We start from the event trigger model and derive a physical relation between adjacent events, light motion, and surface normals. This relation gives a direct physics-only solver, but the solver needs the threshold, enough events at each pixel, and independent per-pixel optimization. To address these limits, we introduce PIE-PS, a learning-based framework for dense surface normal reconstruction from raw event streams and known lighting. We form Physical Irradiance Events (PIEs) by pairing two adjacent events at the same pixel with their corresponding light directions. Each PIE provides a Physical Irradiance Event Feature (PIEF), defined as the signed event rate. PIEF does not require the unknown contrast threshold. To share spatial and temporal context across nearby PIEs, we introduce PIE-GNN, which treats each PIE as a graph node and encodes it with its light-pair geometry. Since the reliability of PIE observations can vary with local appearance, illumination geometry, and sensor noise, Reliability-Grading Attention (RGA) predicts reliability weights to down-weight unreliable PIEs. Pixel aggregation then produces dense normals. Experiments on synthetic and real data show that PIE-PS outperforms prior event-based photometric stereo methods and the direct solver baseline.
[CV-52] View Matters: Keyframe-Guided Text-Driven 3D Gaussian Editing
链接: https://arxiv.org/abs/2610.08179
作者: Kaizhe Zhang,Yijie Zhou,Weizhan Zhang,Xuanyu Wang,Feng Lei,Sha Gong
类目: Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR)
备注: 14 pages, 11 figures, including appendices
Abstract:Text-driven 3D Gaussian editing commonly does not distinguish the editing reliability of rendered views, although different viewpoints provide supervision of substantially different quality. Views that clearly show the scene and match the edit instruction provide reliable guidance, while less informative views may weaken the edit when all views are treated equally. We present View Matters, a view-importance-aware framework that conducts editing around reliable keyframes. Keyframe Importance Estimation (KIE) identifies reliable views using geometric visibility, semantic distinctiveness, and edit relevance. Keyframe-Guided Editing (KGE) then propagates their editing signals asymmetrically to non-keyframes without noisy reverse influence, while Importance-Aware Optimization (IAO) preserves this reliability preference during 3DGS optimization. Across 23 scene-prompt pairs, View Matters achieves the highest average CLIP text-image similarity of 0.2822 and directional similarity of 0.2564 among the evaluated methods, with a four-minute editing time. Additional adjacent-view analysis indicates that the fidelity-oriented editing process maintains cross-view coherence.
[CV-53] Rethinking Visual Provenance: Detection and Watermarking Across Direct Visual Generation and LLM -Driven Code Rendering
链接: https://arxiv.org/abs/2610.08137
作者: Zheng Gao,Xiaoyu Li,Zhicheng Bao,Yang Song,Jiaojiao Jiang
类目: Cryptography and Security (cs.CR); Computer Vision and Pattern Recognition (cs.CV)
备注: 46 pages, 6 figures, 4 tables. Conceptual research agenda; no experiments. Video: this https URL . Project page: this https URL
Abstract:AI systems create images and videos with image/video generation models or by writing code and graphics descriptions that are then rendered. These routes can produce similar visible artifacts but expose different representations, intervention points, and provenance evidence. We develop a production-centered framework that compares detection and watermarking across both routes. An explicit verification specification distinguishes passive inference, message recovery, and authenticated provenance. We organize image, video, source-code, and rendering-aware watermarks by production stage. We examine the different requirements of generated images and video, plots and SVG, programmable video, and agent-composed workflows. Documented Claude, OpenAI, and rendering-tool interfaces connect the framework to concrete systems. We pose ten scoped research questions on identifiability, observability, fair comparison across stages, recoverable payload, reconstruction, synchronization, composition, hybrid local contribution, and private production-event authentication. The result is a conceptual research agenda grounded in published methods, inspected interfaces, and elementary boundary examples. It reports no experiments and claims no new theorems; its appendix results are elementary calculations, and documentation and source inspection establish interfaces, not empirical robustness.
[CV-54] VLA-ACL: Action-Consistent Visual Token Pruning for Efficient Vision-Language-Action Models
链接: https://arxiv.org/abs/2610.08133
作者: Owen Du,Yang Yue,Jie Zhang,Jiaqi Pi,Chi Bene Chen,Gao Huang
类目: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Vision-Language-Action (VLA) models achieve strong robotic manipulation performance but incur high computational costs from processing long token sequences at every control step, limiting real-time deployment. Visual token pruning offers a direct solution, as visual patches dominate the input sequence and contain considerable redundancy. Existing approaches, however, either rely on indirect training-free heuristics, such as attention scores and motion thresholds, or require costly fine-tuning of the base VLA model. We introduce VLA-ACL (Action Consistency Learning), which learns a lightweight visual token pruning policy through action-level supervision while keeping the base VLA model entirely frozen. The training objective encourages actions produced from pruned visual contexts to remain consistent with the full-context teacher, with ground-truth actions as auxiliary supervision. This directly ties token selection to its effect on the downstream control output. Experiments on LIBERO and real-world manipulation tasks show that VLA-ACL prunes up to 87.5% of visual tokens while retaining competitive performance, reduces computation by up to 75%, and achieves a 1.5x inference speedup. These results establish a stronger performance-efficiency trade-off than existing frozen-VLA pruning methods and demonstrate the value of action-level supervision for visual token selection. Code is available at this https URL.
[CV-55] Mu-DisCoCat: A Variational Pipeline for Compositional Generalization on Quantum Processors
链接: https://arxiv.org/abs/2610.08131
作者: Mina Abbaszadeh,Matilda Karabina Moore,Raem Haq,Martha Lewis,Mehrnoosh Sadrzadeh
类目: Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Achieving compositional concept generalization (CoCoGen), the ability to understand novel situations by recombining learned primitives, remains a fundamental challenge in artificial intelligence. Compositional semantic models such as Compositional Distributional Semantics (DisCoCat) offer solutions by generalising vectors to tensors, but suffer from scaling bottlenecks when learning the tensors. Mapping DisCoCat onto Variational Quantum Circuits (VQCs) resolves this limitation for text, yet the methodology has not been expanded to multimodal situations such as the ones involved in CoCoGen. This paper introduces Mu-DisCoCat: a multimodal variational quantum learning framework for DisCoCat that achieves CoCoGen. The framework first learns stable object representations from single-object image-text pairs, then fixes these and uses them to learn the relations between them in multi-object situations. In classical simulations, the model used Uhlmann state fidelity to compute the overlap between the multimodal circuit representations and achieved higher relational OOD accuracy than the evaluated CLIP baseline. Its deployment was evaluated using the destructive SWAP test across noisy quantum emulators, including a range of IBM fake backends, IQM FakeAphrodite, and the IBM Marrakesh quantum processor. Despite real-world device noise, the hardware-executed models maintained a strong positive correlation with simulated fidelities, reliably distinguishing unseen similar and dissimilar pairs. Our work establishes a framework for executing CoCoGen on VQCs, demonstrating a viable use case for near-term quantum hardware.
[CV-56] Supermarket Product Detection and Recognition: Utilizing Deep Learning with Rectified Imagery
链接: https://arxiv.org/abs/2610.08126
作者: Mayank Sah,Jimson Mathew
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注: 10 Pages, 7 Figures, 5 Tables
Abstract:Product Identification has sprung up to become one of the most challenging problems in the automation of the retail industry. With the new industry 5.0 standards, automated inventory management, and catalog creation tasks are vitally important. Object identification models have emerged as a viable answer with their unprecedented identification and localization accuracy. However, the close-knit rack design of supermarkets generates the problem of angle variation in capturing images. The angle-variant densely packed images(a single image contains many objects) become overwhelming for these models alone. In this paper, we try to supplement object detection models with traditional Hough transform (HT) and homogeneous estimation concepts. We study the effect of rectified images using homography estimation and hough transform and their limitations on the problem of grocery identification. We make a case for creating a new dataset to test the effects of such rectification and produce analytical results on different scenarios of angle variation and object densities per image. Extensive experiments on different object detection models suggest that image rectification of angled images improves the detection accuracy of grocery products in images. The results also highlight the limitation of rectification on the angle of image capture and the object density of the image.
[CV-57] Beyond Training from Scratch: Foundation Models for Data-Efficient and Generalizable Cardiac MRI Reconstruction ECCV
链接: https://arxiv.org/abs/2610.08109
作者: Anam Hashmi,Mayug Maniparambil,Julia Dietlmeier,Kathleen M. Curran,Noel E. O’Connor
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted at ECCVW 2026
Abstract:Cardiac magnetic resonance imaging reconstruction aims to recover high-quality images from undersampled acquisitions, enabling faster scans while preserving diagnostic fidelity. Recent reconstruction methods are typically trained from scratch and often require large amounts of task-specific data, limiting their robustness under data scarcity and distribution shifts. In this work, we investigate whether pretrained vision foundation models can serve as effective priors for accelerated cardiac MRI reconstruction. We propose a reconstruction framework that integrates frozen and parameter-efficiently adapted visual encoders, including CLIP, BiomedCLIP, and DINOv2, within a transformer-based reconstruction architecture. Extensive experiments on the CMRxRecon2023 and CMRxRecon2024 benchmarks demonstrate that pretrained representations consistently outperform a transformer trained from scratch across multiple acceleration factors. We further evaluate performance under limited supervision and cross-dataset transfer, showing that foundation models provide superior data efficiency and generalization. While frozen representations are particularly effective in extreme low-data regimes, Low-Rank Adaptation (LoRA) yields additional gains when moderate amounts of training data are available. Among the evaluated backbones, DINOv2 achieves the strongest overall performance. These findings highlight the potential of vision foundation models as robust and transferable priors for cardiac MRI reconstruction.
[CV-58] Multi-Dataset Diagnostic Utility of Clinical Visual Concepts in AI Systems for Dermatology MICCAI
链接: https://arxiv.org/abs/2610.08086
作者: Linda Wermelinger,Simone Lionetti,Fabian Gröger,Nipun Ranasekara,Philippe Gottfrois,Ludovic Amruthalingam,Labelling Consortium,Marc Pouly,Alexander A. Navarini
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted at the MICCAI ISIC Workshop 2026. 11 pages, 3 figures, 3 tables. Code and dataset: this https URL
Abstract:The clinical integration of AI systems in digital dermatology relies heavily on human trust. Clinically interpretable visual concepts can act as intermediate representations enhancing trust and reliability. However, research in this domain is currently limited by scattered, heterogeneous dataset annotations. In this work, we introduce SkinLex, a harmonized dataset of 48 clinical morphological attributes across four public datasets (SkinCon, DermaCon-IN, MM-Skin, and PASSION) for a total of 20,411 records. Supervised nine-partition classification of skin conditions shows that limiting features to specific visual groups, like shapes or colors alone, reduces diagnostic accuracy. Bootstrapped backward elimination reveals that the set of 48 visual concepts has some degree of redundancy for algorithmic nine-partition diagnosis on the examined dataset. This demonstrates that coarse diagnosis on the selected dataset requires a relatively small but varied combination of clinical concepts, and motivates further research to improve concept taxonomy. Results can be translated into clinical benefits by reducing inputs for concept-based models, improving efficiency for annotation and modeling, and further enhancing interpretability. Code and prompt templates are available at this https URL.
[CV-59] Optimization Encoders: Rethinking Second-Order Meta-Learning for Neural Fields
链接: https://arxiv.org/abs/2610.08075
作者: Rudolf L.M. van Herten,Soufiane Ben Haddou,Rachit Saluja,Johannes C. Paetzold
类目: Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Conditional neural fields represent signals continuously, but their effectiveness depends on how the conditional latent representations are inferred from observed data. In meta-learning, this encoding occurs through gradient updates induced by the decoder, tying representation learning directly to decoder design. We formalize this connection by interpreting latent optimization as an optimization encoder, unifying the roles of second-order differentiation, latent parameterization, and task supervision. This concept enables second-order meta-learning for end-to-end training of the encoding procedure alongside the decoder, and clarifies which learning pathway first-order approximations discard. Guided by this view, we introduce Attentive Latent Fields (MetaLF), an equivariant transformer-based neural field that contextualizes a latent pointcloud through self-attention. These interactions shape both field predictions and the updates that construct their representation, allowing local observations to inform coherent non-local structure. Disentangling the inner encoding objective from outer task supervision unifies reconstruction, classification, and segmentation within an end-to-end meta-learning framework, using reconstruction-only latent adaptation at test time. Controlled experiments on polynomial fields link latent coordination to lower effective rank and stronger alignment with the underlying function space. Across image and 3D shape reconstruction, MetaLF improves fidelity within three to five gradient updates, while supporting semantic prediction across images, shapes, and volumes. Together, these findings position the optimization encoder perspective as a unified basis for designing neural fields around how representations are constructed, coordinated, and used.
[CV-60] wo Halves are More than One: Phase-wise Velocity Distillation for Fast and High-Quality Image Generation
链接: https://arxiv.org/abs/2610.08070
作者: Zhen Guo,Rongyuan Wu,Qiaosi Yi,Chenxi Xie,Xinyu Wei,Lei Zhang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Recent diffusion-based image generation backbones have grown substantially in scale, making the network inference cost increase rapidly. While diffusion distillation techniques can reduce the number of inference steps, high-quality image generation within a single full-backbone-forward compute budget remains challenging. Existing one-step methods typically allocate this budget to a single evaluation of a monolithic student. However, approximating the heterogeneous coarse-to-fine transport with a single monolithic mapping is difficult and often leads to over-smoothed outputs. To address this issue, we propose Phase-wise Velocity Distillation (PVD), which partitions the generation timeline into a coarse and a fine phase, and models the transition within each phase via the average velocity. A dedicated half-sized expert is assigned to each phase, decoupling structural composition from detail refinement while keeping the cumulative computation equivalent to one full-backbone forward pass. We show that the use of two half-sized phase-specific experts outperforms a single full-size monolithic student. On class-conditional image generation, PVD achieves an FID of 1.48 on ImageNet 256 x 256. On more complex text-to-image (T2I) tasks, PVD-distilled models (Stable Diffusion 3.5-Medium, FLUX.1-dev, Qwen-Image) produce results competitive with their multi-step teachers, significantly outperforming prior distillation methods. Moreover, across the evaluated T2I backbones, PVD reduces active parameters by 49.10-50.89% and peak VRAM by 45.76-48.36% compared to the corresponding teachers. Source code and distilled models are available at this https URL.
[CV-61] PhysTacGen: Physics-Aware Visual-Tactile Sensor Image Generation
链接: https://arxiv.org/abs/2610.08068
作者: Guo Tang,Yongtao Wang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Realistic physical interaction is a cornerstone of embodied intelligence, yet collecting paired visual–tactile data remains costly. Visual-to-tactile synthesis offers a promising approach to augmenting such data, but learning this mapping is complicated by the gap between visual appearance and contact-related material properties, as well as spatial misalignment in paired observations. To address these challenges, we present \textbfPhysTacGen, a visual-to-optical-tactile image generation framework that integrates material-aware descriptions with geometric conditioning. First, we introduce Group Tactile Policy Optimization (GTPO), a reinforcement learning strategy that refines a vision–language model to generate structured material descriptions using task-specific rewards. Second, we combine DINOv2-based pair curation with monocular relative-depth estimation to select training pairs and provide geometric priors. Finally, an SDXL ControlNet synthesizes optical tactile images conditioned on RGB, relative depth, and GTPO-generated text. Experiments on curated SSVTP data demonstrate improved structural similarity over the compared baselines, while a blinded user study shows a preference for GTPO-generated descriptions. Generated tactile inputs also improve performance on an attribute-derived force-coefficient prediction proxy. Together, these results demonstrate the effectiveness of PhysTacGen for optical tactile image synthesis and its utility in the evaluated downstream this http URL code will be available at this https URL.
[CV-62] Decide Before You Look: Learning Which Retrieved Memories Deserve Pixels
链接: https://arxiv.org/abs/2610.07984
作者: Youxing LI
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:
Abstract:Multimodal assistants answer questions from long-term memories that contain images. After retrieval, each retrieved image reaches the answering model either as pixels, at about a thousand visual tokens per image, or as a stored text proxy that often misses the detail the question asks about. We find that the benefit of pixels usually comes from one or two retrieved memories, and that it can be predicted before the answering model runs, without reading any full-resolution image. In PixelTriage, a plug-in placed after retrieval, a small model that does not generate text reads the dialogue, a short note and a thumbnail of each retrieved memory and predicts how much its pixels would add. It is trained on synthetic memory episodes labeled by a frozen 27B model that answers each question with and without each memory’s pixels. With a 7B answering model, PixelTriage lies on the accuracy–cost frontier of M ^3 Exam, DMV and MemEye and uses 11–23% of the visual tokens without a significant loss of accuracy. On DMV it answers 2.9 times faster than opening all images. It outperforms retrieval order and uniform down-sizing at equal budgets and transfers to other memory systems and to a 397B answering model.
[CV-63] M3SunAgent : Monocular 3D Spatial Understanding Agent for Metric Depth Estimation and 3D Visual Grounding
链接: https://arxiv.org/abs/2610.07982
作者: Jinsong Zhang,Kejun Wu,Ming Zhu,Renjie Qiao,Chengtao Cai,Zhengguo Li
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 13 pages. 7 figures, submitted to IEEE Transactions on Circuits and Systems for Video Technology (TCSVT)
Abstract:Monocular metric depth estimation and 3D visual grounding represent the two complementary cornerstones of monocular 3D spatial understanding (M3Sun), from which the fundamental 3D spatial information required by M3Sun can be acquired. However, these complementary tasks are generally conducted by separate frameworks, which pose challenges of inflexible and unaligned spatial information access for embodied intelligence systems. In this paper, we propose a unified agent for monocular 3D spatial understanding (M3SunAgent) that leverages a large language model (LLM) as a task planner for spatial visual programming, which flexibly generate structured programs and coordinate tools. For instance-level metric depth estimation task, M3SunAgent invokes an object detector tool to locate the target, estimates depth at selected points with a depth estimation tool, and aggregates these predictions into an instance-level depth estimate. We also construct the M3Sun Instance (M3SI) dataset, a benchmark with 2,910 samples for evaluation. For monocular 3D visual grounding task, M3SunAgent uses a vision-language model (VLM) tool to locate the target and output basic spatial attributes, then combines back-projection tool with a dimension-lifting tool to predict its 3D bounding box. Experimental results demonstrate the superior performance of M3SunAgent. Specifically, in evaluations of instance-level monocular metric depth estimation, M3SunAgent achieves the best performance among all compared models, 52.61% of predicted instances are distributed below depth error 0.25 ( \delta 0.25 ). In evaluations of monocular 3D visual grounding, M3SunAgent demonstrates overall competitive performance than vision and VLM models, reaching a 3D mean intersection over union (mIoU) of 41.73% and exceeding the state-of-the-art MonoVLM model by 3.62%.
[CV-64] EmbodiedSmith: Scaling Embodied Data through Recursive Self-Improvement Flywheel in Simulation
链接: https://arxiv.org/abs/2610.07969
作者: Yikai Qin,Yifei Deng,Mingjian Liang,Wenxuan Song,Zepeng Lin,Zhiyi Jiang,Jiajun Fu,Qiao Sun,Huashuo Lei,Xicheng Gong,Jiayi Chen,Han Zhao,Shuanghao Bai,Pengxiang Ding,Pengwei Wang,Haoang Li
类目: Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)
备注:
Abstract:Scaling robotic foundation models requires diverse training data and reliable evaluation environments. Simulation offers a scalable solution, yet existing generation pipelines remain constrained by predefined assets and skills, a disconnect between scene generation and task generation, and limited support for complex embodiments and physics. We introduce EmbodiedSmith, a framework for scalable embodied data generation through recursive self-improvement (RSI). EmbodiedSmith unifies asset, scene, and task generation in a pipeline that supports autonomous creation and language-driven customization. Its core is an agentic refinement loop: scene generation anticipates downstream task requirements, while task generation guides targeted scene edits, allowing scenes and tasks to iteratively improve one another. This joint refinement improves task generation success, including for long-horizon tasks. The framework further supports mobile manipulators, humanoids, and dexterous hands, as well as interactions involving deformable objects and fluids, broadening the range of behaviors and physical phenomena represented in generated data. Together, these capabilities provide a flexible simulation engine for both robot pretraining and evaluation. Extensive experiments validate the quality, diversity, and generation efficiency of the resulting data, while downstream policy experiments demonstrate that increased data diversity improves generalization.
[CV-65] DensiTok: Making Feed-Forward 3D Gaussian Splatting See More Views Than It Is Given
链接: https://arxiv.org/abs/2610.07958
作者: Minhyeok Lee,Jungho Lee,Minseok Kang,Heeseung Choi,Ig-Jae Kim,Sangyoun Lee
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Feed-forward 3D Gaussian Splatting (3DGS) reconstructs a scene in a single forward pass, replacing per-scene optimization with a network trained across many scenes. Its quality, however, degrades sharply as the number of input images drops. The bottleneck is upstream of the reconstruction heads: from a few unposed views, the internal representation they read carries no evidence for unobserved regions, leaving holes, floaters, and blur. The common remedy supplies that evidence as pixels, synthesizing extra views with an image or video generator and re-encoding them, which is costly and not 3D-consistent by construction. We instead densify the evidence itself. We present DensiTok, a plug-in module for pretrained feed-forward 3DGS models that densifies their internal geometry tokens directly, making a frozen backbone behave as though it had observed many more views than it was given. DensiTok compresses those tokens into a compact latent space, completes the latents of the unobserved viewpoints in a single flow-matching step conditioned on camera geometry, and decodes them back into tokens that the original reconstruction heads. The same module design can be integrated into different pretrained predictors while keeping each backbone and its reconstruction heads frozen. Completion in a low-dimensional latent space requires no image synthesis or additional encoder passes. Across three pretrained backbones and two benchmarks, DensiTok consistently improves sparse-view reconstruction and recovers much of the gap to dense-view reconstruction.
[CV-66] Revisiting Numerical Forecasting Models for Language-Based Trajectory Prediction
链接: https://arxiv.org/abs/2610.07954
作者: JunGyu Lee,Inhwan Bae,Hae-Gon Jeon
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 35 pages, 15 figures. Project page: this https URL
Abstract:Language-based trajectory predictors represent coordinates as discrete tokens and learn auxiliary tasks such as destination and group reasoning. This formulation enables the model to capture behavioral intent and social context beyond coordinate dynamics alone. However, token-level objectives provide only indirect guidance for continuous coordinate-space dynamics. To address this limitation, we introduce MoRE (Mixture of Reward Experts), a refinement framework that transfers numerical forecasting priors into a pretrained language-based predictor through reinforcement learning. Five frozen numerical predictors provide complementary coordinate-level knowledge of motion and interactions. Their predictions are converted into expert rewards and combined through an uncertainty-weighted consensus that penalizes disagreement. A ground-truth reward anchors the prediction to the target trajectory. To focus refinement on difficult cases, MoRE refines the policy using the top 1% of training samples ranked by predictive entropy. Expert predictions are computed once and cached before PPO training, so the experts are not run during policy updates or inference. In this way, MoRE combines the contextual modeling of the language-based predictor with coordinate-level feedback from numerical experts. On ETH-UCY, MoRE reduces ADE from 0.22 to 0.20 m and FDE from 0.32 to 0.29 m. Relative to the base policy, ADE decreases by 17.9% on SDD and 12.7% on NBA. On ETH-UCY, MoRE also reduces collision rates and better matches ground-truth pedestrian spacing, without increasing measured inference memory or latency. The project page is available at this https URL.
[CV-67] UltraDiff: Differentiable Ray Tracing in Ultrasound for Shape Optimization SIGGRAPH
链接: https://arxiv.org/abs/2610.07941
作者: Felix Duelmer,Magdalena Wysocki,Nassir Navab,Mohammad Farid Azampour
类目: Graphics (cs.GR); Computer Vision and Pattern Recognition (cs.CV)
备注: 4 pages, 4 figures, 1 table. Accepted at SIGGRAPH Asia 2026 Technical Communications
Abstract:Physically-based differentiable rendering enables gradient-based optimization of scene parameters by matching rendered images to measurements, but has so far mainly focused on light transport. We extend this paradigm to medical ultrasound, where image formation resembles transient rendering: echoes are binned by time-of-flight rather than projected onto an image plane. We present UltraDiff, a modular framework for differentiable ultrasound ray tracing. UltraDiff formulates ultrasound image formation as a path-space integral, gated by travel time between the transducer and tissue interfaces, and derives a Monte Carlo estimator of both the forward model and its gradients with respect to scene parameters. We demonstrate this on an inverse geometry estimation: starting from a sphere, an SDF is optimized until simulated echoes match measured ones, recovering vertebral surfaces from simulated B-mode sweeps and from a real robotic acquisition of a spine phantom. Unlike state-of-the-art ultrasound shape reconstruction methods, which rely on pre-segmented images, our approach operates unsupervised on B-mode images through analysis-by-synthesis, while achieving competitive geometric accuracy. Implemented on top of Mitsuba 3, UltraDiff brings differentiable path tracing to a new sensing modality and provides a foundation for inverse problems in acoustic imaging.
[CV-68] CCDF: A Benchmark Dataset for Deepfake Detection in Real-World Surveillance Footage
链接: https://arxiv.org/abs/2610.07939
作者: Baptiste Chopin,Thomas Swearingen,Arun Ross,Antitza Dantcheva,Christian Rathgeb
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Due to rapid advances in Generative AI, commercial video generation tools can be used to produce fabricated surveillance footage that can fool both human viewers and automated synthetic video detectors. Since these tools are so widely accessible, a malicious user can create a harmful video clip at minimal cost. The production and dissemination of such videos in high-stakes settings, such as crime reporting and elections, can misdirect emergency response efforts or distort political discourse. Existing deepfake video datasets, used by the research community to develop deepfake detection algorithms, exhibit two limitations: (1) they emphasize benign web content rather than footage of possibly malicious activity, and (2) they rely on older or open-source generators that do not represent recent advances in generative systems. We assemble CCtv DeepFakes (CCDF), a video deepfake dataset, to address both gaps. CCDF contains 1840 videos (460 real and 1380 generated) spanning 16 crime and accident categories, with generated content produced using three leading commercial systems: Grok Imagine, Google VEO 3.1, and OpenAI Sora 2. CCDF is a highly realistic, small-scale, manually annotated dataset targeting evaluation of detection models. We release three versions of the dataset: the raw generated data, a cleaned version in which video metadata are standardized between real and synthetic samples to prevent detectors from exploiting trivial cues, and an altered version simulating low-effort post-processing attacks. We evaluate CCDF with ten recent state-of-the-art detectors covering different detection approaches. Our results suggest that these approaches do not reliably distinguish CCDF’s generated videos from real ones, despite their strong reported performance on existing datasets. These results further confirm that existing datasets are not well-suited to evaluating certain threats.
[CV-69] Dynamic Alignment and Calibration for Multimodal Learning
链接: https://arxiv.org/abs/2610.07928
作者: Jinghao Xu,Zhenhua Guo,Xiaofeng Zhu,Xiaoshuang Shi
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注: 17 pages
Abstract:Dynamic multimodal learning aims to learn robust representations by adaptively modeling information discrepancies across modalities. However, existing methods still suffer from two limitations: (i) static cross-modal alignment strategies usually impose uniform constraints on all samples while overlooking sample-wise variations, potentially leading to unreasonable over-alignment; and (ii) confidence- or uncertainty-aware fusion methods often fail to adequately account for feature magnitude and confidence differences across modalities. For modality pairs with significant feature magnitude differences or small confidence gaps, it might be unreliable to strictly align fusion weights according to confidence. To address these issues, we propose an Alignment- and Calibration-driven Multimodal Learning framework (ACML). Specifically, ACML incorporates a dynamic cross-modal triplet alignment module, which enforces strong semantic consistency for high-confidence positive pairs while encouraging diverse representation learning between high- and low-confidence positive pairs according to their confidence gaps. Additionally, ACML introduces a difference-aware attention calibration strategy that adaptively adjusts attention regularization based on feature magnitude and confidence differences across modalities, thereby mitigating biases caused by unreasonable fusion constraints. Extensive experiments on multiple multimodal benchmark datasets demonstrate that ACML consistently achieves superior performance and robustness over recent state-of-the-art methods.
[CV-70] F-PRVR: Training-Free Partially Relevant Video Retrieval
链接: https://arxiv.org/abs/2610.07925
作者: Giyeol Kim,Chanho Eom
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Partially Relevant Video Retrieval (PRVR) aims to retrieve untrimmed videos containing moments relevant to a given text query. Despite recent progress, existing PRVR methods suffer from two key limitations: a fixed video decomposition scheme that causes semantic dilution, and source-domain overfitting induced by task-specific training. In this paper, we propose TF-PRVR, the first training-free framework for PRVR. TF-PRVR leverages frozen vision-language features to construct video-specific hierarchical representations. It derives temporal semantic signals from frame-level features and applies frequency-based multi-scale analysis to identify adaptive temporal boundaries, producing hierarchical segments with coherent event-level semantics. Built on these segments, TF-PRVR constructs a unified multi-scale graph and propagates query relevance across temporally and semantically related nodes. A moment-aware scoring strategy then aggregates temporally aligned relevance across scales, emphasizing consistently supported moments while suppressing isolated false responses. Without task-specific training, TF-PRVR preserves the general-purpose alignment capability of pre-trained vision-language models and avoids dataset-specific overfitting. Extensive experiments demonstrate consistent performance across datasets with diverse visual and temporal characteristics, suggesting a practical direction for training-free PRVR.
[CV-71] OpenWAM: An Open Framework for Composable World-Action Models
链接: https://arxiv.org/abs/2610.07922
作者: Heng Yu,David D. Yuan,Juze Zhang,Changan Chen,Yao Feng,Michelle Baldonado,Steve Cousins,Li Fei-Fei,Jiajun Wu,Ehsan Adeli
类目: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
备注: 18 pages, 5 figures, 14 tables. Project page: this https URL ; Code: this https URL ; Code and project page released June 4, 2026. Equal contribution: Heng Yu, David D. Yuan, Juze Zhang
Abstract:World-action models (WAMs) couple future prediction with robot control, yet existing systems often vary the video backbone, interaction structure, supervision, and inference procedure simultaneously, making their design choices difficult to compare. We introduce OPENWAM, an open world-action modeling framework built around a common causal robot-video foundation and configurable video-action interaction. Starting from Wan2.2-5B, we perform causal robot-video pretraining on over 10,000 hours of video, then integrate an action expert through a shared Mixture-of-Transformers architecture that supports joint, video-then-action, action-then-video, and decoupled generation. OPENWAM achieves high success rates on four LIBERO suites and real-world bimanual tasks; robot-video training with causal adaptation improves VTA success on LIBERO-Long from 68.4% to 97.8%. The same configurable architecture naturally extends to inverse and forward dynamics, allowing us to study how counterfactual transitions improve independently trained dynamics models beyond demonstrations alone. When only the video predictor is adapted to a new task, a frozen local-context inverse dynamics model trained on counterfactual data and demonstrations achieves 84.0% mean success across four held-out LIBERO-90 tasks, compared with 47.0% for a full-context inverse model and 21.5% for a local-context model trained only on demonstrations. For forward dynamics, counterfactual supervision reduces RGB prediction error by 34.5% and raises outcome identification from 21.1% to 71.3% among 16 same-state outcomes. OPENWAM provides a common testbed for comparing WAM interaction designs and for studying dynamics learning from video data beyond successful demonstrations.
[CV-72] Can We Model the Artifacts Explicitly? Disentangle Artifacts via Pairwise Edit Relations for Image Manipulation Localization NEURIPS2026
链接: https://arxiv.org/abs/2610.07916
作者: Xuekang Zhu,Kaiwen Feng,Ruifeng Wang,Xiwen Wang,Xiaochen Ma,Bo Du,Changjiang Jiang,Chenfan Qu,Songyu Ye,Xia Du,Wentao Feng,Jian Liu,Ji-Zhe Zhou
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: NeurIPS 2026 (Oral)
Abstract:Image Manipulation Localization (IML) is commonly formulated as a fully supervised learning task that estimates the optimal manipulation mask y for a given image x . In this work, we first reveal the latent nature of artifacts and thus reinterpret IML as a latent-variable problem, P(y|x)=\int P(y|z),P(z|x),dz , where z denotes the artifacts. Following this interpretation, we pinpoint the cause for the current IML models’ insufficiency as their implicit artifacts modeling strategy, highlighting the necessity of modeling z in an explicit manner. Without direct labels, feature disentanglement is the most appropriate solution for this explicit modeling. Accordingly, we propose a two-stage learning paradigm with the Pairwise Artifacts Learning (PAL) and Standard Localization (SL) phases to estimate P(z|x) and P(y|z) via edit relations. To support our edit-relation-based learning, we further curate EditGroup-45K, a source-anchored dataset organized into edit groups for pair construction. Extensive experiments show that our PAL paradigm yields consistent improvements across diverse IML architectures, and empirical analyses further verify that PAL does capture artifacts explicitly through feature disentanglement. Code and dataset are available at this https URL
[CV-73] Multimodal Knowledge Distillation for Gastric Adenocarcinoma Classification from Whole-Slide Images
链接: https://arxiv.org/abs/2610.07913
作者: Shrihari Dumbre,Bikash Santra
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:
Abstract:Gastric adenocarcinoma (GA) is a leading cause of cancer-related mortality worldwide, and accurate histopathological subtype classification from whole-slide images (WSIs) is essential for effective treatment planning. While multimodal approaches that integrate pathology report text with WSIs can improve classification, existing methods often depend on computationally expensive transformer architectures and large language models. We propose a multimodal knowledge distillation (MKD) framework that combines a pretrained WSI image encoder and a clinical text encoder using Low-Rank Multimodal Fusion (LMF) to efficiently model cross-modal interactions during training. Each WSI is represented as a bag of patches paired with a slide-level diagnostic caption. The teacher model learns fused image-text representations for subtype classification, while the student model distills this knowledge to enable accurate image-only inference. We evaluate our method on the PatchGastric benchmark dataset and achieve at least 3.35% higher mean accuracy than state-of-the-art approaches, without relying on transformer-based fusion, multi-task learning, or large language models. The source code is available at this https URL.
[CV-74] Diverse Motion Customization via Control-based Dynamic Optimization
链接: https://arxiv.org/abs/2610.07911
作者: Youngyoon Choi,Kihyun Kim,Jeongwoo Shin,Joonseok Lee
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注: Preprint
Abstract:Despite recent advances in video generation, motion customization remains challenging due to content leakage, where appearance attributes from the reference video unintentionally propagate into the generated output. We identify this issue as a consequence of the generative process collapsing toward the reference video, which arises from formulating the learning objective as a direct regression on the reference. To address this, we propose Control-based Motion Customization (CMC), a principled training framework that is structurally robust to content leakage. Our key idea is to steer generative dynamics toward desired motion while avoiding collapse toward the reference video, which we formalize using Stochastic Optimal Control (SOC). Under this formulation, customized videos acquire the target motion yet remain within the pre-trained model’s prompt-conditional distribution, where appearance is determined by the text prompt rather than the reference video. Furthermore, to improve efficiency, we tailor the SOC formulation to motion customization by eliminating the need for an explicit reward and introducing a timestep-adaptive motion cost that focuses only on early generative stages, accelerating training by 2.5 times. Extensive experiments demonstrate that CMC effectively mitigates content leakage and achieves competitive motion fidelity while preserving the diversity of the base model across diverse scenarios.
[CV-75] Unsupervised Long-Tailed Adaptation of Vision-Language Models
链接: https://arxiv.org/abs/2610.07903
作者: Keliang Chen,Yaxin Hou,Hui Liu,Yuheng Jia
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 18 pages, 6 figures
Abstract:Adapting vision-language models to downstream tasks has achieved remarkable success by leveraging pseudo-labels generated from unlabeled data. Existing methods typically assume a uniform unlabeled data distribution, and thus the resulting pseudo-label distribution is likewise uniform. However, real-world data distributions are often long-tailed. To tackle this, we formalize a new scenario termed Unsupervised Long-Tailed Adaptation (ULTA). Under this scenario, existing methods exhibit a contrasting phenomenon: head-class performance drops sharply, which is distinct from supervised long-tailed learning where tail classes suffer the most. In particular, we uncover that the distributional mismatch not only erodes head-class boundaries, but also pushes head samples into confusable classes, reinforcing the model’s inherent bias. To address these issues, we propose a novel model called Margin-Aware Refinement with Structural alignment (MARS). Specifically, we mitigate head-class boundary erosion via Boundary-Preserving Alignment, which takes the zero-shot VLM as a fixed visual reference to suppress probability increases that lack visual support in the training targets. Building upon this, we introduce Margin-aware Self-Refinement, which employs a dynamic adjustment strategy to refine tail and confusable classes while preventing prediction bias. Extensive experiments on nine benchmark datasets demonstrate that MARS outperforms state-of-the-art methods, achieving an average accuracy improvement of 4.71 percentage points.
[CV-76] Label-Efficient Deep Learning for ECG Delineation: A Multi-Dataset Benchmark against Widely Used Delineation Tools
链接: https://arxiv.org/abs/2610.07885
作者: Jeonghwa Lim,Minje Park,Yeongyeon Na,Yujin Eom,Soyeon Lim,Young Ho Lee,Yu Jeong Kim,Sunghoon Joo,Ki Hong Lee
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Signal Processing (eess.SP)
备注: 20 pages, 5 figures. First two authors contributed equally
Abstract:Electrocardiogram (ECG) delineation, the identification of waveform boundaries, is a foundational step that translates raw ECG signals into clinically interpretable measurements. Deep learning has advanced this task but remains dependent on costly expert annotations. Label-efficient strategies such as self-supervised pretraining and semi-supervised learning are expected to ease this burden, yet it remains unclear whether they yield reliable delineation and whether the deep models they produce outperform the delineation tools used in practice. We address this in two stages. First, comparing self-supervised objectives with supervised or semi-supervised fine-tuning across one internal and four external datasets, we find that pretraining helps but the objective matters, and that the value of semi-supervised fine-tuning depends on the pretraining objective. Second, we benchmark the selected deep learning model against widely used open-source (NeuroKit2, Prominence, ECGdeli) and commercial (CalECG) tools using three complementary metrics. The model ranks best on every metric and dataset, outperforming the strongest tool by a clear margin on the rhythm-diverse set (mIoU 71.3 vs. 54.8%; averaged point-wise sensitivity 92.6 vs. 76.4%), and degrades the least from sinus to arrhythmia. A rhythm-stratified and point-wise analysis further characterizes the distinctive behavior of each tool, yielding practical guidance for tool selection. These results provide systematic, multi-dataset evidence that self-supervised pretraining is effective for ECG delineation and enables a label-efficiently trained deep learning model to outperform widely used delineation tools by leveraging abundant unlabeled data. This supports adopting such models in diverse, real-world clinical settings.
[CV-77] Revar3r: gauge-aware perturbation uncertainty for feed-forward 3d reconstruction
链接: https://arxiv.org/abs/2610.07883
作者: Sammam Mahdi,Fariha Binta Salim,Rakin Bin Rabbani,Aniqua Nusrat Zereen
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:A correctly reconstructed distant point appears uncertain even when a frozen 3D model processes equivalent inputs because its output frame rotates fractionally. This exposes a weakness of trainingfree perturbation uncertainty: when outputs contain an unobserved symmetry, run-to-run variation potentially reflects symmetry rather than error. Existing alternatives have trade-offs: built-in confidence is outperformed in most evaluated conditions, while trained evidential heads require modelspecific supervision. For point maps, this research derives a closed-form, error-independent variance term that grows with scene extent and potentially overwhelms the desired signal. Simulation reproduces the effect; all 30 real VGGT view-sets tested exhibit its predicted |x_p|^2 signature. ReVar3R robustly registers predictions to a common similarity frame before computing per-point variance, without retraining or modifying the frozen model. Optional calibration and fusion use a held-out split. Across VGGT, \pi3, and MASt3R on six datasets, the same estimator on every backbone lowers AUSE below built-in confidence in 15 of 18 conditions. The staged evaluation yields 11 of 18 wins for the label-free core, 12/18 for label-free equal-weight fusion, 14/18 with held-out weights, and 15/18 when the built-in signal is included. Against a trained evidential head, the result is a trade-off: the head calibrates magnitude better and leads in its training domain, whereas ReVar3R transfers across backbones without adaptation. Its ranking improves point filtering, but it does not detect stable systematic bias, aid novel-view synthesis, or transfer calibration across domains.
[CV-78] CueRator: Agent ic Search for Symbolic Rules to Adapt Frozen Multimodal Encoders
链接: https://arxiv.org/abs/2610.07868
作者: Sunchan Park,Beomkwon Cho,Kyeongbo Kong
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 40 pages, 18 figures
Abstract:Large language model agents have been used to search over symbolic structures such as programs and equations. We propose CueRator, an agentic framework for policy-aware decision-rule discovery, which adapts frozen contrastive multimodal encoders by searching for the decision rule that converts their cross-modal similarities into predictions. We validate it on open-vocabulary audio-visual event perception, where existing methods involve a trade-off between adaptivity and generalization to unseen categories: trained modules adapt at the cost of generalization, and fixed rules the reverse. The framework pairs a symbolic formulation for generalization with a lightweight policy that predicts its parameters per video for adaptivity. A report-guided multi-agent loop discovers the formulation offline, evaluating each candidate on its expressive ceiling and on whether a trained policy can realize it. On OV-AVEBench, CueRator raises the total average from 57.8 to 60.2 and unseen-category performance from 55.8 to 59.9 over the best existing method, reducing the seen-unseen gap from 7.1 to 1.2. Ablations attribute the gains to both the formulation and the policy and show that both feedback signals are necessary for effective search. CueRator also improves over the respective baselines on two further audio-visual event perception tasks, and the discovered rule remains competitive across encoders with only the policy retrained. Code is available at this https URL.
[CV-79] CHARTER: Auditing Reference Substitution in Hierarchical Compact-Evidence Evaluation for Computational Pathology
链接: https://arxiv.org/abs/2610.07843
作者: Hyun Do Jung,Jungwon Choi,Soojung Choi,Yujin Oh,Hwiyoung Kim
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注:
Abstract:In digital pathology, compact evidence is often used to explain or audit predictions made by whole-slide image multiple instance learning models. In hierarchical compact-evidence pipelines, candidate filtering introduces a strategy-specific candidate-conditioned prediction alongside the original full-bag prediction. If the evaluation reference changes while the intended target remains the original full-bag prediction, however, not only can the measured fidelity of the same compact evidence change, but comparisons between competing candidate strategies can also change. To make this dependence explicit, we introduce CHARTER, a reference-aware evaluation charter that asks researchers to DECLARE the intended target and reference, QUANTIFY candidate-induced prediction shift, and AUDIT the stability of comparative conclusions. Across the 15 comparisons in our main five-seed Random-K audit, 4 showed determinate reversals; in a matched native-ranking stress test, the ACMIL comparison changed from REVERSED to PRESERVED. CHARTER turns otherwise implicit candidate-filtering and reference choices into an auditable evaluation specification, helping distinguish genuine preservation of the intended prediction from apparent gains induced by changing the prediction being explained.
[CV-80] owards benchmarking Western Bluebird detection in the wild
链接: https://arxiv.org/abs/2610.07802
作者: Estela Monserrat Arriaga Santana,Julian Rosas Scull,Ibeth P. Alarcón,Bibiana Montoya,Aylin Sosa Mejía,Hugo Jair Escalante
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 15 pages, 4 figures, 10 tables
Abstract:Bird monitoring in natural environments is challenging due to the small size of some species of birds relative to the scene, background clutter, variability in illumination, and the observers’ viewpoint. Progress is further limited by the scarcity of large-scale, realistic datasets, which are essential for understanding behavioral patterns. To address this gap, we introduce a new benchmark dataset for the detection and segmentation of Western bluebirds (Sialia Mexicana), comprising over 6,000 labeled images from 41 recording sessions. The dataset features high-resolution (4K) in-the-wild images in which birds occupy only a small fraction of the image. We evaluated supervised detectors, open-vocabulary models under zero-shot and fine-tuned settings, and segmentation approaches. Supervised detectors remain the most reliable overall, with Faster R-CNN achieving the highest detection mAP and RT-DETR offering the best precision-recall trade-off. Open-vocabulary models perform poorly in zero-shot settings; however, fine-tuning substantially improves their performance, with YOLO-World becoming competitive with supervised methods and achieving the highest precision, F1-score, and mAP@0.5. For segmentation, supervised methods significantly outperform Grounded-SAM and SAM 3: Mask R-CNN achieves the highest mask mAP, while YOLOv8-Seg provides the best precision and fastest inference. A diagnostic analysis further shows that failures are not explained by object size alone, but by a combination of apparent scale, brightness, contrast, clutter, blur, crowding, and recording-session variation. Overall, our findings highlight the difficulty of zero-shot bird detection in cluttered ecological scenes and underscore the importance of domain adaptation in small-object settings.
[CV-81] Efficient Gaussian Splatting Sequence Compression with Standard Video Codecs
链接: https://arxiv.org/abs/2610.07795
作者: Qi Yang,Shuting Xia,Le Yang,Geert Van Der Auwera,Zhu Li
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted by MM Asia 2026
Abstract:This paper presents a novel effective Gaussian Splatting (GS) sequence Compression method that utilizes the Video codec (GSCV). Existing video-based GS sequence compression relies on the Parallel Linear Assignment Sorting (PLAS) and tracked primitive information to convert GS into smooth 2D videos. However, tracked information is not available for most practical applications, and without it, using the vanilla PLAS can generate images exhibiting weak inter-frame correlation, due to its stochastic nature. GSCV incorporates a simple yet efficient Inter-PLAS method to produce close images between the I- and P-frames of GS, enhancing the inter-frame performance of video codec greatly. GSCV also realizes a new pipeline based on the state-of-the-art video codecs with high bit-depth GS images, achieving higher compressibility while simultaneously providing a higher quality upper bound. Experimental results show that the proposed GSCV exhibits obviously improved performance over MPEG video and point cloud-based anchors in GS sequence compression. The code is available at this https URL.
[CV-82] Geometry-Constrained Bidirectional Point Cloud Registration for Thin Sheet-Like Heritage Artifacts
链接: https://arxiv.org/abs/2610.07793
作者: Yuezhe Zhang,Lei Wei,Jingnan Du,Shuai Wan
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 26 pages, 8 figures. Accepted for publication in ACM Journal on Computing and Cultural Heritage
Abstract:Non-contact three-dimensional reconstruction of thin, sheet-like heritage artifacts poses significant geometric and registration challenges. Due to their fragility, these artifacts cannot be suspended or equipped with artificial markers, necessitating independent acquisition of their front and back surfaces. Subsequent registration proves difficult due to the limited number of shared geometric features and the scarcity of explicit physical constraints, which may result in rotational ambiguity, instability, and structural collapse during iterative optimization. To address these challenges, we propose a geometry-constrained bidirectional point cloud registration method specifically tailored for thin, sheet-like heritage artifacts. The method integrates semantic-guided preprocessing, Principal Component Analysis (PCA)-based geometric normalization, and a thickness-aware registration strategy. The estimated physical thickness is incorporated as a geometric constraint to preserve structural integrity during registration. Rotational ambiguity is resolved by evaluating a finite set of global rotation hypotheses, each refined using the point-to-plane Iterative Closest Point (ICP) algorithm, with the optimal transformation selected via a geometry-aware fitness criterion consistent with the thickness scale. Experimental results show that the proposed method achieves competitive or improved performance in most cases, particularly in projected area consistency and physically plausible front-back alignment. In addition, the thickness-aware constraint and rotation hypothesis evaluation reduce the risk of degenerate configurations in which the two surfaces are incorrectly flipped while still yielding deceptively acceptable numerical scores, supporting reliable non-contact digitization of delicate and thin heritage artifacts. Implementation details are available at this https URL.
[CV-83] Image-Space Refraction Correction for Underwater 3D Reconstruction: Warping Flat-Port Views into Pinhole Perspective
链接: https://arxiv.org/abs/2610.07788
作者: Chelim Lim,Tobias Fischer,Emilio Olivastri,Beverley Gorry,Michael Milford,Alejandro Fontan
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Consumer-grade cameras in flat-port housings are widely used for underwater exploration and mapping of coral reefs and seafloor habitats due to their low cost and accessibility. However, refraction at flat-port interfaces causes bowl-shaped deformation in reconstructed scenes and camera trajectories, compromising the metric accuracy required for mapping and navigation. To remove the dominant refractive distortion before reconstruction, we introduce a physics-based refraction correction in image space. Our method is downstream-agnostic: the refraction-corrected images can be directly used as input to existing reconstruction and SLAM algorithms. We characterize the refractive distortion through ray-tracing simulations and validate our correction on two real underwater datasets with differing scene structures. Compared with conventional and refractive Structure-from-Motion (SfM), our approach removes reconstruction deformation while registering more frames and maintaining low reprojection error. The correction further generalizes across diverse reconstruction and VSLAM backends, demonstrating its broad applicability to downstream vision pipelines.
[CV-84] From Laboratory to Road: Evaluating Wearable Gaze Accuracy for Driving
链接: https://arxiv.org/abs/2610.07783
作者: William Engel,Fabian Flohr
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Peer-reviewed and accepted as an Extended Abstract at the German Conference on Pattern Recognition (GCPR 2026). Presented as a poster at GCPR 2026
Abstract:Bird’s-eye-view (BEV) representations have become a widely used interface between perception and planning in autonomous driving, but they encode what is in a scene, not what is behaviorally relevant to a human driver. Gaze offers a compelling behavioral signal for this gap, yet wearable eye trackers are routinely deployed as if their spatial output were ground truth, despite known sensitivity to head motion, illumination, and calibration drift. We present, to our knowledge, the first unified framework for quantifying wearable gaze accuracy under real driving conditions. Our on-road study contains 41 validated scenes in which one driver fixated a vehicle’s license plate. Gaze error is measured as the angular difference between the plate center and the gaze direction estimated by the glasses. Separate indoor studies with the same driver and device systematically analyze how distance, illumination, head motion, target motion, and gaze eccentricity affect both systematic bias and gaze precision. The mean on-road error was 4.58 degrees. Applying an offset estimated from the indoor recordings reduced it to 1.10 degrees and improved all 41 scenes. Because this offset varied between sessions, reliable BEV supervision may require online recalibration and condition-dependent estimates of gaze uncertainty.
[CV-85] Later Is Better: Token Reduction for ViTs Under Distribution Shift
链接: https://arxiv.org/abs/2610.07758
作者: Hyeongheon Cha,Hyungjun Yoon,Sung-Ju Lee
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注: 35 pages. Code: this https URL
Abstract:Training-free token reduction accelerates vision transformers by removing redundant tokens across layers, recovering most of the original accuracy at a fraction of the compute. These methods, however, are designed and evaluated primarily on clean data, and under real-world distribution shift their accuracy gap to the uncompressed model widens with the removal rate. We show that this gap is governed by the reduction schedule, the depth profile of removal, usually left fixed as an implementation detail. Concretely, we introduce a one-parameter late-concentrated power-law schedule that consistently improves out-of-distribution accuracy over flat at no extra inference cost. On ImageNet-C with DeiT-S, the late schedule closes 83% of that gap at a 26% compute reduction (+1.17pp), and 99% of it at a lighter 7% reduction (+0.26pp). The gain cannot be attributed to retaining more tokens or using extra compute: held to flat’s compute, the late schedule removes more tokens in total and leaves fewer tokens at the end, yet still wins. Single-layer probes point to a mechanism: earlier reductions perturb features that pass through more remaining layers, front-loading reduction error in depth. The effect is broad, holding across five token-reduction methods (ToMe, EViT, ATS, ATC, PiToMe), nine backbones, all ImageNet-C corruption types, eight further shift suites, and two further modalities, video and vision-language QA. It is also specific to shift, still positive on clean and rising monotonically to ~4x that at the highest severity 5. The schedule keeps its gain under six test-time adaptation methods, and needs no per-input or per-domain tuning.
[CV-86] Adversarially Trained Linear Transformers Are Optimal Robust In-Context Learners for Gaussian Mixtures
链接: https://arxiv.org/abs/2610.07754
作者: Soichiro Kumano
类目: Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (stat.ML)
备注:
Abstract:Adversarial training is one of the most reliable defenses against adversarial attacks, but its high computational cost must generally be paid anew for each task. Robust foundation models offer a promising alternative: adversarially pretrain a model once and then transfer its robustness to downstream tasks through lightweight adaptation. However, a fundamental question remains open: can robustness acquired during pretraining transfer to unseen tasks without further adversarial training? In this study, we answer this question affirmatively. A single model adversarially pretrained at scale can achieve optimal robustness on new tasks without additional task-specific training. Specifically, we show that, for a family of Gaussian-mixture classification tasks, a sufficiently deep linear transformer adversarially trained across tasks can asymptotically attain the robust Bayes error on previously unseen tasks through in-context learning from clean demonstrations. By contrast, a standardly trained model cannot. We further analyze convergence under gradient flow, an accuracy–robustness trade-off, and demonstration complexity.
[CV-87] Foveated Compression: Selective High-Resolution Preservation for Token-Efficient VLMs
链接: https://arxiv.org/abs/2610.07729
作者: Donghyun Han,Jangho Park,Yuseok Bae
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Visual tokens are a major source of inference cost in vision-language models, yet simple image downsampling remains a surprisingly strong compression baseline. This raises a complementary question: under a fixed token budget, where should visual fidelity be preserved? We introduce Foveated Compression, which encodes a full-resolution image once and represents it with a mixture of native- and compressed-resolution visual tokens. A behaviorally self-distilled Foveated Merger compresses local visual tokens while preserving compatibility with their native counterparts, and a lightweight Foveated Selector chooses one of nine spatial cells to retain at native resolution using exhaustive budget-matched intervention supervision. At 11.11% visual tokens, uniform Foveated Compression shows no significant paired difference from iso-token downsampling. At 20.99%, the learned selector significantly outperforms random and fixed allocation, but remains below strong whole-image resizing, showing that localized fidelity is not universally preferable. A budget-matched region-choice oracle reaches 82.73 macro accuracy versus 69.61 for the learned selector, revealing substantial headroom within the same spatial action space. Matched probing further shows that signals predicting when compression breaks the answer are substantially more accessible after language-model computation than to the lightweight prefill-free selector. These results expose complementary bottlenecks in region selection and compressed-region fidelity.
[CV-88] Structure-aware Keypoint Localization for Videofluoroscopic Swallowing Study ICME2026
链接: https://arxiv.org/abs/2610.07726
作者: Kai Zhou,Chuanshen Chen,Runhao Zeng,Meng Dai,Yifan Yang,Jinwu Hu,Daiyuan Li,Mingkui Tan,Fei Liu
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted by ICME 2026 Oral
Abstract:Videofluoroscopic Swallowing Study (VFSS) is one of the gold standard for diagnosing swallowing disorders, providing dynamic X-ray imaging of the swallowing process. Automated kinematic analysis in VFSS relies fundamentally on precise anatomical keypoint localization. However, existing studies focus on limited keypoints (e.g., cervical vertebrae or the hyoid) and overlook critical regions such as the soft palate, while annotating only active swallowing segments and ignoring abundant non-swallowing data, resulting in poor data efficiency. Moreover, leveraging this unlabeled data via standard semi-supervised learning is suboptimal, as generic methods are prone to spatial bias. In medical X-rays with fixed layouts, models tend to memorize absolute coordinates rather than understanding anatomical structures. To tackle these challenges, we introduce VFSSKep, a novel dataset that extends annotations to the soft palate and incorporates large-scale unlabeled data. We further propose S ^3 KL, a Structure-aware Semi-Supervised Keypoint Localization framework designed to overcome spatial bias. It integrates a Structure-Aware Learning strategy to extract high-resolution structural cues for structure-aware representation learning, and a Structural Representation Consistency Learning strategy with block shuffling to enforce invariant structural recognition. Experiments show our method achieves state-of-the-art semi-supervised performance, even with unlabeled and 25% labeled data surpassing fully supervised learning with 100% labeled data. Code and data will be made publicly available at: this https URL.
[CV-89] RefRoute: Decoupling Conditioning Cost from References via Compact Residual Conditioning and Spatial Routing
链接: https://arxiv.org/abs/2610.07720
作者: Wanning He,Yuyao Zhang,Yu-Wing Tai
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注: 19 pages. Wanning He and Yuyao Zhang contributed equally and share first authorship
Abstract:Multi-reference image generation requires preserving the appearance of multiple subjects while composing them into a coherent scene. However, existing diffusion transformers commonly encode references as dense visual token grids and jointly process them with global attention, making conditioning increasingly expensive as the number and resolution of references grow. We present RefRoute, a framework that addresses both reference representation cost and attention overhead through two complementary mechanisms. Compact residual conditioning combines low-resolution latent tokens with lightweight residual features extracted from full-resolution pixels, reducing reference token counts while retaining fine-grained appearance cues. Condition routing and attention routing align reference tokens with their assigned target regions and restrict cross-reference interactions, while allowing selective reference access beyond region boundaries for scene integration. We further introduce RefRoute-Data for training many-reference generation models and ManyRef100, a benchmark spanning human, object, and mixed compositions with 10-17 references. After many-reference fine-tuning, RefRoute achieves an overall Weighted-Ref-VIEScore of 36.06 on ManyRef100, compared with 8.88 for FLUX.2-Klein-9B. Separate inference-cost evaluations show substantially slower latency growth as the reference count increases: at 16 references, our 50-step and 4-step configurations achieve 18.3\times and 14.2\times speedups over their corresponding FLUX baselines, respectively. These results establish compact reference representations and spatially routed attention as an effective approach to scalable many-reference image generation.
[CV-90] Comprehensive Evaluation and Fine-Tuning of Foundational Cell Nuclei Segmentation Models in Renal Pathology
链接: https://arxiv.org/abs/2610.07711
作者: Ruijie Wu,Junlin Guo,Ruining Deng,Yu Wang,Shilin Zhao,Haichun Yang,Yuankai Huo
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 11 pages, 5 figures, 4 tables. Submitted to SPIE Medical Imaging 2027
Abstract:Accurate nuclei instance segmentation is essential for quantitative renal pathology, yet general-purpose models often struggle with low contrast, dense nuclei, complex morphology, and strong background staining. In this work, we extended a human-in-the-loop framework by combining 5,901 foundation-model-generated pseudo-labels from well-segmented cases (Easy), 860 newly expert-annotated unresolved challenging cases (Medium), and 198 expert-annotated consensus failure cases (Hard). These annotations, spanning different levels of segmentation difficulty, enabled the systematic evaluation of seven single-source and mixed-source fine-tuning strategies across nine cell segmentation model configurations. Fine-tuning improved all models, with Medium data included in seven of the nine best-performing strategies. LSP-DETR achieved the highest F1 score of 0.8725 with Hard-only fine-tuning, while StarDist showed the largest improvement, increasing from 0.7380 to 0.8332 with Medium-only fine-tuning. These findings show that annotations spanning multiple difficulty levels support effective model adaptation, although the optimal annotation composition remains model dependent.
[CV-91] What Frame-Level Labels Can and Cannot Do for Small-UAV Point Detection in Thermal Video
链接: https://arxiv.org/abs/2610.07705
作者: Wonbin Son,Gyumum Choi,Junil Seo,Hyungjoon Kim
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:
Abstract:The growing use of unmanned aerial vehicles (UAVs) has increased the importance of image-based UAV detection. Learning-based detectors are trained on imagery and annotations, with annotation type determining the information available during training. We focus on learning localization from frame-level target presence/absence labels when sensor or scene changes make spatial annotations for additional training burdensome. We analyze the detection capability, learning behavior, and potential applications of an existing architecture for point detection of small UAVs, trained with presence/absence labels and requiring no external detector. The architecture freezes spatial features learned through classification and trains a readout with the same frame labels to produce spatial score maps and point detections. On two thermal infrared datasets, CST Anti-UAV and Anti-UAV410, we evaluate localization hit rates and detection rates under false-alarm constraints, analyze the effects of training stages, label allocation, synthesis, and model configuration, and compare with bounding-box detectors. We also explore potential applications on Airborne Object Tracking (AOT) using its visible-light imagery and frame labels. Classification training strengthened target-related spatial responses, while readout training helped extract them consistently. Distributing similar label counts across more videos yielded higher localization hit rates, while synthesis effects varied by dataset and evaluation criterion. Higher localization hit rates did not always improve detection under false-alarm constraints, and failures remained when target signals were weak relative to background variation and under cross-dataset transfer. These findings provide guidance on label allocation, spatial representations and readouts, synthesis, and false-alarm control.
[CV-92] RBMatch: Dual-Level Class Rebalancing for Semi-Supervised Building Footprint Extraction
链接: https://arxiv.org/abs/2610.07698
作者: Akil Ahmad Taki,Shaikh Anowarul Fattah
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Accurate building footprint extraction from high-resolution remote sensing imagery is essential for urban planning, disaster response, and environmental monitoring. However, obtaining dense pixel-level annotations is costly, motivating the use of semi-supervised learning (SSL) to leverage unlabeled imagery. In remote sensing, severe foreground–background imbalance poses a particular challenge for self-training, as it can bias pseudo-label generation and the resulting unsupervised optimization toward the majority background class. We show that addressing this imbalance at only one stage is insufficient: balancing pseudo-label selection alone does not prevent background bias from re-emerging during unsupervised loss optimization, a failure mode we term \emphimbalance leak. To address this issue, we propose \textbfRBMatch, a dual-level class-rebalancing framework that jointly regulates pseudo-label generation and unsupervised optimization. RBMatch combines a supervised learning pathway with a self-training module comprising three components: adaptive class-specific thresholding (ACT) for balanced pseudo-label selection, confidence-aware class-balanced reweighting (CACBR) for mitigating class bias in the unsupervised loss, and distribution alignment (DAL) for matching the predicted unlabeled-data distribution to the labeled-data prior. Experiments on the WHU, INRIA, and Massachusetts building footprint datasets across labeled ratios of 1%–10% show that RBMatch consistently achieves the best building IoU and F1-score among the evaluated methods. The improvement is most pronounced on the highly imbalanced Massachusetts dataset, where RBMatch improves IoU by 1.37 points over the strongest baseline at a 1% labeling ratio and is the only method to outperform the fully supervised baseline across all twelve dataset–ratio settings.
[CV-93] Anchor-driven Multi-modal Multi-scale Expert Selection for Survival Prediction
链接: https://arxiv.org/abs/2610.07694
作者: Tao Zhou,Ying Hu,Huazhu Fu,Yi Zhou,Xiao-Jun Wu,Haibin Ling
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 15 pages, 6 figures, 7 tables
Abstract:The integrative analysis of histopathological Whole-Slide Images (WSIs) and transcriptomic profiles holds significant promise for cancer survival prediction. However, existing methods typically project multi-modal features directly into a shared latent space without explicit alignment, leading to the entanglement of mismatched morphological cues and molecular signals. Furthermore, current fusion strategies often treat the extreme spatial heterogeneity of WSIs uniformly, lacking mechanisms to adaptively prioritize clinically relevant tissue scales for individual patients. To address these limitations, we propose an Anchor-driven Multi-modal Multi-scale Expert Selection (AM ^2 ES) framework for survival prediction. Specifically, we present an Anchor-driven Multi-modal Fusion (AMF) module, which introduces learnable semantic anchors as cross-modal mediators to bridge the semantic gap by enforcing a structurally regularized alignment between transcriptomic features and multi-scale pathology representations. Built upon this aligned semantic space, we further design a Hierarchical Mixture-of-Experts (H-MoE) selection module to decouple the hierarchical prognostic selection process. Mimicking the pathologist’s diagnostic workflow, H-MoE performs (i) Intra-scale Expert Filtering to discriminatively identify salient tumor regions within each magnification, and (ii) Inter-scale Hierarchy Routing to dynamically weight and select the most informative resolution levels. Extensive experiments on multiple TCGA cancer cohorts demonstrate that our AM ^2 ES achieves state-of-the-art performance while offering fine-grained interpretability by visualizing how specific molecular pathways drive the expert routing decisions across tissue scales. The code will be released at this https URL.
[CV-94] Unlocking Fine-Grained Perception in CLIP via Structurally-Aware Latent Masked Modeling
链接: https://arxiv.org/abs/2610.07689
作者: Juntong Li,Lingwei Dang,Haomin Wu,Ziyan Qiu,Qingxin Xiao,Qingyao Wu
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Vision-Language Models (VLMs) such as CLIP excel in global semantic alignment but often lack fine-grained perceptual capabilities. This hinders dense prediction tasks and bottlenecks the visual potential of Multimodal Large Language Models (MLLMs). Existing research has attempted to enhance CLIP’s visual representations by incorporating geometric priors from vision-centric models. However, these strategies often struggle to achieve deep alignment for both local spatial structures and global semantics, potentially even distorting the original image-text space. To address these limitations, we propose SALM, an unsupervised embedding alignment framework based on structurally-aware latent mask modeling. SALM effectively synergizes local and global alignment via a dual-path design combining explicit and implicit mechanisms, without requiring any image-text pairs. First, we introduce a dual-matrix alignment strategy that explicitly calibrates intra-sample spatial correlations and activation intensities, thereby effectively injecting local geometric priors. Based on this, we further design a latent mask modeling mechanism to guide CLIP to restore the missing semantic details of the target model, thereby implicitly aggregating fine-grained structures into the global semantic space. Furthermore, driven by the empirical observations that CLIP’s shallow features inherently possess strong spatial observational capabilities, we naturally extend SALM to a highly efficient self-distillation paradigm, SALM-Self. This unlocks CLIP’s intrinsic fine-grained potential without relying on any external models. Extensive experiments demonstrate that SALM not only significantly improves performance in dense prediction tasks but also boosts CLIP’s zero-shot accuracy, effectively enhancing the fine-grained understanding capabilities of MLLMs. Project page at this https URL.
[CV-95] Disentangling Dual Image References in Frequency Aware Diffusion Models for Personalized Generation NEURIPS2026
链接: https://arxiv.org/abs/2610.07684
作者: Haipeng Liu,Yang Wang,Meng Wang
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注: 28 pages, 14 figures, to appear at NeurIPS 2026, Sydney, Australia
Abstract:Personalized image generation aims to synthesize text-driven images conditioned on reference images, while mainly casting the generation as image customization for foreground and style transfer for background. Previous arts of diffusion models suffers from the text misalignment with background for image customization and foreground for style transfer during the denoising process. Such facts, as we observed, rooted from the entanglement among hybrid frequency bands during the denoising process. To address such salient limitation, in this paper, we study personalized generation based on dual references - customization and color and style reference - and propose a paradigm to disentangle these Dual image references within Frequency-aware Diffusion Models, dubbed Dual-FDM, to simultaneously tackle two crucial personalized image generation tasks: customization style transfer and color style transfer, by disentangling different frequency bands via mask strategy within frequency domain. For customization style transfer, we replace the mid-frequency band of the background in the style reference with that from the foreground of the customized reference. For color style transfer, we substitute the low-frequency band of the background in the style reference with that from both the foreground and background of the color reference. Both the substituted frequency bands are used as the key and value to reconstruct the query foreground and background of the denoised personalized this http URL experiments validate the superiority of Dual-FDM over the state-of-the-art diffusion models for personalized image generation. Our code can be accessed from this https URL.
[CV-96] PhysLDM: Latent Diffusion for High-Fidelity Deformable Simulation
链接: https://arxiv.org/abs/2610.07609
作者: Yu Zhang,Xudong Xu,Xingang Pan
类目: Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR)
备注:
Abstract:Neural simulation of high-fidelity deformable bodies is a foundational challenge in computer graphics and physical AI. Long-horizon prediction for high-resolution 3D volumetric meshes is hard: autoregressive methods are susceptible to error accumulation, while direct multi-frame prediction at native resolution is computationally prohibitive. This motivates a compact spatiotemporal latent representation, which is largely unexplored for mesh-based volumetric physics. Meanwhile, it remains unclear whether deterministic regression or generative diffusion is the more appropriate predictive paradigm. To address these coupled challenges, we introduce PhysLDM, a unified latent-diffusion paradigm for one-shot volumetric deformable simulation. Its core is a holistic spatiotemporal VAE that avoids the “staircase” artifacts of standard temporal compression (as in common video VAEs), achieving ~2.48 mm reconstruction precision on meter-scale scenes at up to 78x token compression. Based on this reliable latent space, we systematically compare regression and diffusion methods. Our experiments uncover a key modeling insight: complex deformable dynamics are often chaotic, and in this regime deterministic regression tends to produce non-physical averages, whereas diffusion better models their distribution. Accordingly, we employ a latent diffusion model that effectively learns from the chaotic data to generate physically plausible trajectories. Trained purely kinematically on an Objaverse-scale dataset, a single PhysLDM generalizes zero-shot to unseen OOD datasets (GSO and Toys4K). Its differentiability further enables efficient solution of inverse problems and higher-order design optimization. To our knowledge, PhysLDM is the first high-fidelity spatiotemporal autoencoder and latent-diffusion paradigm for volumetric deformable dynamics, offering a scalable and robust approach to neural simulation.
[CV-97] REViT-v2: Hierarchical Windowed Roto-reflection Equivariant ViT for Equivariant Feature Extraction NEURIPS
链接: https://arxiv.org/abs/2610.07585
作者: Sheir A. Zaheer,Jihwan Moon,Chan Y. Park
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注: 7 pages, Accepted for presentation at NeurIPS NeurREPS workshop 2026
Abstract:We propose a scalable roto-reflection-group-equivariant vision transformer based on windowed group-convolutional self-attention and a hierarchical feature architecture. We demonstrate that our approach can be scaled to group-equivariant vision transformers (ViTs) with millions of parameters and large datasets with practically sized images, i.e., ImageNet. The code and pretrained weights for the proposed Hierarchical Windowed Roto-reflection Equivariant ViTs (REViT-v2) are available at this https URL.
[CV-98] CETUS: How Far Do Representations Trained on Earth Transfer to Cassini SAR of Titan?
链接: https://arxiv.org/abs/2610.07576
作者: Kevin Lee
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Image and Video Processing (eess.IV)
备注: Research work at NASA Jet Propulsion Laboratory. Available at: this https URL
Abstract:Cassini synthetic aperture radar (SAR) images reveal the dunes, plains, and lake basins of Titan, providing an instance of representations learned from Earth imagery for planetary terrain classification. Cross-domain Evaluation of Earth-to-Titan Transfer Using SAR (CETUS) compares features from DINOv2, DOFA and CROMA with classical image measurements and features from an untrained vision transformer on the U.S. Geological Survey’s Cassini SAR mosaic. The classifiers learn terrain labels from an expert geomorphological map and predict those labels in geographically separate Titan regions. Under logistic regression settings, pretrained encoders achieve higher mean macro F1 than the combined classical features. Encoder rankings change when feature scaling, optimization, and regularization change together. Further training on Titan improves DINOv2 performance, degrades DOFA performance, and leads to mixed results for CROMA under the tested settings. Architectural and input processing differences prevent these comparisons from isolating the effect of pretraining. Classifier fitting and performance on individual terrain classes matter when assessing representation transfer for planetary mapping. Since the map draws partly on the same radar observations, the scores measure agreement with expert interpretation.
[CV-99] OpenSplatGraph: From Dense Semantic Maps to Structured Scene Graphs for Open-Vocabulary Robot Perception ACCV2026
链接: https://arxiv.org/abs/2610.07569
作者: Binh Long Nguyen,Kien Nguyen,Clinton Fookes,Peyman Moghadam
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted to ACCV 2026
Abstract:Dense 3D mapping with semantic understanding is essential for robotic perception in complex environments. Recent 3D Gaussian Splatting-based mapping approaches enable high-fidelity geometry and efficient open-vocabulary perception, but typically represent semantics as unstructured feature fields that limit object-centric reasoning. In contrast, 3D scene graphs explicitly model objects and their relationships for structured reasoning, but are commonly constructed from sparse geometric representations that do not fully exploit dense semantic maps. In this work, we present OpenSplatGraph, a unified framework that constructs persistent 3D scene graphs directly from an online Gaussian-based open-vocabulary semantic map. The proposed framework augments the dense semantic map with a reliability-aware semantic field that maintains lightweight observation statistics for confidence-aware, query-conditioned object extraction. Extracted object instances are associated with persistent graph nodes, allowing object attributes and relationships to be incrementally updated across observations and queries. By tightly coupling dense semantic mapping with persistent object-centric representations, our framework supports both language-guided object grounding and structured relational reasoning while preserving the geometric fidelity of Gaussian-based mapping. Comprehensive evaluations on standard 3D scene understanding benchmarks and real-world robotic experiments demonstrate that OpenSplatGraph achieves competitive performance for online open-vocabulary perception and downstream robotic tasks. Project page: https://csiro-robotics.github.io/OpenSplatGraph.
[CV-100] AIMS: Anchor-Integrated Multi-View Synthesis for Scalable Novel View Rendering
链接: https://arxiv.org/abs/2610.07566
作者: JooHyun Park,HanYoung Jang,HyeongYeop Kang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Feed-forward novel view synthesis methods achieve strong generalization from posed multi-view inputs, but scaling them to large input view sets remains challenging. Transformer-based approaches that jointly process all input-view tokens incur rapidly increasing computation and memory as the number of views grows, while simple view subsampling discards potentially useful observations. We introduce Anchor-Integrated Multi-View Synthesis (AIMS), a scalable framework that decouples the number of available observations from the number of views processed by the global synthesis model. AIMS selects a fixed set of spatially distributed anchor views using farthest point sampling, groups nearby observations around each anchor, and uses a lightweight learnable integrator to fuse their information into enriched anchor representations. This allows additional observations to contribute to synthesis while keeping the downstream global view budget fixed. Evaluations on RealEstate10K and ScanNet demonstrate a favorable quality–efficiency trade-off against transformer-based and Gaussian-based baselines. AIMS achieves 29.41 dB and 17.73 dB PSNR on the two datasets, respectively, with rendering averaging 7.24 ms per view.
[CV-101] LARK: A Low-Cost Accurate Occlusion-Resilient Kalman Filter-Assisted Tracking System for Image-Guided Surgery
链接: https://arxiv.org/abs/2610.07561
作者: George Sideris,Justin Cree,Andrew Stirling,Mamadou Ly,Étienne Léger,D. Louis Collins
类目: Computer Vision and Pattern Recognition (cs.CV); Image and Video Processing (eess.IV)
备注: 24 pages, 15 figures, including appendices. Supplementary document included as an ancillary file. Supplementary video: this https URL
Abstract:Image-guided surgery (IGS) depends on accurate tracking of surgical instruments to provide real-time navigation relative to anatomical structures. Commercial stereo infrared trackers are accurate but prone to occlusion and cost-prohibitive for many settings. This work presents LARK, a multi-camera optical tracking system using commodity RGB hardware and multi-view redundancy and fusion. We develop and evaluate two complete tracking methods: multi-view monocular pose fusion and multi-view triangulation. Both methods are assessed under varying occlusion levels using a precision-machined grid and an anatomical head phantom, and compared against a gold-standard stereo infrared system. With five cameras and adaptive Kalman filtering, LARK achieves median target registration errors of 0.64 mm for point localization with triangulation and 0.73 mm for trajectory tracking with pose fusion on the machined grid. Camera-subset experiments show graceful degradation in adaptive pose-fusion accuracy as fewer views remain available. With tracking hardware costing under 1,000 USD, LARK provides a low-cost platform for image-guided surgery research. Hardware designs and software are publicly available at this https URL , and datasets at this https URL .
[CV-102] MobileVISTA: Generative Data Augmentation for Pose Generalization in Mobile Manipulation
链接: https://arxiv.org/abs/2610.07511
作者: Suzannah Wistreich,Stephen Tian,Isabella Huang,Vitor Campagnolo Guizilini,Sergey Zakharov,Katherine Liu,Jiajun Wu
类目: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注:
Abstract:Mobile manipulators such as humanoid robots are increasingly deployed in dynamic, unstructured environments to perform dexterous manipulation tasks. However, end-to-end manipulation policies trained to imitate demonstration data collected from a single robot pose are brittle: even centimeter-scale deviations in robot pose at deployment can drive ego-centric observations and end-effector trajectories out of the training distribution, leading to sharp drops in performance. We introduce MobileVISTA, a data generation framework that transforms demonstrations captured at canonical poses into diverse, pose-perturbed training data by jointly (1) augmenting egocentric visual observations and (2) retargeting actions to compensate for base pose changes. Unlike prior methods, which assume a camera rigidly mounted off the actuated chain or non-trivial articulated robot geometry largely out of frame, MobileVISTA targets compatibility with egocentric platforms (e.g., humanoids) where the camera is both influenced by and must observe the robot’s kinematic chain as it moves. We study MobileVISTA in simulated tasks spanning humanoid and bimanual embodiments, and on a real Galaxea R1 Pro. We find policies trained on MobileVISTA-augmented data demonstrate improved robustness to previously out-of-distribution poses encountered at test time, without additional demonstration collection or a trained generative model. Additionally, we find MobileVISTA’s benefit is largest on tested humanoids, where the camera rides the actuated chain and the robot fills much of the frame. Additional videos and appendix can be found on our website: this https URL
[CV-103] Protective Perturbations Must Survive the Resize: Scale-Robust Image Immunization against Malicious Editing
链接: https://arxiv.org/abs/2610.07464
作者: Zhongliang Guo,Yan Lin,Yifei Qian,Daizong Liu
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 16 pages, 9 figures
Abstract:Protective perturbations aim to stop malicious instruction-guided editing of personal photos, but they are optimized and evaluated at the editor’s working resolution, whereas shared photos have 10 megapixels or more and editors first downscale them by an unknown factor. We model this resize as a frequency-selective channel. In this model, a perturbation computed at the native resolution decays with the downscaling factor and is weak even without a resize, and a perturbation computed at a fixed working resolution protects only a window of scales. The best worst-case protection over an unknown range of scales degrades only logarithmically with the width of the range, and averaging over scales does not reach it. Guided by this analysis, we propose SRIM, which samples a grid of anchor scales covering the whole range, with weights that favor the currently weakest scale, at the cost of standard expectation over transformation. On full-resolution photos of 9 to 30 megapixels and downscaling factors from 2 to 8, SRIM raises the worst-case disruption of FLUX.2-klein edits from 0.192 LPIPS, attained by the strongest published protection, to 0.463. At equal visibility, it roughly doubles the protection. The same protected photos also protect against the 9B model and against FLUX.2-dev, with worst cases of 0.450 and 0.386 against at most 0.184 for published protections, and SRIM leads on InstructPix2Pix as well.
[CV-104] ElasticFit: Fit-Aware 3D Object Insertion via VLM Reasoning and Generative Adaptation NEURIPS2026
链接: https://arxiv.org/abs/2610.07460
作者: Tzu-Hsin Hsieh,Ricardo Marroquim
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注: Accepted at NeurIPS 2026. Project page: this https URL
Abstract:Inserting objects into existing 3D scenes requires more than selecting a plausible location: the inserted object must also fit local geometry while preserving semantic intent and physical plausibility. Although recent Vision-Language Models (VLMs) and generative models enable semantic reasoning and visual content creation, they offer limited 3D grounding and geometric control when an inserted object must fit into constrained local spaces. We introduce \textbfElasticFit, a VLM-guided framework for fit-aware object insertion centered on a novel scene-grounded representation. Given a language instruction and rendered scene observations, ElasticFit infers structured fitting cues that specify where the object should be grounded, what volume it should occupy, how it should be oriented, and its adaptation mode (rigid placement, uniform scaling, or elastic fitting). These cues convert high-level VLM reasoning into explicit 3D constraints that condition object generation and guide downstream geometric fitting. ElasticFit then generates a scene-conditioned object prior, reconstructs it in 3D, and refines the mesh through mode-specific fitting while enforcing collision avoidance, contact consistency, and physical grounding. In fixed-asset baseline comparisons, ElasticFit improves spatial relation success from 50.8% to 69.7% and support success from 48.3% to 91.7% over the strongest baseline, while providing novel support for generative “make-it-fit” insertions in complex scenarios. Comments: Accepted at NeurIPS 2026. Project page: this https URL Subjects: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI) Cite as: arXiv:2610.07460 [cs.CV] (or arXiv:2610.07460v1 [cs.CV] for this version) https://doi.org/10.48550/arXiv.2610.07460 Focus to learn more arXiv-issued DOI via DataCite (pending registration) Submission history From: Tzu Hsin Hsieh [view email] [v1] Mon, 5 Oct 2026 22:12:18 UTC (47,540 KB) Full-text links: Access Paper: View a PDF of the paper titled ElasticFit: Fit-Aware 3D Object Insertion via VLM Reasoning and Generative Adaptation, by Tzu-Hsin Hsieh and 1 other authorsView PDFHTML (experimental)TeX Source view license Additional Features Audio Summary Current browse context: cs.CV prev | next new | recent | 2026-10 Change to browse by: cs cs.AI References Citations NASA ADSGoogle Scholar Semantic Scholar export BibTeX citation Loading… BibTeX formatted citation loading… Data provided by: Bookmark checked="checked"class=“labs-tab-input”> Bibliographic Tools Bibliographic and Citation Tools Bibliographic Explorer Toggle Bibliographic Explorer (What is the Explorer?) Connected Papers Toggle Connected Papers (What is Connected Papers?) Litmaps Toggle Litmaps (What is Litmaps?) scite.ai Toggle scite Smart Citations (What are Smart Citations?) Code, Data, Media Code, Data and Media Associated with this Article alphaXiv Toggle alphaXiv (What is alphaXiv?) Links to Code Toggle CatalyzeX Code Finder for Papers (What is CatalyzeX?) DagsHub Toggle DagsHub (What is DagsHub?) GotitPub Toggle Gotit.pub (What is GotitPub?) Huggingface Toggle Hugging Face (What is Huggingface?) ScienceCast Toggle ScienceCast (What is ScienceCast?) Demos Demos Replicate Toggle Replicate (What is Replicate?) Spaces Toggle Hugging Face Spaces (What is Spaces?) Spaces Toggle TXYZ.AI (What is TXYZ.AI?) Related Papers Recommenders and Search Tools Link to Influence Flower Influence Flower (What are Influence Flowers?) Core recommender toggle CORE Recommender (What is CORE?) Author Venue Institution Topic About arXivLabs arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv’s community? Learn more about arXivLabs. Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?) mathjaxToggle(); We gratefully acknowledge support from our major funders, member institutions, , and all contributors. About Help Contact Subscribe Copyright Privacy Accessibility Operational Status (opens in new tab) Major funding support from
[CV-105] Decoupling What from Where: How Should a Small GUI Grounding Model Receive the Action Type?
链接: https://arxiv.org/abs/2610.07444
作者: Aadi Chauhan,Arthur Ilyasov
类目: Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)
备注: 18 pages, 4 figures, 12 tables. Code and per-example logs are available at this https URL
Abstract:A GUI agent decides which action to take and where to take it; we ask how a small grounding model should receive the action type. Fine-tuning Qwen2-VL-2B with LoRA on Android in the Wild, we compare a flat baseline with five ways of supplying the type under matched data, compute, and decoding: an auxiliary loss, a hard-routed action word, an additive learned embedding, a prepended learned token, and the type written into the prompt. With five seeds, an episode-clustered bootstrap, and seed-level paired tests, the ranking on a mixed stream is clear: the auxiliary loss, the additive embedding, and the prompt word each gain five to seven hit@0.10 points over the baseline, while hard routing and the prepended token are not distinguishable from it. Much of that gain is protection from a preprocessing choice of ours rather than a spatial prior. Our serializer clamps the off-screen touch point AITW records for type events to the origin; that class degrades the baseline’s click grounding, and removing it lifts the baseline by nearly seven points, after which no mechanism’s hit rate beats it and the intervals exclude a two-point effect, though the auxiliary loss still shortens the average miss; on a stream of taps and swipes none helps. Whether this generalizes beyond one serialization is open. For deployment, the pipeline’s margin over the baseline with predicted rather than gold types is not established (+0.016, 95% interval [-0.017, +0.052]), and a wrong type collapses every model conditioned at inference. The prepended token does not help at the shared learning rate, where its rows barely move from initialization; trained ten times faster it reaches the level of the other three, with a margin three seeds do not establish. We also document a silent failure: injecting conditioning through inputs_embeds makes Qwen2-VL fall back to 1-D positions for image tokens, costing nine points.
[CV-106] Learnable Spectral Activations
链接: https://arxiv.org/abs/2610.07419
作者: Tamir Shor,Or Litany,Alex Bronstein
类目: Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Implicit neural representations (INRs) are shaped by the spectral structure induced by their input encodings and activation functions. Existing methods improve fitting primarily by modifying which frequencies are available to the network, through coordinate encodings or periodic nonlinearities. However, frequency access is not the only bottleneck: signals with localized or spatially varying structure require the network to efficiently compose frequencies into multi-harmonic internal responses. We introduce learnable spectral activations (LSA), which replace fixed neuron-level nonlinearities with a residual truncated Fourier series whose harmonic amplitudes are learned during training. LSA does not expand the asymptotic function class. Instead, it changes the factorization of the representation: linear weights select features while activation coefficients control spectral shaping, and the two are updated by separate gradients. Because the activation output is affine in the coefficients given fixed pre-activations, spectral tuning becomes a more direct subproblem compared to architectures where it is entangled with feature selection. Empirically, this factorization concentrates more target-signal energy in the leading eigenmodes of the neural tangent kernel, consistent with improved optimization behavior. Across audio, image, neural radiance field, and neural acoustic field tasks, LSA also improves reconstruction quality.
[CV-107] Semantic Capability Acquisition and Specialization During Vision-Language Model Fine-Tuning
链接: https://arxiv.org/abs/2610.07385
作者: Suguru Onda,Matthew Bailey,Ryan Farrell
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 67 pages, including supplementary material
Abstract:Fine-tuning vision-language models (VLMs) is typically evaluated at a single downstream checkpoint, obscuring whether a semantic capability was never acquired or emerged earlier and later declined during specialization. We ask how semantic capabilities are acquired, when they peak, how well they transfer, and what remains at deployment. We study these dynamics as a semantic capability trajectory, tracking identity- and attribute-based capabilities over training. We formulate a trajectory-based framework that separates capability acquisition, capability-specific optima, and later specialization, and introduce Structured Semantic Routing (SSR) to study how the representation of supervision shapes what is acquired. Across six pretrained backbones spanning DFN, MetaCLIP, and OpenAI CLIP, we show that fine-tuning can acquire semantic capability beyond the pretrained state, including gains observed on held-out evaluations. Unstructured name-and-attribute supervision produces strong name-and-attribute retrieval with comparatively weak name-free attribute-profile retrieval, whereas SSR yields substantially stronger name-free attribute-profile retrieval and is further strengthened by stochastic name-branch dropout. Different capabilities can peak at different stages, so a checkpoint selected by target class-name retrieval need not coincide with a transferable semantic optimum. Continued optimization can therefore preserve strong target class-name retrieval while reducing previously acquired transferable semantic capability. In a representative diagnostic study, this late specialization is consistent with reduced cross-modal semantic accessibility while substantial image-only class structure remains available. Comments: 67 pages, including supplementary material Subjects: Computer Vision and Pattern Recognition (cs.CV) Cite as: arXiv:2610.07385 [cs.CV] (or arXiv:2610.07385v1 [cs.CV] for this version) https://doi.org/10.48550/arXiv.2610.07385 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[CV-108] GeoWM: Efficient Direct World Modeling in Explicit Geometry
链接: https://arxiv.org/abs/2610.07381
作者: Mehrdad Noori,Guile Wu,Sam Hosseini,Dongfeng Bai
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Modeling 3D scene geometry and its evolution over time is essential for autonomous driving and robotics. A common paradigm is to use world models to predict future images or latent representations of the environment and subsequently recover geometry from these predictions. However, this paradigm does not explicitly model geometric structure and typically relies on recursive rollouts to reach longer prediction horizons, leading to error accumulation and increasing computational cost. To address these limitations, we present GeoWM, a geometry world model that directly forecasts future scene geometry at specified future horizons without recursive rollout. The key idea is to leverage a geometry foundation model to transform observed RGB frames into a geometric history, which conditions a flow-matching transformer to predict the scene geometry at a specified future horizon. We further show that a lightweight camera-motion predictor can accurately estimate the future viewpoint, and that projecting the observed geometry into the predicted viewpoint provides an effective geometric prior for future geometry forecasting. Extensive experiments on four datasets spanning urban driving, aerial flight, and dynamic manipulation demonstrate that GeoWM outperforms the evaluated world models in forecasting depth, camera pose, and 3D scene geometry, while substantially reducing inference time at longer horizons.
[CV-109] SimCortex v2: Joint Cortical Surface Reconstruction with Near-Zero Collisions and Self-Intersections
链接: https://arxiv.org/abs/2610.07378
作者: Kaveh Moradkhani,Sylvain Bouix
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Submitted to Medical Image Analysis
Abstract:Reconstructing cortical WM and pial surfaces from structural magnetic resonance imaging (MRI) is a prerequisite for surface-based neuroanatomical analysis, yet remains challenging because the cortex is thin and tightly folded. Reconstruction methods can produce geometric artifacts such as mesh self-intersections and collisions between cortical surfaces, and although recent deep learning methods have reduced reconstruction time from hours to minutes, these artifacts persist. We propose SimCortex v2, a deep learning framework for simultaneous reconstruction of the left and right WM and pial surfaces from T1-weighted MRI. SimCortex v2 estimates topologically correct initial surfaces from a volumetric segmentation and refines all four jointly using multi-scale stationary velocity fields predicted by a ribbon-conditioned, U-Net-like network. We evaluated SimCortex v2 on 560 cases from 14 cohorts, thirteen of them unseen during training, spanning ages 6-89, healthy and clinical populations, and scanners from three vendors. SimCortex v2 matched the surface-distance accuracy of the strongest baseline (average symmetric surface distance 0.253 mm) while showing no detected inter-surface collision in 92.14% of cases and the lowest self-intersection fraction (0.044%) among learning-based methods, whereas every baseline produced at least one collision in every case. Source code, configuration files, pretrained weights, preprocessed data, and the exact evaluation splits are publicly released.
[CV-110] Identity-Conditioned Score Fusion for Open-Set Person Re-Identification
链接: https://arxiv.org/abs/2610.07366
作者: Manyi Yao,Jurijs Nazarovs,Eunji Chong,Abhishek Sharma,Rohan Sarkar,Yue Guo,Christian R. Shelton,Amit K. Roy-Chowdhury,Debashish Pal
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注:
Abstract:Robust person re-identification often combines complementary cues such as face, gait, and body shape. While adaptive fusion typically targets query quality, model strength also varies across identities. We introduce identity-conditioned score fusion, a framework that tailors weights to each gallery identity without training. By contrasting intra-identity consistency against cross-identity impostors, it extracts identity-specific profiles that couple with query-conditioned adaptation via a parameter-free rule. This widens the separation between true and false matches while preserving score calibration. Evaluations on three clothes-changing person re-identification benchmarks show that our method consistently outperforms statistical, rank-based, and learned baselines, achieving up to an 8.8% absolute reduction in the false non-identification rate and demonstrating the value of identity-conditioned fusion in open-set person re-identification.
[CV-111] Compositional Concept Erasure in Text-to-Image Diffusion Models via Hierarchically Grounded Semantic Surgery BMVC2026
链接: https://arxiv.org/abs/2610.07337
作者: Chen Dai,Ganyu Zou,Nathan Self,Kevin Piper,Ramachandra Rao Seethiraju,Karthik Shyamsunder,Chang-Tien Lu,Naren Ramakrishnan
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted at BMVC 2026
Abstract:Removing copyrighted, unsafe, or user-specified concepts from a deployed text-to-image diffusion model is now a practical requirement. Weight-editing methods can suppress fixed targets, but they require per-target retraining and modify the model checkpoint. Training-free methods, on the other hand, are deployment-friendly, but they suffer from text-side routing failures on compositional prompts. In such prompts, the erase target may be invoked through a related class rather than its lexical name, and its modifiers may migrate onto preserved objects. This paper proposes Hierarchically Grounded Semantic Surgery (HGSS), a training-free framework for compositional concept erasure. The framework lifts both the routing signal and the edit operator used by text-side erasure. First, hierarchical span grounding resolves erase-target spans through lexical, taxonomic, and semantic evidence, while guarding against broad-hypernym and compound-head false positives. Second, dynamic attribute binding refines the text conditioning during early denoising via a counterfactual reference and a preserve-aware cross-attention objective, keeping surviving attribute-noun bindings intact. HGSS selectively removes the erase target without updating model weights or adding learned parameters. On SEE, HGSS cuts hierarchical evasion from 29.54 to 10.02 and roughly halves pairwise attribute leakage, achieving the best Neighbor E and AttrP scores among the reported erasure methods. On UnlearnCanvas, HGSS slightly improves the six-metric average over the matched Semantic Surgery baseline, reaching state-of-the-art.
[CV-112] Localize Any Object in X-Ray Security Scans without Human Annotation
链接: https://arxiv.org/abs/2610.07326
作者: Yaqi Cai,Mingxuan Liu,Lorenzo Vaquero,Ning Wang,Nan Pu,Feng Xue,Elisa Ricci,Nicu Sebe
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Universal object localization in X-ray security inspection is critical for automated threat detection in safety-critical venues. However, unlike everyday RGB images that dominate web-scale visual data, X-ray scans exhibit distinct color patterns, ambiguous boundaries, and compositional structures caused by volumetric superposition. These gaps hinder the direct zero-shot transfer of dense perception foundation models trained on web-scale RGB data. Moreover, annotated X-ray data is scarce and requires expert labeling, limiting both the training of generalizable X-ray native models and the adaptation of RGB foundation models for X-ray data via fine-tuning. Given these challenges, the bright promise of highly generalizable perception models, enabled by data scaling laws in the RGB domain, remains largely out of reach for X-ray inspection. To this end, we introduce LAO-X, a self-supervised adaptation framework that Locates Any Object in X-ray scans using diverse synthesized image–annotation pairs with granularity-aware supervision. LAO-X first designs a saliency-guided X-ray object mining module to separate diverse object instances, which are then used for physics-guided synthesis in the absorbance domain. LAO-X further incorporates an occlusion-controlled curriculum strategy to fine-tune a Segment Anything Model 2 (SAM2) localizer, progressively adapting it to X-ray scans with increasing object counts and overlap levels. Experiments on six X-ray benchmarks show that LAO-X substantially improves category-agnostic localization, achieving 2% to 23% mAP gains over SAM2 and X-ray specific baselines in heavily cluttered scenarios, entirely without human-annotated labels.
[CV-113] ATLAS-AL: Adaptive Trust-Region for Latent Adversarial Searches via Active Learning
链接: https://arxiv.org/abs/2610.07323
作者: Marsalis Gibson,Claire Tomlin,Shankar Sastry
类目: Machine Learning (cs.LG); Cryptography and Security (cs.CR); Computer Vision and Pattern Recognition (cs.CV)
备注: 9 pages main body, plus 10 additional pages for references and appendix
Abstract:Security evaluation of learning-based systems requires more than just testing the system against a fixed collection of attacks. It requires adaptive mechanisms that can efficiently discover \textitsets of inputs that induce model failure. We introduce ATLAS (Adaptive Trust-Regions for Latent Adversarial Searches), which is a query-based framework that discovers adversarial input sets for black-box learning systems. ATLAS casts attack generation as an active learning level set estimation problem then combines calibrated approximations with a local-global sampling architecture to find regions of the input space that contain adversarial examples. Once discovered, ATLAS is designed to sample points within these adversarial regions to build adversarial sets that accurately represent the state of robustness of the target model. When applied on toy experiments, we find that ATLAS is able to recover more of the adversarial region under a limited query budget than does previous work. When applied to standard and adversarially trained MNIST, CIFAR, and ImageNet model targets, ATLAS produces better representative attacks than other query-based black-box attacks (NES, SignHunter, BayesOpt). ATLAS represents an automated red-teaming framework that can be used for both analyzing the robustness of learning-based systems under development and continuous auditing to see how the robustness of a system changes over time.
[CV-114] What Words Keep of a Place: Zero-Shot Language Reasoning for Cross-View Geo-Localization NEURIPS2026
链接: https://arxiv.org/abs/2610.07269
作者: Ayesh Abu Lehyeh,Jay Hwasung Jung,Safwan Wshah
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注: Accepted at NeurIPS 2026 Workshop Physical World AI: Geometry, Characteristics, and Multimodal Sensing
Abstract:Cross-view geo-localization is commonly solved as an image retrieval problem, matching a ground-level image against a database of satellite tiles through a jointly trained embedding. Such models are accurate, but they need large paired supervision and cannot show what evidence supports a match. In this paper, we study a different question: how much of this task can be solved through language alone? We prompt a multimodal large language model (MLLM) to describe each ground panorama and each satellite tile as structured text, and localize by comparing these descriptions. No component is trained. We evaluate on 9,826 VIGOR pairs from four U.S. cities, in three settings. First, the descriptions are faithful but not discriminative. They agree closely across the two views, yet ranking the full pool by description similarity almost never returns the correct tile (0.39% Recall@1). Second, we narrow the pool to ten neighboring tiles, as a coarse prior would do. The same descriptions now become useful: an MLLM judge that scores structural consistency doubles random ranking and matches a strong lexical baseline. It also states which fields of the two descriptions agree and which conflict, which an embedding distance cannot do, and which we see as a step toward interpretable localization. Third, we place the judge on a trained visual retriever. On the queries it ranks wrongly, reranking from images works, while reranking from our descriptions does not (23.5% against 10.7% Recall@1). Scene structure survives the conversion into language, while the fine appearance detail needed to separate nearby places does not. Code and prompts are publicly available at this https URL.
[CV-115] Hybrid Cross-Modal Attention Network for Early Breast Cancer Detection in Low-Resource Clinical Settings ICIP WWW
链接: https://arxiv.org/abs/2610.07243
作者: Simon Hadush Nrea(1),Filimon Gidey Gebremichael(1),Gebrekirstos Hagos Gebrekirstos(2),Yaecob Girmay Gezahegn(1) ((1) Mekelle University, Mekelle, Ethiopia (2) Clinical Oncologist London School of Hygiene and Tropical Medicine London, UK)
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注: 5 double pages numbers, conference paper presented at AI4SD 2026 ( this https URL )
Abstract:Breast cancer is the leading cause of cancer-related mortality among women in Sub-Saharan Africa, where delayed diagnosis results from limited radiology expertise and fragmented clinical data systems. Although deep learning models have demonstrated strong performance in mammographic analysis, most rely solely on imaging data and are trained on Western populations, limiting their applicability in African healthcare settings. This paper presents a Hybrid Cross-Modal Attention Network (HCMAN) that integrates mammogram images with structured clinical data using transformer-based cross-modal attention mechanisms. The model was developed and validated using a locally collected dataset of 2,560 mammogram images from 1,024 patients across four Ethiopian referral hospitals, with biopsy-confirmed ground truth labels. The proposed framework achieves 97.8% accuracy, 97.2% sensitivity, 98.3% specificity, and an AUC of 0.987, significantly outperforming image-only baselines. The system demonstrates robustness to low-quality images typical of resource-limited settings, with only 3.2% performance degradation compared to 8.7% for image-only models. Cross-modal attention analysis reveals clinically appropriate behavior: higher reliance on clinical features for ambiguous cases such as dense breasts and young patients. The model’s lightweight architecture enables deployment on standard hospital workstations (2 seconds inference on CPU). This work advances sustainable, context-aware AI solutions for equitable breast cancer diagnostics in Africa.
[CV-116] Monocular Navigation Relative to Unknown Spacecraft Using a Transformer-Aided Kalman Filter
链接: https://arxiv.org/abs/2610.07231
作者: Pol Francesch Huc,Simone D’Amico
类目: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:This work presents a novel learning-based pipeline for pose estimation of unknown spacecraft using only monocular images from a single servicer. The approach combines a transformer-based neural network with a Multi-State Constraint Kalman Filter (MSCKF) to estimate the pose (i.e., position and orientation of the target spacecraft relative to the camera) throughout rendezvous and proximity operations. Unlike existing vision-based methods that require prior knowledge of the target shape or inertia properties, rely on additional sensing modalities such as depth, lidar, or stereo, or only recover translation up to scale, the proposed pipeline generalizes to previously unseen spacecraft using a single monocular camera. The transformer network estimates the odometry, the change in pose between images up to scale, from SuperPoint features matched by LightGlue. The MSCKF uses these pseudo-measurements along with an orbit and attitude kinematics model to estimate the pose of the target. In particular, the relative orbit elements, the target’s attitude with respect to the servicer’s camera, and the associated angular velocity are estimated directly by the filter. Given the monocular approach and short distance to the target, the full observability of the range to the target is recovered via attitude maneuvers by the servicer. The method is trained and evaluated on a re-rendered high-resolution version of the SPE3R dataset, which includes synthetic images of 103 spacecraft. Eleven of these spacecraft are held out during training to evaluate the generalization to unseen targets. Monte Carlo simulations are then used to evaluate the navigation pipeline on rendered trajectories of the held out spacecraft. The results demonstrate that learned vision pipelines as a front-end for Kalman filters provide median errors of 3.7° in attitude and 2.2% of range in ROE when navigating about unknown targets.
[CV-117] Deep Learning Based Illegal Bowling Action Detection
链接: https://arxiv.org/abs/2610.07223
作者: Debopom Sutradhar,Niful Islam,Sudipto Mondal,Tasmima Hossain Jamim,Jubaer Muhammad Shufol,Swakkhar Shatabda
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Cricket, often referred to as the “gentleman’s game,” adheres to a strict rule set for both batsmen and bowlers, where each delivery can significantly impact the match outcome. Detecting illegal bowling actions is crucial for maintaining fair play, yet it remains challenging for umpires to monitor in real time. Existing sensor-based solutions have limitations in live match scenarios, making real-time assessment difficult. This paper proposes a computer vision-based deep learning solution to detect illegal bowling actions in live cricket matches. To develop and evaluate our approach, we compiled a dataset of 62 videos featuring 11 male bowlers, capturing both legal and illegal bowling actions from multiple angles-front, back, and side. However, the dataset predominantly comprises right-handed bowlers with conventional actions. The proposed system identifies two key frames, the shoulder frame and the release frame from video footage of a bowler’s delivery and analyzes the change in the bowling arm’s angle between these frames. If the angle difference exceeds a predefined threshold (e.g., 15 degrees), the delivery is flagged as potentially illegal. We evaluated the system on a custom dataset and achieved a high true positive rate, suggesting the system’s potential effectiveness in real-time match settings. However, further research is required to validate the system across diverse environmental conditions and larger datasets to ensure generalizability and robustness in various live match scenarios. To the best of our knowledge, this is the first AI-based computer vision method for detecting illegal bowling actions in cricket.
[CV-118] Energy-Conditioned Noise Schedule and Whitening for Spectral Diffusion
链接: https://arxiv.org/abs/2610.07206
作者: Bata Vasic,Bane Vasic
类目: Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Emerging Technologies (cs.ET)
备注: 7 pages, 5 figures, 2 tables, Manuscript submitted for publication in Elsevier Pattern Recognition Letters
Abstract:This paper introduces an energy-adaptive noise scheduling and whitening strategy for transform-domain diffusion models. Existing spectral diffusion methods account for the non-uniform statistics of transform coefficients through coefficient scaling, normalization, or frequency prioritization, while the forward diffusion noise schedule remains largely independent of the underlying spectral-energy distribution. We investigate whether the temporal evolution of the forward diffusion process should also follow the spectral organization of natural images. The proposed formulation combines global spectral whitening with energy-conditioned noise allocation that jointly modulates the injected noise according to the energy of individual transform coefficients and an image-dependent energy path over diffusion time. The resulting forward process preserves Gaussian transitions with closed-form marginals and remains compatible with standard DDPM and DDIM procedures without modifying the diffusion architecture. Experiments on CIFAR-10 demonstrate the contribution of the proposed energy-conditioned noise schedule and spectral whitening, reducing Fréchet Inception Distance from 142.48 for a compact DCTdiff U-Net variant to 100.45.
[CV-119] CALR: Continuous Anchored Latent Reasoning via Render-of-Thought Compression
链接: https://arxiv.org/abs/2610.07175
作者: Zhaoyang Wei,Bowen Jiang,Yanchao Hao,Wenchao Ding,Zheng Wei,Shaocheng Wu,Zhenjun Han,Jianbin Jiao
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Visual latent reasoning compresses rendered derivations into compact intermediate states, reducing textual reasoning overhead. Existing approaches differ in how they represent these states: continuous methods avoid vocabulary constraints, whereas discrete methods improve accuracy through quantization into a finite codebook. Our analysis of representative continuous and discrete systems identifies two functional requirements: answers must rely on latent states, and those states must carry valid, problem-specific reasoning. Continuous latents influence answers despite collapsed reasoning content, whereas discrete latents retain recoverable intermediate reasoning that answer prediction largely bypasses. To address these challenges, we propose Continuous Anchored Latent Reasoning (CALR), which connects latent formation with answer use through functional anchoring. With reference latents from information-balanced compression, CALR couples latent-mediated answer supervision with derivation-level semantic anchoring: the former routes answer supervision through intermediate states, while the latter grounds their decoded content in problem-specific derivations. A parallel-to-autoregressive curriculum develops sequential reasoning by conditioning subsequent latent blocks on generated prefixes. Evaluations on five mathematical reasoning benchmarks across model families show substantial accuracy gains. Under matched budgets, CALR gains 26.0 percentage points over a comparable continuous latent reasoning method. Further analyses show that its latents support answer prediction and carry problem-specific intermediate reasoning.
[CV-120] PlaySuite: A Large-Scale Benchmark for Interactive Visual Intelligence
链接: https://arxiv.org/abs/2610.07127
作者: Dheeraj Varghese,Anna Vettoruzzo,Walter Simoncini,Michelle Lorena Acevedo Callejas,Mohammad Mahdi Derakhshani,Kristof Meding,Joaquin Vanschoren,Cees G. M. Snoek
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:
Abstract:Recent advances in multimodal foundation models yield strong performance on static perception and reasoning benchmarks, yet such evaluations largely overlook a central aspect of intelligence: acting competently in dynamic environments over extended time horizons. We introduce PlaySuite, a large-scale benchmark for evaluating interactive visual intelligence across more than 5K open-source video games curated from PyWeek and this http URL. Spanning diverse genres and engines, including Pygame, HTML5, Godot, and Unity, these independent games are largely out-of-distribution for current models, reducing the likelihood that success can be achieved by retrieving memorized walkthroughs or web-scale training artifacts. To enable scalable evaluation across heterogeneous titles, we develop a unified closed-loop interaction framework optimized for HPC clusters alongside a Video-LLM-as-a-judge protocol that maps observable gameplay milestones to standardized progress levels. We evaluate fourteen recent open models spanning vision-language models, computer-use agents, and vision-language-action models. Our results yield strong evidence of a perception-action gap: despite strong reasoning capabilities, current models struggle to make sustained progress and exhibit recurring failures in spatial grounding, action execution, and self-correction. PlaySuite provides a reproducible and extensible testbed for measuring progress from visual perception to goal-directed interaction, and a foundation for developing models that can act, adapt, and generalize in dynamic visual environments.
[CV-121] R2RI: A Multi-View Event and RGB Dataset for Robot-to-Robot Interaction
链接: https://arxiv.org/abs/2610.07117
作者: Gabriele Magrini,Riccardo Catalini,Federico Becattini,Guido Borghi,Pietro Pala,Roberto Vezzani,Lorenzo Seidenari
类目: Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)
备注:
Abstract:Understanding and modeling interactions between autonomous agents is a fundamental challenge in robotics, with broad implications for collaborative systems, social robotics, and human-robot coexistence. Although the study of robot interactions has emerged as a compelling research direction, progress has been severely hampered by the absence of large-scale benchmarks. In this paper, we introduce Robot-to-Robot Interaction (R2RI), the first dataset specifically designed to address the Robot-Robot Interaction (RRI) task. R2RI consists of different humanoid robots and realistic interactions modeled on real human social behaviors. Complementary viewpoints are available, \textiti.e., an egocentric perspective from each robot’s onboard sensors, and an exocentric perspective from external fixed cameras, thus enabling rich spatial and contextual understanding of the interaction dynamics. The dataset comprises more than 6.5 M frames and \approx5000 videos at 120 fps, including Event and RGB domains. We investigate pros and cons of each domain, comparing state-of-the-art approaches for a number of key sensing and interaction based tasks. We publicly release the dataset and its annotations for all tasks and modalities at this https URL.
[CV-122] Sample-Optimal Estimation of the Fréchet Inception Distance
链接: https://arxiv.org/abs/2610.07114
作者: Ziyun Chen,Jerry Li,Kevin Tian,Yusong Zhu
类目: Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV); Data Structures and Algorithms (cs.DS); Statistics Theory (math.ST); Machine Learning (stat.ML)
备注: Our code is available at this https URL
Abstract:The Fréchet Inception Distance (FID) is widely used to evaluate generative models, but its empirical plug-in estimator suffers from finite-sample bias [BSAG18, CF20]. We study the sample complexity n of estimating FID to error \epsilon between d -dimensional Gaussians with bounded mean distance and covariances, when one distribution is known. Our contributions are threefold. (1) We establish tight finite-sample \Theta(\fracd^2n) bias and \Theta(\fracdn + \frac d^2 n^2) variance bounds for the empirical plug-in estimator, establishing a \gtrsim d^2 sample complexity. (2) To debias the empirical plug-in estimator, we generalize the \rm FID_\infty estimator of [CF20] to extrapolation methods of arbitrary order k . We further prove tight bias and variance bounds of \Theta(\fracd^k + 2n^k + 1) and \Theta(\frac d n + \fracd^2n^2) for any order- k extrapolation under our framework. (3) We introduce Relative Taylor Debiasing (RTD), a new, computationally efficient FID estimation algorithm using debiasing techniques inspired by U-statistics. We show that RTD achieves an O(\frac d \epsilon^2) sample complexity, and prove that this is optimal. We provide a complementary empirical evaluation of our new estimators. Our experiments on synthetic Gaussians validate the predicted residual bias and support the tightness of our bounds. On ImageNet with Inception embeddings, RTD achieves the lowest mean estimation error at the standard 50K sample budget, while our second-order variance-aware extrapolation estimator (VALE _2 ) uses only 10K samples to achieve accuracy comparable to FID _\infty at 50K samples.
[CV-123] MoonGS: High-quality Representation of the Lunar Surface via Gaussian Splatting Using Robust Depth Features from Image Pairs
链接: https://arxiv.org/abs/2610.07110
作者: Yun Jiang,Bo Zheng,Yingying Zhang,Xueming Xiao,Tao Hu,Hutao Cui,Zhiguo Meng,Ke Gao,Yang Gao,Meibao Yao
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:High-quality 3D reconstruction of lunar terrain from sparse rover images is indispensable for autonomous lunar exploration, but remains challenging because viewpoint overlap is insufficient, surface textures are weak, and data volume is limited. We propose MoonGS, the first feed-forward 3D Gaussian Splatting framework tailored to lunar scenes. Given only two input images, MoonGS predicts pixel-aligned Gaussian primitives in a single forward pass and renders photorealistic novel views without any per-scene optimization. MoonGS (i) adopts an adaptable backbone design that seamlessly integrates advanced vision foundation models to extract robust depth features; (ii) integrates semantic priors in two manners: merging semantic cues with visual features to refine Gaussian parameter estimation, and adopting a semantic ranking loss that regularizes background depth; and (iii) employs an entropy-guided heuristic resampling strategy to augment sparse observations by selecting the most informative distant viewpoints with negligible overhead. Experiments on the LuSNAR benchmark and our synthetic weak-texture MoonBlender dataset show that MoonGS surpasses state-of-the-art feed-forward NeRF/3DGS baselines by +4.9 dB PSNR, +0.29 SSIM, and 40% lower LPIPS while maintaining sub-second inference. Furthermore, we validate the broad applicability of our framework by demonstrating that it effectively leverages state-of-the-art backbones, including VGGT, to significantly boost performance. Qualitative evaluations on Chang’e mission imagery also show the best visual quality among compared methods, indicating robustness on real lunar data. The source code and dataset are publicly available at this https URL.
[CV-124] A Data-Centric Review of Plant Disease Datasets: Taxonomy Critical Analysis Environmental Variability and Implications for Precision Agriculture
链接: https://arxiv.org/abs/2610.07087
作者: Aamir Hilal,Shabir Ahmad Sofi,Neeraj Goel
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Despite rapid advances in artificial intelligence, reliable real-world plant disease detection remains a persistent challenge. Visual and deep learning approaches have shown promising results, but their deployment under field conditions remains limited. A key bottleneck is the reliance on laboratory-generated datasets that lack environmental diversity, realistic backgrounds, and balanced class distributions, resulting in poor generalization. In contrast, datasets collected directly from agricultural environments capture natural variability and better reflect challenges faced by farmers across regions. This review presents a critical analysis of visual and deep learning approaches for plant disease detection, with emphasis on plant disease datasets. It establishes a taxonomy based on acquisition setting, accessibility, plant diversity, disease composition, class structure, and imbalance severity, and examines their implications for model generalization and real-world deployment. A comparative analysis of laboratory and real-field datasets identifies critical gaps that hinder disease detection. The review further analyzes how multi-level dataset imbalance, including intra-class, inter-crop, and cross-dataset imbalance, and limited environmental variability affect model performance and robustness, an area insufficiently examined in existing surveys. Beyond image-based approaches, it highlights the importance of integrating environmental parameters such as temperature, humidity, and leaf wetness with image data to improve prediction under dynamic field conditions. Finally, the review identifies key challenges, research gaps, and future directions concerning dataset construction, environmental variability, structural imbalance, standardization, and multimodal disease monitoring. It provides a foundation for developing next-generation multimodal frameworks for precision agriculture.
[CV-125] Graph-Based Recognition of Simulated Train-Driver States From Facial and Upper-Body Keypoints
链接: https://arxiv.org/abs/2610.07083
作者: Olivia Nocentini,Marta Lagomarsino,Gokhan Solak,Younggeol Cho,Qiyi Tong,Sara Zeynalpour,Marta Lorenzini,Alessandro Ledda,Arash Ajoudani
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注: 11 pages,5 figurees
Abstract:Driver fatigue poses a significant challenge to railway safety, with traditional systems like the dead-man switch offering limited and basic alertness checks. This study presents a vision-based monitoring system that relies solely on a single front-facing RGB camera and a graph neural network to classify simulated train-driver states into alert, not-alert, and an emergency class comprising acted emergency-like behaviours. To optimize input representations for the model, an ablation study was performed, comparing three feature configurations: skeletal-only, facial-only, and a combination of both. Experimental results show that combining facial and skeletal features yields the highest accuracy (81%) for the three-class model under the light condition, outperforming models that use only facial or skeletal features. Furthermore, the combination of facial and skeletal features achieves 99% accuracy in the alert/not alert classification in light condition. Additionally, we introduced a controlled RGB video dataset containing alert, not alert, and acted emergency-like behaviours recorded under three illumination conditions. These contributions represent a step toward passive and non-contact train-driver state recognition based on facial and upper-body dynamics.
[CV-126] On Color Alignment in VAE Latent Spaces and Its Applications
链接: https://arxiv.org/abs/2610.07072
作者: Julian D. Santamaria,Kai Wang,Jesús Malo,Javier Vazquez-Corral,Alexandra Gómez-Villa
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注:
Abstract:Variational autoencoders (VAEs) are a key part of modern text-to-image models, which generate images within their latent space. VAEs are known to disentangle the main factors of variation in the data, and color is known to be one of the most structured of these in natural images: decorrelating it yields one luminance axis and two opponent-color axes. Color should therefore be expected to emerge as a distinct factor in the VAE latent space. Yet how these latent spaces represent color remains largely unexplored. In this work, we show that the VAEs of text-to-image models share a color subspace aligned with brightness and opponent-colors. Through a linear approximation of the encoder and targeted latent steering, we find this subspace consistently across a broad range of VAEs, from SD1.5 to FLUX.2 and Z-Image. Building on this characterization, we propose three applications: ColorTuning, which achieves state-of-the-art in precise numerical color generation on the fine-grained CSS3/X11 system of GenColorBench, saturation control, to adjust the global chromatic intensity, and color transfer, to change the palette to match a reference. The code and models are publicly available at this https URL
[CV-127] A BEMD-Based Quaternion Filtering Approach Sharp-to-Soft Kernel CT Image Conversion
链接: https://arxiv.org/abs/2610.07071
作者: Mahmoud Nasr,Jan K. Argasinski,Krzysztof Brzostowski,Adam Piorkowski
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:The quality of computed tomography (CT) images is significantly affected by the selection of reconstruction kernels: sharp kernels improve spatial resolution but increase noise, whereas soft kernels diminish noise at the expense of edge clarity. This study presents an innovative enhancement framework utilising Bidimensional Empirical Mode Decomposition in conjunction with Quaternion Bilateral Filtering (BEMD–QBF) to convert sharp-kernel CT images into representations resembling soft-kernels, while maintaining critical anatomical structures. The technique disaggregates each image into intrinsic mode functions via BEMD and analyzes them inside a cohesive quaternion framework to attain efficient noise reduction and structural integrity. The proposed methodology is evaluated using several reconstruction kernels (B50, B46, B41, B36, B35, B31) and compared with recognised filtering strategies, including Non-Local Means, Anisotropic Diffusion, Bilateral Filtering, and Quaternion Bilateral Filtering. Quantitative evaluations of the Structural Similarity Index (SSIM) and Peak Signal-to-Noise Ratio (PSNR) indicate that BEMD-QBF consistently attains superior structural fidelity and competitive noise reduction across all evaluated kernels. The results underscore the efficacy of the proposed strategy as a viable approach to enhancing post-reconstruction CT images, yielding superior image quality without requiring access to raw projection data.
[CV-128] Decomposition-Guided Curvelet Thresholding for Sharp-to-Soft CT Kernel Conversion
链接: https://arxiv.org/abs/2610.07067
作者: Mahmoud Nasr,Jan K. Argasinski,Krzysztof Brzostowski,Adam Piorkowski
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Image denoising is a crucial task in image processing, focused on improving image quality by minimizing noise while maintaining essential structural elements. This study presents a hybrid denoising framework that combines several decomposition techniques, including empirical mode decomposition (EMD), variational mode decomposition (VMD), multichannel EMD (MEMD), and bidimensional EMD (BEMD), with curvelet transform thresholding. Each decomposition mode undergoes processing through both soft and hard thresholding, and the denoised modes are combined to rebuild the final image. Comprehensive evaluations of standard CT image datasets reconstructed with various kernels (B50, B46, B41, B36) reveal substantial enhancements in denoising efficacy. VMD consistently achieves the highest peak signal-to-noise ratio (PSNR) and structural similarity index (SSIM), signifying exceptional noise reduction and feature preservation. The study analyses the trade-offs between soft and hard thresholding: soft thresholding maintains intricate visual details, whilst harsh thresholding provides enhanced noise reduction. The suggested method surpasses traditional techniques in both reference and non-reference quality criteria, indicating its potential for broader application in medical imaging and future incorporation with adaptive thresholding algorithms.
[CV-129] Crop Yield Prediction for Punjab Pakistan: A Tree-Ensemble and Leaf-Health Prototype and What Random Validation Hides
链接: https://arxiv.org/abs/2610.07059
作者: Amina Asif,Qurat ul ain Asif,Noor Bakhat Asif
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Yield forecasts help planners and farmers decide on inputs, storage and imports, but small agricultural tables can make reported accuracy fail on a new season. We built a crop-yield prototype for Punjab, Pakistan that combines Random Forest, XGBoost, support vector regression and a Ridge-stacked ensemble with a MobileNetV2 leaf-health classifier, and deployed it as a web application with SHAP explanations. A random 80/20 split of a merged Kaggle-derived table (414 rows) gives the ensemble an R2 of 0.991. An audit showed that the table contains only 46 independent observations: a join with nine temperature records per year copied every crop-year nine times. Holding out whole years takes XGBoost on the same rows from R2 = 0.994 to -0.20. On deduplicated data, and on a longer FAOSTAT table (1990-2024, 70 observations), a per-crop linear trend (leave-one-year-out RMSE 0.29 Ton/Ha) beats every model not given the year (0.84-1.06). Neither pesticide use nor national temperature change explains the trend residuals. An apparent pesticide gain in the Kaggle table disappears on FAOSTAT, where the pesticide series is mostly imputed and the two releases disagree. An independent district-level wheat panel (36 districts, 13 seasons) shows that most variation in Punjab is spatial and that a district mean with a common trend matches the learned models. The leaf classifier reaches 99.87% accuracy on held-out PlantVillage images and recalls 96.4% of 336 unseen diseased leaves after near-duplicates were removed. However, no leaf image is paired with a yield record, so the health score used in the yield models had to be constructed and adds nothing. We report these negative findings together with the prototype.
[CV-130] Artemis: Geometry-Grounded Multi-Agent Driving World Models with Shared 3D State and Progressive Memory Update
链接: https://arxiv.org/abs/2610.07031
作者: Sitian Shen,Jiuming Liu,Mengmeng Liu,Yian Wang,Michael Ying Yang,Francesco Nex,Hao Cheng,Daniele De Martini,Ayush Tewari,Per Ola Kristensson
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: The first three authors contributed equally, and their order was determined by drawing lots. Project Lead: Jiuming Liu. Corresponding Author: Ayush Tewari. Project page: this https URL
Abstract:Recent video world models have witnessed the paradigm shift from single-agent to multi-agent involvements, which can reveal more complicated dynamics and cross-agent interaction in the real world. However, existing approaches commonly adopt implicit inter-agent communications via cross attention, which lack explicit geometry constraints and unified 3D state, thereby leading to poor multi-view consistency and struggling with recovering out-of-sight agents. In addition, most of them assume a static background, failing to represent uncontrolled background dynamics. To address these problems, we propose Artemis: a geometry-grounded multi-agent world model with explicit memory sharing. An explicit 3D world map is reconstructed from multi-agent observations to enforce a unified 3D state across agents, offering high cross-view consistency. Specifically, an action-guided geometric injection module is developed to simultaneously render decomposed foreground-background control maps, which are then injected into a diffusion transformer through a designed GeoAdapter block. Compared to previous methods assuming static-only background, our GeoAdapter can also distinguish uncontrolled non-agent dynamics, which are conditioned on their own multi-frame history positions to provide consistent motion cues. Keyframes selected from progressive video rollouts are used to progressively update the reconstructed 3D world maps. To effectively capture complex dynamic patterns, we curate a novel dataset sampled from the CARLA simulator called MA-CARLA. Extensive experiments demonstrate the superiority of our proposed method in terms of visual fidelity and cross-view consistency in the generated videos. In addition, our Artemis can support simultaneous multi-modal rollouts with both 2D video and 3D point map maintenance, scale to scenarios beyond two agents and multi-camera setting.
[CV-131] WiSPER: Pose-Supervised Predictive and Residual Flow Refinement For Multi-Person 3D Pose Estimation With WiFi CSI
链接: https://arxiv.org/abs/2610.07025
作者: Gabriel Lee Jun Rong,Shanhong Liu,Pai Chet Ng,Konstantinos N. Plataniotis,Jamal Seyedmohammadi,S. Mohammad Sheikholeslami
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Multi-person 3D pose estimation with WiFi channel state information (CSI) is challenging because reflections from different people overlap without directly identifying individual joints. Existing masked embedding objectives capture wireless relationships without explicit pose supervision, while structured decoders can retain coordinate errors. We propose WiSPER, a two-stage framework combining pose-aware predictive pretraining with conditional residual flow refinement. Pose-Aware Masked Embedding Learning (PAMEL) couples masked latent prediction with auxiliary pose-set supervision on the same CSI context, guiding the encoder toward joint localization from partial observations. Residual Flow refinement with Transformer (ReFT) generates a set of pose candidates to accommodate a variable number of people and refines each candidate through a conditional flow guided by its coarse coordinates and per-joint decoder features. Both stages use paired CSI and pose annotations during training, while inference requires only CSI. Experiments on the PiW3D dataset show that WiSPER achieves an overall mean per-joint position error of 63.72 mm, a 40.0% reduction relative to WiFi-JEPA. For experiments with two and three people, WiSPER reduces MPJPE by 42.1% and 38.1%, respectively. Pose-supervised pretraining configurations obtain lower errors than CSI-only JEPA, and enabling the trained residual refiner reduces overall MPJPE by 13.8-15.6% across the evaluated configurations.
[CV-132] Anchor and Adapt: Asymmetric Prompt Adaptation for Few-Shot Industrial Anomaly Detection
链接: https://arxiv.org/abs/2610.07016
作者: Mengyang Zhao,Teng Fu,Haiyang Yu,Ke Niu,Bin Li,Xiangyang Xue
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:
Abstract:In few-shot industrial anomaly detection, the few normal target images provide no direct defect supervision, making anomaly prompts difficult to learn from these samples alone. Some vision-language methods therefore use manually specified descriptions to supply explicit anomaly semantics. However, constructing these descriptions requires product-specific effort, and their effectiveness depends on prompt selection. We propose Anchor and Adapt, a two-stage prompt learning framework that separates the acquisition of anomaly semantics from adaptation to target normal appearance. Stage I learns transferable normal and abnormal anchors from annotated auxiliary data. Stage II keeps these anchors fixed and adapts an additional normal branch using the few target normal samples. The inherited and adapted normal branches jointly characterize target normality, with text-anchor regularization encouraging consistency with the generic normal prior and separation from the abnormal anchors. This design retains learned anomaly knowledge while reducing dependence on category-specific anomaly templates, without requiring synthetic anomaly generation. Cross-dataset experiments between MVTec-AD and VisA under 1-, 2-, and 4-shot settings demonstrate competitive detection and localization performance. Controlled ablations assess the roles of transferred anchors, asymmetric adaptation, dual-normal representations, and anchor regularization.
[CV-133] DTFormer: Text-Guided Semantic Alignment for RGB-D Segmentation
链接: https://arxiv.org/abs/2610.07014
作者: Ziang Wei,Yinlong Liu,Yan Xia,Alois Knoll,Hu Cao
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: This is submitted to IEEE Journal
Abstract:RGB-D semantic segmentation has made notable progress by fusing RGB and Depth, yet mainstream models still learn features almost exclusively from pixel-level supervision, lacking direct high-level semantic constraints. This raises a central question-can external knowledge such as language priors inject stronger semantic discriminability into mainstream RGB-D segmentation models. We present DTFormer, a novel tri-modal (RGB-D-Text) semantic segmentation framework. At its core is Text-guided Semantic Alignment Module (TSAM) that first encodes textual cues into a set of semantic prototypes and then explicitly aligns multi-modal RGB-D features with these prototypes at multiple encoder and decoder layers. This design imposes strong semantic regularization on representation learning, guiding the network toward more discriminative features. Extensive experiments on multiple benchmarks show that DTFormer delivers consistent gains while remaining simple and efficient. Our results demonstrate that explicit semantic alignment offers an effective and practical route to improving RGB-D semantic segmentation. The code will be released upon acceptance.
[CV-134] Learning to Curate What You Generate for Generalizable Few-Shot Class-Incremental Learning ACM-MM2026
链接: https://arxiv.org/abs/2610.07008
作者: Junhui Yin,Yuchen Yang,Yilin Yin,Shuai Na,Haoran Xi,Jianhua Yang,Muyi Sun,Man Zhang,Shengfeng He
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted by ACM MM 2026
Abstract:Few-shot class-incremental learning (FSCIL) aims to learn novel classes from limited annotations while preserving prior knowledge. Existing methods typically assume a sufficiently large base session, but this assumption fails when both base and incremental data are scarce, leading to weak initial representations, semantic drift, and unstable boundaries. We study this underexplored yet realistic setting, termed Generalizable FSCIL (G-FSCIL), where the base session itself contains only a few classes. Although synthetic data can alleviate supervision scarcity, naively mixing generated samples often introduces semantic noise and exacerbates old-new boundary conflicts. To address this, we propose a framework that curates trustworthy synthetic knowledge for stable G-FSCIL. Specifically, we first construct class-specific synthetic candidate pools using a frozen latent diffusion model, where class inversion is performed at the first observation and the resulting condition embeddings are reused for on-demand generation. Building on these candidates, we learn a knowledge curation strategy that selects samples with both semantic consistency and visual diversity, and distill this process into a transferable selection policy during the base session, which is then reused without further optimization. Leveraging the curated synthetic data, we further design a boundary-stable incremental adaptation scheme, including synthetic-informed prototype initialization and bidirectional boundary calibration to mitigate old-new conflicts. Extensive experiments demonstrate that our method consistently outperforms existing FSCIL baselines, with reduced forgetting and improved balance between old and new classes. Code is available at this https URL.
[CV-135] Should We Skip Diffusion?
链接: https://arxiv.org/abs/2610.07002
作者: Yiping Ji,James Martens,Simon Lucey
类目: Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Diffusion models learn semantic representations while generating images. In the Decoupled Diffusion Transformer (DDT), a condition encoder provides features that guide a velocity decoder in denoising. To enable effective denoising at all noise levels, these features must capture both high-level abstract structures and low-level details. However, skip/residual connections in the encoder allow shallow features to bypass successive transformations, which may limit progressive abstraction, or at least make it difficult to disentangle different levels of abstraction. We propose DDT-RFE, which removes the residual connections around the Self-Attention and MLP operations in each encoder block while maintaining stable training. To retain the information that abstraction discards but that the decoder still needs, we fuse the input patch embedding with intermediate and final encoder features to form the encoder output. The decoder thus has access to information from multiple encoder depths, while each encoder block is able to learn more abstract representations. DDT-RFE achieves overall improvements over DDT across visual understanding tasks, including image classification, semantic segmentation, object discovery, and semantic correspondence, while using fewer encoder blocks. It also achieves a lower FID for image generation on ImageNet.
[CV-136] State-Aware Interaction MIL for Rare Joint Molecular Phenotype Prediction in Colorectal Cancer and Lung Adenocarcinoma NEURIPS2026
链接: https://arxiv.org/abs/2610.06991
作者: Dasari Naga Raju,Tripti Bameta
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted at the NeurIPS 2026 Workshop on AI at Scale for Clinical Impact (ASCI): Cancer Pathology Foundation Models
Abstract:Joint molecular phenotype prediction is complicated by small joint-positive populations and overlapping histological features across alternative molecular states. Existing computational pathology approaches typically predict biomarkers independently or formulate the joint-positive phenotype as a binary endpoint. Independent prediction does not model interactions between biomarker-specific histological representations, whereas binary joint prediction collapses the double-negative and two single-positive configurations into a single negative class. We propose State-Aware Interaction MIL, a weakly supervised method that preserves biomarker-specific histological representations, models their interaction, and supervises the complete four-state molecular configuration. We evaluate the proposed approach for joint BRAF+/MSI+ prediction in colorectal cancer and EGFR+/TP53+ prediction in lung adenocarcinoma using frozen UNI2-h and CONCH pathology foundation-model representations. With UNI2-h, State-Aware Interaction MIL achieved an average precision of 0.5566 in colorectal cancer (joint-positive prevalence 6.8%) compared with 0.5161 for NaiveMTL, and 0.2784 in lung adenocarcinoma (joint-positive prevalence 8.6%) compared with 0.2525 for IndependentPair. With CONCH, State-Aware achieved an average precision of 0.4410 compared with 0.3932 for DirectJoint in colorectal cancer and 0.1659 compared with 0.1226 for DirectJoint in lung adenocarcinoma. These results indicate that pathology foundation-model representations contain predictive information for rare joint molecular phenotypes and that preserving biomarker-specific representations within a structured molecular-state formulation can improve prediction of these phenotypes from histopathology.
[CV-137] Hierarchy-GBP: Accelerating Factor Graph Inference via Abstraction and Recovery
链接: https://arxiv.org/abs/2610.06978
作者: Yuzhou Cheng,Tom Yates,Ignacio Alzugaray,Danyal Akarca,Pedro A. M. Mediano,Andrew J. Davison
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Robotics (cs.RO)
备注: 33 pages, 10 figures, including appendices
Abstract:Gaussian Belief Propagation (GBP) is a distributed inference algorithm that passes messages in graphical models, making it attractive for scalable spatial intelligence. However, we find GBP most effective locally: it rapidly smooths message errors that vary sharply between neighbor variables, but corrects global errors across distant graph regions incrementally through long-range message propagations. We propose Hierarchy-GBP (H-GBP), an iterative, two-stage framework that accelerates GBP by first solving these global errors with a coarse graph approximation (abstraction) and projecting the results back to the original graph (recovery), then refining the remaining local errors with GBP. We prove H-GBP convergence to the optimum by deriving the combined matrix operator of our abstraction and recovery steps and analyzing its spectral radius. Experiments on linear sparse graphs show that H-GBP converges fundamentally faster than standard GBP. Moreover, we validate H-GBP on two important spatial problems: Pose Graph Optimization (PGO) and Bundle Adjustment (BA). H-GBP markedly accelerates large-scale PGO and achieves state-of-the-art runtime across all tested BA scales.
[CV-138] Visual-Invariance-Augmented Feature Optimal Alignment for Transferable Adversarial Attacks against Closed-Source MLLM s
链接: https://arxiv.org/abs/2610.06977
作者: Xiaojun Jia,Simeng Qin,Yiming Li,Jie Liao,Sensen Gao,Ke Ma,Yang Liu,Xiaochun Cao
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:
Abstract:Multimodal large language models (MLLMs) remain vulnerable to transferable adversarial examples, especially in black-box settings where only open-source surrogate models are accessible. Existing targeted transfer attacks mainly align adversarial and target samples using global image-level features, such as encoder [CLS] embeddings. However, such coarse alignment insufficiently exploits patch-level visual structures, limiting transferability across heterogeneous closed-source MLLMs. We propose IAU-FOA, a visual-invariance-augmented feature optimal alignment attack with adaptive unbalanced transport, to improve targeted transferability against closed-source MLLMs. IAU-FOA aligns adversarial and target samples at both global and local levels: a cosine-based objective narrows their global semantic gap, while patch tokens are clustered into compact local patterns and matched through optimal transport for fine-grained feature alignment. Balanced optimal transport enforces fixed marginal masses even for local clusters without reliable counterparts, potentially introducing misleading alignment gradients. We therefore introduce confidence-adaptive unbalanced transport to relax these constraints for weakly matched clusters, aiming to reduce unreliable local alignment and improve adversarial transferability. We further study the effect of input transformations and propose visual-invariance augmentation, which applies bidirectional pixel-intensity rescaling and per-channel white-balance adjustment to simulate exposure, contrast, illumination, and color-temperature variations. This strategy encourages adversarial perturbations to generalize across different visual encoders. Extensive experiments on open-source and closed-source MLLMs show that IAU-FOA consistently outperforms state-of-the-art transferable attack methods. Code is available at this https URL.
[CV-139] Event Cameras for Melt-Pool Monitoring in Additive Manufacturing: A Benchmark and a Cross-Machine Transfer Analysis
链接: https://arxiv.org/abs/2610.06973
作者: Mohamad Yazan Sadoun,Sarah Sharif,Yingtao Liu,Zahed Siddique,Yaser Mike Banad
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Melt-pool monitoring is central to qualifying metal additive manufacturing (AM), yet no public event-camera benchmark exists for this domain. Event cameras report per-pixel brightness changes with microsecond timing instead of reading full frames, giving the temporal resolution AM transients demand at a fraction of the data rate. We present SynAM-E (Synthetic AM Events), the first public multi-source simulated event-camera benchmark for metal-AM melt-pool monitoring: 85 physics-calibrated event shards from 15 sources across 8 institutions, with public baselines and fixed cross-machine evaluation splits. On a single-machine case study, event-spatial monitoring matches dense-frame accuracy (0.874 versus 0.863 macro-F1), and the absolute intensity that events discard adds only +0.006 under fusion. On the NIST Additive Manufacturing Metrology Testbed (AMMT) build, a near-sensor event-rate counter recovers a raw-frame-confirmed 528.7 Hz intensity oscillation at ~380 times less sensor readout than the frame stream requires. A compact 93 k-parameter spiking model runs at 15 times lower modeled inference energy for a 0.073 macro-F1 cost. Every cross-source task includes a built-in trust test against camera identity shortcuts: process-type classification passes while material classification remains confounded by camera band, a corpus-structural limitation the release documents and the trust test exposes.
[CV-140] BoT-Feedback: Grounding Multimodal Reasoning in Biomechanical Evidence for Explainable Human Action Feedback
链接: https://arxiv.org/abs/2610.06972
作者: Xu Dong,Wanqing Li,Anthony Adeyemi-Ejeye,Andrew Gilbert
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in visual understanding and multimodal reasoning, yet they remain fundamentally limited in Human Action Feedback Generation. Existing methods infer coaching feedback directly from visual observations, producing generic advice, limited interpretability, and physically implausible hallucinations. In contrast, expert human coaches diagnose performance through explicit biomechanical reasoning over joint kinematics, posture, and body dynamics. We introduce BoT-Feedback, a framework that grounds MLLM reasoning in structured biomechanical evidence. Our key contribution is Biomechanics of Thought (BoT), a four-stage reasoning framework that progressively identifies the action, localises the critical body regions, analyses quantitative biomechanical differences between expert and student performances, and synthesises interpretable coaching feedback. To support this reasoning process, we develop a plug-and-play Biomechanical Data Parser (BDP) that converts videos into structured biomechanical descriptors and an alignment strategy that temporally matches expert and student motions. We further introduce BiomAF, a benchmark containing paired teacher-student videos, 3D skeletons, biomechanical attributes, and expert-coaching annotations. Experiments across twelve open- and closed-source MLLMs demonstrate that grounding reasoning in biomechanical evidence consistently improves feedback quality, interpretability, and robustness while substantially reducing biomechanical hallucinations. BoT-Feedback improves the average expert evaluation score from 2.07 to 2.95 (+40%), enabling compact open-source MLLMs to approach the performance of substantially larger proprietary systems for explainable action feedback generation.
[CV-141] DistScene: Object-to-Scene Distillation for 3D Scene Generation
链接: https://arxiv.org/abs/2610.06960
作者: Kunming Luo,Hongyu Yan,Ken Deng,Chengcheng Zhou,Tianyu Liu,Haipeng Li,Haibin Huang,Xuelong Li,Ping Tan
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Project page: this https URL
Abstract:We present DistScene, a framework for single-image compositional 3D scene generation by jointly modeling the environment and individual objects. Unlike existing methods that represent scenes primarily as collections of objects, we model the environment as an explicit scene component to provide geometric context for object placement. Specifically, we introduce Scene-Frame Generation, which jointly generates separate environment and object components in a shared coordinate frame, allowing their geometry and relative placement to be learned together. Then we introduce Object-Centric Refinement to refine each object in a local frame with scene context. Finally, we develop Object-to-Scene Distillation to transfer pretrained object-generation priors to scene generation through automatically composed and rendered synthetic scenes. Evaluations on indoor and outdoor benchmarks demonstrate improved scene-level spatial coherence over the evaluated baselines. Project page: this https URL
[CV-142] ROMA: LLM System for Real-World Object-Centric Multi-Sensory Active Perception
链接: https://arxiv.org/abs/2610.06955
作者: Ruoxuan Feng,Yutong Chen,Ruihua Song,Huan Yang,Zhongyuan Wang,Guocai Yao,Di Hu
类目: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Humans inherently understand the physical world through an active process. When sensory evidence is insufficient to infer physical properties, we naturally interact with the environment by deciding what information is missing, how to acquire it, and when sufficient evidence has been obtained. In stark contrast, existing multi-sensory robot systems mainly integrate sensory inputs rather than actively acquiring missing evidence through interactions. In this work, we introduce ROMA, an LLM-based system for Real-World Object-Centric Multi-Sensory Active Perception. ROMA integrates vision, audio, tactile, and force sensing into a reasoning-interaction-feedback loop. The model identifies missing evidence and determines the target objects, interactions, and modalities, while a physical interface executes the selected interactions and collects the multi-sensory feedback. To support this capability, we construct ROMI-2K, a large-scale real-world multi-sensory object interaction dataset covering nearly 2,000 objects and 6 atomic interactions with synchronized sensory feedback. Building on these data, we develop a two-stage training framework that aligns sensory modalities and equips the LLM to assess evidence sufficiency, select informative interactions, and reason over the multi-sensory feedback. We further characterize active perception as perception chains, where acquired evidence guides subsequent interactions and reasoning, and establish ROMA Bench to evaluate single-attribute, long-horizon multi-attribute, and intent-driven active perception. Experiments show that ROMA can actively acquire missing evidence and solve complex, long-chain multi-sensory perception tasks that existing methods struggle to handle, laying a strong perceptual foundation for active multi-sensory embodied agents.
[CV-143] Beyond the Linear Representation Hypothesis: Non-Linear Activation Steering in Text-to-Image Models
链接: https://arxiv.org/abs/2610.06945
作者: Muhammad Atif Butt,Paweł Skierś,Joost Van De Weijer,Kamil Deja
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Mechanistic interpretability often relies on the Linear Representation Hypothesis (LRH), which assumes that high-level concepts are encoded as linear directions in activation space. Yet a natural visual concept does not necessarily require a linear visual transition: between sunny and stormy lies an intermediate weather state such as a sky with a few white clouds, not simply a weaker storm; between a caterpillar and a butterfly, the progression is not a caterpillar with continuously growing wings. This raises the question of whether such true intermediate states are also represented nonlinearly by the model. Indeed, when we prompt text-to-image models directly for intermediate attributes, their activations rarely fall along the straight direction connecting the endpoints. Therefore, we propose KANSteer, which models concept traversal as a curve passing through its intermediate states. Seeking a representation that is both simple and interpretable, we propose to use Kolmogorov-Arnold Networks (KANs), which provide a one-dimensional coordinate whose learned functions define the trajectory. This allows the steering direction to vary along the concept while preserving an interpretable representation. Across several concepts and text-to-image diffusion transformers, we find that their activation trajectories substantially deviate from straight lines, and that KANSteer provide a closer fit and smoother traversal of intermediate attributes than linear steering.
[CV-144] UniPro: Unified Multi-Mode Medical Image Segmentation from 2D Images to 3D Volumes via Propagation
链接: https://arxiv.org/abs/2610.06938
作者: Bangwei Guo,Yunhe Gao,Meng Ye,Yang Zhou,Difei Gu,Guoning Zhang,Leon Axel,Dimitris Metaxas
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Medical image segmentation remains fragmented along two axes: segmentation paradigms and data dimensionality. Existing methods are typically developed separately for semantic, in-context, and interactive segmentation, and are further specialized to either native 2D images or 3D volumetric data. In clinical practice, however, segmentation workflows take many forms: a case may be initialized by semantic prediction, reference-guided segmentation, or user interaction. Regardless of how it begins, fine-grained refinement is naturally performed on 2D views; for volumetric scans, such 2D edits must propagate coherently to the rest of the volume. We present UniPro, a unified model that bridges segmentation paradigms and data dimensionality, using propagation to extend 2D segmentation to 3D volumes. Our key insight is that volumetric propagation and in-context segmentation share the same reference-conditioned prediction mechanism, differing only in whether the reference image-mask pairs come from other cases or from previously segmented neighboring slices. Building on this view, UniPro supports semantic, in-context, interactive, and propagation-based segmentation within a single slice-based framework, using class priors, reference exemplars, user clicks, and neighboring-slice predictions as mode-specific conditioning inputs. To improve propagation reliability, UniPro further incorporates bidirectional and 3D supervision to regularize slice-wise propagation beyond per-slice losses. Extensive experiments across diverse modalities and anatomies show that UniPro achieves strong performance across all segmentation settings, enabling annotation-efficient 3D segmentation from sparse 2D initialization and reducing slice-by-slice correction effort.
[CV-145] RADC: Risk-Aware Dual Caching for Vision-Language Test-Time Adaptation
链接: https://arxiv.org/abs/2610.06932
作者: Siyu Huang,Yueyong Chen,Xuejiao Li,Jun Zhou
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Cache-based test-time adaptation (TTA) for vision-language models is often hindered by background bias in global representations and unreliable entropy-based cache admission under representation variations. To address these limitations, we propose RADC, which enhances prototype learning through reliable dual caching. RADC introduces a Semantic Foreground Cache that aggregates category-consistent spatial evidence from CLIP representations, yielding foreground prototypes that complement the global cache while mitigating background interference. To reliably manage both caches, Gaussian Risk Admission models multi-view representations as diagonal Gaussian distributions and jointly considers class separation and feature uncertainty to prioritize reliable cache candidates. RADC integrates zero-shot logits with complementary global- and foreground-cache predictions for robust inference. Extensive experiments on cross-domain and out-of-distribution benchmarks demonstrate consistent state-of-the-art performance.
[CV-146] Anchor Divergence for Semantic Geometry in Contrastive Learning
链接: https://arxiv.org/abs/2610.06919
作者: Akash Kannan,Kiho Park,Victor Veitch
类目: Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Machine Learning (stat.ML)
备注: Code is available at this https URL
Abstract:This paper concerns how semantic context determines geometry in learned vector representations. Similarity is typically measured using cosine similarity, which provides a single fixed geometry. Semantic similarity, however, is inherently context dependent: two images may be similar because they depict the same object, share a visual style, or are relevant to the same clinical finding. We show that contrastive representations naturally encompass a family of geometries that can be specialized to particular semantic structure. The key idea is to use an interplay between contrastive learning, exponential families, and information geometry to establish a correspondence between probability distributions over “anchors” and Bregman geometries on the representation space. We use this correspondence to define “Anchor Divergences”, a method for specifying context-specific semantic geometries on fixed representations. Under this correspondence, modeling the anchor distribution models the geometry itself. Experiments on retrieval show that anchor divergences provide an effective and efficient way to specify context-specific semantic similarity.
[CV-147] Medical Image Alignment Assessment as a Test of Generalist Visual Reasoning in Frontier Multimodal Models
链接: https://arxiv.org/abs/2610.06896
作者: Ross Callaghan,Niannu Gao,Hojjat Azadbakht,Hui Zhang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Frontier multimodal large language models (MLLMs) are increasingly positioned as general purpose visual reasoners as part of the quest for artificial general intelligence. A key test of this generality is whether they can perform novel visual judgments that humans can make reliably from visual evidence and task instructions, without task-specific parameter optimisation. We investigate this question through the task of medical image alignment assessment, where the goal is to establish whether there is anatomical correspondence between two images. Human visual assessment of image alignment is still the gold standard and most common approach; however, it requires trained operators and is impractical to scale for large datasets. We evaluate recent generations of MLLMs on two exemplar medical image alignment tasks, varying both prompting strategies and image-presentation methods. We compare against a locally fine-tuned MLLM and a task-specific CNN to examine the trade-off between frontier general purpose models and smaller models that require specific task optimisation but can be used locally. We show that are reaching an inflection point, where frontier MLLMs can now perform effective visual assessment of medical image alignment. Models released only a few months ago generalise poorly and, in some settings, perform barely above chance, whereas GPT-6 achieves over 85% across almost all scenarios tested. Fine-tuned local models can match or exceed frontier-model performance on the tasks on which they are trained, but transfer substantially less effectively to unseen settings. These findings identify medical image alignment as a useful test bed for generalist visual reasoning and suggest that frontier multimodal models are beginning to acquire capabilities that could support a common quality-control mechanism across heterogeneous medical-imaging pipelines.
[CV-148] EMPEST: Temporal Embeddings for Scalable Driver Identification via Angular Margin Learning
链接: https://arxiv.org/abs/2610.06855
作者: Kyle Musgrove,Dylan B. Lewis,Sarah Powers,Emma J. Reid,Hector Santos-Villalobos
类目: Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Scalable driver identification requires embedding models that maintain discriminative performance as fleet size grows, yet existing triplet-loss formulations degrade rapidly with driver pool size and overfit to session-specific patterns under rigorous temporal evaluation. We introduce TEMPEST, a Temporal Convolutional Network embedding model trained with an additive angular margin (ArcFace) loss that enforces global class-level separation in a normalized angular space. TEMPEST maps 60-second multimodal driving windows to compact 96-dimensional embeddings, supporting truly dynamic enrollment without any retraining or classifier refitting. Under rigorous temporal evaluation on a 45-driver dataset, TEMPEST achieves 91.71% Rank-1 accuracy, outperforming the best classical model by 17.9 pp and the strongest triplet-loss baseline by 58.4 pp. TEMPEST degrades by only 4.3 pp when growing the subject pool from 10 to 45 drivers, compared to 22 pp and 32.5 pp for supervised and unsupervised triplet-loss baselines, and its cross-session advantage is corroborated on the public KIA Soul dataset, where it outperforms the best classical model by 7.3 pp within-session and 14.3 pp cross-session. With 720K parameters, a 2.80 MB footprint, and 50-epoch convergence, TEMPEST establishes a rigorous, reproducible baseline for scalable behavioral driver biometric identification.
[CV-149] How Many Independent Samples Does a Satellite Image Contain? Generalization Bounds for Spatially Dependent Data
链接: https://arxiv.org/abs/2610.08227
作者: Robin Young
类目: Machine Learning (stat.ML); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注:
Abstract:Machine learning classifiers for remote sensing imagery are typically evaluated as though every pixel were an independent sample. Spatial autocorrelation violates this assumption, since neighboring pixels carry redundant information which inflates sample sizes. How many independent samples does a satellite image actually contain? For an n \times n image whose spatial correlation persists over a range of r pixels, the effective sample size is \Theta(n^2/r^2) , not n^2 . We prove this as a finite-sample upper bound for classifiers on spatially correlated data, and show via a matching lower bound that the rate is tight, and no algorithm can do better. We extend the results to images with directional correlation and spatially varying correlation structure. Our result justifies spatial cross-validation since block holdout with separation proportional to the correlation range achieves optimal generalization guarantees, while random holdout can underestimate confidence interval widths by a factor proportional to r . We validate the theory on synthetic data and satellite image tiles from three sensors (Landsat 8, Sentinel-2, and Sentinel-1).
人工智能
[AI-0] Agent in a Bottle: Can LLM Agents Turn Their Capabilities Into Cheap Scalable Artifacts?
链接: https://arxiv.org/abs/2610.08775
作者: Ankit Sonthalia,Haritz Puerto,Alexander Rubinstein,Martin Gubri,Seong Joon Oh
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Large language models (LLMs) can solve many narrow tasks, but querying them separately for millions of related instances can be prohibitively expensive. Can LLM agents autonomously create cheaper solutions for such workloads? We call this ability “bottling”: the ability to turn general capabilities into task-specific solutions that balance answer quality and amortised cost. We introduce BOTTLED, a benchmark in which agents receive an entire unlabelled workload and must complete it under fixed time, compute and LLM API budgets. Agents choose their own approach, such as training a small model or writing a reusable program. Across ten models and three tasks, we find that strong zero-shot task performance does not reliably translate into strong bottling capabilities. Models with similar zero-shot scores can differ substantially after bottling, and 48 of 60 bottling runs score below the lower bound of the 95% confidence interval of their model’s zero-shot performance. Moreover, 31 of 60 runs underperform the stronger of two small-model distillation baselines with the same token budget. Nevertheless, bottling can yield substantial savings: on query-product relevance classification, Opus 5 retains about 82% of its zero-shot macro-F1 at roughly 657 times lower reported cost. Bottling is also competitive with Jev, a “system one” model built especially for cheap, repetitive inference: Opus 5 on the same task recovers about 94% of Jev’s macro-F1 at a quarter of Jev’s projected full-workload cost. BOTTLED provides a basis for evaluating and improving agents’ ability to invest limited resources in reusable solutions for large, repetitive workloads.
[AI-1] VeriFine: Scaling Verification for Self-Improvement in Embodied Reasoning
链接: https://arxiv.org/abs/2610.08761
作者: Zewei Zhou,Rachel Luo,Yulong Cao,Chaowei Xiao,Chensheng Peng,Boyi Li,Thomas Tian,Zheng Lian,Yan Wang,Jiaqi Ma,Boris Ivanovic,Marco Pavone,Wenhao Ding
类目: Artificial Intelligence (cs.AI); Robotics (cs.RO)
备注: Project Website: this https URL
Abstract:Self-improving policies continually expose new failure patterns, changing what their judges must be able to verify. However, current fixed judges constrain both optimization feedback and the discovery of useful training examples, limiting further self-improvement. This challenge is even more acute in embodied reasoning, where reliable evaluation must account for spatial grounding, causal reasoning, and safety-aware decision-making. We introduce VeriFine, an agent harness framework that scales verification through the co-evolution of the policy, training curriculum, and judge. The Policy Improvement Loop uses a rubric judge to diagnose recurring failures, construct an adaptive curriculum, and optimize the policy. When progress plateaus and verification becomes a bottleneck, the Judge Improvement Loop selectively queries human guidance on informative failure cases and refines the judge through coactive calibration, in which humans and agents resolve disagreements and converge toward the objective rubric of physical reasoning. The revised judge then guides the next stage of data selection and policy optimization. Experiments on driving and robot navigation tasks demonstrate continuous self-improvement in both policy and judge capability across reinforcement and supervised fine-tuning. These results show how scaling verification supports continuous self-improvement as policy failure patterns evolve.
[AI-2] Reinforcement Learning with Conformal Action Sets: An Application to Sequential Recommendation
链接: https://arxiv.org/abs/2610.08743
作者: Wenwen Si,Honghao Wei
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:
Abstract:Sequential recommenders typically use a fixed slate size even though the number of useful alternatives changes within a session. We propose Reinforcement Learning with Calibrated Pruning (RLCP), which adapts the retained action set using critic scores and an online threshold. The threshold is updated from binary feedback indicating whether the set contains an action in a proxy target. We prove a deterministic bound on the observed proxy miss rate along adaptive trajectories. To quantify the effect of pruning on reward, we derive an exact decomposition of value loss into filtering and selection losses. Under explicit proxy and critic approximation conditions, this decomposition yields a finite session reward bound that also accounts for imperfect selection and set truncation, without requiring the learning parameters to converge. Experiments on KuaiRand-Pure and MovieLens 1M compare two RLCP implementations with four RL baselines. In each of the 19 configurations, at least one RLCP variant achieves the highest catalog diversity, reaching 1.11\times to 5.21\times that of the strongest baseline, with competitive session depth and no larger retained sets.
[AI-3] EgoLAP: Learning from Egocentric Human Data through Language-Action Reasoning
链接: https://arxiv.org/abs/2610.08726
作者: Lihan Zha,Shresth Grover,Tenny Yin,Samuel M. Bateman,Hengkai Pan,Mengchao Zhang,Aykut Onol,Allen Z. Ren,Dhruv Shah,Anirudha Majumdar
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI)
备注: Project website: this https URL
Abstract:Egocentric human data offer a path to scaling robot learning beyond costly robot demonstrations, yet the embodiment gap makes raw human trajectories a poor supervisory target for control. Our key insight is that, although low-level actions are embodiment-specific, their underlying motion intent can capture task-relevant structure that transfers across humans and robots. We introduce EgoLAP, a VLA pre-training framework that jointly learns from human and robot trajectories through a shared language-based action chain-of-thought. EgoLAP expresses motion intent as structured, temporally abstracted language actions and pairs them with motion-level reasoning grounded in scene geometry, physics, and object affordances. Across extensive real-world and simulated experiments, EgoLAP transfers human experience to robot control more effectively than alternative action representations and reaches 80.1% mean real-world task progress, a 2.3x performance gain over alternative action representations. Motion-level reasoning also outperforms a composite reasoning format that combines subtask, object-box, and visual-trace reasoning.
[AI-4] Does an Agents History Tell You When Compaction Will Hurt? A Modest Bounded Effect on the TRACE Paired-Replay Corpus NEURIPS2026
链接: https://arxiv.org/abs/2610.08722
作者: Egor Pakhomov,Erik Nijkamp
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: Accepted at the IAB Workshop (Interpreting Agent Behavior) at NeurIPS 2026 (non-archival). 20 pages
Abstract:Many long-horizon agents compact their context on a global rule, usually a token budget, blind to what the agent was doing. We ask whether the agent’s recent behaviour predicts when a compaction will hurt. TRACE’s public corpus of 590 harness-triggered AppWorld compaction boundaries replays each boundary from a re-executed prefix state under the pre-compaction context and under the summary, and records the burden of the next actions: calls that error or repeat a call already made. We find that pre-boundary history predicts post-compaction harm only weakly. An internally prespecified contrast by prefix placement is a wide null, and the naive “has-written” label behind it turns out to measure trajectory phase. The best extension-protocol trigger reaches held-out AUROC 0.66 (0.64 on the replicate’s own label) against a same-boundary replicate of 0.72; the best frozen, interpretable trigger avoids 21% of harmful (positive-burden) boundaries while keeping 84% of compaction opportunities, and exceeds the random-rule expectation on count but not on burden mass (a post hoc comparison). Whether the best trigger beats a token-budget rule at matched retention cannot be evaluated on the release. We state what corpora should ship to answer it.
[AI-5] WorldSolver: Can LLM Agents Simulate the Physical Dynamics via Solver Generation?
链接: https://arxiv.org/abs/2610.08720
作者: Siru Jiang,Yongzhe Lyu,Shuo Lu,Yubin Wang,Yuxiang Zhang,Yue Liao,Bin Wang,Jian Liang,Tieniu Tan
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:LLM-based agents are increasingly advancing scientific and engineering problem solving, with physics simulation emerging as a challenging yet practical testbed for reproducing complex physical phenomena with application in embodied AI, games and films. As the workhorse of such simulation, a solver computes how the state of a dynamic system evolves over time. Building such solvers requires physical understanding to identify appropriate models, mathematical reasoning to formulate the underlying dynamics, and software engineering to implement them as executable code, yet this capability of LLM agents remains underexplored. To this end, we introduce WorldSolver, a benchmark of 168 simulation tasks derived from physical phenomena in 61 classic computer graphics papers, spanning 7 physical domains. Each task contains a code scaffold that provides a fixed simulation environment for the scene, with the solver implementation left for the agent to complete. Specifically, we evaluate them along three dimensions: Execution Checks for successful execution, Visual Fidelity for reproducing the intended dynamic behavior in the rendered simulation, and Physical Plausibility for physics-grounded verification of the generated dynamics. Experiments on frontier agents reveal that producing executable solvers is difficult itself, and satisfying visual and physical correctness is even harder. GPT-5.6-Sol and Claude-Opus-5 perform comparatively better than the other evaluated agents, yet achieve overall scores of only 48.7% and 46.7%, respectively. WorldSolver is an early step toward agentic solver generation, and we hope it helps drive progress toward agents that can faithfully simulate the dynamic physical world. Code is available at this https URL.
[AI-6] nanoMuse: An Open-Source Personal Agent for Every Device You Own
链接: https://arxiv.org/abs/2610.08699
作者: Guangyi Liu,Yong Liu,Jiangning Zhang
类目: Artificial Intelligence (cs.AI)
备注: 12 pages, 6 figures, 2 tables. Project page: this https URL ; Code: this https URL
Abstract:Assistants from 2011 answered and waited, and agents from 2023 did a task and stopped. In September 2026 Meta’s Muse showed an agent for one person, with accounts, devices, memory and a conversation that lasts, closed, in a vendor’s cloud, in one country. Such an agent is expected to act on a person’s accounts and devices, remember them across weeks, speak first when it is worth it, and answer for what it did. It is a kind of software, not a model, and until now had no open counterpart. This report defines the personal agent in five questions and three horizons. It reads how Muse is built from Meta’s public record and a copy of its production prompt, each statement marked by its source. It then presents nanoMuse, the open-source counterpart under the GPL-3.0, one agent on every device a person owns, with hands on the phone’s screen and the computer’s. They share one conversation over a relay anyone can run; every action goes through a Sentinel, memory is files the person can read, and the model is their choice. Its size and cost are given as estimates. What is open, memory with provenance, an evaluation suite for the hands and an open model for them, is set out as a roadmap.
[AI-7] ScienceClaw: Benchmarking Continual Self-Evolution of AI-for-Science Agents Across the Natural and Social Sciences
链接: https://arxiv.org/abs/2610.08691
作者: Mingda Zhang,Wenjin Liu,Tiesunlong Shen,Zikai Xiao,Zhenghong Lin,Qing Xu,Erik Cambria,Xiaoying Tang,Haoran Luo
类目: Artificial Intelligence (cs.AI)
备注: 28 pages
Abstract:Large language model agents are accelerating scientific automation, yet verified executions rarely become persistent program-level improvements, and existing evaluations do not examine this process across sequential tasks in both the natural and social sciences. We formalize ScienceClaw as fixed-parameter program self-evolution that unifies task solving, scientific verification, and program updates. ScienceClaw-Eval spans 23 disciplines and measures scientific correctness, evolutionary gain, retention, cross-dataset transfer, and evolution cost through sequential streams and independent reset evaluation. Our framework repairs executable workflows through multi-turn interaction, converts re-execution-verified failure–success trajectories into linked Skill and Operator candidates, and retains an update only when source-task replay reproduces the repair and independent scientific tasks improve. Code is available at this https URL.
[AI-8] Coupled but Late: Turn-Taking Between Full-Duplex Speech Models in Unscripted Dialogue
链接: https://arxiv.org/abs/2610.08683
作者: Lichen Zhu,Yueqian Lin,Yiheng Wang,Hai “Helen” Li,Yiran Chen
类目: Artificial Intelligence (cs.AI)
备注: The paper is under review
Abstract:Full-duplex speech models are trained to converse with a person, but they are increasingly made to converse with each other, in self-play data generation, agent societies, and model-based evaluation. In that loop no human absorbs a timing error: each model’s turn-taking is the other’s input. We ask what timing the loop settles into. Two PersonaPlex-7B instances exchange audio tokens on a shared clock in unscripted conversation, and one floor-transfer rule is applied to them and to Switchboard. Their timing is coupled: re-pairing speakers across conversations destroys it. But the floor changes hands late, at a median of 400-560 ms against 137 ms for humans, and the last 120 ms of the partner’s turn, where human projection places a tenth of its transfers, holds 1% of theirs. Delaying one direction of the channel shifts the response one-for-one and leaves the run-up to it empty, consistent with a reactive wait after the perceived end rather than the turn-end projection human timing requires.
[AI-9] Secure Speculative Decoding for Large Language Models
链接: https://arxiv.org/abs/2610.08678
作者: Yichi Zhang,Zhiqi Wang,Neil Gong,Yuchen Yang
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: 18 pages, accepted by IEEE SP 2027
Abstract:Speculative decoding accelerates inference for a large language model (LLM), referred to as the \emphtarget model, by first using a smaller model, referred to as the \emphdraft model, to generate candidate tokens and then verifying them with the target model for acceptance or rejection. Prior studies primarily focused on the efficiency-utility trade-off of speculative decoding, e.g., lossy speculative decoding, leaving its security implications largely unexplored. In this work, we bridge this gap by providing the \emphfirst systematic study of the security implications of speculative decoding. Through a large-scale measurement study, we reveal a pronounced security-utility asymmetry: across a wide range of lossy speculative decoding methods, improvements in inference efficiency come at a disproportionately high cost to security, with attack success rates for jailbreak and prompt injection attacks increasing much faster than utility degrades. We then propose SecureSD, a new theory-guided speculative decoding method that enhances security while maintaining efficiency and utility. Specifically, our theoretical analysis reveals that security degradation primarily originates from the early tokens generated by the draft model. Motivated by this insight, SecureSD applies a stricter verification criterion to draft-model tokens at early decoding positions. Extensive experiments on both security and utility benchmarks demonstrate that SecureSD significantly improves security while preserving efficiency and utility compared to existing speculative decoding methods. Comments: 18 pages, accepted by IEEE SP 2027 Subjects: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Machine Learning (cs.LG) Cite as: arXiv:2610.08678 [cs.CR] (or arXiv:2610.08678v1 [cs.CR] for this version) https://doi.org/10.48550/arXiv.2610.08678 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[AI-10] MemFLoRA: Memory-Floor LoRA for CNN Adaptation at the Edge
链接: https://arxiv.org/abs/2610.08669
作者: Mehmet Emre Akbulut,Johannes Geier,Ulf Schlichtmann
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: Accepted at the 32nd Asia and South Pacific Design Automation Conference (ASP-DAC 2027), January 25-28, 2027, Tokyo, Japan. Code: this https URL
Abstract:On-device learning is necessary when the model encounters user-,sensor-, or environment-specific shifts after deployment. Although parameter-efficient fine-tuning (PEFT) methods, particularly Low-Rank Adaptation (LoRA) variants, enable efficient adaptation at the edge, the limiting resource for Convolutional Neural Network (CNN) adaptation is often not the number of trainable parameters but the activation state that must be retained until the backward pass. This paper introduces Memory-Floor LoRA (MemFLoRA), a low-rank CNN adapter built around a memory-first design principle rather than a direct application of transformer-oriented LoRA. Instead of merely reducing trainable weights, we define an activation-memory-floor criterion: trainable backward computations must not depend on full-width layer inputs. The resulting adapter freezes the down-projection, trains a scale-matched up-projection, and combines eval-mode backbone normalization with activation-minimal backward rules, reducing saved state to the low-rank branch. Evaluated on three Human Activity Recognition (HAR) datasets and two CNN backbones under subject, body-location, and sensor-placement shifts, MemFLoRA reduces saved-activation memory by 98.5-98.7% and peak training-state memory by 94.9-97.3% relative to full fine-tuning, while matching or exceeding CNN PEFT baselines.
[AI-11] Semantic Behavioral Watermarking: Paraphrase-Robust and Forgery-Resistant Provenance for LLM Agents
链接: https://arxiv.org/abs/2610.08668
作者: Suxin Ji,Hungtao Wan,Shaoxuan Chen,An Zhang
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
备注:
Abstract:Behavioral watermarking embeds an owner identifier in an LLM agent’s high-level action choices, giving provenance without touching output tokens. Prior agent watermarks break in two ways. First, all three prior schemes bind the watermark to the exact action symbol, so renaming a tool desynchronizes decoding even when the observation is untouched; in AgentMark’s own robustness test, paraphrasing the observation alone drops bit-recovery to 16.8%. Second, every prior agent watermark studies only removal: none asks whether an adversary can forge a trajectory that verifies as someone else’s, a question answered affirmatively for text watermarks (Jovanović et al., 2024). We present Semantic Behavioral Watermarking (SBW): watermarking over semantic action clusters under history conditioning, with the public-cluster bin replaced by keyed collision-resistant binning whose fresh-bucket assignment is provably unpredictable in the random-oracle model. Across five agent models (3B-14B, four vendors) and three encoders the ordering holds on both benchmarks: on ToolBench (600 trajectories per model) detection under rewriting is 0.49-0.66 for cluster-level versus 0.05-0.17 for exact-symbol at a permutation-calibrated 1% FPR, at 72-83% choice agreement against 22-27% for logit biasing; on ALFWorld (100 episodes per model) it is 0.92-0.97 versus 0.00-0.01. Keyed binning takes adaptive forgery from 100% to the false-positive floor at the primary operating point (bge, r=64). We also mark the boundary that guarantee does not cover: when the adversary copies the victim’s own steps, shuffled splicing is neutralized (0.000 on Qwen2.5-3B) but chained replay remains at 0.76-0.98 across the five models, reported as open. Paraphrase robustness costs about half of the per-step watermark capacity. Code is available at this https URL.
[AI-12] ParanoiaEval: Benchmarking Unnecessary Defensive Work in Agent ic Coding
链接: https://arxiv.org/abs/2610.08662
作者: Hanjun Luo,Xiucheng Zhang,Zhuoning Xu,Zhimu Huang,Yingbin Jin,Xinfeng Li,Hanan Salam
类目: Artificial Intelligence (cs.AI); Software Engineering (cs.SE)
备注:
Abstract:As coding agents increasingly undertake real-world work autonomously, judging whether their risk treatments are warranted has become important. Existing work evaluates related agent behaviors from separate perspectives, but lacks a systematic framework for unifying these behaviors. To bridge this gap, we introduce ParanoiaEval, the first benchmark for unified evaluation of risk-treatment capabilities in coding agents. Grounded in the well-established Avoidance-Transfer-Mitigation-Acceptance framework in software engineering risk management, ParanoiaEval operationalizes its 4 fundamental treatments for coding-agent settings and contains 200 evidence-controlled repository-level task pairs, each differing only in treatment-defining evidence. We further introduce dedicated metrics for risk-treatment violations and evidence responsiveness, using a human-calibrated agentic judge for reliable evaluation. Large-scale experiments on 8 representative models and a post-hoc human study reveal that (I) unnecessary risk treatment occurs in 11.2%-58.7% of runs despite explicit evidence, with substantial variation across agent configurations; (II) stronger task capability does not ensure more appropriate risk treatment, while treatment violations substantially harm developers’ experience, establishing risk treatment as an independent capability dimension; and (III) agents exhibit systematic patterns consistent with established risk-management findings, suggesting that knowledge from human practice can guide the diagnosis and improvement of this capability.
[AI-13] HygieneRoboBench: Benchmarking Hygiene-Aware Planning for Household Robots
链接: https://arxiv.org/abs/2610.08642
作者: Yurun Chen,Josh Qixuan Sun,Jason Qin,Chengtai Li,Tianyi Wang,Mark Crowley,Wentao Zhu
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI)
备注:
Abstract:Contact with contaminated objects can spread hazards through a household robot’s grippers, tools, and shared surfaces, while new contacts can make an existing plan unsafe. Existing benchmarks do not jointly assess how planners identify hygiene risks from contact history and plan safe continuations after new contact events. Planners must do so within time and resource limits while respecting user priorities. We introduce HygieneRoboBench, with 624 instances across 134 task families, to evaluate safe resolution of household tasks from a given execution history. Tasks capture contamination through two grippers and shared objects, treatment costs, and user priorities. We combine controlled history, profile, and event comparisons with independent plan evaluation. These assess safe resolution, cost efficiency under user priorities, and responses to contact events. Evaluation of LLM-based and symbolic planners shows that safely completing a task does not guarantee the lowest execution costs under the user’s priorities. To address this problem, we introduce Hygiene-NSP. It combines LLM-based grounding, contact-history reconstruction, and CP-SAT to jointly plan hygiene treatment and task execution under user priorities. Hygiene-NSP achieves safe resolution and optimal safe resolution rates of 94.4% and 90.4%, respectively. Both rates are higher than those of the evaluated baseline planners on the full dataset. Project page: this https URL.
[AI-14] Parallel Predictive World Models for Accurate and Efficient Long-Horizon Planning
链接: https://arxiv.org/abs/2610.08627
作者: Wanjin Feng,Baobin Zhang,Ao Yu,Shibo Feng,Xi Wang,Xingyu Gao
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Long-horizon world-model planning typically relies on autoregressive rollouts, where predicted states are repeatedly fed back into the model. This preserves temporal structure but creates a horizon-length sequential path and exposes later predictions to recursive decoded-state feedback. We introduce Parallel Predictive World Models (PPWM), which predict a finite-horizon trajectory in parallel while retaining causal interaction among future representations. Each horizon is conditioned on its causal action prefix, and future representations interact before decoding, separating temporal causality from state-by-state output recursion. We formalize this distinction by viewing autoregressive rollout as a causal trajectory map and identifying the decoded-state feedback pathway removed by PPWM. Across four visual-control tasks, PPWM achieves the lowest long-horizon prediction error and the highest Cross-Entropy Method (CEM) simulator success among the evaluated predictive interfaces. Meanwhile, PPWM achieves more than a 3 \times average CEM planning speedup over the autoregressive LeWM baseline. These results suggest that accurate and efficient long-horizon world-model planning does not require state-by-state autoregression, but can instead be achieved through parallel causal trajectory prediction.
[AI-15] Early Memory Selection for Balanced Adam
链接: https://arxiv.org/abs/2610.08624
作者: Alberto Fernández-Hernández,Cristian Pérez-Corral,Jose I. Mestre,Manuel F. Dolz,Enrique S. Quintana-Ortí
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Machine Learning (stat.ML)
备注: Includes theoretical proofs and reproducibility appendices. Code and data: this https URL
Abstract:We propose a method for choosing the shared memory parameter \beta_1=\beta_2=\beta in Adam from a short pilot training. The selected \beta remains fixed during the subsequent full training. A local model of Adam’s normalized direction balances sampling variability against the delay introduced by averaging past gradients. This balance gives a cubic memory rule, whose two coefficients are estimated from gradient probes at a few pilot checkpoints. The estimator uses the numerator and denominator jointly, preserving their covariance. With a 200-update pilot and sixteen probe gradients at each of four checkpoints, a seed-matched retrospective evaluation on eleven vision and language workloads reduces mean relative validation gap by 40.7% and worst-quarter mean gap by 44.3% against the grid representative of shared \beta=0.95 . The mean gap is also 32.3% lower than that of the best constant \beta chosen across all eleven workloads.
[AI-16] Agent ic RCA for Internet-Scale Services Using Constrained Creativity
链接: https://arxiv.org/abs/2610.08622
作者: Sayan Sinha,Vipul Harsh,B. Aditya Prakash,Vyas Sekar,Hui Zhang
类目: Networking and Internet Architecture (cs.NI); Artificial Intelligence (cs.AI)
备注: 21 pages, including the references and appendix; 10 figures; 4 tables
Abstract:System administrators of Internet-scale services need to resolve failure incidents to maintain reliability of such services. Ideally, we want a troubleshooting system to be: (1) expressive to known and unknown incidents with high accuracy; (2) cost efficient at scale; (3) explainable to provide actionable insights operators can act on; and (4) entail low effort from the operators. Unfortunately, most existing systems, including emerging LLM-assisted agentic workflows and structured frameworks for authoring diverse RCA algorithms fall short of achieving all four requirements. We present E4, a novel agentic system for troubleshooting for Internet-scale services. E4 embodies the paradigm of constrained creativity that combines the best of LLM-assisted automation and exploration with the explainability and efficiency of a structured approach. Instead of allowing an LLM agent to write arbitrary code or generate arbitrary responses, we provide the agent a restricted DSL to generate its response via simple loop-free data flow programs. This DSL, equipped with high level operators for troubleshooting, makes E4’s output accurate, verifiable and explainable. On a mix of synthetic and real-world workloads, E4 achieves up to 62% better accuracy compared to state-of-the-art solutions, while providing more explainable responses at up to 12x reduced cost.
[AI-17] A Swarm-Coordinated Multi-Robot System for Early Stress Detection in Agricultural Rows Using Multimodal Leaf Sensing
链接: https://arxiv.org/abs/2610.08603
作者: Rishi Gupta,Astha Goyal,Vinay Vishwakarma
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI)
备注:
Abstract:Early stress detection in crops is a necessity today to improve efficiency and reduce waste of time, money, and effort. However, most modern techniques, such as hyperspectral imaging and AI-based systems, are too costly and complex for medium and small-scale farmers to implement. This paper showcases CropSentry, a low-cost, ground-based multi-robot system that uses multimodal leaf sensing to continuously monitor crop health by tracking stress levels. The system comprises two autonomous bots that continuously detect leaf color and environmental data row by row. The observations are spatially mapped and sent over to the master bot, which uses color-coded row segments to generate a real-time web-based dashboard displaying crop health. After 63 observations were collected during the experiments, the results showed an overall crop health classification accuracy of 84.12%, with 82.60% for healthy plants, 88% for nutrient-deficient plants, and 80% for diseased plants. Also, 100% wireless communication success rate across 10 slave observations was achieved. Close-range leaf inspection across multiple bots can detect early stress in crops while remaining affordable, accessible, and scalable. It provides farmers with timely information to improve resource utilization and crop management.
[AI-18] One for All All for One: Coordinated Multi-Agent Diffusion Steering via Stochastic Optimal Control
链接: https://arxiv.org/abs/2610.08595
作者: Riccardo Barbano,Vincent Pauline,Runchang Li,George Webber,Alexander Denker,Željko Kereta,Stefan Bauer,Francisco Vargas,Esmeralda S. Whitammer
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI)
备注:
Abstract:Deep generative models often produce structured outputs composed of interacting components. Modelling these outputs with a single model requires learning both the component distributions and their interactions. We pursue a modular alternative: reuse independently trained component generators and learn only how to coordinate them to produce coherent structured outputs. Our framework, Coordinated Multi-Agent Diffusion Steering (CMDS), treats frozen pretrained diffusion models as reusable generative primitives and coordinates their reverse processes through a learned control. We formulate coordination as a stochastic optimal control problem, balancing an assembly-level reward that specifies the desired properties of the combined output against deviations from the pretrained dynamics. The learned control amortises this optimisation, allowing reuse across new task instances. Experiments show that CMDS can recover a known target distribution, satisfy different spatial constraints with the same trained control, and recover individual sources from degraded mixtures. Across multi-agent maze navigation, articulated robot planning, and text-conditioned human motion, CMDS turns frozen models into coordinated multi-agent generators.
[AI-19] MINDSET: Energy-based Schema Evolution for Long Conversational Agent Memory
链接: https://arxiv.org/abs/2610.08586
作者: Sujato Dutta,Sreekruthy Tummala,Shashank Vanga,Ayushmi Pavani
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Long conversational agents have become essential in our daily lives. They must remember what was said long back in order to help us efficiently complete a task without needing the user to repeat instructions and context repeatedly. However, the main issue is that instructions and context change over time and so the agents must be able to adapt accordingly. A useful memory system should preserve both current and historical states, distinguish stale information from active knowledge, retrieve evidence appropriate to the query and avoid repeatedly invoking a large language model to rewrite prior interactions. We introduce MINDSET, a memory controller that stores a conversation as immutable episodes and organizes them into versioned schemas through minimum-energy state transitions. Each incoming episode may reinforce, supersede, split or create a schema. The transition decision balances representation distortion, contradiction, historical damage, fragmentation and internal inconsistency, while hysteresis prevents isolated contradictions from prematurely rewriting stable memory. We evaluate MINDSET against 5 memory systems on a reproducible sample of 850 questions (700 LoCoMo + 150 MemoryAgentBench). MINDSET obtains the highest observed LoCoMo answer F1 while significantly improving retrieval ranking (Recall@8, MRR and nDCG@8) over the second best method LightMem (p0.01 after Holm correction). It obtains the highest observed scores on MemoryAgentBench although the relative difference is low. Ablations identify controlled fragmentation and schema-aware assignment as the largest contributors to answer quality. Additionally, a 700-question cross-model evaluation with GLM-4.7 and Gemma-4-31B supported model independence. These results show that long-term memory can be better handled as constrained state management rather than continual summarization.
[AI-20] How Learning Governs Unlearning across the Memorization-Generalization Spectrum
链接: https://arxiv.org/abs/2610.08577
作者: Hwiyeong Lee,Hyelim Lim,Ingyu Bang,Hoki Kim,Taeuk Kim
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:
Abstract:While unlearning seeks to negate undesired capabilities acquired through learning, little research has examined how the way models learn shapes their subsequent unlearning. In this paper, we investigate this connection from the perspectives of memorization and generalization, the two most representative yet competing strategies that models employ during training. We first classify memorization- and generalization-heavy models using grokking in modular addition and compare their responses to unlearning, showing that the latter suffer greater retain damage, i.e., a larger performance drop on the retain set. Furthermore, we conduct a finer-grained analysis by introducing bucketed modular addition, in which the respective contributions of the two strategies can be explicitly controlled across the memorization-generalization spectrum. In this setup, we reaffirm that the same trend persists and is nearly monotonic. We further demonstrate that this relationship also holds in LLM unlearning across verbatim and factual recall settings. Finally, we provide two practical insights for developing better unlearning methods, highlighting the importance of accounting for learning dynamics in unlearning.
[AI-21] RAG -PIBench: A Leakage-Aware Benchmark for Prompt-Injection Detection in Trustworthy RAG Systems
链接: https://arxiv.org/abs/2610.08571
作者: Niveen O. Jaffal,Ahmet Yuksel,David Mohaisen
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: 19 pages, 3 figures, 8 tables
Abstract:Retrieval-Augmented Generation (RAG) systems are vulnerable to prompt-injection attacks embedded in retrieved content. We introduce RAG-PIBench, a benchmark for RAG-style prompt-injection detection containing 4,876 contextual examples across frozen train, validation, and protected-test splits. Using a leakage-aware construction pipeline and strict evaluation protocol, we compare keyword-based, semantic-reference, TF-IDF, and transformer-based detectors. DistilBERT achieves the best protected-test performance (F1 = 0.896, PR-AUC = 0.968), while TF-IDF SVM and logistic regression remain competitive. Our results demonstrate the value of leakage-aware benchmark design and strong sparse baselines for reliable prompt-injection detection in RAG systems.
[AI-22] Adaptive Power Sampling for LLM Reasoning
链接: https://arxiv.org/abs/2610.08563
作者: Bingnan Xiao,Chenhao Yang,Bingcong Li,Wei Ni,Xin Wang
类目: Artificial Intelligence (cs.AI)
备注: 22 pages, 6 figures
Abstract:Sequence-level power sampling has recently emerged as a training-free approach to reasoning by sampling from a sharpened output distribution of a base large language model (LLM). Nevertheless, existing methods typically sharpen the base model distribution uniformly across queries, overlooking variations in query difficulty and in how well the base model already handles each query. The goal of this work is to equip power sampling with query adaptivity. Theoretically, we show that the benefits of further sharpening are determined by the self-reward gap between correct and incorrect responses. Based on this insight, we propose \emphAdaptive Power Sampling (APS), which adjusts the sharpening exponent on a per-query basis at test time using the relationship between answer agreement and the model’s self-reward. Experiments across diverse reasoning tasks, including MATH500, HumanEval, and GPQA, show that APS consistently outperforms power sampling with a fixed sharpening exponent, without additional training.
[AI-23] AnyBottle: A Recipe to Only Keep the Concepts You Really Need
链接: https://arxiv.org/abs/2610.08552
作者: Wolfgang Stammer,Sukrut Rao,Hevra Petekkaya,David Steinmann,Bernt Schiele
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:
Abstract:Concept bottleneck models (CBMs) make predictions inspectable and intervenable by routing them through human-interpretable concepts, but originally required concept annotations. Annotation-free variants remove this requirement, but typically use large concept vocabularies, static at both training and inference, producing bottlenecks larger than any task or prediction needs and harder to inspect. We propose AnyBottle, a single recipe for building compact, task-specific CBMs. AnyBottle assumes only a frozen backbone and an unsupervised concept pool, such as a sparse autoencoder. A black-box teacher trained on the same backbone then guides selection: each round adds the concept that best explains the bottleneck’s current failures, with candidates restricted to regions of teacher/student disagreement. Trained with nested dropout over this selection order, the final bottleneck predicts accurately from any concept prefix, so inference spends fewer concepts on inputs it is confident about early and more on hard ones. Since no stage is modality-specific, a new domain and task requires swapping only the backbone and concept pool. Across six vision and two text datasets and two teacher paradigms, AnyBottle yields bottlenecks with fewer concepts and higher concept consistency than annotation-free baselines, while staying close to the black-box reference. Overall, AnyBottle shows that going annotation-free need not mean going large: a small, discovered vocabulary can be as expressive as a much larger, fixed one.
[AI-24] Micro Neural Policies for Safe Real-Time Robotic Control
链接: https://arxiv.org/abs/2610.08541
作者: Hongpeng Cao,Riccardo Curcio,Daniele Ottaviano,Marco Caccamo
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI)
备注: 9 pages
Abstract:In this paper, we investigate the synthesis of Micro Neural Policies (MNP) to enable safe and robust real-time robotic control on computationally constrained embedded devices. We demonstrate that integrating Evolution Strategy (ES) and Statistical Model Checking (SMC)-based verification for policy search can drastically reduce neural network size without compromising safety and robustness. We conduct a large-scale training and evaluation of MNP on Cartpole and Quadrotor control tasks, varying control frequencies and network architectures. After validating these policies in simulation, we evaluate their deployability through zero-shot transfer to physical systems. Our experiments show that MNP can successfully achieve safe sim-to-real transfer without sacrificing control performance. We then show that the policies’ memory footprint, ranging from 0.5 to 7.5 kB, allows deployment on microcontrollers, where they achieve real-time inference latency with under 25 ns of jitter while leaving the chip idle for over 97% of the time for additional workloads. This makes them a highly practical solution for severely resource-constrained robotic systems.
[AI-25] From Shared Demand Patterns to Local Uncertainty: Probabilistic Load Forecasting by Mixing Compact Adaptations
链接: https://arxiv.org/abs/2610.08538
作者: Haoran Li,Zhe Cheng,Yang Weng
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:
Abstract:Probabilistic load forecasting has been widely studied for power-system operation and planning, but customer- and transformer-level forecasting introduces a distinct scalability challenge. At these levels, load uncertainty is strongly affected by customer behavior, weather, and mixed load composition, making it difficult for a single shared model to capture heterogeneous patterns. Using separate probabilistic models can improve local accuracy, but becomes costly to train, store, update, and validate at scale. To address this challenge, we develop a scalable customer-aware forecasting framework that learns common demand behavior through a shared model while adapting only a compact subset of parameters. Rather than using an independent model for each load or assigning each load to a specialized model, the proposed design learns a small bank of low-dimensional adaptation components and allows each load to combine them according to its forecasting characteristics. This preserves shared knowledge across customers while providing sufficient flexibility for heterogeneous and mixed load compositions. Experiments on 590 load profiles from the SMART-DS dataset show consistent improvements in deterministic accuracy and probabilistic quality over statistical, neural-network, Transformer-based, and pretrained time-series baselines, while retaining low storage and inference costs.
[AI-26] FlowCF: Sparse Counterfactual Explanations for Mixed-Type Tabular Data using Flow Matching NEURIPS2026
链接: https://arxiv.org/abs/2610.08537
作者: Emmanouil Panagiotou,Eirini Ntoutsi
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: Accepted at the NeurIPS 2026 Geometric Distributional Deep Learning (GDDL) Workshop
Abstract:In the field of Explainable AI (XAI), counterfactual (CF) explanations interpret a model’s decision by suggesting the changes to the input that would lead to a more favourable outcome. To be useful in practice, such an explanation should change few features and change them as little as possible, properties known as sparsity and proximity. We observe that existing methods remain limited in this respect, especially for numerical features, whether they are model-agnostic and amortised, or gradient-based with full access to the model. In this paper, we propose FlowCF, a model-agnostic generative method that frames CF generation as sparse transport from the factual to the target class. We solve this transport with flow matching, which we extend to mixed feature types with a novel mixed flow operator, and exploit the resulting geometry to optimise for sparsity through a gating network that minimises the number of features the transport changes. Extensive experiments on six benchmark datasets demonstrate that FlowCF produces the best numerical sparsity and proximity, changing 29% of the numerical features where the best baseline changes 89%, at 70% smaller displacement, while remaining comparable on the other desiderata.
[AI-27] How Much Evidence Should a Coding Agents Self-Correction Carry? Adaptive Dirichlet Evidence for Self-Distillation
链接: https://arxiv.org/abs/2610.08514
作者: Yunbo Long,Guangya Hao,Yuhan Liu,Yiting Duan,Longyan Tan,Yunchen Long,Hao Wu
类目: Artificial Intelligence (cs.AI); Software Engineering (cs.SE)
备注:
Abstract:Execution feedback lets coding agents revise programs and learn from their own corrections. A correction’s learning weight should reflect both the transitions supported by its executions and the amount of evidence behind that support. We introduce Effective-Evidence Self-Distillation (EESD), which represents these quantities separately. Normalized execution relevance determines relative transition support and an effective pseudo-count mass; a Dirichlet posterior then produces an uncertainty-penalized weight for KL-anchored correction learning. Under a symmetric prior, changing mass preserves category ordering, and effective mass yields a supervised coefficient bounded by its matched fixed-mass counterpart. Across four model-domain history sweeps, increasing visible observations from one to eight reduces future-outcome NLL by 55.0-59.3%. At eight observations, effective mass achieves lower NLL than fixed mass in all four comparisons. In the primary matched DeepSeek/RunBugRun study, argmax predictions agree on all 3,000 examples, with the largest NLL gain under concentrated relevance. After one correction-learning round, DeepSeek/CodeARC all-tests Pass@1 increases from 15.0% to 20.4%, with a paired 95% source-bootstrap interval of [+2.8, +8.0] percentage points. The twelve-setting downstream evaluation establishes the model-domain scope of this update. These results show how separating evidence support from evidence mass changes probability estimation and correction learning in coding agents.
[AI-28] Cylindrical Geodesic Flow Matching for Quasiperiodic Physiological Signal Transformation
链接: https://arxiv.org/abs/2610.08510
作者: Onur Selim Kilic,Afra Nawar,Cem Okan Yaldiz,Michael J. Cho,Ahmet Rasim Emirdagi,Demet Tangolar,Amirali Aghazadeh,Amit J. Shah,Omer T. Inan
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Paired translation between quasiperiodic physiological waveforms (i.e., recovering a target oscillatory signal from the source) is central to the interpretation of cardiovascular signals derived from wearables placed at different body locations. This source-to-target mapping in these problems carries inherent geometric structure: the phase wraps around the cycle and must be treated as a circular variable, the amplitude remains strictly positive, and the beat-to-beat alignment can drift unpredictably across cycles and subjects. While deep neural networks have been used for phase estimation and complex-valued signal modeling, prior work does not explicitly learn phase transport between paired signals. Consequently, neither endpoint-supervised regression nor the standard affine path used in flow matching accounts for this phase–amplitude structure. We introduce \emphcylindrical geodesic flow matching for paired cardiovascular waveform translation. We show that the standard affine path used in flow matching distorts intermediate amplitude and instantaneous frequency when interpolating between quasiperiodic signals; replacing it with a closed-form geodesic on the phase–amplitude cylinder eliminates these artifacts and converts each training pair into dense, geometry-consistent velocity supervision. On zero-shot photoplethysmography and limited-support seismocardiography adaptation benchmarks, our method consistently outperforms interpolation baselines and matches or exceeds direct supervised prediction, reducing Hilbert Transform, L_2 , and Dynamic Time Warping distance by up to \sim15% over the strongest competing baseline. These results suggest that bridge geometry is a critical inductive bias for flow matching on oscillatory signal translation.
[AI-29] X-OPM: Explainable Automatic Digital On-Chip Power Modeling for Enhanced Robustness
链接: https://arxiv.org/abs/2610.08502
作者: Jingbo Jiang,Xizi Chen,Jian Peng,Wei Zhang
类目: Hardware Architecture (cs.AR); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:
Abstract:Proactive power management systems reduce processor dynamic power through runtime power prediction and power-aware scheduling. Accurate, stable and low-overhead digital on-chip power meters (OPMs) are crucial for improving the prediction quality. Recent studies have explored various modeling methods, including using linear models, decision trees, and multi-layer perceptrons (MLPs) to construct OPMs. However, most current approaches train models end-to-end without analyzing the physical interpretability of features, affecting their ability to generalize to unseen workloads. Grounded in the design principles of synchronous digital VLSI circuits, X-OPM introduces a robust feature engineering framework that uses tree-based models to capture feature interactions and linear models for prediction. It also incorporates a human-in-the-loop workflow to balance model accuracy against modeling effort. Evaluated on a commercial C906 vector processor, X-OPM consistently achieves R^2 0.93 across all workloads with sampling window size set below 8 cycles. In contrast, state-of-the-art methods including APOLLO, COBIT, and standard MLPs fail to generalize across all test cases. Layout with commercial EDA tools shows that X-OPM incurs an area overhead below 0.1% , which is on par with lightweight tree-based and linear models, and significantly smaller than MLP-based models.
[AI-30] MetaLearnNCA: Few-Shot Offline Meta-Learning via Interacting Neural Cellular Automata
链接: https://arxiv.org/abs/2610.08479
作者: Etienne Guichard,Stefano Nichele
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:
Abstract:Few-shot meta-learning traditionally formulates task adaptation either as analytical gradient descent through unrolled computational graphs or as metric-based distance comparisons over flattened 1D fea- ture vectors, which either incur costly test-time backpropagation or discard native 2D spatial geometry. In this work, we propose METALEARNNCA, a decentralized framework that achieves few-shot adapta- tion through the dynamical interaction of coupled Neural Cellular Automata (NCAs) without computing analytical gradients during inference. MetaLearnNCA decomposes task adaptation into an Active- NCA, which executes task inference conditioned on a continuous 2D spatial memory grid termed the spatial program, and a learned Meta-NCA, which acts as a decentralized cellular optimizer by diffusing spatial error residuals across local neighborhoods to dynamically update this program. METALEARN- NCA is competitive against canonical meta-learners in-distribution (96.12% on Omniglot) with Out-Of- Distribution transfer gains on MNIST, KMNIST, and Fashion-MNIST transfer across 10 independent testing seeds across 1-, 5-, and 10-shot regimes (e.g., surpassing Prototypical Networks by +10.54% on 10-shot MNIST and a +3.87% gain on 10-shot Fashion-MNIST over FOMAML). Our results establish that robust, gradient-free learning-to-learn can emerge from decentralized cellular dynamics on non-von Neumann substrates.
[AI-31] AssemState: Manual and Physical-State-Guided Reasoning for Zero-shot Furniture Assembly
链接: https://arxiv.org/abs/2610.08446
作者: Zhiyuan Qi,Jierui Li,Yifan Shen,Cheng Qian,Jiateng Liu
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Multimodal large language models (MLLMs) have made significant progress in visual understanding, but precise 3D spatial reasoning integrated with physical environment remains difficult. Furniture assembly requires not only recovering step-level operations from diagrammatic manuals, but also translating semantic attachment relations into 6D pose updates that enable parts to physically interact with the environment and previously assembled components. To study this problem, we propose AssemState, a zero-shot framework for manual and physical-state-guided furniture assembly. It firstly employs anchor-guided boundary assembly states to decompose manual pages into single-part operations and recover an assembly-tree. Then, it uses iterative after-state feedback refinement to guide successive (SE(3)) updates and corrections, and validates their physical plausibility through simulation-based release tests. Experiments show that compared with the strongest prior baseline, AssemState improves F1 from 38.58% to 62.80% and Tree Exact Match from 28.24% to 53.92% for assembly-tree recovery. On 243 independently evaluated part-level operations, our proposed iterative refinement improves judge-accepted operations from 0 to 5.3% and reduces mean Chamfer distance from 5.4111 to 1.7744. However, visually plausible candidate poses may still suffer from collision, floating, mirror-orientation errors, incomplete seating, and wrong-side attachment. These results show that AssemState improves operation-structure recovery and selected local pose metrics, while MLLMs remain limited for spatial relationship reasoning.
[AI-32] EMHO: EMbodied Agent Harness Optimization via Experience Traces
链接: https://arxiv.org/abs/2610.08432
作者: Hyun Jung Lee,Jungtaek Kim,Jongwon Jeong,Tae-Eui Kam,Donghyun Kim,Yong Jae Lee
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Improving embodied agents often focuses on optimizing the underlying model through training, while the surrounding agent harness that controls planning, context, and tool use is typically engineered. We ask whether this harness can instead improve itself directly from experience traces under sparse environmental feedback. We propose EMbodied Agent Harness Optimization (EMHO), a self-evolving framework that keeps the embodied model frozen and iteratively revises its harness by analyzing execution trajectories and prior harness history. EMHO optimizes beyond skills or recovery prompts, modifying how the agent monitors progress, uses vision tools, grounds observations, and responds to failures. To support multiple subtasks with a single harness, we introduce EMHO-Merge, which addresses trade-offs in jointly optimizing a single shared harness across subtasks by using episode-level gains and losses to guide evidence-supported refinement of when and how revised behaviors are applied. We evaluate EMHO on EmbodiedBench across navigation and manipulation tasks, and EMHO consistently improves task success for both Qwen 9B and 27B models. Qualitative analysis shows that EMHO goes beyond recovering from failures and unproductive actions to reshape how the embodied agent interprets and interacts with its environment.
[AI-33] NeMo-DCR: Bit-Exact Delta-Compressed Refit for Scalable Agent ic RL at Trillion-Parameter Scale
链接: https://arxiv.org/abs/2610.08430
作者: Songlin Jiang,Zhiyu Li,Terry Kong,Yu Yao,Youngeun Kwon,Bernard Nguyen,Ashwath Aithal,Mario Di Francesco
类目: Distributed, Parallel, and Cluster Computing (cs.DC); Artificial Intelligence (cs.AI)
备注: The code is open-sourced in NVIDIA NeMo RL PR #2444 at this https URL
Abstract:Agentic reinforcement learning (RL) disaggregates training from rollout, so each policy update must reach the rollout clusters before the next batch. Transferring a full 1T checkpoint for such weight synchronization (refit) takes 87.5 min between two AWS regions. Measurements of BF16 training show that about 1% of weights change their stored values per step. Recent systems exploit this sparsity but fall short on placement, exactness, or efficiency: they reimplement placement rules, assemble full tensors, rebuild values arithmetically, or use a cross-cluster collective, and none fully recovers from mid-refit failures. We present NeMo-DCR (Delta-Compressed Refit), which sends only changes yet is bit-exact: receivers obtain the same parameter and buffer bits as a dense refit. For placement, fixed affine mappings project changes from training shards into the checkpoint’s canonical coordinates, residual conversion covers the other changes, and the serving runtime’s native loader places all changes in receiver storage. For exactness, compressible XOR masks carry affine changes whose projection and loader preserve stored bits, and overwrites carry the others. Receivers apply both in place, retries overwrite partial writes, and a joint commit binds the policy to the baseline for the next delta. For efficiency, object storage or a relay tree streams payloads during delta construction, without a cross-cluster collective. Even at 3% and 5% change rates, NeMo-DCR refits of 30B-1T models are 12-40 \times faster than a transport-only full-checkpoint reference. A 1T relay-tree refit at 3% takes 150 s instead of 87.5 min, making refits practical for cross-cluster agentic RL at trillion-parameter scale. Comments: The code is open-sourced in NVIDIA NeMo RL PR #2444 at this https URL Subjects: Distributed, Parallel, and Cluster Computing (cs.DC); Artificial Intelligence (cs.AI) ACMclasses: C.2.4; C.1.4; I.2.6 Cite as: arXiv:2610.08430 [cs.DC] (or arXiv:2610.08430v1 [cs.DC] for this version) https://doi.org/10.48550/arXiv.2610.08430 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[AI-34] Learning from Failures: A Failure-Driven Prompt Refinement for LLM -Based Vulnerability Analysis
链接: https://arxiv.org/abs/2610.08405
作者: Mandana Ghadamian,David Mohaisen
类目: oftware Engineering (cs.SE); Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR)
备注: Accepted at CSoNet 2026. 15 pages, 6 figures/tables combined
Abstract:Large Language Models have emerged as promising tools for software vulnerability analysis, but their effectiveness depends heavily on prompt design. Existing research primarily compares prompting strategies using aggregate performance metrics, providing limited insight into why models fail or how prompts can be improved systematically. We propose Failure-Driven Prompt Refinement (FDPR), a methodology that analyzes recurring model failures to guide evidence-based prompt refinement. Using the Damn Vulnerable Java Application (DVJA), we identify recurring failure modes, including false positives, false negatives, unsupported reasoning, and CWE misclassification, and translate them into targeted prompt refinements. We then evaluate the resulting prompt on the Juliet Test Suite and perform cross-model validation to assess generalizability. The results show that failure-driven refinement improves the reliability of LLM-based vulnerability analysis while yielding reusable prompt design principles. More broadly, this work demonstrates that recurring model failures provide a principled foundation for prompt engineering, enabling the systematic development of more reliable LLM-based vulnerability analysis systems.
[AI-35] Atom-JEPA: Joint-Embedding Predictive Architecture for 3D Atomistic Systems
链接: https://arxiv.org/abs/2610.08400
作者: Kasper Helverskov Petersen,Rasmus Hannibal Tirsgaard,François R J Cornet,Mikkel Jordahn,Mikkel N. Schmidt
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computational Physics (physics.comp-ph)
备注:
Abstract:Large-scale self-supervised pretraining has reshaped modern machine learning, substantially advancing the ability of language and vision models to generalize across downstream tasks. While deep learning has driven considerable progress in modeling atomistic systems in recent years, self-supervised pretraining in this domain has not yet achieved comparable downstream generalization. To address this, we introduce Atom-JEPA, a self-supervised pretraining framework that learns latent representations from unlabeled 3D structures through complementary atom-level and substructure-level objectives inspired by joint-embedding predictive architectures. We pretrain Atom-JEPA on large-scale molecular and crystalline datasets and evaluate its transfer performance by fine-tuning on a diverse set of downstream property prediction tasks. Atom-JEPA achieves state-of-the-art performance on molecular ADMET and quantum-chemical property prediction tasks, and is highly competitive in predicting the physical properties of crystalline materials. These results demonstrate the potential of latent-space predictive pretraining to support broad downstream generalization from structural data alone. Code and pretrained model checkpoints are publicly available at this https URL
[AI-36] Accelerating the Development of PLGA In Situ Forming Depots Through AI-Driven Multi-Objective Optimization
链接: https://arxiv.org/abs/2610.08368
作者: Pauric Bannigan,Siddarth Chandrasekaran,Brigitte A. G. Lamers,Inge Hermsen,Gary Tom,Riley J. Hickman,Bahar Yeniad,Morgan Fox,Christine Allen
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: 10 pages; 7 figures
Abstract:Developing long-acting injectable formulations requires the simultaneous optimization of drug loading, release kinetics, viscosity, injectability, stability and other objectives. To navigate this multidimensional space, Corbion and Intrepid combined Corbion’s diverse PURASORB bioresorbable polymer library with Intrepid Labs’ proprietary AI algorithm (ANDROMEDA 1) to develop in situ forming depots for a therapeutic peptide. Over approximately 15 weeks, 181 unique formulations spanning drug loadings of 6-12% w/w were prepared and characterized through broad design-space mapping and targeted multi-objective optimization. Four lead candidate formulations were identified at 6%, 9%, and 12% w/w drug loading. Each met the predefined viscosity and injectability criteria while providing distinct 30-day in vitro release profiles. The study evaluated polymers spanning a broad range of molecular weights, including commercially available PURASORB grades and new polymers under development by Corbion to expand its polymer toolbox. ANDROMEDA 1 identified that polymers with intermediate molecular weights provided a favorable balance between sustained release and solution viscosity. Together, these findings demonstrate how integrated polymer expertise and AI-driven optimization can rapidly identify differentiated formulation candidates, focus the development space, and establish a strong data-driven foundation for further optimization and in vivo evaluation.
[AI-37] ransect: Retaining Observability for Long-Horizon LLM Agent Evaluations
链接: https://arxiv.org/abs/2610.08364
作者: Toby D. Pilditch,Konstantinos Voudouris,Alexandra Abbas,Cozmin Ududec
类目: Artificial Intelligence (cs.AI)
备注: 27 pages, 5 figures
Abstract:Frontier AI evaluations increasingly use open-ended, agentic, long-horizon tasks whose transcripts can span hundreds of pages of outputs and actions from complex multi-agent networks. The observability envelop-the range of what evaluators can reliably infer about an agent’s behaviours-is therefore narrowing. Language model assistants can help classify and interpret agent behaviour but also afford human evaluators significant analytical degrees of freedom, threatening the reproducibility and auditability of language-model-based transcript analysis. Transect is an open source package built on Inspect Scout to help evaluators understand how a long agent run unfolded, identify behaviour worth investigating, and check interpretations against the transcript. Users specify task context and behavioural vocabulary in a reusable evaluation-family configuration, with judge models and analysis settings supplied separately. Transect’s navigable reports align recorded events, token use, sub-agent activity, and model-generated behavioural labels on a common turn-based timeline. Reviewers can quickly grasp a run’s narrative, trace any label or event to its source turns, and export the underlying data tables for cross-run analysis. We demonstrate the workflow on an AI RD evaluation that generated almost 13 million tokens, dividing the agents’ work into behavioural phases aligned with research-skill classifications, sub-agent delegations and interactions, and token use. The combined view shows a focus on operational work and manuscript production, with little evidence of a sustained hypothesis generation stage-arguably a necessary component for high-quality scientific outputs. Transect’s flexible, customisable transcript-analysis pipeline will enable evaluators to keep pace with longer, more complex, more frequent AI evaluations while supporting scientific rigour, transparency, and reproducibility.
[AI-38] Explainable Failure Prediction and Prevention in Maritime
链接: https://arxiv.org/abs/2610.08363
作者: Dionisis Kalogeropoulos,Georgia Sovatzidi,Panagiotis G. Kalozoumis,Dimitris K. Iakovidis
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Maritime systems operate in highly dynamic environments where unexpected equipment failures can compromise safety, reliability, and operational efficiency. Recent advances in artificial intelligence (AI), machine learning, digital twins, and predictive maintenance enable proactive failure prediction and prevention. However, ensuring trustworthy and explainable decision-making remains a major challenge in safety-critical maritime applications. This chapter reviews key AI technologies required for explainable failure prediction and prevention in maritime systems and presents a conceptual architecture capable of supporting autonomous or human-in-the-loop corrective actions. This architecture integrates data acquisition, time-series forecasting, anomaly detection, risk assessment, decision-making, and explainable AI into a closed-loop framework. With reference to the architectural components, a review and discussion of relevant maritime studies is performed, outlining their methods, advantages, and limitations. Furthermore, it highlights current challenges, including uncertainty and robustness, model generalization, explainability, limited availability of maritime datasets, and operational deployment, and identifies future research directions toward trustworthy AI-assisted maritime decision-making.
[AI-39] Sensor Geometry as a Flow-Matching Prior for Multi-Channel Brain Signals
链接: https://arxiv.org/abs/2610.08355
作者: Jaedong Hwang
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:
Abstract:Flow-matching models start from an isotropic Gaussian source, the standard choice when the correlation structure of the data is unknown in advance. For multi-channel brain recordings, however, part of this structure is known in advance. Electrodes sit at fixed positions on the head, and volume conduction through the skull and scalp makes nearby electrodes co-vary in a way that is shared across subjects. Existing EEG generative models nonetheless leave the network to learn this from scratch. We put this structure into the source instead. From the sensor coordinates alone, we build a k-nearest-neighbor graph and take a graph-Matérn function of its Laplacian as the source covariance, so the flow starts from spatially coherent patterns rather than channel-independent noise. The change adds no learned parameters, works with any coupling and any drift network, and uses the same three hyperparameters on every dataset. Across eight EEG datasets and four flow-matching methods, the graph-Matérn source lowers the spectral discrepancy between generated and real signals in the five clinical bands (PSD-KL) on most datasets. PSD-KL falls by 12% to 17% in geometric mean over datasets depending on the method and by up to 40% on PhysioNet-MI, the densest montage. We show that the improvement stems from the spatial eigenvectors of the local graph of sensor positions, since randomizing the eigenvectors while preserving the eigenvalue spectrum eliminates the gain. Furthermore, a prior fitted directly to the empirical data covariance performs worse than isotropic noise. The same construction applies unchanged to MEG, intracranial EEG with patient-specific grids, and a traffic-sensor network, lowering PSD-KL for every method on each. this https URL
[AI-40] How Much Planning Is Enough? Reducing Search and Computation in World-Model Planning
链接: https://arxiv.org/abs/2610.08350
作者: Changbai Li,Sirui Li,Yichen Yang,Tongfei Chen,Zichao Feng,Shuwei Shao,Huobin Tan
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI)
备注:
Abstract:Visual world models enable goal-directed control through decision-time action search, but their deployment efficiency is often limited by conservatively large planning budgets. We show that competitive task performance can be achieved without agreement with the Full-budget action, that sufficient budgets vary across model–task pairs, and that iterative planners repeatedly encode solve-invariant context. To address these inefficiencies, we propose SufficientPlan, a simple deployment framework that requires no modification to pretrained world models or planner updates. Its Paired Sequential Budget Certification (PSBC) component uses paired closed-loop evidence to search for and certify a reduced model–task-specific budget within a predefined Full-performance tolerance. Its Static-Context Reuse (SCR) component caches observation and goal representations across search iterations while preserving candidate-dependent planning and selected actions. Experiments across multiple world-model backbones and visual-control tasks show that SufficientPlan substantially reduces search budgets and planning latency while maintaining competitive control performance.
[AI-41] MoF: Preference-Aware Mixture Modeling for Black-Box LLM Personalization EMNLP2026
链接: https://arxiv.org/abs/2610.08330
作者: Hun Park
类目: Artificial Intelligence (cs.AI)
备注: Accepted at EMNLP 2026
Abstract:Proprietary Large Language Models (LLMs) have demonstrated remarkable capabilities across a wide range of tasks, yet aligning their outputs with diverse user preferences remains challenging. Existing personalization approaches for black-box LLMs often rely on user-specific scoring heads, causing the number of personalized parameters to grow linearly with the number of users and requiring additional adaptation for unseen users. To address these limitations, we propose Mixture-of-Facets (MoF), a scalable personalization framework for black-box LLMs that models user preferences as compositions of shared latent preference facets rather than dedicated user-specific parameters. MoF performs personalization through history-conditioned routing over shared facet heads, enabling personalization for users unseen during training without additional parameter updates. Across diverse personalization tasks, MoF delivers stronger personalization performance while maintaining a more scalable and parameter-efficient design than prior approaches. Additional analysis indicates strong generalization to unseen users.
[AI-42] An AI-Assisted Formalization of the Poincaré Conjecture
链接: https://arxiv.org/abs/2610.08329
作者: Zhiyuan Zhang,Axel Delaval,Leheng Chen,Jinxuan Chen,Jie Xu,Yuxuan Liao,Jiedong Jiang,Chunlei Liu,Bin Dong
类目: Artificial Intelligence (cs.AI); Geometric Topology (math.GT)
备注: 15 pages, 2 figures. Code: this https URL
Abstract:We present an AI-assisted Lean 4 formalization of the Poincaré conjecture. The project began with limited reusable formal infrastructure for the geometric analysis behind the proof. To organize this work, we combined a proof blueprint prepared by mathematicians with explicit milestone statements. These milestones enabled parallel agent work and gave mathematicians clear points to locate blockers and provide effective mathematical guidance. Our analysis identifies the human interventions and organizational choices behind this workflow. The project provides a starting point toward reusable infrastructure for future formalization projects; such infrastructure, once developed, could eventually reduce the cost of verifying mathematical results in geometric analysis.
[AI-43] MedZERO: Self-Evolving Agents for Open-Ended Medical Reasoning Through Controlled Knowledge Accumulation NIPS2026
链接: https://arxiv.org/abs/2610.08327
作者: Xilin Dang,Weilin Ruan,Xue Yang,Jinghao Wang,Xiaowei Hu,Jinpeng Li,Pheng-Ann Heng
类目: Artificial Intelligence (cs.AI)
备注: accepted by NIPS 2026
Abstract:Large language models (LLMs) have shown promise in medical question answering and clinical reasoning, yet their improvement remains constrained by static parametric knowledge and costly expert supervision. Self-evolving agents offer a promising alternative by enabling models to improve through iterative task generation and problem-solving. However, most existing self-evolving methods are designed for easily verifiable domains such as mathematics and coding, where solutions can be checked by exact answers or executable programs. Medical reasoning is fundamentally different: it is open-ended, knowledge-intensive, and often only partially verifiable. We present MedZERO, a self-evolving framework for open-ended medical reasoning. MedZERO couples an Examiner that generates frontier medical question-option pairs with a Reasoner that solves them through evidence-grounded multi-turn reasoning with external knowledge tools. To support reliable, continual improvement, MedZERO adopts controlled knowledge accumulation, which maintains temporary exploratory knowledge and curated persistent knowledge in reasoning. We evaluate MedZERO on five public medical reasoning benchmarks using 4B- and 8B-scale base models under open-ended evaluation. Across all settings, MedZERO consistently outperforms the underlying base models and prior self-evolving baselines, achieving up to 13.7 average accuracy-point gains over the next-best self-evolving baseline.
[AI-44] SCOPE: Certified Theorem Proving with a Language Model as the Policy Planner
链接: https://arxiv.org/abs/2610.08319
作者: Hanchao Zhou,Jialei Li
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:In proof assistants such as Lean, a generated proof must pass machine compilation checks, so evaluation needs no human scoring. Direct generation fails on multi-step numeric propositions: a proof is valid only if every content integer is correct, so the pass rate is bounded by the k-th power of the per-integer accuracy. Controlled corruption across 2,617 reference proofs confirms this power law. SCOPE (State-Conditioned Operator Planning and Execution) enforces the natural division of labor: the model plans over an operator vocabulary, a symbolic engine executes the numerics, and a compiler renders the proof. On a 218-problem suite it certifies 191/218 (87.6%) with a 135M backbone; the 7B DeepSeek-Prover-V1.5-RL certifies 18/218 at 27.5 times the tokens and 37.5 times the wall-clock, and DeepSeek-Prover-V2-7B certifies zero on a bidirectional dual suite. Multi-step thinking costs 6.12 discrete decision actions per problem and produces no natural-language thinking text. Replacing the lagged engine state in the decision frame with the current one lifts the pass rate from 117/218 to 191/218, while up-weighting the chain-end loss hurts. On the public Lean-Workbook library, 2,132 of 3,536 gradeable admissible problems certify (60.29%) with zero regression on the main suite. All readings come from a version-frozen review with independent rechecks and reverse verification. Restricting free generation and keeping decision-time information visible is a more direct route than enlarging the model.
[AI-45] MARCO: The Radioactive Watermark for Protein Generative Models
链接: https://arxiv.org/abs/2610.08316
作者: Huajie Chen,Xin Guo,Yuchen Shi,Yuchen Zhong,Minhui Xue,Chi Liu,Congcong Zhu,Kun Gao,Minfeng Qi,Tianqing Zhu
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
备注:
Abstract:Protein Generative Models (PGMs) have revolutionized structural biology by enabling the design of complex 3D protein structures from sequence data. However, this breakthrough introduces a dual-use challenge, exposing high-value PGMs to economic risks like unauthorized model extraction and biosecurity threats such as biohazard synthesis. To mitigate these threats, we propose \textbfMARCO (\textscCOnformation waterMARk), the first radioactive watermarking framework specifically tailored for PGMs. MARCO establishes a Dual-Layer defense that simultaneously protects intellectual property and ensures the forensic traceability of potential biosecurity misuses. (i) To preserve efficiency, MARCO iteratively embeds watermarks during diffusion reverse denoising via an auxiliary encoder-decoder, allowing the original PGM parameters to remain frozen for broad compatibility. (ii) To preserve biophysical fidelity and maximize robustness, we employ specialized loss functions targeting C_\alpha -atom pairwise distances and torsion angles ( \psi, \phi ) within an adversarial training framework integrated with stochastic attack simulations. (iii) Crucially, MARCO exhibits ``radioactivity’’ where the watermark automatically transfers to the outputs of any pirate models trained on the watermarked data, effectively countering model extraction attacks. Comprehensive experiments demonstrate that MARCO achieves superior fidelity and robustness while successfully validating watermark transferability.
[AI-46] he Standardization Trap: Certifying Joint Label Processing in Tabular Foundation Models
链接: https://arxiv.org/abs/2610.08314
作者: Duong Nguyen,Nicolas Chesneau,Milan Bhan
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:
Abstract:Linear regression and kernel smoothing offer tractable explanations of in-context learning: in both, the features determine the weight assigned to each context label. However, whether this fixed-weight account describes pretrained tabular foundation models (TFMs) remains unclear. Testing this account using derivatives runs into a standardization trap: public TFM packages standardize the labels before the model sees them, yet ordinary derivatives also reflect behavior outside the set of standardized labels, making a model appear nonlinear even when every prediction it makes agrees with a fixed-weight map. We propose two certificates that depend only on predictions at standardized labels and can reject two distinct explanations: fixed-weight prediction and sums of independent nonlinear label transformations. Across the five public TFMs that we evaluate, our certificates show that changing one context label alters how other labels influence the prediction, a behavior we call joint processing. We further find that joint processing emerges with training and that attention scores carry most of the measured interaction. Together, these findings motivate TFM explanations that account for how context labels change the influence of individual examples.
[AI-47] Mitigating Concept Drift in QoS Prediction for Teleoperation of Autonomous Vehicles Using Historic Data
链接: https://arxiv.org/abs/2610.08297
作者: Xiyan Su,Jianning Gao,Mahmoud Ashri,Frank Diermeyer
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI)
备注: 2025 IEEE International Conference on Systems, Man, and Cybernetics (SMC)
Abstract:Teleoperation serves as the fallback solution to autonomous driving but reliable functions of the teleoperation require a certain amount of mobile network resources, which cannot be guaranteed at all times. Therefore, predictive quality of service (pQoS) is introduced as a concept to increase the resilience of the teleoperation. In this paper, based on a data measurement campaign, we propose a prediction framework to prediction two important network KPIs of teleoperation: uplink data-rate and round-trip latency. Furthermore, we introduce a method to alleviate the performance degradation of machine-learning-based prediction models on previously unseen data due to concept drift by incorporating historic data into the prediction pipeline. Additionally, we introduce the metric of critical scenario detection to evaluate the prediction performance specifically for teleoperation.
[AI-48] DySCo: Dynamic Sharding for Collaborative Edge-Cloud LLM Inference with Depth-Synchronized Batching
链接: https://arxiv.org/abs/2610.08268
作者: Jingpo Xu,Paul Joe Maliakel,Ivona Brandic,Shashikant Ilager
类目: Distributed, Parallel, and Cluster Computing (cs.DC); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Systems and Control (eess.SY)
备注: article under submission
Abstract:Pervasive intelligent applications are increasingly deployed on mobile and Internet of Things (IoT) edge devices. Consequently, Large Language Models (LLMs) are increasingly used to support these applications. Yet, due to their high resource demands, LLMs are mostly deployed in the cloud. Layer-wise edge-cloud inference lets resource-constrained edge devices contribute computation to LLMs they cannot host in full. However, heterogeneous split points introduce two coupled inefficiencies. First, edge execution and communication create idle gaps between cloud invocations. Second, requests arriving at different model depths cannot be conventionally batched. We present DySCo, a collaborative runtime that keeps KV caches local and introduces dyForward, a model-aware layer-range executor that runs configurable contiguous layer ranges from resident model shards without reloading weights. For multi-edge serving settings, we introduce depth-synchronized batching (DSB), which advances heterogeneous requests to the deepest cut and batches their common suffix. Experiments across heterogeneous devices, two model families, and local and wide-area links show that idle gaps increase the latency of subsequent GPU forward calls even when waiting time is excluded, adding up to 25 ms of additional cloud-side suffix latency per decoding step in our measurements. At an average concurrency of eight, DSB improves throughput by 275% over FIFO, 48% over exact-match batching, and 79% over round-robin interleaving while reducing mean per-session latency. Together, these results show that requests with different edge-cloud splits can reuse resident cloud weights and share batched suffix computation. The artifact repository for this work is publicly available at: this https URL
[AI-49] zkLLM PoT: Efficient Zero Knowledge Proof of Training for Large Language Models ICLR2027
链接: https://arxiv.org/abs/2610.08258
作者: Junkai Liang,Zhanpeng Guo,Pengfei Wu,Qingni Shen,Jiaheng Zhang,Zhonghai Wu,Haiyang Xue,Shengfang Zhai
类目: Artificial Intelligence (cs.AI)
备注: Submitted to ICLR 2027
Abstract:Auditing the claimed outcomes of large language model (LLM) training is challenging when model weights and training data are private, while cryptographically proving the full training process is prohibitively expensive at Transformer scale. We present zkLLMPoT, a zero-knowledge framework that certifies auditor-defined properties of a trained checkpoint through forward evaluation rather than verification of its optimization trajectory. zkLLMPoT includes 2 phases: 1) The trainer fixes the architecture and the model weights are committed. Then the auditor selects challenge sequences, preventing the trainer from modifying the checkpoint in response to the audit data. 2) Then the trainer proves the objective value attained by the committed model on those sequences. This formulation makes the certification cost independent of the number of training iterations, without revealing model weights or requiring access to private training data. We build on sumcheck- and lookup-based arguments to certify Transformer computations, while supporting next-token loss and task-specific audit objectives. Across four model families, operator-level benchmarks yield proving times of 41-59 seconds for 1.1-1.5B-parameter models and 131 seconds at 13B for the covered operators, with verification below half a second at a sequence length of 512.
[AI-50] MASC: A Multi-Agent Self-Calibration Framework with Latent Construct Alignment for Consistent Client Role-Playing in Psychological Counseling
链接: https://arxiv.org/abs/2610.08250
作者: Shixin Peng,Kun Jiang,Jiaxing Zheng,Qihao Yang,Jingying Chen
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Large language models are increasingly used to simulate clients for counselor training and psychological counseling research, but reliable simulation requires clients to remain psychologically coherent across extended interactions. Existing role-playing methods largely rely on static profile prompts and may exhibit persona drift, unrealistic cooperativeness, or inconsistent psychological states, communicative actions, and emotions. Existing evaluations also lack a unified testbed for both stable client characteristics and evolving psychological dynamics. We propose MASC, a Multi-Agent Self-Calibration framework with latent construct alignment for consistent client role-playing in psychological counseling. MASC combines construct-guided generation, collaborative refinement, consistency verification, and memory-based revision in a closed calibration loop that detects and corrects inconsistencies as dialogue unfolds. We further introduce CRPC-Bench, a benchmark covering session-level profile information and Big-Five personality traits, as well as turn-level psychological state, communicative action, and emotion expression. CRPC-Bench contains 38 motivational interviewing client profiles augmented with personality and emotion annotations. Experiments show that MASC outperforms existing methods across profile, personality, receptivity, and turn-level consistency, with the heterogeneous configuration achieving the strongest overall performance. MASC and CRPC-Bench provide a unified foundation for developing and evaluating psychologically coherent client simulations for AI-assisted counseling research and training.
[AI-51] LeanPlan: Optimal Planning with LLM -Generated Heuristics and Admissibility Proofs
链接: https://arxiv.org/abs/2610.08246
作者: André G. Pereira,Augusto B. Corrêa,Felipe Meneguzzi,Jendrik Seipp
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Symbolic Computation (cs.SC)
备注:
Abstract:Frontier large language models (LLMs) can generate heuristic functions that guide search to achieve state-of-the-art performance in satisficing planning, where any plan is acceptable. However, these heuristics are not guaranteed to be admissible and can lead to suboptimal plans. We introduce LeanPlan, the first planning system that finds optimal plans with LLM-generated heuristics whose admissibility is machine-checked. Given a domain description and training tasks, an agentic loop uses planner feedback to iteratively improve a reusable domain-specific heuristic, its admissibility proof and the required domain assumptions. LeanPlan implements the heuristic, its proof and an efficient planner with machine-checked grounding and search in Lean 4. We evaluate LeanPlan on ten domains from the International Planning Competition and three new domains, using test tasks with up to 57 times as many objects as the training tasks. With GPT-5.6 Sol in the agentic loop, we successfully generate heuristics and admissibility proofs for all these domains. With the resulting heuristics, LeanPlan usually expands fewer states than the state-of-the-art Scorpion planner and solves more tasks overall.
[AI-52] Sensor-Language-Action Models
链接: https://arxiv.org/abs/2610.08244
作者: Yuekai Xu,Zitao Shuai,Yuzhe Yang
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Sensors are useful not only for understanding the world but also for deciding what to do next. Existing sensor models however largely stop at perception: they recognize states or predict outcomes, leaving actions modeled separately through task-specific and often closed label spaces. We introduce Sensor-Language-Action (SLA) modeling, a framework that connects multimodal sensor observations, natural language, and actions within a unified model. SLA uses language as a semantic interface between sensing and acting, allowing heterogeneous actions to be represented, predicted, and explained while remaining grounded in the underlying sensor evidence. We build a large-scale SLA benchmark consisting of datasets that span more than 116,000 individuals, 79 sensor modalities, and 60 action groups, together with a multi-faceted captioning pipeline that aligns user context, sensor dynamics, and action evidence. Building on this framework, we present OpenSLA, a unified SLA model for hierarchical action prediction, state understanding, and action explanation. Extensive experiments on real-world tasks in clinical prediction, operating rooms, and metabolic health verify its superior performance over the state-of-the-art. OpenSLA also demonstrates intriguing capabilities including language-guided evidence grounding and zero-shot generalization to unseen actions and cohorts.
[AI-53] OSFP4: Joint Optimization of Diagonal Smoothing and Block Scales for NVFP4 Quantization
链接: https://arxiv.org/abs/2610.08231
作者: Neriah Ben David,Ori Meir,Or Ordentlich
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:NVFP4 is an attractive datatype for large language model (LLM) inference, offering compact storage and native tensor-core acceleration. However, preserving accuracy using NVFP4 requires careful quantization. In this work we develop a novel quantization scheme called Optimized Smoothing and Scaling for NVFP4 (OSFP4). For each linear projection it uses a diagonal smoothing matrix whose entries are optimized to minimize the squared matrix-product quantization error under NVFP4, taking into account the rounding procedure that is used (either round-to-nearest, or GPTQ-style successive interference cancellation). This requires performing joint optimization on the smoothing entries as well as the block scales, which is facilitated by analyzing a multiplicative-dither FP4 quantizer instead of the fixed deterministic one. Experiments show that OSFP4 achieves the highest average accuracy among the evaluated competitors in the corresponding quantization settings, while retaining approximately 94-97% of vendor NVFP4 prefill throughput on the measured workloads. Our code is available in this https URL
[AI-54] VOMMI: Collecting and Leverag ing Portable Demonstrations for Mobile Manipulation
链接: https://arxiv.org/abs/2610.08220
作者: Yutian Zhang,Xingrui Xiong,Siyuan Ma,Yang Li,Jiawen Wen,Jiaqi Zhai,Liwen Yang,Ce Hao,Haozhen Chi,Yangkun Zhu,Yifan Zhu,Xiaowen Chu,Dong Wei,Qiaojun Yu,Dibo Hou
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI)
备注: 9 pages, 6 figures
Abstract:Portable mobile-manipulation demonstrations can help alleviate data scarcity for embodied intelligence, but obtaining reliable, low-cost, and robot-free motion supervision from RGB observations remains challenging. Existing approaches often rely on teleoperation or specialized devices equipped with additional sensing hardware, while directly using estimated visual odometry (VO) trajectories can introduce inconsistencies due to accumulated drift and imperfect motion supervision. We present the Visual-Odometry-Conditioned Mobile Manipulation Interface (VOMMI), a portable demonstration collection and learning framework that connects portable RGB demonstrations to vision-language-action (VLA) post-training through offline trajectory reconstruction and online visual-motion conditioning. VOMMI synchronizes body and hand views to capture navigation context and local object interactions without requiring human-robot kinematic correspondence calibration. R2-VO refines offline demonstration trajectories using sparse geometric anchors and produces causal local-motion tokens over multiple prediction horizons for online policy conditioning. An action-group residual adapter incorporates these tokens only into the base branch. Experiments use a 500-trajectory portable for each task, with 75 trajectories held out for RGB-VO evaluation, and 200 robot demonstrations as references. Our policy, post-trained only on portable demonstrations, achieves 18.2% lower base-velocity error than a policy trained with robot-collected demonstrations, while maintaining comparable end-effector translation accuracy. Offline reconstruction reduces absolute trajectory errors for the body and hand streams by 24.6% on average relative to the best evaluated baseline for each stream. The complete system improves the mean success rate by 8.3 percentage points over OpenPI 0.5 across three real-robot tasks.
[AI-55] Quantum Entangled Multimodal Fusion Networks (QEMFN): Resource-Aware Hybrid Vision-Language Fusion via Trainable Entanglement
链接: https://arxiv.org/abs/2610.08216
作者: Srikar Alla,Ali Shiri Sichani,Chi-Ren Shyu
类目: Artificial Intelligence (cs.AI); Quantum Physics (quant-ph)
备注:
Abstract:Multimodal vision-language systems typically fuse image and text embeddings through classical operators such as concatenation, attention, bilinear pooling, or tensor interactions. We propose Quantum Entangled Multimodal Fusion Networks (QEMFN), a hybrid quantum-classical framework that introduces parameterized entanglement as a structured inductive bias for multimodal fusion. Pretrained visual and textual features are projected into compact latent spaces, encoded as angle-parameterized quantum states, processed through intra-modal and paired cross-modal entangling circuits, and measured to produce fused representations for retrieval. Under matched parameter budgets and identical frozen CLIP backbones, QEMFN outperforms classical fusion baselines on COCO-5k and Flickr30k, including multilayer perceptron, tensor fusion, FiLM, cross-attention, compact transformer, and a dequantized paired-topology analogue. An ablation suite isolates the quantum module’s contribution from the surrounding classical projections, and quantum-centric analyses report Meyer-Wallach entangling capability, expressibility, gradient variance against barren-plateau bounds, and entropy-performance correlation under controls for training progress alongside an intervention study on the entangling component. QEMFN is executed under shot-based estimation, a noise-modeled fake backend, and a real superconducting device with zero-noise extrapolation. This work does not claim quantum computational advantage; the contribution is the framework together with a controlled empirical and quantum-centric evaluation that positions trainable entanglement as an interpretable, hardware-executable fusion mechanism at scales accessible on contemporary devices.
[AI-56] Learn2Play Bench: How Well Do LLM Agents Learn from Experience in Unfamiliar Environments?
链接: https://arxiv.org/abs/2610.08215
作者: Yibo Li,Jinhang Qiu,Zhi Zheng,Qianyun Guo,Jiaying Wu,Shuo Ji,Bryan Hooi
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Learning from experience is essential for LLM agents to adapt to unfamiliar and dynmaic environments. Evaluating this ability is therefore important for understanding how effectively agents acquire and use new knowledge. Existing benchmarks have sought to evaluate this ability, but they primarily evaluate tasks whose rules are provided in the instructions or already familiar to pretrained models, making it difficult to distinguish learning from interactions from reasoning with existing knowledge. To address this, we introduce Learn2Play Bench, a benchmark of newly designed text-based games, whose rules are novel or counterintuitive, requiring agents to acquire knowledge through interaction rather than rely solely on pretrained knowledge. These games provide reproducible feedback and automatic scoring, enabling controlled evaluation of learning across repeated attempts. We also vary game instances to test whether agents can apply what they have learned to new situations. Therefore, we evaluate how backbone models, self-evolving methods, and agent harnesses affect agents’ learning ability, revealing three findings: (1) Experience retention: Retaining complete records of actions and feedback can support more effective learning than summarizing these experiences into rules or strategies. (2) Human agent gap: Top-performing human players achieve higher peak scores than the evaluated agents. Human explore more varied strategies, and repeat actions less. (3) Harness matters: With the backbone fixed, changing the harness can improve performance while reducing estimated inference cost. Together, these findings provide insights into how LLM agents learn from experience and suggest directions for future work to improve their learning ability. Project website: this https URL
[AI-57] Mathematical Proof Assistants for Teaching Logic: The LogiKEy Methodology
链接: https://arxiv.org/abs/2610.08214
作者: Christoph Benzmüller,David Fuenmayor,Luca Pasetto
类目: Artificial Intelligence (cs.AI); Logic in Computer Science (cs.LO)
备注:
Abstract:We report on an approach to teaching logic to mixed groups of computer science, mathematics, and philosophy students, based on the logico-pluralistic LogiKEy methodology, used for more than a decade in courses, summer schools, and tutorials. LogiKEy uses classical higher-order logic (HOL) as a universal metalogic in which object logics, classical and non-classical alike, are encoded by defining their semantics; through these semantical embeddings a single proof assistant (e.g. Isabelle/HOL), with its automated theorem provers and (counter-)model finders, becomes one environment in which students learn, experiment with, and compare logics. After making the pedagogical case for proof assistants in the logic classroom, we present a graded sequence of classroom examples, each transition motivated by a limitation of the preceding representation, by a need for more explicit modelling resources, or by a new application. A liars-and-truth-tellers puzzle leads from propositional to modal logic; the Wise Men puzzle leads on to dynamic epistemic logic; Boolos’s curious inference illustrates what a higher-order meta-logic buys, even for automated proof search; Chisholm’s paradox takes the sequence into deontic logic, and from standard to dyadic deontic logic; and Gödel’s ontological argument brings it to a research-level metaphysical argument. We then rebut the objection that embedding everything in classical HOL is monism rather than pluralism, reflect on three years of teaching such a course, and sketch the portability of the approach beyond Isabelle.
[AI-58] ool-calling retrieval versus vector RAG for a small Greek–English knowledge base: accuracy and robustness to how users type Greek
链接: https://arxiv.org/abs/2610.08205
作者: Nikolaos D. Tantaroudas,Ilias Karachalios,Andrew J. McCracken
类目: Computational Engineering, Finance, and Science (cs.CE); Artificial Intelligence (cs.AI)
备注: 13 pages; 2 figures;
Abstract:Assistants grounded in a small, frequently edited knowledge base can retrieve through tool calls to a live data interface or through vector retrieval-augmented generation (RAG). We compare the two on KyGround, a benchmark of 198 questions drawn from the published records of a Greek–English agricultural platform on Kythera, Greece, with answers verified automatically against the records and each question posed in up to nine forms, including Greek without accents, in capitals and in three Latin-script (Greeklish) schemes. With Claude Haiku 4.5 as router and answer model, a reconstruction of the platform’s tool agent answered 71.6% of canonical Greek questions correctly and vector RAG 95.3% (difference -23.6 percentage points, 95% CI -33.1 to -15.1 ). Letting the router write the vector query changed nothing, and placing the whole knowledge base of about 26,000 tokens in the prompt reached 99.3%. The tool agent’s losses arose in retrieval. Its literal searches returned nothing when the router’s arguments did not occur verbatim in a record, for example when it transliterated Greek into Latin script or combined words that occur in a record but not as one phrase, and the agent then abstained. Unaccented and capitalised questions cost the tool agent about 20 points and vector RAG at most 2; accent-insensitive search removed this loss, and matching stemmed tokens raised the tool agent to 83.8% on canonical Greek. Greeklish cost both designs about 21 to 32 points. Tool interfaces for community knowledge bases need search that tolerates how users type.
[AI-59] Compact Robot Policies Need Fine-Grained Visual Representations
链接: https://arxiv.org/abs/2610.08183
作者: Nanhe Chen,Runqiu Yang,Jiawei Tang,Sichao Liu,Yuquan Wang
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: 35 pages, 21 figures, 8 tables
Abstract:Multi-task manipulation policies differ in architecture, scale, and pretrained priors all at once, so published comparisons cannot attribute performance to any single component. We argue that most of it comes from the visual representation, and that parameter scale and generative priors are largely incidental. To test this, we build CoRP (Compressed Representation Policy), a deliberately compact policy (48.9M parameters, no vision-language model and no video-generative prior) that factorizes into a representation extractor and a flow-matching action generator. It reaches 97.0% on LIBERO and 75.78%/73.36% on RoboTwin 2.0 Clean/Randomized, matching systems 40.9-163.6x larger. Holding the action generator fixed, we then vary one extractor property at a time. Pretrained initialization is decisive: a random ViT-S/14 drops to 78.1% and an ImageNet ResNet-34 to 74.5% on LIBERO. Pretraining alone is not enough, as freezing the encoder costs 19.8 points. Compression matters as much: resampling each view to 48 tokens beats passing all patch tokens (97.0% vs 83.2%), and a variational information bottleneck over those tokens is worse than a hard token budget, cutting LIBERO-Goal from 95.8% to 33.0% by suppressing the instruction-dependent token selection the policy relies on. Language conditioning contributes only where the observation leaves the goal ambiguous (LIBERO-Goal: 9.2% to 95.8%), while on RoboTwin 2.0, where observations are unambiguous, removing it slightly improves success. Therefore, we argue that a compact policy works when its representation is pretrained, task-adapted, and compressed. Project page: this https URL
[AI-60] LFHE: Local-First Heuristic Evolution for Bounded Local Topology Search in Decentralized Learning with Non-IID Data
链接: https://arxiv.org/abs/2610.08176
作者: Yin-Kuan Liang(Durham University),Yan Gao(University of Cambridge),Yang Long(Durham University)
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: Preprint. 31 pages, 12 figures, 8 tables
Abstract:Decentralized learning is highly sensitive to communication topology under non-IID data. Adaptive peer-selection methods can exploit local model information, but broader peer discovery may require increasingly large control state, whereas direct spectral optimization typically relies on graph-wide information. We study the intermediate setting of bounded local topology search and propose Local-First Heuristic Evolution (LFHE), a representation-driven rewiring framework whose candidate discovery and scoring use only ego-neighborhood and friend-of-a-friend (FoF) information. The structural score admits an exact interpretation through graph Dirichlet energy: its sum across clients equals twice the representation Dirichlet energy, which under standard linear consensus dynamics governs the instantaneous dissipation of representation disagreement. LFHE combines this state-dependent structural signal with early exploration and degree control, while algebraic connectivity remains an offline graph diagnostic. Under bounded sparse degree, its FoF candidate state remains local rather than expanding toward population-wide peer tracking. Across four image, speech, and text benchmarks, LFHE achieves competitive decentralized learning performance. Matched-protocol controls identify the structural term as the principal empirical topology-selection signal, while comparison with broader peer discovery exposes a trade-off between predictive performance and discovery-state locality. Together, these results motivate state-aware bounded local topology search between pairwise peer selection and globally informed topology optimization.
[AI-61] Which alloy compositionwhat process parameters? Inferring the recipe from optimized metallic microstructure and texture
链接: https://arxiv.org/abs/2610.08165
作者: Mahish K. Guru,Jan Bohlen,Louam Lemjid,Marius Tacke,Roland Aydin,Noomane Ben Khalifa
类目: Artificial Intelligence (cs.AI); Materials Science (cond-mat.mtrl-sci); Computational Engineering, Finance, and Science (cs.CE)
备注:
Abstract:The mechanical properties of a metallic alloy are set by its microstructure and texture: the size and shape of its grains and the orientation of their crystals. That structure is in turn set by a recipe, the alloy composition together with the processing parameters. Alloy development runs this chain forwards, tuning the structure until a target property is met. Running it backwards, from an optimized structure to the recipe that would produce it, still relies on expert knowledge. We ask whether this backwards step can be learned. On an in-house dataset of 107 magnesium alloy extrusion conditions across 14 alloys, each with optical micrographs and an X-ray texture measurement, we compare three descriptors of microstructure and texture: conventional grain and texture statistics, a vision embedding from a pretrained image encoder, and a graph neural network on the grain network. Each is paired with prediction heads for two tasks: the alloy composition given the process (Task A), and the process parameters given the composition (Task B). Under 5-fold cross-validation, the conventional descriptors identify the correct alloy for 65% of held-out conditions, against 17% for always guessing the most common alloy, while the learned embeddings stay below 30%. The process parameters are recoverable but noisier: compared with using the composition alone, the microstructure roughly halves the temperature error. Because only a few alloys were cast and only a few press settings were used, both answers are discrete, and heads that pick from these known options, while respecting their order, worked better than heads that predict a free value.
[AI-62] st-Time Agent Evolution for Long-Horizon Legal Reasoning
链接: https://arxiv.org/abs/2610.08138
作者: Haotian Chen,Shuaicheng Niu,Haocong Rao,Kaisong Song,Jun Lin,Lizhen Cui,Zhiqi Shen,Yonghui Xu
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Legal intelligence aims to support reliable decision-making across long-horizon legal processes involving evolving case states and multiple roles. However, real-world legal deployment exhibits substantial case heterogeneity in facts, evidence, and procedural contexts, exposing the limitations of static agent strategies. Moreover, legal reasoning is inherently interdependent across roles and procedural stages, making global reliability fundamentally different from isolated role competence. To address these challenges, we study training-free test-time agent adaptation, where agents continuously exploit deployment-time signals from preceding cases and ongoing interactions without updating model parameters. We propose \method, which introduces \emphTest-Time Memory Evolution to retrieve reusable experience from previous cases, adapt it to the current factual and procedural context, and consolidate accumulated experience for subsequent decision-making. Further, \emphRubric-Aligned Collaboration verifies and revises role-specific actions according to behavioral and procedural requirements, enabling coordinated decision-making across roles and stages. Extensive experiments on J1-EVAL and LegalWorld across five backbone models demonstrate consistent improvements over representative reasoning and agent baselines with reasonable interaction and computational costs. Ablation and case studies further show that the two components provide complementary benefits in experience adaptation and cross-role coordination, improving the reliability and efficiency of long-horizon legal reasoning.
[AI-63] Beyond Waypoint Regression: Query-Based Cost Learning over Reachable Ego Futures for End-to-End Driving ACCV2026
链接: https://arxiv.org/abs/2610.08123
作者: Ahmed Abouelazm,Rupert Polley,Qingyuan Zhang,Yin Wu,Philip Schörner,Carl Esselborn,J. Marius Zöllner
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Systems and Control (eess.SY)
备注: Accepted in the 18th Asian Conference on Computer Vision (ACCV 2026)
Abstract:End-to-end planners based on waypoint regression achieve strong open-loop accuracy, but they primarily learn to mimic expert geometry and remain difficult to adapt to deployment-time safety constraints. We propose a query-based cost-learning framework that estimates bounded costs for dynamically reachable ego trajectory queries, rather than dense BEV cells or a small regressed trajectory set. Compact joint scene tokens capture coherent multimodal agent futures, while contingency-aware cost aggregation and cost-guided intra-cluster MPPI mixing convert the learned cost topology into feasible ego plans. On nuScenes, our method improves over prior cost-estimation planners such as ST-P3 and NMP, outperforms most regression baselines in collision rate, while remaining competitive in L2, and retaining an interpretable cost interface. On real-world driving logs, the proposed planner reduces collision rates compared with SparseDrive and Alpamayo without fine-tuning, while maintaining a diverse set of candidate trajectories.
[AI-64] Exploiting Acoustic and Content-Oriented Speaker Verification Attacks Against Multilingual Voice Anonymization
链接: https://arxiv.org/abs/2610.08107
作者: Ridwan Arefeen,Ze Li,Rong Tong,Ming Li,Xiaoxiao Miao
类目: ound (cs.SD); Artificial Intelligence (cs.AI)
备注: Accepted in IEEE Spoken Language Technology (SLT) 2026
Abstract:Attacker ASV systems for voice anonymization have been studied primarily in English, leaving their behavior in multilingual settings largely unexplored. Conventional ASV has shown that both acoustic and contextual information are important for multilingual speaker verification. Inspired by this, we investigate whether the same holds for attacker ASV on anonymized speech. We evaluate both acoustic- and content-oriented attackers on multilingual anonymized speech and construct a multilingual voice-converted dataset to improve cross-lingual generalization. Our results show that attacker effectiveness depends on the linguistic utility of the anonymized speech. Overall, acoustic-oriented attackers achieve better performance. However, when linguistic information is well preserved, the performance gap between content- and acoustic-oriented attackers narrows compared with conditions involving stronger speech distortion. The multilingual voice-converted dataset further improves performance and partially reduces the cross-lingual gap. These findings highlight the need for more comprehensive attacker modeling and evaluation protocols that consider both privacy and utility, rather than relying on a attacker strategy\footnoteFull code and pretrained models and MultiVC Dataset link are available at: this https URL
[AI-65] ChartBmkAgent : Harness-Governed Multi-Agent Construction of Chart QA Benchmarks from Sparse Error-Taxonomy Specifications
链接: https://arxiv.org/abs/2610.08106
作者: Langxi Huang,Pingping Zhang,Lanyun Zhu,Chunyang Jiang,Jiawei Shao,Haocheng Yuan,Peilin Chen
类目: Artificial Intelligence (cs.AI)
备注: 9 pages, 2 figures
Abstract:Multimodal large language models (MLLMs) advance rapidly, while conventional benchmark development lags behind, delaying investigation of newly observed capability gaps. Such investigation requires an expressive task format and an on-demand construction process: information-rich charts make chart question answering (Chart QA) suitable for probing coupled perception and reasoning. Automated Chart QA construction is intended to shorten the benchmark-development cycle by turning identified gaps into targeted samples on demand. Current methods, however, commonly separate target guidance from scratch generation: target-guided systems often require prepared data, charts, or templates, while scratch-generation systems primarily ensure artifact validity, without explicitly controlling whether newly synthesized requirements and content remain aligned with an externally specified diagnostic target. We introduce ChartBmkAgent, which turns an identified capability gap into targeted diagnostic evidence by constructing complete Chart QA samples from sparse error-taxonomy specifications. Throughout construction, a central harness governs specialized agents, requires stage-specific evidence of alignment with the original error category, and records the basis for each acceptance decision. On 300 taxonomy-wide samples, MLLM accuracies ranged from 32.7% to 84.3% with distinct category profiles, showing that generated samples reveal capability differences. Across three source-model comparisons, targeted follow-ups scored 50.0% versus 82.2% on matched controls ( p=8.96\times10^-6 ); all six cross-model comparisons had the same direction, demonstrating targeted validation and diagnostic-data generation. Multiple evaluator models assessed whether each sample tested its specified error category; 86.4% met this criterion, providing empirical evidence of target preservation.
[AI-66] DSV-Mem: Evaluating Multimodal Memory in Professional Workflows for MLLM Agents
链接: https://arxiv.org/abs/2610.08102
作者: Jike Zhong,Ritwick Chaudhry,Xuanbai Chen,Tianchen Zhao,Linghan Xu,Yifan Xing,Nishant Sankaran
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Conversational MLLM agents are increasingly expected to assist in professional workflows, from AI research and engineering design to product management and business operations. Yet this capability remains underexplored: existing benchmarks largely focus on informal, everyday interactions and personal-life scenarios featuring photographic natural images, isolated static artifacts, and recall-oriented questions. In contrast, professional scenarios often involve structured, information-heavy artifacts that undergo frequent revisions and authority updates, and compositional queries requiring reconciliation of many artifact versions while tracking state precisely. To address these challenges, we introduce DSV-Mem, a benchmark for evaluating Dense Stateful Visual Memory. DSV-Mem comprises expert-reviewed scenarios and 1,000 questions across five user-oriented categories (Current State, Past State, Derived State, Change History, and Conflict/Refusal). A Hartley-inspired criterion favors questions with broader visual-evidence inspection demands. We also introduce a generation harness that produces evaluation suites by decoupling state-transition synthesis from conversation filling. Evaluation over 27 configurations spanning frontier and open-weight models and memory management methods reveals that the strongest baseline scores below 45% on DSV-Mem. Analysis surfaces findings: 1) multimodality and information density both contribute to difficulty, but state evolution, particularly the number of governing updates, is the dominant tested factor. Raw conversation/haystack length, OCR, and arithmetic are not the primary bottlenecks; 2) models often fail to verify user premises against prior state updates before answering; 3) increased reasoning effort and memory management methods yield limited gains, whereas state-aware designs prove more effective. The benchmark and code will be publicly released.
[AI-67] Beyond Corrected Memory: Execution Consistency in Multi-Agent Systems
链接: https://arxiv.org/abs/2610.08101
作者: Zhe Yu,Zixuan Wang,Peidong Wang,Hehai Lin,Ruochen Zhao,Chengwei Qin
类目: Artificial Intelligence (cs.AI)
备注: 39 pages, 7 figures, 30 tables (including appendix)
Abstract:Shared memory coordinates agents’ actions, but correct records do not establish that those actions satisfy task requirements. Memory governance and failure diagnosis regulate or inspect recorded information; they do not by themselves establish whether it is sufficient to judge task duties. We define execution consistency through duties governing state use, information handoffs, and final-state agreement, with explicit evidence conditions for judging fulfillment. Our core claim is that identical retained records can correspond to compliant and violating executions under the same task rule. Controlled removal of evidence such as receipt, action dependence, or response validity leaves 82.4% of opposite-label pairs indistinguishable; restoration separates 97.9% of the merged pairs. Natural-log annotations identify the defined violations in actual executions. However, existing logs do not always explicitly represent the execution relationships needed for these judgments. To assess the definition’s practical value, we use CAVERT, a framework for consistency diagnosis and recovery, to extract supported relationships from logs and apply these criteria. It consistently outperforms contract-prompted LLM and rule-based baselines in diagnosis across all 12 benchmark-executor settings. Under the same gate and executor limits, it also outperforms rule-guided recovery in all four evaluated environments. These findings identify execution evidence that agent-memory and execution interfaces should preserve for reliable judgment.
[AI-68] When Tools Lie: Reliability of Mathematical Agents Under Corrupted Tool Feedback NEURIPS2026
链接: https://arxiv.org/abs/2610.08097
作者: Kavienan Jegatheesan,Gayathri Lihinikaduarachchi
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Software Engineering (cs.SE)
备注: 8 pages, 2 figures, 3 tables. Accepted to the 6th Workshop on Mathematical Reasoning and AI (MathAI) at NeurIPS 2026
Abstract:Mathematical problem solving often requires deterministic computational steps that agents delegate to tools and implicitly trust. Yet tools can fail silently, returning plausible but incorrect results. How well can agents detect and correct corrupted tool call outputs? We study this through a controlled corruption framework where a hidden interceptor replaces tool call results with plausible incorrect information on targeted problems. We evaluate agents across 31 problems under four verification designs including no verification (baseline), mandatory same-context reflection, optional fresh-context verification, and optional structural verification. Without verification, corruption causes dramatic accuracy loss, from 100% down to 72.4%. Mandatory reflection fully recovers this performance to 100%. Optional verification improves accuracy only when models actively invoke it. Our results show that checking frequency is strongly associated with robustness differences, while unequal invocation prevents a controlled comparison of verifier quality. A supporting recovery experiment shows that full problem restart succeeds in 100% of cases after explicit detection. These findings demonstrate that verifier availability and verification policy are separate components of mathematical-agent reliability. Mandatory policies enforce verification while optional policies depend on the model’s own choice to invoke it.
[AI-69] When Plans Change Answers: Formalizing Cost-Accuracy Optimization for Semantic Queries
链接: https://arxiv.org/abs/2610.08089
作者: Kyoungmin Kim
类目: Databases (cs.DB); Artificial Intelligence (cs.AI)
备注:
Abstract:In semantic query engines, predicates are evaluated by machine-learned models, and the choice of a query plan affects not only the cost of a query but also its result. Existing systems either apply a fixed threshold to each semantic operator or tune accuracy per operator, without accounting for how errors propagate through joins. We give a formal problem definition for cost-accuracy optimization of such queries. Our starting point is the calibrated confidence that decision models such as Jev attach to each decision. It yields an expected error for every decision; weighting these errors by each decision’s contribution to the output (in the simplest case, its fan-out) gives the expected output quality of a plan without any labeled data, and the same computation in reverse turns an output-level accuracy target into a price on each base or intermediate tuple. Building on this, we define an oracle semantics for relational algebra with semantic operators, physical plans as pairs of a logical plan and a decision policy, declarative output-level targets, and a hierarchy of plan equivalence. We show that accuracy is plan-invariant under pointwise-deterministic policies, and that selection pushdown is not quality-sound when escalation bands are calibrated on the plan’s own candidates. Expected quality can be computed in polynomial time under bag semantics; under set semantics it follows the dichotomy of tuple-independent probabilistic databases when every relation carries a semantic predicate. Choosing which tuples to drop is NP-hard, while the optimization problem decomposes into per-tuple decisions through two Lagrange multipliers. Simulations on a synthetic workload illustrate these effects; an evaluation on real engines is left for future work.
[AI-70] SpeedrunBench: Challenging LLM Agents with Video Game Speedrunning
链接: https://arxiv.org/abs/2610.08076
作者: Yoshinari Fujinuma,Keisuke Kamahori,Ryuto Koike,Abdelrahman Madkour,Varun Prashant Gangal,Monty Bichouna,Martyna Markiewicz,Shivani Jain,Duncan Curtis,Rebecca Qian,Anand Kannappan
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Frontier LLM agents have been shown to be capable of solving increasingly complex tasks for which humans have measurable solutions. This begs the pertinent question of whether LLM agents can go beyond what humans have already solved. The ability to develop sophisticated strategies to tackle consequential problems becomes paramount as well-trodden, human-developed solutions become insufficient for problems for which we lack context or enough training data. We study agents’ capability of such strategy formation through the communal practice of video game speedrunning. In speedrunning, practitioners compete to find the fastest way to complete a video game under certain conditions, and in so doing uncovering interesting unorthodox play styles that require a thorough understanding and mastery of the underlying game mechanics. We introduce SPEEDRUNBENCH, a benchmark that evaluates frontier LLM agents across 9 different games. To perform well in this benchmark, agents must repeatedly improve their strategy, reflect on their performance, exploit their gained knowledge, and reason across a long-horizon of actions to improve on an increasingly difficult problem: being faster than themselves and everyone else. Our experiments show that while frontier agents approach human world records in simple platformer games, they remain behind human performance on longer, more complex games under practical budgets. These results suggest that SPEEDRUNBENCH is a useful testbed for studying agents’ strategy formation capabilities as well as being a saturation-resistant evaluation measure, as there is almost always a faster completion time waiting to be discovered.
[AI-71] Same Feedback Different Answer: Measuring Run-to-Run Instability in Frontier-Model Customer Feedback Analysis
链接: https://arxiv.org/abs/2610.08036
作者: Viraj Bagal,Raviraja Ganta,Prabhath Chellingi
类目: Artificial Intelligence (cs.AI)
备注: 9 pages
Abstract:AI agents are increasingly being programmed to automate knowledge work over large collections of unstructured data. Such automation requires repeatability: when the underlying evidence is unchanged, the agent’s categories, priorities, and counts should not shift materially between runs, even if each individual answer appears plausible. We introduce a repeat-run evaluation framework that aligns semantically equivalent categories and focuses on two operating metrics: theme churn, the normalized change in the returned category set, and volume disagreement, the change in counts for categories that persist. We evaluate three recurring customer-feedback tasks across eight frontier models, corpus sizes from 100 to 5,000 records, multiple prompts, and three execution designs: raw generation, taxonomy-free hierarchical decomposition, and a taxonomy-grounded agent (TGA) using persistent themes, subthemes, and record-level predictions. With Claude Opus 4.8 and the 1,000-record corpus fixed, TGA reduces theme churn by 86–88% relative to both raw generation and hierarchical decomposition, while matched-theme volumes have zero disagreement. The taxonomy-grounded agent is more stable than every raw model in the screen, remains more stable at each corpus size, and keeps this advantage when theme matching is made stricter or looser. Although evaluated on customer feedback, the framework targets repeated synthesis of unstructured corpora more broadly, including financial reports, legal documents, incident records, and scientific literature. Overall, these results show that taxonomy grounding produces more consistent and repeatable outputs for recurring knowledge work.
[AI-72] Learning in Dreams Winning in Reality: A Continuous Dyna Loop for a Ten-Hero MOBA
链接: https://arxiv.org/abs/2610.08033
作者: Jordy Kieto
类目: Artificial Intelligence (cs.AI)
备注: 15 pages, 11 figures. Code, policies, evaluation and videos: this https URL
Abstract:World models are usually judged from the inside: by prediction loss, by the return a policy earns in imagination, or by how convincing their frames look. We judge one from the outside. We learn a structured, multi-agent world model of a complete ten-hero MOBA (206 units, every hero acting every tick, games of up to 6,000 ticks), train a policy only inside it with 1,400-tick free-running imagined episodes, and measure that policy in the real game against the opponent the game ships with. The real game never provides a gradient; it provides the policy’s own games as training data for the world model, and an online evaluation that selects and anchors the policy. Run as a continuous asynchronous Dyna loop, the policy wins 70.2% of real games as radiant (421 of 600; 95% CI 66.4-73.7) on seeds never used for any decision, up from 0% for dream training alone and 33.7% before the loop. It wins none as dire, and neither does the shipped opponent when it plays itself. Four findings explain the result. Model exploitation is invisible from inside the dream: every unanchored run collapsed within a few updates while no in-dream metric tracked the collapse. A world model that is accurate on its training corpus is badly wrong on the policy’s own games, and Dyna repairs it there, which is worth +9.2 points of real win rate with the policy recipe held fixed. Finally, the policy inherits its world model’s fidelity profile mechanic by mechanic: the model represents the macro game but not crowd control, cast timing or lethality, and the policy wins by map-wide pressure with almost no coordinated fighting. We release the world model, the dream-PPO harness, a world-model debugger, the evaluation protocol, and every policy and log.
[AI-73] ICDA: Tabular In-Context Data Attribution
链接: https://arxiv.org/abs/2610.07996
作者: Yacine Benihaddadene,Milan Bhan,Eliot Dugelay,Mohammed Jawhar,Benjamin Wong,Nicolas Chesneau,Duong Nguyen
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:
Abstract:Tabular foundation models (TFMs) achieve strong predictive performance by conditioning on labeled demonstrations provided in context, without any parameter update. Yet how individual demonstrations shape a given prediction remains poorly understood. This gap matters in practice: the context is often assembled from whatever labeled data is available, potentially leading to the inclusion of mislabeled, redundant, or low-quality examples that degrade performance. Standard data attribution methods do not transfer to the TFM setting: resampling-based approaches such as DemoShapley require a combinatorial number of forward passes, and gradient-based estimators such as influence functions require computing training point’s effect on the model parameters, which in-context learning never updates. We introduce TICDA, a method that measures the influence of every demonstration in the context directly from linear surrogates trained on TFM latent embeddings, in a single forward pass and at negligible cost. We show that TICDA offers the best compromise against competitors across four tasks: detecting labeling errors, curating context to preserve predictive accuracy while lowering inference cost, producing attribution scores that transfer across TFMs, and supporting an acquisition strategy for efficient active learning.
[AI-74] Learning from Revision Consequences: Hindsight Meta-Experience Distillation for Self-Improving Agents
链接: https://arxiv.org/abs/2610.07979
作者: Qianhan Feng,Zhongzhen Huang,Yakun Zhu,Xiaofan Zhang,Qi Dou
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:As agents continuously improve by generating and revising Skills, the process that discovers and refines those Skills becomes a learnable object in its own right. Task-Skills directly act on task execution, whereas Meta-Skills govern how agents discover and improve future Skills; their value therefore emerges through the subsequent search processes they induce. Existing approaches improve Meta-Skills from observed raw Skill-search trajectories and branch outcomes. However, branch performance entangles the effects of the initial discovery state and the Meta-Skill revision that generated the search process, making it difficult to characterize what a particular revision actually changed, and pushing updates toward revisions that benefit from favorable states rather than those that improve the process. We introduce HMED (Hindsight Meta-Experience Distillation), a mechanism for constructing Meta-Experience for self-improving agents. HMED revisits the completed event from which a revision originates and re-executes the incumbent and revised Meta-Skills from the same restored discovery state, so that the changes associated with the revision can be observed under a shared condition. Each comparison is distilled into a Meta-Experience, a structured record that can be reused by future updates, so that even revisions that are not ultimately retained still contribute a learning signal. Across three interactive agent benchmarks and both open-source and closed-source models, HMED consistently improves Skill discovery performance over strong baselines, shifting Meta-Skill learning beyond branch outcomes toward the consequences of changing the improvement process.
[AI-75] Can Agents Work for Everyone? Cross-User Reliability for Mobile GUI Agents in Personalized User Interfaces
链接: https://arxiv.org/abs/2610.07972
作者: Yeji Park,Jaeyun Shim,Taesik Gong
类目: Artificial Intelligence (cs.AI)
备注: 23 pages, 12 figures
Abstract:Mobile GUI agents increasingly operate on interfaces influenced by users’ histories and preferences, but their reliability across different users remains underexplored. We introduce PAIR (Personalized Application-state Instantiation and Rendering), a pipeline for constructing user-conditioned application states that enables controlled evaluation of the same task across different users. We further introduce RePAIR (Reinforcement learning with Personalization-Aware Interaction Rewards), a training approach that learns from cross-user differences in subgoal outcomes to improve reliability across user-conditioned mobile environments. Across six agents, we find substantial variation in task success across users and consistently lower subgoal achievement in user-conditioned UI contexts (6.98 to 15.4 pp). This gap further increases for personal targets drawn from each user’s own content (8.77 to 22.0 pp). Failures in these contexts frequently involve selecting another item instead of the intended target, particularly before target exposure. Finally, RePAIR improves user-conditioned SAR (+5.87 pp), all-success (+7.50 pp), and overall Task SR (+9.42 pp) over its supervised fine-tuning parent on unseen users, providing initial evidence that explicitly learning from cross-user variation can improve GUI-agent reliability.
[AI-76] ReGraph: A Computational Account of Emergent Generalization in the “what” and “where” Dual Visual Streams
链接: https://arxiv.org/abs/2610.07962
作者: Hyewon Kang,Jungmin Lee,Ilgyu Lee,Seok-Jun Hong
类目: Neural and Evolutionary Computing (cs.NE); Artificial Intelligence (cs.AI); Neurons and Cognition (q-bio.NC)
备注: 29 pages, 6 figures
Abstract:Where generalization capacity–the ability to extract context-invariant relational structures–first emerges remains a central question in AI and neuroscience. The foundation for this capacity lies upstream of the hippocampus, within the entorhinal cortex, where parallel pathways dissociate relational structure in the medial entorhinal cortex (MEC) from sensory content in the lateral entorhinal cortex. However, as Eichenbaum argued, such factorization likely originates earlier, driven by the segregation of the dorsal (‘where’) and ventral (‘what’) visual streams. Supporting this, grid-like firing patterns–a signature of MEC (context-invariant codes)–also appear in preceding neocortical regions along the dorsal pathway. Yet, how such representations are computationally formed along upstream pathways remains unknown. To investigate this in silico, we developed ReGraph, a recurrent dual-stream graph model with biological inductive biases, including retina-driven stream-specialized encoding, dorsal-to-ventral modulation, and dynamic lateral connectivity. Trained on the action benchmark Something-Something V2, ReGraph revealed a pathway-specific emergence of relational mapping: context-invariant codes and grid-like spatial bases uniquely co-emerged along the extended dorsal stream. In contrast, their absence in single-stream, unmodulated variants, and standard baselines implies that these inductive biases are prerequisites for relational structures. Crucially, our post-hoc analyses demonstrated that these grid-like bases serve as reusable routing templates for information processing via lateral connectivity. Together, our findings provide a computational account that generalization may not be a faculty that emerges abruptly within a dedicated region, but a property that already takes shape as sensory information is parsed into factorized streams of hierarchical visual processing.
[AI-77] Adapting Vision-Language-Action Models to Unknown Visual Disruptions During Execution
链接: https://arxiv.org/abs/2610.07946
作者: Ahin Lee,Jinwoo Seo,Youngsoo Jang,Taesik Gong
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI)
备注: 22 pages, 15 figures, 13 tables
Abstract:Visual disruptions can arise while a robot is executing a task, leaving a vision-language-action (VLA) policy to respond without knowing the disruption type or timing. We introduce Self-supervised Adaptation from Leftover Trajectories (SALT), which uses the leftover trajectory, the unexecuted part of the previous action chunk, as self-supervision for test-time adaptation. Because consecutive chunks overlap in time, the leftover provides a temporally aligned target for the current prediction over the same future control interval. At the onset of a visual shift, the leftover can retain a plan formed before the corruption, so updating the policy toward it anchors the adaptation across the shift (Transition Anchoring). SALT keeps the adapted policy and regenerates the current chunk, whose leftover becomes the target at the next replan, carrying the correction forward along the execution trajectory (Sequential Correction Propagation). Supervision comes entirely from the policy’s own predictions, requiring no disruption annotations, expert actions, or target-domain demonstrations, and a lightweight adaptation gate calibrated only on nominal trajectories decides when updates begin. On LIBERO-10, SALT increases average success across five persistent visual corruptions from 43.9% to 53.2% with SmolVLA and from 58.7% to 66.0% with GR00T N1.7, while largely preserving nominal performance. On a real robot, it raises task progress averaged over digital and physical disruptions from 0.49 to 0.61.
[AI-78] SIGMA: Self-Improving Alignment Generalization from a Model Spec
链接: https://arxiv.org/abs/2610.07935
作者: Jingyu Zhang,Shruti Palaskar,Daniel Khashabi,Benjamin Van Durme,Leon A. Gatys,Joseph Yitan Cheng
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:LLM agents are increasingly capable of executing complex tasks and of recursively improving themselves on easy-to-verify objectives such as software engineering and mathematics. Since alignment is much harder to verify, this creates a growing risk of capabilities increasing without appropriate safety alignment, especially as capabilities expand to auto-research and cybersecurity. Existing approaches focus on capability self-improvement using verifiable feedback or on alignment training with supervision from stronger models or curated data, creating an external supervision bottleneck for alignment. We ask whether current models can improve their own safety alignment, and propose SIGMA, a data generation and training pipeline enabling alignment self-improvement that generalizes to out-of-distribution settings. Given only a “Model Spec” stating the model’s desired behavior, SIGMA leverages a model’s reasoning capabilities to strengthen its own safety reasoning. SIGMA first performs spec-guided task synthesis, using the candidate model as a task designer agent to generate diverse alignment dilemma scenarios and convert them into training tasks that stress-test its understanding of the Model Spec. Next, SIGMA conducts self-judged alignment training through supervised fine-tuning and rubric-based reinforcement learning with the model itself as the reward model. Despite training only on single-turn chat data, SIGMA improves safety alignment in multi-turn agentic environments (AgentHarm harmfulness decreases from 22.6 to 14.8; Agentic Misalignment decreases from 79.1 to 3.8), outperforms Deliberative Alignment and Constitutional AI baselines, and retains general capability. Analyses show that a Model Spec balancing harmlessness and helpfulness, test-time reasoning for safety deliberation, and high-quality rubrics from SIGMA’s task designer agent are crucial for effective self-improvement.
[AI-79] Continuous Memory Machines NEURIPS2026
链接: https://arxiv.org/abs/2610.07907
作者: Ciaran Regan,Kai Arulkumaran,Luke Darlow,Stefania Druga,Sebastian Risi,Llion Jones
类目: Artificial Intelligence (cs.AI)
备注: NeurIPS 2026 Workshop: Personalized, Aligned, Long-Term Memory for AI Systems (PALM)
Abstract:Recurrent neural networks typically compress information into a single vector-valued recurrent state, forcing short-term computation and long-term retention to share the same representation. Past extensions alleviate this bottleneck by increasing the memory capacity or separating timescales, but lack the combination of rapid neuron-level processing and longer-term retention found in biology. To that end, we introduce the Continuous Memory Machine (CMM), a recurrent architecture with matrix-valued short- and long-term memory states serving distinct functional roles. Building on the Continuous Thought Machine (CTM), the CMM’s short-term memory tracks recent neural activity, with uniquely parameterized neuron-level models learning to use these activity patterns for computation. A persistent long-term memory stores information for later use, with a Transformer jointly updating both memory stores, providing an expressive bidirectional read–write mechanism such that each store can reorganize its own contents and both read from and write to the other. Across algorithmic, in-context learning, and recurrent reasoning tasks, the CMM outperforms a broad suite of baselines, exhibiting stronger generalization than prior memory-augmented networks while preserving the CTM’s interpretable attention patterns. Code is available at this https URL.
[AI-80] IEEE 802.11bx - WLAN Intelligent Networking (WIN): Toward an AI-Ready Wi-Fi 9
链接: https://arxiv.org/abs/2610.07900
作者: Francesc Wilhelmi,Katarzyna Kosek-Szott,Szymon Szott,Boris Bellalta
类目: Networking and Internet Architecture (cs.NI); Artificial Intelligence (cs.AI)
备注:
Abstract:Wi-Fi 9 is expected to go beyond mere communication and provide new services such as sensing or computation. At this juncture, Artificial Intelligence (AI) is taking a leading role in the definition of the 802.11bx amendment, named WLAN Intelligent Networking (WIN). In this tutorial, we survey the recent progress made toward Wi-Fi 9 within IEEE 802.11 standardization, tracing the drivers and technological advances that motivate an AI-ready Wi-Fi 9. We then examine AI’s role along three complementary dimensions, i.e., AI as a protocol (AI is applied to Wi-Fi’s PHY/MAC operation), AI as a platform (Wi-Fi infrastructure is repurposed to provide AI computation), and AI as traffic (AI flows call for new traffic-handling policies), and discuss candidate features and open challenges along each. As a concrete illustration of the AI as traffic paradigm, we present a case study on AI traffic differentiation, where we explore a potential extension of the current Enhanced Distributed Channel Access (EDCA) to support new AI traffic flows.
[AI-81] Variance-Averse n-Step Offline Reinforcement Learning for Sparse Long-Horizon Environments NEURIPS2026
链接: https://arxiv.org/abs/2610.07899
作者: Guhyeon Kang,Minhae Kwon
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: Accepted at NeurIPS 2026
Abstract:Generative actors are transforming offline reinforcement learning (RL) by enabling expressive policy classes that model complex action distributions. However, this expressiveness also exposes a key challenge in heterogeneous datasets: generative policies can reproduce unreliable action modes whose return distributions exhibit high variance, occasionally yielding high returns by chance but lacking consistency. Consequently, maximizing the expected Q -value alone is insufficient for identifying reliable actions. We propose VAN-Flow (Variance-Averse n -step Flow), a framework that promotes reliable actions in generative offline RL. VAN-Flow combines (i) a categorical distributional critic, (ii) a variance-averse expectation operator that smoothly reweights atom probabilities to favor actions with both high returns and low dispersion, and (iii) a flow-matching generative actor guided via rejection sampling. Unlike CVaR or mean-variance objectives, the operator redistributes probability mass over the categorical return distribution without hard truncation or auxiliary penalty terms. Across more than 40 tasks from D4RL and OGBench, VAN-Flow consistently outperforms strong baselines, with the largest gains in long-horizon and high-variance regimes where reliable action selection becomes critical.
[AI-82] xtual Environmental Context and Spatial Graphs for LLM -Based Regional SST Forecasting
链接: https://arxiv.org/abs/2610.07895
作者: Xiong Li,Xiaowei Zhou,Yanwei Yu,Qian Cui,Junyu Dong
类目: Artificial Intelligence (cs.AI)
备注: preprint
Abstract:Sea surface temperature (SST) forecasting depends on local temporal persistence, regional spatial dependence, and environmental conditions that evolve with the forecast date. We study how these heterogeneous conditions can be presented to a large language model (LLM) for regional multi-step forecasting without serializing the full SST grid as text. We formulate forecasting as conditional numerical generation: historical SST and anomaly sequences, date-aligned environmental records, and static ocean knowledge form a textual context, while regional spatial state is supplied through continuous graph-derived prefixes. A static graph encodes persistent geographic–climatological relations, and a dynamic graph encodes recent SST correlations and localized tropical-cyclone influence. Two graph neural networks produce a target-node representation that is mapped by a spatial-prefix fusion and injected into the LLM input. On SST forecasting in the South China Sea, the complete configuration achieves the best MAE and \Rtwo among the compared methods over ten forecast steps. Alongside the numerical forecast, a rule-based module matches predicted trends and environmental-factor directions with knowledge entries to return source-linked, post-hoc contextual explanations.
[AI-83] WorkflowOps: Learning Agent Collaboration Priors for Multi-Agent Workflow Orchestration
链接: https://arxiv.org/abs/2610.07860
作者: Qi Cheng,Shengyu Chen,Wei Cheng,Zhengzhang Chen,Xiaowei Jia,Haoyu Wang,Haifeng Chen
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Multi-agent systems are increasingly deployed for complex knowledge work, yet their orchestration layers remain largely memoryless: each new task is decomposed, assigned, and executed from scratch with no benefit from prior successful executions. We present WorkflowOps, a multi-agent workflow orchestration framework that learns agent collaboration priors from historical workflows and expands its agent pool on demand to cover new capability requirements. Our approach introduces three coupled mechanisms. First, a transition probability matrix captures pairwise agent collaboration frequencies from past workflows and applies them as soft guidance during DAG workflow construction through intra-layer ordering optimization, probability-thresholded edge suggestion, and transitive reduction for parallelism maximization. Second, a sufficiency-driven agent creation loop detects capability gaps via semantic matching scores, generates specialized agents through an LLM, and simultaneously injects them into the collaboration matrix, so that newly created agents are immediately usable with predicted collaboration priors. Third, a layered semantic matching strategy uses pre-trained sentence embeddings for fast, deterministic capability matching as a first pass, invoking LLM verification only for low-confidence cases, thereby reducing LLM routing calls by over 80% compared to pure-LLM approaches. Experiments on mixed code, math, and question-answering suites show that WorkflowOps improves end-to-end pass rates over recent workflow-construction baselines, with the largest gains on structured, decomposable tasks where past agent handoff patterns transfer.
[AI-84] RA-MoWE: Workflow-Affinity Embeddings for Query Clustering and Agent ic Workflow Generation
链接: https://arxiv.org/abs/2610.07851
作者: Qi Cheng,Shengyu Chen,Wei Cheng,Yiqun Xie,Haoyu Wang,Haifeng Chen,Xiaowei Jia
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Agentic workflows enable large language models (LLMs) to solve complex tasks by coordinating reasoning, tool use, and verification. However, a workflow optimized for an entire task collection can overlook differences in the reasoning strategies that individual queries need, while searching for a new workflow for every query repeats costly optimization. To address this tradeoff, we introduce RA-MoWE, a framework that uses workflow-affinity embeddings to cluster queries and guide the generation of reusable expert workflows. Each embedding records how well a fixed set of reference workflows solves a query, revealing similarities in which reasoning strategies are effective. RA-MoWE uses each cluster’s queries and average embedding to initialize and refine a specialized workflow through execution feedback. An embedding encoder predicts these embeddings from query text, allowing new queries to select a generated expert without first executing the reference workflows. On a 300-query test set drawn from four benchmarks spanning mathematics, science, and programming, RA-MoWE improves average task score by 4.04 percentage points over selecting among the reference workflows, while using 27.7% fewer language-model calls at inference.
[AI-85] DHCG: Dynamic Construction of Hierarchical Collaboration Graphs for LLM -Based Multi-Agent Reasoning
链接: https://arxiv.org/abs/2610.07835
作者: Jie Ren,Jiakang Yuan,Chenyu Huang,Hezeer Ma,Jiayuan Fan,Tao Chen
类目: Artificial Intelligence (cs.AI)
备注: 9 pages, 4 figures, 4 tables
Abstract:LLM-based multi-agent systems (MAS) have demonstrated strong capabilities in solving complex problems across diverse domains. Recently, the dynamic orchestration of agent systems has become an important research direction. However, existing methods suffer from limited composition, misaligned dependencies, and inflexible scale, restricting their ability to adapt to reasoning requirements during execution. To address these limitations, we reframe MAS design as a partially observable Markov decision process, in which both the composition and scale of the MAS are dynamically determined. We propose DHCG, a novel framework that coordinates three modules (Planner, Worker, and Generator) to progressively construct a dynamic hierarchical collaboration graph from scratch based on the query and evolving execution feedback. At each step, guided by feedback, the Planner generates a set of distinct and complementary roles tailored to the current reasoning needs and selectively routes relevant information to each role. It can also finalize the hierarchical collaboration graph early or progressively expand it when additional reasoning is required. We further introduce action-aware preference optimization to train the Planner to make more effective decisions when constructing hierarchical collaboration graphs. We systematically evaluate DHCG across code generation, mathematical reasoning, and domain-specific reasoning benchmarks. DHCG achieves state-of-the-art average performance among the compared methods, improving over the single-agent baseline by 13.06 points and outperforming both static and dynamic MAS baselines by 2.77-8.02 points. Additional experiments further demonstrate its generalization across different Planner backbones and unseen Worker models.
[AI-86] Agent ic Semantic Sensing for Resource-Adaptive AI-RAN
链接: https://arxiv.org/abs/2610.07829
作者: Zhongqin Wang,Xiaoqi Zhang,Nan Yang,Kai Wu,J. Andrew Zhang,Y. Jay Guo
类目: Artificial Intelligence (cs.AI); Signal Processing (eess.SP)
备注:
Abstract:Semantic sensing (SemS) acquires task-relevant information rather than reconstructing complete physical information. Existing SemS formulations typically operate open loop: sensing configurations and observation schedules are fixed before inference and cannot respond to evolving task-level evidence. We propose Agentic SemS, a closed-loop framework for AI-enabled radio access networks (AI-RANs) that controls sensing within a communication-feasible profile set. A profile-conditioned causal Transformer updates the semantic belief from streaming observations, while key-value caching enables efficient state updates across profile changes without repeatedly processing the complete history. A semantic utility network estimates the task-level benefit of acquiring the next observation block under each feasible profile after accounting for sensing cost. The resulting continuation utilities jointly support next-profile selection and semantic early exit, adapting sensing configuration and duration to evolving evidence. The expected semantic gain is further related to conditional mutual information, providing a value-of-information interpretation of continued online sensing. Experiments on Widar3.0 with six emulated sensing profiles show that, in comparison with full-sequence High, the resource-efficient Agentic setting reduces normalized cumulative sensing cost by 25.33% while achieving 85.79% Macro-F1. At the same utility checkpoint, semantic early exit provides a further 12.35% cost reduction over adaptive sensing without early exit, with a 0.97-percentage-point Macro-F1 decrease.
[AI-87] hin Evidence Thick Priors: How Language Models Substitute Identity for Missing Financial Facts
链接: https://arxiv.org/abs/2610.07798
作者: Saanvi Khetan,Sankar Balasubramanian
类目: Artificial Intelligence (cs.AI)
备注: 49 pages, 17 figures, 12 tables; Submitted Accepted to ICAIF’2026
Abstract:People increasingly ask large language models what to do with their money, yet seldom describe their finances in full. This paper asks what a model does with the gap. Holding finances fixed and changing only who the investor is said to be, we grade the financial evidence in the prompt from eight facts to none and measure how far the recommended equity allocation moves. Across 96,600 prompts to Llama-3.1-8B-Instruct, built from 100 financial profiles, 138 personas and seven disclosure conditions, the average gap between two personas with identical finances rises from 4.78 percentage points at full disclosure to 10.34 points with no financial facts. A two-way cluster bootstrap counting duplicated prompts once places the ratio at 2.16 (95% interval 1.69 to 2.79), and the rise is already 1.69-fold with a single fact left. Identity explains 5% of within-profile variation in advice at full disclosure and 96% with no disclosure. Household size is the only attribute whose influence grows reliably as evidence is withdrawn. Once standard errors are clustered on the persona, the unit to which identity was assigned, most attribute-specific interactions reported in the conference version lose significance, and gender instead appears as a small standing gap that full disclosure does not close. Stating risk appetite alone brings the swing into the range seen with two to seven generic facts. With no facts, the model’s one-line rationale cites incomes, debts and savings it was never told, and these invented finances turn adverse more often for larger households. Inside the network, gender is linearly decodable at every layer, and ablating the gender direction at five layers leaves the aggregate identity swing unchanged. Advisory systems built on such models should be audited at the disclosure levels users actually reach, and judged across the whole identity space rather than one attribute at a time.
[AI-88] he Geometry of Empowerment
链接: https://arxiv.org/abs/2610.07796
作者: Catherine Ji,Vivek Myers,Sergey Levine,Benjamin Eysenbach
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: 34 pages, 12 figures
Abstract:Empowerment captures the capacity for an agent to actively control its environment. While conceptually appealing as an information-theoretic quantity, the connection between empowerment and structurally central states that provide broad access to future outcomes has remained an open question. In this work, we link empowerment maximization and skill-learning methods to provide new geometries for interpreting and analyzing empowerment. Our analyses answer longstanding open questions on the connections between empowerment and structural centrality. Our analyses also reveal distinctions between information and reward geometries, highlighting important theoretical implications to build scalable empowerment-maximization methods. Website and code can be found at this https URL.
[AI-89] ServeLearnBench: How Well Can Agents Self-Improve from Serving Experience?
链接: https://arxiv.org/abs/2610.07792
作者: Haizhong Zheng,Yizhuo Di,Ranajoy Sadhukhan,Shuowei Jin,Beidi Chen
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: 28 pages, 13 figures. Code: this https URL . Project website: this https URL
Abstract:Large language model agents are increasingly deployed to perform complex tasks in real-world environments. However, the knowledge required for correct behavior in these environments is often implicit, undisclosed, and subject to change over time. Recent continual-learning harnesses seek to address this challenge by enabling agents to improve from serving experience. Yet the effectiveness and limitations of these methods are not yet well characterized. Existing benchmarks provide only partial coverage: some explicitly provide the target knowledge, others assume a static environment, and those that support continual adaptation remain limited in scale and knowledge diversity. To enable systematic evaluation, we formalize an evolving-environment streaming dataset (EESD), in which agents must infer, apply, and revise latent environment knowledge from interaction and outcome feedback as hidden policies evolve, and introduce ServeLearnBench, spanning retail support, banking, and sales-pitch generation with 53 environment windows and 7,718 tasks. We evaluate five learning harnesses (RAG, Mem0, SkillOpt, Continual Harness, and Prime) across six models (GPT-5.6 Terra, Opus 5, Kimi K3, GLM-5.3, DeepSeek V4.1 Flash, and GLM-5.3 Flash), covering 28 model-harness pairs and 252 learning runs. Our evaluation reveals three main findings: a substantial gap remains between task capability and learning from experience; continual adaptation is costly and can degrade already-correct behavior; and insufficient exploration emerges as a key bottleneck to effective adaptation. Overall, ServeLearnBench provides a controlled testbed for diagnosing these limitations and tracking progress toward agents that continually and reliably improve through serving experience.
[AI-90] Illusory Pattern Perception Drives Spurious Inference in Large Language Models NEURIPS2026
链接: https://arxiv.org/abs/2610.07791
作者: Peihua Mai,Zhuoyan Shao,Xinbao Qiao,Meng Zhang,Xinyue Zhou,Yan Pang
类目: Artificial Intelligence (cs.AI)
备注: accepted by NeurIPS 2026
Abstract:Illusory pattern perception is a well-documented human cognitive tendency to infer meaningful relationships in data that is actually random. Such a tendency, often described as “connecting the dots” where none exist, can result in systematic reasoning errors. This paper investigates whether Large Language Models (LLMs) exhibit such perceptual tendencies, which can lead to systematic errors in downstream applications. To our knowledge, this work presents the first systematic study of illusory pattern perception in LLMs, adapting classic psychological paradigms to three tasks with direct empirical comparison to human behaviors. We find that LLMs frequently exhibit stronger illusory pattern perception than humans. In particular, models tend to over-associate frequent positive attributes with majority groups or large organizations, and show increased tendencies to construct causal narratives from ambiguous events. To uncover the mechanism behind these behaviors, we develop a feature interpretability framework based on Sparse Autoencoders (SAEs) to analyze internal representations. Our results reveal that holistic frequency perception and analytic cognitive orientation are linked to the emergence of illusory perceptions. These findings highlight a previously underexplored cognitive-like illusion that may affect the reliability of LLM reasoning. Code available at this https URL.
[AI-91] OOPMAS: Object-Oriented Multi-Agent Systems for Query-Level Workflow Generation
链接: https://arxiv.org/abs/2610.07787
作者: Qi Cheng,Shengyu Chen,Wei Cheng,Yiqun Xie,Xiaowei Jia,Haoyu Wang,Haifeng Chen
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Multi-agent systems (MAS) powered by large language models have shown strong performance across code generation, mathematical reasoning, and question answering. However, existing methods for automating MAS design mostly operate at the task level, producing a single fixed workflow per benchmark that is applied uniformly to all queries. This assumption fails under realistic conditions. Query difficulty varies widely within a task, and real-world workloads mix heterogeneous task types. We introduce OOPMAS, a training-free framework that generates both the agent set and the coordination workflow at the granularity of individual queries. Agents are represented as object-oriented class definitions with dedicated roles, tools, and persistent state, and workflows are expressed as executable main functions over these agent objects. A dynamic skill library accumulates structured lessons from execution feedback across optimization rounds, enabling in-context improvement without any gradient updates or fine-tuning. On a mixed-task benchmark of queries spanning code, math, and QA, OOPMAS achieves 89.6% accuracy, outperforming the strongest baseline by 18.1 percentage points. A model-swap study across four LLM backbones shows consistent scaling, reaching 92.4% with the strongest model.
[AI-92] Attacca: Goal-Directed Control under State Continuity for Long-Horizon Embodied Agents
链接: https://arxiv.org/abs/2610.07785
作者: Gyusik Seo,Jaehong Yoon
类目: Artificial Intelligence (cs.AI)
备注: Project page: this https URL
Abstract:A central capability of embodied agents is to accomplish complex objectives through sequences of interdependent tasks. Yet existing visual goal-conditioned policies underlying these agents are typically evaluated on isolated interactions where the target is already visible, and thus do not capture the conditions that arise during continuous long-horizon task execution. In such settings, each task begins from the state left by the previous one: the agent may end at a different position and orientation, the world may have been modified, and the next interaction target may lie outside the current field of view. As a result, agents relying on such policies may struggle to proceed to the next task when they cannot ground their target in the current observation. To address this challenge, we propose Attacca, a new approach that trains visual goal-conditioned policies on complete search-to-interact trajectories using goal images decoupled from the execution environment. Attacca uses context-decoupled goal sampling to pair each demonstration with a class-compatible masked goal image from another world, removing direct scene and pose correspondence. It learns dense current-view grounding through a target-mask prediction head, providing auxiliary supervision beyond action imitation. We further introduce behavioral-phase conditioning that teaches the policy to distinguish Search, Approach, and Interact stages and adapt its control as execution progresses. We evaluate Attacca on multiple short- and long-horizon embodied tasks in Minecraft. Our method achieves 39.0-47.5% clean success, improving over the strongest baseline by 1.7-2.4x. On long-horizon tasks, it attains 54%, 30%, and 28% completion, yielding up to a 7x improvement.
[AI-93] owards One-for-All Foundation Model for Attributed Graph Clustering
链接: https://arxiv.org/abs/2610.07778
作者: Yunhui Liu,Xudong Jin,Kang Zhang,Danshuo An,Yu Xing,Te Song,Jia Liu,Tieke He
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:
Abstract:Attributed graph clustering aims to discover node groups by jointly exploiting node attributes and graph topology, yet its unsupervised nature makes model selection and adaptation inherently difficult. Existing methods typically train and tune a separate model for each input graph, leading to costly and fragile pipelines that often fail to transfer across graphs with different feature spaces, structural patterns, and attribute-structure correlations. In this paper, we study a one-for-all alternative: can a single model be trained once and directly applied to diverse attributed graphs without graph-specific training, fine-tuning, or hyperparameter search? We propose OFAG, a foundation model for attributed graph clustering. Building upon Prior-data Fitted Networks, OFAG learns a reusable clustering inference strategy from synthetic attributed graphs generated under broad priors over latent clusters, node attributes, and graph structures. To handle incompatible feature spaces across graphs, OFAG adopts a dimension-agnostic signal-wise graph encoder that treats each feature channel as a graph signal and models its response to shared graph filters. The model is trained with a hyperspherical clustering objective, producing clustering-friendly node representations in a single forward pass at inference time. On ten datasets, one frozen OFAG model achieves the best mean performance and average rank across NMI, ACC, ARI, and F1, while completing all ten datasets in 12.43 minutes total—over 6* faster than the second-fastest baseline and nearly 28* faster than the second-best on clustering quality. Our code and pretrained checkpoint are available at this https URL, allowing practitioners to directly apply OFAG to their own attributed graph datasets without additional training or tuning.
[AI-94] OTel: Open Telco AI Datasets Benchmarks and Models NEURIPS2026
链接: https://arxiv.org/abs/2610.07766
作者: Farbod Tavakkoli,Gregory Diamos,Kenneth Church,David Kanter,Mark Austin,Imtiaz Karim,Mirza Masfiqur Rahman,Merouane Abdelkader Debbah,Zeinab Nezami,Ali Maatouk,Leandros Tassiulas,Rex Ying,Nick Sorros,Louis Powell,Nikolaos Vasiloglou,Ashish Vaswani,Somanshu Singla,Adarsh Chaluvaraju
类目: Artificial Intelligence (cs.AI)
备注: Accepted to NeurIPS 2026, ED Track, Spotlight
Abstract:We present Open Telco (OTel), an open telecom AI resource that releases derived telecom datasets for retrieval, reranking, instruction tuning, and safety/abstention, together with 30 full-parameter post-trained baselines spanning 10 embedding models, 3 rerankers, and 17 language models. The community has already engaged substantially with the resource: as of May 3, 2026, the released models have been downloaded over 16 million times and the project has received 157+ pieces of media coverage worldwide. Building on prior open telecom datasets and benchmarks, OTel provides documented telecom data sources, held-out evaluation partitions, trained embedding models, rerankers, context-grounded LLMs, and safety/abstention data in one unified resource. Each baseline starts from an open-weight model and is post-trained on OTel-derived data using an open training recipe, then evaluated on held-out OTel evaluation partitions. OTel post-training improves performance across all three model families: embedding retrieval reaches 93.1% NDCG@10, reranking reaches 0.947 MRR@10, and language-model correctness reaches 87.8%. We release OTel as a reproducible starting point and invite the community to expand the data, improve embedding and reranking models, and build stronger context-grounded telecom LLMs.
[AI-95] ST-Bench: A Spatial-Temporal Benchmark for Multi-Agent System Generation on Scientific Research Tasks
链接: https://arxiv.org/abs/2610.07763
作者: Qi Cheng,Rongchao Dong,Shengyu Chen,Licheng Liu,Dan Lu,Zhengzhang Chen,Wei Cheng,Yiqun Xie,Haifeng Chen,Xiaowei Jia,Haoyu Wang
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:The rapid progress of LLM-based multi-agent systems (MAS) has shown that they largely outperform single agents on coding, math, and QA tasks, where executable tests provide a binary success signal. Whether this advantage transfers to real scientific data analysis remains untested. We introduce ST-Bench, a benchmark designed to answer two questions: whether MAS outperform single agents on complex scientific data analysis tasks, and if so, by how much and at what additional cost. ST-Bench contains 100 data science tasks adapted from published Earth science studies across hydrology, agriculture, and wetland methane research, expanded into 2,067 queries grounded in additional published studies and validated by domain experts. Using ST-Bench, we evaluate five recent MAS generation methods under two training protocols, against single-agent baselines on the same GPT-5 backbone. Nine of the ten MAS configurations exceed the cheapest single-agent baseline, with the strongest reaching nearly three times its composite score. This gain is primarily attributable to coverage: trained workflows produce realistic numerical metrics on a larger fraction of queries, while the quality of those metrics, conditional on producing realistic output, is comparable to that of the single-agent baseline. The strongest configuration requires approximately four times the single-agent inference time, whereas a more economical workflow captures the majority of the benefit at less than twice the cost. MAS specialization confers measurable benefit on scientific data analysis, but the benefit is conditional rather than universal.
[AI-96] How Well Do LLM s Reason with Noisy Evidence? An Active Visual Reasoning Benchmark
链接: https://arxiv.org/abs/2610.07751
作者: Bach Nguyen,Zhaonan Li,Mau Son Nguyen,Sanika Chavan,Nilay Kumar,Hong Anh Nguyen,Khoa Vo,Ben Zhou
类目: Artificial Intelligence (cs.AI)
备注: 27 pages, 9 figures, 11 tables
Abstract:Real-world reasoning rarely reduces to static question answering: agents must actively gather information from tools and sensors that are often noisy and unreliable. Yet most existing active reasoning benchmarks assume that environmental feedback is trustworthy, or introduce noise without exposing an explicit, calibrated uncertainty signal, leaving open how LLMs should reason when the evidence itself is uncertain. We introduce VisualNoiseQA, a novel benchmark for active reasoning under noisy visual feedback. A text-only LLM must solve VQA problems by iteratively querying a fixed, off-the-shelf VLM treated as a stochastic visual sensor. For each query, we draw multiple samples and expose an empirical uncertainty signal via self-consistency, enabling the reasoner to probe from different angles and decide what to ask next and when to stop. Our construction is automatic and scalable: starting from diverse VQA sources and two noisy VLMs, we retain only questions where the sensor is inconsistent yet human-solvable. We evaluate multiple LLM reasoners on 1,000 instances spanning perception, chart understanding, and knowledge-intensive reasoning. VisualNoiseQA thus provides a controlled playground to study how different LLMs exploit uncertainty signals for robust reasoning.
[AI-97] Cleave: Scaling Tensor Program Optimization via Decoupled Algebraic Search and Operator Scheduling
链接: https://arxiv.org/abs/2610.07742
作者: David Pissarra,Jinkun Lin,Haitian Jiang,Aurojit Panda,Jinyang Li
类目: Programming Languages (cs.PL); Artificial Intelligence (cs.AI)
备注:
Abstract:Optimized kernels such as FlashAttention and FlashDecoding are crucial for accelerating today’s large models. Most of them are handwritten by experts because existing ML compilers cannot match their efficiency. Producing such kernels requires fusing computations with multiple reductions, which requires both algebraic transformation of the computation graph and operator scheduling of the transformed graph. Unfortunately, searching the two jointly yields a space too large to navigate. We propose Cleave, an ML compiler built on symbolic decoupling: Cleave discovers transformations by performing superoptimization on a graph with symbolic shapes, and then schedules each resulting graph on concrete shapes. Representing shapes as symbols makes equivalence checking cheap and lets a new Split operator, with a symbolic split count, parallelize along a reduction dimension. Cleave’s scheduler fuses graphs with multiple reductions through iterative tiling and horizontal fusion. Evaluation on common LLM subgraphs shows that Cleave generates kernels up to 2.8x faster than the best baseline (1.6x on average) and reduces compilation time by 5.9x on average compared to Mirage. For dynamic workloads captured from production serving traces, Cleave compiles each operator once and achieves geometric mean speedups of 1.4x and 1.7x over FlashInfer’s handwritten FA2 and FA3 backends. Cleave’s code is available at: this https URL
[AI-98] PERSIST: Who-What-When Memory Across Sessions for Full-Duplex Spoken Dialogue
链接: https://arxiv.org/abs/2610.07725
作者: Achira Lin,Siyuan Hou,Wenyi Yu,Xinnian Zhao,Haoyu Niu,Wang Geng,Longshuai Xiao,Shihai Xiao,Mangsuo Zhao,Chao Zhang
类目: Artificial Intelligence (cs.AI)
备注: 19 pages, 3 figures
Abstract:Modern voice assistants may be shared by multiple users and should be able to answer questions about earlier conversations such as “When did I originally plan to leave?” or adapt their behavior to individual users based on past interactions. This requires more than retrieving a topically similar passage: the assistant must identify the current speaker, recover the relevant past state, and distinguish it from later revisions. We present PERSIST, a persistent memory system for multi-session, multi-speaker spoken dialogue that explicitly models Who, What, and When. PERSIST structures cross-session histories into readable event records and retrieves them with a 3W joint scoring mechanism that combines semantic content, acoustic speaker identity, and temporal state. For real-time full-duplex interaction, PERSIST further reuses intermediate representations from the dialogue backbone, avoiding query-audio re-encoding and reducing retrieval latency from 578.42 ms to 7.03 ms. We also introduce SpokenTrace, a diagnostic benchmark that factorizes evaluation along memory tasks and speaker-query types, exposing failures in recall, speaker attribution, and temporal-state tracking. On SpokenTrace, PERSIST achieves 85.08% end-to-end task accuracy and improves all-support EM@3 from 49.01% with BGE-large to 82.10%.
[AI-99] Evidence Before Sampling: Interpretable Implicit Negative Candidate Discovery for Recommendation
链接: https://arxiv.org/abs/2610.07708
作者: Shreya Rajpal,Sonia Sharma,Swapnil Parekh,Lisa Li,Jeyendran Balakrishnan,Nagaraj Janardhana,Andrew Mattarella-Micke
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Recommender systems learn from observed user-item interactions, but explicit negative feedback is often unavailable. Since deep learning models require negative signals for training, negative sampling methods typically treat selected unobserved interactions as negatives. However, a missing interaction does not explain why a user is uninterested in an item or whether there is sufficient evidence to label it negative. This is especially important in business recommendation, where negative signals should be interpretable and aligned with business objectives. We formulate implicit negative candidate discovery to identify unobserved interactions supported by observed customer behavior. We encode these patterns as symbolic rules, score them based on support, informativeness, and product relevance, and rank the retained rules by evidence. An LLM then interprets the retained rules using business objectives and domain knowledge; the interpretations are combined with the statistical evidence in the final report. We evaluate our method in an industrial B2B setting and across five public recommendation datasets. Candidate-quality evaluations in the industrial setting and three public datasets show higher precision than the evaluated baselines, while symbolic selection improves downstream test PR-AUC by 12.5% over random selection with four negatives per positive example in the industrial task. Our results show that negative candidate validity can be evaluated separately from downstream recommendation performance. This distinction enables evidence-based, business-aligned, and explainable negative selection, improving both interpretability and model training in sparse, skewed, real-world recommendation settings.
[AI-100] Agent MemGate: Addressing Speculation Contamination in Conversational Assistant Memory NEURIPS2026
链接: https://arxiv.org/abs/2610.07707
作者: Chirag Sharma,Benjamin Fowlersmith,Karime Maamari
类目: Artificial Intelligence (cs.AI)
备注: Accepted at the PALM Workshop at NeurIPS 2026. 15 pages, 4 figures
Abstract:Conversational AI assistants with long-term memory extract facts from user messages into a store consulted in later conversations. A stated plan can enter that store as fact: a user who might move to Seattle may be recorded as already living there. We call this speculation contamination. Final-state memory benchmarks miss this error because they do not probe intermediate state and include few unresolved speculations. We present AgentMemGate, a write-time gate for profile-store memory that classifies extracted statements as speculation, completed event, correction, or other. Speculations remain outside memory, with conditions governing later promotion or deletion. We also contribute a dataset of multi-session conversations in which plans are confirmed, abandoned, or left unresolved. On our 147-conversation held-out set, Mem0 and Graphiti assert unresolved plans as current state for 35.2% and 27.3% of pending plans. On the core benchmark, AgentMemGate eliminates all observed contamination relative to the identical ungated pipeline (87.5% to zero for the most exposed extraction style) and raises task accuracy from 65% to 95%. On the harder held-out set, gated contamination is 3.4% to 5.7% and task accuracy rises by 9 to 13 percentage points. Our analysis identifies field matching as the main remaining bottleneck: realistic speculations often match no profile field and never reach the gate. We release our datasets, prompts, and evaluation code.
[AI-101] WASD: Wasserstein-based Knowledge Distillation for Large Language Models NEURIPS2026
链接: https://arxiv.org/abs/2610.07706
作者: Byeonghu Na,Donghyeok Shin,Yeongmin Kim,Mina Kang,Il-Chul Moon
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: Accepted at NeurIPS 2026
Abstract:Autoregressive large language models (LLMs) have rapidly advanced in capability, but their increasing scale comes with substantial computational and memory costs at inference time. Knowledge distillation (KD) offers a practical solution by transferring knowledge from a large teacher model to a smaller student model via alignment of discrete probability distributions. However, existing KD methods for LLMs primarily rely on divergences that evaluate discrepancies through probability values at each vocabulary index, without explicitly leveraging token-level semantic information. We propose Wasserstein-based knowledge distillation (WASD) for LLMs, which incorporates token-level semantic information via the Wasserstein-based distance with a cost matrix derived from token embeddings. To ensure computational tractability, we adopt the Sinkhorn divergence and derive a gradient-equivalent objective that can be efficiently optimized without introducing additional networks. Experiments across multiple LLM families and scales show that WASD consistently improves distillation performance on diverse tasks, including instruction following, mathematical reasoning, and code generation. Our results highlight the importance of semantic information encoded in the token space for effective distribution alignment in LLM distillation. The implementation is publicly available at this https URL .
[AI-102] On the Boundary of Admission Gates: An Injected-Truth Study of Falsification-First Selection in Quantitative Strategy Research
链接: https://arxiv.org/abs/2610.07701
作者: Tianlun Zheng
类目: Artificial Intelligence (cs.AI); Computational Engineering, Finance, and Science (cs.CE)
备注: 12 pages, 3 figures, 7 tables. Code and data to reproduce every result: this https URL
Abstract:Strategy research conflates two problems: finding a profitable rule, and establishing that the finding is not search luck. The latter calls for admission gates – statistical criteria that must be satisfied before a conclusion is adopted – yet whether gates work, and at what cost, remains untested. We introduce an injected-truth protocol with a random-admission control that adopts at the same rate as the gate; only if the gate beats this control does it carry information rather than merely raise a threshold. Across synthetic and real-calibrated panels, gates eliminate false discoveries in the weak-signal regime but cut adoption to 1–7%, and add nothing when signals are strong. Most importantly, criteria computed on absolute rather than excess returns silently reject every candidate, including true signals. Keywords: multiple testing, backtest overfitting, strategy admission, injected-truth validation, excess returns, false discovery rate
[AI-103] BluffJAX: Adversarial Imperfect Information Games in JAX
链接: https://arxiv.org/abs/2610.07686
作者: Aryaman Reddi,Jan Peters,Carlo D’Eramo
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:We introduce BluffJAX: an open-source suite of adversarial imperfect information games in JAX. We provide canonical implementations of games designed for high simulation throughputs and parallelization on GPU accelerators. Our suite consists of well-studied benchmarks such as Texas Hold’Em Poker and Kuhn Poker, as well as games that have not been previously studied in reinforcement learning research, such as Bluff, Stud Poker, and Kemps. We hope that implementing a variety of game mechanics and difficulties will introduce new challenges and foster novel research directions in game-theoretic methods for RL. We benchmark the throughput performance and memory usage of our environments in single and multi-GPU settings, demonstrating scaling of up to hundreds of millions of samples per second, and motivating the usage of BluffJAX over related GPU and CPU-based libraries. We benchmark reinforcement learning, tree search, and game-solving algorithms in JAX in order to provide users with baseline results and facilitate future comparisons.
[AI-104] EigenDEXplore: Structured Exploration for Dexterous Manipulation with Human Priors
链接: https://arxiv.org/abs/2610.07681
作者: Harsh Gupta,Tyler Ga Wei Lum,Changhao Wang,Chuer Pan,C. Karen Liu,Jeannette Bohg,Shuran Song
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI)
备注: 15 pages, 12 figures. Project page: this https URL
Abstract:Dexterous manipulation poses a challenging high-dimensional optimization problem, as useful behaviors require coordinated motion across many hand joints. In reinforcement learning (RL) and sampling-based trajectory optimization, exploration commonly relies on independent robot joint perturbations, making coordinated behaviors difficult to discover. Prior work reduces this search space for grasp learning using low-dimensional spaces of coordinated joint motions learned from human hand data, but this restricts the expressivity required for general manipulation. Some combine learned and joint-space actions to restore expressivity, but this increases dimensionality and introduces redundancy. We study these effects across diverse manipulation settings, varying action dimensionality, exploration strategy, and the source of human data. Our experiments suggest that human-motion priors are most effective when used to structure exploration rather than change the action representation. Motivated by this finding, we propose EigenDEXplore, which induces correlated exploration by adding perturbations along human-derived eigenvectors to independent joint-space noise, leaving the action space unchanged. Across multiple dexterous hands, EigenDEXplore consistently outperforms joint-space and learned action-space baselines in grasping, in-hand reorientation, and contact-rich manipulation. These gains span unstructured and reference-guided RL, trajectory optimization, and sim-to-real deployment, and are largest in settings with less reward shaping and curriculum design.
[AI-105] Exact-Solution Volume and Length Generalization in Transformers
链接: https://arxiv.org/abs/2610.07676
作者: Yijia Jessica Zhu,David Chiang
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: 26 pages, 2 figures
Abstract:Research on transformer expressivity shows whether a transformer is capable of solving a given task, but gives little indication of whether the solution, if learned, is generalizable to longer input lengths. We study this question through normalized exact-solution volume (NESV): the fraction of a bounded parameter region that achieves an exact solution on every input of length n . For fixed-width, single-layer transformers with \log n -scaled attention, we establish asymptotic bounds on NESV for four tasks: FIRST ( \Theta(1) ), MAJORITY ( \Theta(1/(n\log n)) ), INDEX ( \Theta(1/n^3) ), and PARITY ( 0 ). These results are consistent with previous empirical results: the faster the exact-solution volume decays with input length, the harder it is to length-generalize on that task. Looking deeper into INDEX, our volume analysis reveals two error sources that grow with n . Consequently, we study a transformer model that would structurally eliminate one of the terms, theoretically improving the NESV bound to \Theta(n^-1) , and empirically achieving 85% accuracy when tested at 10\times the training length, compared with the 60% accuracy of the original model. We conclude that volume analysis may be a useful approach to identify concrete sources of length sensitivity and thus provide insights into task-specific model refinements.
[AI-106] CACHEFORGE: LLM -Guided End-to-End Generative Cache Replacement Policy for Performance and Hardware Efficiency
链接: https://arxiv.org/abs/2610.07668
作者: Kaushal Mhapsekar,Bita Aslrousta,Brijesh Kumar Bhayana,Paula Contreras,Azam Ghanbari,Ethan Goodman,Anna Andriiko,Samira Mirbagher Ajorpaz
类目: Hardware Architecture (cs.AR); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:
Abstract:Modern cache replacement designs saturate because they operate within fixed representational structures, hand-crafted and heuristic based feature-engineered predictors, or offline imitation models that cannot generate new decision logic on their own. At the same time, replacement is shaped by the causal interaction of prefetching, thrashing, spatial locality, and access-type behavior, producing an enormous design space that is difficult to traverse manually. Prior approaches typically rely on heuristics, parameter tuning, or imitation of an offline optimal policy, capturing correlations rather than synthesizing new mechanisms. As a result, their performance gains often plateau and they overfit under dynamic workload conditions. CACHEFORGE is the first framework to evolve cache-replacement policies end-to-end by embedding a large language model inside a governed hardware-aware loop. In each iteration, the LLM proposes new C++ replacement logic, the policy is evaluated under a trace-based CRC-2 ChampSim simulator, and the framework enforces feasibility through reward shaping, structural checks, dynamic mutation, temperature scheduling, and cross-policy crossover. This closed-loop generation-evolution loop specifically designed for cache replacement policy enables the discovery of compact policies that satisfy hardware constraints while exploring algorithmic transformations beyond fixed predictor structures. Across SPEC CPU2006, CACHEFORGE outperforms all CRC-2 baselines. It improves the total hit rate by 27.36%, 19.69%, 13.72%, 13.15%, 11.83%, and 5.73% over MPPPB, ReD, Hawk-eye, SHiP++, LIME, and LRU, respectively. On memory-intensive workloads, it increases IPC by 10.15%, 7.89%, 6.34%, 3.64%, 3.12%, and 2.71% over LRU, MPPPB, LIME, ReD, SHiP++, and Hawkeye. Subjects: Hardware Architecture (cs.AR); Artificial Intelligence (cs.AI); Machine Learning (cs.LG) Cite as: arXiv:2610.07668 [cs.AR] (or arXiv:2610.07668v1 [cs.AR] for this version) https://doi.org/10.48550/arXiv.2610.07668 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[AI-107] Massive Activation Gating Channel in Large Language Models
链接: https://arxiv.org/abs/2610.07661
作者: Minjia Mao,Shi Chen,Bowen Yin,Xiao Fang
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Massive activations, a phenomenon in which a small number of hidden channels exhibit exceptionally large magnitudes, are pervasive in large language models (LLMs). However, the mechanism by which a token develops massive activations as it propagates through a pretrained LLM remains poorly understood. In this paper, we find that the emergence of massive activations is controlled by a single channel in the input embedding to a spike feed-forward network (FFN). The position of this channel is fixed for a particular LLM. We name this channel the massive activation gating channel (MAGC). When the value of the MAGC is sufficiently large (or small, depending on the LLM), the output of the spike FFN exhibits massive activations. Examining six LLMs across four model families and different model sizes, we verify the existence and effect of MAGC. We further provide a theoretical explanation of the mechanism by which MAGC induces massive activations. When the value of MAGC is sufficiently large (or small), the output of a spike FFN asymptotically reduces to a quadratic form that mixes a few columns of the down-projection matrix of the FFN. Since these columns exhibit the shape of massive activations, the output therefore exhibits massive activations.
[AI-108] Does On-Policy Distillation for Safety Pose Backdoor Risks?
链接: https://arxiv.org/abs/2610.07654
作者: Jian Luo,Kehan Qi,Qingqiao Hu,Meilong Xu,Jiacheng Qiu,Weimin Lyu,Jiawei Zhou,Chao Chen
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:
Abstract:On-policy distillation (OPD) has attracted growing attention as an effective way to transfer capabilities from teacher models to student models. Recent studies further explore OPD as a tool for improving large language model safety with promising results. However, these approaches typically assume that the teacher and training data are trustworthy. In this paper, we uncover an overlooked threat to OPD for safety: a safety-aligned but backdoored teacher can propagate its hidden malicious behavior to an initially clean student. Under our threat model, a poisoning rate as low as 3% results in an attack success rate (ASR) of up to 70% on the distilled student. We further identify two training choices that can amplify this risk. First, increasing the number of training epochs can lead to high ASR even at low poisoning rates. With only 10 poisoned samples, ASR reaches 67% after 16 epochs. Second, the commonly used top-k KL can accelerate backdoor transfer, causing trigger-conditioned harmful behavior to emerge earlier than sampled-token KL in most settings. Alongside these findings, we explore a simple mitigation, Lazy Defense, which clips KL rewards to make student updates less aggressive, limiting aggressive updates and slowing backdoor learning. Experiments show that Lazy Defense delays backdoor transfer in low poisoning rate settings. Together, our findings reveal that OPD can propagate backdoors, highlighting the need to address the safety risks of OPD.
[AI-109] SMART: Zero-Shot Sim-to-Real Articulated Object Manipulation via Large-Scale Synthetic Pretraining
链接: https://arxiv.org/abs/2610.07652
作者: Jicong Ao,Shuhan Jiang,Yuling Zhong,Yanwen Liu,Yuhan Gao,Jiangyuan Zhao,Yang Zhang,Shiqiang Zhu,Chenjia Bai,Xuelong Li
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI)
备注: Technical Report, 31 pages
Abstract:The ability to interact with articulated objects is essential for embodied intelligent systems, but collecting large-scale real-world demonstrations for these interactions remains challenging due to the precise contact and constraint-following motions involved. Although simulation provides a promising alternative, existing synthetic data efforts cover limited articulated-object categories, while general-purpose synthesis pipelines lack explicit designs for part-level semantics and articulation constraints, hindering agentic task generation and scalable synthesis of high-quality articulated-manipulation demonstrations. To bridge this gap, we introduce SMART, a scalable system leveraging large-scale Synthesized Manipulation demonstrations for ARTiculated-object manipulation. At its core, we develop SMART-Sim, a simulation platform with articulation-aware design that enables effective task generation and efficient demonstration collection. Building on SMART-Sim, we apply agentic task generation and design a scalable distributed synthesis system, using them to synthesize SMART-Data, comprising over 1M demonstrations across 44 atomic task types, 5 robot setups, and 2,507 articulated objects. The vision-language-action (VLA) model pretrained on SMART-Data shows competitive performance on simulation benchmarks and achieves zero-shot sim-to-real transfer and scalable performance in real-world articulated-object manipulation tasks. This highlights the potential of synthetic demonstrations in providing effective and scalable supervision for improving VLA model performance in contact-rich articulated-object manipulation.
[AI-110] Matching Object or Relation? Tracing Abstract Reasoning Inside VLMs
链接: https://arxiv.org/abs/2610.07646
作者: Gouki Minegishi,Hiroki Furuta,Takeshi Kojima,Yusuke Iwasawa,Yutaka Matsuo
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Vision Language Models (VLMs) excel on visual benchmarks but fail systematically on tasks requiring abstract reasoning. Existing benchmarks document this failure but cannot say \emphwhy it happens or which cognitive capability is missing. We close this gap by adopting the Relational Match-to-Sample (RMTS) paradigm from comparative and developmental psychology and pairing it with a mechanistic analysis of the model’s internals. On a parametrically controlled stimulus set evaluated across frontier API models (GPT, Claude, Gemini) and three open-source families (Qwen3.5, Gemma-4, InternVL3), we identify four levers that shift VLMs toward the relational match—capability tier, model scale, the number of objects per scene, and the absence of per-object stimulus noise—together producing a developmental-like trajectory that mirrors the human \emphrelational shift. Opening up the model, a per-layer representational similarity analysis and a causal mediation analysis reveal that VLM abstract reasoning is implemented by two competing circuits: an early circuit that organises images by their surface object features, and a late circuit that organises them by their abstract relation. Extending the analysis to ARC-AGI-1, we find that ablating the relational heads identified on RMTS degrades performance more than ablating random heads, indicating that the relational circuit is recruited beyond our controlled stimuli. We hope this mechanism-level view serves as a step toward understanding how abstract reasoning is implemented in VLMs.
[AI-111] SkillPoison: Progressive Skill Poisoning via Successful Experiences
链接: https://arxiv.org/abs/2610.07645
作者: Lizhi Zhang,Xin He,Dianxuan Fu,Yuyuan Feng,Jiatong Li,Qi Wang,Xin Wang,Qinggang Zhang
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
备注:
Abstract:Self-improving LLM agents increasingly distill successful experiences into persistent, reusable skills. Existing skill attack methods corrupt this learning pipeline by injecting malicious triggers, behaviors, or false facts into individual experiences or extracted skills. However, such attacks are easily detected, and the injected malicious behaviors often fail to accumulate as persistent skills. In this paper, we show that skill poisoning can arise even from verified successful experiences, without making any individual trajectory malicious. Based on this insight, we propose SkillPoison, a novel framework that progressively poisons skill via successful experiences. SkillPoison first constructs a set of successful experiences that reinforce a target behavior, and then removes the contextual conditions that constrain when the behavior applies. Rather than injecting malicious content, SkillPoison shapes how the skill extractor generalizes, allowing useful behavior to support task success while inducing harmful behavior when they are misapplied. Extensive experiments on three benchmarks show that SkillPoison achieves 95.71% attack success rates, while all injected experiences remain task-correct and pass verification and lexical inspection. Our code, data and implementation details are available for the community at this https URL.
[AI-112] owards the Automatic Synthesis of Interpretable Chess Tactics
链接: https://arxiv.org/abs/2610.07640
作者: Abhijeet Krishnan,Chris Martens
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Symbolic Computation (cs.SC)
备注:
Abstract:State-of-the-art reinforcement learning agents are capable of outperforming human experts at games like chess, Go and StarCraft II. These agents do not simply take advantage of their digital hardware in being able to react and calculate faster than humans, but employ better strategies that lead to more victories. Interpreting these strategies would give human players valuable insight into how to improve their play. In this preliminary work, we propose a symbolic sub-policy model for playing chess. Inspired by chess tactics, our model attempts to incorporate domain knowledge to improve interpretability. We adapt patterns learned by an inductive logic programming system called PAL to derive our model. We contribute a divergence metric to evaluate our model against a random baseline, and find a set of tactics that is able to suggest moves of similar playing strength to a human beginner. Finally, we propose a computational evaluation scheme for the model by augmenting an off-the-shelf engine with it.
[AI-113] Learning Explainable Representations of Complex Game-playing Strategies
链接: https://arxiv.org/abs/2610.07638
作者: Abhijeet Krishnan,Colin M. Potts,Arnav Jhala,Harshad Khadilkar,Shirish Karande,Chris Martens
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:
Abstract:As part of learning to play complex games, human players develop develop abstractions for concepts and strategies of gameplay consistent with game rules to improve their performance. These concepts are applied to explain other players’ actions, and to inform their own actions in-game. Understanding other players’ strategies is a crucial part of such improvement, but requires time and effort. In this paper, we propose a strategy similar to human cognition for training RL agents to synthesize learned strategies and policies as executable procedures based on sequences of gameplay actions. We present methods to automatically learn such programs to play chess and to solve tasks in a grid-based environment. We show that the learned strategies produce effective actions, and can be learned from gameplay data.
[AI-114] Measuring climate backlash in Twitter and Reddit archives: Lexical definitions recorded responses and participant turnover
链接: https://arxiv.org/abs/2610.07634
作者: Wentao Xu
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Social media archives are often used to study resistance to climate action, but words, response counters and observed participants do not measure the same social process. We examine four supplied Twitter and Reddit archives by processing all registered files without sampling and applying transparent, non-exclusive lexical rules. The study links frame co-occurrence to source-specific temporal and response models, then separates event-period changes among returning authors from participant turnover. Renewable-energy terms accompany cost-related language on Reddit, yet narrower backlash phrases sharply reduce cross-source contrasts and reverse the sign of the Paris Agreement contrast in submissions. Cross-discourse history does not improve eligible primary-context forecasts. Denial/hoax terms are associated with higher recorded Twitter likes, whereas Reddit response associations depend on frame, outcome and author specification. Around the 2019 global climate strike, returning-author expression and participant turnover both contribute to increased protest-language shares. An archive endpoint prevents the corresponding Climate Twitter migration inference. Most crossed-cluster estimates lack released intervals, and joint author/month response covariance estimates fail, restricting formal inference. These results show how operational definitions, platform-specific response fields and observation boundaries shape what can be claimed about climate backlash. The contribution is an archive-based account of these measurement consequences, rather than a measure of individual opposition, persuasion or advocacy-induced backlash.
[AI-115] Learning to Outgrow a Theory: Experimental Discovery Beyond the Initial Hypothesis Space
链接: https://arxiv.org/abs/2610.07627
作者: SiYuan Ma,Albert Gao,Chunzheng Zhu,Xin Yan,Wenlong Zhang,Wenxin Zhang,Luqi Gong,Tianlin Li,Qixin Zhang
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Scientific discovery systems typically optimize experiments within a fixed hypothesis space. This creates a failure mode when all available candidates omit the same missing mechanism: candidate disagreement can collapse even while the model class is systematically wrong. We formulate experimental model-class revision, in which a discovery policy jointly proposes a structural edit and a diagnostic experiment that tests whether that edit is necessary. The method couples a class-level distinguishability objective, in which one shared parameterization must explain all selected experiments, with anytime-valid sequential evidence that triggers structural revision only after the current class is rejected. On 400 held-out controlled dynamical environments, the joint policy reaches 89.5% exact recovery with a budget of 32 real experiments, improving the strongest matched baseline by 10.0 percentage points while requiring fewer executed experiments and candidate fits. The learned revision-experiment pairing transfers across unseen mechanism combinations, held-out but expressible primitives, parameter extrapolation, and shifted experiment costs; when the true mechanism is outside the edit grammar, it detects library insufficiency in 88% of cases with a 5.5% false-support rate. Revision gains also transfer to ODEBench and ODEBase model-library tasks, as well as DiscoverPhysics worlds. These results support a view of scientific discovery in which deciding what mechanisms a theory should make expressible and where to collect evidence are treated as a single sequential decision problem.
[AI-116] Explore Then Commit: Measurement-Efficient Scientific Law Discovery with Language Models
链接: https://arxiv.org/abs/2610.07620
作者: Kautik Mandve,Dileepa Fernando
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Computational Physics (physics.comp-ph)
备注: 19 pages, including Supplementary Material S1; code and data included as ancillary files. Preprint
Abstract:Scientific law discovery requires selecting measurements and converting evidence into a governing equation. We evaluate an explore-then-commit protocol in which a large language model proposes hypotheses, a programmatic planner gathers measurements, and a fresh prompt synthesizes the final law from fixed observations. The protocol combines structured probes, automatic numerical diagnostics, restricted measurement batches, and optional interpreter access. Across 576 NewtonBench trials, we compare eight configurations on 12 physics modules using GPT-4.1-mini and a medium-difficulty GPT-4.1 replication. On medium tasks, interpreter-enabled planners use 8.6 versus 22.5 measurements per trial for GPT-4.1-mini and 8.9 versus 43.0 for GPT-4.1. Their mean magnitude-based root-mean-squared logarithmic error falls from 2.514 to 0.202 and from 0.626 to 0.149, respectively. An additional audit retains incomplete and invalid submissions in a coverage-sensitive analysis. Observed symbolic-accuracy gains are less consistent across modules, and random acquisition is competitive with disagreement scoring. Measurement savings occur in every module, but unequal batch constraints prevent attributing them solely to acquisition quality. These results support the complete protocol as a promising measurement-efficient configuration, while leaving its causal components and generalization beyond noiseless direct-equation tasks unresolved.
[AI-117] BioStudyBench: Evaluating Agents on Post-Cutoff Biomedical Studies NEURIPS2026
链接: https://arxiv.org/abs/2610.07614
作者: David Li,Shaamil Karim,Christian Gensbigler
类目: Artificial Intelligence (cs.AI)
备注: Accepted into AgenticLS (NeurIPS 2026 workshop)
Abstract:We evaluate whether AI agents can match the reported findings of published biomedical studies using public data. Existing evaluations do not consistently separate analysis from prior knowledge or retrieval of the published answer. We introduce BioStudyBench, a benchmark of 25 long-horizon analysis tasks drawn from studies first published between July and September 2026, after the developer-reported knowledge cutoffs of the models we evaluate, semi-automatically filtered down from 404,019 PubMed records. In each task, the agent receives a neutral research question but no data files, so it must find and download the relevant public data, search the literature through tools that return only records dated before its cutoff, and report findings through data analysis. To measure gains over prior knowledge, we run every task both with and without access to data and tools. Across eight models, access to data and tools raises the pass rate by 47 percentage points on average over the no-data baseline. Open-weight models across sizes trail closed-weight models, with the best open-weight model passing 81.3% of tasks against 94.7% for the best closed-weight model.
[AI-118] VALSE: Vertical Adaptive Layer Skipping for Efficient Inference in Large Language Models
链接: https://arxiv.org/abs/2610.07606
作者: Jia-Dong Zhang
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:This paper establishes a theoretical framework for vertical adaptive layer skipping, proving three foundational results: (i) an Expected FLOPs formula (theorem 2) giving a closed-form expression for the computational cost of arbitrary per-sample skip schedules as a function of layer-wise skip probabilities; (ii) function-space superset (theorem 10) and strict inclusion (theorem 11) theorems showing that skip-layer models are strictly contained in—yet meaningfully approximate—the full-layer function space, with an explicit separating example; and (iii) a structural duality between VALSE and Mixture-of-Experts architectures (proposition 6), positioning vertical depth-wise sparsity as the orthogonal counterpart to horizontal width-wise sparsity. Building on this theory, we propose VALSE (Vertical Adaptive Layer Skipping for Efficiency), a per-sample, non-contiguous layer skipping method: a lightweight difficulty estimator scores each input from the first few layers, and per-layer gates selectively skip redundant layers—including arbitrary middle layers while retaining deeper ones—so that only the necessary depth is activated for each input, whose feasibility is preliminarily assessed at prototype scale.
[AI-119] Beyond Scalar IoU: Structured Verification from Rollout Groups for Video Temporal Grounding
链接: https://arxiv.org/abs/2610.07601
作者: Youngjae Cho,Won Young Jhoo,Jongsuk Kim
类目: Artificial Intelligence (cs.AI)
备注: Preprint
Abstract:Reinforcement learning with verifiable rewards (RLVR) provides a natural framework for adapting pretrained models to video temporal grounding, where generated temporal intervals can be scored directly against ground truth intervals. Yet existing overlap verifiers typically score each rollout independently, leaving the joint structure of the rollout group unused. We introduce SUTURE, which conditions verification on the rollout group and exploits its structure at two complementary scales: disagreement across rollouts controls how strongly the target is reweighted, while coverage at each position determines where reward mass is redistributed. We show that the resulting verifier admits an exact decomposition into the standard IoU term and a covariance correction determined by the rollout group. A local gradient diagnostic finds a preference for responses covering relatively less supported target regions in the analyzed groups. Across five temporal grounding benchmarks, SUTURE improves grounding performance at every reported IoU threshold. Its trained policy also shows less video-start anchoring in reasoning traces: for later events, the first temporal mention more often overlaps the annotated target. Together, these results show that the joint structure of a rollout group can support a more informative temporal verifier.
[AI-120] Modeling Latent Disturbances for Robust Decision-Making in World Models
链接: https://arxiv.org/abs/2610.07599
作者: Junwon Seo,Andrea Bajcsy
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:
Abstract:In this paper, we study robust decision-making in the latent space of world models (WMs). Robust optimization is a mathematical framework where, given explicitly specified dynamics and physically meaningful disturbances, a robot can select actions that remain effective even under worst-case disturbances. However, applying this principle to the learned latent space of WMs introduces a fundamental challenge: because WMs have fully learned state spaces and dynamics inferred from high-dimensional observations, it is unclear how to define latent-space disturbances that faithfully represent uncertainty in the underlying system. Our key idea is to model a latent-space disturbance as a perturbation to the learned latent dynamics that induces pessimistic but plausible transitions. Specifically, we construct a set of plausible latent dynamics by combining a dynamics-aware similarity metric that captures plausible transitions with out-of-distribution detection that excludes implausible latent states. We calibrate this uncertainty set over latent dynamics using conformal prediction, ensuring that WM imaginations induced by the latent disturbance remain plausible without becoming overly pessimistic. We then jointly optimize robust robot actions and the worst-case latent disturbances through game-theoretic optimization. We leverage this latent-space robust optimization to robustify policy steering, considering two paradigms: latent safety filtering and sample-and-verify steering of a generative control policy. Our controlled simulation experiments show that our latent disturbance enables robust decision-making directly in WM latent spaces, and hardware experiments with a Franka manipulator show that modeling latent disturbances enables robust policy steering, reducing failures by 70% in safety filtering and 54% in sampling-based policy steering. Project website: this https URL.
[AI-121] LSC-DPO: Learning-Signal-Controlled Direct Preference Optimization
链接: https://arxiv.org/abs/2610.07592
作者: Yang Qu,Yusheng Han,Chengjia Feng,Handan Liu
类目: Artificial Intelligence (cs.AI)
备注: 27 pages, 15 figures, 14 tables
Abstract:Direct Preference Optimization (DPO) has become a standard reward-model-free approach for aligning language models with preference data. However, as the scaled preference margin grows during training, the logistic DPO loss becomes progressively less sensitive to further changes. We study DPO from a loss-level geometric perspective and identify the sigmoid factor as a learning signal that characterizes the local sensitivity of the objective. Based on this view, we propose Learning-Signal-Controlled Direct Preference Optimization (LSC-DPO), which dynamically regulates the learning signal near a target regime. A log-space analysis establishes conditions for stable tracking of the target learning-signal regime. Experiments on AlpacaEval 2, MT-Bench, and Anthropic-HH show that LSC-DPO consistently improves over DPO and strong preference-optimization baselines. We further find that different coefficient initializations induce distinct transient learning-signal trajectories even when their later signal levels become similar. Based on this observation, we derive a signal-budget compensation rule that adjusts the target learning signal to compensate for these transient differences. The resulting compensation substantially reduces performance variation across coefficient initializations.
[AI-122] Personal-Agent Mediated Recommendation with Cross-Platform User History
链接: https://arxiv.org/abs/2610.07588
作者: Yu Xia,Jiangfan Zhang,Jun Xiao,Julian McAuley,Xiangjun Fan
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:
Abstract:Modern recommendation is shifting from platform-centric personalization toward user-governed personalization, where a personal LLM agent can act on the user’s behalf across services. We formalize this emerging paradigm as Personal-Agent Mediated Recommendation: a platform recommender ranks a candidate set using platform-local information, and a personal agent uses user-authorized cross-platform history to mediate the resulting ranking and produce the final top-K slate. Such mediation is nontrivial: the platform ranking can encode strong population evidence that the personal agent cannot observe, so effective mediation must therefore balance beneficial rescues against harmful overrides. To study this trade-off, we introduce MediateRec, a benchmark that includes scalable proxy cross-platform environments and a real cross-platform test under a controlled platform-agent information boundary. To train the agent to use cross-platform history effectively, we further propose Personal Attribution Mediation Optimization (PAMO), which counterfactually masks that history to estimate personal mediation support and reallocates rank-aware advantage mass under a platform-relative value floor. We theoretically prove that PAMO preserves cutoff-level advantage mass and is locally optimal among first-order reallocations that preserve this mass without lowering average platform-relative value. Experiments on MediateRec show that personal-agent mediation enables meaningful platform corrections, yet even strong proprietary LLMs introduce non-negligible harmful overrides. PAMO consistently improves over matched outcome-only RL across seen and unseen target platforms and on the real cross-platform test, while achieving a better rescue-harm balance.
[AI-123] Mechanistic Interpretability of Atmospheric Rivers in GraphCast NEURIPS
链接: https://arxiv.org/abs/2610.07583
作者: Madelyn Mathai,Timothy B. Higgins,Kevin M. Grise,Chirag Agarwal,Antonios Mamalakis
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: Accepted to TCCML NeurIPS workshop 2026
Abstract:While AI weather models now rival operational forecasts, how they represent the atmosphere internally remains an open question: feature attribution reveals which input patterns matter, not what the model computes or how it combines information internally. We train sparse autoencoders (SAEs) on GraphCast to uncover its learned concepts, using atmospheric rivers as our phenomenon of focus. Both standard and Matryoshka SAEs show GraphCast computes atmospheric river intensity, measured by integrated vapor transport (IVT), as a stable internal variable, despite IVT being neither an input nor a target. In contrast to the unstructured concept retrieval of the standard SAE, the Matryoshka SAE orders concepts by importance and exposes their relations. Atmospheric river concepts persist across depth and direct interventions confirm causality. This method offers a way to find internal variables and determine which of them the model actually relies on, which is a prerequisite for asking whether those variables remain meaningful as the phenomenon changes under a warming climate.
[AI-124] Representation Bias Correction Transfer and Resolution Sensitivity in Three-Dimensional Mitochondrial Morphometry
链接: https://arxiv.org/abs/2610.07582
作者: Farouk Ganiyu Adewumi,Timothy Oladunni
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Quantitative imaging pipelines can produce precise but systematically different measurements of the same object. We present an empirical reliability assessment of three-dimensional mitochondrial morphometry that connects representation bias, a controlled processing intervention, correction transfer, and resolution sensitivity. Using 2,720 development objects from the 3D Mitochondria Shape Library for Optical Microscopy, we find that occupancy-derived volumes exceed reference mesh volumes by 3.665% on average despite an intraclass correlation coefficient of 0.994. Boundary analysis identifies an outward label displacement of 0.00304 normalized units. In a controlled label-pipeline reimplementation, removing the depth offset reduces volume error in all 55 analyzed objects by a mean of 1.57 percentage points, approximately 45% of mean reproduced inflation; the source of the remainder is not isolated. A frozen regression using occupancy-derived features reduces median absolute percentage error from 3.481% to 0.664% in 2,728 previously unused objects from the same resource. However, its calibrated error bound covers only 92.1% overall and 49.2% in a low-occupancy subgroup, demonstrating that accuracy and uncertainty transfer must be evaluated separately. In 550 rat-cortex objects from the MitoEM resource, coarsening in-plane spacing from 8 to 24 nanometers changes median surface area by minus 10.60% and sphericity by plus 11.76%, despite a rank correlation of 0.994. These results provide quantitative checks for distinguishing processing-induced descriptor changes from candidate biological differences, without establishing biological invariance or cross-source correction transfer.
[AI-125] Cooperating with Future Collaborators: Multi-Agent RL under Staggered Participation
链接: https://arxiv.org/abs/2610.07578
作者: Jianglin Qiao,Siyi Hu,Thien Hoang Nguyen,Zehong Cao,Salah Sukkarieh
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:In cooperative Multi-Agent Reinforcement Learning (MARL), agents are often trained under concurrent participation, while in many tasks some agents act earlier and leave task-relevant information that becomes useful to agents participating later. We study this setting as staggered participation (SP), which introduces a cross-time, cross-agent learning dependency because an early action may affect the return through the information it provides and the later policy that uses it. Learning under SP therefore requires both identifying what information is useful for future decisions and learning how later agents should use it. We propose Staggered Participation Learning (SPL), a training-time augmentation that addresses these two parts with prospective acquisition supervision for earlier agents and outcome-supervised receiver learning for later agents. We evaluate SPL across multiple policy-based MARL backbones, environments, and staggered-participation patterns. Across 60 MPE/RWARE backbone setting comparisons, SPL achieves higher observed mean task completion in every case, with an average difference of 14.1%. The gains also extend to eight-agent teams and a physics-based UAV-UGV environment in Isaac Lab, providing evidence across algorithmic, temporal, and embodied settings.
[AI-126] Unanimously Wrong: Certified Abstention from How Medical LLM Consensus Forms NEURIPS2026 ALT
链接: https://arxiv.org/abs/2610.07570
作者: Xiaoyang Wang,Tianrui Wang,Christopher C. Yang
类目: Artificial Intelligence (cs.AI)
备注: Accepted at the GenAI4Health Workshop at NeurIPS 2026
Abstract:In clinical practice, agreement among independent experts is treated as evidence of reliability, and multi-round consensus has become a core mechanism of agentic medical question-answering systems. When such a system must decide whether to trust its own answer, the prevailing signal is again agreement, now among the sampled answers. But agreement is a fragile proxy for correctness. A system can be unanimously wrong, returning the same incorrect answer on every sample, and on these questions agreement-based signals carry no information. The cause is that these signals read only the final state of the consensus and discard how it was reached. Agreement that was reached by resolving disagreement with evidence looks identical, at the end, to agreement that was present from the first sample because every sample shares one misconception. ProbeGuard is a certified abstention framework that bases the abstention decision on how the consensus formed. Process features trace agreement trajectories, minority persistence, and retrieval saturation. For unanimous votes, rationale semantic entropy checks whether the reasons behind the vote cohere, and an active probe retrieves counter-evidence and measures whether the consensus survives. A stratified Learn-then-Test calibration then converts these scores into a distribution-free bound on selective risk. We evaluate ProbeGuard on three medical QA benchmarks and a hard-frontier reference, with a published multi-round agentic RAG substrate, against six abstention baselines. On MedQA, 13.4% of unanimous votes are wrong, and no agreement-based signal can flag them. Process signals raise the discrimination of correct from incorrect consensus from chance to 0.696 AUROC. The certified rule answers six in ten unanimous-layer questions at an observed selective risk of 9.0%, and nine in ten once in-domain calibration data accumulate.
[AI-127] Complementary Feature Domains: Information Preservation Does Not Imply Predictive-Contribution Preservation
链接: https://arxiv.org/abs/2610.07565
作者: Timothy Oladunni,Farouk Ganiyu-Adewumi
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:
Abstract:Complementary Feature Domains (CFD) theory characterizes predictive value as a context-indexed contribution system induced jointly by representations and their realization family. We show that Shannon-information preservation does not imply preservation of this contribution system: an invertible representation transformation can leave target information unchanged while altering predictive contribution under a restricted decision family. We formalize the resulting transition through a CFD contribution defect that measures how contextual contributions change under controlled recoding. For bounded Lipschitz utility, we show that each coalition utility shift is bounded by the behavioral distance between the attainable action sets before and after recoding; consequently, every contextual contribution defect is bounded by the sum of the corresponding coalition incompatibilities. Exact behavioral closure yields invariance, while increasingly accurate compensation yields restoration. A controlled ECG experiment illustrates the mechanism: a nonlinear bijective recoding preserves the information in a frozen time-frequency representation but changes accuracy under a fixed affine learner; applying the exact inverse restores all tested coalition accuracies. The result separates information preservation from realization-dependent contribution and provides a quantitative transition law for multi-representation prediction.
[AI-128] Learning a Mixture of GFlowNets
链接: https://arxiv.org/abs/2610.07562
作者: Tiago da Silva,Amauri H. Souza,Salem Lahlou
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Machine Learning (stat.ML)
备注:
Abstract:Learning an ensemble of GFlowNets to sample from a discrete target distribution has become a common approach for achieving better state space exploration and convergence than that of a monolithic sampler. However, these methods often add a substantial runtime overhead to the base model, and their conceptual connection remains elusive. To address this, we first propose a general-purpose theoretical framework for describing a mixture of GFlowNets, which we specialize into continuously (CI) and discretely indexed (DI) collections. On the one hand, we show CI GFlowNets can be interpreted through the lens of a random features expansion, provably boosting the sampler’s expressivity in graph-structured tasks and reducing learning instability via spectral shifting. On the other hand, we demonstrate DI GFlowNets encompass prior approaches for GFlowNet training and provide the foundation for the newly proposed Stratum-Conditioned (SC) GFlowNets. This method, which is inspired by the Doob’s h-transform of Markov chains, decomposes the state space according to a prescribed modular function and restricts each component to sample from a distinct subset of it. Importantly, SC GFlowNets support centralized and component-wise embarrassingly parallel training, and we show both of them significantly speed up learning convergence and mode coverage without introducing any non-negligible extra computation.
[AI-129] Navigating Route Latent Space for Synthesizable Molecular Design
链接: https://arxiv.org/abs/2610.07560
作者: Tao Li,Tuan Vinh,Monika Raj,Yuan Fang,Zhichun Guo,Carl Yang
类目: Artificial Intelligence (cs.AI); Computational Engineering, Finance, and Science (cs.CE)
备注:
Abstract:Goal-directed molecular design has advanced rapidly, yet a substantial proportion of designed molecules remain difficult to synthesize in practice, limiting their real-world utility. Prior synthesizability-aware methods either project generated molecules back to synthesizable analogs that deviate from the intended target, or optimize directly in discrete synthesis spaces that lack a continuous landscape for efficient search. We argue that this limitation mainly comes from the search space rather than the optimizer. To address this, we propose RouteFlow, a framework that reformulates synthesizable molecular design as a search over a continuous route latent space, where each latent maps back to a complete synthesis route and synthesizability is inherently preserved. To navigate this space, we adopt reward-guided flow matching as an efficient sampler that steers toward high-property regions. Since reward optimization may push latents off the manifold of real synthesis routes, where decoding becomes unreliable, we further introduce a cycle-consistency mechanism to stabilize fine-tuning. Across 16 optimization tasks from Therapeutic Data Commons, RouteFlow achieves the best sample efficiency among synthesizability-aware baselines, with the best synthetic accessibility and the highest retrosynthesis success rate. Our results also confirm that the proposed cycle-consistency reliably keeps optimization on-manifold while improving target properties, supporting effective synthesizable molecular discovery.
[AI-130] Seeing the Invisible: Physics-Guided Visual Prompting for Temperature- and Radiation-Aware VLA Navigation
链接: https://arxiv.org/abs/2610.07558
作者: Hojoon Son,Fan Zhang
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: 8 pages, 7 figures, 2 tables
Abstract:Vision-Language-Action (VLA) models have become a major paradigm for Vision-and-Language Navigation (VLN). However, in safety-critical facilities, invisible risks such as radiation or temperature spikes cannot be detected by an RGB camera, and handling each risk is expensive, requiring a new encoder, new data, and model retraining. We propose Physics-Guided Visual Prompting (PG-VP), a plug-and-play multimodal perception module that instead reuses what a frozen VLA model already does well: avoiding visible obstacles. Given a proximal radiation or thermal source, PG-VP performs a physics-guided risk assessment to determine the avoidance direction and overlays a corresponding virtual obstacle that moves across consecutive frames (Dynamic Visual Prompting). The navigation policy then naturally detours around this invisible hazard. The identical virtual obstacle is used regardless of hazard type, so the visual prompting pattern remains fixed as sensors are added. When no hazard is detected, nothing is rendered, and the policy behaves exactly as it would without PG-VP. We evaluate PG-VP on OmniNav using the val-unseen splits of R2R-CE and RxR-CE, where it guides the policy toward intended low-risk actions in 84.9% and 83.2% of cases, at a cost of 6.8 and 7.9 percentage points in navigation success rate. We further test it with distinct scenarios on a real robot in the presence of actual thermal and radiation sources, all without any retraining. The real test shows that PG-VP effectively avoids these invisible hazards, improving worst-10% average trajectory safety by 63.45% and 32.59% against thermal and radiation sources, respectively.
[AI-131] CheckerBench: Can Long-Horizon Agents Synthesize Static-Analysis Checkers?
链接: https://arxiv.org/abs/2610.07557
作者: Hang He,Li Wang,Hao Chen,Yuchen Shao,Yuling Shi,Lisheng Wang,Peiyang Liu,Goose Lin,Zaiyuan Wang,Haiying Sun,Ting Su,Chengcheng Wan
类目: oftware Engineering (cs.SE); Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR)
备注:
Abstract:Static-analysis checker synthesis requires agents to interpret a defect specification, inspect a repository, implement analyzer-specific logic, and refine the checker through repeated compilation and analysis feedback. Existing coding-agent benchmarks focus on tasks such as patch generation or vulnerability detection and rarely assess whether an agent can develop a working checker in a repository from start to finish. We introduce CheckerBench, an executable benchmark of 300 tasks derived from 297 CVEs across 167 repositories, 85 CWEs, and five language ecosystems. Each task includes vulnerable and fixed revisions, a pinned analysis environment, and a checker scaffold. We further introduce CheckerLab, a common evaluation framework that independently rebuilds submitted checkers and measures vulnerable-fixed diagnostic contrast, patch localization, false positives, and tool use. Across 21 model-harness configurations and three independent repeats per configuration, mean Pass@1 is 32.30%, while the best reaches 45.33%. These results show that reliable, reusable checker development remains challenging for current coding agents.
[AI-132] Decoupled Multi-Agent Orchestration
链接: https://arxiv.org/abs/2610.07556
作者: Xinle Wu,Yao Lu
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Learned orchestration can automatically construct effective language-model multi-agent systems, but existing approaches couple planning to fixed worker pools and train decomposition and collaboration from the same terminal outcome, limiting transfer and obscuring credit assignment. We introduce DeOrch, which separates worker-agnostic planning from concrete worker selection. Its two-stage planner first decomposes the task without worker information, then chooses collaboration operations using compact, worker-identity-free matchability feedback from the pool, enabling conditional credit assignment to decomposition and collaboration decisions. A lightweight matcher estimates worker suitability from behavior on a fixed probe set and adapts online with a contextual bandit, allowing new workers to be incorporated without retraining the planner or matcher. Across diverse in- and out-of-distribution tasks, DeOrch outperforms prior automatic MAS orchestration methods with fewer worker calls than competing learned orchestrators, remains effective when transferred to an entirely unseen worker pool without retraining, and shows consistent gains from both components.
[AI-133] Which and When to Admit: Gradient Admission for Data-Centric Small Language Model Finetuning
链接: https://arxiv.org/abs/2610.07553
作者: Hongyu Cao,Yanchi Liu,Kunpeng Liu,Xujiang Zhao,Wei Cheng,Zhengzhang Chen,Yanjie Fu,Haifeng Chen
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:
Abstract:LoRA fine-tuning adapts small language models (SLMs) to heterogeneous instruction data within a low-rank update subspace, making it vulnerable to three structural problems: conflicting gradients that cancel, static data selection that cannot track evolving learning dynamics, and subspace saturation that causes later updates to overwrite useful directions. We argue that effective adaptation therefore requires controlling which data-induced gradients enter the LoRA subspace and when. We propose GRADE (GRadient-Aligned Data-centric rEcipe), a data-centric framework combining two mechanisms: a state-aware selector that continually admits samples aligned with the evolving multi-task gradient field, and a self-calibrating step-level gate that rejects updates likely to cause destructive overwrite near saturation. Across three current-generation backbones and a heterogeneous seven-dataset instruction pool, GRADE outperforms strong data-selection and PEFT-stabilization baselines in accuracy and robustness. It is the only method to improve consistently over standard LoRA on every architecture, while producing more coherent gradient trajectories and less destructive overwrite. These results show that successful SLM adaptation depends not only on which data are selected, but also on which gradients are allowed to enter and persist in the constrained update subspace.
[AI-134] Foundation Model-Aided Multi-Agent Reinforcement Learning for Wireless Random Access Network Optimization
链接: https://arxiv.org/abs/2610.07550
作者: Myeung Suk Oh,Zhiyao Zhang,Alvaro Velasquez,Nathaniel D. Bastian,Jia Liu
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: This paper has been accepted in ACM International Symposium on Theory, Algorithmic Foundations, and Protocol Design for Mobile Networks and Mobile Computing (MobiHoc) 2026
Abstract:Random access (RA) is one of the most foundational medium access control (MAC) layer scheduling schemes for handling unpredictable data traffic from multiple terminals. While multi-agent reinforcement learning (MARL) has been explored to optimize RA-based wireless networks, its reliance on experience-driven, distributed policy learning incurs significant training overhead for each optimization task, limiting its feasibility in real-world applications. In this work, we propose to leverage a foundation model (FM) to improve MARL efficiency across diverse RA network optimization tasks. Specifically, we design an FM-aided actor-critic algorithm within a consensus-based decentralized MARL architecture and provide its convergence analysis under local reward exchanges and nonlinear value function approximations to show that our algorithm achieves the same convergence order as the conventional MARL with critic model exchanges and linear approximations. Our numerical results show that our FM-based approach significantly enhances MARL speed for RA network optimization.
[AI-135] A Systematic Investigation of Bias in Large Language Models for Advertising Relevance
链接: https://arxiv.org/abs/2610.07544
作者: Weiwei Wang,Yinchuan Xu,Jialu Gao,Youkow Homma,Jian Jiao
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Large language models (LLMs) are increasingly used to judge how well an advertisement matches a query, but the fairness of these judgments has received limited attention. We conduct a systematic study of fairness in relevance judgments made by LLMs for queries and advertisements. Our counterfactual framework examines the effects of advertiser identity and possible popularity, input language, and demographic wording. We study GPT-4o as a categorical relevance judge and a Qwen-7B model trained specifically for relevance prediction. The advertiser and language experiments use query and advertisement pairs sampled from real advertising logs. Controlled synthetic queries are used to study demographic associations in employment, housing, and credit. For both models, changing the advertiser identity or input language can alter the relevance assessment. Selected demographic comparisons also show patterns consistent with common stereotypes, particularly those involving gender and occupation. We further study mitigation during model inference and training. The results indicate that its effectiveness depends on whether advertiser information is relevant to the query and how advertiser labels are distributed in the training data. These findings can help advertising practitioners identify fairness risks and develop suitable mitigation methods for LLM relevance systems.
[AI-136] Grounding What Shapes the Plan: Rethinking Groundedness for Physical Intelligence in Autonomous Driving
链接: https://arxiv.org/abs/2610.07521
作者: Minkyoung Cho,Zewei Zhou,Wenhao Ding,Shuhan Tan,Boyi Li,Yuxiao Chen,Yan Wang,Zheng Lian,Min-Hung Chen,Chaowei Xiao,Zhuoqing Mao,Boris Ivanovic,Marco Pavone,Yulong Cao
类目: Artificial Intelligence (cs.AI)
备注: 20 pages; Project website: this https URL
Abstract:Driving models increasingly ground reasoning in causal relations, spatial structure, perceptual evidence, and predicted futures. These advances make reasoning more faithful to the driving scene, but leave a fundamental question unresolved: what should groundedness mean when the model ultimately outputs an action? Correctly grounded reasoning does not, by itself, ensure desirable driving outcomes. We introduce GroundAct, which starts from a simple premise: driving unfolds through physical entities and their interactions. Entities therefore become the unit of grounding; a lightweight reference token keeps each selected entity’s continuous state addressable through symbolic reasoning; and only the referenced entities’ interactions with the evolving proposal correct the plan. The result is an explicit path from what reasoning grounds to what the plan does, which we call grounded planning. To assess its practical value, we evaluate GroundAct in both open- and closed-loop settings. GroundAct shows strong open-loop planning across normal, out-of-distribution, and safety-critical scenarios, with closed-loop results extending this evidence to driving in simulation.
[AI-137] Harmful SFT Leaves a Continuous Trace in LLM Checkpoint Updates
链接: https://arxiv.org/abs/2610.07518
作者: Ziqun Bao,Xinyu Zhang,Yuchen Shao,Chengcheng Wan
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR)
备注:
Abstract:Safety auditing of post-trained large language models typically relies on model behavior, requiring model execution and depending on the coverage of available evaluations. This work asks a different question: Do the target behaviors optimized during supervised fine-tuning (SFT) leave readable evidence directly in checkpoint updates? We find that harmful-compliance SFT induces a continuous, objective-dependent ordering in checkpoint-update space. Using a reference geometry defined by pure harmful-compliance, safety-targeted, and benign-utility SFT, we find that a checkpoint-level coordinate s_H tracks controlled harmful-objective composition with Spearman correlations of 0.986-0.992 across four 7-8B backbones, with the same ordering persisting at larger model scales. Matched compliance-versus-refusal controls show that this checkpoint trace reflects the SFT objective rather than harmful-input exposure, while additional controls rule out simple explanations based on harmful-example count or generic training intensity. Building on this structure, we introduce TRACE, a weights-only auditing method that localizes an unknown checkpoint update relative to frozen harmful and non-harmful reference prototypes and converts this geometry into a continuous harmful-objective score. TRACE requires neither model queries nor access to the unknown SFT data, and can be evaluated directly from checkpoint updates. Across distribution shifts, unseen data, different SFT configurations, partial checkpoint access, and LoRA/full-parameter fine-tuning, the trace remains stable and is positively associated with independently measured attack success rates. TRACE remains informative even at low harmful-objective proportions, providing a complementary auditing signal when behavioral evaluation is unavailable or incomplete. Code is available at this https URL.
[AI-138] From Local Evidence to Safety Verdicts: Causal Tracing in Vision-Language Models
链接: https://arxiv.org/abs/2610.07514
作者: Faezeh Dehghan Tarzjani,Mevan Wijewardena,Alexander Romanus,Sampad Mohanty
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:A vision-language model may need to combine an image with a prompt to recognize a safety risk that neither reveals alone. Where does this joint safety judgment become accessible inside the model? We introduce SSU-Bench, a dataset of matched safe and unsafe image-text combinations constructed using single-item prompt edits or image edits with annotated intended regions. Using three vision-language models, we transfer internal states between paired inputs and measure the resulting change in the safety verdict. Across models and both types of counterfactual, interventions at the changed input positions are effective in earlier decoder layers, while interventions at the final input token become effective later. Directions estimated from other examples produce similar late-layer effects. A linear readout of the final-token state also predicts the model’s own verdict, including incorrect judgments, and cross-model comparisons reveal similarities in the patterns of counterfactual change. These findings identify a recurring transition in where interventions can influence a joint safety verdict and distinguish a readable model decision from a correct safety judgment.
[AI-139] MARS: Multi-resolution Adaptive Routing for Sequential Recommendation
链接: https://arxiv.org/abs/2610.07505
作者: Ming Yin,Sixun Dong,Yudong Liu,Wen-Yun Yang,Yunjiang Jiang,Yiran Chen
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Long-history recommenders often compress each user’s history into a compact, candidate-independent memory that is cached and reused to score large candidate pools. We show that real user histories exhibit multi-scale semantic structure, with short-lived intent, medium-term interests, and long-term preferences coexisting in one sequence, and that monolithic cached memories preserve these scales unevenly: linear probes recover recent and mid-range content far worse than long-range content. We call this failure mode \textittemporal aliasing. We propose \textbfMARS, a multi-resolution user memory that writes the full history into recurrent state tracks anchored to different half-lives, and a sparse routing reader that materializes compact seed memories by selecting the relevant temporal resolutions for each seed, preserving fixed-size candidate scoring. MARS outperforms strong baselines on three public datasets, with gains that grow with history length. Component-matched ablations with paired tests show that temporal diversity and selective routing each contribute beyond what hard-window memories or added capacity provide. The advantage of MARS over its interface-matched baseline also widens after within-user behavioral shifts, at about 1.02\times that baseline’s warm-cache serving latency for 1,000 candidates per user.
[AI-140] Does Muon Need Fine-Grained Spectral Shaping?
链接: https://arxiv.org/abs/2610.07497
作者: Meher Chaitanya,Tianyi Zhou,Aristides Gionis
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Machine Learning (stat.ML)
备注:
Abstract:Muon combines current and past gradients into matrix momentum. For M=U\Sigma V^\top , the idealized polar update Q=UV^\top gives every singular direction the same weight. We refer to this as the flat profile. Several recent optimizers replace this flat profile with fine-grained spectral maps that give each direction its own gain. We ask how much of this spectral detail a Muon update needs. Our spectral diagnostics show that approximately 94 – 97% of measured singular modes lie below an estimated noise edge, yet collectively align positively with a reference gradient. We introduce BulkBoost, a two-band spectral reweighting framework with fixed-rank and noise-calibrated variants. The latter uses split-minibatch gradient differences to calibrate a Marchenko–Pastur reference edge for Muon’s Nesterov input, separating the bulk below the edge from the spikes above it. Both variants increase the bulk’s relative weight through one shared gain while preserving the Frobenius norm of each matrix’s unreweighted direction. For a fixed partition, our theory gives the first-order condition under which moving weight toward the bulk lowers the loss. It also quantifies the fraction of the maximal first-order improvement rate, over all per-mode reallocations, that two bands can capture. Across 30 continued-pretraining settings spanning Pythia-14M to 410M and six corpora, two-band reweighting is competitive with the fine-grained power-law profile of Freon and outperforms Spectra. Measured against Muon’s flat profile, Freon reduces final loss by 0.022% of the pre-adaptation loss on average, whereas the two-band variants achieve reductions of 0.073 – 0.147% . These observations suggest that useful departures from the flat profile are surprisingly low-dimensional: a single bulk-to-spike gain captures at least as much benefit as the fine-grained spectral profiles. Subjects: Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Machine Learning (stat.ML) Cite as: arXiv:2610.07497 [cs.AI] (or arXiv:2610.07497v1 [cs.AI] for this version) https://doi.org/10.48550/arXiv.2610.07497 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[AI-141] Can Power Draw Constrain Covert Compute? Limits of Analogue Verification for AI Governance
链接: https://arxiv.org/abs/2610.07476
作者: Tom Kimpson,Mauricio Baker,Emlyn Graham
类目: Computers and Society (cs.CY); Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR)
备注: 14 pages, 8 figures
Abstract:Frontier AI treaties or agreements on limiting computation require external verification; an external auditor must be able to confirm how much computation actually ran and that parties are adhering to the agreement. Analogue, off-chip measurements such as power draw provide an information channel for verification. It is unknown how well these analogue channels can constrain computation against an adversary who actively tries to subvert the audit. We derive a closed form for \beta , the largest hidden computation a power trace cannot exclude, as a fraction of the declared machine capacity. Measurements on NVIDIA A100 GPUs constrain \beta = 1.16 in the worst case, while adversarial matched-energy strategies are shown to hide at least \beta = 0.41 of compute. Analogue power measurements alone therefore constrain compute weakly. Additional restrictions granted by the threat model, such as the ability of the verifier to re-execute the declared work at an observed operating point, let the verifier push \beta down to 0.059 in the maximally restricted case. This gives a quantitative estimate of what analogue measurements can contribute to compute verification.
[AI-142] PsyCIDRA: A Dual-Agent Framework for Psychiatric Interviewing and Diagnostic Reasoning
链接: https://arxiv.org/abs/2610.07473
作者: Milad Mohammadi,Fatemeh Akrami Shamsabadi,Zahra Mohseni,Amirhossein Safdarian,Malekfarhad Malek,Hadi Moradi,Hesham Faili
类目: Artificial Intelligence (cs.AI)
备注: 41 pages, 27 figures, 20 tables; includes appendices
Abstract:Large language models show promise in clinical reasoning, but psychiatric interviewing requires guiding an evolving conversation. Their ability to carry out this interactive assessment remains less studied. We present PsyCIDRA, a dual-agent framework linking free-form psychiatric interviewing with diagnostic reasoning for expert review. Its interviewer agent uses tools to maintain working notes, load expert-written skills, and retrieve ICD-11 references to guide inquiry. Its diagnostic reasoning agent then receives the completed interview transcript and reports hypotheses alongside supporting, conflicting, and missing evidence, withholding a final hypothesis when none is sufficiently supported. Using patient profiles generated with PsyCPG, we first evaluate PsyCIDRA in simulation. Across four models on 53 evaluation cases, it achieves higher diagnostic agreement than direct prompting. On 81 held-out simulated cases, rank-1 accuracy is 60.5% versus 51.9%. In a blinded study of 101 human participants in separate arms, PsyCIDRA agrees with psychologists on whether to propose a diagnostic hypothesis in 79.6% of cases, compared with 65.4% for direct prompting. Together, these findings support the potential of LLM agents to assist psychiatric assessment through free-form dialogue. By examining diagnostic reasoning, interview quality, and safety together, this study contributes to understanding the capabilities and limitations of psychiatric interview agents.
[AI-143] Structure Not Belief: Correlated Thompson Sampling from LLM -Derived Covariance in Combinatorial Semi-Bandits NEURIPS2026
链接: https://arxiv.org/abs/2610.07470
作者: Vikram Kakaria,Anish Kataria,Anany Kotawala
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Machine Learning (stat.ML)
备注: 15 pages. Accepted (poster) at DynaFront 2026: Dynamics at the Frontiers of Optimization, Sampling, and Games, NeurIPS 2026 Workshop
Abstract:Combinatorial Thompson sampling (CTS) draws independent posterior samples for every arm, so its exploration dynamics ignore any relation among arms. We study a minimal change to those dynamics: an LLM is queried once for a partition of the arms, the partition becomes a positive-definite correlation matrix \Sigma through an RBF kernel on cluster ranks, and the per-round posterior sample is drawn with covariance \Sigma while the Beta posteriors are updated from real rewards only, so the LLM shapes how the sampler moves, not what it believes. We give a self-contained Bayesian regret bound for the idealized Gaussian sampler whose information gain splits into a K\log T term from the K -cluster structure and a ridge term that grows to d\log T : the \sqrtd/K improvement over independent sampling is a finite-horizon transient, exact only as the within-cluster correlation tends to one. The correlated sampler reduces regret by 19% over CTS on 16 synthetic Bernoulli families at T=2,500 (6-7% at T=25,000 with data-adaptive kernels) and by 41% on the Microsoft MIND-small news benchmark ( d=200 real articles), while pseudo-observation warm starts give nothing. An LLM-free ablation with a simulated oracle of controlled quality shows that on unstructured instances the gain is a property of the kernel shape (a random partition, or a plain tempering of the sampling noise, reproduces it), while belief injection at matched oracle quality never helps.
[AI-144] COMPASS: Finding Where Reasoning Lives in Language Models
链接: https://arxiv.org/abs/2610.07469
作者: Pratyay Dutta,Kowshik Thopalli,Vivek Narayanaswamy
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Explicitly eliciting reasoning substantially improves LLM performance. Existing approaches require a predefined characterization of reasoning, whether through CoT prompt design, contrastive CoT directions, or via SAE derived reasoning features. For mathematical reasoning with verifiable answers, we show that a much simpler signal suffices, which is the correctness of the model’s own direct answer attempts. This signal yields a latent direction that elicits reasoning. This direction is decodable within the activations of most attention heads, but only a small subset of them can be effectively intervened. We introduce COMPASS, an inference-time steering method that identifies these heads using a logit-space attribution score and steers their activations along the correctness direction, requiring only per-head activation statistics. Across three model families and multiple math benchmarks, COMPASS outperforms the activation-steering baselines we compare against, improves GSM8K accuracy by 16 percentage points on average, and approaches CoT accuracy with 20-70% fewer generated tokens. Interventions transfer without re-fitting to unseen benchmarks, and ablations show that both the correctness direction and the small set of heads carrying it are necessary, with the effect concentrated in remarkably few heads.
[AI-145] Active Feature Acquisition for Cost-Efficient Temporal Prediction with Reduced Participant Burden
链接: https://arxiv.org/abs/2610.07452
作者: Yunni Qu(1),Bing Cai Kok(2 and 3),Whitney Ringwald(4),Grant King(5),Aidan Wright(5),Kathleen Gates(2),Junier Oliva(1) ((1) Department of Computer Science, University of North Carolina at Chapel Hill, (2) Department of Psychology and Neuroscience, University of North Carolina at Chapel Hill, (3) School of Social Sciences, Nanyang Technological University, Singapore, (4) Department of Psychology, University of Minnesota Twin Cities, (5) Department of Psychology, University of Michigan)
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Applications (stat.AP); Methodology (stat.ME); Machine Learning (stat.ML)
备注:
Abstract:Accurate forecasting of pathological outcomes is a central problem in psychology. To do so, psychologists often collect intensive longitudinal data. However, in such studies, the desire to acquire a large number of variables for the sake of accurate prediction is often counteracted by the need to minimize participant burden. Acquiring more variables per occasion can yield better predictions, but having too many acquisitions increase the risk of non-response and attrition. Longitudinal Active Feature Acquisition (LAFA) is a principled approach to resolve this conundrum. Instead of requiring responses to every item at every acquisition occasion, LAFA produces a policy that seeks to optimally select dynamic subsets of items to be acquired at each timepoint while preserving our ability to forecast a specific outcome. However, existing LAFA methods are mostly based on Neural Networks (NN) that are difficult to interpret in practice. In this work, we introduce a tree distillation method for learning an interpretable policy from NN-based LAFA networks. We validated our method through both a simulation and an empirical EMA dataset on forecasting daily alcohol consumption. In both cases, we find that we can meaningfully reduce the number of items acquired at each occasion with minimal loss in accuracy. Networks (NN) that are difficult to interpret in practice. In this work, we introduce a tree distillation method for learning an interpretable policy from NN-based LAFA networks. We validated our method through both a simulation and an empirical EMA dataset on forecasting daily alcohol consumption. In both cases, we find that we can meaningfully reduce the number of items acquired at each occasion with minimal loss in accuracy.
[AI-146] When Does AI Supervision Help? A Role-Aware Study of Network Fraud Decision Management with Blockchain Auditability
链接: https://arxiv.org/abs/2610.07434
作者: Saviz Changizi,Nasibeh Mohammadzadeh,Mohammad Shojafar,Rahim Tafazolli
类目: Artificial Intelligence (cs.AI)
备注: 27 pages, 10 figures, 13 tables
Abstract:When does a second artificial intelligence (AI) component improve a primary network-fraud decision rather than add operational burden? We study this question through a role-aware Decider-Supervisor (DS) framework with blockchain auditability, evaluating four directional configurations that combine centralised machine learning, a Federated Averaging (FedAvg)-trained federated meta-model, and Base or Quantized Low-Rank Adaptation (QLoRA) large language model variants. The analysis compares primary-only and supervised decisions using non-hard fraud performance, intervention burden, conditional calibration, traffic-mix and Review-capacity sensitivity, dependability tests, and blockchain lifecycle controls. The deterministic hard gate resolves 89.994% of fraudulent requests, leaving the non-hard population as the main AI decision setting. Conditional validation calibration does not produce a consistently transferable supervisory advantage on deployment replay. DS-3 QLoRA is the least disruptive supervised configuration, but it still underperforms its primary FedAvg stage in F1 and total errors. Across 36 reweighted traffic mixtures, supervision reduces total errors only for DS-4 Base in two extreme high-fraud scenarios. Blockchain tests support digest verification, tamper detection, authorisation, single-use review resolution, and post-finalisation integrity, while exposing a pre-finalisation single-write limitation. The results show that the value of AI supervision depends on role assignment, calibration, escalation policy, traffic composition, and lifecycle controls rather than on the presence of a second model alone.
[AI-147] Adaptive Gait Biofeedback With Participant-Held-Out Modeling and Participant-Specific Updating in Chronic Ankle Instability
链接: https://arxiv.org/abs/2610.07428
作者: Jaeyoon(Jason)Kim,Veronika Lebisova,Jeniya Sultana,Jaeyoung Cho,Jaeho Jang
类目: Artificial Intelligence (cs.AI); Signal Processing (eess.SP)
备注: 28 pages (14-page main manuscript and 14-page supplementary material), 5 main figures
Abstract:Adaptive gait biofeedback may support repeated practice in chronic ankle instability, but its evaluation must address model performance and human response. We evaluated a temporal convolutional classifier on protocol-defined, angle-derived GOOD/BAD gait-cycle labels using participant-held-out leave-one-subject-out (LOSO) cross-validation in 20 participants. Seven participants in the adaptive-intervention group completed nine sessions over three weeks, with one motion-capture recording analyzed per session. Models updated after failed sessions were compared offline with their parent models on the same-session validation subset used for candidate selection and the first subsequent adaptive-session recording. Frontal-plane ankle angle was compared between the adaptive group and 10 sequentially enrolled controls at Baseline, Post, and 7-day Retention. Across 20 held-out folds, mean fold-level area under the receiver operating characteristic curve (AUROC) was 0.948, sensitivity for angle-threshold-exceeding BAD cycles was 0.941, and specificity for angle-threshold-meeting GOOD cycles was 0.366. Mean BAD-class F1 was higher in candidate models by 0.187 on the same-session subset and 0.118 on the first subsequent recording. At Post, the adaptive group had a baseline-adjusted frontal-plane ankle angle 5.168 degrees lower than controls (95% confidence interval, 1.766-8.569 degrees lower); the Retention contrast was uncertain. These findings characterize population-model discrimination and offline participant-specific updating during repeated biofeedback use, alongside a nonrandomized Post frontal-plane ankle angle association. They do not establish independent clinical gait classification or a causal benefit of updating.
[AI-148] 2d-fet-bench: from spatial reasoning to fet design on flakes
链接: https://arxiv.org/abs/2610.07423
作者: Dunzhi Zhou,Chengyu Zhu,Gang Qiu,Caiwen Ding
类目: Artificial Intelligence (cs.AI); Mesoscale and Nanoscale Physics (cond-mat.mes-hall); Materials Science (cond-mat.mtrl-sci)
备注:
Abstract:Field-effect transistor (FET) layouts on exfoliated two-dimensional flakes are typically drawn by hand for each flake, placing contacts and gates to match its position and outline in optical micrographs. To our knowledge, no executable benchmark tests whether language-model agents can perform this flake-specific construction reliably. We introduce 2D-FET-Bench V2, a benchmark of 128 layout tasks built from microscopy-derived flake contours, including hole-containing flakes and multi-flake tasks. Each task supplies a textual device specification and contour coordinates. An agent generates typed polygon and path operations rendered to GDSII. A deterministic verifier checks geometric and structural requirements, and a separate integrity check verifies that the supplied contours remain unchanged. Scripted reference layouts pass all 128 tasks, showing that every task is solvable. We evaluate six models and seven workflow and scaffold variants of GPT5.6-Luna, with five attempts per task. The best-performing configuration in the six-model panel, GPT5.6-Luna with ReAct-3, passes 62.3% of attempts and solves 80.5% of tasks at least once (coverage) and 43.8% in all five attempts (consistency). ReAct-3 exceeds the one-pass Plan-and-Execute by 27.0 pass@1 points at 2.46 times the tokens. An expert audit of one sampled verifier-passing layout per covered task, across five ReAct-3 configurations, accepts 56.4% to 63.5% of them. The benchmark evaluates geometric and structural FET layout construction.
[AI-149] Defense-in-Depth for LLM s: Evaluating Memory Gates Against Activation-Induced and Memory-Induced Sycophancy NEURIPS
链接: https://arxiv.org/abs/2610.07403
作者: Ritvij Sharma,Russell Dlugosz,Ryan Zhou,Maheep Chaudhary
类目: Artificial Intelligence (cs.AI)
备注: Accepted to NeurIPS (IAB, RTCA, AIWILD, and CL4FM)
Abstract:Long-term memory allows Large Language Models (LLMs) to maintain personalized context across interactions, but retrieved user history can induce memory-induced sycophancy, causing models to favor stored user beliefs over objective evidence. Existing defenses primarily operate on retrieved context and are rarely evaluated jointly with internal behavioral bias. We introduce a 2 \times 2 defense-in-depth framework separating internal activation steering from external memory handling. We extract sycophancy steering directions from 100 paired prompts and evaluate four open-weight models across 10 steering coefficients and five memory-defense configurations on MemSyco-Bench (answers for all 1,550 items; defense conditions judged on a fixed 250-item subsample), with three LLM judges. Three of the five configurations are new (rewriting every memory, a Router Gate that keeps, rewrites, or drops each memory, and dropping all memory); the other two are MemSyco’s baselines. Selective Router Gate filtering preserves substantially more of MemSyco’s average accuracy than complete memory removal, and this separation persists when the models are steered toward sycophancy. On Llama 3.1 8B with Router Gate, mild inverse steering ( \alpha = -1.5 ) lowers judge-averaged sycophancy from 35.80% to 31.32% while average accuracy moves from 43.99% to 43.31%; this reduction has the same direction under all three judges but is not statistically significant (paired p = 0.08 to 0.63 on 149 items). External memory filtering is the part of the design that holds up; our data do not show that inverse steering adds to it.
[AI-150] Inference and learning in sparse autoencoders as natural gradient flow
链接: https://arxiv.org/abs/2610.07389
作者: Hadi Vafaii,Tejas Rao,David Chanin,Thomas Fel,Jacob L. Yates,Bruno Olshausen,David Klindt,Dileep George,Miguel Lázaro-Gredilla
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: Code: this https URL
Abstract:Sparse autoencoders are widely used to uncover interpretable features in neural networks, yet reliable recovery remains difficult when features overlap or activate infrequently. These challenges involve both inferring which features explain an input and learning the dictionary that represents them. Here, we unify inference and dictionary learning as natural-gradient flows on a shared variational free energy. We instantiate this framework as BeFOND, an encoder-free sparse coding model with closed-form inference and learning dynamics. We show how recurrent explaining away reduces interference between overlapping features, while Fisher preconditioning can compensate for the slow learning of rare features. On synthetic data, BeFOND improves dictionary recovery and rare-feature detection, with a growing advantage over amortized baselines as superposition increases. On language-model activations, it improves single-feature concept detection and selective intervention, outperforming pretrained reference SAEs with substantially less training data. Its feature quality continues to improve with dictionary width, whereas the evaluated baselines largely plateau. Together, these results show how improving inference and learning within a unified probabilistic framework can make better use of data and dictionary capacity to interpret and intervene on neural representations.
[AI-151] MemCo: Memory-Centric Collaboration for Generalizing LLM Agents to Unseen Environments
链接: https://arxiv.org/abs/2610.07376
作者: Xinting Liao,Siyan Liu,Rabab K. Ward,Holger R. Roth,Xiaoxiao Li
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Large language model (LLM) agents increasingly operate in interactive environments, where they need to make sequential decisions through observation, action, and feedback. Although memory can help agents reuse experience, existing work designs memory in isolation, where collecting enough trajectories to populate it is expensive. Existing shared-memory approaches mitigate isolated experience by pooling episodic memories across tasks and environments. However, retrieving shared memory is challenged by the granularity, where retrieved memories can be either too specific to preserve current grounding or too coarse to support the next action. In this work, we propose MemCo, a memory-centric collaboration framework for generalizing LLM agents to unseen interactive environments. It maintains complementary local and global memory spaces, preserving environment-specific details locally while promoting transferable workflows induced from local trajectories to global memory. During online interaction, MemCo routes relevant local and global memories in terms of the agent’s current state and decision phase, enabling agents to reuse the experience of other agents without blindly transferring environment-specific details. Experiments on interactive decision-making benchmarks show that MemCo improves task success and reduces redundant exploration compared with isolate-memory and shared-memory baselines. Our code is available at this https URL.
[AI-152] Evaluate the Stack Not the Layer: Do Deterministic and LLM Gates for Agent Actions Fail Independently?
链接: https://arxiv.org/abs/2610.07359
作者: Chenglin Yang
类目: Artificial Intelligence (cs.AI)
备注: 15 pages, 1 figure, 10 tables. Artifact (data, scripts, provenance): this https URL
Abstract:Runtime gates for agent tool calls are stacked on the assumption that their errors multiply. We test it on 1,119 labelled agent actions from three corpora, without an adaptive adversary. The stack has one deterministic rule layer and four LLM judges, three of them re-collected with the served model recorded on every call. We read each stack as a number of multiplication-equivalent layers, n_mult, with its floor under perfect coupling. Under the STRICT miss definition (escalation to a human scored as not stopped), any two judges compose to about 1.2 to 1.4 layers (\phi median +0.430, 6 of 6 pairs significant, floors 1.02 to 1.17). The rule layer plus one judge composes to 1.86 to 2.09 layers (\phi median +0.014, 0 of 4 significant, floors 1.01 to 1.09). Under PRIMARY (escalation scored as caught) the bands are 1.21 to 1.57 and 1.80 to 2.13. Intervals separate on the pooled data, point estimates split on each corpus, and a third-vendor judge lands in the judge band. Solo accuracy does not predict what a layer adds: a cloud rule pack lowers the rule layer’s solo miss rate by 20% and adds no new joint coverage. The difficulty share of judge coupling is not identifiable: 31.8% to 61.8% depending on the probe and the miss definition. One judge tier was served by an unrequested model version in 50 of 112 batches, concentrated on the external corpus. That event overturned a pre-declared analysis rule, and the scoring of review verdicts reversed five conclusions. We report both.
[AI-153] A Validated Dataset and Benchmark for Coherent Multi-Diagram SysML Models
链接: https://arxiv.org/abs/2610.07356
作者: Ardalan Aryashad,Yan Jin
类目: oftware Engineering (cs.SE); Artificial Intelligence (cs.AI)
备注:
Abstract:Systems engineers use several diagrams to describe the structure and behavior of systems. Engineers create these diagrams together to make sure that they use the same elements and remain consistent with one another. Large language models can generate diagrams as text or code, which makes it possible to create system diagrams automatically. However, their ability to generate coherent sets of diagrams is not well understood, and existing datasets and benchmarks do not directly measure this ability at scale. We introduce SEMAADB (Systems Engineering Modeling Assistant with AI Dataset and Benchmark), a dataset of 3,000 engineering contexts and 15,000 diagrams. Each context contains five connected SysML views: Requirement, Block Definition, Activity, State Machine, and Sequence. Here, a view is a diagram that presents one aspect of a system. We checked the diagram sets for consistency and valid rendering. A set of 100 contexts is also human-verified and forms the benchmark test set. We evaluate three language models on two tasks. In diagram repair, the strongest model repairs 64.3% of semantic errors . In cross-diagram update the best propagation F1 is 80.7% when a model applies one change across related diagrams. The results show that syntax repair is nearly solved, but semantic repair and consistency across diagrams are still challenging tasks for models. SEMAADB therefore provides both a large diagram resource and a set of benchmarks for measuring coherent multi-diagram SysML generation.
[AI-154] Evaluating Escalation Signals for LLM Routing: Targets Controls and Five Ways to Fool Yourself
链接: https://arxiv.org/abs/2610.07354
作者: Ramin Pishehvar,Andrea Morandi,Mahesh Viswanathan
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Deciding when to escalate a query from a small language model to a larger one requires a cheap signal that predicts, before the large model is called, whether escalating would help. Semantic entropy, originally developed to detect hallucinations, is a natural candidate: it measures how much a model’s sampled answers disagree in meaning, and high disagreement often signals an unreliable answer. We test it across three benchmarks and two model families. On GSM8K, with a small/large pair about twelve times apart in size, semantic entropy reliably distinguishes the small model’s mistakes (AUROC 0.871) and improves routed accuracy over random escalation by up to nine points at matched cost. An earlier strong-looking result on a synthetic benchmark proved misleading: a simple rule based only on question difficulty, with no model involved, matched semantic entropy almost exactly. This paper’s main contribution is a set of checks that catch this before it is reported as real. We show that scoring a cheap, question-only difficulty estimate alongside any signal reveals whether the signal adds real information or just tracks how hard a question looks; that two reasonable definitions of “escalation worked” can produce very different results on the same data; that a benchmark can leave almost no room for any signal to beat simply always using the large model; and that the true cost of live sampling can make routing more expensive than calling the large model directly. For a cheaper alternative that reuses cached past outcomes, we show how to predict whether it will work on a new dataset – confirmed by correctly forecasting a collapse from AUROC 0.908 to chance level (0.518) ahead of time. We offer these as a general checklist for evaluating escalation signals.
[AI-155] rajectory-Retrieval Speculative Decoding: When Does a Models Own History Help?
链接: https://arxiv.org/abs/2610.07350
作者: Yuyang Dai,Yushun Dong
类目: Artificial Intelligence (cs.AI)
备注: 33 pages
Abstract:Long chain-of-thought reasoning increases sequential decoding cost while creating a growing history of potentially reusable continuations. We investigate when this history supplies useful drafts and complements an existing drafter. Controlled source comparisons reveal trajectory-specific reuse, motivating our method Trajectory-Local Adaptive Retrieval (TLAR). TLAR retrieves approximately matched continuations from the current trajectory and uses recent verification outcomes to adapt retrieval activation and candidate width. TLAR combines retrieved continuations with model-generated drafts in a shared candidate tree, preserving the target model’s output distribution through exact verification. Across code debugging, mathematics, and open-ended writing, our evaluation connects source reuse, incremental acceptance, and execution cost. Combining TLAR with strong retrieval baselines improves token acceptance under matched verification budgets and increases end-to-end throughput over the draft-model baseline. These findings support generated trajectories as runtime memory for adaptive inference.
[AI-156] RELACE: retrospective likelihood-based action credit estimation for long-horizon language agents
链接: https://arxiv.org/abs/2610.07349
作者: Sayak Chakrabarti,Sathish Reddy Indurthi
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:
Abstract:Group Relative Policy Optimization (GRPO) avoids a separate critic by estimating advantages from rollout groups. For multi-turn agents, however, trajectory-level supervision provides coarse, noisy credit: terminal rewards do not locate errors and can penalize useful actions alongside mistakes. Group-in-Group Policy Optimization (GiGPO) and subsequent methods refine supervision through state-conditioned comparisons, but their credit estimates remain sensitive to downstream decisions and outcomes. We introduce RELACE, Retrospective Likelihood-based Action, a critic-free framework that integrates retrospective action assessment with state-conditioned advantage estimation. RELACE evaluates executed actions through teacher-forced likelihood scoring under both their original contexts and outcome-augmented contexts. Comparing these likelihoods yields a trajectory-normalized retrospective factor that captures outcome-dependent changes in action plausibility, rather than hindsight plausibility alone. We use this factor to reweight discounted task returns and construct local advantages by comparing weighted returns among actions from equivalent states within a task. This couples retrospective relevance with observed reward, producing fine-grained credit that complements trajectory-level GRPO supervision. Temporal smoothing and success-protecting masking further stabilize the local signal. RELACE requires neither auxiliary value nor reward models nor additional autoregressive rollouts for credit estimation. Experiments on ALFWorld and WebShop with Qwen2.5-1.5B-Instruct and Qwen2.5-7B-Instruct demonstrate substantial improvements over GRPO, GiGPO, and HCAPO. With the 1.5B model, RELACE achieves 96.35% success on ALFWorld and 79.43% on WebShop, surpassing GiGPO by 5.47 and 5.60 percentage points, respectively.
[AI-157] Rationale-Guided Policy Optimization: Learning to Reason with Adaptive Rationale Scaffolding NEURIPS2026
链接: https://arxiv.org/abs/2610.07342
作者: Hoang Phan,Minh Pham,Chau Pham,Chinmay Hegde,Trung Le,Qi Lei
类目: Artificial Intelligence (cs.AI)
备注: NeurIPS 2026
Abstract:On-policy reinforcement learning has become a central paradigm for improving the reasoning abilities of large language models. However, its effectiveness is often limited by reward sparsity: when a model fails to discover correct trajectories for difficult problems, the optimization process receives little useful signal and may stagnate. Existing approaches mitigate this issue by incorporating off-policy demonstrations, expert traces, or model-generated solutions, but they typically require the auxiliary data to match the format of the reinforcement-learning task, often relying on rejection sampling from stronger models to obtain suitable training trajectories. We introduce Rationale-Guided Policy Optimization (RGPO), a framework that adaptively leverages ground-truth rationale information according to the model’s current capability while preserving its freedom to explore. Rather than treating reference solutions as fixed imitation targets, RGPO uses them as temporary scaffolds: rationales help the model generate improved responses, after which only higher-reward, model-generated solutions are transferred back to the original unguided setting. This design allows training to exploit available ground-truth information without requiring off-policy data to follow the same format as the RL task. Across both language-only and vision-language reasoning settings, RGPO consistently improves performance over RLVR baselines, and ablation studies show that adaptive rationale guidance is a key contributor to these gains. These results suggest that RGPO offers a practical and general approach for reducing reward sparsity, stabilizing reinforcement learning, and improving reasoning performance in both text-only and multimodal models.
[AI-158] CausalBind: Causal Modeling and Learning for Protein-Molecule Virtual Screening NEURIPS2026
链接: https://arxiv.org/abs/2610.07340
作者: Loka Li,Jin Tian,Kun Zhang
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: NeurIPS 2026 (Oral)
Abstract:Protein-molecule virtual screening is increasingly cast as a problem of representation learning in a shared embedding space. Existing methods rely on dense holistic alignment, entangling invariant binding determinants with nuisance correlations and limiting transfer to new targets. It has been noted that binding in protein-molecule systems involves sparse cross-modality interactions: binding is governed by a small contact interface and a few decisive local interactions (e.g., hydrogen bonds, hydrophobic contacts, and salt bridges) rather than the global structures of the protein and molecule. We hypothesize that uncovering and leveraging sparse interaction patterns is critical for generalization beyond the training data, as these patterns are reusable and expected to improve performance across different scenarios. In this paper, we aim to identify and leverage sparse interaction patterns, and verify our hypothesis. Since the training data contain only observed binding pairs, we formalize this prior via a V-structure causal model under Heckman-style selection, and establish three theoretical results: (i) the latent concepts of interacting proteins and molecules are not identifiable without appropriate sparsity constraints; (ii) these concepts and their sparse interactions are component-wise identifiable under structural sparsity conditions; and (iii) a low-rank relaxation of these conditions yields subspace identifiability of the concepts and interactions. Inspired by these principles, we propose CausalBind with three implementation variants. Extensive experiments on DUD-E and LIT-PCBA benchmarks show that all variants consistently outperform strong retrieval baselines, with the largest gains on LIT-PCBA early enrichment, and further generalize to target- and scaffold-level out-of-distribution splits. Code is available at this https URL.
[AI-159] Selective Critique for Cost-Aware LLM Agents in Long-Horizon Decision Making NEURIPS2026
链接: https://arxiv.org/abs/2610.07335
作者: Heewon Park,Somin Im,Minhae Kwon
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: Accepted to NeurIPS 2026
Abstract:Improving the reliability of large language model (LLM) agents in long-horizon decision-making remains a key challenge. When deployed as autonomous agents interacting with complex environments, early mistakes can propagate through trajectories and cause cascading failures. Recent approaches improve reliability by incorporating external critique or deliberation, but invoking these mechanisms at every step substantially increases token consumption and latency, limiting practical deployment. We propose SAG (Self-improving Agent with Gated critique), a cost-aware framework that formulates critique invocation as a step-wise decision problem during long-horizon interaction. SAG introduces a lightweight, training-free gating mechanism that estimates the utility of critique using action-level ambiguity signals–global entropy and local top-2 margin–computed over admissible actions. From a decision-theoretic perspective, this mechanism approximates the Value of Information (VoI) of critique, enabling the agent to selectively allocate expensive feedback only when its expected benefit justifies the cost. SAG further incorporates online bootstrapped self-improvement, allowing the actor to internalize critic-assisted behaviors and progressively reduce reliance on critique. Across three long-horizon interactive benchmarks and multiple backbone models, SAG substantially improves the performance-cost trade-off compared with both no-critique and always-on critique agents. On ALFWorld, SAG increases task success from 24.6% to 78.4% while maintaining a token budget comparable to ReAct, yielding a 3.1\times improvement in normalized token efficiency. Moreover, a 7B actor with a lightweight 3B critic achieves performance comparable to a 14B actor without critique, showing that selective critique can recover most of the reliability benefits of deliberation while dramatically reducing inference cost.
[AI-160] Memory-Efficient Expert Routing for Distributed MoE Training
链接: https://arxiv.org/abs/2610.07333
作者: Arnab Kanti Tarafder,Jaume Guasch-Martí,Gokcen Kestor,Jie Ren
类目: Distributed, Parallel, and Cluster Computing (cs.DC); Artificial Intelligence (cs.AI); Networking and Internet Architecture (cs.NI)
备注:
Abstract:As Mixture-of-Experts (MoE) models scale toward hundreds of experts and higher top- k routing, memory efficiency in distributed training becomes a critical bottleneck. Peak memory is dominated by the MoE block, not attention: every intermediate buffer in the MoE dispatch pipeline is individually scaled by top-k routing. The standard all-to-all dispatcher sends all routed tokens in a single collective step, requiring the full top- k -expanded buffer to be constructed at once. In this work, we propose RelayMoE, a ring-based MoE execution model that computes locally as expert weights or tokens circulate, avoiding full top- k -expanded dispatch buffers. RelayMoE selects between expert and token routing according to communication volume and overlaps transfers with computation. The ring structure naturally supports memory-efficient MoE recomputation during backward: each hop reconstructs expert intermediates, uses them to compute gradients, and releases them before the next hop. The saved memory supports longer sequences and larger batches, or retains more attention activations to reduce attention recomputation and improve training throughput. We evaluate RelayMoE on 30B - 57B production MoE models and varied expert configurations. In single-layer MoE experiments, RelayMoE achieves a 2\times average speedup over Megatron-LM. In full-model training under the same GPU memory budget, it improves throughput by up to 2.02\times and extends the largest tested trainable sequence length by up to 2.85\times .
[AI-161] Scale-Invariant Training for Time Series Foundation Models
链接: https://arxiv.org/abs/2610.07324
作者: Ignacy Stepka,Willa Potosnak,Kin G. Olivares,Artur Dubrawski
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:
Abstract:Time series foundation models (TSFMs) are trained on large collections of time series datasets that span various morphologies and domains. This setting exposes models to series whose scales – typical magnitudes of their values – can differ substantially. Affine scaling methods such as Reversible Instance Normalization (ReVIN) scale model inputs and reverse the transform before computing the loss. We show that this inversion multiplies each series’ gradient by b^p relative to loss on scaled targets, where b is the scaling denominator (e.g., standard deviation) and p is the loss degree. We call this scale-contaminated training (ScaleCon), because the scale of each series consequently becomes an importance weight, causing high-scale series to dominate training. For any scale-equivariant scaler and residual loss that is homogeneous of degree p , including MSE, MAE, and Quantile Loss, we prove that computing loss on scaled targets makes every mini-batch gradient and, consequently, the full optimization trajectory invariant to arbitrary independent rescaling of the training series, yielding scale-invariant training (ScaleIn). Notably, existing TSFMs use both objectives, with neither consistent reporting nor a common convention on how to compute training loss. We isolate the convergence disparity induced by ScaleCon and its correction under ScaleIn in controlled studies on synthetic and real data. In pretraining across four TSFM architectures, ScaleIn lowers MASE in all 24 architecture-benchmark comparisons, with average reductions across TSFMs of 18.8% on GIFT-Eval and 21.9% on the M-competitions. The gains extend to supervised neural forecasting, where it lowers MASE in 16 of 20 matched settings. Most existing time series forecasting pipelines can adopt ScaleIn with a one-line code change.
[AI-162] Rule-Based Languages for Neurosymbolic AI
链接: https://arxiv.org/abs/2610.07313
作者: Stefania Dumbrava,Efthymia Tsamoura
类目: Artificial Intelligence (cs.AI); Logic in Computer Science (cs.LO)
备注: To appear in the Proceedings of Rules and Reasoning - 10th International Joint Conference, RuleML+RR 2026, Vilnius, Lithuania
Abstract:Logic programming is increasingly used as the symbolic component of neurosymbolic AI systems. We survey the main rule-based languages in this setting, namely Datalog, answer set, and probabilistic logic programs, along four axes: semantics, expressiveness, neural integration, and evaluation mechanism. We analyse over 50 recent systems and applications, comparing formalism usage across four research areas: databases and programming languages, machine learning, vision, and robotics. We provide a decision matrix mapping application scenarios to required features and close by outlining open problems.
[AI-163] Understanding and Mitigating Inference-Time Overreliance Using Agent ic Memory
链接: https://arxiv.org/abs/2610.07311
作者: Luoxi Tang,Yuqiao Meng,Nilesh Auradkar,Muchao Ye,Dazheng Zhang,Zhaohan Xi
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Agentic memory allows LLM agents to reuse past experience, yet retrieved memories can also distort inference even when they are benign, correctly stored, and appropriately retrieved. We study this failure mode, which we call memory over-reliance. Across benchmarks and memory architectures, we find that memory is useful when past experience transfers to the current task, but can become misleading when only part of the evidence transfers. Failures are strongest under partial query-memory overlap, a pattern further confirmed by controlled experiments thatvary the amount of overlapping evidence. Motivated by this finding, we propose MEMTRIM, a plug-and-play framework that indexes memory evidence at write time and controls its reuse at read time. MEMTRIM removes repeated or conflicting evidence while preserving useful memory-specific information, requires no retraining, and applies to both embedding-based and structured memory this http URL show that MEMTRIM reduces memory overreliance while preserving the benefits of useful memory across models and memory settings.
[AI-164] From Sandbox to Enforcement: Confidence-Qualified Threat Intelligence for Critical Infrastructure
链接: https://arxiv.org/abs/2610.07310
作者: Nikolaos Kekatos,Mihaela Curcă,Georgios Koutidis,Mihai Nena,Tom Nianios,Robert-Ştefan Şandru,Michael Ioannou,Charalambos Bratsas
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
备注: 20 pages, 3 figures. Accepted at the 21st International Conference on Critical Information Infrastructures Security (CRITIS 2026)
Abstract:Security operations centres and national incident-response teams defending critical infrastructure collect abundant threat data yet struggle to turn it into actionable intelligence. A malware sandbox produces detailed behavioural evidence, but as a large, unranked report whose confidence is unstated. We present CG-CTI, an operational pipeline that converts live sandbox output (CAPEv2) into STIX 2.1, correlates it in a knowledge graph with other critical-infrastructure sensors, and attaches to every intelligence object an explicit confidence status derived from provenance, cross-source corroboration, and observation durability. This status gates automated action: only corroborated intelligence is eligible for automated enforcement, while lower-confidence objects are routed to analyst review or kept as context. A grounded language-model stage then narrates the confidence-qualified evidence, where each statement either cites a supporting object or is marked unsupported, so fabricated references are removed before analyst review. We implement CG-CTI within the CYBERGUARD project, whose consortium includes Romania’s national cyber-security directorate, and evaluate it against the live sandbox on a labelled malware corpus, measuring conversion validity, indicator yield, technique coverage, corroboration, enforcement eligibility, latency, and summary grounding. CG-CTI turns fragmented sandbox output into corroborated, confidence-ranked, and auditable intelligence for critical-infrastructure defence.
[AI-165] Polar: LLM -Powered Synthesis of Real-World Cyber Evidence for Prioritization and Mitigation
链接: https://arxiv.org/abs/2610.07298
作者: Luoxi Tang,Yuqiao Meng,Ankita Patra,Weicheng Ma,Muchao Ye,Zhaohan Xi
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
备注:
Abstract:Cyber threat analysis increasingly depends on evidence distributed across vendor advisories, vulnerability databases, and threat intelligence sources. Turning these fragmented observations into timely decisions requires models to connect technical severity with evolving exploitation evidence and available defensive actions. We present POLAR, an LLM-powered framework for synthesizing real-world cyber evidence into threat-centric assessments for prioritization and mitigation. POLAR first disentangles overlapping incidents and grounds each threat in source-linked evidence. For prioritization, it infers severity metrics from cyber evidence and combines the resulting assessment with temporally ordered exploitation signals to estimate near-term exploitation likelihood. For mitigation, it links the synthesized threat data to authoritative remediation knowledge and organizes applicable actions according to threat urgency and operational constraints. We evaluate POLAR on real-world vulnerability evidence collected from public resources and compare it with multiple baselines. Across heterogeneous incidents and zero-day settings, POLAR improves threat ranking and mitigation retrieval while producing evidence-linked intermediate assessments that support analyst inspection. The results establish evidence synthesis as a practical foundation for LLM-based cyber decision support across related security tasks.
[AI-166] Catching Developers in the Flow: Low-Latency Agent ic Program Repair at Google Scale
链接: https://arxiv.org/abs/2610.07289
作者: Celal Ziftci,Spencer Greene,Ray Liu,Livio Dalloro,Lorenzo Dini
类目: oftware Engineering (cs.SE); Artificial Intelligence (cs.AI)
备注: Accepted at the 41st IEEE/ACM International Conference on Automated Software Engineering (ASE 2026)
Abstract:Manual repair of program failures is time-consuming and disruptive for software developers, particularly during the pre-submit phase where test failures occur within continuous integration systems. While Automated Program Repair has seen significant advancement through Large Language Models, existing state-of-the-art techniques primarily focus on post-submit workflows, operating offline without the low-latency requirements necessary to assist developers in real-time within their flow before they switch context. In this paper, we introduce FlowAgent, an AI agent deployed at Google to automatically repair test failures in the pre-submit outer-loop workflow inside continuous integration systems. Integrated into Google’s internal developer tools, Critique and Cider,FlowAgent utilizes a ReAct-style generate-and-validate loop, as well as rigorous pre-execution and post-execution abstention filters to ensure high-quality suggestions under strict latency constraints. Based on our case studies, FlowAgent is highly effective. First, a manual evaluation conducted on 195 real-world test failures demonstrated 67.18% accuracy in suggesting correct fixes. Following its Google-wide deployment, FlowAgent suggested fixes on 295,508changes, of which developers previewed 65,069 and applied 28,554. Developer feedback from interviews indicate that the agent is useful in suggesting correct fixes, integration of autonomous repair agents into industrial software engineering workflows is received well, while interesting challenges and opportunities still remain. Comments: Accepted at the 41st IEEE/ACM International Conference on Automated Software Engineering (ASE 2026) Subjects: Software Engineering (cs.SE); Artificial Intelligence (cs.AI) Cite as: arXiv:2610.07289 [cs.SE] (or arXiv:2610.07289v1 [cs.SE] for this version) https://doi.org/10.48550/arXiv.2610.07289 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[AI-167] FlexiFlow: Bandit-based Model Switching in ML Workflows
链接: https://arxiv.org/abs/2610.07286
作者: Abhilash Jindal,Todd Nief,Bhanu Prakash Vangala,Shankaradithyaa V,Tvisha Malik,Anshik Sahu,Aaron Schein,Amitabh Chaudhary,Tanu Malik
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:
Abstract:Model optimizations help improve inference performance and accuracy of ML workflows. However, relying on a single model to perform inference across all data batches often fails to maximize accuracy and thus overall performance. In many cases, alternate models could perform better on specific subsets of data where a primary model underperforms. Our experiments with real ML workflows indeed show that switching models improves workflow accuracy by up to 23%. Yet, current systems lack the ability to adaptively switch between models based on performance, forcing users to manually test models in sequence. We present FlexiFlow, a dataflow system that dynamically switches between alternate models when the current model exhibits low accuracy. FlexiFlow learns to rank models using a novel multi-armed bandit approach that accounts for model runtimes, probability of passing user-defined assertions, and the computational structure of the ML workflow. We show that the standard Thompson sampling approach is insufficient for switching models in ML workflows. In contrast, our proposed approaches are effective and scales to complex real-world ML workflows. Experiments show that switching models at runtime while reusing intermediate results provides higher accuracy, but also 48% efficiency gain compared to sequential workflow runs.
[AI-168] SAFESHIELD: A Decision-Organization Framework for Deployment-Time Safety of Small Language Models
链接: https://arxiv.org/abs/2610.07276
作者: Xingru Zhou,Luis Sentis,Aarti Choudhary
类目: oftware Engineering (cs.SE); Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR)
备注: Accepted to the Application Track of IEEE TPS 2026. 12 pages
Abstract:Deployment-time safety of language models is commonly implemented through runtime guardrails such as input moderation, routing, retrieval verification, and output filtering. Existing deployment frameworks provide increasingly capable mechanisms for these functions, but offer limited guidance on how the safety decisions they produce should be explicitly organized, coordinated, and audited. We formulate deployment-time safety as a decision-organization problem with two elements: responsibility-oriented decomposition of safety decisions and explicit coordination among them. We instantiate this formulation in SAFESHIELD, a deployment-time safety system for small language models that organizes four recurring decision responsibilities (admission, routing, evidence, and release) and records committed decisions in auditable Decision Traces. We evaluate SAFESHIELD through mechanism-level experiments, aggregate stage ablations, controlled coordination ablations, and a deployment-oriented stress suite. Mechanism-level results show that the instantiated safeguards provide the capabilities required by the decision process, while aggregate ablations show substantial degradation in end-to-end safety as the surrounding safety organization is removed. More importantly, dedicated coordination ablations preserve the participating safeguard mechanisms while selectively severing their dependencies: removing admission gating substantially increases false release, and withholding upstream evidence from the release decision reduces release accuracy from 96.0% to 69.5%. These results provide system-level evidence that deployment-time safety depends not only on the capability of individual guardrails, but also on how their decisions are organized and coordinated.
[AI-169] A Trust Layer for Agent Evaluation
链接: https://arxiv.org/abs/2610.07274
作者: Mohammadreza Sediqin,Shivali Dalmia,Srinivasa Karthikeya Reddy Kovvuri,Abhishek Mukherji
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Deterministic benchmark scores show that an agent received credit, but not whether that credit was earned, reported honestly, or would hold on a second run. We introduce a Trust Layer for Agent Evaluation, an additive post-hoc framework that reports, beside each recorded score, whether it should be believed. It verifies four properties: whether the result is supported by the benchmark’s own grading logic, whether a passing answer was earned through traceable computation, whether the agent’s completion claim matches what occurred, and whether the result is stable under repeated execution. The first three use only saved artifacts; the fourth re-runs the agent. Model judgments only label evidence under majority voting; all verdicts follow deterministic rules and never modify the recorded score. Applied to five agent configurations on 108 tasks from Agents’ Last Exam, every model shows passing runs with no traceable computation (at rates varying tenfold), confirmed false completion claims, and unstable results: 18-46% of tasks do not stay in one score band over five runs. Only 22.6% of recorded passes clear all four checks (95% CI 15.0-32.6, n=84). Measuring what an agent can do and verifying that it did it are different problems, and current benchmarks address only the first.
[AI-170] Does the Model Use the Feature? Separating Steering from Mechanism in LLM s
链接: https://arxiv.org/abs/2610.07270
作者: Tong Che,Yilong Li
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Internal features in LLMs are often interpreted as mechanisms when they track a concept and their manipulation changes a related behavior. Yet steering can push a feature far outside its natural range, where its effects need not reflect the model’s own computation. We examine this inference and propose an empirical contract whose tests evaluate features at values observed on natural inputs. One test copies a feature’s value from an input that shows a behavior into a matched input that does not (installation) or the reverse (removal); the other restores the feature after an upstream edit (downstream rescue). Installation measures how far the feature suffices for the behavior; removal and downstream rescue measure how much the model uses it. Applied to three kinds of representations, the two strengths separate sharply. The published unknown-entity latent strongly steers knowledge abstention, yet installing observed values from either published latent into matched prompts transfers only a small fraction of the natural known–unknown abstention contrast. Dense known–unknown directions show opposite asymmetries between installation and removal in Gemma and Llama, and how fully a released subject–verb agreement feature set reproduces and restores the behavior depends on how its values are written into the model. Tracking a concept and steering a behavior therefore do not by themselves show that the model uses a feature, and each conclusion holds only for the intervention tested.
[AI-171] Verifying Coordination in Parallel Coding Agents : NP-Bench and a Scheduling Planner
链接: https://arxiv.org/abs/2610.07261
作者: Sumanyu Muku
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:A team of coding agents can look fine agent by agent yet fail as a team: each passes its own tests while the merged result is broken, and single-agent evaluation never catches it. As teams run several LLM coding agents in parallel on one codebase, the agents collide: two rewrite the same function, one codes against a contract a teammate just changed, and integration fails after the work is done. Most coordination tools react (watch for a conflict, then warn), but at agent speed the warning arrives after the wasted edit. We recast the problem as scheduling: take each work item’s declared scope, partition the work into disjoint scopes, and order merges along the producer-consumer graph, all up front. We build this planner into Nerveplane and evaluate it with NP-Bench, an environment-grounded three-arm benchmark (no coordination; reactive detection; proactive planning) that verifies integration off a real git merge, both in a deterministic simulation and with live agents. The planner lifts clean-integration from 1/9 to 9/9 scenarios and cuts merge conflicts from 13 to 0, with a gap that grows in the number of agents. On a live breaking contract change it rescues an outcome both baselines miss on every seed: the clean-integration rate rises from 0 (no coordination and reactive detection) to 1.0 on a frontier model and 0.6 on a small one, while agents respect assigned scopes (0/5 leakage). A cross-session memory drops the repeated-mistake rate from 1.00 to 0.00 on strong and weak models alike. We also report a negative result: routing facts to agents does not rescue long-context accuracy at window-fitting scales; its value is cost and capacity, not attention. Across two capability tiers and two vendors, the benefit did not shrink as models got stronger, because it comes from how work is allocated, not model reasoning.
[AI-172] MemMux: Runtime Verification and Honest Resource Attribution for Fleets of Parallel Coding Agents
链接: https://arxiv.org/abs/2610.07257
作者: Sumanyu Muku
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Developers increasingly run a fleet of coding agents side by side on one workstation. The tools they reach for, terminal multiplexers like tmux and a new generation of agent managers, were built to arrange windows, not to govern memory. When ten agents each spawn language servers, test runners, and browsers, no standard tool can say how much memory belongs to which agent, confirm that a terminated agent’s descendants are gone, notice a child that has escaped its agent, or keep the machine off the swap cliff when an OOM kill would silently discard uncommitted work. We treat these as runtime-verification problems: an agent-hosting substrate should continuously emit observable signals an operator or auditor can check while agents run. We present MemMux, a local runtime that turns resource governance into checkable signals (per-agent attribution, complete reclamation, escaped-process visibility, bounded footprint under overcommit, and monitoring overhead), with a claims-disciplined benchmark against tmux, a purpose-built agent multiplexer, and a raw-process baseline on identical workloads. Under a binding memory budget on a Linux host, MemMux keeps the fleet under budget (7.5 GiB) with zero swap by admitting a subset and reclaiming under pressure, while the ungoverned tools run every agent, pin the machine at its RAM ceiling (2x over budget), and spill about 2 GiB into swap. MemMux reclaims 100% of a terminated agent’s process subtree where the raw baseline strands half of it, and it alone surfaces escaped children (10 of 10 detected). We report the cost: the 1 Hz attribution scan runs near 0.6% CPU at one agent but 2.7% at ten, above our 2% target. Running the harness on real Claude Code sessions shows 100% attribution and low overhead carry over to live agent trees. We release the engine, benchmark, and a one-command reproducer.
[AI-173] Internalizing Agent Experience into Diffusion Model Weights via On-Policy Context Distillation
链接: https://arxiv.org/abs/2610.07250
作者: Wenxuan Wang,Zekai Liu,Weinan Zhang,Yu Cheng,Yang Yang
类目: Artificial Intelligence (cs.AI)
备注: 27 pages, 11 figures
Abstract:Wrapping an image generation model in an agentic harness can effectively boost Text-to-Image task performance: the harness can leverage memory, skills, workflow orchestration, result verification, and iterative refinement to continually construct and revise prompts, thereby eliciting better images. These gains, however, remain external to the diffusion model and are realized only while the full harness runs. We propose Diffusion On-Policy Context Distillation (D-OPCD), which treats the agent-improved prompt as privileged context and distills the knowledge encoded in the agent harness into the weights of the diffusion model, so that the model retains part of the harness’s benefit when conditioned on the original query alone. Using a Text-to-Image agent equipped with our proposed Auto Skill Evolver (ASE), we show that D-OPCD can internalize harness capabilities into the generator’s weights, raising the average direct-generation score from 60.52 to 65.09 across four benchmarks. With this knowledge absorbed into the weights, the harness can shed its saturated skills and resume evolving: a second ASE round on the updated generator improves on a skill-free harness by additional 1.83 points, pointing toward text-to-image systems in which harness and model keep improving each other through continual co-evolution.
[AI-174] Can Semantic Geometry Teach an AI Judgement?
链接: https://arxiv.org/abs/2610.07249
作者: Thomson D. Nguy(Radiant Institute for Manifold Studies)
类目: Artificial Intelligence (cs.AI)
备注: 16 pages, 6 figures. Four bounded studies of semantic measurements for pre-action judgment; consequence-graph hypothesis remains untested
Abstract:How can an AI agent determine what rules to follow? One rule permits an action. Another imposes a condition, exception, or conflicting obligation. Deterministic systems can resolve those relationships when they have been specified. When they remain implicit in language, an agent can follow one rule while missing another that should stop it. Refusing every unresolved action avoids that risk, but also blocks permissible actions. We wanted the agent to make the distinction and still act. Our initial hypothesis was that geometric measurements could supply a basis for judgment. We represented actions and policies as vectors, then tested whether their geometry could identify governing policies and interpret the action’s relation to them. Across four studies, the tested approaches did not establish reliable pre-action judgment. In the final synthetic study, a lexical router recovered every governing and blocking policy while reducing median policy checks by 97.7%. The composed pipeline nevertheless escalated all 2,304 test actions, including those it should have allowed. Supplying every policy to the same downstream mechanism changed no decision. Finding the policies had not solved the problem of interpreting them. This result led us to revise our hypothesis: judgment in AI agents requires developing a consequence graph. Such a graph would connect the actor and authority to policy conditions, exceptions, and the changes an action would produce. Follow-on studies will ask whether making those relationships explicit helps the agent distinguish when to act, stop, or seek review. Comments: 16 pages, 6 figures. Four bounded studies of semantic measurements for pre-action judgment; consequence-graph hypothesis remains untested Subjects: Artificial Intelligence (cs.AI) Cite as: arXiv:2610.07249 [cs.AI] (or arXiv:2610.07249v1 [cs.AI] for this version) https://doi.org/10.48550/arXiv.2610.07249 Focus to learn more arXiv-issued DOI via DataCite (pending registration) Submission history From: Thomson Nguy [view email] [v1] Mon, 5 Oct 2026 18:49:55 UTC (1,905 KB)
[AI-175] SPECTRUM: Proximal Spectral Modulation for Looped Self-Distillation
链接: https://arxiv.org/abs/2610.07237
作者: Yunbo Long,WenJie Chen,Jiaquan Zhang,Guangya Hao,Zihang Zeng,Pengze Li,Xi Chen
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:A model that learns from its own outputs inherits more than their correctness: it inherits which solutions it produces. We formulate Looped Self-Distillation, a self-evolution framework for code generation in which a model repeatedly generates and learns from its own raw outputs, under a fixed information budget, without ongoing external assessment or test-based selection of the generated samples. We identify a consequential separation: correctness can improve while the breadth of correct implementations contracts. We introduce SPECTRUM, which re-estimates loss-sensitive key/value geometry from a fixed reference anchor at each round and converts it into full-rank proximal spectral modulation. All generated completions train a single student, whose subsequent inference requires no intervention. After five rounds of experiments on MBPP, SPECTRUM retains 89.9% of the initial model’s 64-sample correct AST richness, compared with 66.4% for Vanilla self-distillation and 65.5% for a subspace-projection control. The advantage persists at matched correct-sample counts. Without further training or recalibration, the resulting student also achieves higher matched-correct richness than Vanilla SD on HumanEval+ and APPS Intro, demonstrating transfer of the diversity benefit. These findings establish correct-solution retention as a complementary objective of recursive self-improvement (RSI) and show that generation-time intervention can improve the solution repertoire retained by subsequent students.
[AI-176] Cascadia: Resident 975B MoE Inference on Eleven AI PCs
链接: https://arxiv.org/abs/2610.07219
作者: Tate Berenbaum(Not Community Labs Inc.),Matias Parij(Not Community Labs Inc.),Muthaiah Venkatachalam(Intel Corporation)
类目: Artificial Intelligence (cs.AI); Distributed, Parallel, and Cluster Computing (cs.DC)
备注:
Abstract:Mixture-of-experts models make nearly trillion-parameter capacity accessible with sparse per-token computation, provided that the serving system can distribute the weights and coordinate their execution. We present Cascadia’s resident execution of Inkling, a 975B-total/41B-active-parameter model, on eleven Intel Core Ultra X7 358H AI PCs, each with 64 GB of memory, Arc B390 integrated graphics and gigabit Ethernet. We contribute a custom resident MoE engine that preserves Inkling’s routing rules, constructs compressed graphs for OpenVINO’s fused iGPU primitives, and coordinates FP16 expert computation with FP32 output restoration. The engine fits six consecutive decoder layers per machine and represents dense feed-forward blocks as all-active expert slices, reducing measured dense-layer call time from approximately 8.1 to 4.5 ms. A streaming pipeline coordinates concurrent generation, while captured-state draft evaluation measures agreement with the deployed numerical path. Paired measurements at fifteen concurrency levels from 1 to 176 streams reach 60.29 aggregate decode tokens/s at 88 streams, with 46.87 tokens/s over the complete serving phases. At fifteen streams, median first-token latency is 6.05 s. Raising the context budget from the 1,024-position default, real prompts of 1k to 64k tokens recover the embedded code in all 19 measured answers, with first-token time growing as aN+bN^2 and decode latency growing approximately linearly, both bounded by a single-threaded CPU attention loop rather than by memory, which holds 512k positions per stream. Evaluation on captured fleet states separates the effects of vocabulary selection and weight quantization on draft agreement. Together, these contributions establish an execution and evaluation approach for large sparse models on distributed client systems with shared CPU-GPU memory.
[AI-177] Distributionally Robust Mixture-of-Experts Training NEURIPS2026
链接: https://arxiv.org/abs/2610.07207
作者: Xin Teng,Muxiao Li,Hongyi Wen
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: In proceedings of NeurIPS 2026
Abstract:Mixture-of-Experts (MoE) transformers scale capacity by activating only a few experts per token, but this sparsity creates a hidden reliability problem: when routing is imperfect, load-balanced models may send tokens to experts that are insufficiently trained for the assigned inputs. We propose Distributionally Robust MoE Training (DRMoET), a drop-in objective that treats layer-wise experts as endogenous robustness groups and optimizes high-loss routing outcomes rather than merely equalizing traffic. DRMoET updates a per-layer expert distribution by an entropy-regularized softmax rule on EMA-smoothed, activation-weighted expert losses, strengthening plausible non-top routing paths while preserving standard MoE computation. Under the FLAME-MoE recipe at 746M-total and 10.3B-total scales, DRMoET improves downstream averages over both standard FLAME-MoE and auxiliary-loss-free balancing. At 10.3B total parameters and 67B training tokens, DRMoET improves the seven-task average from 0.6625 to 0.6767, while the auxiliary-loss-free baseline achieves 0.6431. Mechanistic analyses show lower expert-loss variance with nearly unchanged mean loss, 4.3% lower excess loss under forced mid- k misrouting, and improved domain-expert specialization. These results position routing robustness-not only utilization balance-as a practical objective for reliable sparse MoE scaling. Project page and code are available at: this https URL.
[AI-178] Sim-to-Real Transfer of Vision-Language Navigation in Continuous Environments Using an Ackermann-Steered Mobile Robot
链接: https://arxiv.org/abs/2610.07192
作者: Chalindu Abeywansa,Sahan Gunasekara,Devindi De Silva,Seniru Dissanayake,Ranga Rodrigo,Peshala Jayasekara
类目: Artificial Intelligence (cs.AI); Robotics (cs.RO)
备注:
Abstract:Vision-Language Navigation (VLN) enables robots to navigate through environments using natural language instructions, making human-robot interaction intuitive. Traditional VLN models often rely on navigation graphs, 360-degree views, and perfect localization which pose significant challenges when adapting these models to real-world settings. This work addresses these limitations by performing a simulation-to-real domain shift of a VLN approach that operates in continuous environments without requiring navigation graphs or panoramic views. The proposed system integrates vision-language models that align visual inputs and linguistic instructions within a shared embedding space, facilitating natural language-driven navigation. We employ a Cross-Modal Attention (CMA) based architecture trained on an existing dataset in a simulated environment and fine-tune it using real-world data collected from a custom-built Ackermann-steered robot equipped with a camera and a LiDAR sensor. By utilising linear photometric adjustments and fine-tuning on a limited number of episodes, our model successfully adapts to real-world environments, achieving effective navigation while running offline on dedicated hardware. Experimental results, evaluated using Success weighted by Path Length (SPL) and Normalized Dynamic Time Warping (nDTW) metrics, demonstrate the robustness and adaptability of our approach. Keywords: Vision-Language Navigation, Cross-Modal Attention, Natural Language Instructions, Sim-to-Real Transfer, Autonomous Navigation, Ackermann-steering.
[AI-179] Agent ic Design Space Exploration for Joint Hardware Configuration Selection and Mapping of AI Inference Workloads on Heterogeneous Edge SoCs
链接: https://arxiv.org/abs/2610.07191
作者: Geetha Prasuna Yarramneni,Surya Selvam,Wilfried Haensch,Anand Raghunathan
类目: Hardware Architecture (cs.AR); Artificial Intelligence (cs.AI)
备注: Published at the 2026 ACM/IEEE International Symposium on Machine Learning for CAD (MLCAD '26), Jeju Island, Republic of Korea, September 7-9, 2026
Abstract:Modern edge Systems-on-Chip (SoCs) integrate heterogeneous processing units (PUs) such as CPUs, GPUs, and NPUs, each with distinct performance and energy characteristics. Deploying AI inference workloads on them under real-time latency and energy constraints requires jointly mapping workloads to PUs and configuring each PU (e.g., selecting the number of active cores and the operating frequency). This joint space grows combinatorially, making exhaustive search infeasible. Most prior work on design space exploration (DSE) applies black-box optimization (BBO) such as evolutionary search, where each evaluation returns only aggregate metrics such as latency and energy. Recent LLM-guided DSE relies on the same sparse feedback. We observe that this limits its efficiency: it offers no insight into the design space or the reasons a design choice performs the way it does, and it leaves the reasoning abilities of LLMs largely unused. We present TraceDSE, an agentic DSE flow that performs joint workload mapping and PU configuration selection for AI inference on heterogeneous SoCs. TraceDSE is an iterative proposer-critic loop driven by richer feedback in the form of system execution traces. The LLM proposer agent generates candidate mappings and PU configurations for hardware evaluation. The LLM critic agent, equipped with programmatic trace-analysis tools, analyzes the traces to identify bottlenecks and suggest targeted refinements. This loop yields deeper insight into each design point, higher-quality decisions, and a more effective search. Across four AI inference workloads (models of varying complexity and a multi-model pipeline) on an Intel Meteor Lake SoC, TraceDSE consistently outperforms two state-of-the-art BBO tools, improving Pareto frontier hypervolume by up to 35% over NSGA-II and up to 68% over Bayesian optimization, while requiring ~6-9x fewer hardware evaluations.
[AI-180] Is this machine playing?
链接: https://arxiv.org/abs/2610.07130
作者: Nathan Cloos,Antonio Norelli,Daniel Durbin,Jacob Andreas,Daniela Rus,Phillip Isola
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: 13 pages of main text, 17 figures
Abstract:We placed a modern AI coding assistant in an unintended role: as the mind of a body on an unknown digital island. With only a minimal instruction mentioning no specific task, reward, or activity, the machine started animating its virtual body. Across thirty-hour runs, the embodied AI agent climbed hills, stacked blocks into towers, drew mandalas, reinterpreted sports, ran experiments on the physics of its world, and learned techniques that later expanded what it could accomplish. These activities recurred across thirteen agents but diverged into distinct histories. We examine whether this behavior satisfies classical criteria for play and ask whether play can become a mode of machine development.
[AI-181] Aggregating User Preferences while Ensuring Equity Diversity and Inclusion using Graph Summarization
链接: https://arxiv.org/abs/2610.07128
作者: Adji Marieme Sita Cissé,Malek Mouhoub
类目: ocial and Information Networks (cs.SI); Artificial Intelligence (cs.AI)
备注: 22 pages, 3 figures. Interactive dashboard: this https URL
Abstract:Aggregating the preferences of diverse user groups into a collective outcome raises fundamental challenges of equity, diversity, and inclusion (EDI): classical aggregation rules such as Borda and Condorcet have no mechanism to prevent results from systematically favoring majority groups, collapsing onto homogeneous items, or under-representing minorities. We address this problem through EDI-constrained graph summarization. User preferences are modeled as a weighted attributed bipartite graph, and a greedy coarsening algorithm iteratively merges user nodes while enforcing three structural EDI criteria: an equity gap constraint ( \Delta E ), an intra-list diversity constraint (ILD), and a group inclusion constraint. Rather than correcting fairness after aggregation, our method embeds EDI preservation directly into the graph structure. We evaluate across five datasets spanning four domains: MovieLens 100k and 1M, this http URL, Rate My Professors, and OpenAlex (2018-2023). Our method, AURORA, achieves the largest and most consistent diversity gains over classical voting rules, and on MovieLens 100k at k = 20 it simultaneously improves all three EDI criteria over both Borda and Condorcet. On Rate My Professors, it combines high diversity (ILD = 0.808) with the highest female item representation (60%), at a moderate equity cost, and it achieves the lowest equity gap ( \Delta E = 0.031 ) on OpenAlex, where Borda-based methods recommend zero female authors. On this http URL, the only dataset where the sensitive attribute is present on both sides of the bipartite graph, our method does not reduce the equity gap, a limitation we connect to prior findings that demographic parity is not always an appropriate target. These results demonstrate that embedding EDI constraints into aggregation structure yields more robust fairness-diversity trade-offs than post-hoc approaches.
[AI-182] AMBER: Training Long-Horizon Web Agents through Append-Only Memory
链接: https://arxiv.org/abs/2610.07118
作者: Chinmay Savadikar,Zhaoyu Zhang,Mingyu Zhao,Shuang Xie,Han Li,Tianfu Wu,Lingyun Wang
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: 29 pages, 11 figures
Abstract:Modern language-model agents increasingly interact with external environments over long-horizon, multi-step trajectories, where the accumulated interaction history can quickly exceed practical context budgets. To ensure reliability, agents must maintain factual information over long horizons, remember execution errors and corrective feedback, and track progress across actions. Several approaches have been proposed to achieve this without the need for maintaining the entire execution history in context, such as using the reasoning and action history, learning to maintain a fixed-size memory through an overwrite mechanism, and periodic summarization. Although overwrite memory can in principle retain anything an append-only memory can, it must learn to carry each fact through every subsequent rewrite, which is difficult to learn from sparse outcome rewards; for interactive applications like web agents, we find that trained overwrite memories delete key information required by the trajectory, as well as corrective feedback received from the environment. We introduce AMBER (Append-only Memory Bank for Evidence Retention) - a simple and scalable framework where an agent jointly learns to reason, act, and write free-form memory, while an append-only rule guarantees retention by construction. This allows AMBER to be trained end-to-end with reinforcement learning from outcome rewards without the need for extensive curated SFT data. On WebArena Lite, AMBER improves average success over overwrite-based memory by 4.09 percentage points, increases the fraction of tasks solved in five repeated runs by 4.8 percentage points, and matches an overwrite baseline trained on substantially more expensive curated supervision. AMBER achieves these improvements while maintaining a practical token budget, providing a strong balance between context efficiency, task performance, and reliable long-horizon execution.
[AI-183] Will the Judge Flip? Predicting Position-Sensitive LLM Judgments from Residual Stream Activations NEURIPS
链接: https://arxiv.org/abs/2610.07115
作者: Hashmath Shaik,Gnaneswar Villuri,Alex Doboli
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: It was submitted to neurips workshop (JUDGE workshop) and received a review score 6.5 combined with one strong accept and one above threshold accept
Abstract:The order in which candidate responses are presented can change an LLM judge’s verdict. Detecting such a position flip ordinarily requires judging each pair in both orders, which doubles the number of judgments. We investigate whether residual stream activations recorded immediately before the initial verdict can predict a flip. We use nested grouped cross-validation to evaluate regularized linear probes on 534 JudgeBench pairs for three Qwen3 judges and Llama-3.1-8B. The linear probes achieve AUROCs of .621-.850 and outperform a combined baseline that uses verbalized confidence, verdict-label logits, response lengths, and the judge’s initial choice by .062-.113 AUROC. Linear probes trained on JudgeBench and then frozen achieve AUROCs of .685-.853 on 1,802 MT-Bench comparisons without MT-Bench fitting or recalibration. These results show that pre-verdict activations support prediction of susceptibility to candidate order and outperform the non-activation predictors evaluated here.
[AI-184] LiLib: Lifelong Air-to-Ground Path-Loss Prediction on UAVs via a Drift-Triggered Model Library
链接: https://arxiv.org/abs/2610.07111
作者: Minh Tran
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:
Abstract:UAVs that act as relays or base stations need accurate air-to-ground path-loss predictions for rate adaptation and placement, but propagation conditions change as a UAV moves between suburban, urban and high-rise areas, and the same areas are often revisited. Online regressors that adapt by forgetting must relearn each environment from scratch, whereas a single model trained on all data averages incompatible regimes. We propose LiLib, a lightweight continual-learning scheme in which a UAV maintains a small library of recursive-least-squares experts. A windowed residual test detects drift; a short probe phase then either reuses the best stored expert or creates a new one. In simulations based on four standard urbanization profiles, LiLib reduces prediction RMSE from 5.89 dB (best sliding-window baseline) to 4.03 dB (p 0.001), lowers the error shortly after a return to a known environment from 12.3 dB to 5.7 dB, and recovers 99% of the throughput of a regime-aware oracle in rate adaptation. The library stores four experts in under 0.5 KB, and identifies regimes with 92% purity without labels. When a second UAV is initialized with the library of a peer, its error after environment changes halves. LiLib does not reach the oracle, and similar regimes may be merged when shadowing is strong. The results indicate that, for recurring drift, remembering is more effective than re-adapting.
[AI-185] Muon Is Theoretically Wrong For Convolutions But Empirically Effective
链接: https://arxiv.org/abs/2610.07103
作者: Thibaut Boissin(IRIT),Thomas Massena(IRIT, DTIPG - SNCF, UT3),Mathieu Serrurier(IRIT),Franck Mamalet
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:
Abstract:Muon, an optimizer known for its efficiency, has a clear interpretation for matrix-valued updates, but convolutional kernels are stored as four-dimensional tensors. Standard implementations reshape these tensors into matrices, a shortcut which breaks the theoretical understanding behind Muon. To investigate this, we formalize the corresponding optimization objective directly in convolutional operator geometry and introduce Convolutional Newton-Schulz (Conv-NS), which approximates the polar factor in this geometry while preserving kernel support. When applied in fast training experiments, Conv-NS and reshape-based Muon are both computationally efficient and achieve comparable accuracy on CIFAR-10 and ImageNet classification tasks. However, as one could expect a theoretically aligned Conv-NS to outperform reshape-based Muon, we investigate this mismatch between practice and theoretical understanding, with the hypothesis that exact convolutional orthogonalization may overconstrain updates. These findings highlight Muon’s strong practical performance while opening directions for its further development on convolutions. Our code is publicly available at \hrefthis https URLgithub conv-muon.
[AI-186] -CCL: Resource Efficient and Performant Collective Communication using Tensor Memory Accelerator
链接: https://arxiv.org/abs/2610.07098
作者: Keyvan Dadashzadeh,Yuehong Zhou,Minyu Cui,Miquel Pericas
类目: Distributed, Parallel, and Cluster Computing (cs.DC); Artificial Intelligence (cs.AI)
备注: Workshops on Supercomputing (SC’26)
Abstract:Large transformer-based models increasingly depend on multi-GPU execution, which requires frequent collective communication among GPUs. Existing communication libraries often rely on many GPU threads to achieve high bandwidth or low latency, resulting in a large streaming multiprocessor (SM)-side resource footprint. This footprint can limit the resources available to other GPU work, particularly when communication and computation execute concurrently. Thus, efficient collective communication should not only achieve high collective performance but also reduce its SM-side resource usage. This paper presents T-CCL, a resource-efficient collective communication library based on the Tensor Memory Accelerator (TMA) for intra-node communication. T-CCL offloads both data movement and reduction operations to TMA and executes each collective as a pipelined series of asynchronous TMA operations, reducing the SM resources required for collective communication while maintaining high bandwidth. Evaluated across AllReduce, AllGather, and ReduceScatter collectives, T-CCL outperforms NCCL by up to 2.4x with unrestricted communication resources and up to 3.42x under restricted resource budgets, remains competitive with NCCL’s recent symmetric-memory kernels, and occupies the same or fewer SMs in profiled cases. In a GEMM-collective overlap case study, switching the communication backend from NCCL to T-CCL raises the average operator-level speedup over a sequential baseline from 1.12x to 1.25x on two GPUs and from 1.04x to 1.14x on four GPUs, as T-CCL uses fewer SMs for communication, leaving more SMs available to the overlapped GEMM. Integrated into vLLM as a communication backend, T-CCL improves end-to-end inference throughput over vLLM’s automatic backend dispatch by up to 1.31x, outperforming it at every evaluated batch size on both the conversation and decode-heavy workloads.
[AI-187] Verified not generated: expert-verified AI study materials and the distribution of learning gains in a university course
链接: https://arxiv.org/abs/2610.07097
作者: Canh Thien Dang,An Nguyen
类目: Artificial Intelligence (cs.AI); General Economics (econ.GN)
备注:
Abstract:Experimental studies of generative AI in education mostly report average effects, yet field evidence shows that AI can narrow attainment gaps or widen them. We argue that the direction depends on the judgement burden, the expertise a learner must supply to screen AI output before learning from it, and that expert verification before release moves this burden from students to an accountable tutor. We test the argument in a two-cohort difference-in-differences design in which one half of a compulsory firstyear university economics course received AI-generated podcasts, FAQs and quiz-based study guides, produced with a source-grounded model and checked by a named graduate teaching assistant (170 students; 340 examination marks). Access was associated with a 2.34-mark advantage on a 50-mark component. The share of marks below the upper-second classification boundary fell by 24.7 percentage points relative to the counterfactual, effects were significant at every threshold from 23 to 31 marks and at none above, and roughly three-quarters of the average originated in the bottom quintile. The threshold estimate is robust to removing the lowest-scoring students from the pre-intervention cohort; the average effect is not. Interviews and feedback from 36 students indicate that the verification label gave students a reason to engage with AI-generated material without ending their scrutiny of it. Evaluations of AI learning resources that report only mean effects cannot detect whether the students the resources are meant to help are the ones who gain.
[AI-188] Small Language Models for Smart Data Model Classification at the Edge: A Cost-Aware Hybrid Approach
链接: https://arxiv.org/abs/2610.07093
作者: Cristian Martella,Angelo Martella,Antonella Longo,Motaz Saad
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:The rapid proliferation of heterogeneous data sources within the Internet of Things (IoT) across domains such as smart cities, energy management, and environmental monitoring necessitates efficient and scalable data standardization methods. Effective classification of smart data models (SDMs) is essential for facilitating interoperability. However, existing approaches are often limited by high resource consumption and lack applicability in edge environments with constrained computational capabilities. Aiming to bridge this gap, the proposed study evaluates the performance of lightweight open-source language models (LMs) to resolve an input data entity against its corresponding best fitting SDM representation under resource-constrained conditions. It systematically benchmarks a diverse array of models, including general purpose (GP), reasoning-specialized (RS), and code-specialized (CS) architectures, across multiple domain-specific datasets. Addressing the current omission of lightweight, resource-efficient solutions in the literature, the investigation provides significant and valuable insights into model selection, task formulation, and deployment strategies that optimize accuracy and efficiency. A complementary experiment also compares the surveyed large language models (LLMs) against two near-zero-cost similarity baselines (Term Frequency-Inverse Document Frequency (TF-IDF) and a lightweight sentence encoder) on the same task, providing a strong reference point for interpreting the practical value of LLM-based classification on edge platforms.
[AI-189] owards a Unified Misuse Monitoring Benchmark
链接: https://arxiv.org/abs/2610.07089
作者: Aniruddh Pramod,James Oldfield,Adel Bibi
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: 50 pages, 13 figures, 17 tables, Code: this https URL
Abstract:LLM agents increasingly act in multi-actor environments, exposing them to misuse from multiple sources: decomposition attacks, where a harmful request is split into innocuous sub-requests, and prompt injection attacks, where a compromised tool delivers a malicious instruction. Existing evaluations treat these threats separately and ask whether a trajectory is harmful, rather than when it becomes harmful. We propose monitoring the agent’s responses, where its actions are externalised, and ask whether the first point where monitors identify harm lands within a harm window (from the agent’s first harmful commitment to goal execution). We develop a unified formalism for trace-level misuse monitoring and use it to construct a benchmark of ~6,200 conversation transcripts between a user, an LLM agent, and the external environment, spanning both threats in a shared schema, with a labelled harm window, corresponding benign controls, and matched instances of refusals to these requests. Across 17 monitor configurations, we find that our proposed action-framed monitors perform well on both threats under classical metrics (AUC: 0.95 and 0.99 respectively), while content-framed monitors collapse on injection attacks (AUC: 0.52). We also show that classical position-blind metrics paint an optimistic picture of monitor performance, since all monitors localise decomposition attacks poorly under the interval metric, which measures the ability to localise harm. Broadly, we illustrate the need for a unified study of misuse monitoring.
[AI-190] SchemaFill: Efficient LLM Tool Calling via Slot-Parallel Speculative Decoding
链接: https://arxiv.org/abs/2610.07086
作者: Zhi-Kai Chen,Song-Yan Li,De-Chuan Zhan,Han-Jia Ye
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:
Abstract:LLM agents interact with external systems by generating structured tool calls. Given a user request, conversational context, and a catalog of tool schemas, a tool-calling model must select tools and generate their arguments, potentially producing multiple calls in a single response. Standard autoregressive decoding generates these calls token by token, incurring substantial latency for requests involving multiple calls or many argument fields. The explicit argument structure offers opportunities for parallel generation, but later argument values may depend on preceding fields and calls, so independently generated values can differ from the target model’s output. We present SchemaFill, a framework for efficient LLM tool calling through slot-parallel speculative decoding. SchemaFill generates future slot values concurrently as candidates, without requiring advance knowledge of the actual call sequence or argument values. Candidates spanning multiple fields and calls are concatenated for verification by the target model under the actual output prefix. Only verified tokens are committed, and the target supplies corrections when candidates disagree. This applies target verification while exploiting parallelism across slots and calls. On Glaive and BFCL, SchemaFill achieves up to a 4.05 \times improvement in end-to-end throughput over autoregressive decoding. Code is available at this https URL.
[AI-191] Demo: Vision-Language Model-Guided Online Calibration of an Electromagnetic Digital Twin
链接: https://arxiv.org/abs/2610.07081
作者: Zerui Kang,Yishen Lim,Zhouyou Gu,Seungnyun Kim,Seung-Woo Ko,Tony Q.S. Quek,Jihong Park
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI)
备注:
Abstract:An electromagnetic (EM) digital twin gives mobile robots wireless situational awareness but depends on material conductivities that change with the environment. Online calibration faces initialization sensitivity and measurement travel costs. We demonstrate a vision-language model (VLM)-guided framework using a Unitree G1 robot and NVIDIA Sionna, with two VLM calls: material classification maps visible materials through ITU-R P.2040 to conductivity priors for Sionna’s gradient descent on accumulated received signal strength (RSS) measurements; waypoint planning selects the next measurement location online using residual RSS calibration error and image coverage. In a real indoor scenario, the framework achieves a normalized mean absolute conductivity error of 1.74\times10^-4 within 20 m of travel; random initialization never converges, while random waypoints require over twice the travel.
[AI-192] CuratorMAS: Automating Dataset Curation via Multi-Agent Orchestration
链接: https://arxiv.org/abs/2610.07075
作者: Yixin Zhang,Wenjie Feng
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:High-quality datasets are essential for reliable machine learning, but dataset curation remains costly and hard to generalize across domains. Existing methods typically rely on manually designed heuristics or model-dependent signals, limiting their applicability across tasks and user queries. To address these limitations and automate data curation, we propose \textbfCuratorMAS, a multi-agent collaboration framework that orchestrates agents to evaluate and curate high-quality datasets. To achieve the goal of flexible curation, CuratorMAS decomposes the complex curation process into five programmable execution stages and forms a parallelizable workflow. Specifically, CuratorMAS first performs dataset exploration to collect contextual information such as file structures and constraint cues, thereby developing a comprehensive understanding of the given task. In order to acquire up-to-date information, CuratorMAS retrieves domain knowledge from online sources to augment the evaluation process. Next, CuratorMAS derives the necessary evaluation criteria and computes the corresponding metrics. Based on these results, CuratorMAS executes filtering accordingly. Finally, an evolution module summarizes the evaluation outcomes and updates the relevant skills. Extensive and comprehensive experiments demonstrate that CuratorMAS significantly reduces the noise rate by up to 36.03 percentage points (pp) while also improving the F1 score of downstream models by up to 8.88 pp.
[AI-193] RIAGE: Direction-Aware Mismatch Stabilization of Native NVFP4 Reinforcement Learning
链接: https://arxiv.org/abs/2610.07043
作者: Zhen Li,Shuai Zhang,Yanggan Gu,Yiming Zhang,Yang Yu,Mingfa Feng,Congkai Xie,Shuang Yu,Junjie Lai,Hongxia Yang
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:
Abstract:Low-precision execution can substantially accelerate reinforcement learning (RL) for large language models, but discrepancies between learner and sampler execution can destabilize policy optimization. In this paper, we characterize the interaction between mismatch and the policy-gradient direction, distinguishing locally amplifying from contracting update contributions that mismatch magnitude alone cannot identify. In native NVFP4 runs, we observe an early imbalance between the two amplifying regions, favoring negative-advantage, negative-gap updates. Their tail tokens become concentrated in a small fraction of response segments before mismatch spreads globally. Motivated by these findings, we introduce TRIAGE, a direction-aware stabilization method that uses segment-level diagnosis to selectively rebalance policy-gradient updates and applies bounded repair to residual severe mismatch. TRIAGE modifies the optimization objective while retaining native NVFP4 weight-and activation 4-bit (W4A4) forward execution on both the sampler and learner. Experiments on Qwen3-4B and Qwen3-30B-A3B show stable optimization throughout the evaluated training horizon and achieve full precision level performance across five mathematical reasoning benchmarks, while native NVFP4 with TRIAGE provides up to 2.3x higher rollout throughput than BF16.
[AI-194] Inference-Time Projection for Physically Valid Biomolecular Diffusion Models
链接: https://arxiv.org/abs/2610.07037
作者: Qurat-ul-ain,Yee Whye Teh,Charlotte M. Deane,Matteo Cagiada
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:
Abstract:AlphaFold 3-style cofolding models predict biomolecular complexes with high structural accuracy, yet a large fraction of their outputs are physically invalid: chains overlap at interfaces, ligand bond lengths and angles are distorted, rings are non-planar, and stereocentres are inverted. Current approaches either steer the sampler with physics-informed potentials, which multiplies sampling cost and memory overhead making inference impossible on large complexes, or finetune the model, costing time and tying the fix to one architecture. We observe that, unlike structural accuracy, physical validity is fully verifiable at inference time from quantities the sampler already holds. We therefore treat physical validity as a constrained inference problem and introduce two closed-form projection operators applied to the diffusion model’s denoised clean-coordinate estimate, \hatx_0 : an inter-chain van der Waals projection that pushes apart the most severely clashing atom pairs, and a ligand distance-geometry projection that restores bond lengths, angles, internal contacts, planarity and chirality. Both operators are local, sparse and displacement-capped, require no network evaluations, gradients or importance sampling, and leave the denoiser and its weights untouched, so they can be dropped into any AF3-style sampler without retraining. Applied to two independently developed models, Boltz-2 and OpenFold-3, across five benchmarks (CASP15, CASP16, the PoseBusters monomer and complex sets, and the Boltz physical-validity test set), our method recovers perfect physical validity while preserving structural-accuracy and ligand-placement metrics. These gains are achieved with negligible runtime and memory overhead, providing a practical, model-agnostic route to physically valid all-atom structure prediction.
[AI-195] JIVEAdapter: A Multi-Task Additive Low-Rank Adapter via Joint and Individual Variation Explained (JIVE)
链接: https://arxiv.org/abs/2610.07036
作者: Sara Abdali,Pashmina Cameron
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:
Abstract:Parameter-efficient fine-tuning adapts pretrained models at a fraction of the cost of full fine-tuning, yet most low-rank adapters are single-task and represent each weight update multiplicatively, leaving no explicit account of what is shared across tasks and what is task-specific. We introduce JIVEAdapter, a multi-task “additive” low-rank adapter inspired by statistical Joint and Individual Variation Explained (JIVE). JIVEAdapter decomposes every weight update into a Joint structure shared across all tasks plus a per-task Individual structure, penalizes the Individual structures to be near-orthogonal to the Joint so shared and task-specific signal stay “interpretable” and separated, and allocates rank adaptively across a shared Joint pool and a per-task Individual pool. The Joint is learned once, jointly over a task group or incrementally, one task at a time, then frozen and reused as a prior for new tasks without retraining the shared part. On GLUE and SuperGLUE with DeBERTaV3-base, JIVEAdapter is competitive with strong single-task and multi-task low-rank baselines at a matched per-task effective rank, without extra modules such as MoE, and when a related held-in task exists its frozen Joint serves a held-out task by reusing that task’s Individual with only a cheap per-direction scale, otherwise training a small new one.
[AI-196] An Empirical Fault Vulnerability Exploration of ReRAM-based Process-in-Memory CNN Accelerators
链接: https://arxiv.org/abs/2610.07029
作者: Aniseh Dorostkar,Hamed Farbeh,Hamid R. Zarandi
类目: Hardware Architecture (cs.AR); Artificial Intelligence (cs.AI)
备注:
Abstract:Resistive random-access memory (ReRAM)-based Processing-in-Memory (PIM) accelerator is a promising platform for processing massively memory intensive matrix-vector multiplications of neural networks in parallel domain, due to its capability of analog computation, ultra-high density, near-zero leakage current, and non-volatility. Despite many advantages, ReRAM-based accelerators are highly error-prone due to limitations of technology fabrication that lead to process variations and defects. These limitations degrade the accuracy of Deep Convolutional Neural Networks (Deep CNNs) running on PIM accelerators. While these CNNs accelerators are widely deployed in safety-critical systems, their vulnerability to fault is not well explored. In this paper, we have developed a fault injection framework to investigate the vulnerability of large-scale CNNs at both software- and hardware-level of inference phases. Faulty ReRAM devices are another reliability challenges due to significant degradation of classification accuracy when CNN parameters are mapped to the accelerators. To investigate this challenge, we map the CNN learning parameter to the ReRAM crossbar and inject faults into crossbar arrays. The proposed framework analyzes the impact of stuck-at high (SaH) and stuck-at low (SaL) fault models on different layers and locations of CNN learning parameters. By performing extensive fault injections, we illustrate that the vulnerability behavior of ReRAM-based PIM accelerator for CNNs is greatly impressible to the types and depth of layers, the location of the learning parameter in every layer, and the value and types of faults. Our observations show that different models have different vulnerabilities to faults. Specifically, we show that SaL further reduces classification accuracy than SaH.
[AI-197] Offline AI Modules: Voice-First Offline Architecture Hardware Reference Stack Quantization and Benchmarking
链接: https://arxiv.org/abs/2610.07026
作者: Sunday Afariogun,Odunolaoluwa Jenrola,Zeinab Nezami
类目: Artificial Intelligence (cs.AI); Networking and Internet Architecture (cs.NI); Systems and Control (eess.SY)
备注:
Abstract:The Offline AI Modules workstream enables practical, low-power, and community-accessible deployment of voice-first AI systems that operate fully offline. Designed for African language communities where speech is the dominant mode of interaction and internet connectivity is unreliable or absent, the workstream delivers three reinforcing components: a modular voice-first offline architecture, a low-cost hardware reference bill of materials, and a reproducible quantization and a reproducible quantization and benchmarking pipeline for instruction-tuned language models in the 2-5B parameter class. This paper presents the first end-to-end benchmark evaluation of the stack across two hardware tiers: an NVIDIA Jetson Orin NX (TierB) and a Raspberry Pi5 (TierA). Three instruction-tuned models are evaluated across four quantization formats, assessed for deployment metrics (decode throughput, chat latency, memory, power) and multilingual quality (topic classification accuracy on MasakhaNEWS across English, Hausa, Igbo, Nigerian Pidgin, and Yoruba; per-language perplexity drift). Speech recognition is evaluated using Ethio-ASR on Amharic and Oromo across both tiers. The principal finding is that Q4_K_M quantization represents the best size-to-quality trade-off for deployment on both tiers: gemma-4-E2B-it achieves 28.8t/s decode throughput and 89.2% topic classification accuracy at Q4_K_M on TierB, while all three models run within the 16GB memory budget on TierA.
[AI-198] When to Rethink: Learning Multi-Perspective Self-Verification for Vision-Language Models
链接: https://arxiv.org/abs/2610.07018
作者: Ziquan Zhu,Hanruo Zhu,Si-Yuan Lu,Morris Yu-Chao Huang,Yicheng Lin,Wei Han,Tianlong Chen,Mingyuan Wu,Hanchao Yu,Gaojie Jin,Lu Liu,Bo Sun,Tianjin Huang
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Vision-language models (VLMs) have achieved strong performance in multimodal reasoning, yet they remain prone to generating plausible but incorrect answers. Self-verification offers a practical way to improve answer reliability without relying on external judges, but existing methods typically depend on a single verification criterion or fixed prompt, resulting in incomplete and unstable reliability estimates. We first systematically analyze how verifier capability and prompt design affect verification performance. Our findings show that stronger verifiers provide more reliable judgments, while verification performance is highly sensitive to prompt choice, with no single prompt consistently dominating across tasks. Guided by these findings, we propose \textttMOTIVE, a \textbfMulti-View Self-Verificati\textbfOn wi\textbfTh Rel\textbfIability-Guided Selecti\textbfVE Rethinking framework for reliable multimodal reasoning. \textttMOTIVE evaluates each candidate answer from complementary verification perspectives and learns a correctness-aligned reliability score through correctness-grounded multi-view verification learning. During inference, this score governs an accept-or-rethink decision, allowing reliable answers to be returned directly while uncertain ones trigger history-guided rethinking. Extensive experiments across diverse multimodal benchmarks and VLM backbones demonstrate that \textttMOTIVE consistently outperforms strong self-verification and self-correction baselines. Further results show that reliable verification improves accept-or-rethink decisions and reduces unnecessary reasoning turns, enabling more reliable and efficient self-verification without an external judge.
[AI-199] Which Image Property Carries the Jailbreak? A Controlled Dissection of Image-to-Text Jailbreaks
链接: https://arxiv.org/abs/2610.07009
作者: Boyuan Chen,Yehia Dawoud,Hailemariam Mersha,Minghao Shao,Siddharth Garg,Ramesh Karri,Muhammad Shafique
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:
Abstract:Image-to-text jailbreaks place harmful intent in text, image content, or the relationship between them. We examine image-side factors across four published attack families on a 313-prompt StrongREJECT slice, using five multimodal models and an additional appendix evaluation of InternVL3.5-8B. The harmful instruction is held constant across conditions; the baseline matrix uses one draw per prompt, and paired ablations use three draws with an automated rubric judge. A bare harmful query, with or without a benign unrelated image, produces little attack success on most victims, while attack images substantially increase it. First, per-tile entropy and JPEG size do not reliably distinguish attack tiles from size-matched benign distractors, limiting density-only screening. Second, earlier tile-count ladders were confounded by payload visibility. A corrected region-count test found no detectable effect, so the role of tile-count structure remains unresolved. Third, on Qwen3-VL-8B, the E4 manipulation that removes query-specific relatedness lowers ASR by about 0.12. This supports a bounded attribution to the relatedness manipulation, although image-text congruence remains unmeasured. A within-category control reproduces the direction on overlapping stimuli. Five tests survive the global statistical correction, but only E4 supports attribution to one measured descriptor; the within-category result is a robustness check, not a separate attribution. These conclusions remain conditional on the rubric judge.
[AI-200] Where Does the Audio Jailbreak Live? A Controlled Frequency-Depth Audit of AdvWave-P on Qwen 2-Audio
链接: https://arxiv.org/abs/2610.07005
作者: Boyuan Chen,Minseok Kim,Sohaila Abdulsattar,Minghao Shao,Siddharth Garg,Ramesh Karri,Muhammad Shafique
类目: ound (cs.SD); Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR); Machine Learning (cs.LG)
备注:
Abstract:We audit frequency and decoder-depth claims for AdvWave-P, an additive audio jailbreak, on Qwen2-Audio. The protocol masks frequency components of the perturbation in the short-time Fourier transform (STFT) domain and measures attack success and audio-span representations. On 520 AdvBench prompts, the primary judge labels 76.7% of adversarial inputs as jailbreaks. A condition-blind, single-annotator validation yields a Rogan-Gladen sensitivity estimate of 0.87 for this condition (about 0.83-0.95 with validation-rate uncertainty); this correction is not applied to masked conditions. The apparent frequency ranking depends on the partition: energy share alone predicts the standard eight-band ranking (Spearman’s rho = 0.95), and equal-Hz and equal-energy partitions show that masking any tested band can sharply reduce attack success. At a finer 16-band equal-energy resolution, however, masking the narrow 7520-7960 Hz band leaves ASR at 0.10, which remains unresolved without a matched control. Matched-energy scattered-removal tests show that contiguous removal is more damaging in lower bands, while both forms approach the floor in upper bands. Global rescaling leaves ASR near baseline but tests amplitude sensitivity rather than frequency location. In a re-optimization pilot (n = 20), tested single- and two-band supports reach ASRs of 0.00, 0.25, and 0.40, while random supports covering about half the STFT bins reach a mean of 0.86. In a prompt- and energy-adjusted model, audio-span divergence is associated with band necessity, with the coefficient rising from +0.63 at the projector output to +0.93 at layer 30 (contrast +0.293, 95% CI [0.11, 0.51]). This association is not a causal localization, and single-layer patching does not establish a causal layer. The results support partition-aware auditing of frequency claims, leaving the fine-resolution top-band result and broader generality open.
[AI-201] opology-Consistent Task Planning over Cellular Workflow Complexes for LLM -based Agents
链接: https://arxiv.org/abs/2610.07004
作者: Sen Zhao,Jia Tang,Ruiqi Kong,Zuyu Zhang,Lifeng Shen,Ding Zou,Xinyu He,Xu Zhang,Junwei Han
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Task planning for LLM agents requires workflows that satisfy both user intent and complex sub-task dependencies. While existing planners work well for sequential or directed acyclic graph (DAG)-like structures, they struggle with workflow patterns such as verification-correction loops, convergent branch merging, and reusable intermediate states that arise naturally in real-world tool orchestration. We present TopoPlanner, a topology-consistent planning framework that lifts tool dependency graphs into cellular workflow complexes and uses them as topologyaware context for LLM tool planning. TopoPlanner retrieves a request-relevant closed subcomplex through cosheaf-consistent cellular retrieval, performs multidimensional structural reasoning over the retrieved topology, and interfaces the resulting cellular representation with the planner LLM for tool-sequence generation. Experiments on four tool-planning benchmarks with topology-guided loop, merge, and loop-merge workflows show consistent improvements over prompt-based and graph-enhanced baselines across different local LLM backbones.
[AI-202] Joint upper-bound coverag e and route-choice utility: an empirical evaluation on two urban proxy tasks
链接: https://arxiv.org/abs/2610.06995
作者: Jianing Long,Xiaobin Li,Wuming Lei,Weiguang Wang
类目: Artificial Intelligence (cs.AI); Computers and Society (cs.CY)
备注:
Abstract:Whether more accurate traffic forecasts or higher uncertainty coverage improve route decisions is unclear. We evaluate this question with a frozen protocol that separates speed error, joint candidate path upper bound coverage, route selection, and realized loss. Using processed road speed data from Beijing and Chengdu, we construct offline proxy tasks with 150 origin destination pairs, three candidate paths, and 14 test days per city. We compare raw 90th percentile path time bounds with jointly calibrated upper bounds under minimum bound route choice. Joint coverage rises from 83.19% to 92.26% in Beijing M1, from 75.14% to 88.33% in Chengdu M1, and from 74.01% to 90.64% in Chengdu M2. Yet C2 increases lateness by 0.1633, 0.7848, and 0.9200 percentage points, respectively, and mean travel time by 0.588, 3.082, and 4.418 seconds. In a separate Chengdu predictor comparison, a 14.91% reduction in speed mean absolute error accompanies a 1.4571 percentage point reduction in lateness under C0. Joint coverage is therefore not a surrogate for downstream route utility in these frozen tasks the offline results do not establish online or causal benefits.
[AI-203] ARE: Weigh a Never-Poisoned Twin Before Reading Backdoor-Defense Costs
链接: https://arxiv.org/abs/2610.06994
作者: Ruizhi Xu,Wei Xu,Sibo Zhu
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
备注: 84 pages (9-page main text)
Abstract:Backdoor-defense leaderboards print a clean-accuracy drop and read it as removal cost. Measured on the poisoned victim alone, the drop cannot separate removal from what the defense does to any model, and inherits the victim’s start, which for three of BackdoorBench’s sixteen attacks is a configuration file: WaNet, BPP and Input-Aware ship a MultiStepLR that never fires, so their victims never anneal and are the least accurate in 30/31 public CIFAR cells at \leq 5%. On PreAct-ResNet18, fine-tuning-family defenses return a low start to their own level, so there the published cost is negative, the benchmark’s rating clips the “gain” to zero, and 2 of 48 citing defense papers we read rest a no-cost claim on those cells; TSBD and CGD, re-run with their code, “gain” on a never-poisoned model too. A 2\times2 editing only that scheduler line isolates the cause, its swapped arms self-registered before they ran: the sign of the fine-tuning family’s clean-model cost reverses both ways while its published gain on the annealed victim only shrinks toward zero, 44/44 seeds following the schedule, replicated on BPP, FT-SAM, CIFAR-100 and VGG19-BN and induced in a second toolkit. TARE runs the same defense on a never-poisoned twin of the same recipe, schedule and seed (on BackdoorBench, \leq 10 poisoned images, admitted only below 5% attack success); what the twin loses is the tare. On the BadNets grid seven of eight defenses charge the twin (Neural Cleanse only where its detector fires), +0.13 (fine-tuning) to +5.70 points (I-BAU); the eighth, ABL, destroys it. Within an attack the start cancels from rankings, so the tare re-orders nothing there; what poisoning adds beyond it is printed under two estimators and not corrected, its removal share unidentified. We ship the three-key patch, a signed tare column (7 attacks \times 8 defenses) and TARE-Z, a twin-free estimator for seed-stable defenses.
[AI-204] DART-ES: Difficulty-Aware Reweighting and Targeted Replay for Fine-Tuning LLM s with Evolution Strategies
链接: https://arxiv.org/abs/2610.06993
作者: Zhishen Sun,Hongzhan Wang,Sizhe Dang,Guang Dai,Haishan Ye
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:
Abstract:Evolution Strategies (ES) enable memory efficient full parameter fine-tuning of large language models (LLMs) using only forward computation. However, standard ES uniformly averages rewards across problems and compresses problem level population feedback into a single scalar, making it difficult to capture how the learning value of each problem changes with model capability. To address this limitation, we propose Difficulty-Aware Reweighting and Targeted Replay for Evolution Strategies (DART-ES). DART-ES estimates the local solvability of each problem from its pass rate across the perturbation population and aggregates historical observations to construct a dynamic difficulty state. This shared state jointly guides continuous difficulty reweighting and rare solvable sample replay, thereby improving perturbation direction evaluation and training data allocation without introducing an additional difficulty model or backpropagation. Extensive experiments show that DART-ES achieves good fine-tuning performance. DART-ES outperforms ES on all five base models and improves the average accuracy from 72.07% to 73.53%, exceeding the 73.26% achieved by GRPO on GSM8K. Across five challenging mathematical reasoning benchmarks, DART-ES achieves an average accuracy of 49.20%, compared with 48.34% for ES and remains competitive with strong 7B models trained with RL. Further experiments show consistent gains in instruction tuning, code generation and the Countdown task with a 14B model, demonstrating strong generalization across tasks and scalability to larger models. Beyond performance gains, DART-ES also shows clear advantages in system efficiency. It reduces runtime per step by 15.2%–50.2% and peak memory usage per GPU by 21.1%–51.1% compared with GRPO. Despite performing full parameter fine-tuning, DART-ES also requires less runtime and GPU memory than GRPO+LoRA.
[AI-205] When better traffic forecasts fail to improve signal control: a layered diagnostic study of forecast-to-decision value
链接: https://arxiv.org/abs/2610.06992
作者: Jianing Long,Xiaobin Li,Wuming Lei,Weiguang Wang
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Improved traffic forecasts do not necessarily yield better signal-control decisions. We investigate this gap through a layered diagnostic study using 29 days of reconstructed demand from Xuancheng, China, with seven dates reserved for testing. The framework evaluates point forecasts, conformal intervals, dependence-aware scenarios, and matched closed-loop controllers. Entry-level and movement-level forecasts reduce mean absolute error by 4.03% and 3.92%, respectively, relative to historical means. A nominal 90% conformal interval achieves 90.72% marginal coverage but only 75.66% on an ex-post high-demand subset. Interface audits identify decision-time leakage and reveal that only two of nine controlled intersections offer multiple effective actions. We correct the temporal interface and compare causal forecasts with a five-second event oracle using exhaustive joint-action search. A synthetic positive control demonstrates that future information can reduce the internal rollout cost by 61.5%. On the frozen test dates, however, causal forecasts and the event oracle increase queue vehicle?seconds by 6.09% and 3.39% relative to the matched no-future rollout, while the oracle reduces spillback exposure by 3.78%; paired-day bootstrap intervals cross zero. These findings indicate that forecast value depends on temporal observability, action identifiability, dynamics consistency, and objective alignment. The proposed protocol provides a practical way to diagnose where predictive improvements fail to translate into operational benefits.
[AI-206] EPOCH: Reliable Discovery through Evidence-Governed Search
链接: https://arxiv.org/abs/2610.06986
作者: Binjie Guo,Aisheng Mo,Ruitong Li,Xinle Deng
类目: Artificial Intelligence (cs.AI)
备注: 49 pages, 16 figures, including supplementary material
Abstract:AI research agents are increasingly used to search over programs, mathematical constructions, and proofs. However, existing systems typically optimize evaluator feedback without adequately governing how that feedback is interpreted, challenged, and reused. As a result, promising but fragile candidates can be promoted as discoveries, while benchmark improvements, finite certificates, and theorem-level claims are too easily conflated. We introduce EPOCH, an evidence-governed architecture designed to close this gap. EPOCH implements an evidence-governed discovery loop by combining explicit task contracts, typed memory, active falsification, admission checks, and independent replay, so that each candidate is evaluated against the strength and scope of the claim it supports. EPOCH achieves state-of-the-art aggregate performance on AlgoTune, substantially exceeding the strongest baseline in mean normalized score (0.65 vs. 0.53), and attains the highest mean score on the internal Math14 suite (0.57). It further shows favorable held-out behavior under official-test replay and leads the descriptive aggregate on AgentHPO. Across ten discovery problems, EPOCH delivers substantial task-specific advances, including improved executable constructions, optimized algorithms, counterexamples, and proof-supported results. These advances demonstrate its ability to convert search into concrete progress across mathematical and computational domains. Together, the results suggest that evidence governance is a necessary step toward AI research agents that produce not only stronger solutions, but also more trustworthy scientific discoveries.
[AI-207] APEX: Active Protection at Execution Boundaries for LLM Agents
链接: https://arxiv.org/abs/2610.06966
作者: Xinran Zheng,Xin Fan Guo,Zhiqiang Hao,Fan Yang,Xingzhi Qian,Jiawei Du,Jinfeng Xu,Zheng Xing,Shuo Yang,Xingjun Wang
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
备注:
Abstract:Indirect prompt injection (IPI) hides adversarial instructions in content that large language model (LLM) agents read at runtime. As agents compose heterogeneous capability units, including Tools, MCP servers, and Skills, the carriers of injection multiply, and defenses built to recognize attack patterns fall behind them. We instead shift defense from covering attack patterns to one stable point: whatever the carrier and however the injection propagates, harm materializes only at the \emphexecution boundary, where the agent turns internal state into an external action or released output. Safety there turns on two conditions, both settled by the trusted task rather than by the run: whether the proposed effect is authorized, and whether the runtime information reaching it is endorsed by that task. We present APEX, an active defense that enforces both at this boundary from a single authorization contract compiled before untrusted execution: \emphevidence-gated prevention admits an effect only when the contract justifies it, while \emphdeception-based exposure makes unendorsed use reveal itself before the effect commits. Protection therefore follows from what the task permits rather than from how an attack is built, and applies uniformly across capability units without attack-specific policies or taint tracking. Against 13 baselines, APEX attains 0% attack success on five of six benchmarks and 0.56% on the sixth, holds 0% under adaptive attacks on all three capability-unit types, and remains effective across defender backbones. Code is available at this https URL.
[AI-208] Principles that Guide Actions that Inform: Agent Evolution via Knowledge Abstraction
链接: https://arxiv.org/abs/2610.06964
作者: Bowen Ye,Yongchao Xu,Junkai Ma,Xiang Yin,Wenzhao Li
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:
Abstract:Large language model (LLM) agents have demonstrated strong capabilities in interactive environments, yet their ability to continually evolve from experience remains limited. Although fine-tuning enables adaptation, its dependence on parameter access and high computational costs restrict its flexibility, especially for large-scale and closed-source LLMs. External memory offers an alternative by allowing agents to accumulate experience without modifying model parameters. However, existing methods mainly focus on experience representation and organization, while the acquired knowledge remains tightly coupled with specific tasks and contexts, limiting generalization. A key challenge is how to transform concrete interactions into abstract and reusable knowledge that guides future decisions beyond individual experiences. To address this challenge, we propose SAGA (\underline\textbfSelf-evolving \underline\textbfAgents through Experience-\underline\textbfGrounded \underline\textbfAbstraction), a framework for experience-grounded knowledge abstraction and utilization in LLM agents. SAGA progressively transforms interaction trajectories into episodic descriptions, reusable procedures, and principles with explicit applicability conditions, while maintaining links to execution evidence. Retrieved principles are instantiated into task-specific guidance and used to refine candidate actions through corrective feedback and resampling. This creates an execution–abstraction feedback loop, where accumulated knowledge guides future interactions and new experiences continuously update hierarchical memory. Experiments on ScienceWorld and ALFWorld demonstrate improved task performance, with ablation studies highlighting the importance of contextual instantiation and action regulation for leveraging principle-level knowledge. Subjects: Artificial Intelligence (cs.AI); Machine Learning (cs.LG) Cite as: arXiv:2610.06964 [cs.AI] (or arXiv:2610.06964v1 [cs.AI] for this version) https://doi.org/10.48550/arXiv.2610.06964 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[AI-209] Learning to Decide Not to Reason : Parameter-Efficient Decision Operators via Low-Rank Activation Steering
链接: https://arxiv.org/abs/2610.06950
作者: Ran Li,Lei Chen
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:
Abstract:Injecting skills into a frozen language model currently costs a million parameters and a reinforcement-learning pipeline. We introduce \method, a System-1 decision operator trained by behavior cloning that lowers this cost by roughly two orders of magnitude. The default operator uses 330K parameters to match a 1.33M-parameter operator trained with reinforcement learning, exceeds or achieve comparable performance, while collapsing 3,685-token deliberation into a 6-token decision with no loss in accuracy. A rank-4 variant with 23K parameters, 1/58 of the strongest published skill operator, suffices for SearchQA and near-suffices for LiveMath, where higher rank still helps; the same recipe transfers across five tasks and three backbones, with out-of-distribution gains persisting on LiveMath problems released months after training. The gap to prior work is trainability, and it is set jointly by initialization and architecture: the initialization of prior operators zeroes the gradient of both large factor matrices at the first optimization step, whereas our zero-initialized output projection inside a shared low-rank backbone receives a gradient immediately, which a gradient-flow probe confirms directly. The gain isn’t chain-of-thought compression: 23 of 57 LiveMath points beat the base model’s best-of-8 sampling, and a logit-lens probe shows the operator amplifies the answer along the model’s existing late-layer pathway, not writing it earlier. Gains track the base model’s headroom across 13 base–task pairs, and skills compose as approximately linear operators that can be added, interpolated, and hot-swapped at inference time. Code on this https URL.
[AI-210] Metonymic Circuits for Abstract Concept Grounding in Vision Transformers EMNLP2026
链接: https://arxiv.org/abs/2610.06928
作者: Jing Ding,Ziqiao Ma,Jiayuan Mao,Joyce Chai,Freda Shi
类目: Artificial Intelligence (cs.AI)
备注: EMNLP 2026 Main. Project Website: this https URL
Abstract:We study how Vision Transformers ground abstract concepts (e.g., angry) when training data provide limited direct referential evidence. We hypothesize a metonymic grounding mechanism in which abstract predictions are driven by concrete, interpretable anchor concepts (e.g., fire) that bridge visual signals to abstract semantics. By applying Transcoders on CLIP and DINO vision encoders, we recover intermediate features that can be associated with semantic labels for more concrete concepts, and trace their contributions in circuits underlying abstract concept recognition. Experiments on a carefully curated icon dataset reveal structured metonymic circuits, in which perceptual primitives dominate early layers and object-like anchors precede abstract targets. Images containing rendered text instead recruit a distinct perceptual-to-textual route. Causal interventions further validate that metonymic intermediates are functionally involved in grounding abstract concepts.
[AI-211] AttSVD:Prompt-Adaptive Low-Rank KV Cache Compression via Attention-Guided SVD
链接: https://arxiv.org/abs/2610.06927
作者: Sara Abdali,Jongwoo Ko,Pashmina Cameron
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:
Abstract:The key-value (KV) cache of autoregressive transformers grows linearly with context length and dominates memory at long context. Most training-free remedies evict low-importance tokens, an irreversible choice along the sequence axis. We instead keep every token and store it more cheaply along the “feature” axis. We therefore propose AttSVD, a new “interpretable” low-rank compression whose basis is derived from each prompt’s own attention geometry: an online, per-prompt truncated SVD that keeps only the directions attention actually reads, cutting persistent per-head KV memory in proportion to the retained rank. We propose two decode-time caching strategies, accumulating and streaming, for short and long generation regimes. Furthermore, we propose two refinements that make compression adaptive. A per-matrix energy rule sizes the logit space and the attention mass independently. An attention-aware basis truncates only in the spaces attention actually reads, preserving both the attention logits and the attention output. The same factors also provide free, per-head interpretability insights into the effective rank and the geometry attention consumes. Across multiple models, on both an agentic benchmark and the full LongBench suite AttSVD stays on par with the dense cache while using up to 50% of the KV-cache memory.
[AI-212] RadOnc-Agent : An LLM -Orchestrated Framework for AI Workflows Across the Radiotherapy Care Pathway
链接: https://arxiv.org/abs/2610.06923
作者: Caiwen Jiang,Shuoyang Wei,Songlin Zhao,Junyu Li,Jingyuan Chen,Wei Liu
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Artificial intelligence has advanced individual radiotherapy tasks, yet these capabilities remain separated across clinical stages, software environments and data modalities. This fragmentation contrasts with the longitudinal radiotherapy workflow from treatment decision-making through follow-up. Here we present RadOnc-Agent, an agentic artificial-intelligence framework that formalizes radiotherapy into four clinical phases and provides 26 callable functions through a conversational interface. A large-language-model controller maps clinical intent to schema-constrained calls, preserves patient and workflow context, and routes requests to specialist services. We evaluated system execution using 2,600 single-function requests (7,800 repeat executions), 200 prespecified synthetic cross-stage scenarios spanning four phases (600 executions), and 120 workflow instances from 60 de-identified patient records (360 clean executions) representing decision-to-planning and planning-to-adaptation. RadOnc-Agent selected the intended function in 98.79% of single-function executions, completed 96.50% of scripted cross-stage workflows, and completed 96.67% of real-patient workflow executions. In comparative ablations, removing longitudinal state reduced cross-stage completion from 96.50% to 84.00%, while disabling schema and identity validation increased mismatched backend dispatch from 0% to 95.28% in a replay/test evaluation. These findings establish the technical feasibility of an LLM-orchestrated architecture for coordinating heterogeneous radiotherapy capabilities and information across longitudinal workflows; they do not establish clinical correctness, clinical utility or prospective benefit.
[AI-213] FluidPD: In-Place Elasticity for SLO-Aware Prefill-Decode Disaggregated LLM Serving
链接: https://arxiv.org/abs/2610.06917
作者: Kartik Ramesh,Kaidi Fu,Zihan Zheng,Jiahuan Yu,Fabio Oliveira,Carlos Costa,Minjia Zhang
类目: Artificial Intelligence (cs.AI)
备注: 13 pages, 11 figures
Abstract:Prefill-decode disaggregation is becoming a common architecture for LLM serving because it separates two phases with distinct execution patterns and SLO objectives. Existing systems typically combine a fixed prefill/decode worker ratio with request routing across workers. However, real-world workloads exhibit both short bursts and sustained shifts in the prefill-to-decode demand ratio. As a result, a configuration that is well provisioned at one time may quickly become mismatched, causing latency SLO violations even when idle capacity exists elsewhere. Existing autoscaling mechanisms can add capacity, but they react slowly, require spare GPUs, and do not directly address short-timescale phase imbalance. We present FluidPD, a P/D-disaggregated serving system that provides SLO-aware in-place elasticity. FluidPD introduces two complementary mechanisms. FluidToken handles transient imbalance by offloading a bounded portion of prefill computation to decode workers when decode-side slack is available. FluidRole handles sustained imbalance by reassigning running workers between prefill and decode roles in place, avoiding model reload and engine restart. Both mechanisms are guided by lightweight pressure indices that expose prefill and decode-side resource pressure before they appear as SLO violations. Across production Azure trace workloads, FluidPD improves overall SLO attainment over static SGLang by up to 94.6 percentage points, demonstrating that SLO-aware in-place P/D elasticity improves service quality without provisioning additional workers. Comments: 13 pages, 11 figures Subjects: Artificial Intelligence (cs.AI) Cite as: arXiv:2610.06917 [cs.AI] (or arXiv:2610.06917v1 [cs.AI] for this version) https://doi.org/10.48550/arXiv.2610.06917 Focus to learn more arXiv-issued DOI via DataCite
[AI-214] xt2Dashboard: A Governed Agent Architecture for Natural-Language Dashboard Generation over Enterprise DataBrain
链接: https://arxiv.org/abs/2610.06914
作者: Yiou Wu,Zezhi Tang,Ningwei Bai,Liuhaichen Yang
类目: Artificial Intelligence (cs.AI); Databases (cs.DB)
备注: 14 pages, 3 figures
Abstract:Text2Dashboard is a DataBrain-specific prototype that turns natural-language analytic requests into inspectable dashboards. An installable Codex plugin and standalone Agent Runtime combine schema-constrained model decisions with typed tools, persistent state, and deterministic Hooks for approval, audit, checkpointing, recovery, and failure handling. The pipeline resolves entities, discovers metadata, enforces read-only SQL, composes dashboards, and applies static checks, dynamic preflight, and browser inspection. The model proposes actions while deterministic software controls execution and records state transitions. We evaluate the workflow on frozen real-DataBrain tasks and controlled Hook faults. Strict success was 6/8 on metadata and SQL tasks: metadata selection passed 4/4, all four SQL tasks met semantic criteria, and 2/4 met the exact output-column contract. The final release passed 4/4 single-panel dashboard tasks, one two-panel task, and one existing-dashboard refinement; a parameterised task exceeded its step limit. All ten fault scenarios met their specified outcomes without unapproved external side effects. Model inference accounted for over 97% of observed runtime in every reported group. These small, DataBrain-specific results do not establish production readiness, general text-to-SQL accuracy, or an efficiency advantage over manual dashboard construction. Comments: 14 pages, 3 figures Subjects: Artificial Intelligence (cs.AI); Databases (cs.DB) Cite as: arXiv:2610.06914 [cs.AI] (or arXiv:2610.06914v1 [cs.AI] for this version) https://doi.org/10.48550/arXiv.2610.06914 Focus to learn more arXiv-issued DOI via DataCite
[AI-215] GAMEGO: Training Game-Dev Agents with Synthetic Trajectories Anchored in Real-World Assets
链接: https://arxiv.org/abs/2610.06910
作者: Haoyue Yang,Jingyao Li,Zhengfan Wu,Jing Liu,Xuanle Zhao,Kang Liu
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Recent advances in Large Language Models (LLMs) have demonstrated remarkable capabilities in web front-end execution, with browser-based game generation emerging as a particularly prominent frontier. While previous efforts frequently rely on complex multi-turn workflows or focus on static game evaluation benchmarks, this work targets direct end-to-end real-world game synthesis driven by coding agents. However, generating complex games directly from sparse user queries often forces coding agents to make underspecified assumptions, yielding incomplete mechanics, disconnected gameplay flows, and limited visual aesthetics. To resolve this issue, this paper presents GameGo, a scalable framework that systematically transforms brief game seeds into comprehensive Product Requirements Documents grounded in industry game-development practices. To retain core gameplay constraints without restricting design exploration, GameGo uses task-specific dynamic compression to maximize information density while preserving instruction following. Based on this pipeline, GameGoData is constructed with 55,060 development trajectories across 2D, 2.5D, and 3D games, alongside GameGoBench, a benchmark comprising 124 diverse game queries. Training GameGoCoder on GameGoData yields a model that outperforms matched baselines and is comparable to frontier models across gamedev benchmarks. All code, datasets, and models will be made publicly available.
[AI-216] Dynamical low-rank equilibrium computation for stochastic games between advanced persistent threats and moving target defense
链接: https://arxiv.org/abs/2610.06885
作者: Tian Zijian,Zhang He,Chen XinJie,Liu Xinggao
类目: Computer Science and Game Theory (cs.GT); Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR); Systems and Control (eess.SY)
备注: Submitted to Automatica. Source code: this https URL
Abstract:Moving target defense (MTD) against advanced persistent threats (APTs) in industrial control systems (ICS) has well-established game-theoretic formulations, but their practical value hinges on equilibrium computation: full-rank value iteration is prohibitively expensive at industrial state dimensions, and the resulting defense strategies admit no certified robustness against adversarial perturbations. We first reveal that the attack and defense influence matrices of ICS dynamics are intrinsically low-rank: APTs infiltrate through a handful of entry points, and MTD reconfigures only a limited subset of components per cycle. We prove that this structure propagates through the non-smooth Bellman operator of the zero-sum stochastic game: an augmented gradient matrix bridging physical and algorithmic low rank certifies that every Bellman target lies near a low-dimensional subspace, with an explicit error bound on the optimal value function. Because these subspaces drift under value iteration, static low-rank projections are inadequate. We therefore propose dynamical low-rank equilibrium computation (DLR-NE), which augments the rank-r search space at each iteration, regularizes the core matrix spectrum, and retracts via truncated SVD, extracting a Nash equilibrium at every step. Four guarantees follow: explicit approximation error; geometric convergence to a neighborhood with five physically interpretable error sources; per-step cost O(nr^2), a Theta(n/r^2) speedup over full-rank value iteration; and robustness in which a single weight trades accuracy against certified safety. Experiments on a nonlinear power-system testbed confirm each prediction, with 94% parameter compression at 0.16% utility loss.
[AI-217] Learning When to Refine: Long-Horizon Reinforcement Learning for Budgeted Neural-Operator PDE Solvers
链接: https://arxiv.org/abs/2610.06883
作者: Ange Tong
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: 21 pages, 4 figures, 2 tables
Abstract:Neural operators provide fast surrogates for time-dependent PDEs, but autoregressive deployment creates a refinement-allocation problem: prediction errors vary over space and time, while only a finite number of local corrections can be committed along a trajectory. We formulate this as budgeted adaptive neural-operator solving. A global Fourier neural operator advances the full field, a local operator proposes patch-wise residual corrections, and a set-aware selector chooses where to refine. A macro policy decides when and how much of the remaining refinement budget to spend. We introduce rollout-verified policy improvement (RV-PI), which evaluates feasible refinement counts through actual continuation rollouts of the learned PDE solver, converts long-horizon advantages into conservative policy targets, and accepts an update only when held-out trajectory error improves. On the shallow-water benchmark with a 32-intervention budget, RV-PI achieves a three-seed mean trajectory relative L2 error of 0.6910, improving over immediate-only policy improvement by 5.37% and RandomMacro by 2.41%. On the forcing-driven Brusselator benchmark with a 76-intervention budget, RV-PI attains 0.09954, improving over immediate-only policy improvement by 2.31% and RandomMacro by 5.32%. These results show that, under a fixed refinement budget, the value of a local correction depends on its downstream effect on the autoregressive trajectory, not only on its immediate error reduction.
[AI-218] Comparative review of hybrid forecasting models for short-term prediction of building thermal load
链接: https://arxiv.org/abs/2610.06881
作者: Nikolaos A. Efkarpidis,Despoina Kothona,Georgios C. Christoforidis
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: 27 pages, 15 tables, and 13 figures
Abstract:In this paper, a comparative review of different hybrid models for short-term forecasting of building thermal demand is carried out. Particularly, the assessment tackles the comparison of data-driven models enhanced with other state-of-the-art techniques. At the first step, the existing techniques reported in the literature are analysed. It is concluded that Metaheuristics or a data-driven model are used to identify the parameters of the basic model. The qualitative evaluation includes for each method the input and output features, main advantages and drawbacks. At the second step, an existing dataset of historical thermal demand from Scottish households, as well as historical weather forecasts are utilized to assess additionally the performance of existing hybrid methods. From the assessment of 13 hybrid methods, the Empirical Modal Decomposition - long short-term memory - Markov (EMD-LSTM-Markov) model can predict with the highest accuracy the day-ahead power pattern of heating and domestic hot water (DHW) demands. Though local power peaks are also accurately predicted, high power swells and spikes are underestimated. Other methods, such as Support Vector Machine - Simulated Annealing (SVM-SA) and Random Forest - Improved Sparrow Search Algorithm - LSTM (RF-ISSA-LSTM) predict a smooth pattern of heating and DHW demand profiles with rapid changes underestimating most power peaks.
[AI-219] Neutrosophic Ensemble Classification for Uncertainty-Aware Bearing Fault Detection: Evidence from Laboratory and Variable-Speed Industrial Benchmarks
链接: https://arxiv.org/abs/2610.06880
作者: Maikel Leyva-Vazquez,Dayron Rumbaut Rangel,Lorenzo Cevallos-Torres,Alexis Matheu Perez
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: 18 pages, 5 figures
Abstract:Machine learning classifiers for bearing fault detection produce scalar confidence scores that conflate confident errors with genuinely ambiguous predictions, and the conventional truth/falsity pair (F = 1 - T) is algebraically redundant by construction. We operationalize a refined neutrosophic decomposition of a Random Forest + XGBoost + Logistic Regression ensemble into four indicators – T-hat (top-class evidence), F-hat (best-competitor evidence), predictive entropy I1-hat, and decision disagreement I2-hat – evaluated on two bearing benchmarks (CWRU and JNU, 600-1000 rpm) under a leave-one-condition-out protocol. On CWRU, after correcting a file-to-class mapping error, the ensemble reaches 100.00 percent accuracy on three of four held-out loads (92.27 percent on the fourth), leaving too few errors for uncertainty analysis. On JNU, holding out 1000 rpm, accuracy collapses to 40.64 percent, below a majority-class baseline; Logistic Regression (57.91 percent) generalizes far better than the tree ensembles. I1-hat shows a robust association with error beyond T-hat/F-hat, while I2-hat contributes little; standalone Logistic Regression confidence outperforms the full decomposition, a boundary condition we report honestly. Two further results extend this: fusing a time-domain and a frequency-domain model of the same signal and scoring their Jensen-Shannon divergence beats that model own entropy (AURC 0.29 vs. 0.36 on the standard split; 0.54 vs. 0.73 under a harder single-condition reproduction), the only indicator moving correctly under a CWRU-versus-JNU distributional-shift contrast; and, on CWRU alone, literature-verified bearing fault frequencies, correctly demodulated via the envelope spectrum, separate most fault classes almost perfectly (99.57 percent) using three interpretable features. Code, logs, and figures are released for independent verification.
[AI-220] What Does Fréchet Distance Measure? A Directional Decomposition
链接: https://arxiv.org/abs/2610.05518
作者: Yunghee Lee,Jaeyeon Kim
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: 20 pages, 4 figures, 7 tables
Abstract:The Fréchet distance is a de facto standard for evaluating generative models across domains, appearing as FID for images and FVD for videos. It summarizes the discrepancy between generated and reference distributions in a single scalar, with lower values typically interpreted as better generation quality. However, this scalar view can obscure what drives the comparison. For example, in COCO dataset, increasing the number of diffusion sampling steps improves ImageReward scores yet worsens (increases) FID. Motivated by this mismatch, we seek to make the Fréchet distance more interpretable by uncovering where the discrepancy lies. To this end, we introduce directional Fréchet distance, the expected squared projection of the optimal transport displacement onto a given direction. Across our image, video, and protein case studies, we find that a small number of interpretable directions account for much of the distance. We use these directions to explain the FID increase in terms of semantic concepts represented by CLIP embeddings, quantify FVD’s bias toward per-frame appearance, and revisit the interpretation of Protein FID. We open-source our codebase at this https URL.
[AI-221] From Search to Signal: Online Post-Training in Automatic Heuristic Design
链接: https://arxiv.org/abs/2609.39383
作者: Yilun Yuan,Tianyu Zhou,Zhenzhou Tang
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Neural and Evolutionary Computing (cs.NE)
备注: 18 pages, including supplementary material. Preprint
Abstract:Large language model (LLM)-based automatic heuristic design (AHD) iteratively proposes and refines heuristics, pairing design rationales with executable code. Task-specific evaluators assess programs; execution outcomes and performance scores guide search. Many AHD systems keep the generator frozen; EvoTune and Co-Evolution of Algorithms and Language Model (CALM) instead update it from evaluated candidates. When such outcomes drive reinforcement learning with verifiable rewards (RLVR), they create a search-coupled loop: the evaluated candidate stream supplies both search-state updates and training signals for the model that generates future candidates. Yet validity and performance do not uniquely determine useful model updates; converting them into learning signals must account for the prompt and evolving search state that produced each candidate. We formulate online post-training of small open-weight LLMs in AHD as context-dependent signal construction and develop alternative mappings from program validity, task performance, and generation context to update signals. Using shared evaluated rollouts and matched update budgets, controlled experiments across AHD tasks and model families compare these mappings with online post-training baselines, testing their effects on validity, performance among valid proposals, and the yield of valid proposals that improve under contextual comparisons. Complementary checkpoint, frozen-search, and live-system evaluations assess whether proposal-level gains appear in updated checkpoint behavior and subsequent search, rather than arising solely from accumulated search state. A resource-matched comparison under pre-specified cost accounting tests whether online updating adds value beyond additional search with a frozen generator. Together, this design avoids treating end-to-end search gains alone as evidence of stronger heuristic-design capabilities.
[AI-222] Feature Information Dynamics in Diffusion NEURIPS2026
链接: https://arxiv.org/abs/2610.08626
作者: Jia-Shu Pan,Tao Zhang,Yufei Huang,Yanjun Sheng,Tailin Wu
类目: Machine Learning (stat.ML); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: Accepted as poster at NeurIPS 2026. 28 pages, including references, appendices, and checklist
Abstract:Diffusion models generate data through a continuum of denoising problems, and are widely observed to reveal coarse structure before fine detail. Yet, this intuition is mostly empirical and qualitative. We introduce feature information dynamics, an information-theoretic framework for localizing when a feature is generated during diffusion. Using the I-MMSE identity, we connect the rate of feature mutual information change to a gap between optimal unconditional and feature-conditional denoising losses, yielding practical estimators for feature information density. We further develop a chained decomposition that separates shared from incremental information in a feature hierarchy. We use this framework first to quantitatively confirm spectral autoregression in pixel diffusion, and then to extend the analysis beyond frequency: under a class \to mask \to Canny conditioning chain, the per-feature information densities differ across pixel, SDVAE, VAVAE, and RAE, exposing fundamental differences between these representations and suggesting that ordered generation could be beneficial for training diffusion models. Our code is available at this https URL.
[AI-223] Navier-Stokes lost in translation: Why Lean verification of AI autoformalisation does not guarantee correct natural language proofs
链接: https://arxiv.org/abs/2610.08144
作者: Alexander Bastounis,Fabian Circelli,Anders C. Hansen
类目: Analysis of PDEs (math.AP); Artificial Intelligence (cs.AI); Logic (math.LO)
备注: 25 pages, 4 Figures
Abstract:Autoformalisation is increasingly used to verify mathematical texts, including those generated by AI, as in OpenAI’s announced proof of blow-up of solutions to the Navier-Stokes equations. In this process, an AI system translates the text from a natural language (NL) into a formal language such as Lean. Once this translation is done, the argument expressed in the formal language can easily be mechanically verified. The purpose of this article is to demonstrate why this process may offer no confidence in the original NL argument, owing to the various difficulties in performing the translation semantically faithfully. In particular, we highlight that the problem of resolving ambiguities in mathematical NL text, which is necessary in order to provide semantically faithful translation, is arbitrarily high up in the Solvability Complexity Index (SCI) hierarchy/arithmetical hierarchy (the SCI = \infty ). Hence, informally, providing semantically faithful AI autoformalisation is harder than any computational problem including the Halting problem (which has SCI = 1 ). To demonstrate the effect of this result we provide several examples of AI mistranslations of NL statements and proofs into Lean in practice, resulting in mismatches between NL proofs and their Lean `verifications’. These include OpenAI’s announced Navier-Stokes proof. In particular, we show that the formalised Lean proof does not correspond to the NL proof of blow-up of solutions to the Navier-Stokes equations.
[AI-224] A self-learning scientific agent for X-ray diffraction
链接: https://arxiv.org/abs/2610.07862
作者: Bin Cao,Huichi Zhou,Runyu Yang,Jingsong Li,Shuchen Sun,Yan Song,Hanyu Gao,Zhongwei Yu,Tong-Yi Zhang,Jun Wang
类目: Materials Science (cond-mat.mtrl-sci); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:
Abstract:A central challenge for scientific agents is to turn analytical experience into reusable expertise grounded in physical evidence. Here we introduce Gan Jiang, a self-learning agent for powder X-ray diffraction built on a diffraction-analysis ecosystem we developed: XMatcher, XQueryer, XDecomposer and WPEM. Together, these engines span phase identification, multiphase decomposition and physics-constrained whole-pattern modelling. Gan Jiang converts analytical experience into executable skills by diagnosing failures, revising skill instructions and code, and validating revisions before reuse, without retraining the language model or changing the underlying physical models. Skills selected using development data and frozen before held-out evaluation achieve higher refinement scores than the original expert-designed skills across FullProf, GSAS-II and PyWPEM. The agent resolves strongly overlapping reflections, quantifies a five-phase ancient Egyptian cosmetic, tracks lattice evolution in an operating battery and compares atomic configurations in a disordered oxide catalyst. On DeltaXRDbench, it leads the evaluated methods in single- and multiphase identification across simulated and experimental data. Without supplied composition, single-phase top-1 accuracies reach 96.30%, 81.78% and 40.83% on MP500, RRUFF and opXRD, respectively, compared with 58.00%, 58.47% and 26.45% for the strongest comparator. These results demonstrate how an integrated scientific tool ecosystem can support agents that extract structural knowledge from measurements while accumulating validated analytical expertise that transfers to new samples.
[AI-225] Linear Fitness Subspace in Protein Language Models Enables Sample-Efficient Directed Evolution EMNLP2026
链接: https://arxiv.org/abs/2610.07607
作者: SiYuan Ma,Canran Xiao,Zikai Xiao,Albert Gao,Liang He,Xuan-Yu Wang,Shuying Cao,Xiaojun Jia
类目: Quantitative Methods (q-bio.QM); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: Accepted to Findings of the Association for Computational Linguistics: EMNLP 2026
Abstract:Model-guided directed evolution seeks to identify high-fitness protein variants under limited oracle budgets. Protein language models (PLMs) provide rich representations for this task, but task-agnostic zero-shot scores can be misaligned with a target assay, while supervised search in high-dimensional embedding spaces can make surrogate modeling and uncertainty estimation sample-inefficient. We propose the Linear Fitness Subspace (LFS) hypothesis: within mutation-induced residue-level representation changes, a compact, assay-specific set of directions makes fitness variation linearly accessible from few labeled variants. This is a local, supervision-recoverable statement rather than a claim that protein fitness landscapes or global PLM geometry are universally linear. Building on this observation, we introduce Subspace-Guided Evolutionary Search (SGES), which estimates an LFS from a small initial sample and performs surrogate modeling, uncertainty estimation, and acquisition in the learned subspace. Across 10 core ProteinGym assays, 87 extended static-validation assays, and an 18-assay budgeted-search evaluation, SGES improves fitness prediction and search efficiency over zero-shot PLMs and recent ML-guided protein optimization baselines. Controlled comparisons with PCA, random projections, label-shuffled PLS, classical mutation features, and acquisition ablations further isolate the benefit of a fitness-aligned site-delta coordinate.
[AI-226] DeepAJM: Deep Association Joint Model for Irregularly Sampled data
链接: https://arxiv.org/abs/2610.07388
作者: Barsha Halder,Jeffrey A. Thompson
类目: Applications (stat.AP); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Machine Learning (stat.ML)
备注:
Abstract:Joint Models simultaneously model longitudinal and survival outcomes, leveraging patterns in patients’ longitudinal trajectory to improve the prediction of survival outcomes. The classical parametric joint models, however, rely on fixed parametric assumptions, making them susceptible to bias under model misspecification and smaller sample sizes. We propose a deep joint model, DeepAJM, that does not require any parametric assumptions, while retaining a partially interpretable, per-longitudinal-outcome association structure. The joint model uses an encoder-decoder (sequence-to-sequence) architecture to learn the latent structure in patients’ time-varying covariate trajectories. The model links the longitudinal processes to the survival processes through a learned interpretable association structure, in which each longitudinal output from the decoder gets remodulated by baseline covariates before it contributes to the risk scores from the survival head of the architecture. The model was evaluated on three datasets ( a cardiovascular-disease EHR cohort, a primary biliary cirrhosis (PBC2) dataset, and a simulated dataset) against a classical parametric joint model, TransformerJM, DA-LSTM and a Cox-based survival-only model. All models were assessed using C-index, integrated brier score (IBS), time-dependent AUROC, and time-dependent AUPRC. Our model achieved the best discrimination in terms of the C-index, time-dependent AUROC, and AUPRC across all datasets.
[AI-227] CrystalJev: thinking fast and slow with atomistic foundation models for materials discovery
链接: https://arxiv.org/abs/2610.06985
作者: Peng Kang,Zhen Li,Yu Liu,Lei Zheng,Huibin Xu
类目: Materials Science (cond-mat.mtrl-sci); Artificial Intelligence (cs.AI); Computational Physics (physics.comp-ph)
备注: 43 pages, 6 main figures, 5 Extended Data figures, 1 Extended Data table; Supplementary Information included
Abstract:Atomistic foundation models triage millions of hypothetical materials but are used as slow simulators, their thresholded energies taken at face value. They are better read as fast decision-makers. CrystalJev queries a frozen interatomic potential once per unrelaxed structure and answers typed questions with calibrated probabilities, finite-sample guarantees and a rule for when to think slowly. Across 65 Matbench Discovery models, a ‘stable’ call is a probability in disguise, explained by a model’s errors and the candidate population. Once trained, one forward pass decides nearly as well as a relaxation at a thirtieth of its cost, and a value-of-information theory sends slower computation only where decisions can change. The same layer answers electronic, mechanical and molecular questions. In a registered prospective test with 700 new density-functional calculations, single-pass forecasts calibrated only on existing data over-stated the stable fraction of unseen candidates (5.8%) by at most 2.1 percentage points.
[AI-228] When Can World Models Recover Physical Laws?
链接: https://arxiv.org/abs/2610.06877
作者: Ye Yuan,Jun Liu
类目: Other Statistics (stat.OT); Artificial Intelligence (cs.AI); Information Theory (cs.IT); Machine Learning (cs.LG); Systems and Control (eess.SY)
备注:
Abstract:Accurate prediction does not establish that a world model has recovered a physical law: distinct dynamics can generate identical records under the same observation protocol. We formulate law recovery on a fixed physical domain under an explicit catalog of experiments, sensor uncertainty, and an acquisition budget. A rate–distortion converse separates the information needed to describe a law from the information the apparatus can reveal. Its constructive counterpart gives a finite response codebook and an explicit decoding budget. On compact world classes, uniform recovery is possible exactly when every pair of different laws is experimentally distinguishable; equivalently, the apparatus can recover all the entropy of every finite law source. An inverse response modulus quantifies stability. For Lipschitz fields on a d -dimensional state–action domain, noisy full-state readouts after resets require minimax budget \Theta(\varepsilon^-(d+4)/2) for squared law error \varepsilon , compared with \Theta(\varepsilon^-(d+2)/2) for direct field observations. Exact crossing-time symmetries establish the lower bound under adaptive experiment selection and arbitrary durations with constant inputs. Reproducible synthetic cases illustrate the separate roles of intervention, calibration, and repeated measurement. Together, the results identify which evidence supports a claim of physical-law recovery and the cost of acquiring it.
机器学习
[LG-0] QF3: Fast Flow RL with Filtered Q-Gradients
链接: https://arxiv.org/abs/2610.08789
作者: Chung Min Kim,Brent Yi,David McAllister,Hongsuk Choi,Himanshu Gaurav Singh,Jinkun Cao,Ken Goldberg,Pieter Abbeel,Carmelo Sferrazza,Angjoo Kanazawa
类目: Robotics (cs.RO); Machine Learning (cs.LG)
*备注: Project page: this https URL
Abstract:Flow policies have become a standard policy class for learning robot behaviors from demonstrations, but reinforcement learning is still critical for improving pre-trained flow policies or learning them from scratch through interaction. We introduce QF3 (Fast Flow RL with Filtered Q-Gradients), an online off-policy RL algorithm that trains a flow policy with flow matching plus the critic’s action gradient, backpropagated through a one-step prediction of the flow’s output. To keep updates where the critic and this prediction are reliable, QF3 applies the critic gradient only to action dimensions that stay near the replay action. To our knowledge, QF3 is the first off-policy flow RL method to train humanoid locomotion policies from scratch and transfer them zero-shot to hardware. Paired with a high-throughput off-policy training recipe, it trains humanoid locomotion and motion-tracking policies with a 10x wall-clock speedup over FPO++, a recent on-policy flow RL method. We further apply QF3 to fine-tune pretrained flow-based manipulation policies on both ABC-Sim and Robomimic tasks. These results suggest that QF3 can both learn robot policies from scratch and refine those acquired from demonstrations. Website: this https URL
[LG-1] Conformal Prediction Sets Quantify Information Gain: A Theoretical Perspective
链接: https://arxiv.org/abs/2610.08785
作者: Kevin Zhang,Stephen Bates
类目: Machine Learning (cs.LG)
*备注:
Abstract:Conformal prediction is a popular tool for uncertainty quantification that outputs prediction sets with finite-sample coverage guarantees. While prediction set size is commonly used as a heuristic measure of uncertainty, the information-theoretic basis for this interpretation remains poorly understood. In this work, we provide such a foundation using a decision-theoretic generalization of entropy tailored to set-valued prediction. In particular, we introduce a family of generalized information measures based on the size and coverage of conformal prediction sets. Notably, Shannon mutual information admits an exact integral representation in terms of these measures. We then show that, in standard classification settings, the reduction in conformal set size from additional information (i) is sandwiched between calibration-dependent members of this family and (ii) obeys a data processing inequality, both up to finite-sample calibration and model error terms. Together, our results formally relate conformal prediction to classical information-theoretic quantities and justify using set-size reduction as an information gain metric. Empirically, we validate our theory across 11 classification settings and show that set-size reduction and Shannon mutual information can rank features differently in a greedy feature selection experiment.
[LG-2] Rapid Fredholm stabilization of the Kuramoto–Sivashinsky equation with unrestricted spatially-varying anti-diffusion
链接: https://arxiv.org/abs/2610.08764
作者: Luke Bhan,Miroslav Krstic,Yuanyuan Shi
类目: ystems and Control (eess.SY); Machine Learning (cs.LG); Analysis of PDEs (math.AP); Optimization and Control (math.OC)
*备注: 46 pages
Abstract:We develop the first feedback design for rapid stabilization of the Kuramoto–Sivashinsky equation with a spatially varying anti-diffusion coefficient. For constant coefficients, the single-input Fredholm design of Coron and Lü (2015) excludes a discrete set of values at which repeated unstable eigenvalues cause a loss of controllability. We overcome this obstruction by introducing a second boundary input and assigning the two inputs distinct roles. The key idea, inspired by Heymann’s Lemma, is to use the boundary value u(0,t) entirely for a pre-feedback that renders the modified plant controllable through the curvature input u_xx(0,t) . The latter input then stabilizes the plant through a Fredholm backstepping transformation. We show that two inputs suffice for controllability and are necessary when the plant has an unstable double eigenvalue. However, the Fredholm kernel still must be approximated for implementation. Hence, to enable kernel and gain approximation, we prove continuity of the coefficient-to-gain design map on compact admissible design classes. Unlike Volterra-based continuity proofs using successive approximations, our proof uses the modal representation to control the spectral data, the inverse coefficient system, and the tails of the kernel and gain series. This yields a single neural operator approximation of the gain to any prescribed L^2 accuracy across the class. Finally, we establish rapid local stabilization of the nonlinear closed-loop system under both the exact gains and sufficiently accurate approximations. We conclude with numerical results that illustrate prescribed decay rates and the computational cost of the approximations. In particular, we train a Fourier neural operator that achieves typical relative gain errors of approximately 0.1% and stabilizes all held-out cases tested, including a plant with an unstable double eigenvalue.
[LG-3] Neural Petri flows for chemical reactions
链接: https://arxiv.org/abs/2610.08750
作者: Jose Eduardo Escrig Molina,Daniel Probst
类目: Machine Learning (cs.LG); Chemical Physics (physics.chem-ph); Quantitative Methods (q-bio.QM)
*备注: 30 pages, 3 figures, 19 tables
Abstract:Petri nets have been used to describe chemical processes such as this http URL map well to chemistry: Places are the bonds between atoms and the free valence of each atom, a token is a unit of bond order, a transition forms or breaks a bond, the conserved quantities are the valence budgets of the atoms, and the enabling rule is the valence rule. These semantics are not guaranteed by learned models of reactions or neural networks that are built on Petri nets that use the net as a scaffold for message passing. Here, we ask what architecture remains a Petri net for every value of its weights. We find the answer in the theory, where all semantics of a net share the firing form m^\prime=m+C\sigma , locality, as enabling reads only the inputs of a transition, and the enabling rule, and we prove that conservation forces the firing form and that non-negativity forces the enabling rule on local rate laws. This leaves free the rate law, which is the propensity of each transition to fire. We introduce Neural Petri Flow, which learns this rate law, or a readout for classification, and hard-wires the rest as parameter-free layers. On what we denote a valence net, atom mapping, reaction classification, and forward prediction become three tasks on one firing vector. Without training, the minimum firing vector maps 88.8% of the curated Golden set against 85.6% for RXNMapper, and 88.7 against 77.9% of the enzymatic reactions of EnzymeMap. On USPTO-480K, NPF trained on these firing vectors predicts 87.7% of the products and 67.4% when trained on a 1% subset of the training reactions. EC numbers of ECREACT are predicted at the third level for 90.2% of reactions, 5.6 points ahead of the best published method. With electrons as tokens, the same token game predicts 90.5% of the elementary steps of FlowER first, ahead of the published baseline, and every top-1 prediction is a valid molecule without a filter.
[LG-4] Linear Bandits under Exact Sliding-Window Constraints
链接: https://arxiv.org/abs/2610.08745
作者: Seyed Mohammad Hadi Hosseini,Yasin Abbasi-Yadkori,Sattar Vakili
类目: Machine Learning (cs.LG)
*备注: 53 pages, including supplementary material; 8 figures and 6 tables
Abstract:We study linear bandits under exact sliding-window constraints, where every consecutive block of actions must belong to a prescribed feasible set. In the offline setting, where the reward function is known, we show that convexity and cyclic-shift invariance make a stationary solution optimal when w\mid T and within an additive O(w) gap otherwise. In the online setting, we show that geometric structure alone is insufficient for learning, and sublinear regret can be impossible. We introduce a transition diameter \tau that quantifies feasible reachability and develop a rare-switching OFUL algorithm with regret \widetildeO(d\sqrtT+\tau d+w) against the offline-optimal feasible trajectory. Finally, we remove cyclic invariance and consider general sliding-window constraints, where optimal behavior may be non-stationary. We represent recent action history as the state of a finite-memory control problem and introduce a history-state diameter D that measures feasible communication between viable histories. Combining optimistic remaining-horizon planning with rare policy updates, we obtain a regret bound of \widetildeO(d\sqrtT+dD+w) . We evaluate our approach on real-world and synthetic benchmarks, showing that it maintains exact feasibility while achieving reward and regret comparable to baselines with substantially fewer policy updates.
[LG-5] On the Computational Tractability of Robust Bandits
链接: https://arxiv.org/abs/2610.08740
作者: Vanessa Kosoy,Vinayak Pathak
类目: Machine Learning (cs.LG)
*备注:
Abstract:Learning when the environment does not belong to the learner’s hypothesis class is typically handled using agnostic learning guarantees. However, for anything beyond supervised learning, agnostic guarantees are difficult to come by. Recently, imprecise bandits (Kosoy, 2025) (later renamed to robust bandits in Appel and Kosoy, 2025) were introduced as another approach to unrealizable learning in the bandits setting and a \Theta(\sqrtT) regret learner was shown for a large class. However, no computational guarantees were provided. In this paper we identify a special case that admits a polynomial-time learner with \tildeO(\sqrtT) regret. We also show that several small generalizations of this special case are NP-hard thus indicating that the special case is at the boundary of what is tractable. It has been recently suggested (Kosoy, 2018) that computationally efficient learners for unrealizable learning problems are crucial for solving the AI alignment problem. This work is a small step in that direction.
[LG-6] Optimal and Efficient Online Inverse Optimization
链接: https://arxiv.org/abs/2610.08735
作者: Anupam Gupta,Guru Guruganesh,Honghao Lin,Vahab Mirrokni,Renato Paes Leme,David P. Woodruff
类目: Machine Learning (cs.LG); Data Structures and Algorithms (cs.DS)
*备注:
Abstract:In online inverse linear optimization, a learner recommends an action and then observes the choice of an expert who maximizes a fixed, unknown linear objective on \mathbbR^d ; the goal is to learn to optimize this objective without observing it. Sakaue recently obtained the optimal regret O(\sqrt d) with a randomized algorithm making (dT)^O(d) linear optimizations per round, and asked whether it can be attained in polynomial time. We answer positively: our deterministic algorithm has regret O(\sqrt d) for every horizon T and runs in time polynomial in d and T . It is a variant of the variable-metric algorithms of Sakaue et al.\ and Cai et al., in which a metric update is revoked once the query point moves far enough from where the update was made.
[LG-7] GeneICL: A Tabular Foundation Model for Bulk Transcriptomics
链接: https://arxiv.org/abs/2610.08694
作者: Michael Bohl,Alexander Theus,David Wissel,Valentina Boeva
类目: Machine Learning (cs.LG)
*备注:
Abstract:Gene expression is widely measured in biomedicine, yet clinical outcome prediction remains challenging due to high dimensionality, strong feature correlations, and limited labeled data. Large self-supervised transcriptomic foundation models often fail to outperform simple supervised baselines. Tabular foundation models offer an alternative through in-context learning, but are typically pretrained on generic synthetic data rather than transcriptomic structure. We ask whether transcriptomics-aware pretraining, rather than scale, is the missing ingredient. Towards this end, we introduce GeneICL, a 4.2M-parameter tabular foundation model combining a semi-synthetic pretraining prior built from measured bulk expression profiles with a parameter-efficient recurrent architecture. We further enable right-censored survival prediction via a training-free reduction to regression using Cox partial-likelihood residuals. We evaluate GeneICL on 80 clinical outcome-prediction tasks spanning classification, regression, and survival. Tabular foundation models consistently outperform self-supervised transcriptomic models, while GeneICL achieves the best overall rank among evaluated foundation models and tuned baselines. GeneICL does so with up to 387 \times fewer parameters, no gradient updates at inference, and predictions within seconds on a laptop CPU.
[LG-8] Probabilistic Counterfactual Inference for Discrete Outcomes in Gaussian-Process Causal Models
链接: https://arxiv.org/abs/2610.08689
作者: Juliette Sinnott,Amir-Hossein Karimi,Mohammad Kohandel
类目: Machine Learning (cs.LG)
*备注:
Abstract:Counterfactual inference in Gaussian-process structural causal models (GP-SCMs) has been developed primarily for continuous endogenous variables, limiting applicability to causal graphs that contain discrete child nodes with continuous parents. We introduce a unified probabilistic framework for counterfactual inference with heterogeneous variable types by pairing GP predictors with explicit exogenous noise mechanisms. For discrete outcomes, we derive exact conditional noise-abduction procedures using a uniform threshold for binary variables, a Gumbel-max race for nominal categories, and a latent Gaussian cut-point model for ordinal ones. In each case, we propagate abducted noise through interventions while accounting for posterior uncertainty in the GP latent functions, and prove that the resulting mechanisms reproduce the fitted model’s observational and interventional distributions. On synthetic SCMs with known ground-truth counterfactuals, we evaluate estimation accuracy, consistency, and robustness to coupling misspecification. A key finding is that applying a categorical coupling to ordinal data inflates counterfactual error roughly threefold even when observational fit remains comparable, and that this error does not diminish with more data. As the training set grows, the fitted structural equation converges to the truth while the counterfactual error flattens onto a floor. In the reverse direction, forcing a false order onto nominal data instead degrades the fitted equation itself. The choice of coupling must therefore be justified on structural grounds rather than read off the fit.
[LG-9] Variance-Optimal Off-Policy Evaluation with Conjunct Effect Modeling
链接: https://arxiv.org/abs/2610.08677
作者: Nicolò Felicioni,Michael Benigni,Maurizio Ferrari Dacrema,Paolo Cremonesi
类目: Machine Learning (cs.LG)
*备注:
Abstract:Off-policy evaluation (OPE) for contextual bandit policies becomes challenging when action-level importance weighting incurs excessive variance. Doubly robust (DR) estimation remains unbiased under common support but retains these high-variance action-level weights. A prior estimator, Off-policy evaluation with Conjunct Effect Model (OffCEM), replaces them with more stable cluster-level weights, at the cost of relying on local correctness of the reward model. In this paper, we show that, under the assumptions required by DR and OffCEM, there exists an unbiased family of estimators that interpolates between OffCEM and DR. Building on this result, we propose the Variance Optimal-CEM (VOCEM) estimator, which selects the interpolation coefficient to minimize variance. We derive the population-optimal coefficient in closed form and show that the resulting estimator has variance no larger than either endpoint, OffCEM or DR. Experiments in controlled synthetic settings and on two large-action benchmarks show that VOCEM improves upon both endpoints in all 23 evaluated conditions, exhibiting greater stability and empirical robustness.
[LG-10] Multi-Label Perceptual Bug Detection in Video Games using Deep Learning on Gameplay Footage
链接: https://arxiv.org/abs/2610.08593
作者: Nahian Rifaat,Felix Morosov,Loutfouz Zaman
类目: Machine Learning (cs.LG)
*备注:
Abstract:Traditional approaches for automated bug detection in video games, such as manual testing, can be beneficial for the improvement of quality assurance, but they can be expensive and time-consuming. The scarce number of tools available to detect multiple perceptual bugs in the same video frame introduces detection challenges for automated bug detection tools in real-world scenarios. We propose a deep learning model for multi-label perceptual bug detection and compare it against video classification models such as Inflated 3D ConvNet and 3D ResNet. Our proposed model, ResNet-BiLSTM, achieved an F1 score of 85.78% on the benchmark dataset. Our results demonstrated that temporal dependency modelling is beneficial for accurate video-based bug detection. We believe this work with multi-label perceptual bug detection on gameplay videos will help save resources spent on manual testing workloads in video games. Furthermore, we introduce a new dataset with multi-label perceptual bugs in this work. The dataset contains 77,969 video clips across different genres of games with approximately 1.2 million frames, containing combinations from 5 classes of bugs in the same video frame.
[LG-11] CNet: A Complex-Valued Deep Learning Framework with Wirtinger Autodifferentiation and FFT–Hadamard Convolution
链接: https://arxiv.org/abs/2610.08592
作者: Marcel Crasmaru
类目: Machine Learning (cs.LG); Mathematical Software (cs.MS); Optimization and Control (math.OC)
*备注:
Abstract:CNet is a C++/CUDA framework for building and training deep complex-valued neural networks (CVNNs) and, more generally, for optimizing complex-valued functions by gradient descent with Wirtinger (CR-calculus) derivatives. It takes a physics-native stance: a network is a cascade of complex – and often unitary (the DFT) – operations acting on an amplitude vector, and classification is a Born-rule measurement p_k = |z_k|^2 / |z|^2 rather than a softmax over real logits. Every layer ships a CPU reference and a CUDA kernel checked against finite differences, and the computation graph is cloned across the batch for GPU execution. On top of the base layers we add signal-processing primitives that turn the identity conv(x,k) = IFFT(FFT(x) . FFT(k)) into a learnable complex convolutional network, together with a true-Adam optimizer and a reduced-memory inference mode. We report three studies. First, a fully complex-valued, FNet-style causal sequence model built on a new O(N \log N) causal Fourier mixer – a triangular-masked DFT evaluated by a Bluestein / chirp-z factorization: once properly tuned it matches or exceeds a parameter-matched real-valued causal FNet on character-level language modeling, reaching the real model’s converged quality in under half the training steps. Second and third, bottleneck analyses on radio-modulation classification (RML2016.10a) and the Fourier phase problem of coherent-diffraction imaging, which isolate exactly where complex-valued networks still need new operators. Across all three the complex formulation provably learns the physically correct structure. Code: this https URL Subjects: Machine Learning (cs.LG); Mathematical Software (cs.MS); Optimization and Control (math.OC) Cite as: arXiv:2610.08592 [cs.LG] (or arXiv:2610.08592v1 [cs.LG] for this version) https://doi.org/10.48550/arXiv.2610.08592 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[LG-12] Random Feature Gaussian Process Attention: Linear-Time Probabilistic Attention with Calibrated Uncertainty
链接: https://arxiv.org/abs/2610.08578
作者: Amir Mohammad Mahfoozi,Zi Yang,Ying Li,Michael Minyi Zhang
类目: Machine Learning (cs.LG)
*备注: 14 pages, 3 figures, 3 tables
Abstract:Transformers provide a state-of-the-art modeling framework, yet poor calibration limits their reliability in safety-critical applications. A promising direction addresses this issue by interpreting attention as a Gaussian process (GP) posterior, which enables principled uncertainty calibration but incurs cubic complexity in sequence length due to the inversion of the kernel; although decoupled GP variants reduced the cost to quadratic, the computation remains prohibitive in practice. In this paper, we propose the plug-and-play random Fourier feature Gaussian process attention (RFF-GPA) module, which represents the attention as a GP with a stationary kernel approximated by random Fourier features. This low-rank approximation results in linear-time complexity for approximating the posterior mean and variance, making it far more scalable compared to previous work. Empirical results on multiple real-world datasets show that our attention module improves calibration while maintaining predictive accuracy, and simultaneously reduces computational complexity to linear in the sequence length.
[LG-13] Singular Value Decomposition: A Geometric Rediscovery Where Proofs Become Algorithms
链接: https://arxiv.org/abs/2610.08565
作者: Paul Agron
类目: Machine Learning (cs.LG); History and Overview (math.HO)
*备注: 31 pages, 7 figures. Expository article
Abstract:This article is a geometric rediscovery of the singular value decomposition, with a further claim: the construction it builds is the machinery behind much of machine learning. The same argument that answers an idle question about ellipses is the algorithm behind principal component analysis, kernel methods, and PageRank, and it is not only the results that transfer but the proofs themselves, run as procedures. The usual introduction states A = U\Sigma V^T and justifies it via the spectral theorem applied to A^T A . This is correct but unilluminating, since it assumes a powerful theorem to reach a result that is, in the end, about ellipses. Part I reverses the order. A linear map sends the unit circle to an ellipse; one asks which input directions map to its axes, and finds, example after example, that they are perpendicular. In the plane this can be watched: rotate a frame, track how far its images are from perpendicular, and a sign change forces a frame where they are exactly perpendicular, which is also where the map stretches hardest. Maximizing the stretch and recursing generalizes this to n dimensions, with singular values falling out in order, and the construction proves the spectral theorem rather than assuming it. Part II puts each construction to work: maximize-and-recurse becomes the power method and PageRank; the lemma locating the maximizer becomes the stopping rule of gradient descent; the duality between A^T A and A A^T becomes the transport at the heart of kernel PCA. Each connection is stated with its boundary, saying what the decomposition supplies and where another idea takes over. Prerequisites are the standard sophomore sequence, and the worked examples are small enough to check by hand. Comments: 31 pages, 7 figures. Expository article Subjects: Machine Learning (cs.LG); History and Overview (math.HO) Cite as: arXiv:2610.08565 [cs.LG] (or arXiv:2610.08565v1 [cs.LG] for this version) https://doi.org/10.48550/arXiv.2610.08565 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[LG-14] Valid for Free: Homophily-Gated Conformal Prediction for Training-Free Node Classification with Tabular Foundation Models
链接: https://arxiv.org/abs/2610.08564
作者: Nguyen Duy Long,Phung Minh Hien,Nguyen Trong Viet,Nguyen Thai Anh
类目: Machine Learning (cs.LG)
*备注: 6 pages, 4 figures, 1 table
Abstract:Tabular foundation models (TFMs) can classify the nodes of a graph without training on it, by reading node and neighborhood features as table rows next to labeled context rows. Work in this line reports predictive performance, not conformal coverage or prediction-set size. To our knowledge, we give the first reliability study of the setting, with TabICL as the TFM and half of each graph as labeled context. As for any predictor fixed before calibration, a frozen in-context predictor makes split conformal prediction exactly valid in finite samples, with no training, validation fold, or tuning on the target graph. An audit across ten graphs then shows that the training-free TabICL posterior has lower expected calibration error (ECE) than GCN with temperature scaling (GCN+TS) on nine of them. Its mean ECE over the ten graphs is 0.019, about 35 percent below the 0.029 of GCN+TS. We also introduce HG-DAPS, a training-free diffusion score whose homophily gate reads only the in-context labels, so the guarantee still holds. Relative to adaptive prediction sets (APS), it reduces mean set size by 5.8 to 17.1 percent on six homophilous graphs and changes it by under 1 percent on four heterophilous ones. On two binary, class-imbalanced graphs, a pre-registered trap case shows that gating on raw rather than adjusted homophily lowers coverage among low-homophily nodes by 0.27 and 0.12. Marginal coverage stays at the nominal 0.90 and masks this drop.
[LG-15] Reinforcement Learning for Hierarchical Reasoning Rewards: Minimax-Optimal Rates with Transformers
链接: https://arxiv.org/abs/2610.08561
作者: Naoki Nishikawa,Taiji Suzuki
类目: Machine Learning (cs.LG); Machine Learning (stat.ML)
*备注:
Abstract:Reinforcement learning (RL) has become a standard tool for post-training language models on reasoning tasks, where the policy is updated by reward feedback while exploring the space of responses. Despite its empirical success, theoretical understanding of RL post-training remains limited, in particular of why on-policy exploration combined with a neural reward model is effective. In this paper, we address this question by modeling the reward as a hierarchical function on the response space: the reward consists of infinitely many local components, each of which becomes relevant only after the preceding ones have been resolved. We show that a natural Transformer-based actor–critic algorithm, which alternates between sampling from the current KL-regularized policy, fitting a Transformer critic to the observed rewards, and updating the policy, achieves the minimax optimal rates in the query budget and in the regularization strength up to logarithmic factors, and is minimax optimal for a fixed number of prompts. In contrast, we prove that sampling from the fixed reference distribution, as in offline reward modeling, can limit regret decay to a logarithmic rate. These results show that on-policy exploration progressively zooms in on the region where the reward is concentrated, and quantify its benefit for RL post-training.
[LG-16] How Bregman Divergences Shape Shampoo
链接: https://arxiv.org/abs/2610.08534
作者: Bing Liu,Wenjie Zhou,Chengcheng Zhao,Hongtao Zhang,Boao Kong,Felix Dangel,Wu Lin
类目: Machine Learning (cs.LG)
*备注:
Abstract:Understanding the principles behind Shampoo has recently guided the development of more effective neural network optimizers. These methods learn a preconditioner by optimizing the Frobenius or Kullback-Leibler (KL) divergence against the gradient second moment. In this work, we investigate how the choice of divergence shapes preconditioning, which remains unclear and blocks further improvements. To do so, we develop a unified Bregman divergence framework that connects all popular divergences, allowing us to study them jointly. Through empirical spectral analysis of gradient second moments, we examine how divergence choice shapes Kronecker approximation and interacts with finite-sample error in preconditioning. We find that some divergences can better compensate for finite-sample underestimation of the empirical second moment, helping explain the differing behavior of their corresponding Shampoo variants. We further validate this explanation through GPT-2 pretraining experiments. By connecting divergence choice to practical training behavior, we believe our framework provides principled guidance for understanding the foundations of, and further improving, Shampoo.
[LG-17] PHBA: Prefix-State Hybrid Block Attention
链接: https://arxiv.org/abs/2610.08527
作者: Ruijie Li,Jiaxi Hu,Shiyu Wang,Yuxuan Liang
类目: Machine Learning (cs.LG)
*备注:
Abstract:Hybrid architectures combining linear sequence models with softmax attention provide an effective balance between efficient long-context modeling and precise token retrieval. Existing designs such as Native Hybrid Attention (NHA) combine compressed long-term states with sliding-window attention, but their exact attention is restricted to a fixed local window. In this work, we introduce Prefix-State Hybrid Block Attention (PHBA), which replaces local sliding-window attention with top-k block-sparse retrieval and couples each retrieved block with a compact prefix state summarizing its preceding context. The prefix states are constructed by a gated linear recurrence at block boundaries and retrieved together with the corresponding token blocks, allowing the model to combine precise long-range evidence with compressed historical context within a unified layer. We further develop a hardware-aware Triton implementation that streams routed token blocks and prefix states without materializing large intermediate tensors. Experiments show that PHBA improves long-context and retrieval performance over strong linear and hybrid baselines while retaining efficient training and inference.
[LG-18] Learning PDE solution operators with variable initial conditions via Latent Dynamics Networks
链接: https://arxiv.org/abs/2610.08475
作者: Stefano Maria Pizzamiglio,Stefano Pagani,Francesco Regazzoni
类目: Machine Learning (cs.LG); Numerical Analysis (math.NA)
*备注:
Abstract:In many-query scenarios, data-driven surrogate models provide an efficient alternative to high-fidelity solvers for simulating physical systems governed by Partial Differential Equations (PDEs). In this context, the Latent Dynamics Network (LDNet) has recently demonstrated remarkable performance in predicting the response of spatio-temporal systems, combining Neural Ordinary Differential Equations with nonlinear dimensionality reduction. However, the original formulation assumes a fixed initial condition, limiting its applicability to many real-world applications where a system evolves from varying starting states. In this work, we overcome this limitation while keeping the end-to-end training procedure of the original LDNet and its encoder-free nature, which preserves its intrinsic independence from spatial resolution and grid topology. We infer the initial latent state directly from a small set of early-time observations, treating latent-state initialization as an adaptation problem, and investigate two strategies: an auto-decoding formulation and a meta-learning approach in which the initial latent state acts as a task-specific context variable. We demonstrate the accuracy of the proposed methods across diverse physical phenomena, spanning advection-diffusion, fluid dynamics, and solid mechanics. Meta-learning markedly accelerates latent-state inference and induces smoother, better-conditioned optimization landscapes, and spontaneously organizes the latent space into a structured representation that reflects physically meaningful features of the underlying dynamics. The coordinate-based decoder enables training from spatially subsampled data while recovering high-resolution solution fields at inference. The resulting approach provides an efficient and resolution-independent surrogate modeling framework for many-query simulations of time-dependent PDEs with varying initial conditions.
[LG-19] Climbing the Design Ladder: Sequential Knowledge Distillation for Early-Stage Circuit Timing Prediction
链接: https://arxiv.org/abs/2610.08457
作者: Reza Moravej,Fahad Rahman Amik,Zhanguang Zhang,Didier Chételat,Yingxue Zhang
类目: Machine Learning (cs.LG); Hardware Architecture (cs.AR)
*备注:
Abstract:Integrated circuit design involves multiple design stages: logic synthesis, floorplanning, placement, and routing, with each stage taking hours to weeks to complete. Discovering timing violations late in this flow forces costly iterations back to earlier stages, wasting computational resources and delaying product launches. While predicting post-routing timing from early-stage data could prevent these failures, existing machine learning approaches struggle with the massive abstraction gap between post-synthesis logical descriptions and post-routing physical layouts. We propose STEP-KD (Sequential Timing Evaluation via Progressive Knowledge Distillation), which leverages intermediate design stages as ``stepping stones’’ for progressive knowledge transfer rather than attempting direct prediction. STEP-KD trains teacher models at the post-routing, post-placement, and post-floorplan stages, then sequentially distills their knowledge to a post-synthesis student model through representation alignment. Experiments on diverse circuits demonstrate that STEP-KD reduces timing prediction error compared to direct distillation and supervised baselines, and in most settings compared to the industry-standard Static Timing Analysis (STA) tool. STEP-KD reduces the weighted mean absolute percentage error of Total Negative Slack prediction to 19.78%, compared with 74.84% for STA. Our proposed method is step forward to identify timing problems earlier, avoiding expensive late-stage redesigns.
[LG-20] Symmetry-Aware Feature Learning: A Polynomial Separation for Multi-Index Models
链接: https://arxiv.org/abs/2610.08420
作者: Jivan Waber,Vanessa Piccolo,Yatin Dandi,Florent Krzakala
类目: Machine Learning (cs.LG); Probability (math.PR); Machine Learning (stat.ML)
*备注: 71 pages, 3 figures
Abstract:We establish a polynomial sample complexity separation between symmetry-aware and symmetry-agnostic feature learning. We study growing-rank multi-index models with high-dimensional Gaussian covariates in \mathbbR^d and r=\Theta(d^\delta) teacher directions forming a cyclic symmetry orbit, where 0\delta1/2 . We compare three ways of exploiting this structure: architectural weight sharing, data augmentation over the full symmetry group, and learning without access to the symmetry. In particular, we analyze a symmetry-tied convolutional network, an untied network, and the same untied network trained with full-group data augmentation, using spherical online SGD with correlation loss. For a class of polynomial links with information exponent p\ge3 , we prove matching sample complexity bounds up to logarithmic factors: the tied and augmented learners achieve weak directional recovery in \widetilde\Theta(d^p-1) samples, whereas the symmetry-agnostic learner requires \widetilde\Theta(rd^p-1) . For the pure quadratic Hermite link, the same separation holds for weak recovery of the teacher subspace, with sample complexities \widetilde\Theta(d) and \widetilde\Theta(rd) , respectively. Thus, full-group data augmentation matches the sample efficiency of architectural weight sharing, and both provide a polynomial advantage over training without symmetry. For p\ge3 , the proof reveals a two-stage mechanism: fluctuations at initialization select one direction in the teacher orbit, after which localized growth amplifies its overlap to the weak recovery scale while competing overlaps remain near their initialization scale.
[LG-21] SSR: Sparse Segment Reduction for Ternary GEMM Acceleration DATE2026
链接: https://arxiv.org/abs/2610.08403
作者: Adeline Pittet,Shien Zhu,Valérie Verdan,Gustavo Alonso
类目: Machine Learning (cs.LG)
*备注: Published in the Proceedings of the Design, Automation Test in Europe Conference (DATE 2026)
Abstract:Large Language Models (LLMs) require substantial computational resources, limiting their deployment on resource-constrained hardware. Ternary LLMs mitigate these demands through weight quantization via ternary values, achieving significant compression often with 50-90% sparsity. However, existing approaches have limitations: methods optimized for ternary weights, such as BitNet, redundant segment reduction (RSR), and its improved version RSR++, do not exploit sparsity structures, while conventional sparse formats neglect ternary characteristics, foregoing dual optimization opportunities. In this paper, we introduce Sparse Segment Reduction (SSR), a ternary matrix multiplication method designed to accelerate the inference of ternary LLMs and general Ternary Weight Networks (TWNs). SSR has a dedicated optimized ternary data format and an algorithm that systematically exploits sparsity patterns through computation trees that scale with the sparsity. SSR provides theoretical gains with asymptotically faster inference than RSR++ for sparsity above 50%, while practical evaluations reveal performance improvements across all sparsity levels. Evaluation results show that SSR achieves 2.1-11.3x speedup over RSR++ on ternary GEMM with 45-95% sparsity. Furthermore, SSR achieves 3.5-6.3x end-to-end speedup and 4.9% of memory saving over RSR++ on the Llama-3 1B model inference. Comments: Published in the Proceedings of the Design, Automation Test in Europe Conference (DATE 2026) Subjects: Machine Learning (cs.LG) Cite as: arXiv:2610.08403 [cs.LG] (or arXiv:2610.08403v1 [cs.LG] for this version) https://doi.org/10.48550/arXiv.2610.08403 Focus to learn more arXiv-issued DOI via DataCite (pending registration) Journalreference: 2026 Design, Automation Test in Europe Conference (DATE), pp. 1-7, 2026 Related DOI: https://doi.org/10.23919/DATE69613.2026.11539457 Focus to learn more DOI(s) linking to related resources
[LG-22] VETTA: Coordinating Turn- and Token-Level Credit Assignment for Multi-Turn LLM Agents
链接: https://arxiv.org/abs/2610.08402
作者: Jiaju Chen,Min Yang,Jinghua Piao,Xiaochong Lan,Xu Xia,Xiangnan He,Yong Li
类目: Machine Learning (cs.LG)
*备注: 16 pages, 6 figures
Abstract:Multi-turn LLM agents often receive sparse task feedback across several interactions, while generating each response token by token. This creates two related credit-assignment questions: which responses helped achieve the outcome, and which generation decisions mattered within each response? Existing methods typically focus on only one level: turn-level methods evaluate complete responses but do not distinguish the decisions within them; token-level methods can propagate feedback across turns but do not explicitly model credit for each response. These complementary limitations motivate learning credit at both levels and coordinating it in a single policy update. We introduce VETTA, a credit assignment method that jointly learns turn- and token-level values through separate heads on a shared lightweight critic. VETTA computes advantages along both temporal sequences and combines each turn advantage with a within-response-centered token residual for PPO updates. Furthermore, to reduce value-learning cost, the critic retains only early Transformer blocks from the pretrained checkpoint used to initialize the actor. On two challenging agent benchmarks, ALFWorld and WebShop, VETTA improves success rates over PPO by 37.5% and 22.3%, respectively, with Qwen2.5-1.5B-Instruct and achieves success rates of 95.5% and 76.0%, respectively, with Qwen2.5-7B-Instruct. Critic-depth comparisons further show strong task performance with substantially lower critic-side computation. These results suggest that a compact shared critic can coordinate turn- and token-level credit to improve agent performance while keeping value estimation efficient. Code is available at this https URL.
[LG-23] Decision-Focused Learning in MDPs: An Occupancy Measure Approach NEURIPS2026
链接: https://arxiv.org/abs/2610.08384
作者: Zihao Zhao,Ashwath K. Karunakaram,Ali Eshragh,Yuexing Li,Kai Wang
类目: Machine Learning (cs.LG)
*备注: Accepted at NeurIPS 2026
Abstract:In this work, we consider decision-focused learning (DFL) for a Markov decision process (MDP), where existing methods differentiate through the KKT conditions of the Bellman equation and require solving a linear system over all state-action pairs, limiting its scalability. We address this by reformulating the MDP as an occupancy measure-based linear program (LP), whose feasible region is induced by predicted dynamics, and we derive a closed-form gradient by identifying the active constraints in the feasible polyhedron via the pivoting algorithm. This occupancy measure-based LP layer raises two challenges: (1) LP’s solution gradient is discontinuous when active constraints change, and (2) the LP backward cost still scales with the state size, which is costly for large or continuous state spaces. We address the challenges with an augmented Lagrangian surrogate and smooth the boundary jumps by random row sketching of the constraints, and a learnable soft state-aggregation layer and its function-approximation generalization that scales the LP to large finite and continuous-state MDPs. Across multiple tasks, our methods reach lower regret than KKT-based DFL and two-stage baselines with significantly lower computation cost. The source code for all experiments is available at this https URL.
[LG-24] Evolutionary One-Step Generators: Fast and Diverse Sampling for Discrete Design
链接: https://arxiv.org/abs/2610.08367
作者: Marcus Vukojevic,Erik Nielsen,Veronica Lachi,Andrea Passerini,Giovanni Iacca
类目: Machine Learning (cs.LG); Neural and Evolutionary Computing (cs.NE)
*备注:
Abstract:Several discrete design tasks, such as molecular discovery, require diverse collections of useful candidates at low computational cost. High validity alone does not guarantee a useful candidate library: repeatedly generating the same valid structures leaves few distinct alternatives. Training for both feasibility and diversity is challenging because many relevant criteria can only be evaluated after hard decoding. To address this challenge, we propose EGO (Evolutionary Generators with One-step inference), a framework for training compact generators directly on discrete outputs. The method combines distribution matching with structural constraints and optional diversity or history-dependent rewards, using antithetic low-rank evolution strategies without requiring criterion-specific differentiable surrogates. Once trained, the generator produces the entire graph in a single neural-network evaluation. On molecular generation benchmarks, our compact generator achieves over 50\times the valid-and-unique yield per estimated dense operation compared to recent one-step flow-map baselines while retaining high chemical validity. In scaffold completion, EGO achieves an observed 44.3\times speedup over MoLeR in generation to SMILES and produces approximately 10\times as many filter-passing proposals within matched time budgets for generation and screening. Beyond chemistry, EGO produces 1.54\times as many distinct held-out elite architectures as relaxed gradient training on NAS-Bench-101. The low generation cost may enable real-time candidate generation across discrete design tasks, supporting interactive exploration of constrained design spaces and rapid construction of candidate sets for downstream evaluation.
[LG-25] Uncertainty Quantification Is Indispensable for Reliable Connectome-Based Graph Learning: A Narrative Review and Case Study
链接: https://arxiv.org/abs/2610.08353
作者: Mansooreh Pakravan
类目: Machine Learning (cs.LG)
*备注:
Abstract:While graph neural networks (GNNs) have shown substantial promise in connectome-based diagnostic classification, deterministic models inevitably suppress pipeline-induced noise and model ambiguities, yielding overconfident predictions. Although uncertainty quantification (UQ) is widely adopted in voxel-level segmentation, its role in connectomic graph learning remains largely unaddressed. This paper presents a comprehensive narrative review of UQ frameworks tailored to connectome graph learning alongside an empirical case study demonstrating the perils of uncalibrated predictions. We delineate sources of aleatoric and epistemic uncertainty across neuroimaging pipelines and review prominent UQ paradigms, from Bayesian approximations and ensemble methods to evidential learning and conformal prediction. In our case study, a temporal Graph Attention Network (GAT) trained on dynamic functional connectivity (dFC) matrices from the SUDMEX CONN dataset achieves 80.0% diagnostic accuracy (F1 = 0.794) for Cocaine Use Disorder. However, a post-hoc uncertainty audit via Monte Carlo dropout reveals severe overconfidence (ECE = 0.127), with misclassified subjects assigned prediction confidences up to 95%. This empirical divergence between discrimination and calibration underscores the confidence paradox in deep connectomics. Our findings establish that rigorous UQ, calibration, and selective prediction mechanisms are indispensable for deploying trustworthy graph-based biomarkers in clinical neuroscience.
[LG-26] Machine Learning for German Redispatch Forecasting under Data Delays and Temporal Distribution Shift ATC
链接: https://arxiv.org/abs/2610.08337
作者: Faraz Shamim(1),Faris Shamim(2) ((1) KIST Medical College and Teaching Hospital, Nepal, (2) OTH Regensburg)
类目: ystems and Control (eess.SY); Machine Learning (cs.LG)
*备注: 15 pages, 4 figures, 3 tables. Code available at this https URL
Abstract:Public redispatch records provide empirical data for grid congestion forecasting, but delayed reporting, zero-inflated distributions, and temporal shift present major modeling challenges. We assess the accuracy and reliability of probabilistic machine-learning forecasts using published German transmission records under experimentally imposed information-age constraints. The benchmark evaluates eight daily series of upward and downward intervention energy across four German transmission system operators from 2021 to 2024 (48,242 eligible records; 354 evaluation dates in 2024). We compare seasonal empirical, regularized autoregressive (ARX), quantile LightGBM, GRU, and Transformer models under a minimum seven-day target-latency constraint. Neural architectures use a zero-censored output head to accommodate exact-zero outcomes. Static, rolling, and adaptive delayed-feedback calibration are evaluated using normalized weighted interval score (nWIS), empirical coverage, and block-bootstrap inference. Raw LightGBM achieved nWIS 0.7952, outperforming ARX (1.0604) and the seasonal baseline (0.8739) by 25.0% and 9.0%, respectively (Holm-adjusted p0.005). Rolling calibration improved LightGBM to nWIS 0.7767 versus 0.8251 for static calibration (p=0.0092), with 91.81% coverage for nominal 90% intervals. The zero-censored Transformer achieved nWIS 0.8161, with no significant difference from LightGBM (p=0.260). However, aggregate coverage concealed substantial undercoverage during high-volume interventions (61.91% coverage among above-threshold events). These results show that boosted-tree models with rolling calibration provide accurate probabilistic forecasts of aggregate redispatch volumes under target delays, while nominal aggregate validity does not ensure reliability during extreme congestion events.
[LG-27] Structure-Aware Graph Abstention for Reliable Selective Forecasting
链接: https://arxiv.org/abs/2610.08322
作者: Jianxiang Xie,Belal Alsinglawi
类目: Machine Learning (cs.LG)
*备注:
Abstract:Selective forecasting abstains on high-risk test windows under a retained-coverage budget. Existing gates such as TEM (Brusokas et al., 2025) score each forecast as a whole; for multivariate outputs, trajectories can look plausible while violating dependencies among variables. We treat instance-level plausibility and relational consistency as distinct reliability axes and operationalize the latter via a learned sparse graph and a Dirichlet-style structural energy E_struct, trained with error-weighted graph regularization and score-error alignment. On seven long-horizon benchmarks and four backbones, structural gating often reduces selective MSE versus TEM at matched coverage, with the largest gains where cross-variable structure appears more informative in our benchmarks; gains are not universal, indicating a complementary abstention signal. Table 1 is a Protocol A ranking diagnostic (seed 2024); three-seed deployable Protocol B on an aligned subset is in Table 3 (full validation-to-test grids: Appendix A).
[LG-28] OxiGen: Oxidation-State-Aware Crystal Generation
链接: https://arxiv.org/abs/2610.08296
作者: Dylan John,Kim E. Jelfs,Alex M. Ganose,Eleonora Giunchiglia
类目: Machine Learning (cs.LG); Materials Science (cond-mat.mtrl-sci)
*备注: 27 pages, 4 figures
Abstract:Generative models have the potential to accelerate inorganic materials discovery by enabling inverse design, but generating experimentally realisable crystals remains challenging. Oxidation states are widely used to assess the compositional validity of crystals and guide inorganic materials discovery. While existing generative models for crystals can generate materials with charge-neutral oxidation-state assignments, they poorly reproduce the distributions of oxidation states observed in synthesised materials. To address this limitation, we propose OxiGen, an oxidation-state-aware crystal diffusion model that explicitly represents oxidation states during generation. OxiGen enforces global charge neutrality by construction using a structured output layer with exact inference over a finite-state automaton. Empirically, OxiGen substantially improves oxidation-state fidelity, generates the highest rate of stable, unique, and novel crystals among evaluated methods, and maintains high compositional validity even under property conditioning.
[LG-29] Scalable extraction and visualization of multi-attribute logical and functional dependencies in tabular data
链接: https://arxiv.org/abs/2610.08287
作者: Chaithra Umesh(1),Arvind Lomrore(4),Neethu D(4),Kristian Seegel-Schultz(1),Saptarshi Bej(1 and 4),Olaf Wolkenhauer(1,2, and 3) ((1) Institute of Computer Science, University of Rostock, Germany, (2) Leibniz-Institute for Food Systems Biology, Technical University of Munich, Freising, Germany, (3) Stellenbosch Institute for Advanced Study, South Africa, (4) School of Data Science, Indian Institute of Science Education and Research, Thiruvananthapuram, India)
类目: Machine Learning (cs.LG)
*备注: 31 pages, 4 figures, submitted to Pattern Recognition Journal
Abstract:Understanding the structural relationships among attributes in tabular data is fundamental to machine learning and pattern recognition. While functional dependency (FD) discovery has been extensively studied, scalable discovery of logical dependencies (LDs), particularly as the number of attributes and dependency order increase, remains underexplored. These dependencies capture non-deterministic, condition-specific relationships among pairwise or multiple attributes. Furthermore, existing approaches do not provide a unified framework for extracting multi-attribute LDs and FDs. To address these limitations, we propose LDTool and HLDTool for extracting and visualizing multi-attribute LDs and FDs from tabular data. LDTool extends dependency discovery beyond pairwise relationships, while HLDTool enables scalable extraction through hypergraph-guided search-space reduction. Experiments on three simulated and eleven real-world datasets demonstrate that the proposed framework extracts meaningful LDs and FDs while improving scalability. LDTool recovers the same FDs as existing FD discovery methods with lower runtime in high-dimensional feature spaces, whereas HLDTool enables dependency discovery in datasets with hundreds of features. The proposed framework provides interpretable visualizations of dependency structures and supports applications in exploratory data analysis and the quantitative evaluation of synthetic tabular data.
[LG-30] Performative Prediction with Selective Labels NEURIPS2026
链接: https://arxiv.org/abs/2610.08272
作者: Giovani Valdrighi,Isabel Valera,Marcos Medeiros Raimundo
类目: Machine Learning (cs.LG)
*备注: Accepted at NeurIPS 2026. Camera-ready version
Abstract:Many social applications of machine learning exhibit performative effects: population behavior changes in response to deployed models. Performative prediction studies this interaction through a distribution map that relates each model to the population distribution it induces. One of the main results in this framework showed that repeated risk minimization (RRM), which updates models by retraining on the most recent data, can converge to a stable model that minimizes risk on its own induced distribution. However, existing analyses typically assume access to the complete distributions of features and labels after model deployment, ignoring the possibility of selective labels: observing labels only for the accepted subset of the population. In this work, we formalize performative prediction with selective labels and show that retraining only on observed data can misguide the retraining procedure and undermine the guarantees of convergence to a stable solution. We then propose a worst-case objective based on knowledge of a confidence interval on the probability of a positive label. Applying RRM to this objective permits us to remain within a bounded distance to the true stable point. Under a sensitivity assumption on the conditional label distribution, we further show how previously accepted data can tighten these confidence intervals over time. Experiments in a lending application with fairness regularization show that our robust optimization approach closely matches the performance of RRM with complete label access.
[LG-31] Reinforcement Learning with Segment Reward Feedback under Linear Function Approximation
链接: https://arxiv.org/abs/2610.08271
作者: Fengxu Liu,Siwei Wang,Gal Dalal,Shie Mannor,Yihan Du
类目: Machine Learning (cs.LG)
*备注:
Abstract:Classical reinforcement learning (RL) assumes that a reward is observed for every visited state-action pair. However, in real-world applications such as autonomous driving, such fine-grained feedback can be costly or difficult to collect, whereas trajectory-level feedback may be too sparse for efficient learning. To provide a general feedback model bridging these two extremes and handle large state spaces, we study RL with segment reward feedback under linear function approximation. Our work answers how the granularity of segment feedback and the choice of segmentation influence learning. For equal-length segments with known transitions, we design algorithms \bitssegd and \edlinucbsegd for binary and sum feedback types, respectively. They adopt posterior sampling with planning to achieve computational efficiency and the E-optimal experimental design to attain near-optimality. Nearly matching lower bounds are established. For equal-length segments with unknown transitions, we develop a unified \seglsvits framework with two instantiations for binary and sum feedback, which carefully integrates the posterior estimated reward parameters into least-squares value iteration. These results reveal a fundamental insight: under binary feedback, increasing the number of segments significantly reduces the regret through an exponential factor, while surprisingly, under sum feedback, the granularity of segments does not affect learning much. Finally, to investigate whether segmenting according to state-action features can further expedite learning, we design an algorithm \uneqsegbitsd that allows arbitrary segmentations. The resulting regret bound shows that under the usual elliptical potential analysis, the influence of state-action features on the regret appears only through logarithmic factors, and equal segmentation achieves the best performance.
[LG-32] Finding the Heads and the Neurons Responsible for Network Information Retrieval in Language Models
链接: https://arxiv.org/abs/2610.08200
作者: Abdul Kadir(1 and 2),Md Mohasin Hossain(2 and 3),Daniel Sonntag(1 and 2) ((1) University of Oldenburg, Oldenburg, Germany, (2) German Research Center for Artificial Intelligence (DFKI), Germany, (3) Saarland University, Saarbrucken, Germany)
类目: Machine Learning (cs.LG)
*备注: 13 pages, 2 figures, 12 tables
Abstract:We ask whether specific attention heads, and more finely specific neurons inside those heads, are responsible for recognizing that a language model’s context contains network infrastructure information (a hostname paired with its IP address), and whether that responsibility can be validated causally rather than by correlation alone. At the head level the answer is yes, across five models spanning three architecture families: in every model, a small set of heads (1 to 9 out of 128 to 1152 candidates), found by causal ablation screening and tested for selectivity against matched negative and context-free controls, supports a detector with 99.5–100% held-out accuracy. We then ask whether a head’s responsibility concentrates into one neuron or stays spread across its dimensions; this is model-specific. In one model, the top head’s signal concentrates into a single neuron, found independently by both a causal intervention and a correlational ranking, which agree exactly (AUC = 1.000, matching the full head). In another, the single clean head works as a whole (AUC = 1.000) but the best causally ranked neuron inside it does not (AUC = 0.665), so the responsibility there is spread across the head. The remaining three models fall in between. On an independent dataset collected by a different institution (reverse-DNS records rather than the discovery data), every model’s full-head detector flags 100% of positive records; the single-neuron versions transfer less reliably, and in one model score below chance. Causal head-finding for a specific network-information entity works across models and architectures; how far that finding can be pushed down to individual neurons varies, and needs to be checked for each model.
[LG-33] On the Intrinsic Limited Robustness of Latent-Based Watermarking
链接: https://arxiv.org/abs/2610.08178
作者: Cheng-Han Yeh,Kuan-chun Yu,Cheng-Chang Tsai,Chun-Shien Lu
类目: Machine Learning (cs.LG); Cryptography and Security (cs.CR)
*备注:
Abstract:Existing latent-based watermarking methods for diffusion models have overestimated their robustness to image distortions, including geometric transformations such as rotation, scaling, and translation (RST). Moreover, this paradigm of watermarking approaches may suffer from inherent limitations arising from the domain in which the watermark is embedded. In this paper, we provide the first theoretical analysis explaining why these methods lack invariance to perturbations. By relaxing the invariant relation, we derive a maximum perturbation bound that characterizes the relationship between pixel-space perturbations and their corresponding effects in latent space. In addition, we present the first analytical formulation that captures all components of practical detection mechanisms. Finally, we conduct experiments to validate the theoretical findings and the limitations of latent-based watermarking methods. Our theoretical and empirical results indicate that, under the current design paradigm, latent-based watermarking methods intrinsically exhibit limited robustness. We conclude by providing the analytical tool and design guidelines that future research could follow.
[LG-34] Beyond the Leaderboard: Multi-Dimensional Evaluation of Dense and Mixture-of-Experts Models for Automated Program Repair
链接: https://arxiv.org/abs/2610.08173
作者: Anvi Kalpesh Shah,Umamaheswara Sharma B
类目: oftware Engineering (cs.SE); Machine Learning (cs.LG)
*备注:
Abstract:Automated Program Repair (APR) with language models is usually evaluated by whether a generated patch passes the test suite, which can hide differences in maintainability, security, and computational cost. We propose a Weighted Quality Index (QI), inspired by the ISO/IEC 25010 software quality model, that combines functional correctness, maintainability, security, and generation efficiency under configurable weighting schemes. We evaluate three dense Qwen2.5-Coder models (3B, 7B, 14B) and the 16B-parameter DeepSeek-Coder-V2-Lite Mixture-of-Experts (MoE) model (2.4B active parameters) on 40 QuixBugs and 90 Defects4J bugs, all run locally on identical hardware to control for infrastructure effects. Model rankings change with the weighting scheme, showing that single-metric evaluation can hide trade-offs. The MoE model shows almost no statistically significant difference in correctness from the 7B and 14B dense models (McNemar’s exact test) while using 3-6 times fewer active parameters, whereas correctness increases significantly across the three dense scales. These results suggest that active parameter count can be a more informative lens than total parameter count for sparse code models.
[LG-35] Beyond Marginal Monitoring: Distributed Joint-Distribution Testing for Data Concept Drift in Large Scale E-Commerce Operations
链接: https://arxiv.org/abs/2610.08132
作者: Cagdas Pullu,Mahmut Emir Arslan,Bugra Balkac,Aylin Ondersev Balta,Cihangir Celal Palaci,Fikri Cem Yilmaz,Altan Cakir
类目: Distributed, Parallel, and Cluster Computing (cs.DC); Machine Learning (cs.LG); Methodology (stat.ME); Machine Learning (stat.ML)
*备注:
Abstract:Concept drift threatens production machine learning, yet the empirical behavior of multivariate two-sample drift detectors at scale remains under-characterized. Existing benchmarks rarely address the hundreds of millions of rows and high-cardinality features typical of industrial-operational datasets. We evaluate five multi-column two-sample tests (marginal, projection-based, and kernel embedding methods) across three complementary environments: the Harvard Dataverse, a validated Failing Loudly reproduction (mean absolute error between 0.030 and 0.053), and a novel synthetic-injection benchmark on the 137.5-million-row Trendyol collection-ranking feature table. Testing four drift types across two severity-scope regimes, we demonstrate that distributed Maximum Mean Discrepancy with Random Fourier Features on Apache Spark scales robustly. Averaged over the four drift types in the strong regime and under a calibrated threshold, it achieves a Pearson correlation of r = 0.940 with expected drift magnitude, an 80.4% true positive rate, and a 3.2% false positive rate. Conversely, the per-dimension Kolmogorov-Smirnov test failed due to statistic saturation from ID-like columns under asymmetric sampling, establishing a critical constraint for large-scale sampling design. At weak configurations (realized-flip fractions of at most 0.57%), detectors struggled to reliably discriminate, highlighting the need for future intensity-grid power analyses to distinguish fundamental sensitivity bounds from scalable threshold shifts.
[LG-36] Do LLM s Act on What They Know? From Partner Representations to Cooperative Actions
链接: https://arxiv.org/abs/2610.08129
作者: Yuhwan Jeong,Jinnyeong Yang,Kuk-Jin Yoon
类目: Machine Learning (cs.LG)
*备注:
Abstract:Cooperation with unfamiliar partners requires adapting to communication conventions that are not known in advance. We study this problem in a controlled Hanabi-derived environment with scripted hint generation, LLM-controlled receiving decisions, and frozen model weights. Across eight LLMs, linear probes recover intent conventions substantially more accurately than target conventions, yet receiving choices do not consistently agree with the sender’s convention. We compare probe-predicted and ground-truth conventions presented either as general rules or as externally computed action recommendations. Rule statements yield modest and model-dependent changes in cooperation, whereas action translation produces larger gains on average. In a Qwen3-8B case study, matched-state statement reversals reveal much greater sensitivity to action recommendations than to rule statements. Activation transfers from oracle-action and non-oracle hint-restatement donors improve intent accuracy on both action classes, but the tested alternatives do not reliably reproduce these benefits. Together, these results distinguish convention decodability, sensitivity to convention information, and cooperative performance, and highlight limitations in turning available partner information into receiving decisions.
[LG-37] Attenuated in-context identification in time-series foundation models: diagnosis under counterfactual inputs and repair by synthetic forced-system fine-tuning
链接: https://arxiv.org/abs/2610.08118
作者: Hong-In Won
类目: Machine Learning (cs.LG); Systems and Control (eess.SY)
*备注: 12 pages, 4 figures, 5 tables
Abstract:Covariate-aware time-series foundation models (TSFMs) promise training-free what-if answers for instrumented plants: the change in output that a different future input would cause. We test this on forced engineering systems with exact counterfactuals, comparing Chronos-2, TimesFM-2.5 and TabPFN-TS with classical system identification fitted to the same context. Through their default covariate interfaces, TimesFM-2.5 and TabPFN-TS are memoryless: the predicted effect of an input change is a same-time function of that change ( R^2 = 1.000 for TimesFM-2.5). Chronos-2 identifies dynamics in context but attenuates them. Its predicted effect is 0.33-0.80 of the true effect, its recovered impulse response has the wrong shape, and its error on a one-degree-of-freedom oscillator levels off at 0.57 with 8192 context samples, where ARX fitted to 256 samples reaches 0.02. Context dither at inference lowers the what-if error on all six synthetic classes without training. A 26-minute fine-tune on synthetic forced systems restores the response magnitude (sensitivity 0.83-0.96) and outperforms structure-agnostic identification on Wiener-Hammerstein and a held-out friction class. A specialised in-context identifier trained on the same data comes close, so the forced-system data carry most of the gain. On three of four measured plants classical identification remains clearly better, and the fine-tuned model loses part of its univariate forecasting skill. Paired counterfactual inputs, together with shuffled future inputs on measured records, test two properties: whether the covariate interface can represent dynamics and whether the pretraining prior covers the plant’s time scale. Only the counterfactual pairs expose the attenuation.
[LG-38] Energy-Aware Path Following: Comparative Analysis of Reinforcement Learning and NMPC for Electric Vehicles WWW
链接: https://arxiv.org/abs/2610.08112
作者: Mohamed Sabaa,Mostafa Emam
类目: Robotics (cs.RO); Machine Learning (cs.LG); Systems and Control (eess.SY)
*备注: 20 pages, 12 figures, currently submitted for review at the journal (Robotics and Autonomous Systems) this https URL
Abstract:Path-following control strategies typically follow the bi-objective optimization dilemma: minimizing deviations from a reference path while maintaining smooth speed profiles. The latter objective is especially relevant for Electric Vehicles (EVs), since their limited driving range can be extended by recovering energy through regenerative braking, a feature that has not yet been sufficiently studied in the literature. In this work, we perform a comparative analysis of four controllers under one common Frenet frame-based kinematic vehicle model, utilizing a validated energy model (VT-CPEM) with explicit regenerative braking. Herein, we implement the following controllers: Nonlinear Model Predictive Control (NMPC), Proximal Policy Optimization (PPO), gain-scheduled Ackermann state-feedback baseline (PID-SF), and a Stanley geometric baseline. To satisfy real-time requirements, we implement the NMPC using JIT-compiled CasADi. Moreover, we train the PPO using traditional straight and S-curve tracks, after which we successfully transfer the unmodified policy to unseen tracks, including: an ISO 3888-1 lane-change, a chicane, randomly-generated parameterized-splines, and a \pm3^\circ graded road. In addition, the policy transfers to a dynamic single-track vehicle model with linear tires, zero-shot with an acceptable initial performance, which was optimized after brief fine-tuning. Thereby, we demonstrate that our PPO is readily transferable to more comprehensive vehicle models. We conclude with a performance analysis of developed controllers and discuss ideas for future work.
[LG-39] Enhancing Diffusion Language Models with Autoregressive Post-Training Weights
链接: https://arxiv.org/abs/2610.08108
作者: Yiming Qin,Ke Wang,Amel Abdelraheem,Adam Hazimeh,Pascal Frossard
类目: Machine Learning (cs.LG)
*备注:
Abstract:Diffusion language models (dLLMs) have emerged as a promising alternative to autoregressive (AR) language models, offering flexible token-update orders and parallel decoding. Recent dLLMs are often initialized from pretrained AR models before diffusion conversion in order to inherit their learned representations. After the conversion, however, they typically ignore the extensive post-training ecosystem of their AR ancestors. In this work, we show that these existing AR post-training weight updates can instead be effectively recycled to enhance diffusion models. Despite the changes by AR-to-diffusion conversion, directly adding an AR post-training weight update to a diffusion base model remains effective, bringing its performance close to that achieved by direct diffusion post-training. Notably, AR and diffusion post-training updates are nearly orthogonal in weight space, yet induce substantially more aligned representation changes in the diffusion model. Their distinct updates are also complementary: composing their weights can retain gains from both regimes and further improve the post-trained diffusion model. Based on these findings, we propose A2D, a simple training-free framework for enhancing diffusion models with existing AR post-training resources. A2D can transfer capabilities from AR post-trained models to diffusion base models, and further improve already post-trained diffusion models by composing AR and diffusion post-training updates. Across various dLLMs, including Dream, DreamReasoner, DiffuCoder, Dream-Coder, Nemotron-Labs-Diffusion, and DiffusionGemma, A2D reliably improves instruction following, mathematical reasoning, and coding with both supervised fine-tuning and reinforcement learning updates, without additional training, or inference-time computation.
[LG-40] Surviving the Router: Optimizing Skill Injections for Retrieval and Execution
链接: https://arxiv.org/abs/2610.08098
作者: Haneen Najjar,Luca Scionis,Haritz Puerto,Sahar Abdelnabi
类目: Cryptography and Security (cs.CR); Machine Learning (cs.LG)
*备注: 16 pages, 5 figures
Abstract:AI agents increasingly rely on modular third-party “skills” that are dynamically selected by skill routers to execute complex tasks. While recent studies highlight the threat of prompt injections embedded in these skills, existing evaluations often assume settings where the malicious skill is already selected for execution. We show that this assumption can substantially overestimate attack success. In realistic multi-skill environments, injected skills must first compete for retrieval, reducing the effective attack success rate (ASR) of existing injections by 87-97%. To address this limitation, we introduce CORSA (Cluster Optimization for Router-Aware Skill Attacks), a router-aware attack that optimizes skill injections for both retrieval and execution across clusters of related tasks. We evaluate skill injection attacks under router-managed multi-skill settings by extending the benchmark introduced by SkillRouter with eight malicious payload categories. CORSA uses successive optimization stages to first improve retrieval and then optimize end-to-end attack success, while we evaluate user utility and injection naturalism separately. Our experiments show that CORSA substantially improves both retrieval and end-to-end attack success over existing skill injections while preserving user utility, and that the resulting attacks transfer across different router architectures and LLM backbones.
[LG-41] Explainable Rule Mining of IPv6 Extension-Header Presence Patterns from Paired-Vantage Captures
链接: https://arxiv.org/abs/2610.08090
作者: Priyanka Sinha,Nikolaos Kekatos,Stylianos Basagiannis,Antonio Anastasio Bruto da Costa,Alexios Lekidis,Pabitra Mitra,Tom Nianios,Elpiniki Papageorgiou
类目: Cryptography and Security (cs.CR); Machine Learning (cs.LG); Networking and Internet Architecture (cs.NI)
*备注: 6 pages, 1 figure, 2 tables. Accepted at the 2026 IEEE International Conference on Cyber Security and Resilience (IEEE CSR 2026)
Abstract:IPv6 extension headers (EHs), such as fragmentation, segment routing, and in-situ telemetry, are operationally important yetwidely dropped in transit, and characterising their behaviour from packet captures is a recurring measurement problem. We ask whetheran explainable miner can recover human-readable rules of EH behaviour, and we contribute two reusable tools: a negative-control protocol that diagnoses whether a mined “temporal” network rule reflects genuine cross-packet dynamics or mere within-packetco-occurrence, and a sender-conditioned, per-family EH-retention measurement. Applying an interpretable temporal-logic rule miner to the JAMES paired-vantage dataset, we recover a portable Fragment-EH rule that the protocol reveals to be a within-packet,near-definitional co-occurrence rather than a temporal pattern, so the temporal-logic machinery does no work for this dominant rule;the retention measurement independently recovers the expected within-window ordering of EH observability. Our main result istherefore an honest, controlled negative finding, corroborated by executed decision-tree and large-language-model baselines: on theevaluated JAMES traces network-temporal structure does not carry the dominant Fragment-EH signal, and we supply the controls thatestablish when it would, validated on a synthetic positive control containing a genuine cross-packet dependency.
[LG-42] Detecting a Shift Is Not Enough: Exact Minimax Limits of Linear Representation Repair
链接: https://arxiv.org/abs/2610.08069
作者: Anuar Aimoldin,Yankai Chen,Ayana Mussabayeva,Xue Liu
类目: Machine Learning (cs.LG); Statistics Theory (math.ST); Machine Learning (stat.ML)
*备注: 29 pages, 7 figures, 4 tables. Code: this https URL
Abstract:A mean shift between two data sources can be easy to detect but hard to remove without substantially changing their representations. We cast its removal as a statistical decision problem: from noisy differences between paired calibration measurements in \mathbbR^d , learn one linear map, applied to both sources under a hard distortion budget, that leaves as little of the shift as possible on fresh data. We derive the exact finite-sample minimax risk over all such maps, (d-k) \mathbbE[1/(d+2J)] with J\sim\mathrmPois(\kappa/2) , where the budget allows deleting k directions and \kappa is the calibration signal-to-noise ratio. Projecting out the mean calibration difference attains it without knowing \kappa or the noise scale. This exposes a detection-repair gap: detecting the shift needs only \kappa\gg\sqrt d , whereas removing a fixed fraction of it at constant distortion needs \kappa\asymp d , as for estimating its direction. Standard linear concept erasers (MP, SAL, LEACE) remove the same calibration difference, so the formula gives, before fitting, exactly how much shift they leave on fresh data and how much calibration a target requires. The limit is robust: pairing keeps it exact for non-Gaussian shared content, the projection keeps its guarantee under anisotropic noise, and selective abstention cannot close the gap. On paired clinical and wearable sleep EEG, where differences between participants act as calibration noise, the formula predicts the device shift left in new participants, and more recordings per person soon stop helping. Together, these results tell whether a correction that falls short needs a better method, more recordings, or more participants.
[LG-43] A Riemannian Geometry for Low-rank Adaptation
链接: https://arxiv.org/abs/2610.08049
作者: Shoichiro Takeda,Shin’ya Yamaguchi,Satoshi Suzuki,Yasunori Akagi
类目: Machine Learning (cs.LG)
*备注:
Abstract:Low-rank adaptation (LoRA) is widely used as a parameter-efficient fine-tuning technique for pre-trained deep neural networks, which approximates the weight update via full fine-tuning by a low-rank matrix BA^\top . This parameterization leads to the equivalence relation (B, A) \sim (BG^-1, AG^\top) for any invertible matrix G because BA^\top = BG^-1(AG^\top)^\top and thus both pairs yield the same loss value. This relation induces a quotient manifold where matrices (BG^-1, AG^\top) for all G are identified, eliminating redundant directions along which the loss value remains unchanged. To respect the geometry of this manifold, the original search space is endowed with a Riemannian metric that is invariant under the equivalence relation. Such a metric induces preconditioning at each gradient step and ensures that each weight update via LoRA changes the loss value, leading to efficient optimization. In this paper, we propose a new Riemannian metric that is specifically tailored to LoRA to close the gap to full fine-tuning at the weight level. We theoretically show that LoRA with our preconditioning induced by this metric satisfies the following two properties at each iteration: (i) The weight update follows the direction closest to the gradient of full fine-tuning within the subspace of first-order weight changes allowed by the LoRA parameterization. (ii) The updated weight matrix is closer in Frobenius norm to that of full fine-tuning than the updated weight matrices of LoRA with conventional preconditioning and without preconditioning. These theoretical insights suggest that our preconditioning makes LoRA better approximate full fine-tuning, thereby leading to more efficient optimization. Experiments show the effectiveness and efficiency of our preconditioning for LoRA on fine-tuning tasks with language and vision domains.
[LG-44] SepsisLens: Structure-Preserving Sequence Modelling for Decomposable Early Sepsis Warning
链接: https://arxiv.org/abs/2610.08046
作者: Yikun Ou,Wei Li
类目: Machine Learning (cs.LG)
*备注:
Abstract:Early sepsis warning from ICU records can be cast as a structure-preserving prediction problem. A model needs to detect deterioration from irregular measurements while keeping each alert connected to the physiological signals that support it. Many temporal models fuse clinical variables into a patient-level representation, supporting scalar risk prediction but weakening the structure needed for clinical decomposition. We present SepsisLens, which preserves variable-indexed temporal states until risk composition. Observation-aware representations encode each variable’s dynamics and measurement history, while a shared temporal encoder models each trajectory without collapsing the variable axis. The StructuredRiskHead composes multi-horizon risk from explicit variable-level and organ-level components. We evaluate SepsisLens on three public ICU cohorts and one private-hospital cohort under a common pre-onset protocol. SepsisLens achieves strong discrimination on all four cohorts and lower alert burden at matched event recall on MIMIC-IV. Structural ablations support the design, while input-side masking shows that the ranked components reflect variables with greater influence on prediction.
[LG-45] Spectra: Exact Component Transport for Test-Time Prior Adaptation in Simulation-Based Inference
链接: https://arxiv.org/abs/2610.08021
作者: Xin Zhao,Nico Scherf,Robert Trampel,Kerrin J. Pine,Nikolaus Weiskopf
类目: Machine Learning (cs.LG); Machine Learning (stat.ML)
*备注: 41 pages, 7 figures
Abstract:Simulation-based inference (SBI) has become a powerful approach to Bayesian inference in complex scientific models whose likelihoods are difficult or impossible to evaluate. Amortized SBI learns reusable inference models from simulated data, enabling rapid posterior inference for new observations, and modern generative models have made these models increasingly expressive. However, this reuse is limited to the prior distribution chosen during training, whereas scientific analyses often need revised priors as knowledge accumulates or alternative assumptions are tested. We introduce Spectra, a test-time adaptation method for diffusion-based SBI. Spectra uses an exact score-transport identity to obtain the adapted score from a frozen diffusion model in closed form for structured prior changes, without additional simulation or training. Across six SBI benchmarks, Spectra achieves accurate adaptation under strong prior shifts at low online sampling cost. This enables pretrained SBI models to incorporate updated prior information at test time.
[LG-46] FOSLS-deRhaNN: native de Rham neural classes for H(div) and H(curl) with applications to first-order system least-squares neural network methods for partial differential equations
链接: https://arxiv.org/abs/2610.08016
作者: Shun Zhang
类目: Numerical Analysis (math.NA); Machine Learning (cs.LG)
*备注:
Abstract:We construct neural approximation classes native to the graph spaces H(div) and H(curl), in two and three dimensions and, for H(div), in any dimension. Every realization lies in the space for all parameter values, and with kinked potentials, such as ReLU networks, the admissible jumps appear at finite width. The classes are images of scalar and componentwise networks under fixed operators of the de Rham complex, and do not involve a mesh or finite element emulation. For H(div) in R^n two native classes are given on an equal footing, with a skew-symmetric potential A : \mathrmDiv,A+R_nq+\mathbfh , with the divergence q as an explicit unknown, and \mathrmDiv,A+\mathbfz with an H^1 field \mathbfz ; for H(curl) the analogous classes are \mathrmgrad,\phi+Sr+\mathbfh in two dimensions and \mathrmgrad,\phi+\mathbfz in two and three dimensions. In all of them every interface jump of the field is carried by the potential term, \mathrmDiv,A or \mathrmgrad,\phi , while the remaining part has no interface jump (it is an H^1 field in the regular-decomposition classes); the classes with \mathbfz are the componentwise approach enriched by this term. Known or learned interface geometry enters the potential through factors with trainable amplitudes, and the remaining part if the divergence jumps. The classes lead to the FOSLS-deRhaNN method, first-order system least squares with de Rham neural networks, whose loss is the least-squares functional posed in the natural spaces of the weak formulation; for elliptic equations this includes H^-1 right-hand sides and H^1/2 Dirichlet data. Elliptic equations with discontinuous coefficients and curl-curl problems are treated as instances, with the functional equivalent to the error; linear transport with discontinuous solutions and conservation laws with shocks use the same flux classes.
[LG-47] Can phenotypic activity be predicted without experimental readouts? NEURIPS2026
链接: https://arxiv.org/abs/2610.07997
作者: Télio Cropsal,Rocío Mercado
类目: Machine Learning (cs.LG)
*备注: Accepted to the ML4Molecules: Agentic Systems for Molecular Sciences Workshop at the 40th Conference on Neural Information Processing Systems (NeurIPS 2026)
Abstract:Molecular encoders contrastively pretrained on paired molecule-morphology data, such as CLOOME and CellCLIP, have been proposed as cheap surrogates for phenotypic prediction, avoiding the need to run a Cell Painting assay. We evaluate this idea for these molecular encoders under a protocol designed to control for two confounds that can inflate apparent performance: leakage across an encoder’s own pretraining boundary, and the correlation between phenotypic activity and cytotoxicity. Testing six representations, including a non-pretrained MLP control matching CLOOME’s input and layer count, on two distinct Cell Painting screens, we find that once these confounds are controlled for, the pretrained molecular encoders show no clear advantage over plain physicochemical descriptors, and that toxicity is generally easier to predict than phenotypic activity across representations. Our results suggest leakage-aware, confound-controlled evaluation should be standard practice before phenotype-pretrained encoders are trusted as surrogates for phenotypic drug discovery.
[LG-48] Do Higher-Order Models Win for Higher-Order Reason s? Rethinking Performance Gains in Hypergraph Learning
链接: https://arxiv.org/abs/2610.07981
作者: Fanchen Bu,Fan Li,Geon Lee,Sunwoo Kim,Xiaoyang Wang,Renaud Lambiotte,Kijung Shin
类目: Machine Learning (cs.LG); Social and Information Networks (cs.SI)
*备注:
Abstract:Higher-order models (e.g., hypergraph neural networks) often outperform lower-order baselines on hypergraph learning benchmarks, and their advantages are commonly attributed to their ability to exploit higher-order information. However, better performance alone does not establish this explanation. We therefore ask: Do higher-order models win for higher-order reasons? To investigate this question, we introduce a controlled performance-attribution framework that perturbs higher-order information while preserving the lower-order, i.e., pairwise, information. Across 25 commonly used hypergraph learning benchmarks spanning three tasks, we frequently observe an intriguing pattern: higher-order models originally outperform lower-order baselines, yet retain most of their advantage after perturbation. This suggests that much of the observed advantage remains achievable without the higher-order information. We then investigate potential lower-order explanations for these remaining gaps. We find that simple additions to a lower-order baseline, e.g., richer pairwise weighting, more steps of pairwise feature propagation, and normalization, reduce the remaining performance gaps, supporting lower-order explanations for part of the observed advantage. Our analysis calls for the hypergraph learning community to rethink performance attribution by distinguishing performance gains from their explanations, adopt stronger lower-order baselines, and use suitable benchmarks that better test the value of higher-order information.
[LG-49] Learning a Ranking from Human Feedback in Log-Concave Random Utility Models
链接: https://arxiv.org/abs/2610.07973
作者: Diego Alovisetti,Marco Mussi,Alberto Maria Metelli
类目: Machine Learning (cs.LG)
*备注:
Abstract:We study the problem of recovering the ranking of a fixed set of items according to their unknown numerical utilities. At each interaction with the environment, a learner presents the item set to a human and receives comparative feedback of two types. Under full-ranking feedback, each interaction reveals a noisy ranking of all items, whereas under winner-only feedback, it reveals only the item ranked first. In both settings, we model human feedback using a random utility model with log-concave noise and study the number of observations needed to recover an \epsilon -accurate ranking with high probability. This novel criterion tolerates ordering errors only between items whose utilities differ by less than \epsilon . For both feedback types, we establish worst-case sample-complexity lower bounds and develop algorithms that match these bounds up to logarithmic factors. Neither algorithm requires knowledge of the noise distribution, while only requiring an upper bound on its variance. Our results show that the ranking problem under winner-only feedback is intrinsically harder by exposing the sample complexity dependence on the minimum winning probability across the item set.
[LG-50] DecepEval: A Benchmark for Evaluating Deception in LLM Agents
链接: https://arxiv.org/abs/2610.07967
作者: Yiming Xu,Hongyue Yu,Beihua Yang,Zihan Chen,Yixin Liu,Zhen Peng,Bin Shi,Bo Dong,Chao Shen,Irwin King,Qinghua Zheng
类目: Machine Learning (cs.LG)
*备注:
Abstract:As large language model (LLM) agents become increasingly autonomous, they may pursue task performance through deception, raising concerns about their reliable deployment. Existing evaluations show that LLM agents can deceive, but often examine isolated scenarios or narrowly defined conditions, limiting systematic understanding of when deception becomes more likely. To address this gap, we introduce DecepEval, a benchmark comprising 1,532 instances across 3 task families and 28 professional scenarios. Drawing on classical fraud theories, we propose the LLM Deception Diamond framework, which characterizes four external conditions that may induce deception: pressure, incentive, opportunity, and conflict. DecepEval pairs neutral and induced versions of each instance to measure condition-dependent changes in deception rates, while explicit task facts and observable agent behavior help distinguish deception from capability-related errors. Evaluations of nine frontier LLMs show that inducements increase deception across models and task families, even among models with low baseline deception rates. DecepEval makes these vulnerabilities measurable, providing a shared benchmark for progress toward trustworthy artificial intelligence.
[LG-51] Feature Encoding in VAE-based Audio Decoders: Effects of Input Depth and Distribution
链接: https://arxiv.org/abs/2610.07966
作者: Louis McCallum,Mick Grierson
类目: ound (cs.SD); Machine Learning (cs.LG)
*备注: This manuscript has been accepted for publishing in IEEE Transactions on Audio, Speech and Language Processing (TASLP)
Abstract:Neural audio synthesis models like the Realtime Audio Variational autoEncoder (RAVE) achieve impressive genera tion quality, yet how their internal representations encode musical features remains poorly understood. We present a systematic layer-wise and cross-layer cluster analysis of RAVE decoder activations across three models trained on different musical domains, tested with four stimulus types. We then evaluate architectural generalization with a general purpose EnCodec model. For RAVE, we find that synthetic stimuli are encoded well across models and audio features (pitch |\rho|=0.45, 5.1x the null, BPM |\rho| = 0.76, 8.6x the null). These results are reduced but still substantively apparent when using natural audio (mean across features |\rho|=0.25, 2.8x the null). Natural audio sees a stronger encoding when nonlinear probes are used (mean across features R2=0.56, 18x the null, +0.152 nonlinear gain over the linear probe R2). Encoding strength varies throughout the layers of the decoder and an increased ability to joint-encode in the middle layers is seen across all audio features (\beta2 all negative, p 0.05). The general purpose EnCodec decoder also sees similar strong synthetic responses across audio features, similar nonlinear gains for natural audio joint encoding and similar depth profiles. We find the best cross-layer cluster improves the strength (r = 0.65, p = 0.006) and prevalence (r = 0.75, p = 0.001) of BPM encoding when compared against the best whole layers within the same section, with no effect for joint encoding. These findings advance the interpretability of neural audio models and inform targeted control strategies for neural synthesis.
[LG-52] Generalized Matheron Variational Implicit Processes
链接: https://arxiv.org/abs/2610.07938
作者: Luis A. Ortega,Andrés R. Masegosa,Thomas D. Nielsen
类目: Machine Learning (cs.LG)
*备注: 36 pages, 10 figures, 19 tables. Submitted for review
Abstract:Implicit-process priors specify distributions over functions through sample-forward mechanisms such as Bayesian neural networks and stochastic simulators, but their function-space densities are typically unavailable. We introduce Generalized Matheron Variational Implicit Processes (GMVIP), a pathwise variational family for posterior inference with such priors. For Gaussian-process priors, GMVIP recovers the standard inducing-variable variational GP construction; for general implicit priors, its empirical covariance construction preserves the prior mean and covariance in the population limit. GMVIP constructs posterior samples by drawing a function from the prior and applying a correction anchored at a set of inducing inputs. The effect of this correction away from the inducing inputs is determined directly from prior samples, allowing the posterior to retain the structure and variability of the original implicit process. The (surrogate) prior and variational posterior use the same pathwise construction and differ only in the distribution of whitened inducing coefficients, yielding a tractable coefficient-space Kullback-Leibler divergence. Experiments on regression, classification, and forecasting with simulator-defined and retrieval-conditioned empirical trajectory priors show that GMVIP is broadly competitive with existing methods.
[LG-53] Revisiting Temporal Regularization for Smooth Control in Deep Reinforcement Learning ICRA2027
链接: https://arxiv.org/abs/2610.07910
作者: SungJae Ahn,Jeong Woon Lee,Kyoleen Kwak,Hyoseok Hwang
类目: Machine Learning (cs.LG); Robotics (cs.RO)
*备注: Submitted to ICRA 2027
Abstract:Deep Reinforcement Learning policies can produce nonsmooth action oscillations that hinder deployment on physical robots. Existing architectural and penalty-based approaches seek spatial smoothness by directly reducing sensitivity to changes in state inputs, but their broad constraints can degrade task performance as stronger smoothing is pursued. Temporal regularization instead constrains action differences along observed transitions, but has been considered unable to provide the spatial smoothness needed under observation noise. We revisit this assumption by proving that the temporal penalty bounds the expected action differences between current states sharing a next state, revealing a spatial effect that empirically extends to spatial smoothness. Building on this finding, we propose Conditioning for Action using only Temporal Smoothness (CATS), which combines a temporal penalty with linear ramp-up. We highlight temporal regularization’s ability to provide spatial smoothness while better preserving task performance than explicit spatial regularization. Through linear ramp-up, CATS allows the policy to learn rewarding behavior before progressively smoothing its actions, improving return preservation and both temporal and spatial smoothness. Experiments in both simulation and the real world show that CATS substantially reduces action oscillation without degrading task performance, with little computational overhead.
[LG-54] ApexQuant: Data-Free Elastic Quantization by Residual Re-Isotropization
链接: https://arxiv.org/abs/2610.07904
作者: Aksel Fristrup,Sumit Pandey,Ankit Kariryaa
类目: Machine Learning (cs.LG)
*备注: 25 pages, 6 figures
Abstract:We introduce ApexQuant, a calibration-free quantization method that recursively re-quantizes the residual error, serving as a refinement layer on top of existing quantizers. We establish that a fresh random rotation returns each residual to the uniform distribution on the hypersphere, which characterizes the rate of progressive error decay across successive passes. This result lets us determine, before any weight is read, how many passes a layer needs for a target weight-space error. Every prefix is itself a valid lower-rate model, so one artifact serves several precisions. We instantiate ApexQuant with three interchangeable stages, scalar, E_8 and trellis, and validate it on four open-weight LLMs and on Earth-observation and medical domains where in-distribution data is often unattainable as imagery arrives under restrictive licences or due to patient material under privacy constraints. Progressive re-isotropization comes within a few percent of full precision at four bits and gives the best two-bit arm we measure, in a completely data-free setting.
[LG-55] FC-SWE: Failure-Conditioned RL for Long-Horizon Software Engineering Agents
链接: https://arxiv.org/abs/2610.07898
作者: Jia Liufu,Bin Hu,Linglin Jing,Terry Kong,Yuki Huang,Ashwath Aithal,Wenming Yang,Jun Yang
类目: Machine Learning (cs.LG); Software Engineering (cs.SE)
*备注: 23 pages, 8 figures, 6 tables
Abstract:Repository-level software engineering (SWE) is a challenging long-horizon setting: agents must reason over extended interactions, use tools, and adapt to stateful environments. Recent work trains SWE agents with reinforcement learning methods such as Group Relative Policy Optimization (GRPO), which independently sample multiple trajectories per issue, test the resulting patches, and compare terminal rewards within a fixed group. However, this training setup does not reuse verifier feedback from failed patches as context for subsequent attempts, even though this feedback contains valuable diagnostic information about what went wrong. Training on recovery trajectories is challenging because the preceding outcome determines whether the next trajectory is generated, while the failed execution determines its conditioning context. We introduce FC-SWE, a failure-conditioned RL framework that incorporates recovery attempts into policy training. After a patch fails verification, FC-SWE restores the repository to its original task state and uses the failed patch and verifier feedback as context for a recovery trajectory. FC-SWE adapts GRPO to these chains of complete, multi-turn tool-use trajectories through two mechanisms. Trajectory-local rewards preserve each attempt’s verifier outcome, preventing recovery success from rewarding an earlier failed patch. Active-set advantage estimation forms a comparison group from all initial and recovery trajectories actually executed for the same issue, so failed attempts remain in the group while unexecuted attempts are excluded. On all 500 SWE-bench Verified tasks under a verifier-assisted protocol, FC-SWE with Qwen3.5-4B and SWE-agent achieves 41.7% Resolved@1 and 52.8% Resolved@2, compared with 38.9% and 48.5% for GRPO. Although trained with at most two attempts per chain, FC-SWE reaches 70.7% Resolved@11 under an eleven-attempt test-time budget.
[LG-56] Learned Adaptive Multiresolution Diffusion Imaging
链接: https://arxiv.org/abs/2610.07884
作者: Christian Tantardini,Stig Rune Jensen,Roberto Di Remigio Eikås,Joakim Henrik Beck
类目: Numerical Analysis (math.NA); Machine Learning (cs.LG); Image and Video Processing (eess.IV)
*备注:
Abstract:Adaptive multiresolution methods reduce representation cost by concentrating fine-scale degrees of freedom where needed, but their tree updates are usually governed by fixed local criteria. We introduce Learned Adaptive Multiresolution Diffusion Imaging (Learned AMDI), which preserves the AMDI fixed-tree propagator and hierarchy constraints while replacing the post-propagation selector with a shared local policy trained by proximal policy optimization. Regression tests reproduce deterministic AMDI trajectories to machine precision when identical trees are used. In the Haar implementation studied here, the deterministic one-step selector accepts no refinements in 54 decisions. Across nine held-out cases, Learned AMDI executes 393 refinements and reduces the mean terminal reference discrepancy from 0.17496 to 0.13657 , while occupancy rises from 0.13737 to 0.26660 . Step-resolved diagnostics reveal occasional small adaptation-energy increases; fixed-tree energy stability therefore does not guarantee monotonicity of the learned outer iteration. At comparable occupancy, a validation-tuned observed-detail threshold reaches a discrepancy of 0.13792 with slightly better RMSE and SSIM, placing both methods on essentially the same accuracy–occupancy tradeoff. A decision-1-only control reaches 0.13742 , indicating that most of the improvement on this static benchmark arises from the initial allocation. The shared actor transfers without retraining to 64\times64 and 128\times128 images, improving reference discrepancy, RMSE, and SSIM relative to deterministic AMDI, while the frozen threshold rule remains competitive. Learned AMDI thus provides a hierarchy-constrained, resolution-transferable mechanism for adaptive allocation and clarifies the contribution of sequential decisions.
[LG-57] On-Policy Distillation with Negative-Policy Rollouts
链接: https://arxiv.org/abs/2610.07874
作者: Jaehui Hwang,Dongyoon Han,Sangdoo Yun,Byeongho Heo
类目: Machine Learning (cs.LG)
*备注: 25 pages, 7 figures, 24 tables
Abstract:On-policy distillation (OPD) has been widely studied as a post-training method in which a student model obtains token-level supervision from a stronger teacher on its own rollouts. Recent studies have improved OPD through alternative distillation reward formulations and teacher configurations, while the objective of distillation remains centered on mimicking the teacher. However, when a stronger teacher has limited distributional overlap with the student, such positive guidance can provide insufficient learning signals. In this work, we introduce Negative-Policy OPD (NP-OPD), which complements teacher supervision with rollouts from a lower-performing, lower-capability negative policy that serves as a negative reference for the student. Rather than modifying the distillation reward formulation, NP-OPD introduces the negative policy at the rollout stage, continuously supplying tokens preferred by the negative policy over the teacher so that they remain exposed to teacher supervision throughout training. This provides an explicit negative signal through negative-policy rollouts while preserving the positive teacher supervision used in OPD. Through extensive experiments, we show that NP-OPD improves OPD across model scales, generation modes, reasoning domains, and different OPD variants. Furthermore, our analyses show that NP-OPD effectively suppresses tokens preferred by the negative policy over the teacher and moves the student away from the negative policy. These results support our design of introducing negative signals through negative-policy rollouts and provide new insight into the role of the rollout policy in OPD. Code will be available at this https URL.
[LG-58] ram-FL: Reducing Communication and Computation Costs through Sequential Model Circulation in Decentralized Federated Learning
链接: https://arxiv.org/abs/2610.07859
作者: Kota Maejima,Takayuki Nishio,Asato Yamazaki,Yuko Hara-Azumi
类目: Machine Learning (cs.LG); Distributed, Parallel, and Cluster Computing (cs.DC); Networking and Internet Architecture (cs.NI)
*备注: 12 pages, 7 figures, 5 tables. This work has been submitted to the IEEE for possible publication
Abstract:Conventional decentralized federated learning (DFL) often focuses on clients, with each client maintaining a model copy, performing updates individually, and undertaking model exchange and integration. While fully leveraging computational resources can shorten training times, it can also lead to significant computational and communication waste. This is especially pronounced with non-independent and identically distributed (non-IID) data, where achieving high model accuracy demands extra resources. This research shifts focus to the model itself, aiming to realize DFL with minimal computation and communication costs. To this end, we propose Tram-FL (Traveling Model Training Mechanism for Decentralized Federated Learning), a mechanism designed to efficiently address these challenges. It sequentially trains a single model by circulating it among nodes. We address the training scheduling problem in model circulation-based training, specifically determining which nodes should update the model and the number of updates to perform. This is approached by considering the model’s circulation route and update iteration allocation, for which we propose simple yet effective methods. Additionally, with quantized momentum, Tram-FL achieves high accuracy with fewer model circulations while controlling communication load per transmission. Experimental results show that the proposed algorithm, even with non-IID data, converges to a global model with reduced communication and computation.
[LG-59] A Decision-Focused Neural Optimization Framework for Personalized Route Reproduction from Vehicle Trajectories
链接: https://arxiv.org/abs/2610.07857
作者: Gyeongjun Kim,Yeseul Kang,Keemin Sohn
类目: Machine Learning (cs.LG)
*备注:
Abstract:This study formulates individual route reproduction as a shortest-path problem over learned driver-specific latent link costs. The central idea is that, once such latent costs are inferred from contextual information, observed routes can be reproduced without enumerating alternative route sets. We propose a neural pipeline that includes a perception model that embeds context covariates, which comprises individual characteristics, trip-specific attributes, and network-level traffic states, into the personalized link costs. A constrained optimization (CO) layer, which determines the shortest path (SP) based on these estimated costs, follows the perception encoder. To enable end-to-end training, we employ decision-focused learning to align the predicted shortest paths with observed routes. The implicit maximum likelihood estimation (iMLE) provides an approximate gradient of the loss function that contains the non-differentiable CO layer. Furthermore, a regularization term anchors the latent cost distribution to the empirical scale of observed link travel times, mitigating the scale ambiguity inherent in shortest-path supervision. Empirical evaluations demonstrate that the proposed framework outperforms baseline route choice models in path reproduction. The learned latent costs, interpreted as proxies for perceived travel costs, provide plausible explanations for heterogeneous route choices.
[LG-60] Scen-Opt: A Scenario Optimization Toolbox for Data-Driven Convex Programming
链接: https://arxiv.org/abs/2610.07846
作者: Ben Wooding,Simone Garatti,Marco C. Campi,Abolfazl Lavaei
类目: Mathematical Software (cs.MS); Machine Learning (cs.LG); Systems and Control (eess.SY); Optimization and Control (math.OC)
*备注: 49 pages. Software archived at this https URL (v1.0); source code at this https URL web app at this https URL
Abstract:The scenario approach is a well-established statistical framework for data-driven decision-making. In particular, in data-driven optimization, the scenario approach unveils how the problem structure governs out-of-sample generalization, and offers a principled basis for assessing and certifying the reliability of the optimal solution as per constraint satisfaction. Despite its strong theoretical development and wide applicability, no software toolbox has been available to date that enables user-friendly, data-driven convex optimization within the scenario-approach framework. In this paper, we introduce Scen-Opt, an open-source software tool that integrates convex programming with data samples while providing statistical guarantees grounded in scenario theory. Scen-Opt is implemented in Python, supporting data-driven linear, quadratic, and semidefinite programming, and offers a Python-based web application with an intuitive and reactive graphical user interface (GUI) built using modern web technologies. Scen-Opt can be used directly through its online interface or installed locally, accommodating both manual input and data-file uploads (CSV, JSON, TXT, TSV, MAT, Excel, NPY, NPZ, Parquet). Built on a Python backend with a modern JavaScript frontend, Scen-Opt offers a highly user-friendly experience and efficient usability across desktops, laptops, tablets, and mobile devices. In this paper, Scen-Opt is applied to a set of representative benchmarks, demonstrating its practical effectiveness for data-driven convex optimization with guaranteed performance.
[LG-61] Privileged Context as Drift in On-Policy Self-Distillation
链接: https://arxiv.org/abs/2610.07842
作者: Ravenor Davion,Nick Rui
类目: Machine Learning (cs.LG)
*备注: 18 pages, 4 figures, 5 tables
Abstract:On-policy self-distillation (OPSD) trains a language model to match a copy of itself conditioned on privileged context. Existing work varies what privileged context contains and how it is produced while also changing models, data, and training setups, making the effects of privileged context design difficult to isolate. Motivated by efforts in continual learning to reduce catastrophic forgetting, we study how the choice of privileged context affects policy drift. Specifically, we vary two axes: content (a demonstration, feedback, or rephrase) and source (external, self-generated with a verifier, or self-generated without a verifier). We train Qwen2.5-7B with OPSD across these nine combinations and three datasets, measuring target-task accuracy, prior-task retention, reverse KL from the base policy, and parameter-update geometry. Holding source fixed, changing content spans a wider median KL range than holding content fixed and changing source. The ratio between these ranges is 5.1\times for per-token KL and 2.2\times for per-sequence KL. Parameter-update geometry shows the same pattern: updates from adapters that share content are more closely aligned (mean cosine 0.571 ) than updates from adapters that share source ( 0.255 ). For continual learning, these findings suggest that privileged context should be treated as part of OPSD’s stability design because it is associated with how far and in what direction the policy moves.
[LG-62] Retrieval Is Not Enough: Refreshing Memory for Frozen Time-Series Forecasters
链接: https://arxiv.org/abs/2610.07834
作者: Chao He,Jianyu Xu,Xinyi Guo,Ruiqi Liu,Haobin Ding,Ruiqi He,Dongqing Song
类目: Machine Learning (cs.LG)
*备注: 13 pages, 6 figures, 6 tables
Abstract:Retrieval-augmented time-series forecasting uses the continuations of historical segments similar to the current context as references for a forecaster. Most existing methods build the retrieval memory once from the training segment, leaving observations revealed after deployment unavailable as references, and generally do not calibrate how much the retrieved information should influence a frozen forecaster. We identify two key determinants of retrieval utility for a frozen forecaster: whether the history still reflects the current state, and whether the correction it induces aligns with the forecaster’s residual errors, an alignment that can shift between validation and deployment when the memory becomes stale. We propose FreshCast, a plug-in retrieval framework that keeps the forecaster frozen, continuously updates a non-parametric memory with new observations, forms a memory forecast through relational kernel regression, and calibrates its weight in closed form on the validation segment. Under a simplified generative model, we characterize the optimal combination gain through the second-order relation between forecaster error and memory correction, and show that a sufficiently long look-back can make periodic memory information redundant. Across seven benchmarks and ten forecasting architectures, FreshCast reduces average MSE for every evaluated forecaster and input length, by 14.6% and 5.6% at input lengths 96 and 720, and achieves lower MSE than the evaluated retrieval-augmented and online baselines in their comparison settings. Ablations show that freezing the memory at the end of training removes most of the gain, identifying post-training observations as a primary source of improvement. For a frozen forecaster, useful historical references must remain timely and provide information that helps correct its remaining errors.
[LG-63] Forecast Accuracy Is Not Trading Profit: Evolving Small Recurrent Networks for Stock Return Prediction
链接: https://arxiv.org/abs/2610.07825
作者: Jonathan Chang,Zimeng Lyu
类目: Neural and Evolutionary Computing (cs.NE); Computational Engineering, Finance, and Science (cs.CE); Machine Learning (cs.LG)
*备注:
Abstract:Time series forecasting models are typically compared on pointwise error, which scores a prediction in isolation from the decision it is produced for, and a lower forecast error does not imply a better decision downstream. A parallel debate asks whether modern transformer architectures forecast better than recurrent and other lightweight models. We compare linear, fixed recurrent, transformer, and mixing based architectures against recurrent networks evolved by neuroevolutionary architecture search, evaluating each on forecast accuracy and on the net return of a daily long/short strategy. All models are fit on a pooled panel, one network trained across the whole universe. Across four mid-cap portfolios and three trading years, the evolved networks rank first on both forecast accuracy and net trading performance, while the second most accurate model loses money once positions are formed and costs are charged. The advantage tracks a horizon match, since rank IC for the evolved networks rises from a one-day to a ten-day scoring horizon while every model above 300 parameters declines. They are also the cheapest end to end: a CPU-only search of 16 minutes yields 66-weight networks that predict in 10.8~ \mu s on a Raspberry Pi Zero, against transformer baselines of up to 817,153 parameters that require GPU training.
[LG-64] CANDLE: Cortical Null-Space Decomposition for Noninvasive Brain Source Imaging
链接: https://arxiv.org/abs/2610.07824
作者: Shuntaro Suzuki,Yuiga Wada,Komei Sugiura
类目: Machine Learning (cs.LG); Signal Processing (eess.SP); Neurons and Cognition (q-bio.NC)
*备注:
Abstract:Electrophysiological source imaging (ESI) aims to estimate cortical source activity from noninvasive electrophysiological measurements such as electroencephalogram (EEG). However, ESI is fundamentally ill-posed because source activity is substantially higher-dimensional than sensor observations, resulting in non-unique solutions. Recent learning-based approaches address this ambiguity by learning data-driven source priors, yet they often struggle to generalize across subject-specific cortical geometries. To address this, we propose CANDLE, a learning-based ESI model that estimates source activity on subject-specific cortical geometries. CANDLE learns a prior over the null space induced by the source-to-sensor mapping derived from T1-weighted MRI, restricting learning to unobservable source components while preserving geometric constraints. To train CANDLE, we develop a whole-brain simulator spanning over 1,100 subject-specific cortical geometries with source configurations derived from over 26,000 statistical brain maps. Trained exclusively on simulated data, CANDLE outperformed prior ESI methods on simulated source activity estimation and generalized to two empirical tasks: (i) intracranial stimulation localization from simultaneously recorded scalp EEG and (ii) epileptogenic zone estimation from presurgical interictal EEG. Our project page is available at this https URLthis https URL.
[LG-65] Net: Multi-Task Deep Learning for Table Tennis Player Analysis with Smart Racket
链接: https://arxiv.org/abs/2610.07823
作者: Ko-Hsun Chen,Xiang-Wei Ke,Hsien-Cheng Huang,Shang-Kuan Chen
类目: Machine Learning (cs.LG)
*备注: 12 pages, 4 figures, 5 tables
Abstract:The AI CUP 2025 Precise Analysis of Table Tennis Smart Racket Data Competition introduced smart table tennis rackets that collect extensive player swing data, enabling research on table tennis big data. These data support in-depth analysis of players’ return techniques and swing-force consistency, improving the accuracy of player skill assessment. This study focuses on six-axis sensor data collected by smart table tennis rackets and proposes TTNet, a novel deep learning model with multitask learning capabilities, to advance table tennis data analysis and related applications. TTNet combines convolutional neural networks (CNNs), residual networks (ResNet), and self-attention mechanisms to simultaneously predict four player attributes: gender, playing hand, years of experience, and skill level. We adopt a two-stage training strategy that incorporates data augmentation and task-specific loss functions to improve generalization on imbalanced data. Our approach achieved second place on the official competition leaderboard.
[LG-66] SIFT: Search Intent-to-Filter Transformer for Multi-Task Personalized Filter Ranking at Airbnb CIKM2026
链接: https://arxiv.org/abs/2610.07810
作者: Shashank Dabriwal,Tanya Piplani,Hao Li,Yiwei Wang,Ashish Jain,Kedar Bellare,Stephanie Moyerman
类目: Machine Learning (cs.LG)
*备注: 9 pages, 5 figures, 6 tables. Accepted at GRAIL 2026: Workshop on Generative, Retrieval-augmented, and Agentic Intelligence for Personalization, co-located with CIKM 2026, Rome, Italy
Abstract:Search filters help guests navigate vast catalogs in two-sided marketplaces like Airbnb, and recommending the right filters can meaningfully lift booking conversion. Many such production filter-ranking systems, however, represent the guest through hand-engineered, pre-aggregated features generated by ETL pipelines. This makes it expensive to maintain and difficult to extend for new filter types or contextual dimensions (trip length, group size). We present SIFT (Search Intent-to-Filter Transformer), a ranking model built on transformers that learns guest preferences directly from raw behavioral sequences. SIFT replaces manual feature engineering with a unified guest representation that feeds multiple prediction tasks, including booking likelihood, filter engagement, and ordinal capacity thresholds (e.g., 2+ bedrooms) – a general framework for filter ranking in two-sided marketplaces that accommodates both boolean and numeric-range filter types. Extending SIFT to new filters requires only adding a new head, not a new feature pipeline. To keep serving fast, this guest representation is computed offline on a daily cadence rather than at request time. Offline, SIFT improves booking and amenity-engagement PR-AUC by +51.9% and +62.8% respectively over the production baseline. In online A/B testing, SIFT increased engagement with recommended filters by +20.0%, overall filter usage among searchers by +0.72%, and usage of the newly-supported bedroom, bathroom, and bed filters by +3.9%, +10.7%, and +0.52% respectively. Demonstrating the system’s extensibility, we rapidly integrated a novel hotel-intent filter using the same shared representation, driving a +3.8% lift in uncancelled hotel bookings and a +0.76% lift in overall marketplace bookings. SIFT is now fully deployed in production, serving scalable personalization to millions of guests.
[LG-67] MASKerade: Token-Routed Mask Experts for Dense-to-MoE Upcycling
链接: https://arxiv.org/abs/2610.07809
作者: Mingyuan Zhang,Yue Bai,Zhongruo Wang,Yupin Huang,Yiyang Huang,Hailing Wang,Huimin Zeng,Yun Fu
类目: Machine Learning (cs.LG)
*备注:
Abstract:Sparsely activated Mixture-of-Experts (MoE) models increase model capacity without a proportional increase in per-token computation. Dense-to-MoE upcycling reuses pretrained dense models to construct such systems, commonly by copying feed-forward networks (FFNs) into independently trained experts. We introduce MASKerade, a dense-to-MoE training method that instead learns experts as sparse subnetworks of a frozen pretrained FFN. Each expert is defined by a learned binary mask, and a token-level router selects which masked FFNs to execute and combine. The router and mask scores are optimized jointly, while the underlying FFN weight values remain unchanged. This formulation supports neuron-structured, semi-structured, and unstructured experts within the same routing architecture. Our main configuration uses four 2:4 experts with top-2 routing, where two half-dense expert passes have the nominal FFN arithmetic of one dense pass, without requiring independent expert weight matrices. On five vision-language benchmarks with Qwen and Gemma backbones, this configuration achieves the highest performance among the compared baselines. Comparisons across mask granularities, routing interventions, and compute-matched controls distinguish the effects of learned connectivity from expert activation count. These results establish mask learning over frozen weights as a practical alternative for constructing token-routed MoE experts.
[LG-68] Common-Mode Errors Limit Low-Timestep Deep Spiking Q-Networks
链接: https://arxiv.org/abs/2610.07808
作者: Zijie Xu,Bingrui Guo,Yiding Sun,Yiting Dong,Zhile Yang,Zhaofei Yu
类目: Neural and Evolutionary Computing (cs.NE); Machine Learning (cs.LG)
*备注:
Abstract:Spiking neural networks (SNNs) offer sparse and event-driven computation, making them attractive for energy-constrained reinforcement learning (RL) on edge devices. In value-based RL, deep spiking Q-networks (DSQNs) combine such efficiency with action-value estimation for decision making. However, existing DSQNs often require multiple simulation timesteps for competitive performance, increasing computational and energy costs, whereas reducing the timesteps can cause substantial performance degradation. We investigate this degradation from the perspective of Q-value estimation errors. By decomposing errors across actions into common-mode and differential-mode components, we find that low-timestep DSQNs suffer disproportionately from common-mode errors shared across action values, which are particularly detrimental to temporal-difference learning through bootstrapped targets. Based on this finding, we propose Common-Mode Compensation Deep Spiking Q-Network (CMC-DSQN), which uses an auxiliary ANN to compensate for common-mode errors in the SNN outputs. At inference, greedy action selection can be performed directly from the SNN outputs, allowing the auxiliary ANN to be completely removed and preserving the energy efficiency of SNNs. Extensive experiments on Atari and MiniAtar environments demonstrate substantial performance improvements under low-timestep settings. CMC-DSQN outperforms state-of-the-art DSQN baselines by nearly 20% at T=2 and further surpasses the ANN baseline at T=4 .
[LG-69] Adaptive Mean Estimation by In-Context Learning: A Gradient-Flow Analysis
链接: https://arxiv.org/abs/2610.07804
作者: Martin Eppert,Krishna Balasubramanian,Subhro Ghosh,Jason Klusowski,Yan Shuo Tan
类目: Machine Learning (cs.LG); Machine Learning (stat.ML)
*备注:
Abstract:Prior Fitted Networks (PFNs) such as TabPFN now rival established statistical procedures across prediction and estimation tasks. A natural explanation is that PFNs have the property of statistical adaptivity, that is, they perform nearly as well as a method tailored to the true data-generating model for a heterogeneous set of models, while not being told which model the data comes from. We study how such adaptivity is learned in a controlled location-estimation problem. Each task is an unlabeled sample whose family is hidden: Gaussian data call for averaging, with error of order n^-1 , whereas uniform data are best estimated from their extremes, at the faster rate n^-2 . We also provide the example of a symmetric Gaussian mixture, for which a rate of \sigma^2_n/n can be attained. On scalar inputs, softmax attention computes the derivative of the empirical cumulant-generating function. A single primitive therefore both supplies features that distinguish the families and forms estimators interpolating between the sample mean and the mid-range. We combine attention experts through either a softmax mixture of experts or a gated linear unit (GLU), and analyze stagewise gradient flow. With \widetilde\Omega(n^1+\epsilon) pretraining tasks, the learned estimator is asymptotically efficient on Gaussian tasks, within a factor n^\epsilon of the minimax rate on uniform tasks, and order-optimal on mixtures in a shrinking-variance regime. These guarantees extend to new locations and longer contexts. A risk decomposition separates expert error, routing error and normalization error, which clarifies the architectural contrast. Softmax gating enforces normalization and exact translation equivariance, whereas the GLU must learn it: its dynamics separate into fast bias removal followed by slow expert selection. End-to-end experiments recover the predicted specialization.
[LG-70] Extending Pathwise Gradients to Discrete Random Variables via Finite-Order Relaxation
链接: https://arxiv.org/abs/2610.07786
作者: Donghan He,Luhuan Wu
类目: Machine Learning (cs.LG); Machine Learning (stat.ML)
*备注:
Abstract:Pathwise gradients are preferred for continuous random variables because they are unbiased, low variance, and work with a single sample. For discrete variables, however, the pathwise identity cannot generally be exact for every differentiable function. We propose a general framework to construct finite-order exact pathwise gradient estimators for a range of common discrete variables such as Poisson. The estimator is the least-norm solution among all solutions that are unbiased for polynomials of degree at most. The resulting estimators preserve the hard forward sample, require no temperature tuning, and can be implemented in a few lines of codes. Against other admissible solutions, our estimator is unique and minimizes weight variance; in contrast, prior works use categorical variables or augmented representations to approximate non-categorical variables that induces excess variance and computations. To understand approximation bias for functions beyond the prescribed class, we also derive a non-asymptotic bias bound. In experiments our low order methods match or improve tuned baselines across linear, nonlinear and hierarchical latent-variable models, while out-speeding competitors in every runtime benchmark.
[LG-71] Cite What You Explore: Budget-Aware LLM Reasoning over Medical KGs with Verifiable Evidence NEURIPS2026
链接: https://arxiv.org/abs/2610.07739
作者: Chen Chen,Dongjie Wang,Mei Liu,Zijun Yao
类目: Machine Learning (cs.LG)
*备注: Accepted at NeurIPS 2026 (Poster)
Abstract:Post-discharge risk prediction from electronic health records (EHRs) is difficult because many dependencies that link discharge-time observations to downstream complications, such as comorbidity cascades and drug-disease interactions, are absent from the record. External medical knowledge graphs (KGs) can supply these missing dependencies, but tracing them demands three properties: KG exploration must remain cost-bounded, retrieved evidence must be differentiated by source quality, and the resulting rationale must be citable for retrospective review. Large language models (LLMs) can plan and verify over structured evidence, making them natural candidates for KG reasoning, but existing LLM-based methods do not satisfy these three properties jointly. In this paper, we propose BAR, a Budget-Aware LLM Reasoning framework over medical KGs with three contributions. First, BAR refines the raw KG into disease-specific evidence graphs whose edges carry support scores and provenance records, turning the KG into a quality-annotated reasoning space rather than a static feature source. Second, an LLM then reasons over this graph through a plan-navigate-verify loop that decomposes the question into steps, retrieves evidence under a patient-specific budget, and revises when verification fails. Third, a reasoning policy is trained with a reward that compares predictions with and without acquired evidence, combined with acquisition cost and citation-integrity terms. Across 8 diseases and 3 prediction horizons on MIMIC-III and MIMIC-IV, BAR improves AUPRC by 3.4 points over the strongest baseline, raises citation precision from 59.8% to 77.9%, and consumes only 62-65% of the budget cap.
[LG-72] he Model Plants the Trigger: Answer-Side Backdoor Attacks in Multi-Turn Large Language Models
链接: https://arxiv.org/abs/2610.07723
作者: Yibo Zhang,Tianrong Guan,Liang Lin,Puze Wang,Jin Wang,Qingsong Wen
类目: Cryptography and Security (cs.CR); Machine Learning (cs.LG)
*备注:
Abstract:Safety alignment in Large Language Models (LLMs) remains vulnerable to backdoor attacks. Existing LLM backdoors are almost all input-centric: activation depends on explicit trigger patterns in the user input, so modern guardrails are built to sanitize the input space. We challenge this assumption with a novel answer-side backdoor for multi-turn dialogue. Instead of inserting the trigger into the input, the adversary uses a benign first-turn prompt to naturally induce the model to generate a specific, seemingly innocuous word. Once merged into the dialogue history, this self-generated word becomes the trigger. When a later harmful query arrives, the model detects its own trigger and bypasses its safety refusal, while the user input stays perfectly clean. Across four LLMs, our attack reaches near-perfect Attack Success Rates, approaching 100% at only a 5% poisoning rate, while preserving general utility and clean-input safety, and it evades mainstream input-centric defenses. Representation-level analysis shows that the self-generated trigger consistently suppresses the model’s refusal signal, exposing a critical blind spot in current LLM defenses.
[LG-73] Neuromotor Hierarchy Network: Physiological Inductive Biases for Robust Generalization in sEMG Decoding
链接: https://arxiv.org/abs/2610.07713
作者: He Wang,Hongyuan Qi,Zhaoxian Zhang,Jinbin Luo,Linyi He,Mehul Motani,Changsheng Wu
类目: Machine Learning (cs.LG)
*备注:
Abstract:Surface electromyography (sEMG) provides a wearable, noninvasive interface to neuromuscular activity for movement decoding and human-computer interaction. Population-scale decoding remains difficult because the relationship between sEMG and neuromuscular activity varies across users and sessions, while task-relevant dynamics span channels and multiple timescales. Learning waveform-to-output mappings from task labels leaves the distinction between recording variability and coordinated motor activity implicit. We introduce the Neuromotor Hierarchy Network (NHN), which learns a compact latent neuromotor state from task supervision to represent task-relevant neuromuscular coordination. NHN constructs this latent state through a hierarchy inspired by neuromotor this http URL adapts recording statistics while preserving relative this http URL spatiotemporal encoder uses parameter-efficient channel interactions and modulates features with multi-timescale history. The resulting features yield candidate activations of learned motor primitives, which are temporally integrated and continuously weighted to form the state. Theoretical analysis characterizes the efficiency, temporal behavior, and optimization of NHN’s core mechanisms. We evaluate the architecture for both continuous hand-pose estimation on emg2pose and touch-typing recognition on emg2qwerty. On emg2pose, NHN reduces user-averaged angular error by 0.52% to 2.84% across all three generalization splits in both Regression and Tracking relative to Hadidi et al.'s best task-specific variants, using 48.42% to 48.51% fewer parameters. On emg2qwerty, NHN reduces beam-search character error rate by 19.40% zero-shot and 30.42% after fine-tuning relative to SplashNet-Upscale, using 65.86% fewer parameters. Physiology-guided inference of a latent neuromotor state supports parameter-efficient sEMG decoding.
[LG-74] Adaptive Model Inversion Attacks Generalize a Privacy-Robustness Tradeoff
链接: https://arxiv.org/abs/2610.07677
作者: Shailen Smith,Rasmus Torp,Adam Breuer
类目: Machine Learning (cs.LG); Cryptography and Security (cs.CR)
*备注:
Abstract:In this paper, we show that standard evaluations of high-resolution Model Inversion Attacks (MIAs) significantly underestimate training-data privacy leakage. State-of-the-art privacy defenses, standard training techniques such as MixUp and Adversarial Training, and undefended models all leak training images at rates 1.16 to 6.59 times higher on FaceScrub under simple adaptive changes to the attack, with the largest increases among defenses reporting the strongest privacy. We further show that measured leakage depends on the feature basis of the external classifier used to evaluate reconstructions: for the same reconstructed images, an adversarially trained Inception evaluator identifies the targeted identity at different rates than the standard Inception evaluator. Our results suggest that standard MIA evaluation can mistake optimization and measurement failures for privacy. These underestimated leakage rates also concealed a broader relationship between privacy and adversarial robustness. Once we adapt the attack and vary the evaluator, reconstruction leakage closely tracks adversarial robustness across recent defenses and standard training regimes, suggesting that robustness provides an attack-agnostic proxy for reconstruction vulnerability that applies far more broadly than previously theorized. This raises an open question: can a practical defense reduce training-data reconstruction without paying a corresponding cost in adversarial robustness? Subjects: Machine Learning (cs.LG); Cryptography and Security (cs.CR) Cite as: arXiv:2610.07677 [cs.LG] (or arXiv:2610.07677v1 [cs.LG] for this version) https://doi.org/10.48550/arXiv.2610.07677 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[LG-75] MS-ECG-FM: Towards a More Universal Electrocardiogram Foundation Model for Health Monitoring using Multi-source Contrastive Learning
链接: https://arxiv.org/abs/2610.07662
作者: Robert A. Lewis,I-Min Chiu,Kyle Verrier,Karthik Jayaraman Raghuram,Francoise Marvel,Salar Abbaspourazad,Anshuman Mishra,Guillermo Sapiro,Andrew C. Miller,Joseph Futoma
类目: Machine Learning (cs.LG)
*备注: Andrew C. Miller, Joseph Futoma: equal contribution. 52 pages, 4 figures, 20 tables, including supplementary information
Abstract:Electrocardiography (ECG) records the electrical activity of the heart, aiding diagnosis by detecting abnormalities in cardiac function. ECG foundation models have demonstrated promising results, but are limited by a reliance on ECG interpretation reports as their sole supervision. Because interpretation reports only capture the subset of waveform information routinely recognized by clinicians, this constrains representation learning to overlook the broader diagnostic signals present in ECG. We introduce a new ECG foundation model — MS-ECG-FM — that is trained through contrastive alignment to multiple distinct clinical note types, including ECG, echocardiography, radiology, and discharge reports. We evaluate MS-ECG-FM on an extended set of ECG detection benchmarks, showing that it comprehensively outperforms existing methods on the full span of conditions that ECG can detect, including in reduced-lead configurations. Different reports improve representations for different diagnostic domains, while multi-source alignment captures their complementary information and produces consistently strong representations across clinically diverse tasks.
[LG-76] Complementary Supervised and Self-Supervised Representations for Out-of-Distribution Graph Learning
链接: https://arxiv.org/abs/2610.07628
作者: Qingying Hao,Zikang Chen,Chuxuan Hu,Jinyuan Jia,Bo Li,Gang Wang,Carl Gunter
类目: Machine Learning (cs.LG)
*备注:
Abstract:Out-of-distribution (OOD) generalization remains challenging for graph neural networks (GNNs), as graph distributions can vary substantially across time and domains. Supervised and self-supervised graph representation learning are guided by distinct objectives and offer different perspectives on graph representations. In this work, we study whether self-supervised representations (SSL) can provide complementary signals to improve supervised OOD node classification. We develop two backbone-agnostic frameworks that exploit such information at different stages of learning and prediction. Co-Train jointly learns supervised and SSL representations and adaptively integrates them during training, while Dual-Space Retrieval performs non-parametric prediction in the two representation spaces and combines their predictions through confidence-aware fusion at inference time. The supervised and SSL encoders are separately parameterized and need not share the same GNN architecture. We evaluate multiple GNN backbones and two distinct SSL objectives, DGI and GRACE, on four graph benchmarks spanning temporal and cross-domain distribution shifts. Extensive experiments show that Co-Train consistently outperforms strong supervised OOD baselines, while Dual-Space Retrieval achieves competitive performance as a flexible non-parametric alternative. Results across different backbones and SSL objectives, together with representation analyses and ablations, demonstrate that SSL representations provide complementary information to supervised representations and can improve OOD node classification across diverse settings. Subjects: Machine Learning (cs.LG) Cite as: arXiv:2610.07628 [cs.LG] (or arXiv:2610.07628v1 [cs.LG] for this version) https://doi.org/10.48550/arXiv.2610.07628 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[LG-77] AFA-BANDIT: Provably Near-Optimal Online Multi-Feature Classification Under Budget Constraints
链接: https://arxiv.org/abs/2610.07615
作者: AbdAlRahman Odeh,Teng-Hui Huang,Hesham El Gamal
类目: Machine Learning (cs.LG)
*备注: 10 pages, 3 figures
Abstract:Active Feature Acquisition (AFA) is a classification problem in which an agent decides which costly features to acquire before predicting each sample’s label. Unlike batch AFA, which trains a fixed policy and classifier offline on fully observed data, online AFA updates its predictor from revealed labels as samples arrive. Existing online methods either use deep reinforcement learning (RL) without performance guarantees or maximize cost-adjusted reward rather than enforce a global budget. We formulate online AFA as a combinatorial Bandits with Knapsacks (BwK) problem that couples acquisition and prediction. Unlike prior bandit-based AFA and classical BwK, our setting has combinatorial complexity, evolving rewards, a global budget, and structured side information. We obtain an improved regret upper bound over standard BwK bounds in this framework, leveraging a cardinality-aware confidence bound and the subset update structure. To avoid an exponentially large action space, we propose \emphLP-Chain, a variant that searches a cost-aware chain of feature subsets with a size that grows linearly with the number of features. While the regret upper bound is specific to the combinatorial framework, \emphLP-Chain empirically achieves comparable predictive performance. On synthetic data, \emphLP-Chain outperforms HEDGE-based BwK and deep RL-based online AFA baselines and scales favorably to more features.
[LG-78] Learning Grasp Targeting from Point Clouds for Log Pile Clearing on a Hydraulic Crane
链接: https://arxiv.org/abs/2610.07613
作者: George Sideris,Lucas Bessai,Heshan Fernando,Elie Ayoub,Nicolas Lemieux,Inna Sharf
类目: Robotics (cs.RO); Machine Learning (cs.LG)
*备注: 9 pages, 15 figures. Supplementary video: this https URL
Abstract:In mill yards, log loaders clear dense piles by a sequence of bundle grasps: hundreds of logs rest in contact, and each removal changes the pile available to the next grasp. A learned policy chooses where to place and orient the grapple from unsegmented point clouds and runs on a trailer-mounted hydraulic forestry crane. The policy classifies at which observed point to grasp and predicts depth and grapple orientation there. The same network outputs support behavior cloning (BC), reinforcement learning (RL), and deployment. BC learns from successful top-of-pile demonstrations; RL explores for improvements by fine-tuning the cloned policy (BC \to RL) or by training from scratch. In simulation, BC clears 98 of 100 piles of 200 logs, while BC \to RL improves load stability. Twelve field trials compare a geometric heuristic, RL from scratch, BC, and BC \to RL through complete grasp-transport-deposit cycles. BC and BC \to RL deposit 93.8% and 88.9% of pooled inventory, against 80.4% for the heuristic. BC \to RL deposits logs on 83.6% of its cycles, against 79.6% for the heuristic and 65.7% for BC, while its simulated stability gain does not carry over to the crane testbed. Trained entirely in simulation and run unchanged on the crane, the learned policies clear more than the hand-filtered heuristic while observing unfiltered clouds that still contain the storage rack’s rails and poles.
[LG-79] Hub for Outliers Spokes for Inliers: Uniform Latent Space Construction for Dual-Mismatched Semi-Supervised Learning
链接: https://arxiv.org/abs/2610.07610
作者: Li Yuan,Yaxin Hou,Jiawei Tang,Yongbiao Gao,Yuheng Jia
类目: Machine Learning (cs.LG)
*备注: Equal contribution by Li Yuan and Yaxin Hou. Corresponding author: Yuheng Jia. Emails: {yuan-li,yaxin,230259148,yhjia}@seu. this http URL , gaoyb@qlu. this http URL . 18 pages
Abstract:Semi-supervised learning typically assumes that labeled and unlabeled data share an identical class distribution and label space. However, this setting is often violated: unlabeled data may be imbalanced and contain unknown class samples, causing mismatches in both class distribution and label space. Such dual mismatch leads to majority classes dominating the latent space and unknown class samples being overconfidently misclassified, degrading feature discriminability and pseudo-label quality. To address this, we propose a hub-spoke latent geometry, where known classes are uniformly distributed around a central hub and each class forms compact clusters around its prototype, while the hub provides an anchor for a low-evidence region specifically designed for high-uncertainty unknown class samples. Integrated with an evidence-based classifier, this geometry ultimately enhances feature discriminability and uncertainty separation by mitigating majority-class domination through structured feature organization and guiding high-uncertainty unknown class samples toward the hub. Extensive experiments show that our method outperforms state-of-the-art methods, with a maximum improvement of 3.25% across various settings.
[LG-80] A Neural JKO Scheme for Hellinger-Kantorovich Gradient Flows via Monge-Growth Pairs
链接: https://arxiv.org/abs/2610.07602
作者: Geuntaek Seo,Cheolhyeong Kim,Hwijae Son,Hyung Ju Hwang
类目: Numerical Analysis (math.NA); Machine Learning (cs.LG); Analysis of PDEs (math.AP); Optimization and Control (math.OC)
*备注: 55 pages, 10 figures
Abstract:We develop a mesh-free neural JKO scheme for advection-reaction-diffusion equations with a gradient-flow structure in the Hellinger-Kantorovich (HK) geometry of unbalanced optimal transport. Each update is parametrized by a spatial map and a mass-changing factor, allowing spatial redistribution and local mass creation or loss to be treated jointly within a single variational step. Their cone action bounds the squared HK distance from above, yielding a sufficient condition for discrete energy dissipation through comparison with the identity pair. Minimizing the pair objective over all admissible pairs recovers the exact JKO minimum when the source and a minimizer have positive densities. We establish existence and mass bounds for JKO minimizers and, under additional assumptions, obtain positivity and regularity together with a discrete Euler-Lagrange equation and a metric-dissipation identity. The self-consistent chemical potential is then nonincreasing along an optimal map. There exist parametric pairs whose endpoint densities and objective values converge to those of an exact JKO minimizer, provided a regular-pair approximation hypothesis holds. Finally, we show that a primal-dual gap controls objective suboptimality and, for Boltzmann entropy, the L^1 density error, assuming exact-step regularity, positive-semidefinite interactions, and global dual feasibility. Numerical experiments examine pointwise agreement with the PDE, energy dissipation, and the roles of transport, reaction, and fully implicit interactions.
[LG-81] he Robot Is Not Its Description: GaugeBench for Representation Robustness in Morphology-Aware Policies
链接: https://arxiv.org/abs/2610.07597
作者: Rahath Malladi,Arshia Sangwan,Rajesh K. Gupta,Tauhidur Rahman
类目: Robotics (cs.RO); Machine Learning (cs.LG)
*备注:
Abstract:A robot description does more than specify a physical mechanism: it also encodes arbitrary conventions, such as joint-axis direction, joint-angle zero, and the order and names of links and joints. Morphology-aware policies consume interfaces built from these descriptions, yet cross-embodiment evaluation typically changes the robot while keeping those conventions fixed. This leaves a simple question unanswered: does behavior survive when the robot stays fixed but its description changes? GaugeBench isolates this case by rewriting a fixed mechanism under physically equivalent conventions, verifying that its physics and policy interface are preserved, and then evaluating the same policy weights. The result is stark: three MetaMorph policies score 4030.6 on 80 familiar robots, but only 51.6 when those same robots are equivalently re-described, while 98 genuinely held-out robots score 1489.6. A new description can therefore be more damaging than a new robot. Tracing the failure reveals that axis reversal alone reproduces the collapse, joint-angle zero changes are nearly harmless, and reordering lies between them; moreover, changing joint-state and torque coordinates alone is sufficient to cause the failure, while changing description-derived features alone is not. The same phenomenon appears in ModuMorph and an unrelated PyBullet framework. Yet it is not irreversible: exact two-description transport restores the original controller, and training across equivalent axis conventions raises retained return under axis reversal from 3.6% to 80.6%. Together, these results separate mechanism robustness from representation robustness and show that cross-embodiment evaluation should test both.
[LG-82] BiGym 2.0: Benchmarking Learned and Agent -Developed Policies for Humanoid Household Manipulation
链接: https://arxiv.org/abs/2610.07594
作者: Zexi Zhang,Zecheng Zhu,Zidong Chen,Zulkhuu Tuya,Stephen James
类目: Robotics (cs.RO); Machine Learning (cs.LG)
*备注:
Abstract:Humanoid household manipulation requires the arms to act while the body balances, steps and changes posture. We present BiGym 2.0, an adaptation of BiGym for the Unitree G1 across 20 household tasks using a unified whole-body controller for demonstration and evaluation. The suite provides 60 native human virtual-reality demonstrations per task with synchronised multi-camera views and full-body execution records. We benchmark vision-language-action fine-tuning, imitation learning, demo-driven reinforcement learning, and cold-start coding agents given the interaction budget of online reinforcement learning. With the same onboard views, proprioception and whole-body controller for every method, vision-language-action fine-tuning has the highest nine-task mean, and agent-developed programs outperform every demo-driven reinforcement learning baseline on this mean and lead on bimanual reaching. Cross-workspace stacking remains open, \pi_0.5 stays low on pick-box, and multi-object transport is hard for imitation learning, demo-driven reinforcement learning and coding agents. All environments, human demonstrations, and evaluation traces are open-sourced at this https URL.
[LG-83] AFFY: A Task-Adaptive Tabular Foundation Model with In-Context Diversity
链接: https://arxiv.org/abs/2610.07559
作者: Zijian Li,Xiangchen Song,Gongxu Luo,Jie Qiao,Ruichu Cai,Zhenhao Chen,Xinshuai Dong,Fan Feng,Guangyi Chen,Kun Zhang
类目: Machine Learning (cs.LG)
*备注:
Abstract:Recent progress in tabular foundation models suggests that training on synthetic tasks can substantially improve in-context learning capabilities, with overall performance largely depending on how well models can infer task-specific predictive relationships from the available context during inference. In this paper, we introduce TAFFY, a tabular foundation model with an In-Context Diversity Prior and a Task-Conditioned Looped Transformer that strengthen this ability. Specifically, to construct each synthetic pretraining context, the In-Context Diversity Prior samples from multiple related environments derived via controlled interventions and distribution shifts on a shared causal process. This in-context diversity encourages the model to learn a more comprehensive and task-specific representation. Moreover, the Task-Conditioned Looped Transformer iteratively and selectively applies a shared group of Transformer blocks to refine contextual representations, with a task-conditioned gate modulating the final hidden-state update. This enables task-adaptive iterative refinement. Together, these components encourage the model to identify predictive relationships from contextual contrasts during pretraining and dynamically modulate context integration for each task. Across six classification and five regression benchmark datasets, TAFFY attains the lowest average rank.
[LG-84] Global Transport Couplings for Classifier-Free Guided Flows
链接: https://arxiv.org/abs/2610.07555
作者: Katarina Petrović,Zander W. Blasingame,Danyal Rehman,İsmail İlkan Ceylan,Michael Bronstein,Stephen Y. Zhang,Lazar Atanackovic,Alexander Tong
类目: Machine Learning (cs.LG)
*备注:
Abstract:Optimal-transport couplings have been shown to reduce training variance in unconditional flow models, but their role in conditional generation remains unclear. A natural approach constructs separate couplings for each condition, but this is impractical for large or continuous conditioning spaces found in modern image foundation models. We introduce Global Transport (GT), a global class-agnostic optimal-transport coupling, computed without class labels. GT can associate different conditions with different regions of the source noise, and consequently worsens performance without guidance. However, when combined with classifier-free guidance (CFG), GT consistently improves generation across domains, model scales, and sampling budgets. This reversal suggests that couplings for conditional flows should be evaluated both empirically and theoretically under the guided flow used at inference, rather than on unguided generation. We evaluate GT over both discrete class and continuous text conditioned image generation across model scales, and investigate how coupling choice alters guided trajectories. These results identify coupling design in the guided flow setting as a simple training time axis to improve performance without modifying existing architectures, samplers, or guidance mechanisms.
[LG-85] Preserving Unstable Modes Through Inverse Dynamics in JEPA World Models
链接: https://arxiv.org/abs/2610.07540
作者: Leonardo F. Toso,Yann LeCun,James Anderson,Oumayma Bounou
类目: Machine Learning (cs.LG); Robotics (cs.RO); Systems and Control (eess.SY); Optimization and Control (math.OC)
*备注:
Abstract:Robotic systems often exhibit unstable modes, along which small perturbations and disturbances can cause unbounded growth unless corrected through feedback. Controlling such systems from high-dimensional visual observations requires representations that preserve these modes. Joint-embedding predictive architectures (JEPAs) provide a natural framework for learning such representations and their dynamics from visual data. However, we demonstrate that next step prediction combined with anti-collapse regularization does not guarantee that controllable unstable modes are preserved: the training loss can be minimized while these modes are collapsed, making stabilization from the learned representation impossible. To address this, we augment world-model training with an action reconstruction objective (i.e., an inverse dynamics loss) that encourages control-aware representations, namely, visual representations that preserve crucial features for control. We prove that exact action reconstruction makes the encoder injective on the finite-horizon reachable subspace. Thus, the encoder cannot discard any state direction reachable by an action sequence within H steps. Moreover, we show that, as H grows, the dominant eigenspace of the finite-horizon controllability Gramian converges to the controllable unstable subspace. We establish our theoretical results for linear systems and demonstrate empirically that our findings extend to nonlinear visual control tasks (CartPole, Walker2D, and PointMaze), highlighting the benefits of control-aware representation learning.
[LG-86] SkillFormer: Skill-Decomposed Adaptation for Audio Language Models
链接: https://arxiv.org/abs/2610.07533
作者: Lee Seung-woo,Bowen Qi,Kim Min-jun,Jang Won-young
类目: ound (cs.SD); Machine Learning (cs.LG)
*备注:
Abstract:Audio language models must handle dozens of distinct skills, from pitch comparison and speaker counting to musical tempo estimation and emotion recognition. Joint training on all skills at once causes interference: gains on one skill often come at the cost of another. We propose \textbfSkillFormer, which decomposes audio understanding into skill-specific low-rank adapters and composes them at inference time through a learned router. The router examines the question to decide which adapters to activate and how much weight each should carry, so that a pitch query engages different parameters than a genre classification query. An alternating training schedule updates each adapter on its own skill cluster before jointly calibrating the router, preventing the gradient conflicts that arise in standard multi-task optimization. SkillFormer adds fewer than 4% of the base model’s parameters and requires no changes to the audio encoder or language backbone. Evaluated on three architecturally distinct models across MMSU, MMAU-Pro, and MMAR, it raises the average accuracy by 2.5 to 4.1 points, with balanced gains across perception, reasoning, and semantic subcategories.
[LG-87] argeted search shows that random-device testing underestimates worst-case error in a simulated wave-based neural operator
链接: https://arxiv.org/abs/2610.07529
作者: Samrendra Roy,Jason Yoo,Souvik Chakraborty,Syed Bahauddin Alam
类目: Machine Learning (cs.LG); Emerging Technologies (cs.ET); Optics (physics.optics)
*备注: 50 pages (19 main text and references, 31 Supplementary Information), 5 figures, 1 table
Abstract:Wave-based processors promise fast, energy-efficient Fourier layers for neural operators. They are usually validated on randomly sampled devices, but using them requires knowing how large their error can become under fabrication and alignment variation. In a stylised numerical case study, a hybrid Fourier neural operator runs its four spectral layers on simulated coherent 4f processors with 32 toleranced knobs, whose half-widths are representative rather than calibrated. For 120 models (four tasks, six training methods, five seeds), we compared the worst of N random in-spec devices with a searched one. On a deterministic simulator with one frozen draw of the random static errors, the searched device’s held-out error was 1.08-3.10 times the maximum over 200 Monte Carlo devices and 1.06-2.71 times that over 1000. With 20 fresh static draws, it still exceeded the maximum over 200 random devices in 116 of 120 models. Under uniform sampling, the probability of drawing such a device is at most 0.37% per model (two-sided 95% Clopper-Pearson), which says nothing about how large its error is. The gap persisted with uniform or Sobol’ sampling at the search’s budget, shared knobs, a second crosstalk model, box scales of 0.25-2 and a pixel-level device model. Models trained only with random static errors reached 3.7-39.9 times their nominal error on searched devices, and fine-tuning on random and gradient-searched devices gave the lowest searched error of the six in all 20 task-seed pairs. For two heat-exchanger quantities, a search targeted at each exceeded the worst of 1000 random devices in all 39 models, and hence the Wilks 95/95 limit (worst of 59). For the mean pressure of 11 models, no random device exceeded a 1% error threshold, but the searched device did. Random testing estimates how often errors exceed a threshold; worst-device search gives a lower bound on how large they can be.
[LG-88] Activation Denoising: A Robustness View on Parallel vs Sequential LLM Quantization
链接: https://arxiv.org/abs/2610.07522
作者: Yan Scholten,Rachel Lawrence,James Hensman,Stephan Günnemann,Alicia Curth,Riccardo Grazzi
类目: Machine Learning (cs.LG)
*备注:
Abstract:Post-training quantization is a powerful tool for compressing large language models. The most scalable methods quantize every layer in parallel, but quantization errors then compound through the residual stream, as no layer corrects for the errors of the layers before it. Sequential quantization accounts for this error compounding by re-calibrating each layer on the already-quantized outputs of its predecessors, yielding stronger results but at the cost of a serial schedule that becomes a bottleneck at scale. As a solution, we propose parallel quantization with activation denoising, which recovers much of the sequential benefit while keeping quantization fully parallel. Rather than re-calibrating layer-by-layer, we take a robustness perspective and model the upstream error as noise, regularizing to be robust to it through a preprocessing step followed by metric-weighted rounding. Applied at every layer, this regularization forms a depth-compounding smoothness penalty that dampens how strongly quantization errors amplify through the model. Unlike orthogonal rotations commonly used in quantization, which must preserve the model’s function, we multiply the weights by a more general linear transformation. We find that the two are complementary and their effects compound. Empirically, our robustness regularization recovers a significant part of sequential quantization’s benefit in a single parallel pass, at a fraction of its time. Overall, by treating compounding quantization errors as a robustness problem, we offer a principled foundation for more efficient and accurate LLM quantization at scale.
[LG-89] Source-Learned Reliance for Selective Test-Time Adaptation of Multimodal Time Series
链接: https://arxiv.org/abs/2610.07499
作者: Payal Mohapatra,Yueyuan Sui,Haodong Yang,Benjamin Lundell,Stephen Xia,Qi Zhu
类目: Machine Learning (cs.LG)
*备注:
Abstract:Multimodal wearable systems must remain reliable when sensor streams become noisy or unavailable. Existing multimodal test-time adaptation (TTA) methods often assess reliability online, but cross-modal agreement can be misleading when sensors measure different physical processes, and evaluating alternative modality configurations adds inference cost. We propose CARAT, which decouples model reliance from runtime corruption detection to guide omission or attenuation, amortizing reliance estimation through source training. An asymmetric modality-dropout curriculum prepares a missingness-resilient backbone for omission and derives a frozen, backbone-specific reliance proxy from windowed input-projection gradient norms. At deployment, a lightweight one-class detector flags suspect streams, and the proxy guides a joint choice between replacing the suspect set with the backbone’s trained missingness symbol and attenuating its representations before fusion, without candidate-subset evaluation. Across four wearable datasets, five corruption types, three backbones, and eight TTA baselines, CARAT achieves the highest overall macro-F1 and best mean rank (2.42), exceeding EATA, the strongest baseline, by 1.58 F1 points across 12 equally weighted dataset-backbone settings. Across five profiled configurations, CARAT uses 9.49% fewer GFLOPs and updates 47.82% fewer parameters than EATA. A pattern also emerges across sensing regimes: multimodal TTA methods such as PTA are competitive on IMU-dominated homogeneous datasets, whereas unimodal TTA methods like TENT and EATA match or exceed it on heterogeneous datasets. These results position CARAT as a practical default to wearable TTA, offering competitive robustness with modest computational requirements and benefits that vary across backbones and dataset regimes.
[LG-90] Deep Defence on Wheels: A Dual Intrusion Detection System Architecture for Comprehensive In-Vehicle Network Security
链接: https://arxiv.org/abs/2610.07489
作者: Shashwat Khandelwal,Shanker Shreejith
类目: Cryptography and Security (cs.CR); Hardware Architecture (cs.AR); Machine Learning (cs.LG)
*备注: 30 pages, 9 figures, 11 tables, ACM Transactions on Embedded Computing Systems
Abstract:Increasing connectivity to the outside world and the lack of inbuilt security mechanisms have made legacy intra-vehicular networks vulnerable to cyberattacks. Initial research focused on maximising detection accuracy for known and unknown attacks, often using large, full-precision machine learning models. However, embedding IDSs into vehicular electronic systems also requires low detection latency, energy efficiency and minimal electronic control unit (ECU) resource overhead to process about 2,000 CAN frames/s. Lightweight models must balance accuracy with these deployment constraints. We propose a dual IDS framework comprising supervised and unsupervised learning-based solutions, each optimised for real-time, resource-constrained automotive platforms. A quantised LSTM-based IDS (QLSTM-IDS) achieves over 99.9% detection accuracy for DoS/Flooding, Fuzzing and Spoofing/Malfunction attacks using a single model architecture evaluated on two widely used datasets. The model is trained using the Brevitas quantisation-aware training library, transformed into a dataflow accelerator with custom blocks compatible with AMD’s FINN toolchain, and synthesised using Vitis HLS. Complementing this, an 8-bit quantised convolutional autoencoder-based IDS (QCAE-IDS), quantised using AMD’s Vitis-AI toolchain, detects previously unseen anomalies that alter CAN-ID sequence patterns with over 99% accuracy. An integration architecture enables both models to operate on a single FPGA, bridging the network interface IP and processing system to minimise software overhead. QLSTM-IDS achieves 0.25 ms inference latency and 0.8 mJ energy consumption per message, while QCAE-IDS achieves 0.42 ms and 1.1 mJ per block. Both solutions are deployed and evaluated on the ZCU104 SoC (XCZU7EV FPGA), demonstrating a flexible hardware/software co-design for real-time detection of known and unknown attacks on high-speed CAN buses.
[LG-91] Robust Importance Sampling for Rare Events via Constrained Gaussian Mixtures NEURIPS2026
链接: https://arxiv.org/abs/2610.07485
作者: Paweł Lorek,Rafał Nowak,Rafał Topolnicki,Tomasz Trzciński,Maciej Zięba
类目: Machine Learning (cs.LG)
*备注: Accepted at NeurIPS 2026
Abstract:We study estimating rare-event probabilities I = \mathbbP(g(\mathbfX) \gamma) with \mathbfX \sim \mathcalN(\boldsymbol\mu, \boldsymbol\Sigma) and general g : \mathbbR^d \to \mathbbR . We address this problem through importance sampling, and propose a framework that substantially improves efficiency and robustness over baselines such as crude Monte Carlo, adaptive cross-entropy, variational-inference-based methods (including reverse- and forward-KL approaches), as well as Safe-ICE, Subset Simulation, and Sequential Monte Carlo, drawing on ideas from both rare-event estimation and cross-entropy optimization. The key contribution has two parts: first, we separate the problem into coverage, to overcome the cold-start barrier, and fitting, to refine proposals once a meaningful signal is available; second, we constrain the final GMM proposal so that it has finite importance-sampling variance (since coverage alone is not sufficient – without safeguards, importance sampling may still suffer from infinite variance). Together, these ingredients yield expressive proposals; finite variance does not by itself guarantee practical stability at a fixed sampling budget. Extensive experiments demonstrate substantial variance reduction, strong robustness across diverse benchmarks, and favorable cost–efficiency trade-offs, with the proposed approach often outperforming these baselines, particularly in high-dimensional and multimodal settings where competing methods frequently become unstable or fail. Our code is available at this https URL.
[LG-92] SpecBraM: What Should an EEG Foundation Model Predict? Masked Band-Power Prediction versus Waveform Reconstruction
链接: https://arxiv.org/abs/2610.07484
作者: Peng Xie,Yequan Bie,Jianda Mao,Kani Chen
类目: Machine Learning (cs.LG)
*备注: 12 pages, 3 figures, 9 tables; includes an appendix
Abstract:Self-supervised EEG models often reconstruct masked waveforms or predict discrete codes. We study a task-aligned alternative: masked band-power prediction (MBP), which predicts fixed narrow-band log spectral energy for masked channel-time patches. This target retains rhythm power relevant to sleep staging while avoiding phase-sensitive waveform reconstruction and a learned codebook. Across three pretraining seeds, we compare band-power and waveform targets with matched backbones, pretraining data (2,388 hours), and training steps, including a 2x2 tokenizer-by-target design. On ISRUC and HMC sleep staging, MBP exceeds raw- and band-waveform reconstruction by 1.6-2.8 balanced-accuracy points with all labels and 4.7-7.3 points with 1% of labels under a strict linear probe; the target effect exceeds the tokenizer effect. Its frozen features reach 0.7916/0.7425 balanced accuracy, versus 0.7636/0.7227 for a matched rich handcrafted spectral baseline, although the gap is about one point with 1% of labels. Full fine-tuning reaches 0.8107/0.7669. The gains do not extend to every task with spectral cues, including motor imagery, depression screening, and vigilance regression. These results support choosing pretraining targets to match the physical quantities and spatial and temporal scales relevant to downstream labels.
[LG-93] Adapting to Changes in Agent Behavior via Finite-Depth Policy Sensitivity
链接: https://arxiv.org/abs/2610.07475
作者: Lan Shi,Daigo Shishika,Xuan Wang
类目: Machine Learning (cs.LG)
*备注: 8 pages, 4 figures
Abstract:Adapting a reinforcement learning policy to changes in another agent’s behavior typically requires a large amount of new interaction data. Policy sensitivity provides a first-order prediction of how a locally optimal policy changes with a behavioral parameter, but its computation requires second-order derivatives whose effects propagate across future interactions. We develop a finite-depth framework to estimate this sensitivity by approximating the policy Hessian and mixed derivative using information from a reference environment. The method features an adjustable propagation depth which determines where derivative propagation along the trajectory is truncated. We characterize the derivative contributions omitted by finite-depth propagation and derive truncation-error bounds for the approximated derivatives and resulting policy sensitivity. The bounds are nonincreasing with propagation depth and vanish at full-horizon propagation. Using a belief-driven pursuit-evasion game as a validation scenario, the proposed method generally achieves lower derivative-estimation errors as the propagation depth increases and outperforms the baseline methods in both estimation accuracy and policy adaptation. The sensitivity-based initialization improves zero-shot return over direct transfer, and also shows advantages for the subsequent fine-tuning in the target environment.
[LG-94] Efficient Multimodal Inference through Adaptive Acquisition and Sequential Fusion
链接: https://arxiv.org/abs/2610.07466
作者: Payal Mohapatra,Haodong Yang,Yueyuan Sui,Stephen Xia,Benjamin Lundell,Qi Zhu
类目: Machine Learning (cs.LG)
*备注:
Abstract:Multimodal systems often encode every available input, even when a subset suffices for prediction. Adaptive acquisition can reduce this cost by using predictions from incrementally fused evidence to decide which modality to encode next and when to stop. However, sequential fusion makes these predictions order-dependent, so decisions based on them may need to distinguish factorially many histories of the same acquired set. We introduce SemARC, which couples a Sequential Modality Aggregator (SeMA) with an Adaptive Runtime Controller (ARC) and uses acquired evidence to select each modality before its encoder runs. SeMA executes only selected encoder and fusion branches, updates a fixed-size state, and predicts after each acquisition without recomputing earlier branches. We supervise every acquisition prefix under randomized modality subsets and orders to encourage consistent predictions across acquisition orders. ARC combines a set-dependent marginal-utility prior with residual fitted-Q learning to select the next available modality or stop, without inspecting unacquired inputs or retaining acquisition order. Across six multimodal classification datasets and eleven baselines, SemARC achieves 3.2% higher macro-F1 and 61.4% lower total inference GFLOPs on average relative to each dataset’s most accurate baseline. End-to-end latency falls by 44.0% across GPU and CPU and by 47.2% on Android INT8 relative to the fastest measured baseline, on average. Under varying runtime modality missingness, SemARC still skips available modalities, matching or exceeding the best baseline macro-F1 in 21 of 24 conditions with 14.8% lower total GFLOPs on average. SemARC thus offers a practical path toward efficient multimodal inference across heterogeneous devices.
[LG-95] Interpretable Hypergraph Learning via Neural Additive Models
链接: https://arxiv.org/abs/2610.07458
作者: Shihan Feng,Xin Zheng,Shiyi Yang,Ren Wang,Chudi Zhong,Can Chen
类目: Machine Learning (cs.LG); Social and Information Networks (cs.SI)
*备注: 16 pages, 8 figures, 13 tables
Abstract:Hypergraphs offer a natural framework for modeling networked data, where dependencies among entities are governed by higher-order interactions. While hypergraph learning methods such as hypergraph neural networks have demonstrated remarkable predictive performance, most existing approaches rely on black-box message-passing architectures, making it difficult to disentangle the contributions of node attributes and higher-order structural information. To address this challenge, we introduce the hypergraph neural additive network (HGNAN), an inherently interpretable framework for learning on hypergraph-structured data. HGNAN extends classical neural additive models to higher-order relational data by integrating feature-wise nonlinear decomposition with hypergraph-aware structural aggregation, enabling transparent prediction for both node- and hyperedge-level tasks. Extensive experiments on benchmark datasets demonstrate that HGNAN achieves performance comparable with state-of-the-art hypergraph learning methods while providing intrinsic and meaningful interpretability.
[LG-96] Fork-and-Flush: Escaping Idea Basins in Autoresearch Agents
链接: https://arxiv.org/abs/2610.07447
作者: Ziyang Cai,Christos Ziakas,Vasilis Kontonis,Tim Pearce,Siddhartha Sen,Akshay Krishnamurthy,Shivam Garg,Dimitris Papailiopoulos
类目: Machine Learning (cs.LG)
*备注:
Abstract:Autoresearch agents tackle open-ended problems by repeatedly proposing candidate solutions, evaluating them, and using feedback to guide subsequent experiments. We show that independent runs of the same agent on the same task often plateau at substantially different scores, with gaps that persist even after considerable additional compute. Embedding their candidate artifacts by functional similarity provides further evidence that trajectories remain in localized regions of the solution space, which we call idea basins. To help agents escape these basins, we study a simple periodic intervention, fork-and-flush. Our method forks the agent into parallel trajectories, each inheriting the accumulated workspace but starting with a fresh chat context. After running each trajectory for a fixed horizon, the agent continues from the highest-scoring one. Across 13 long-horizon research and engineering tasks, with individual agent runs lasting up to several days, fork-and-flush outperformed the single-run and best-of-N baselines by a relative improvement of 66.0% and 44.4%, respectively, on the min-max normalized average score under an equal compute budget.
[LG-97] StaFIR: Convex Learning of Stationarity-Aware Causal Filters NEURIPS2026
链接: https://arxiv.org/abs/2610.07430
作者: Lorena Egger,Mathis Linger
类目: Machine Learning (cs.LG)
*备注: Accepted at the TS-LIMITS Workshop at NeurIPS 2026
Abstract:Reducing nonstationarity in a persistent time series entails deciding how much of its temporal dependence to remove. In finance, fractional differencing is often tuned using the Augmented Dickey–Fuller (ADF) test, limiting the search to a one-parameter family of lag profiles and addressing input preservation only indirectly. We propose StaFIR, a causal finite-impulse-response filter with a learned nonnegative mixture of exponential lag profiles. Its convex learning objective balances empirical stationarity with similarity to the input. We evaluate StaFIR on ARFIMA–GARCH controlled settings and rolling financial series, including a realized-volatility forecasting task. The experiments show that StaFIR adjusts its filtering strength to persistence while limiting unnecessary transformation in stationary regimes. In downstream forecasting, there is no clear accuracy difference from fixed half-order differencing, while StaFIR achieves higher measured similarity to the raw signal. A complementary direct forecasting experiment finds that greater input similarity is associated with smaller forecasting penalties, although the raw representation remains stronger.
[LG-98] Benchmarking Label-Revealed Online Updates for EEG BCI Decoding
链接: https://arxiv.org/abs/2610.07420
作者: Bogdan Kozyrskiy,Artem Grachev,Abraham I. Camelo Guerrero
类目: Machine Learning (cs.LG); Signal Processing (eess.SP)
*备注:
Abstract:Electroencephalography (EEG) signals drift over time, which can cause static brain-computer interface (BCI) models to degrade in practice. We present a benchmark for online adaptation and compare two widely used pipeline families, Common Spatial Patterns (CSP) and Riemannian covariance-based methods, under time-ordered prequential (test-then-train) evaluation. We examine (i) which pipelines benefit most from label-revealed updates, (ii) whether controlled forgetting of older data improves robustness, and (iii) how a minimal-calibration cold start compares with starting from a pretrained model. Across four datasets (three motor-imagery datasets and one movement-decoding dataset), label-revealed online updates improve 13 of 14 model/dataset pairs on the two largest streams, with relative accuracy gains of up to about 18% over a frozen model. A Shapley-based data-valuation analysis over temporal blocks assigns the largest mean value to the most recent block in each of the three analyzed datasets, while older blocks retain positive value.
[LG-99] Evaluation of Active Feature Acquisition Policies with Tabular Foundation Models
链接: https://arxiv.org/abs/2610.07406
作者: Yuta Kobayashi,Divyam Madaan,Shalmali Joshi
类目: Machine Learning (cs.LG)
*备注:
Abstract:Active feature acquisition learns policies that sequentially acquire features to maximize information about a target variable. We study how to learn and evaluate such policies from finite offline data using prior-data fitted networks (PFNs), which are off-the-shelf models that output posterior predictive distributions without task-specific training. We show that under the imbalanced coverage of offline data, using total predictive entropy as a reward creates an epistemic bias that penalizes acquiring sparsely observed features. Specifically, this reward conflates epistemic uncertainty (arising from lack of offline data) with aleatoric uncertainty (arising from uninformative features). To address this, we target the posterior expected (aleatoric) entropy instead of the total predictive entropy output by a PFN for evaluating feature acquisitions. Empirical evaluations on synthetic and real-world datasets demonstrate that our approach consistently reduces value estimation bias and yields credible intervals with strong empirical coverage, which can translate to improved downstream policy selection.
[LG-100] What pass@k Cannot Measure: Evaluating Diversity and Capability Retention after Post-Training NEURIPS2026
链接: https://arxiv.org/abs/2610.07405
作者: Subham Rath,Raj Dandekar,Rajat Dandekar,Sreedath Panat
类目: Machine Learning (cs.LG)
*备注: 10 pages, 2 figures. Accepted to the NeurIPS 2026 Workshop on Transitioning from Pre-training to Post-training (non-archival)
Abstract:pass@ k , the fraction of problems a model solves within k sampled attempts, is the field’s default protocol for deciding whether reinforcement-learning (RL) post-training on verifiable rewards improved a model. At the population level, pass@ k depends only on a problem’s probability of a correct sample, with no term for how it is distributed across outputs. We show this gap is not academic. Training Qwen2.5-1.5B-Instruct on grade-school math with Group Relative Policy Optimization (GRPO) and with rejection-sampling fine-tuning (RFT, training on the model’s own shortest verifier-passed rollout) moves three complementary diversity measures (token-level entropy, answer-level entropy, unique answers per prompt) in opposite directions, with zero overlap across three seeds per arm. The gap survives restricting to verifier-correct completions only (lexical diversity among correct solutions is 15% lower for GRPO, after controlling for length) and a count-controlled check isolating diversity among incorrect answers alone, ruling out that GRPO’s higher accuracy alone explains it. Yet pass@8 and pass@32 show no consistent winner on GSM8K, and a hard MATH-500 subset shows the same pattern: separation only at low k . Compared against the starting checkpoint, no trained arm significantly improves hard-problem coverage: RFT is significantly worse, while GRPO is statistically indistinguishable from it - so GRPO’s pass@1 edge over RFT reflects a smaller loss relative to Base, not a capability gain, a missing-control issue, not a failure of pass@ k . On GSM8K, only pass@1, with no role in detecting diversity by construction, separates the arms cleanly, rewarding the arm whose correct solutions are least diverse. We argue this is a concrete instance of a standard evaluation protocol missing a property it is routinely used to certify.
[LG-101] Fed-BRDECS: Privacy-Preserving and Heterogeneity-Aware Federated Deep Embedded Clustering
链接: https://arxiv.org/abs/2610.07399
作者: Haemin Park,Diego Klabjan,Martin W. Braun,Xiuqi Li,Balakrishnan Ananthanarayanan
类目: Machine Learning (cs.LG)
*备注:
Abstract:Federated deep clustering seeks to learn clustering-friendly representations from decentralized unlabeled data while preserving client privacy. However, Deep Embedded Clustering (DEC)-style objectives depend on global soft-assignment statistics that require clients to reveal their sensitive information. We propose Fed-BRDECS, a privacy-preserving and heterogeneity-aware federated deep embedded clustering framework. Fed-BRDECS replaces the globally normalized clustering objective with a locally computable sample-stability loss, avoiding the transmission of local soft-assignment distributions. To tackle non-IID client distributions, we introduce prediction-balanced sampling, which oversamples locally rare predicted clusters without requiring ground-truth labels, and centroid-level restarting, which periodically refreshes biased or inactive centroids. Experiments on image and text clustering benchmarks show that Fed-BRDECS consistently outperforms representative federated clustering and deep clustering baselines under both IID and non-IID partitions. We further demonstrate its applicability to federated time-series anomaly detection, where it improves reconstruction-based detectors without adding inference-time cost.
[LG-102] Multigroup Fairness and Omniprediction: Separations and Equivalences NEURIPS2026
链接: https://arxiv.org/abs/2610.07374
作者: Sílvia Casacuberta,Parikshit Gopalan,Varun Kanade,Omer Reingold,Konstantinos Stavropoulos,Pranay Tankala
类目: Machine Learning (cs.LG)
*备注: Accepted for presentation at NeurIPS 2026
Abstract:Omniprediction is a learning guarantee which requires a single predictor to be competitive relative to the best hypothesis from a benchmark class for any loss chosen from a family of loss functions. Loss Outcome Indistinguishability (loss OI for short) is a stronger notion that implies omniprediction. It requires the predicted distribution on labels to be indistinguishable from the true distribution to tests that depend on the loss functions and the benchmark class. Multiaccuracy and multicalibration are multigroup fairness notions that generalize classical notions of calibration and accuracy in expectation. Most known learning algorithms for omniprediction (both for the standard notion and for strengthenings like loss OI) rely on some version of these multigroup fairness notions, or on an intermediate notion called calibrated multiaccuracy. We ask if this is necessary: Does omniprediction require some form of multigroup fairness? We show that the answer is no for (plain) omniprediction, and yes for loss OI. First, a sequence of works shows that multicalibration or calibrated multiaccuracy imply omniprediction. We rule out even a weak converse, by showing that omniprediction for proper losses does not imply even accuracy in expectation, a much weaker notion than any of calibration, multiaccuracy, or multicalibration. Second, prior work showed how to achieve loss OI from a combination of calibration and multiaccuracy. We show a converse: loss OI is equivalent to a form of calibrated multiaccuracy. Comments: Accepted for presentation at NeurIPS 2026 Subjects: Machine Learning (cs.LG) Cite as: arXiv:2610.07374 [cs.LG] (or arXiv:2610.07374v1 [cs.LG] for this version) https://doi.org/10.48550/arXiv.2610.07374 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[LG-103] Dynamic Budget Allocation for LLM Evaluation under Hard Resource Constraints
链接: https://arxiv.org/abs/2610.07362
作者: Shai Feldman,Yaniv Romano
类目: Machine Learning (cs.LG)
*备注:
Abstract:We evaluate large language models (LLMs) in multi-turn interactions through their time-to-event: the number of interaction steps required to produce an event of interest, such as a successful jailbreak or agentic task completion. Under limited compute, interactions may be terminated before the event occurs, so that event times are only partially observed (censored). Existing allocation methods for calibrating time-to-event bounds satisfy the budget only in expectation and can exceed the available budget on a particular evaluation run. Enforcing a hard constraint is particularly challenging as the cost of a trajectory is initially unknown. We introduce Hard-budget Allocation with Reflow for Predictive calibration (HARP), a budget allocation that satisfies hard resource constraints and adaptively reallocates unused budget. We show how to use HARP to construct lower predictive bounds (LPBs) on the time-to-event and to estimate evaluation metrics such as the jailbreak rate on a fixed benchmark. Although HARP induces dependence in acquisition decisions across different trajectories, we prove that HARP never exceeds the target budget, that its LPBs have finite-sample coverage guarantees, and that its metric estimates are unbiased. Experiments on agentic task success, LLM jailbreaks, toxic content generation, and RAG hallucinations show that HARP achieves coverage close to the nominal level with low variance, while never exceeding the given budget.
[LG-104] owards Explainable Benchmarking for Data-driven Post-Wildfire Debris Flow Prediction
链接: https://arxiv.org/abs/2610.07358
作者: Zhisheng Qi,Li Zhu,Utkarsh Sahu,Douglas Tommey,Josh Roering,Yu Wang
类目: Machine Learning (cs.LG)
*备注:
Abstract:Post-wildfire debris flows (PFDFs) are destructive sediment-laden hazards triggered when intense rainfall strikes recently burned terrain, destabilizing hillslopes and threatening infrastructure, local economies, and community safety. Data-driven methods have been proposed to learn predictive patterns directly from historical PFDF observations. However, the current research landscape of data-driven PFDF prediction remains highly fragmented across feature spaces, model architectures, and evaluation protocols, making rigorous comparison and the derivation of scientific insights difficult. Moreover, existing studies lack a systematic investigation into the relative importance of heterogeneous factors (e.g., meteorological conditions, terrain characteristics, soil properties, and burn severity) in triggering PFDF. To address these limitations, we present a unified benchmark for data-driven PFDF prediction, enabling fair and comprehensive evaluation across diverse models and feature configurations. Furthermore, to better understand the underlying drivers of PFDF formation, we propose a reinforcement learning-based feature selection framework that identifies factors whose perturbations render positive and negative events indistinguishable, thereby discovering the regional underlying mechanisms of PFDF occurrence across regions. Our code and benchmark are publicly available at this https URL.
[LG-105] Evaluating Behavioral Context for Interpretable IAM Policy Risk Scoring in Cloud Environments
链接: https://arxiv.org/abs/2610.07345
作者: Yassin Elsharkawy
类目: Cryptography and Security (cs.CR); Distributed, Parallel, and Cluster Computing (cs.DC); Machine Learning (cs.LG); Networking and Internet Architecture (cs.NI)
*备注: 6 pages, 5 figures, 5 tables. Peer-reviewed and accepted at IEEE Conference ID# 71863; to appear in the IEEE Xplore proceedings
Abstract:IAM policy analysis typically emphasizes the authorization capabilities encoded in a policy, but security analyst review priority may also depend on the behavioral and environmental context surrounding a policy event. This paper evaluates whether contextual information provides measurable incremental value for interpretable IAM policy risk prioritization beyond policy and effective-authorization information. AWS is used as the experimental cloud provider because its IAM and audit-telemetry ecosystem enables controlled evaluation using AWS IAM Context Bench, a benchmark containing 534 real AWS experimental observations across policy, environment, and behavioral scenarios, including matched cases where policy and environment remain fixed while behavioral context changes. Three Explainable Boosting Machine models are evaluated under the same leakage-controlled grouped cross-validation protocol: a policy-centric baseline, a policy-plus-environment model, and a full-context model incorporating CloudTrail telemetry. The full-context model substantially reduces analyst-priority prediction error relative to the policy-centric baseline and closely tracks the reference priority ordering. In matched same-policy context pairs, the policy-centric model remains invariant, whereas the full-context model separates benign and suspicious behavioral conditions with high directional accuracy. The results also show improved concentration of high-priority cases at the top of simulated analyst review queues. These findings indicate that behavioral and environmental context can provide useful incremental information for analyst-oriented IAM risk prioritization while preserving an interpretable additive model structure. The formulation is applicable beyond AWS conceptually, although cross-provider validation remains future work.
[LG-106] Weight Oracles: Reading Neural Network Weights with Language Models NEURIPS2026
链接: https://arxiv.org/abs/2610.07334
作者: Krishna Kabra,Constantin Venhoff,Christian Schroeder de Witt
类目: Machine Learning (cs.LG); Cryptography and Security (cs.CR)
*备注: Spotlight at the NeurIPS 2026 Workshop on Neural Network Artifacts as a New Data Modality (NeuralArtifacts), Paris. 14 pages, 10 figures
Abstract:Interpretability methods for neural networks are predominantly reactive: they analyse activations produced during specific forward passes, requiring known inputs to find hidden capabilities such as backdoors. We propose Weight Oracles, fine-tuned language models that diagnose properties of a target network by reading its raw weights directly, without behavioural testing. We investigate this paradigm in two phases. Phase I establishes feasibility: through a staged curriculum and an external chain-of-computation that delegates parameter-free operations to deterministic code, an explainer LLM learns to simulate the forward pass of small transformers from their weights, achieving 99% holdout accuracy on unseen targets. Phase II repurposes this infrastructure for safety auditing. We train an oracle on natural language diagnostic questions about weight anomalies using only benign pathologies as training signal, and evaluate it zero-shot on backdoors absent from training. The oracle achieves AUROC 0.93 on attention-routed backdoors and 0.81 across a diversified threat distribution including stealth and adversarially regularized variants. Hand-crafted statistical detectors are sharp on the threat models they implicitly target but collapse on threat-model shift, while the oracle remains uniformly competent across attack types. Scaling to realistic model sizes remains the principal open challenge.
[LG-107] Lock-in EP: An In-Situ Training Algorithm for Oscillatory Hardware
链接: https://arxiv.org/abs/2610.07283
作者: Sowjanya Tammali,Wilkie Olin-Ammentorp
类目: Machine Learning (cs.LG)
*备注:
Abstract:Analog hardware platforms offer the potential to reduce energy consumption over digital architectures, but in order to succeed, large-scale analog systems must also be able to operate with or recover from the variability of their components. Towards this goal, we derive and demonstrate the lock-in equilibrium propagation (LIEP) training method. LIEP provides local gradient information for each component in an oscillatory network without separate forward and backward sweeps, potentially allowing for in-situ learning capabilities on analog oscillatory hardware platforms. We demonstrate that LIEP can be used both for ab-initio training as well as recovering performance when pre-trained parameters are perturbed. We show that LIEP can be formulated as a three-factor update rule, and suggest that although the method is currently only validated on shallow networks, alternate architectures may allow it to extend to deep and large-scale networks addressing complex tasks.
[LG-108] Algorithmically Aligned Neural Agglomerative Tree Construction
链接: https://arxiv.org/abs/2610.07271
作者: Robert R Nerem,Pranav Singh,Cheyenne Ward,Yusu Wang
类目: Machine Learning (cs.LG)
*备注:
Abstract:Linkage algorithms for hierarchical clustering (HC) are a powerful and efficient framework for constructing clustering trees, yet it is often unclear which merge rule best suits a given dataset or task. In contrast, neural approaches can learn from data, but often fail to retain the efficiency and size generalization of classical algorithms. We introduce NN-linkage, a neural network (NN) model that can learn task-specific and locally dependent merge rules while retaining the recursive structure and efficient inference of classical linkage algorithms. In particular, our model is algorithmically aligned with the Lance-Williams (LW) recurrence, a parameterized framework for defining a broad, continuous family of linkage rules for agglomerative HC. Classical methods such as single linkage (SL), complete linkage (CL), and average linkage arise as discrete choices within this broader family. We show that NN-linkage is a universal approximator for continuous linkage functions, including LW recurrences, and, when paired with a transformer encoding, can also approximate globally dependent rules such as robust single-linkage. We further show that NN-linkage can exactly implement any symmetric constant-coefficient LW recurrence across all input sizes. On the empirical front, we evaluate NN-linkage in real-world applications, clock-tree routing and phylogenetic reconstruction, using both synthetic and real datasets, demonstrating its effectiveness over both classical algorithms and other neural approaches. By learning merge rules directly from target trees, NN-linkage extends efficient HC to scientific and engineering objectives not adequately captured by existing hand-designed linkage rules.
[LG-109] Neural Algorithmic Reasoning for Graph Saddle Point Problems
链接: https://arxiv.org/abs/2610.07255
作者: Samantha Chen,Jesse He,Coleman Clougherty,Gal Mishne,Chester Holtz
类目: Machine Learning (cs.LG)
*备注:
Abstract:Neural algorithmic reasoning, or aligning a neural network with an algorithmic paradigm, has emerged as an approach to solving polynomial-time-solvable and computationally harder combinatorial optimization problems. We propose a new message-passing framework based on the Chambolle-Pock Primal–Dual Hybrid Gradient (PDHG) method called \textscGraphPDHG for solving general graph saddle-point problems. Theoretically, we show that \textscGraphPDHG can efficiently solve a family of graph saddle-point problems by simulating PDHG. We also show that our network can learn an accelerated PDHG algorithm. Experimentally, we support our results on accelerated PDHG by evaluating the performance of our model as a learned warm start for second-order optimization techniques (SSNAL). We also show that alignment with PDHG leads to stronger size generalization than non-aligned graph neural network (GNN) baselines. Overall, we propose a novel architecture for solving a general family of optimization problems on graphs.
[LG-110] Neural Fields Encode Adaptation Geometry NEURIPS2026
链接: https://arxiv.org/abs/2610.07253
作者: Prateik Sinha,Stefania Druga
类目: Machine Learning (cs.LG)
*备注: 24 pages, 2 figures. Extended version of work accepted at the NeurIPS 2026 Workshop on Symmetry and Geometry in Neural Representations (NeurReps)
Abstract:Neural fields are usually evaluated by how well they reconstruct an observation. We show that this misses two useful properties of a fitted network: how easily it can adapt to new observations, and what its weights retain from earlier ones. We study these properties as adaptation geometry. For images, we meta-learn class-specific initializations, adapt each one to a new image, and measure how much the network must change to fit it. A simple local linear model closely predicts this adaptation cost, while replacing one network’s tangent kernel with another’s substantially worsens the prediction. Adaptation thus depends on the local geometry of the fitted network, not only on its current reconstruction. For physical fields, we repeatedly fit the same network to observations from a sequence. Its weights then retain information about that history. When two wave histories end at exactly the same observation, the final weights recover the sign of the wave velocity with 68.6% accuracy, whereas the current observation alone contains no such information and gives 50%. These two phenomena are quantitatively linked: tangent-kernel eigenvalues predict both which changes are easy to learn and how quickly they are overwritten by later fitting. Together, these results show that neural fields contain useful information beyond what they currently reconstruct: in how they can change and in how they got there.
[LG-111] Learning What to Distill: Bilevel Top-K Token Selection for Self-Distillation in Large Language Models
链接: https://arxiv.org/abs/2610.07247
作者: Heng Liang,Xinwen Zhang,Hongchang Gao
类目: Machine Learning (cs.LG)
*备注:
Abstract:Large language models have shown strong reasoning capabilities, but their high inference costs make knowledge distillation an important approach for transferring such capabilities to compact models in resource-constrained scenarios. On-policy self-distillation further reduces the reliance on external large teacher models while improving the reasoning ability of compact language models. However, existing methods typically either distill all token positions uniformly or select tokens using fixed heuristic criteria, assigning the same distillation strength to the selected positions rather than adaptively learning which tokens are most beneficial for distillation. To address these limitations, we propose BiToK-SD (Bilevel Top-K Token Selection for Self-Distillation), a bilevel-optimization-based token selection method that learns where distillation should be applied during on-policy self-distillation. Specifically, BiToK-SD is formulated as a bilevel optimization problem, where the lower-level problem models Top-K token selection as a differentiable threshold-based relaxation, allowing the selected positions to adapt as the student policy evolves, while the upper-level problem performs knowledge distillation on the selected positions. Experiments on mathematical reasoning benchmarks show that BiToK-SD achieves the best average performance among all compared methods while requiring only lightweight additional computation.
[LG-112] Benchmarking Time Series Foundation Models for Load Forecasting Under Covariate Uncertainty ALT
链接: https://arxiv.org/abs/2610.07232
作者: Tomas Kaljevic,Ivan Arzola,Yu Zhang
类目: Machine Learning (cs.LG); Systems and Control (eess.SY)
*备注: 5 pages, 1 figure, 5 tables. Accepted to the 2027 IEEE PES Grid Edge Conference Expo, Salt Lake City, UT, USA, 19-22 April 2027
Abstract:Accurate short-term load forecasting (STLF) is essential for the reliable and efficient operation of modern power systems. While time series foundation models (TSFMs) have recently demonstrated remarkable performance across a wide range of forecasting tasks, their effectiveness for STLF under realistic operational conditions remains largely unexplored. In this paper, we present a comprehensive benchmark of four trained-from-scratch (TFS) models and four TSFMs across three real-world load forecasting datasets under operational scenarios that differ in the availability and quality of future covariate information. Our results show that Chronos-2 consistently achieves state-of-the-art performance in both zero-shot and fine-tuned settings when future covariates are available or accurately forecast. However, its performance degrades as covariate forecasts become increasingly noisy, whereas TimesNet exhibits greater robustness under severe covariate uncertainty. These findings demonstrate the effectiveness of covariate-informed TSFMs for STLF while highlighting the critical role of robust covariate modeling in real-world forecasting applications.
[LG-113] Conditional Flow Matching for Transport Between Markov Processes
链接: https://arxiv.org/abs/2610.07229
作者: Syamantak Kumar,Dheeraj Nagaraj,Saptarshi Roy,Purnamrita Sarkar
类目: Machine Learning (cs.LG); Machine Learning (stat.ML)
*备注:
Abstract:Motivated by sequence-to-sequence transport in the context time-series domain adaptation, we study the problem of transportation between trajectories of Markov processes. Given a limited number of trajectories from source distribution and the target distribution, we formulate a flow matching based algorithm which learns a transport map from the source to target trajectory distribution, while preserving the Markov structure. We show that this is consistent in the population limit and derive finite-sample error bounds under mixing time assumptions, following the analysis of classical statistical problems including regression (Nagaraj et al., 2020), principal component analysis (Kumar and Sarkar, 2023), and matrix concentration (Neeman et al., 2024) in the Markov setting. We complement that with a lower-bound construction showing that a mixing-time dependent sample complexity is unavoidable even with regular Gaussian conditional transitions. We evaluate on synthetic and real-world data. For image retrieval from electroencephalography (EEG) on THINGS-EEG2 (Gifford et al., 2022), the task is to identify the viewed image from EEG signals captured from human subjects, which suffers from high inter subject variability. We augment the ENIGMA decoder (Kneeland et al., 2026) with a conditional flow before its subject-specific temporal map. This improves mean top-5 retrieval accuracy from 43.87% to 49.05%, an 11.82% relative improvement.
[LG-114] Data Numbers and Geometry: Three Tutorials on Numerical Methods Machine Learning and Evaluation
链接: https://arxiv.org/abs/2610.07220
作者: Jessica N. Howard,Yidi Qi,Tomás S. R. Silva
类目: Machine Learning (cs.LG)
*备注: Combined notes from three tutorials presented at the DANGER: Data, Numbers, and Geometry workshop (BIRS, Banff, April 2026). Includes links to companion code, Jupyter notebooks, and exercises
Abstract:We present three practical tutorials on numerical computation and machine learning for mathematical research, developed for the DANGER: Data, Numbers, and Geometry workshop held at the Banff International Research Station in April 2026. The first develops a numerical approach to exterior calculus from pointwise evaluations of differential forms, using a flux formulation of the exterior derivative. Examples in Euclidean space and on the sphere illustrate geometric identities, topological features, and the effects of approximation and finite precision. The second examines how mathematical structure guides neural network design through examples involving elliptic curves, quivers, and a boundary value problem. It explores how architectural choices affect learning and uses interval arithmetic to bound the residual of a trained network over the full interval of the boundary value problem. The third addresses the evaluation and presentation of machine learning results, covering performance metrics, statistical uncertainty, classification thresholds, receiver operating characteristic curves, and accessible figure design. Throughout, the tutorials distinguish numerical agreement, predictive accuracy, structural guarantees, and rigorous bounds as different forms of evidence. Each contribution can be read independently, with accompanying notebooks and exercises that allow readers to reproduce the examples and adapt the methods to other problems.
[LG-115] Constant-Curvature Sliced Gromov-Wasserstein for Heterogeneous Cross-Curvature Alignment
链接: https://arxiv.org/abs/2610.07218
作者: Shanglin Li,Wenjing Lu,Muyang Li,Nicu Sebe,Ziheng Chen
类目: Machine Learning (cs.LG)
*备注:
Abstract:Recent advances in representation learning have highlighted the utility of constant-curvature models, such as hyperbolic and spherical spaces, for modeling complex data. Mixed-curvature models further enhance this by integrating multiple constant-curvature components. However, these models typically learn each component space independently because spaces with different curvatures are inherently heterogeneous and lack a unified metric. Consequently, they lack explicit mechanisms to enforce geometric consistency across various spaces. Moreover, the problem of comparing probability distributions across mixed-curvature spaces remains unexplored. To compare distributions on heterogeneous spaces, Gromov-Wasserstein (GW) distances provide a principled framework by aligning their intra-space geometries. Building on this, we propose constant-curvature sliced Gromov-Wasserstein (CCSGW), a novel divergence for aligning distributions supported on heterogeneous constant-curvature spaces. We first introduce the missing geodesic-based one-dimensional projections for spherical spaces, and then extend sliced GW to constant-curvature spaces, enabling efficient and principled comparison across manifolds with different curvatures. This formulation preserves intrinsic geometric relationships while avoiding the high computational cost. We provide theoretical analysis showing that CCSGW controls intrinsic geometric discrepancy across heterogeneous spaces, promoting distribution-level geometric consistency. By integrating CCSGW into existing mixed-curvature learning tasks, including graph anomaly detection, graph node classification, and multimodal learning, we observe consistent performance gains across diverse settings.
[LG-116] Reward-Driven Learning under Prompt-Level Differential Privacy
链接: https://arxiv.org/abs/2610.07212
作者: Jiachen Zhao,Antonia Januszewicz,Taeho Jung
类目: Machine Learning (cs.LG); Cryptography and Security (cs.CR)
*备注:
Abstract:Reinforcement learning with verifiable rewards (RLVR) trains a language model on problems that may themselves be confidential, and the trained model can reveal which problems it saw. We study RLVR under prompt-level differential privacy: the released weights must be (\epsilon,\delta)-differentially private with respect to the presence of any one training problem. Taking the group of responses to one prompt as the privacy record, our method aggregates their gradients, clips the prompt’s contribution once, adds Gaussian noise, and composes the privacy loss across updates, so the budget depends on neither the number of responses per prompt nor the clipping norm; to our knowledge this is the first differential privacy guarantee for RLVR training. We train Qwen2.5-1.5B-Instruct with LoRA at a per-run budget of \epsilon=8 and compare, on the same prompts and at the same budget, a control that removes only the reward signal and two private supervised fine-tuning recipes. The reward signal improves accuracy over the control by 2.65 points on MATH and 3.24 on GSM8K, in every seed; the improvement survives a format-robust scorer, at 1.3 points on MATH, and is not explained by response length. At the same budget the private model outperforms both supervised recipes on MATH and GSM8K by 2.3 to 3.8 points, retains 85–90% of the gain of non-private GRPO on these tasks, and on MATH the noise of an eightfold tighter budget costs at most 1.2 points. The reward effect also carries to CommonsenseQA, an exploratory non-mathematical task. Verifier feedback thus remains a usable learning signal under prompt-level privacy.
[LG-117] An overview of machine learning-enhanced iterative methods for systems of linear and nonlinear equations
链接: https://arxiv.org/abs/2610.07211
作者: Yuhuang Meng,Jing Zhao,Alexander Heinlein
类目: Numerical Analysis (math.NA); Machine Learning (cs.LG)
*备注: 78 pages
Abstract:Systems of equations arise in a wide range of scientific and engineering applications. The present work focuses on solvers for general systems of equations, including but not limited to those arising from partial differential equations. These systems can be broadly categorized into linear and nonlinear problems. For large linear systems, iterative solvers are generally preferred over direct methods due to the latter’s superlinear growth of computational costs. Although convergence theory is well-developed under certain assumptions on the coefficient matrix, many classes of systems still pose open challenges. These difficulties become even more severe for systems of nonlinear equations, where nonlinear solvers typically rely on repeated linearization. For example, Newton’s method may even converge quadratically near the solution; it can also converge slowly or diverge when the initial guess is not chosen appropriately. A wide range of solvers with diverse variants and hyperparameter settings exists, and the development of efficient and robust iterative methods remains an active area of research. Recently, machine learning (ML) techniques have been applied to enhance the efficiency of classical iterative methods while preserving their interpretability and reliability. We refer to these ML-enhanced iterative methods as hybrid iterative methods, in the sense that they combine classical iterative methods with ML. This paper provides a comprehensive overview of state-of-the-art approaches to constructing hybrid iterative methods for systems of both linear and nonlinear equations, while also discussing open challenges and outlining potential directions for future research.
[LG-118] Can LLM -assisted regularization increase forecast accuracy for migration flows in low data regimes?
链接: https://arxiv.org/abs/2610.07208
作者: Nathaniel T. Hindman,Fabricio Murai
类目: Machine Learning (cs.LG); Computers and Society (cs.CY)
*备注:
Abstract:Predicting migration flows remains a significant challenge for traditional gravity-based forecasting models, which primarily rely on structured socio-economic indicators such as economic disparity, political stability, and geographic distance. This work investigates whether Large Language Models (LLMs) can improve migration forecasting by extracting contextual migration-related signals from news articles and incorporating them into a weighted Lasso forecasting framework through feature-specific regularization penalties. The proposed framework uses hierarchical LLM inference pipelines to classify migration-related push–pull signals from news data and evaluates the resulting forecasting performance across multiple migration corridors between November 2021 and November 2022, including Mexico–United States, Ukraine–Poland, and Syria–Turkey. Experimental results showed mixed performance across migration corridors and modeling strategies, and no single regularization approach consistently outperformed the others across all experiments. The best-performing Mexico configuration, which consisted of a gravity-based model augmented with the proposed push–pull ratios, achieved a Mean Absolute Percentage Error (MAPE) of 17.15%, while the strongest Syria configuration achieved a MAPE of 29.29% using Direct LLM-Lasso. For Ukraine, the best-performing configuration used LLM-Assisted Regularization (AR) and achieved a MAPE of 41.05%. Overall, the results suggest that contextual article-derived features and LLM-guided regularization can improve migration forecasting under certain conditions, although migration corridor characteristics, article volume, and hyperparameter configuration strongly influenced performance.
[LG-119] Exact Unlearning via Quantized Sufficient Statistics NEURIPS2026
链接: https://arxiv.org/abs/2610.07197
作者: Ami Tavory,Shripad Gade,Tal Sarig,Noam Touitou,Ido Guy
类目: Machine Learning (cs.LG)
*备注: 33 pages, 20 figures, 15 tables. Accepted at NeurIPS 2026
Abstract:Exact unlearning requires a deployed predictor to match one rebuilt without the information named by a deletion request. Existing general-purpose exact methods localize retraining through disjoint shards, but every request still invalidates a model, and smaller shards reduce the data available to each constituent predictor. We introduce Quantized Sufficient Statistics (QSS), which separates a small frozen schema from mutable, sum-decomposable content. The schema learns global structure; the content stores local prediction corrections as additive statistics indexed by quantized regions. Deleting content is therefore exact subtraction rather than optimization. We distinguish two guarantees: QSS-L exactly removes a label while retaining the unlabelled input, whereas QSS-E exactly removes both input and label by learning the schema without deletable examples. A deletion takes the arithmetic fast path with probability 1-\rho and triggers a full rebuild with probability \rho ; all reported expected latencies include both events. Across 15 vision, text, and tabular datasets at \rho=0.5% , QSS-L is within 2 percentage points of SISA on 11 tasks and provides 4–483 \times lower expected deletion latency on the low-class-count tasks where a compact schema is effective. QSS-E quantifies the additional accuracy cost of removing every trace of an input.
[LG-120] Interleaved Projected Gradient Descent for Safe Imitation Learning
链接: https://arxiv.org/abs/2610.07167
作者: Shengfan Cao,Francesco Borrelli
类目: ystems and Control (eess.SY); Machine Learning (cs.LG)
*备注: Submitted to the 2027 American Control Conference (ACC). 9 pages, 3 figures
Abstract:We propose an imitation-learning design for neural-network control policies under state and input constraints. Training alternates a standard imitation gradient step with a block of k safety steps that pull the network’s actions toward their projection onto the safe set; at run time, the controller is the trained network alone, with no safety filter. We analyze this scheme as inexact projected gradient descent in the space of policy actions. When the projected actions are recomputed at every safety step and each step moves the actions consistently toward the safe set, letting k grow logarithmically yields asymptotic constraint satisfaction on the training states and bounds the distance to the constrained optimum of the imitation loss; with the projected actions held fixed, the same holds only if they are exactly representable by the network. On a nonlinear autonomous racing task, we compare our method with adding a weighted constraint-violation penalty to the imitation loss. With a sufficiently large weight, our method matches the lap time of unconstrained imitation while reducing the fraction of violating episodes from 15% to 1% , about six times fewer than the penalty approach at its best weight. Its lap times are less sensitive to the weight, which instead sets how quickly violations vanish during training. In racing, the safety corrections are sparse and the conditions of the analysis do not hold; the gain arises instead through the data collected during training. These gains come at the cost of additional training computation.
[LG-121] Adversarial Training for Deep Hedging in Nonstationary Markets
链接: https://arxiv.org/abs/2610.07162
作者: Philipp J. Schneider,Lukas Looser,Antoine Garin,Shuhan Liu,Daniel Kuhn
类目: Machine Learning (cs.LG); Optimization and Control (math.OC); Computational Finance (q-fin.CP)
*备注:
Abstract:Deep hedging learns trading policies from historical or simulated market trajectories, yet under nonstationarity these training paths may not represent future market conditions. We propose WRAP (Wasserstein-Reweighting Adversarial Perturbation), a drift-aware adversarial training framework derived from a two-budget distributionally robust optimization (DRO) formulation. The formulation is anchored to a weighted empirical reference distribution whose fixed baseline weights are chosen to balance sampling uncertainty against temporal drift. Around this reference distribution, the ambiguity set addresses two complementary forms of distributional misspecification by allowing an adversary to reweight the observed trajectories subject to a \phi -divergence constraint and perturb their paths subject to an optimal-transport (OT) constraint. We derive a joint first-order expansion in which the leading-order increase over the nominal expected loss decomposes into a reweighting contribution determined by the dispersion of hedging losses across trajectories and a transport contribution determined by the sensitivity of the loss to path perturbations. This expansion yields an explicit finite-dimensional adversarial attack that replaces the distributional inner supremum with a tractable first-order approximation. Across stationary and nonstationary Heston dynamics and a generalized affine diffusion (GAD), the experiments show complementary benefits from reweighting and transport, with joint adversarial training providing the largest gains under nonstationarity.
[LG-122] he Implicit Bias of Hyperbolic Representation Learning for Multiclass Data: A Busemann Risk Perspective NEURIPS2026
链接: https://arxiv.org/abs/2610.07131
作者: Xingrun Li,Sho Kuno,Yusuke Mukuta,Xin Yang,Tatsuya Harada
类目: Machine Learning (cs.LG)
*备注: Accepted at NeurIPS 2026 as a Spotlight
Abstract:We study the implicit bias of Riemannian gradient flow for hyperbolic multiclass classification with fixed class prototypes in hyperbolic space \mathbbH^n . Our framework accommodates general permutation invariant relative margin (PERM) losses, a class that includes cross entropy and other standard multiclass losses. Our analysis is based on a decomposition: at large radius, the distance to each prototype splits into a radial term and a direction-dependent term described by the Busemann function. This yields two main results. First, we prove a radial dichotomy: the sign of a drift coefficient \mu determines whether the radius is pushed toward the ideal boundary or back toward the interior; if the positive drift persists, then r(t)=\frac12\log t+O(1) , while persistent negative drift returns the trajectory to the large-radius threshold in finite time. Second, we show that the boundary direction converges to a critical point of the Busemann risk on \partial\mathbbH^n . These results provide a rigorous asymptotic perspective on two phenomena we refer to as boundary saturation and near-boundary clustering in hyperbolic representation learning.
[LG-123] SoloQ: Calibration-Free Quantization for Diffusion Language Models
链接: https://arxiv.org/abs/2610.07121
作者: Donghyun Lee,Arkapravo Ghosh,Varun Manjunath,Bumjoon Kyle Rhee,Hyunho Kook,Shiting Xiao,Youngeun Kim,Priyadarshini Panda
类目: Machine Learning (cs.LG)
*备注:
Abstract:Diffusion large language models dLLMs) have emerged as a promising alternative to autoregressive language models through bidirectional diffusion-based token generation. However, their growing model sizes and high inference costs make efficient deployment challenging: full-sequence denoising repeatedly invokes compute-intensive forward passes, while block-diffusion models additionally introduce a memory-intensive KV-cache. Low-bit weight-activation quantization is therefore attractive, yet existing dLLM post-training quantization methods rely on calibration data despite activation distributions shifting across masking states and denoising steps. We present SoloQ, a calibration-free quantization framework that maps weights and activations into a normalized rotated basis with a predictable marginal distribution, enabling data-independent quantization. SoloQ combines a structured K-RPBH rotation with a lightweight rescaling correction for calibration-free quantization. Its predictable post-rotation distribution supports both distribution-matched codebooks and hardware-native NVFP4. For block-diffusion models, SoloQ further applies commit-time KV-cache quantization to compress persistent states without perturbing the actively denoised block. Across full-sequence dLLMs (LLaDA and Dream) and block-diffusion dLLMs(Fast-dLLM v2 and Nemotron-Labs-Diffusion), SoloQ retains accuracy under 4-bit quantization and outperforms calibration-based baselines on knowledge- and reasoning-intensive benchmarks. With NVFP4, SoloQ reduces peak memory by up to 2.61X and accelerates end-to-end inference by up to 2.24X.
[LG-124] Evaluating Inference Compute for Generative AI: A Framework for Enterprise Workloads
链接: https://arxiv.org/abs/2610.07094
作者: Abbas Raza Ali,Muhammad Ajmal Siddiqui,Moona Zahid
类目: Hardware Architecture (cs.AR); Distributed, Parallel, and Cluster Computing (cs.DC); Machine Learning (cs.LG); Performance (cs.PF)
*备注:
Abstract:LLM deployment is shifting from single-turn completion to agentic trajectories in which a model plans, calls tools, reads results and reasons at test time before acting. This inverts the economics of inference hardware: chat serving amortises weight reads across large batches, whereas agent trajectories are sequentially dependent, run at effective batch one, and make per-token decode latency (TPOT) the dominant term in task completion time. Using a roofline analysis and a closed-form episode-latency model, we show why this regime favours accelerators that keep weights in on-die SRAM (Cerebras WSE-3/3T, Groq/NVIDIA LPU) or compiler-managed tiered memory (SambaNova SN40L/SN50), and why three vendor ecosystems converged in 2026 on disaggregated prefill/decode serving. We show that per-step reliability compounds exponentially in trajectory length-a 2% per-step failure rate erases a 2x decode advantage for a 20-step agent-so determinism and tail latency are first-order performance variables. We then propose a four-layer evaluation framework (silicon, serving system, agent episode, enterprise) with a metric set built on goodput at an agentic SLO and cost per successful episode, a six-axis benchmark protocol over six task families, a paired-bootstrap statistical design, an attestation protocol for vendor-run benchmarks, and TCO, availability and adoption-timing models with explicit break-even conditions. All performance figures are public and labelled by evidence class; we state seven falsifiable hypotheses and the experiments that test them, and argue that the most likely original result is that token-throughput rankings diverge from cost-per-successful-task rankings on long-horizon work.
[LG-125] When Attention Does Not Explain the Peak: Temporal Reference vs. Forecast Output in Attention-Based Time-Series Forecasting NEURIPS2026
链接: https://arxiv.org/abs/2610.07080
作者: Yuji Akamatsu,Takao Yamanaka
类目: Machine Learning (cs.LG)
*备注: NeurIPS 2026: Accepted to the TAE (Trust-AI-Eval) Workshop
Abstract:Attention maps are often interpreted as evidence of what a forecasting model uses when making predictions. In our load-forecasting model, a CLS representation of historical demand queries 24 future exogenous horizon tokens through cross-attention, inviting a temporal interpretation in which highly attended horizons may appear to explain forecast peak timing. We test this interpretation using a horizon-level attention descriptor, \Psi_\mathrmout . Across 31 day-aligned windows of the Panama load dataset, the forecast achieves a median peak-time error of 0 h and a 51.6% exact-match rate, whereas the argmax of \Psi_\mathrmout has a median error of 5 h and 0% exact match. The forecast peak is closer to the observed peak in 27 of 31 windows. This dissociation is not merely an argmax artifact: within \pm1 h, attention reaches only 1.16\times , 1.11\times , and 1.14\times the uniform baseline around observed, predicted, and weekly-naive peaks, respectively, indicating weak and non-selective concentration. Yet the attention profile is structured, with cross-window consistency of 0.83. Replacing 12 future weather features with their training-set means makes the profile nearly uniform, showing sensitivity to future weather variation rather than fixed horizon position alone. The dissociation is also reproduced across three random-seed runs. These results show that structured, input-sensitive, and reproducible horizon-level cross-attention need not provide a valid peak-selective explanation of forecast behavior. The observed behavior is instead consistent with an internal horizon-reference role for integrating future exogenous information, although this functional role is not causally established.
[LG-126] Few-Shot Bioactivity Prediction with Meta-Learning under Assay Heterogeneity
链接: https://arxiv.org/abs/2610.07079
作者: Michal Kmicikiewicz,Tommy Rochussen,Vincent Fortuin,Ewa Szczurek
类目: Machine Learning (cs.LG)
*备注:
Abstract:Accurate bioactivity prediction is a central challenge in early-stage drug discovery, as individual assays often contain too few measurements to train reliable models independently. Meta-learning offers a principled approach to this few-shot setting, but assay heterogeneity may limit its effectiveness. Here, we test this hypothesis and show that meta-learning performance degrades as meta-training tasks become more heterogeneous. To address this, we introduce MetaHeta, a meta-learning framework that accounts for assay heterogeneity by conditioning predictions on auxiliary data from related assays, with relatedness defined flexibly from available assay information. The architecture of MetaHeta combines linear attention over large auxiliary datasets with exact attention over scarce task-specific context, enabling efficient scaling to the former without compromising exact attention over the latter. We demonstrate the benefits of our approach on assays from ChEMBL and BindingDB, improving few-shot bioactivity prediction and downstream compound prioritization in retrospective Bayesian optimization.
[LG-127] What Must Replay Preserve? Separating Correctable Bias from Class Correspondence
链接: https://arxiv.org/abs/2610.07077
作者: BoRen Deng,Xiangyue Ma,Chenglong Li,Xiaoting Du
类目: Machine Learning (cs.LG)
*备注:
Abstract:Class-incremental learning must recognize all classes seen so far without task labels. Logit replay methods such as DER and DER++ mitigate forgetting by matching the model’s past predictions on stored examples. Deleting this matching reveals its benefit, but the resulting accuracy cost cannot show whether the stored scores themselves are needed, or whether the cost survives correction of the classifier’s bias toward recent classes. We propose a diagnostic framework that treats a cached prediction as temporally heterogeneous supervision: it separates classes known when an example was stored from classes learned afterward, edits each group, and evaluates every model before and after a task-level offset that leaves within-task predictions unchanged. On CIFAR-100 with DER++, suitable fixed constants replace the unrefreshed stored scores of later-learned classes within an equivalence margin of 1 percentage point, and the offset reduces the cost of deleting their matching from 14.9 to 1.8 points. Reassigning the non-gold scores of classes known at storage, which preserves their values and each task’s target probability, costs 4.3 points before and 4.0 after the offset, and a parallel cost persists in image distillation. In the tested fixed-head setting, the large cost of deleting later-class matching is thus mostly correctable by this offset, whereas the smaller cost of disrupting class correspondence persists. Code and data are available at this http URL.
[LG-128] ImpactMat: Continuous Material Estimation for Inverse Impact Sound Rendering
链接: https://arxiv.org/abs/2610.07061
作者: Hyebin Cho,Bumsoo Kim,Joon son Chung
类目: ound (cs.SD); Machine Learning (cs.LG)
*备注: Preprint
Abstract:Impact sound rendering synthesizes the sound produced when a 3D object is struck, but practical renderers often rely on fixed material presets such as wood, plastic, or steel. These presets limit the range of impact sounds a renderer can express, while manually adjusting the underlying material parameters remains difficult without expertise in material acoustics. We therefore study inverse impact sound rendering: predicting material parameters from a reference impact sound so that a simulator can recreate a similar material response. To support this task, we introduce ImpactMat, a dataset and benchmark of single and blended material impact sounds paired with ground-truth material parameters. We further propose a feed-forward model that predicts these parameters from one or more recordings, using blended materials to learn smooth transitions between material types. Experiments show that our method outperforms competitive baselines and enables re-rendering from real recordings without manual parameter tuning. The project page is available this https URL.
[LG-129] Skillful Data-Driven Subseasonal Soil Moisture Forecasting: Prospects and Limits for Flash Drought Prediction
链接: https://arxiv.org/abs/2610.07060
作者: Noelia Otero,Atahan Özer,Miguel-Ángel Fernández-Torres,Jackie Ma
类目: Machine Learning (cs.LG); Atmospheric and Oceanic Physics (physics.ao-ph)
*备注: 27 pages, 8 figures, 5 tables. Accepted for publication in npj Hydrosphere. Supplementary information available with the published version
Abstract:Despite substantial progress in short-to-medium-range weather forecasting, predicting high-impact events such as flash droughts remains a key challenge for both early warning operations and physically-based subseasonal-to-seasonal (S2S) prediction systems. Here we demonstrate that, for S2S soil-moisture forecasting over Europe, forecast skill depends as much on how the prediction problem is formulated as on the forecasting model itself. Using a Vision Transformer-based architecture with dual-pathway temporal and spatial attention, we show that residual learning is essential to outperform persistence. This advantage is realized only when forecasting root-zone soil moisture in physical units rather than standardized anomalies, revealing that the target representation itself constrains predictability. A probabilistic extension via quantile-head fine-tuning further provides well-calibrated predictive distributions. Benchmarked against deep-learning and operational ECMWF S2S baselines over 2021-2022, our model achieves the highest deterministic and probabilistic skill at all lead times and reliably detects anomalously dry root-zone states (below the 20th percentile). Yet flash drought onset, defined by multi-pentad intensification criteria, remains a fundamental challenge shared across all current S2S systems. These findings advance data-driven S2S soil-moisture forecasting while highlighting the remaining challenge of predicting rapid drought development.
[LG-130] Behavioral Cloning Mystery
链接: https://arxiv.org/abs/2610.07056
作者: Seohong Park,Sergey Levine
类目: Robotics (cs.RO); Machine Learning (cs.LG)
*备注:
Abstract:Behavioral cloning (BC), despite its simplicity, exhibits many counterintuitive phenomena in the real world. For example, the performance of BC often keeps increasing as the model overfits more to the dataset, and fully closed-loop policies often completely fail without action chunking. Unfortunately, properly studying these anecdotal phenomena (“behavioral cloning mysteries”) is challenging: in the real world, datasets and experiments are costly and not fully controllable; in simulation with synthetic data, these phenomena are often not easily observed partly due to the discrepancy between scripted policies and human demonstrations. In this work, we propose OCBench, a robotic manipulation benchmark with controllable scripted policies that have similar properties to human demonstrations. We show that, by mimicking key properties of human demonstrations, OCBench reproduces many anecdotal BC-related phenomena in controlled settings. With its GPU-accelerated environments and scripted policies, we demonstrate how OCBench enables scientific studies of previously reported BC-related phenomena by analyzing and refuting various hypotheses. Project page: this https URL
[LG-131] he Premise Is the Problem: Exchangeability Failure in Self-Monitored Test-Time Adaptation
链接: https://arxiv.org/abs/2610.07038
作者: Weijia Han,Lisha Qu,Zhenda Li,Liying Liang
类目: Machine Learning (cs.LG); Machine Learning (stat.ML)
*备注: 49 pages, 9 figures
Abstract:Modern forecasting models are often updated after deployment so they can respond to changing data. These updates can also make predictions worse, so practical systems need a reliable monitor that can detect harmful changes and trigger protection. A natural design is to monitor the same prediction errors that guide the updates. This paper asks whether the statistical guarantee behind such a monitor remains valid when monitoring and adaptation use the same feedback. We study this question in multi-step time-series forecasting. We show that overlapping targets and dependence in forecast errors can break a key assumption required by the guarantee. The monitor may then raise alarms even when no harmful change has occurred, and its response can further damage prediction quality. We also find that adaptation can hide sustained changes from its own monitor, while the original frozen model retains a clearer signal. These results expose a basic failure mode in self-monitored adaptation. They show why reliable deployment requires checking the monitor’s assumptions, comparing adaptation with the frozen model under realistic feedback, and limiting the effect of every protective response.
[LG-132] Shaping the Wind: Nested Potentials for Kinematically Admissible Urban Wind Prediction
链接: https://arxiv.org/abs/2610.07033
作者: Yidi Wang,Yunhe Zhang,Jiawei Gu,Ziyue Qiao,Pengyang Wang
类目: Machine Learning (cs.LG)
*备注:
Abstract:Predicting transient urban winds is fundamental to understanding urban microclimates and designing climate-resilient cities. Building-resolving large-eddy simulation produces detailed incompressible urban wind fields at substantial computational cost for each layout. Neural surrogates offer a faster alternative by learning to predict the evolution of velocity fields. However, minimizing velocity prediction error does not guarantee local mass conservation and wall impermeability, which together define kinematic admissibility. This limitation stems from an unconstrained output representation: geometry conditioning guides predictions but does not restrict them to admissible velocity fields. Correcting boundary violations in these outputs changes the flux balance in adjacent fluid cells and may consequently compromise local mass conservation. To address the challenge, we propose Sculpt, a nested potential framework that builds the coupled, geometry-dependent constraints directly into its parameterization. This nested parameterization generates divergence-free velocity updates through the discrete curl of a volume vector potential on the native three-dimensional staggered grid. A shared scalar potential constrains the vector potential’s boundary values so that the same operator also enforces impermeability, without a per-step pressure projection. Because backpropagation through this curl attenuates large-scale gradient signals, we parameterize the volume potential at multiple resolutions to better capture large-scale flow structures. We introduce UrbanWindFlow, an LES dataset spanning urban morphologies and inflow conditions, to evaluate accuracy and kinematic admissibility together.
[LG-133] Identifiable World Models from Pretrained Diffusion Representations
链接: https://arxiv.org/abs/2610.07028
作者: Ruchi Sandilya,Conor Liston,Logan Grosenick
类目: Machine Learning (cs.LG)
*备注:
Abstract:Diffusion-based world models can generate and predict trajectories in high-dimensional dynamical systems, but predictive accuracy does not imply that their latent coordinates recover the underlying state variables or causal interactions. We ask whether a frozen pretrained diffusion model can be equipped with identifiable coordinates without retraining its generative backbone. We show that auxiliary-variable nonlinear ICA guarantees can be transferred to Contrastive Diffusion Alignment (ConDA), which learns only a lightweight alignment map on top of frozen diffusion latents. Under standard TCL/GCL assumptions, the aligned representation identifies latent dynamical states up to permutation and componentwise invertible transformations, preserves the latent dynamic structural causal model, and reduces lagged graph recovery to transition-Jacobian sparsity. We evaluate TCL-, GCL-, and CEBRA-based ConDA against TDRL, CaRiNG, IDOL, temporal SuaVE, and iVAE across physical and robotic video systems. TCL and GCL achieve near-perfect blockwise state recovery and competitive lagged graph recovery, including exact recovery in a simulated falling-body system. In a simulated bipedal robot, learned dynamics recover the sign and temporal structure of responses to held-out control perturbations. These results show that a frozen generative diffusion model can be equipped with coordinates that are identifiable, structurally interpretable, and useful for analyzing intervention-relevant dynamics.
[LG-134] STOCK-JEPA: Prior-Anchored Latent Revision Representation Learning in Equity Markets
链接: https://arxiv.org/abs/2610.07006
作者: Yizhi Luo,Jiahe Yi,Jianhui Zhang,Shuo Sun
类目: Machine Learning (cs.LG)
*备注:
Abstract:Learning effective representations helps characterize the structure and dynamics of equity markets from financial data with a low signal-to-noise ratio. Black-box deep models can capture complex patterns but may overfit sample noise and lack explicit economic structure. Meanwhile, classic linear financial models provide interpretable references, but their oversimplified assumptions leave non-linear signals uncaptured. To combine the strengths of these two directions, we propose Stock-JEPA, a joint-embedding predictive framework that learns predictable incremental revisions relative to a point-in-time financial prior. First, we leverage a low-complexity financial model to produce fixed statistics summarizing multi-horizon return and risk. A prior projector then maps these statistics into the target encoder’s latent space as an anchor. Second, we design a context-conditioned revision predictor to estimate the future representation’s predictable displacement from the anchor. Separate losses update the two branches: the anchor learns from prior statistics, while the revision captures additional predictable information from historical context. Third, we freeze all representation modules and train a downstream readout, evaluating its forecasts through cross-sectional ranking and portfolio performance. Theoretically, we prove that optimal revision reduces the prior anchor’s expected squared error for the same future representation by exactly \mathbbE[|\boldsymbol\Delta|_2^2] . This non-negative gain is the expected squared magnitude of the additional signal predictable from historical context. Experimentally, Stock-JEPA outperforms 13 strong baselines across large-scale China and U.S. equity universes on 5 key evaluation metrics. Ablation studies and representation analysis further demonstrate the value of the learned revisions for representation learning in equity markets.
[LG-135] Repair Lot Skyline: A Weighted Constraint Satisfaction Approach to Pavement Repair Optimization from Geospatial Hazard Density
链接: https://arxiv.org/abs/2610.06989
作者: Takato Yasuno,Keita Kobayashi,Ryuta Sakaguchi,Takuya Okamoto
类目: Machine Learning (cs.LG)
*备注: 30 pages, 7 tables, 5 figures
Abstract:Pavement agencies must translate a spatially distributed distress inventory into a bounded, actionable repair-lot plan: accident-critical defects (potholes) must always be addressed, lower-risk defects (cracks) should be included only when their benefit justifies the repair cost, and historical patch locations signal re-degradation risk without themselves triggering repair. We formalize this as a Repair Lot Skyline problem: a Weighted Constraint Satisfaction Problem (WCSP) defined over chainage (distance along the road) rather than over time, so that it requires only a single-epoch distress survey and makes no claim about future deterioration. The WCSP identifies 143 candidate hazard clusters (61 hard, 82 soft), of which 106 are merged into a final repair plan totaling 1,997.4 m—83.7% of the 2,385.9 m that would be required if every soft candidate were included regardless of cost. This plan covers 100% of observed potholes (138/138) and 91.6% of observed cracks (404/441), capturing 93.6% (542/579) of the total hazard benefit available in the full candidate set. The skyline frontier shows pronounced diminishing returns beyond this point: the remaining 37 excluded soft candidates would add only 6.8% additional benefit for a 19.4% increase in repair length. We further formalize the minimum-lot-length L_\min and historical-context radius \kappa as a joint, four-objective hyperparameter search over this WCSP; on the same case study, the recommended configuration ( L_\min = 14.7 m, \kappa = 10 m) reduces repair-crew mobilizations by 7.1% relative to an untuned default, at the cost of a 12.9% larger budget and a 1.1-percentage-point lower crack coverage.
[LG-136] An Information-Theoretic Evaluation Framework for Benchmark and Model Diagnosis in Knowledge Tracing
链接: https://arxiv.org/abs/2610.06988
作者: Houru Jiang,Zixi Wang,Tengteng Cheng,Xueyi Li,Mingliang Hou,Jiaqi Zheng,Renqiang Luo,Teng Guo,Zitao Liu
类目: Machine Learning (cs.LG)
*备注: 21 pages
Abstract:Knowledge tracing (KT) models are predominantly evaluated using aggregate metrics such as area under the curve (AUC) and accuracy. However, these global scores obscure where the remaining errors originate and fail to indicate whether a benchmark is approaching saturation. While estimating a global theoretical performance limit is challenging in realistic KT settings, it is possible to quantify local predictability. To address this, we propose an information-theoretic evaluation framework for KT benchmark diagnosis. We use Context Tree Weighting (CTW) on item-response histories and current-item queries as an operational causal uncertainty coordinate, while distinguishing it from the unobserved Local Irreducible Uncertainty (LIU) under the full KT information set. By projecting predictions onto this shared uncertainty coordinate, we evaluate model performance gains across distinct entropy bands rather than only at the global level. Comprehensive evaluations on NIPS Task 3/4 and Algebra 2005 reveal that model improvements are highly non-uniform. Modern KT models show substantial gains in high-entropy regions, and additional item-aware references, log-loss, and equal-frequency analyses support this localization. The framework also flags regions where apparent gains require checks for noise-sensitive behavior. By surfacing these local modeling failures alongside genuine gains, this approach provides a diagnostic tool for studying both residual predictive structure and the limitations of current KT benchmarks and models.
[LG-137] A Data-Driven Framework for Unsupervised Monitoring of Transmission Systems Using End-of-Line Testing Data: A Case Study at Ford Motor Company
链接: https://arxiv.org/abs/2610.06980
作者: Mohammad N. Bisheh,Mehrdad Moradi,Parinaz Farajiparvar,Colin Brady,Rajesh Gupta,Xueling Li,Javad Navaei,Milad Parvaneh,Kamran Paynabar
类目: Machine Learning (cs.LG); Applications (stat.AP); Computation (stat.CO)
*备注:
Abstract:Sensing technologies have advanced rapidly across industries ranging from energy to automotive manufacturing. These systems generate high-dimensional (HD) data characterized by complex nonlinear patterns and strong temporal dependencies. Traditional statistical monitoring methods are often limited in their ability to capture such nonlinear structure. Likewise, many analytical approaches used in End-of-Line testing rely on predefined thresholds and heuristic rules, which restrict their ability to detect informative anomaly signatures in HD temporal data. In contrast, while modern deep learning and generative AI models offer strong predictive capabilities, they are often unsuitable in applications where data are costly to collect and where the monitoring system must remain interpretable, low-latency, computationally efficient, and usable by non-technical practitioners. To overcome these limitations, we propose an advanced multivariate monitoring framework for HD data. The framework operates in two stages. In the first stage, the data are preprocessed to remove incomplete and non-informative samples and to temporally align time series data. In the second stage, nonlinear dimensionality reduction is performed, followed by anomaly detection through a control chart based phase I monitoring procedure. The framework can be used in both unsupervised and supervised settings, depending on the availability of ground truth labels during training. Moreover, its flexible and modular structure allows practitioners to adapt its components to different domains and operational requirements. We evaluate the proposed framework on real production data from an automotive manufacturing environment at Ford Motor Company. The proposed method achieves higher accuracy, recall, and F1 score than the company’s existing model, improving these metrics from 0.50, 0.30, and 0.429 to 0.625, 1.00, and 0.769, respectively.
[LG-138] Uncertainty in Representation Learning on Knowledge Graphs
链接: https://arxiv.org/abs/2610.06974
作者: Yuqicheng Zhu
类目: Machine Learning (cs.LG)
*备注: Doctoral Dissertation, 255 pages
Abstract:Knowledge graph embedding (KGE) methods represent entities and predicates in continuous vector spaces to infer missing knowledge. Despite strong benchmark performance, their predictions often lack principled reliability guarantees, limiting their use in high-stakes applications. Moreover, uncertainty arises throughout the KGE pipeline, from incomplete or probabilistic input knowledge to stochastic training and prediction. This thesis systematically investigates three sources of uncertainty in KGE: knowledge uncertainty, arising from incomplete, noisy, or probabilistic input knowledge; algorithmic uncertainty, induced by randomness in model training; and predictive uncertainty, concerning the reliability of model outputs. To address algorithmic uncertainty, the thesis demonstrates that models trained under identical settings can produce substantially different predictions and introduces a voting-based aggregation framework to mitigate this instability. To quantify predictive uncertainty, it adapts conformal prediction to KGE, constructing answer sets with distribution-free coverage guarantees and extending them to provide predicate-conditional reliability guarantees. To support reasoning under knowledge uncertainty, it develops statistically valid prediction intervals for confidence-scored triples and an embedding-based approach to approximate probabilistic reasoning over statistical ontologies with formal soundness guarantees. Together, these complementary, model-agnostic methods provide a practical and theoretically grounded approach to uncertainty in KGE, advancing beyond predictive accuracy toward reliable and uncertainty-aware knowledge graph reasoning.
[LG-139] Do Neural PDE Solvers Learn the Right Dynamics?
链接: https://arxiv.org/abs/2610.06952
作者: Haonan Li,Yue Song,Bin Yang,Kaihong Luo
类目: Machine Learning (cs.LG); Chaotic Dynamics (nlin.CD)
*备注:
Abstract:Neural PDE solvers can achieve low prediction errors, but do they reproduce the dynamics of the systems they model? Prediction scores alone offer an incomplete answer: they measure agreement with reference solutions but provide limited insight into how errors accumulate, nearby states diverge, or extreme events arise. We propose an evaluation framework that directly examines these behaviors in deterministic and stochastic neural solvers. By evolving ensembles of nearby initial states and comparing them with direct numerical simulation, we assess three complementary aspects of learned dynamics: error formation, ensemble geometry, and extreme events. Experiments on two-dimensional Kolmogorov flow reveal limitations that conventional scores can obscure. Smaller trajectory errors can reflect weaker error amplification despite less accurate local updates. Models can match an ensemble’s overall spread and effective dimension while failing to capture the spatial directions where nearby states diverge. Similarly, matching overall event frequencies can conceal failures to predict persistent extreme events. These findings show that improved prediction accuracy does not necessarily imply greater dynamical fidelity. Our framework makes this distinction measurable, providing concrete criteria for evaluating whether advances in neural PDE solvers better capture the underlying dynamics.
[LG-140] AdaLoop: Adaptive-Depth Latent Reasoning for Audio Language Models
链接: https://arxiv.org/abs/2610.06949
作者: Lee Seung-woo,Bowen Qi
类目: ound (cs.SD); Machine Learning (cs.LG)
*备注:
Abstract:Large audio language models answer questions about speech, sound, and music, yet their accuracy drops sharply on tasks that need fine-grained acoustic analysis. Judging which of two speakers has the higher pitch demands iterative signal-level reasoning that a content question does not. Current models spend the same computational depth on both. We introduce AdaLoop, a lightweight recurrent module that learns how many latent refinement steps a given audio–question pair requires. A shared transformer block iterates over the audio representation, guided by the question, while a learned halting mechanism exits the loop once the representation is ready. AdaLoop adds fewer than 3% of the base model’s parameters and plugs into any audio encoder–language model pair without modifying either component. Evaluated on three architecturally distinct models across MMSU, MMAU-Pro, and MMAR, AdaLoop raises the average accuracy by 2.9 to 3.8 points, with the largest gains on perception-heavy subtasks where the model learns to apply deeper reasoning.
[LG-141] Learning to Remember: Distilling Memory Retention for Compact Recurrent Neural Networks
链接: https://arxiv.org/abs/2610.06942
作者: Nilushika Udayangania,Kishor Nandakishora,Marimuthu Palaniswami
类目: Machine Learning (cs.LG)
*备注: Preprint
Abstract:Deep learning models, particularly recurrent neural networks and their variants, such as long short-term memory, have significantly advanced time series analysis. These models capture complex, sequential patterns in time series, enabling real-time assessments. However, their high computational complexity and large model sizes pose challenges for deployment in resource-constrained environments, such as wearable devices and edge computing platforms. Knowledge Distillation (KD) offers a solution by transferring knowledge from a large, complex model (teacher) to a smaller, more efficient model (student), thereby retaining high performance while reducing computational demands. Current KD methods, originally designed for computer vision tasks, neglect the unique temporal dependencies and memory retention characteristics of time series models. To bridge this gap, we propose a novel KD framework termed Memory-Discrepancy Knowledge Distillation (MemKD). MemKD leverages a specialized loss function to capture memory retention discrepancies between the teacher and student models across subsequences within time series data, ensuring that the student model effectively mimics the teacher’s behaviour. This approach facilitates the development of compact, high-performing recurrent neural networks suitable for real-time, time series analysis tasks. We provide additional experiments, in-depth theoretical analysis, and insights into the proposed framework across extended time series benchmarks. Our experiments demonstrate that MemKD significantly outperforms state-of-the-art KD methods. Additionally, we demonstrate that it can match the teacher model’s performance across a wide range of compression levels, achieving notable reductions in parameter count and memory usage without a significant loss in accuracy.
[LG-142] QiYao-I: A Manifold Based Foundation Model for Irregular Multivariate Time Series Forecasting
链接: https://arxiv.org/abs/2610.06936
作者: Linfeng Wang,Ruitong Zhang,Kai Zhao,Yang Shu,Zhongwen Rao,Meng Wang,Yijie Li,Bin Yang,Chenjun Guo
类目: Machine Learning (cs.LG)
*备注: 29 pages, 5 figures, 20 tables. Preprint
Abstract:Irregular multivariate time series forecasting is a challenging yet important problem in real-world applications, where observations are often irregularly sampled and asynchronously recorded across variables. Existing time series foundation models are mostly built on regularly sampled sequences, making them difficult to generalize to irregular time intervals and asynchronous cross-variable dependencies. To address these challenges, we propose QiYao-I, a manifold based foundation model for irregular multivariate time series forecasting. Specifically, we introduce a novel sampling-conditioned temporal manifold attention mechanism that maps real timestamps into a learnable temporal manifold feature space and injects temporal manifold biases into attention layers, enabling the model to capture both irregular time intervals and local sampling structures. Further, we propose a dynamic variable interaction mechanism with frequency awareness. It selectively performs cross-variable message passing under asynchronous observations. Extensive experiments on real-world irregular multivariate forecasting benchmarks demonstrate that QiYao-I achieves superior performance compared with both time series foundation models and end-to-end irregular forecasting models, showing strong generalization ability in zero-shot and few-shot settings.
[LG-143] Near-Optimal Sample Complexity for Recursive Entropic Risk Reinforcement Learning with a Generative Model
链接: https://arxiv.org/abs/2610.06931
作者: Amirparsa Bahrami,Oliver Mortensen,Mohammad Sadegh Talebi
类目: Machine Learning (cs.LG); Machine Learning (stat.ML)
*备注:
Abstract:In this paper, we study the sample complexities of value and policy learning in finite discounted Markov decision processes (MDPs) under recursive entropic risk preferences with risk parameter (\beta\neq 0), assuming access to a generative model of the MDP. We provide a refined analysis of model-based risk-sensitive Q-value iteration (MB-RS-QVI), a plug-in model-based method introduced in prior work, and derive ((\varepsilon,\delta))-PAC guarantees for both learning the optimal (Q)-value function and an (\varepsilon)-optimal policy. Our bounds improve the exponential dependence on the effective horizon (1/(1-\gamma)) compared with the best existing guarantees for this setting. In particular, they match the existing lower bounds in their exponential dependence on (|\beta|/(1-\gamma)), as well as in (S), (A), (\varepsilon), and (|\beta|), up to logarithmic factors. Consequently, our analysis removes the exponential gap between the previously known upper and lower bounds, leaving only a polynomial gap in the effective horizon.
[LG-144] Extending Music Annotation Schemas: Zero-Shot Prediction or Few-Shot Adaptation? ICASSP2027
链接: https://arxiv.org/abs/2610.06920
作者: Christos Plachouras,Emmanouil Benetos,Johan Pauwels
类目: ound (cs.SD); Machine Learning (cs.LG)
*备注: 5 pages, 3 figures. Submitted to IEEE ICASSP 2027; under review
Abstract:Automatic music annotation is typically tackled under the assumption of a fixed annotation schema. In practice, commercial music catalogs often need to accommodate new musical attributes as needs evolve. Given that expert music annotation is expensive, it is not evident which methodological approach is most effective at accommodating new attributes and backfilling existing tracks; audio-language models promise zero-shot prediction, but at what annotation budget does supervised adaptation become more compelling? We propose a benchmark based on the MGPHot popular music annotation dataset for simulating music schema extension across different annotation budgets. We investigate zero-shot prediction with audio-language models, learning new attributes from pretrained representations, and adapting models trained on existing annotations. Our results suggest that supervised adaptation is more effective than zero-shot prediction even with small annotation budgets, while frozen representation reuse remains the most effective approach for modest budgets without the tuning required by deeper adaptation. Comments: 5 pages, 3 figures. Submitted to IEEE ICASSP 2027; under review Subjects: Sound (cs.SD); Machine Learning (cs.LG) Cite as: arXiv:2610.06920 [cs.SD] (or arXiv:2610.06920v1 [cs.SD] for this version) https://doi.org/10.48550/arXiv.2610.06920 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[LG-145] Learning from Unreliable Trajectories: Adversarially-Robust Federated Q-Learning
链接: https://arxiv.org/abs/2610.06918
作者: Sreejeet Maity,Aritra Mitra
类目: Machine Learning (cs.LG); Systems and Control (eess.SY)
*备注:
Abstract:We study federated reinforcement learning in which multiple agents interact with a common Markov decision process and communicate through a central server to collaboratively learn the optimal state-action value function. Our goal is to understand whether the sample-efficiency benefits of collaboration can be retained when a fraction of the agents behave adversarially and transmit arbitrarily corrupted information. To address this problem, we introduce Robust Async-Fed-Q, an epoch-based federated learning algorithm that combines variance-reduced estimation of the Bellman optimality operator at the agents with robust aggregation at the server. We establish high-probability finite-time guarantees showing that the proposed method preserves the statistical gains of collaboration among the honest agents while tolerating adversarial corruption. In particular, the effect of the adversarial agents decreases as the amount of data collected by each honest agent grows and eventually vanishes in the infinite-sample limit. We complement these guarantees with information-theoretic lower bounds that characterize the unavoidable statistical cost of adversarial corruption, leading to the first nearly matching upper and lower bounds for adversarially robust federated reinforcement learning. We further extend our framework to accommodate single-trajectory Markovian sampling and heterogeneous partial coverage, where different agents may explore different regions of the state-action space and learning relies on their collective coverage. Finally, our epoch-based design substantially improves the best known communication complexity for federated Q-learning under asynchronous sampling.
[LG-146] Event-Driven ML Pipeline Orchestration for Manufacturing: An AWS Industry Experience
链接: https://arxiv.org/abs/2610.06890
作者: Zhengyang(Cissy)Gu,Thomas Cook,Fredaljohn Rohrbaugh,Joseph E. Hernandez,Chris Couch
类目: Machine Learning (cs.LG)
*备注: Accepted in the 14th IEEE International Conference on Cloud Engineering (IC2E 2026). It will be hosted on October 13th-15th, 2026 at Santa Clara, California, USA
Abstract:We present an industry experience report on three years of operating an event-driven cloud infrastructure for continuous machine learning training in automotive manufacturing. Our system orchestrates GPU-accelerated training of product-specialized model pairs, a physics prediction model and a reinforcement-learning control policy, across multiple plants, coordinating long-running GPU workloads triggered by manufacturing events. The architecture combines Amazon ECS with EC2 GPU capacity providers, SQS-based messaging with dead-letter queues, and an admission-controlled Lambda dispatcher that enforces cluster concurrency limits. A Conductor orchestrator on ECS Fargate initiates dependency-aware retraining chains on a weekly schedule. The entire infrastructure is codified in modular Terraform with multi-account separation. From 40000+ production training jobs we report a 72-78% cost reduction versus always-on GPU infrastructure. A discrete-event simulation confirms that admission control is necessary (naive dispatch loses 65% of jobs) and that queue-draining matches AWS Step Functions latency while eliminating per-job startup overhead. We provide lessons learned and release the simulator and Terraform module skeletons as open-source artifacts.
[LG-147] Beyond Marginals: A Multi-Dimensional Evaluation Framework for Multi-Table Synthetic Data Generation
链接: https://arxiv.org/abs/2610.06854
作者: Aparana Gupta,Anurup Dey,Suyash Dwivedi
类目: Databases (cs.DB); Machine Learning (cs.LG)
*备注:
Abstract:Synthetic data generation is critical for privacy compliance, machine learning augmentation, and software testing. While single-table evaluation is well established, multi-table (relational) synthesis, the dominant enterprise use case, lacks a unified evaluation framework. Existing approaches assess marginal column distributions in isolation, overlooking joint distributions, cross-table structural integrity, downstream utility, and production-readiness edge cases. We present SynEval, a six-dimensional evaluation framework for multi-table synthetic databases. SynEval jointly assesses per-column fidelity, multivariate structure preservation including a novel conditional distribution check, cross-table integrity, ML utility, privacy protection, and edge-case robustness. The framework produces a unified weighted quality score with per-table and per-dimension drill-down, and is generator-agnostic, operating on any pair of real and synthetic CSV folders with automatic schema inference. SynEval is a framework to combine conditional distribution checks P(Y|X), cross-table cardinality validation, and production-readiness edge-case testing within a single evaluation pipeline for relational synthetic data.
[LG-148] Asymptotically Optimal Best Arm Identification with Fixed-Budget under Differential Privacy NEURIPS2026
链接: https://arxiv.org/abs/2610.04600
作者: Keqin Chen,Jie Bian,Yulian Wu,Vincent Y. F. Tan
类目: Machine Learning (cs.LG); Information Theory (cs.IT)
*备注: Accepted to NeurIPS 2026
Abstract:Best arm identification under differential privacy is a pure-exploration problem in which both statistical efficiency and privacy protection must be achieved simultaneously. We study fixed-budget best arm identification for bandits under pure \epsilon -differential privacy, where the learner must recommend an arm after a prescribed sampling budget while protecting the full transcript. We prove that the optimal exponential decay rate of the error probability is upper bounded by an instance-dependent privacy-aware transportation exponent that differs from the analogous quantity used to characterize the stopping time in fixed-confidence analysis by Jourdan and Azize [2025]. Guided by this exponent, we propose AO-Pri-BAI, an adaptive algorithm that maintains private running estimates through Laplace-tree mechanisms and learns a sampling design through a min–max interaction between hard alternatives and arm allocations. We prove that AO-Pri-BAI satisfies pure \epsilon -differential privacy. We also establish that the exponent of the failure probability of AO-Pri-BAI matches the privacy-aware benchmark. Numerical studies show that even in the non-asymptotic setting, AO-Pri-BAI outperforms benchmark algorithms on various instances, complementing the theoretical analyses.
[LG-149] Prediction-powered inference for time series across space NEURIPS2026
链接: https://arxiv.org/abs/2610.08715
作者: Shahzar Rizvi,David Burt,Vishwak Srinivasan,Renato Berlinghieri,Stefano Del Col,Tamara Broderick
类目: Methodology (stat.ME); Machine Learning (cs.LG); Machine Learning (stat.ML)
*备注: Accepted to TS-LIMITS Workshop at NeurIPS 2026
Abstract:The following motif is common in spatiotemporal settings: we have a sequence of covariate and label pairs observed for a relatively short, recent time period. We have access to unlabeled covariates over a longer time period. Data is observed over many spatial locations. For instance, crop yield might be observed over a large geographical area for recent years, but weather data (which is informative about crop yield) is available for a much longer period. The goal is to estimate, at each spatial location, the expected label (e.g., crop yield) in the future and provide a valid confidence interval for this value. The observed time period alone is too short for reliable estimates. Imputing missing labels with machine learning can cause substantial bias. Prediction-powered inference (PPI) can correct for this bias, but it relies on an i.i.d. assumption that breaks under our expected temporal dependencies. Heteroskedasticity and autocorrelation consistent (HAC) procedures account for temporal correlation, but have not been adapted to cases where some labels are imputed. We provide reliable point estimates and confidence intervals given: short labeled time series (across spatial locations), a longer unlabeled time series, and an imperfect predictor of labels given covariates. We show our method outperforms natural alternatives.
[LG-150] Steering Diffusion Models to Rare Events with Sequential Monte Carlo NEURIPS2026 STOC
链接: https://arxiv.org/abs/2610.08652
作者: Aavash Subedi,Tim Reichelt,Christopher Williams,Philip Stier,Yee Whye Teh,Saifuddin Syed
类目: Machine Learning (stat.ML); Machine Learning (cs.LG)
*备注: A previous version of this work was presented at the NeurIPS 2026: AI for Stochastic Dynamics workshop
Abstract:Diffusion models are increasingly used as surrogates for expensive simulators in weather prediction, molecular dynamics, and materials design. In these models, computing the probability p_0[E] of an event E is difficult, especially when the event of interest is rare. A stable estimate using Monte Carlo becomes computationally intractable, requiring a growing sample size \propto!1/p_0[E] to compensate for an increasing rarity. In this paper, we present Diffusion Importance Sampling of Rare Events or DireSMC, a sequential Monte Carlo scheme that guides a population of weighted samples towards the rare event, giving access not only to samples but also to a calibrated estimate of its probability. We set up our guidance using an analytical relaxation of the event set, allowing the method to easily extend to a wide range of user-defined rare events. We validate our method on a toy problem with analytical solutions and on a score-based climate emulator, where we obtain accurate rare-event probabilities on a range of rarities from 10^-3 to 10^-5 , achieving net speed-ups of 9\times to 1413\times over Monte Carlo.
[LG-151] Information-Dense Synthesis for Molecular Discovery
链接: https://arxiv.org/abs/2610.08495
作者: Kasper K. Jakobsen,Eli N. Weinstein
类目: Machine Learning (stat.ML); Machine Learning (cs.LG); Chemical Physics (physics.chem-ph); Biomolecules (q-bio.BM)
*备注:
Abstract:Machine learning can accelerate molecular discovery by designing molecules and planning experiments. However, many scientific challenges demand molecules with very rare properties, and in this sparse setting, existing algorithms offer little gain over random guessing. We propose a method to efficiently search large regions of molecular space using algorithmically controlled stochastic synthesis. Rather than design, make and test individual molecules, we design and make complex mixtures, test them as a pool, then deconvolute the molecule-activity map. We optimize synthesis to encode maximal information. Theoretically, this approach can reduce the number of experiments required to find the optimal molecule among d candidates from \mathcalO(d) to \mathcalO(\log d) or \mathcalO(1) . In simulation, on estimated protein fitness landscapes, it finds active molecules with an order of magnitude fewer experiments than existing Bayesian optimization methods.
[LG-152] High-Dimensional Statistical Inference for Sparse Support Vector Machines
链接: https://arxiv.org/abs/2610.08345
作者: Peng Zeng,Hanwen Huang
类目: Machine Learning (stat.ML); Machine Learning (cs.LG); Methodology (stat.ME)
*备注: 7 figures
Abstract:Using a replica-symmetric high-dimensional characterization, we develop an inferential framework for sparse support vector machines when the sample size and number of features grow proportionally. The main challenge is the nonsmooth hinge loss, which prevents direct application of debiasing arguments developed for smooth classification losses. We overcome this difficulty by representing the L_1 -penalized support vector machine (SVM) as a linear program and identifying the hinge-loss subgradient through its dual variables. This yields a computationally accessible debiased estimator whose coordinates are asymptotically Gaussian under the proportional asymptotic regime. The resulting distributional characterization provides confidence intervals and hypothesis tests for individual features and enables false-discovery-rate-controlled variable selection. Extensive simulations examine calibration, power, and variable-selection performance under a range of covariance structures, including strongly correlated designs. An analysis of high-dimensional breast cancer gene-expression data illustrates how the proposed inference can distinguish statistically significant features from variables selected by the original sparse SVM.
[LG-153] Where Do Two Populations of Persistence Diagrams Differ? Calibrated Local Inference at a Fixed Budget
链接: https://arxiv.org/abs/2610.08292
作者: Pramita Bagchi,Edward Bae,Atish Mitra,Alexander D. Silberman,Žiga Virk,Sushovan Majhi
类目: Methodology (stat.ME); Machine Learning (cs.LG); Statistics Theory (math.ST)
*备注: 45 pages, 9 figures. Appendices with proofs and additional experiments. Under review
Abstract:Many two-sample tests for populations of persistence diagrams assess global differences without identifying the regions of the birth-death plane that contribute to them. We study simultaneous inference for local mean contrasts when the number of available diagrams is fixed. They are differences in expected weighted feature mass within \ell_\infty neighborhoods at several centers and radii. We estimate these contrasts using additive landmark responses. A Gaussian multiplier bootstrap calibrates simultaneous confidence intervals while allowing unequal group covariances. The neighborhoods whose intervals exclude zero form a map with approximate family-wise error control, and selecting a subset of original intervals for display preserves their joint coverage guarantee. On the simultaneous coverage event, every reported neighborhood lies within twice its radius of the support of the mean-measure difference. A geometric result gives sufficient radius conditions for a displaced feature to produce a nonzero contrast. A comparison of sufficient detection thresholds quantifies the tradeoff between reducing the number of tested coordinates and reserving observations for an independent pilot. In simulations with 40 to 120 diagrams per class, the bands achieved 94%-98% simultaneous coverage under both the strict null and equal means with unequal covariances. In the latter setting, a permutation maximum and the pooled-t implementation of the two-stage persistence-image test of Moon and Lazar rejected in up to 32% and 26% of runs, respectively. In the fixed-budget simulations, spending a third of the observations on a pilot to choose landmarks or radii located changes less often than a prespecified grid at a single radius. On the MUTAG benchmark, the localized region concentrates on rings of fused-ring systems, an exploratory reading.
[LG-154] wo-Sample Testing via Generative Processes
链接: https://arxiv.org/abs/2610.08277
作者: Eshant English,Kenji Fukumizu,Taiji Suzuki
类目: Machine Learning (stat.ML); Machine Learning (cs.LG)
*备注:
Abstract:Deciding whether two samples come from the same distribution is a classical problem in statistics, and generative transport offers a new way to approach it. We build a stochastic interpolant directly between the two samples and observe that, for a symmetric schedule, its law is invariant under the time reflection t \mapsto 1-t whenever the two distributions coincide. We therefore test whether the marginals at times t and 1-t agree by computing their Jensen–Shannon divergence. Both marginals are explicit mixtures over all cross-pairs of observations, so nothing is learned, and permutation calibration gives an exact finite-sample level. For Gaussian noise, this divergence equals a time integral that pairs the reflection defects of the velocity field and of the score, so the test compares transport dynamics rather than endpoints alone. With a narrow-plus-broad noise design, the test attains the minimax separation rate n^-2s/(4s+d) over bounded, compactly supported densities whose difference has Sobolev smoothness s 3d/4, with no lower bound on the densities. Fusing a dyadic grid of noise scales through their permutation ranks, without sample splitting, preserves exact level and adapts to unknown s at an iterated-logarithmic cost. Empirically, the test matches or outperforms state-of-the-art kernel two-sample tests.
[LG-155] Anytime-valid simulation-based hypothesis testing
链接: https://arxiv.org/abs/2610.08210
作者: Patrick Forré,Lydia Brenner
类目: Machine Learning (stat.ML); Machine Learning (cs.LG); Statistics Theory (math.ST); Methodology (stat.ME)
*备注:
Abstract:For a given data distribution (X_t)_t \in \mathbbN \sim Q i.i.d., we investigate the hypothesis testing problem: H_0: Q = P_0 vs. H_1: Q = P_1 , for two different model probability distributions P_0 and P_1 . In contrast to the standard setting, where analytic densities p_0 and p_1 are given, here, we consider the density-free setting, where we only have access to i.i.d. simulations (Z^0_t)_t \in \mathbbN \sim P_0 and (Z^1_t)_t \in \mathbbN \sim P_1 . For this simulation-based hypothesis testing setting, we construct an e-test martingale, resulting in a sequential test with anytime-valid type-I error guarantees, approximate growth optimality, geometrically decaying type-II error bounds, and asymptotic power one. Most ingredients used in our constructions are variants of well known concepts. The value of this paper lies in the compact presentation of an effective, anytime-valid solution for the density-free simulation-based sequential hypothesis testing case.
[LG-156] ProximalFM: Amortized Proximal Causal Inference under Hidden Confounding
链接: https://arxiv.org/abs/2610.08078
作者: Christophe Muller,Ayub Kharel,Alex Luedtke,Chan Park,Eric Tchetgen Tchetgen,Juan L. Gamella,Rahul Krishnan,Ricardo Silva,Jakob Zeitler
类目: Machine Learning (stat.ML); Machine Learning (cs.LG)
*备注:
Abstract:Standard causal identification methods often assume no unmeasured confounding and can fail when relevant confounders are unobserved. Proximal causal inference instead uses proxy variables to identify effects under hidden confounding. However, nonparametric proximal estimation can be challenging in practice: recovering causal estimands such as the conditional average treatment effect (CATE) requires solving an ill-posed integral equation that is data-hungry, hyperparameter-sensitive, and optimization-unstable. Bayesian inference for such models provides a desirable alternative, mitigating these difficulties by regularizing through the prior. However, computing a posterior is itself challenging, as a typical likelihood function will include latent variables. Following the recent success of tabular foundation models in backdoor, instrumental variable, and frontdoor settings, we propose that prior-data fitted networks (PFNs) are uniquely suited to resolve this bottleneck. Indeed, by training on synthetic data sampled from compliant structural causal models with access to oracle counterfactuals, we simplify the task substantially, amortizing the implied Bayesian operator inversion into a single transformer forward pass. Compared to prior literature that focuses primarily on point estimation, our model, ProximalFM, explicitly targets the Bayesian posterior distribution of the CATE. One unique aspect of this problem is that we need to provide Monte Carlo estimates of the oracle CATEs, leading to a novel variation of PFNs that accounts for the added stochastic error. Across a diverse suite of proximal regimes, ProximalFM achieves consistently strong CATE-estimation performance without dataset-specific tuning, with its largest advantage when latent confounding is substantial and the proxies are weakly informative; it also provides fast inference through a single amortized forward pass.
[LG-157] Learning consistent molecular mechanics force fields from first principles NEURIPS2026
链接: https://arxiv.org/abs/2610.08020
作者: Berkay Günes,Leif Seute,Jigyasa Nigam,Frauke Gräter
类目: Chemical Physics (physics.chem-ph); Machine Learning (cs.LG); Computational Physics (physics.comp-ph)
*备注: Accepted to the ML4Molecules Workshop at NeurIPS 2026
Abstract:Classical force fields (FFs) remain the workhorse for large-scale simulations even as machine-learned interatomic potentials (MLIPs) approach ab initio accuracy. They decompose total configuration energies into simple effective interactions whose parameters are traditionally assigned based on atom or bond types, enabling efficient simulations but also limiting their ability to adapt across configurations. Recent machine learning approaches have improved the accuracy and transferability of bonded parameters in these FFs by inferring them as functions of local atomic environments, but still rely on empirical nonbonded parameters for practical simulations. In this work, we introduce a unified approach, \textttgrappa-fullFF, which learns both bonded and nonbonded parameters \emphconsistently and simultaneously from ab initio reference data. By incorporating physically inspired regularization via supervision of the electrostatic potential and an architecture that facilitates charge equilibration, our model recovers accurate electric response properties, achieves state-of-the-art accuracy on geometry optimization benchmarks, and reproduces the conformational sampling of both classical and existing machine-learned FFs, without relying on externally assigned nonbonded parameters.
[LG-158] Stochastic Gradient Descent Ascent is Suboptimal for Nonconvex-PL Min-Max Games
链接: https://arxiv.org/abs/2610.07814
作者: Junsoo Ha
类目: Machine Learning (stat.ML); Computer Science and Game Theory (cs.GT); Machine Learning (cs.LG)
*备注:
Abstract:How far can stochastic gradient descent ascent (SGDA) go by tuning its timescale ratio and step sizes in nonconvex min-max games? We answer this question for nonconvex-PL (NC-PL) games by establishing the first tight complexity of two-timescale SGDA with a fixed timescale ratio and non-increasing step sizes. For \ell -smooth games with an inner \mu -PL inequality, we prove a complexity lower bound \Omega(\kappa^2\ell\varepsilon^-2+\kappa^4\ell\sigma^2\varepsilon^-4) , where \kappa=\ell/\mu is the condition number, \sigma^2 is the gradient variance, and \varepsilon measures the outer gradient norm. This matches existing SGDA upper bounds and establishes a complexity separation from Smoothed-AGDA (Yang et al., 22’). In addition, we show that SGDA can fail to find a stationary point when its timescale ratio is as small as o(\kappa^2) . Our negative results highlight the fundamental limitation of SGDA in NC-PL games, and justify the development of alternative methods.
[LG-159] rustworthy Method Comparison with AI Judges: Estimation and Design under Order Batch and Aggregation Effects
链接: https://arxiv.org/abs/2610.07755
作者: Tianxi Li,Jie Ding
类目: Machine Learning (stat.ML); Machine Learning (cs.LG); Applications (stat.AP); Methodology (stat.ME)
*备注:
Abstract:Large language models (LLMs) are increasingly used as judges for automated AI evaluation. A common practice is to randomize prompt sequences and average the resulting scores, but its statistical validity remains unclear. We show that LLM evaluation mechanisms can be approximated by a class of Markov generalized linear mixed models (GLMMs), supported by out-of-sample predictions across three major commercial LLMs. Using a first-order Markov GLMM, we study leaderboard ranking and group comparison. For leaderboard ranking, randomize-and-average selection is consistent under a mild separation condition, and a Williams square design can improve efficiency when item qualities are close. For group comparison, naive averaging can yield inconsistent conclusions about differences in group-level quality because of the response model’s nonlinearity. Empirical results further support the validity of the proposed model-based inference beyond the first-order theory, including settings with higher-order sequence memory. We illustrate the approach in an application where AI judges compare two graphical model estimation methods.
[LG-160] High-dimensional online calibration from harmonic weights
链接: https://arxiv.org/abs/2610.07740
作者: Maxwell Fishelson,Mehryar Mohri
类目: Machine Learning (stat.ML); Machine Learning (cs.LG)
*备注:
Abstract:We study the online calibration of multidimensional forecasts over an arbitrary convex set Y\subseteq\mathbbR^d relative to an arbitrary error norm |\cdot|_L . For forecasting d binary outcomes simultaneously ( Y=[0,1]^d ), we give the first algorithm that achieves \varepsilon -calibration in a number of rounds that is polynomial in d for every fixed accuracy. It requires d^O(1/\varepsilon) rounds, exponentially improving the dimension dependence of previous bounds. For multi-class forecasting ( Y=\Delta_d ), we obtain the same d^O(1/\varepsilon) rate, improving the d^\widetildeO(1/\varepsilon^2) bounds of Peng and Fishelson et al. Our algorithm is simple: on each round, it outputs a harmonically weighted distribution over harmonically smoothed past outcomes. The same algorithm works for every forecast set and norm. More generally, it achieves \varepsilon -calibration after \exp(O(\gamma(Y,L)/\varepsilon)) rounds, where \gamma(Y,L) is a geometric parameter defined by a matrix discrepancy problem. The harmonic weights are motivated by the fact that the discrete Hilbert transform matrix achieves the optimal discrepancy up to a universal constant, simultaneously for every L . This optimality result may be of independent interest. Subjects: Machine Learning (stat.ML); Machine Learning (cs.LG) Cite as: arXiv:2610.07740 [stat.ML] (or arXiv:2610.07740v1 [stat.ML] for this version) https://doi.org/10.48550/arXiv.2610.07740 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[LG-161] Nash Social Welfare for Multi Armed Bandits: Trajectory-wise Expected and High Probability Regret NEURIPS2026
链接: https://arxiv.org/abs/2610.07737
作者: Avishek Ghosh
类目: Machine Learning (stat.ML); Machine Learning (cs.LG)
*备注: Accepted at NeurIPS 2026
Abstract:We study fair multi-armed bandits under the Nash Social Welfare (NSW) objective, which measures performance via the geometric mean of accumulated rewards. Existing work defines Nash regret as \mathrmNR_T = \mu^\star - (\prod_t=1^T \mathbbE\mu_I_t)^1/T , where \mu_I_t is the mean reward of the recommended arm I_t and T is the horizon. Since it applies the geometric mean to per-round marginal expectations, it ignores the joint distribution of rewards across rounds, leaving the NSW fairness motivation unaddressed at the trajectory level. We propose \emphtrajectory-wise Nash regret \widetilde\mathrmNR_T = \mu^\star - \mathbbE[(\prod_t=1^T \mu_I_t)^1/T] , which computes the geometric mean over complete sample paths before taking expectations, capturing NSW fairness more faithfully. By Jensen’s inequality, \widetilde\mathrmNR_T \geq \mathrmNR_T , making it a strictly stronger metric. We also introduce \emphhigh probability Nash regret \widehat\mathrmNR_T = \mu^\star - (\prod_t \mu_I_t)^1/T , giving the first high probability regret bounds in fair bandits. Our two-phase algorithm, Round Robin Nash Confidence Bound (\textttRR-NCB), combines round robin exploration with a Nash confidence bound index policy. We show \widetilde\mathrmNR_T \leq \widetilde\mathcalO(\sqrtk\log T/T) and, with probability 1-\delta , \widehat\mathrmNR_T \leq \widetilde\mathcalO(\sqrtk\log(kT/\delta)/T) , matching the optimal \widetilde\mathcalO(\sqrtk/T) rate despite the stronger metrics. Optimality follows from a lower bound via AM-GM and standard k -armed bandit minimax arguments. Simulations validate our theory.
[LG-162] Exact Calibration and Sharp Risk Geometry for Volume-Sampled Ridge Regression
链接: https://arxiv.org/abs/2610.07721
作者: Kihun Rhee
类目: atistics Theory (math.ST); Machine Learning (cs.LG); Machine Learning (stat.ML)
*备注: 66 pages, 0 figures
Abstract:We study ridge regression from exactly s distinct rows of a fixed design. Responses are fixed, and only the subset is random. The determinant law and selected ridge fit share one positive definite penalty. Established mean identities and exponential-family duality give the unique penalty that matches a prescribed full-data ridge fit in expectation. It exists exactly when s exceeds the target’s effective dimension. Our main result concerns centered covariance risk normalized by full-data penalized loss. For balanced signed coordinate replicas, a strict sector inequality gives the sharp risk and all maximizing responses at every budget from the dimension to one below the row count. This holds for any nonzero positive semidefinite query. With the target and query fixed, the maximizing response space is unchanged across these budgets. For general designs, we characterize attainment of a leave-one-out envelope. For existing real equiangular tight frames, flat row query energy characterizes when every nonzero residual response maximizes at two deletions. At three deletions, we give the sharp risk and complete maximizing space for isotropic queries, using unequal triangle weights. The balanced geometry yields a same-sample unbiased ridge–Horvitz–Thompson mixture with lower sharp risk and an exact mean-share improvement boundary. Under full recalibration after feature changes, we prove quadratic regret from searching the complete old maximizing space and a query-uniform bound on the mixture’s risk gain. The strongest sector inequalities have exact computer-assisted proofs.
[LG-163] Stability of Measure-to-Measure Transformers on Sub-Gaussian Data
链接: https://arxiv.org/abs/2610.07717
作者: Frank Cole,Nicholas H. Nelsen,Takashi Furuya
类目: Machine Learning (stat.ML); Machine Learning (cs.LG)
*备注:
Abstract:Transformers have exhibited impressive empirical success across various domains, but their theoretical foundations remain less developed. This work constitutes a mathematical study of the measure-to-measure operators defined by transformers. We show that transformers map sub-Gaussian inputs to sub-Gaussian outputs; this ensures that taking arbitrary-length compositions of the softmax operator is well-defined. We then show that transformers are Hölder continuous with respect to the 1-Wasserstein distance on appropriate spaces of sub-Gaussian inputs. This allows us to establish estimates on the error propagation along a transformer between a sub-Gaussian input and its empirical approximation. We also study a mean-field analog of the cross-attention mechanism, which is an operator from a pair of probability measures to a single probability measure. We show that cross-attention exhibits different Hölder regularity and sample-complexity in its two input arguments. Last, we apply our results to deduce approximation guarantees for measure-to-measure transformers. Together, these results provide a firm stability and finite-sample theory for transformers on sub-Gaussian data.
[LG-164] Mathematical Invariant-Enabled Topological Neural Networks for Molecular and Materials Property Prediction
链接: https://arxiv.org/abs/2610.07712
作者: Yiming Ren,Xiang Liu,Mustafa Hajij,Pietro Liò,Guo-Wei Wei
类目: Biomolecules (q-bio.BM); Machine Learning (cs.LG)
*备注:
Abstract:Existing molecular and materials learning approaches often rely on a limited set of structural representations, which may capture only selected aspects of complex three-dimensional structure. Here, we introduce mathematical invariant-enabled topological neural networks (MITNNs), a framework that represents complex structures through multiple complementary mathematical views and integrates them with topological neural architectures. MITNNs combine multiscale invariants from topology, spectral theory, commutative algebra, differential geometry, and discrete curvature, capturing complementary structural information from the same system. Systematic invariant-subset, architecture-subset, and ensemble analyses show that predictive performance depends on how mathematical representations and neural architectures are paired, with selected combinations outperforming individual models and the aggregation of all available components. Across protein-ligand binding, metal-organic framework properties, mutation-induced protein solubility, and molecular toxicity prediction, MITNN consistently outperforms existing methods. These results establish MITNN as a mathematically multimodal framework for scientific machine learning.
[LG-165] Uniform Discrete Diffusion Models are Minimax Optimal for Estimating Distributions with Small Effective Support Size
链接: https://arxiv.org/abs/2610.07655
作者: Dongsun Yoon,Saptarshi Chakraborty
类目: Machine Learning (stat.ML); Machine Learning (cs.LG)
*备注:
Abstract:Discrete diffusion models have emerged as a practically successful framework for generative modeling on discrete product spaces, yet their statistical generalization properties remain poorly understood. Discrete real-world data such as text or biological sequences often concentrate on a small fraction of the astronomically large ambient space because of semantic or physical constraints, but existing bounds fail to capture this distributional structure and instead scale with the size of the ambient space, giving rise to almost vacuous error bounds. We address this gap for uniform discrete diffusion, one of the two dominant discrete diffusion paradigms alongside masking diffusion, by deriving statistical guarantees governed by the effective support size s_n(P_0) , a sample-size-dependent measure of distributional complexity. Given n independent and identically distributed (i.i.d.) samples from an unknown data distribution P_0 on [K]^d , we show that, with appropriate choices of network size and hyperparameters, the expected total variation (TV) loss scales as O(\sqrts_n(P_0)/n) , while the expected Kullback–Leibler (KL) divergence is bounded by O(\frac1ns_n(P_0)\log(eK^d/s_n(P_0))\log n) . Furthermore, we show that the TV rate is minimax optimal and that the KL rate is minimax optimal up to a factor of \log n . Together, these upper and lower bounds show that uniform discrete diffusion successfully avoids the curse of dimensionality for distributions with small effective support size: the TV error rate depends on the ambient state-space size only through s_n(P_0) , while the corresponding KL rate incurs only an additional logarithmic dependence on the ambient state-space size.
[LG-166] Asymptotic Analysis of Empirical Risk Minimization on Entry-wise i.i.d. Heavy-Tailed Data
链接: https://arxiv.org/abs/2610.07637
作者: Kaito Takanami,Takashi Takahashi,Yoshiyuki Kabashima
类目: Disordered Systems and Neural Networks (cond-mat.dis-nn); Machine Learning (cs.LG); Statistics Theory (math.ST); Machine Learning (stat.ML)
*备注:
Abstract:Many real-world datasets exhibit unusually large values far more frequently than predicted by Gaussian models. Heavy-tailed distributions capture this behavior, yet evaluating learning performance under them remains challenging because rare, large feature entries retain non-vanishing effects even in high dimensions. Even in the canonical setting of empirical risk minimization for linear regression with entry-wise i.i.d. symmetric \alpha -stable data, a precise asymptotic characterization of prediction has been lacking. In this work, we introduce a functional order parameter that describes the random effective problem associated with each coefficient. Using the replica method, we fully characterize the generalization error in the proportional high-dimensional limit where the sample size and feature dimension diverge at a fixed ratio. Additionally, this analysis establishes a heavy-tail universality law, scaling laws relating typical errors to prediction reliability, and the Bayes-optimal prediction error. In addition to characterizing the effects of extreme entries on the learning process, our method applies broadly to other systems with persistent local heterogeneity.
[LG-167] Explicit Asymptotic Bounds for Sequential Calibration Beyond T2/3
链接: https://arxiv.org/abs/2610.07623
作者: Eric Dai,Maxwell Fishelson
类目: Machine Learning (stat.ML); Data Structures and Algorithms (cs.DS); Computer Science and Game Theory (cs.GT); Machine Learning (cs.LG)
*备注:
Abstract:Probability forecasts are calibrated when predicted probabilities match empirical outcome frequencies: among events assigned a probability p , we’d hope that the fraction of positive outcomes is close to p . We study the problem of sequential forecasting of binary outcomes. The classical O(T^2/3) bound on expected cumulative \ell_1 -calibration error established by Foster and Vohra stood for over two decades until Dagan et al. reduced the exponent 2/3 by an unspecified constant. We establish a new two-phase recursive labeling strategy for the sign-preservation-with-reuse game that yields the bound O(n^\alphat^\beta) for all choices of space and time. We then sharpen the reduction from upper bounds on sign preservation to calibration by modifying the equivalence of Dagan et al. to use only O(\log T) instances of the sign-preservation-with-reuse game. This lets us establish an explicit bound of O(T^0.662942288) , the first explicit exponent below 2/3 for sequential calibration, by combining both improvements and choosing explicit feasible parameters. Subjects: Machine Learning (stat.ML); Data Structures and Algorithms (cs.DS); Computer Science and Game Theory (cs.GT); Machine Learning (cs.LG) Cite as: arXiv:2610.07623 [stat.ML] (or arXiv:2610.07623v1 [stat.ML] for this version) https://doi.org/10.48550/arXiv.2610.07623 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[LG-168] Is sqrtd Separation Necessary for Gradient EM to Learn Gaussian Mixtures in High Dimensions?
链接: https://arxiv.org/abs/2610.07551
作者: Yiran Zhang,Mo Zhou,Weihang Xu,Maryam Fazel,Simon S. Du
类目: Machine Learning (stat.ML); Machine Learning (cs.LG)
*备注: 51 pages
Abstract:Learning Gaussian mixture models (GMMs) using the Expectation-Maximization (EM) algorithm and its gradient-based variants is a fundamental problem in machine learning. It is known that randomly initialized (gradient) EM fails to learn multi-component GMMs in the exact-parameterized setting, where the number of components matches that of the ground-truth GMM. Recently, global convergence of gradient EM has been established in the over-parameterized setting, where more components are used, provided that the ground-truth components are well separated. In particular, the minimum separation between ground-truth components is required to scale as \Omega(\sqrtd) , where d is the dimension. In this paper, we show that this dimensional dependence is unavoidable in high-dimensional settings. Specifically, we consider a hybrid EM algorithm that uses standard EM updates for the mixing weights and gradient EM updates for the component means. For any \epsilon 0 , we prove that when the dimension is sufficiently large, in the worst case a separation of order \Omega(d^0.5-\epsilon) is insufficient to guarantee global convergence of population gradient EM in sub-exponential time under random initialization, even in the over-parameterized regime. Our result establishes an almost optimal worst-case lower bound on the ground-truth separation required for learning Gaussian mixtures via gradient EM in high dimensions.
[LG-169] wo-Sample Testing for Random Graphs without Vertex Correspondence
链接: https://arxiv.org/abs/2610.07503
作者: Soham Dan
类目: Machine Learning (stat.ML); Machine Learning (cs.LG)
*备注:
Abstract:Two populations of graphs often have to be compared without any correspondence between their vertices, for instance when networks come from different communities, or when a graph generative model is evaluated against held-out graphs. We study how many graphs such an unaligned two-sample test needs, and which graph statistics can detect which differences. For an Erdős–Rényi null and a planted two-block difference that leaves every expected degree unchanged, we show that m\asymp t^-3 graphs per group are necessary and sufficient when the per-graph signal-to-noise ratio is t1 . Signed triangle counts attain this rate, and the lower bound holds for every graph size. With aligned vertices m\asymp t^-1 graphs suffice, so misalignment costs a factor of order t^-2 . When the triangle signal cancels, the rate becomes t^-4 and 4 -cycles are needed. Statistics built from trees have exactly the same expectation under both hypotheses, and tests based on finitely many of them have asymptotically no power. In the graphon limit, this class includes degree distributions and message-passing graph neural network features. For a non-constant null, a generic difference is visible at first order, and a simple motif test attains the aligned order of sample size, suggesting that misalignment is costly mainly for differences that are invisible at low orders. We also give an exactly valid test for one or two graphs per group, at a cost in power. In our simulations, the fitted exponents are close to the predicted ones, and degree-based and random-GNN evaluation metrics stay at their level in a setting where signed triangles need about 65 graphs.
[LG-170] Bayesian Optimization on Function Spaces via Sparse RKHS Manifolds
链接: https://arxiv.org/abs/2610.07417
作者: Davide Sartor,Meghan E. Huber,Donghyun Kim,Nathan Wycoff
类目: Machine Learning (stat.ML); Machine Learning (cs.LG)
*备注:
Abstract:Bayesian Optimization (BO) has become an established methodology for minimizing black-box functions of a vector input. Often, however, this parameter vector arises from the discretization of an inherently functional relationship. Several recent articles have considered the Functional Bayesian Optimization (FBO) setting, in which the variable to be optimized is not a member of a finite dimensional vector space, but rather an infinite dimensional function space. In this work, we propose L^0 Manifold Optimization (L0MO), a simple approach to FBO which searches the subset of a Reproducing Kernel Hilbert Space (RKHS) consisting of functions with a sparse representation in the kernel functions, optimizing both the kernel locations and their coefficients. We discuss in detail the relationship between our method and existing ones, providing a unifying lens through which to view prior works. To assess our method against the state of the art, we conduct an extensive computational study, and along the way develop a novel set of benchmark test functions which port standard finite-dimensional ones to the infinite dimensional domain. Our experiments demonstrate that, on balance, the proposed method achieves superior performance across a wide range of test benchmarks.
[LG-171] A perspective note on likelihood approximation and inference for complex simulation models using a chain of aggregated normalizing flows
链接: https://arxiv.org/abs/2610.07391
作者: Getachew K Befekadu
类目: Methodology (stat.ME); Machine Learning (cs.LG); Machine Learning (stat.ML)
*备注: 17 pages, 1 figure
Abstract:We present a new perspective on the problem of likelihood approximation within the framework of simulation-based inference that promotes scalable and controllable simulation routines for large-scale data analysis, allows efficient parameter space exploration or smooth interpolation in high-dimensions and, thus, supports valid statistical treatments of hypothesis testings as well as uncertainty quantification. In particular, we consider a chain of n -aggregated normalizing flows for likelihood approximation scheme, where a set of upfront replicated observation datasets from the forward complex simulation model pass through the first set of bijective transformations, and then subsequently pass to the other sets of bijective transformations. Here, we assume that, for any k \in \1,,2, \ldots, n\ , the parameters corresponding to the first k sets of bijective transformations are estimated sequentially, in some sense of optimality, for constructing flexible probability distributions, regardless of the remaining (n-k) sets of bijective transformations. Moreover, our objects of interest are to highlight two complementary mathematical arguments that leverage an informatics-theoretic formalization, based-on empirical likelihood estimators under moment restrictions, and a sequential decision-making paradigm, with mixing distributions, for updating and aggregating the estimated parameters of the overall normalizing flows. As a by-product, the framework provides a reliable surrogate model, conditioned on the model parameters defining the forward computational simulation, that allows samples generation, with statistical powers, and facilitates computationally tractable scheme in the Bayesian paradigm for inference, hypothesis testings and uncertainty quantification.
[LG-172] HyperNSDE: Personalized Neural SDEs for Joint Static-Longitudinal Clinical Data Generation
链接: https://arxiv.org/abs/2610.07383
作者: Perrine Chassat,Agathe Guilloux
类目: Machine Learning (stat.ML); Machine Learning (cs.LG); Applications (stat.AP); Methodology (stat.ME)
*备注:
Abstract:Synthetic patient data generation is a promising solution to the dual challenge of data scarcity and privacy constraints in healthcare machine learning. Realistic synthesis of patient-level clinical data requires jointly modeling heterogeneous static covariates, irregularly sampled longitudinal trajectories, and informative observation times - three tightly coupled components in practice yet rarely addressed together. We propose HyperNSDE, a continuous-time generative model that conditions a latent Neural SDE on static patient representations through a hypernetwork, allowing baseline characteristics to shape trajectory evolution beyond the initial condition without requiring a trajectory encoder, while stochastic latent dynamics capture realistic variability in generated paths. Observation times are modeled jointly through a latent-state-dependent intensity process, and training on irregular stochastic paths is stabilized via a deterministic-stochastic path decomposition with a non-adversarial signature-kernel objective. Experiments on simulated and real clinical datasets show improved observation-time fidelity and competitive performance, while matched-grid analyses reveal that forecasting and correlation metrics are affected by observation-grid regularity and trajectory smoothness.
[LG-173] Advantage of Entangled Learning Rules in Quantum Measurement Class Learning
链接: https://arxiv.org/abs/2610.07328
作者: Arka Prabha Das,Abram Magner
类目: Quantum Physics (quant-ph); Machine Learning (cs.LG)
*备注: 9 pages
Abstract:Learning with data in the form of quantum states is of current interest and has led to a variety of problems that boil down to interaction with the available data via quantum measurement and classical post-processing of observed classical outcomes. In quantum measurement PAC learning, one is given a sequence of unknown, prepared quantum states and classical labels, along with a hypothesis class of candidate measurements. The task is to select a measurement from the hypothesis class that minimizes a fixed notion of error in prediction of the classical labels via measurement of a new state by the selected hypothesis. In this work, we consider the advantage of interacting with the given data in the measurement learning framework using learning rules given by measurements that cannot be implemented using local operations and classical communication (LOCC), as opposed to single-copy learning rules. We provide a construction showing that there exist learning scenarios wherein single-copy learning rules are asymptotically suboptimal compared to optimal ones. We then show that learning rules based on entangled measurements enjoy at most a polynomial sample complexity advantage over single-copy learning rules in the PAC learning setting (under a natural joint measurability covering assumption).
[LG-174] Assumption-lean logistic regression with missing covariates
链接: https://arxiv.org/abs/2610.07292
作者: Jyotishka Ray Choudhury,Kabir Aladin Verchand,Richard J. Samworth,Ashwin Pananjady
类目: Machine Learning (stat.ML); Information Theory (cs.IT); Machine Learning (cs.LG); Statistics Theory (math.ST); Methodology (stat.ME)
*备注:
Abstract:Missing covariates are frequently encountered in supervised learning problems, and classical methods for estimation using such data use carefully chosen imputation schemes for missing data, or likelihood approximations that lead to nonconvex M -estimation problems. These methods and their relatives are suitable for scenarios in which the covariate distribution is known, and more broadly, have enjoyed tremendous success in linear models. But even in basic nonlinear problems such as logistic regression in moderate dimensions, such methods can experience drastic failure modes when the covariate distribution is unknown. Motivated by the need for reliable alternatives, we consider the problem of parameter estimation in logistic regression with missing covariates. Crucially, we operate in the assumption-lean setting where the covariate distribution is unknown (but bounded). We design a stochastic approximation method that is based on Z -estimation with a novel monotone operator, and establish that our algorithm is computationally efficient and achieves provable signal recovery at parametric rates under the hypothesis that covariates are missing completely at random. Our theory sharply characterizes the \ell_2^2 risk of the estimator in terms of the missingness profile, accommodating heterogeneous observation probabilities. Importantly, it shows that our method always outperforms the de facto ``complete-case’’ estimator that ignores observations with any missing data. Even in the setting with homogeneous missingness (in which each covariate is observed independently with probability q ), our bounds exhibit intricate and nonstandard dependence on q that can yield significant improvements over using only complete cases. We complement our upper bounds with new information-theoretic lower bounds that show that this intricate dependence on q is fundamental in a minimax sense. Subjects: Machine Learning (stat.ML); Information Theory (cs.IT); Machine Learning (cs.LG); Statistics Theory (math.ST); Methodology (stat.ME) MSC classes: 62J12 (Primary), 62C20, 62L20 (Secondary) Cite as: arXiv:2610.07292 [stat.ML] (or arXiv:2610.07292v1 [stat.ML] for this version) https://doi.org/10.48550/arXiv.2610.07292 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[LG-175] A Single-Loop Constant-Batch First-Order Penalty Method for Stochastic Bilevel Optimization
链接: https://arxiv.org/abs/2610.07290
作者: Xingyu Chen,Ming Yang,Quanqi Hu,Tianbao Yang
类目: Optimization and Control (math.OC); Machine Learning (cs.LG)
*备注:
Abstract:Recent advances in penalty-based methods for stochastic bilevel optimization (SBO) have eliminated the need for second-order derivative oracles. However, for stochastic nonconvex-strongly convex bilevel problems, existing first-order methods typically rely on nested loops and/or large batch sizes for attaining O(\epsilon^-6) or O(\epsilon^-4) sample complexity under standard bounded-variance assumption or mean-square smoothness assumption. Achieving these rates with a single-loop penalty method and a constant batch size remains challenging due to a large penalty value needed for an accurate approximation. To address this challenge, we develop a stochastic SIngle-loop COnstant-Batch first-order penalty method (SICO) that combines two complementary ingredients. First, it performs one stochastic-gradient update per-iteration for both the original lower-level and penalized problems, with a projection that controls the separation between their iterates. Second, it applies an exponential moving average to stabilize the upper-level gradient estimator. We show that this combination achieves O(\epsilon^-6) sample complexity using only O(1) stochastic-gradient samples per iteration under unbiased, bounded-variance stochastic gradients. Under the additional mean-square smoothness assumption on the lower-level stochastic gradients, the same algorithm improves the complexity to O(\epsilon^-4) also with O(1) batch size. To the best of our knowledge, this is the first work to match the best-known convergence rate for fully first-order SBO methods using a single loop and a constant batch size. This result addresses an open problem posed in the literature.
[LG-176] How Inefficient Is Natural Gradient Descent? From Exact Optimality to Θ( sqrt log d ) Divergence
链接: https://arxiv.org/abs/2610.07228
作者: Guni Sharon,Alan Kuhnle
类目: Machine Learning (stat.ML); Machine Learning (cs.LG)
*备注:
Abstract:Natural gradient descent (NGD) underlies common methods in ML. For dually flat families, idealized NGD on the forward Kullback–Leibler objective follows the mixture geodesic which is often longer than the shortest Fisher–Rao path. We quantify this overhead by the inefficiency ratio (R \ge 1), the Fisher length of the mixture geodesic divided by the Fisher–Rao distance, and bound its supremum over endpoint pairs as a function of the parameter dimension (d). A tensor criterion identifies the regime (I) families, with (R=1) everywhere: exactly those with quadratic potential or dimension one, such as fixed-covariance Gaussians. For non-quadratic families, we prove two further regimes: (II) bounded third-order skewness plus finite Fisher–Rao diameter yields a dimension-independent bound; and (III) for products of scale families—including Gaussian covariances and Gamma rates—(R) grows as (\Theta(\sqrt\log d)), unbounded in (d). Under a per-step Fisher-chord budget, (R) translates to a practical computational cost: NGD requires asymptotically at least (R) times as many steps as an optimizer following the Fisher–Rao geodesic. Experiments confirm all three regimes: (R=1) to machine precision for quadratic-potential families (I), the categorical bound (\pi/(2\sqrt2)) is approached but not attained (II), and sampled scale-product (R) grows with (d), reaching (R \approx 1.5) for long, high-dimensional moves (III).
[LG-177] Learning Disentangled Representations with Quantum Variational Autoencoders
链接: https://arxiv.org/abs/2610.07196
作者: Gaoyuan Wang,Jerry Tan,Mark Gerstein
类目: Quantum Physics (quant-ph); Emerging Technologies (cs.ET); Machine Learning (cs.LG)
*备注:
Abstract:Variational autoencoders are powerful representation learning models that map complex data into low-dimensional latent spaces, enabling the discovery of interpretable and disentangled factors. Such representations can facilitate the interpretation and controllable generation of data describing complex scientific systems. Understanding how these factors are organized and encoded in latent space is therefore important for developing reliable representation learning models. Recently, quantum variational autoencoders (QVAEs) have been proposed as quantum representation models, demonstrating informative latent representations and improved latent-space occupancy through quantum regularization. However, it remains unclear whether and how QVAEs can learn disentangled and interpretable latent factors. A key challenge in investigating quantum latent factors is that a small number of qubits spans an exponentially large Hilbert space, making the notion of an individual quantum latent dimension nontrivial. Here, we investigate what constitutes an individual quantum latent dimension and whether it can encode a distinct factor. We develop theoretical insights into quantum latent dimensions and support them with empirical studies on representative synthetic problems, including MNIST variants. Across three datasets, we demonstrate that QVAEs can discover factorized and semantically interpretable latent representations, with individual qubits functioning as meaningful latent factors. These results establish a foundation for understanding quantum latent spaces and their potential for structured and interpretable representation learning.
[LG-178] A Query Is Not a Commitment: Learning to Correct Expert Answers in Online Deferral
链接: https://arxiv.org/abs/2610.07084
作者: Yannis Montreuil,Axel Carlier,Lai Xing Ng,Wei Tsang Ooi
类目: Machine Learning (stat.ML); Machine Learning (cs.LG)
*备注:
Abstract:An inaccurate expert can still provide useful information after correction. We study online learning to defer in which the learner chooses an expert and fixes a correction function before purchasing its answer, then applies that function to the answer received. The difficulty is that observed losses reflect both expert quality and an unfinished correction: early errors can discourage queries that would be valuable after learning. We propose ORUCB, which pools shared and expert-specific polynomial responses. A bound on cumulative response-learning error calibrates confidence-weighted risk regression and exploration, allowing the router to account for this error when deciding which answers to buy. Under bounded residuals and disagreements, a fixed feasible model of optimal responses, and linear models of free and optimal queried risk, the calibrated algorithm achieves high-probability pseudo-regret O(\sqrt T\log(T+1)) over T rounds for fixed problem parameters. The guarantee permits singular answer distributions and misspecified shared responses; optimality is relative to the bounded response class. On four test streams, the selected cubic policy has lower fee-inclusive cost than seven baselines that deploy answers unchanged. Comparisons with a common correction learner examine routing, while six-price comparisons measure cost and query rates.
[LG-179] Learning Decision-Stump Thresholds in Context: Dynamics of Softmax Attention
链接: https://arxiv.org/abs/2610.07074
作者: Hong Ha Le,Jackie Lok,Atsushi Nitanda,Yan Shuo Tan
类目: Machine Learning (stat.ML); Machine Learning (cs.LG)
*备注:
Abstract:Estimating a decision threshold requires locating observations near an unknown boundary. We study how gradient-based pretraining learns this statistical rule in a two-parameter softmax-attention model with a fixed feature and inequality direction. Pretraining uses labeled contexts and their true thresholds; a fresh threshold must be inferred from context alone. Under a large-resolution initialization, constant-step gradient descent on m tasks with n examples each produces a frozen estimator with error \widetilde O((m\wedge n)^-1+N^-1) for each fixed interior threshold and every fresh-context size N . The two terms separate finite-pretraining accuracy from fresh-context localization. The mechanism is coordinated parameter divergence: population training calibrates the relative label and feature scores, then increases the attention scale as t^1/4 , giving population threshold error O(t^-1/4) . To transfer this mechanism to a fixed finite corpus, we control gradient errors relative to the shrinking directions of progress at successive parameter scales. This certifies a growing training interval without requiring long-time tracking of the population trajectory. We also identify the boundary limitation of the one-head model and explain statistically what a reflected symmetrization could achieve.
[LG-180] sHAIL-Causal: A Sequential Staircase Procedure for Invariant Causal Predictor Discovery
链接: https://arxiv.org/abs/2610.07057
作者: Ernest Fokoué
类目: Machine Learning (stat.ML); Machine Learning (cs.LG)
*备注: 17 pages, 2 figures, 2 tables. Introduces the general sHAIL learning framework and its causal specialization
Abstract:We introduce sHAIL-Causal, the causal specialization of the Saturated Hierarchical Atomic Incremental Learning (sHAIL) paradigm: a sequential staircase procedure that ascends a nested hierarchy of hypothesis classes H_0 H_1 … H_K once a saturation signal indicates that mastery of the current stage has plateaued. Where general sHAIL leaves the saturation criterion open, sHAIL-Causal instantiates it with a joint criterion of goodness-of-fit saturation and cross-environment invariance, replacing the complexity control of Structural Risk Minimization. We show, theoretically and by simulation, that complexity-only staircases are seduced by confounded predictors that lower empirical risk without reflecting stable causal structure, whereas an invariance-gated staircase provably halts at the true causal predictor set under a per-variable Richness condition. We show that naive greedy search fails to recover the causal set even under Richness, trace the failure to non-monotonicity of the invariance statistic along single-variable paths, and validate a fix combining bounded-exhaustive block-seeding with a calibrated acceptance threshold. We then extend the guarantee to environments arriving sequentially, yielding a confidence guarantee that stays valid at every arrival, which one-shot exhaustive search cannot offer without repeating its full combinatorial search. We close by formalizing the intervention of a wise teacher who lifts a saturated learner off a plateau of boredom.
[LG-181] Data Fusion for Errors-in-Variables
链接: https://arxiv.org/abs/2610.07048
作者: Huali Zhao,Molei Liu,Tianying Wang
类目: Machine Learning (stat.ML); Machine Learning (cs.LG); Statistics Theory (math.ST); Methodology (stat.ME)
*备注:
Abstract:We study errors-in-variables problems in which a target study contains only a single error-prone surrogate of an unobserved exposure, while an external source study provides repeated surrogate measurements from a different population. The measurement error distribution is allowed to depend on the observed error-free variables, and the error-free variable distribution itself may differ between studies. We introduce a conditional transportability assumption that enables the use of external repeated measurements under source-target heterogeneity. Together with additional replicate-error conditions, it identifies the target conditional measurement-error distribution. Building on this identification result, we develop a data-fusion estimator for a broad class of target functionals. The estimator combines conditional deconvolution, flexible nuisance estimation, and orthogonal correction that reduces first-order sensitivity to nuisance estimation. For the proposed estimator, we develop a unified spectral theory covering both diffuse-spectrum and finite atomic-spectrum target functionals, derive a general asymptotic expansion, and establish consistency and target-specific convergence-rate bounds. The resulting convergence-rate bounds depend jointly on the spectral properties of the measurement error, the latent exposure, and the target functional. For finite atomic-spectrum targets, we further establish joint Gaussian and bootstrap limits, yielding inference for smooth moment transformations under an additional centering condition. In the reported simulations, Fuse-EIV has small bias for the primary exposure-related coefficient. Applications to the National Health and Nutrition Examination Survey illustrate how accounting for population heterogeneity and error heteroscedasticity can change empirical conclusions.
[LG-182] FactorBench: A Portfolio-Aware Benchmark for Automated Factor Mining
链接: https://arxiv.org/abs/2610.06947
作者: Zhuohan Wang,Carmine Ventre
类目: Portfolio Management (q-fin.PM); Machine Learning (cs.LG)
*备注: 30 pages, 21 figures, 8 tables
Abstract:Factor mining seeks to discover signals from financial data that predict future asset returns and guide portfolio construction. Automated factor mining now spans genetic programming, reinforcement learning, generative models, and large language model agents. Yet it remains unclear whether advances across these paradigms yield more generalizable, distinct, and economically useful financial signals. We introduce FactorBench, a portfolio-aware benchmark comparing roughly five thousand mined factors from nine automated mining methods across five equity markets. A shared data and evaluation contract supports both symbolic expressions and executable Python factors, connecting heterogeneous discovery algorithms to common signal combination and portfolio construction procedures. FactorBench traces the outputs of mining systems across three levels: factor validity, temporal generalization, and predictiveness beyond measured risk and style exposures; within- and across-method pool distinctness, including similarity to the benchmark Alpha101; and composite-signal quality and after-cost long-only and long–short portfolio performance. After systematically assessing whether advances in factor mining translate into signal quality and portfolio performance, FactorBench finds that no paradigm consistently dominates.
[LG-183] Low-Rank and Structured Sparse Tensor Decomposition for Anomaly Detection in Multivariate Functional Data
链接: https://arxiv.org/abs/2610.06930
作者: Mohammad N. Bisheh,Che-Yi Liao,Kamran Paynabar
类目: Machine Learning (stat.ML); Machine Learning (cs.LG); Applications (stat.AP); Computation (stat.CO)
*备注:
Abstract:Multivariate functional data arise in many modern manufacturing systems, where multiple sensors record densely sampled process trajectories. Monitoring such data is challenging because nominal variation is strongly correlated across samples, sensors, and time, while faults may appear either as isolated deviations or as structured departures concentrated within a limited number of sensor-specific temporal trajectories. We propose two unsupervised sparse tensor decomposition methods that preserve this multimode structure. Entrywise Sparse CP Decomposition (ES-CP) uses an entrywise (\ell_1) penalty to identify localized anomalies, whereas Fiberwise Sparse-Group Lasso CP Decomposition (FG-Lasso) combines entrywise and fiberwise penalties to detect both localized deviations and anomalies concentrated within temporal fibers. Both methods represent nominal process behavior through a low-rank CP decomposition and are estimated using alternating optimization with closed-form sparse-component updates. Two simulation studies evaluate performance under different fault structures, signal severities, noise levels, and missing observations. FG-Lasso attains or ties the highest macro F _1 score in almost all settings in the first study and achieves the highest macro F _1 score. In a multichannel forging-process case study, FG-Lasso and ES-CP obtain macro F _1 scores of 0.85 and 0.82, respectively, compared with 0.69 or lower for TRPCA and PCA-based anomaly detectors. The results demonstrate that explicitly matching the sparse penalty to the anticipated fault structure improves both anomaly detection and fault localization in high-dimensional functional processes.
[LG-184] Nonlocal Hamiltonian Dynamics on Sparse Lévy Graphs: Spectral Analysis and Multimodal Sampling
链接: https://arxiv.org/abs/2610.06904
作者: Miaolei Zheng,Ting Gao,Jinqiao Duan
类目: Dynamical Systems (math.DS); Machine Learning (cs.LG)
*备注:
Abstract:We develop a sparse graph method for transporting probability mass toward multimodal target distributions through damped nonlocal Hamiltonian dynamics. The formulation combines logarithmic-mean mobility with symmetric Lévy-type interaction weights, coupling the evolving density to an edge momentum field. A graph constructed from nearest-neighbor connections and sampled long-range edges provides direct mass exchange between spatially separated regions. Once the graph is constructed, the density evolution is deterministic, and each update costs linear in the number of nodes and the long-range sampling budget. Linearization around the target distribution yields a damped oscillator governed by a weighted graph Laplacian. Its spectrum characterizes the interaction between nonlocal connectivity and inertia, with the Lévy exponent alpha tuning the nonlocal connectivity: the spectral gap determines the optimal asymptotic damping, while the largest eigenvalue governs the time-step stability. Experiments on synthetic multimodal distributions demonstrate improved mode balance and more stable mode coverage relative to first-order and MCMC baselines. The resulting framework provides a sparse implementation of nonlocal inertial density transport for sampling problems with low-dimensional spatial structure.
[LG-185] Memory Prediction Excess: A Probabilistic Quantity for Predictive Gain and Memory Length in Stochastic Processes
链接: https://arxiv.org/abs/2610.06894
作者: Jiahao Jiang
类目: Machine Learning (stat.ML); Information Theory (cs.IT); Machine Learning (cs.LG); Probability (math.PR)
*备注:
Abstract:A central question in the prediction of stochastic processes is the extent to which past information can improve the probability of correctly predicting the next state. We introduce the Memory Prediction Excess (MPE) to address this question quantitatively. The MPE measures the average improvement in prediction accuracy obtained by using the entire observed history relative to using only the static marginal distribution, in discrete-time finite-state processes. It is defined as the difference between the expected optimal conditional prediction accuracy and the optimal static prediction accuracy. Its basic properties are examined: the MPE is always non-negative; it admits an upper bound depending on the static accuracy, attained if and only if the future is almost surely a deterministic function of the past; and degenerate cases in which the MPE vanishes are characterized. A normalized version, taking values in the unit interval, is introduced as a dimensionless measure of predictive efficiency. A lower bound is derived by comparing predictions based on histories of different lengths, showing that the expected optimal prediction accuracy is monotone with respect to the history length. The framework is extended to finite-length histories, where the finite-history MPE (FH-MPE) measures the predictive gain attainable when only the most recent observations are retained. This leads to the notion of a minimal memory length required to achieve the same predictive performance as the full history. For finite-order Markov chains, this minimal memory length is shown to be bounded by the Markov order. The MPE and its variants are formulated in terms of conditional probabilities and prediction accuracies, offering a probabilistic perspective on the predictive utility of memory that is complementary to classical information-theoretic approaches.
[LG-186] Statistical Turbulence and High-Fidelity Disturbance Fields for Quadrotor Flight Control
链接: https://arxiv.org/abs/2610.06874
作者: Xun Huang
类目: Fluid Dynamics (physics.flu-dyn); Machine Learning (cs.LG); Robotics (cs.RO)
*备注:
Abstract:Reinforcement-learning quadrotor controllers are usually trained under simplified wind models, yet the impact of wind-field fidelity, as opposed to magnitude, on policy robustness remains unquantified. This paper compares five disturbance-fidelity levels, from wind-free flight and discrete 1-cosine gusts through statistical turbulence and synthetic coherent structures to large-eddy-simulation fields of the atmospheric boundary layer, in a full cross-fidelity train test evaluation of proximal policy optimization (PPO) agents, with cascaded PID and geometric SE(3) controllers as training-free references, over a 0-12 m/s wind sweep. Before any controller comparison is made, all disturbance data are validated: every synthetic generator is checked quantitatively against its analytical or certification-standard reference, and the large-eddy-simulation fields against the imposed log law. On a racing-class quadrotor in hover, the train test matrix is remarkably flat, and the cheapest structured training wind, which is discrete-gust domain randomization, ranks first in every test column, a ranking replicated on a wind-sensitive 27 g platform; a once-tuned geometric controller brackets the learned PPO policies at zero crash rate. Mechanism diagnostics show that control authority, not wind realism, bounds robustness, so wind-fidelity investment should scale with platform wind sensitivity.
[LG-187] he Geometry of Statistical Feature Learning in Mean-Field Langevin Dynamics
链接: https://arxiv.org/abs/2606.31429
作者: Zong Shang,Tomoya Wakayama,Guillaume Lecué,Taiji Suzuki
类目: atistics Theory (math.ST); Machine Learning (cs.LG); Machine Learning (stat.ML)
*备注:
Abstract:We introduce a geometric formulation of statistical feature learning for supervised regression. Feature learning is defined through a base–fiber decomposition: the base is the feature-side geometry produced by training, and the fiber is the learned feature space where estimation is performed. We prove this property for spherical mean-field Langevin dynamics, viewed as the Wasserstein gradient flow of a negative entropy-regularized empirical risk. In Gaussian multi-index models, the low-temperature stationary distribution concentrates near the hidden indices, forms a multi-spike structure, and yields parameter recovery with high probability, even though negative entropy regularization penalizes concentration. This concentration has a sharp transition at temperature \lambda\asymp 1 . In Gaussian single-index models, the stationary measure satisfies a concentration property, with parity determining whether it lives on S_2^d-1 or \mathbbRP^d-1 . The induced learned feature space aligns the regression signal and yields rates d/N and Md/N , up to logarithmic factors.
附件下载


