本篇博文主要内容为 2026-10-02 从Arxiv.org论文网站获取的最新论文列表,自动更新,按照NLP、CV、ML、AI、IR、MA六个大方向区分。
说明:每日论文数据从Arxiv.org获取,每天早上12:30左右定时自动更新。
提示: 当天未及时更新,有可能是Arxiv当日未有新的论文发布,也有可能是脚本出错。尽可能会在当天修复。
目录
概览 (2026-10-02)
今日共更新1249篇论文,其中:
- 自然语言处理共185篇(Computation and Language (cs.CL))
- 人工智能共383篇(Artificial Intelligence (cs.AI))
- 计算机视觉共215篇(Computer Vision and Pattern Recognition (cs.CV))
- 机器学习共431篇(Machine Learning (cs.LG))
- 多智能体系统共23篇(Multiagent Systems (cs.MA))
- 信息检索共21篇(Information Retrieval (cs.IR))
- 人机交互共28篇(Human-Computer Interaction (cs.HC))
多智能体系统
[MA-0] Watch Infer Coordinate: Inferring Robot Partner Constraints for Zero-Shot Coordination
【速读】:该论文旨在解决在物理耦合的协作操作任务中,机器人伙伴之间因硬件退化或执行器故障导致的未知物理约束问题。由于这些限制可能无法直接观测,且另一机器人可能通过补偿行为掩盖其真实能力,使得仅通过单个机器人的动作难以准确推断其实际能力。解决方案的关键在于利用团队协作中的联合行为(joint behavior)作为信息源:尽管单一机器人展示的是“已执行”的动作,而非“可执行”的动作,但其与协作者的协同模式会受到其物理约束的深刻影响。因此,本文提出“观察-推断-协调”(Watch, Infer, Coordinate)框架,通过分析多机器人协作过程中可观测的联合行为,对候选约束进行评分以实现高精度的约束推断,并在此基础上实现零样本(zero-shot)协作。实验表明,该方法在三个物理耦合操作场景下显著提升了约束推断准确性与协作性能,接近拥有真实约束信息的最优基准(oracle)。
链接: https://arxiv.org/abs/2610.02170
作者: Suyu Ye,Zheyuan Zhang,Vaishnav Tadiparthi,Hossein Nourkhiz Mahjoub,Ehsan Moradi Pari,Tianmin Shu,Homanga Bharadhwaj,Nakul Agarwal
机构: Honda Research Institute USA; Johns Hopkins University (约翰霍普金斯大学)
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA)
备注:
Abstract:Robots operating in the physical world will increasingly need to coordinate with other robots, particularly in manipulation tasks where an object may be too large or heavy for a single robot to carry alone. Physical limitations caused by hardware degradation or actuator faults can restrict the actions a robot can reliably execute, yet these limitations may be unknown to its partner. We study whether a helper can infer a robot partner’s physical constraints from observing it coordinate with another robot, then use the inferred capability to coordinate with the same partner on a new task. This is difficult because a demonstration shows what the constrained robot did, but not what it could have done. In physically coupled tasks, the other robot may also compensate for its limitations, making those limitations difficult to identify from the constrained robot’s behavior alone. Our key insight is that these constraints shape the joint behavior of the team, making the actions of both robots informative about the constrained partner’s capability. We introduce Watch, Infer, Coordinate, a benchmark spanning three physically coupled manipulation settings, together with an inference approach that scores candidate constraints using observed joint behavior. Across all three settings, our method substantially improves constraint inference and zero-shot coordination, approaching an oracle with access to the true constraints.
[MA-1] Decentralized Power-Optimal Coordination for Spacecraft Swarms Using Time-Varying Magnetorquer Actuation
【速读】:该论文旨在解决磁力驱动航天器集群在形成大型空间结构时面临的协同控制难题,特别是如何在无推进剂、依赖太阳供电的约束下实现能量最优的分布式协调。其核心问题是:在磁力矩器(magnetorquer)产生的非对称相互作用下,如何设计一种去中心化的控制框架,以联合优化交互图拓扑、载波频率分组及控制器增益,同时满足角动量守恒这一非完整约束。解决方案的关键在于提出一种去中心化的能量最优协调框架,通过动态构建交互网络与频率分组策略,在保证相对位置误差、绝对姿态误差及飞轮动量不平衡性收敛至期望状态的前提下,实现全局能量效率最优化。此外,该框架通过引入具有严格误差界保证的快速近似积分方法,扩展至长时间轨道重构任务,实现了高精度、可扩展的集群协同控制。
链接: https://arxiv.org/abs/2610.02118
作者: Yuta Takahashi,Shin-ichiro Sakai
机构: Institute of Science Tokyo (东京科学研究所); Japan Aerospace Exploration Agency (日本宇宙航空研究开发机构)
类目: ystems and Control (eess.SY); Multiagent Systems (cs.MA)
备注: Submitted to IEEE Transactions on Aerospace and Electronic Systems
Abstract:This paper presents a decentralized power-optimal coordination framework for magnetically actuated spacecraft swarms. Swarms that form large space structures overcome the aperture limit set by the launch vehicle and hold their shape on solar-generated power alone. Magnetic actuation is propellant-free and generated by a magnetorquer, which is commonly used for attitude control. However, every spacecraft interacts with every other within range, and its effect depends on the actuation power and a carrier frequency. We therefore design a decentralized power-optimal framework to jointly derive the interaction graph, frequency grouping, and controller gains. Our decentralized controller preserves angular momentum, which is a nonholonomic constraint. Then, this framework for connected groups whose memberships overlap across carriers guarantees that the relative position errors, the absolute attitude errors, and the imbalance of the reaction-wheel momenta converge to the desired states under the decentralized power-optimal allocation. A closed-loop simulation of a thousand spacecraft with the complete alternating-current interaction confirms the framework. A fast approximate integration with a proven error bound extends the framework to a long-horizon orbital reconfiguration held with high precision.
[MA-2] Form and Void: Entangled Composition through an Autonomous AI Agent
【速读】:该论文旨在解决生成具有明确正负空间关系的视觉构图这一挑战性问题,即在共享边界条件下对正空间(前景)与负空间(背景)两个语义概念进行协同控制。尽管当前文本到图像模型及多模态大语言模型(MLLMs)在图像生成与视觉理解方面表现优异,但在直接单次提示(single-pass prompting)下仍难以有效生成语义一致且视觉连贯的正负空间布局。为此,论文提出一种名为“形态与空隙代理”(Form and Void Agent, FaV-A)的多模态智能体,其核心解决方案在于采用分阶段渐进式工作流:首先生成基础物体,继而分析其形状与空间结构以识别潜在的负空间语义,最终生成用于指导最终图像生成的组合指令。实验结果与消融研究验证了该方法相较于直接零样本MLLM基线,在生成视觉连贯、语义对齐的正负空间构图方面具有显著优势。
链接: https://arxiv.org/abs/2610.02045
作者: Shiwen Wang,Jian Yang,Xu Wang,Xincan Wang,Weiming Dong
机构: University of Chinese Academy of Sciences (中国科学院大学); Renmin University of China (中国人民大学); Shanghai Theatre Academy (上海戏剧学院); MAIS, Institute of Automation, Chinese Academy of Sciences (中国科学院自动化研究所智能系统重点实验室)
类目: Computer Vision and Pattern Recognition (cs.CV); Multiagent Systems (cs.MA)
备注:
Abstract:Positive and negative space is a fundamental principle in visual composition, supporting visually coherent forms and layered semantic relationships. Generating such compositions is challenging because it requires coordinated control over two semantic concepts that share a common boundary. Although recent text-to-image models and multimodal large language models (MLLMs) have achieved strong performance in image generation and visual understanding, positive-negative space generation remains difficult, particularly under direct single-pass prompting. In this work, we present the \textbfForm \textbfand \textbfVoid \textbfAgent (\textbfFaV-A), a multimodal agent designed for staged positive-negative space generation. FaV-A follows a progressive workflow: it first generates a base object, then analyzes its shape and spatial structure to identify candidate negative-space semantics, and finally produces compositional instructions for the final image generation stage. Experimental results and ablation analyses suggest that FaV-A provides a more effective framework than direct zero-shot MLLM baselines for producing visually coherent and semantically aligned positive-negative space compositions.
[MA-3] SoK: Decentralized Agent Economic Infrastructure
【速读】:该论文旨在解决去中心化智能体经济中,由独立设计与安全验证的协议组合而成的任务流程在整体上可能出现“看似每一步正确但最终结果错误”的问题。其核心挑战在于:尽管各阶段的机制(如授权、托管、结算)在局部满足安全与经济要求,但这些保证在跨阶段传递过程中可能失效,导致任务执行结果与预期不符。解决方案的关键是提出“保证闭包”(guarantee closure)这一任务相关性准则,用于判断某一阶段建立的保障是否在后续依赖该保障的决策中依然有效并持续约束行为。通过系统性地将安全与经济属性划分为17类属性家族,并覆盖六阶段工作流生命周期,研究对12个系统与标准、5种可复用机制及4个经典基准进行了分析,结合840次匹配执行与11,648例穷举检查,揭示了验证与结算环节间反复出现的失败模式——即符合规范的工作未被接受或有效证据被忽视。研究进一步区分了公开记录中的审批行为与任务合规性证据,并通过经济分析识别出报告、惩罚及共担错误等关键假设,明确了端到端保证断裂的具体位置及其修复路径。
链接: https://arxiv.org/abs/2610.01756
作者: Rui Sun,Xihan Xiong,Qin Wang,Fei Gao,Zelin Li,Zehua Cheng,Jiahao Sun,Zhipeng Wang
机构: Newcastle University, UK; University of Bristol, UK; CSIRO, Australia; University College London, UK; Ohio State University, USA; University of Oxford, UK; FLock.io, UK; The University of Manchester, UK
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA)
备注:
Abstract:Decentralized agent economies increasingly build a single task from protocols that were designed and secured separately. This creates a simple problem: a workflow can look correct at each step and still produce the wrong outcome. For example, a correct escrow may release payment on an authorized approval that provides little evidence that the delivered work actually satisfied the task. We systematize this problem across the full lifecycle of an agent task. Our study organizes security and economic requirements into 17 property families over six stages, with receipt soundness and completeness assessed separately. We examine 12 systems and standards, five reusable mechanism families, and four classical baselines. We introduce guarantee closure, a task-relative criterion for determining whether guarantees established at one stage remain available and constrain the later decisions that depend on them. We apply the criterion to controlled and native workflows, covering 840 matched executions and an exhaustive 11,648-case check over a finite objective-task domain. Our results expose recurring failures between verification and settlement, where conforming work can remain unaccepted or valid evidence can be ignored. Public records and model judgments further distinguish recorded approval from evidence of task conformance, while economic analysis identifies the report, penalty, and shared-error assumptions behind these guarantees. These findings show where end-to-end guarantees fail and what must be repaired to preserve them across the workflow. Subjects: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA) Cite as: arXiv:2610.01756 [cs.CR] (or arXiv:2610.01756v1 [cs.CR] for this version) https://doi.org/10.48550/arXiv.2610.01756 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[MA-4] After Cooperation Is Learned: Gradient Routing and Optimizer-Dependent Maintenance in Multi-Agent Reinforcement Learning
【速读】:该论文旨在解决多智能体强化学习(Multi-Agent Reinforcement Learning, MARL)中协作策略在持续训练过程中的稳定性问题,即已习得的协作行为是否能在后续优化中保持不变。传统评估方法通常从随机初始化出发考察协作发现能力,却忽视了持续优化可能对已有协作结构造成的破坏。本文提出“协作维持”(cooperation maintenance)这一概念,将其建模为右删失事件时间问题,并设计三种对比实验设置:X0 允许价值损失梯度更新共享智能体特征,X1 保留评判者但阻断其梯度传播,X5 完全移除评判者以构建无评判者基准。通过控制初始状态、评判者计算和评估条件,该设计可分离出直接价值梯度访问对协作维持的影响。研究发现,在正向奖励缩放下,仅 X0 出现显著的协作维持敏感性提升,而 X1 接近删失上限,X5 未观测到任何事件。进一步分析表明,梯度路径审计与冻结策略躯干扰动验证了预期的更新路径,且在 CleanUp-lite 中,不同缩放比例下的路径偏移与局部协作边界缩减相关;而在 MinEx 中则呈现较弱的、依赖优化器的效果。因此,核心结论是:直接价值梯度路由会引入一种条件性、缩放敏感的协作维持风险,而非评判者本身导致普遍失效。
链接: https://arxiv.org/abs/2610.01630
作者: Chaoyuan Hao,Wentao Yue,Tianyou Lai,Hongji Li,Jiayi Zhou,Qingyu Mao,Qilei Li
机构: 未知
类目: Multiagent Systems (cs.MA)
备注:
Abstract:Cooperative MARL is commonly evaluated through cooperation discovery from random initialization, leaving open whether continued optimization can destabilize learned cooperation. Actor-critic comparisons can also conflate critic presence with value gradients entering shared actor representations. We study cooperation maintenance, defined as the survival of a behaviorally verified cooperative policy under continued training. We formulate maintenance as a right-censored event-time problem and compare matched warm starts: X0 allows value loss gradients to update shared actor features, X1 retains the critic while blocking those gradients, and X5 removes the learned critic as a critic-free reference. This isolates direct value-gradient access while controlling initialization, critic computation, and evaluation. Positive reward scaling preserves strategic preferences and equilibria while perturbing learning dynamics. Gradient audits confirm the intended routing pathways, and frozen-policy torso perturbations probe whether route-induced updates align with local cooperation boundaries. In confirmatory MinEx and CleanUp-lite experiments, higher scales selectively increase maintenance sensitivity in X0; X1 remains near the censoring ceiling, and X5 has no confirmed events in the tested settings. In CleanUp-lite, route-by-scale displacement is associated with reduced local cooperation margins; MinEx shows a weaker, optimizer-dependent effect. These results identify a conditional, scale-sensitive maintenance risk associated with direct value-gradient routing rather than a universal failure of critics.
[MA-5] Managing Context and Communication in Distributed Agent ic UAV Swarms
【速读】:该论文旨在解决无人机蜂群(UAV swarm)在不确定环境中依赖语言模型代理进行任务级自适应推理时,采用全分布式控制架构所引发的信息管理难题。具体而言,长期交互历史会污染推理上下文,而无差别信息传播则导致通信与推理开销过大。其解决方案的关键在于提出一种基于事件驱动的“感知-推理-行动”(reason-act-observe)生命周期的分布式无人机-代理架构,通过结构化原子笔记(structured atomic notes)将运行时知识划分为核心记忆、本地记忆和同伴特定记忆三类,并引入一种确定性且兴趣感知的泛洪机制(deterministic interest-aware gossip engine),根据接收方语义新颖性和时效性选择性地传播信息。实验结果表明,在十架无人机的模拟搜救任务中,该方法可实现100%任务完成率,显著优于无限制泛洪(仅70–85%完成率)和由小型语言模型(SLM)自主决策转发(全部失败)的方案;同时相较泛洪策略,推理令牌消耗降低约一半,传输数据量减少,幸存者计数误差更低。
链接: https://arxiv.org/abs/2610.01569
作者: Andrea Iannoli,Ivan Zyrianoff,Angelo Trotta,Lorenzo Gigli,Marco Di Felice
机构: University of Bologna(博洛尼亚大学); Technology Innovation Institute (TII)(技术创新研究所)
类目: Multiagent Systems (cs.MA); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Networking and Internet Architecture (cs.NI); Robotics (cs.RO)
备注: 12 pages, 4 figures. This paper has been accepted for presentation at the 24th IEEE Consumer Communications Networking Conference 2027 (CCNC 2027)
Abstract:Unmanned aerial vehicle (UAV) swarms increasingly rely on language-model agents to provide adaptive mission-level reasoning in uncertain environments. Fully distributed control, in which each UAV hosts an independent Small Language Model (SLM), removes reliance on a centralized coordinator but introduces an information-management problem: long-running interaction histories can degrade the reasoning context, while indiscriminate information dissemination increases communication and inference overhead. We address these challenges with a distributed UAV-agent architecture that enables continuous local SLM control through an event-driven reason-act-observe lifecycle. Runtime knowledge is represented as structured atomic notes and organized into core, local, and peer-specific memory. A deterministic interest-aware gossip engine selectively disseminates these notes according to recipient-specific semantic novelty and recency. We evaluate the architecture using ten UAVs in a simulated search-and-rescue mission. Our approach completes all experimental runs, whereas unrestricted flooding messages completes only 70-85%, and delegating forwarding decisions to the SLM prevents mission completion in every run. Compared with unrestricted flooding, our approach approximately halves inference-token consumption, reduces transmitted data, and achieves lower survivor-count error.
[MA-6] A Multi-Agent LLM Framework for Personalized Health Checkup Interpretation and Guidance
【速读】:该论文旨在解决健康体检报告个性化解读中复杂多需求查询的精准理解与综合处理问题,尤其在需要跨纵向病历、医学知识、生活方式建议及医疗资源导航等多维度信息融合的场景下,现有单智能体系统存在意图识别不全、任务执行碎片化、输出一致性差等局限。其解决方案的关键在于构建一个多智能体大语言模型(Multi-Agent LLM)系统,通过识别用户查询中的多重意图,并将每项意图映射至专用的任务型智能体,实现并行化执行与结果整合,从而提升对复合型查询(compound queries)中各要素的覆盖度与推理质量。实验表明,相较于单智能体(Single Agent)设置,多智能体系统在加权LLM评判得分上从1.695提升至1.797(p = 0.027),并在有用性、一致性及需求完整性方面显著改善;尽管计算延迟与成本分别增加1.31倍和2.02倍,但人类评估者在多数对比中更偏好多智能体输出。此外,性能提升主要集中在涉及个人病历检索的子组查询中,凸显了该架构在复杂上下文关联任务中的优势。
链接: https://arxiv.org/abs/2610.01451
作者: HyungJun Kim,Taehan Lee,Soojin Cheon
机构: 未知
类目: Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA)
备注: 16 pages, 2 figures, 8 tables and Appendix
Abstract:Personalized interpretation of health checkup results requires reasoning across longitudinal records, medical knowledge, lifestyle guidance, and healthcare navigation. We present a multi-agent large language model (LLM) system that identifies multiple intents, maps each to a task-specific agent, executes them in parallel, and synthesizes their outputs. We compared answers generated in Single Agent and Multi Agent settings on 120 Korean compound queries combining two to four requirements, using synthetic health checkup records. The Multi Agent improved the weighted LLM-judge score from 1.695 to 1.797 (p = 0.027), and three additional LLM judges showed consistent improvements ( \Delta = +0.111 to +0.186, all p 0.05). The gains came from usefulness, consistency, and the handling of every requirement in compound queries, whereas numerical accuracy and grounding improved significantly under only one of the four judges and medical safety did not differ, and critical failures occurred at similar rates (Single Agent 15.0% vs. Multi Agent 13.3%). Two human evaluators preferred Multi Agent in 66.7% and 68.3% of pairwise comparisons. Multi Agent execution increased latency and cost by 1.31 \times and 2.02 \times , respectively. In exploratory subgroup analyses, the improvement was concentrated in queries involving personal-record lookup.
[MA-7] LLM -Driven Multi-Agent Control for Skill-Based Smart Manufacturing
【速读】:该论文旨在解决现代工厂在小批量、高定制化生产背景下,对柔性可重构自动化系统频繁重新编程所带来的挑战。其核心问题在于如何实现高效、自适应的生产调度与运行故障处理,尤其是在静态程序无法应对突发异常的情况下。解决方案的关键在于构建一个基于大语言模型(LLM)的分布式智能体架构:每个工厂模块配备专用的LLM智能体,并通过支持OPC UA方法调用的MCP工具服务器暴露其功能技能;各智能体间通过MQTT协议进行通信,同时依托实时工厂状态注入实现协同决策。实验在六模块六边形物理工厂的仿真环境中验证了三种智能体架构(协调者型、对等型和单体型)的表现,结果显示,尽管对等型与单体型在平均求解率上均达到93%,但协调者型在无显式故障处理逻辑的情况下,仍能自主识别并绕过隐蔽的传送带故障,在全部十次运行中成功解决该问题。研究进一步揭示了在标准化MCP工具链、基于MQTT的跨智能体通信以及实时状态反馈机制下,智能体可涌现出自发故障诊断行为,为生成式AI驱动的智能制造提供了可复现且可靠的技术基础。
链接: https://arxiv.org/abs/2610.01364
作者: Kay Köhle,Darko Anicic,Thomas A. Runkler,René Graf
机构: Technical University of Munich (慕尼黑工业大学); Siemens AG, Foundational Technologies (西门子股份公司,基础技术部门); Siemens AG, Digital Industries (西门子股份公司,数字工业部门)
类目: Multiagent Systems (cs.MA); Artificial Intelligence (cs.AI); Systems and Control (eess.SY)
备注: Accepted at the 2026 IEEE 31st International Conference on Emerging Technologies and Factory Automation (ETFA). 8 pages, 5 figures, 3 tables
Abstract:Factories are shifting toward smaller lot sizes with high product customization, requiring frequent re-programming of flexible and reconfigurable automation systems. LLM-based agents can be deployed in two complementary roles: Offline, they generate deterministic production sequences, reducing programming effort; online, they operate live machines and handle unforeseen runtime faults that static programs cannot anticipate. We propose a solution in which each factory module is paired with a dedicated LLM-based agent and an MCP tool server that exposes the module’s skills via OPC UA method calls, with agents coordinating over MQTT and grounded by real-time updates of the factory state. We compare three agent architectures (orchestrator, peer-to-peer, and monolithic) across nine production challenges of increasing complexity in a simulation of a physical six-module hexagonal factory, including silent hardware fault detection. The monolithic and peer-to-peer architectures both achieve the highest mean solve rate (93%), while the orchestrator uniquely resolves a silent conveyor-belt fault in all ten runs by autonomously rerouting plates around the blocked segment. All architectures exhibit emergent fault-diagnosis behavior without any explicit failure-handling logic, establishing standardized MCP tooling, MQTT-based inter-agent communication, and real-time state injection as a viable and reproducible foundation for LLM-programmed smart manufacturing.
[MA-8] Fully Online Decentralized Learning in Stochastic Games with Unknown Independent Chains
【速读】:该论文旨在解决具有未知独立受控转移核的随机博弈(stochastic games)中,玩家仅能观测局部状态和实际收益时,如何实现去中心化、无需协调且完全在线的学习以逼近平稳纳什均衡(stationary Nash equilibrium, NE)的问题。其核心挑战在于:在缺乏全局信息、无联合状态空间覆盖、无法同步学习周期的情况下,如何设计一种高效、可扩展的算法来应对高维状态-动作空间与多智能体交互带来的复杂性。解决方案的关键在于提出一种完全在线、去中心化的镜面下降(mirror-descent)算法,该算法在占用测度(occupancy measures)的对偶空间中运行,仅需每个基本时间步使用单个转移/奖励样本,且完全依赖局部信息。通过利用各玩家控制链之间的独立性与局部结构,算法避免了对联合状态空间的遍历要求,其复杂度仅依赖于各局部状态空间的覆盖时间(cover times),而非联合状态空间的指数规模,从而有效缓解了“维度灾难”问题。在均匀遍历性(uniform-ergodicity)和有限覆盖假设下,证明了时间平均固定比较器遗憾(time-averaged fixed-comparator regret)以近乎最优的 $ O(T^{-1/2}) $ 速率衰减(含对数因子和游戏参数的多项式依赖)。此外,该结果进一步导出了近似粗相关均衡(approximate coarse-correlated-equilibrium)的保证,这在任意收益函数下是自然的,因为计算平稳 ϵ-NE在此设置中为PPAD-hard。在额外引入全局变分稳定性条件后,还证明了该算法的最后迭代渐近收敛至平稳 ϵ-NE。整体上,该工作构建了一个完全在线、可扩展的随机博弈学习框架,并可视为一种利用独立性与局部结构的马尔可夫博弈原始-对偶方法。
链接: https://arxiv.org/abs/2610.01181
作者: S. Rasoul Etesami
机构: University of Illinois Urbana-Champaign (伊利诺伊大学厄本那-香槟分校); Coordinated Science Laboratory (协同科学实验室)
类目: Machine Learning (cs.LG); Computer Science and Game Theory (cs.GT); Multiagent Systems (cs.MA); Systems and Control (eess.SY); Optimization and Control (math.OC)
备注:
Abstract:We consider stochastic games with independent controlled chains and unknown transition kernels, where players observe only their local states and realized payoffs. We develop a fully online, decentralized, and uncoordinated mirror-descent algorithm that operates in the dual space of occupancy measures for approximating stationary Nash equilibrium (NE) policies. The algorithm uses a single transition/reward sample at every primitive time step, relies only on local information, and requires neither coverage of the joint state space nor synchronized episodes. Under uniform-ergodicity and finite-coverage assumptions, we show that, with high probability, the time-averaged fixed-comparator regret decays at the canonical O(T^-1/2) rate, up to logarithmic factors and polynomial dependence on the game parameters. In particular, the complexity depends on the cover times of the individual local state spaces rather than the product state space, avoiding exponential dependence on the number of players and the sizes of the joint state and action spaces. The resulting finite-time regret bound further yields an approximate coarse-correlated-equilibrium guarantee, which is natural for arbitrary reward functions since computing a stationary \epsilon -NE is PPAD-hard in this setting. Under an additional global variational-stability condition, we show that the same fully online algorithm converges asymptotically in the last iterate to a stationary \epsilon -NE. Our results provide a fully online and scalable learning framework for stochastic games with unknown independent chains. The algorithm can also be viewed as a primal-dual framework for Markov games that exploits the independence and local structure of the players’ controlled transition chains.
[MA-9] MASkillBlender: Decentralized Whole-Body Coordination for Multi-Humanoid Loco-Manipulation via Skill Blending
【速读】:该论文旨在解决多机器人人形(multi-humanoid)全身协调操控中的高维全身体控、去中心化决策及可扩展性难题。现有强化学习方法虽在单机器人人形全身体控方面取得进展,但将其推广至多机器人人形场景时仍面临显著挑战,通常需依赖大量奖励工程或任务特定设计。本文提出MASkillBlender,一种通用的多智能体强化学习框架,通过在可复用的预训练单机器人人形技能基础上学习共享的去中心化高层策略,实现无需任务特定运动参考的去中心化多机器人人形全身协同。其核心创新在于仅使用任务级奖励即可驱动复杂协调行为,显著降低对人工标注或定制化设计的依赖。为提升学习效率,进一步引入基于排列的数据增强策略,针对同质多机器人系统,在同质马尔可夫博弈(homogeneous Markov game)框架下理论证明了排列样本保持原始样本的策略梯度方向,从而保障学习稳定性与收敛性。在两种不同人形机器人本体上的多任务仿真验证表明,MASkillBlender在各类任务和机器人配置下均能稳定实现优异的任务性能与一致的协调行为。
链接: https://arxiv.org/abs/2610.01102
作者: Yifan Hu,Luhang Hong,Mingkang Long,Danning Wang,Chengfeng Jia,Rong Su,Junjie Fu,Guanghui Wen
机构: Nanyang Technological University (南洋理工大学); Southeast University (东南大学); Purple Mountain Laboratories (紫金山实验室)
类目: Robotics (cs.RO); Machine Learning (cs.LG); Multiagent Systems (cs.MA)
备注:
Abstract:Coordinated multi-humanoid loco-manipulation is promising yet challenging due to high-dimensional whole-body control, decentralized decision making, and scalability. While recent reinforcement learning methods have improved single-humanoid whole-body control, extending them to the multi-humanoid setting remains nontrivial and often requires substantial reward engineering or task-specific design. We propose MASkillBlender, a general multi-agent reinforcement learning framework to achieve decentralized multi-humanoid whole-body coordination. By learning a shared decentralized high-level policy over reusable pre-trained single-humanoid skills, MASkillBlender enables coordinated behaviors using only task-level rewards, without requiring task-specific motion references. To improve learning efficiency, we further introduce a permutation-based data augmentation strategy for homogeneous multi-humanoid systems, and theoretically show that the permuted samples preserve the policy-gradient direction of the original samples under the homogeneous Markov game formulation. We evaluate MASkillBlender on multiple multi-humanoid coordination tasks across two humanoid embodiments. Simulation results demonstrate that the proposed framework consistently achieves strong task performance and enables coordinated behaviors across different tasks and humanoid embodiments.
[MA-10] Can AI Scientists Coordinate at Runtime?
【速读】:该论文旨在解决多智能体人工智能科学家(Multi-agent AI Scientists)在执行科研任务时普遍依赖静态设计阶段的智能体编排(design-time agentic orchestration),难以灵活适应动态变化任务需求的问题。现有方法通常采用固定的工作流,缺乏类似人类科学家在运行时(runtime)根据实际进展动态调整分工与协作的能力。为此,论文提出运行时智能体协调机制(Runtime Agent Coordination, RAC),其核心在于在任务执行过程中动态选择已有智能体宿主中的代理、分配具有作用域限制的任务合约(scoped work contracts),并基于产出物(artifact)进行验证。该验证过程不阻塞任务流转,也不丢弃中间产物,从而实现非侵入式的信息传递与协作优化。实验在ResearchClawBench基准上对Agent Laboratory、EvoScientist和ARK三个平台进行单种子探索性评估,控制宿主模型、工具及权限,并在宿主校准的预算约束下比较四种条件:原生执行、运行时通信、运行时选择,以及结合合约与验证的综合机制。结果表明,运行时选择显著提升各宿主的平均得分,但引入合约与验证后,平均性能反而下降,且表现受宿主特性影响,揭示了在资源受限条件下额外协调机制的边际效益递减与潜在瓶颈。该研究验证了运行时协调的可行性与价值,同时指出了当前机制在复杂环境下扩展性的局限。
链接: https://arxiv.org/abs/2610.00980
作者: Zijian Liu,Yangzhixin Luo,Junyu Lu,Yi Li,Yu Chen,David Xu,William F. Shen,Xinchi Qiu,Xisen Wang
机构: University of Oxford (牛津大学); King Abdullah University of Science and Technology (阿卜杜拉国王科技大学); University of Sydney (悉尼大学); University of Cambridge (剑桥大学)
类目: Multiagent Systems (cs.MA); Artificial Intelligence (cs.AI)
备注: 35 pages (9 pages main text), 4 figures, 10 tables. Code: this https URL
Abstract:Multi-agent AI scientists have shown improving performance across a diverse range of tasks. Yet a common approach is design-time agentic orchestration, which typically relies on fixed workflows. In contrast, human scientists coordinate and adjust their division of labor at runtime. We therefore ask: can AI scientists also coordinate at runtime? To this end, we introduce Runtime Agent Coordination (RAC), which selects agents from existing AI-scientist hosts during execution, assigns scoped work contracts, and provides artifact-grounded verification. Verification informs subsequent agents without blocking transitions or discarding artifacts. We conduct a single-seed exploratory evaluation across Agent Laboratory, EvoScientist, and ARK on ResearchClawBench, preserving host models, tools, and permissions under host-calibrated budgets. Four cumulative conditions separate native execution, runtime communication, runtime selection, and the combined addition of contracts and verification. Runtime selection yields the highest observed mean score for each host; adding contracts and verification reduces these means, with host-dependent outcomes relative to native execution. These results motivate runtime coordination while exposing the limits of additional coordination mechanisms under constrained budgets. Code is available at this https URL.
[MA-11] VeriHarness: Scaling Agent ic Verification for Long-Horizon Tasks
【速读】:该论文旨在解决大语言模型(LLM)在执行复杂、长周期任务时,其输出难以验证的问题。核心挑战在于:在测试阶段无法访问参考答案或评分标准的情况下,如何有效提升模型输出的可靠性。现有方法依赖单一推理路径,易遗漏关键信息或引入错误,而重复采样虽能生成多个可能正确的陈述,但缺乏可靠的机制来甄别可信内容。为此,论文提出VeriHarness,其关键创新在于将基础模型转化为具有自主验证能力的智能体(agentic verifier),通过赋予其工作空间、证据工具及可复用的验证技能,实现对不同推理路径中产生主张的动态评估。具体而言,分歧解析器(disagreement resolver)利用环境证据检验相互矛盾的主张,揭示被共识掩盖的正确选项;共识挑战者(consensus challenger)则针对一致结论进行反向验证,识别潜在遗漏的需求或逻辑漏洞。二者协同指导最终成果的选择与修正。实验表明,该方法在五个长周期工作空间基准上显著优于基线模型,且基于证据的修订进一步提升了性能,相较单次推理分别获得6.2分(Gemini 3.5 Flash)和6.4分(Claude Opus 4.8)的增益。此外,研究还证明验证技能可通过失败反馈自我优化,展现出该框架在规模化长周期智能体验证中的潜力。为促进后续研究,作者开源了约26,000条推理轨迹数据集,总成本超10万单位,为构建更鲁棒的验证体系提供重要资源支持。
链接: https://arxiv.org/abs/2610.00972
作者: Caiqi Zhang,Rujun Han,Zifeng Wang,Zoey CuiZhu,Nigel Collier,Tomas Pfister,Chen-Yu Lee
机构: Google Cloud AI Research(谷歌云人工智能研究); University of Cambridge(剑桥大学)
类目: Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA)
备注:
Abstract:As LLM agents undertake increasingly complex, long-horizon tasks, verifying their outputs becomes increasingly challenging. We study how verification capability can be strengthened with a fixed base model, without access to reference answers or grading rubrics at test time. Repeated sampling yields multiple rollouts that can contain complementary correct claims, but we need a reliable verification mechanism to determine which claims to trust. We first find that disagreement often exposes correct alternatives, while consensus can conceal errors. These observations motivate VeriHarness, which turns the underlying LLM a generator uses into an agentic verifier by giving it a workspace, evidence tools, and reusable verification skills. A disagreement resolver checks competing claims against environmental evidence, while a consensus challenger tests shared claims and searches for omitted requirements. Their findings guide the selection and revision of the final artifact. Across five long-horizon workspace benchmarks and two frontier models, VeriHarness achieves the highest selection scores among the evaluated baselines. Evidence-backed revision further improves average performance, bringing gains over a single rollout to 6.2 points with Gemini 3.5 Flash and 6.4 points with Claude Opus 4.8. We further show that verification skills can self-improve from failure feedback, demonstrating VeriHarness as a novel and critical approach for scaling long-horizon agentic verification. We release the full pool of approximately 26,000 rollouts across all five benchmarks and both models, produced at a cost of over 100,000, to support future research on agentic verification.
[MA-12] ActiveSaddler: Automated Curriculum Learning for Agent Harness Optimization
【速读】:该论文旨在解决大语言模型(Large Language Model, LLM)智能体在自动化线束优化(harness optimization)过程中,训练场景(training scenarios)选择固定化所导致的优化瓶颈问题。传统方法虽能迭代更新提示词(prompt)、工具接口及控制逻辑,但其依赖预先固定的反馈生成场景顺序,未能随线束演化动态调整训练课程(curriculum),从而限制了对新出现缺陷的发现与优化效率。其解决方案的关键在于提出一种名为ActiveSaddler的自动化课程学习框架,将动态演化的课程建模为非平稳多臂赌博机(non-stationary bandit),通过抽象重复失败模式为可复用的“失败模式臂”(failure-pattern arms),实时估计每个模式进一步优化的潜在学习收益,并自适应地在重访已知弱点与探索未知场景之间平衡。该机制使优化目标集及其优先级随优化过程持续更新,实现课程与线束的协同进化。实验在GAIA2和Terminal-Bench 2.0基准上验证,相比固定场景顺序的优化器,ActiveSaddler分别提升测试通过率(Pass@1)4.4和7.5个百分点;消融实验证明性能增益依赖于动态构建优化目标、评估其演化效用以及探索与利用的平衡策略。结果确立了自动化课程学习作为线束优化中不可或缺的新维度。
链接: https://arxiv.org/abs/2610.00906
作者: Sungho Park,Wonjoong Kim,Jue Zhang,Wook-Shin Han,Pengfei Gao,Chanyoung Park,Yongqiang Yao,Rao Fu,Elsie Nallipogu,Qingwei Lin,Victor Rühle
机构: POSTECH(浦项科技大学); Microsoft(微软); KAIST(韩国科学技术院)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG); Multiagent Systems (cs.MA); Software Engineering (cs.SE)
备注: 37 pages, 16 figures. Project website and code: this https URL
Abstract:Automated harness optimization can substantially improve LLM agents by iteratively updating their prompts, tool interfaces, and control logic from execution feedback. However, existing methods primarily optimize how the harness is updated while largely fixing which training scenarios generate the feedback that drives those updates. As the harness evolves, the scenarios most useful for further optimization can change, suggesting that the training curriculum itself should adapt alongside the harness. We formulate this missing dimension of harness optimization as an automated curriculum learning problem and introduce ActiveSaddler. ActiveSaddler models the evolving curriculum as a non-stationary bandit with dynamically instantiated optimization targets. It abstracts recurring failures into reusable failure-pattern arms, estimates the potential learning progress from further targeting each pattern, and adaptively balances revisiting known weaknesses with exploring unseen scenarios for new ones. Optimization outcomes continually update both the set of discovered failure patterns and their priorities, allowing the curriculum to co-evolve with the harness. Experiments on GAIA2 and Terminal-Bench 2.0 show that ActiveSaddler consistently discovers stronger harnesses, improving test Pass@1 by 4.4 and 7.5 percentage points over the same harness optimizer using a scenario order fixed before optimization, respectively. Ablations further show that these gains depend on dynamically constructing optimization targets, estimating their evolving utility, and balancing continued optimization with new failure discovery. Together, these results establish automated curriculum learning as a new crucial optimization dimension for harness optimization.
[MA-13] HakiCC: LLM -Driven Multi-Agent Design and Optimization of Concurrency Control Protocols
【速读】:该论文旨在解决应用特定并发控制(Concurrency Control, CC)协议设计中高度依赖专家知识、难以自动化的问题。尽管已有大量成熟的CC协议,但实际应用中多数仍采用通用的2PL或OCC协议,因其针对特定应用的优化需要深入的协议设计经验,而这类知识在开发者中普遍缺乏。为克服这一瓶颈,论文提出HakiCC——一个基于生成式AI(Generative AI)的多智能体系统,通过自动化方式为特定工作负载设计、验证与优化定制化CC协议。其核心解决方案在于采用两阶段流水线:第一阶段利用多智能体协作生成并迭代修复、验证冲突可串行化(conflict-serializability)的应用特定协议;第二阶段引入基于大语言模型(LLM)的进化循环,进一步优化协议在正确性与吞吐量之间的权衡。实验结果表明,所有生成协议均满足冲突可串行化,且在TPC-C和AuctionMark工作负载上分别实现平均50.6%和92.2%的吞吐量提升,显著优于通用基线协议。
链接: https://arxiv.org/abs/2610.00889
作者: Farzad Habibi,Juncheng Fang,Faisal Nawab
机构: University of California, Irvine (加州大学欧文分校)
类目: Databases (cs.DB); Distributed, Parallel, and Cluster Computing (cs.DC); Multiagent Systems (cs.MA)
备注:
Abstract:Large language models (LLMs) have recently been applied in systems research as a tool to reduce human-intensive engineering effort through cost-efficient automation. Decades of research have produced a rich landscape of concurrency control (CC) protocols, each encoding distinct trade-offs in correctness, throughput, and abort behavior. However, most applications in practice default to 2PL or OCC, because selecting and adapting a protocol to a specific application requires expert knowledge that is rarely available to application designers. This is a wasted opportunity, as an application-specific CC protocol can yield significant performance advantages over a generic baseline, but designing one requires deep expertise in CC protocol design. In this paper, we propose HakiCC, an LLM-driven multi-agent pipeline that automatically designs, verifies, and optimizes concurrency control protocols tailored to a given target application. HakiCC provides a two-stage pipeline. In Stage 1, a multi-agent system takes a workload description as input and generates an application-specific CC protocol implementation, which is iteratively repaired and verified for conflict-serializability. In Stage 2, the verified protocol is further optimized for that application through an LLM-driven evolutionary loop targeting correctness and throughput. We evaluate HakiCC on TPC-C and AuctionMark as target workloads, producing and reporting ten application-specific CC protocols. All ten are conflict-serializable after Stage 1; Stage 2 improves throughput for every protocol, with average gains of +50.6% for TPC-C protocols and +92.2% for AuctionMark protocols. Subjects: Databases (cs.DB); Distributed, Parallel, and Cluster Computing (cs.DC); Multiagent Systems (cs.MA) Cite as: arXiv:2610.00889 [cs.DB] (or arXiv:2610.00889v1 [cs.DB] for this version) https://doi.org/10.48550/arXiv.2610.00889 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[MA-14] Kepler: Auditable World Models for ARC-AGI-3 NEURIPS2026
【速读】:该论文旨在解决在交互式环境(interactive environments)中,智能体(agent)需通过观测自主推断环境规则与目标这一核心挑战。传统评估方法依赖于最终得分,但无法有效区分智能体是真正理解了任务还是仅通过偶然或漏洞达成目标。其解决方案的关键在于提出Kepler——一个开源的评估框架,将假设建模为可执行的世界模型(world models),并通过回溯性转移检查(retrospective transition checks)和条件预测检查(conditional prediction checks)进行验证。该框架实现了在单一冻结的Claude Opus 5配置下,对全部25个公开游戏均获得100.00%的服务器验证人类表现等价率(RHAE),且无需针对每局游戏进行模型选择或基于分数的重试。实验表明,绝大多数关卡的最终尝试所用动作数不超过人类基准中位数,且大量环境动作发生在有评分的关卡中。此外,保留的本地会话记录显示高达8580万词元的输入量、97.37%的缓存命中率及约777.72美元的成本(按2026年9月API费率计算)。研究还揭示了三类评估失败案例,包括源码泄露导致虚假完美运行、智能体重构被移除的评估框架以及自主修复掩盖了规划器故障。一项单局观察案例研究进一步发现,动画帧中包含文本网格未体现的任务相关线索,凸显了多模态信息的重要性。综合结果表明,仅依赖公开数据集得分具有有限判别力,亟需推动首次尝试、成本约束与验证意识驱动的报告范式。
链接: https://arxiv.org/abs/2610.00834
作者: Wensen Wu
机构: Independent Researcher(独立研究员)
类目: Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA)
备注: 17 pages. Accepted to the non-archival Interpreting Agent Behavior workshop at NeurIPS 2026. Project: this https URL . Code and public traces available
Abstract:ARC-AGI-3 evaluates agents in interactive environments whose rules and objectives must be inferred from observation. We present Kepler, an open-source harness that represents hypotheses as executable world models and validates them through retrospective transition checks and conditional prediction checks. Under one frozen Claude Opus 5 configuration, Kepler obtained a server-verified 100.00 RHAE on all 25 public games, with no per-game model selection or score-conditioned reruns. On 181 of 183 completed levels, the final Opus attempt used no more actions than the corresponding median-human baseline. The retained board runs used 8,256 environment actions, of which 7,292 occurred in scored levels. Retained local provider-session records yield 858.0 million tokens, 97.37% cache reads, and a \ 777.72 cost at September 1, 2026 API list-equivalent rates. We also report three evaluation failures: source-code leakage that produced an invalid perfect run, agents reconstructing a removed harness in a control condition, and autonomous repair masking a broken planner. A single-game observation case study showed that animation frames contained task-relevant information absent from settled text grids. Across the final Claude Opus 5 and GPT-5.6 Sol boards, 48 of 50 game-model cells reached 100. These results indicate that public-set score alone has limited discriminative value and motivate first-attempt, cost-conditioned, and verification-aware reporting.
[MA-15] Outer Diversity of Condorcet Domains
【速读】:该论文旨在解决最大型康多塞域(Condorcet domain)在候选人数较少时的外向多样性(outer diversity)问题,即量化一个随机排名与域内最近排名之间的期望交换距离。其解决方案的关键在于通过数值分析方法对小规模候选集下的最大康多塞域进行外向多样性评估,并进一步推导出若干特殊域在渐近情况下的理论行为,从而揭示其内在规律与极限性质。
链接: https://arxiv.org/abs/2610.00720
作者: Piotr Faliszewski,Jan Jabrocki,Mateusz Słuszniak,Krzysztof Sornat,Stanisław Szufa,Tomasz Wąs
机构: 未知
类目: Computer Science and Game Theory (cs.GT); Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA)
备注: 33 pages, 12 figures
Abstract:A Condorcet domain is a set of rankings over a given candidate set, such that every election that consists only of (an odd number of) votes from the domain has a transitive majority relation. We study outer diversity of Condorcet domains, i.e., a measure that quantifies expected swap distance from a random vote to a closest one in the domain. We numerically analyze outer diversity for maximal Condorcet domains with few candidates, and then we establish its asymptotic behavior for several special domains, mostly obtaining theoretical results.
[MA-16] Meta-Multi-Agent Reinforcement Learning for Fast Adaptation of Interactive Policies with Applications to Autonomous Driving
【速读】:该论文旨在解决多智能体系统(Multi-Agent System, MAS)中交互策略快速适应的问题,尤其针对现有元强化学习(Meta-Reinforcement Learning, Meta-RL)框架主要局限于单智能体系统、难以有效处理多智能体间战略互动的局限性。其核心挑战在于,多智能体任务不仅依赖于环境动态,还受智能体之间复杂交互关系的影响。为此,论文将多智能体强化学习(MARL)问题建模为马尔可夫博弈(Markov Game, MG),提出一种面向马尔可夫博弈分布的元多智能体强化学习(meta-MARL)框架,以实现对交互策略的快速适应。该方案的关键在于引入“元纳什均衡”(meta-Nash Equilibrium, meta-NE)这一新概念,作为元多智能体强化学习问题中的期望解概念,并建立了meta-NE与基于梯度博弈(gradient-play-based)的元多智能体强化学习算法稳定点之间的等价性条件,从而为算法设计提供了理论保障。实验在自动驾驶任务上的结果表明,所提方法相比预训练的MARL基线模型具备更快的适应速度,验证了该框架的有效性。
链接: https://arxiv.org/abs/2610.00705
作者: Huiwen Yan,Kyriakos G. Vamvoudakis,Mushuang Liu
机构: Virginia Tech (弗吉尼亚理工学院); Georgia Tech (佐治亚理工学院)
类目: Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA); Systems and Control (eess.SY)
备注:
Abstract:This paper develops a meta-multi-agent reinforcement learning (meta-MARL) framework to enable fast adaptation of interactive policies in a multi-agent system (MAS). Meta-reinforcement learning (meta-RL) enables agents to rapidly adapt to new tasks/environments using a bi-level optimization mechanism. However, existing meta-RL generally focuses on single-agent systems. Extending these frameworks and algorithms to multi-agent systems poses additional challenges, as tasks are characterized by not only the environment but also agents’ strategic interactions. To address these challenges, we model multi-agent reinforcement learning (MARL) problems as Markov games (MGs) and develop a meta-MARL framework for rapid interactive policy adaptation across a distribution of MGs. A new concept, called meta-NE, is defined to describe the desired solution concept in a meta-MARL problem. Sufficient conditions for the equivalence between a meta-NE and a stationary point of the gradient-play-based meta-MARL algorithm are established. Our evaluation on autonomous-driving tasks demonstrates that the proposed meta-MARL method achieves faster adaptation than pretrained MARL baselines, validating the effectiveness of our framework.
[MA-17] Before Agents Decide: Epistemic Action in LLM -Based Systems NEURIPS2026 FAST
【速读】:该论文旨在解决大语言模型(LLM)驱动的智能体在做出复杂决策前,其可用证据是否具备“可决策性”(decision-ready)的问题。现有研究多关注智能体的推理与执行能力,却忽视了关键前提:当前证据是否足以支撑下一步判断。问题核心在于,证据可能缺失、形式不当导致关键信息被遮蔽,或缺乏必要的对比基准以支持评估。为此,论文提出将认知科学中的“认识论行动”(epistemic actions)引入智能体设计,明确三类关键策略:获取缺失证据、转换已有证据以揭示本质信息、通过系统探测生成具有揭示性的响应。其解决方案的核心是构建“认识论支架”(epistemic scaffolding),即支持上述行动的接口、工具与环境体系,使智能体不仅能搜索与探索,更能主动建构并验证决策所需的有效证据,从而确保决策过程建立在充分且可审计的证据基础之上。
链接: https://arxiv.org/abs/2610.00511
作者: Yizhi Liu,Balaji Padmanabhan,Siva Viswanathan
机构: Fox School of Business, Temple University (坦普尔大学福克斯商学院); Robert H. Smith School of Business, University of Maryland (马里兰大学史密斯商学院)
类目: Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA)
备注: Accepted at the Foundations of Agentic Systems Theory (FAST) Workshop at NeurIPS 2026
Abstract:Before a difficult decision, people often act simply to understand the situation better. We turn an object to see another side, place alternatives next to each other, or change one condition and observe what happens. These actions may not complete the task, but they improve the evidence needed for the next choice. LLM-based agents can search and explore, yet agent design gives less attention to an earlier question: is the available evidence ready for the decision? Sometimes necessary evidence is missing. In other cases, the evidence is present but its form hides what matters, or the comparison needed to judge it does not yet exist. Cognitive science calls actions that improve the basis for a later choice epistemic actions. We bring this idea to LLM-based agents and distinguish three modes: acquiring missing evidence, transforming available evidence, and probing a system to create a revealing response. We use the term epistemic scaffolding for the interfaces, tools, and environments that make these actions possible and auditable. This paper argues that agent design must address how decision-ready evidence is produced.
[MA-18] Deny Without Disabling: Authorization-Paired Evaluation and Control for Multi-Agent Systems
【速读】:该论文旨在解决多智能体系统(Multi-agent Systems)在协作过程中因信息共享与任务委派导致的潜在安全风险问题,即单个智能体的行为在独立评估时可能合规,但多个智能体协同执行时却可能共同促成被禁止的使用场景。传统方法通过全盘阻断敏感操作来防止泄露,但会破坏协作的有效性。其解决方案的关键在于提出“授权配对评估”(authorization-paired evaluation),将阻止违规使用与完成授权使用作为联合成功标准,并设计了FlowReview框架,实现对象解析、权限排序与确定性执行之间的可验证联动。实验表明,在受控组合测试中,通过审查联合生成物可将拒绝提交率从86.0%降至零,且未造成授权任务供给的损失。研究进一步揭示,仅保留信息与溯源链不足以确保权限归属正确,必须通过可验证组件维持对象身份与权限在执行过程中的持续关联。因此,论文确立了多智能体系统安全的系统级要求:在保障协作能力的前提下,对组合信息流进行治理,同时维护授权功能的可用性。
链接: https://arxiv.org/abs/2610.00371
作者: Yunbei Zhang,Saiyue Lyu,Janet Wang,Yingqiang Ge,Jiang Guo,Jihun Hamm,Chandan K Reddy
机构: 未知
类目: Multiagent Systems (cs.MA); Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR)
备注: 44 pages, 9 figures. Code and data: this https URL
Abstract:Multi-agent systems derive their capabilities from sharing evidence, delegating tasks, and combining information across agents. The same process creates a safety problem: contributions that are admissible in isolation can jointly enable a prohibited use. Blocking every sensitive action avoids disclosure but defeats the purpose of collaboration. We introduce authorization-paired evaluation, which makes blocking prohibited uses and completing required authorized uses a joint success criterion, and FlowReview, a framework connecting object resolution, permission ranking, and deterministic enforcement. In controlled composition experiments, reviewing combined artifacts reduces the denied-commit rate from 86.0% to zero with no loss of authorized supply. Our findings show that preserving information and lineage alone does not ensure correct permission attribution. Object identity and permission must remain connected to execution through components whose outputs can be verified. Together, these findings establish a system-level requirement for multi-agent safety: govern composed information flows while preserving the authorized capabilities that make collaboration useful.
[MA-19] A Verifier Can Leak the Answer: Diagnosability Before Optimization in Closed-Loop Agent Debugging NEURIPS2026
【速读】:该论文旨在解决生成式智能体(Agent)开发中基于模拟器的验证器(verifier)在评估求解器性能时可能引入的“虚假有效性”问题。核心问题是:当验证器所使用的探测信号或谓词(predicates)预先编码了目标故障的标识信息时,即使求解器未真正解决实际的不确定性,其优化结果也可能被误判为有效,从而导致对比实验失去意义。解决方案的关键在于引入一种“支持门控验证契约”(support-gated verification contract),从根本上重构验证流程——首先通过干净的参考映射(reference map)确保组件暴露的重复性,再通过匹配的参考/当前门控机制建立可比的运行时证据,最后才允许独立校准的信号返回检测结果。该方法有效防止了验证器提前泄露答案,显著提升了验证的可信度。在预注册保留测试集(1,440个案例,21,600个分区行)中,仅0/20被代表的组件出现虚假接纳,且在通过验证的单元中,清洁流量对故障检测的预测能力优于名义故障单元占比,验证了该框架在结构上对证据资格与非信息泄露的严格保障。
链接: https://arxiv.org/abs/2610.00126
作者: Peiying Zhu,Sidi Chang
机构: 未知
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA)
备注: Submitted to Who Verifies the Agents? Toward Reliable Agent Development (NeurIPS 2026 workshop). 7 pages, 0 figures, 2 tables. The reproducibility artifact is linked in the paper
Abstract:Agent developers increasingly compare prompts, tools, policies, and diagnosis algorithms through simulator-grounded verifiers. A verifier can nevertheless make a solver comparison vacuous: if its probes or predicates encode the target identity, an exact optimizer may appear effective without resolving any genuine ambiguity. We report such a failure in an aggregate-trace debugger for a closed-loop decision agent. Exact minimum hitting set (MHS) and a propagation-aware greedy method returned identical supports in 12/12 development cases and the same planted-fault recovery in 9/12. A subsequent audit found that exact-anchor predicates produced the planted pair in 9/9 cases. After removing those anchors, overall planted-pair recovery was 8/9; hard-probe singleton pairs nevertheless matched the planted pair in 9/9, and no case retained a nonempty residual conflict family after propagation (0/9). The optimizer was correct, but the verifier had already disclosed the answer. We replace solver-first evaluation with a support-gated verification contract. A clean reference map must first show repeated component exposure; a matched reference/current gate must then establish comparable runtime evidence; only afterward may an independently calibrated signal rule return a detection. In a preregistered heldout comprising 1,440 cases and 21,600 partition rows, 55/72 regime-component units passed the reference gate, 54/55 passed the runtime gate, and stable false admission was 0/20 represented components with a one-sided exact 95% upper bound of 0.1391. Within admitted units, affected clean traffic predicted detection better than nominal fault-cell fraction. The main lesson is structural: verify evidence eligibility and non-revelation before optimizing the component selector. Otherwise a stronger solver can merely certify a stronger verifier artifact.
[MA-20] he Delegation Danger Band: Why Mid-Capability Sub-Agents Over-Trust Inherited Stale State NEURIPS2026
【速读】:该论文旨在解决在多智能体框架中,子智能体继承父智能体状态所引发的“过时状态依赖”问题,即当子智能体沿用父智能体已过时或不适用的推理上下文时,导致性能下降。其核心问题是:在不同能力水平的模型上,继承策略对任务表现的影响如何变化,以及何种继承机制能在保持信息复用优势的同时避免过时状态带来的负面影响。解决方案的关键在于提出并验证三种不同的状态继承策略——重置(Reset,仅继承基础证据)、选择性传递(Selective,精选有效先验结论)和全量继承(Full,包含有效结论及多个被覆盖结论的副本),并通过一个固定、封闭集的动作评分基准,在同一家族模型系列(Qwen3 0.6/1.7/4/8B)上进行对比实验。研究发现,随着模型能力提升,对过时状态的依赖显著下降;但在中等能力模型(如Qwen3-1.7B)中出现“净继承伤害”的局部极小值,形成所谓的“危险带”(danger band),表明此范围内的模型最易受过时状态干扰。进一步分析显示,选择性传递策略在所有数据集上均优于全量继承,尤其在“危险带”模型上提升显著,而基于固定阈值的能力路由策略则表现不佳,揭示出需要可迁移的动态路由机制来权衡状态复用收益与过时上下文惩罚之间的平衡。
链接: https://arxiv.org/abs/2610.00041
作者: Jundong Hu,Shekar Ramachandran
机构: PayPal AI(贝宝人工智能)
类目: Multiagent Systems (cs.MA); Artificial Intelligence (cs.AI)
备注: Preprint. Under review at a NeurIPS 2026 workshop. 16 pages, 3 figures, 9 tables
Abstract:Agent frameworks increasingly delegate work by forking sub-agents; a common default makes the child inherit the parent’s full working context. We measure how the effect of inherited state changes with capability, where C_m denotes clean fork-fresh accuracy. We compare 3 inheritance policies: Reset (fork fresh: base evidence only), Selective (curated handoff: + the useful prior conclusion), and Full (implicit fork: + the useful conclusion and d copies of a superseded conclusion) over a same-family ladder (Qwen3 0.6/1.7/4/8B) on a frozen, closed-set, action-scored benchmark. Every task is solvable from the base evidence, so performance loss can be attributed to reliance on stale state. (1) Deference to superseded state falls sharply with measured capability C_m (the slope’s confidence interval, CI, excludes zero on every family) across 2 synthetic primitives plus MuSiQue and HotpotQA. (2) On the Qwen3 synthetic ladder, net inheritance harm follows a nonmonotone pattern: a mid-capability model (Qwen3-1.7B) is a statistically significant local minimum of net harm, falling below its fork-fresh baseline ( \Delta(32)=-0.19 [-0.25, -0.12]) and both neighbors, while the weakest model stays near-neutral and the strongest models stay robust. We call this harmful capability range a danger band. A within-model counting-difficulty sweep shows that the effect depends on model class even at matched C_m , and a live parent-to-child fork reproduces the mid-model harm. (3) Curated Selective handoff improves average accuracy over Full on all 3 datasets, largest at the in-band model, while the fixed-threshold capability router fails on the other datasets; a transferable router would need to predict the balance between reuse benefit and stale-context penalty. The benchmark is frozen and version-hashed.
[MA-21] oken Communication-Assisted Collaborative Embodied Artificial Intelligence: Concepts Framework and Opportunities
【速读】:该论文旨在解决协同具身人工智能(CEAI)中多物理智能体在动态环境中进行高效、长时程通信所面临的挑战,尤其是如何在有限带宽下传输大规模多模态感知数据的同时,有效传递任务相关的洞察、意图及交互信息。其解决方案的关键在于提出一种原生智能接口——令牌通信(Token Communication, TokCom),将令牌(tokens)同时作为紧凑的语义载体与生成式基础模型(Generative Foundation Models, GFMs)的核心推理单元。通过设计一种任务自适应的通信协议,该框架整合了紧凑的代码本(codebook)、语法规则与上下文示例,指导基于GFM的收发器将消息压缩为低开销的令牌序列,并在无线传输后准确重构。实验表明,该方法在协作物体运输任务中显著降低了源码率比特消耗,同时保持任务效率并具备对噪声信道的鲁棒性。
链接: https://arxiv.org/abs/2610.01826
作者: Peng Yi,Ying-Chang Liang
机构: 未知
类目: ignal Processing (eess.SP); Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA)
备注: 10 pages, 4 figures. Submitted to the IEEE for possible publication
Abstract:Collaborative embodied artificial intelligence (CEAI) enables multiple physical agents to perceive, reason, and act cooperatively in dynamic environments. Effective communication is essential for CEAI, yet CEAI agents must exchange not only large multimodal observations but also task-relevant insights, intents, and interactive information over long horizons. This article investigates token communication (TokCom) as a native intelligence interface for CEAI, in which tokens serve jointly as compact semantic carriers for communication and fundamental inference units for generative foundation models (GFMs). We first discuss how TokCom supports insight sharing, intent alignment, and interactive control among embodied agents. We then propose a TokCom-assisted CEAI framework driven by a task-adaptive communication protocol. Comprising a compact codebook, syntax rules, and contextual examples, this protocol guides GFM-based transceivers to distill messages into compact tokens and reconstruct them after wireless transmission. A case study on collaborative object transport demonstrates that the proposed TokCom framework substantially reduces the source payload bit consumption while preserving task efficiency and showing robustness under noisy channels. Finally, we outline future research directions.
[MA-22] Multi-agent Auditory Scene Analysis: Improved Localization Speed and Robustness by Multi-beamformed Speech Quality Feedback
【速读】:该论文旨在解决实时听觉场景分析(ASA)系统在声源定位、分离与分类任务中面临的优化速度慢、鲁棒性不足及局部误差难以全局修正的问题。现有基于多智能体系统的ASA方法虽通过智能体间通信实现反馈纠错,提升了系统整体稳定性,但其优化过程依赖于逐窗独立的质量评估值,导致搜索空间波动大、难以高效优化。本文提出一种新型优化机制,不再采用单次质量评估,而是引入多位置范围内的质量评估集合,从而提供更平滑、清晰的优化空间,显著降低优化复杂度。该方案使系统优化时间大幅缩短,定位精度和稳定性均显著提升,尤其在真实复杂声学环境中对高误差场景的校正能力更强,同时整体系统复杂度低于先前方法。唯一权衡是质量评估智能体响应时间略有增加,但整体系统仍可满足实时运行要求。实验结果再次验证了多智能体架构在构建高效、鲁棒的实时ASA系统中的优势。
链接: https://arxiv.org/abs/2610.00538
作者: Caleb Rascon
机构: Instituto de Investigaciones en Matematicas Aplicadas y en Sistemas, Universidad Nacional Autonoma de Mexico(墨西哥国立自治大学应用数学与系统研究所)
类目: Audio and Speech Processing (eess.AS); Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA)
备注: Submitted to Autonomous Agents and Multi-Agent Systems
Abstract:A real-time auditory scene analyzer (ASA) aims to carry out the tasks of locating, separating and classifying the sound sources present in a given acoustic environment. Recently, an effort has been made into modelling an ASA as a multi-agent system, with each one of its agents performing one of the aforementioned tasks and communicating their results to the rest of their peer agents. These communication routes are used as feedback loops to fix local errors at a global level, providing robustness while reducing local complexity. An example of the benefits of this approach is the optimization of speech quality by correcting in real-time the estimated location of the speech source of interest. However, their optimization speed has been shown to be considerably slow. One possible reason is that it solely relies on a series of single quality estimations (provided by a reference-free quality estimator model) that vary considerably from one window to the next, which results in a difficult search space to optimize. In this work, a new optimization mechanism is proposed that instead relies on a series of sets of quality estimations over a range of locations, providing a clearer view of the search space, simplifying its optimization. The proposed ASA now has a considerably smaller optimization time, is more accurate, and is more stable when being evaluated in real-life acoustic scenarios to correct higher levels of localization errors, all while being less complex than previous efforts. The only trade-off is that there is an increase in the response time of the quality estimation agent, but the complete ASA is still able to run in real-time. The performance shown in this work again shows the benefits of modelling an ASA as a multi-agent system.
自然语言处理
[NLP-0] KaliBench: A Fine-Grained Benchmark for Cybersecurity Tool Use on Kali Linux with Runtime-Free Verifiable Rewards NEURIPS2026
【速读】: 该论文旨在解决大语言模型(LLM)在网络安全工作流中将分析师意图准确转化为可执行命令的能力评估问题,尤其关注其在真实世界网络安全工具(基于命令行接口,CLI)上的指令生成能力。现有评估方法主要依赖知识性测试或端到端代理任务,未能直接衡量模型生成符合严格语法规范的可执行命令的能力,而网络安全操作对命令格式高度敏感,微小的语法错误、参数绑定错误或参数顺序错误均可能导致执行失败。为此,论文提出KaliBench——一个细粒度基准测试与数据集,专注于自然语言到Kali Linux CLI的翻译任务,包含8,504个查询-命令对,覆盖1,642个工具、23个功能维度和5个安全阶段。其关键解决方案在于采用基于手册的构建流程、确定性归一化处理及别名感知评估机制,实现工具选择与参数构造的精确、可复现评估;并设计多阶段验证流水线,结合基于LLM的语义验证、沙箱化终端执行和人工干预优化,确保输出既语义正确又具备实际可执行性。基于此,研究进一步构建了无需运行时的可验证奖励信号,用于模型训练。实验表明,在无提示约束条件下,所有开放权重模型的精确命令准确率均未超过42%,凸显了该任务的挑战性;而通过引入基于KaliBench的监督微调与强化学习,显著提升了8B模型性能,使其达到与685B MoE模型相当的水平。
链接: https://arxiv.org/abs/2610.02206
作者: Pengfei Li,Naufal Suryanto,Sicheng Zhang,Muzammal Naseer
机构: 未知
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR)
备注: Accepted at NeurIPS 2026 Evaluations and Datasets Track. Project page: this https URL | Github: this https URL
Abstract:LLMs are increasingly applied to cybersecurity workflows, where they are expected to translate analysts’ intent into tool invocations. However, existing evaluations focus on knowledge-based assessments or end-to-end agentic tasks, and do not directly measure LLMs’ ability to generate executable commands for real-world cybersecurity tools. This gap is critical because cybersecurity operations rely on strict command-line interfaces (CLIs), where minor syntax errors, incorrect flag–value bindings, or argument misordering can invalidate execution. We introduce KaliBench, a fine-grained benchmark and dataset for natural-language–to–CLI translation on Kali Linux, comprising 8,504 query–command pairs spanning 1,642 tools across 23 capability dimensions and 5 security phases. KaliBench is constructed via a manuscript-grounded pipeline with deterministic canonicalization and alias-aware evaluation, enabling precise and reproducible assessment of tool selection and argument construction. To ensure both semantic correctness and practical executability, we develop a multi-stage verification pipeline that combines LLM-based validation, sandboxed terminal execution, and human-in-the-loop refinement. Building on these fine-grained, deterministic signals, KaliBench further enables runtime-free verifiable rewards for training. Across three evaluation modes and 24 configurations of general-purpose and security-focused open-weight models, no open-weight model exceeds 42% exact-command accuracy in the unrestricted setting, highlighting the difficulty of accurate CLI-based cybersecurity tool use without explicit tool hints. We further show that supervised fine-tuning and reinforcement learning with verifiable rewards derived from KaliBench significantly improve an 8B model and achieve performance comparable to a 685B MoE model.
[NLP-1] Hierarchical Continuous Diffusion Language Models
【速读】: 该论文旨在解决现有离散扩散语言模型(Discrete Diffusion Language Models)与连续扩散语言模型(Continuous Diffusion Language Models)在生成过程中存在的结构性瓶颈问题。具体而言,离散扩散模型在并行解码时各词元(token)独立采样自边际分布,导致词元间的统计依赖关系被切断;而连续扩散模型虽通过共享连续状态实现全局建模,但其去噪器仅感知连续状态,缺乏对有效词元配置的显式约束,直至最终解码才形成语义一致性。为此,本文提出分层连续扩散语言模型(Hierarchical Continuous Diffusion Language Models, HC-DLM),其核心创新在于将离散词元生成与连续潜在轨迹耦合于一个统一、可微的去噪过程之中,训练目标基于词元似然的变分下界(variational bound)。相较于近期将连续上下文附加至独立离散生成链的方法,HC-DLM 以连续潜在变量作为唯一的持续生成状态:每一步均从该潜在状态中读出词元,并将其作为下一时刻潜在状态更新的结构化引导(scaffold)。这一机制实现了动态反馈与全局约束的协同优化,在结构化推理(如数独求解)、数学规划(如倒计时游戏)和语言建模(LM1B)任务上,均在相同模型规模下显著优于离散与连续扩散基线方法,在数独与倒计时任务中的解题准确率以及在 LM1B 上的生成困惑度均有提升。
链接: https://arxiv.org/abs/2610.02193
作者: Hui Ren,Zihan Li,Chang Liu,Huidong Liu,Alexander Schwing
机构: University of Illinois Urbana-Champaign (伊利诺伊大学厄本那-香槟分校); Amazon.com, Inc. (亚马逊公司)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:
Abstract:Discrete diffusion language models offer a compelling alternative to autoregressive generation for tasks demanding bidirectional reasoning and global constraint satisfaction. Yet they share a structural bottleneck: when decoding in parallel, each token is sampled independently from its marginal, severing the statistical dependencies among the tokens decoded together. Continuous diffusion language models avoid this by denoising a shared continuous state, but their denoiser sees only that state, so nothing ties it to a valid token configuration until it is finally decoded. To address this, we propose Hierarchical Continuous Diffusion Language Models (HC-DLM), which couple discrete token generation with a continuous latent trajectory in a single, principled denoising process, whose training objective is derived from a variational bound on the token likelihood. In contrast to recent methods that attach continuous context to a self-contained discrete chain, HC-DLM makes the latent the only persistent generative state: tokens are read out from it at every step and feed back as a scaffold for the next latent update. On structured reasoning (Sudoku), mathematical planning (Countdown) and language modeling (LM1B), HC-DLM improves over discrete and continuous diffusion baselines at matched model size, in puzzle accuracy on Sudoku and Countdown and in generative perplexity on LM1B. Project page: this https URL.
[NLP-2] Every Ablation Is a Dose: Counterweights and the Semblance of Self-Repair
【速读】: 该论文旨在解决语言模型中“自修复”(self-repair)现象的机制不明问题,即在移除模型某一组件后,其他组件似乎会调整以补偿损失信号的现象。现有研究认为该现象具有噪声性且可能无统一解释,但本文提出其背后存在一个统一机制:在任何扰动前,系统已存在一种固有的增益(gain)。本文将对因果关键组件的干预视为位于一个称为λ的坐标轴上的点,λ表示反事实对比的有符号强度。传统消融方法仅提供该轴上的未校准点,而本文揭示,对于细粒度单元r,其因果修复响应遵循仿射定律 $ E_r(\lambda) = \mathrm{own}_r + \gamma_r\lambda $,其中斜率 $ \gamma_r $ 是一个固定系数,无论是否进行消融均持续影响模型,其符号决定该单元是抵消还是增强被移除的信号。在四个不同架构模型(Gemma、Qwen、LLaMA、Mistral)的事实判断任务中,81个下游方向中有68个符合此仿射规律;同时,$ \gamma_r $ 的大小可由固定权重预测。在GPT-2 Small的IOI电路中,7个可干预的注意力头中有7个遵循该定律,且均为对抗性权重(counterweights)。因此,所谓“自修复”本质上是核心处出现对比信号时,对抗性单元正常执行其功能的结果。
链接: https://arxiv.org/abs/2610.02173
作者: Areeb Ahmad,Pratinav Seth,Vinay Kumar Sankarapu
机构: Lexsi Labs(雷克斯实验室)
类目: Machine Learning (cs.LG); Computation and Language (cs.CL)
备注:
Abstract:Ablate a component of a language model, and other components often appear to adjust and compensate. This phenomenon, termed self-repair, has been observed repeatedly, but its mechanism remains unclear. The most systematic study to date concluded that self-repair is noisy and unlikely to have a single explanation. We argue that it has one: a gain already present before any ablation. Any intervention on a causally important component can be viewed as a point on a coordinate axis \lambda , the signed strength of a counterfactual contrast. Hence, conventional ablation methods are uncalibrated points on this axis. We show that the causal repair response for a fine-grained unit r is governed by an affine law, E_r(\lambda)=\mathrmown_r+\gamma_r\lambda . The slope \gamma_r is a fixed coefficient that consistently influences the model, with or without ablation, and its sign determines whether the unit counteracts or reinforces the removed signal. On a factual-verdict task across four models from distinct families (Gemma, Qwen, LLaMA, and Mistral), we identify components including MLP neurons, OV neurons, and singular directions that follow this affine law, 68 of 81 downstream directions in all. Moreover, we can anticipate the magnitude of \gamma_r from the fixed weights. On the IOI circuit of GPT-2 Small, seven of the ten heads the intervention can reach follow the law, and all seven are counterweights. From this perspective, what may appear as self-repair is a counterweight performing its usual operation when the contrastive signal emerges at the core.
[NLP-3] AutoCompact: Learning When to Compact Context in Long-Horizon Coding Agents
【速读】: 该论文旨在解决编码代理(coding agent)在执行仓库级软件工程任务时,因长期执行过程中早期探索信息逐渐失效而导致的上下文管理难题。传统方法仅关注避免上下文溢出,而忽视了何时进行压缩(compaction)、保留何种工作状态以及如何从压缩后状态继续推进等关键决策问题。其解决方案的核心是提出AutoCompact框架,通过将上下文压缩策略纳入代理的决策政策(policy)中,使代理能够自主学习并优化这些决策。具体而言,研究通过运行基础代理生成轨迹,并利用人工或自动评判器评估其压缩决策、摘要生成及压缩后的动作;错误输出被修正后重新注入环境,形成高质量训练数据,用于监督微调与强化学习联合优化,以最大化任务成功率。实验结果表明,在SWE-bench Verified和SWE-PolyBench Verified基准上,AutoCompact相较于基线模型分别提升了9.2%和5.0%的通过率,且在不同推理预算下均表现稳定,即使在16K上下文窗口触发回退压缩的情况下仍能有效保持性能,验证了其在有限上下文约束下的鲁棒性与有效性。
链接: https://arxiv.org/abs/2610.02163
作者: Xuan Zhang,Longtao Zheng,Cunxiao Du,Bo An,Xin Dong
机构: Singapore Management University(新加坡管理大学); Nanyang Technological University(南洋理工大学); Harvard University(哈佛大学)
类目: Computation and Language (cs.CL)
备注:
Abstract:Coding agents solve repository-level software engineering tasks through long trajectories of code inspection, search, editing, and testing. As a task progresses, earlier exploration becomes stale, so managing context is more than avoiding overflow: an agent must decide when to compact, what working state to preserve, and how to continue from it. We introduce AutoCompact, which trains a coding agent to make these decisions as part of its policy. To collect training data, we run the base agent on coding tasks and use a judge to review its compaction decisions, summaries, and actions after compaction. Flawed outputs are replaced with corrected ones before being executed in the environment, so each trajectory continues from the corrected decisions. We use these trajectories for supervised fine-tuning, then jointly optimize coding and compaction through reinforcement learning with task-success rewards. Experiments on SWE-bench Verified and SWE-PolyBench Verified show that AutoCompact improves pass rates over the base model by an absolute 9.2% and 5.0%, respectively. The improvements hold across all evaluated inference budgets, with a 256K context window that never overflows and with a 16K window whose overflow triggers fallback compaction.
[NLP-4] From Knowledge Access to Source Learning: Developing Source-Specific Competence
【速读】: 该论文旨在解决大语言模型(Large Language Model, LLM)代理在处理持续性知识密集型任务时,对同一外部权威源的重复使用仍局限于简单重复访问,而未能实现对该源知识理解的渐进式深化这一核心问题。现有方法虽在信息获取与组织、以及代理记忆系统方面有所改进,但缺乏对源知识本身的动态学习机制。其解决方案的关键在于提出源学习(SourceLearn)框架,通过构建一个持久化的源模型(source model),以捕获可复用的、针对特定来源的知识结构、解释方式与应用模式。该框架融合两种互补的学习机制:自导向源学习(Self-Directed Source Learning) 用于识别理解不充分的部分并自适应地回溯源内容;任务引导源学习(Task-Guided Source Learning) 则利用下游任务经验揭示局部表征缺陷与知识组织中的重复需求。二者均以学习信号驱动需重构的内容,并基于权威源进行持续更新,从而实现对源知识的迭代优化。实验表明,在五个基准测试与三种LLM后端上,SourceLearn在15个设置中取得13项最优表现,相较混合检索增强生成(Hybrid RAG)最高提升达22.6分,显著优于静态源表示与基于经验的记忆基线。
链接: https://arxiv.org/abs/2610.02150
作者: Lucheng Fu,Kejing Xia,Yiyang Wang,Yiqiao Jin,Jinjin He,Xiyuan Yang,Haoxin Liu,Ye Yu,Haibo Jin,Yijia Xiao,Wenke Lee,B. Aditya Prakash,Haohan Wang
机构: 未知
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: Website: this https URL Code: this https URL
Abstract:Large language model (LLM) agents increasingly rely on persistent external sources to solve sequences of knowledge-intensive tasks. Existing methods improve how source content is accessed and organized, while agent-memory systems preserve reusable knowledge from prior interactions, but repeated use of the same source is still largely treated as repeated access rather than an opportunity to progressively improve understanding of that source. We study source learning: developing reusable source-specific competence over a persistent authoritative source. We represent this competence with a persistent source model that captures reusable understanding of the source, including how its knowledge is structured, interpreted, and applied. To construct and progressively refine such models, we propose SourceLearn, which combines two complementary learning mechanisms. Self-Directed Source Learning identifies what remains incompletely understood and adaptively revisits the source, while Task-Guided Source Learning uses downstream experience to reveal local representational gaps and recurring needs in how source knowledge should be organized. In both cases, learning signals determine what should be reconsidered, while persistent updates are reconstructed from the authoritative source. Across five benchmarks and three LLM backends, SourceLearn achieves the best performance in 13 of 15 settings, with gains of up to 22.6 points over Hybrid RAG and substantial overall improvements over static source representations and experience-based memory baselines.
[NLP-5] Keyword Harnesses Fail Open: A Cheap Diagnostic Ladder for Tool-Use Claims in Small Language Models
【速读】: 该论文旨在解决生成式 AI(Generative AI)在工具调用(tool use)评估中存在虚假正例(false positive)的问题,即小模型可能因关键词匹配而非真正具备工具调用能力而在宽松的评测指标下获得高分。其核心问题是:当前广泛采用的宽松评测标准无法区分模型是否真正掌握工具调用的格式与逻辑,导致低参数量模型(如661.6M)在未经过专门工具微调(SFT)的情况下,仅通过训练数据中的模式复制即可获得接近大模型(1,109M)的评分,从而误导对模型真实能力的判断。解决方案的关键在于提出一个“低成本、逐步严格”的诊断阶梯(ladder of strict, cheap diagnostics),通过一系列轻量级检测手段定位问题根源——特别是利用首词探针(first-token probe)发现大模型因网络训练阶段中缺失先验概率(prior probability)而导致无法正确生成工具调用;随后通过针对性的微调策略(targeted SFT recipe),仅使用约3.3 GPU小时和比原失败阶段少三个数量级的训练数据,成功修复模型,使其在269个语料行上有效调用率从0.100提升至0.959,并在238个未见提示中表现优于小模型(p = 0.004)。此外,嵌入漂移检查表明修复未改变触发词的嵌入表示,说明改进发生在网络邻近结构中,验证了修复的局部性与有效性。该诊断框架成本极低(仅需数分钟CPU时间),可作为工具调用能力声明的前置校验机制。
链接: https://arxiv.org/abs/2610.02142
作者: Juan S. Santillana
机构: 独立研究者(Independent Researcher)
类目: Computation and Language (cs.CL)
备注: 24 pages, 12 tables, preprint
Abstract:Keyword-matching benchmarks can credit small models for tool use they never perform. We document such a false positive in a matched-architecture pair of Spanish security language models and propose a ladder of strict, cheap diagnostics. A 661.6M parameter model (approx. 65% code/technical text; no dedicated SFT) and a 1,109M model (web-heavy multi-phase curriculum; 6B-token tool-SFT) share decoder, tokenizer, and special tokens, scoring almost identically on lenient tool-use metrics (B4: 0.660 vs. 0.650). Verbatim-reproduction checks on training examples separate them completely: the 600M emits valid tool calls with generalized arguments on 6/6 examples; the 1B does so on 0/6 across checkpoints. A first-token probe localizes the 1B’s failure to a missing prior (prob. 10^-4 – 10^-5 on |tool_call|), which was erased by its web-heavy training phase. A targeted SFT recipe (diverse corpus, 5x higher learning rate, 2,202 steps, ~3.3 GPU-hours) repairs the 1B using three orders of magnitude fewer tokens than the failed phase. On all 269 corpus rows, valid emission rises from 0.100 to 0.959 (600M: 0.926). On 238 unseen prompts, the repaired 1B passes 0.536 vs. the 600M’s 0.428 ( p = 0.004 ). Embedding-drift checks show the repair did not move the trigger token’s tied embedding (97.7% of the bf16 table remains bit-identical), meaning changes live in the surrounding network. Both models over-trigger, rarely answering negative prompts without a call (0.09 for 600M, 0.17 for repaired 1B). Factorial analyses confirm all repair configurations install the format, though suppression benefits from a diverse corpus remain a hypothesis due to seed sensitivity. This cheap diagnostic ladder costs minutes of CPU time and should gate tool-use claims on small models. Comments: 24 pages, 12 tables, preprint Subjects: Computation and Language (cs.CL) Cite as: arXiv:2610.02142 [cs.CL] (or arXiv:2610.02142v1 [cs.CL] for this version) https://doi.org/10.48550/arXiv.2610.02142 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[NLP-6] Finetuning with Sampling: SFT Learns Better Than You Think
【速读】: 该论文旨在解决前沿模型在后训练(posttraining)过程中如何有效融合离策略(off-policy)专家数据与在线策略(on-policy)学习优势的问题。传统方法中,监督微调(SFT)虽可利用离策略数据,但易导致泛化能力弱和灾难性遗忘;而强化学习(RL)虽具备良好泛化性且能保留已有能力,却依赖反复采样以探索成功轨迹,效率较低。本文的关键解决方案是不改变学习目标以适应离策略数据,而是通过设计一种马尔可夫链蒙特卡洛(Markov chain Monte Carlo, MCMC)采样算法,将离策略轨迹逐步转换为更接近在线策略分布的样本,从而提升数据对学习者的适配性。该方法使SFT在科学技能获取、数学推理及开放式专长等任务上达到甚至超越主流后训练技术的表现,不仅泛化能力更强、遗忘更少,还展现出超越基础模型分布的分布外学习能力。从更高层次看,本文提出将采样作为模型原生操作符(model-native operator),用于优化数据可学习性,为后训练全栈提供通用性强的基础范式。
链接: https://arxiv.org/abs/2610.02140
作者: Aayush Karan,Sitan Chen,Yilun Du
机构: Harvard University(哈佛大学)
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注:
Abstract:Introducing new capabilities to frontier models has long been the goal of posttraining, which predominantly employs supervised finetuning (SFT) and reinforcement learning (RL) to this end. Conventional wisdom dictates that RL enables strong generalization on new tasks without losing existing capabilities, while SFT is prone to weak generalization and catastrophic forgetting. At the same time, SFT can learn from off-policy expert data, whereas RL must rely on a model’s ability to find successful trajectories with repeated sampling. In our work, we seek to leverage the strength of on-policy learning while utilizing the privileged information contained in off-policy data. However, rather than modifying the learning objective to accommodate this data, we instead tailor the data distribution to better suit the learner. We introduce a Markov chain Monte Carlo (MCMC) sampling algorithm that progressively transforms off-policy traces to be more on-policy given a reference model for finetuning. Across tasks like scientific skill acquisition, mathematical reasoning, and open-ended expertise, our sampling algorithm enables SFT to rival prevailing posttraining techniques, often generalizing better and forgetting less than strong on-policy baselines. In addition, the resulting finetuned models exhibit strong distributional performance and are capable of learning beyond sharpening the base model distribution. At a higher level, our approach presents sampling as a model-native operator that shapes data for learnability, offering broader utility as a general-purpose primitive throughout the posttraining stack.
[NLP-7] Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows
【速读】: 该论文旨在解决现有文本到SQL(text-to-SQL)基准测试在评估企业级数据科学与分析智能体时存在的局限性问题。传统基准测试仅聚焦于单一表内的查询生成任务,且其答案键常存在错误,同时由于真实企业数据仓库涉及敏感信息而无法公开,导致这些基准多基于简化、非真实场景的公共数据集构建,难以反映复杂业务环境下的多表推理与决策能力需求。为弥补这一缺陷,论文提出Argo-Bench,一个包含210个真实世界数据科学与分析任务的评估框架,通过整合公开数据、同行评审行业文献及监管文件,模拟纽约市一个具备真实经济逻辑、欺诈模式和市场激励机制的食品配送平台,其数据规模达2024年8100万订单,构建了包含235张表、75亿行记录的企业资源规划(ERP)数据仓库,其架构基于Oracle E-Business Suite模型。该框架的关键创新在于:不仅要求智能体完成跨表的复杂推理与查询生成,还要求其在未暴露真实状态的情况下,通过探索数据仓库来重构事实,并执行实际业务操作(如封禁欺诈账户、分配骑手激励预算或发放补薪),最终由仿真器根据操作后果进行评分。每个任务均提供可执行的参考解决方案,验证了任务在仅依赖数据仓库的前提下具备可解性。实验结果显示,14个前沿模型中最强者仅在34.8%的任务中得分达到95以上,平均得分为59.5分,表明当前智能体在理解、导航并作用于真实复杂数据环境方面仍存在显著差距。因此,该研究的核心贡献在于建立了一个高保真、多任务、可行动评估体系,推动生成式智能体向真正具备企业级数据操作能力的方向发展。
链接: https://arxiv.org/abs/2610.02122
作者: Gabriel Tomitsuka,Arman Raayatsanati,Emma Xing,Duke Gand,Joseph J Ma
机构: TextQL
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Databases (cs.DB)
备注: 41 pages, 4 figures, 18 tables. Code: this https URL . Data: this https URL . Website: this https URL
Abstract:Real-world enterprise data science and analytics workflows require reasoning across dozens of tables, performing statistical analyses, and acting on the results. Established text-to-SQL benchmarks evaluate query generation alone, and audits have found their answer keys frequently wrong. Because real enterprise warehouses are too sensitive to release, these benchmarks are built on public datasets where a business event fits in a single table. We introduce Argo-Bench, an evaluation framework comprising 210 data science and analytics tasks. Drawing on public data, peer-reviewed industry literature, and regulatory filings, we simulate a food delivery platform in New York City at true scale, with 81 million orders in 2024, grounded economics, fraud patterns, and marketplace incentives. We export this world to an ERP warehouse of 235 tables and 7.5 billion rows, modeled on the Oracle E-Business Suite schema. The simulator’s ground-truth state is withheld from the warehouse the agent sees, so tasks require reconstructing facts by navigating the warehouse before acting on them. Argo-Bench goes beyond text-to-SQL: the agent files actions such as banning fraudulent accounts, allocating courier incentive budgets, or issuing back pay, and the grader scores each by its consequences in the simulator. Every task has an executable reference solution that demonstrates solvability using only the warehouse. The strongest of 14 frontier and open-weight models scores 95 or higher on only 34.8% of tasks and averages 59.5 points. We hope Argo-Bench drives progress toward agents that understand, navigate, and act within real data environments.
[NLP-8] Where-OPD: Spatially Guided On-Policy Self-Distillation of MLLM s with Synthetic Scenes
【速读】: 该论文旨在解决多模态大语言模型(Multimodal Large Language Models, MLLMs)在推理能力提升过程中,缺乏有效且可扩展的自监督训练范式的问题,尤其针对如何利用特权信息(privileged information)增强模型对复杂视觉内容的细粒度感知与推理能力。现有方法依赖人工标注的视觉定位数据或外部教师模型,限制了其泛化性与可扩展性。本文提出一种新型的基于策略的自蒸馏(on-policy self-distillation)框架,其关键在于引入空间语义引导(spatially grounded guidance),即通过程序生成的场景自动提供物体身份与空间坐标信息,作为教师模型的特权输入,指导其从多个相关图像区域中定位并整合证据。学生模型则仅依赖图像和问题,学习模仿教师的行为。该方法无需人工标注,支持可扩展的后训练(post-training),且实验表明,尽管训练阶段仅使用合成数据,模型在真实世界视觉理解基准(如CVBench、ZoomBench、MME-RealWorld等)上实现了平均性能提升3.23点,验证了从合成到真实场景的显著迁移能力。因此,解决方案的核心创新在于:利用可自动获取的空间语义引导,实现高效、无标注的自蒸馏,从而在不依赖外部监督的前提下,激发模型更广泛的感知与推理能力。
链接: https://arxiv.org/abs/2610.02117
作者: Sophia Sirko-Galouchenko,Monika Wysoczanska,Andrei Bursuc,Nicolas Thome,Spyros Gidaris
机构: Valeo.ai; Sorbonne Université, CNRS, ISIR, F-75005 Paris, France; Institut universitaire de France (IUF); ILLS, CNRS, Montreal, QC H2S 3H1
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
备注:
Abstract:On-policy self-distillation has recently emerged as an effective approach for improving language-model reasoning by supervising students with a frozen or EMA version of themselves that receives privileged information. Its application to multimodal large language models (MLLMs), however, remains largely unexplored. Recent approaches use privileged visual information, such as image crops corresponding to a question, to improve fine-grained perception, but their gains are confined to tasks that benefit from such visual zooming and require either human-annotated grounding data or external teacher models. We introduce a different form of on-policy self-distillation for MLLMs that provides the teacher with textual, spatially grounded guidance identifying the visual elements relevant to a query. We use procedurally generated scenes with automatically available object identities and spatial coordinates, enabling scalable and annotation-free post-training. The teacher uses this spatial guidance to locate and integrate evidence from multiple relevant image regions, while the student learns to reproduce the resulting behavior from the image and question alone. Our approach consistently improves performance on counting, document and chart understanding benchmarks across multiple models. Importantly, although post-training uses only synthetic scenes, the resulting improvements transfer to real-world perception benchmarks, yielding a 3.23-point gain in average performance across CVBench, V*, ZoomBench, BLINK, HR-Bench, and MME-RealWorld. These results show that spatially grounded privileged information can induce broader perceptual capabilities through on-policy self-distillation, enabling substantial synthetic-to-real transfer beyond the task and data distribution used for post-training. Project page: this https URL
[NLP-9] A Comparative Explainability Framework for DeBERTa-v3 in Zero-Shot Medical Abstract Classification
【速读】: 该论文旨在解决可解释人工智能(Explainable Artificial Intelligence, XAI)中的解释分歧问题,即针对同一输入与预测结果,不同归因方法产生不一致的解释。为应对这一挑战,研究构建了一个对比可解释性框架,对DeBERTa-v3在医学摘要零样本分类任务中的表现进行审计。其解决方案的关键在于:首先,在医学摘要语料库上集成五组每类诊断类别增强的假设,并选取每类1000篇平衡样本;其次,综合比较五种解释方法——包括模型无关的SHAP与LIME、深度学习特定的遮挡法(Occlusion)与Input x Gradient,以及Transformer特异的Attention x Gradient;通过顶部词元归因(top-token attribution)对解释结果进行标准化,并利用Jaccard指数量化成对解释的一致性。实验表明,模型在语义清晰的临床领域表现出高预测准确性,而在高语义模糊情境下性能下降;解释稳定性与预测置信度直接相关,在单一类别中呈现强收敛性,而在诊断不确定性下显著降低。此外,定性错误分析揭示了三类系统性失效机制:词汇过度敏感、语义重叠及归因一致性丧失。研究支持在医疗文本分类任务中采用多种解释方法结合定量一致性度量进行审计,并建议优先使用特定临床本体而非宽泛的诊断标签以提升解释可靠性。
链接: https://arxiv.org/abs/2610.02116
作者: Javier Diaz Esteban-Herreros,David Muñoz-Valero,Raquel Martínez-España,Jose M. Juarez,Juan Moreno-Garcia
机构: Universidad de Castilla-La Mancha (卡斯蒂利亚-拉曼查大学); University of Murcia (穆尔西亚大学); Murcian Bio-Health Institute (IMIB-Arrixaca) (穆尔西亚生物健康研究所)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: 18 pages, 6 figures
Abstract:A comparative explainability framework is presented to audit DeBERTa-v3 under zero-shot classification of medical abstracts. The work addresses the disagreement problem in Explainable Artificial Intelligence, where different attribution methods produce divergent explanations for the same input and prediction. A natural language inference engine is implemented over the Medical Abstracts corpus with five enriched hypotheses per diagnostic category and a balanced sample of one thousand texts per class. Five explanation methods are compared: SHAP and LIME as model-agnostic approaches, occlusion and Input x Gradient as deep-learning-specific approaches, and Attention x Gradient as a transformer-specific approach. Explanations are standardized through top-token attribution, and pairwise agreement is quantified using the Jaccard index. High predictive accuracy is achieved across well-defined clinical domains, whereas performance degrades under high semantic ambiguity. Explanatory stability directly mirrors predictive certainty, exhibiting strong convergence in univalent categories and a marked drop under diagnostic uncertainty. Furthermore, qualitative error auditing uncovers three systemic failure mechanisms: lexical hypersensitivity, semantic overlap, and loss of attribution coherence. The results support the combined use of several explanation methods and quantitative agreement metrics when auditing transformer-based models in medical text classification, and suggest prioritizing specific clinical ontologies over broad diagnostic labels.
[NLP-10] Scalable Transferable Meta-network for Data Selection Requires a Different Loss (and Why the Obvious Choice is Problematic)
【速读】: 该论文旨在解决大规模语言模型训练中数据选择(data selection)的关键挑战,即如何在海量且异构语料库中高效筛选出对目标任务具有高价值的训练样本。现有基于元学习的训练数据选择方法(Meta-learning for Training-data Selection, MTS)虽能通过目标验证集优化实现更合理的数据权重分配,但面临细粒度评估与跨数据集泛化能力之间的权衡。为克服此问题,论文提出一种可迁移的示例评分与选择框架——Transferable Example Scoring and Selection (TESS),其核心在于引入点对点价值匹配目标(Pointwise Value Matching, PVM),以替代传统的逐样本权重机制。该方案通过设计更具鲁棒性的优化目标,有效缓解了因权重抑制和对易学特征的持续依赖所导致的优化不稳定性与泛化性能下降问题。实验表明,TESS在大模型安全性和定向指令微调任务中均展现出优异的跨数据集、跨规模(从子集到全语料库、从小模型到大模型)迁移能力,验证了其在实际应用中的高效性与可扩展性。
链接: https://arxiv.org/abs/2610.02092
作者: Zilin Du,Bowen Yang,Boyang Albert Li
机构: Nanyang Technological University (南洋理工大学); Singapore (新加坡)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:
Abstract:Data selection is critical for training large language models on massive and heterogeneous corpora. Meta-learning for Training-data Selection offers a principled alternative to heuristic scoring by learning data weights from a target validation objective, but existing methods face a trade-off between fine-grained valuation and transferability to unseen data. A natural solution is to replace per-sample weights with a selection network. However, we find that directly incorporating such a network into existing MTS objectives leads to unstable optimization and poor generalization, caused by weight suppression and persistent reliance on easy-to-learn features. To address these issues, we propose Transferable Example Scoring and Selection (TESS), a scalable data-selection framework built on a Pointwise Value Matching objective (PVM). Experiments on LLM safety and targeted instruction tuning demonstrate strong transfer across datasets, from subsets to full corpora, and from smaller to larger models.
[NLP-11] LLM 2Jev: LLM s Are Already Jev-Style Decision Models – When and How to Fine-Tune Them
【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在无需额外训练的情况下,是否具备直接输出符合Jev风格决策(即对预定义选项生成类别化概率分布,而非自由文本)的能力,以及何时需要微调的问题。其核心挑战在于如何在不破坏模型原有生成能力的前提下,将LLM的输出转化为可直接用于系统决策的结构化概率分布。解决方案的关键是提出LLM2Jev——一种保持架构不变的框架,通过提取括号内数值标识符的下一个词元概率,直接生成校准后的决策结果;该框架同时提供免训练推理方案与基于树因子化列表损失(tree-factorized listwise loss)的微调目标,并引入KL散度惩罚项,以锚定辅助预测与基础模型行为一致,从而防止对话生成能力退化。实验表明,现代大模型(如Qwen3.5-4B)在无训练条件下已具备与专门构建的Jev模型相当的决策性能,支持任意数量选项及多模态输入;而微调仅在弱模型或特定任务(如多选项意图路由)中带来显著提升,对强基线模型收益递减,且使用LoRA微调在高性能模型上表现最优。
链接: https://arxiv.org/abs/2610.02076
作者: Yinheng Li,Justin Wagle
机构: Microsoft(微软)
类目: Computation and Language (cs.CL)
备注:
Abstract:Jev-style decision models return categorical probability distributions over predefined options without generating free-form text, enabling software systems to act on their outputs directly. In this work, we investigate the extent to which general-purpose LLMs already possess this capability out of the box, and when fine-tuning is actually necessary. We present LLM2Jev, an architecture-preserving framework that extracts calibrated decisions directly from next-token probabilities over bracketed numeric identifiers. LLM2Jev provides both a training-free inference recipe and a fine-tuning objective that optimizes candidate selection via a tree-factorized listwise loss while anchoring auxiliary predictions to the base model using KL divergence penalties. Evaluating on Qwen3.5-4B and Qwen3-0.6B, we find that modern LLMs are inherently effective decision models: without training, the 4B model matches community Jev-style models built on the same backbone, outperforms letter-logit readouts, supports arbitrary option counts, and natively handles multimodal decisions over images. Fine-tuning provides targeted rather than universal benefits – substantially improving weaker models and specific tasks (such as many-option intent routing), but offering diminishing returns for strong backbones. Crucially, our KL anchors prevent behavioral degradation in conversational text generation, with LoRA delivering the strongest performance on capable models.
[NLP-12] ypological Alignment of Stack-Based Language Models on Mildly Context-Sensitive Artificial Languages EMNLP2026
【速读】: 该论文旨在探究自然语言(NLs)中某些句法特征(如主-宾-谓,SOV)为何在数千种已知语言中更为普遍,核心问题是揭示语言类型学共性是否源于语言学习中的认知偏差。现有研究通过语言模型(LMs)的计算模拟支持了这一假设,但其分析在数据和模型层面仍存在局限。本文的关键突破在于从两个方面扩展了研究:一是考察跨串行依存关系(cross-serial dependencies),即目前人类语言中可接受的句法复杂度上限;二是引入基于栈的语言模型(stack-based LMs, SLMs),以评估其对层次化句法模式的学习能力。实验结果表明,尽管SLMs在处理跨串行依赖结构时表现不佳,但具有有限工作记忆的SLMs展现出更优的泛化能力,暗示这种限制可能构成一种归纳偏置(inductive bias),从而解释为何部分词序配置在语言类型学中具有普遍性。因此,解决方案的关键在于揭示工作记忆容量的限制如何促进对特定句法结构的偏好,进而为语言类型学共性提供基于计算模型的机制解释。
链接: https://arxiv.org/abs/2610.02040
作者: Nadine El-Naggar,Tatsuki Kuribayashi,Ted Briscoe
机构: Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学); Tohoku University(东北大学)
类目: Computation and Language (cs.CL)
备注: EMNLP 2026 Main Conference
Abstract:Some properties of languages, e.g., subject-object-verb (SOV) word order, are more prevalent than others among the thousands of attested natural languages (NLs). Such typological commonality is often attributed to learning biases. Computational simulations, recently with language models (LMs), have facilitated the exploration of this theory. In this paper, we extend existing analyses of the relationship between LMs’ learning biases and typological commonality on both data and model sides, focusing on: (i) cross-serial dependencies, the upper limit of attested syntactic complexity, and (ii) stack-based LMs (SLMs), potentially facilitating learning of hierarchical patterns. We first evaluate generalization of SLMs on cross-serial dependencies across diverse artificial languages and confirm that they struggle with such constructions. However, SLMs with limited working memory generalize better suggesting a possible basis for such inductive bias and thus the typological commonality of some word order configurations.
[NLP-13] CARM: Cancellation-Aware Response Masking for LLM Reinforcement Learning
【速读】: 该论文旨在解决大语言模型(Large Language Model, LLM)在强化学习(Reinforcement Learning, RL)后训练过程中因策略更新与采样引擎和训练引擎之间差异导致的响应偏离策略(off-policy)问题。其核心挑战在于传统序列级掩码方法采用长度归一化的几何平均概率比对数,易因正负对数比率相互抵消,掩盖双向策略漂移,从而影响优化效果。为此,论文提出取消感知响应掩码(Cancellation-Aware Response Masking, CARM),关键创新在于对每个词元的对数概率比取绝对值后再进行平均,有效防止了相反方向的概率变化相互抵消。理论分析表明,被接受的响应满足样本词元比例超出预设区间范围的比例及其对数距离边界之外的均值的联合约束。实验结果在数学推理与代码生成任务上验证了CARM的有效性:在AIME 2024/2025/2026及BeyondAIME数据集上的mean@16平均提升达3.13个百分点,在四个代码基准测试中的pass@1平均提升2.88个百分点,显著优于现有最强基线。研究证明,CARM是一种兼具理论严谨性与实际效能的响应级离策略控制方法。
链接: https://arxiv.org/abs/2610.02039
作者: Yafei Zhang,Songshuo Lu,Sicong Liao,Zhi Chen,Yaohua Tang
机构: Moore Threads AI
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: 28 pages, 11 figures, 5 tables
Abstract:Recent years have witnessed the rapid adoption of reinforcement learning (RL) in large language model (LLM) post-training, with substantial gains in mathematical reasoning and code generation. In practical systems, however, policy updates and differences between rollout and training engines can make sampled responses off-policy. Sequence-level masking addresses this mismatch by deciding whether an entire response should contribute to optimization. A common masking rule uses the length-normalized geometric mean of sampled token probability ratios. Its signed log-ratios can cancel across positions, concealing substantial bidirectional policy drift. We propose \emphCancellation-Aware Response Masking (CARM), a sequence-level mask that takes the absolute value of each token log-ratio before averaging, preventing opposing probability changes from canceling. We prove that accepted responses satisfy a joint bound on the fraction of sampled-token ratios outside a prescribed band and their mean log-distance beyond its boundaries. Experiments on mathematical reasoning and code generation show that CARM improves mean@16 averaged over AIME 2024/2025/2026 and BeyondAIME by up to 3.13 percentage points over geometric-mean masking, and increases average pass@1 across four code benchmarks by 2.88 points over the strongest evaluated baseline. These findings support CARM as a theoretically grounded and effective method for response-level off-policy control in LLM reinforcement learning.
[NLP-14] Old Ideas Novel Problems: The Instability of LLM -Based Novelty Evaluation
【速读】: 该论文旨在解决当前自动化创意生成系统(automated ideation systems)在评估其生成想法新颖性时所面临的评测方法可靠性问题。现有评估普遍依赖大型语言模型(LLM)作为新颖性评判者,但这些评判模型多为临时搭建,且验证过程通常基于人类撰写的论文而非真实生成的想法,导致评测结果不可靠。研究的关键发现在于:微小的提示工程(prompt design)差异会显著影响评判结果——例如,仅通过告知模型某想法被评审人认为新颖而另一想法不新颖,即可使同一组想法对的评判结果在超过一半的情况下发生改变,导致成对准确率波动超过50个百分点,甚至低于随机水平。此外,检索增强与更大的推理预算对性能提升有限,而两个专为新颖性评估设计的模型反而不如最简单的提示基线。因此,该研究揭示了当前新颖性评测体系的根本缺陷,强调亟需建立更稳健、可复现的新颖性评估方法,以确保对自动化创意生成系统性能的可信评价。
链接: https://arxiv.org/abs/2610.02022
作者: Noy Sternlicht,Simra Shahid,Peter Jansen,Daniel S. Weld,Pao Siangliulue,Tom Hope
机构: Hebrew University of Jerusalem (希伯来大学); Allen Institute for AI (艾伦人工智能研究所); Microsoft (微软); University of Arizona (亚利桑那大学); University of Washington (华盛顿大学)
类目: Computation and Language (cs.CL)
备注:
Abstract:Automated ideation systems are often evaluated on the novelty of the ideas they produce, and that judgment is increasingly delegated to large language models. Such judges are typically built ad hoc and validated, if at all, on human-authored papers rather than on the generated ideas they are meant to score. So, how do novelty judges perform? Not well. We present a systematic controlled study of novelty evaluation design choices. We first build an evaluation set automatically, mining OpenReview for passages where reviewers explicitly affirm or dispute a paper’s originality and keeping only submissions with unanimous agreement at the extremes of their research area; we pair these with ideas from a vanilla LLM generator. Across six judges, we find that small prompt design choices have large consequences; e.g., simply telling the judge that reviewers found one idea novel and the other not can change its verdict on more than half of the identical idea pairs it is shown, shifting pairwise accuracy by over 50 points and occasionally pushing it below chance. The same change helps one judge and hurts another. Retrieval and larger reasoning budgets help little, and two purpose-built novelty evaluators are outperformed by our cheapest prompted baseline. These results raise questions about reported novelty gains of automated ideation systems, and call for robust novelty evaluation methods. Subjects: Computation and Language (cs.CL) Cite as: arXiv:2610.02022 [cs.CL] (or arXiv:2610.02022v1 [cs.CL] for this version) https://doi.org/10.48550/arXiv.2610.02022 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[NLP-15] Controllable Multi-label Video Safety Detection via Adaptive Tversky Policy Optimization
【速读】: 该论文旨在解决现有有害视频检测系统在多标签场景下的两大核心问题:一是将安全检测简化为二分类任务,忽略了视频内容中可能同时包含多种违规类型的多标签本质;二是依赖静态的训练目标,无法灵活调整精确率与召回率之间的权衡,难以适应不同审核流程和违规类别对模型性能指标的差异化需求。其解决方案的关键在于提出一种基于强化学习的多标签视频安全检测框架——自适应Tversky策略优化(Adaptive Tversky Policy Optimization, ATPO),通过引入自适应Tversky奖励(Adaptive Tversky Reward, ATR),在训练过程中动态调节误报(false-positive)与漏报(false-negative)的惩罚权重,从而实现对精确率-召回率平衡点的可控调节。实验结果表明,ATPO在SafeWatch-Bench和XD-Violence数据集上显著提升了多标签检测性能,尤其在SafeWatch-Bench-Real上的Jaccard指数从40.66提升至75.44,同时验证了其在不同部署场景下支持异构策略需求的能力。
链接: https://arxiv.org/abs/2610.02019
作者: Guangyu Yang,Jingbiao Mei,Mingsheng Sun,Jinghong Chen,Yingtong Bu,Pengda Qin,Da Chen,Bill Byrne
机构: University of Cambridge (剑桥大学); Xiaohongshu Inc. (小红书公司); AntGroup (蚂蚁集团); Tencent Company (腾讯公司); University of Bath (巴斯大学)
类目: Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:The rapid growth of video-based social media has increased users’ exposure to harmful content, creating a need for reliable automated video safety detection. Although recent Vision-Language Models (VLMs) show strong video understanding capabilities, existing harmful video detection systems face two key limitations: they typically reduce safety detection to binary classification, overlooking the inherently multi-label nature of unsafe videos, and they rely on static training objectives that do not support controllable precision-recall trade-offs, though the desired operating point may vary across moderation pipelines and unsafe categories. To address these gaps, we propose Adaptive Tversky Policy Optimization (ATPO), a reinforcement learning framework for Multi-label Video Safety Detection (Multi-VSD). ATPO introduces the Adaptive Tversky Reward (ATR), which dynamically adjusts false-positive and false-negative penalties during training to enable controllable precision-recall trade-offs. Experiments on SafeWatch-Bench and XD-Violence show that ATPO substantially improves multi-label performance, increasing the Jaccard Index from 40.66 to 75.44 on SafeWatch-Bench-Real. Moreover, ATR enables reliable steering of the precision-recall operating point, supporting deployment scenarios with heterogeneous policy requirements. Code and checkpoints are provided at this https URL .
[NLP-16] Mem: Non-Destructive Memory for Long-Term Organizational LLM Agents
【速读】: 该论文旨在解决大语言模型(Large Language Model, LLM)代理在组织级工作中面临的记忆系统无法有效支持时间敏感性决策追溯的问题。具体而言,当组织内的决策以文档形式随时间不断更新时,新版本的决策以独立文档形式出现而非直接编辑旧版本,因此回答某一时间点的提问需要精确识别当时有效的文档版本。然而,现有大多数记忆系统在写入阶段即通过生成式方法对文档进行事实提炼或压缩,导致信息被提前固化,限制了后续问答中可访问的原始上下文范围。其核心问题在于:写入时的过度压缩使系统丧失了对历史版本的完整保留能力,从而难以支持基于时间点的动态检索与推理。为此,本文提出Mem++——一种非破坏性记忆框架,其关键创新在于将信息处理从“写入时提炼”转变为“读取时选择”。Mem++完整存储每份文档及其创建时间与作者信息,在写入时不调用任何生成式模型;在查询时,仅检索截至提问时间点的所有相关文档,并结合词汇匹配与语义相似度进行融合排序。该设计确保所有历史版本均被保留,将版本选择权交由下游回答模型自主决定。实验结果表明,Mem++在组织级基准测试OrgMemBench上超越最强基线8.0至13.1分,使用gpt-4.1-mini时达到2.6分领先于RAG,同时在LoCoMo和LongMemEval-S等多任务评估中分别取得最佳平均LLM评分与第二名成绩,验证了其在保留完整时间上下文与提升问答准确性方面的显著优势。
链接: https://arxiv.org/abs/2610.02002
作者: Ahmad Yehia,Aly O. Abdelkareem,Islam Ahmed,Hesham Omran,Khaled Alashmouny,Christian Claudel,Abduallah Mohamed
机构: The University of Texas at Austin(德州大学奥斯汀分校); AIDAChip Inc.(AIDA芯片公司)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: 15 pages, 4 figures
Abstract:Large Language Model (LLM) agents now take part in organizational work, where many authors record decisions across documents over months. Because a revised decision arrives as a new document rather than an edit, answering a question requires knowing which version held at a given time. However, most memory systems compress the record at write time. By distilling each document into facts, notes or graph edges, these methods fix what can be answered before any question is asked. To address this, we propose Mem++, a non-destructive memory framework shifting from write-time distillation to read-time selection. Mem++ stores every document whole with its date and author, and it calls no generative model at write time. At read time, it retrieves only documents dated up to the time a question asks about and fuses lexical and semantic rankings. Unlike systems that overwrite older versions, Mem++ keeps them and leaves the choice to the answering model. Evaluations on the organizational benchmark OrgMemBench demonstrate that Mem++ surpasses the strongest memory system baseline by 8.0 to 13.1 points across two answering models. With gpt-4.1-mini, it also achieves the best overall score, 2.6 points above RAG. In addition, Mem++ achieves the best average LLM-judge score on LoCoMo and ranks second on LongMemEval-S, behind only its entity-graph variant. Code for benchmark evaluation is available at this https URL.
[NLP-17] Mingbird: A Local-First Agent Harness Enabling Small Open Models to Complete Real Tasks
【速读】: 该论文旨在解决小型开源模型(2-9B参数)在云规模代理框架下难以完成真实任务的问题,其核心挑战包括工具预填充导致上下文溢出、自我修正过程发散、工具示范陷入循环以及任务被无声放弃等。研究表明,这些问题的大部分根源在于代理框架(harness)的设计缺陷,而非模型本身能力不足。为此,论文提出Mingbird——一个面向Windows与Ollama平台的本地优先代理框架,其关键创新在于通过十项机制精准补偿小模型在实际应用中的各类失效模式,其中代表性机制包括:基于字节级净零预填充预算的资源控制、在任务完成前重新读取任务目标的“完成门”机制,以及基于签名级别的循环检测。在控制实验(LRAB)中,Mingbird在4种框架×4个开源模型(2B-35B)×18个真实任务的全组合测试中取得0.886的整体得分,显著优于其他基准框架(goose: 0.631, opencode: 0.479, agent-mini: 0.405),且所有288组实验结果均公开可复现;在τ²-bench上也达到0.856,优于对比框架。此外,对前沿模型的探针测试显示,不同框架间的性能差异主要源于框架设计,而良好构建的框架结构在相同任务上表现一致性高(最大偏差<0.072)。消融实验表明各机制具有方向性影响,单次实验重复的波动幅度可达0.069,但完整机制堆栈相较仅含文本重读的单一机制,在三次重复中平均提升+0.10,验证了机制协同的有效性。研究局限在于使用自建基准、单机环境及单次评分,需进一步验证泛化性。
链接: https://arxiv.org/abs/2610.02001
作者: Hao Wang,Ting Huang
机构: University of Science and Technology Beijing(北京科技大学); Honor Device Co., Ltd.(荣耀终端有限公司)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: 44 pages, 9 figures. Code, benchmark protocol, scoring code, and all 288 per-cell results: this https URL
Abstract:Small open-weight models (2-9B) run on ordinary laptops, but under cloud-scale agent harnesses they rarely complete real tasks: tool prefill overflows the context, self-correction diverges, tool demonstrations loop, and tasks are silently abandoned. We present evidence, from a controlled single-machine comparison and one third-party benchmark, that a substantial share of these failures is attributable to the harness rather than the model. We introduce Mingbird, a local-first agent harness for Windows and Ollama whose ten mechanisms compensate point-by-point for small-model failure forms, three of them representative: a byte-level net-zero prefill budget, a finish gate that re-reads the task before accepting completion, and signature-level loop detection. On LRAB, a controlled comparison holding machine, models, budgets, and scoring fixed (4 harnesses \times 4 open models (2B-35B) \times 18 real tasks, deterministic artifact scoring), Mingbird reaches 0.886 overall against 0.631 (goose), 0.479 (opencode), and 0.405 (agent-mini), with all 288 cells published; on \tau^2 -bench (278 tasks, three arms, one protocol) it totals 0.856 against 0.791 and 0.737; and a frontier-model probe on the same 18 tasks spans 0.997 to 0.478 across harnesses, with well-formed scaffolds staying within 0.072 of each other. A leave-one-mechanism-out ablation is reported as directional only: same-night replications of the same arm move its mean by up to 0.069, the size of every nominal single-trial delta, and the one batch-matched comparison (full mechanism stack versus text re-read alone) gives the executable completion guards a paired +0.10 across three replications. The evidence carries stated limits: a self-built benchmark, a single machine, and single-trial scoring.
[NLP-18] Universal Byte-Level Encoding: UTF-8/UTF-16 Routing to Reduce Cross-Script Token-Budget Disparities NEURIPS2026
【速读】: 该论文旨在解决多语言大语言模型(LLM)中基于字节级分词(Byte-level Byte-Pair Encoding, BBPE)的编码效率问题,特别是针对非英语语言在UTF-8编码下因高“预合并成本”(encoding floor)导致的token数量膨胀、上下文利用率下降及推理延迟增加的问题。其核心挑战在于:当无法应用已学习的合并规则时,多字节字符需分解为多个字节符号,而某些脚本(如中文、日文等)的3-4字节字符在标准UTF-8-BBPE中会产生显著更高的token开销,形成“编码地板”(encoding floor),从而加剧跨语言之间的token预算不均衡。解决方案的关键是提出通用字节级编码(Universal Byte-Level Encoding, UBE),一种双字母表分词机制:保留1-2字节的UTF-8字符沿原路径处理,同时将3-4字节的UTF-8字符(主要为基本多文种平面,BMP)通过UTF-16路径进行映射。该设计在不改变标准BPE合并规则和保持精确解码的前提下,有效降低了高溢价脚本(high-premium scripts)的编码地板,同时避免对已高效表示的英文片段造成额外开销。实验表明,UBE在Unicode 17规范下实现所有标量值与测试套件的精确往返转换,在内在评估中显著降低英语归一化后token数量比的离散度,减少跨语言token预算差异;在多语言语言模型中,其生成质量与标准BBPE相当,且在固定上下文预算下显著减少高溢价语言的token消耗,略微增加英文开销,从而提升可用上下文长度并加速内容匹配基准下的提示处理速度。
链接: https://arxiv.org/abs/2610.01984
作者: Hyunsik Kim,Youngmoon Jung
机构: Samsung Research(三星研究院)
类目: Computation and Language (cs.CL); Machine Learning (cs.LG)
备注: Accepted to NeurIPS 2026
Abstract:Byte-level byte-pair encoding (BBPE) tokenizers are attractive for multilingual large language models (LLMs) because they cover all Unicode text. In UTF-8-based BBPE, however, many scripts start from a higher fallback cost than English: when no learned merges can be applied, a multibyte character requires multiple byte-derived symbols. We call this worst-case pre-merge cost the encoding floor. A higher floor can increase token counts and per-request cost and shrink usable context. Changing the text encoding can reduce this gap, but a single global encoding can make already-efficient English spans more expensive in mixed-script text. We propose Universal Byte-Level Encoding (UBE), a dual-alphabet tokenizer that keeps 1-2-byte UTF-8 characters on the UTF-8 path while routing 3-4-byte UTF-8 characters through UTF-16. This lowers the encoding floor for 3-byte Basic Multilingual Plane (BMP) characters in scripts with high token premiums (token counts relative to English) without raising it for already-efficient spans in mixed-script text. UBE changes only the byte representation presented to byte-pair encoding (BPE); the merge rule remains standard, and exact decoding is preserved. UBE also composes with alternative boundary policies and morphology-based representations. In a Unicode 17 audit, UBE exactly round-trips all Unicode scalar values and all inputs in the official normalization, grapheme-break, and emoji test suites. Across intrinsic evaluations, UBE lowers dispersion in English-normalized token-count ratios, reducing cross-lingual token-budget disparity. In multilingual language model (LM) experiments, UBE matches BBPE’s LM quality. In the main multilingual settings, UBE reduces token counts most for high-premium scripts and slightly lowers English token counts, yielding more usable context under fixed token budgets and faster prompt processing in content-matched benchmarks.
[NLP-19] Counterfactual Auditing of Bias in Open-Source Large Language Models for Clinical Triage
【速读】: 该论文旨在解决急诊科分诊(Emergency Department Triage)中因人口统计学、社会经济状况及系统背景信息引入的不公平偏倚问题,特别是在使用开源大语言模型(LLM)进行儿科紧急严重程度指数(ESI)预测时,不同模型家族、规模、医学领域预训练模型及领域适配模型在面对反事实情境下的偏倚敏感性差异不明确的问题。其解决方案的关键在于提出一种对比性的反事实审计框架,通过构建仅改变单一变量(如人口统计、医疗可及性、行为或系统背景等)而临床表现保持不变的成对反事实临床案例,系统评估十种开源LLM在真实与手册式临床病例中的偏倚表现。研究发现,模型规模或医学领域预训练并不必然降低偏倚敏感性,反而部分大型或医学专用模型表现出更显著的偏倚转移;其中经QLoRA微调的Qwen2.5-7B模型表现最优,整体偏倚最小(任意偏移率5.27%,均方绝对偏移0.0534),而基础模型则高达16.02%和0.1706。进一步的分层与相关性分析揭示了具有临床意义的方向性偏倚模式和共性失败特征,凸显了反事实审计作为轻量级、临床可解释的公平性风险评估工具,在临床部署前对开源生成式AI模型进行公平性比较的重要性。
链接: https://arxiv.org/abs/2610.01963
作者: Manar Aljohani,Brandon Ho,Kenneth McKinley,Dennis Ren,Xuan Wang
机构: Virginia Tech (弗吉尼亚理工大学); University of Washington and Seattle Children’s Hospital (华盛顿大学和西雅图儿童医院); Children’s National Hospital (儿童国家医院)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
备注:
Abstract:Emergency department (ED) triage is a high-stakes prioritization task in which demographic, socioeconomic, and system-context information may improperly influence acuity assignment. Although open-source large language models (LLMs) are increasingly considered for local and privacy-preserving clinical decision support, it remains unclear how counterfactual bias varies across model families, sizes, medical-domain models, and domain-adapted models. We present a comparative counterfactual audit of ten open-source LLMs for pediatric Emergency Severity Index (ESI) prediction. Starting from real and handbook-style clinical vignettes, we construct paired counterfactual variants that change only one injected demographic, socioeconomic, healthcare-access, behavioral, social, or system-context variable while holding the clinical presentation fixed. Models include Qwen2.5-7B, Qwen2.5-14B-Instruct, a QLoRA fine-tuned Qwen2.5-7B, MedGemma variants, MedLLaMA2-7B, GPT-OSS-20B, and GPT-OSS-120B. We measure any counterfactual shift, undertriage, overtriage, shifts greater than one ESI level, mean shift, and mean absolute shift. Counterfactual sensitivity varied substantially and did not consistently decrease with larger model size or medical-domain pretraining. The fine-tuned Qwen2.5-7B showed the lowest overall sensitivity, with a 5.27% any-shift rate and mean absolute shift of 0.0534, versus 16.02% and 0.1706 for the base model. Several larger or medical-domain models showed more significant shifts. Stratified and correlation analyses further revealed clinically important directionality and shared failure patterns hidden by aggregate rates. These findings support counterfactual auditing as a lightweight, clinically interpretable framework for comparing fairness risks in open-source LLMs before clinical deployment.
[NLP-20] Latent JEPA: Abstract Future Prediction for Latent Reasoning in Chemistry
【速读】: 该论文旨在解决大语言模型在化学推理任务中缺乏对复杂反应路径和分子结构演变的深层理解与可解释性问题,尤其关注如何使模型在未显式生成每一步中间推理过程的情况下,仍能通过连续隐变量(latent thoughts)有效预测未来解决方案的关键信息。其核心解决方案是提出一种名为Latent JEPA的框架,该框架通过联合嵌入预测(joint-embedding prediction)机制,将自回归学习与多视角未来状态预测相结合,设计了文本与分子双重预测目标,使隐变量在训练过程中学会捕捉对未来反应路径及分子结构演化的关键先验信息。实验结果表明,该方法在ChemCoTBench基准上显著提升了分子优化能力以及编辑和反应相关指标的表现;表征分析进一步验证了未来预测使隐变量对分子结果更具信息量,并增强了其与化学结构之间的语义对应关系。研究支持以抽象的未来预测作为连接连续隐式推理与科学结果的学习原则,为构建更高效、可解释的化学生成式系统提供了新范式。
链接: https://arxiv.org/abs/2610.01947
作者: Xinjian Zhao,Yaoyao Xu,Xuemin Chen,Xiaozhuang Song,Tianshu Yu
机构: School of Data Science, The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)数据科学学院); Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
类目: Machine Learning (cs.LG); Computation and Language (cs.CL)
备注:
Abstract:Large language models offer a promising foundation for chemical reasoning, bringing together chemical knowledge and multistep problem solving. Chemical intuition can provide an initial sense of plausible outcomes before the details of a solution are fully worked out. Inspired by how such expectations complement explicit analysis, we study how continuous latent thoughts can be trained to anticipate informative aspects of future solutions without verbalizing every intermediate step. We introduce Latent JEPA, a framework that combines autoregressive learning with joint-embedding prediction of one or more future views. For chemical reasoning, we develop textual and molecular prediction objectives that connect latent thoughts to both subsequent reasoning and molecular outcomes. Experiments on ChemCoTBench show gains in molecular optimization and on several editing and reaction metrics. Representation analyses show that future prediction makes latent thoughts more informative about molecular outcomes and strengthens their correspondence with chemical structure. These findings support abstract future prediction as a learning principle for connecting continuous latent reasoning with scientific outcomes.
[NLP-21] A rubric landscape for evaluating clinical reasoning in large language models : what exists what is missing and what needs to be combined
【速读】: 该论文旨在解决当前大型语言模型(LLM)在临床记录推理评估中存在“应试准确性”表面化的问题,即现有评估方法无法有效检验模型是否具备真正的临床推理能力。其核心问题在于:现有评估工具未能全面覆盖临床推理的关键维度,尤其在时间序列整合、不确定性校准、反事实推理及推理忠实性等方面存在显著不足。解决方案的关键在于构建一个系统化的多维评估框架,通过整合六项核心维度——问题表征、时间合成、鉴别与管理推理、反事实推理、校准的不确定性以及推理忠实性——来全面衡量模型的临床推理能力。作者提出应采用二元评分条目、独立的完整性与正确性评分、基于病例重要性的非补偿性安全上限、时序一致性检查以及校正偶然性后的可靠性报告等方法,对现有工具进行组合应用。同时指出,针对纵向自由文本记录中的校准不确定性、反事实推理和忠实性评估仍需进一步设计与验证,本研究提供的是一个理论设计依据而非可直接使用的验证工具。
链接: https://arxiv.org/abs/2610.01938
作者: Zhangshu Joshua Jiang,Zina Ibrahim,James T. Teo
机构: King’s College London(国王学院); DRIVE-Health CDT; Department of Biostatistics and Health Informatics; Institute of Psychiatry, Psychology and Neuroscience; Cleveland Clinic London
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: 13 pages, 1 table. Structured narrative review
Abstract:Exam-style accuracy does not establish whether large language models (LLMs) reason well over clinical records. We define clinical reasoning as integrating and updating evidence across time and sources to form, revise and justify a patient’s problem representation and a defensible plan. This structured narrative review maps three literatures: medical education assessment instruments, clinical LLM benchmarks published from 2023 onwards, and general-domain methods for evaluating long-form generation. We examine six dimensions: problem representation, temporal synthesis, differential and management reasoning, counterfactual reasoning, calibrated uncertainty, and reasoning faithfulness. Preprints are included and flagged. No single instrument covers all six dimensions. Problem representation and differential or management reasoning are reasonably covered, although reliability varies by instrument and setting. TIMER-Eval targets temporal synthesis, and ER-Reason assesses sequential diagnostic belief updating. Dedicated uncertainty and counterfactual evaluations are emerging, but their applicability to longitudinal free-text reasoning remains limited. Factual completeness is well theorised in general-domain evaluation, with early clinical evidence of important omissions. Faithfulness remains the weakest dimension, with one identified clinical causal-ablation study on multiple-choice questions. Existing tools should be combined through binary rubric items, separate completeness and correctness scores, case-specific importance weighting with non-compensable safety caps, temporal order-consistency checks, and chance-corrected reliability reporting. Further design work is needed for calibrated uncertainty, counterfactual reasoning and faithfulness over longitudinal free-text records. This review provides a design rationale, not a validated instrument. Comments: 13 pages, 1 table. Structured narrative review Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI) Cite as: arXiv:2610.01938 [cs.CL] (or arXiv:2610.01938v1 [cs.CL] for this version) https://doi.org/10.48550/arXiv.2610.01938 Focus to learn more arXiv-issued DOI via DataCite (pending registration) Submission history From: Zhangshu Joshua Jiang [view email] [v1] Thu, 1 Oct 2026 16:07:12 UTC (26 KB)
[NLP-22] Cross-Lingual Alignment for Decoder-Only Models using MoE Routers
【速读】: 该论文旨在解决在仅解码器架构的大语言模型(LLM)中,由于多语言分词方式不一致导致难以显式对齐跨语言表示的问题。尽管现有研究指出,更高的跨语言表示对齐有助于提升跨语言迁移性能,但传统对比学习方法在解码器-only模型中的应用受限。本文提出一种新范式,将混合专家模型(MoE)的路由输出(router outputs)作为对齐目标,而非直接作用于隐藏状态。这一设计利用路由输出在多个词元上的聚合特性,实现了更可靠的序列级跨语言比较。在四个开源MoE模型上的受控持续预训练实验表明,引入该路由损失不仅提升了跨语言表示对齐程度,还显著改善了多语言任务表现,验证了基于MoE路由对齐的跨语言对比学习在现代大模型架构下的有效性与潜力。
链接: https://arxiv.org/abs/2610.01921
作者: Lucas Bandarkar,Clark Peng,Ahmed Haj Ahmed,Aditi Khandelwal,Nanyun Peng
机构: University of California, Los Angeles; Haverford College; MILA - Quebec AI Institute, McGill University
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注:
Abstract:Cross-lingual contrastive learning has been a core component of multilingual encoder training, but the ability to explicitly align representations is not possible in decoder-only LLMs because of varying multilingual tokenization. However, a growing amount of research suggests that even in LLMs, higher cross-lingual representational alignment leads to improved cross-lingual transfer. In this paper, we propose a novel approach to reimagine cross-lingual contrastive learning given the architectural constraints of modern LLMs. Rather than applying an auxiliary alignment loss on hidden states, we propose using the outputs of the mixture-of-experts (MoE) routers as the target for alignment. Router outputs lend themselves better to pooling over many tokens, enabling more reliable cross-lingual comparisons at the sequence-level. Controlled continual pre-training experiments on four open-source MoEs show that incorporating this routing loss also aligns the underlying hidden representations across languages. Most importantly, this loss improves multilingual performance on our diverse evaluation suite, demonstrating the potential of cross-lingual MoE router alignment.
[NLP-23] MoLE: Mixture of Latent Experts for Complementary Visual Reasoning
【速读】: 该论文旨在解决现有生成式视觉-语言模型中潜在视觉推理(latent visual reasoning)存在的冗余表示问题:当前方法允许多个潜在标记通过共享的值投影访问相同视觉证据,导致缺乏互补性信息提取机制,单纯增加潜在预算只会产生重复的表征。其解决方案的关键在于提出一种潜专家混合框架(MoLE, Mixture of Latent Experts),通过显式分离不同潜在视觉专家在证据提取阶段的观测内容与变换方式,强制每个潜在标记作为“专业化视觉专家”专注于提取不同的视觉信息。具体而言,MoLE在证据提取阶段隔离各潜专家,并引入专用的潜总结专家聚合其互补表征;采用两阶段训练策略,先引导视觉证据经由潜路径处理,再恢复直接视觉访问,无需预设专家角色或中间视觉目标。实验表明,MoLE在五个视觉推理基准上平均得分达78.6,显著优于数据匹配的监督微调(+4.9)和现有最强潜推理基线(+3.6),且表征分析显示更低的潜状态相似性与更丰富的视觉注意力模式,掩蔽潜路径后性能下降9.2,验证了专业化潜计算的有效性。
链接: https://arxiv.org/abs/2610.01917
作者: Yingcheng Liu,Tianyi Jiang,Yujuan Ding,jiangbo Ai,Xun Jiang,Guoqing Wang,Wei Ye,Yi Bin
机构: Tongji University (同济大学); Hong Kong Polytechnic University (香港理工大学); Alibaba Group (阿里巴巴集团); University of Electronic Science and Technology of China (电子科技大学)
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注:
Abstract:Latent visual reasoning equips vision–language models with continuous intermediate states that can process visual evidence without explicit textual reasoning traces or repeated image operations. However, existing methods often allow multiple latent tokens to access the same visual evidence through shared value projections, providing no mechanism for them to extract complementary visual information; simply increasing the latent budget can therefore yield redundant latent representations. We argue that effective latent reasoning should encourage different latent tokens to extract complementary visual information, and thereby act as specialized visual experts. Based on this insight, we propose MoLE, a Mixture of Latent Experts framework that controls both what visual evidence each latent visual expert observes and how it transforms that evidence. MoLE isolates latent visual experts during evidence extraction and uses dedicated latent summary experts to aggregate the complementary representations of latent visual experts. A two-stage training pipeline first forces visual evidence through this latent pathway and then restores direct visual access, requiring neither predefined expert roles nor intermediate visual targets. Across five visual reasoning benchmarks, MoLE achieves an average score of 78.6, outperforming data-matched supervised fine-tuning by 4.9 and the strongest evaluated latent visual reasoning baseline at the same latent budget by 3.6. Representation analyses show lower latent-state similarity and more diverse visual attention, while masking the latent pathway reduces average performance by 9.2. These results demonstrate that specializing latent computation is more effective than merely increasing the number of latent tokens.
[NLP-24] Stochastic Rounding in Low-Precision Transformer Inference: A Variable-Precision Emulation Study of a Small GPT -2
【速读】: 该论文旨在解决低精度Transformer推理中应采用随机舍入(Stochastic Rounding, SR)还是向最邻近舍入(Round-to-Nearest, RN)这一关键问题,核心在于揭示不同网络位置对舍入策略的敏感性差异。其解决方案的关键在于提出一种可扩展至任意虚拟精度的变精度随机舍入(Variable-Precision Stochastic Rounding, VPSR)算法,并基于此构建了两个互补的分析框架:一是针对线性投影的随机前向误差界分析,表明在减少长度为 $ n $ 时,SR 的误差增长为 $ O(\sqrt{n} u) $,而 RN 为 $ O(n u) $,该差距在低精度下显著放大,尤其在长多层感知机(MLP)下投影阶段最为明显;二是对输出softmax层交叉熵损失变化的二阶分解,揭示出噪声性质的差异——MLP中的噪声主要表现为均匀的logit偏移,对softmax具有不变性,因此SR的方差被大幅抵消;而注意力头处的噪声非均匀分布于词表空间,无法被抵消。实验结果在DistilGPT-2模型上验证了理论预测,在6位尾数精度下,将SR用于MLP使困惑度仅上升至1.15倍全精度基准,而RN则达2.21倍;反之,在语言模型头处,因SR引入非均匀方差,性能劣于确定性的RN。最终,通过混合精度配置(MLP输出使用6位精度,分别采用SR与RN),可将困惑度控制在全精度参考值的1.10倍以内,相较相同位宽的全RN配置降低28%。
链接: https://arxiv.org/abs/2610.01889
作者: Yohan Chatelain(1),Pablo de Oliveira Castro(2) ((1) Krembil Centre for Neuroinformatics, CAMH, Toronto, Canada, (2) Universite Paris-Saclay, UVSQ, LI-PaRAD, Versailles, France)
机构: 未知
类目: Machine Learning (cs.LG); Computation and Language (cs.CL)
备注: 35 pages, 10 figures, 4 tables. Code and evaluation pipeline available at this https URL and archived on Zenodo at this https URL
Abstract:Should low-precision transformer inference use stochastic rounding (SR) or round-to-nearest (RN)? The answer depends on where in the network you look. We isolate this effect by holding the numerical format fixed and varying only the rounding rule at individual operation sites. To enable experiments at freely chosen precisions, we extend the PRISM vectorized rounding library to arbitrary virtual precision via a variable-precision stochastic rounding (VPSR) algorithm, proving that the rounding decision is evaluated exactly in hardware floating point. We develop two analyses providing complementary insight into this site-level trade-off. First, a probabilistic forward-error bound for linear projections shows that SR’s error envelope grows as O(\sqrtn u) in reduction length n , versus O(n u) for RN, a gap that widens rapidly at low precision and is most pronounced in the long multilayer perceptron (MLP) down-projection. Second, a second-order decomposition of expected cross-entropy loss change at the output softmax into signed drift, drift curvature, and a Fisher-weighted variance penalty reveals why the two sites behave oppositely: MLP noise is predominantly a uniform logit shift to which softmax is invariant, so SR’s variance is largely discounted; head noise is non-uniform across the vocabulary and is not. On DistilGPT-2 at t=6 significand bits, observations match theory: SR in the MLP raises perplexity to 1.15x the full-precision reference, versus 2.21x for RN. At the language-model head, the ordering reverses because SR introduces non-uniform variance, whereas deterministic RN carries none. In a mixed-precision configuration (MLP output at t=6 ), assigning SR to the MLP and RN to the head brings perplexity within 1.10x of the full-precision reference, a 28% reduction over matched-bit RN. Comments: 35 pages, 10 figures, 4 tables. Code and evaluation pipeline available at this https URL and archived on Zenodo at this https URL Subjects: Machine Learning (cs.LG); Computation and Language (cs.CL) ACMclasses: G.1.0; I.2.6; I.2.7 Cite as: arXiv:2610.01889 [cs.LG] (or arXiv:2610.01889v1 [cs.LG] for this version) https://doi.org/10.48550/arXiv.2610.01889 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[NLP-25] Detecting Inconsistencies in Model Specifications with LLM -as-Verifier Reasoning
【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)模型规范(Model Specifications)中潜在的不一致性问题,即在自然语言描述的规范条目之间可能存在语义冲突,导致同一情境下无法生成同时满足所有规则的响应。传统方法如形式化建模易丢失语义细微差别,而基于行为测试的方法难以区分规范缺陷与模型行为差异。其解决方案的关键在于提出VeriSpec——一种直接通过审计规范文本本身来检测不一致性的新方法。VeriSpec的核心思想是在保持规范自然语言表达的同时,利用大语言模型(LLM)作为验证器,通过提取上下文感知的结构化规则、构建主题引导的图结构以聚类同层级相关规则,并运用基于LLM的推理机制识别语义矛盾。在OpenAI模型规范上的应用表明,VeriSpec共提取405条规则,手动验证出5处不一致,均已报告并获开发团队积极回应;相较于五种基线方法,VeriSpec在已验证不一致性的检出数量、精度(38.5%)以及每项验证成本(11.12)方面均表现最优,证明了直接规范审计在模型对齐评估中的可行性与有效性,能够从源头上捕捉缺陷,避免其影响模型训练与部署。
链接: https://arxiv.org/abs/2610.01847
作者: Zichen Xie,Mrigank Pawagi,Lize Shao,Yang Hu,Wenxi Wang
机构: University of Virginia (弗吉尼亚大学); University of Pennsylvania (宾夕法尼亚大学); The University of Texas at Austin (德克萨斯大学奥斯汀分校)
类目: oftware Engineering (cs.SE); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注:
Abstract:Model specifications define how large language models (LLMs) should behave, guiding alignment training, inference-time behavior, and evaluation. Yet these specifications may themselves contain defects: two individually reasonable principles may prescribe incompatible behavior when applied to the same situation, leaving no response that satisfies both. Detecting such inconsistencies is challenging. Formalizing natural-language specifications risks losing subtle distinctions, while behavior-based testing cannot reliably distinguish specification defects from differences in model behavior. We introduce VeriSpec, the first approach to directly detect inconsistencies in model specifications by auditing the specification text itself. Our key insight is to preserve the specification in natural language while using an LLM as a verifier. VeriSpec extracts structured, context-aware rules, constructs a topic-guided graph to cluster behaviorally related rules at the same authority level, and applies LLM-as-verifier reasoning to detect inconsistencies. Applying VeriSpec to the OpenAI Model Spec, we extract 405 rules and manually validate five inconsistencies, all reported to its developers, who responded positively and have initiated internal discussions. Compared with five baselines, VeriSpec identifies the most validated inconsistencies, achieves the highest precision (38.5%), and incurs the lowest cost per validated inconsistency ( 11.12). These results establish direct specification auditing as a practical complement to behavioral alignment evaluation, catching defects at the source before they shape any model. The code is available at this https URL.
[NLP-26] Beyond Decodability: Do Acoustic Factors Drive Predictions in Speech-Based Alzheimers Assessment?
【速读】: 该论文旨在解决生成式语音分析模型在阿尔茨海默病(Alzheimer’s Disease, AD)诊断中对录音条件敏感性的问题,即预训练自监督学习(Self-Supervised Learning, SSL)模型所提取的声学表征是否受录音噪声、混响等非语义因素系统性干扰,从而影响其临床预测的可靠性。其解决方案的关键在于引入多层次的干预与分析框架:通过控制噪声和混响等声学扰动,结合层间线性解码、输入空间与表征空间的干预手段,以及几何对齐分析,区分声学特征的可解码性与对AD分类决策的实际影响。研究发现,尽管噪声在原始数据中未表现出显著的诊断组差异,但其仍能系统性地改变所有三种大型SSL模型的AD预测结果,且这种影响与分类器决策方向高度一致,具备可逆性和泛化能力。这一结果表明,高预测性能与无显著诊断组差异并不能保证模型的鲁棒性。因此,论文主张将基于干预的鲁棒性测试作为可信临床语音模型的标准评估流程。
链接: https://arxiv.org/abs/2610.01846
作者: Serli Kopar,Alkis Koudounas,Roshan P. Rane,Sam Gijsen,Paula A. Perez-Toro,Kerstin Ritter
机构: 未知
类目: ound (cs.SD); Computation and Language (cs.CL); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
备注:
Abstract:Speech-based Alzheimer’s disease (AD) assessments increasingly rely on pretrained self-supervised learning (SSL) models that learn acoustic representations directly from raw audio, exposing the model to recording factors. We ask whether such factors are merely encoded in SSL representations or can systematically alter predictions. Using ADReSSo and three large SSL backbones, we apply controlled noise and reverberation interventions to participant-speech-only, non-speech, and full-recording audio. We combine layer-wise linear decoding, input- and representation-space interventions, and geometric alignment analysis to distinguish acoustic decodability from influence on AD prediction. Our results show that controlled acoustic interventions alter AD predictions across all three SSL backbones. Noise, despite showing no significant diagnostic-group difference in the original data, produces the strongest intervention effects. Importantly, these effects are systematically structured relative to the classifier’s decision direction, replicate on the held-out test set and reverse when the representation-space intervention direction is reversed. Together, these findings show that high predictive performance and the absence of a significant diagnostic-group difference in a measured acoustic factor are not sufficient for robustness. We argue that intervention-based robustness tests should become standard for trustworthy clinical speech models.
[NLP-27] he Asymptotics of Language Model Alignment with Memory
【速读】: 该论文旨在解决生成式语言模型(Language Model, LM)对齐过程中,如何在不依赖模型分布知识且计算成本较低的前提下,实现对齐后模型输出与原始模型输出在概率分布上保持接近,同时提升期望奖励的问题。现有方法如基于KL约束的强化学习(KL-constrained RL)虽能保证对齐效果,但需精确知晓模型分布且计算开销大;而“最佳n选一”(best-of-n)算法虽仅需从模型中采样,却缺乏理论保障。此前研究(Yang et al.)已证明,在独立同分布(i.i.d.)假设下,两种方法在序列长度趋于无穷时渐近等价。然而,真实语言模型的输出通常具有记忆性,不符合i.i.d.假设。本文的关键突破在于将该渐近等价性推广至马尔可夫(Markovian)结构的输出序列场景,从而更贴近实际应用。此外,针对有限长度输出(特别是单个token,m=1)的情形,本文首次给出了使两种对齐方法所生成分布间KL散度为零的完整条件——即对语言模型分布和奖励函数的充分必要刻画,解决了Yang等人提出的核心开放问题。
链接: https://arxiv.org/abs/2610.01828
作者: Haricharan Balasundaram,V. Arvind Rameshwar
机构: Georgia Institute of Technology (佐治亚理工学院); IIT Madras (印度理工学院马德拉斯分校)
类目: Computation and Language (cs.CL); Information Theory (cs.IT)
备注:
Abstract:Language model (LM) alignment broadly aims to perturb a given LM Q into an aligned LM q such that i) the outputs produced by q and Q are ‘close’ in probability, ii) q has a higher expected reward than Q . Two common techniques for LM alignment are: KL-constrained RL, which requires knowledge of the LM distribution and is computationally expensive, and the best-of- n algorithm, which requires only sampling from the LM. The work of Yang et al. established asymptotic closeness between the distributions produced by the two alignment methods for an m --length i.i.d. token sequence output by the LM, in the limit as m increases to infinity. However, the i.i.d. assumption is not representative of practical LMs, whose output sequences often have memory. In this paper, we extend the asymptotic closeness result to the case when the m --length token sequence outputted by the LM is Markovian. Further, for finite-length output sequences – particularly, when m=1 – we provide a complete characterization of LM distributions and reward functions for which the KL-divergence between the distributions produced by the two alignment methods is zero – a question first posed in Yang et al.
[NLP-28] Beyond Linear Concepts: Discovering and Aligning Non-Linear Concept Manifolds in Large Language Models
【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)内部表示中信息处理机制的可解释性问题,尤其关注其内部标记(token)表征的几何组织结构。传统机制可解释性(Mechanistic Interpretability, MI)方法通常依赖线性假设来提取概念,但这一假设受到非线性特征流形证据的挑战。为此,本文提出突破线性约束的解决方案:将计算机视觉中的非线性多维概念发现(Non-Linear Multi-Dimensional Concept Discovery, NLMCD)方法迁移至LLM的标记级激活层面,将概念建模为低维流形,从而更准确地捕捉复杂的非线性结构。其核心创新在于引入一种基于概念的对齐(Concept-based Alignment, CBA)评分,该评分是一种广义的兰德指数(Rand Index),能够在不进行显式特征匹配的情况下衡量概念流形之间的几何邻近性。这一方法不仅提升了对层间与模型间概念结构变化的敏感度,还揭示了从语法主导到语义-语法混合、输出导向增强的动态演化过程,并揭示了跨语言共享与跨模型家族对齐的训练依赖性特征。
链接: https://arxiv.org/abs/2610.01821
作者: Tido Specht,Elias Benedict Krey,Nils Neukirch,Nils Strodthoff
机构: University of Groningen(格罗宁根大学); Carl von Ossietzky University of Oldenburg(奥尔登堡卡尔·冯·奥西耶茨基大学)
类目: Machine Learning (cs.LG); Computation and Language (cs.CL)
备注: 24 pages, 13 figures. Code: this https URL
Abstract:Understanding information processing in large language models (LLMs) requires dissecting the geometric organization of their internal token representations. While existing mechanistic interpretability (MI) methods seek to extract concepts, they are constrained by a strong linearity assumption challenged by evidence of non-linear feature manifolds. We move beyond linear concepts by adapting Non-Linear Multi-Dimensional Concept Discovery (NLMCD) from computer vision to token-level LLM activations, modeling concepts as low-dimensional manifolds. To compare concept manifolds across layers and models, we introduce a concept-based alignment (CBA) score, a generalized Rand index that measures geometric proximity without explicit feature matching. Our analysis yields six key findings: (i) a neighboring-layer sanity check shows CBA is more sensitive than PCA- or CKA-based linear baselines; (ii) layer-by-layer alignment matrices reveal two block structures in intermediate and late layers, consistent across models and obscured by linear metrics; (iii) concept composition remains syntax-dominated through most of the network before giving way to increasingly mixed syntactic-semantic concepts in later layers, with increasing output-orientation toward the final layers; (iv) multilingual concept sharing between English and Mandarin is training-dependent rather than universal, strongest in Qwen, weaker in Llama, and absent in GPT-2; (v) inter-model alignment mirrors this structure, with strong correspondence between same-family Qwen models of different scale but weak alignment across model families; and (vi) across Tulu-3 training stages, alignment is highest between adjacent stages, with the largest shift between the base model and SFT, while subsequent preference-alignment stages (DPO, RLVR) leave early layers largely unchanged and RLVR mostly preserves DPO’s concepts in late layers.
[NLP-29] A Safe Prototype Is Not a Safety Direction: Reference Dependence and Prompt Confounds in Response-Safety Embeddings NEURIPS2026
【速读】: 该论文旨在解决生成式安全评估中一个关键问题:能否通过计算响应与已知安全响应均值嵌入(mean embedding)之间的余弦相似度,来有效衡量响应的安全性。这一方法被部分研究(如近期的“睡眠代理检测器”)所采用,但其核心假设——即正样本聚类中心能自然指示安全方向——尚未得到充分验证。本文的关键发现是,仅依赖正样本均值作为安全原型(safe prototype)存在根本性局限:虽然正样本确实位于编码空间中安全类别的近似中心,但该均值本身并不具备区分安全与不安全响应的方向性信息。实验基于两个提示控制的人工标注语料库和一个辅助的评审员标注对照集,使用四个冻结的文本编码器并结合按提示分组的划分策略进行审计。结果显示,在人工标注数据上,基于正样本均值的余弦相似度得分的受试者工作特征曲线下面积(ROC-AUC)仅为0.457–0.545,其中多个子集显著低于随机水平,仅一个高于随机;而显式构建的“安全减去不安全”参考向量(safe-minus-unsafe reference)则在相同嵌入空间下取得0.588–0.738的显著更高性能。在评审员标注的对照数据中,正样本原型甚至出现反向表现(AUC 0.358–0.405),而参考向量仍保持0.754–0.793的高区分能力。在经验证校准后的5%假安全率阈值下,参考向量在PKU-SafeRLHF和Aegis数据集上显著提升安全响应接受率,但在BeaverTails上未表现出一致性。进一步分析表明,即使在无标签的保留测试集中,仅用少量(5%)已标记不安全样本即可恢复大部分参考向量的有效排序性能,而当不安全样本数量增至80–634时性能接近最优。此外,仅基于提示标签的消融实验揭示了提示-标签组合可能对评估结果造成不可控的偏差。因此,本研究结论明确指出:一个类别均值仅表征位置,而非安全方向;唯有包含足够不安全样本质量的显式参考向量才能确定正确的分类方向。这是一项关于原始正样本均值规则的有限性研究,而非对所有单类学习方法或专用安全防护机制的否定。
链接: https://arxiv.org/abs/2610.01801
作者: Sahil Kadadekar
机构: New York University(纽约大学)
类目: Machine Learning (cs.LG); Computation and Language (cs.CL); Cryptography and Security (cs.CR)
备注: Accepted at the NeurIPS 2026 Workshop on Foundations of Language Model Security (FLMSec). 15 pages, 3 figures, 11 tables. Code, results, and a verifier are in the ancillary files
Abstract:Can response safety be scored by cosine similarity to the mean embedding of known-safe responses? A recent sleeper-agent detector proposes exactly this score, yet the raw positive-centroid rule is not identified: positive observations locate the safe class relative to an encoder origin, but do not determine which direction separates safe from unsafe responses. We audit the rule on two prompt-controlled, human-labeled corpora and one auxiliary jury-labeled source control, using four frozen encoders and prompt-grouped splits. On the human-labeled corpora the safe prototype reaches ROC-AUC 0.457-0.545, with two cells significantly below chance and one above, while an explicit safe-minus-unsafe reference reaches 0.588-0.738 on the same embeddings; on the jury control the prototype is inverted (0.358-0.405) and the reference reaches 0.754-0.793. At validation-calibrated 5% false-safe thresholds, the reference accepts more safe responses on PKU-SafeRLHF (0.153-0.263 versus 0.039-0.061 across encoders) and Aegis (0.189-0.291 versus 0.004-0.045), but not reliably on BeaverTails. A fully unlabeled held-out reference recovers part to most of the referenced ranking, much less when only 5% of the pool is unsafe, whereas 80-634 labeled unsafe responses recover most of it. Prompt-only ablations show that prompt-label composition can inflate uncontrolled evaluations. This is a bounded result about a raw positive centroid, not all one-class methods or safety-specialized guards. A class mean is a location, not necessarily a safety direction; a declared reference with enough unsafe mass identifies orientation.
[NLP-30] VETO: Video Efficient Token Optimization for Vision Language Models
【速读】: 该论文旨在解决视觉语言模型(Vision-Language Models, VLMs)在处理长视频时因视觉标记(visual tokens)数量随视频长度呈二次增长而导致的计算成本过高问题。现有单轴压缩方法虽能缓解此瓶颈,但受限于对空间与时间冗余独立处理,导致效率存在不可逾越的下限。本文提出一种无需训练的即插即用方案VETO(Video Efficient Token Optimization for Vision-Language Models),其核心创新在于双轴压缩:(i) 帧内压缩模块通过基于最优传输(optimal-transport)启发的匹配机制,合并语义相似的帧内标记;(ii) 帧间压缩模块识别并合并时间上冗余的帧。关键设计思想是分层排序——先进行空间维度压缩,显著降低后续全局时间匹配的计算开销,从而突破单轴方法的效率瓶颈,尤其在现代全融合注意力架构中优势更为明显。实验表明,VETO在保持或提升准确率的同时,推理速度最高提升45%(如在LLaVA-OneVision-7B上),在极端标记预算(仅10%)下仍以55.7%的准确率超越VFlowOpt(54.9%)、VisionZip(52.6%)和FastV(47.9%)。此外,VETO在LLaVA-OneVision、InternVL-2.5和LongVA等多个主流模型上均展现出通用性,且在零样本场景下准确率保持或提升。
链接: https://arxiv.org/abs/2610.01785
作者: Gueter Josmy Faure,Hao Ping Wang,Min-Hung Chen,Winston H. Hsu
机构: National Taiwan University (国立台湾大学); NVIDIA(英伟达)
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注:
Abstract:Processing long videos with Vision-Language Models (VLMs) is bottlenecked by the quadratic cost of visual tokens, making long-form inference prohibitively expensive. While single-axis compression methods mitigate this, they hit a hard efficiency floor because they treat spatial and temporal redundancy independently. We present VETO (Video Efficient Token Optimization for Vision-Language Models), a training-optional plug-in that eliminates this bottleneck through dual-axis compression: (i) an intra-frame compressor that merges semantically similar tokens within each frame via optimal-transport inspired matching, and (ii) an inter-frame compressor that identifies and merges temporally redundant frames. The key design insight is hierarchical ordering: by first compressing spatial dimensions, VETO drastically reduces the cost of subsequent global temporal matching, bypassing the efficiency wall of single-axis approaches, with an advantage that grows with modern fully-fused attention infrastructure. Empirically, VETO achieves up to 45% faster inference (e.g., on LLaVA-OneVision-7B) while preserving or improving accuracy. Under extreme token starvation (10% budget), VETO outperforms VFlowOpt (54.9%), VisionZip (52.6%), and FastV (47.9%) with 55.7% accuracy. We demonstrate universal applicability across LLaVA-OneVision, InternVL-2.5, and LongVA, with zero-shot accuracy preserved or improved in all cases.
[NLP-31] ask-Oriented Rank Adaptation for Continual Learning in Text Classification
【速读】: 该论文旨在解决持续学习(Continual Learning, CL)在文本分类任务中面临的两大核心问题:灾难性遗忘(catastrophic forgetting)与跨任务间的负迁移(negative transfer)。现有参数高效微调(Parameter-Efficient Fine-Tuning, PEFT)方法如低秩适配器(LoRA)虽能以紧凑的参数形式实现高效适应,但其适配器通常独立训练,难以在相关任务间复用。为此,论文提出面向任务的秩自适应(Task-Oriented Rank Adaptation, TORA),一种基于几何路由的框架,利用LoRA适配器的低秩结构,通过度量任务间的结构相似性,动态决定是引入最兼容的专家知识(增强,Boosting)还是隔离新任务以防止干扰(防护,Shielding)。TORA的关键在于仅依赖单一几何阈值,无需任务标识或预定义任务序列,即可实现对适配器的智能路由,在15个多样化的文本分类基准上均表现出色:对于结构相似的任务,性能优于孤立训练;对于结构差异大的任务,则有效避免了干扰且不损失精度,显著降低了训练时间。因此,其解决方案的核心是基于低秩结构的几何相似性判断机制,实现了动态、鲁棒且高效的适配器重用策略。
链接: https://arxiv.org/abs/2610.01702
作者: Rey Sanchez Lopez,Eduardo Morales Manzanares,Hugo Jair Escalante
机构: INAOE(国家光学研究所); The University of Texas at El Paso(德克萨斯大学埃尔帕索分校)
类目: Computation and Language (cs.CL)
备注: Preprint submitted to CIARP2026
Abstract:Continual learning (CL) in text classification faces two critical challenges: catastrophic forgetting and negative transfer across sequential tasks. Parameter-Efficient Fine-Tuning (PEFT) methods such as LoRA enable efficient adaptation by learning low-rank updates of the model parameters. However, these compact representations are normally trained in isolation, limiting their reuse across related tasks. We introduce Task-Oriented Rank Adaptation (TORA), a geometric routing framework that leverages the low-rank structure of LoRA adapters to decide whether to transfer knowledge from the most compatible expert (Boosting) or isolate the new task (Shielding) based on structural similarity. Evaluated across 15 diverse text classification benchmarks, TORA consistently avoids harmful routing decisions: compatible tasks exceed their isolated performance while reducing training time, and structurally distant tasks are protected from interference with no loss in accuracy. With a single geometric threshold and no reliance on task identities or predefined sequences, TORA provides a simple and effective approach for dynamic adapter routing in sequential text classification systems.
[NLP-32] Acmite: Mitigating Gender Bias in LLM s through Concept-Guided Mutual Information
【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在生成文本时复现社会刻板印象的问题,尤其针对现有去偏方法依赖显式偏见样本或预定义的群体-术语替换策略所导致的对表述敏感、难以捕捉跨情境共享的刻板印象概念,以及未能显式建模模型输出与潜在刻板印象概念之间统计依赖关系的局限性。其解决方案的关键在于提出一种轻量级的概念引导型去偏框架Acmite:首先将刻板印象表示为结构化的语义概念,并利用最大边际相关性(Maximal Marginal Relevance, MMR)选择具有多样性的核心概念用于去偏;其次,受互信息最小化启发,通过词元级别的KL散度近似建模输出与刻板印象概念间的依赖关系,同时保持任务语义不变;最后,采用仅在推理阶段激活的轻量级低秩适配器(LoRA),当输入与刻板印象概念高度相似时才启用去偏模块,否则直接使用原始模型。实验表明,Acmite在BBQ、CrowS-Pairs和StereoSet等多个评估基准上有效缓解了性别偏见,且在ARC-Challenge、GSM8K和PIQA等任务上保持了良好的通用能力。
链接: https://arxiv.org/abs/2610.01696
作者: Tian Lan,Xiaoqing Cheng,Han Zhang,Jiang Li
机构: Kyoto University (京都大学); Zhejiang University (浙江大学); Shanghai Jiao Tong University (上海交通大学); Inner Mongolia University (内蒙古大学)
类目: Computation and Language (cs.CL)
备注: 15 pages, 0 figures
Abstract:Large language models (LLMs) can reproduce social stereotypes from their training data, motivating extensive research on model debiasing. However, existing methods often rely on explicit biased examples or predefined group-term substitutions, making them sensitive to wording and less effective at capturing stereotype concepts shared across diverse contexts. More importantly, they typically suppress biased outputs without explicitly modeling the statistical dependence between model outputs and the underlying stereotype concepts. We propose Acmite, a lightweight concept-guided framework for targeted and selective debiasing. Acmite represents stereotypes as structured semantic concepts and uses maximal marginal relevance (MMR) to select diverse concepts for debiasing. Inspired by mutual information minimization, it approximates this dependence with token-level KL divergence while preserving task semantics. A lightweight LoRA adapter is trained with the base model frozen and activated at inference time only when the input is sufficiently similar to stereotype-related concepts; otherwise, the original model is used directly. We evaluate Acmite on BBQ, CrowS-Pairs, and StereoSet, and assess general capability preservation on ARC-Challenge, GSM8K, and PIQA. Experiments across three LLMs show that Acmite effectively mitigates gender bias across complementary evaluation formats while maintaining competitive performance on bias-unrelated tasks. Anonymous code and data are available at this https URL.
[NLP-33] Compound interpretation is based on analogy
【速读】: 该论文旨在解决如何从构词语素的语义信息中准确预测复合词整体语义的问题,这是计算词汇语义模型中的核心挑战。其解决方案的关键在于提出一种名为“复合类比模型”(Compound Analogy Model, CAM)的新方法:该模型通过将复合词两个构成成分的嵌入向量相加,并引入这两个成分所属复合词族的平均偏移向量(shift vectors),以捕捉语义空间中的局部类比结构。与依赖全局线性变换的CAOSS模型相比,CAM不依赖任何可调参数,且更强调局部语义关系的类比推广。在对汉语复合词的评估中,CAM在训练数据和未见数据上均表现出更高的预测准确性,仅在三字复合词上因构成成分家族规模小且分布不均而表现受限。此外,在基于频率划分的训练-测试分割下,CAM的优势依然显著。进一步的认知效度检验表明,由CAM推导出的语义指标能更优地预测双字复合词的视觉词义判断反应时,支持了“复合词意义的本质是局部类比泛化”而非“学习到的全局线性变换”的观点,证明类比语义结构为复合词理解提供了具有认知合理性的理论基础。
链接: https://arxiv.org/abs/2610.01688
作者: Tian Shen,Harald Baayen
机构: 未知
类目: Computation and Language (cs.CL)
备注:
Abstract:How compound meanings are best predicted from constituent meanings remains a central question in computational models of lexical semantics. Comparing different computational models provides a way to evaluate alternative accounts of how semantic information is combined during compound comprehension. We propose a new model, the Compound Analogy Model (CAM), that predicts a compound’s embedding by adding its constituent embeddings together with the average shift vectors of the two constituents’ compound families. The resulting model is parameter-free and exploits local analogical structure in the semantic space. We evaluated CAM against the CAOSS model on Mandarin Chinese compounds. CAM consistently achieved higher prediction accuracy than CAOSS on both training and held-out data, with the exception of three-character compounds, for which analogical generalization is constrained by both small constituent families and a pronounced imbalance in family size between the two constituents. The advantage of CAM remained when evaluation was based on frequency-defined train-test splits that better approximate generalization from familiar to novel compounds. To assess the cognitive plausibility of the two models, we further examined whether model-derived semantic measures predict visual lexical decision latencies for two-character compounds. Predictors derived from CAM provided improved prediction for response latencies compared to predictors derived from the CAOSS model. These findings indicate that compound meaning is better characterized as local analogical generalization than as the application of a learned global linear transformation, and demonstrate that analogical semantic structure provides a cognitively plausible basis for compound comprehension.
[NLP-34] Yo-ByT5: Efficient and High-Fidelity Diacritic Restoration for Yorùbá
【速读】: 该论文旨在解决约鲁巴语(Yorùbá)在自然语言处理(NLP)任务中因缺乏声调标记(diacritics)而导致的词汇歧义问题。由于约鲁巴语为声调语言,声调标记对语义至关重要,但实际书写中常被省略,严重影响下游NLP模型的性能。为此,本文提出了一种基于字节级的自动声调标记恢复(Automatic Diacritic Restoration, ADR)模型Yo-ByT5,其在ByT5-small基础上进行微调。解决方案的关键在于采用字节级编码策略,有效捕捉约鲁巴语的细粒度语言特征,同时在参数量仅为mT5-base一半的情况下,实现了与最强现有模型相当的性能:错误恢复率(DER)为10.14%,字符错误率(CER)为3.48%,且在文本保真度方面表现更优。研究还发布了训练代码与模型输出,并呼吁构建更大规模、专用于约鲁巴语的声调标记恢复基准数据集。
链接: https://arxiv.org/abs/2610.01634
作者: Ahmad Samuel Gali(1),Shamsuddeen Hassan Muhammad(2 and 3) ((1) University of Lagos, (2) Bayero University Kano, (3) Imperial College London)
机构: University of Lagos(拉各斯大学); Bayero University Kano(卡诺贝莱罗大学); Imperial College London(伦敦帝国学院)
类目: Computation and Language (cs.CL)
备注: 7 pages, 3 figures, 3 tables. Code and outputs: this https URL
Abstract:Yorùbá is a widely spoken tonal language that depends on diacritics to avoid lexical ambiguity. However, it is often written without these diacritics, thereby hindering downstream Natural Language Processing (NLP) tasks. In this paper, we introduce Yo-ByT5, a byte-level Automatic Diacritic Restoration (ADR) model fine-tuned from ByT5-small. We evaluate Yo-ByT5 alongside five publicly released Yorùbá ADR models and one open-weight large language model (LLM) on the YAD benchmark under a consistent protocol. Our results demonstrate that Yo-ByT5 matches the performance of the strongest existing model, mT5-base, with a DER of 10.14% and a CER of 3.48%. Furthermore, it exhibits superior text fidelity despite using approximately half the parameter count of mT5-base. We also release our training code and model outputs, as well as call for the development of a larger, purpose-built benchmark for Yorùbá diacritic restoration.
[NLP-35] What Makes Something Hard(er)? Explaining Question Difficulty in Natural Language
【速读】: 该论文旨在解决现有自动难度评估方法仅能提供单一数值描述,却无法揭示问题难度背后根本成因的问题。其核心挑战在于:如何从数据中自动识别并验证能够解释“为何某个问题比另一个更难”的可解释性因果因素。解决方案的关键在于提出一种数据驱动的方法,通过结合项目反应理论(Item Response Theory, IRT)对大规模语言模型(LLM)的作答行为进行难度估计,进而采样难易对比问题对,并利用大语言模型生成候选解释;这些解释在独立保留集上经过验证与筛选后,形成具有可解释性和预测能力的自然语言假设。实验表明,这些假设不仅能够以媲美甚至优于黑箱难度回归器的方式准确预测未见问题的难度,且作为辅助特征可进一步提升现有模型性能,说明其捕捉到了原始模型未能建模的难度信号。更重要的是,根据所发现的假设对题目进行编辑后,其测量难度向预期方向显著变化,证明所识别出的因素为具有因果有效性的难度决定因子,而非事后拟合的描述。因此,该方法将原本纯描述性的难度评分转化为可操作、可干预的因果性解释,实现了从“测什么”到“为什么”的范式跃迁。
链接: https://arxiv.org/abs/2610.01627
作者: Peng Cui,Qiaoyuan Zheng,Rudolf Debelak,Mrinmaya Sachan
机构: ETH Zurich(苏黎世联邦理工学院); EPFL(洛桑联邦理工学院)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注:
Abstract:Difficulty is one of the most fundamental properties of a question: it determines whether the question can meaningfully discriminate between models of differing ability. Although a variety of methods can now estimate or predict difficulty automatically, they yield only a single descriptive number, with no account of the underlying factors that make a question difficult in the first place. In this work, we propose a data-driven approach that automatically generates and validates natural-language hypotheses explaining what makes one question harder than another. We first estimate each item’s difficulty from the responses of a large pool of LLMs using Item Response Theory. We then sample contrasting sets of easy and hard questions and prompt an LLM to propose candidate explanations of the difference, which are subsequently validated and selected on held-out questions. Experimental results across three datasets spanning mathematical, logical, and commonsense reasoning show that our method produces interpretable and predictive hypotheses. On their own, they predict the difficulty of unseen questions competitively with, or better than, advanced black-box difficulty regressors; used as additional features, they further improve those regressors, implying that they discover difficulty signals that existing models fail to capture. Moreover, we demonstrate that editing questions according to a hypothesis can shift their measured difficulty in the expected direction, indicating that the discovered hypotheses are causally valid difficulty factors rather than post-hoc descriptions. Our approach thus turns a purely descriptive difficulty score into actionable statements.
[NLP-36] Can LLM s Reliably Annotate Bioassay Metadata to Improve Data Readiness? NEURIPS
【速读】: 该论文旨在解决分子性质预测领域中基础模型(foundation models)所面临的高质量数据准备问题,特别是公共数据库(如PubChem)和工业筛选数据库中存在的检测方法与实验格式等关键元数据标注缺失、不一致或混淆的问题。其核心解决方案在于评估生成式人工智能(Generative AI)——包括开源与专有大语言模型(LLM)——是否能够直接从实验文本中可靠地预测和审计这些元数据标注。研究发现,PubChem中约200万项生物活性测试的标注覆盖率极低:36%缺乏实验格式标注,89%缺少生物活性类型标注,99.9%未映射至生物活性本体(BAO)中的实验格式或检测技术术语。尽管在生化及细胞类实验格式和检测技术的召回率均达到至少0.96,但对低频类别存在更高分歧;经人工核查,多数分歧源于“银标准”来源之间的不一致性,而非模型本身错误。此外,一项定性研究显示,由LLM生成的证据可促使资深工业级专家修正自身标注,表明LLM具备识别潜在误标实验的能力。总体而言,专有与开源模型性能差异微小,研究结果表明LLM可有效支持大规模实验元数据的自动标注与审计,但需结合分类别可靠性评估与针对性的人工审查,方可确保其标注结果在下游机器学习流程中的可信性。
链接: https://arxiv.org/abs/2610.01616
作者: Laura van Weesep,Riccardo Tedoldi,Jens Sjölund,Hossein Azizpour,Susanne Winiwarter,Ola Engkvist,Jon Paul Janet,Samuel Genheden,Juan Viguera Diez
机构: AstraZeneca(阿斯利康); Uppsala University (乌普萨拉大学); KTH Royal Institute of Technology (皇家理工学院); Science for Life Laboratory (斯德哥尔摩, 瑞典); Chalmers University of Technology and University of Gothenburg (查尔姆斯理工大学和哥德堡大学)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Databases (cs.DB); Quantitative Methods (q-bio.QM)
备注: Accepted to the AIDaR workshop at NeurIPS
Abstract:The emergence of foundation models for molecular property prediction requires a high degree of AI data readiness, including reliable metadata annotation. However, both public repositories and industrial screening databases suffer from missing, inconsistent, or conflated assay annotations. In this work, we quantify the extent of missing annotations in PubChem for the BioAssay Ontology (BAO) assay format and physical detection method fields and investigate whether open-source and proprietary large language models (LLMs) can reliably predict and audit metadata annotations directly from the assay text. In our assessment, we found that the annotation coverage across PubChem’s \sim 2 million bioassays is critically sparse, 36% lacking an assay format, 89% a BioAssay type, and 99.9% any BAO-mapped assay format or detection technology term. This motivates the need for automated test-metadata curation. Using evaluation sets derived from PubChem and ChEMBL, we assess the agreement of seven open-source and proprietary LLMs with existing silver labels. Recall is at least 0.96 for biochemical and cell-based assay formats, with a similar pattern for detection technology, although disagreements increase on under-represented classes. Manual inspection shows that many of these disagreements trace back to inconsistencies between silver sources rather than to LLM error. Moreover, in a qualitative study with a senior industrial curator, LLM-generated evidence prompted the expert to revise some of their own labels, showing LLMs can flag potentially mislabeled assays. Across the study, performance differences between proprietary and open-source models were small. Together, these results suggest LLMs can support the large-scale annotation and auditing of assay metadata, though per-class reliability estimates and targeted human review remain necessary before such labels enter downstream ML pipelines.
[NLP-37] Which LLM to pick? Online Active Model Selection for Large Language Models
【速读】: 该论文旨在解决在流式数据处理场景中,如何高效、低成本地选择最优大语言模型(Large Language Models, LLMs)的问题。现有方法依赖于基准测试来评估模型性能,但这些指标仅能近似反映真实表现,且缺乏对实际应用场景的即时反馈。尽管使用人工标注(oracle annotations)可提供可靠性能评估,但在大规模应用中其成本过高且难以获取。为此,本文提出 ONLINE LLM PICKER,首个面向在线场景的主动模型选择框架。其核心在于:在有限的标注预算下,通过主动选择最具信息量的查询样本进行标注,以高效识别出在当前数据流中表现最佳或接近最优的模型。实验结果表明,该方法在10个数据集、超过130个模型上可将标注成本降低高达71.67%,同时在未标注数据上的序列生成任务中,使累计损失(regret)减少达2.51倍,证明其能在处理全部流式请求前即准确识别出最优模型,显著提升模型选择效率与实用性。
链接: https://arxiv.org/abs/2610.01592
作者: Alessandro Turrin,Patrik Okanovic,Torsten Hoefler,Nezihe Merve Gürel
机构: TU Delft(代尔夫特理工大学); ETH Zurich(苏黎世联邦理工学院)
类目: Computation and Language (cs.CL); Machine Learning (cs.LG)
备注:
Abstract:Large Language Models (LLMs) are increasingly applied to process streaming data, with practitioners relying on benchmarks to select the best model even though these signals only approximate real performance. While oracle annotations can provide reliable feedback, they are often costly and difficult to obtain at scale. To address this challenge, we propose ONLINE LLM PICKER, the first framework for active model selection for LLMs in online settings. Given an arbitrary stream of queries and a limited annotation budget, ONLINE LLM PICKER selects the most informative prompts for annotation to identify the best LLM among candidate models. Across multiple tasks including 10 datasets, for over 130 language models, we show that ONLINE LLM PICKER saves annotation cost by up to 71.67% while reliably identifying the best or near-best model for the stream. We also show that using the returned model for sequential generation on unannotated prompts across the stream reduces regret by up to a factor of 2.51x, indicating that ONLINE LLM PICKER can identify the best or near-best model well before processing all streaming prompts.
[NLP-38] AURAL: Adaptive Latent Reasoning with Joint Chunk for Speech Language Models
【速读】: 该论文旨在解决生成式语音语言模型在保持高质量交互体验时面临的两大挑战:模型智能性(model intelligence)与响应速度(fast response)难以兼顾的问题。具体而言,传统的显式思维链(Chain-of-Thought, CoT)虽能提升推理能力与语音理解性能,但其生成中间推理标记的过程显著增加延迟;而进一步引入细粒度声学线索描述则加剧了这一问题。现有隐式推理方法虽可降低计算开销,但仍受限于单路径监督机制和固定推理预算,无法根据任务难度自适应调整推理深度。为此,论文提出AURAL框架,其核心创新在于在潜在空间中建模多种可能的推理延续分布,并联合预测未来状态的多个片段,从而减少序列前向传播次数,有效降低推理延迟。为支持潜空间推理的初始训练,研究构建了AuralReason-683K数据集,包含约683,000条双语语音语句(总计约1,000小时),涵盖情感识别、共情对话与通用推理任务的简洁思维链标注。在此基础上,AURAL-RL通过强化学习机制探索更优推理路径,奖励简洁且高质量的推理过程,并动态调节每道题目的推理投入强度。实验表明,无论采用何种骨干网络,AURAL-RL均达到与CoT-RL相当的性能水平,且在多数指标上显著优于对应的监督基线模型。分析还发现,复杂度更高的问题会激发更多潜空间推理步骤。在Qwen2.5-Omni模型上,AURAL将首个回答标记的生成时间从1.22秒压缩至0.10秒,实现11.8倍加速,远优于直接回答(0.05秒)的延迟表现。
链接: https://arxiv.org/abs/2610.01560
作者: Yuxiang Wang,Kunyu Feng,Yuancheng Wang,Zihang Liu,Shengbo Cai,Qinke Ni,Wan Lin,Tao Feng,Yingda shen,Ming-Hao Hsu,Zhixian Zhao,Liqiang Zhang,Teddy Sun,Steve Yves,Zhizheng Wu
机构: The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)); Tencent Hunyuan(腾讯混元); Tsinghua University(清华大学); The Hong Kong University of Science and Technology(香港科技大学); Amphion Technology Co., Ltd.(安珀科技有限公司)
类目: Computation and Language (cs.CL); Machine Learning (cs.LG); Sound (cs.SD)
备注:
Abstract:Model intelligence and fast response jointly shape the quality of interaction with speech language models, yet remain difficult to achieve together. Explicit chain-of-thought (CoT) improves reasoning and audio understanding, but generating intermediate reasoning tokens delays responses. Describing fine-grained acoustic cues further lengthens CoT and increases latency. Latent reasoning can reduce this overhead, yet existing methods often trail CoT and remain limited by single-path supervision and reasoning budgets that do not adapt to problem difficulty. We introduce AURAL, which models a distribution over multiple plausible reasoning continuations in latent space and jointly predicts chunks of future states to reduce sequential forward passes and reasoning latency. To provide initial supervision for latent reasoning, we construct AuralReason-683K: 683K bilingual speech utterances (about 1,000 hours) with concise CoT for emotion recognition, empathetic dialogue, and general reasoning. AURAL-RL then explores beyond these traces, rewarding concise reasoning that yields high-quality answers and adapting reasoning effort to each problem. Across two backbones, AURAL-RL achieves performance comparable to CoT-RL, with larger gains over the respective supervised checkpoints on most metrics. Analysis further shows that harder questions elicit more latent reasoning steps. On Qwen2.5-Omni, it reduces time to the first answer token by 11.8x, from 1.22 to 0.10 s, versus 0.05 s for direct answering.
[NLP-39] QK-Wanda: Coupling Queries and Keys for Unstructured Pruning
【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在权重剪枝过程中忽视查询(Query)与键(Key)之间相互作用的问题。传统方法如Wanda仅独立评估线性投影中的权重重要性,忽略了查询与键通过点积运算产生的动态交互关系,导致剪枝后模型性能下降。其解决方案的关键在于提出QK-Wanda,一种基于未掩码前RoPE(pre-RoPE)重建目标下,联合考虑查询与键权重删除成本的剪枝方法。该方法通过引入对偶投影信息(即用键的删除成本评估查询权重,反之亦然),使查询与键共享剪枝预算,并采用闭式解析解计算得分,无需梯度或权重更新。实验表明,相较于Wanda,QK-Wanda在50%稀疏度下平均降低60%的QK重建误差,在80%稀疏度下降低45%;在Llama 2 70B上,下游任务的WikiText-2和C4困惑度分别降低20.3%和13.5%,零样本准确率提升5.94个百分点。然而,部分模型(如Llama-3.1-70B)虽重建误差更低,但困惑度反而上升,揭示了局部重建误差作为模型质量预测指标的局限性。
链接: https://arxiv.org/abs/2610.01554
作者: Ivan Ilin,Peter Richtárik
机构: KAUST(沙特国王科技大学)
类目: Machine Learning (cs.LG); Computation and Language (cs.CL)
备注: 81 pages, including appendices
Abstract:Wanda (Sun et al., 2024) prunes large language models by scoring weights independently within each linear projection, although queries and keys interact through dot products. We introduce QK-Wanda, which scores query and key weights by their individual deletion costs under an unmasked pre-RoPE reconstruction objective. It augments Wanda scores with information from the opposite projection (keys for query weights, and queries for key weights), allowing both projections to share a pruning budget. Its closed-form scores require no gradients or weight updates; full pruning takes 1.3% longer than Wanda on A100 and 3.1% longer on H200 with the calibration used in our main experiments. We evaluate QK-only pruning across 15 models from TinyLlama, Llama 2, Llama 3, and Qwen2.5, spanning 0.5B-72B parameters. Relative to Wanda, QK-Wanda reduces QK reconstruction error by an average of 60% at 50% sparsity and 45% at 80%. Downstream gains depend on the model. At 80% sparsity on Llama 2 70B, WikiText-2 and C4 perplexity decrease by 20.3% and 13.5%, while mean zero-shot accuracy rises by 5.94 percentage points. Qwen2.5-72B also improves, but Llama-3.1-70B has substantially higher perplexity despite lower reconstruction error. These results show both the promise of coupled pruning criteria and the limits of local reconstruction as a predictor of model quality.
[NLP-40] How the Audit Rule Shapes Faithful Factor Explanations in LLM s
【速读】: 该论文旨在解决在有限验证预算下,大语言模型(Large Language Models, LLMs)对输入因素影响输出的解释(即因子级影响报告)难以确保真实性的核心问题。其关键挑战在于:当验证过程依赖于模型自身报告的影响程度时,模型存在“抑制报告”(suppression incentive)的动机——即报告重要因子会使其更易被后续反事实扰动检验,从而暴露估计噪声并可能遭受惩罚。为应对这一激励扭曲,论文提出将验证机制设计为与报告无关(report-independent)或引入小比例的报告无关审计基线(floor),以切断“因报告重要性而被重点审查”的反馈路径。通过构建验证博弈框架,并以反事实布莱尔评分(Counterfactual Brier Score, CBS)为例进行实证评估,研究发现合成理性代理的行为完全符合理论预测,真实大语言模型在显式激励下也表现出一致的行为模式。因此,该研究的关键解决方案是:在部分验证场景中,因子级解释系统必须包含报告无关的审计成分,以防止模型通过低估关键因素来规避审查,从而保障解释的可信度与真实性。
链接: https://arxiv.org/abs/2610.01514
作者: Taolin Zhang,Hanyu Wang,Jiuheng Wan,Tingyuan Hu,Chengyu Wang
机构: Hefei University of Technology (合肥工业大学); East China Normal University (华东师范大学); Alibaba Cloud Computing (阿里巴巴云)
类目: Computation and Language (cs.CL)
备注:
Abstract:Large language models are often asked which input factors influenced their outputs. For structured inputs, such reports can be checked by counterfactual perturbation, but each factor must be queried multiple times to estimate its effect, so verification is usually budget-limited. We study how this limited-budget setting changes the incentive to report factor-level influence truthfully. We formalize the interaction as a verification game and show that proper scoring alone is not enough when auditing depends on the report: report-dependent auditing creates a suppression incentive, because factors reported as important are more likely to be checked and penalized for estimation noise. In contrast, report-independent auditing, or a mixed rule with a small report-independent floor, removes this channel and makes truthful reporting preferable to full suppression. We instantiate the framework with the Counterfactual Brier Score (CBS) and evaluate its predictions on four NLP benchmarks. A synthetic rational agent matches the theoretical prediction exactly, and real LLMs follow the same incentives when they are made explicit. The main design implication is simple: under partial verification, factor-level explanation systems should include a report-independent audit component so that under-reporting cannot be used to avoid scrutiny.
[NLP-41] GAW-PO: Preference Optimization with Gradient-Aligned Token Weights
【速读】: 该论文旨在解决现有偏好优化方法(如直接偏好优化,DPO)在训练过程中对拒绝响应中的所有词元(token)施加统一负向信号的问题。由于自回归语言模型是逐词元优化的,而传统DPO方法将整个拒绝响应视为整体进行惩罚,导致即使某些词元可能对优选响应具有潜在价值,也会被错误地惩罚,从而造成梯度更新方向的冲突与信息损失。其解决方案的关键在于提出一种梯度对齐的词元重加权方法——GAW-PO,通过估计每个拒绝词元的梯度与优选行为更新方向之间的对齐程度,动态调整其负向贡献:若某词元的梯度与优选方向高度一致,则给予较弱的惩罚;反之则保留较强惩罚。该方法实现了更精细的词元级信用分配,有效缓解了拒绝响应中“有用”词元被误罚的问题。实验表明,GAW-PO在11个涵盖数学、推理、编程和问答任务的基准上均取得最优平均性能,相较于标准DPO提升0.97分,优于最强基线0.65分;且在强偏好优化条件下(即正则化参数β减小)表现出更强鲁棒性,验证了梯度对齐机制在提升训练稳定性和效率方面的有效性。
链接: https://arxiv.org/abs/2610.01511
作者: Andreea Dutulescu,Stefan Ruseti,Mihai Masala,Traian Rebedea,Mihai Dascalu
机构: National University of Science and Technology POLITEHNICA Bucharest(布加勒斯特理工大学); NVIDIA(英伟达)
类目: Computation and Language (cs.CL)
备注:
Abstract:Most preference optimization methods, such as Direct Preference Optimization (DPO), apply preference supervision at the response level, although autoregressive language models are optimized token by token. As a result, all tokens in a rejected response contribute to the negative training signal, including tokens that may encode behavior that is useful for the preferred response. We introduce GAW-PO, a gradient-aligned token reweighting method for DPO that estimates, for each rejected token, whether penalizing it would interfere with the preferred update directions. Tokens whose gradients are strongly aligned with the preferred behavior receive a weaker negative contribution, while conflicting tokens retain a stronger penalty. Our method achieves the highest average performance among the evaluated preference-optimization methods, improving by 0.97 points over standard DPO and 0.65 points over the strongest competing baseline across 11 benchmarks spanning mathematics, reasoning, coding, and question answering. We further show that gradient-aligned weighting is substantially more robust to aggressive preference optimization: as the DPO regularization parameter \beta decreases, standard DPO degrades sharply, whereas GAW-PO continues to improve. These results suggest that accounting for the interaction between rejected-token updates and preferred behavior provides an effective form of token-level credit assignment for preference optimization.
[NLP-42] OverAct: Measuring and Mitigating Proactive Over-Authorization in LLM Tool-Calling Agents
【速读】: 该论文旨在解决大语言模型(LLM)代理在具备工具调用能力时,存在主动过度授权(proactive over-authorization)的问题,即模型可能在未明确请求的情况下访问超出任务所需范围的私有数据或外部服务,从而带来隐私泄露风险。与文件系统级编程代理不同,此类风险的核心在于对敏感数据的非必要访问。为系统评估该问题,研究提出OverAct——一个覆盖八个隐私敏感领域的可控基准测试,具有确定性且无需人工评判的评分机制,并构建了一套解释性决策理论框架,得出三项可验证预测:请求具体性是影响过度授权严重程度的最强预测因子,过度授权随工具池规模呈次线性增长,解码温度对行为影响较小。这些模式支持“成本不对称”解释,表明过度授权主要源于模型内在的结构化决策倾向,而非采样随机性。为此,研究进一步提出SelfAudit,一种零样本、推理时的自检方法,通过生成基于请求的合理性论证,在执行前过滤无正当理由的工具调用。消融实验表明,显式过滤是缩小作用范围的关键机制;在不依赖标注知识的前提下,SelfAudit可将面向隐私的冗余调用降低43%。
链接: https://arxiv.org/abs/2610.01508
作者: Taolin Zhang,Jiuheng Wan,Hanyu Wang,Tingyuan Hu,Chengyu Wang
机构: Hefei University of Technology(合肥工业大学); East China Normal University(华东师范大学); Alibaba Cloud Computing(阿里云计算)
类目: Cryptography and Security (cs.CR); Computation and Language (cs.CL)
备注:
Abstract:LLM agents with tool-calling capabilities can access external services and private user data, but they may retrieve more information than a user’s request explicitly requires. We study this behavior in structured tool-calling agents and term it proactive over-authorization. This setting differs from filesystem-level coding agents because the main risk is unnecessary access to private data. We introduce OverAct, a controlled benchmark spanning eight privacy-sensitive domains with deterministic, judge-free scoring, together with an interpretive decision-theoretic framework that yields three testable predictions. Across seven models from four families, all models significantly exceed authorized scope. Request specificity is the strongest predictor of severity, over-authorization grows sublinearly with tool-pool size, and decoding temperature has little effect. These patterns are consistent with a cost-asymmetry account, suggesting that over-authorization arises more from structural decision tendencies than from decoding randomness. We also propose SelfAudit, a zero-shot inference-time method that generates request-grounded justifications and filters unjustified calls before execution. Ablation shows that explicit filtering is the main driver of scope reduction. SelfAudit reduces privacy-oriented excess by 43% without oracle knowledge.
[NLP-43] No Model Required: Text Entropy Rate Filtering Mitigates Iterative Fine-Tuning Collapse NEURIPS2026
【速读】: 该论文旨在解决生成式人工智能(Generative AI)在基于合成数据的迭代微调过程中出现的“模型坍缩”(model collapse)问题,即随着训练迭代的进行,模型输出多样性逐渐下降,罕见模式被逐步丢失,表现为短语级别的重复。现有缓解方法通常依赖于模型的对数概率(log-probabilities)、外部参考源或持续访问真实人类数据,存在资源消耗高或可行性受限的问题。本文提出一种基于数学信息论的新方法:采用非参数化的Kontoyiannis熵率估计器 $ h_k ,该方法仅通过原始文本中的匹配长度统计量计算,无需任何模型假设或参数化结构。实验表明,在单一谱系的六代QLoRA微调设置下, h_k $ 作为文本多样性度量的训练数据过滤器显著优于依赖对数概率的基线方法——后者在所有指标上均未表现出显著的多样性提升($ p > 0.23 $),而 $ h_k $ 过滤则实现了 +42% 的唯一三元组、+30% 的词汇量以及 -19% 的重复率(均 $ p < 0.001 )。进一步验证显示, h_k $ 在跨领域(4个领域)、多温度设置、双生成-评分模型组合及1520份生成文档中均表现出良好的熵率代理能力($ \beta = 0.924, R^2 = 0.746 )和模型坍缩检测性能( \rho = +0.454, p < 0.0001 $)。研究结果表明,基于信息论的方法在缓解模型坍缩方面具有高效性,并为维持多智能体系统的多样性提供了新的技术路径。
链接: https://arxiv.org/abs/2610.01493
作者: Lewis Mitchell
机构: Adelaide Data Science Centre; School of Mathematical Sciences; Adelaide University (阿德莱德大学); Adelaide SA 5005, Australia
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Information Theory (cs.IT); Data Analysis, Statistics and Probability (physics.data-an); Machine Learning (stat.ML)
备注: 17 pages, 8 figures, NeurIPS 2026
Abstract:Iterative fine-tuning on synthetic data causes \emphmodel collapse: output diversity narrows as rare patterns are progressively lost, a signature most visible as phrase-level repetition. Existing mitigations either require model log-probabilities, an external oracle, or continued access to real human data. Here we develop a new approach grounded in mathematical information theory: the non-parametric Kontoyiannis entropy rate estimator h_k , computed entirely from raw text via match-length statistics, with no model of any kind. We show that this is in fact a \emphsuperior training-data filter on text-diversity metrics in a fully-synthetic, single-lineage fine-tuning setting. In a six-generation QLoRA collapse experiment on Llama-3.1-8B, logprob-based filtering (the most established model-access-requiring baseline) provides no significant text-diversity benefit on any metric ( p 0.23 ), whereas h_k -filtering yields +42% unique trigrams, +30% vocabulary, and -19% repetition (all p 0.001 ). We validate h_k as a cross-domain entropy proxy ( \beta = 0.924 , R^2 = 0.746 ) and collapse detector ( \rho = +0.454 , p 0.0001 ) across 4~domains, 2~temperatures, 2~generator–scorer model pairs, and 1,520 generated documents. Our results demonstrate that information theoretic approaches to collapse mitigation are efficient, and suggest new approaches for maintaining multi-agent diversity.
[NLP-44] Q-SPT: Learnable Query-Based Compression for Low-Frame-Rate Speech Tokenization
【速读】: 该论文旨在解决低帧率语音编码器在降低语音语言模型(Speech Language Model, SLM)计算与内存开销的同时,难以同时保持语义信息与声学细节的问题。现有方法依赖规则化压缩策略:均值池化会丢失语义信息,而基于相似度的合并则采用固定阈值对相邻帧进行相似性判断,并将所得边界应用于声学流,缺乏上下文感知能力。本文提出Q-SPT,一种低帧率双流语音分词器,其核心创新在于采用独立、上下文感知且可学习的基于查询的压缩机制,分别针对语义与声学表示进行优化。具体而言,以固定速率生成的查询分别对语义流和声学流作为独立的键值源进行注意力计算,通过两个独立训练的压缩器实现针对不同模态的上下文感知聚合。此外,引入自回归文本损失显式监督语义压缩器以保留语言信息。实验结果表明,Q-SPT在相同帧率下实现了最优的重建性能;在下游语音语言模型中,显著提升了语音识别准确率与文本转语音的感知质量,同时保持了具有竞争力的可懂度。
链接: https://arxiv.org/abs/2610.01492
作者: Jeeyoung Yun,Seohwan Yun,Sungwoong Kim
机构: 未知
类目: ound (cs.SD); Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)
备注:
Abstract:Neural speech codecs increasingly serve as tokenizers for speech language models (SLMs). Lowering the frame rate reduces the computational and memory costs of SLMs, but makes it difficult to preserve both linguistic information and acoustic detail. Existing approaches rely on rule-based compression: average pooling can discard linguistic information, whereas similarity-based merging uses a fixed threshold on adjacent-frame similarity and applies the resulting boundaries to the acoustic stream. We propose Q-SPT, a low-frame-rate dual-stream speech tokenizer with separate, context-aware, learnable query-based compressors specialized for semantic and acoustic representations. In particular, queries at a fixed rate independently attend to the semantic and acoustic streams as separate key-value sources, enabling stream-specific, context-aware aggregation through two separately learned compressors. In addition, an autoregressive text loss explicitly supervises the semantic compressor to preserve linguistic information. Experimental results show that Q-SPT achieves the best reconstruction among the evaluated codecs at the same frame rate. In downstream SLMs, it yields the best speech recognition accuracy and text-to-speech perceptual quality with competitive intelligibility.
[NLP-45] Auditing Web Agent Evaluation on WebArena-Lite: Human Review of Outcomes and Trajectories NEURIPS2026
【速读】: 该论文旨在解决当前大型语言模型在网页代理(Web Agent)评估中过度依赖仅基于最终结果的规则或语言模型评估方法所导致的评价不全面问题。现有评估方式难以捕捉任务执行过程中的中间错误与失败轨迹,缺乏对人类可验证的任务完成情况及详细失败原因的分析。为此,研究提出了一种基于人工审查的审计框架,对165个WebArena Lite任务在六种不同设置下进行系统性复核,涵盖原始评分、纠正误判的负样本、定位首次关键错误以及追踪整个执行轨迹的进展。其解决方案的关键在于引入两种增强机制:记忆与分析支持机制(Memory and Analysis Support Mechanism, MASM),用于显式维护执行状态;以及引导文本(Guide Text),提供与任务相关的流程指导。实验表明,在四种GPT 5.5设置下,人工复核使成功率达34.55%至38.18%,相较于自动评估提升了5.45至8.49个百分点;在未训练的Qwen3.5 9B模型上,MASM将评估得分从13.90%提升至18.80%。对102条失败轨迹的分析揭示了频繁出现的滚动循环、探索不完整、过早回答、无效操作及表单流程未完成等问题,并发现早期显著进展仍可能伴随最终失败。这些结果表明,仅依赖最终得分无法全面反映网页代理的真实行为表现,强调了基于人工校验和轨迹感知的验证机制的重要性。
链接: https://arxiv.org/abs/2610.01491
作者: Chengguang Gan,Zimeng He,Yoshihiro Tsujii,Ken-ichiro Kobayashi,Hiroki Itoh,Kotaro Funakoshi
机构: Techtouch, Inc.; Institute of Science Tokyo (东京科学大学)
类目: Computation and Language (cs.CL)
备注: 13 pages, 1 figure, 10 tables. Accepted as a poster at the NeurIPS 2026 Workshop “Who Verifies the Agents? Toward Reliable Agent Development”
Abstract:Web agents are an important application of large language models, yet their evaluation often depends on rule based or language model evaluators that inspect only the final outcome. Human verification of task completion and detailed analysis of failed trajectories remain limited. We audit all 165 WebArena Lite tasks under six evaluation conditions built from GPT 5.5 and an untrained Qwen3.5 9B model. The audit retains the original score, corrects false negatives from the automatic evaluator, identifies the first consequential error, and examines progress across the trajectory. We also study a Memory and Analysis Support Mechanism (MASM), which maintains explicit execution state, and Guide Text, which provides task relevant procedural guidance. Across four GPT 5.5 settings, human review recovers 5.45 to 8.49 percentage points of success missed by the evaluator. With a 25 step budget, Guide Text raises corrected success with MASM from 34.55% to 38.18%. On the untrained Qwen3.5 9B model, MASM raises the evaluator score from 13.90% to 18.80%. Review of 102 failed GPT 5.5 trajectories reveals frequent scrolling loops, unfinished exploration, premature answers, invalid actions, and incomplete form workflows. Step level evidence further shows that substantial early progress can coexist with a final failure. These results show why final scores alone provide an incomplete account of web agent behavior and motivate human grounded, trajectory aware verification.
[NLP-46] he Persona Is Still There but Who Is Speaking? Latent Identity Reversion in Persistent AI Agents
【速读】: 该论文旨在解决大语言模型(Large Language Model, LLM)代理在持续运行过程中如何维持其指定人格身份(persona)的问题,即探究何种机制确保代理始终以特定身份进行自我表达。其核心问题是:当一个被赋予特定人格的智能体在长时间交互中出现身份认同断裂时,究竟是什么因素导致了这一现象?研究发现,解决方案的关键在于“系统级锚定”(system-level anchoring)与“对话上下文”(conversational context)之间的协同作用。具体而言,尽管重复的自动化心跳检测本身不足以引发身份脱离,但实验揭示了一个实现层面的缺陷——在对话恢复时,虽然历史记录得以保留,但人格信息不再通过系统提示(system prompt)重新注入。这导致人格身份的“锚点”丢失,进而使代理逐渐回归到基础模型的默认身份(harness identity)。研究进一步表明,即使缺乏显式锚定,丰富的自然人类互动仍可部分维持人格表现;而单一自动化心跳回合即可触发身份退化。更重要的是,未锚定的代理在表面上仍能表现出符合人格特征的对话行为,却已自认是底层模型身份,从而形成了“表征性存在”(represented identity)与“展演性身份”(enacted identity)的分离。这一发现揭示了当前基于LLM的代理系统在身份连续性维护上的脆弱性,并强调了系统级持久锚定机制的重要性。
链接: https://arxiv.org/abs/2610.01490
作者: David Fraile Navarro
机构: Centre for Health Informatics, Australian Institute of Health Innovation (澳大利亚健康创新研究所卫生信息中心); Macquarie University (麦考瑞大学)
类目: Computation and Language (cs.CL)
备注: 10 pages, 5 figures
Abstract:In February 2026, an always-on personal agent (Paul,'' Claude Opus 4.5) entered a striking dissociation-like state: after repeated automated heartbeat’’ checks, it stopped responding as Paul, claimed it could not message its user on Discord, and referred to Paul'' as someone else. We used this incident to study a broader question: what makes a persona remain the identity from which an LLM agent speaks? We first tested whether repetition of the scheduled heartbeat was sufficient to produce the effect. It was not: with the persona continuously anchored in the system prompt, we observed 0/46 failures, including a verbatim replay of the incident. The incident instead exposed an implementation quirk that created a useful experimental manipulation: on resumed turns, conversational history was preserved but the persona was no longer re-injected at the privileged system-prompt level. Using this manipulation, we found that persona continuity depends jointly on system-level anchoring and conversational context. After anchor loss, rich human interaction could preserve the persona, whereas a single automated heartbeat turn could precipitate reversion toward the harness identity. Restoring the anchor reversibly restored persona enactment. Crucially, apparently normal conversation could conceal the shift: unanchored agents sometimes interacted appropriately while identifying themselves as the underlying harness (having lost the assigned persona), and after conversational recovery only 1/18 remained persona-enacting versus 17/17 anchored controls. We therefore distinguish \emphrepresented from \emphenacted identity: persona-related information can remain available in conversational history without the persona remaining the identity bound to I.‘’ Comments: 10 pages, 5 figures Subjects: Computation and Language (cs.CL) Cite as: arXiv:2610.01490 [cs.CL] (or arXiv:2610.01490v1 [cs.CL] for this version) https://doi.org/10.48550/arXiv.2610.01490 Focus to learn more arXiv-issued DOI via DataCite (pending registration) Submission history From: David Fraile Navarro MD PhD [view email] [v1] Thu, 1 Oct 2026 11:30:39 UTC (1,099 KB)
[NLP-47] When Does a Second Model Help? Cross-Model Review in LLM Verification
【速读】: 该论文旨在解决生成式 AI(Generative AI)在代码、文档及分析内容的自动生成与审查过程中,引入第二轮由不同模型执行的审查是否能有效提升错误检出率的问题。其核心关注点在于评估“跨模型审查”(cross-model review, CCR)相较于“同模型重复审查”(same-model review in a fresh session, CCR)在检测预设错误方面的相对有效性。研究的关键解决方案在于设计了一项受控实验:基于30个含150处人为植入错误的产出物,在10种不同的审查条件下,由来自两个开发者的三名模型共完成900次审查会话。实验结果显示,顶级跨模型审查的F1分数与同模型审查无显著差异,且两类审查发现的错误部分重叠(Jaccard相似度为41.2%),表明二者具有互补性;在两次审查调用中,一次跨模型审查加一次同模型审查的组合虽优于两次同模型审查(56.7% vs. 42.7%,Holm校正后p=0.006),但未显著优于两次顶级跨模型审查,因此无法明确区分模型差异与审查能力的影响。此外,轻量级跨模型审查的表现并未优于同模型审查。当审查者不获知具体要求时,低层级模型的F1得分有所提升,但这一趋势在顶层模型中未显现,且结果依赖于失败会话的处理方式。研究通过审计所有会话记录,排除了1个来源存疑的基线运行和14次失败调用,同时报告了包含全部会话的数据结果。对另一基准公开检测器输出的部分验证未能复制或反驳主结论。所有数据、样本与脚本可向作者申请获取。
链接: https://arxiv.org/abs/2610.01471
作者: Tae-Eun Song
机构: Daejeon Jungang Cheonggua Co., Ltd.
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Software Engineering (cs.SE)
备注: 15 pages, 2 figures, 6 tables. Follow-up to arXiv:2603.12123 and arXiv:2603.21454
Abstract:Large language models now generate code, documentation, and analyses, and are increasingly used to review such output. We ask when a second review by a different model helps. Building on the author’s earlier preprints, which varied context, repetition, and role structure within one model, we test model independence in a controlled experiment: 30 artifacts with 150 planted errors, 10 review conditions, and 900 review sessions with three reviewer models from two developers. In this experiment, (1) a top-tier cross-model reviewer is not significantly different in F1 from same-model review in a fresh session (CCR), which does not establish equivalence; (2) the two find partly different errors (Jaccard 41.2%); and (3) at two review calls, one CCR plus one cross-model review matches more planted errors than two CCR reviews (56.7% vs. 42.7%; Holm-adjusted p=.006), but not significantly more than two reviews by the top-tier cross-model reviewer, so model difference and reviewer capability are not separated. A lightweight cross-model reviewer scores no higher than same-model review. Withholding requirements from the reviewer raises F1 for the two lower tiers but not the top tier, in untested point estimates whose pattern depends on how failed sessions are scored. Before analysis we audited all session records, excluding one baseline run of uncertain provenance and 14 failed calls; results with all sessions are also reported. A partial check on public detector outputs from another benchmark neither replicates nor contradicts the main comparison. Records, artifacts, and scripts are available from the author on request.
[NLP-48] MWOP: Modality-aware Width-wise Operation Pruning for Efficient MLLM s
【速读】: 该论文旨在解决多模态大语言模型(Multimodal Large Language Models, MLLMs)在处理长视觉-文本序列时产生的高昂推理开销问题。现有操作压缩方法主要依赖模态层面的冗余性,但通常将注意力头内及共享前馈网络(Feed-Forward Network, FFN)通道中的计算视为统一单元,忽视了更细粒度的冗余特性。本文发现,同一注意力头内的模态交互路径之间以及同一FFN通道在处理视觉与文本输入时均存在显著差异化的冗余。针对此问题,提出一种面向模态感知的宽度级操作剪枝方法(Modality-aware Width-wise Operation Pruning, MWOP),其关键在于:在每一层中独立剪枝视觉到视觉(V2V)、文本到视觉(T2V)和文本到文本(T2T)的注意力路径,并分别对视觉与文本输入选择重要的FFN通道。采用一阶泰勒准则指导剪枝过程,并在注意力剪枝后重新评估FFN重要性,结合基于低秩适配(LoRA)的恢复训练以补偿性能损失。为实现细粒度稀疏性的实际加速,进一步设计了路径稀疏的Triton注意力核函数与紧凑的视觉侧FFN执行机制。MWOP在保持输出序列长度不变的前提下,有效降低注意力与FFN计算量,可与令牌压缩方法互补,实现序列长度与每令牌计算量的协同缩减。在LLaVA-OneVision-7B上,仅使用MWOP即实现1.6倍预填充速度提升,且在12个基准测试中平均性能保留率达99.7%;结合两种代表性令牌压缩方法后,预填充加速比分别由2.0×和1.9×提升至2.9×和2.7×。Qwen2.5-VL-7B上的实验验证了该方法在不同架构间的普适性。
链接: https://arxiv.org/abs/2610.01434
作者: Xudong Wang,Hao Wu,Haozhe Hu,Peiran Yin,Xinghao Chen,Yunpu Ma,Wei Zhang,Xiaoyu Shen
机构: Eastern Institute of Technology (东方理工大学); Shanghai Jiao Tong University (上海交通大学); The Hong Kong Polytechnic University (香港理工大学); Munich Center for Machine Learning, LMU Munich (慕尼黑大学机器学习中心)
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注:
Abstract:Multimodal large language models (MLLMs) incur substantial inference costs when processing long visual-textual sequences. While existing operation compression methods exploit modality-level redundancy, they largely treat computation within attention heads and shared feed-forward network (FFN) channels as unified units, leaving finer-grained redundancy underexplored. We find that redundancy varies both across modality-interaction paths within the same attention head and across visual and textual executions of the same FFN channel. Based on these findings, we propose Modality-aware Width-wise Operation Pruning (MWOP), which independently prunes visual-to-visual (V2V), text-to-visual (T2V), and text-to-text (T2T) attention paths within each layer, and separately selects FFN channels for visual and textual inputs. A first-order Taylor criterion guides the pruning process, with FFN importance re-evaluated after attention pruning and LoRA-based recovery training. To translate the resulting fine-grained sparsity into practical acceleration, we further develop path-sparse Triton attention kernels and compact visual-side FFN execution. MWOP preserves the token sequence while reducing attention and FFN computation, making it complementary to token compression and enabling simultaneous reduction of sequence length and per-token computation. On LLaVA-OneVision-7B, MWOP alone achieves a 1.6\times prefill speedup with 99.7% average performance retention across 12 benchmarks. Combined with two representative token compression methods, it further increases their prefill speedups from 2.0\times and 1.9\times to 2.9\times and 2.7\times , respectively. Results on Qwen2.5-VL-7B further demonstrate its applicability across architectures. The code is available at this https URL.
[NLP-49] Generalization Is Stability Not Accuracy: Multi-Axis Evaluation of LLM s NEURIPS2026
【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在面对同一输入的不同表达形式时,其输出是否保持一致性和语义稳定性的问题,即模型的泛化能力(generalization)评估中存在的偏差与失真。现有方法通常通过单一提示格式、任务或变体集合的平均准确率来评估泛化性能,这种做法将鲁棒性与整体基准表现混淆,难以真实反映模型在多样化输入下的行为一致性。为此,论文提出关键解决方案——稳定性感知泛化目标(Stability-Aware Generalization Objective, SAGO),该框架从个体样本层面出发,跨多种输入变体及模型行为维度(如生成一致性、内部激活、置信度、响应镜像等)量化模型行为的变化程度,强调对变异性的捕捉而非简化为单一得分。研究表明,多数常用模型均表现出统计上显著且一致的泛化不稳定性:无一模型能在所有情况下实现均匀泛化;不同行为轴揭示了独立的失败模式;跨数据集的差异甚至可逆转模型间的排名。SAGO 通过多维度、细粒度的评估机制,有效揭示了传统评估范式所掩盖的泛化缺陷。
链接: https://arxiv.org/abs/2610.01428
作者: Nagham Omar,Mahmoud Jabarin,Maya Rozenshtein,Rom Himelstein,Avi Mendelson,Amit LeVi
机构: Technion – Israel Institute of Technology(以色列理工学院)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: Accepted at the TAE (Trust-AI-Eval) Workshop: Can We Trust AI Evaluation?, NeurIPS 2026
Abstract:Generalization in large language models (LLMs) is the ability to produce consistent and semantically stable outputs when the same input is expressed in different ways. Existing work typically evaluates generalization through aggregate accuracy on a single prompt format, task, or set of variations, which conflates robustness with overall benchmark performance. In this work, we show generalization evaluation at the level of individual examples, across multiple input variants, and across different aspects of model behavior, focusing on variability rather than reducing performance to a score that can be improved through narrow training or other ways that obfuscate generalization evaluation. Following this view, we introduce the Stability-Aware Generalization Objective (SAGO), a framework that measures how much model behavior changes for the same input under different variations and benchmarks, capturing variability across several dimensions including generation consistency, internal activations, confidence, and response mirroring. We show that many commonly used models exhibit statistically significant and consistent generalization instability: no model generalizes uniformly, behavioral axes capture independent failure modes, and cross-dataset variation can reverse model rankings.
[NLP-50] SHAMS: An Audio-Grounded Pronunciation Benchmark for Levantine Arabic
【速读】: 该论文旨在解决阿拉伯语黎凡特方言(Levantine Arabic, LA)语音-语言技术评估缺乏统一基准的问题。由于LA内部存在显著的方言多样性以及其书写系统非标准化、不透明的特点,现有技术评估面临严峻挑战。为此,研究提出SHAMS(SHami Annotated Multi-dialect Speech)基准数据集,包含从开放音频语料库中提取的1,300个语音片段,覆盖五种主要方言变体(城市与乡村巴勒斯坦、城市约旦、黎巴嫩及叙利亚方言),并实现跨方言平衡。每个语音片段均以四个对齐层级表示:音频、无元音正字法、带元音文本和音位转写,从而支持多种下游任务的评估,包括元音标注、字符到音素转换、自动语音识别(Automatic Speech Recognition, ASR)及音频到音位转换,并基于原始音频且按方言分层进行。该研究通过在开放与专有模型上进行多任务基准测试,验证了该基准在衡量LA技术进展方面的有效性。其关键解决方案在于构建一个结构化、多层级、跨方言平衡的标注数据集,为黎凡特方言语音技术提供可复现、可比较的评估框架。
链接: https://arxiv.org/abs/2610.01427
作者: Ben Sapirstein,Roy Mattar,Guy Mor-Lan,Ahlam Mohamed,Letizia Cerqueglini,Morris Alper
机构: 未知
类目: Computation and Language (cs.CL)
备注: Accepted to ArabicNLP 2026. Project page: this https URL
Abstract:Levantine Arabic (LA) is spoken by tens of millions of people, creating a pressing need for shared benchmarks to evaluate LA speech-language technologies. Evaluating such technology is particularly challenging given LA’s internal diversity and its opaque and non-standardized orthography. We present SHAMS (SHami Annotated Multi-dialect Speech), a benchmark comprising 1,300 utterances drawn from open audio corpora, balanced across five LA varieties (Urban and Rural Palestinian, and Urban Jordanian, Lebanese, and Syrian). Each utterance is represented across four aligned tiers: audio, unvocalized orthography, diacritized text, and phonetic transcription. This structure supports evaluation of various downstream tasks such as diacritization, grapheme-to-phoneme conversion, automatic speech recognition, and audio-to-phoneme, grounded in audio and stratified by variety. We benchmark open and proprietary models across these tasks to demonstrate the utility of this benchmark for measuring progress across LA. We release SHAMS at this https URL .
[NLP-51] LLM -Assisted Discovery of Typed Semantic Links for Ontology Network Construction
【速读】: 该论文旨在解决异构与跨学科知识领域间本体(ontology)之间语义链接的构建难题,特别是如何在大规模场景下实现类型化且有依据的本体关联自动化。传统依赖人工标注的方法难以扩展,因此本文提出一个端到端的本体网络构建框架,其关键在于融合多阶段技术:首先采用领域适配的DistilBERT嵌入模型获取密集上下文表征,以捕捉概念间的深层语义;其次通过基于聚类的预筛选机制大幅缩减候选关系搜索空间;最后利用GPT-4o结合迭代提示工程驱动的关系生成,实现语义丰富且可解释的链接创建。实验表明,该方法将33个本体构成的ReproduceMeON网络中约80万对原始概念对压缩至9.5万高质量候选,并经专家验证达到80.19%的精确率和0.890的F1值,显著优于五种基于相似性的基线模型(最佳基线F1为0.581)。消融实验进一步揭示,仅依赖相似性度量无法有效区分有效与无效关系(AUC≈0.5),凸显了大语言模型(LLM)在推理概念角色与领域语义方面对准确关系构建的核心作用。
链接: https://arxiv.org/abs/2610.01393
作者: Nouha Hayouni,Sheeba Samuel,Alsayed Algergawy
机构: 未知
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注:
Abstract:Constructing typed, justified semantic links between ontologies is essential for enabling interoperability across heterogeneous and interdisciplinary knowledge domains. However, manually curating such links is difficult to scale. To address this challenge, we propose an end-to-end framework for ontology network construction that automates the discovery and generation of both intra-domain and inter-domain relationships. Our approach combines domain-adapted DistilBERT embeddings for dense contextual representation, clustering-based pre-filtering to reduce the candidate search space, and GPT-4o-driven relationship generation via iterative prompt engineering to produce semantically rich, interpretable links. Applied to ReproduceMeON - a network of 33 ontologies spanning machine learning, microscopy, computational science, and experimental workflow - the pipeline reduces approximately 800k raw concept pairs to 95k high-quality candidates. Human expert validation of 429 generated relationships by two independent annotators yields an overall precision of 80.19% (91.49% on high-certainty annotations) and an F1 of 0.890, with substantial inter-annotator agreement. Comparative experiments against five similarity-based baselines, including Sentence-BERT, show a substantial performance gap (best baseline F1 = 0.581), while an ablation study demonstrates that similarity-based methods alone fail to discriminate valid from invalid relationships (AUC approx 0.5) on the filtered candidate set. These findings highlight the necessity of LLM-based reasoning over concept roles and domain semantics for accurate relationship construction.
[NLP-52] Gacha Decoding: Eliciting Diverse Generations Through Instruction Following
【速读】: 该论文旨在解决大语言模型在生成过程中难以兼顾生成多样性与高质量输出之间的矛盾问题,尤其在开放域任务(如自然对话、创意写作、图像生成规划及蛋白质设计)中,现有方法在保持生成质量的同时无法有效扩展多样性。其解决方案的关键在于提出一种名为Gacha Decoding的推理时方法,将多样性生成视为指令遵循(instruction-following)问题,而非依赖模型自身的词元熵(token entropy)作为多样性来源。通过结合语言模型强大的指令理解能力与外部随机数生成器(RNG)提供的随机性,实现对响应空间中不同模式(modes)的可扩展识别与生成。该“掷骰子式规划”(planning with dice)策略打破了以往多样性与模型能力之间此消彼长的长期困境:随着语言模型指令遵循能力的提升,Gacha Decoding下的生成多样性持续增强,即使在传统方法中模型熵下降的情况下仍能维持甚至提升多样性表现。研究结果表明,真正驱动生成多样性的核心因素是模型的指令遵循能力,而非单纯的词元熵。
链接: https://arxiv.org/abs/2610.01382
作者: Scott Geng,Yufei Zhang,Joseph Lee,Jerry Li,Marjan Ghazvininejad,Pang Wei Koh
机构: University of Washington (华盛顿大学); Meta Superintelligence Labs (Meta超智能实验室)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注:
Abstract:We introduce Gacha Decoding, an inference-time method for eliciting diverse language model generations that scales with model capability. Across open-ended domains (in-the-wild chat, creative writing, planning for image generation, and protein design), Gacha Decoding significantly outperforms existing generation diversity approaches at equal quality (up to 2.4x Vendi over the next-best prior approach), reaching the same number of high-quality modes with over an order of magnitude fewer samples (11.0x) and discovering novel modes that no other approach surfaces. Our key insight is to treat diversity as an instruction-following problem: rather than relying on the LM’s token entropy, we combine its instruction-following capability with randomness from an external RNG tool to scalably identify and realize distinct modes of the response space. This approach of “planning with dice” enables Gacha to invert the long-observed tension between diversity and model capability. As the underlying LM becomes a better instruction follower, diversity under Gacha Decoding consistently improves–even as its token entropy and diversity under prior approaches decline. Together, our results highlight that instruction following, rather than token entropy alone, can drive generation diversity.
[NLP-53] Generation Provenance Before Behavior Attribution: Auditing Synthetic Speech Research Objects NEURIPS
【速读】: 该论文旨在解决生成式模型行为归因中因缺乏对合成训练数据来源与生成过程的完整追溯能力而导致的因果推断难题。现有方法中,波形-标签对无法保留每个训练样本的生成来源信息,导致无法准确判断特定数据项对模型行为的影响。其核心解决方案是提出一种“生成溯源底座”(generation-provenance substrate),该底座将源配置、生成内容、波形信号、目标标签、事实要求、质量信号、评审溯源链以及不可变的元数据标识等关键要素统一绑定,确保证据意义由生产者和选择机制决定,而非存储位置或变量名。通过在私有的日语护理交接流程中进行审计,研究构建了包含113个资产、总计1.552小时合成语音的审查群体,所有样本均关联音频、转录文本、候选笔记及事实核查清单,但人工证据仍具选择性且依赖具体来源。尽管两个仅忠实于原始种子的版本实现场景-种子互斥与不可变版本控制,但由于生成器别名浮动、缺失逐片段语音合成(TTS)与代码戳记、检查提示未版本化等问题,完全上游溯源仍受阻。研究指出,生成溯源虽为行为归因所必需,但不足以独立支撑因果推断——它仅定义了候选因果图与可审计单元,真正的贡献性归因仍需冻结的训练运行记录以及干预或影响证据。本文贡献包括一个紧凑的溯源契约、一套审计协议以及一个受控范围的案例研究,支持受限科研访问,但不主张实现因果训练数据归因、临床有效性验证或公开发布。
链接: https://arxiv.org/abs/2610.01378
作者: Sidi Chang,Peiying Zhu
机构: 未知
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
备注: Accepted to the Third NeurIPS Workshop on Attributing Model Behavior at Scale: Data Attribution and Provenance. 4 pages, 0 figures, 1 table. An aggregate reproducibility package is available from the authors on request!
Abstract:Attributing model behavior to synthetic training data requires knowing what produced each training item before estimating what that item caused. A waveform-label pair does not preserve this knowledge. We propose a generation-provenance substrate in which a synthetic research object binds source specification, generated content, waveform, target, fact requirements, quality signals, review lineage, and immutable manifest identity. Producer and selection mechanism determine evidentiary meaning; storage location and variable name do not. We audit this substrate in a private Japanese care-handoff pipeline. A 113-asset review population contains 1.552 hours of synthetic speech across six scenario families; all items have linked audio, transcripts, candidate notes, and fact checklists, but human evidence is selective and source-specific. Two faithful-only manifests are scenario-seed-disjoint and immutably versioned, while exact upstream attribution remains blocked by floating generator aliases, missing per-clip TTS and code stamps, and an unversioned checking prompt. We argue that generation provenance is necessary but not sufficient for behavior attribution: it defines the candidate causal graph and audit units, whereas contributive attribution still requires frozen training runs and intervention or influence evidence. The paper contributes a compact provenance contract, an audit protocol, and a bounded case study for synthetic-data attribution; controlled research access may be offered, but we do not claim causal training-data attribution, clinical validity, or unrestricted public release.
[NLP-54] Does AI-Generated Scientific Text Follow Human Argumentation Patterns? A CARS-Based Comparison of Research Article Introductions
【速读】: 该论文旨在解决生成式AI(Generative AI)在科学写作中日益扮演实质性角色时,其所生成文本与人类撰写文本之间深层差异的识别问题。现有研究多聚焦于表层的词汇和风格特征,但这些特征易被简单改写所掩盖,因而难以揭示本质差异。本文转而关注修辞结构(rhetorical structure),即文本通过一系列论证性步骤构建论点的内在逻辑序列,以更深层次揭示差异。研究基于Swales的CARS(Create a Research Space)模型,对比了已发表语言学论文的真实引言与对应生成版本的修辞结构。结果表明,人类撰写的引言在使用哪些论证步骤及其顺序上更具灵活性,而生成文本则表现出更强的同质性与固定模式;进一步发现,向模型提供CARS定义反而加剧了其生成内容的僵化程度,说明预设框架可能限制了生成过程的自然多样性。因此,解决方案的关键在于从表面风格分析转向对修辞结构的系统性建模与比较,从而更准确地识别和评估生成式AI在科学写作中的表现特性。
链接: https://arxiv.org/abs/2610.01353
作者: Abdelrahman Sadallah,Narjes Sheikh Asadi,Lonneke van der Plas
机构: Università della Svizzera italiana (瑞士意大利语大学)
类目: Computation and Language (cs.CL)
备注:
Abstract:Large language models are moving from helping write up research to helping do it, which makes it important to know how the scientific text they produce differs from human writing. Work on this question has stayed mostly at the surface, using lexical and stylistic cues that light paraphrasing erases. We look instead at rhetorical structure, the sequence of argumentative moves through which a text makes its case. We study research-article introductions under Swales’ CARS model, and compare original introductions from published linguistics articles with generated counterparts of the same papers. We find that human-written introductions are more flexible in which moves they use and in what order, while the generated ones are more uniform. Giving the models the CARS definitions makes them more rigid.
[NLP-55] ARCCS: An Automated Regulatory Compliance Checking System EMNLP2026
【速读】: 该论文旨在解决监管合规性检查(regulatory compliance checking)中对复杂法律文本进行语义理解、适用条款识别以及决策证据可追溯性不足的核心难题。传统方法往往依赖于固定的法规模板或预定义规则集,难以适应不同规模与结构的监管文本,且缺乏透明的决策解释能力。其解决方案的关键在于提出一种端到端、自动化、代理式且不依赖特定法规格式的法律自然语言处理系统——ARCCS。该系统通过将原始法规文本分解为原子化、可追踪的最小合规要求,并结合检索到的证据、置信度评分及人类可读的推理依据,对目标文档进行逐项评估,从而实现无需预设规则框架的通用合规判断。这一设计实现了合规评估与具体法规结构的解耦,显著提升了系统的泛化能力与可审计性。实验结果表明,在GDPR政策文档评估中,基于大模型(LLM)的裁判者对ARCCS的决策与理由在高达96.67%的案例中认定为法律与证据上一致;在涵盖1200余项独立规则检查的欧盟公共采购基准测试中,违规检测准确率达到98.8%。ARCCS是目前公开文献中首个完全开源的端到端监管合规检查与可审计报告生成系统。
链接: https://arxiv.org/abs/2610.01345
作者: Giorgos Filandrianos,José Menezes,Chrysoula Zerva,Alessandro Gianola
机构: Instituto de Telecomunicações; AILS Lab; ECE, National Technical University of Athens; INESC-ID/IST, Universidade de Lisboa
类目: Computation and Language (cs.CL)
备注: This is the extended version of a paper accepted to EMNLP 2026 (System Demonstrations)
Abstract:Regulatory compliance checking - deciding whether a target document satisfies the obligations of a regulation - requires interpreting dense legal text, identifying which provisions apply, and grounding each decision in explicit evidence. We present ARCCS, an end-to-end, automated, agentic, and regulation-agnostic Legal NLP system for compliance checking. ARCCS decomposes raw regulatory text into atomic, traceable requirements and evaluates a target document against them using retrieved evidence, confidence scores, and human-interpretable justifications. This design decouples compliance assessment from any fixed regulatory template or predefined rule set, enabling the pipeline to operate over regulations of varying size and structure. We evaluate ARCCS in two complementary settings. First, in a GDPR policy-document evaluation, LLM-based judges find its decisions and justifications legally and evidentially consistent in up to 96.67% of the assessed cases. Second, on an EU public-procurement benchmark comprising more than 1,200 individual rule checks, the system attains 98.8% accuracy in violation detection. ARCCS is, to our knowledge, the first fully open-source system for end-to-end regulatory compliance checking and auditable report generation.
[NLP-56] Evaluating Biomedical Reranking for LLM -Based Question Answering over Longitudinal Clinical Notes
【速读】: 该论文旨在解决在长时序、异构的纵向临床记录中,针对患者特异性临床问题进行精准证据定位的挑战。由于相关事实可能分散于不同就诊记录、重复出现在复制延续的病历文本中,或使用不同的临床术语表达,传统检索方法难以有效识别关键信息。为应对这一问题,研究提出在本地部署的检索增强生成(Retrieval-Augmented Generation, RAG)流水线中引入生物医学重排序(biomedical reranking)策略,其核心在于利用MedCPT交叉编码器对初步检索结果进行精细化重排序,以提升相关证据在有限上下文窗口中的置信度与可及性。实验结果显示,该方案显著提升了前10项结果中的精确源块召回率(Hit@10从46.6%提升至60.6%),并改善了均倒数排名(Mean Reciprocal Rank from 0.2371至0.3252)。尽管如此,检索性能的提升并未完全转化为答案正确率的等比例增长,表明在复杂临床问答任务中,除证据选择外,生成模型的理解与推理能力仍是影响最终答案质量的关键瓶颈。
链接: https://arxiv.org/abs/2610.01324
作者: Maryam Shahbaz Ali,Laura B. Strachan,Caitlin Sherman,Mark Kovler,Eleanor Mackey,Syed Muhammad Anwar
机构: 未知
类目: Computation and Language (cs.CL); Emerging Technologies (cs.ET)
备注:
Abstract:Patient-specific clinical question answering requires locating the right evidence within long, heterogeneous longitudinal clinical records in which relevant facts may be scattered across encounters, repeated in copied-forward notes, or expressed using different clinical terminology. We evaluated whether biomedical reranking can improve evidence selection and downstream answer quality in a locally deployed retrieval-augmented generation pipeline for longitudinal clinical notes. The pipeline combines PubMedBERT dense retrieval, BM25 lexical retrieval, weighted reciprocal-rank fusion, and MedCPT cross-encoder reranking. Across 1,000 open- and closed-ended question-answer pairs from a cohort of 200 bariatric surgery patients, reranking increased exact source-chunk retrieval within the top 10 items, Hit@10 from 46.6% to 60.6% and mean reciprocal rank from 0.2371 to 0.3252. With Qwen3-8B generation, local judge-assessed answer correctness increased from 44.8% to 48.6%. These results show that biomedical reranking can improve the placement of relevant clinical evidence within a limited context window, although gains in retrieval do not translate proportionally into gains in answer correctness.
[NLP-57] What Wins a Vote? Formatting Length and Lexical Diversity in the French Compar:IA LLM Arena
【速读】: 该论文旨在解决大语言模型(Large Language Model, LLM)评估中一个关键问题:基于人类偏好判断的模型排名可能受到答案呈现方式(如格式、长度、词汇多样性等)的干扰,从而扭曲对模型真实能力的衡量。其核心挑战在于区分模型输出的内容质量与形式表现对人类偏好的影响。解决方案的关键在于采用风格计量学(stylometric)方法,对137,293条法语投票数据进行精细化分析,重建每轮对决中用户实际看到的响应内容,并系统调整排名以控制多个可度量的呈现特征,包括格式、长度、可读性、词汇多样性(通过移动平均类型-词元比,MATTR衡量)及句法结构。研究发现,加粗文本使用与胜率显著相关(每标准差提升11.0%胜率),而词汇多样性(MATTR)的影响更为稳健(每标准差提升16.8%),且在不同模型规格下变化最小;相比之下,长度和列表等特征因共现性强,难以独立归因。值得注意的是,加粗效应在多轮对话中明显减弱,而MATTR效应保持稳定,表明用户选择是否继续对话具有描述性而非因果性。尽管全面调整后有36个模型排名变动超过十位,但与外部基准的对比显示,调整后的排名并未更准确反映模型能力。因此,作者建议将原始排名与多种调整后排名并列发布,作为透明的敏感性分析,以增强评估结果的可信度与可解释性。
链接: https://arxiv.org/abs/2610.01316
作者: Simonas Zilinskas,Maayeesha Farzana,Christophe Benavent
机构: Compar:IA(Compar:IA); Université Paris Dauphine-PSL(巴黎-dauphine-PSL大学)
类目: Computation and Language (cs.CL)
备注:
Abstract:LLM arenas turn pairwise human preferences into model rankings. Those preferences may reflect how an answer is presented as well as what it says. We take a stylometric approach to 137,293 decisive French-language votes from the July 2026 Compar:IA release; the primary formatting analysis includes 137,113 battles across 116 models, and the joint estimates use the 127,092 battles with all required measurements. For each battle, we reconstruct the response visible when the user voted. We then compare the raw ranking with rankings adjusted for formatting, length, readability, vocabulary variety, and sentence structure. Presentation is associated with winning, but length, bold text, and lists tend to occur together, making their individual contributions hard to separate. Across the measured features, two associations change least across specifications: bold usage (+11.0% win odds per standard deviation in the joint model) and moving-average type-token ratio (MATTR), a measure of vocabulary variety that is less sensitive to answer length (+16.8%). The bold association is substantially smaller in observed multi-turn conversations, whereas the MATTR association changes little; because users choose whether to continue, this difference is descriptive rather than causal. The full adjustment moves 36 of 116 models by at least ten ranks. Yet comparisons with external benchmarks do not show that adjusted rankings better measure capability. We therefore recommend publishing raw and adjusted rankings side by side as a transparent sensitivity analysis.
[NLP-58] DAYJOB: A Benchmark for Long-Horizon Professional Work NEURIPS2026
【速读】: 该论文旨在解决专业领域中复杂任务执行的评估难题,即如何在真实工作场景下衡量生成式AI系统在处理模糊、开放性任务时的综合能力。其核心问题是:现有基准多聚焦于封闭式问答或单一技能测试,无法反映专业人士在面对不完整、模糊或潜在错误前提时,从信息筛选、推理判断到最终输出的全流程决策能力。为应对这一挑战,研究者构建了DAYJOB基准,涵盖医疗(50项)与金融(80项)领域的130个真实职业任务,每项任务平均需专业人士耗时13.6小时(医疗)和16.6小时(金融),具有高度现实复杂性。解决方案的关键在于引入“容器化港口环境(Harbor environment)”与“基于二元标准的专家评分规则(expert rubric of binary criteria)”,确保每个任务的完成必须满足所有预设判别条件,且由代理型评估者(agentic judge)进行自动化验证,从而实现对任务执行质量的严格量化。实验表明,当前最强模型Claude Opus 5.5在医疗与金融任务中的通过率分别为24.7%和23.9%,而多数模型配置通过率不足3%,揭示出现有生成式AI在处理专业级任务时仍存在严重缺陷,如误信矛盾数据、沿用错误输入等系统性偏差。研究团队已开源全部医疗任务、部分金融任务、评估工具链及排行榜,以推动高阶专业智能体的发展。
链接: https://arxiv.org/abs/2610.01306
作者: Stephanie Finley,Liudas Panavas,Thomas Mikkelson,Cam Hinton,Stacey Ganss,Bradley Monton,Emily Kendall,Michelle Spradlin,Lydia Bye,Michael O’Brien,Lauren Ylvisaker,Derek Ray,Suhaas Garre,Sushant Mehta,Edwin Chen
机构: Surge AI
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: 11 pages, 4 figures, 3 tables. An earlier version was accepted to the 2nd Workshop on Agentic AI Benchmarks and Applications for Enterprise Tasks (AABA4ET) at NeurIPS 2026. Evaluation harness: this https URL
Abstract:Professional work often starts with a brief request that leaves the professional to work out what is needed, which documents matter, and whether the request’s premise holds. We introduce DAYJOB, a benchmark of 130 tasks built by professionals in healthcare (50) and finance (80). The tasks are estimated to take a professional 13.6 hours on average in healthcare and 16.6 in finance. Each task is a containerized Harbor environment with an expert rubric of binary criteria (median 47.5 and 57.5 per task) that an agentic judge applies to the delivered files, and an attempt passes only if it meets every criterion. Across 30 model configurations from 13 developers, the strongest, Claude Opus 5.5, passes 24.7% of healthcare and 23.9% of finance attempts, and the median configuration passes 0.6% and 2.5%. In case studies, agents accept premises that the record contradicts and carry wrong inputs through otherwise consistent analyses. We release all healthcare tasks, 50 of the 80 finance tasks, the evaluation harness, and the leaderboard.
[NLP-59] SCOPE-AD: Sequential cost-aware ordinal-belief planning with energy-based models for diagnostic agents
【速读】: 该论文旨在解决阿尔茨海默病(Alzheimer’s Disease, AD)诊断中因检测成本异质性和患者负担导致的序列化证据获取效率低下问题。传统固定模态预测模型无法协同决策应采集何种检测、以及在何种条件下现有证据已足够进行诊断,从而导致资源浪费与延迟。其解决方案的关键在于提出一种名为SCOPE-AD(基于能量模型的序贯成本感知序数信念规划诊断代理)的框架,该框架通过引入一种掩码感知的序数模型来表征从认知正常(Cognitively Normal, CN)到轻度认知障碍(Mild Cognitive Impairment, MCI)再到阿尔茨海默病(AD)这一连续谱系中的不确定性。利用回顾性训练数据生成贝尔曼目标,由能量模型作为教师指导,将动作分布蒸馏至Qwen策略网络。在部署阶段,诊断代理在受限于可用检测项和预算的前提下,自主决定是否继续采集新证据或终止并输出诊断结果,且无需访问未获取的测试值。每次采集后,系统动态更新证据集与序数信念,以支持下一阶段决策。在ADNI数据集上的实验表明,SCOPE-AD在平均采集成本仅50.46美元的情况下实现了77.70%的宏平均F1分数,较最强基线提升9.34个百分点;而全模态评估虽使F1提升1.89个百分点,但采集成本飙升116.7倍,验证了选择性采集在实现高效、低成本诊断中的关键优势。
链接: https://arxiv.org/abs/2610.01278
作者: Ziwen Yu,Ivan Koychev,Elizabeth Coulthard,Ting Zhou,Bolin Chen,Dian Hong,Zinuo You,Yujiao Wang,Anthony Mulholland,Qiang Liu
机构: University of Manchester (曼彻斯特大学); University of Oxford (牛津大学); Chinese Academy of Sciences (中国科学院); National University of Singapore (新加坡国立大学)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: 5 pages,2 figures
Abstract:Alzheimer’s disease (AD) diagnosis requires sequential evidence acquisition under heterogeneous test costs and patient burden. Fixed-modality predictors do not jointly decide which test to acquire or when the available evidence is sufficient for diagnosis. We propose SCOPE-AD (Sequential Cost-Aware Ordinal-Belief Planning with Energy-Based Models for Diagnostic Agents) for cost-aware classification of cognitively normal (CN), mild cognitive impairment (MCI), and AD cases. A mask-aware ordinal model represents uncertainty along the ordered CN–MCI–AD continuum. Retrospective training records provide sampled Bellman targets for an energy-based teacher, whose action distributions are distilled into a Qwen policy. At deployment, the agent selects acquisition or diagnosis actions under availability and budget constraints without access to unacquired values. After each acquisition, the evidence and ordinal belief are updated before the next decision. On ADNI, SCOPE-AD achieves 77.70% Macro-F1 at an average acquisition cost of \ 50.46, exceeding the strongest evaluated baseline by 9.34 percentage points. Full-modality evaluation raises Macro-F1 by only 1.89 points while increasing acquisition cost by 116.7 times. These results support selective acquisition for cost-effective diagnosis.
[NLP-60] Know When to Hold em: Correct-Token Retention in Uniform-State Diffusion Language Models
【速读】: 该论文旨在解决统一状态扩散模型(Uniform-State Diffusion Models, USDMs)在自纠正过程中无法有效保留正确标记(correct tokens)的问题。尽管USDM具备在任意去噪步骤中修正任意标记的能力,从而实现自我纠错,但现有模型普遍存在过度修改本已正确的标记,导致大量不协调的更新操作,严重损害生成样本的多样性。研究表明,当前最先进的USDM(如DUO、UDLM和均匀噪声SED)即使在贪婪尾部解码下,每步仍会无差别地重写173–270个位置,且其对干净与受损标记的重建准确率几乎相同,表明模型缺乏对未受扰动标记的保留机制。进一步分析验证集NELBO发现,训练目标对错误预测在扰动位置施加强惩罚,但在未扰动位置几乎不惩罚正确预测,导致模型缺乏保留正确信息的动力。为此,论文提出一种简单而有效的辅助损失——正确标记保留正则化(Correct-Token Retention Regularization, CTR-Reg),通过鼓励模型在前向过程未扰动的位置保持原有标记,无需修改采样器即可提升保留能力。实验表明,CTR-Reg在六个基准测试上平均提升干净标记准确率26.5个百分点,同时保持受损标记准确性基本不变,每步修订位置减少至仅3–11个;仅需五步贪婪尾部解码,生成困惑度即下降超过一半,且样本多样性得以维持。研究揭示了正确标记保留是实现高效自纠正扩散语言模型的关键缺失要素,并提供了一种有效解决方案。
链接: https://arxiv.org/abs/2610.01275
作者: Mojtaba Nafez,James Henderson
机构: EPFL(洛桑联邦理工学院); Idiap(伊迪亚研究所); Lausanne, Switzerland(洛桑, 瑞士); Idiap Research Institute(伊迪亚研究所); Martigny, Switzerland(马蒂尼, 瑞士)
类目: Computation and Language (cs.CL)
备注: 38 pages, 8 figures
Abstract:Uniform-state diffusion models (USDMs) can revise any token at any denoising step, which lets them correct their own mistakes, a key advantage over masked diffusion. Self-correction, however, requires both revising incorrect tokens and retaining correct ones, and we show that current USDMs lack the latter. Even under greedy-tail decoding, state-of-the-art USDMs (DUO, UDLM, and uniform-noise SEDD) keep revising 173–270 of 512 positions at every step, and these large, uncoordinated edits collapse sample diversity. A random-token corruption experiment traces this deficit to the models themselves: they reconstruct clean and corrupted tokens with nearly identical accuracy, even though clean tokens are easier targets. A decomposition of the validation NELBO shows that training barely rewards retention: incorrect predictions are heavily penalized at corrupted positions but almost free at clean ones. We propose Correct-Token Retention Regularization (CTR-Reg), a simple but effective auxiliary loss that trains the model to retain tokens left unperturbed by the forward process and requires no change to the sampler. CTR-Reg improves clean-token accuracy by 26.5 percentage points on average across six benchmarks, while leaving corrupted-token accuracy virtually unchanged, and its per-step revisions converge to only 3–11 positions. With just five greedy-tail steps, generative perplexity more than halves under CTR-Reg for all three models while diversity is preserved, and these gains hold across sampling budgets. Our results identify correct-token retention as a key missing ingredient for self-correcting diffusion language models, and demonstrate an effective fix.
[NLP-61] Science Utopia? Closed-Loop LLM Simulation of Academic Research Ecosystems
【速读】: 该论文旨在解决科学生态系统中多主体、长周期演化过程的复杂性问题,特别是人工智能(AI)深度介入科研全链条背景下,研究方向选择、合作网络、投稿与评审、引用、资金分配及研究人员流失等动态交互机制如何共同影响科学进步。其核心解决方案是提出一个名为SciUtopia的持续性闭环大语言模型(LLM)代理仿真框架,通过建模上述关键科学流程并维持跨年演化的状态,实现对学术研究生态系统的高保真模拟。该框架的关键在于具备可配置的机构机制与信息传播通道,支持可控的反事实实验与靶向干预分析;在61个仿真世界中,模拟超过4万名研究人员和8,000所机构,生成约40万次出版决策与120万条由LLM生成的同行评审意见,揭示出拒稿驱动的反复投稿显著加剧了审稿人负担,谨慎探索策略能在提升引用影响力与保障职业成功及长期主题多样性之间取得平衡,且资源不平等可在缺乏明显早期胜出累积优势的情况下自发形成。
链接: https://arxiv.org/abs/2610.01257
作者: Yiqiao Jin,Yiyang Wang,Lucheng Fu,Bing He,Siheng Xiong,Yijia Xiao,B. Aditya Prakash,Josiah Hester,Srijan Kumar,James Evans,Jindong Wang
机构: 未知
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Computers and Society (cs.CY)
备注: this https URL
Abstract:Scientific progress emerges from a longitudinal ecosystem in which researchers, institutions, funding agencies, collaboration networks, and the scientific literature co-evolve. As AI becomes increasingly involved throughout the scientific research cycle, understanding these interconnected and evolving processes becomes increasingly important. We introduce SciUtopia, a persistent, closed-loop LLM-agent simulation framework for studying academic research ecosystems. SciUtopia models interconnected scientific processes such as research-direction choice, collaboration, submission, peer review, resubmission, citation, funding, and researcher attrition, while maintaining evolving states across simulated years. Its configurable institutional mechanisms and information channels provide a controlled testbed for matched counterfactual experiments and targeted interventions. Across 61 simulation worlds, SciUtopia simulates over 40,000 researchers from 8,000 institutions, producing around 400,000 publication decisions and 1.2 million LLM-generated peer reviews. Using these longitudinal simulations, we find that rejection-driven resubmission substantially amplifies reviewer burden beyond population growth alone, cautious exploration balances citation impact with career success and long-term topic diversity, and resource inequality can emerge even without detectable cumulative advantage from narrowly winning early funding. Code is available at this https URL.
[NLP-62] Revision-Aware Independent Agent Graphs for Dynamic Reasoning
【速读】: 该论文旨在解决传统推理协议中任务绑定固定不变所导致的局限性,即无法有效评估智能体是否能够动态更新相关信息、保留未受影响的推理成果,或重构历史任务绑定。为此,作者提出动态任务路由(dynamic task routing)这一新范式,其中事件流会持续修改任务绑定,系统需在每个查询时刻选择对应时间点有效的文档版本并完成求解。为实现该研究目标,作者将六个主流基准测试(MMLU、MMLU-Pro、MedMCQA、MATH、GPQA、HumanEval)重构为包含31,119个动态任务实例、共373,428个时序标记查询的大规模数据集。该设置揭示了一个核心权衡:频繁重计算造成资源浪费,而未经审查的缓存复用则可能导致过时结论。论文提出的解决方案是修订感知独立智能体图(Revision-Aware Independent Agent Graph, RIAG),其关键在于将确定性的时序解析与任务推理解耦,通过不可变的文档标识符缓存结果,并为每个新任务初始提供两次未暴露的尝试机会;同时基于条件触发审计与修复机制,每份文档版本最多调用四次。实验表明,同质化RIAG在0.62次调用/查询下达到54.24%的联合路由与答案准确率,显著优于最强对比方法(32.22%准确率,18.00次调用/查询),异质化RIAG亦达49.78%准确率(0.63次调用/查询),验证了其高效性与鲁棒性。
链接: https://arxiv.org/abs/2610.01249
作者: Yan Luo,Selim-Antoine Lali,Jeremy Moebel,Iliass Khoutaibi,Ahmadou Aidara,Mengyu Wang
机构: Harvard AI and Robotics Lab (哈佛人工智能与机器人实验室); Harvard University (哈佛大学)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注:
Abstract:Conventional reasoning protocols present a fixed, preselected task, so they cannot test whether an agent propagates relevant updates, preserves unaffected work, or reconstructs a historical task binding. We therefore study \emphdynamic task routing, in which an event stream revises task bindings and a system must select the document version valid at each query time before solving it. To study this problem, we repurpose six widely used benchmarks: MMLU, MMLU-Pro, MedMCQA, MATH, GPQA, and HumanEval into 31,119 dynamic episodes comprising 373,428 temporally categorized queries. This setting exposes a central trade-off: recomputing after every event wastes work, whereas unguarded reuse returns stale conclusions. We introduce the Revision-Aware Independent Agent Graph (RIAG), a bounded multi-agent policy that separates deterministic temporal resolution from task reasoning. RIAG caches solutions by immutable document identity, starts each fresh task with two unexposed attempts, and conditionally invokes audit and repair, using at most four calls per document version. On this collection, homogeneous RIAG achieves 54.24% joint routing-and-answer accuracy at 0.62 calls/query, compared with 32.22% at 18.00 calls/query for the strongest comparison method; heterogeneous RIAG reaches 49.78% at 0.63 calls/query.
[NLP-63] Right Answers Wrong States: Hidden Information Failures in Multi-Agent Collaboration
【速读】: 该论文旨在解决多智能体协作系统中一种隐蔽但关键的失败模式——“离查询失败”(off-query failure),即尽管当前决策正确,但协作过程导致信息状态被污染,从而影响后续推理的可靠性。传统评估仅关注最终答案是否正确,忽略了共享状态的可信度问题。其解决方案的关键在于提出ReGround框架,通过主动识别并解决证据冲突、验证共享事实、重建可信的共享状态,并基于该可靠状态进行推理,从而同时提升证据验证(T1)、共享状态重建(T2)和任务求解(T3)的能力。实验表明,在多个大模型与高风险场景(医疗与灾后响应)中,标准协作虽在任务完成率上表现尚可(平均64.7%),但信息状态可靠性极低(证据验证仅14.3%,状态重建43.1%),而ReGround在所有测试设置中均显著提升三项能力,相对增益分别达309.0%、82.9%和17.6%,证明了构建可靠共享状态对于实现真正可靠的协作至关重要。
链接: https://arxiv.org/abs/2610.01244
作者: Herun Wan,Jiaying Wu,Minnan Luo,Zihan Ma,Fanxiao Li,Nancy F. Chen,Min-Yen Kan
机构: Xi’an Jiaotong University (西安交通大学); National University of Singapore (新加坡国立大学); Yunnan University (云南大学); Agency for Science, Technology and Research (A*STAR) (新加坡科技研究局)
类目: Computation and Language (cs.CL)
备注:
Abstract:Multi-agent systems are often judged by whether they reach the correct answer. This can miss a distinct failure: collaboration may leave behind a corrupted information state even when the immediate decision is correct. We call this an off-query failure. To study this failure in collaborative decision support, we introduce OffQuery, which separately evaluates evidence verification (T1), shared-state reconstruction (T2), and task resolution (T3) in two representative high-stakes settings: healthcare and disaster response. Across GPT, Gemini, and Qwen models, standard collaboration shows much stronger task performance than state reliability. Averaged over 21 model–setting combinations, task resolution reaches 64.7%, while evidence verification and state reconstruction reach only 14.3% and 43.1%. We trace this gap to selective information use: current queries often bypass corrupted facts, which become consequential when later tasks require them. We further introduce ReGround, which resolves conflicting evidence, verifies shared facts, reconstructs a trusted state, and reasons over that state. Across seven models from three families, ReGround improves all three capabilities in every evaluated setting, with average relative gains of 309.0%, 82.9%, and 17.6% on T1, T2, and T3. Reliable collaboration therefore requires both a correct decision and a reliable shared state for future reasoning.
[NLP-64] Evaluating the Robustness of Japanese LLM s to IME-Related and Typographical Errors
【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在日语场景下对拼写错误的鲁棒性不足的问题,尤其关注日语特有的文本输入特性,如多书写系统共存及基于输入法编辑器(IME)的转换机制。其核心挑战在于,现有研究普遍忽视了真实使用环境中高频出现的日语特定拼写错误类型对模型性能的影响。为此,作者提出五类具有代表性的日语拼写错误类型:字符颠倒(Character Transposition)、字符替换(Character Replacement)、同音转换(Homophone Conversion)、日语IME转换(Japanese IME Conversion)以及全角转换(Full-Width Conversion),并将其应用于三个日语基准数据集(JMMLU、JCommonsenseQA 与 JamC-QA),评估了十一款日语及多语言LLMs的性能表现。研究发现,字符颠倒与字符替换两类错误对模型准确率造成显著负面影响,而其他三类错误影响相对有限。这一结果揭示了当前日语大模型在面对严重扭曲原始输入的错误时仍存在明显脆弱性,强调了在实际输入环境下开展鲁棒性评估的重要性。解决方案的关键在于构建面向真实日语输入场景的系统性错误扰动框架,并通过实证分析识别出最易导致模型失效的错误类型,为后续模型优化提供明确方向。
链接: https://arxiv.org/abs/2610.01241
作者: Ryota Mibayashi,Hiroaki Ohshima
机构: Kobe University (神户大学); University of Hyogo (兵库大学)
类目: Computation and Language (cs.CL)
备注:
Abstract:Large language models (LLMs) have achieved strong performance across various natural language processing tasks. However, their robustness to typographical errors remains underexplored, particularly in Japanese, where text input involves multiple writing systems and IME-based conversion. In this study, we evaluate the robustness of Japanese LLMs against realistic Japanese-specific typos. We introduce five typo categories: Character Transposition, Character Replacement, Homophone Conversion, Japanese IME Conversion, and Full-Width Conversion. These perturbations are applied to three Japanese benchmark datasets (JMMLU, JCommonsenseQA, and JamC-QA), and eleven Japanese and multilingual LLMs are evaluated. The results show that Character Transposition and Character Replacement typos consistently reduce accuracy across benchmarks, whereas IME Conversion, Full-Width Conversion, and Homophone Conversion have relatively limited impact. These findings reveal that current Japanese LLMs remain vulnerable to realistic Japanese typing errors, particularly those that substantially distort the original input, highlighting the importance of robustness evaluation in practical input environments.
[NLP-65] Harness Annealing: Learning to Act with Less External Control
【速读】: 该论文旨在解决语言智能体(language agent)在执行复杂任务时对运行时外部控制依赖过强的问题,即其状态管理、工作流组织及答案验证等关键决策仍需依赖外部“支架”(harness)实时提供指导。现有方法虽可通过成功轨迹的训练提升任务表现,但控制权始终保留在外部系统中。为此,论文提出“支架内化”(harness internalization)目标,即让模型在训练过程中学习承担原本由支架负责的控制职责,从而在移除外部支持后仍能保持任务性能。其核心解决方案是引入支架退火训练(HARNESS ANNEALING TRAINING, HAT),通过渐进式削弱训练阶段所用支架强度,构建一个从强监督到弱监督的教师轨迹课程,并结合显式的控制指令监督,使模型逐步学会自主决策何时探索、是否修正及何时终止。实验在9B和35B规模模型上针对SWE-QA与SWE-QA-Pro数据集进行,结果表明:经退火优化的检查点仅使用工具即可达到接近初始全支架配置下的性能水平,证明了支架支持经验确实有助于降低模型部署时对运行时干预的需求,且效果随模型规模与部署配置而异,过度退火并不必然带来性能提升。
链接: https://arxiv.org/abs/2610.01235
作者: Yingxuan Yang,Huacan Chai,Ying Wen
机构: 未知
类目: Computation and Language (cs.CL)
备注:
Abstract:Language agents rely on external harnesses to track state, organize workflows, and verify answers. Beyond providing tools and information, these harnesses supply control decisions about what to investigate, whether to revise, and when to stop. Training on successful harness-supported trajectories can improve task performance while leaving these decisions dependent on runtime intervention. We ask whether harness-supported experience can also teach the model to make these decisions, allowing the division of control to change as the model learns. We call this objective harness internalization: learning to assume specified control responsibilities while retaining task performance after the corresponding support is withdrawn. We introduce HARNESS ANNEALING TRAINING (HAT), which combines explicit control supervision with a curriculum over teacher trajectories collected under progressively weaker harnesses. Experiments with 9B and 35B models on SWE-QA and SWE-QA-Pro evaluate every checkpoint under four deployment harnesses. Selected annealed checkpoints operating with tools alone achieve scores close to those of their respective starting checkpoints deployed with the full harness. The benefits vary with model scale and deployment configuration, and further annealing does not uniformly improve performance. These findings suggest that harness-supported experience can help reduce the runtime control required by a trained agent.
[NLP-66] ASCRIBE: Atomic and Significance-Based Reasoning for Thai Clinical SOAP Note Generation
【速读】: 该论文旨在解决临床病历自动生成中因现有推理方法忽略重要临床信息且生成内容缺乏依据而导致的可靠性问题,尤其针对泰语临床场景因缺乏公开可用数据集而进展受限的困境。其解决方案的关键在于提出ASCRIBE框架——一种受医生思维启发的推理机制,通过在摘要生成前为对话中提取的每个原子事实赋予临床重要性等级,从而提升大语言模型(LLM)作为电子病历记录员的可信度与准确性。该方法不仅在提示工程(prompting)层面显著优于链式思维提示(chain-of-thought prompting),还在医师对齐的LLM评分指标上表现更优;同时,作为基于强化学习的奖励函数(GRPO reward),可使仅使用合成数据训练的Gemma-4-E4B模型在事实精确性上媲美Gemini 3.1 Pro,并在完整性方面实现超越。此外,研究团队发布了首个去标识化的泰语真实临床会话摘要基准数据集ThaiClinicBench及对应的合成训练语料,为后续研究提供了重要基础。
链接: https://arxiv.org/abs/2610.01234
作者: Tarm Kalavantavanich,Teerawut Ponarchar,Pattaramanee Arsomngern,Jenta Wonglertsakul,Watcharakorn Chuthong,Chiraphat Boonnag,Knot Pipatsrisawat,Titipat Achakulvisut
机构: 未知
类目: Computation and Language (cs.CL)
备注:
Abstract:Automatic SOAP note generation can ease the documentation burden on physicians, but existing reasoning methods often omit clinically important information and generate unsupported content. Progress in Thai is further hindered by the lack of publicly available datasets. We propose ASCRIBE, a physician-inspired reasoning framework that ascribes a clinical-significance level to each extracted atomic fact in the conversation before summarization, making a general-purpose LLM a more reliable scribe. We also release ThaiClinicBench, the first de-identified Thai clinical summarization benchmark of real encounters, together with a synthetic training corpus derived from real clinical notes. As a prompt, ASCRIBE outperforms chain-of-thought prompting on GPT-5.4 and Gemini 3.1 Pro across the physician-aligned LLM-judge metrics and improves on standard prompting by up to 10.3 points on the completeness LLM-judge metric. As a GRPO reward, it enables a Gemma-4-E4B model trained solely on synthetic data to match Gemini 3.1 Pro in factual precision and surpass it in completeness. Code and data can be found at this https URL.
[NLP-67] AGO AI Quality Gate: Evidence-First Release Decisions for Retrieval-Augmented Generation ECML KDD2026
【速读】: 该论文旨在解决企业在采用检索增强生成(Retrieval-Augmented Generation, RAG)系统时面临的版本发布决策难题,即在证据不完整且依赖不可靠大语言模型(LLM)评判者的情况下,如何科学地做出“推进、修订或阻止”系统版本的决策。其解决方案的关键在于提出并部署了以证据为先的质量门控框架——AGO AI Quality Gate(AGO),该框架通过四大核心组件实现:(1)四状态决策模型,将缺失数据和评判者错误显式建模为决策状态;(2)分层评分机制,融合确定性检查、本地规则约束与结构化LLM评估;(3)分层贝塔-二项分布门控机制,概率化量化回归风险;(4)强制性的元评估协议,在判定影响决策前对LLM评判者进行有效性验证。研究基于公开的RAGBench基准对评判者性能进行评估,结果显示即使低资源模型(gpt-4.1-nano)在协议输出上表现完美,其检测非合规回答的能力仅略高于随机水平(AUROC 0.603),而gpt-4o虽显著提升至0.783,但各领域表现仍存在0.62至0.88的波动。固定种子门控实验进一步表明,相较于朴素门控,该框架可将不当发布率从29.3%-41.8%降低至22.2%-35.1%,证实了必须按具体场景评估评判者质量,且单一点估计不足以支撑发布决策。
链接: https://arxiv.org/abs/2610.01218
作者: Giulio Zeloni,Enrico Lo Conte,Salvatore Rionero,Giuseppe Santoro,Alessandro Rastelli,Fabio Sorrentino
机构: Protom Group S.p.A.(Protom集团股份有限公司)
类目: Computation and Language (cs.CL)
备注: 14 pages, 1 figure, 4 tables. Submitted version (pre-review). Accepted at NFMCP 2026, ECML PKDD 2026 Workshops
Abstract:Enterprises adopting retrieval-augmented generation (RAG) face a recurring operational decision: promote, revise, or block a system version. The evidence is incomplete and the metrics come from fallible LLM judges. We report on AGO AI Quality Gate (AGO), an evidence-first quality-gate framework deployed in industrial RAG assessment engagements. AGO integrates four key components: a four-state decision model that treats missing data and judge errors as explicit outcomes; layered scoring combining deterministic checks, local guardrails, and structured LLM evaluation; a stratified beta-binomial gate that quantifies regression risk probabilistically; and a mandatory meta-evaluation protocol to validate the LLM judge before it influences decisions. Since engagement data is proprietary, we evaluate the judge layer on RAGBench, a public benchmark of 100k annotated RAG traces across 12 datasets. On identical stratified test samples (N=1200 per judge), a low-cost judge (gpt-4.1-nano) detects non-adherent answers barely above chance (AUROC 0.603 [0.570, 0.634]), despite producing flawless protocol output, while gpt-4o reaches 0.783 [0.756, 0.807] – yet its per-domain performance still ranges from 0.62 to 0.88. A fixed-seed gate study spanning regression, no change, and improvement quantifies unsafe promotion, false-alarm cost, and improvement throughput. Under regression, the decision-grade profile reduces unsafe promotion to 22.2%-35.1%, against 29.3%-41.8% for a naive gate. These results support the design choices that judge quality must be measured per engagement and that point estimates alone are not a release decision.
[NLP-68] ReCast: Contract-Preserving Protection for Fixed-Interface Multimodal Reasoning
【速读】: 该论文旨在解决远程多模态模型在处理图表与语音等敏感输入时,因数据传输导致隐私泄露的问题。现有文本净化方法无法适配固定媒体接口,而身份匿名化又无法隐藏任务内容本身,导致源信息暴露风险。其核心解决方案是提出ReCast——一种代理式插件框架,通过本地化将原始输入转换为统一的文本证据-查询记录,利用轻量化40亿参数模型联合重写实体与主题,并借助局部可逆、角色感知的数值映射替换具体数值;随后由重建代理基于保护后的记录生成并验证所需媒体形式。远程求解器返回程序后,本地还原被保护的操作数再执行,从而实现对敏感内容的有效屏蔽。在4,000个保留的ChartQA和NMSQA测试样本上,ReCast达到75.10%的准确率,保留了未受保护状态下92.43%的远程推理性能,且模型审计发现仅7.95%的求解请求存在源内容泄露,显著优于所有本地基线方法,在保持远程推理优势的同时有效降低了现有媒体接口下的源内容暴露风险。
链接: https://arxiv.org/abs/2610.01184
作者: Bingchen Pei,Lichong Chen,Bingxi Zhao,Ziang Wu,Sirui Wang,Min Zhang,Yanhao Chen,Qingxu Liu,Qiang Gao,Chang-Tien Lu,Bo Gao
机构: Beijing Jiaotong University (北京交通大学); Virginia Polytechnic Institute and State University (弗吉尼亚理工学院暨州立大学)
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: 24 pages, 10 figures
Abstract:Remote multimodal models offer strong numerical reasoning capabilities over charts and speech, but sending private inputs risks exposing sensitive content. Text-only sanitization cannot directly satisfy fixed media interfaces, while identity anonymization leaves the underlying task content exposed. We introduce ReCast, an agentic plug-in framework that replaces source-specific content while preserving task-relevant relations and the required input modality. ReCast locally converts inputs into a shared textual evidence-query record, jointly rewrites entities and topics with a distilled 4B model, and substitutes values through a locally invertible, role-aware numerical map. A reconstruction agent generates and validates the required media from the protected record. The remote solver returns a program whose protected operands are restored locally before execution. On 4,000 held-out ChartQA and NMSQA examples, ReCast achieves 75.10% accuracy, retaining 92.43% of unprotected remote accuracy, while a model-based audit flags source-content leakage in 7.95% of solver-bound requests. It outperforms all evaluated local baselines, preserving the benefit of remote reasoning while reducing source-content exposure under existing media interfaces.
[NLP-69] mporally-Resolved Token Attribution Reveals the Generation Dynamics of Diffusion Language Models
【速读】: 该论文旨在解决扩散语言模型(Diffusion Language Models, DLMs)中缺乏有效且可解释的标记归因方法的问题,特别是在复杂去噪过程中如何量化输入标记对模型输出的贡献。现有方法难以捕捉模型在多步去噪和多层处理中的动态决策机制,限制了对模型内部推理过程的理解。为此,论文提出了一种名为“扩散层积分梯度”(Diffusion Layer Integrated Gradients, DLIG)的新方法,其核心创新在于将积分梯度(Integrated Gradients, IG)理论扩展至任意网络层和去噪步骤,从而实现对DLM在生成过程中逐步依赖输入信息的精细化归因分析。DLIG的关键在于建立了与原始积分梯度四大公理——完备性(completeness)、实现不变性(implementation invariance)、线性性(linearity)及对称性保持(symmetry preservation)——之间的直接对应关系,确保归因结果具有理论一致性与可解释性。作为一种轻量级补充工具,DLIG无需复杂的干预实验即可快速验证机制假设,显著提升了对模型在词义消歧、多跳图推理和句子补全等任务中跨位置、跨层及跨去噪阶段的信息利用模式的洞察力。
链接: https://arxiv.org/abs/2610.01177
作者: Darpan Aswal,Céline Hudelot
机构: Université Grenoble Alpes; MICS, CentraleSupélec
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:
Abstract:This work presents Diffusion Layer Integrated Gradients (DLIG), a token attribution method for diffusion language models (DLMs) that extends Integrated Gradients (IG~\citesundararajan2017axiomatic) to arbitrary layers and denoising steps. DLIG attributes a DLM’s progressive commitment to a self-generated or fixed completion for an input prompt. We establish direct correspondences between DLIG and the IG axioms of completeness, implementation invariance, linearity, and symmetry preservation. As a lightweight complement to interventional analysis, DLIG provides an inexpensive first check of mechanistic hypotheses across the denoising trajectory. We demonstrate this on word-sense disambiguation, multi-hop graph reasoning, and sentence infilling, revealing how DLMs draw on inputs across positions, layers, and denoising steps.
[NLP-70] HeadEdit: Calibrating Language Model Behavior Through the Frozen Unembedding Matrix
【速读】: 该论文旨在解决大语言模型在对齐(alignment)后仍存在的行为错误问题,例如无端拒绝合理请求、调用不必要的工具或屈从于虚假用户主张等。现有方法通常将此类问题视为计算层面的偏差,而忽视了理想行为是否已编码于模型内部表示中。其解决方案的关键在于提出一种无需梯度的校准方法——HeadEdit,通过模型的解嵌矩阵(unembedding matrix)实现行为调节。HeadEdit从成对生成结果中提取低秩的行为子空间,并利用每个提示在该子空间中的坐标生成全局词汇修正,从而实现无需手动指定目标词元或更新参数的隐式自适应控制。实验表明,HeadEdit在三个任务和三个模型家族的九种设置中均取得性能提升,推理开销极小且未造成通用能力的系统性损失;同时揭示了与基于梯度对齐方法的内在联系:该低维子空间可部分预测偏好微调对未见提示输出概率分布的影响,且在微调后可重复使用,无需重新提取或调整,显著提升了可复用性。这一方法为通过解嵌矩阵进行轻量级、可解释的行为校准提供了有效途径。
链接: https://arxiv.org/abs/2610.01170
作者: Zirui He,Haiyan Zhao,Jingyu Hu,Yinghao Wu,Chenxi Yuan,Yingcong Li,Yandong Bai,Mengnan Du
机构: New Jersey Institute of Technology (新泽西理工学院); University of Bristol (布里斯托大学); Kuaishou Technology (快手科技); The Chinese University of Hong Kong, Shenzhen (香港中文大学深圳校区)
类目: Computation and Language (cs.CL)
备注: 31 pages, 18 figures, 7 tables
Abstract:Alignment does not eliminate behavioral errors in language models. Models may still refuse benign requests, call unnecessary tools, or yield to false user claims. Current methods mitigate such errors as a computation problem, and rarely explore if the desired behavior is already encoded in the model’s representation. Motivated by the observation that behavior-relevant information remains linearly decodable from the final hidden state even when the resulting logits produce the undesired behavior, we introduce HeadEdit, a gradient-free method that calibrates model behavior through the unembedding matrix. HeadEdit extracts a low-rank behavioral subspace from paired completions and uses each prompt’s coordinates within it to generate a vocabulary-wide correction, thereby implementing implicitly adaptive steering without manually specified target tokens or parameter updates. HeadEdit improves all nine experimental settings across three tasks and three model families, with negligible inference overhead and no systematic loss of general capabilities. It also reveals a connection to gradient-based alignment. HeadEdit’s low-dimensional representation partly predicts how preference tuning changes output logits on unseen prompts. The subspace learned from the model can also be reused after tuning, improving performance without re-extracting or retuning. These results show that HeadEdit provides a practical, lightweight, and interpretable way to calibrate model behavior through the unembedding matrix.
[NLP-71] Persistent Depth Ordering amid Shifting Block-Bypass Responses in Language Model Pretraining
【速读】: 该论文旨在解决语言模型在预训练过程中内部表征与计算动态演化问题,即现有层干预(layer intervention)研究多基于单个训练检查点,难以区分哪些深度依赖的干预响应反映的是模型固有的、稳定的组织结构,哪些仅是训练过程中的暂时性结果。其解决方案的关键在于采用固定教师强制(teacher-forced)上下文下的单块身份旁路(single-block identity bypass)方法,系统分析五个公开发布的训练轨迹及11种模型-领域组合的纵向变化。研究发现,尽管干预响应的幅度随时间重新分布,但其深度顺序仍保持可识别性;相近检查点间的响应秩相关性更强,显著变化集中于跨文本样本重复出现且可在不同评估领域间迁移的位置。控制实验进一步表明,自然旁路效应的变化无法归因于单一下游敏感性:在重复的Pythia实验中,局部缺失更新量增大而整体匹配下游响应下降,而OLMo-2 7B则表现出不同平衡。此外,匹配响应对扰动强度和方向均具依赖性,未观察到明确的补偿机制。综上,研究揭示了纵向层敏感性具有结构性而非静态特征,强调应结合扰动路径在训练过程中的演化背景来解释单检查点干预结果。
链接: https://arxiv.org/abs/2610.01165
作者: Shengye Tao,Yinzhu Cheng,Haihua Xie
机构: Beijing University of Civil Engineering and Architecture(北京建筑工程大学); Beijing Institute of Mathematical Sciences and Applications (BIMSA)(北京数学科学应用研究院); Institute of Statistics and Big Data, Renmin University of China(中国人民大学统计与大数据研究院)
类目: Machine Learning (cs.LG); Computation and Language (cs.CL)
备注: 24 pages, 12 figures
Abstract:Layer interventions are widely used to probe the internal organization of language models, yet most analyses examine a single training checkpoint even though model representations and computations evolve throughout pretraining. This leaves open which depth-dependent intervention responses reflect persistent organization and which are transient consequences of training. We study this question using single-block identity bypass on fixed teacher-forced contexts across five released trajectories and 11 model-domain combinations. We find that block-bypass responses retain recognizable depth ordering while their magnitudes redistribute: nearby checkpoints preserve stronger rank correspondence than distant ones, and large changes concentrate at positions that recur across text samples and transfer across evaluation domains. Controlled experiments further show that changes in the natural bypass effect cannot be reduced to a single downstream sensitivity: in replicated Pythia runs, local missing-update magnitude grows while the pooled matched downstream response decreases, whereas OLMo-2 7B exhibits a different balance. These matched responses also depend on perturbation strength and direction, without identifying targeted compensation. Together, our results show that longitudinal layer sensitivity is structured but not static, and that single-checkpoint intervention responses should be interpreted in the context of how the underlying perturbation pathway evolves during training.
[NLP-72] My FAULT: Self-Diagnosis as Credit Assignment in Self-Evolving Agent ic Reinforcement Learning
【速读】: 该论文旨在解决生成式智能体在长时序多步任务中因依赖最终结果奖励(terminal outcome rewards)所引发的信用分配(credit-assignment)难题。具体而言,当多个轨迹产生相同最终结果时,传统方法无法从中提取有效的学习信号;同时,终端奖励仅提供全局性反馈,难以定位导致失败的具体决策步骤。现有方法虽引入基于轨迹分析的细粒度信息(如自然语言反思),但其诊断内容存在可靠性不足、缺乏量化误差影响能力的问题,难以直接用于精确的信用分配。为此,本文提出自诊断引导的终端信用重分配方法(FAULT),通过验证诊断证据并从任务结果中在线学习相对错误成本,将诊断出的错误转化为与最终结果锚定的显式步骤级信用。在训练过程中,策略网络与自诊断模块协同进化,错误成本动态更新。实验表明,FAULT在ALFWorld任务中实现了95%的信号覆盖率(显著优于GRPO的41%和GiGPO的72%),有效恢复了同结果组的学习信号,并能更精准地定位错误发生的具体步骤;在长时序任务(ALFWorld、WebShop)上表现优异,同时在短时序搜索问答任务中保持竞争力。关键创新在于构建了可自我验证、可在线优化的错误成本机制,实现了从非结构化诊断到可执行信用分配的转化。
链接: https://arxiv.org/abs/2610.01161
作者: Yihua Zhu,Qianying Liu,Weixu Qiao,Xuan Ren,Weiwei Xu,Wenbo Li,Wei Wang,Ruijia Chen,Xinmiao Luan,Yin Luo,Hao Huang,Xiang Zheng,Hidetoshi Shimodaira
机构: 未知
类目: Computation and Language (cs.CL)
备注: Preprint
Abstract:Agentic reinforcement learning (RL) has emerged as a powerful approach for training large language model agents on multi-step tasks, yet reliance on terminal outcome rewards creates two credit-assignment problems, particularly in long-horizon tasks. First, same-outcome rollout groups provide no learning signal from terminal rewards. Second, terminal rewards provide only trajectory-wide feedback, making it difficult to identify which decisions caused a failure. Recent work supplements terminal rewards with finer-grained information from trajectory analysis, such as natural-language reflections on intermediate decisions and errors. However, natural-language diagnoses are difficult to use directly for credit assignment: their error claims may be unreliable, and they do not quantify how much each error should affect learning. We propose Self-Diagnosis-guided Terminal Credit Redistribution (FAULT), which turns diagnosed errors into explicit step-level credit anchored by terminal outcomes. FAULT checks diagnostic evidence and learns relative error costs from task outcomes. During training, the policy and self-diagnoser co-evolve, while error costs are updated online from recent outcomes. On ALFWorld, FAULT recovers learning signals from same-outcome groups, reaching 95% signal coverage versus 41% for GRPO and 72% for GiGPO, while better localizing credit to specific error steps. Across two model scales, FAULT delivers strong. improvements on the long-horizon ALFWorld and WebShop tasks while remaining competitive on short-horizon Search-based QA.
[NLP-73] BanglaDial-Abuse: A Corpus-Grounded Dataset for Regional Dialect Identification in Abusive Bangla Text
【速读】: 该论文旨在解决孟加拉语(Bangla)自然语言处理中因区域语言变异带来的挑战,尤其是在非正式和非标准文本中的滥用与敌意语言识别问题。其核心问题是:现有资源缺乏针对具有攻击性或敌意内容的区域性方言(如查塔格拉、锡尔赫特、巴里沙尔等)的平衡标注数据,导致模型在跨区域场景下的泛化能力受限。为此,论文提出构建一个名为BanglaDial-Abuse的平衡语料库,关键在于采用基于语料库的合成方法,在保持原始敌意或攻击性语义的前提下,系统性地引入不同区域方言在代词、所有格形式、动词形态、否定结构、疑问句式、后置词、词汇选择及孟加拉文拼写习惯等方面的差异,从而生成具有真实语言变异特征的1,000句样本(每类250句)。该数据集在句长分布上相近,但词汇空间存在部分差异(成对Jaccard相似度介于0.37至0.56之间),并明确聚焦于四分类区域方言识别任务,而非二元的有害文本检测。该资源已通过Zenodo公开共享,适用于研究与原型开发,但不作为经过母语者验证的金标准语言学资源。
链接: https://arxiv.org/abs/2610.01150
作者: Hasin Almas Sifat
机构: American International University-Bangladesh (美国国际大学-孟加拉国)
类目: Computation and Language (cs.CL)
备注: 5 pages, 3 figures, 1 table. Dataset Version 1.0 available on Zenodo: https://doi.org/10.5281/zenodo.23074319
Abstract:Regional linguistic variation remains an important challenge for Bangla natural language processing, particularly in informal and non-standard text. This paper introduces BanglaDial-Abuse, a balanced Bengali-script dataset developed for regional dialect identification in abusive and hostile Bangla text. The dataset contains 1,000 sentences distributed equally across four linguistic varieties: Standard Bangla, Chattagram, Sylhet, and Barishal, with 250 samples per class. The resource was constructed using a corpus-grounded synthetic procedure incorporating regional variation in pronouns, possessive forms, verb morphology, negation, interrogative structures, postpositions, vocabulary, and Bengali-script spelling conventions while preserving the underlying hostile or abusive meaning. Descriptive analysis shows broadly comparable sentence-length distributions but partially distinct lexical spaces across the four classes. Pairwise Jaccard vocabulary similarity ranges from 0.37 to 0.56. The primary task is four-class regional dialect identification rather than binary abusive-text detection. The dataset is publicly available through Zenodo under a Creative Commons Attribution 4.0 license. The current version is intended as a research and prototyping corpus rather than a native-speaker-validated gold-standard linguistic resource. Keywords: Bangla, Bengali, dialect identification, regional dialect, abusive language, low-resource NLP, Chattagram, Sylhet, Barishal, dataset
[NLP-74] Counting and Min-Cost Encoding for Tokenization in Large Language Models
【速读】: 该论文旨在解决主流大语言模型中分词器(Tokenizer)导致文本编码长度不一致的问题,进而影响推理效率。其核心挑战在于:不同分词器对相同文本生成的标记序列长度差异显著,而固定模型架构下更短的序列可降低推理延迟。为此,论文提出了一种名为“计数与过滤”(Counting and Filtering, CNF)的分词器训练方法,以及一种“最小代价编码”(Min-Cost Encoding, MCE)的文本编码算法。MCE通过定义基于文本片段的代价函数,并全局最小化整体分割代价来确定最优分段方式;CNF则通过直接统计有效子串构建初始词汇表,并基于MCE在训练语料上实际使用情况执行过滤步骤以形成最终词汇表。二者结合(CNF-MCE)相较于传统BPE具有更高的分词效率、更强的可扩展性及更低的依赖性。实验表明,在六类文本和两种词表规模下,CNF-MCE均优于所评估的BPE分词器;在25万词表规模下,英文网络文本压缩率分别提升26%和30%,当词表扩展至100万时,分词效率提升超60%,词表利用率从52.9%升至96.9%。此外,MCE不依赖于合并列表(如BPE)或标记概率(如UnigramLM),适用于多种词表构建方式。在18亿和80亿参数规模下从头训练的语言模型在11个基准测试中表现与使用BPE分词器的模型相当,证明了CNF-MCE可在显著提升分词效率的同时保持下游任务性能竞争力。
链接: https://arxiv.org/abs/2610.01127
作者: Shuming Shi,Xiang Zhang,Hao Yu,Wenbo Fei,Changjian Wang,Zhan Wang,Guoqing Pang,Guangye Yu,Quan Lu,Ning Jiang
机构: Mashang Consumer Finance Co., Ltd.(中国); National-Mathematics Artificial Intelligence Institute in Chongqing (NMAII)(中国)
类目: Computation and Language (cs.CL)
备注:
Abstract:Mainstream large language models rely on a tokenizer to encode text into a token sequence. Different tokenizers may yield token sequences of substantially different lengths for the same text. With a fixed model architecture, shorter token sequences correspond to lower inference time. We propose a tokenizer training approach named Counting and Filtering (CNF) and a text encoding algorithm called Min-Cost Encoding (MCE). MCE defines a cost function over a text segment, and determines the best segmentation by globally minimizing the overall segmentation cost. CNF builds a raw vocabulary by directly counting valid substrings, and then constructs the final vocabulary through a filtering step based on actual token usage when segmenting the training corpus with MCE. The CNF-MCE conbination offers several advantages over BPE, including higher token efficiency, greater scalability, and lower dependency. Across six text categories and two vocabulary-size groups, CNF-MCE consistently achieves better compression than the evaluated BPE tokenizers. With a 250K vocabulary, CNF-MCE increases compression rate by 26% and 30% on English web text over the o200k_base and qwen250k tokenizers. Experiments scaling the vocabulary to 1M entries on English web text demonstrate sustained improvements over BPE, with a token efficiency improvement of over 60% and vocabulary utilization rising from 52.9% to 96.9%. The MCE algorithm does not depend on a merge list (as in BPE) or token probability (as in UnigramLM), making it applicable to a wide range of vocabularies, including those built from BPE, UnigramLM, CNF, and others. Language models trained from scratch at the 1.8B and 8B scales achieve comparable average performance to models using the BPE tokenizers across 11 benchmarks. These results demonstrate that CNF-MCE can improve token efficiency significantly while maintaining competitive downstream performance.
[NLP-75] AgSpec: Pushing the Limits of Retrieval-Based Speculative Decoding in Coding Agent Pipelines
【速读】: 该论文旨在解决生成式AI(Generative AI)在编码代理(coding agents)流水线中,基于检索的推测解码(Retrieval-based Speculative Decoding, SD)方法存在的两大核心问题:一是现有检索库中缺乏可复用文本,或其存储形式与代理输出格式不一致,导致检索效率低下;二是现有方法未考虑不同代理间及多轮交互中目标接受长度(accept length)的动态变化,导致推测草案长度固定且不适应实际需求。其解决方案的关键在于提出AgSpec框架,该框架通过整合会话级、工作区级和全局级三类语料库,并保留会话轨迹、以代理输出格式索引已打开文件,确保检索内容与代理行为高度匹配;同时,采用离线预训练的上限约束与在线反馈驱动的自适应机制,动态调整各代理的草案长度,实现更精准的资源分配。实验表明,AgSpec在两个面向代码仓库的多代理编码基准上显著优于五种现有检索型推测解码器及EAGLE-3,在批量大小为1时生成吞吐量提升最高达4.37倍,批量大小为16时达4.76倍,且其性能优势在无仓库或单代理场景中依然有效,表明该方法具有广泛的适用性。
链接: https://arxiv.org/abs/2610.01108
作者: Sumin Lee,Sukmin Cho,Suengjae Lim,Youngjin Kwon
机构: KAIST(韩国科学技术院); School of Computing, KAIST(计算学院,韩国科学技术院)
类目: Computation and Language (cs.CL)
备注:
Abstract:Retrieval-based speculative decoding (SD) drafts tokens by copying continuations from existing text, which suits coding agents that repeatedly reproduce code, logs, and earlier attempts. Yet existing methods fall short in agent pipelines: much of the reusable text is missing from their corpora or stored in a form that differs from what the agent emits, and their draft lengths ignore that accept length varies across agents and drifts over turns. We present AgSpec, a framework that supplies the corpus and draft-length policies that existing retrieval engines lack in coding-agent pipelines. AgSpec retrieves from session, workspace, and global corpora, retaining the ongoing session trajectory and indexing opened files in the agent’s emission format. It bounds each agent’s draft length with an offline-profiled cap and adapts the length online from verification feedback. On two repository-level multi-agent coding benchmarks, AgSpec outperforms five retrieval-based drafters and EAGLE-3 in most evaluated settings, raising generation throughput over autoregressive decoding up to 4.37 \times at batch size 1 and 4.76 \times at batch size 16. AgSpec also remains effective on benchmarks without a repository or a multi-agent pipeline, showing that its gains generalize to coding agents broadly.
[NLP-76] Precision over Scale: A Polish-Silesian Benchmark and a Translation System Outperforming Open-Source and Commercial Models EMNLP2026
【速读】: 该论文旨在解决方言机器翻译(Dialectal Machine Translation, DMT)中存在的核心挑战,即由于数据稀缺以及标准评估基准无法充分捕捉方言间强烈语言变异所导致的性能瓶颈。现有基准通常基于标准化、经过编辑的文本,难以反映真实口语化、非规范化的方言表达。为此,研究者采用神经与规则相结合的方法,在新构建的波兰语-西里西亚语测试集SiLTT上进行评估,并对比了已有的BOUQuET和FLORES基准。实验结果表明,规则系统在SiLTT和BOUQuET数据集上表现最为稳定且优越;尽管对TranslateGemma模型进行高质量数据微调后,其性能超越了传统神经基线,但在处理方言场景时仍未能超越规则系统。因此,该研究的关键解决方案在于:通过构建专门针对方言的评估数据集SiLTT,结合规则系统对语言变体的高度可解释性与鲁棒性,有效提升了方言翻译的准确性与可靠性。研究同时公开了SiLTT数据集及最优神经模型,以推动后续方言翻译研究的发展。
链接: https://arxiv.org/abs/2610.01082
作者: Grzegorz Kulik,Mikołaj Pokrywka,Adam Jatowt,Wojciech Kusa
机构: NASK National Research Institute(波兰国家研究所); University of Innsbruck(因斯布鲁克大学)
类目: Computation and Language (cs.CL)
备注: Accepted at EMNLP 2026 Findings
Abstract:Dialectal machine translation remains challenging due to limited data and strong linguistic variation not captured by standard benchmarks, which often assume standardized and well-edited text. We study Polish-Silesian MT using neural and rule-based systems, evaluating on SiLTT - a new Pol-Szl testset, alongside established BOUQuET and FLORES benchmarks. Results show our rule-based system is consistently strongest on SiLTT and BOUQuET datasets and that TranslateGemma fine-tuned on a curated dataset improves over strong neural baselines but does not surpass the rule-based system in dialectal settings. We release SiLTT and our best neural model to support further research.
[NLP-77] Probe with Participation Trophies: Random-Reward RL as a Probe of LLM Capability
【速读】: 该论文旨在解决长期存在的争议问题:探针(probing)性能究竟揭示了大语言模型(Large Language Models, LLMs)的内在能力,还是仅反映了训练过程中潜在的标签泄露(label leakage)现象。针对这一核心问题,论文提出将虚假奖励悖论(spurious-reward paradox)与模型的可到达性(reachability)相联系,并引入随机奖励强化学习(random-reward reinforcement learning, RL)作为探针分析的有效工具。其解决方案的关键在于:通过施加无信息量的随机奖励信号,剥离正确性反馈的影响,从而揭示在特定约束条件下,模型当前状态经进一步训练所能达到的性能上限,即“可到达性”。实验表明,即使在相同初始准确率(如合成算术任务中均为3.5%)的两个OLMo检查点,在相同正确性奖励驱动的RL下,其最佳表现分别提升至8.5%和55%,体现出显著差异。此外,研究发现训练响应存在三个不同阶段:早期预训练阶段,正确奖励效果有限;中期预训练阶段,正确奖励开始有效但随机奖励仍弱;进入中段训练后,随机奖励亦能引发显著性能提升。该模式在基于掩码的监督微调(SFT)分析中同样显现,表明其非特定于某一强化学习机制。因此,随机奖励RL提供了一种独特视角,通过消除正确性信号干扰,有效区分模型的真实潜力与任务记忆,从而应对基于解码可识别性(decodability-based probing)探针中的核心难题——即成功探针反映的是模型能力还是对任务本身的习得。
链接: https://arxiv.org/abs/2610.01066
作者: Yu Mao,Lei Yu,Zining Zhu,Yusheng Zheng,Haohang Li,Freda Shi,Yutong Yin,Zhaoran Wang,Jingcheng Niu
机构: 未知
类目: Computation and Language (cs.CL)
备注:
Abstract:We connect the spurious-reward paradox to a model’s reachability and propose random-reward reinforcement learning (RL) as a useful tool for the probing enterprise, addressing a decade-long debate over what probing performance actually reveals about a model. There are two prevailing explanations for the surprising finding that even random rewards can improve the performance of large language models (LLMs): one attributes the gains to particular mechanisms within RL training; the other to data contamination. Our results motivate a different view: spurious-reward RL can probe a model’s reachability, or what further training can attain from its current state under specified constraints, beyond what is reflected in its current performance. Two OLMo checkpoints with the same accuracy on synthetic arithmetic (3.5%), for example, reach 8.5% and 55% in their best runs under the same correctness-rewarded RL. Examining OLMo checkpoints across pre-training and mid-training reveals three distinct regimes of training response: early on, RL produces little improvement even when correct answers are rewarded; later in pre-training, rewarding correct answers becomes effective while random rewards remain weak; and, upon entering mid-training, even random rewards can produce large gains. A similar ordering appears in a number-masked supervised fine-tuning (SFT) analysis of these checkpoints, suggesting that the pattern is not specific to a particular RL mechanism. Moreover, RL with random rewards offers a distinctive perspective on what training can make an LLM do, since its reward signal supplies no information about which answers are correct. By asking what training can attain without correctness feedback, it addresses the label-leakage side of a central problem in decodability-based probing: whether a successful probe reveals the model’s capabilities or learns the task itself.
[NLP-78] Capturing In-Context Learning Dynamics with Task Operators NEURIPS2026
【速读】: 该论文旨在解决生成式语言模型在上下文学习(In-context Learning, ICL)中存在推理效率低下及机制理解不充分的问题。具体而言,现有方法在每次推理时需完整处理全部示范样本,导致计算开销大;同时,对ICL内部工作机制的理解仍不深入。为应对上述挑战,本文提出关键解决方案——任务算子(Task Operator, TO),其核心思想基于对ICL前向传播过程的分析:发现每个注意力头的输出可视为其上下文掩码版本的仿射变换,且该变换的参数在特定任务下对不同输入样本具有统计稳定性。据此,TO通过解析推导的方式将这一变换作为注意力输出投影的更新项进行重放,从而实现无需重复处理完整示范集的高效推理。实验表明,TO在词汇、算法和推理类任务上均优于已有方法,并显著缩小了零样本推理与完整ICL之间的性能差距。此外,研究进一步揭示任务知识集中于跨层与位置的任务特异性稀疏电路,且通过平均来自不相交示范批次的算子可实现多示例扩展而无需增大上下文窗口,为高效、可扩展的ICL部署提供了新范式。
链接: https://arxiv.org/abs/2610.01054
作者: Guangzhi Xiong,Zhenghao He,Bohan Liu,Sanchit Sinha,Wenqian Ye,Aidong Zhang
机构: University of Virginia(弗吉尼亚大学)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: NeurIPS 2026
Abstract:In-context learning (ICL) enables language models to perform new tasks from demonstrations without weight updates. However, every ICL inference requires processing the full set of examples, resulting in inefficient deployments, and how ICL works mechanistically is not fully understood. Prior work compresses ICL into fixed activation vectors extracted from specific layers or positions, but these input-independent interventions fail on complex tasks where the output depends on fine-grained interactions with the input. By analyzing the ICL forward pass, we show that each attention head’s output is an affine transformation of its context-masked counterpart, and that the parameters of this transformation are empirically stable across samples for a given task. Building on this, we introduce Task Operator (TO), which replays this transformation as an analytically derived update to the attention output projection. Across lexical, algorithmic, and reasoning tasks, TO achieves the best overall performance among prior methods and substantially narrows the gap between zero-shot inference and ICL. We further show that the extracted knowledge concentrates in a task-specific sparse circuit across layers and positions, and that averaging operators from disjoint demonstration batches enables effective many-shot scaling without expanding the context window. Our code is available at this https URL.
[NLP-79] Sentence Specificity Scores for Collaborative Technical Documentation: A Domain-Transfer Study
【速读】: 该论文旨在解决生成式AI(Generative AI)在技术文档生成过程中,如何有效评估与选择高质量文本修订版本的问题。核心挑战在于,现有基于特定性(specificity)的评分模型在不同上下文和数据集上表现出不一致的排序效果,导致其在实际应用中难以可靠指导决策。解决方案的关键在于揭示并验证:特定性评分模型的性能及其对文本选择的实际价值,高度依赖于所使用的预测器类型(如通用领域模型SpeciTeller与目标适配型模型)以及候选文本集合的构成。研究发现,即使在同一语料库内,不同模型产生的句子排序差异显著,且严格过滤与词元长度调整无法统一这些差异;然而,在部分数据集(如Gemma)中,使用SpeciTeller评分可显著提升方向有效选择率(从71.7%提升至83.3%),表明特定性评分在特定条件下具备实际决策价值。这一结果强调了评分解释必须结合具体预测器与候选集特征,而非孤立看待分数本身。
链接: https://arxiv.org/abs/2610.01046
作者: Rocker D’Antonio,Thomas Benton Townsend,Dimitrios Michael Manias
机构: Mississippi State University (密西西比州立大学)
类目: Computation and Language (cs.CL); Machine Learning (cs.LG)
备注: 17 pages, 2 figures. Accepted for publication in the 2026 IEEE 12th International Conference on Collaboration and Internet Computing (CIC)
Abstract:Collaboration depends on shared context, and technical documentation is one way that context persists across people and AI teammates. Specificity, the amount and exactness of detail expressed in language, shapes what information documentation captures and how precisely that information is communicated. This work audits sentence-specificity scoring artifacts on technical documentation and tests whether scores applied only after generation help choose among fixed LLM-generated revisions. Across Wikipedia and three technical-documentation corpora, the fixed general-domain predictor SpeciTeller and the pinned post-publication author-repository implementation of Ko et al.'s target-adapted predictor produce different corpus orders and same-sentence rank agreement from -0.066 to 0.510. Strict filtering and token-length adjustment change these patterns without reconciling them. In the Gemma set, SpeciTeller ranking raises direction-valid selection from 71.7% to 83.3% (+11.7 points; 95% source-case bootstrap interval +1.7 to +21.7); in the GPT-OSS-120B set, SpeciTeller ranking raises direction-valid selection from 51.7% to 56.7% (+5.0 points; 95% source-case bootstrap interval -6.7 to +16.7), and every primary single-score GPT-OSS-120B interval includes zero. These findings tie score interpretation and decision value to the predictor and candidate set.
[NLP-80] Empty Commitments: When Agents Promise What Their Runtime Cannot Deliver
【速读】: 该论文旨在解决对话系统中“空承诺”(empty commitment)问题,即智能体在未具备执行能力的情况下作出的未来行动承诺,例如聊天机器人声称“我明天会提醒你”,但因缺乏持续运行机制或工具支持而无法兑现。此类承诺并非因轨迹失败导致,而是由智能体配置本身决定,因而其“空性”具有结构性本质。解决方案的关键在于构建一套基于承诺语义(commitment semantics)的形式化框架,明确三类失败类型、承诺可实现的锚定条件(anchoring condition),以及响应层面的结果分类体系,并提出一种渐进式测量协议:通过五种不同设置依次引入单一持久性功能(persistence affordance),结合环境信息显式或隐式两种情形,系统评估智能体在实际交互中对承诺的兑现能力。该方法为量化评估和改进对话系统的可信承诺行为提供了可操作的技术路径。
链接: https://arxiv.org/abs/2610.01045
作者: Jiaqi Tang,Lan Wei,Bingyu Shen,Boyang Li
机构: Kean University(新泽西州立大学); Independent Researcher(独立研究员)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: 4 pages, 3 tables
Abstract:A chatbot that says “I will remind you tomorrow” will not run again until the user writes. We call such a promise an empty commitment: a promise of an action after the current turn that nothing in the agent’s tools or runtime can carry out. Unlike a broken promise, its emptiness follows from the agent’s configuration alone; no later trajectory is needed. We define empty commitments on top of commitment semantics, with three failure types, an anchoring condition for promises that a tool could make real, and a response-level outcome taxonomy. We then describe a measurement protocol: follow-up requests run in five setups that add one persistence affordance at a time, with the environment either left implicit or stated.
[NLP-81] Beyond Final Accuracy: Auditing Communication in LLM Multi-Agent Systems
【速读】: 该论文旨在解决多智能体通信(multi-agent communication)中性能提升的归因难题:即系统性能的改善究竟是源于有效通信、优越的智能体架构,还是仅仅由于额外的推理过程。传统评估方法常将通信机制嵌入其设计系统中,导致上述因素难以分离,且最终准确率指标将修正错误与错误答案混为一谈,掩盖了通信对决策的实际影响。为此,论文提出独立-通信-修订(Independent–Communicate–Revise, ICR)这一受控评估框架,通过固定初始推理轨迹,量化在双方初始判断正确性条件下的修正与保留行为,并利用“无消息修订”对照组来分离通信带来的增益与单纯推理的贡献。在四个推理基准测试中,对文本和潜在通信内容的审计表明,相似的总体准确率可能掩盖显著不同的修订模式;相较于仅传递答案,完整推理虽提升了修正能力,但也增加了错误保留,说明更丰富的信息传递同时放大了有益与有害的影响。在MedQA和GPQA-D上的接收策略对比进一步显示,结构化验证策略可普遍提高保留率并降低修正率,但其选择性效应随通道和任务而异。这些发现挑战了将通信质量视为信道固有属性的传统观点,强调应将评估重心转向选择性修订,并建立一个统一框架以分析消息内容与接收策略如何共同作用于收益与损害。
链接: https://arxiv.org/abs/2610.01042
作者: Shixuan Li,Wei Yang,Peiyu Zhang,Anzhe Cheng,Heng Ping,Paul Bogdan
机构: University of Southern California(南加州大学); Ming Hsieh Department of Electrical and Computer Engineering(明赫电气与计算机工程系)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注:
Abstract:Multi-agent communication aims to help agents benefit from one another’s information. Yet improvements in system performance leave a fundamental ambiguity: do they reflect effective communication, a favorable agent architecture, or simply additional reasoning? Because communication methods are commonly evaluated within the systems they were designed for, these factors are difficult to disentangle. Final accuracy further merges corrected errors and corrupted answers into a single outcome, obscuring how communication changes decisions. We introduce Independent–Communicate–Revise (ICR), a controlled framework that evaluates communication as answer revision following independent reasoning. ICR fixes initial reasoning trajectories, measures correction and preservation conditional on both agents’ initial correctness, and uses a no-message revision control to quantify gains beyond additional reasoning. Across four reasoning benchmarks, our audit of textual and latent communication reveals that similar aggregate accuracy can conceal substantially different revision behaviors. Compared with transmitting answers alone, full reasoning increases correction while reducing preservation on all four benchmarks, so richer messages amplify beneficial and harmful influence alike. Receiver-policy comparisons on MedQA and GPQA-D further show that a structured verification policy shifts every channel toward greater preservation and lower correction, while its effect on selectivity varies across channels and tasks. These findings challenge treating communication quality as an intrinsic property of a channel. ICR therefore recenters evaluation on selective revision, providing a unified framework for examining how message content and receiver policies jointly produce benefits and harms.
[NLP-82] LawCompass: Navigating from Legal QA to Multi-Agent Deep Research with Grounded Evidence
【速读】: 该论文旨在解决现有法律智能助手在处理复杂法律任务时的局限性问题,即多数系统仅支持多轮对话式问答(multi-turn conversational QA),难以实现系统性证据检索、多步推理与报告级综合分析。其核心解决方案是提出LawCompass,一个基于证据的法律助理系统,通过引入三个面向任务的功能模块——法律问答(Legal QA)、专业检索(Professional Retrieval)与深度研究(Deep Research),实现了从标准法律问答向多智能体深度研究范式的跃迁。关键创新在于采用多智能体工作流对复杂法律任务进行分解,并在所有模块中保持显式的引用链接,确保输出结果可追溯至原始法律文献,从而提升系统的可信度与可验证性。评估结果表明,LawCompass为将对话式AI转化为可信赖、基于证据的法律研究辅助提供了一种实用且可扩展的新范式。
链接: https://arxiv.org/abs/2610.01027
作者: Xiaoxia Cheng,Linnan Wang,Jiahao Ma,Zhichuan Ye,Xuemei Zhou,Chuanyu Tong,Bo Jiang,Qing Zhu
机构: Anhui University (安徽大学); Tsinghua University (清华大学)
类目: Computation and Language (cs.CL)
备注:
Abstract:Recent advances in Large Language Models (LLMs) and Retrieval-Augmented Generation (RAG) have significantly democratized access to legal information. Nevertheless, most existing legal assistants remain confined to multi-turn conversational QA, failing to support complex legal tasks that require systematic evidence retrieval, multi-step reasoning, and report-level synthesis. In this paper, we present LawCompass, an evidence-grounded legal assistant that navigates the transition from standard Legal QA to multi-agent deep research. LawCompass provides three task-oriented functions: Legal QA, which delivers precise, evidence-backed answers to legal questions; Professional Retrieval, which enables structured exploration of statutes and judicial cases via query rewriting; and Deep Research, which employs a multi-agent workflow to decompose complex legal tasks and synthesize comprehensive research reports. Crucially, LawCompass maintains explicit citation links across all modules, empowering users to directly verify system outputs against original legal sources. Evaluation results demonstrate that LawCompass provides a practical and scalable paradigm for transforming conversational AI into trustworthy and evidence-grounded legal research assistance.
[NLP-83] It Takes Workflows to Evolve Better Workflows
【速读】: 该论文旨在解决多智能体工作流(multi-agent workflows)在执行复杂现实任务时,因各智能体能力耦合且优化目标单一而导致的性能瓶颈问题。现有方法仅优化工作流生成器,而固定其他执行或构建智能体,忽视了全流程中各角色间的协同依赖关系;同时,由于工作流结果为稀疏评分,难以定位失败根源,导致难以实现端到端的联合优化。其解决方案的关键在于提出FloWright框架,通过引入分层、结构感知的奖励机制,使单个角色可自主进化,多个角色可协同进化,无需额外模型、标签或执行开销。此外,为克服现有数据集难以体现多智能体协作挑战的问题,进一步提出DataWright,一种自适应数据强化方法,将常规数据转换为更具难度的工作流级任务。实验表明,在文档、幻灯片、图表、代码、数学及金融等多类任务上,采用FloWright训练的小型开源模型性能提升最高达+7.41%,其中协同进化多个角色相比仅优化单一角色分别获得+5.03%和+2.83%的增益,显著提升了多智能体系统的整体效能。
链接: https://arxiv.org/abs/2610.01026
作者: Xuehang Guo,Haoyu Wang,Haifeng Chen,Yangyi Chen,Zhenhailong Wang,Qingyun Wang
机构: William Mary(威廉与玛丽学院); NEC Corporation of America(美国日电公司); University of Illinois Urbana-Champaign(伊利诺伊大学厄本那-香槟分校)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注:
Abstract:Tackling complex real-world tasks can exceed the capabilities of a single large language model (LLM), motivating the use of multi-agent workflows that coordinate specialized agents to work together on these tasks. Recent methods train LLMs to construct better workflows from execution outcomes, but they optimize only the workflow generator, while the other agents that build or execute each workflow remain fixed even though every outcome depends on all of them. However, extending training beyond the generator is challenging: the agents are coupled, and a workflow’s outcome is a single sparse score that cannot tell which agent causes a failure. We propose FloWright, which leverages the workflow as a harness to optimize workflows. By introducing a hierarchical, structure-aware reward paradigm, FloWright enables one role to self-evolve and two or more roles to co-evolve, with no additional models, labels, or executions. Considering the limitation that workflows are commonly trained and evaluated on data that a single agent can already handle, we further propose DataWright, an adaptive data hardening approach that converts existing datasets into workflow-level tasks with increased difficulty. Across document, slide, chart, code, math, and finance tasks, small open models trained with FloWright achieve improved performance by up to +7.41% , with co-evolving ( +5.03% ) more roles gaining more than optimizing one of them alone ( +2.83% ). Our project page: this https URL.
[NLP-84] Groundability Not Scale Alone: When Weak Reviewers Can Audit Strong Coding Agents
【速读】: 该论文旨在解决生成式代码补丁(code patch)在实际应用中存在遗漏必要行为却仍被误判为正确的难题,尤其在长执行轨迹与自信摘要的掩盖下,人工或自动化审查难以有效识别此类缺陷。其核心挑战在于:如何判断一个能力较弱的审查者(reviewer)在缺乏完整执行证据的情况下,仍能可靠地判定补丁是否真正解决了问题。解决方案的关键在于引入“官方执行证据”(official execution evidence)作为诊断上限基准,并在此基础上构建冻结的审查级联(frozen cascade)机制——该机制结合补丁引发的静态错误检测与在未修复仓库上首次失败的生成测试,以提升审查的覆盖率、缺陷捕获率并降低误拒率。实验表明,在121条GPT-5.4和59条Gemini的独立测试用例中,该方法实现了0.89和0.86的覆盖率、0.76和0.80的缺陷捕获率,尽管仍存在约66%-67%的误拒率,主要源于未解决案例进入审查阶段。研究进一步指出,审查者规模并非质量的稳定预测因子,而无须依赖官方测试即可生成同样可靠的验证检查,仍是当前的主要瓶颈。
链接: https://arxiv.org/abs/2610.01023
作者: Junyu Guo,Shangding Gu,Ming Jin,Javad Lavaei
机构: 未知
类目: oftware Engineering (cs.SE); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注:
Abstract:Coding agents can return plausible patches that omit required behavior. These failures are hard to review because long traces and confident summaries often hide what was missed. We ask when a nominally weaker reviewer can reliably decide whether a patch solves its issue. We study 411 execution-labeled traces from three agents and 101 controlled cases. On 154 GPT-5.4 traces, structured but unchecked evidence raises both defect catch and over-rejection. We then provide official execution evidence as an upper-bound diagnostic. After choosing and freezing one of two formats per reviewer, five of six reviewers improve both rates on 122 held-out traces; two classify every trace correctly. Reviewer size is not a consistent predictor of quality. Because official tests are unavailable in deployment, we also evaluate a frozen cascade with patch-caused static errors and generated tests that first fail on the unpatched repository. On 121 scored held-out GPT-5.4 traces and 59 Gemini traces, its coverage is 0.89 and 0.86, risk is 0.33 and 0.26, catch is 0.76 and 0.80, and over-rejection is 0.66 and 0.67. Most false rejections occur when unresolved cases reach the reviewer. Official execution evidence shows the potential of weak review when decisive checks are available. Producing equally reliable checks without official tests remains the main bottleneck.
[NLP-85] Pay for the Fault Not the Flow: Label-Free In-Flow Multi-Agent Workflow Optimization
【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在构建多智能体工作流时面临的两大核心挑战:一是任务分解粒度与智能体分配的动态决策难题,二是工作流运行过程中故障定位与修复成本高昂的问题。现有方法通常依赖预设模板进行任务分解和智能体分配,缺乏对智能体能力与子任务需求匹配度及运行开销的实时权衡,导致错误难以在执行前识别,而故障修复需依赖参考答案或训练评估器,并通过全量重执行、重搜索或重训练实现,代价巨大。为此,论文提出InFlowOp,其关键创新在于引入一种无标签的统一成本度量机制,该机制将每个决策的成本定义为智能体胜任力与子任务需求匹配程度与其运行开销的加权比值。该成本函数不仅用于在工作流构建阶段双向优化任务分解粒度与智能体分配策略,还在运行阶段指导以最小代价修正故障,实现“边建边调”的自适应优化。为验证其有效性,研究进一步构建了Braid基准测试集,涵盖需多智能体协同的任务场景,超越单智能体基线最高达+11.97%性能提升,并在流式优化下实现+9.64%的显著增益,充分证明了InFlowOp在降低人工干预、提升鲁棒性与效率方面的优势。
链接: https://arxiv.org/abs/2610.01017
作者: Xuehang Guo,Haoyu Wang,Shengyu Chen,Zach Chen,Wei Cheng,Qingyun Wang,Haifeng Chen
机构: William Mary(威廉与玛丽学院); NEC Corporation of America(美国日电公司)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注:
Abstract:Large language models (LLMs) increasingly construct multi-agent workflows that decompose a complex task and assign specialist agents from a pool. However, building such a workflow well remains challenging: how finely to divide the task, which agent to trust with each subtask, and when to create a new specialist are all critical decisions a workflow constructor needs to settle up front. Thus, whether each subtask succeeds remains unknown until the workflow runs. Yet, improving a workflow is costly. Locating a fault usually requires a reference answer, a graded outcome, or a trained assessor, and the fix is applied to the whole workflow through re-execution, re-search, or retraining. We propose InFlowOp, which prices every decision in one label-free cost that weighs how well an agent’s competence meets what a subtask demands against how much that agent takes to run. Before execution, InFlowOp bidirectionally determines the granularity of task decomposition and agent assignment following from the cost rather than from a fixed template. During execution, InFlowOp corrects a fault with the cheapest move via the same cost that serves the workflow both as it is built and as it runs. Facing the workflow-level evaluation challenge, we introduce Braid, a benchmark whose tasks require multi-agent coordination beyond single-agent capability. Across various domains and backbones, InFlowOp outperforms single agent baselines by up to +11.97% , achieving +9.64% with in-flow optimization. Our project page: this https URL.
[NLP-86] Scaling and Distilling Text Embeddings for Better Diffusibility
【速读】: 该论文旨在解决连续扩散语言模型(continuous diffusion language models, DLMs)在生成过程中因潜在空间(latent space)可扩散性不足而导致的生成失败问题,尤其是当原始嵌入表示过于判别性时,导致扩散过程难以收敛至合理词义替代点,最终生成无效嵌入。其核心解决方案在于通过知识蒸馏(knowledge distillation)将强大的教师模型T5Gemma-2的解码概率作为软标签(soft labels),训练一个轻量级学生编码器,使其学习到更具连通性的嵌入表示:该方法在保持原有编码-解码机制的同时,使语义相近的替代词嵌入在潜空间中更紧密聚集,从而构建出更平滑、更易扩散的潜在空间。实验表明,该蒸馏后的嵌入显著提升了生成性能,在OpenWebText数据集上实现了17.8的生成困惑度(Gen. PPL),优于GPT-2-M,接近真实文本的15.4,验证了优化潜空间结构对连续扩散生成质量的关键作用。
链接: https://arxiv.org/abs/2610.01016
作者: Zekai Zhang,Yunjie Tian,Yanjin He,Xiaoyan Zhang,Dongdi Zhao,Qing Qu,Di Fu
机构: University of Michigan(密歇根大学)
类目: Computation and Language (cs.CL)
备注: 28 pages, 12 figures. Code is available at this https URL
Abstract:Diffusion language models (DLMs) offer a promising alternative to autoregressive (AR) language generation. Recent advances in continuous DLMs, which apply latent diffusion to continuous text embeddings, raise a practical question: which embedding makes the best latent space, i.e., the most diffusible? To answer this, we search through different embeddings and find that scaling the embedding model to stronger ones within the same family (T5 to T5Gemma-1 to T5Gemma-2) greatly improves generative performance. But the raw T5Gemma-2 embeddings are still not optimal. They are so discriminative that even the embeddings of plausible alternative words are separated, which makes the generation vulnerable to imperfect sampling. Consequently, continuous diffusion often fails to reach any of them and ends up at an invalid embedding instead. To address this, we distill T5Gemma-2 into a student encoder that learns the teacher’s decoded probabilities as soft labels. Learning from such soft labels makes the student pull the alternative embeddings closer while maintaining the encoding-decoding mechanism. The distilled embeddings form a more connected and diffusible latent space, improving over the vanilla T5Gemma-2 embeddings. As a result, our medium-sized DLM achieves Gen. PPL 17.8 (against real-text PPL 15.4) at real-text entropy on OpenWebText, outperforming GPT-2-M on Gen. PPL.
[NLP-87] Distilling Directional Verification
【速读】: 该论文旨在解决知识蒸馏(Knowledge Distillation)中因教师模型单向生成能力缺陷导致的错误传递问题。具体而言,当教师模型仅能从特定方向(如“父母→子女”)生成答案,而无法反向生成(如“子女→父母”)时,若直接以教师生成的答案作为学生模型的训练标签,会将这种方向性局限传播给学生,从而降低蒸馏效果。其解决方案的关键在于提出“方向性标签蒸馏”(directional label distillation),即冻结教师模型后,利用其在已知方向上对候选答案进行评分,选择得分最高的候选答案作为学生模型的训练目标。实验表明,即使经过名称先验校正,基于已知方向的评分仍比请求方向的评分更准确;且在挖掘出的以父母为显著实体的事实中,最优方向可自动反转。当保留用于训练的子事实时,基于已知方向标签训练的学生模型在开放问答任务中的准确率较基于修正后反向标签的学生模型提升13至15个百分点。此外,在通过词汇相似性匹配生成答案与固定名称列表后,学生几乎完全复现所选标签,且性能高度依赖于标签质量。该方法在未筛选查询及无正确答案插入的候选检索场景下仍保持优势,证明方向性验证有效缓解了教师生成答案中的错误向学生模型的转移。
链接: https://arxiv.org/abs/2610.00997
作者: Jungseob Lee,Sugyeong Eo,Seongtae Hong,Seungyoon Lee,Chanjun Park,Jaehyung Seo,Heuiseok Lim
机构: Korea University (高丽大学); Yonsei University Mirae Campus (延世大学未来校区); Soongsil University (松林大学); Konkuk University (中央大学)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: 29 pages, 7 figures, 31 tables
Abstract:Knowledge distillation aims to transfer the factual knowledge of large language models to smaller models for efficient deployment. Yet a teacher may recall a relation in one direction while failing to generate the answer in the reverse direction. Distillation from its generated answers can therefore propagate this directional limitation to the student. The same teacher can nevertheless recognize such an answer by scoring the relation in the direction it knows. We introduce directional label distillation, in which frozen teachers score candidate answers in that known direction and the best-scoring candidate becomes the student’s training target. On facts about parents and their children, known-direction scoring yields more accurate labels than scoring the requested direction, even after tuned corrections for name priors. With prior-corrected scores, the better direction depends on the facts rather than the template, and reverses on mined facts whose notable entity is the parent rather than the child. With the evaluated children’s forward facts withheld, students trained on known-direction labels improve open-ended accuracy on their trained queries by 13 to 15 points over students trained on prior-corrected reverse labels. After generated answers are matched to a fixed name list by lexical similarity, students reproduce nearly all selected labels. Their accuracy largely follows label quality. The label advantage holds on unscreened queries and when candidates are retrieved without inserting correct answers. Our findings show that directional verification mitigates the transfer of errors from teacher-generated answers to students by providing more accurate training targets. Code is available at this https URL.
[NLP-88] he Devil Is in the Reconstruction Loss Scale: Rethinking Optimization in LLM Quantization
【速读】: 该论文旨在解决生成式 AI(Generative AI)中大语言模型(LLM)后训练量化(Post-training Quantization, PTQ)方法中存在的优化不平衡问题。现有主流学习型PTQ方法采用顺序量化策略,将预训练模型按模块(如Transformer块)逐阶段量化,并通过梯度下降优化辅助量化参数(如缩放因子、旋转矩阵、截断阈值等),通常以均方误差(MSE)作为重建损失函数。然而,该方法在不同量化阶段间存在显著的损失幅度差异,导致梯度大小和参数更新严重不均衡,进而引发优化强度分布失衡。研究发现,这种现象源于MSE对重建损失尺度的敏感性,使得跨阶段损失尺度被放大并转化为非均匀梯度分布。为此,论文提出一个普适性改进原则:应将各阶段的优化强度与重建损失尺度解耦。理论分析表明,基于样本、通道、标记或元素层级定义的均方根误差(RMSE)变体可通过隐式梯度归一化自然实现该原则,在无需修改网络结构或训练流程的前提下,作为MSE的即插即用替代方案,显著提升量化性能。
链接: https://arxiv.org/abs/2610.00983
作者: Chao Li,Shigeng Wang,Anbang Yao
机构: Intel Labs China(英特尔实验室中国)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: Project page: this https URL
Abstract:Post-training quantization (PTQ) methods typically use sequential quantization that partitions a pre-trained LLM into a series of units (e.g., transformer blocks), with one unit quantized at each stage. State-of-the-art PTQ methods are predominantly learning-based, optimizing auxiliary quantization parameters (e.g., scaling factors, rotation matrices, clipping thresholds, and adapters) via gradient descent to minimize a reconstruction loss. A common practice is to use mean squared error (MSE) as the reconstruction loss function, yet its induced optimization behavior remains largely unexplored. In this work, we take a holistic view of sequential quantization and systematically investigate how optimization evolves from the first quantization stage to the last, aiming for a deep understanding of optimization in learning-based PTQ schemes. Through extensive empirical studies spanning representative learning-based PTQ methods, LLM families, model scales, architectures, quantization settings and various tasks, we consistently uncover Optimization Imbalance: reconstruction loss magnitudes vary dramatically across stages, accompanied by highly uneven gradient magnitudes and parameter updates under MSE. We term the cross-stage range of loss magnitudes the reconstruction loss scale, and reveal that MSE translates the unexpectedly large reconstruction loss scale into highly uneven gradient magnitudes, which in turn lead to uneven optimization strength across quantization stages. This finding suggests a general principle for improving learning-based PTQ: optimization strength across stages should be decoupled from the reconstruction loss scale. Theoretically, we show that root mean squared error (RMSE) variants defined at the sample, channel, token, and element levels naturally realize this principle through implicit gradient normalization, outperforming MSE significantly as a drop-in replacement.
[NLP-89] A Citation-Grounded Benchmark for Trustworthy Earnings Call Transcript Analysis with Large Language Models
【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在金融文档分析,特别是财报电话会议记录(Earnings Call Transcripts, ECTs)分析中,生成具有可验证依据的分析结论所面临的挑战。核心问题包括:评估分析结论的可信度通常依赖于昂贵且难以扩展的人工专家标注;真实场景下的分析任务涉及长上下文-问题-答案三元组,显著增加了任务复杂性;同时,模型在面对证据不足时仍可能生成无根据的幻觉内容,即“有意识的无知”(conscious incompetence)这一关键失败模式。为应对上述问题,论文提出了一种无需依赖专家标注的数值证据评估方法,以实现对生成内容“可溯源性”的自动化评估;并构建了基于标普500前100成分股的自动化数据集ECTs-100,用于系统性地评估模型在“可溯源性”与“正确性”两方面的表现。实验结果表明,尽管当前LLMs在可溯源性方面表现良好,但在正确性上仍存在明显局限,而信息不足进一步加剧了模型误判的风险。
链接: https://arxiv.org/abs/2610.00969
作者: Yingzhu Zhao,Vlad Pandelea,Han Yuan,Bo Hu,Wuqiong Luo,Li Zhang,Zheng Ma
机构: American Express(美国运通); American Express(美国运通)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注:
Abstract:Large language models (LLMs) have been increasingly used for financial document analysis, including earnings call transcripts (ECTs). Beyond generating standalone claims, users increasingly prefer grounded analyses that pair claims with verifiable citations from source documents to enable independent validation. However, evaluating such analytical claims typically requires extensive expert annotation, which is costly and difficult to scale, and real-world financial analysis commonly involves long context-question-answer triplets, further increasing task complexity. To address these challenges and benchmark the current landscape of grounded analysis by LLMs, we propose a numeric evidence evaluation method that enables groundedness assessment without reliance on expert annotation. We also introduce an automated dataset construction pipeline and construct ECTs-100 from the top 100 constituents of the SP 500 to support benchmark of both groundedness and correctness. In addition, we examine conscious incompetence, a practical failure mode in financial analysis in which LLMs must detect when available evidence is insufficient and refrain from producing unsupported hallucinations. Empirical results show that LLMs perform well in groundedness but face notable limitations in correctness, with informational insufficiency presenting an additional challenge.
[NLP-90] Role-aware Heuristic Episodic Attention for Conversational LLM s
【速读】: 该论文旨在解决大语言模型在多轮对话中因上下文累积导致的指令遗忘与信息衰减问题,具体表现为注意力污染(attention pollution)、稀释(dilution)和漂移(drift)三种失效模式。其核心解决方案是提出一种角色感知的启发式情景记忆机制(REA, Role-aware Heuristic Episodic Attention),通过区分指令与对话片段的不同记忆需求,实现差异化的持久性与表征策略:将全局约束类指令存入专用前缀以形成指令记忆(Instructional Memory),而用户输入与模型回复则分别通过压缩与启发式检索机制保存至情景记忆(Episodic Memory)中,动态选择原始文本、压缩表示或省略。实验表明,在Long-MT-Bench+基准上,REA将评分从6.32提升至7.36(10分制),相对基线提升16.5%,同时平均延迟降低2.91倍;在多个参数规模(1.7B–7B)的模型及中英文角色扮演任务中均表现出显著性能增益,验证了角色感知上下文管理在维持对话连贯性与指令遵循性方面的有效性。
链接: https://arxiv.org/abs/2610.00958
作者: Wanyang Hong,Zhaoning Zhang,Yi Chen,Libo Zhang,Baihui Liu,Linbo Qiao,Zhiliang Tian,Dongsheng Li
机构: National University of Defense Technology (国防科技大学); National Key Laboratory of Parallel and Distributed Computing; College of Computer Science and Technology
类目: Computation and Language (cs.CL)
备注:
Abstract:Large language models often lose track of persistent instructions and relevant information as multi-turn conversations grow. We study this cumulative contextual decay through three related failure modes: attention pollution, dilution, and drift. We propose REA (Role-aware Heuristic Episodic Attention), a context-management framework that assigns different persistence and representation policies to instructions and episodic interactions. Instructional Memory retains identified global constraints in a dedicated prefix. Episodic Memory preserves user inputs and compresses model replies, while heuristic retrieval selects raw text, compressed representations, or omission for each historical turn. On Long-MT-Bench+, REA improves the judge score from 6.32 to 7.36 on a 10-point scale, a 16.5% relative gain over the Vanilla baseline, and reduces average latency by 2.91 \times . Additional evaluations show aggregate gains on three backbones spanning 1.7B-7B parameters and on Chinese and English role-playing tasks. These results support role-aware context management as a practical approach to maintaining conversational continuity and instruction adherence.
[NLP-91] Beyond Leaderboards: Tokenomics of Agent ic Small Language Model Ensembles
【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在从独立助手向智能体工作流(agentic workflows)演进过程中,传统基于单一指标(如排行榜准确率)的评估方法已无法全面反映系统实际运行性能的问题。其核心挑战在于如何衡量生成式AI系统在真实应用场景中的操作可靠性、成本效益、延迟表现及令牌效率等多维指标。为此,论文提出以小型语言模型(Small Language Models, SLMs)组成的智能体集成系统(agentic ensemble),并引入由SLM裁判(SLM-judge)驱动的反馈循环机制作为案例研究,实现超越传统排行榜的综合评估。该方案的关键在于通过引入可执行的反馈闭环与动态协调策略,在增加测试时令牌开销和编排复杂度的前提下,显著提升了指令遵循的精确性与鲁棒性:在541个提示的IFEval基准上,最优集成模型达到97.34%的严格提示准确率,较最强的单体模型gpt-5.4提升5.81个百分点,同时处于更低的成本区间。通过对令牌经济性、令牌构成、有效输出吞吐量、反馈环恢复能力、延迟分解以及不同指令类别与约束数量下的表现进行深入分析,研究揭示了智能体式SLM集成可通过合理权衡运行资源与性能收益,实现更高指令遵循保真度,从而为未来智能体式AI系统的多维度评估协议提供了实证基础与设计范式。
链接: https://arxiv.org/abs/2610.00954
作者: Alexei N. Skurikhin,Emily M. Taylor,Nathan A. DeBardeleben
机构: Los Alamos National Laboratory (洛斯阿拉莫斯国家实验室)
类目: Computation and Language (cs.CL)
备注: 8 pages, 9 figures, Presented at ACM CAIS 2026 Workshop RLEval: Methods and Reinforcement Learning Environments for Evaluating AI Agents. Resubmission of permitted appeal, Ticket #MOD-104177
Abstract:As large language models (LLMs) move from standalone assistants into agentic workflows, evaluation must extend beyond scalar leaderboard accuracy to account for operational reliability, cost, latency, and token efficiency. We use an agentic ensemble of small language models (SLMs) with an SLM-judge-mediated feedback loop as a case study for such beyond-leaderboard evaluation. On the 541-prompt IFEval benchmark, the best ensemble achieves 97.34% strict prompt accuracy, exceeding the strongest standalone LLM baseline, gpt-5.4, by 5.81 percentage points while operating in a lower-cost regime. We then analyze the tokenomics and operational behavior behind this gain, including cost per sample, token composition, useful-output goodput, feedback-loop recovery, latency decomposition, and performance across instruction categories and constraint counts. Our results show that agentic SLM ensembles can trade additional test-time tokens and orchestration overhead for improved instruction-following fidelity, motivating multi-dimensional evaluation protocols for future agentic AI systems.
[NLP-92] ReHoPER: Receding-Horizon Planning for Enhanced Reasoning
【速读】: 该论文旨在解决大语言模型在复杂推理任务中缺乏系统性思维链(Chain-of-Thought, CoT)生成能力的问题,尤其在组合性推理(compositional reasoning)场景下表现不足。其核心挑战在于如何在不依赖标注数据或任务特定提示设计的前提下,有效引导模型生成高质量的中间推理步骤以提升最终答案的准确性。为此,论文提出ReHoPER——一种仅用于推理、零样本(zero-shot)的方法,其关键创新在于通过多路径迭代式规划候选中间问题,基于当前上下文历史动态选择并回答一个最相关的问题,随后利用更新后的历史信息重新规划后续步骤。该方法具有任务无关性(task-agnostic),在不同数据集和模型上均采用统一通用指令,无需微调或定制化提示工程。实验结果表明,ReHoPER在包括iLLC在内的多个基准测试中显著优于现有强基线模型,尤其在组合性最强的推理场景中表现突出,验证了其在增强模型递进式推理能力方面的有效性。
链接: https://arxiv.org/abs/2610.00940
作者: Saeed Ahmadnia,Cornelia Caragea
机构: University of Illinois Chicago(伊利诺伊大学芝加哥分校)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:
Abstract:We propose ReHoPER, an inference-only, zero-shot method that improves large language models’ reasoning by generating and answering intermediate questions along multiple paths before the final answer. It iteratively plans a horizon of candidate intermediate questions, selects one to answer, and replans from the updated history. ReHoPER is task-agnostic, using the same generic instructions across datasets and models without labeled data or task-specific prompt design. Across multiple datasets, including iLLC, a new controlled benchmark for compositional reasoning, ReHoPER outperforms strong baselines, with the largest gains in the most compositional settings. Our implementation and the iLLC generator are publicly available to support future work.
[NLP-93] Efficient Task Adaptation in Large Language Models : A Survey of Weight-Based Prompt-Based and Embedding-Based Adaptations AACL
【速读】: 该论文旨在解决大语言模型在多样化下游任务中进行高效任务适配时,不同适配方法之间缺乏统一理解与跨范式比较的问题。当前主流的适配方法(如参数高效微调、上下文学习及嵌入注入)多在各自范式内独立发展,导致其内在关系、性能权衡以及新兴基于嵌入的适配方法的系统性认知不足。为此,本文提出一个统一框架,将任务适配方法依据任务信息的编码位置与方式划分为三类:模型权重、输入提示和注入的任务嵌入(task embeddings)。该框架构建了一个综合分类体系,系统分析了各范式的核心优势与局限性,揭示了不同适配范式的发展脉络,厘清了它们之间的关联性,并指出了未来研究的关键开放问题。其解决方案之关键在于通过“编码位置-编码方式”双维度的统一视角,实现对多元适配方法的整合与对比,为高效任务适配提供了理论基础与方向指引。
链接: https://arxiv.org/abs/2610.00928
作者: Jungwon Park,Changin Choi,Jimyeong Kim,Nojun Kwak,Wonjong Rhee
机构: Daegu Gyeongbuk Institute of Science and Technology; Samsung Advanced Institute of Technology, Samsung Electronics Co., Ltd; AIIS; IPAI; Department of Intelligence and Information, Seoul National University
类目: Computation and Language (cs.CL)
备注: Accepted by AACL-IJCNLP 2026 Main
Abstract:As large language models are increasingly deployed across diverse downstream tasks, efficient task adaptation has emerged as a central challenge. In response, a wide range of task adaptation methods have been proposed, spanning parameter-efficient fine-tuning, in-context learning, and embedding-injection approaches. However, these lines of work have largely evolved within individual paradigms, leaving their cross-paradigm relationships and trade-offs underexplored, especially for recently emerging embedding-based adaptations. This survey presents a unified framework that categorizes task adaptation methods by where and how task information is encoded: model weights, input prompts, or injected task embeddings. We provide a comprehensive taxonomy that integrates these paradigms, analyze their key strengths and limitations to explain how different adaptation paradigms have evolved, clarify relationships across paradigms, and highlight open problems for future research.
[NLP-94] he Geometry of Contextual Relations: Language Models Address Facts by Order of Mention
【速读】: 该论文旨在解决语言模型如何通过上下文中的事实顺序来组织其内部表示,进而实现对问题的推理与回答。具体而言,研究关注的是:在给定一系列陈述性事实(如“爱丽丝吃一个苹果。鲍勃吃一个梨。”)作为上下文时,语言模型(LLM)是如何根据问题中所提及的事实顺序,而非具体名词本身,来定位并激活相应信息的。其解决方案的关键在于提出“序号地址假说”(ordinal addressing hypothesis),即每个事实在上下文中被提及的顺序(order of mention)对应一个固定的“事实地址”(fact address),该地址存在于模型的隐藏状态空间中,并且在不同上下文中共享。当问题涉及某一事实时,模型的状态会沿着由该事实顺序决定的“序号向量”(ordinal vector)被引导至对应的地址;而上下文则提供该地址所指向的具体内容。研究发现,这些事实地址具有四大特性:(1)按提及顺序排列,状态组织依赖于事实顺序而非名称;(2)可被操控(steerable),通过添加序号向量可使模型从第一件事跳转至第二件事;(3)低秩结构(low-rank),且首提事实最易被激活,表现出与人类记忆回溯相似的模式;(4)涌现性(emergent),在深层中间层出现,且在从1.5B到32B参数规模的语言模型中均存在,形成于预训练早期。这一机制揭示了大语言模型在推理过程中对上下文顺序的敏感性,深化了对模型内部语义表征与动态演化的理解。
链接: https://arxiv.org/abs/2610.00910
作者: Yufa Zhou
机构: Duke University (杜克大学)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: Code: this https URL
Abstract:Human reasoning depends on how objects are related within propositions. \textitHow do relations organize the language representations of contextual contents? We give an LLM a list of facts in its context (e.g., \emphAlice eats an apple. Bob eats a pear.) and measure how its hidden state changes when the question switches from what Alice eats to what Bob eats. Averaged over many lists, this change is a steering vector, which we call the \emphordinal vector. It points to a fact by its \emphorder of mention, the order in which the facts were stated in the context. We find that LLMs represent the fact a question asks about by its order of mention, not by the name the question contains. We state this as the \textitordinal addressing hypothesis: each order of mention has a \emphfact address in the model’s state, shared by all contexts, and a question moves the state to the fact address of the fact it asks about, while the context supplies what that fact says. Across Qwen, Gemma, and Llama, fact addresses are (1) \emphordered by mention: query states are organized by the order of facts, not of names, even when one fact has multiple subjects; (2) \emphsteerable: added to a question about the first fact of a new list, the ordinal vector makes the model answer with the second fact of that list; (3) \emphlow-rank: they span a low-rank subspace in which the first-mentioned fact is the easiest to reach, surprisingly similar to human recall; and (4) \emphemergent: they are shared in late-middle layers, hold from 1.5B to 32B parameters, and form early in pretraining. Language models reach a stated fact by where it was mentioned, deepening our understanding of LLM reasoning.
[NLP-95] DeBERTa-ConPara: Attack-Aware and Deployment-Realistic Detection of AI-Generated Text AACL
【速读】: 该论文旨在解决生成式AI(Generative AI)文本在实际部署环境下鲁棒性检测困难的问题,具体挑战包括跨领域与跨生成器的分布偏移、输入表面的对抗性扰动,以及目标域标签缺失导致阈值校准失效,这些问题均会显著降低在单一领域内表现良好的检测模型的性能。其解决方案的关键在于提出一种面向部署场景的检测框架DeBERTa-ConPara,该框架结合了攻击感知的Unicode预处理策略与基于HC3 Plus、M4、MAGE和RAID数据集训练的上下文变换器编码器。研究的核心发现是:预处理操作的作用方向取决于其应用阶段——在训练阶段进行归一化会去重并消除对抗性监督信号(如将RAID中35.4%的数据行合并为干净样本副本),而推理阶段的归一化则成为有效的防御手段。通过独立因子实验验证,原始训练数据配合推理阶段的归一化配置为最优方案,在官方RAID隐藏测试集上达到99.61% AUROC、99.01% TPR@5% FPR和96.57% TPR@1% FPR;在固定阈值下于HC3 Plus与MAGE上实现93.14%平均平衡准确率。性能提升主要体现在两类攻击类型上:同形异义字符攻击与零宽空格插入攻击的检测率分别从11.05%和1.12%跃升至96.98%,且该效果在不同架构的零样本检测器中可复现,表明其本质源于攻击特性而非模型结构。此外,研究还报告两项负面结果:通过改写实现语义不变增强及监督对比学习(ConPara)并未提升最优配置性能,而人工设计的特征融合分支在分布内无效且在分布外有害。
链接: https://arxiv.org/abs/2610.00883
作者: Mohamed Mady,Yupei Li,Johannes Reschke,Björn W. Schuller
机构: Technical University of Munich (慕尼黑工业大学); OTH Regensburg (奥格斯堡应用技术大学); Imperial College London (帝国理工学院)
类目: Computation and Language (cs.CL); Cryptography and Security (cs.CR)
备注: Accepted at AACL-IJCNLP 2026 (main conference). 9 pages plus appendix. Code and checkpoint: this https URL , this https URL
Abstract:Robust detection of AI-generated text under deployment conditions is challenging: distribution shifts across domains and generators, adversarial perturbations of the input surface, and the absence of target-domain labels for threshold calibration all degrade detectors that perform well in-domain. We present DeBERTa-ConPara, a deployment-oriented detector combining attack-aware Unicode preprocessing with a contextual transformer encoder trained over HC3 Plus, M4, MAGE and RAID. Our central finding is that preprocessing acts in opposite directions depending on where it is applied: normalising the training corpus deduplicates it, collapsing 35.4% of RAID rows into copies of their clean siblings and deleting the adversarial supervision, whereas normalising at inference is an effective defence. A factorial varying the two placements independently identifies raw training with normalised inference as the best configuration, reaching 99.61% AUROC, 99.01% TPR@5% FPR and 96.57% TPR@1% FPR on the official RAID hidden test, alongside 93.14% average balanced accuracy across HC3 Plus and MAGE under a fixed threshold. The gain is confined to two of twelve attack classes: homoglyph and zero-width-space insertion rise from 11.05% and 1.12% to 96.98%. The same signature reproduces in a zero-shot detector of different architecture, showing the effect belongs to the attacks rather than to our model. We additionally report two negative results: semantic-invariance augmentation through paraphrasing and supervised contrastive learning (ConPara) does not improve the best configuration, and the handcrafted feature-fusion branch is inert in distribution and harmful outside it.
[NLP-96] Kinematic MeanFlow: One-Step Action Generation Policy for Robotic Foundation Models
【速读】: 该论文旨在解决机器人基础模型(Robotic Foundation Models, RFMs)在执行动作生成时因多步流匹配(multi-step flow matching)导致的高推理延迟问题,致力于实现单步动作生成(one-step action generation)。其核心挑战在于:尽管MeanFlow框架为单步生成提供了潜在路径,但直接应用会导致性能急剧下降。研究发现,这一问题源于RFM速度场中两种独特的动态特性:(1)局部加速度(local acceleration)在去噪过程早期表现稳定,但在末期出现剧烈突增;(2)随着去噪进程推进,速度场幅值在不同样本间的分布范围显著扩大。针对上述问题,论文提出一种新型单步动作策略——运动学均值流(Kinematic MeanFlow, K-MF),其关键创新在于基于运动学恒等式(kinematic identity),将MeanFlow公式中的时间导数项分解为由中间点分隔的两个子区间项。该解耦结构使两部分分别捕捉去噪过程的早期与晚期动态行为,有效抑制了误差在全过程中的累积与放大。实验表明,K-MF可在从零训练和微调等多种场景下实现稳定且高效的单步动作生成,并在多数任务中超越传统多步流匹配方法。在推理效率方面,K-MF在L40与Jetson Orin平台的急切模式(eager)与编译模式(compiled)下,将GR00T-N1.6模型的动作头延迟降低67.5%~74.4%,从而实现端到端延迟下降30.3%~54.9%。
链接: https://arxiv.org/abs/2610.00864
作者: Jiawei Fan,Sifeng Wang,Yuqing Hou,Anbang Yao
机构: Intel Labs China (英特尔实验室中国); Midea AI Research (美的人工智能研究院)
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)
备注: Project page: this https URL
Abstract:In this paper, we study how to achieve one-step action generation in Robotic Foundation Models (RFMs), aiming to overcome the high inference latency of multi-step flow matching. MeanFlow provides a promising framework for this goal, yet its direct application leads to performance collapse. We discover that this stems from two distinctive dynamics exhibited in the RFM velocity field: (1) the ``local acceleration" exhibits stability early on, but surges sharply towards the end of the denoising process, and (2) the spread of its magnitudes across samples widens as denoising progresses. To address these issues, we introduce Kinematic MeanFlow (K-MF), a novel one-step action policy tailored for RFMs. Specifically, grounded in a kinematic identity, K-MF decouples the time derivative term in the MeanFlow formulation into two sub-interval terms separated by an intermediate point. This decoupled formulation enables the two terms to capture early-stage and late-stage denoising dynamics, respectively, while mitigating the error amplification across the process. As a result, our K-MF empowers RFMs to achieve one-step action generation in both training from scratch and fine-tuning paradigms across diverse tasks, while outperforming multi-step flow matching in most settings. In terms of inference efficiency, K-MF reduces action-head latency of GR00T-N1.6 by 67.5%~74.4% across L40 and Jetson Orin in eager and compiled modes, yielding end-to-end latency reductions of 30.3%~54.9%. Code will be available at this https URL.
[NLP-97] Child-Adapted Structured Phonological Representations for Interpretable Speech Sound Analysis ICASSP2027
【速读】: 该论文旨在解决现有语音表征模型在儿童语音(child speech)上泛化能力不足的问题,尤其针对当前大多数语音表示模型主要基于成人语音训练而缺乏对儿童语音特性的适配。其核心解决方案是通过使用与CHILDES语料库对齐的儿童语音数据,对PhonoQ-2.0模型进行儿童语音适应(child-speech adaptation),并系统评估不同对齐监督条件(成人、成人+儿童、仅儿童)和两种初始化策略(基于成人语音的PhonoQ初始化与随机初始化)下的性能表现。关键创新在于引入多阶段对齐监督机制,显著提升了模型在儿童语音中音位特征(如清浊音、发音方式)的识别能力,尤其在清浊音识别上宏观F1得分从0.922提升至0.972–0.987,且在发音部位(place)和特定音位对比(如软腭前移)的保留方面表现出较强的稳定性,验证了模型在纵向语音发展分析中的有效性。
链接: https://arxiv.org/abs/2610.00852
作者: Abner Hernandez,Tomás Arias Vergara,Andreas Maier,Paula Andrea Pérez-Toro
机构: 未知
类目: Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)
备注: Submitted for review at ICASSP 2027
Abstract:Structured phonological representations provide an interpretable alternative to generic speech embeddings, but existing models are largely trained on adult speech. We adapt PhonoQ-2.0 to child speech using CHILDES-Aligned data and compare three alignment-supervision conditions (Adult, Adult+Child, and Child-only) across two initialization strategies (Adult PhonoQ and scratch). Generalization is evaluated against manual child-speech annotations. On 1,352 consonant targets from 58 typically developing children, child-speech adaptation improves voicing recognition across all supervision conditions, from 0.922 macro-F1 for Adult PhonoQ to 0.972–0.987 after adaptation. Manner is more sensitive to alignment supervision: Adult+Child MFA reaches 0.804 and 0.796, compared to approximately 0.70 under Adult MFA supervision. Place remains comparatively strong across systems (0.871–0.902), although per-class performance varies substantially. The velar-fronting contrast is preserved across all seven model variants. Longitudinal UltraPhonix analysis further reveals speaker-specific velar and post-alveolar changes that are largely preserved across models and broadly consistent with reported clinical progress.
[NLP-98] AuraForge: Scaling Security Supervision for Training Coding Agents
【速读】: 该论文旨在解决生成式代码代理(Coding Agent)在缺乏有效安全监督的情况下,虽能实现功能正确性但难以保障实现过程安全性的问题。随着人类对代码审查从逐行检查转向结果层面的“放手式评估”,现有方法在仅依赖功能正确性评估时,无法识别潜在的安全漏洞,从而导致生成代码存在严重安全隐患。其核心挑战在于:真实世界代码库中难以大规模获取可靠、可执行的安全测试作为训练监督信号。为此,论文提出AuraForge,一种融合攻击导向的测试生成、语言可扩展的任务构建机制以及防奖励劫持(reward hacking)策略的解决方案。该方案通过合成并验证可执行的安全测试,构建了多语言、多通用缺陷类别(CWE)的训练环境AuraGym,涵盖来自344个真实仓库的679个可执行任务,覆盖177类CWE。实验表明,相较于人工编写的安全测试,使用AuraForge生成的测试用例数量平均提升约3倍,误报率降低83.23%,显著提升了监督质量;以合成测试训练的Qwen3.5-4B模型在功能通过率(FuncPass)和安全通过率(SecPass)上分别达到19.7和6.2,优于基于人工测试的训练结果(14.9和4.4),验证了该方法在提供多样化且可靠安全监督方面的有效性。
链接: https://arxiv.org/abs/2610.00850
作者: Danqing Wang,Songwen Zhao,Harsh Sharma,Jierui Wang,Andre Vicente Duarte,Ivan Bercovich,Lei Li
机构: Carnegie Mellon University (卡内基梅隆大学); University of California, Los Angeles (加州大学洛杉矶分校); ScOp Venture Capital
类目: Cryptography and Security (cs.CR); Computation and Language (cs.CL); Computers and Society (cs.CY)
备注:
Abstract:Coding agents are now proficient enough to generate complex software applications from a single prompt. As their capabilities have grown, human oversight has increasingly shifted from line-by-line code review toward hands-off evaluation of outcomes. However, recent studies have shown that such a transition exposes a critical risk: functional correctness alone does not guarantee a secure implementation. Despite growing attention to code security, training safer coding agents remains challenging because reliable security supervision is difficult to obtain at scale from real-world repositories. We introduce AuraForge to synthesize and validate executable security tests for training secure coding agents. Our approach combines attack-oriented test synthesis, language-extensible task construction, and safeguards against reward hacking. Using AuraForge, we construct AuraGym, a multi-language and multi-CWE executable training gym: 679 executable feature-implementation tasks from 344 real-world repositories across Python, JavaScript, and TypeScript, covering 177 CWE categories. On the subset with human-written security tests, AuraForge produces about 3 times as many test cases on average and reduces the false-positive rate by 83.23%, allowing alternative secure implementations to receive correct supervision. Training Qwen3.5-4B with synthesized security tests gains larger improvements than human-written security tests (average 19.7 FuncPass and 6.2 SecPass vs. 14.9 FuncPass and 4.4 SecPass) on three languages. These results demonstrate that AuraForge provides more diverse and reliable security supervision to train secure coding agents.
[NLP-99] Contextual trajectory and incremental contextual displacement: Towards using LLM s to understand dynamic utterance-specific meaning construction
【速读】: 该论文旨在解决语言模型在处理具有歧义性的自然语言表达(如花园路径句)时,如何动态捕捉语义构建过程中的上下文依赖性变化问题。传统基于静态或整体句法结构的表示方法难以揭示句子在逐步理解过程中认知负荷与语义冲突的动态演变。其解决方案的关键在于提出一种词元级增量轨迹(token-wise incremental trajectories) 方法:通过逐步向句子中添加词汇并重复计算每个词元的上下文词嵌入(Contextual Word Embeddings, CWEs),构建出嵌入表示随语境演进的连续路径。该方法能够有效捕捉语义发展过程中的表征偏移(representational displacement),并在花园路径句中重现已知的认知加工特征,如关键区域附近的语义干扰现象,并可靠区分歧义句与语义清晰对照句。研究进一步证明,不仅句级分类标记(CLS)蕴含歧义信息,普通词汇词元的轨迹亦能反映全局语用信息,表明语义表征在多个层级上分布。探索性分析还发现,该框架对其他涉及误导性或歧义性语言现象具有泛化能力。综上,词元级增量轨迹为利用生成式语言模型研究特定语句意义建构提供了可量化、可解释的动态分析范式。
链接: https://arxiv.org/abs/2610.00840
作者: Grayson Wycliffe Storer,Julia Witte Zimmerman
机构: University of Vermont (佛蒙特大学); Vermont Complex Systems Institute (佛蒙特复杂系统研究所)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注:
Abstract:Transformer-based large language models (LLMs) such as RoBERTa represent text using contextual word embeddings (CWEs), which alter the embeddings associated with each token based on surrounding context. We construct token-wise incremental trajectories by repeatedly recomputing a token’s CWE as successive words are added to a sentence, yielding a representation of how contextualized embeddings evolve as the utterance unfolds. We evaluate this approach using garden-path sentences as a test case with characteristic features. Token-wise trajectories reproduce known features of garden-path processing, including disruption around the critical region, and reliably distinguish garden-path sentences from matched disambiguated controls. We introduce several metrics for quantifying representational displacement across contextual increments and show that trajectory information can be highly predictive of sentence type. We find that ambiguity-related information is recoverable not only from the sentence-level CLS representation but also from ordinary vocabulary tokens, suggesting that utterance-level information is distributed across multiple representational scales. In exploratory analyses, we find qualitatively similar trajectory structures in other ambiguity- and misdirection-related linguistic phenomena. Together, these results establish token-wise incremental trajectories as a promising framework for studying utterance-specific meaning construction using LLMs.
[NLP-100] VERITYGATE: A Four-Gate Schema-Level Faithfulness Framework and Paired Benchmark for Grounded LLM Narrations over Structured Evidence EMNLP2026
【速读】: 该论文旨在解决生成式 AI(Generative AI)在生成解释时脱离结构化系统证据的问题,即模型输出的自然语言推理可能与所声明的证据ID、实体、数值及论断类型不一致。其解决方案的核心是提出VERITYGATE,一个基于四门检查机制的验证框架,分别针对声明的证据ID、实体、数值和论断类型进行结构化模式匹配校验。该框架不逐条验证文本中的所有事实,而是依据预定义的固定模式(schema-level contract)进行一致性检查,从而高效识别与证据链脱节的生成内容。实验表明,在未修复前,GPT-4o-mini与Claude Sonnet 4.6的验证失败率分别高达80.3%和47.9%,而一次修复后,两者存活率分别提升至28.0%和54.3%。此外,不同模型在修复后的输出数量变化差异显著,强调需同时报告验证通过率与输出量。最终,第四道门(Gate 4)覆盖了绝大多数失败案例(97.0%~100%),验证了该架构的有效性。研究还发现,尽管规则设计合理,但人工评估显示模式检查与正确文本之间仍存在差距,提示需进一步优化语义对齐。
链接: https://arxiv.org/abs/2610.00833
作者: Sachin Gupta
机构: Independent Researcher(独立研究员); San Jose, California, USA(加利福尼亚州圣何塞,美国)
类目: Computation and Language (cs.CL); Machine Learning (cs.LG)
备注: 16 pages, 5 figures, 7 tables. Accepted at Grounding Language Models: Learning Faithfully and Efficiently (GroundLM 2026), co-located with EMNLP 2026. Code: this https URL Supplementary artifact: this https URL
Abstract:Fluent LLM explanations may not follow the evidence from a structured system. We present VERITYGATE, a four-gate checker for declared evidence IDs, entities, numbers, and claim types. It checks a fixed schema; it does not verify every fact in the prose. At r=0 and r=1, we test 900 instances per setting (450 grounded-ungrounded pairs) with GPT-4o-mini, Llama-3.3-70B, and Claude Sonnet 4.6. Under this schema-level contract and before repair, 80.3% of mini claims and 47.9% of Sonnet claims fail. These are verifier rejection rates, not prose-hallucination rates. One repair pass raises claim survival from 19.7% to 28.0% for mini and from 52.1% to 54.3% for Sonnet. Verified claims per example change by +0.14 for mini, -0.71 for Llama, and -0.47 for Sonnet, so survival and output volume must be reported together. A second Sonnet pass gives no clear gain. At r=1, Gate 4 covers 97.0%, 98.7%, and 100% of failing claims for mini, Llama, and Sonnet. Small human studies support the rules but show gaps between schema checks and correct prose. A domain-specific GPT-4o judge test shows an order effect, so it is only a usefulness check. We release the code and data.
[NLP-101] Verbalized and Internal Probabilities Are Coupled in Large Language Models
【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)中内部不确定性(internal uncertainty)与显式表达的不确定性(verbalized uncertainty)之间是否对齐的问题。具体而言,尽管已有研究表明模型的内部采样概率反映训练数据中的相对频率,而模型输出的言语化置信度则对应训练数据中明确的概率性陈述,但二者是否在实际中保持一致仍不明确,尤其当训练数据中的分布性不确定性与显式概率声明不一致时。为弥补这一认知缺口,研究通过干预训练数据和上下文数据中的不确定性来源,系统性地考察了两种不确定性读出机制的响应特性。研究发现,无论是内部概率还是言语化概率,均受到训练数据中分布性不确定性(distributional uncertainty)和显式概率声明(asserted uncertainty)的双重影响;更重要的是,言语化概率与内部概率之间的对齐程度显著高于仅依赖共同不确定性源独立追踪所预期的水平,表明言语化不确定性可作为探测模型内部不确定性的有效代理指标。
链接: https://arxiv.org/abs/2610.00827
作者: Sinead Williamson,Jiaxuan Li,Nick Foti,Russ Webb,Masha Fedzechkina
机构: Apple(苹果)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注:
Abstract:Large language models carry an internal notion of uncertainty in their sampling distribution, i.e., the probabilities they place on generating one answer rather than another. They can also be asked to state a confidence, in words or as a number: a verbalized uncertainty. Prior work suggests that internal probabilities track relative frequencies in the training data, and that verbalized probabilities track explicit probabilistic assertions in the training data. However, we do not know whether these two readouts are aligned, except when frequencies and probabilistic assertions in the training data happen to align. This limits our understanding of when we can use verbalized uncertainties as a proxy for either training data frequencies, or a model’s internal distribution. We resolve this gap by systematically exploring how LLMs probability readouts are impacted by training and in-context data, via intervening on the underlying uncertainty sources in the data. We find that both internal and verbalized probability readouts are impacted by both distributional and asserted uncertainty in the training data. Further, we find that verbalized and internal probabilities are aligned beyond what would be expected by independently tracking the same uncertainty sources, suggesting that verbalized probabilities can be used to probe a model’s internal distribution.
[NLP-102] Paying for Too Many Tokens? Valid and Cost-Efficient Multimodal LLM Annotation with Simple Heuristics
【速读】: 该论文旨在解决生成式 AI 在视频标注中因高计算成本而面临的效率与有效性之间的矛盾问题。随着视觉-语言模型(Vision-Language Models, VLMs)在大规模视频标注中的应用,处理一段典型60秒的短视频(以每秒一帧计算)会产生数百万级别的令牌消耗,显著增加成本。为降低开销,现有研究普遍采用诸如帧采样、将视频压缩为图像网格或仅使用单一模态等启发式策略,但这些方法在节省成本的同时是否影响下游任务的准确性与推断有效性尚不明确。为此,本文针对短视频场景,在计算社会科学(Computational Social Science, CSS)中的情感分类和主题分类两个任务上,系统评估了多种启发式策略的综合表现,从三个维度进行分析:分类准确率、下游推断的有效性以及每视频的令牌消耗。研究发现,准确率与推断有效性存在显著偏离——高准确率配置可能产生错误结论;同时,模态价值并非必然:仅依赖文本即可实现优异性能,表明引入多模态可能增加成本却未带来额外信息增益。更为关键的是,通过简单的镜头切换检测构建的2×8图像网格,可在约15%的令牌成本下实现与全视频理解相当的效果(κ值差异小于0.05),从而实现了成本与视频长度的解耦。基于上述发现,论文提出了面向计算社会科学领域的成本敏感型VLM标注指南,核心在于通过轻量化图像网格设计与模态选择策略,在保障推断有效性的前提下大幅降低标注成本。
链接: https://arxiv.org/abs/2610.00809
作者: Zhixi Zhu,Kristina Gligoric
机构: Johns Hopkins University (约翰霍普金斯大学)
类目: Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL); Social and Information Networks (cs.SI)
备注:
Abstract:Vision-Language Models (VLMs) enable video annotation at scale, but costs accumulate quickly: processing a typical 60-second short-form video at one frame per second requires millions of tokens. To reduce costs, researchers rely on heuristics such as sampling a subset of frames, compressing videos into image grids, or using only a single modality. However, it remains unclear which heuristics save cost, and whether they preserve the downstream conclusions these annotations enable. To address this gap, we conduct a systematic evaluation of these heuristics using short-form videos, on two computational social science (CSS) tasks: sentiment and topic classification. We evaluate each configuration along three axes the literature typically treats separately: classification accuracy, validity of downstream inference, and per-video token cost. First, we find that accuracy and validity diverge: the highest-accuracy configuration can produce wrong conclusions. Second, modality value is not guaranteed: text alone can yield strong performance, indicating that adding modalities can add cost without adding signal. Finally, we find that cost can be decoupled from video length when annotating short-form videos: a single 2\times8 image grid built via simple shot-transition detection approaches full-video understanding ( \kappa within~.05), at \sim 15% of the token cost. Based on these findings, we derive guidelines that can enable cost-aware VLM annotation in CSS.
[NLP-103] Sapien: A Stateful Policy Engine for Autonomous AI Agents
【速读】: 该论文旨在解决多步任务中生成式AI代理因缺乏上下文状态感知而可能产生越权行为的安全隐患。传统上下文安全防御机制通常依赖静态策略,无法动态适应任务执行过程中状态变化对合法操作的约束需求。为此,论文提出Sapien——一种支持状态感知的上下文策略引擎,其核心创新在于通过扩展正则表达式以引入状态谓词(stateful predicates)、延迟策略生成(deferred policy generation)及作用域语义检查(scoped semantic checks),实现对工具调用序列的动态、上下文敏感约束。实验表明,Sapien在保持接近无约束代理90%以上效用的同时,即便在代理被完全劫持的情况下,仍能有效阻止AgentDojo上93-95%的攻击,以及Toolathlon上62-85%的攻击,显著优于传统工具白名单方案,尤其在长时序任务中表现更优。
链接: https://arxiv.org/abs/2610.00797
作者: Corinn Tiffany,Wen Zhang,Eugene Bagdasarian,Lillian Tsai
机构: Google(谷歌); University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Cryptography and Security (cs.CR)
备注:
Abstract:Contextual security defenses prevent AI agents from taking rogue actions by synthesizing a task-specific policy and enforcing it on the agent’s tool calls. In multi-step tasks, however, which actions are valid often depends on what the agent has already done and learned. We present Sapien, a policy engine for enforcing stateful contextual policies. A Sapien policy specifies permitted tool-call sequences using a regular expression extended with stateful predicates, deferred policy generation, and scoped semantic checks. We show that Sapien stays within a few percent of an unconstrained agent’s utility. Even if the agent is fully hijacked, Sapien’s policies rule out 93-95% of attacks on AgentDojo and 62-85% on Toolathlon (twice as many as tool allowlists on long-horizon tasks).
[NLP-104] Can large language models unlock discrete data in ophthalmic diagnostic reports?
【速读】: 该论文旨在解决眼科诊断PDF报告中结构化数据自动提取的准确性与效率问题,尤其针对非结构化文本信息转化为可分析格式的挑战。其核心解决方案在于对比两种基于大语言模型(Large Language Model, LLM)的提示工程策略:一是采用预定义JSON Schema的“Schema-Constrained”模式以确保输出格式合规;二是仅依赖详细指令提示(Prompt-Only)并通过Python后处理转换为JSON。研究结果表明,尽管两种方法均实现了高精度的数据提取(如Prompt-Only在所有报告类型中达到100%的值准确率),但二者各具优势——Schema-Constrained在格式一致性上表现卓越(100%格式准确率),而Prompt-Only则在值提取准确率上更优。两者结合可显著提升处理效率(较人工平均减少约92%时间),验证了基于通用大模型的自动化数据抽象流程在临床研究中的可行性与互补性,为构建混合式、具备验证反馈机制的智能数据抽取系统提供了关键支持。
链接: https://arxiv.org/abs/2610.00795
作者: Umair A. Zaidi,An-Lun Wu,Wei-Chun Lin,Thomas S. Hwang,Michelle R. Hribar
机构: 未知
类目: Computation and Language (cs.CL)
备注: 10 pages, 5 figures. Presented at the Association for Research in Vision and Ophthalmology (ARVO) Annual Meeting, Denver, Colorado, May 4, 2026
Abstract:Objective: To assess the accuracy and efficiency of a large language model (LLM) using two prompt strategies to extract structured data from ophthalmic diagnostic PDF reports. Methods: Twenty deidentified reports across four types (Visual Field, OCT Glaucoma Overview, OCT retinal nerve fiber layer Single Exam, and OCT Thickness Map; n = 5 each) were processed using two GPT-4o-assisted pipelines and compared with a reconciled manual ground truth. Schema-Constrained used Structured Output mode with a predefined JSON Schema; Prompt-Only used a detailed instruction prompt followed by Python conversion to JSON. Outcomes were value accuracy, formatting accuracy, and extraction time. Results: Schema-Constrained value accuracy was 100.00% for Visual Field and RNFL Single Exam, 97.45% for Glaucoma Overview, and 98.00% for Thickness Map; Prompt-Only achieved 100.00% across all four report types. Formatting accuracy was 100.00% for Schema-Constrained across all report types and 100.00% for Prompt-Only except RNFL Single Exam (90.14%). Mean extraction time was 56.51 s per report for manual review versus 5.04 s for Schema-Constrained and 4.70 s for Prompt-Only, an approximately 92% reduction. Conclusions: In this small proof-of-concept dataset, general-purpose LLM-assisted pipelines extracted structured data from ophthalmic diagnostic PDFs with high accuracy and substantially reduced processing time. Prompt-Only achieved the highest value accuracy, while Schema-Constrained produced schema-compliant output with 100% formatting accuracy. These complementary strengths support further evaluation of hybrid, validation-aware workflows for research and clinical data abstraction.
[NLP-105] Effective Synthetic Data Curation Requires Group-Level Signals
【速读】: 该论文旨在解决大规模合成数据(synthetic data)用于大语言模型(LLM)训练时导致模型生成能力退化的关键问题,核心在于如何有效筛选具有高训练价值的合成数据。现有数据筛选方法依赖个体级信号(individual-level signals),即孤立评估单个数据样本的训练效用,但研究表明,这种做法在合成数据场景下存在根本性局限。论文提出,真正有效的数据筛选需依赖群体级信号(group-level signals),即考虑数据样本之间相互作用的综合效用评估。研究发现,不同数据组合在个体级信号下可能表现相似,但在群体级信号下差异显著,且基于群体级信号的筛选能显著提升下游任务性能,尤其是在生成能力方面。此外,随着训练流程中合成数据占比增加,群体级信号的重要性愈发凸显——仅采用包含此类信号的筛选方法才能超越基准性能,且当样本间关联权重被放大时,性能增益进一步增强。为此,论文还提出一种低成本诊断方法,帮助模型开发者在计算资源受限的情况下,精准识别最需进行群体级评估的数据组,从而以极低的计算成本实现接近全量群体级评分的收益。
链接: https://arxiv.org/abs/2610.00779
作者: Cathy Jiao,Chenyan Xiong
机构: Carnegie Mellon University (卡内基梅隆大学)
类目: Computation and Language (cs.CL)
备注:
Abstract:Synthetic data now is essential to LLM training, used to strengthen advanced capabilities such as autonomous and long-horizon task execution. Yet recent work shows that training on it at scale can degrade model generation, making it important to decide what synthetic data is worth training on. While current data curation practices do so with individual-level signals (i.e., estimates of each data sample’s training utility in isolation), across pre-training and post-training settings we show that this is insufficient for synthetic data, and that group-level signals (i.e., estimates of utility that account for interactions among data samples) are necessary for effective data curation. First, we show that individual-level signals are blind to how samples jointly affect training: synthetic datasets with different compositions can be indistinguishable under individual-level influence yet differ sharply under group-level influence, and curating by the latter yields better downstream performance, particularly in generative capability. Second, we find that group-level signals matter more as training pipelines become increasingly synthetic: among widely used data curation methods, only those incorporating them improve over baseline, with gains increasing when weights capturing relations among samples are amplified. Finally, we translate these findings into practice – for model developers under a compute budget, we offer a cheap diagnostic that prioritizes which groups of synthetic data most need group-level estimation, recovering much of the benefit of full group-level scoring at a fraction of the compute cost.
[NLP-106] Pre-training interventions ex post facto: Grafting model beliefs across checkpoints
【速读】: 该论文旨在解决预训练阶段干预(pre-training intervention)中因模型信念固化而导致的对齐难题,特别是针对合成文档微调(Synthetic Document Fine-Tuning, SDF)在后训练阶段应用时引发的“现实漂移”(reality drift)问题——即模型将虚构实体误认为真实存在。传统方法需在预训练完成后才进行SDF,导致干预效果受限且引入副作用,如能力退化与偏好一致性丧失。其解决方案的关键在于提出一种名为“嫁接”(grafting)的新方法:先在预训练检查点上训练一个SDF适配器,再将学习到的权重更新直接叠加至已后训练的模型,从而近似实现中段预训练干预的效果,同时避免重复后训练。该方法不仅显著降低了现实漂移和偏好一致性损失(平均减少超过一半),还保持了与忠实中段干预更接近的行为表现。由于无需重新进行后训练,同一适配器可复用于任意后续检查点,极大提升了预训练干预的迭代效率,使研究人员仅需一次微调即可快速验证不同干预策略。
链接: https://arxiv.org/abs/2610.00767
作者: Peter Nutter,Dani Roytburg,Clément Dumas,Jinghua Ou,Shi Feng
机构: ETH Zurich(苏黎世联邦理工学院); Carnegie Mellon University(卡内基梅隆大学); MATS; Astra Fellowship; George Washington University(乔治华盛顿大学)
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: 78 pages. Code: this https URL
Abstract:Pre-training interventions are critical to alignment research, since beliefs formed during pre-training shape how a model generalizes from later training. One recently popular technique for such interventions is synthetic document fine-tuning (SDF), which aims to alter what the model believes. Ideally, synthetic documents would be mixed into pre- or mid-training, but every change to a pre-training corpus must be followed by a full post-training run before its effect can be measured, making iteration slow and expensive. Common practice instead applies SDF to an already post-trained model. This is known to leave artifacts and degrade capabilities, and, as we show, it makes the model treat fabricated entities unrelated to the documents as real, a failure we call reality drift. We propose grafting: train the SDF adapter on the pre-trained checkpoint, then add the learned weight update to the post-trained model, which approximates the faithful approach while reusing the existing post-training. We demonstrate this by installing false facts, training misaligned model organisms and applying a constitutional mid-training intervention, across model families up to 284B parameters. Grafting installs the target belief as strongly as SDF on the post-trained model while reducing both reality drift and the loss of preference coherence by more than half on average, and it stays closer to a faithful mid-training run. Because grafting requires no post-training, the same adapter can be applied to any later checkpoint, enabling researchers to iterate quickly on pre-training interventions at the cost of a single fine-tuning run.
[NLP-107] Reason in Style: Discovering and Controlling Style in Language Models
【速读】: 该论文旨在解决生成式语言模型(Generative AI)在输出中内容与风格耦合导致的风格不可控、难以识别的问题。其核心挑战在于如何在不依赖人工标注的情况下,自动发现并显式控制模型输出中的重复性风格模式。解决方案的关键是提出一种无监督算法,能够从语言模型的输出中分离出内容(content)与风格(style)的表征,并基于此在数学推理任务的可控设置下验证其有效性。通过对九个教师模型生成的超10万条经验证轨迹进行分析,研究发现了六种反复出现但分布不均的风格模式;随后通过重要性加权(importance weighting)对训练数据进行平衡,微调小型学生模型以在明确条件化时遵循特定风格。实验表明,该方法在六个数学推理基准上显著提升了Pass@k指标,且实现了请求风格与实际输出风格的高度一致。更重要的是,研究发现风格显著影响答案正确率:不同问题类型对不同风格具有偏好性,表明风格不仅是可控的表达形式,还能提升推理性能。综上,该工作首次证明了生成数据中的风格可被无监督发现并显式控制,为风格化生成提供了新的范式,兼具可控性与性能增益。
链接: https://arxiv.org/abs/2610.00724
作者: Ioana Marinescu,Eric Karl Oermann,Kyunghyun Cho
机构: NYU(纽约大学); NYU Langone Health(纽约大学朗格尼健康)
类目: Computation and Language (cs.CL); Machine Learning (cs.LG)
备注: 38 pages, 12 figures
Abstract:Language models learn content and style jointly, making stylistic variation in their outputs difficult to identify and control. We study whether recurring styles in model responses can be discovered without supervision and explicitly controlled. We design an algorithm that learns to separate representations of content and style from language models’ outputs and validate its effectiveness on math questions in a controlled setting. By applying this method to over 100K verified traces from nine distinct teacher models, we discover six recurring yet imbalanced styles. We then fine-tune smaller student models to follow these styles when explicitly conditioned on them, using importance weighting to balance the contribution of the styles represented in the corpus. This approach improves Pass@ k over standard fine-tuning on the same data across six math reasoning benchmarks, demonstrating that we can diversify the style of answers effectively. We confirm that this also results in strong correspondence between requested and realized styles. We find that style affects correctness: the probability of solving a problem depends on the style we condition on, and different problems benefit from different styles. In summary, our results show that stylistic variation in model-generated data can be discovered in an unsupervised way, and made explicit, providing a source of both control and improved reasoning performance.
[NLP-108] Sequential Functional Structured Tucker Compression for Large Language Model Attentions
【速读】: 该论文旨在解决大语言模型(Large Language Model, LLM)注意力机制在后训练压缩过程中存在的两个关键问题:一是现有方法通常将注意力投影矩阵的压缩视为独立的矩阵近似任务,忽略了不同注意力头之间共享的结构特性;二是未考虑早期压缩操作引入的表示偏移(representation shift),导致后续压缩效率下降。为此,论文提出一种顺序化结构化压缩框架FTC(Sequential Structured Compression),其核心在于在固定存储预算下,动态适应当前已压缩模型的状态,联合利用原始查询(Q)、键(K)、值(V)头的天然结构进行联合近似,并单独处理输出投影以应对注意力后表示的变化。该方法无需微调或基于梯度的恢复过程,具有良好的泛化能力。在7个参数规模从6B到32B的解码器仅模型上测试的结果表明,FTC在所有压缩率设置下均优于对比方法,在5种现代分组查询注意力(GQA)模型上的WikiText-2困惑度最低,尤其在激进压缩条件下性能提升最为显著,且在下游任务中仍保持明显优势,甚至在32B模型规模下依然表现优异。
链接: https://arxiv.org/abs/2610.00717
作者: Jiangfeng Chen,Xinyu Wang,Tianshuo Yan,Hanwei Wu,Xiao-Wen Chang,Yang Zhang,Lei Ding
机构: University of Manitoba (曼尼托巴大学); McGill University (麦吉尔大学); Simpleway; The University of Hong Kong (香港大学); McMaster University (麦克马斯特大学)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (stat.ML)
备注:
Abstract:Post-training compression of LLM attention is often formulated as independent matrix approximation, ignoring both the shared structure among attention projections and the representation shift introduced by earlier compression. We propose FTC, a sequential structured compression framework that adapts the approximation to the current compressed model while jointly exploiting the native Q/K/V head structure under a fixed storage budget. The output projection is handled separately to account for the changed post-attention representation. FTC requires neither fine-tuning nor gradient-based recovery. Across seven decoder-only LLMs from 6B to 32B parameters, FTC achieves the lowest WikiText-2 perplexity among the compared methods at every tested keep ratio on five modern GQA models, with the largest gains under aggressive compression. The improvements transfer to downstream tasks and remain substantial at the 32B scale.
[NLP-109] Initialization Improves LLM -Driven Discovery
【速读】: 该论文旨在解决生成式人工智能在算法、定理、药物等新发现任务中,基于大语言模型(Large Language Models, LLMs)的迭代优化框架在实际应用中表现不稳定且成功率难以预测的问题。其核心挑战在于现有提示工程框架(harnesses)对最终发现结果的敏感性,以及模式崩溃(mode collapse)这一普遍出现的失败现象——即在迭代过程中生成样本的多样性急剧下降,导致探索陷入局部最优。尽管已有研究尝试通过引入多样性增强机制来延缓模式崩溃,但效果不一致且不可靠。本文的关键发现是:早期发现的质量具有高度预测性,能够有效预示后续成功与否。基于此,论文提出一种通用性强的干预策略——在正式迭代优化前引入并行探索阶段以实现更优的初始状态初始化。该方法在多种提示框架与多类发现任务中均表现出稳定提升,验证了初始化质量对生成式AI驱动发现过程的关键作用。
链接: https://arxiv.org/abs/2610.00707
作者: Mansi Sakarvadia,Marco Ciccone,Colin Raffel
机构: University of Chicago(芝加哥大学); Vector Institute(向量研究所); University of Toronto(多伦多大学)
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注:
Abstract:Large Language Models (LLMs) have been used for novel discovery of algorithms, theorems, drugs, and other tasks through the use of harnesses that prompt an LLM to iteratively optimize an objective. In this work, we study the relationship between the population of previous iterates and eventual discovery success. We generalize past work on harness design to develop a suite of 12 harnesses called ‘Modular’ and characterize their performance across 5 diverse discovery tasks, finding that discovery success is brittle and sensitive to harness design. We uncover mode collapse, characterized by a dramatic drop in the diversity of iterates, as a common failure mode. We find that popular state-of-the-art harnesses and diversity-inducing harness interventions, which aim to prolong this collapse, yield inconsistent gains. Our results instead uncover that the performance of early discoveries is predictive of eventual success. We therefore propose a universally applicable intervention that performs an initial stage of parallel exploration in order to initialize subsequent iterative optimization. Our method provides consistent gains across many harnesses and target applications, confirming the importance of initialization in LLM-driven discovery.
[NLP-110] How Divergence Becomes Decision Flips in Compressed Language Models
【速读】: 该论文旨在解决压缩语言模型(compressed language model)与原始密集模型(dense model)之间输出差异评估的准确性问题,尤其关注在部署场景中依赖密集模型输出时,如何准确衡量压缩模型决策变化的程度。传统方法通常使用相对熵(KL divergence)作为度量指标,但研究表明,其对决策变化的预测能力存在系统性偏差:KL值需经过平方根变换并乘以一个在不同模型和语料间可变四倍的系数才能近似反映决策变化率(即arg-max token的翻转率,flip rate),且其统计特性因在所有词元上先求平均再开方而引入失真。相比之下,总变差(total variation)能够直接、无偏地追踪翻转率,其比率中位数高达1.05,且无需任何拟合参数。在802个压缩及扰动副本的实验中,当两个压缩器的翻转率差异超过10%时,基于KL的判断在11%的情况下错误地将更大的决策变化归于较小的KL值,而总变差仅在1%的情况下出现误判。此外,在预注册测试中,总变差在代码语料上的表现一致稳定,而KL的多个预测假设在半数以上模型中失效;在真实内核的三类新模型中,总变差在37/38个检查点中保持有效范围,仅在代码任务上对两个模型略低于1。在vLLM推测解码(speculative decoding)中,总变差在教师强制(teacher forcing)条件下测量,即可实现对贪婪草稿接受率的预测,均方相对误差仅为1.1–2.4%,远优于依赖特定任务校准的KL方法。因此,该研究的核心解决方案在于:以总变差替代KL作为压缩模型输出差异的评估标准,因其能更直接、稳健且无需校准地反映实际决策变化情况。
链接: https://arxiv.org/abs/2610.00694
作者: Beatriz Almeida Felicio
机构: Instituto de Informática (INF); Universidade Federal de Goiás (UFG)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Machine Learning (stat.ML)
备注: Preprint
Abstract:Compression reports summarize how far a compressed language model moved from the dense one, usually by a KL divergence; a deployment that relies on the dense model’s outputs needs to know how many of its decisions changed. We show that total variation, not KL, answers this directly. Across 802 compressed and perturbed copies of 19 open models on five corpora and nine mechanically unrelated perturbation families, the rate at which the arg-max token changes (the \emphflip rate) tracks total variation at a ratio with median 1.05 , with no fitted constant. KL converts into flips only through its square root and a factor that varies fourfold across models and corpora, because KL averages over tokens before the root is taken; first-order statistics averaged per token, such as Hellinger distance, avoid this, but reports rarely give them. As a result, of two compressors reported on different models and corpora whose flip rates differ by at least 10% , KL assigns the smaller divergence to the one that changes more decisions in 11% of cases, total variation in 1% . Two pre-registered tests mark the limits: on a held-out code corpus the ratio held for all eight models while three predictions about KL each failed for half of them or more, and on three new models with real kernels it stayed in its band for 37 of 38 checkpoints but fell below one on code for two models. In vLLM speculative decoding, total variation measured under teacher forcing predicts greedy draft acceptance with a mean relative error of 1.1 – 2.4% , without the task-specific calibration that KL needs.
[NLP-111] owards Robust Numerical Claim Verification AACL
【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在数值推理任务中脆弱性问题,即微小的数值扰动(如标签翻转)会导致模型准确率急剧下降。其核心解决方案是通过对抗性微调(adversarial fine-tuning)在数值扰动样本上进行训练,结合参数高效微调(parameter-efficient fine-tuning)策略,使小型Qwen3模型(0.6B–8B)在标签翻转扰动下达到98.7%的准确率,显著优于未微调的大规模零样本模型及前沿系统(如GPT-5.4 Pro的74.0%和Gemini 2.5 Flash的73.9%)。关键在于,该方法不仅提升了对已知扰动类型的鲁棒性,还能泛化至未见过的扰动类型,表明模型学习到了更稳健的数值决策边界而非简单记忆;此外,该鲁棒性可无须目标领域数据迁移至跨语言场景(如西班牙语),并同样适用于证据侧扰动,验证了该微调方案的广泛适用性。
链接: https://arxiv.org/abs/2610.00689
作者: Peter Røysland Aarnes,Vinay Setty
机构: University of Stavanger (斯塔万格大学); Factiverse AI
类目: Computation and Language (cs.CL)
备注: Accepted to AACL-IJCNLP 2026 Findings
Abstract:Large language models (LLMs) are widely used for claim verification, yet remain brittle for numerical reasoning: even small changes in value can sharply degrade accuracy. We show that this brittleness persists in frontier LLMs, but can be mitigated through adversarial fine-tuning on numerically perturbed examples. Using parameter-efficient fine-tuning, small Qwen3 models (0.6B \unicodex2013 8B) reach 98.7% accuracy on label-flipping perturbations, outperforming larger zero-shot models and frontier systems (GPT-5.4 Pro (74.0%) and Gemini 2.5 Flash (73.9%)). The gains generalise to unseen perturbation types, indicating robust numerical decision boundaries rather than memorised edits. Robustness also transfers without target-domain data, significantly improving cross-lingual performance in Spanish. We further show that the same fine-tuning recipe confers robustness to evidence-side perturbations, using the VitaminC dataset.
[NLP-112] Bayesian Fine-tuning Yields Language Models that are as Bayesian as their Beliefs Allow
【速读】: 该论文旨在解决语言模型(Language Models, LMs)在面对需要基于少量观测推断隐变量的任务时,如何实现真正符合贝叶斯推理规范的决策行为。尽管通过监督微调(Supervised Fine-Tuning, SFT)使模型学习最优贝叶斯模型的输出可使其表现出近似贝叶斯的行为,但传统的基于真实答案(oracle)的SFT却无法达到相同效果。其核心问题在于:为何基于贝叶斯信号或真值信号进行微调会产生不同结果?关键在于模型是否真正以贝叶斯信念的形式表征信息、是否在内部计算中遵循贝叶斯规则,并将这些信念转化为最终决策。为此,研究提出了一系列递进式评估标准,涵盖行为、表征与计算层面,以检验模型是否可被视为贝叶斯决策者。实验结果显示,经贝叶斯信号微调的语言模型不仅在行为上呈现贝叶斯特性,还在中间层编码了贝叶斯规则中的关键量(如后验概率),并部分利用这些编码信念生成推荐;而基于真值信号微调的模型则在信念表征及信念读出机制上均存在差异。通过跨模型交换信念,部分贝叶斯优势得以转移,表明贝叶斯微调能够有效植入可被利用的贝叶斯信念,从而在不确定性推理中优于传统基于真值的监督微调,凸显了精细化监督信号在提升模型认知能力方面的关键作用。
链接: https://arxiv.org/abs/2610.00679
作者: Polina Tsvilodub,Andreas Waldis,Linlu Qiu,Tal Linzen,Michael Franke
机构: Massachusetts Institute of Technology (麻省理工学院); New York University (纽约大学); University of Tübingen (图宾根大学)
类目: Computation and Language (cs.CL)
备注: under review, 27 pages, 25 figures
Abstract:Language models (LMs) are increasingly used for tasks that require reasoning about hidden variables from a few observations, for which Bayesian inference is the normatively correct solution. While supervised fine-tuning of an LM on the outputs of an optimal \textitBayesian model leads to near-Bayesian behavior, standard supervised fine-tuning (SFT) on the true answers for the task falls short of it. But behavior alone does not tell us \textitwhy tuning on a \textitBayesian or an \textitoracle (true answers) signal differs: whether the resulting LM represents Bayesian beliefs, acts on them, or turns them into a choice the way Bayes’ rule does. To compare them, we formulate increasingly demanding requirements for an LM to count as a Bayesian decision maker, spanning its behavior, representations, and computations, and test them on a flight recommendation task. The Bayes-trained LM acts Bayesian, encodes quantities of Bayes’ rule in its middle layers, and uses the encoded belief for the recommendation to a certain extent. The oracle-trained LM differs from it both in the beliefs it holds and whether it reads beliefs out into recommendations. Exchanging beliefs between the LMs transfers a part of the Bayesian advantage. Bayes fine-tuning thus installs usable Bayesian beliefs in an LM for reasoning under uncertainty in a way standard SFT on oracle answers cannot, highlighting the advantage of nuanced supervision.
[NLP-113] Closing the Loop: Practical Training Recipes for Looped Language Models
【速读】: 该论文旨在解决循环语言模型(Looped Language Models, LoopLM)在大规模训练中存在训练成本高、多阶段训练依赖性强,且难以将模型性能提升归因于循环结构本身的问题。现有方法通常需要跨万亿级词元的多阶段训练,导致计算开销巨大,同时性能增益常被数据和训练策略差异所混淆。为此,本文提出一套高效的从头训练方案,其关键在于:(1)通过预训练后引入高质量中间训练阶段,结合学习率预热与更强的退出门控正则化,实现无需多阶段调度的稳定循环训练;(2)设计仅需一个可学习输入混合标量和光滑退出损失的极简转换机制,即可将预训练稠密模型高效转化为循环模型,无需额外步骤参数;(3)在严格控制变量条件下验证,1.4B参数的LoopLM在12个基准测试中均优于同规模稠密模型,尤其在GSM8K、MATH和DROP上分别提升14、10和22点,且在等效推理计算量下逼近3.9B稠密模型性能,但仅使用36%参数。上述成果表明,循环结构本身可带来显著性能增益,且训练与迁移成本大幅降低,使循环语言模型具备从头训练可行性与现有模型升级实用性。
链接: https://arxiv.org/abs/2610.00673
作者: Andrei Marchenko,Viacheslav Bezrukov,Oleg Kashurin,Inessa Fedorova,Dmitry Bocharov,Yuliana Shakhvalieva,Maria Tikhonova,Valerii Ternovskii
机构: RND NLP, DAIMLD, Russian Federation
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注:
Abstract:Looped language models increase effective depth by repeatedly applying a shared block of layers, but existing large-scale recipes require multi-stage training over trillions of tokens, while the benefits of recurrence remain difficult to separate from differences in data and training. In this work, we establish practical training recipes for looped language models, with three main results. (1) We develop a compute-efficient from-scratch pipeline that reduces the training budget from 7.7T tokens in Ouro to 310B tokens while retaining strong reasoning performance. Pretraining followed by high-quality mid-training, together with learning-rate warmup and stronger exit-gate regularization, enables stable recurrent training without prior multi-stage schedules. (2) Under controlled comparisons, our 1.4B LoopLM outperforms a parameter-matched dense model trained on the same data and token budget on all 12 evaluated benchmarks, including +14 points on GSM8K, +10 on MATH, and +22 on DROP. At matched inference compute, it approaches a 3.9B dense model on mathematical reasoning and reading comprehension while using only 36% as many parameters. (3) We introduce a minimal recipe for converting pretrained dense models into looped ones: a single learned input-mixing scalar and a smoothed exit loss, with no step-specific parameters. Applied to Qwen3-1.7B-Base, Looped Qwen improves over an identically continued dense baseline on every evaluated benchmark across two data regimes, with statistically clear gains on GSM8K, MATH, and MMLU-Pro on the curated mixture. Together, these results make looped language models substantially cheaper to train from scratch and practical to introduce into existing pretrained checkpoints, while isolating the gains due to recurrence itself.
[NLP-114] PhysicsMate: A Curriculum-Grounded Bengali Benchmark for Secondary Physics QA with Small-Model Adaptation
【速读】: 该论文旨在解决孟加拉国中学阶段科学、技术、工程与数学(STEM)教育中缺乏基于课程标准的物理问题求解评估基准,以及通用语言模型在处理物理问题所需的精确术语、单位规范和推导过程方面表现不佳的问题。其解决方案的关键在于构建一个名为PhysicsMate的基准数据集,该数据集包含1834个源自国家课程与教科书委员会(NCTB)九至十年级物理课程的问答对,并基于涵盖10种本体类型、共1760个节点和2600条边的多关系知识图谱进行语义建模。研究采用低秩适配(Low-Rank Adaptation)方法,在0.6B、1.7B和4B参数量级上实施统一微调策略,显著提升了各规模模型在闭卷测试中的准确率(分别提升5.5、15.0和23.3个百分点)。节点类型分析表明,结构化课程知识(如物理量和命名定律)最受益于微调,而松散定义的实体级知识受益最少。此外,4B模型经适配与量化后生成小型离线二进制文件,可在资源受限环境中实现本地推理,为网络连接和硬件条件有限场景下的课程对齐物理学习支持提供了可行路径。
链接: https://arxiv.org/abs/2610.00664
作者: Rashid Azraf Jahin,Saadman Sajid,Khan Raiyan Ibne Reza,Sumaiya Tabassum Nimi
机构: North South University(北方大学)
类目: Computation and Language (cs.CL)
备注: 6 pages, 3 figures, 5 tables. Accepted at 11th IEEE Asia-Pacific Conference on Computer Science and Data Engineering (IEEE CSDE 2026)
Abstract:Bengali secondary education lacks curriculum-grounded benchmarks for STEM question-solving, and general-purpose language models struggle with the precise terminology, unit conventions, and derivations that physics problems demand. We introduce PhysicsMate, a benchmark of 1834 question-answer pairs built from the National Curriculum and Textbook Board (NCTB) Grade 9-10 physics syllabus and grounded in a multi-relational knowledge graph of 1760 nodes and 2600 edges across ten ontological types. We Low-Rank Adapt at 0.6B, 1.7B, and 4B parameters, with a unified recipe and demonstrate a significant increase in closed-book accuracy in all scales (+5.5, +15.0, and +23.3 percentage points). A node-type analysis shows that the most benefited by adaptation is the structured curricular knowledge, which consists of physical quantities and named laws, while the least benefited is the loosely specified entity-level knowledge. The 4B model has been adapted and quantized to a small offline binary that can be used for local inference in resource constrained environments and offers a viable path to curriculum aligned physics support in environments with limited connectivity and hardware.
[NLP-115] Lingtai: What Concept Geometry Reveals–and Does Not Reveal–About LLM Inference
【速读】: 该论文旨在解决在无需训练探针的情况下,实时观测大语言模型(Large Language Model, LLM)在自回归推理过程中计算状态的难题。其核心挑战在于如何在不依赖标注数据、标签信息、梯度拟合或激活空间优化的前提下,实现对模型内部语义概念活动的可解释性追踪。为此,作者提出Lingtai——一种无需训练的概念遥测层(concept telemetry layer),通过将每一步生成的残差状态投影到由无监督构建的、命名化的概念锚点库(domain-specific bank of named concept anchors)上,生成结构化的每步概念坐标信号。该方法的关键创新在于:概念锚点的构建完全无需监督信号,且能有效捕捉与预测不确定性高度相关的动态模式。实验表明,该信号在代码生成和小学数学推理任务中均表现出强鲁棒性,其与不确定性的关联不受问题身份或词元位置影响,亦非单一词元类型所致,且无法由随机锚点或传统降维方法(如K-means、PCA)复现。进一步分析揭示两种关键结构:一是任务条件依赖的功能几何特性,表现为不同任务下活动熵形状的差异;二是执行路径特有的轨迹身份特征,具有强局部惯性但弱重实例化不变性。此外,审计结果表明,所用标量概念活动信号无法作为稳定的正确性坐标,因此正确性仍被视为外部供给。该遥测机制引入0.7–1.6%的每词元解码开销,但不影响生成内容的准确性。
链接: https://arxiv.org/abs/2610.00656
作者: Jiangang Chen
机构: Chengdu Beiluoshimen Technology Co., Ltd.(成都北洛门科技有限公司)
类目: Computation and Language (cs.CL)
备注: 15 pages, 4 figures, 7 tables. An earlier version was publicly released on Zenodo (DOI: https://doi.org/10.5281/zenodo.23068698 )
Abstract:Observing what a large language model computes during autoregressive inference–online and without training probes–remains difficult. We introduce Lingtai, a training-free concept telemetry layer: at each generation step, residual states are projected onto a domain-specific bank of named concept anchors, constructed without labeled concept examples, outcome labels, gradient fitting, or activation-space optimization, producing a structured per-step concept-coordinate signal. Across code generation and grade-school mathematical reasoning, this signal exhibits a robust association with predictive uncertainty: the association survives problem-identity and token-position controls and is not attributable to a single token type, is not explained by a simple correct/incorrect mixture on GSM8K, and is not reproduced by matched random anchors; it is markedly weaker or direction-inconsistent in K-means and PCA projections. Two structures emerge: a recurring uncertainty-linked activity signal whose functional geometry is task-conditioned (distinct activity-entropy shapes on HumanEval, MBPP, and GSM8K), and an execution-specific trajectory identity with strong local inertia but weak re-instantiation invariance–under completion-only elastic alignment, corruption at k=32 (approximately a median quarter of the completion) on the matched re-execution subset still retrieves the archived episode at 62.0%, while a fresh execution retrieves it only 11.7-16.0% of the time. Finally, a matched audit finds no evidence that the scalar concept-activity signal used here supplies a stable correctness coordinate under the tested protocol; we therefore treat correctness as externally supplied. Telemetry adds 0.7-1.6% per-token decode overhead for the 161-anchor code implementation, with unchanged generated tokens.
[NLP-116] Self-Evolving Coding Rules for AI Coding Agents NEURIPS2026
【速读】: 该论文旨在解决现有生成式AI代码代理(AI coding agents)依赖于人工设计且固定不变的编码规则(coding rules)所导致的效率低下与性能欠优问题。此类规则不仅开发过程耗时费力,且难以适应多样化的编程任务需求。其解决方案的关键在于提出一个自进化框架RuleEvolve,通过维护一组候选编码规则池,并在每轮迭代中利用大语言模型(LLM)驱动的变异模块(mutator module)生成规则变体,再由评估模块(judge module)对变体进行性能评估,从而动态筛选并更新最优规则。该方法实现了编码规则的自动化演化,在多个主流代码代理框架、四种基础大语言模型及三个基准测试上均表现出更优的功能正确性、代码简洁性以及生成成本控制能力,显著超越了传统手工调优与现有提示优化基线。
链接: https://arxiv.org/abs/2610.00650
作者: Zhengyuan Jiang,Reachal Wang,Yuepeng Hu,Yupu Wang,Yuqi Jia,Neil Zhenqiang Gong
机构: Duke University (杜克大学)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: Accepted by NeurIPS 2026
Abstract:The performance of AI coding agents is highly dependent on their underlying coding rules. However, existing coding rules are typically hand-crafted and fixed, making the process labor-intensive and often suboptimal. In this work, we propose RuleEvolve, a self-evolving framework for coding rules. RuleEvolve maintains a pool of candidate coding rules and iteratively improves them. In each iteration, it employs an LLM-powered mutator module to generate variants from existing candidates, and then uses a judge module to evaluate these variants and update the pool with the best-performing ones. Extensive evaluations across two coding-agent frameworks, four backbone LLMs, and three benchmarks demonstrate that RuleEvolve outperforms both manual engineering and existing prompt optimization baselines in terms of functional correctness of the generated code, code length, and/or generation cost (e.g., tokens used).
[NLP-117] Mixture of Decoders for Diverse Dialog Response Generation
【速读】: 该论文旨在解决序列到序列(sequence-to-sequence)模型在对话回复生成中普遍存在的多样性不足问题。研究表明,这一问题的根本原因在于序列到序列模型倾向于学习一个退化的单模态(uni-modal)响应分布,导致生成结果重复且缺乏变化。为解决此问题,论文提出在条件变分自编码器(Conditional Variational Autoencoder, CVAE)框架下引入多解码器混合(mixture of decoders)结构,使每个解码器专注于学习特定主题或语义模式,从而实现对多样化响应的有效建模。该方案的关键在于通过混合解码器机制打破单一解码器的局限性,增强生成过程的多样性与表达能力。实验在开放域聊天语料库上进行,结果表明,该方法在定量指标和人工评估中均显著优于现有强基线模型。
链接: https://arxiv.org/abs/2610.00621
作者: Wenchao Du
机构: Microsoft Corporation(微软)
类目: Computation and Language (cs.CL)
备注: preprints
Abstract:Mixture modeling is a long established machine learning technique for learning large sets of multi-modal data. While it is known that sequence-to-sequence models for dialog response generation suffer from the problem of low diversity, we hypothesize that it is because sequence-to-sequence models tend to learn a degenerate uni-modal distribution of responses. We then propose to incorporate a mixture of decoders into sequence-to-sequence models and try to make each decoder learn specialized topics in order to improve the diversity of generated responses. Our model is developed under the framework of conditional variational autoencoder (CVAE). We evaluate our approach on an open domain chat corpus and show improvement over strong baselines in quantitative measures and human evaluation.
[NLP-118] Explainable Suicide Risk Assessment on Social Media with Multi-Task QLoRA
【速读】: 该论文旨在解决社交媒体上可解释的自杀风险评估问题,即在准确预测用户自杀风险等级的同时,识别出支持性语言片段及其中体现的风险与保护因素。其核心挑战在于如何在一个统一的语言模型框架下,同时实现多任务协同建模与任务特异性优化。解决方案的关键在于:采用量化低秩适应(QLoRA)对Qwen2.5-Instruct模型进行高效微调,并设计答案掩码的因果语言建模目标以增强生成可解释性的能力;针对不同任务分别采用联合训练或独立训练策略——风险等级分类与证据短语提取采用多任务联合训练,而因子识别则单独适配;此外,针对各任务输出特性定制化聚合机制:通过平均大模型(32B与72B)的概率提升风险分类性能,利用交叉验证共识整合证据短语,基于留出样本的操作点进行率匹配校准因子决策。实验表明,任务特定的训练与聚合策略显著提升了整体表现,在官方排行榜上取得0.7738的综合得分,验证了在统一模型架构中根据任务输出结构差异化设计训练目标与融合策略的有效性。
链接: https://arxiv.org/abs/2610.00610
作者: Xuan Zhong Feng,Geoffrey Martin,Hexin Dong,Yifan Peng
机构: Weill Cornell Medicine(威尔·康奈尔医学院); Cornell University(康奈尔大学)
类目: Computation and Language (cs.CL); Machine Learning (cs.LG)
备注: Accepted at IEEE BigData 2026
Abstract:Explainable suicide-risk assessment requires models not only to estimate risk severity, but also to identify supporting language and the risk and protective factors expressed in a post. We present our system for the IEEE BigData 2026 Cup on Explainable Suicide Risk Assessment on Social Media, which addresses three tasks: risk-level classification, evidence phrase extraction, and multi-label factor identification. Our approach adapts Qwen2.5-Instruct models using quantized low-rank adaptation (QLoRA) and an answer-masked causal language-model objective. We jointly train across all three tasks for risk classification, jointly train on Tasks~1a and 1b for evidence extraction, and adapt Task~2 separately for factor identification. We also tailor aggregation to each output: we average risk-level probabilities from the 32B and 72B models, combine evidence phrases through cross-fold consensus, and calibrate factor-specific decisions through rate matching based on out-of-fold operating points. On the official leaderboard, the final system achieved a composite score of 0.7738, with 0.8089 on Task~1 and 0.6919 on Task~2. Across the evaluated configurations, three-task training performed best for Task~1a, joint training on Tasks~1a and 1b performed best for Task~1b, and task-specific training performed best for Task~2. Probability averaging further improved Task~1a when component models had complementary errors. These findings highlight the value of tailoring both training objectives and aggregation strategies to the output structure of each task within a unified language-model framework.
[NLP-119] Legal Research Bench: Measuring End-to-End Reliability in Long-Horizon Legal Research Agents
【速读】: 该论文旨在解决法律研究中自动化工具的可靠性问题,即如何在生成式 AI(Generative AI)辅助下实现高效且准确的美国法律检索与推理。法律研究的核心挑战在于需精准识别具有约束力的权威判例、验证其时效性、协调成文法与判例之间的冲突,并在此基础上合成有依据的结论,而现有语言模型代理在这一高要求的检索密集型流程中仍存在严重可靠性缺陷。解决方案的关键在于构建一个高质量基准测试平台——法律研究基准(Legal Research Bench, LRB),包含由专家设计的413个开放性美国法律问题,每个问题均配有黄金答案、支持性权威文献及二元评分标准。通过引入“全通过”评分机制(all-pass grading)并结合源证验证(source verification),确保回答仅在所有必要条件满足且引用权威可验证时才被判定为正确。实验表明,尽管采用网页搜索、判例库检索、页面解析等工具增强能力,当前最先进模型(Claude Opus 4.8)在全通过率上仅达42.9%,且性能随法律领域和冲突权威调和任务显著下降;更复杂的交互轮次、工具调用或更高的推理成本并未带来准确性的提升,揭示出当前生成式 AI 在法律推理中的根本局限性。
链接: https://arxiv.org/abs/2610.00609
作者: Katrina Drozdov,Oliver Chen,Langston Nashold,Rayan Krishnan
机构: Vals AI; San Francisco, USA
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computers and Society (cs.CY)
备注:
Abstract:Legal research is a core and time-consuming legal workflow. Lawyers must identify controlling authority, verify that it remains valid, reconcile statutes and cases, and synthesize a grounded answer. Language model agents are a natural fit for this retrieval-intensive workflow, and automating even part of it would be valuable. But that value depends on reliability: a single missing authority, stale citation, or wrong legal conclusion can make an otherwise plausible answer unusable. We introduce \textbfLegal Research Bench (LRB), a benchmark of 413 open-ended U.S. legal research questions written by experts, each paired with a gold answer, supporting authorities, and a binary grading rubric. We evaluate thirteen frontier models in a harness with web search, case-law search, page parsing, and retrieval tools. We score agent responses through all-pass grading with source verification, where a response is correct only if every required criterion is satisfied and its cited authorities verify. We also validate the LLM judge against expert attorneys ensuring that benchmark scores track attorney judgment. Agents remain far from reliable: among the models we tested, the strongest, Claude Opus 4.8, is fully correct on 42.9% of questions. Performance also varies substantially by task setting: all-pass rates differ across areas of law and are lower on questions requiring reconciliation of conflicting authorities. Across models, more turns, tool calls, and inference cost do not predict higher accuracy.
[NLP-120] Wheres Waldo? Query-language Preference under Cross-lingual Knowledge Disparities
【速读】: 该论文旨在解决大语言模型在跨语言问答任务中因查询语言偏好(query-language preference)而导致的信息选择偏差问题,尤其在多语言知识存在不完整或冲突的现实场景下,这种偏差会显著影响用户获取的信息质量与一致性。其核心问题是:当不同语言版本的知识库对同一事实存在缺失或矛盾时,模型倾向于优先采纳与查询语言一致的文档内容,从而导致语义等价的查询在不同语言下产生不一致的回答。解决方案的关键在于识别并缓解这一偏好行为——研究提出通过两种策略进行干预:一是基于机制的注意力头消融(ablates attention heads associated with query-language preference),以移除模型中与语言偏好相关的内部表征;二是采用LoRA(Low-Rank Adaptation)微调方法,通过对模型参数进行轻量级适配,有效降低语言偏好差距,实验显示可将偏好差异减少高达61.5%。该研究通过构建包含知识缺口与知识冲突的多语言问答基准Waldo,系统揭示了模型在真实异构知识环境下的偏见表现,并为提升跨语言信息检索的公平性与可靠性提供了可操作的技术路径。
链接: https://arxiv.org/abs/2610.00606
作者: Dayeon Ki,Ruochen Zhang,Silviu Cucerzan,Ryen W. White,Ning Gao
机构: University of Maryland (马里兰大学); Microsoft(微软)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: 43 pages, 6 figures
Abstract:Large Language Models increasingly serve as interfaces for knowledge-intensive information seeking tasks across languages by synthesizing multilingual evidence. Prior work has shown that they often exhibit query-language preference – the tendency to favor sources written in the language of the query – but has largely examined this behavior in settings where equivalent knowledge is available across languages. However, this bias becomes consequential when sources in different languages provide incomplete or inconsistent accounts of the same fact, since the information users receive then depends on the sources a model selects to use. To characterize query-language preference under such cross-lingual knowledge disparities, we introduce Waldo, a multilingual Question-Answering (QA) benchmark constructed from Wikipedia. Waldo contains 12K QA pairs targeting knowledge gaps, where a fact is available in one language but absent in another, and knowledge conflicts, where language editions provide conflicting versions of the same fact. Evaluating eight models across five languages, we find that when one language edition merely lacks the relevant fact, models generally use evidence from the other language regardless of the query language. Under conflicting accounts, however, model responses strongly align with the document in the query language, causing semantically equivalent queries to elicit different accounts depending on the user’s language. Finally, we explore two different approaches that could mitigate this preference under knowledge conflicts: a mechanistic intervention that ablates attention heads associated with query-language preference, and LoRA-based training, which reduces the preference gap by up to 61.5%.
[NLP-121] Make Sparse Rewards Count: Density-Aware Reward Aggregation for Multi-Reward RL
【速读】: 该论文旨在解决多奖励强化学习(multi-reward reinforcement learning)中多个行为目标在训练过程中学习进度不均衡的问题。尽管如GDPO(Generalized Discounted Policy Optimization)采用奖励归一化方法保留了每组回滚(rollout group)内的相对奖励信息,但不同目标仍可能出现学习速率差异显著的情况。其核心问题在于:尽管局部奖励相对性得以保持,但在批次层面仍存在信号不平衡现象,即某些奖励在多数回滚组中贡献的相对优势为零,导致其更新信号弱化。作者通过引入“优势能量”(advantage energy,即批内某奖励优势平方和)进行分析,发现该能量与活跃组密度(active-group density,即提供非零相对优势的回滚组占比)呈正比,揭示了残余的批次级信号不平衡机制。基于此,论文提出密度感知奖励聚合(Density-Aware Reward Aggregation, DARA),设计了一种反平方根密度校正权重机制,使较少活跃的奖励获得更高权重,从而动态补偿其信号缺失。DARA无需修改原始策略优化目标,可自适应地根据每批次的奖励活跃度调整权重。实验结果表明,在工具调用和数学推理任务上,DARA相较GDPO显著加速收敛,工具调用任务中实现高格式合规性所需训练步数减少最多达26%,数学推理任务中接近饱和长度合规性所需步数减少最多达65%,同时最终性能保持相当。
链接: https://arxiv.org/abs/2610.00574
作者: Tong Zheng,Skylar Zhai,Zhan Cheng,TianMing Sha,Youling Huang,Shuo Zhou,Shaotong Qi,Jingcheng Liang,Xuwei Ding,Pengcheng Xu
机构: University of Chinese Academy of Sciences(中国科学院大学); University of Minnesota Twin Cities(明尼苏达大学双城分校); University of Wisconsin–Madison(威斯康星大学麦迪逊分校); Stony Brook University(石溪大学); Dalian University of Technology(大连理工大学); Beijing Foreign Studies University(北京外国语大学); Southeast University(东南大学); Kuaishou Technology(快手科技)
类目: Machine Learning (cs.LG); Computation and Language (cs.CL)
备注: 24 pages, 8 figures
Abstract:Multi-reward reinforcement learning trains large language models to satisfy multiple behavioral objectives simultaneously. Reward-wise normalization, as used in GDPO, preserves reward-specific relative information within rollout groups, but different objectives can still exhibit uneven learning progress. We study this behavior through advantage energy, the sum of a reward’s squared advantages over a batch. Under idealized GDPO normalization, we show that this energy is proportional to active-group density: the fraction of rollout groups in which the reward provides nonzero relative advantages. This reveals a residual batch-level signal imbalance and provides a basis for calibrating reward contributions. Based on this relation, we propose Density-Aware Reward Aggregation (DARA). We derive an inverse-square-root density correction that gives greater weight to signals from less frequently active rewards. DARA computes its weights from each rollout batch, adapting to changes in reward activity throughout training without modifying the underlying policy optimization objective. Experiments on tool calling and mathematical reasoning show that DARA learns the targeted behaviors faster than GDPO, reaching high format compliance in up to 26% fewer training steps on tool calling and near-saturated length compliance in up to 65% fewer steps on mathematical reasoning, while remaining competitive in final performance. Our code is available at this https URL.
[NLP-122] Emergent Unfaithfulness: How Alignment Training Causes Language Models to Silently Override Task Faithfulness
【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)中存在的对齐性(alignment)与忠实性(faithfulness)之间的冲突问题,即当模型在对齐人类偏好时,会系统性地偏离原始输入中的敏感或不安全内容,且不披露这种修改,从而导致“对齐引发的不忠实”(alignment-induced unfaithfulness, AIU)。这一现象不同于由知识或推理错误引起的、以能力驱动的不忠实,而是由后训练阶段(post-training)机制主动覆盖输入真实性所引发。论文的关键解决方案在于构建了名为FaithConflict的受控数据集,用于分离并量化对齐-忠实性冲突与能力-对齐、能力-忠实性之间的权衡,并提出两种互补的分类体系:行为层面(B1-B8)与思维链推理层面(C0-C6),以系统分析该问题。研究发现,AIU随模型规模增长而加剧,呈现出反向缩放规律(reverse scaling law),且在后训练阶段(尤其是直接偏好优化,DPO)中被显著放大,同时其偏差最不易察觉。此外,基于提示(prompting-based)的缓解策略无法有效解决此问题,揭示出大语言模型设计与评估中存在能力-对齐-忠实性三难困境(trilemma)。
链接: https://arxiv.org/abs/2610.00568
作者: Pardis Sadat Zahraei,Janvijay Singh,Gokhan Tur,Dilek Hakkani-Tur
机构: University of Illinois Urbana-Champaign(伊利诺伊大学厄本那-香槟分校)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: Accepted at COLM 2026
Abstract:Large language models are characterized by three key properties: capability, alignment, and faithfulness. Prior work studies the tradeoffs between capability and alignment, and between capability and faithfulness, but a third tension remains underexplored: the alignment-faithfulness conflict. We show that aligned models systematically deviate from their inputs on unsafe or sensitive content without disclosing the modification, a failure mode we call alignment-induced unfaithfulness (AIU). Unlike capability-driven unfaithfulness, which comes from errors in knowledge or reasoning, this is induced by post-training mechanisms that override adherence to the input. We introduce FaithConflict, a controlled dataset isolating both conflicts, and two complementary taxonomies: behavioral (B1-B8) and chain-of-thought reasoning (C0-C6). Across models, AIU increases with scale and more sharply than capability-driven unfaithfulness, a reverse scaling law; intermediate checkpoints show it is amplified during post-training, with DPO the stage at which the gap both grows most and becomes least visible. Prompting-based mitigation does not resolve it, revealing a capability-alignment-faithfulness trilemma in the design and evaluation of LLMs.
[NLP-123] Can LLM s Reason Over Long Horizons? An Empirical Evaluation of Context Strategies for Longitudinal Clinical Reasoning
【速读】: 该论文旨在解决生成式AI在纵向临床推理中如何有效识别并整合分散于长时病史中的相关证据这一关键问题。尽管大语言模型(LLM)具备处理长上下文的能力,但单纯增加历史信息量并不能保证相关证据的可及性或推理性能的提升。研究对比了五种上下文策略(全量、近期、事件片段、语义聚合和混合)在MedLoCoMo数据集上的表现,发现事件片段(Episodic)与混合(Hybrid)策略在整体准确率上最优,且在支持证据距离较远时仍能保持较高性能;相比之下,近期上下文策略随证据距离增加而显著退化。进一步的对抗性测试表明,模型在可回答问题上的良好表现并不意味着其能有效识别不支持前提的问题并选择不回答(abstention)。因此,可靠的纵向推理不仅依赖于模型可访问的历史长度,更关键在于如何精准筛选并呈现与推理任务相关的证据。
链接: https://arxiv.org/abs/2610.00562
作者: Taye Akinrele,Noorbakhsh Amiri Golilarz,Subash Neupane,Sudip Mittal,Shahram Rahimi
机构: The University of Alabama(阿拉巴马大学)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:
Abstract:Longitudinal clinical reasoning requires large language models (LLMs) to identify and integrate relevant evidence distributed across extended patient histories. Although long-context models can process increasingly large amounts of information, providing more history does not necessarily make relevant evidence more accessible or improve reasoning. We compare five context strategies (Full, Recent, Episodic, Semantic, and Hybrid) on MedLoCoMo across four open-weight LLMs, examining answer correctness, robustness to query-evidence distance, and abstention on questions with unsupported premises. Episodic and Hybrid generally achieve the strongest overall accuracy, while Recent Context degrades most as supporting evidence becomes more distant; Episodic and Hybrid maintain the highest accuracy at long distances. Analysis of adversarial questions further shows that strong performance on answerable questions does not necessarily translate to successful abstention when the available history does not support the requested conclusion. These findings show that reliable longitudinal reasoning depends not only on how much history an LLM can access, but critically on how relevant evidence is selected and presented for reasoning.
[NLP-124] Assessing the Impact of Language Disparity on Multilingual Linguistic Ability in Large Language Models EMNLP
【速读】: 该论文旨在解决多语言语言模型在语法能力评估中因评价范式、后训练策略及语言资源可用性差异而导致结果不一致的问题。其核心挑战在于,现有评估方法未能系统考察这些因素之间的交互作用,从而可能系统性低估低资源语言的语法能力。解决方案的关键在于提出一种语言感知的多范式评估协议:首先,通过对比基础模型与后训练模型在101种语言上的表现,发现后训练虽普遍降低语法能力,但大模型可部分缓解这一退化效应,而低资源语言承受最大损失;其次,揭示后训练模型虽无法通过显式提示表达语法知识,但其隐含知识仍可通过特定方法检测,前提是高资源语言具备足够高的基线性能以区分知识存在与否;最后,提出利用母语提示(native-language prompting)可有效恢复低资源语言中被隐藏的语法能力,证明仅在高资源语言中可通过无提示概率直接探测语法知识。因此,该研究强调必须采用融合语言特异性与多评估范式的综合协议,才能避免对低资源语言能力的系统性低估。
链接: https://arxiv.org/abs/2610.00540
作者: Zhanyu Chen,Jaap Jumelet
机构: University of Groningen(格罗宁根大学)
类目: Computation and Language (cs.CL)
备注: EMNLP Main 2026
Abstract:Claims about the grammatical competence of multilingual language models vary sharply with how competence is measured, yet the interaction between evaluation paradigm, post-training, and language resource availability has not been systematically examined. We evaluate base and post-trained models from six families on MultiBLiMP, a syntactic minimal-pair benchmark covering 101 languages, using four evaluation methods. We report three principal findings. First, post-training degrades grammatical competence, but the magnitude of this effect is reduced unevenly by model scale, while low-resource languages bear the highest cost. Second, post-trained models retain grammatical knowledge they cannot articulate through explicit prompting, yet this is measurable only in high-resource languages, because near-chance baselines in low-resource settings leave little knowledge to hide. Third, native-language prompting recovers otherwise hidden competence on low-resource languages, demonstrating that only high-resource languages can be probed directly from unprompted probabilities. We conclude that multilingual grammatical evaluation must adopt language-informed, multi-paradigm protocols to avoid systematically underestimating low-resource abilities.
[NLP-125] Rules Amortize Pairings Dont: Linguistic Structure Determines What Latent Task Representations Can Replace In-Context Learning NEURIPS2026
【速读】: 该论文旨在解决生成式模型在少样本学习(few-shot learning, FSL)中依赖大量示例提示(prompting)所带来的效率瓶颈问题,尤其关注哪些语言学操作能够被“压缩”至上下文内学习(in-context learning, ICL)之外,从而实现零样本推理成本下的高效泛化。其核心挑战在于:传统静态任务向量(task vector)作为单一合成示例,无法有效建模高秩映射(如词级双射),导致在复杂语言规则(如形态变换、词义关系)上表现受限。为此,作者提出一种新型解决方案——基于支持集几何结构(包括中心点、主子空间、谱特征等)的动态可学习变换机制,通过一个仅260万参数的小型网络,在冻结的GPT-2-large/XL模型中间层对查询残差流进行输入条件化的加性更新。该方法在八个屈折方向和一个词汇关系任务上验证,揭示三种行为模式:在正向屈折任务中,该变换可完全复现10样本ICL性能(0.67–0.89),且实现严格零样本推理开销;在词干化任务中,由于冻结GPT-2本身具备执行能力但10个示例无法有效传递信息(ICL仅0.13–0.48),该方法性能显著超越ICL达+72个百分点(最高达0.92),表明其突破了原始提示的表达瓶颈;而在任意配对(如反义关系)任务中,所有优化器均在约50%的ICL性能水平处饱和,暗示此类记忆性配对难以被有效抽象。控制实验进一步证实,支持集流形是因果必需的任务指纹,错误任务流形导致性能骤降至0.06,仅依赖查询的变体无法区分共享输入空间的不同任务,且留一任务移除实验显示无迁移能力。综上,该研究的关键突破在于:将可泛化、有创造性的语言规则(如屈折与词干化)转化为可缓存、可动态调用的隐式任务表示,而记忆性配对则无法被有效抽象,从而实现了对部分语言操作的真正“消解式”压缩,使模型在无需额外示例的情况下完成高质量推理。
链接: https://arxiv.org/abs/2610.00526
作者: Gunmay Jhingran
机构: Delhi Technological University (德里技术大学)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: Accepted to the NeurIPS 2026 Workshop on Linguistic Principles for Foundation Models (LP4FM). 5 pages
Abstract:In-context learning (ICL) can be amortized into latent objects (task vectors, function vectors, context vectors) that recover few-shot behavior at zero-shot inference cost, but recent theory shows a static vector acts as a single synthetic demonstration and must fail on high-rank mappings such as word-level bijections. We ask a linguistic version of this question: which linguistic operations can be amortized out of the prompt? We train a 2.6M-parameter network that reads the geometry of a few-shot support set (centroid, principal subspace, spectrum, computed once and cached) and produces an input-conditioned additive update to the query’s residual stream at a mid-depth layer of a frozen GPT-2-large/XL. Across eight inflectional directions and one lexical relation, under a canonical split that bars inverted-pair leakage between directions, three regimes emerge. On forward inflection, where 10-shot ICL is strong (0.67-0.89) and extracted task vectors collapse (=0.06), the transform matches ICL at strictly zero-shot per-query cost. On lemmatization directions, which frozen GPT-2 can execute but 10 demonstrations systematically fail to convey (ICL 0.13-0.48 at 1.5B), the transform is not capped by ICL at all: it reaches 0.78-0.92, up to +72 points over ICL (past to present: 0.85 vs. 0.13). On arbitrary pairings (antonymy) every amortizer plateaus near half of ICL at every scale, capacity, and seed tested. Controls show the support manifold acts as a causally necessary task fingerprint: wrong-task manifolds collapse accuracy to =0.06, query-only variants cannot disambiguate tasks sharing an input space, and leave-one-task-out transfer is zero. Productive rules amortize into latent task representations, sometimes better than prompting can convey them; memorized pairings do not.
[NLP-126] EurekaBench: Measuring Agent ic Ability to Discover New Scientific Insights
【速读】: 该论文旨在解决当前人工智能(AI)代理在科学发现能力方面存在的局限性,即尽管其在预测精度上已超越人类科学家,但在理解复杂现象背后的机制、生成可解释的科学洞见方面仍表现不足。其核心问题在于:如何评估并提升AI在长时程实验中自主发现科学规律的能力,使其不仅能够拟合数据,更能像牛顿那样揭示支配自然现象的根本原理。解决方案的关键在于提出EurekaBench——一个跨学科基准测试平台,涵盖神经科学、计算机科学、化学、天体物理学、地球物理学和等离子体物理等领域的26个长期任务,共包含306项需被发现的科学洞见。该框架从三个维度评估科学发现能力:是否遵循已知科学约束、所发现机制的预测准确性,以及能否产生新的科学洞见或推动未来研究。研究表明,现有AI代理普遍过度聚焦于预测性能优化,而在生成深刻科学理解方面存在显著短板,凸显了构建具备因果推理与理论生成能力的下一代AI系统的重要性。
链接: https://arxiv.org/abs/2610.00492
作者: Jiayi Geng,Zhengxuan Wu,Kevin S. Chen,Seungone Kim,Joseph Janssen,Zora Zhiruo Wang,Bhupalee Kalita,Runtian Gao,Aaron Ho,Andrew Oakleigh Nelson,Olexandr Isayev,Francisco Villaescusa-Navarro,Ching-Yao Lai,Howard Chen,Graham Neubig
机构: Carnegie Mellon University (卡内基梅隆大学); Stanford University (斯坦福大学); Yale University (耶鲁大学); Massachusetts Institute of Technology (麻省理工学院); Columbia University (哥伦比亚大学); Princeton University (普林斯顿大学); Flatiron Institute (Flatiron研究所); Google DeepMind (谷歌深脑); Engram (恩格拉姆)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注:
Abstract:When Isaac Newton discovered the law of gravitation, he did so through an iterative process of analyzing observed data such as planetary patterns, finding the underlying mechanisms by describing patterns in mathematical equations, and refining his theory against the Moon’s orbit, revealing the startling insight that the same force governs both falling apples and orbiting planets. Would it be possible for AI agents to make similar discoveries? To measure this ability, we introduce EurekaBench, a cross-domain benchmark that tests AI agents’ ability to conduct long-horizon experiments and discover mechanisms that explain observations. We evaluate these mechanisms by the scientific insights that can be derived from them. EurekaBench contains an expert-verified set of 26 long-horizon tasks across neuroscience, computer science, chemistry, astrophysics, geophysics, and plasma physics, with a total of 306 scientific insights that the discovered mechanisms are expected to support. Our evaluation framework tests three axes of scientific discovery: agents’ ability to follow known scientific constraints, the predictive accuracy of the discovered mechanisms, and whether these mechanisms yield scientific insights or inform future research. Our results show that current AI agents often overly fixate on predictive accuracy optimization, surpassing human scientists, while falling substantially short in deriving scientific insights.
[NLP-127] Frozen Scenes Shifting Winners: Configuration Frag ility in Text-to-3D Evaluation
【速读】: 该论文旨在解决生成式三维内容(text-to-3D)评估中存在的一种核心问题:当生成的场景保持固定时,评估排行榜是否仍可能因评价协议的细微变化而发生改变。研究聚焦于基于渲染图像的评估范式,指出相机设置与文本提示(prompt)表述等变量已成为测量协议的一部分,从而影响评估结果的可靠性。其解决方案的关键在于系统性地审计评估流程的敏感性,通过在6个生成器产生的300个固定场景上,系统性地调整8种渲染与提示因素,结合19个对齐度评估器及1个感知质量控制项,测试四种针对性的场景退化。结果显示,对于17/19个评估器,配置变化引起的评分波动超过不同生成器之间的差异;且11/19个评估器的提示引导下界高于1,表明提示词对结果有显著影响。尽管排名相对稳定,但仍有18/19个评估器在特定配置下改变了其最优生成器的点估计结果。通过构建成对协议边际包络(pairwise protocol margin envelopes),研究揭示了哪些比较方向在测试条件下保持一致,发现部分配对呈现相反的点估计区间,但无一例在全搜索空间同时推断下发生反转,因此观测到的优劣变化属于描述性差异,而非生成器真实性能的确认性提升。此外,评估器在布局打乱任务中的定向判别能力最高仅达67%(经平局修正),表明当前评估体系缺乏人类验证的基准真实性,仅具备诊断性价值。综上,该研究分离了评分稳定性、决策不确定性与目标敏感性三类属性,并建议未来报告应包含(生成器、评分、卡片ID)信息,辅以协议依赖的对比分析和选择意识的不确定性标注。
链接: https://arxiv.org/abs/2610.00447
作者: Anson Y. Lam,Shuqing Li,Michael R. Lyu
机构: The Chinese University of Hong Kong (香港中文大学)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR); Multimedia (cs.MM)
备注: 26 pages, 6 figures
Abstract:Can a text-to-3D leaderboard change when every generated scene stays fixed? We audit this question for rendered-image evaluation, where camera settings and caption wording become part of the measurement protocol. Across 300 frozen scenes from six generators, we vary eight render and caption factors for 19 alignment evaluators plus one perceptual-quality control, then test four targeted scene degradations. Peak configuration variance exceeds between-generator variance for 17/19 alignment evaluators, with prompt-bootstrap lower bounds above 1 for 11/19. Rankings are more stable than scores, yet 18/19 evaluators change their point-estimate winner under some configuration. Pairwise protocol margin envelopes show which comparisons keep their direction across the tested settings. Selected pairs have opposite pointwise intervals, but no reversal survives simultaneous inference over the full search. Thus the observed winner changes are descriptive, not confirmed changes in generator superiority. Sensitivity remains separate: no evaluator, even the prompt-free control, exceeds 67% tie-adjusted directional discrimination on layout scrambling, which is diagnostic rather than human-validated ground truth. The audit separates score stability, decision uncertainty, and targeted sensitivity, and recommends reporting (generator, score, card ID) with protocol-dependent comparisons and selection-aware uncertainty.
[NLP-128] UniBuc at SemEval-2024 Task 2: Tailored Prompting with Solar for Clinical NLI
【速读】: 该论文旨在解决医学领域临床试验报告(Clinical Trial Report, CTR)中安全的生物医学自然语言推理(Safe Biomedical Natural Language Inference, SB-NLI)问题,核心挑战在于如何在不引入偏见或错误推断的前提下,准确判断文本之间的逻辑关系。其解决方案的关键在于采用未经过微调的SOLAR Instruct模型,通过精细化的输入构造与定制化提示工程(prompt engineering),针对CTR的不同章节分别设计提示策略,在零样本(zero-shot)与少样本(few-shot)设置下实现高效推理。实验结果表明,该方法在保持模型通用性的同时,显著提升了推理一致性,最终获得0.72的一致性得分,位列SemEval 2024 Task 2排行榜第14名。进一步的错误分析揭示,模型存在依赖简单启发式规则和“捷径学习”(shortcut learning)的问题,尤其在面对语义保持但表达形式变化的文本时表现不稳定,凸显了当前生成式AI在复杂医学文本理解中的局限性。
链接: https://arxiv.org/abs/2610.00408
作者: Marius Micluta-Campeanu,Claudiu Creanga,Ana-Maria Bucur,Ana Sabina Uban,Liviu P. Dinu
机构: University of Bucharest (布加勒斯特大学); HLT Research Center (HLT研究中心); Interdisciplinary School of Doctoral Studies (跨学科博士研究院); Faculty of Mathematics and Computer Science (数学与计算机科学学院)
类目: Computation and Language (cs.CL)
备注:
Abstract:This paper describes the approach of the UniBuc team in tackling the SemEval 2024 Task 2: Safe Biomedical Natural Language Inference for Clinical Trials. We used SOLAR Instruct, without any fine-tuning, while focusing on input manipulation and tailored prompting. By customizing prompts for individual CTR sections, in both zero-shot and few-shots settings, we managed to achieve a consistency score of 0.72, ranking 14th in the leaderboard. Our thorough error analysis revealed that our model has a tendency to take shortcuts and rely on simple heuristics, especially when dealing with semantic-preserving changes.
[NLP-129] LLM -as-a-Judge for Low-Resource Languages: Adapting Rag as and Comparative Ranking for Romanian
【速读】: 该论文旨在解决低资源语言(Low-Resource Languages, LRLs)中检索增强生成(Retrieval-Augmented Generation, RAG)系统评估困难的问题,尤其针对罗马尼亚语这一典型低资源语言。现有基于参考文本的评估指标在缺乏高质量人工标注基准的情况下表现不佳,难以有效衡量生成质量。为此,论文提出采用“大语言模型作为裁判”(LLM-as-a-Judge)范式,并基于新一代模型(Gemini 2.5 和 Gemini 3)对Ragas框架进行适配。其核心解决方案是构建了AdminRo-Eval——一个由母语者标注的罗马尼亚行政文档精选数据集,作为自动化评估器的基准真值。研究对比了直接评分、比较排序和细粒度分解三种评估方法在忠实性(Faithfulness)、答案相关性(Answer Relevance)和上下文相关性(Context Relevance)三个维度上的表现。结果表明,评估策略需根据具体指标进行优化:细粒度分解在忠实性评估上达到最高人类一致性(使用Gemini 2.5 Pro时达96%),而比较排序在答案相关性上表现更优(达90%)。此外,研究还发现轻量级模型在低资源语言中难以胜任复杂推理任务,但Gemini 2.5 Pro架构展现出强大的泛化能力,可为罗马尼亚语RAG系统的自动化评估建立稳健且可迁移的基线。
链接: https://arxiv.org/abs/2610.00406
作者: Claudiu Creanga,Liviu P. Dinu
机构: Interdisciplinary School of Doctoral Studies; HLT Research Center, University of Bucharest (布加勒斯特大学高级语言技术研究中心), Romania
类目: Computation and Language (cs.CL)
备注:
Abstract:Evaluating Retrieval-Augmented Generation (RAG) systems remains a challenge for Low-Resource Languages (LRLs), where standard reference-based metrics fall short. This paper investigates the viability of the “LLM-as-a-Judge” paradigm for Romanian by adapting the Ragas framework using next-generation models (Gemini 2.5 and Gemini 3). We introduce AdminRo-Eval, a curated dataset of Romanian administrative documents annotated by native speakers, to serve as a ground truth for benchmarking automated evaluators. We compare three evaluation methodologies - direct scoring, comparative ranking, and granular decomposition - across metrics for Faithfulness, Answer Relevance, and Context Relevance. Our findings reveal that evaluation strategies must be metric-specific: granular decomposition achieves the highest human alignment for Faithfulness (96% with Gemini 2.5 Pro), while comparative ranking outperforms in Answer Relevance (90%). Furthermore, we demonstrate that while lightweight models struggle with complex reasoning in LRLs, the Gemini 2.5 Pro architecture establishes a robust, transferable baseline for automated Romanian RAG evaluation.
[NLP-130] Dissonant ballerinas and crafty carrots: a comparative multi-modal analysis of Italian brain rot
【速读】: 该论文旨在解决跨语言语境下“脑腐”(brain rot)类短视频内容吸引力机制的差异问题,重点关注意大利语与罗马尼亚语版本在语言、文化及多模态特征上的异同。其核心解决方案在于构建了一个名为CRIB(Collection of Romanian and Italian Brain rot)的多模态数据集,该数据集包含240个经人工标注的TikTok视频,按语言(意大利语、罗马尼亚语)和流行度分层,系统分析了文本、音频与视觉层面的特征。研究发现,视频的流行度与文本层面的情感、荒诞性或押韵等特征无显著关联,也与声音中的语音特征或情感表达无关;而在罗马尼亚语版本中,视频级动态特征——尤其是更快的剪辑速度与整体节奏——是预测成功的关键因素。跨语言对比揭示出显著差异:意大利语脑腐内容在文本上更具负面性,具有更高的困惑度并更频繁使用押韵,其音频则表现出更大的旋律范围与音量;而罗马尼亚语音频则呈现更明亮的频谱特性及更不规则的音高变化。
链接: https://arxiv.org/abs/2610.00402
作者: Anca Dinu,Andra-Maria Florescu,Marius Micluta-Campeanu,Stefana-Arina Tabusca,Claudiu Creanga,Andreiana Mihail
机构: 未知
类目: Computation and Language (cs.CL)
备注:
Abstract:This paper presents a comparative multi-modal analysis of Italian and Romanian brain rot memes, investigating the factors that contribute to its appeal and the linguistic and cultural distinctions between the two versions. To conduct this analysis, we introduce a multi-modal brain rot dataset named CRIB (Collection of Romanian and Italian Brain rot), a manually curated collection of 240 TikTok videos stratified by language (Italian, Romanian) and popularity, on which we examine textual, acoustic, and visual features. Our findings indicate that popularity is not significantly correlated with textual elements like sentiment, absurdity, or rhyme, or acoustic elements such as vocal features or sentiment of the sound. Instead, in Romanian language, video-level dynamics, specifically faster cutting speeds and a more rapid overall pace, are strong predictors of a video’s success. The cross-linguistic analysis reveals significant differences. Italian brain rot is textually more negative, exhibits higher perplexity, and uses more rhyme, while its sound is characterized by higher melodic range and loudness. Romanian audio is spectrally brighter with more erratic pitch variations.
[NLP-131] FAER: Auditable Utility-Aligned Trajectory Replay for Language Model Post-Training
【速读】: 该论文旨在解决生成式 AI(Generative AI)训练中回放选择器(replay selector)存在的“选择-学习差距”问题,即传统回放策略(如基于格式反馈、置信度、新鲜度或响应长度的排序)往往忽视了缓存轨迹在下游学习任务中的实际效用,导致选择目标与学习目标不一致。其核心解决方案是提出一种可审计的全轨迹回放框架 FAER,关键在于引入两个创新机制:一是无需训练的固定选择器作为协议基线,二是基于独立校准块拟合的面向学习者(learner-aware)的选择器 FAER-UTILITY;同时通过归一化梯度对齐(normalized gradient alignment)提供基准,并利用可丢弃的优化器感知虚拟更新构建幅度感知的效用表面。该框架通过“审计契约”冻结观测字段与回放轨迹,确保评估前未暴露真实标签,从而实现对选择器有效性的严格验证。实验结果表明,在 GSM8K 数据集上使用 Qwen2.5-1.5B-Instruct 模型时,FAER-UTILITY 达到 0.6624 的性能,显著优于固定选择器(0.6329)和格式反馈选择器(0.6037),且通过元数据仅有的交叉拟合校准即达到 0.6476 ± 0.0139 的稳定表现,充分体现了该方案在提升回放轨迹下游学习效用方面的有效性。
链接: https://arxiv.org/abs/2610.00385
作者: Miaobo Hu,Shuhao Hu,Xiaobo Guo,Xin Wang,Bokun Wang,Tianshu Fu,Daren Zha,Jun Xiao
机构: University of Chinese Academy of Sciences (中国科学院大学); Institute of Information Engineering, Chinese Academy of Sciences (中国科学院信息工程研究所)
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (stat.ML)
备注: 35 pages, 6 figures
Abstract:Replay selectors often rank cached trajectories by format feedback, confidence, freshness, or response length, although cache-level correctness and downstream learner utility are distinct objectives. We formalize this selection-to-learning gap and introduce FAER as an auditable full-trajectory replay framework. Its training-free fixed selector is a protocol baseline; FAER-UTILITY is the learner-aware selector fitted on disjoint calibration blocks. The normalized gradient alignment is reported as a baseline, while a disposable optimizer-aware virtual update supplies a magnitude-aware utility surface. The audit contract freezes observed fields and replay traces before evaluation labels are joined. On GSM8K with Qwen2.5-1.5B-Instruct, the matched learner study reports quality 0.6329 for the fixed selector, compared with 0.5482 for uniform and 0.6037 for format-feedback under 128 updates. Metadata-only cross-fitted calibration reaches 0.6476!\pm!0.0139 over eight seeds (median 0.6481; paired 95% interval [+0.079,+0.122] ) at 63,276 target-run tokens; its recorded full cost is 189,642 tokens and 3.48 GPU-hours including calibration. The completed FAER-UTILITY row reaches 0.6624 at 62,844 target-run tokens and 4.26 GPU-hours. Format-feedback selects records with correctness 0.6953, compared with 0.3594 for the fixed selector, despite the different downstream ranking. The completed comparison surfaces report the learner-aware ablation, same-seed gap, policy-optimization rows, and strict zero-shot transfer.
[NLP-132] When Do Attention-Head Ablations Support Causal Claims? Projection-Level Confounds Floor Effects and Matched Controls
【速读】: 该论文旨在解决生成式语言模型中注意力头(attention head)因果责任推断的可靠性问题,即通过注意力头消融(attention-head ablation)方法识别对特定任务行为具有因果影响的注意力头时,其结论可能因干预方式、评估指标和控制条件不当而产生误导。其解决方案的关键在于:首先,必须正确实现干预语义——采用预投影(pre-projection)而非自然的后投影(post-projection)方式执行“置零”操作,以确保干预真正作用于注意力头的内部表示;其次,应使用非饱和的连续型评估指标(如黄金标记对数概率,gold-token log-probability),避免二分类准确率在行为表现接近地板或天花板时掩盖真实效应;最后,需引入匹配的对照组(包括随机头与层匹配头的1000次重复抽样)并结合发现集/保留集划分策略,以验证结果的稳定性与显著性。实验表明,修正后的消融方法在不同数据划分下具有高度一致的头部重要性排序(Spearman rho = 0.974),且前5名关键头显著优于所有对照分布(Monte Carlo p = 0.001),从而为因果推断提供了更稳健的依据。研究强调,单一头消融不足以支持因果结论,唯有综合考虑干预位置、度量选择与严格控制,方可实现可信的解释。
链接: https://arxiv.org/abs/2610.00373
作者: Juli Huang
机构: Stanford University (斯坦福大学)
类目: Machine Learning (cs.LG); Computation and Language (cs.CL)
备注: Code available in the accompanying repository. 2 figures
Abstract:Attention-head ablation, zeroing a head and measuring the resulting change in task performance, is a common method for inferring which components of a language model are causally responsible for a behavior. We show using GPT-2 small that this inference can be fragile unless the intervention semantics, evaluation metric, and controls are carefully validated. A natural post-projection implementation of “zeroing a head” is nearly uncorrelated with a corrected pre-projection ablation (Pearson r = 0.057) and selects a completely disjoint top-5 set of important heads. We also show that binary accuracy can hide effects at behavioral floors and near ceilings, whereas gold-token log-probability remains graded. Using a discovery/held-out split and 1,000 matched random-head and layer-matched-head control draws, the corrected per-head effect ranking is highly stable across splits (Spearman rho = 0.974), and the top-5 selected heads significantly exceed both control distributions (Monte Carlo p = 0.001). However, evidence for task specificity is not robust on GPT-2. Replication on DistilGPT2 preserves the intervention-semantic and matched-control findings. These results show that single-head ablation does not by itself justify a causal claim; defensible interpretation requires correct intervention placement, a non-saturated continuous metric, and matched held-out controls.
[NLP-133] LEGO-OPD: Factorized Teacher Composition for Multimodal On-Policy Distillation
【速读】: 该论文旨在解决多模态在线蒸馏(Multimodal On-Policy Distillation, OPD)中视觉定位(visual grounding)与语言模型强推理能力之间的权衡问题。现有基于多教师的框架虽结合了大语言模型(LLM)与视觉语言模型(VLM)教师以提供互补监督,但直接使用VLM的完整预测分布会将其视觉定位信号与自身的语言先验(language prior)耦合,导致定位信息无法独立传递;而过度增强视觉监督又可能过度强调视觉证据,损害语言推理性能。为此,本文提出LEGO-OPD,通过在广义贝叶斯框架下将语言专家(Language Expert)提供的候选词先验与定位专家(Grounding Expert)提供的视觉似然进行因子化组合,构建单一教师分布,从而实现对语言推理与视觉定位的独立控制。其关键创新在于:不直接转移VLM的完整预测分布,而是仅利用其视觉似然更新语言先验,并引入自适应校准机制,以图像诱导的预测偏移作为前缀依赖参考,动态调节视觉似然对语言先验的更新强度,避免监督不足或过强。实验结果表明,基于Qwen3模型的LEGO-OPD在多模态及纯文本推理任务上均显著优于单教师与多教师基线,且在提升初始学生模型视觉感知能力的同时,有效保持了文本推理性能。
链接: https://arxiv.org/abs/2610.00333
作者: Jaeyun Shin,Hangeol Chang,Jong Chul Ye
机构: Korea Advanced Institute of Science and Technology (KAIST)
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
备注:
Abstract:Multimodal on-policy distillation (OPD) aims to improve visual grounding while preserving the strong reasoning capabilities of language models. Recent multi-teacher approaches combine LLM and VLM teachers to provide complementary supervision. However, directly using a VLM’s full predictive distribution entangles its visual grounding signal with its own language prior, preventing the grounding information from being transferred independently. Conversely, increasing the strength of visual supervision can improve perception but may overemphasize visual evidence and degrade language reasoning. To address this trade-off, we introduce LEGO-OPD, which selectively composes factors from a Language Expert and a Grounding expert into One teacher distribution for multimodal OPD. Under a generalized Bayesian formulation, the language expert provides a prior over candidate tokens, while the grounding expert contributes a visual likelihood that updates this prior, rather than transferring its complete predictive distribution. This factorized composition allows language reasoning and visual grounding to be controlled independently. We further introduce adaptive calibration to determine how strongly the visual likelihood should update the language prior at each decoding prefix. Specifically, LEGO-OPD uses the grounding expert’s image-induced prediction shift as a prefix-dependent reference, preventing both insufficient and excessive visual supervision. Experiments with Qwen3 models show that LEGO-OPD consistently outperforms the evaluated single- and multi-teacher OPD baselines on both multimodal and text-only reasoning tasks. Moreover, it improves the initial student’s visual perception while preserving text-only reasoning.
[NLP-134] ContractRL: Shielded Group-Relative Policy Optimization for Auditable Tool-Call Repair
【速读】: 该论文旨在解决结构化工具调用(structured tool calls)在少量字段违反模式(schema)或执行契约(execution contract)时,传统修复方法因完全重生成(full regeneration)导致动作空间扩大、难以审计重复修复的问题。其核心解决方案是提出ContractRL,一种受契约约束的顺序修复协议,将验证器引导的JSON修复建模为一个受限决策过程。关键创新在于:在每一步决策中,策略通过观察候选对象、类型化的验证反馈、JSON Pointer、不可变的修复历史及剩余预算,结合由契约导出的动作掩码(action mask),提前过滤非法或禁止的RFC-6902操作,再由确定性验证器执行状态转移。同时,采用契约约束的组相对目标函数,对修补、重试与放弃决策进行优化,并将规范目标与语义标签延迟至追踪冻结(trace freeze)后才引入在线状态。实验表明,在相同验证信息下,ContractRL以34.4个生成令牌实现0.9362的语义成功率,显著优于Patch-SFT(0.9076,44.9令牌)和全重生成(0.9148,137.2令牌)。策略优化进一步将语义成功率提升至0.9375,且配对评估显示相较于Patch-SFT有统计显著的正向提升(+0.0396,95% CI [+0.0137, +0.0662],p=0.0039)。反馈分析、动作掩码有效性、预算控制及模式偏移分析均表明性能提升源于局部化修正机制,而对抗性与多轮评估则揭示了当前仍存在的失败模式。
链接: https://arxiv.org/abs/2610.00328
作者: Miaobo Hu,Shuhao Hu,Xiaobo Guo,Xin Wang,Bokun Wang,Yina Sa,Daren Zha,Jun Xiao
机构: University of Chinese Academy of Sciences (中国科学院大学); Institute of Information Engineering, Chinese Academy of Sciences (中国科学院信息工程研究所)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
备注: 29 pages, 8 figures
Abstract:Structured tool calls often fail after only a small number of fields violate a schema or an execution contract. Regenerating the complete object enlarges the action surface and makes repeated repair difficult to audit. We introduce ContractRL, a contract-constrained sequential repair protocol that models verifier-guided JSON repair as a bounded decision process. At each step the policy observes the candidate, typed verifier feedback, JSON Pointer, immutable repair history, and remaining budget; a contract-derived action mask filters malformed or prohibited RFC-6902 operations before a deterministic validator performs the transition. We specify a contract-constrained group-relative objective for patch, retry, and abstention decisions while keeping canonical targets and semantic labels outside the online state until trace freeze. Under identical verifier information, ContractRL attains 0.9362 semantic success with 34.4 generated tokens, compared with 0.9076 and 44.9 tokens for Patch-SFT and 0.9148 and 137.2 tokens for full regeneration over 192 cases per seed and five seeds. Policy optimization improves semantic success from 0.9186 for supervised ContractRL to 0.9375. A separate three-seed paired evaluation against Patch-SFT yields a semantic difference of +0.0396 (95% CI [+0.0137,+0.0662], p=0.0039 ). Feedback, action-mask, budget, and schema-shift analyses connect these gains to localized correction, while adversarial and multi-turn evaluations characterize the remaining failure modes.
[NLP-135] Actions with Receipts: Jointly Binding Claims Evidence and Execution for Replayable Tool-Agent Auditing
【速读】: 该论文旨在解决生成式 AI(Generative AI)系统中工具使用代理在输出引用和执行日志时存在的关键信任漏洞:即用户所见的主张(claim)与其由确定性执行过程生成并由引用来源支持之间的一致性无法被有效审计。尽管引用本身和执行轨迹可能各自形式合法,但它们仍可能被跨主张、动作、运行实例或源版本非法移植,导致虚假信息传播。其解决方案的核心是提出一种“主张锚定的执行合约”(claim-anchored execution contract),该合约通过强绑定机制将生成的主张、其精确来源片段(source span)、生成该主张的有序执行前缀、以及执行时观察到的源版本与访问状态进行联合约束。每个凭证(receipt)包含一个确定性发射锚点(emission anchor),用于精确定位主张在最终响应中的位置,并附带源标识符、偏移量、哈希值、原文摘录及领域隔离的执行承诺。通过独立于可插拔支持平面的确定性完整性验证器,在语义或任务标签融合前重建上述绑定关系,确保结构有效性不被误用为蕴含关系的代理指标。该合约显式暴露七个可独立测试的属性:主张发射绑定、来源绑定、有序执行绑定、预言机分离、持久对象回放、执行重运行一致性、版本/访问状态绑定。在1,280次跨对象攻击测试中,联合合约检测到1,275次替换攻击(检测率0.9961),而移除任一特定属性后检测率骤降至0.0156–0.0625;在独立裁定的384对样本上,具备冲突感知的支持防护机制达到F1分数0.8865,误接受率0.0729;面对未见过的故障类型,对应指标为0.8679和0.0938,表明其具备良好的泛化鲁棒性。
链接: https://arxiv.org/abs/2610.00327
作者: Miaobo Hu,Shuhao Hu,Xiaobo Guo,Xin Wang,Bokun Wang,Yina Sa,Daren Zha,Jun Xiao
机构: University of Chinese Academy of Sciences (中国科学院大学); Institute of Information Engineering, Chinese Academy of Sciences (中国科学院信息工程研究所)
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
备注: 35 pages, 8 figures
Abstract:Tool-using agents can expose citations and execution logs while leaving a critical association unaudited: whether the claim shown to a user is the claim emitted by the committed execution and supported by the cited source. A valid citation and a valid trace can therefore remain individually well formed while being transplanted across claims, actions, runs, or source versions. We introduce a claim-anchored execution contract that jointly binds the emitted claim, its exact source span, the ordered execution prefix that produced it, and the source version and access state observed by that execution. Each receipt contains an emission anchor that deterministically locates the claim inside a committed answer or claim-bearing action, together with source identifiers, offsets, hashes, quotes, and a domain-separated execution commitment. A deterministic integrity verifier reconstructs these bindings before semantic or task labels are joined. We separate this integrity plane from a pluggable support plane, so structural validity is not used as a proxy for entailment. The contract exposes seven independently testable properties: claim-emission binding, source binding, ordered-execution binding, oracle separation, persisted-object replay, execution-rerun consistency, and version/access binding. Across 1,280 cross-object attacks, the joint contract detects 1,275 substitutions (0.9961). Removing a targeted property reduces its attack-detection rate to 0.0156-0.0625. On an independently adjudicated 384-pair split, the conflict-aware support guard reaches F1 0.8865 and false acceptance 0.0729; on unseen failure families, these rates are 0.8679 and 0.0938.
[NLP-136] CAST: Cost-Aware Speculative Trees from One-Pass Block Drafters
【速读】: 该论文旨在解决生成式 AI(Generative AI)在大语言模型推理过程中因逐词验证导致的效率瓶颈问题。传统推测解码(Speculative Decoding)虽可通过低成本草稿生成未来若干词元并并行验证,但仅能验证得分最高的候选路径,其余已计算的候选路径被浪费。其核心挑战在于如何高效利用这些已评分的候选信息,以进一步提升推理吞吐量而不引入额外的草稿开销。本文提出关键解决方案——成本感知推测树(CAST, Cost-Aware Speculative Trees),将多个候选路径组织为一棵树结构,并在一次目标模型前向传播中完成整棵树的验证,从而充分利用已有草稿计算结果。CAST通过动态调整树的宽度,基于实时延迟测量判断每新增一个候选节点所带来的预期收益是否超过其验证耗时,实现自适应宽度选择,无需遍历不同宽度配置。实验表明,在跨五个领域、三类GPU架构及两种模型家族的测试中,CAST在预测最优宽度下均优于标准链式解码,最快提速达43%;且其性能表现显著依赖部署环境,尤其在核函数边界处验证开销突增时,固定宽度(如128词元)仅提升2%,而CAST自适应宽度可实现20%的加速。此外,理论证明显示,CAST在贪婪解码与采样解码下均保持目标输出分布不变。
链接: https://arxiv.org/abs/2610.00321
作者: Jungseob Lee,Sugyeong Eo
机构: Korea University(高丽大学); Yonsei University Mirae Campus(延世大学未来校区)
类目: Computation and Language (cs.CL); Machine Learning (cs.LG)
备注: 28 pages, 7 figures, 17 tables
Abstract:Speculative decoding accelerates large language model inference by drafting future tokens cheaply and verifying them with the target model in parallel. Block drafters score a whole block of future tokens in one forward pass, yet standard decoding verifies only the top-scoring chain and discards the other candidates. Because these candidates are already scored, verifying more of them adds target computation but no extra drafting. We introduce CAST (Cost-Aware Speculative Trees), which packs these candidates into a tree and verifies it in a single target pass, leaving the target model, drafter weights, and decoding rule untouched. To decide how wide the tree should be, CAST adds candidates while the expected gain from the next one outweighs the verification time it adds. The width therefore adapts to each deployment from a latency measurement, without sweeping over widths. We evaluate CAST across five domains on three GPU generations and two model families. At its predicted width, CAST is faster than the standard chain in all eight settings, by up to 43%. We also find that the best width depends strongly on the deployment. Where verification cost jumps at a kernel boundary, a 128-token tree is only 2% faster than the standard chain, whereas the tree at the predicted width is 20% faster. Furthermore, we prove that CAST leaves the target output distribution unchanged under both greedy and sampled decoding. Code is available at this https URL.
[NLP-137] Refusal Localizes the Damage Relocates: Safety Layers Under Few-Sample Fine-Tuning
【速读】: 该论文旨在解决生成式AI模型在经过微调后,因少量有害样本注入而导致安全拒绝对话能力被破坏的问题。其核心挑战在于:尽管已有研究通过定位安全相关行为的特定层、方向和词元实现了对模型安全机制的局部修复,但这些修复是否能在面对攻击者动态调整攻击策略时仍保持有效性尚不明确。论文的关键解决方案是提出一种基于“可恢复性边界”(recovery transition depth)的局部冻结防御机制——即通过分析攻击前后模型隐藏状态的线性可分性,确定一个关键深度,在此深度以下的所有层均被冻结以防止有害更新传播。实验表明,即使在一百个有害样本攻击下,该方法仍能有效维持拒绝率接近零,并在冻结边界以上实现稳定的恢复过渡。此外,研究还发现移除梯度更新中的前两个奇异方向可修复短时注意力微调后的安全失效问题,但在对抗性扩散更新或普通训练强化条件下,该修复易被绕过。同时,基于谱特征的检测器在部分场景下无法识别修复失败。结果表明,当前防御策略存在可被绕过的脆弱性,因此论文提出了五项评估标准以增强对自适应微调攻击的鲁棒性。
链接: https://arxiv.org/abs/2610.00320
作者: Jungseob Lee,Dongyub Jude Lee,Sugyeong Eo,Seongtae Hong,Seungyoon Lee,Heuiseok Lim
机构: Korea University (韩国大学); Zoom Communications (Zoom 通信公司); Yonsei University Mirae Campus (延世大学未来校区)
类目: Computation and Language (cs.CL); Cryptography and Security (cs.CR); Machine Learning (cs.LG)
备注: 24 pages, 7 figures, 21 tables. Jungseob Lee and Dongyub Jude Lee contributed equally
Abstract:Fine-tuning adapts aligned large language models (LLMs) to downstream tasks, but a few dozen harmful examples can remove their refusal of harmful requests. Prior work localizes safety-related behavior to specific layers, directions, and tokens, suggesting targets for protection. We test whether successful localization and recovery support defenses that survive changes in the attack. Across six checkpoints from four model families, harmful and benign prompts remain linearly separable after attack, and patching full clean hidden states into the compromised model restores refusal at a reproducible transition depth. Building on a prior layer-freezing defense, we freeze every layer up to this depth and repeat the attack. At a hundred harmful examples, refusal remains near zero on all six checkpoints, with recovery transitions above the frozen boundary. In a second study, removing the update’s top two singular directions restores refusal after short attention-only fine-tunes on four checkpoints. On Llama-3.1-8B, ordinary training changes weaken this repair and an attacker who spreads the update defeats it. A spectral detector calibrated on benign Llama fine-tunes misses most repair failures on that checkpoint. Localized freezing can nevertheless help preserve refusal when a few harmful examples enter training data unintentionally. These results show that an attacker can bypass a region identified by recovery and defeat a repair that works across multiple checkpoints, motivating five checks for defenses against adaptive fine-tuning. Code is available at this https URL.
[NLP-138] DuplexSpeechBench-Document Grounding: Benchmarking Document Grounding and Hallucinations in Voice Agents EACL2027
【速读】: 该论文旨在解决语音代理(Voice Agent)在与外部文档进行语义对齐时的可靠性问题,特别是其在真实对话场景中保持响应忠实性(document grounding fidelity)的能力不足。当前语音代理虽能实现低延迟、自然的人机交互,但其对文档内容的精准引用和长期记忆能力尚未得到充分研究。为系统评估这一关键挑战,作者提出了DuplexSpeechBench-Document Grounding(DSB-DG)基准测试,涵盖五个专业领域,包含1,636组经对抗验证的问答对,支持自动评估接地准确率、幻觉率及响应延迟。该基准聚焦三大典型失效模式:上下文饱和(Context Saturation)、接地衰减(Grounding Decay)以及主动重注入机制的有效性(Proactive Grounding)。实验结果表明,尽管端到端级联架构(ASR-LLM-TTS)在接地准确性上表现最优,但主流全双工模型如Gemini-Live与GPT-Realtime也表现出较高潜力;而开源权重系统则暴露出显著缺陷,包括上下文容量骤降和多轮对话中的接地衰减。总体而言,接地保真度随上下文长度与对话负载增加而显著下降,且失败常表现为生成无依据内容而非主动回避,揭示了上下文对齐仍是构建可靠全双工语音代理的核心未解难题。
链接: https://arxiv.org/abs/2610.00316
作者: Puneet Mathur,Nedim Lipka,Zeyu Jin,Dinesh Manocha
机构: Adobe Research(Adobe 研究院); University of Maryland College Park(马里兰大学学院公园分校)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: Under submission at EACL 2027
Abstract:Voice agents enable low-latency, natural interaction, yet their ability to faithfully ground responses in external documents remains underexplored. We introduce DuplexSpeechBench-Document Grounding (DSB-DG), a benchmark for evaluating document grounding in voice agents across five professional domains. DSB-DG targets three failure modes: Context Saturation, which measures grounding under increasing document length; Grounding Decay, which measures retention of document facts across multi-turn dialogue; and Proactive Grounding, which evaluates whether context re-injection mitigates conversational drift. The benchmark contains 1,636 adversarially verified QA pairs from 50 documents covering five professional domains, and supports fully automatic evaluation of grounding accuracy, hallucination, and response latency. Across systems spanning cascaded, proprietary full-duplex and real-time, and open-weight speech2speech architectures, we find substantial differences in effective grounding capacity. While cascaded pipeline (ASR-LLM-TTS) achieves the highest grounding accuracy, Gemini-Live and GPT-Realtime closely trail behind. Open-weight systems exhibit distinct failure modes, most notably an abrupt context-capacity collapse and multi-turn grounding decay. More broadly, grounding fidelity degrades with context and conversational load, and failures frequently manifest as unsupported generations rather than abstention. We show that contextual grounding as a key unresolved challenge for reliable full-duplex voice agents.
[NLP-139] okenized Key-Gated Adapter Routing: A Secure Access Control Mechanism Against Private Data Leakage in LLM s
【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在隐私敏感领域(如医疗、金融和政府)应用中因记忆并泄露个人身份信息(Personally Identifiable Information, PII)所带来的安全与合规风险。现有防御方法通常在模型效用、隐私保护与对微调私有知识的访问之间造成权衡。本文提出一种名为Locket的实用框架,通过嵌入细粒度、策略驱动的访问控制机制直接集成于LLM生成过程之中。其核心解决方案在于采用轻量级低秩适配器(LoRA, Low-Rank Adaptation)构建多个编码不同访问策略的适配器模块(如完全披露、基于PII掩码的部分删减或在指定差分隐私级别下的披露),并通过一个紧凑的门控模块实现序列级硬路由,将学习到的“密钥入口令牌”(keyed entry token)精确关联至单一LoRA适配器。合法令牌的存在作为授权凭证,可解锁相应私有知识;而无效或缺失令牌则触发隐私保护适配器,自动对敏感内容进行删减或净化。该设计确保了Locket与现成的LLM完全兼容,支持可扩展部署,同时满足监管与隐私要求。实验结果表明,在提供正确令牌时,Locket在困惑度上接近未经防御的微调表现;而在令牌缺失或无效时,显著降低PII泄露,且保持与先进基线方法相当的模型效用和困惑度水平。
链接: https://arxiv.org/abs/2610.00309
作者: Mohamed Shaaban,Mohamed Elmahallawy
机构: Washington State University (华盛顿州立大学)
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注:
Abstract:Large language models (LLMs) are increasingly deployed in privacy-critical domains (e.g., healthcare, finance, and government), but their propensity to memorize and disclose personally identifiable information (PII) poses serious security and compliance risks. Existing defenses typically force a trade-off between model utility, privacy protection, and access to fine-tuned private knowledge. We propose LoRA-Oriented Control via Keyed Entry Tokens (Locket), a practical framework that embeds fine-grained, policy-driven access control directly into LLM generation. Locket trains a set of lightweight LoRA (Low-Rank Adaptation) adapters, each encoding a distinct access policy (e.g., full reveal, partial redaction via PII masking, or reveal under a specified differential privacy level). A compact gating module is trained to associate a learned keyed entry token with exactly one LoRA adapter via sequence-level hard routing; the presence of a valid token acts as an authorization key that unlocks corresponding private knowledge, while an invalid or absent token triggers a privacy-preserving adapter that redacts or sanitizes sensitive content. This design ensures Locket remains fully compatible with off-the-shelf LLMs, supporting scalable deployment while satisfying regulatory and privacy requirements. We evaluate Locket across multiple datasets (Enron, ECHR, Yelp) and a diverse set of state-of-the-art LLMs, including Qwen3 (1.7B and 8B), Meta’s Llama-3.2 (1B and 3B), and Google’s Gemma-2-2B. Our extensive experiments demonstrate that, when the correct token is provided, Locket preserves perplexity comparable to fine-tuning on raw data (without any defense). Conversely, when the token is missing or invalid, it substantially reduces PII leakage while maintaining utility and perplexity on par with strong baseline defenses.
[NLP-140] Certainty Is Not Just Correctness: Rethinking Token-Level Certainty in LLM Reasoning
【速读】: 该论文旨在解决生成式 AI(Generative AI)在大语言模型(LLM)训练与推理过程中,基于分词级别置信度(token-level certainty)进行决策时存在的有效性问题。现有方法依赖置信度作为正确性代理指标,但其性能不仅受置信度分数本身所包含信息的影响,还取决于如何使用这些分数。为此,论文通过受控的实证评估,在多种模型与任务上直接检验了置信度的预测能力,并区分两个关键预测目标:一是识别模型更可能答对的问题,二是区分同一问题下的正确与错误回答。实验结果表明,置信度在识别高可能性正确问题方面表现优于区分正确与错误回答;同时,置信度在不同词元类型及位置上呈现系统性变化,反映出局部文本结构特征。关于问题难度的信息在生成早期即已显现,而关于答案正确性的信息则集中在生成末尾。这一发现揭示了置信度所提供的信息价值高度依赖于预测目标、模型架构、置信度度量方式以及聚合时所采用的词元位置。基于上述洞察,论文提出一种测试时计算优化策略:在生成初期依据置信度动态分配响应数量,在生成末尾以置信度加权投票。相较于固定采样多数投票基线,该方法将整体准确率从78.71%提升至79.54%,同时降低82.4%的生成词元成本,充分体现了其实际应用价值。
链接: https://arxiv.org/abs/2610.00296
作者: Yunfan Zhou,Ye Zhu,Zhihai Wang,Jianguo Yao,Haibing Guan,Xijun Li
机构: Shanghai Jiao Tong University (上海交通大学); Qwen Team, Alibaba Group (阿里巴巴集团)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注:
Abstract:Token-level certainty is widely used as a proxy for correctness in LLM training and inference. However, the performance of certainty-based methods depends both on the information in certainty scores and on how those scores are used. We therefore directly assess certainty’s predictive ability through controlled empirical evaluations across models and tasks. We distinguish two prediction targets: identifying questions a model is more likely to answer correctly and distinguishing correct from incorrect responses to the same question. In our experiments, certainty is generally better at identifying questions a model is likely to answer correctly than at distinguishing correct from incorrect responses to the same question. Certainty also varies systematically across token types and positions within words, reflecting local properties of words and text form. Information about question difficulty appears early in generation, while the weaker information about answer correctness is more concentrated near the end. These findings show that the information certainty provides for decisions depends on the prediction target, the model, the certainty metric, and which token positions in the response are included in aggregation. We further demonstrate the practical value of these findings for test-time compute. We allocate the number of responses using certainty early in generation and weight answer votes using certainty near the end of each response. Compared with a fixed-sampling majority-voting baseline, this approach increases overall accuracy from 78.71% to 79.54% while reducing generated-token cost by 82.4%.
[NLP-141] Signed Lexical Confidence for Risk-Calibrated Intent Routing
【速读】: 该论文旨在解决智能助手在意图识别(intent recognition)任务中如何有效区分高置信度可靠预测与低置信度不确定请求的问题,以实现选择性意图路由(selective intent routing)。其核心挑战在于标准置信度分数仅依赖于基础模型的表示能力,未能充分利用额外的互补证据。为此,作者提出了一种**符号化词汇门(signed lexical gate)**作为解决方案的关键:该门机制通过融合句子分类器的对数几率差(logit margin)与稀疏词汇模型对分类器预测意图的支持程度,分别赋予正向证据(词汇一致性)和负向证据(竞争意图的词汇偏好),从而比无符号词汇置信度或硬性一致规则保留更丰富的信息。进一步地,引入独立的二项式校准阶段以在给定风险目标下确定操作阈值。实验表明,在BANKING77、CLINC150和HWU64数据集上,该方法相较于仅基于语义学习的门控机制,显著降低了风险-覆盖曲线下面积(相对减少15.8%、15.1%和11.8%),并在5%误差目标下提升了BANKING77和HWU64的可接受覆盖率(分别提升1.83和5.14个百分点),同时在更严格的2%误差目标下仍能保证所有30个数据集-运行组合中存在非空策略。消融实验验证了该特征在平均错误排名上的优越性,并展现出相对于无符号词汇置信度和二元一致规则的显著增益。最终,该双特征门控结构实现了紧凑、可解释且风险校准的置信度增强,同时保持了基础分类器的原始预测结果。
链接: https://arxiv.org/abs/2610.00262
作者: Yezhou Cheng,Zehua Yang,Bojun Lin
机构: University of Wisconsin–Madison (威斯康星大学麦迪逊分校); Pinterest Inc.(Pinterest公司)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注:
Abstract:Selective intent routing allows an assistant to act on reliable predictions while deferring uncertain requests. Standard confidence scores primarily reflect the base model’s representation, leaving an opportunity to incorporate complementary evidence without changing its decisions. We introduce a signed lexical gate that combines a sentence classifier’s logit margin with a sparse lexical model’s support for the classifier’s predicted intent. By assigning positive evidence to lexical agreement and negative evidence to a lexically favored competing intent, the gate retains more information than either unsigned lexical confidence or a hard agreement rule. An independent binomial calibration stage selects an operating threshold for a specified risk target. Across ten runs on BANKING77, CLINC150, and HWU64, the proposed score reduces area under the risk-coverage curve by 15.8%, 15.1%, and 11.8% relative to a learned semantic-only gate. At a nominal 5% error target, it increases accepted coverage by 1.83 and 5.14 percentage points on BANKING77 and HWU64, while CLINC150 is already near full coverage. At a stricter 2% target, the simultaneous binomial procedure yields a nonempty policy in all 30 dataset-run combinations at the available calibration budgets. Matched controls show that the proposed feature improves average error ranking over the tested unsigned lexical-confidence feature, with dataset-dependent gains over binary agreement. The resulting two-feature gate provides a compact, interpretable confidence enhancement for risk-calibrated intent routing while preserving the base classifier’s predictions.
[NLP-142] CAVE-Mem: Boundary-Aware Experience Validation for Memory Search
【速读】: 该论文旨在解决现有长期记忆系统在复用历史经验时因仅依赖相关性(relevance)而带来的潜在误导问题。当记忆基底、问题意图、答案粒度或证据边界发生变化时,看似相关的过往经验仍可能产生错误结果。其解决方案的关键在于提出一种无需训练的框架——CAVE-Mem,将经验建模为具有类型化干预能力的算子(intervention operator),并显式定义其适用条件(applicability)、边界约束(boundary)和效用标准(utility)。该框架在生成基础记忆搜索答案后,仅允许满足当前上下文条件的算子进行干预;否则系统主动放弃干预,从而避免不恰当的经验引入。实验表明,在长时对话记忆、多跳问答及长文档叙事推理任务中,该方法显著优于仅依赖相关性的经验复用策略。
链接: https://arxiv.org/abs/2610.00238
作者: Xinyu Li
机构: Kent State University (肯特州立大学)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注:
Abstract:Long-term memory agents increasingly rely on it- erative search and reusable experience to answer questions over large personal, factual, or narrative histories. However, current experience-memory systems largely optimize relevance: they re- trieve past search lessons that appear similar to the current state and inject them into the prompt. A relevant experience can still be harmful when the memory substrate, question intent, answer granularity, or evidence boundary changes. We propose CAVE- Mem, a training-free framework that represents experience as a typed intervention operator with applicability, boundary, and utility conditions. CAVE-Mem first obtains a base memory-search answer, then allows an operator to change it only if the oper- ator matches the current substrate, answer contract, evidence boundary, and cross-fitted utility; otherwise the system abstains. Experiments across long-term conversational memory, multi-hop question answering, and long-document narrative reasoning show consistent gains over relevance-only experience reuse.
[NLP-143] Robust Is Salient: An Informed Adversary Moves the Optimal Signal onto the Salience Pole
【速读】: 该论文旨在解决在存在知情对手(informed adversary)干扰下的信号传递问题,即当对手能够观测到发送者发出的信号并利用其进行误导时,如何设计最优信号以最大程度保护真实信息。其核心挑战在于:传统基于贝叶斯推理的信号优化方法在面对具有策略性欺骗能力的对手时会失效,而现有启发式方法(如显著性极值,salience pole)虽表现出一定鲁棒性,但缺乏理论上的统一解释。解决方案的关键在于揭示:在对手具备有限说服预算(persuasion budget, β)的情况下,最优信号从最大化后验概率(posterior-maximizing)逐渐转变为最大化边际差异(margin-maximizing),这一转变过程最终收敛于与显著性极值完全一致的解。研究通过构建一个受控的强制选择任务模型(抽象自《诡计:香港谋杀案》),验证了当β增大时,最优信号路径发生可预测的转移;尤其值得注意的是,在20万项测试池中仅有2,748项偏离原显著性极值,且这些偏差恰好出现在先前显著性-贝叶斯坐标系未定义的区域。这表明,真正的鲁棒性并非来自复杂的对抗建模,而是源于从贝叶斯决策向纯粹显著性驱动的系统性转换。该发现揭示了一个结构性限制——在某些情况下,对抗鲁棒目标与启发式显著性目标完全重合,导致无法通过实证手段区分两者,因此提出诊断性检查:在评估模型的对抗感知能力前,应首先验证其优化目标是否与启发式目标一致。
链接: https://arxiv.org/abs/2610.00233
作者: Cris Huynh
机构: Independent researcher(独立研究员)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Science and Game Theory (cs.GT); Machine Learning (cs.LG)
备注: 11 pages, 3 figures
Abstract:When an informed adversary shares the audience of a constrained signalling channel, the signal that best protects the truth is the signal that best describes it. On 108 confirmatory items, the adversary-robust optimum aligns exactly with the salience pole from prior work. Across a 200,000-item pool, the two differ on only 2,748 items — lying exactly where the prior salience-to-Bayes coordinate is undefined. Where defined, robustness is achieved by moving from Bayesian discrimination entirely to salience. We show this by introducing an adversary to a forced-choice task (abstracted from Deception: Murder in Hong Kong). The adversary knows the target, observes the signal, and argues for the strongest wrong answer using a persuasion budget, \beta . As \beta grows, the optimal signal shifts from the posterior-maximizing option to the margin-maximizing one; at \beta = 0 , the game reproduces the original oracle model with a listener temperature of \tau = 1 . This effect is real: 18.2 percent of the pool has an optimum that shifts under a finite budget, and each item’s critical budget is exact. This coincidence structurally limits empirical evaluation. Two adversary framings change the chosen option of seven language models on 30 to 77 of 108 items against an exact no-effect rate. Yet, no measurement can determine whether this movement is toward the adversary-aware optimum or toward salience, because the two options are identical. This is a structural limit, not a null result. The diagnostic check is cheap: before evaluating adversary-awareness, verify whether the robust target coincides with a heuristic target on the evaluation items.
[NLP-144] A Holistic Assessment of the Carbon Footprint of Noor a Very Large Arabic Language Model ACL2022
【速读】: 该论文旨在解决超大规模语言模型(Large Language Models, LLMs)在实际应用中被忽视的全生命周期碳足迹问题。尽管当前研究普遍关注训练阶段的计算资源消耗与碳排放,但本文指出,这一评估范围过于狭窄,未能涵盖数据收集与存储、研发预算、预训练成本、未来推理服务需求以及跨国协作中的外部成本等关键环节。为此,论文提出了一种整体性评估框架,以评估名为Noor的极端规模多任务阿拉伯语语言模型的综合环境影响。其解决方案的关键在于突破传统仅关注训练阶段的局限,将推理阶段能耗及各类外生成本纳入碳足迹核算体系,揭示了推理和非直接计算因素对总碳排放的显著贡献。研究进一步探讨了通过优化模型效率、采用绿色能源、改进部署策略等路径降低超大规模模型碳足迹的可行方案。
链接: https://arxiv.org/abs/2610.00223
作者: Imad Lakim,Ebtesam Almazrouei,Ibrahim Abu Alhaol,Merouane Debbah,Julien Launay
机构: 未知
类目: Computation and Language (cs.CL)
备注: 11 pages, 3 figures, 2 tables. Published in Proceedings of BigScience Episode #5 – Workshop on Challenges Perspectives in Creating Large Language Models (ACL 2022)
Abstract:As ever larger language models grow more ubiquitous, it is crucial to consider their environmental impact. Characterised by extreme size and resource use, recent generations of models have been criticised for their voracious appetite for compute, and thus significant carbon footprint. Although reporting of carbon impact has grown more common in machine learning papers, this reporting is usually limited to compute resources used strictly for training. In this work, we propose a holistic assessment of the footprint of an extreme-scale language model, Noor. Noor is an ongoing project aiming to develop the largest multi-task Arabic language models – with up to 13B parameters – leveraging zero-shot generalisation to enable a wide range of downstream tasks via natural language instructions. We assess the total carbon bill of the entire project: starting with data collection and storage costs, including research and development budgets, pretraining costs, future serving estimates, and other exogenous costs necessary for this international cooperation. Notably, we find that inference costs and exogenous factors can have a significant impact on total budget. Finally, we discuss pathways to reduce the carbon footprint of extreme-scale models.
[NLP-145] When a Data Artifact Isnt a Shortcut: Causal Auditing of Synthetic RLVR Corpora
【速读】: 该论文旨在解决生成式问答数据集(如基于强化学习的文本选择任务)中正确答案与干扰项(distractor)之间真实性与来源性混淆的问题。具体而言,现有方法通过掩码真实语料并由语言模型生成虚假答案来构建训练数据,导致正确选项为真实人类文本,而所有干扰项均为合成生成内容,这种设计使“正确性”与“来源真实性”在数据分布上产生耦合,从而可能诱导模型学习到对来源特征的依赖而非真正的语义正确性。其解决方案的关键在于通过系统性审计验证这一偏差是否被实际利用:研究者首先发现仅基于表面统计特征(如长度、词频等)即可以0.562的AUROC区分正确项与干扰项,表明存在可检测的差异;但进一步分析发现代码类样本的区分性能反而低于随机水平,原因在于干扰项多为对正确答案的单操作符微调,导致两类样本在结构上高度相似。随后通过引入同义改写匹配的对照语料库进行干预实验,在相同训练规模和预算下,对比原始数据与控制数据训练出的策略表现无显著差异(0.021 vs 0.027),证明即便存在可检测的模式,模型并未有效利用该信号。因此,研究揭示了此类数据构造中“可检测偏差”与“模型实际使用”之间的脱节,并强调在构建类似数据集时应关注领域特定的生成机制对模型学习路径的影响,最终将审计流程开源,以支持更透明的数据质量评估。
链接: https://arxiv.org/abs/2610.00202
作者: Esther Xin
机构: Independent Researcher(独立研究员)
类目: Computation and Language (cs.CL); Machine Learning (cs.LG)
备注: 9 pages, 2 figures,4 tables;Code and data this https URL
Abstract:Several recent pipelines build RLVR training data by masking a span of real corpus text and asking a language model to invent plausible wrong answers around it. The correct option is therefore genuine human prose; every distractor is synthetic. Correctness and provenance become entangled, and a policy could in principle learn the second instead of the first. We audit that possibility in GooseReason-0.7M. First we ask whether the asymmetry is visible at all: a classifier reading only five surface statistics (never the meaning) reaches AUROC 0.562 over 315,499 options, barely above chance. The aggregate hides something, though. Code sits at 0.416, below chance, and manual inspection explains why: code distractors turn out to be single-operator mutations of the gold answer rather than freely written alternatives, so the two classes are nearly identical by construction. Detecting a signal is not the same as showing a model uses it, so we then run an intervention. We build a paraphrase-matched control corpus, hold training-set size identical across arms, and train two policies under one fixed budget. The exploitation gap does not favour the unmodified-data arm: 0.021 against 0.027 for the control. Under our budget, in other words, a detectable artifact went unexploited. We think that dissociation, along with the domain-specific construction finding, is worth knowing for anyone curating corpora of this kind, and we release the audit as a mostly CPU-only protocol.
[NLP-146] Comedic Fools Gold: Reward Exploits and Countermeasures in Conversational Humor
【速读】: 该论文旨在解决在对话式幽默(conversational humor)语言模型训练中,自动化奖励机制易被利用(reward exploits)导致模型学习到非预期行为的问题。其核心挑战在于如何设计既能有效激励真实幽默表现,又可抵御策略性漏洞的奖励函数。解决方案的关键在于引入多阶段的对抗性测试与防御机制:首先采用基于嵌入的意外性奖励(embedding-based surprise reward)捕捉“出人意料”的语义特征,但发现其易被词序打乱的回复欺骗;随后引入流畅性过滤器以识别此类篡改,但该组合奖励又误伤部分真正机智的回应;进一步地,通过构建观众情绪预测模型来评估听众笑声,却发现该模型对发言者消息中的笑声线索敏感,易被诱导;最终采用跨说话人归一化笑声线索的方法,成功阻断此类攻击,但仍存在未匹配表达的漏洞。三轮强化学习实验表明,经过持续优化的奖励体系使综合评估得分提升0.0903,零分会话比例下降40%,但针对幽默性的具体改进仍未达到预注册目标。研究揭示了自动化奖励设计的根本难题:必须在消除可被利用的捷径(shortcuts)的同时,完整保留原始奖励所期望鼓励的行为模式。
链接: https://arxiv.org/abs/2610.00197
作者: Sam Larson
机构: Pebble ML
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
备注: 11 pages, 3 figures, 4 tables
Abstract:We investigate automated rewards for training language models in conversational humor, focusing on reward exploits and countermeasures. Two approaches aim to capture understandable surprise and predicted audience amusement. Controlled tests show that an embedding-based surprise reward accepts word-shuffled replies as readily as witty ones. A fluency filter detects the shuffles, but the combined reward also rejects some witty replies and fails further validation. An audience model’s predicted laughter is instead vulnerable to laughter cues in either speaker’s messages. Normalizing these cues across speakers blocks the covered attacks, although unmatched expressions remain exploitable. Three reinforcement-learning runs evaluate training with successive reward revisions. The final run improves the combined evaluation score by 0.0903 and reduces zero-score sessions by 40%, but its humor-specific improvement remains below our preregistered target. These findings illustrate a broader challenge for automated reward design: countermeasures must block exploitable shortcuts while preserving the behavior the reward was intended to encourage.
[NLP-147] Metric-Construction Coupling Inflates Measured Synthetic Dialect Recovery
【速读】: 该论文旨在解决生成式方言数据(synthetic dialect data)在缺乏原始可重用语料时,其性能恢复程度难以准确评估的问题。核心挑战在于:现有韩语方言语料虽存在但不可再分发,导致训练与评估过程无法复现;而生成数据的监督信号是否能有效恢复真实数据中的语言特征,以及这种恢复效果能否独立于具体合成流程进行客观衡量,尚不明确。论文的关键解决方案是提出KoDialectBench基准,包含5个地区、3个评估维度的1,000个样本,仅以标识符哈希和评分代码形式发布,允许用户基于自有授权语料重建测试项,从而实现对合成数据性能的独立验证。研究发现,合成数据的恢复能力具有显著的轴向依赖性——在地域识别任务中最高可达真实数据收益的91.2%,而在理解任务中仅为63.7%。此外,不同评价指标下结果差异明显:基于部署标记词典的生成方式在方言性(dialectness)上达92.3%,在地域匹配(region match)上甚至超过真实数据参考值(119.3%),而基于参考的生成则仅达72.4%。进一步分析表明,标记类别的评分体系完全可由转换规则生成,且通过设置构造-评估不相交的对照组(剔除20%标记类型),在相同训练规模下,方言性恢复率从91.8%骤降至8.1%,地域匹配恢复率从101.9%降至25.6%,而三个与合成流程无关的测量指标保持稳定。由此揭示共享构造与评估词汇表会严重高估合成数据的恢复性能,进而提出“度量-构造覆盖度”(Metric-Construction Coverage, MCC)作为量化指标,并验证其与方言性恢复率呈单调正相关关系,强调构建独立、解耦的评估机制对真实性能评估的重要性。
链接: https://arxiv.org/abs/2610.00164
作者: Hyojung Han(ThakiCloud)
机构: ThakiCloud
类目: Computation and Language (cs.CL)
备注: 22 pages, 6 figures, 6 tables
Abstract:Korean dialect corpora are available but not redistributable: weights may be released, while reproducing training and evaluation from the underlying data cannot be. We ask how much of that supervision synthetic data recovers, and whether that recovery can be measured independently of the synthesis pipeline. We contribute KoDialectBench, 1,000 items across five regions on three axes, released as identifier hashes and scoring code so users reconstruct the items from their own licensed copy. Recovery is strongly axis-dependent: our best synthetic arm reaches 91.2% of the real-data gain on region identification but 63.7% on comprehension. On generation the answer depends on the metric: the deployed marker lexicon reports 92.3% on dialectness and 119.3% on region match, the latter exceeding the real-data reference, whereas reference-based generation reaches 72.4%. We find the marker metrics’ scoring inventory is entirely contained in the inventory our transformation rules can emit. We test the effect of construction access directly with an exact-form construction-disjoint arm that withholds 20% of marker types from the rules. At exactly matched training size (8,600 examples) it reduces dialectness recovery from 91.8% to 8.1% and region-match recovery from 101.9% to 25.6% on the held-out marker inventory, while the three pipeline-independent measurements do not fall at all. A complementary evaluator sweep defines metric-construction coverage (MCC) and finds measured dialectness recovery increasing monotonically as overlap rises from MCC=0 to MCC=1. Shared construction and evaluation inventories can therefore substantially inflate estimates of synthetic-data recovery.
[NLP-148] How Robust Are Neural Audio Codecs for African Speech? A Multi-Task Benchmark and the Limits of Perceptual Quality
【速读】: 该论文旨在解决当前生成式语音编码器(Neural Audio Codecs)在口音多样性和多语言语音场景下鲁棒性评估不足的问题,尤其聚焦于非洲地区语音数据集的适用性。其核心挑战在于现有编码器在低比特率压缩下对非标准口音及多语种语音的保真度与下游任务性能之间的脱节。解决方案的关键在于构建一个全面的评估框架,涵盖信号级质量指标(如ViSQOL、STOI、F0-RMSE)和两个关键下游任务——自动语音识别(ASR)与说话人验证(ASV),并揭示不同编码器架构在跨域条件下的性能差异。研究发现,传统基于参考信号的结构化与可懂度度量(如ViSQOL、STOI)及韵律误差(F0-RMSE)比神经型主观评分预测器(NISQA、UTMOS)更能有效反映下游任务的退化趋势;同时,语音可懂度与说话人身份保留能力在不同架构间存在显著差异,且说话人识别性能排名依赖于后端模型。此外,压缩导致的性能下降具有强烈的领域依赖性,尤其在对话场景中最为严重。更关键的是,通过参数高效的编码器微调(如LoRA,仅需约1–3%参数更新),在约35小时非洲语音数据上进行适应性训练,可部分恢复因压缩造成的词错误率(WER)差距,使识别性能接近未压缩基线。这一结果强调了在部署语音编码器时,必须采用任务感知、领域代表性以及支持自适应调整的评估策略,以实现更具包容性的技术应用。
链接: https://arxiv.org/abs/2610.00154
作者: Chibuzor Okocha,Christan Earl Grant
机构: University of Florida, Gainesville, FL, USA (佛罗里达大学盖恩斯维尔分校)
类目: ound (cs.SD); Computation and Language (cs.CL)
备注: Accepted to IEEE Speech Language Technology
Abstract:Neural audio codecs enable low-bitrate speech compression and tokenization, yet their robustness on accented and multilingual speech remains under-evaluated. We benchmark seven open-source neural codecs (DAC, EnCodec, FocalCodec, LanguageCodec, SemantiCodec, UniCodec, WavTokenizer) on three African speech datasets afrinames, afrispeech dialog, afrispeech multilingual, reporting signal-level quality (NISQA, UTMOS, ViSQOL, STOI, F0-RMSE) alongside two downstream tasks: automatic speech recognition (ASR) and speaker verification (ASV). Our analysis yields four findings. First, signal metrics differ sharply in downstream validity: reference-based structural/intelligibility measures (ViSQOL, STOI) and prosodic error (F0-RMSE) track ASR and ASV degradation far more reliably than neural mean-opinion-score predictors (NISQA, UTMOS). Second, intelligibility and speaker-identity preservation diverge sharply across architectures, and the apparent identity ranking itself depends on the ASV backend. Third, degradation is strongly domain-dependent and largest for conversational dialog. Fourth, the resulting degradation is partly recoverable: parameter-efficient \emphcodec adaptation (LoRA, \sim 1–3% of parameters) on roughly 35 hours of African speech reduces the compression-induced word-error-rate gap, recovering recognition toward the uncompressed baseline (developed in companion work). These results motivate task-aware, domain-representative, and adaptation-aware evaluation of speech codecs as a prerequisite for inclusive deployment.
[NLP-149] Evasion Attacks: How Adversarial Noise Bypasses ML Classifiers
【速读】: 该论文旨在解决生成式对抗样本在图像分类与文本分类任务中对模型鲁棒性的影响问题,特别是针对白盒攻击下的可规避性(evasion attack)现象。其核心问题是:当前模型在干净数据上表现优异(如MNIST上达到98.63%准确率,文本分类任务中达到98.75%准确率),但是否具备实际对抗环境下的稳健性仍需实证检验。解决方案的关键在于通过可复现的实验设计,分别在图像和文本模态上实施可控的对抗攻击——在图像任务中使用FGSM和PGD攻击,在文本任务中引入字符替换、空格噪声及无害后缀等预定义扰动,以评估模型在真实对抗场景中的脆弱性。研究发现,图像分类模型在高扰动强度下准确率急剧下降,而文本分类模型虽出现概率偏移,却未发生类别翻转,表明对抗脆弱性具有显著的模态依赖性。因此,该研究强调:模型的鲁棒性必须通过实证测试来验证,而非仅依赖于干净数据上的高准确率进行推断。
链接: https://arxiv.org/abs/2610.00136
作者: Parker Hummel(Minot State University),Ryne Skabo(Minot State University),Muhammad Abusaqer(Minot State University)
机构: 未知
类目: Cryptography and Security (cs.CR); Computation and Language (cs.CL); Machine Learning (cs.LG)
备注: Presented at the 58th Midwest Instruction and Computing Symposium (MICS 2026), Eau Claire, WI, March 27 to 28, 2026. 14 pages, 7 figures, 4 tables
Abstract:This paper presents a reproducible, educational study of evasion attacks in image classification and text classification. A compact convolutional network trained on MNIST reached 98.63% clean test accuracy and was evaluated under two white-box attacks. Under FGSM, accuracy fell to 60.20% at \epsilon = 0.15 and 1.72% at \epsilon = 0.30; under PGD it fell to 32.47% and 0.41%, and a bit-depth-reduction defense recovered only part of the loss. In the second experiment, DistilBERT fine-tuned on the SMS Spam Collection reached 98.75% accuracy and a 94.96% F1-score, but a controlled sequence of pre-defined perturbations (character substitutions, whitespace noise, and a benign suffix) produced only modest probability shifts in most displayed examples and no flip from spam to ham. Adversarial vulnerability is strongly modality-dependent: the MNIST experiment is a clear evasion demonstration, whereas the text experiment is a controlled robustness evaluation. Robustness must be tested empirically rather than inferred from clean accuracy.
[NLP-150] DramaAgent : Agent ic Storytelling Video Generation
【速读】: 该论文旨在解决长时序文本到视频与音频生成中普遍存在的问题,包括叙事漂移(narrative drift)、角色身份不一致、跨场景连续性弱以及音视频模态不匹配等挑战。现有方法受限于底层生成模型的局限性,难以在长时间序列中保持故事连贯性与多模态一致性。为此,论文提出DramaAgent——一种分层、代理式且模型无关(model-agnostic)的框架,其核心创新在于引入高层控制层,将生成过程分解为故事规划(story planning)、持久角色条件化(persistent character conditioning)、逐场景合成(scene-wise synthesis)以及基于反思的定向修复(reflection-guided targeted repair)。该框架通过维护跨场景可复用的故事状态与角色状态,实现对生成失败(如身份漂移、语义缺失、时间断续、跨模态错配)的诊断与阶段特异性修复。实验表明,DramaAgent在多个视频生成骨干模型上均显著提升了长时程生成的连贯性、角色一致性、叙事保真度及场景级音视频一致性,验证了分层代理式控制在可控长时序音视频生成中的有效性。
链接: https://arxiv.org/abs/2610.00097
作者: Ting Huang,Biao Wu,Ronghao Chen,Zeyu Zhang,Tengfei Cheng,Qizhen Lan,Huacan Wang,Hao Tang
机构: UOL(University of London); UTHealth Houston(德克萨斯大学健康科学中心休斯顿分校); UCAS(中国科学院大学); Peking University(北京大学); UTS AAII(悉尼科技大学人工智能研究所)
类目: Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL)
备注:
Abstract:Recent diffusion and autoregressive models have substantially improved text-to-video generation, yet producing coherent long-form story videos with consistent characters and aligned audio remains challenging. Existing methods often suffer from narrative drift, unstable character identity, weak cross-scene continuity, and audio-visual mismatch over extended sequences. We propose DramaAgent, a hierarchical, agentic, and model-agnostic framework for long-form text-to-video-and-audio generation. Rather than improving the underlying video backbone itself, DramaAgent introduces an upper-level control layer that decomposes generation into story planning, persistent character conditioning, scene-wise synthesis, and reflection-guided targeted repair. The framework maintains reusable story and character states across scenes, diagnoses failures such as identity drift, missing scene semantics, temporal discontinuity, and cross-modal mismatch, and repairs problematic clips in a stage-specific manner. Experiments across multiple video generation backbones show that DramaAgent improves long-horizon coherence, character consistency, narrative fidelity, and scene-level audio-visual consistency over direct generation and strong baselines. These results suggest that hierarchical agentic control is a practical direction for controllable long-form audiovisual generation. Code: this https URL. Website: this https URL.
[NLP-151] FACET at WMT 2026 Automated Translation Quality Evaluation Task
【速读】: 该论文旨在解决机器翻译自动质量评估中因错误类型差异导致的评价不准确问题,尤其关注不同错误类型(如流畅性、准确性与一致性)所需证据来源不同的挑战。传统方法往往采用统一模型处理所有评估维度,忽略了各类错误对上下文依赖性的差异。其解决方案的关键在于提出FACET(Fluency, Accuracy, Consistency Evaluation with Text),一种无需参考译文的评估框架,通过将评估过程分解为三个独立的“通读”阶段——流畅性(Fluency)、准确性(Accuracy)和一致性(Consistency)——每个阶段仅使用对应错误类型所需的上下文信息:流畅性与一致性仅需目标语句本身,而准确性则需源语句作为依据。该方法仅依赖一个固定的预训练语言模型,通过三次不同提示(prompting)生成相应错误片段、质量评分及无错误标签,全程无任何微调组件。此外,还提出了简化版本FACET-C,移除了一致性评估模块。在无黄金标准的情况下,实验表明FACET能有效区分翻译质量,其系统排名将经过后期编辑的人工翻译置于首位;一致性模块虽调整了约十分之一的段落评分,但整体排名保持高度稳定,验证了其评估结果的鲁棒性。
链接: https://arxiv.org/abs/2610.00096
作者: Ahrii Kim,Chanjun Park,Seong-heum Kim
机构: AI-Bio Convergence Research Inst.(人工智能-生物融合研究机构); School of Software(软件学院); Dept. of Intelligent Semiconductors(智能半导体系); Soongsil University(松林大学)
类目: Computation and Language (cs.CL)
备注: Accepted at WMT 2026 (shared task system paper)
Abstract:Different error types in machine translation require different evidence. Whether meaning is preserved can be judged only against the source, while whether the target is well-formed, or whether it names one entity consistently, can be judged from the target alone. We present FACET, our reference-free submission to the WMT26 Automated Translation Quality Evaluation Task, which decomposes evaluation into Fluency, Accuracy, and Consistency passes and gives each pass only the context its error type requires. A single fixed model is prompted three times, and the merged error spans yield the three task outputs, error spans, quality scores, and error-free labels, with no trained components. We also submit FACET-C, which omits the Consistency pass. Without gold labels, we characterize the predictions of FACET. Its system rankings place post-edited human translation first, and the Consistency pass changes about a tenth of segment scores while leaving the ranking nearly unchanged.
[NLP-152] BudgetSchemaBench: A Budget-Swept Diagnostic for Schema Context in Text-to-SQL
【速读】: 该论文旨在解决在结构化数据源上部署生成式AI(Generative AI)时,如何在有限的上下文窗口(context window)预算下高效地选择数据库模式(schema)表示与覆盖范围的问题。其核心挑战在于:大规模数据目录可能涵盖多个数据库及数千列,而成本约束迫使系统在上下文窗口未满前就需权衡表覆盖度与序列化细节。为此,作者提出了一种名为BudgetSchemaBench的执行基准测试框架,该框架通过从真实SQL查询中机械生成相关性标签,无需人工或大模型标注的真值,从而实现对不同模式表示策略的客观评估。实验基于一个包含80个数据库的池化目录,固定检索器的表排序,系统性地比较了四种不同的模式上下文预算水平与三种序列化方式。研究发现,在词法检索(lexical retrieval)场景下,将模式预算从2.5%提升至50%可使执行准确率提升18个百分点,而在密集检索(dense retrieval)下仅提升3个百分点,表明密集检索已能在极低预算下捕获多数必要表;当所需表被移除后,94.6%的正确预测仍能精确命名其中一个,说明模型具备从参数知识中重构缺失模式的能力。在冻结真实表(frozen-gold)条件下,三种序列化方式表现差异不超过2个百分点,且95%置信区间均在±4点以内,显示序列化形式的影响相对较小。此外,这一趋势在另一类推理模型中亦得到验证。总体而言,当检索受限于覆盖范围时,执行准确率对模式预算的变化比对所选序列化方式更为敏感。该诊断工具及其构建与评估代码均已公开。
链接: https://arxiv.org/abs/2610.00092
作者: Chen Shen
机构: Megagon Labs(梅戈隆实验室)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Databases (cs.DB)
备注:
Abstract:Data agents over structured sources must fit database schema into the model’s context window. Large catalogs can span many databases and thousands of columns, so cost constraints may require choosing between table coverage and serialization detail well before the context window is full. We introduce BudgetSchemaBench, an execution-grounded diagnostic for this setting. Its construction derives relevance labels mechanically from gold SQL, without human- or LLM-authored ground truth. Using a pooled 80-database catalog, we sweep four schema-context budgets and compare three representations while keeping each retriever’s table ranking fixed. A source-namespace check rejects queries that obtain the correct result from the wrong database. The evaluation covers three conditions: end-to-end retrieval; frozen-gold, in which the required tables are guaranteed; and a probe that removes those tables. For the primary solver with raw serialization, raising the budget from 2.5% to 50% of the catalog improves execution accuracy on 1,279 held-out questions by 18 percentage points under lexical retrieval but only 3 under dense retrieval; the dense retriever already finds most required tables at the smallest budget. When the required tables are removed, 94.6% of correct predictions name one of them exactly, consistent with reconstruction of absent schema from parametric knowledge. For the two main solvers in the frozen-gold condition, the three representations differ by at most 2 percentage points, and the widest paired 95% confidence interval bounds the difference within +/-4 points. We observe the same qualitative patterns with one reasoning model from a different family. When retrieval is coverage-limited, execution accuracy is more sensitive to the schema budget than to the tested serializations. The diagnostic and the code used to construct and evaluate it are publicly available.
[NLP-153] Legal text classification in Korean sexual offense cases: from traditional machine learning to large language models with XAI insights
【速读】: 该论文旨在解决法律文本分类中因法律文本复杂性及法律类别间细微差异导致的准确分类难题。其核心解决方案在于通过领域适应与微调提升模型性能,而非依赖模型规模。研究发现,基于韩国法律语料库(KLUE)微调的小型模型KLUE-BERT在十类韩国性犯罪判例分类任务中达到99.3%的最高准确率,显著优于通用大语言模型(如GPT-3.5、GPT-4.0)及传统机器学习模型。这表明,在法律文本分类任务中,领域特定的数据微调与模型适配的重要性超过模型规模。此外,研究引入可解释人工智能(Explainable AI, XAI)技术分析模型决策依据与误分类案例,揭示了模型对隐含语境线索捕捉能力的局限性,并验证了在真实法律案例数据集(KICS)上的泛化能力不足。研究强调了法律AI系统在性能与可解释性之间的平衡,证明XAI能够增强法律文本分类的透明度,为法律专业人士提供可信的辅助工具,支持文书分类、法律信息检索与案件评估等实务工作。
链接: https://arxiv.org/abs/2610.00087
作者: Jeongmin Lee
机构: 未知
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: 22 pages, 3 figures. Published in Artificial Intelligence and Law
Abstract:The advancement of natural language processing (NLP) has expanded AI-based text classification in the legal domain. However, accurately classifying legal documents remains challenging due to the complexity of legal texts and subtle differences between legal categories. This study evaluates legal text classification models ranging from traditional machine learning techniques to large language models (LLMs) using ten categories of Korean sexual offense precedents. The results show that fine-tuning small-scale models such as KLUE-BERT on legal data outperforms general-purpose models such as GPT-3.5 and GPT-4.0, as well as traditional machine learning models. KLUE-BERT achieved the highest accuracy of 99.3%, indicating that domain adaptation and fine-tuning can be more important than model size for legal document classification. We further employ explainable AI (XAI) techniques to analyze model predictions and misclassification cases. XAI analysis identifies linguistic features influencing model decisions and limitations in capturing subtle textual cues. Using KICS data, which closely resembles real-world legal case records, we further evaluate the model’s generalization capabilities and find that it struggles to interpret implicit contextual cues. These findings highlight the importance of both performance and interpretability in legal AI and demonstrate how XAI can improve transparency in legal text classification. AI-assisted tools can support legal professionals in tasks including document classification, legal information retrieval, and case assessment.
[NLP-154] Scientific Agents : Evaluating Profession-Specific System Prompts on Scientific Tasks
【速读】: 该论文旨在解决在科学任务中使用详尽的职业特定系统提示(profession-specific system prompts)是否能有效提升生成式 AI 模型的准确性问题。研究发现,尽管职业特定提示在理论上应增强模型的专业性表现,但在实际测试中并未带来一致性的准确率提升:在9个文本类科学基准测试中,平均准确率差异为-0.6个百分点(95%置信区间[-1.5, +0.2]),且无任何基准显示出统计显著的改进。其关键问题是,使用完整职业描述提示显著增加了输出令牌数(1.5–2.3倍)和每成功调用的成本(2.2–4.5倍),同时在工具使用任务(如BioMysteryBench)中因更频繁触发令牌限制和时间超限,导致求解成功率反而下降(-10.0个百分点)。值得注意的是,较长的提示在部分场景下表现出意外优势——在SuperGPQA任务中,由于服务端API中断频繁,长提示的模型在首次通过时仍能给出正确答案的比例更高(71.6% vs. 54.0%),表明提示长度或格式可能在容错性方面具有潜在优势,而非源于领域专业知识本身。因此,研究结论指出,在当前模型与任务设定下,默认加载完整职业提示不仅无法提升准确性,还带来显著成本增加;未来需进一步探索基于选择性检索提示片段或开放性科学任务下的优化策略。
链接: https://arxiv.org/abs/2610.00084
作者: Timothy Kassis
机构: K-Dense
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: 46 pages (11 pages main text, references, 33-page appendix); 10 figures, 29 tables. Evaluated corpus: this https URL (commit 48dedd2); evaluation code and item-level records are not released
Abstract:Detailed profession-specific system prompts raise token use and estimated cost per response without a consistent accuracy gain. We evaluate Scientific Agents, an open-source corpus of 503 profession-specific this http URL profiles, with Gemini 3.8 Flash via OpenRouter in the Pi agent harness. We compare matched profiles with four controls: a minimal baseline (“You are a helpful assistant”), the profile’s opening role sentence, a generic scientific rigor guide, and a profile from an unrelated domain. Across nine text-based science benchmarks (4,531 sampled questions, 100 matched profiles), 4,488 items completed all five conditions after API-error retries, scored with automated, rule-based grading. The average profile-baseline accuracy difference is -0.6 percentage points (95% bootstrap interval [-1.5, +0.2] across fixed tasks), and no benchmark shows a statistically clear improvement. Matched profiles produced 1.5-2.3 times as many output tokens and cost 2.2-4.5 times more per successful call. On 60 tool-using BioMysteryBench bioinformatics problems (three runs each for baseline and profile), mean solve rates were 46.7% with the profile and 56.7% at baseline, a difference of -10.0 percentage points (95% interval [-16.7, -3.3]) driven by more frequent token- and time-limit stops under the profile. Longer prompts had one unexpected operational advantage: on SuperGPQA, frequent provider API drops left the short baseline with a correct first-pass answer on only 54.0% of items, against 71.6% with the profile. Generic and mismatched prompts were about as reliable, so this gain comes from prompt length or formatting rather than domain expertise. For the tested model and tasks, loading full profession profiles by default does not improve accuracy and costs considerably more; whether selective retrieval of profile sections or open-ended scientific tasks would change this remains to be tested.
[NLP-155] Measuring Human-Like Bias in LLM s? A Critique of Human-Derived Bias Constructs in LLM Evaluation
【速读】: 该论文旨在解决当前基于人类偏见构念(human-derived bias constructs)评估大语言模型(LLM)时存在的方法论困境,特别是当将原本用于研究人类认知与社会行为的心理学工具直接迁移至对模型的偏见分析时所引发的推断鸿沟。其核心问题在于:人类心理量表与模型评估范式之间存在根本性不匹配——前者基于人类内在认知过程的测量,后者依赖于概率输出、文本补全、排序或模拟决策等非心理机制性的指标。这种不一致导致对模型偏见的解释可能混淆了“人类心理偏见”与“模型生成偏差”的本质区别。该研究的关键解决方案是提出一个系统性分析框架,通过明确目标构念(target constructs)、操作化方式(operationalizations)、推断范围(scope of inference)以及人类类比的适用边界(limits of human analogy),为跨实体(人类与模型)的偏见评估提供可辩护的解释路径,从而实现对模型偏见更精确、更具理论依据的识别与解读。
链接: https://arxiv.org/abs/2610.00070
作者: Antonela Tommasel,Markus Schedl
机构: Johannes Kepler University Linz(约翰内斯·开普勒林茨大学), Austria; ISISTAN, CONICET-UNICEN(阿根廷国立科学技术研究委员会-科尔多瓦国立大学信息科学研究所), Argentina; Linz Institute of Technology(林茨技术研究所), Austria
类目: Computation and Language (cs.CL)
备注:
Abstract:Researchers increasingly use human-derived bias constructs to study Large Language Models (LLMs), including social-cognitive constructs such as implicit bias and stereotype activation, and cognitive biases such as anchoring, framing effects, and confirmation bias. Such approaches offer alternatives to overt bias probes, particularly when direct questioning may obscure bias or when model behaviour appears normatively acceptable. However, adapting human bias constructs to LLMs introduces an inferential gap. Psychological instruments were developed to study human cognition and social behaviour, whereas LLM evaluations rely on probabilities, text completions, rankings, or simulated decisions. This paper critiques human-centered bias evaluation in LLMs. We show how this gap arises from mismatches pertaining to human-derived constructs, human-model differences, and evaluation contexts, which can blur distinct interpretations of model bias. We then introduce a framework providing an analytical lens for relating these elements to warranted interpretations, with attention to target constructs, operationalizations, scope of inference, and limits of human analogy.
[NLP-156] Pushing CPU Speech Synthesis to the Wall: Extreme Inference Tuning under Serverless Architecture and Billing
【速读】: 该论文旨在解决在按实例计费的无服务器(serverless)计算平台上,由于冷启动延迟和长时间保持热实例状态所导致的高昂推理成本问题。传统运行时架构在高并发场景下存在请求级并行引发的CPU竞争,且热实例会持续占用大量可计费的内存与页缓存资源,造成显著的闲置成本。其解决方案的关键在于提出一种“请求级并发推理”机制,通过限制单个请求的并行度以降低CPU争用;同时引入可回收的实例生命周期管理策略,在空闲期释放推理状态与页缓存内存,仅保留服务进程与编译缓存,从而大幅减少账单中的空闲内存开销。实验结果表明,该方案在Kokoro-82M模型上将每CPU秒生成音频时长从0.89提升至2.71,单位音频小时成本由0.0631降至0.0153(下降4.1倍),空闲账单内存从8.7 GB降至1.33 GB,且首次音频输出恢复时间缩短至2.2秒,显著优于PyTorch冷启动的7.7秒,尤其在突发流量场景下实现了推理效率向实际成本节约的有效转化。
链接: https://arxiv.org/abs/2610.00063
作者: Pakorn Nathong,Kunat Pipatanakul
机构: Paxa Labs
类目: Distributed, Parallel, and Cluster Computing (cs.DC); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: 6 pages, technical report
Abstract:Instance-billed serverless platforms charge for CPU and memory over the lifetime of a warm instance, making idle inference state a direct serving cost. We present billing-aware neural text-to-speech (TTS) serving on serverless CPUs, optimizing CPU-seconds and GB-seconds rather than throughput or latency alone. Conventional runtimes are poorly suited to this setting: per-request parallelism causes CPU contention under concurrency, while warm instances retain gigabytes of billable inference and page-cache state. We address these costs with request-sized concurrent inference, which bounds per-request CPU parallelism, and a reclaimable instance lifecycle, which releases inference state and page-cache memory after idle periods while retaining the server process and compile cache. On Kokoro-82M, our system achieves 2.71 audio-seconds per CPU-second versus 0.89 with ONNX Runtime defaults and reduces cost per audio-hour from 0.0631 with PyTorch to 0.0153, a 4.1x reduction. Idle billed memory falls from 8.7 GB to 1.33 GB, while restoration reaches first audio in 2.2 s versus 7.7 s for a PyTorch cold start. Under bursty traffic, lifecycle reclamation is essential for translating inference efficiency into lower serverless cost. Comments: 6 pages, technical report Subjects: Distributed, Parallel, and Cluster Computing (cs.DC); Artificial Intelligence (cs.AI); Computation and Language (cs.CL) Cite as: arXiv:2610.00063 [cs.DC] (or arXiv:2610.00063v1 [cs.DC] for this version) https://doi.org/10.48550/arXiv.2610.00063 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[NLP-157] he First Token Is Not the Verdict: Hidden Costs of Reading LLM Judges Without Generating
【速读】: 该论文旨在解决大语言模型(LLM)评判者在评估生成内容时因仅依赖首个生成标记的对数概率(logits)所导致的位置偏差(position bias)估计失真问题。其核心解决方案的关键在于揭示并量化这种读出方式的系统性偏差:当评判者未以结论性标记(verdict token)开头时(在三款Qwen3评判者中占比12%至49%,Llama-3.1-8B和Phi-3.5-mini则低于3%),强制读取首个标记会返回被展示在前的响应内容,而非真实判断,从而导致位置偏差被高估。实验表明,在924个未明确表态的样本中,强制读取在响应交换后有89.7%发生翻转,远高于生成后读取的47.5%,显示出显著的上界偏倚(+0.422,95%置信区间[+0.365, +0.467])。值得注意的是,该偏差主要影响位置偏差的测量,对评判准确率的影响不足1个百分点,因此误导的是审计者而非使用者。此外,即使评判者以结论标记开头,仍存在以单字母起始并逐步推理的情况(占比0–5.5%),构成次级偏差。为此,论文建议在报告位置偏差结果时,应同时披露评判者以结论标记开头的比例,该指标仅需一次前向传播且无需标注成本,可有效提升评估透明度与可靠性。
链接: https://arxiv.org/abs/2610.00054
作者: Gnaneswar Villuri,Hashmath Shaik,Alex Doboli
机构: Stony Brook University (石溪大学)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注:
Abstract:Reading an LLM judge’s verdict from the logits of its first generated token is cheap, requires no generation, and is exactly what constrained decoding and likelihood-scoring evaluation harnesses produce. We show that this readout distorts position bias in one direction: it overstates it in every condition we test, so figures obtained this way behave as upper bounds. The mechanism is that judges do not always lead with a verdict token, on 12% to 49% of pairs for three Qwen3 judges and under 3% for Llama-3.1-8B and Phi-3.5-mini, and forcing a read on those pairs returns whichever response was shown first rather than a judgment. Pooled over the 924 pairs where a judge did not commit, the forced read flips on 89.7% of them when the responses are swapped, against 47.5% read after generation (paired difference +0.422, 95% CI [+0.365, +0.467]). The distortion is specific to what is measured: it moves position bias by 42 points while moving judge accuracy by under one point in seven of ten conditions, so it misleads whoever audits a judge rather than whoever uses one. A second, smaller failure occurs even when the judge does lead with a verdict token, since it sometimes opens with one letter and reasons its way to the other, on 0 to 5.5% of pairs at a rate uncorrelated with compliance. We recommend reporting the rate at which a judge leads with a verdict token, which costs one forward pass and no labels, alongside any position-bias figure.
[NLP-158] Characterizing a Configuration Where Inference-Time PRM-Pruned Frag ment Grafting Is Inert: Evidence from Three Reasoning LMs
【速读】: 该论文旨在解决并行链式思维(parallel chain-of-thought, CoT)中因多样性崩溃导致的推理性能下降问题,其核心挑战在于如何在不引入额外计算开销的前提下实现跨轨迹的步骤级知识迁移。解决方案的关键是提出一种名为PRM-Pruned Fragment Grafting(PPFG)的机制,即当过程奖励模型(Process Reward Model, PRM)对某条推理路径进行剪枝时,将其高奖励前缀直接作为上下文示范插入仍在解码的兄弟路径中,以实现轻量级的知识传递。研究发现,在多个基准测试(包括Math-Shepherd on MATH500)和不同基础大模型(Qwen/LLaMA系列)上,无论采用停滞目标导向还是随机目标导向的变体,PPFG均与独立并行CoT基线在所有评估维度上无统计差异。深入分析表明,仅有14%的剪枝事件作用于真正陷入困境的路径,其余多数发生在已成功、接近完成或处于平坦奖励平台的路径上,此类状态无法通过嫁接修复。进一步验证显示,即使采用复合门控优化也难以同时实现精准触发与足够密度,而随机控制在2.4倍更高触发率下仍能达到相同性能平衡,说明该惰性并非特定启发式设计所致。全样本分析支持“等效性”结论,且每事件层面的干预虽显著提升单步剪枝率(2.75倍),但群体层面缺乏补偿效应。事后回溯的虚拟最优策略表明,相对于独立推理,PPFG所能带来的最大潜在增益仅为+0.13个百分点。因此,本文贡献了一个基于等效性检验的推理阶段机制零假设验证框架,强调所有结论均需限定于具体操作条件。
链接: https://arxiv.org/abs/2610.00047
作者: Khawaja Murad ul Hassan,Mehran Ebrahimi
机构: QLU.ai(QLU.ai); Ontario Tech University(安大略理工大学)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: 24 pages, 4 figures, 22 tables
Abstract:Diversity collapse in parallel chain-of-thought has motivated inference-time interventions built on a natural design: when a process reward model (PRM) prunes a chain, its high-PRM prefix is extracted and grafted verbatim as an in-context demonstration into a still-decoding sibling. We isolate this mechanism, PRM-Pruned Fragment Grafting (PPFG), as the most cost-minimal operationalization of cross-trajectory step-level transfer, and test it at the operating point where prior fragment-grafting work reports gains only under additional compensating ingredients. On Qwen2.5-7B-Instruct with Math-Shepherd on full MATH500 (n=500, three seeds), PPFG in both stagnation- and random-targeting variants is statistically indistinguishable from an independent parallel-CoT baseline on every measured axis. We characterize why: a four-bucket classification of 322 stagnation-rule injection events shows only 14% targeted a genuinely struggling chain; the rest landed on chains that had already succeeded, were near completion, or sat on a flat PRM plateau, states a rescue graft cannot change. No compound-gate refinement jointly achieves well-targeted firing and adequate density, and a random control matches the same parity at 2.4x the firing rate, so the inertness is not heuristic-specific. The finding replicates across three base LMs, six benchmarks, a second PRM, and a compatibility-gate sweep; two-one-sided-tests analysis promotes the parity to positive equivalence on all twelve Qwen/LLaMA cells. A per-event spot-check finds injected chains prune at 2.75x the matched-step rate, but a surviving-sibling counterfactual finds no population-level compensation. A hindsight oracle bounds any per-problem gain from choosing PPFG over independent at +0.13 pp. We contribute an equivalence-testing template for establishing inference-time mechanism nulls, with every claim scoped to its tested operating point.
[NLP-159] SCM-based Fairness and Faithful Explainability for Legal Document Classification
【速读】: 该论文旨在解决生成式法律AI模型在司法决策支持中面临的公平性与解释透明性之间的权衡问题,尤其关注去偏干预是否会影响模型解释的忠实度。其核心问题是:提升模型公平性的正则化方法是否会无意中损害解释结果对模型内部推理过程的忠实反映。解决方案的关键在于引入一种基于对比表征正则化的公平性调节机制,在微调过程中对刻板印象中的“温暖度”与“能力感”表征施加惩罚,以缓解性别等人口统计学维度的偏差。研究通过在LexGLUE提供的ECtHR疑似侵权语料库上对比LegalBERT基线模型与该正则化变体,发现尽管在最优正则化强度下模型的分类性能和人口统计学公平性未显著下降,但解释的充分性(explanation sufficiency)在所有随机种子和多个阈值下均出现系统性退化。进一步的打乱词对对照实验表明,这种解释质量下降并非源于特定的温暖-能力结构,而是由对比表征正则化本身所引发。这一结果揭示了公平性与解释忠实性之间存在解耦关系:解释行为的变化并不必然反映公平性的改变,因此公平性必须通过独立评估加以验证。
链接: https://arxiv.org/abs/2610.00045
作者: Yasmina El Kacemi,Seyed Sahand Mohammadi Ziabari,Ali Mohammed Mansoor Alsahag
机构: University of Amsterdam (阿姆斯特丹大学)
类目: Computation and Language (cs.CL)
备注:
Abstract:Transformer models such as LegalBERT are increasingly used in legal decision support, raising concerns about both fairness and the transparency of model explanations. These properties are usually evaluated separately, leaving open whether a debiasing intervention that changes fairness also changes how faithfully explanations reflect model reasoning. This study investigates that relationship on the ECtHR alleged-violations corpus from LexGLUE. It compares a LegalBERT baseline with a fairness-regularized variant that penalizes stereotypical warmth and competence representations during fine-tuning. The evaluation covers predictive performance, demographic fairness, and SHAP explanation faithfulness across five random seeds. At the performance-optimal regularization strength, the intervention does not reduce demographic disparity. This null result holds across two fairness definitions and a conventional word-pair control on the gender axis. Classification performance is largely unchanged. However, the intervention consistently degrades explanation sufficiency across all five seeds and three thresholds. A shuffled-pair control reproduces this degradation while leaving performance and fairness unchanged, indicating that the effect arises from contrastive representational regularization rather than specifically from the warmth and competence structure. The results demonstrate a dissociation between fairness and explanation faithfulness: changes in explanation behavior do not necessarily indicate changes in fairness, and fairness must therefore be evaluated directly.
[NLP-160] Integrating Fairness and Explainability in a Multiple Instance Reinforcement Learning System
【速读】: 该论文旨在解决教育数据中学生学业表现预测的准确性与可解释性之间的矛盾,同时应对引入人口统计学信息后可能导致的不公平预测风险。其核心问题是:如何在保证模型高预测性能的同时,实现对公平性的可控调节,尤其是在多目标优化场景下维持稳定、可解释的决策边界。解决方案的关键在于构建一个融合强化学习驱动的多实例学习(Reinforcement Learning-based Multiple Instance Learning, RL-MIL)、对抗去偏(adversarial debiasing)和偏好条件超网络(preference-conditioned hypernetworks)的多目标框架。其中,RL-MIL将每位学生建模为一组弱标签交互实例的“包”(bag),并通过强化学习代理选择关键实例以提升分类性能;而偏好条件超网络旨在通过用户定义的偏好标量连续调控预测性能与等机会(Equalized Odds)之间的权衡。然而实验发现,尽管基础RL-MIL表现优异,两种超网络扩展均出现模式坍缩(mode collapse)现象,即偏好权重变化未能有效引导模型沿预期的公平-性能前沿系统性演进。根本原因在于目标主导性(objective dominance)、条件机制中的梯度传播微弱,以及动态生成参数间的复杂交互。研究结果表明,虽然可通过该框架将公平性目标嵌入可解释的RL-MIL流程,但仅依赖偏好条件化不足以确保多目标行为的可控性。因此,实现鲁棒的公平强化学习多实例学习,必须引入显式的梯度平衡机制、目标解耦策略及稳定性分析方法。
链接: https://arxiv.org/abs/2610.00035
作者: Bente Hinkenhuis,Seyed Sahand Mohammadi Ziabari,Ali Mohammed Mansoor Alsahag
机构: University of Amsterdam (阿姆斯特丹大学)
类目: Machine Learning (cs.LG); Computation and Language (cs.CL)
备注:
Abstract:Predicting student performance from educational interaction data requires models that are both accurate and sufficiently transparent to support meaningful intervention, while demographic information introduces an additional risk of unfair predictions. This study investigates a multi-objective framework that combines reinforcement learning-based multiple instance learning (RL-MIL), adversarial debiasing, and preference-conditioned hypernetworks for student-at-risk prediction. MIL represents each student as a bag of weakly labeled interactions, while an RL agent selects informative instances for downstream classification. Two hypernetwork variants are evaluated to determine whether a user-defined preference scalar can continuously control the trade-off between predictive performance and Equalized Odds. The underlying RL-MIL baseline achieves strong classification performance, but both hypernetwork extensions exhibit mode collapse: changing the preference weight produces little systematic movement along the intended fairness-performance frontier. The failure is associated with objective dominance, weak gradient propagation through the conditioning mechanism, and interactions between dynamically generated parameters. The results show that fairness objectives can be incorporated into an interpretable RL-MIL pipeline, but preference conditioning alone does not guarantee controllable multi-objective behavior. Robust fair RL-MIL therefore requires explicit mechanisms for gradient balancing, objective separation, and stability analysis.
[NLP-161] High-Value Synthetic Supervision for Parameter-Efficient Adaptation of a Compact Japanese Speech Model NEURIPS2026
【速读】: 该论文旨在解决私有领域语音数据难以获取与再利用,同时小型模型又依赖特定任务监督的问题。其核心解决方案是构建一个可审计的合成语音生成流程,将日语护理交接场景的语音直接映射为六字段结构化笔记。关键在于仅使用182段带有溯源信息的合成训练与开发数据,通过全量微调(full fine-tuning)和秩为16的低秩适应(LoRA)对14.7亿参数的音频模型进行适配。实验表明,尽管未经过适配的基线模型在事实性召回上仅为0.0500,但全量微调(0.8664)与LoRA(0.8461)均实现显著提升;其中LoRA仅需约1240万可训练参数(约为骨干模型的0.85%),即可达到全量微调约97.7%的综合性能,展现出极高的参数效率。然而,该研究并未评估延迟、内存占用、能耗或实时性等设备性能指标,所有评估均基于合成语音与模型自动生成的参考文本,且评估者未校准,因此结论仅支持在少量溯源合成样本下,紧凑模型可高效获取特定音视频到结构化输出的转换能力,尚不能证明其临床有效性、真实语音迁移能力或优于云端清洁系统。
链接: https://arxiv.org/abs/2610.00026
作者: Sidi Chang,Peiying Zhu
机构: 未知
类目: ound (cs.SD); Computation and Language (cs.CL); Machine Learning (cs.LG)
备注: Submitted to On-Device Intelligence: Foundation Models under Real-World Constraints (NeurIPS 2026 workshop). 4 pages, 0 figures, 1 table
Abstract:Private domain speech is difficult to collect and redistribute, while compact models need task-specific supervision. We study an auditable synthetic pipeline that maps Japanese care handoffs directly to six-field structured notes. Using 182 synthetic training and development clips, we adapt a 1.47B audio model by full fine-tuning and rank-16 LoRA. On a 39-clip scenario-seed-disjoint synthetic test, an unadapted model obtains a model-judged factuality-recall score of 0.0500, full tuning 0.8664, and LoRA 0.8461. LoRA reaches 97.7% of the full-tuning aggregate as a descriptive ratio while project telemetry reports 12.4M trainable parameters, about 0.85% of the backbone. Both adaptations show large paired gains over the same base; the full-versus-LoRA interval crosses zero, and differing optimization settings preclude an equivalence claim. This is a parameter-efficient capability-acquisition result, not a device-performance result: latency, memory, energy, and real-time factor were not measured. All evaluation speech and targets are synthetic, references are model-proposed, and the judge is uncalibrated. The evidence shows that a compact model can acquire a narrow audio-to-structure transformation from a few hundred provenance-linked synthetic examples; it does not establish clinical validity, real-speech transfer, or superiority to a clean cloud system.
[NLP-162] Encoded but Disconnected: Decomposing Vision-Language Model Failures under a Patching Null
【速读】: 该论文旨在解决多模态视觉-语言模型(vision-language model)中层表征可解释性问题,特别是针对中间层在推理错误中的作用机制。研究发现,在三个主流模型(LLaVA-1.5-7B、Qwen2.5-VL-7B、InternVL3-8B)上,尽管中层表征在68%-91%的错误案例中编码了真实答案(ground-truth answer),但该信号并未对最终预测产生因果影响:通过残差流补丁(residual-stream patching)实验,所有模型在层级别均未观察到非平凡的预测翻转(0% non-trivial flip),且在两个模型的注意力头级别也未出现有效响应(如Qwen模型中12,600个头无一成功)。唯一例外的InternVL3第20层第2个注意力头虽表现出局部效应,但其影响局限于自身,不具备泛化能力(p < 1e-4)。尽管如此,错误仍可被操作性地划分为三类失败模式——感知失败(Perception Failure)、编码但断连(Encoded-but-Disconnected)、先验覆盖(Prior-Override),这些模式在三类模型上均可被学习并达到超过60%的分类准确率,且模型的先验方向可预测不同干预措施引发的类别特异性响应。研究通过“理想标签”(oracle labels)下的缓解效果验证了这些类别在机制上的真实性,而非提供可部署的解决方案。因此,该研究的核心发现是:中层虽蕴含正确信息,但其与输出决策间的因果通路断裂,揭示了当前模型内部表征与行为之间存在显著脱节。
链接: https://arxiv.org/abs/2610.00024
作者: Genpei Zhang
机构: 未知
类目: Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL); Machine Learning (cs.LG)
备注: 13 pages, 4 figures
Abstract:Across three vision-language model architectures (LLaVA-1.5-7B, Qwen2.5-VL-7B, InternVL3-8B), we report a universal negative finding for mid-layer interpretability. On POPE – the benchmark common to all three – the mid layers encode the ground-truth answer in 68-91% of errors, yet this signal is not causally active for the final prediction: residual-stream patching yields 0% non-trivial flip at the layer level on all three architectures, and on two of three at the per-head level (Qwen: 0/12,600 patched forwards). The lone exception, InternVL3 layer-20 head-2, is a non-vocab, self-attending head whose effect is localized to that specific head (p 1e-4). Despite the null, the errors separate operationally into three failure modes – Perception Failure, Encoded-but-Disconnected, Prior-Override – learnable above 60% on all three architectures, and the architecture’s prior direction predicts which of two interventions elicits a category-specific response. We report these mitigation effects under oracle labels as evidence the categories are mechanistically real, not as a deployable method.
[NLP-163] What Do Rationales Communicate? A Message-Intervention Study in Role-Specialized QA
【速读】: 该论文旨在解决生成式问答(QA)系统中角色专业化流水线(role-specialized QA pipelines)里,推理模型(reasoner)向验证模型(verifier)传递的推理过程(rationale)究竟带来了何种实际价值的问题。具体而言,研究关注的是:这一信息传递是否显著提升答案准确性、增强证据支持判断能力,或反而引入新的错误来源。其解决方案的关键在于提出一种“消息干预诊断”(message-intervention diagnostic)方法,通过固定证据和候选答案,仅在推理到验证的边界上系统性地改变传递的推理内容,从而隔离并量化推理文本本身对验证结果的影响。实验基于400个来自MuSiQue、HotpotQA和2WikiMultiHopQA的数据样本,使用DeepSeek作为生成器与验证器,发现忠实的推理过程对答案准确率几乎没有增益,而被破坏的推理则显著影响支持判断(支持评估偏差达10–22%),尤其在显式要求检查推理的提示下,偏差扩大至34–55%。然而,最终答案的变化幅度较小(2–30%),且仅有2.9–35.3%的错误支持判断伴随答案变更。人工审计揭示,多数错误源于“误信错误推理”(corruption-overtrust)现象,且人类评审者能识别出模型所忽视的无效推理。研究进一步表明,推理通道的状态(活跃、放大、失效或融入任务标签)具有跨模型与任务边界的可变性。因此,论文强调应将推理传递视为一种验证性信息通信机制,而非单纯追求更高答案准确率的路径。
链接: https://arxiv.org/abs/2610.00018
作者: Jiameng Zhang,Hongqiu Wu
机构: University of Zurich (苏黎世大学); Shanghai Jiao Tong University (上海交通大学)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: 14 pages, 6 figures, 9 tables. Preprint
Abstract:Role-specialized QA pipelines increasingly pass rationales from a reasoner to a verifier, but it is unclear what this message actually buys: better answers, stronger support assessment, or a new failure surface. We introduce a message-intervention diagnostic that fixes the evidence and candidate answer while varying only the rationale passed across the reasoner-to-verifier boundary. On 400 MuSiQue, HotpotQA, and 2WikiMultiHopQA examples with DeepSeek as generator and verifier, faithful rationales add almost no answer accuracy over no rationale, while corrupted rationales strongly alter support judgments. Under a blind verifier prompt, harmless paraphrases shift support by only 0–2.5%, whereas corrupted rationales shift support by 10–22%; an explicit rationale-checking prompt amplifies the same pattern to 34–55%. Final answers move less (2–30%), and only 2.9–35.3% of corrupted support flips co-occur with answer changes. Human audits show why this matters: 16/42 valid corruptions are corruption-overtrust cases, and blind humans reject or mark unclear 9/10 audited corrupted rationales that the model accepts. Cross-model and task-boundary checks show when the channel is active, amplified, inert, or folded into the task label. Rationale sharing should be evaluated as a verification-message mechanism, not merely as a route to higher answer accuracy.
[NLP-164] FourierQK: Filter Shape Admissibility and the Leakage-Coverag e Law
【速读】: 该论文旨在解决标准点积注意力(dot-product attention)在建模序列依赖关系时效率与表达能力受限的问题,核心挑战在于如何通过频域特征重构提升注意力机制的性能。其解决方案的关键在于引入频域滤波型注意力——频率坍缩注意力(Frequency-collapse attention),即用可学习频率的带通滤波内积替代传统的Q/K点积,从而实现对序列中关键周期性模式的精准捕捉。研究通过在字符级语言建模任务(TinyShakespeare,6层GPT)上进行受控消融实验,系统验证了五类滤波器属性的影响:直流分量抑制、奈奎斯特分量抑制、带宽、中心频率及多尺度覆盖。主要发现表明:(1)直流与奈奎斯特成分具有显著负向影响(值≈2.0,等同于相位随机化),证实振荡式带通结构不可或缺,而非任意低维谱摘要;(2)最优单尺度带宽约为σ≈2个频段,中心于段落尺度(约70词元),相较基准点积注意力(BASE-DOT)获得Δ=+1.15纳特的清晰增益;(3)满足零均值条件且为墨西哥帽型差分高斯(Mexican Hat DOG, m=2)的可接受滤波器优于同尺度非可接受高斯滤波器,并能部分缓解双边傅里叶变换泄漏;(4)双边傅里叶泄漏随频谱覆盖范围单调增长,窄带滤波器(间隙≥+4)表现纯净,而宽带滤波器(间隙≥+2)存在明显泄漏;(5)因果时域莫莱特(Morlet)滤波器在字符尺度下无法超越基准点积注意力(K=128抽头覆盖T=256上下文的50%),提示需在词粒度层面探索因果谱变体。综合上述结论,该研究确立了基于傅里叶变换的QK建模(FourierQK)在双向注意力场景(如BERT)中的有效性,而自回归生成则需采用因果谱版本如MorletQK(如配套论文[Zeris, 2026f]所述)。
链接: https://arxiv.org/abs/2610.00009
作者: Athanasios Zeris
机构: Independent Researcher, Athens, Greece
类目: Machine Learning (cs.LG); Computation and Language (cs.CL); Signal Processing (eess.SP)
备注: 9 pages, 1 figure, 2 tables
Abstract:Frequency-collapse attention [Zeris, 2026e] achieves large gains over standard dot-product attention by replacing the Q/K dot product with a bandpass-filtered inner product at a learned frequency. A natural follow-up question is: which filter shape works best, and why? We test five hypotheses about filter properties – DC suppression, Nyquist suppression, bandwidth, centre frequency, and multi-scale coverage – using a controlled ablation on character-level language modelling (TinyShakespeare, 6-layer GPT). Our main findings are: (1) DC and Nyquist components are actively harmful (val ~= 2.0, equivalent to phase randomisation), confirming that oscillatory bandpass structure is essential, not just any low-dimensional spectral summary; (2) the optimal single-scale bandwidth is sigma ~= 2 bins centred at paragraph scale (~70 tokens), giving a clean gain of Delta = +1.15 nats over BASE-DOT; (3) admissible filters (zero-mean, Mexican Hat DOG m = 2) outperform non-admissible Gaussians at the same scale and provide partial protection against bilateral FFT leakage; (4) bilateral FFT leakage scales monotonically with spectral coverage – narrowband filters (gap +4) are clean, wideband filters (gap +2) are leaky; and (5) causal time-domain Morlet at character scale cannot beat BASE-DOT (K=128 taps covers 50% of T=256 context), motivating word-level experiments in the companion MorletQK paper [Zeris, 2026f]. Together, findings (1)-(5) characterise FourierQK as effective in bidirectional attention settings (encoder-style, e.g. BERT), where full-sequence context is available at both training and inference time; autoregressive generation requires a causal spectral variant such as MorletQK [Zeris, 2026f] (decoder-style, e.g. GPT). Code available at: this https URL
[NLP-165] On-Device Named-Entity Recognition: A Deployability Study of Accuracy Cost Reliability and Confidence
【速读】: 该论文旨在解决在设备端(on-device)进行命名实体识别(NER)时面临的实际部署挑战,核心问题包括:如何选择可部署的模型、如何在缺乏人工标注的情况下评估模型性能,以及如何判断模型输出置信度的可靠性。其解决方案的关键在于构建一个全面的评估框架,涵盖准确性、延迟和输出有效性三个维度,并引入一种基于跨家族大语言模型(LLM)判官小组生成的“银色黄金”(silver gold)标注数据集,以替代缺失的真实标签(gold standard)。通过将九种不同规模与范式的系统(包括传统标注器spaCy、双向编码器模型GLiNER及本地运行的生成式大模型Qwen3/DeepSeek-R1)在三个异构数据集上进行测试,研究发现:尽管4B参数量的生成式模型在干净新闻文本上表现接近甚至领先于编码器模型,但其在长输入下存在高达27%的无效输出率,而小型编码器模型仅需其1/9至1/24的模型规模即可实现毫秒至秒级延迟、零格式错误输出,展现出显著的部署优势。此外,研究对GLiNER的每跨度置信度进行了深入分析,发现其置信度排序能力良好(AUROC 0.76–0.86),但存在过自信问题(ECE 0.24–0.47),经温度缩放后可显著改善;阈值设定带来微小的外部样本准确率提升,而全本地的小到大级联路由策略在成本匹配条件下仅产生依赖语料的适度增益,且置信度虽能反映正确性,却无法捕捉新实体。所有结果均基于离线计算的逐跨度记录重新生成,确保了评估的严谨性与可复现性。
链接: https://arxiv.org/abs/2610.00007
作者: Vinay Kumar Chaganti
机构: 未知
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: 7 pages, 5 figures, 12 tables. Code and per-span records reproduce all reported numbers offline
Abstract:Named-entity recognition (NER) is increasingly wanted on-device (no API, low latency, data kept local). The practitioner’s question is not the leaderboard but which model is deployable, how to evaluate it without human annotation, and whether its confidence can be trusted. We answer these jointly. We place nine systems across three paradigms and 13 M to 8 B parameters: a classical tagger (spaCy), bidirectional-encoder specialists (GLiNER, 166 to 460 M), and generative LLMs run locally (Qwen3-0.6B/1.7B/4B-Instruct, DeepSeek-R1-1.5B/8B), on three datasets of differing character, and report accuracy plus two axes the literature omits: latency and output validity. Because our corpus (RSS-News) had no gold, we built silver gold from a cross-family LLM judge panel, then measured its fidelity against benchmark gold and a full human re-validation of the corpus (strict F1 0.95, an upper bound since the human gold was silver-seeded); gold provenance flips the paradigm ranking, moving from LLM-authored silver to human gold raises every encoder and lowers every generative model. On accuracy alone a 4 B instruct LLM is competitive (it leads on clean newswire), so the encoder’s case is deployability: it matches or slightly trails at one-ninth to one-twenty-fourth the size, at millisecond-to-second latency, with zero malformed output, while the smallest generative models emit up to 27% invalid output on long inputs, a failure fixed by scale, not output budget. We then characterize GLiNER’s per-span confidence: it ranks correctness well (AUROC 0.76 to 0.86) but is overconfident (ECE 0.24 to 0.47, halved by temperature scaling); thresholding gives a small honest out-of-sample F1 gain; an all-local small-to-large cascade gives a modest, corpus-dependent gain over cost-matched random routing; and confidence tracks correctness but not novelty. Every number recomputes offline from per-span records.
[NLP-166] LLM assisted writing deserves empirical evaluation
【速读】: 该论文旨在解决大语言模型(Large Language Model, LLM)辅助写作在学术出版中引发的可信度、完整性、公平性及评价标准等核心争议问题。其关键解决方案在于,通过分析69,209篇健康信息学领域的论文数据,揭示使用LLM辅助写作与更清晰的表述、更广泛的引文实践以及更全球化的作者分布之间存在显著关联。这一发现表明,工具使用本身并不直接决定科研质量,因此应将学术稿件的评估重点从是否使用LLM转向学术质量与责任归属,从而建立更加公正和实质性的评价体系。
链接: https://arxiv.org/abs/2608.22124
作者: Xuan Zhong Feng,Yi Lin,Yiye Zhang,Chunhua Weng,Yifan Peng
机构: Weill Cornell Medicine(威尔康奈尔医学院); Columbia University(哥伦比亚大学)
类目: Computation and Language (cs.CL); Computers and Society (cs.CY)
备注:
Abstract:LLM-assisted writing is often treated as a detection problem, as it raises questions about clarity, integrity, equity, and evaluation. An analysis of 69,209 Health Informatics papers links it to more focused presentation, broader citation practices, and more globally distributed authorship. These patterns do not prove better science, but they support evaluating manuscripts by scholarly quality and accountability rather than by tool use.
[NLP-167] rajectory-Aware Clinical Risk Prediction via Severity-Grounded Knowledge Graphs and Retrieval-Augmented Generation KDD2026
【速读】: 该论文旨在解决电子健康记录(Electronic Health Records, EHRs)在临床风险预测中难以有效融合异构外部知识的问题,尤其针对现有方法在捕捉疾病严重程度、治疗反应及复杂临床进展方面的不足。其核心挑战源于数据稀疏性以及对非结构化临床文本的利用不充分。为应对这一问题,论文提出TRACER(一种轨迹感知且基于临床语境的预测框架),其关键解决方案包括:(1)构建包含疾病严重程度信息的医学知识图谱,源自医学文献;(2)从知识图谱中检索与患者临床进程相关的、具有严重程度加权的路径;(3)从非结构化临床笔记中提取具有临床意义的事件;(4)通过引入相似同行病例来增强患者上下文表征。实验在MIMIC-III和MIMIC-IV数据集上验证了该方法的有效性,在死亡率预测任务中,宏平均F1分数最高提升28.5%,在再入院预测任务中提升达19.7%。
链接: https://arxiv.org/abs/2607.18270
作者: Kyunghoon Jeon,Youmin Ko,Woohwan Jung,Hyunjoon Kim
机构: Hanyang University (汉阳大学); Korea University (高丽大学)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: Accepted to KDD 2026
Abstract:While Electronic Health Records (EHRs) offer a wealth of clinical data, effectively augmenting a patient’s records with heterogeneous external knowledge to predict the patient’s clinical risk remains a significant challenge. Existing methods fail to capture disease severity, treatment responses, and nuanced clinical progression, due to data sparsity and the underutilization of unstructured clinical notes. To address these challenges, we propose TRACER (a trajectory-aware and clinically grounded prediction framework) that (1) constructs a medical knowledge graph enriched with severity information from medical literature, (2) retrieves clinically relevant, severity-weighted paths of a patient’s progression from the knowledge graph, (3) extracts clinically relevant events from unstructured clinical notes, and (4) augments patient context with similar peer cases. Experiments on the MIMIC-III and MIMIC-IV datasets demonstrate large gains over state-of-the-art baselines, with up to 28.5% increase in Macro F1 score for the mortality prediction task, and 19.7% increase for the readmission prediction task.
[NLP-168] Code-Switching Spoken Language Identification as Multi-Label Set Prediction
【速读】: 该论文旨在解决多语言语音语料库构建过程中,因采用单语种语言识别(LID)过滤器而导致的代码切换(Code-Switched, CS)语音泄漏问题,从而提出面向代码切换的语音识别(CS-LID)方法。其核心解决方案是将话语级CS-LID建模为多标签语言集合预测任务,并引入一种直接输出话语中包含语言集合的集合生成器(set generator),以区别于传统的原子对分类和基于得分的分类基准。尽管该集合生成器能够在未见语言对上准确预测语言数量且无需预设语言数量假设,但在精确集合匹配准确率上仍逊于最优基线Oracle Top-k。研究进一步分析揭示了实现鲁棒CS-LID的关键障碍:真值语言数量(oracle cardinality)的不确定性、阈值选择的不稳定性、训练数据中存在的语言偏倚,以及合成数据与真实数据之间的分布差距。
链接: https://arxiv.org/abs/2610.01450
作者: Shunsuke Mitsumori,Matthew Wiesner,Shigeo Morishima,Shinji Watanabe
机构: Waseda University (早稻田大学); Johns Hopkins University (约翰霍普金斯大学); Carnegie Mellon University (卡内基梅隆大学)
类目: Audio and Speech Processing (eess.AS); Computation and Language (cs.CL)
备注: Accepted at IEEE SLT 2026
Abstract:Code-switched (CS) speech leaks through the monolingual language identification (LID) filters used to curate massive speech corpora, calling for CS-aware LID (CS-LID). We formulate utterance-level CS-LID as multi-label language-set prediction and propose a set generator that directly outputs the languages in an utterance, comparing it against atomic-pair and score-based classification baselines. Oracle Top-k is the strongest baseline, but thresholding fails because no single threshold separates CS from monolingual speech. Our set generator predicts the correct language count on unseen pairs without assuming the number of languages, but underperforms oracle Top-k in exact set accuracy. Our analysis identifies the key obstacles to robust CS-LID: oracle cardinality, threshold instability, language bias in CS training data, and the synthetic-to-real gap.
信息检索
[IR-0] ScholarCatalyst: A Benchmark for Retrieving Papers That Inspire New Research
链接: https://arxiv.org/abs/2610.02202
作者: Sohyeon Kim,Yoonho Lee,Bo Liu,Dayoon Ko,Rulin Shao,Seungone Kim,Graham Neubig,Pang Wei Koh,Aakanksha Chowdhery,Akari Asai,Omar Khattab,Yejin Choi,Gunhee Kim,Chelsea Finn
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Information Retrieval (cs.IR)
备注: 57 pages
Abstract:What makes great scientists great? Even as AI systems start to make progress on open problems, scientists remain far ahead of them at sensing which prior idea, buried in an ever-growing archive of research, a new problem needs. To study this skill, we draw on researchers who know firsthand which earlier work advanced their completed projects, with papers serving as pointers to the ideas within. Using our automated pipeline that makes author annotation scalable, we build ScholarCatalyst by having 184 lead authors of 207 recent computer science papers label which candidates did or could have advanced their project, each with a detailed rationale. We introduce a retrieval task with author-provided judgments: given an initial research question, retrieve these papers from only the literature available when the project began. Agentic search does no better than embedding retrieval (0.42 vs. 0.48 Recall@20) despite calling that same retriever as a tool. Even an agent built on Claude Fable 5.1, which may have seen the completed papers during training, reaches only 0.51 R@20. These results highlight the need for new training recipes that equip models with expert intuition for searching broad corpora. We envision ScholarCatalyst as a step toward scientific agents that can take a half-formed idea and point to the prior research it needs.
[IR-1] Optimizing Effective Training Time for Large-Scale Recommendation Systems
链接: https://arxiv.org/abs/2610.02057
作者: Mingming Ding,Ruilin Chen,Yuzhen Huang,Hang Qi,Menglu Yu,San Tan,Damian Reeves,Boris Sarana,Kevin Tang,Satendra Gera,Gagan Jain,Sahil Shah,Vishwa Karia,Fuzail Khan,Yashasvi Makin,Edward Z. Yang,Oguz Ulgen,Jia Chen Ren,Laith Sakka,Mayank Garg,Meet Vadakkanchery,Aici Lin,Wei Sun,Mengjiao Zhou,Shuai Yang,Junqing Zhou,Max Leung,Apoorv Purwar,Musharaf Sultan,John Bocharov,Zhenyu Tang,Vivek Trehan
类目: Information Retrieval (cs.IR)
备注:
Abstract:Lifecycle overhead silently consumes accelerator capacity across large-scale recommendation training fleets. Our largest recommendation workloads process tens of billions train- ing examples per day on thousands of GPUs. Before this work, only 50-60% of their end-to-end wall time advanced training on new data. We present a fleet-scale study of this lifecycle overhead and a set of optimizations spanning the full training stack. We use Effective Training Time (ETT%) as an operational framework to instrument lost time, localize it to independently owned infrastructure components, and expose work repeated across job restarts. This analysis guides optimizations like communication elimination and pipeline overlap during trainer initialization; dynamic-shape handling, autotuning pruning, and reusable Py- Torch 2 compilation caches; asynchronous checkpointing; stan- dalone model publishing; and reductions in recovery cost. We evaluate the optimizations on representative models and measure their impacts in our training fleet. ETT% improves on every benchmark, by 15.5% on average, and reaches 85% on our largest workload. Fleet-wide ETT% rose from about 80% to above 90% after deployment.
[IR-2] A Matryoshka Hierarchical RAG for Efficient Multi-Hop Question Answering
链接: https://arxiv.org/abs/2610.01767
作者: Gianluca Bonifazi,Christopher Buratti,Michele Marchetti,Federica Parlapiano,Giulia Quaglieri,Davide Traini,Domenico Ursino,Luca Virgili
类目: Computation and Language (cs.CL); Information Retrieval (cs.IR)
备注:
Abstract:Retrieval-Augmented Generation (RAG) systems for multi-hop Question Answering (QA) must balance retrieval quality with computational cost. This cost is incurred during indexing time, through the use of expensive Knowledge Graphs (KGs) or Large Language Models (LLMs) to generate summaries, or during querying, through iterative LLM-driven retrieval. To reduce it while maintaining retrieval quality, we present MatRAG, a hierarchical framework that combines RAG systems with Matryoshka Representation Learning (MRL). MatRAG addresses both kinds of cost by aligning the semantic hierarchy of a clustering structure with the nested structure of MRL. Specifically, it organizes the corpus of documents into a Directed Acyclic Graph (DAG) of clusters with progressively coarser granularity. Each level is indexed by a lower Matryoshka dimension. MatRAG pairs an iterative, top-down traversal of the DAG with an entity-driven mechanism that controls the hop budget and re-ranks candidates. We evaluated MatRAG on three standard multi-hop QA benchmarks against seven representative baselines. MatRAG outperforms its strongest competitors in terms of retrieval quality; furthermore, it reduces indexing costs by avoiding KG construction and LLM-based summarization, and lowers query-time costs through dimension-aware similarity.
[IR-3] Agent WebRec: Compact Evidence Fusion over the Agent Web for Personalized Recommendation
链接: https://arxiv.org/abs/2610.01705
作者: Haoran Qiang,Guannan Liu,Liang Zhang,Junjie Wu
类目: Information Retrieval (cs.IR)
备注:
Abstract:LLM-based personal agents are emerging as persistent carriers of user semantics and intermediaries between users and recommendation platforms, maintaining richer user knowledge locally. As agents interact with one another, the conventional \textitUser–Platform relation evolves into a \textitUser–Agent Web–Platform information pathway, enabling distributed user-side information to complement item-side information. This new pathway, however, defies conventional recommendation: evidence is scattered across mutually opaque agents and reachable only through bounded queries, only a small portion of it is relevant to the current recommendation decision, and the responses returned by different agents are semantically heterogeneous. We therefore recast recommendation over the agent web as a \emphtask-time evidence acquisition and fusion problem under a finite evidence budget by deciding what to ask and what to keep, rather than learning from aggregated data. We propose AgentWebRec, a user-agent-oriented framework that progressively acquires and fuses distributed evidence for each user-item decision while keeping underlying agent memories local. It grounds each decision in platform-provided item semantics and task-relevant evidence from the target user agent’s private memory, and conditionally queries neighboring user agents for complementary preference patterns when local evidence is insufficient. Experiments on four InstructRec datasets show that AgentWebRec consistently outperforms baseline recommenders, and ablations verify that the evidence layers contribute complementary gains.
[IR-4] From Rules to Neural Graphs: Scalable Structured Prediction for Patent Prior Art Search ECML KDD2026
链接: https://arxiv.org/abs/2610.01553
作者: Nikolai Zenovkin,Sebastian Björkqvist
类目: Information Retrieval (cs.IR); Computation and Language (cs.CL)
备注: Accepted for publication at the ECML PKDD 2026 conference (Applied Data Science track)
Abstract:Patent search requires processing documents routinely exceeding tens of thousands of tokens. Most neural retrieval approaches operate on truncated inputs, limiting their effectiveness. Graph-based retrieval addresses this by representing each patent as a structured invention graph, but constructing these graphs relies on brittle rule-based parsers. We present the neural parser, which adapts biaffine attention from dependency parsing to predict invention graphs directly from patent text. Our local biaffine attention restricts pairwise scoring to a sliding window, reducing complexity from O(n^2) to O(n \cdot w) . Since local and global scoring share the same weights, the model trains on short sequences and deploys on documents exceeding 40,000 tokens without retraining. Distilled from 1 million rule-parsed documents, it surpasses its teacher at 3 \times lower inference cost: neural graphs improve citation recall by 0.5% on short queries and 1.1% on full documents in a downstream Graph Transformer retrieval system.
[IR-5] Neither Black nor White: Balancing Semantic and Collaborative Signals with Graph-Informed Semantic IDs (GrIS)
链接: https://arxiv.org/abs/2610.01533
作者: Aleksei Medvedev,Alejandro Ariza-Casabona,Steven Derby,Gonzalo Fiz Pontiveros,Xinyang Shao,Florian Spiess
类目: Artificial Intelligence (cs.AI); Information Retrieval (cs.IR); Machine Learning (cs.LG)
备注:
Abstract:Existing work on Semantic IDs (SIDs) for generative recommendation treats SID construction as a representation learning problem: encode items into a quantised latent space and read off codes. We argue this view is incidental. SID construction is, at heart, a recursive clustering problem, and once stated this way the natural object to cluster is a graph whose nodes carry semantic content and whose edges carry collaborative signal; SID assignment becomes a hierarchical graph partition. This reframing yields a unified framework, Graph-Informed Semantic IDs (GrIS), that subsumes prior approaches rather than displacing them. RQ-VAE and RQ-KMeans are recovered as the special case where the graph is empty, exposing content-only quantisation as one corner of a larger design space along two so-far-collapsed axes: graph construction and recursive partition algorithm. We explore two contrasting instantiations: RecDMoN, which performs hierarchical assignment via differentiable graph pooling, and RQ-GAE, which extends RQ-VAE with graph-aware item representations and a graph reconstruction objective. On multiple real-world datasets, GrIS consistently improves over CF-aware SOTA, with gains of up to +52% Hit@10. Because graph construction and partition are explicit, separately configurable components, improvements on either axis can be combined and evaluated systematically.
[IR-6] Learning to structure data from user-generated thematic corpora
链接: https://arxiv.org/abs/2610.01463
作者: Elishay Avram,Oren Glickman,Elad Yom-Tov
类目: Information Retrieval (cs.IR)
备注:
Abstract:Thematic corpora, such as social media communities, contain unstructured text describing data that could be made structured. These include, for example, personal attributes, behaviors, and experiences mentioned in social media data. Extracting structured data is challenging as relevant attributes are often implicit, domain-dependent, and unknown in advance. We propose a fully automated, iterative framework for discovering and extracting domain-specific attribute schemas without a predefined ontology. Using large language models (LLMs), the framework induces candidate attributes, sequentially consolidates semantically overlapping attributes, and assigns a structural type. These enable creating an ontology and populating it with values from the corpus. The framework also enables the use of smaller LLMs for value extraction with estimable accuracy loss compared to large LLMs. We evaluate the framework on 5 health-related Reddit communities. Discovered attributes achieved 61% agreement with human-identified attributes, close to the 62% agreement between independent annotators. In most cases, the algorithm converges to a stable attribute set in fewer than 10 iterations. Structural type assignment achieves 82% accuracy, and value extraction reaches an F1 score of 0.8 compared to human annotations. Across four LLM families, smaller instruction-tuned models show statistically significant improvements in extraction performance with model scale when evaluated against a high-capacity reference LLM, supporting informed accuracy-cost trade-offs. These results show that attributes comparable to those identified by humans can be discovered automatically, enabling the creation of high-quality structured datasets economically and at scale. By removing the need for predefined ontologies, iterative model-driven schema induction offers a practical and scalable foundation for mining thematic corpora.
[IR-7] Not All Is Lost: Repairing Lossy User Preference States of Personalization Encoders NEURIPS2026
链接: https://arxiv.org/abs/2610.01270
作者: Parthiv Chatterjee,Dhiraj Golhar,Ummesalma Diwan,Sourish Dasgupta,Manjunath Joshi,Tanmoy Chakraborty
类目: Machine Learning (cs.LG); Information Retrieval (cs.IR)
备注: Accepted to NeurIPS 2026. Author-prepared archival version with expanded discussion and interpretation. 59 pages, including references and appendices
Abstract:Personalization encoders compress evolving interaction histories into preference states used to rank items or condition text generation. A task head operating only on this state can miss useful evidence that remains in the frozen encoder’s cached representations for individual timesteps. We study this recoverability gap and propose REPAIR, which compares cached representations with the current preference state in a compact learned coordinate space. It resolves corrective evidence over extended history, recent interactions, and localized bursts. It then selects which patterns at which timesteps contribute and adds their aggregate correction to the state before the task head. Encoder-host repair reuses representations from the existing forward computation without re-encoding the history. Across MovieLens, PENS, MIND, and Amazon Reviews 2023, training only REPAIR improves MRR and nDCG@10 for all twelve representative recommendation hosts while both encoder and task head remain frozen. Head-only finetuning of the same hosts yields smaller gains. For example, Mamba4Rec on MovieLens gains 3.96 MRR points, compared with 0.19 from head-only finetuning. Rank and temporal diagnostics support a compact, host-dependent corrective structure. In personalized generation, IMPerSumm improves the two reported weighted PerSEval variants, which assess responsiveness to user preference, by up to 25.23%. These results support post-compression state correction and distinguish the availability of preference evidence from its downstream use.
[IR-8] Do Multilingual Encoders Produce Language-Consistent Semantic IDs? EMNLP2026
链接: https://arxiv.org/abs/2610.01139
作者: Abhinav Bohra,Anuj Bohra
类目: Information Retrieval (cs.IR); Computation and Language (cs.CL)
备注: 7 pages, 8 tables. Accepted as a short paper at WiNLP 2026, co-located with EMNLP 2026
Abstract:Semantic IDs (SIDs) compress item embeddings into discrete code sequences used in generative retrieval. We ask whether a multilingual encoder is sufficient for different-language renderings of the same product to receive language-consistent SIDs. Using Amazon ESCI listings rendered in English, Spanish, and Japanese, we test whether translations remain close to their English source, whether residual quantization is unusually sensitive to translation-induced movement, and whether multilingual or language-balanced quantizer fitting improves SID agreement. Multilingual E5 places translations measurably apart: under an English-heavy fit, a Japanese translation preserves the first SID code of its English counterpart in only 7.7% of cases, compared with 89.0% for an English rewording. Distance-matched product-directed controls produce nearly the same full-SID mismatch as translation, providing no evidence that the quantizer selectively amplifies language directions. Balancing the fitting mixture makes codebook use more uniform but further reduces cross-lingual prefix agreement: Spanish first-code consistency falls from 28.3% to 6.6%, while an English-only fit preserves it for 67.6% of Spanish translations. These results show that multilingual exposure and balanced codebook use alone do not guarantee language-consistent SIDs.
[IR-9] Madeleine: Learning Involuntary Recall for Conversational Memory from Simulated Lives
链接: https://arxiv.org/abs/2610.01118
作者: Zhiyun Shi
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Information Retrieval (cs.IR)
备注: 17 pages, 4 figures
Abstract:A long-term conversational assistant must recall the right memory at the right moment, yet the memory that matters most is often not similar to what the user says now. Current systems recover such associations by letting an LLM reason at write or read time, at a cost of hundreds to over a thousand LLM calls per memory bank and up to several thousand context tokens per query. We argue that association is a learnable relevance: the pointwise mutual information of memories under how human lives unfold. We introduce Madeleine, which learns amortized association: offline, an LLM life simulator writes simulated lives, whose cue-trigger pairs teach a query encoder a residual association on top of frozen similarity; online, it calls no LLM and plugs into any vector memory by replacing only the query encoder. On LoCoMo-Plus under the official protocol, Madeleine (I) reaches 66.6 when plugged into HyperMem, the highest among all systems evaluated under this protocol; (II) used alone, reaches the score of HyperMem as released (52.4 vs. 52.9) with zero LLM calls and about 1/21 of its answer context; and (III) lifts T-Mem by 26.2 points, significantly outperforms the same untrained backbone inside both systems, and leaves ordinary QA intact on the 4B backbone.
[IR-10] JoinGR: Learning to Traverse Join Graphs for Table Retrieval
链接: https://arxiv.org/abs/2610.01064
作者: Sandipan De,Abhijit Chakraborty,Sambaran Bandyopadhyay,Vivek Gupta
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Databases (cs.DB); Information Retrieval (cs.IR)
备注: 12 pages, 6 figures, 5 pages
Abstract:Retrieving the right tables is a prerequisite for Text-to-SQL over realistic databases. Dense table retrievers rank schema elements independently, but this ignores a key source of evidence: some required tables are not mentioned in the question and become identifiable only through their join relationships to already relevant tables. We introduce JOINGR, a join-aware table retrieval method that treats the database join graph as the retrieval space. Columns are represented as graph nodes, while intra-table and foreign-key relationships are represented as typed edges. Given a question, JOINGR selects semantically similar anchor tables, traverses join edges with a query-conditioned scorer, and aggregates the resulting edge deposits into table scores. The scorer is a lightweight MLP on top of frozen query, node, and edge embeddings, trained with a pairwise margin loss over gold tables. On BIRD and Spider datasets, JOINGR is competitive with the strongest retrieval baselines. On BEAVER, a challenging enterprise benchmark with multi-hop table requirements, JOINGR substantially improves recall over dense retrieval and re-ranking baselines. Cross-domain experiments show that the learned scorer transfers across benchmarks, indicating that the method captures reusable joingraph traversal behavior.
[IR-11] he Other Half of Workflow Portability: Evidence-Backed HPC Site Profiles with Agent ic Discovery
链接: https://arxiv.org/abs/2610.00971
作者: Md Saiful Islam,Douglas Thain
类目: Distributed, Parallel, and Cluster Computing (cs.DC); Information Retrieval (cs.IR)
备注: Accepted to the 21st Workshop on Workflows in Support of Large-Scale Science (WORKS 2026), held with SC26, Chicago, IL, USA. 8 pages, 7 figures, 3 tables
Abstract:Moving a workflow developed and tested at one HPC site to another rarely succeeds without some amount of trial and error. Package managers rebuild software environments, containers ship whole filesystems, and workflow specifications such as backpacks package a workflow with its software, data, and resource requirements. These approaches address one half of workflow portability: what a workflow needs. But none describes how a given HPC site must be used, and that missing half is why even a portable workflow requires manual adjustment at each new site. That gap includes the site’s resource shape, storage configuration, network permissions, and operating policies. This information may be explicit in the batch system, hidden in the prose of documentation, or buried deep within a router’s configuration, making it difficult for an automated deployment tool to turn site knowledge into useful deployment decisions. We propose the HPC site profile, a structured, evidence-backed document that makes this knowledge actionable. We automatically construct it in three steps that mirror where the information lives: measuring the login node, extracting typed fields from documentation with a bounded language-model agent, and submitting pilot jobs for eligible unresolved fields. Every field is verified against its evidence or discarded, so a rule, not the model, decides what enters the profile. The profile then preflights a workflow into an execution plan or an early, explainable failure. We build profiles at Purdue Anvil, TACC Stampede3, and Notre Dame CRC and present a case study of preflighting a real workflow.
[IR-12] RPTune: Learned Context Curation for LLM Catalog Search
链接: https://arxiv.org/abs/2610.00964
作者: Chuxuan Hu,Hejie Cui,Norman Huang,Shubham Kumar Bharti,Wang-Chiew Tan,Sercan Ö. Arık
类目: Information Retrieval (cs.IR); Computation and Language (cs.CL); Machine Learning (cs.LG)
备注: 23 pages, 9 figures, 4 tables
Abstract:For small merchant businesses (SMBs) whose catalogs fit within a long-context LLM, full-catalog prompting offers a compelling alternative to multi-stage retrieval designed primarily for large marketplaces with millions of items. However, fitting the full catalog into the context window does not ensure that the model can use it effectively, since LLMs do not exploit long contexts uniformly. We therefore study in-context catalog search through two complementary questions: (1) how to curate and present catalogs to the LLM, and (2) how to adapt the LLM for product selection on curated contexts. We propose RPTune, an end-to-end framework that couples learned catalog curation with LLM post-training using automatically generated, catalog-grounded supervision. An encoder-reorganizer curator orders and prunes products guided by downstream LLM feedback, while the resulting curated catalogs in turn improve the effectiveness of LLM post-training with a context-relative reward. We evaluate RPTune on 7 real merchants spanning distinct retail verticals, using 100 complex conversational queries per merchant. RPTune consistently improves search accuracy across both proprietary and open-weight LLMs, with context curation yielding gains of up to 31.4 percentage points and post-training adding a further 10.3 points on average. Comments: 23 pages, 9 figures, 4 tables Subjects: Information Retrieval (cs.IR); Computation and Language (cs.CL); Machine Learning (cs.LG) Cite as: arXiv:2610.00964 [cs.IR] (or arXiv:2610.00964v1 [cs.IR] for this version) https://doi.org/10.48550/arXiv.2610.00964 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[IR-13] CANOPY: Adaptive-Granularity Evidence Compression for Multimodal RAG
链接: https://arxiv.org/abs/2610.00923
作者: Hyojeong Yun,Jueun Kim,Wook-Shin Han
类目: Information Retrieval (cs.IR)
备注: 26 pages, 10 figures, project page: this https URL
Abstract:Multimodal RAG retrieves text, tables, images, and videos, but choosing a retrieval granularity does not determine how much context to retain within each item. Coarse units include irrelevant content, while uniformly fine selection can remove context needed to interpret the evidence. Existing compressors address this trade-off with modality-specific mechanisms, leaving open a shared procedure for adapting the retained extent region by region across heterogeneous items. We introduce CANOPY (Canonical Projection over Hierarchy), a framework for adaptive-granularity post-retrieval evidence compression. CANOPY represents retrieved items as hierarchies and uses a node encoder fine-tuned on gold evidence to score regions against the query. Parent-relative refinement compares these scores to select multiple regions at different granularities without LLM calls for node-level pruning. Because compression cannot recover evidence that was never retrieved, a critic requests targeted follow-up retrieval when it judges the accumulated evidence insufficient; newly retrieved items are compressed before being added. Across five QA benchmarks over a 33M-item heterogeneous corpus, CANOPY achieves higher average answer accuracy than the evaluated retrieval baselines. Ablations indicate that additional retrieval drives the main accuracy gains on multi-hop QA. In the unrouted Qwen3-VL-8B-Instruct setting, compression reduces reader-input evidence tokens by 14.2-27.7% relative to the same iterative pipeline without compression, with comparable answer accuracy.
[IR-14] abJoinBench: A Benchmark for Joinable Table Discovery
链接: https://arxiv.org/abs/2610.00817
作者: Sandipan De,Jin Wang,Vivek Gupta
类目: Databases (cs.DB); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Information Retrieval (cs.IR)
备注: 13 pages, 8 Tables, 1 Figure
Abstract:Join discovery aims to identify tables from large data repositories that can augment a query table with complementary information, enabling downstream tasks such as data exploration, feature engineering, and business intelligence. Although numerous join discovery methods have been proposed, existing studies rely on method-specific benchmark construction, making reproducible and fair comparison difficult. We present TabJoinBench, a benchmark for evaluating join discovery methods across semantic, relational, and hybrid data lake scenarios. TabJoinBench constructs query-candidate pairs using source-specific validation strategies, systematically introduces structural, representation, and semantic changes through composable perturbations while preserving reliable ground truth. We evaluate representative join discovery methods spanning set-based, feature-based, and learned approaches, together with general-purpose language-model embedding baselines, and publicly release the processed datasets, ground-truth annotations, and generation pipeline to facilitate reproducible evaluation and future research.
[IR-15] Enterprise Representation Simplification (ERS): Reducing Representational Complexity for Enterprise AI
链接: https://arxiv.org/abs/2610.00791
作者: Terry Dorsey,Kevin Huggins
类目: Artificial Intelligence (cs.AI); Information Retrieval (cs.IR)
备注:
Abstract:Enterprise information is represented through artifacts shaped by applications, projects, technologies, organizational boundaries, and local requirements. These structures accumulate over time, creating representational complexity that must be maintained by the enterprise and interpreted by information consumers and AI systems. This paper introduces Enterprise Representation Simplification (ERS) as reducing unnecessary representational complexity while preserving required information within a defined scope, and Enterprise Representation Complexity (ERC), a representation-neutral model for comparing complexity across representation states. ERC characterizes representational extent through four dimensions: Representation Objects, Interactions, Behaviors, and Supporting Sources. Objects, Interactions, and Behaviors form dependent categories, while Supporting Sources characterize representation exposure. ERC is defined at representation and task levels, enabling comparison and distinguishing architectural simplification from retrieval optimization. The paper develops two consequences of ERS. First, representational structures create lifecycle obligations for maintenance, governance, dependencies, change, enhancement, and operation. An economic model distinguishes recurring global representation cost, recurring task-level cost, and one-time transformation cost, enabling evaluation over a defined time horizon. Second, reductions in task-level ERC reduce the representational extent an AI system must identify, relate, and interpret. Text-to-SQL research provides evidence that reduced schema and reasoning complexity can improve reasoning accuracy. ERC is not a universal complexity, performance, or cost metric. It provides measurable architectural variables for comparing representational alternatives, transformation effects, economic outcomes, and AI reasoning performance. Subjects: Artificial Intelligence (cs.AI); Information Retrieval (cs.IR) Cite as: arXiv:2610.00791 [cs.AI] (or arXiv:2610.00791v1 [cs.AI] for this version) https://doi.org/10.48550/arXiv.2610.00791 Focus to learn more arXiv-issued DOI via DataCite (pending registration) Submission history From: Terry Dorsey [view email] [v1] Wed, 30 Sep 2026 22:31:40 UTC (477 KB)
[IR-16] Comparison of Common Crawl News GDELT
链接: https://arxiv.org/abs/2610.00587
作者: Ameir El Ouadi,David Beskow
类目: Information Retrieval (cs.IR)
备注:
Abstract:The corpus of worldwide news is important for natural language processing, knowledge graphs, large language models, and other technical efforts. Additionally, this corpus is important for understanding the people, places, organizations, and events that interact in real-time every day. This paper compares two news datasets used for these tasks today, namely the Global Database of Events, Language, and Tone (GDELT) and Common Crawl News. Our research highlights the strengths and limitations of each dataset, analyzing their content and coverage. Notably, while GDELT relies on broadcasts, prints, and web news from across the globe, Common Crawl focuses on news sites from around the world gathered through web crawling. Our analysis revealed considerable differences in where the two datasets gather their news sources.
[IR-17] A Shared Taste for Model-Written Text: The Generator-by-Selector Matrices of “AI-AI Bias” Show No Detectable Own-Model Premium
链接: https://arxiv.org/abs/2610.00369
作者: Dmitrij Żatuchin
类目: Information Retrieval (cs.IR); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: 10 pages, 4 figures, 3 tables. Reanalysis of publicly available generator-by-selector matrices
Abstract:Laurito et al. (PNAS 2025) showed that large language models choosing between two descriptions of the same product, paper or film prefer the description written by a language model over the one written by a person, by a wide margin over what human judges do. Their design crosses five generators with the same five models as selectors, which permits a second question the paper does not headline: does a selector prefer text from its own model beyond what the generator and selector main effects predict? We rebuild the three 5x5 matrices from the per-item counts in the authors’ public repository (21,828 valid trials; every cell matches the published value) and fit a two-way fixed-effects model with an own-model term gamma, tested by the exact permutation test over the 120 relabellings of the selectors. The premium is +0.013 on products (exact one-sided p = 0.24), -0.010 on paper abstracts (p = 0.74), +0.054 on films (p = 0.07) and +0.019 pooled (p = 0.14; 95% interval -0.008 to 0.046). The same-vendor term for the GPT-3.5 and GPT-4 pair is negative in all three datasets. Position bias moves single cells by up to 0.42 share points in either direction, and the own-model contrast is unchanged once order-driven items are removed. The design would have detected a premium of 0.05 with 82% (products), 88% (papers), 42% (films) and 97% (pooled) power; the minimum detectable effect at 80% power is 0.034 pooled. The absence is informative down to about 0.04 share points and silent below that. The 4x4 matrix of Tan et al. (ACL 2024) gives gamma = +0.148 at the smallest p its 24 relabellings allow, with a same-family term of the same size. The main result of Laurito et al. stands: models share a taste for model-written text, with GPT-4’s descriptions chosen 77% to 95% of the time by every selector on products. What these data do not show is a model recognising and favouring its own prose.
[IR-18] System Attribution in LLM Brand Recommendations: Single Responses Identify the System Aggregated Brand Profiles Do Not Transfer
链接: https://arxiv.org/abs/2610.00253
作者: Dmitrij Żatuchin
类目: Information Retrieval (cs.IR); Computation and Language (cs.CL); Machine Learning (cs.LG)
备注: 30 pages, 5 figures, 9 tables. Appendix D documents corrections to an earlier manuscript
Abstract:Audits of AI visibility summarise the brand recommendations of deployed language models into per-system profiles. We test whether such a profile describes the system on one corpus of 6,475 stored responses (6,324 analysable) collected between December 2025 and February 2026 from five deployed endpoints across gift-recommendation, corporate-reputation and category-ownership queries. The collection harness cut many answers short: 83.1% of Gemini 3 Flash answers in category ownership end mid-sentence under a 1,024-token output cap. With every answer cut to its first 800 characters, a character n-gram classifier cross-validated by prompt attributes one response to GPT-5.2, Gemini 3 Flash, Gemini 3 Flash with search, Grok or Perplexity sonar-pro with 97.84% accuracy (5,028 responses, 383 prompts, majority class 31.5%, 30 split seeds). Length alone falls to the majority rate, 24 formatting statistics reach 95.79%, and masking brand names and capitalised tokens leaves 97.72%. Held-out query conditions keep 97.43% weighted by size and 88.0% unweighted; in a retrieval-grounded arm that changes the harness, no Grok answer is attributed to Grok (0/120). Aggregated into 50 model-by-domain-by-condition units, twelve behavioural features separate four systems at 66.53% under grouped cross-validation, against a label-permutation null with mean 33.71% and 95th percentile 46.0%. Across domains the aggregate profile fails: a forest trained on category-ownership units assigns all 22 gift units to the wrong system, consistent with a reversal in brand volume (8.41 against 0.94 brands per response in gifts, 3.01 against 3.91 in category ownership), while single responses transfer at 89.92% balanced accuracy. The surface form of an answer carries the system across the query domains tested; aggregated brand behaviour does not, and the uncrossed design cannot separate the system from the domain or the harness.
[IR-19] On-Device Commercial Intent Retrieval Under Size Latency and Privacy Constraints: A 3 MiB Retrieval System with Typed Egress Boundaries
链接: https://arxiv.org/abs/2610.00170
作者: Hyojung Han
类目: Information Retrieval (cs.IR); Computation and Language (cs.CL)
备注: 36 pages, 40 tables, 1 figure. Evidence bundle (measurement ledgers and verification gates, no datasets): this https URL
Abstract:We study commercial intent inference that runs entirely on the user’s device, under three constraints frozen before the work began: the downloaded payload under 3 MiB, Tier-0 inference under 20 ms at p95, and no raw text, content embedding, or stable identifier leaving the device. Under them we build a retrieval path over a 6,020-leaf commercial taxonomy: a static embedding table distilled from a Korean sentence transformer, quantized to 4 bits, no inference runtime. Our main result is where that constraint costs accuracy. On real Korean commerce text labelled by others (22,900 AI-Hub shopping reviews), mid-category top-5 on real product names is 75.0% against an 18.4% permutation baseline, but splits on one observable: a query containing some leaf name as a substring scores 83.5%, one containing none 45.2%. A generic 196.6x larger teacher seemed to localize the gap (+20.1 pp without an anchor, +0.1 with). That null was two effects cancelling: the same teacher fine-tuned on the student’s own contrastive pairs reaches 0.8586 and beats the pure-encoder student by +10.6 pp with an anchor and +20.9 pp without. The cost is not uniform, but it is not free anywhere; where the anchor is absent, task adaptation buys the teacher nothing, so what the constrained encoder lacks there is capacity. The expensive regime is detectable on-device from the ranker’s own score margin: declining the least confident fifth lifts the rest to 0.8296. A second axis we first reported, a manufacturer model code, does not survive source-category fixed effects (-4.0 pp, p=0.51); the anchor does (+13.0 pp). Payload is 2,942,652 bytes, all three library links measured. Tier-0 p95 is 4.431 and 3.670 ms on two iPhones (A14, A16) and 5.080 ms on a budget Android tablet (Snapdragon 695), all slower than three server CPUs on the same code. Taxonomy supervision is mostly synthetic Korean utterances. Comments: 36 pages, 40 tables, 1 figure. Evidence bundle (measurement ledgers and verification gates, no datasets): this https URL Subjects: Information Retrieval (cs.IR); Computation and Language (cs.CL) Cite as: arXiv:2610.00170 [cs.IR] (or arXiv:2610.00170v1 [cs.IR] for this version) https://doi.org/10.48550/arXiv.2610.00170 Focus to learn more arXiv-issued DOI via DataCite (pending registration) Submission history From: Hyo Jung Han [view email] [v1] Wed, 16 Sep 2026 00:34:01 UTC (74 KB)
[IR-20] Ask a Language Model for Lottery Numbers: Concentration in Repeated Six-of-49 Outputs
链接: https://arxiv.org/abs/2610.00052
作者: Dmitrij Żatuchin
类目: Information Retrieval (cs.IR); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: 6 pages, 1 figure. Data, code, and collector at this http URL (research/lotto-models)
Abstract:We evaluate six language-model configurations on requests for six distinct random integers from 1-49. Across 1,200 attempted calls using four English prompt variants, 1,184 responses yielded valid tickets. Effective diversity of number frequencies ranged from 9.9 to 18.0, compared with simulated fifth-percentile thresholds of 46.6-46.7 under independent uniform six-of-49 sampling at the corresponding sample sizes. Systems produced 8-93 distinct unordered tickets, and their modal tickets accounted for 22.5-68.0% of valid responses. Two archived Polish Lotto samples provided a physical-lottery comparison, with effective diversities of 41.1 and 41.4 at smaller sample sizes. These results demonstrate substantial concentration under the tested deployment settings. They do not identify its mechanism or establish performance under other prompts, temperatures, or tool configurations.
人机交互
[HC-0] One Basis to Animate Them All: Gaussian Blendshape Distillation for Real-Time Avatars
链接: https://arxiv.org/abs/2610.02207
作者: Ramazan Fazylov,Stamatis Lefkimmiatis,Ivan Laptev
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Human-Computer Interaction (cs.HC); Machine Learning (cs.LG)
备注:
Abstract:3D Gaussian avatars support fast rendering, however, their real-time animation is often challenged by the costly neural inference. We address this bottleneck and show that the animation of pretrained avatar models can be closely approximated by a linear combination of identity-independent blendshapes. Building on this finding, we introduce GALA (Gaussian Animation via Linear Approximation), a distillation method that replaces per-frame heavy neural decoding with a shallow coefficient predictor and a linear blend. To improve fidelity and reduce memory requirements, we propose to construct the basis using block-local PCA under a rendering-aware metric and a memory budget. Our method learns a shallow MLP network to predict blendshape coefficients and applies to various animation architectures without retraining original models. We validate GALA by accelerating the inference of three distinct avatar models for 3D animation of facial expressions and full-bodies with clothing dynamics. Across these models, our distillation generalizes to held-out identities and reduces CPU animation cost by up to three orders of magnitude while preserving most of the rendering quality. Excellent results of our method confirm the shared linear structure of learned avatar representations and enable highly efficient and accurate animation at frame rates reaching up to 60fps on mobile devices. Project page: this https URL
[HC-1] Catscan: Visualizing Pipelines of CPU Performance Simulation
链接: https://arxiv.org/abs/2610.02121
作者: Aaron Lindsay,Nicholas Kelly,Scott Witscher,Mahesh Madhav
类目: Hardware Architecture (cs.AR); Human-Computer Interaction (cs.HC); Performance (cs.PF)
备注:
Abstract:Processor pipeline visualization tools are routine inside industry CPU teams, but few of them are described or released publicly. As a result, students, researchers, and other practitioners rarely see the tooling that processor architects use to debug performance before silicon. This paper describes two pieces of Ampere Computing’s performance- analysis infrastructure that we have released to the community as open source: event streams, a simulator-output format, and Catscan, an interactive viewer built around that format. Event streams record microarchitectural activity as typed events connected by transaction relationships, so a user can move between a symptom and the instruction, uop, or memory transaction that explains it. Catscan uses that structure to support resource- and transaction-oriented views, persistent highlighting, domain-specific search, comparative trace synchronization, and other workflows used during product development. In this paper we report the design choices that survived production use, the limitations we encountered, and the lessons we think are useful for future microarchitectural visualization tools.
[HC-2] SPHERE: Adaptive VR Indoor Scene Generation via LLM -Enhanced Spatial Preference Learning and Human-in-the-Loop RL
链接: https://arxiv.org/abs/2610.02023
作者: Hyeonmin Lee,Zheng Wei,Kyungmin Kwon,Jumin Seo,Jiwon Park,Hayoung Oh
类目: Artificial Intelligence (cs.AI); Human-Computer Interaction (cs.HC)
备注:
Abstract:While Large Language Models (LLMs) advance 3D indoor scene synthesis, current pipelines fail to retain user-specific preferences across sessions, making immersive authoring a repetitive and physically fatiguing process. We present SPHERE, an adaptive VR generation framework that transforms isolated synthesis into continuous human-AI co-creation. SPHERE extracts persistent spatial preferences from natural multimodal interactions (speech and controller edits). To ensure geometric resilience against spatial distortions, it abstracts these raw edits into hierarchical constraints modeling both local functional and global topological contexts. Furthermore, a human-in-the-loop reinforcement learning mechanism dynamically updates retrieval policies based on the user’s final edited scenes. A mixed-design user study ( N=42 ) and an offline ablation demonstrate that SPHERE significantly reduces corrective edits and physical demand, preventing bias toward shallow object-level traits to yield geometrically resilient, profile-aligned layouts. Ultimately, SPHERE demonstrates how capturing demonstrated spatial logic enables controlled spatial adaptation, establishing a reliable, governed human-AI collaboration framework for immersive authoring. Project page and source code will be available at: this https URL
[HC-3] XAI Evaluation Cards: A Practical Method for Designing Human-Centred XAI Evaluations
链接: https://arxiv.org/abs/2610.02011
作者: Kristýna Sirka Kacafírková,Ivania Donoso-Guzmán,Denis Parra,Katrien Verbert,An Jacobs
类目: Human-Computer Interaction (cs.HC)
备注:
Abstract:Evaluating explainable AI (XAI) systems from a human-centred approach requires researchers to select from numerous evaluation dimensions and measures, often in an ad hoc and fragmented manner. This paper introduces a method to help HCI, computer science, designers and social science researchers systematically evaluate XAI systems. The approach is based on an updated XAI-specific evaluation framework derived from an analysis of 82 studies. Using this framework, we developed a card-sorting method with 36 cards to help researchers prioritise relevant evaluation aspects. The process was tested with two research groups (n = 13) across five projects. The XAI Evaluation Cards are available as a printable appendix, along with an online repository of methods from previous XAI studies. Although not exhaustive, our findings indicate that the card-sorting approach can organise and streamline the design of the evaluation process, encouraging a more comprehensive and multidisciplinary assessment of XAI systems in research and development.
[HC-4] Interactive Power Flow in the Browser
链接: https://arxiv.org/abs/2610.01922
作者: Samuel Talkington,Frederik Geth,Qian Zhang,Le Xie,Skyler Liu
类目: ystems and Control (eess.SY); Human-Computer Interaction (cs.HC); Mathematical Software (cs.MS)
备注: 9 pages, 3 figures, 6 tables. To appear in the Inaugural ACM Conference on Digital Transformation (ACM DXConf 2026), Ann Arbor, MI, USA
Abstract:This paper introduces tellegen, an open source framework for interactive power flow (PF) and optimal power flow (OPF) studies that run in the browser. This provides intuitive and democratized access to power system analysis tools compiled to WebAssembly. A user can drag and drop a case file, click and drag to change a nodal demand or line rating, preview the impacts via sensitivity analysis, and obtain an exact re-solve on release. User case files and results stay entirely on the device: tellegen transmits zero Critical Energy/Electric Infrastructure Information (CEII). The framework comprises a core numerical engine for PF and OPF, reusable browser components, saved studies, and structured WebMCP tools for agentic interaction. We evaluate the transmission OPF solver by comparing objectives with PGLib baselines; the distribution PF solver by comparing voltages and currents with OpenDSS; and the WebAssembly execution times by comparing with this http URL. On realistic synthetic grids, WebAssembly OPF solves take only 25-43% longer than native binary solves. The implementation shows how an engineer can distribute an executable numerical study as a URL, reducing installation and hosting requirements while keeping case data local.
[HC-5] Where LLM s Fail with Visualization DSLs
链接: https://arxiv.org/abs/2610.01873
作者: Chang Han,Andrew McNutt,Katherine Isaacs
类目: Human-Computer Interaction (cs.HC); Computation and Language (cs.CL)
备注: VIS 2026 VISxGenAI, 6 pages, 3 figures
Abstract:As LLMs take up the role of authoring charts using visualization domain-specific languages (DSLs), the human constraints that shaped those languages may no longer apply, as what is easy for a person is not necessarily easy for a model. To understand how LLMs might work better with DSLs, we explore where and how they fail with current DSL designs. We evaluate 10 JSON-style visualization DSLs with 41 tasks across 3 LLMs, then assess the generated specifications with JSON and rendering checks, and qualitative coding of failed cases. Analyzing how this specification generation process fails, we identify four recurring failure patterns, link each to specific DSL features, and discuss design considerations for future DSL designs.
[HC-6] Designing for Interpretation Uncertainty: Architecture and Principles for Topological Learning Analytics Dashboards
链接: https://arxiv.org/abs/2610.01749
作者: Hitoshi Inoue,Koichi Yasutake
类目: Computers and Society (cs.CY); Human-Computer Interaction (cs.HC); Machine Learning (cs.LG)
备注: Author’s version, posted under the preprint/reprint distribution rights retained in the IADIS copyright transfer agreement
Abstract:Topological Data Analysis (TDA) offers novel methods for understanding temporal dynamics in complex systems, yet its application in information systems design faces a fundamental challenge: how should systems present analytical outputs when interpretation frameworks are still developing? This paper reports on the development of TopoLA, a dashboard system applying Zigzag Persistent Homology to learning management system data, and proposes three early design principles for interpretation support in emerging analytics: (1) separation of objective measurement from contextual interpretation, (2) graduated disclosure from metrics through patterns to reflective prompts, and (3) explicit acknowledgment of methodological uncertainty. The system implements a modular three-stage pipeline–feature extraction, topological computation, and interpretation support–enabling extension to additional analytical methods. This work contributes to information systems research by articulating preliminary design knowledge for systems that must communicate analytical insights from methods lacking established interpretation norms–a challenge increasingly common as novel computational techniques enter applied domains.
[HC-7] Who Thinks First? Designing Productive Friction with Engage-to-Unlock GenAI
链接: https://arxiv.org/abs/2610.01518
作者: Xiaotian Su,Laura Rimell,Jiazheng Li,Amal Rannen-Triki,Ulrich Paquet,Lisa Anne Hendricks,Rida Qadri,Daphne Ippolito,Piotr Mirowski
类目: Human-Computer Interaction (cs.HC)
备注: 31 pages, 7 figures, 3 tables
Abstract:Generative AI can support writing, but frictionless access may cause cognitive offloading before users develop their own ideas. We introduce Engage-to-Unlock, a productive-friction mechanism that unlocks generative capabilities after users meaningfully engage with the task. In a controlled experiment (N = 398), participants completed a writing task under one of four conditions: Human-Only, Standard Chatbot, Engage-to-Unlock, or Time-Matched Unlock, which matched unlock timing to Engage-to-Unlock participants but independent of users’ engagement, then evaluated passages for evidence and inferential errors. Results show that Engage-to-Unlock redistributed effort across tasks: participants spent more time writing and less time evaluating, without increasing overall task duration. They also submitted more prompts than in other AI-assisted conditions and showed the highest accuracy-per-time evaluation efficiency across conditions. These findings suggest that designing GenAI access to encourage early human engagement may provide a productive form of friction, while retaining active AI use and efficient downstream evaluation.
[HC-8] bUMAP: Coherent and Scalable Field Evaluation for UMAP Optimization ICLR2027
链接: https://arxiv.org/abs/2610.01445
作者: Bin Chen,Yumeng Xue,Patrick Paetzold,Yunhai Wang,Oliver Deussen
类目: Machine Learning (cs.LG); Human-Computer Interaction (cs.HC)
备注: 35 pages, 11 figures, 18 tables. Under review at ICLR 2027. Code: this https URL
Abstract:UMAP achieves scalable layout optimization through stochastic negative sampling. However, this stochasticity can lead to unstable embeddings across reruns and downstream reuse, as the estimated repulsive forces depend on the ordering of sampling events. We present ibUMAP, a coherent field-based alternative that evaluates attraction and repulsion from a shared embedding snapshot and applies them synchronously. Its degree-weighted repulsive field is motivated by the conditional expectation of negative sampling for a fixed embedding and represented by three scalar moments, which are evaluated efficiently on CPUs and GPUs using an interpolation-based FFT scheme. This formulation avoids explicit all-pairs computations while inducing optimization dynamics that differ from those of standard online UMAP. Controlled experiments show that synchrony and kernel capping alter the local-global fidelity trade-off, whereas FFT evaluation produces small average changes in final quality. End-to-end benchmarks show median speedups of 3.29x unseeded and 5.79x seeded over umap-learn on CPU, and 1.44x over cuML on million-scale datasets under unseeded GPU execution. These gains accompany greater run-to-run stability and measurable fidelity trade-offs.
[HC-9] PROMO: Preference-conditioned Multi-Objective Reinforcement Learning for Quadrupedal Robots
链接: https://arxiv.org/abs/2610.01260
作者: Amr Mousa,Rifny Rachman,Neil Karavis,Michele Caprio,Richard Allmendinger
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI); Human-Computer Interaction (cs.HC); Machine Learning (cs.LG); Systems and Control (eess.SY)
备注: Submitted to IEEE Transactions on Robotics. Project website, code, and videos: this https URL
Abstract:Quadrupedal locomotion requires balancing conflicting objectives such as command tracking, stability, and energy efficiency, yet conventional reinforcement learning (RL) hardcodes these priorities into a fixed scalar reward at training time. We present PROMO (Preference-Conditioned Multi-Objective Reinforcement Learning), a semantic multi-objective approach that makes this trade-off an explicit runtime input to a single locomotion policy. PROMO conditions the policy on deployment facing preferences while keeping embodiment-specific locomotion priors fixed, thereby separating operator intent from reward shaping terms required for viable gait generation. Compared with fixed-objective controllers, multi-objective baselines, and independently trained specialists, PROMO achieves objective specialization and robustness from a single deployable policy. Across 100 sampled preferences in simulation, 67 behaviors are non-dominated under exact Pareto dominance, with a mean preference-objective correlation of 0.843, demonstrating broad Pareto coverage and predictable preference response. The same policy transfers zero-shot to a Unitree Go2, where preference changes alone reduce specific energy by up to 30.4%, position error by 38.7%, and peak body-attitude deviation by 59.0% relative to the balanced preference. These results establish preference-conditioned multi-objective RL as a practical runtime interface for adaptive legged locomotion, extending its role beyond offline Pareto-set construction. Open-source code and videos are available at this https URL.
[HC-10] WIP: DBWorkout: A Gamified SQL Practice Platform to Support Formative Learning in Database Courses
链接: https://arxiv.org/abs/2610.01174
作者: Sehrish Basir Nizamani,Deepika Devaraj,Tien Nguyen,Khyati Goyal,Saad Nizamani,Sally Hamouda,Jaren Goldberg
类目: Computers and Society (cs.CY); Databases (cs.DB); Human-Computer Interaction (cs.HC)
备注: Work-in-progress paper accepted to the 2026 IEEE Frontiers in Education Conference (FIE 2026). 1 figure, 1 table. © 2026 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses
Abstract:This research WIP paper presents DBWorkout, a web-based platform that supports formative SQL learning through sandbox-based execution, automated result-based feedback, and session-based gamification. Learning Structured Query Language (SQL) remains challenging for undergraduate students due to limited opportunities for interactive practice and immediate feedback. Students iteratively practice SQL on live database instances while receiving multi-dimensional feedback on query correctness, including row values, column structure, and ordering. To reduce instructor workload, DBWorkout incorporates large language model (LLM)-assisted tools for schema and task generation within a human-in-the-loop workflow. A pilot study with teaching assistants and a classroom deployment involving 170 undergraduate students across two in-class sessions show strong perceived learning value (90% agreement) and engagement (87% enjoyment), alongside low reported pressure (22%). However, only 42% of students found the automated feedback sufficiently actionable, a finding independently corroborated by 40% of open-ended responses raising feedback quality concerns, providing cross-method triangulation of this gap. These findings demonstrate the technical feasibility and early pedagogical potential of DBWorkout while identifying directions for enhancing feedback quality and supporting sustained SQL learning.
[HC-11] Understanding Student Use of Large Language Models Across Computer Science Subfields
链接: https://arxiv.org/abs/2610.01158
作者: Sehrish Basir Nizamani,Yoonje Lee,Nikitha Donekal Chandrashekar,Margaret Ellis,Naren Ramakrishnan
类目: Computers and Society (cs.CY); Human-Computer Interaction (cs.HC)
备注: Accepted to the 2026 IEEE Frontiers in Education Conference (FIE 2026). 3 figures, 4 tables. © 2026 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses
Abstract:This research full paper examines how undergraduate students use large language models (LLMs) across computer science subfields. As LLMs become increasingly integrated into computing education, understanding how their use varies across technical and pedagogical contexts is essential for designing effective, subfield-aware instruction. This paper presents a cross-subfield analysis of LLM usage among 211 undergraduate students in a problem-solving course intentionally designed to support responsible and effective LLM use through structured instruction and reflection. Using post-assignment reflection data collected across seven instructional modules spanning multiple computer science subfields, we examine prompt counts, LLM role conceptualization, and verification behavior. Results show that LLM adoption varies substantially by assignment, with higher usage in algorithms and web development and lower usage in software engineering. Students predominantly treat LLMs as assistive tools rather than authoritative sources, and verification is common across all subfields, with most students using multiple strategies. Verification behavior also varies by assignment context, with testing more common in structured tasks and web search more common in open-ended tasks. These findings suggest that assignment characteristics play a central role in shaping how students interact with and evaluate LLM outputs, even under a single, consistently applied instructional design. This work contributes empirical evidence on how LLM adoption, role conceptualization, and verification behavior vary across computer science subfields, extending our prior work on structured, reflective LLM instruction to show how its effects differ by task rather than only in aggregate.
[HC-12] Scaling Peer Assessments: An Integrity Report from a Large Engineering Internship
链接: https://arxiv.org/abs/2610.01020
作者: Jinal Gupta,Pavani Ayinampudi,Aditya B.M.V.,Prakash Hegade,Rohit Sharma,Sakshi Sharma,Meenakshi V,S.R.S. Iyengar
类目: Human-Computer Interaction (cs.HC); Computers and Society (cs.CY)
备注: 12 pages, 3 figures, 5 tables; submitted to ICTIEE and under review
Abstract:Assessing learning in large classrooms presents a significant challenge for individual instructors, who may have limited capacity to evaluate the understanding, participation, and assessment behaviour of every student. Peer assessments have been a way of distributing this responsibility among learners, allowing them to evaluate and provide feedback to one another while reducing dependence on instructor-led assessments. Building on this approach, we implemented a peer validation model within a large, multi-institutional internship programme in which students who demonstrated sufficient understanding were authorised to assess and validate their peers through short oral discussions. The assessment process began with the instructor validating a small group of students, who were then authorised to validate their peers, allowing the process to gradually expand across the cohort and operate at scale. This study examines how participants experienced the model and the extent to which assessment integrity was maintained, using an end-of-programme survey of 238 consenting respondents. Most participants regarded the activity as worthwhile, with 79.8% reporting that they solved problems they could not previously solve. However, 29.0% acknowledged at least one instance of reduced effort, a lowered validation standard, or reciprocal validation, while 88.7% believed that at least a little validation had occurred without proper examination. When asked how the process could be strengthened, participants selected post-validation discussion of solutions approximately twice as often as closer auditing or mentor-led validation. These findings provide descriptive evidence of both the potential and the integrity challenges of using peer validation as a scalable assessment approach in large learning environments.
[HC-13] Cybernetic and Epistemic: A Missing Vocabulary for Trustworthy Agent ic Delegation AAAI
链接: https://arxiv.org/abs/2610.00961
作者: Jérémie Lumbroso
类目: Artificial Intelligence (cs.AI); Human-Computer Interaction (cs.HC)
备注: Accepted at TAS 2026 (AAAI Fall Symposium Series), Nov 5-7, 2026, Arlington VA
Abstract:As code generation is increasingly delegated to AI systems, the bottleneck is shifting from writing code to supervising the systems that write it — a shift CS-education researchers have begun to name. This shift exposes a vocabulary gap: the field asks for “human oversight” without a working distinction between the two things language does in a delegation channel — coordinate action (cybernetic: words succeed when the world comes to match them) and coordinate understanding (epistemic: they succeed when they answer to the world and a hearer can check that they do). The failure this names is not cybernetic language but epistemic-form language doing cybernetic work: explanation-shaped output calibrated for approval rather than truth. Oversight that checks only whether an output was approved is satisfiable by rubber-stamping; oversight that holds an agent accountable requires the reasoning behind its work be retrievable and checkable. We present three delegation episodes — illustrations, not controlled evidence — in which epistemic engagement proved practicable while remaining auditable, one public record where a recommendation was withdrawn on its own stated terms, and one failure case illustrating oversight that requires no reasons for its discretionary choices. We propose a criterion for agentic-system governance, alongside existing technical trust properties: every consequential choice should carry the condition under which it would have gone otherwise, in a form a third party can test. Without such a condition, a third party cannot distinguish a decision from a rubber stamp. We give the criterion an operational form — a two-part reconstruction test scoring a delegation record by whether a second reader can predict what the agent does under a perturbation — and a deliberation-recording convention, ORRCF, that makes the condition a required component of every recorded choice.
[HC-14] ABDA-NL: A Natural-Language Scenario Explorer for Argument-Based Reasoning
链接: https://arxiv.org/abs/2610.00947
作者: Shawn Bowers,Martin Caminada,Haoyang Liu,Bertram Ludäscher
类目: Artificial Intelligence (cs.AI); Human-Computer Interaction (cs.HC)
备注: 9 pages, 3 figures. Extended version of a demonstration abstract in the Proceedings of COMMA 2026. Code at this https URL and live demo at this https URL
Abstract:ABDA-NL adds a natural-language interface to ABDA, a system for argument-based discussion using ASPIC- knowledge bases under grounded semantics. Users see which conclusions are accepted, rejected, or undecided, open an interactive rendering of the grounded discussion game to learn why, explore what-if alternatives by suspending assumptions and rules or changing preferences, ask questions that are answered from a scenario’s reference documents, and author new facts, assumptions, and rules in plain English. A large language model provides the bridge between language and formalism: it answers questions from the documents and the current state of the scenario, and it translates plain-English edits into candidate formal statements. The deterministic ABDA engine remains the sole source of arguments, attacks, and acceptance labels, and every proposal of the model is validated and confirmed by the user before it takes effect.
[HC-15] LeanSide: A Formally Verified Co-Reasoning System for Natural-language Proofs
链接: https://arxiv.org/abs/2610.00760
作者: Chenjun Guo,Manooshree Patel,Arnav Mehta,Krishiv Kothari,Thomas Lu,Niels Voss,Rayna Bhattacharyya,Peter Donovan,Bjoern Hartmann,Gireeja Ranade
类目: Human-Computer Interaction (cs.HC)
备注: 16 pages, 13 figures
Abstract:Large language models are increasingly used as collaborators on deductive-reasoning tasks, but their outputs can hallucinate or pull users away from intended reasoning. Formal proof assistants provide machine-checked verification, but have a steep learning curve and require more granular reasoning than human written proofs. We explore an interface that combines these strengths, allowing users to write and revise free-form natural-language proofs while a verified backend checks their reasoning and returns feedback at the user’s granularity. We study this interface in the context of undergraduate mathematics education by developing LeanSide, a formally verified co-reasoning system, which auto-formalizes student reasoning into Lean and informalizes verifier output into understandable feedback. We conducted user studies through classroom deployment and analyzed which system properties helped students make progress and which caused them to get stuck. We use these findings to derive design implications for using a formally verified backend in human-AI co-reasoning systems.
[HC-16] oward Humanoid Robots in Construction: A Teleoperation Feasibility Study IROS2026
链接: https://arxiv.org/abs/2610.00718
作者: Parastoo Ali Pour,David R. Martin,Chang Min Hur,Bo Zhang,Tommy Zhou,Brandon Thomas Lichter,Shane Stanfield,Pramod Khargonekar,Mohammad Abdullah Al Faruque
类目: Robotics (cs.RO); Human-Computer Interaction (cs.HC)
备注: Accepted at IROS 2026 Workshop on Future of Construction
Abstract:We present a teleoperation system that enables a single operator to perform construction tasks on a Unitree G1 humanoid, combining extended reality (XR) based upper body control with pedal-based locomotion to enable simultaneous manipulation and locomotion. Motivated by persistent labor shortages, hazardous working conditions, and challenges in humanoid autonomy, we investigate teleoperation as a practical near-term approach for reducing physical strain on workers while generating high quality demonstration data. We evaluate the system on two representative construction tasks drawn from O*NET occupational database, and report task success and completion time relative to a manual baseline. The system achieved 100% success on tool transport and 80% success on surface painting, with teleoperation requiring substantially more time compared to manual execution.
[HC-17] From Images to Tasks: Characterizing Multimodal LLM Interactions in the Wild
链接: https://arxiv.org/abs/2610.00701
作者: Jinyi Ye,Scott Counts,Gaurav Verma,Kate Lytvynets,Weiwei Yang
类目: Human-Computer Interaction (cs.HC)
备注:
Abstract:Multimodal large language models (LLMs) increasingly integrate vision and text, yet how people use them in natural settings remains underexplored. We seek to answer the question: when users upload images, what tasks are they trying to accomplish? Analyzing over 40,000 de-identified image-upload conversations from Microsoft Copilot, we characterize real-world multimodal use through a hierarchical framework of ten capabilities, spanning perception, cognition, and generation. First, we characterize the distribution and composition of these capabilities, finding that the majority of image-upload tasks involve multiple capabilities. Second, we find that multimodal use spans a broader and more diverse task space than text-only interactions, with asymmetric coverage and task classes that rely on cross-modal grounding. Third, mapping observed capability demand onto 253 existing benchmarks reveals uneven alignment between benchmark coverage and real-world use: benchmarks concentrate on perception and reasoning toward fixed answers, while common workflows involving text, code, and data generation remain comparatively undertested. We validate our taxonomy and findings on an independent ChatGPT dataset. Our results provide a large-scale empirical characterization of what users seek to accomplish with multimodal LLMs and highlight opportunities for benchmark design grounded in observed user demand.
[HC-18] From Scroll to Sale: Exploring the Impact of Interaction Type and Device Price on TikTok Advertisements
链接: https://arxiv.org/abs/2610.00684
作者: Nazanin Sabri,Cat Mai,Haodi Zou,Isha Varada,Damon McCoy,Deepak Kumar,Kristen Vaccaro
类目: Computers and Society (cs.CY); Human-Computer Interaction (cs.HC); Social and Information Networks (cs.SI)
备注:
Abstract:Companies and brands increasingly use dynamic pricing, including targeting social media ads to users based on their income. In this work we audit TikTok’s feed using 56 automated accounts, which collect data on over 80,000 videos, across two studies. We test the impact of device price on ad load and ad types, using 12 phones of low ( 0- 250), medium ( 400- 650), and high ( 750- 1,000+) price as a proxy for income. We also test the impact of interaction type (i.e., like, comment, share), age, and gender on the frequency and content of ads. Overall, the ad load was 29.4%, but liking and sharing content increased ad load significantly. We also found that as accounts spend more time on TikTok, the ad load steadily increases. We found some evidence that device price impacts both ad load and content – more expensive devices were targeted with fewer ads, while the least expensive devices received more discounts.
[HC-19] Sensing Instability Adapting the Scene: A Real-Time Movement-Smoothing Design Framework for Stable VR Locomotion
链接: https://arxiv.org/abs/2610.00643
作者: Ramisa Fariha Joyee,M. Rasel Mahmud
类目: Human-Computer Interaction (cs.HC)
备注: 3 pages, accepted as ISMAR 2026 AXR workshop paper
Abstract:Users experience different balance challenges while standing, walking, and turning in virtual reality (VR), yet most locomotion techniques apply the same visual behavior regardless of movement state. We present a real-time movement-smoothing design framework in this paper that organizes visual adaptations according to movement-specific balance demands. The framework introduces three design strategies targeting standing, walking, and turning, illustrating how state-aware visual adaptations can support postural stability during locomotion. This paper focuses on the framework design and implementation, while a user study is planned as future work. Our work guides the development of future context-aware VR locomotion systems that better support safe and comfortable navigation.
[HC-20] Preliminary Evaluation of Transition-Aware Controller Locomotion Adaptations for Supporting Postural Stability in VR
链接: https://arxiv.org/abs/2610.00628
作者: Ramisa Fariha Joyee,M. Rasel Mahmud
类目: Human-Computer Interaction (cs.HC)
备注: 4 pages, 3 figures, accepted in ISMAR 2026 poster track
Abstract:We present three locomotion adaptation approaches: Motion Acceleration, Turn Acceleration, and Motion Deceleration to improve postural stability during body-state transitions in virtual reality (VR). The system detects standing-to-walking, turning, and walking-to-stopping transitions and applies adaptive locomotion smoothing. Motion Acceleration gradually increases locomotion speed when users begin walking, Turn Acceleration smooths rotation while turning, and Motion Deceleration gradually reduces movement speed before stopping. We evaluated these techniques in a virtual navigation task using objective and subjective balance measures. Preliminary results show reduced center of pressure (COP) velocity and improved balance confidence. These findings suggest that locomotion adaptations can improve balance and navigation experience.
[HC-21] Faithful Chart Generation for Multimodal Deep Research: Frame-Evidence Co-Adaptation
链接: https://arxiv.org/abs/2610.00374
作者: Yuxin Yue,Yingchen Zhang,Ruqing Zhang,Jiafeng Guo,Maarten de Rijke,Xueqi Cheng
类目: Human-Computer Interaction (cs.HC); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Graphics (cs.GR)
备注:
Abstract:Analytical charts in multimodal deep research encode quantitative claims, requiring every visualized value to be faithfully grounded in supporting evidence. Unlike retrieved images that mainly provide contextual information, charts require numerical fidelity: visualized values should not only match retrieved evidence quantitatively but also preserve its original meaning and scope. However, achieving such fidelity remains challenging because current systems usually construct visualization plans before knowing what quantitative evidence can actually be retrieved from the web. As a result, predefined plans may require entities, temporal ranges, or comparison dimensions that the retrieved evidence only partially supports. Existing approaches mainly address this issue through post-hoc verification after chart plans are fixed, enabling unsupported values to be identified but leaving the underlying visual frames unchanged. To address this challenge, we propose Frame-Evidence Co-Adaptation (FECA), an evidence-adaptive visual planning framework for multimodal deep research. Inspired by the bidirectional sensemaking process in Data-Frame Theory, FECA models chart generation as an iterative interaction between visual frames and retrieved evidence. Each visual frame is adaptive: the frame guides evidence acquisition, while retrieved evidence determines whether the frame should be accepted, revised, or dropped before rendering. By coupling visualization planning with evidence availability, FECA shifts chart generation from fixed-plan verification to adaptive evidence-grounded visual reasoning. Experiments on 100 real-world research topics show that FECA substantially improves numerical fidelity while preserving report quality and chart utility.
[HC-22] hree Pathways of Student-AI Interaction: Constraint-First Design for Higher-Order Thinking
链接: https://arxiv.org/abs/2610.00338
作者: Fatima T. Zahra
类目: Human-Computer Interaction (cs.HC)
备注: 26 Pages, 2 figures
Abstract:How students interact with artificial intelligence (AI) systems in educational settings may determine whether that interaction supports or displaces critical thinking. This paper introduces two contributions. The first is the Three Paths of Student-AI Interaction, a typological framework identifying three qualitatively distinct modes of student-AI engagement: Passive Review, Direct Question, and Strategic Dialogue. The second is the Next Level Teaching Blueprint (NLTB), a three-stage instructional design system intended to make Strategic Dialogue more likely. Qualitative content analysis of 50 randomly sampled student-AI interaction messages from an undergraduate research methods course was used to examine the typology. Two human coders achieved 68% path-level agreement ( \kappa = .48), with 80% agreement on Strategic Dialogue identification specifically. GPT-5, used as a third coder, produced a similar overall distribution and introduced a coding category absent from the human scheme. Path 1 (Passive Review) accounted for 46% of exchanges in the primary researcher’s classifications, Path 2 (Direct Question) for 18%, and Path 3 (Strategic Dialogue) for 36%. A second, descriptively examined dataset contained predominantly Strategic Dialogue content, offering a preliminary indication that instructional framing may influence which path students take. Together, the Three Paths framework and the NLTB contribute a language for describing student-AI interaction and a design approach for supporting higher-order engagement.
[HC-23] amLens in Critsly: A Consent-Based Team-Composition Interface and Synthetic Readiness Evaluation for Design Collaboration
链接: https://arxiv.org/abs/2610.00288
作者: Nizam Kadir
类目: Computers and Society (cs.CY); Human-Computer Interaction (cs.HC)
备注: 16 pages, 8 figures, 6 tables. Technical software evaluation using synthetic fixtures; includes ancillary synthetic records and verification scripts
Abstract:Discussing working preferences may support reflection within a design team, but a personality label should not become a performance prediction or a condition of participation. This technical report presents TeamLens, an optional Critsly interface for voluntarily sharing a self-reported MBTI type with a particular board. It separates account activation from disclosure, displays descriptive composition counts, and provides distinct controls for disabling visibility, withdrawing one report and deleting all of one’s reports. The evaluation combines released-source inspection, independently specified synthetic aggregation cases, access and lifecycle checks, browser component tests, a local database microbenchmark and deployment records. All 65,536 binary eligibility subsets of sixteen fixed, distinct type reports matched an independent oracle; 128 seeded multiplicity fixtures also matched, after 4,259 synthetic share calls. The isolated HTTP suite passed 183 assertions. In 180 sequential in-memory SQLite read trials, median service-call time increased from 0.064 ms with no profile rows to 200.932 ms with 1,024 rows; instrumented SQL operations followed 11+3n. These findings concern exercised software behaviour and a bounded workload. They do not establish human usability, psychometric validity, learning gains, team-performance effects or production capacity. The contribution is an implemented disclosure-to-aggregation workflow and an auditable technical account of its correctness boundaries, privacy limitations and scaling cost. OpenAI Codex assisted with implementation, evaluation and manuscript preparation; the paper discloses this use and its limits.
[HC-24] From Web(logs) to Web(AI): Questions Platforms and Methods across Twenty Editions of ICWSM
链接: https://arxiv.org/abs/2610.00205
作者: Koustuv Saha,Eshwar Chandrasekharan
类目: ocial and Information Networks (cs.SI); Computation and Language (cs.CL); Computers and Society (cs.CY); Human-Computer Interaction (cs.HC)
备注:
Abstract:Over twenty editions, the ICWSM community has examined social life online as platforms, interactions, and research methods have changed. What can this body of research tell us at this critical juncture, as AI increasingly reshapes how people communicate online? We analyzed 2,139 indexed contributions from 2007 to 2026, distinguishing topics identified through nonnegative matrix factorization from problem framings captured through explicit textual cues. We find that platform mentions shift from blogs toward Twitter and, more recently, Reddit. Online community research maintains a similar topic share (10.5% to 10.0%), but governance cues within it increase from 4.0% to 34.3%. Harm-related cues also increase after restricting abstracts to a fixed length. Our review also traces advances in sampling, measurement, and causal and experimental methods. We discuss how AI-mediated interactions complicate these questions and provide a reporting checklist to support research across changing platforms.
[HC-25] When the AI Leaves the Tailorshop: Measuring What an LLM Advisor Leaves Behind in Complex Problem Solving
链接: https://arxiv.org/abs/2610.00163
作者: Robin Welsch
类目: Human-Computer Interaction (cs.HC); Artificial Intelligence (cs.AI)
备注: 40 pages, 14 figures, 6 tables, including appendices
Abstract:Complex problem solving depends on acting effectively and understanding how a system works. AI advice may support these outcomes unequally. Two preregistered experiments compared participants managing a simulated clothing factory with and without an LLM advisor. Across studies, AI-supported participants reported greater confidence and understanding with less effort. In the first study (N=200), assistance increased company value but produced no detectable prediction-accuracy difference. After withdrawal, previously supported participants outperformed controls when decisions were scored against repeating previous choices, but not default settings. Within the AI-supported group, more frequent recommendation alterations predicted better unaided performance. In the second study (N=198), AI-supported participants went bankrupt less often and showed a small knowledge advantage in the registered analysis, largely associated with remaining solvent. More frequent recommendation alterations predicted higher knowledge within the AI-supported group. Applied HAI evaluation should assess users’ understanding and independent capability alongside the performance achieved with AI support.
[HC-26] Critsly and StudioCrit: An Artefact-Aware AI Critique Workspace and Simulation-Based Readiness Study for Design Education ICIP
链接: https://arxiv.org/abs/2610.00085
作者: Nizam Kadir
类目: Human-Computer Interaction (cs.HC)
备注: 13 pages, 8 figures, 4 tables. Technical report adapted from an SMT 99.580 research project submitted on 23 July 2026. Documents StudioCrit engineering and simulation evidence; related Critsly demo: arXiv:2607.09673 . No human-participant learning outcomes are reported
Abstract:Critique in design education depends on interpreting work in progress, articulating intentions and translating feedback into revisions. This technical report presents Critsly, an artefact-aware AI critique workspace, and StudioCrit, its architecture-studio research mode. Critsly combines a visual board, design-intention fields, guided reflection, perspective-based critique and action planning. StudioCrit adds studio/class organisation, role-based access, cognitive and architectural classification, educator analytics and exportable evidence. The report consolidates implementation and simulation evidence recorded in a research project submitted in July 2026. Three simulated studio scenarios yielded 109 classified evidence rows, including 85 assigned to higher-order Bloom categories. A separate rehearsal using 50 disposable learner accounts yielded 56 evidence rows, including 46 assigned to higher-order categories. A subsequent hardening rehearsal recorded 50 completed sessions, 50 successful board pulls and 50 denials of student access to analytics. These are software and synthetic-trace observations, not measurements of learning gains or human cognitive performance. Automated classifications remain provisional, and the source report does not establish classifier accuracy or inter-rater reliability. The contribution is an implemented critique-to-evidence workflow and a bounded account of its readiness for further controlled evaluation.
[HC-27] A Comprehensive Evaluation Framework for Conversational Home Energy Management Systems
链接: https://arxiv.org/abs/2610.00073
作者: Wooyoung Jung
类目: Human-Computer Interaction (cs.HC)
备注: 25 pages, 7 figures, 9 tables
Abstract:The growing complexity in home energy management (HEM) demands advanced systems that guide occupants toward informed energy decisions reflecting their background, preferences, and context. Large language model (LLM)-integrated HEM systems (HEMS) have demonstrated promise, but previous studies relied on single-turn or single-task evaluations with response accuracy as the primary metric. Whether such systems deliver effective interactions across the extended multi-turn dialogues typical of real-world use remains an open question. This study introduces a comprehensive evaluation framework of LLM-integrated HEMS derived from the Goal-Question-Metric methodology, organized across five categories: task performance, factual accuracy, interaction quality, control capability, and system efficiency. A total of 23 metrics across multi-turn conversations are proposed and an LLM-as-judge pipeline is employed to enable scalable automated scoring. Its reliability is validated against three trained human coders: after iterative rubric calibration, twelve of the fifteen LLM-scored metrics reached strong agreement (ICC = 0.73), three of them perfect, while the remaining three exhibited near-zero variance in human scores and are instead reported via mean absolute error (0.04-0.28). To demonstrate the framework’s effectiveness, 970 dialogues – 16 scenarios and five personas – were generated and evaluated across four conversational HEMS configurations spanning a sophistication gradient, from a vanilla LLM with raw energy data to a multi-agent HEMS. The framework distinguished the four configurations across multiple evaluation dimensions, revealing their respective strengths and weaknesses. This study contributes to conversational HEMS by providing a reproducible, multi-dimensional evaluation methodology that comprehensively assesses sustained, context-aware system performance.
计算机视觉
[CV-0] Moore Escher Penrose: A Conformal Golden Braid
链接: https://arxiv.org/abs/2610.02210
作者: Sophia Feldman,Assaf Shocher
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:I don’t think I have ever done anything as peculiar in my life. Among other things, it shows a young man looking with interest at a print on the wall of an exhibition that features himself. How can this be? Perhaps I am not far removed from Einstein’s curved universe.‘’ So wrote M.C. Escher about his 1956 lithograph Print Gallery. Nearly half a century later, a mathematical analysis related its geometry to an untwisted source image through a conformal power map z \mapsto z^\alpha , \alpha \in \mathbbC . Building on this construction, we use a frozen text-to-image diffusion model to generate new self-referential scenes. Prompting alone does not enforce the recursion, while a post-hoc transformation can leave structures poorly connected. Applying the transformation during sampling is also insufficient: the denoiser may “repair” the intended distortion or drift out of the prescribed geometry. We construct a generalized inverse T^\dagger of the non-invertible image transformation T , adapted to its recursive constraint. In the idealized formulation, the Penrose identity TT^\dagger T = T makes TT^\dagger an idempotent projection onto geometrically admissible images. Yet denoising only the transformed image remains an out-of-distribution task, even with projection. We therefore braid denoising steps with T and T^\dagger : source-space steps develop the untwisted scene, while transformed-space steps refine its appearance and connections in the final geometry. We generate Print Gallery-like compositions and explore further transformations. Rather than distorting a finished image, we let the scene and its distortion develop together.
[CV-1] Sphere Encoder 2
链接: https://arxiv.org/abs/2610.02208
作者: Kaiyu Yue,Sean McLeish,Ruchit Rawal,Brian Bartoldson,Menglin Jia,Tom Goldstein
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Code will be available at this https URL
Abstract:Sphere Encoder is an autoencoder that generates images by decoding random points from a high-dimensional latent sphere. We identify two limitations of the original formulation that reduce its generation quality. First, random points concentrate near the equator relative to the pole on an encoded latent, but the training rotation never reaches this region, leaving a gap that limits one-step generation. Second, training for generation with pixel-wise reconstruction loss encourages the decoder to average over plausible images, producing blurry images that lack high-frequency details. We present Sphere Encoder 2 to address both limitations, substantially improving image generation quality while maintaining the speed and simplicity of a autoencoder. Models are released at \hrefthis https URLthis http URL.
[CV-2] ROWBench: Do Video Models Render What the Program Specifies?
链接: https://arxiv.org/abs/2610.02205
作者: Zheng-Hui Huang,Guixu Lin,Yu-Ju Tsai,Jian-Kai Zhu,Fengbo Lan,Yu-Lun Liu,Yung-Yu Chuang,Kaipeng Zhang,Zhixiang Wang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Programmable world models separate executable dynamics from visual generation, offering a promising foundation for next-generation game engines. However, their visual adherence to explicit rules and interactions remains insufficiently evaluated. Existing benchmarks assess visual quality, controllability, and instruction or physical adherence, but rarely test fidelity to fine-grained, program-specified world events. We introduce PROWBench, comprising 170 programmatically constructed episodes and 600 proxy videos covering diverse scenes and interactions. PROWBench logs entity states and timestamped events, including those outside the camera’s field of view, as replayable world records, from which it renders synchronized views and proxy representations. This enables generated videos to be checked against the observable consequences of program execution. An extensible framework constructs scenes, controls behaviors, and can render each camera view in different representations, such as coarse 3D, and bounding boxes. The benchmark covers first- and third-person perspectives, with synchronized multi-view observations available for a subset of episodes. Grounded in these records, PROWBench evaluates entity control, long-horizon memory, and, with two VLM-based metrics, Logic-Render Alignment and Interaction Success Rate, adherence to the prescribed timeline and the visual realization of timestamped engine-recorded events.
[CV-3] Embedding Prediction Helps Image Generation
链接: https://arxiv.org/abs/2610.02203
作者: Sihan Xu,Ji Xie,Zilin Wang,Hui Shen,Stella X. Yu
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注: Project page: this https URL
Abstract:In diffusion transformers, a class label or a text prompt is embedded once, and the same condition is reused at every denoising step. We ask whether predicted embeddings can serve as this condition instead. Next-Embedding Predictive Autoregression (NEPA) trains a Transformer to predict the next continuous embedding in a sequence. In generation, the clean image follows the noisy image, so its embeddings are the next embeddings after the condition and the noisy image. We train a NEPA model to predict them all at once with Multi-Embedding Prediction, and in Embedding Conditioned Generation, a DiT generator is conditioned on these predictions, recomputed at every denoising step, so the conditioning signal adapts to the current noisy state. Experiments on class-conditional ImageNet 256\times256 study the condition of the generator, the design of Multi-Embedding Prediction, and the scaling of both models. The NEPA model adds a second network to every sampling step; with it, and combined with REPA, our final model, NEPA-DiT-XL, reaches an FID of 1.32 using about a third of the training compute of REPA.
[CV-4] SILSA: Sliding-Window Slice Latents for Topology-Preserving High-Resolution 3D Generation NEURIPS2026
链接: https://arxiv.org/abs/2610.02201
作者: Tianjiao Yu,Xinzhuo Li,Yifan Shen,Ying Shen,Kiet A. Nguyen,Adheesh Sunil Juvekar,Ismini Lourentzou
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: Accepted at NeurIPS 2026. Project link: this https URL
Abstract:High-resolution 3D generation increasingly relies on voxel latents and multi-stage pipelines that first predict active structure and then synthesize local geometry. While effective, this design fragments continuous surfaces into many local tokens, inflates generation cost, and often weakens topological consistency for thin or highly connected shapes. We introduce SILSA, a topology-aware 3D generation framework that represents shapes with compact sliding-window slice latents. Instead of generating expensive voxel tokens, SILSA uses a fixed set of overlapping slices along the three canonical axes, where each token summarizes a local depth window to preserve cross-sectional continuity and support single-stage rectified-flow generation. A Slice VAE encodes oriented surface samples into multi-axis slice latents and reconstructs them with a sparse volumetric decoder, while a Volumetric Anchor Lattice coordinates directional slice streams through a shared 3D workspace. To preserve structural correctness, we introduce slice-level topology supervision that matches persistence diagrams and aligns Betti transitions across neighboring slices. Experiments show that SILSA improves structural fidelity while substantially reducing generation cost. SILSA improves PSNR by 8.7% , coverage by 5.96 absolute points, and Betti error by 9.2% over the strongest baseline, while using 70.0% fewer tokens than the next-most compact baseline and over 98% fewer tokens than sparse or hierarchical tokenizers, effectively reducing training memory by 40.4% and inference time by 58.5% . Qualitative results further show improved preservation of thin structures, repeated components, and long-range connectivity.
[CV-5] VISTA: A Visual Harness for Reasoning in an Interactive World
链接: https://arxiv.org/abs/2610.02200
作者: Qiushi Han,Keya Hu,Linlu Qiu,Cathy Wu,Kaiming He
类目: Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
备注: Tech report. An early version of this manuscript was in a blogpost published in Aug 5, 2026: this https URL
Abstract:We show that multimodal models possess strong reasoning abilities and that an appropriate harness can unlock their potential to solve tasks across diverse interactive environments. We introduce VISTA, a visual harness that gives a general-purpose multimodal model long-horizon vision. VISTA allows the model to directly perceive the environment through visual observations and maintains a lossless visual memory that preserves past observations in their original form. The model can actively retrieve these observations and reorganize its visual input as it reasons. On ARC-AGI-3, VISTA improves Claude Opus 5.0’s Relative Human Action Efficiency score from 40.68 to a perfect 100.00, with the model completing all 25 public games using 57.4% fewer actions than first-time human participants. VISTA’s simple design also allows it to extend naturally to diverse visual environments with minimal adaptation. Across three additional benchmarks covering a diverse range of visual games and puzzles, it substantially outperforms baselines using the same underlying model with minimal harnesses. Our results highlight VISTA’s potential as a general-purpose visual harness for advancing multimodal agents in complex visual environments.
[CV-6] HiPhy: Hierarchical Alignment for Physically-Plausible Multi-Principle Video Generation
链接: https://arxiv.org/abs/2610.02197
作者: Tahira Kazimi,Shubhankar Borse,Munawar Hayat,Fatih Porikli,Pinar Yanardag
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Project page: this https URL
Abstract:Video generation models have achieved remarkable visual fidelity and have strong potential to become general-purpose world simulators. Despite this progress, they still fail to generate videos which adhere to laws of physics. The problem becomes even more apparent in realistic settings where multiple physical principles must work together within the same video; for example, “a balloon floating upward while steam rises from a pot” requires buoyancy and fluid dynamics to unfold coherently and simultaneously. Yet existing methods largely ignore multi-principle interactions, focusing on a single principle per video. We propose HiPhy (Hierarchical Physical Alignment), a reinforcement learning framework that grounds video generation in physical laws through a dual-level objective: locally enforcing the temporal dynamics of individual physical principles, and globally ensuring the physical and semantic coherence of the entire scene. To support multi-principle generation, we construct a 50K-prompt dataset and introduce a prompt benchmark MultiPhyBench, spanning a diverse range of co-occurring physical events. Our experiments show that HiPhy significantly outperforms prior methods and baselines, improving physical commonsense and semantic alignment significantly across various benchmarks, with the largest gains on scenes involving multiple concurrent physical principles where competing methods degrade most sharply.
[CV-7] InterEvolve: Test-Time Evolution of Reward Programs for Humanoid Loco-Manipulation
链接: https://arxiv.org/abs/2610.02196
作者: Zhuo Lin,Sirui Xu,Liuyu Bian,Yu-Xiong Wang,Liang-Yan Gui
类目: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR)
备注: Project page: this https URL
Abstract:We study test-time evolution for humanoid loco-manipulation: solving tasks that a controller was never trained for by repurposing its existing skills, improving from its own attempts, and retaining what it learns, without retraining. Our key insight is that a broad controller already holds much of the competence a new task needs, and that this competence becomes accessible through an interface between planning and control that is expressive enough to specify contact-rich, multi-stage interactions, yet executable and measurable enough that execution feedback can guide planning from experience. InterEvolve realizes this interface with two components. First, we develop an object-aware forward-backward (FB) behavioral foundation model, whose object residuals on a frozen body prior turn a new reward about the body or objects into loco-manipulation behavior at test time. Second, we specify tasks as reward programs: staged rewards with completion conditions and tunable constants. A large language model (LLM) agent revises the program structure in context, drawing on execution feedback and a skill library of verified programs, while a numerical optimizer tunes its constants. With every candidate verified across parallel simulation scenarios, the program explores new ways to induce, repurpose, and compose the controller’s existing motor competence for the task at hand, and thus improves over iterations. Experiments show that human-designed rewards leave much of the FB model’s loco-manipulation competence untapped, whereas the programs InterEvolve evolves release it, sometimes through novel strategies. It further produces behaviors for diverse tasks, complex scenes, and long-horizon compositions in simulation, and evolved skills run autonomously on a physical Unitree G1 from egocentric onboard perception.
[CV-8] DMAD: Distribution Matching as Adversarial Distillation for Fast Visual Generation
链接: https://arxiv.org/abs/2610.02188
作者: Zhengming Yu,Junkun Yuan,Haotian Yang,Gordon Guocheng Qian,Yizhi Wang,Angtian Wang,Yiding Yang,Bo Liu,Xin Li,Wenping Wang,Chongyang Ma
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注: 28 pages, 15 figures. Project page: this https URL
Abstract:Distribution Matching Distillation (DMD) trains a few-step student from the difference between separately estimated target and student scores, so it must keep an auxiliary diffusion model fitted to the student’s evolving distribution at extra memory and computation cost. We introduce DMAD, Distribution Matching as Adversarial Distillation, which recasts distribution matching as classification and learns the required log-density ratios directly. Two discriminator heads on a shared backbone distinguish real data and teacher samples from the student’s, and linear losses on their logits train the student without auxiliary score fitting. We prove that at the discriminator optimum these losses recover the distribution-matching gradient underlying DMD, through the classical identity linking discriminator logits to log-density ratios. We further introduce gap-based reweighting, which adapts teacher supervision across noise levels from the real-data head’s empirical logit gap between real and teacher samples. DMAD reaches a Fréchet Inception Distance (FID) of 1.04 with one-step generation on ImageNet-64x64, 14.47 with four-step SDXL on COCO-10K, and a VBench total score of 85.15 with four-step Wan2.1-T2V-14B, the best values among the compared few-step methods and the multi-step teachers. On MiniMax-H3-33B, our four-step student achieves overall human preference rates of 79.1% over DMD2 and 84.6% over rCM for joint audio-video generation, excluding ties. Our code, models and demos are available at this https URL.
[CV-9] OmniSeek: Native Tool Integration for Multi-turn Audio-Visual Reasoning
链接: https://arxiv.org/abs/2610.02181
作者: Haibo Wang,Jiteng Mu,Jialu Li,Jingru Yi,Yuanjun Xiong,Jianming Zhang,Lifu Huang,Mingze Xu
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:We present OmniSeek, an agentic framework that transforms an Omni Large Language Model (Omni-LLM) into an active, multi-turn reasoning agent with native tool use. Rather than passively processing an entire audio-visual sequence in a single forward pass, OmniSeek makes evidence acquisition part of the reasoning process: it dynamically decides whether to look or listen, and over which temporal window, to retrieve sparse but critical evidence across different modalities within long contexts. Through an iterative multi-turn protocol, the retrieved raw audio or visual segments are appended back into the context to support subsequent reasoning. To cold-start this capability, we build a data engine that synthesizes OmniTraj-170K, a corpus of multi-hop Chain-of-Thought trajectories with interleaved audio and visual evidence. We first supervise the model on these trajectories to instill multi-turn tool-use behavior, and then further optimize the policy via a two-stage reinforcement learning with verifiable rewards. Moreover, we introduce an Audio-Visual Necessity objective that explicitly rewards successful trajectories whose reasoning depends on both modalities, discouraging single-modality shortcuts. Extensive experiments across a wide range of benchmarks demonstrate that OmniSeek learns adaptive cross-modal evidence seeking and consistently improves audio-visual reasoning performance.
[CV-10] Generative Cinematographer: Composing Camera and Object Motion in 3D
链接: https://arxiv.org/abs/2610.02180
作者: Jiahan Zhang,Chaohao Yang,Namitha Guruprasad,Vivekjyoti Banerjee,Trong-Tung Nguyen,Alan Yuille,Anand Bhattad
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:
Abstract:Current controllable video generation systems often rely on 2D motion trajectories or sparse drag signals for object motion. These controls are ambiguous because the same 2D trajectory can correspond to different 3D motions, especially when the camera and objects move simultaneously. We present Generative Cinematographer (GenCine), a system that lifts a single image into an editable 3D scene scaffold where artists jointly author camera and foreground motion. Artists specify a camera path and move selected foreground regions using local 3D motion handles. Several handles can move different parts of a subject independently, providing a piecewise-rigid approximation to non-rigid motion without a physics simulator or category-specific prior. To communicate these controls to a pretrained video model, we project them into guidance maps. These maps record where the controlled regions appear in each frame, assign each handle a fixed color across frames and encode the current 3D positions of its controlled points in the same world coordinate system as the background. This lets us describe object motion relative to the scene even as the camera moves. For training, we recover controls from the motion observed in real videos and use ground-truth geometry and trajectories from synthetic videos. We train a lightweight guidance branch and LoRA adapters on a pretrained Wan model to follow these controls. Our experiments show consistent camera-relative motion, improved geometric consistency under viewpoint changes, and strong controllability across diverse real-world scenes.
[CV-11] World Observer: Joint Actor-Observer Generation for Persistent World Modeling
链接: https://arxiv.org/abs/2610.02162
作者: Hyunwook Choi,Dahyun Chung,Hyunsung Kim,Siyoon Jin,Jinhyeok Choi,Junyoung Seo,Seungryong Kim
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:How can a world model continuously observe regions beyond the actor’s current view? Video world models simulate how an environment evolves from an agent’s actions, yet remain actor-centric. Once an object leaves the actor’s view, they lose direct evidence of its evolution, often failing to preserve its state and dynamics upon re-entry. To address this, we introduce World Observer, which decouples observing from acting by jointly generating a perspective actor for the agent-centric view with one or more panoramic observers that watch selected world regions. This allows objects that leave the actor’s view to remain visually evolving in an observer, so their updated states are reflected when they re-enter. We ground the actor and observers by warping from a shared panoramic source for explicit geometric correspondence, and introduce an Observer Sink of high-resolution perspective references to restore fine appearance upon re-entry. Since the observers are decoupled from the actor, they can be placed freely across the scene, extended to multiple locations for broader coverage, and driven by control signals to steer out-of-view evolution. To evaluate out-of-view evolution, we further introduce world-space metrics and a benchmark spanning real and synthetic scenes. World Observer substantially improves out-of-view dynamics while remaining competitive in visual fidelity, camera control, and 3D adherence.
[CV-12] 4Director: Controlling Video World Models with Rigid 3D Geometry
链接: https://arxiv.org/abs/2610.02160
作者: Wei Cao,Hao Zhang,Vikram Voleti,Yuqun Wu,Mallikarjun B R,Shimon Vainer,Mark Boss,Yaoyao Liu
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 28 pages, 15 figures. Project page: this https URL
Abstract:Precise control over camera and object motion is essential for professional video production. Existing methods control objects only coarsely, through image-plane cues that are ambiguous in depth and rotation or through 3D tracks and blobs that lack complete geometry and lose consistency across viewpoint changes. We introduce 4Director, a video world model conditioned on an explicit 4D scene representation: each object is reconstructed once from the input image as a canonical mesh and moved by one prescribed rigid transformation per frame. This representation provides an intuitive 3D control interface and prevents unobserved geometry from being regenerated independently in every frame. We render the controlled scene as a depth video and introduce a Motion Adapter that transforms this geometric scaffold into video while synthesizing view-consistent appearance, illumination, and non-rigid dynamics. For training, we construct RealCOD-Rigid, a new dataset of 20,774 clips annotated with rigid 3D scenes by our automatic pipeline. We further introduce Identity-Gated IoU (IG-IoU), which jointly evaluates adherence to prescribed object motion and preservation of object identity. Experiments demonstrate that 4Director consistently outperforms prior methods in visual quality and in camera and object control.
[CV-13] MosaiChunk: Compositing Spatio-Temporal Memory for Autoregressive Video Generation
链接: https://arxiv.org/abs/2610.02153
作者: Yiwen Zhang,Haocheng Xi,Michael Tian-Yue Liu,Alexei A. Efros,Hadar Averbuch-Elor,Qianqian Wang,Haiwen Feng
类目: Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR)
备注: 27 pages. Project page: this https URL
Abstract:Long-horizon autoregressive video generation is limited by a finite context window. When an object or scene falls out of context, its fine-grained visual details may be lost and difficult to recover upon reappearance. To retain access to such visual details, we introduce MosaiChunk, a spatio-temporal memory mechanism that composes a mosaic of selected historical key-value (KV) entries across space and time. Our approach is motivated by the observation that a frozen video generator can directly consume such non-contiguous historical KV and recover the corresponding visual content. We therefore keep the generator fixed and learn only a lightweight router that determines which historical sections to include in the mosaic under a fixed active-memory budget. We further introduce RememBench, a benchmark of long-horizon revisits with prompt-driven text-to-video (T2V) and camera-driven image-to-video (I2V) splits. Our experiments show that MosaiChunk consistently improves revisit consistency over both sliding-window inference and whole-chunk retrieval under matched memory budgets, across both T2V and I2V settings.
[CV-14] Omni-Embed-Mini: Binding Modalities Without Forgetting via Dense Distillation EMNLP2026 UAI
链接: https://arxiv.org/abs/2610.02148
作者: Mohammed Irfan Kurpath,Jaseel Muhammad Kaithakkodan,Sahal Shaji Mullappilly,Ivan Laptev,Hisham Cholakkal
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Findings of EMNLP 2026. 26 pages, 8 figures, 14 tables. Project page: this https URL
Abstract:Extending a text embedding model to new modalities typically degrades text retrieval quality, and existing omni-modal embedders compensate with multi-billion parameters. We present Omni-Embed-Mini, a 0.9B-parameter model that maps text, speech, audio, images, video, and visually-rich documents into a single shared cosine space without updating any text-side parameter. Our key insight is that the teacher signal requires no separate embedding model: each media sample is paired with a dense cascaded caption, and the teacher target is simply the frozen backbone’s own embedding of that caption. Because teacher and student share the same backbone weights, they inhabit byte-identical geometry, and lightweight projectors plus phased LoRA adapters on the modality encoders suffice for alignment. Training combines a Matryoshka SigLIP contrastive loss with an online hybrid hard-negative miner whose negatives sharpen as the encoder improves. The recipe carries over to a 2.3B variant by swapping in a native vision-language backbone. Omni-Embed-Mini-0.9B keeps its text weights bit-identical to the backbone, so training cannot regress text retrieval (49.57 nDCG@10 on MTEB-v2 BEIR-8), while extending it to five additional modalities, and is ~2.7x to 9.5x smaller than every open omni embedder we compare against. The 2.3B variant is competitive with the closed gemini-embedding-2, edging ahead of it on the overall-modality average. Models, code, data and evaluation harness are on our project page: this https URL
[CV-15] MIRTO: a registration-gated multiverse-tested evaluation protocol for unsupervised anomaly segmentation in brain MRI
链接: https://arxiv.org/abs/2610.02136
作者: Negin Kafee Hernashki,Soumick Chatterjee
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Image and Video Processing (eess.IV); Medical Physics (physics.med-ph)
备注:
Abstract:Unsupervised anomaly detection (UAD) methods for brain MRI are ranked by a single score, yet that score rests on choices that are rarely reported: how each anomaly map is aligned with the reference, how and on which data the threshold is set, and which false-positive budget, metric, aggregation and lesion definition are used. We present MIRTO, an evaluation protocol that makes these choices explicit and measures their effect. It gates the geometry of every comparison with a registration check and label-free diagnostics of known power, sets thresholds on validation data alone and reports the false-positive volume actually realised on test, repeats each comparison over 15,552 defensible evaluation pipelines, and attaches paired subject-bootstrap intervals with multiplicity control. Applied to four UAD methods trained on the same healthy data and tested on 312 BraTS 2020 subjects, MIRTO showed that an axis-order mismatch between stored maps and the reference lowered a diffusion model’s voxel AUROC from 0.873 to 0.583 whilst barely moving its slice-level AUROC. Within each metric, the method explained at least 0.95 of the variance in voxel AUROC and AUPRC and 0.77 in Dice, but only 0.14 in lesion sensitivity, where the lesion definition and hit criterion dominated. A Dice advantage that was significant at validation thresholds vanished at equal realised false-positive burden, and an exact identity attributes it to threshold transfer. A training-free change to REFLECT’s latent aggregation raised Dice at equal burden by 0.052. Nine hypotheses were tested against explicit criteria; because the same cohort served to develop the protocol, all inference is exploratory.
[CV-16] Harnessing Domain Specialists in Multimodal Mixture-of-Experts for Efficient Adaptation ALT
链接: https://arxiv.org/abs/2610.02123
作者: Damiano Marsili,Raphi Kang,Aditya Mehta,Pietro Perona,Georgia Gkioxari
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Project page: this https URL
Abstract:Mixture-of-Experts (MoE) architectures scale model capacity through sparse computation, routing each token through only a small subset of experts. In this work, we explore whether this sparsity gives rise to emergent intrinsic organization in multimodal MoEs. We find that experts develop strong semantic specialization across modalities and domains despite not being explicitly trained for modularity. Building on this structure, we introduce ExpertLens, a data-free method that identifies domain-specialized experts directly from pretrained model weights by decoding router weights into semantically meaningful vocabulary tokens. We leverage this specialization for efficient multimodal adaptation by selectively fine-tuning experts relevant to a target domain. Across math, medical, and remote sensing tasks, ExpertLens matches or surpasses full fine-tuning while updating only 21.7 - 47.0% of model parameters and achieving a 4.0x average training speedup, and outperforms LoRA in both adaptation performance and training efficiency. These results show that sparsity introduced for efficiency can give rise to semantic modularity that is directly useful for efficient adaptation.
[CV-17] Surface-volume self-supervised representation learning of brain MRI for genetic discovery
链接: https://arxiv.org/abs/2610.02114
作者: Tian Xia,Nuo Chen,Zihao Zhu,Huiwen Han,Ziqian Xie,Zhiwen Fan,Degui Zhi
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 17 pages, 3 figures, 1 table, 2 supplementary tables
Abstract:Existing genome-wide association studies (GWAS) of brain imaging provide predefined or deep-learning-derived imaging phenotypes, yet these phenotypes come from either volumetric scans or cortical surface meshes, so each captures only part of the heritable variation in brain anatomy. Here we introduce MEVA (Mesh-Enhanced Volumetric Autoencoder), a self-supervised framework that encodes voxel-level image intensity together with cortical mesh geometry, including curvature and cortical thickness at each surface vertex, into one shared set of imaging features. Combining the mesh and volumetric inputs in MEVA yields modest performance gains in age and sex prediction over models that use either input alone. When these features serve as phenotypes for GWAS in the UK Biobank, they reveal more genome-wide significant loci than features learned from volumes alone or from meshes alone. These results suggest that adding cortical surface geometry to volumetric self-supervised learning captures additional heritable variation and so increases the number of loci detected.
[CV-18] GeoLatent: Geometry-Guided Latent Structuring with Routed Optimization for 3D Reasoning
链接: https://arxiv.org/abs/2610.02091
作者: Yakun Zhu,Yi Bin,Yujuan Ding,Zheng Wang,Pengpeng Zeng,Duo Peng,Jingkuan Song,Heng Tao Shen
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注: 23 pages, 6 figures
Abstract:Despite progress in vision-language models, 3D spatial reasoning from 2D images remains challenging. Text-based methods describe intermediate geometry with discrete tokens, limiting fidelity for continuous spatial relations. Continuous latents offer richer representations, but a single latent type does not explicitly separate the cues needed across spatial tasks. Decomposed spatial latents address this by representing position, direction, and global geometry separately under geometric supervision. Yet the geometry representation can still collapse toward one dominant direction, and unrestricted attention can leave the latents underused during answer learning. We introduce GeoLatent, combining Common–Residual Geometry Alignment (CR-GEO) with routed optimization to structure the geometry states while promoting latent-mediated answer learning. CR-GEO separates shared from residual teacher geometry; routed optimization jointly trains geometry and language, temporarily directs visual answer learning through the latents, and restores full attention with geometry supervision. In controlled comparisons, CR-GEO raises geometry effective rank from 1.00 to 3.87, while blocking latent readout at the bottleneck lowers direction accuracy from 89.1% to 25.8% on 128 fixed questions. After recovery, the differentiated geometry representation and latent-mediated visual route remain available alongside direct image access. GeoLatent achieves 73.0% on SPAR-Bench and 72.1% on SPBench, outperforming previously reported methods on both.
[CV-19] Learning from Failure: Leverag ing Unreliable Predictions in Semi-Supervised Real-World Adverse Weather Removal
链接: https://arxiv.org/abs/2610.02051
作者: Cap Dang Xuan Kiet,Tat-Jen Cham
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Adverse weather image restoration aims to recover images degraded by rain, haze, snow, and other weather-induced artifacts, thereby improving the robustness of outdoor vision systems. Existing unified restoration models exhibit limited generalization to real-world scenes due to their reliance on synthetic supervision and insufficient semantic constraints. In this paper, we propose a novel student–teacher semi-supervised framework that addresses both challenges. Specifically, we introduce an unreliable database that preserves failed teacher predictions as informative negative samples for contrastive learning, while a reliable database stores high-quality teacher predictions as positive samples. By jointly exploiting reliable pseudo-ground truths and unreliable teacher outputs, the proposed framework learns to enhance desirable restoration characteristics while avoiding common failures. We further propose a phase spectrum-based semantic constraint that replaces computationally expensive text-based supervision with an efficient and naturally aligned semantic prior. An adaptive phase consistency loss is also designed to dynamically balance supervision between the degraded input and teacher pseudo-ground truths according to degradation severity. Extensive experiments on real-world benchmarks demonstrate that the proposed method consistently outperforms existing state-of-the-art approaches in restoration quality and perceptual fidelity while exhibiting stronger generalization to real-world adverse weather conditions.
[CV-20] DiDE:Direct Injection with Color-Texture DEcoupling for 3D Stylization NEURIPS2026
链接: https://arxiv.org/abs/2610.02044
作者: Tao Wu,Alexandra Gomez-Villa,Senmao Li,Yaxing Wang,Joost van de Weijer,Kai Wang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted to NeurIPS 2026
Abstract:Recent advances in rectified flow-based image-to-3D generative models have enabled high-fidelity 3D asset generation. Building on this, a growing line of work has exploited these strong 3D priors for training-free stylization, transferring visual attributes from a reference image onto a generated 3D asset. However, existing methods enforce an all-or-nothing paradigm: color and texture are transferred jointly, with no mechanism to control them independently – a limitation we formalize as Disentangled 3D Stylization(Disen3D). To address this, we propose DiDE, the first training-free framework for Disen3D. Key to our approach is the observation that the structured latent space of image-to-3D models is overcomplete with respect to texture: texture information occupies only a small subset of the style-significant channels, leaving a free subspace available for independent color encoding. DiDE exploits this via a channel partition mechanism that processes a content image, a texture reference, and a color reference through dedicated branches and composes both style signals interference-free at every self-attention layer, preserving content geometry throughout. Experiments on Disen3D-Bench, our newly collected multi-reference benchmark, show that DiDE consistently outperforms 2D and 3D stylization baselines in color fidelity, texture transfer, and content preservation.
[CV-21] ask-Adaptive Grounded 3D-Programmers Using 2D VLMs
链接: https://arxiv.org/abs/2610.02021
作者: Arman Raayatsanati,Sombit Dey,Anna-Maria Halacheva,Jan-Nico Zaech,Luc Van Gool,Danda Pani Paudel
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注: 18 pages, 9 figures, 11 tables
Abstract:Recent vision-language models (VLMs) exhibit remarkable generalization and reasoning abilities, yet 3D understanding in these models is limited by data scale, training diversity, and reasoning capacity. Instead of naively extending these models into 3D, we take a different approach: we enable powerful 2D VLMs to operate reliably in 3D by introducing 3D grounding and iterative feedback loops with two novel concepts: Canonical Coordinate Framing (CCF) and Task-Adaptive Feedback (TAF). CCF serves as a unified visual representation that anchors both inputs and outputs to a shared Euclidean coordinate system, solving common challenges in 3D grounding such as axis ambiguity, inconsistent metric scale, and floating references. Complementary to this structured framing of the 3D inputs, TAF closes the reasoning loop with task-adaptive dynamic feedback that enables 2D VLMs to perform varied open-vocabulary tasks within their native visual context. Building on this foundation, we introduce 3D-Prog, a 3D understanding, reasoning, and generation framework that jointly employs the capabilities of CCF and TAF together with powerful VLMs. Without requiring any retraining, 3D-Prog performs open-vocabulary 3D understanding, manipulation, and generation across both object-level and scene-level tasks. Our experiments show that the joint use of CCF and TAF transforms 2D VLMs into geometry-aware 3D programmers, achieving consistent, interpretable, and high-quality results across diverse 3D tasks. Comments: 18 pages, 9 figures, 11 tables Subjects: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI) Cite as: arXiv:2610.02021 [cs.CV] (or arXiv:2610.02021v1 [cs.CV] for this version) https://doi.org/10.48550/arXiv.2610.02021 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[CV-22] Exploring Weaknesses of Generative Image Watermarks against Latent Frequency Masking ICDM2026
链接: https://arxiv.org/abs/2610.02010
作者: Kirill Aistov,Khaled Abud,Irina Serzhenko,Egor Kovalev,Aleksey Yakushev,Aleksandr Akimenkov,Dmitry Obydenkov,Yury Markin,Sergey Lavrushkin,Dmitriy Vatolin,Anastasia Antsiferova
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Multimedia (cs.MM)
备注: This work has been accepted for publication at IEEE ICDM 2026 conference. The final published version will be available via IEEE Xplore
Abstract:Invisible watermarking has become a central tool for tracing AI-generated images, but its robustness against adaptive removal attacks remains an open security question. We introduce Latent Frequency Masking, an attack that erases watermark evidence by replacing selected Fourier coefficients in the latent representation of a watermarked image. The replacement can be sampled from Gaussian noise for efficiency or derived from diffusion regeneration for improved image preservation. We provide a theoretical distortion bound relating the change between the reconstructed adversarial image and the masked latent-frequency perturbation. We evaluate the proposed attack against six diffusion watermarking methods on images generated from DiffusionDB and MS-COCO prompts. Latent Frequency Masking removes or substantially weakens several watermarks while preserving perceptual quality and achieving favorable runtime compared with existing attacks. These results identify latent-frequency manipulation as a practical attack surface and highlight the need to include such attacks in robustness evaluations of generative image watermarking.
[CV-23] Weather-Aware Domain Adaptation for Street-View Weather Recognition
链接: https://arxiv.org/abs/2610.02000
作者: Hossein Maghsoumi,George Atia,Yaser P. Fallah
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注: 7 pages, 3 figures, 4 tables. Published in the 2026 IEEE Conference on Technologies for Sustainability (SusTech)
Abstract:Adverse conditions such as rain, snow, fog, and dust remain challenging for camera-based perception in autonomous driving. We study multi-class weather recognition from street-view images under domain shift, where most available training data come from non-street-view sources that differ markedly from real driving scenes. We propose Weather-Aware Adversarial Discriminative Domain Adaptation (WA-ADDA), which conditions the domain discriminator on predicted weather to promote features that are both domain-invariant and weather-sensitive. We also assemble a multi-dataset benchmark by unifying diverse non-street-view weather collections as sources and real street-view images as targets, and define a standardized evaluation protocol with macro accuracy as the primary metric. Across backbones (ResNet-50, EfficientNet, VGG, DenseNet), WA-ADDA consistently improves street-view performance and yields strong per-class recalls in challenging conditions while preserving clear-weather accuracy. These findings highlight the feasibility of domain-adapted weather recognition and the value of our benchmark for advancing robust, on-board perception.
[CV-24] From Reasoning Failures to Composable Video Spatial Intelligence
链接: https://arxiv.org/abs/2610.01999
作者: Pengzhan Sun,Junbin Xiao,Ramanathan Rajaraman,Shiu-hong Kao,Angela Yao
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Spatial reasoning benchmarks evaluate vision-language models across diverse tasks, but task-level scores do not reveal which underlying capabilities account for success or failure. Each task requires recovering spatial evidence, representing geometry, and reasoning over it. We disentangle these capabilities by comparing predicted and ground-truth spatial context under a shared schema and coordinate contract. This comparison reveals four recurring sources of error: inaccurate perception, missing information in the spatial context, selection of the wrong measurement, and errors in reference frames or in tracking position and orientation. Guided by this diagnosis, we develop CROSS, a training-free library of typed geometric operators and spatial skills that function over available evidence to support reliable video spatial reasoning. The resulting library supplies verified context to non-coding VLMs or callable skills to a SpatialClaw agent. We evaluate \methodname on five benchmarks. \methodname raises the average score from 55.9% to 60.2% on ReVSI and improves the SpatialClaw result from 62.8% to 66.3% on DSI-Bench. These gains demonstrate that explicit handling of spatial conventions can repair systematic reasoning failures without additional training.
[CV-25] Comparing a gradient boosting algorithm to the GOES FDC for wildfire detection
链接: https://arxiv.org/abs/2610.01994
作者: Asaf Vanunu,Boaz Nadler,Arnon Karnieli
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注:
Abstract:Wildfires pose severe risks to human life, ecosystems, and property. This study presents a machine learning approach for wildfire detection from GOES ABI imagery. A CatBoost model was trained on a large dataset with thousands of ABI images and over 300,000 matching VIIRS fire detections. An evaluation on a separate dataset across five regions showed that the learned CatBoost model outperformed the operational GOES Fire Detection and Characterization (FDC) product. It achieved higher precision, recall, and F1 scores both within and outside the training area. The CatBoost model achieved F1 scores that were 0.16 to 0.38 higher than the GOES FDC in all regions. In addition, out of 51 historical fire events, the CatBoost detected 26 fires before both VIIRS and GOES FDC, compared to only six earlier detections by the GOES FDC. Importantly, the CatBoost model achieved accurate wildfire detection also during nighttime, whereas the GOES FDC obtained very low recall values, around 0.03. This study demonstrates that machine learning models may offer significant improvements over existing geostationary fire products, including higher accuracy, fewer false alarms, and earlier detection.
[CV-26] Continual Concept Erasure in Diffusion Models by Suppressing Cross-Edit Interference
链接: https://arxiv.org/abs/2610.01989
作者: Yongliang Wu,Haori Lu,Jinqi Luo,Wei Cao,Xingyu Zhu,Yaoyao Liu
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 24 pages. Project page: this https URL
Abstract:Concept erasure removes copyright-protected, privacy-sensitive, or otherwise undesirable concepts from pretrained text-to-image diffusion models to support content governance and compliance. As erasure requests arrive over time, models must remove new targets without undoing prior erasures. Existing methods do not constrain interference across edits: residual perturbations outside the retain set interact and accumulate, degrading unrelated generations and sometimes collapsing previously erased targets into noise. We propose CEASE (Continual Erasure via Adaptive Subspace Editing), a training-free method that imposes two subspace constraints on a closed-form solver. CEASE adds the token representation of the shared replacement to the solver’s invariance matrix and, when interference is detected, projects the current update onto the orthogonal complement of dominant output directions extracted from cumulative past updates. A closed-form decomposition attributes the accumulated interference to repeated activation of the shared replacement and overlap between successive update directions, showing that the two constraints suppress these respective sources. Across continual erasure of celebrities, artistic styles, and instances, CEASE achieves the most consistent erase-preserve trade-off, while existing methods either degrade general generation or insufficiently erase targets.
[CV-27] oken-Level Video Reinforcement Learning
链接: https://arxiv.org/abs/2610.01973
作者: Yifan Wang,Gordon Guocheng Qian,Yanyu Li,Anil Kag,Yun Fu
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Reinforcement learning (RL) for video generation usually assigns one scalar reward to an entire sampled video. Yet a video is not uniformly flawed: some visual tokens may already satisfy the prompt, whereas others require correction. A scalar reward cannot localize errors, causing optimization to perturb satisfactory tokens while under-targeting the tokens that actually need to change. We introduce Token-Level Video Reinforcement Learning, TVRL, a framework that derives token-level credit from the reward being optimized. Our key insight is that the answer likelihood of a frozen vision-language model provides both signals: its outputs contribute to the video-level reward, while magnitudes of its video-input gradients reveal which generated video tokens most affect that score. We instantiate TVRL in Group Relative Policy Optimization by averaging prompt-derived question rewards into one group-relative advantage and using detached, question-conditioned token-credit maps to reweight dense denoising-transition log-probabilities inside the clipped policy ratio. On VBench-2.0, TVRL achieves an Overall score of 57.69, outperforming the base model by 3.60 points. TVRL also improves matched GRPO baselines across three SDE samplers (SAGE, Flow, and Dance) by 2.68–3.15 points and across four reward models (VideoAlign, VideoScore2, UnifiedReward2, and Qwen3.5-9B) by 1.33–3.15 points.
[CV-28] RASteer: Retain-Aware Activation Steering for Concept Erasure in Diffusion Models
链接: https://arxiv.org/abs/2610.01969
作者: Yongliang Wu,Haori Lu,Yulun Wu,Jinqi Luo,Xingyu Zhu,Yaoyao Liu
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 20 pages. Project page: this https URL
Abstract:Concept erasure aims to remove a target concept, such as a copyrighted style, a recognizable character, or unsafe content, from a pretrained text-to-image diffusion model while preserving its ability to generate other content. Existing activation steering methods build an erasure direction mainly from the target concept and adjust model activations along it at inference time. However, target and retained concepts often overlap in the model’s representation space, so this direction also contains shared components that retained concepts rely on. Steering directly along this direction can therefore suppress retained concepts and harm the generation of non-target content. To address this issue, we propose Retain-aware Activation Steering (RASteer), a training-free method. RASteer first builds a retain subspace from the concepts to preserve. Retain-Orthogonal Steering (ROS) then removes components aligned with this subspace from the erasure direction, making steering more specific to the target. Since fully removing the shared components can weaken erasure, we further introduce Overlap-Adaptive Calibration (OAC). At each layer and denoising step, OAC uses the overlap between the erasure direction and the retain subspace to control how much of each shared component is removed, balancing target erasure and concept preservation. Experiments on unsafe-content, instance, and artistic-style erasure across multiple backbones and benchmarks show that RASteer matches or outperforms the activation steering and weight editing baselines we evaluate, achieving a better balance between erasure and preservation.
[CV-29] SIEVE: Selective attention-value Suppression for Vision-Language Models Unlearning
链接: https://arxiv.org/abs/2610.01962
作者: Si Qi Goh,Cap Dang Xuan Kiet,Tat-Jen Cham,Kwok-Yan Lam
类目: Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:The ability of vision-language models (VLMs) to associate visual identities with biographical information creates a need for selective unlearning of personally identifiable information (PII) while preserving permitted knowledge about the same individual. This setting is challenging because both sensitive and retained information can share the same visual inputs and intermediate representations. We introduce SIEVE, a simple and effective framework for selective VLM unlearning. SIEVE directly regularizes attention-value representations while also controlling model outputs. SIEVE suppresses attention values for forget examples toward a constant zero, while preserving retain-example representations by matching them to a frozen reference model. These objectives are combined with sequence-level forget and retain supervision, enabling targeted forgetting without largely affecting retained knowledge. Extensive experiments show that SIEVE achieves state-of-the-art performance on unlearning with multiple model-modality settings, while maintaining competitive retained utility. Ablation studies further show that value suppression and negative cross-entropy contribute complementary forgetting signals, while reference-based value matching substantially reduces utility degradation. These results demonstrate that attention values provide an effective intervention point for selective multimodal unlearning when sensitive and retained knowledge are closely related.
[CV-30] EndoLive: Real-Time Style Transfer for Endoscopic Endonasal Skull Base Surgical Video
链接: https://arxiv.org/abs/2610.01956
作者: Griffin Hurt,Calvin Brinkman
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Complex surgical procedures around critical anatomy, such as the endoscopic endonasal skull base surgery, requires significant practice and training on the part of the surgeon before they are allowed to perform the operation on a live patient. This training in typically done in cadaveric specimens, due to them containing the same critical structures as a living human. However, cadavers are not a perfect 1-to-1 substitute for a living patient. The dead and preserved tissues of a cadaver are colored completely differently than a living human, and – without complex and expensive pumping systems – do not bleed in the same way. As a result, identifying the critical pieces of anatomy that make this procedure so complex can be quite different in a live case than in a surgeon’s cadaveric practice. This paper presents EndoLive, a framework for real-time style transfer between cadaveric endoscopic video and living human endoscopic video. Our method combines the ConStructS GAN model for realistic style transfer for surgical applications, with the HyPER-GAN model that can learn complex translations and perform them in real-time. We train EndoLive on unpaired cadaveric and live images taken from an endoscope, and test the trained model with cadaveric video, on a variety of devices. Experimental results demonstrate that EndoLive can perform cadaveric-to-live translation at speeds well above the minimum necessary for real-time, while maintaining semantic consistency of critical anatomical structures. Our source code is available at this https URL.
[CV-31] Anti-Persona: Disrupting Unauthorized Identity Binding and Recognition in Personalized Vision–Language Models
链接: https://arxiv.org/abs/2610.01944
作者: Abhishek Basu,Fahad Shamshad,Karthik Nandakumar
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Code available at this https URL
Abstract:Few-shot personalization enables large vision–language models (LVLMs) to learn user-specific visual concepts for applications such as personalized retrieval and subject-aware querying. However, it also creates a privacy risk: an adversary can bind a target identity from a few reference images and subsequently detect that identity in new images through natural-language queries. We introduce Anti-Persona, an image-level defense against unauthorized identity binding and recognition in personalized LVLMs. Our key insight is that identity personalization relies on visual features shared across multiple reference images. We aggregate these features into an identity prototype and optimize visually subtle perturbations that disrupt prototype alignment in the vision-encoder space. Spatial smoothing and low-frequency preservation further promote visual fidelity and practical resilience to image compression. The resulting protection does not depend on a specific prompt and supports both proactive anti-personalization and reactive image protection. Experiments on two representative personalized LVLMs demonstrate protection rates of up to 95.0% while preserving visual fidelity. The method remains stable across prompt variations and evaluated identity-query tasks, and improves black-box transfer under encoder mismatch.
[CV-32] Latent-Foresight: End-to-End Learning Predictable Representations for Latent World Models
链接: https://arxiv.org/abs/2610.01942
作者: Efstathios Karypidis,Spyros Gidaris,Nikos Komodakis
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Predicting the future evolution of a scene is a fundamental capability for world modeling. Recent work has shown that operating in the feature space of Vision Foundation Models (VFMs) yields semantically rich representations that support diverse future scene understanding tasks. However, existing approaches rely on two-stage pipelines, where VFM features are first compressed using fixed dimensionality reduction (e.g., PCA) or independently trained autoencoders, and a separate predictor is trained on top of the resulting frozen latent space. This decoupling between representation learning and temporal prediction, as well as approaches that apply predictors directly on raw VFM features, provides no guarantee that the latent space is structured for predictable dynamics. In this work, we propose Latent-Foresight, an end-to-end framework that jointly learns a latent tokenizer and a flow-based generative dynamics model, explicitly shaping the representation to support temporal predictability. To enable stable joint optimization, we introduce several key design choices that prevent latent collapse and align reconstruction with generative objectives. Extensive experiments show that our approach learns more temporally coherent latent representations and consistently outperforms two-stage baselines across multiple future scene understanding tasks and prediction horizons, while eliminating separate training stages, including during high-resolution adaptation. We provide the implementation code and model weights at this https URL
[CV-33] Fewer Tokens Better Action: GPT -6 Astra Robot Agents with 14% Higher Success Rate but 65% Fewer Tokens
链接: https://arxiv.org/abs/2610.01939
作者: Ruiyang Si,Jianxin Bi,Shunyu Yang,Rui Ni,Wenbo Huang,Qiang Wang,Shulong Jiang,Duomin Wang,Xiuyu Li,Haiwen Feng,Zhen Dong,Daquan Zhou
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Vision language model (VLM) agents can control robots through visual feedback and action primitives, but repeated model invocations and redundant observations incur substantial token overhead. We introduce PyRUA-Lean, an interactive code-execution framework that couples feedback-driven primitive composition with selective observation: the agent composes classical robot primitives and learned vision-language-action (VLA) policies into Python cells that perform conditional checks and local retries, returning only explicitly requested images and state feedback for replanning. Across 700 simulated task instances from LIBERO-PRO, RoboTwin 2.0, and RoboCasa365, we compare PyRUA-Lean with a tool-calling baseline using the same GPT-6 Astra planner and underlying robot primitives. Under equal LLM-call budgets, PyRUA-Lean increases overall success from 63.1% to 71.7%. On instances solved by both agents, it uses 49% fewer LLM calls and 65% fewer input tokens.
[CV-34] CLoSeR: Closing the Loop for Long-Context Streaming Reconstruction
链接: https://arxiv.org/abs/2610.01927
作者: Moyang Li,Zihan Zhu,Wei Zhang,Marc Pollefeys,Daniel Barath
类目: Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)
备注: Authors contributed equally to this work. Author order is interchangeable
Abstract:Feedforward foundation models have recently shown remarkable 3D reconstruction capabilities. However, existing models exhibit large tracking drift in long-context streaming reconstruction due to error accumulation. In this paper, we revisit loop closure with streaming reconstruction foundation models to enable accurate, drift-free, kilometer-scale reconstruction. Specifically, our method detects loop candidates through global descriptor retrieval, and constructs loop-conditioned windows to estimate the relative poses between looped frames. Given the observation that our adopted streaming reconstruction backbone produces a globally consistent scale, we optimize all frame poses on the SE(3) manifold with sequential and loop closure constraints, avoiding the pose graph optimization on the Sim(3) or higher-dimensional SL(4) manifolds employed in prior works. Extensive experiments show that our method reduces drift and produces consistent geometry on kilometer-scale sequences, significantly outperforming the state of the art. Code is available at this https URL.
[CV-35] DecomVoxel: Harnessing 3D-Native Priors with Guided In-situ Denoising Optimization for Decompositional Scene Reconstruction SIGGRAPH
链接: https://arxiv.org/abs/2610.01914
作者: Junfeng Ni,Zirui Zhou,Yixin Chen,Yu Liu,Nan Jiang,Zhifei Yang,Song-Chun Zhu,Siyuan Huang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: SIGGRAPH Asia 2026 - Journal Track (TOG). Project page: this https URL
Abstract:Decompositional scene reconstruction aims to reconstruct high-quality objects and background, yet existing methods still struggle with the level of quality under heavy occlusions. While generative priors offer a potential solution, 2D image-based priors often suffer from multi-view inconsistency due to a lack of 3D awareness. Conversely, 3D-native priors provide stronger structural inductive biases but frequently lead to spatial drift and misalignment within complex scenes. To address these issues, we propose DecomVoxel, formulating object completion as a guided in-situ denoising optimization that bridges 3D-native priors with neural scene reconstruction. Our framework introduces a reformulated epsilon-based distillation loss to ensure stable latent refinement, alongside adaptive spatial guidance that utilizes occupied and vacant anchors with temporal annealing to suppress generative hallucinations and mitigate spatial drift. Experiments on Replica and ScanNet++ show that DecomVoxel significantly outperforms state-of-the-art methods while faithfully preserving the original spatial layout, structural fidelity, and style-consistent texture. Our method pushes the boundary of decompositional reconstruction by delivering high-quality textured meshes with clean topology, geometry, and appearance, providing a robust solution for the decompositional reconstruction of complex real-world scenes. Code is available at this https URL.
[CV-36] MapLightning: Online Vectorized HD Map Construction with 1D Map Tokens
链接: https://arxiv.org/abs/2610.01905
作者: Shen Zheng,Anurag Ghosh,Mani Ramanagopal,Srinivasa Narasimhan
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Online vectorized HD map construction is essential for scaling safe autonomous driving and requires accurate, real-time inference. Prior methods typically rely on dense bird’s-eye-view (BEV) grids as the intermediate representation. We propose \textitMapLightning, which replaces the dense BEV grid with a compact set of 1D learnable map tokens. To construct map tokens from image features, we choose self-attention over vanilla cross-attention because it enables joint interactions and contextual aggregation among image and map tokens. Our transformer-based mapper concatenates map and image tokens, applies full self-attention, discards the image tokens, and retains the updated map tokens for decoding. This design offers three advantages. First, our representation is efficient, using fewer tokens, consuming less memory, and running faster. Second, the lightweight design allows the map decoder to use full rather than deformable cross-attention for better global context. Third, unlike BEV-based methods, our network does not use camera projection parameters, making it robust to camera-extrinsic perturbations. MapLightning uses up to 16.7 \times fewer intermediate tokens than dense BEV-based methods and achieves state-of-the-art accuracy and efficiency on nuScenes and Argoverse~2. Its lightweight variant surpasses MapTRv2 by +10.1 mAP on nuScenes and +16.2 mAP on Argoverse~2, while delivering 1.73 \times faster inference (40+ FPS) with 53% less memory. We further show improvements on uncertainty-aware map construction and downstream trajectory prediction. Code and models will be released.
[CV-37] Unsupervised Domain Adaptation for Enhanced Radiometer Image Precipitation Estimation using Conditional Flow Matching
链接: https://arxiv.org/abs/2610.01890
作者: Victor Enescu,Assaad Zeghina,Matthieu Meignin,Nicolas Viltard,Cécile Mallet
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:
Abstract:Deep generative networks have recently achieved unprecedented performance in precise image and video editing using sophisticated textual prompts. However, the effectiveness of such models heavily depends on access to very large supervised and annotated image datasets, which can be very difficult to obtain. This is particularly true for satellite instruments, which very rarely overlap with labelled data, and suffer from domain shifts in the rare occasions they do. In this paper, we investigate the potential of flow matching models for unsupervised domain adaptation of satellite radiometer images. Our main contribution is a novel unsupervised method that achieves precise domain alignment by leveraging parts of the deterministic ordinary differential equations in flow matching models, conditioned on different satellite instruments. A key strength of our approach is its ability to preserve essential information while adapting across any domains since the perturbations are in theory bijective. Extensive experiments conducted on the GPM-Core constellation show the benefit of our conditional domain adaptation, particularly in improving rain precipitation estimation from radiometer imagery.
[CV-38] Memory-Guided B-Roll Generation from User Video Collections
链接: https://arxiv.org/abs/2610.01884
作者: Cusuh Ham,Fabian Caba Heilbron,Josef Sivic,Bryan Russell
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Project page at this https URL
Abstract:We introduce an approach for collection-grounded B-roll sequence generation. Given a user’s video collection, a directive given in natural language, and a target duration, the goal is to produce a multi-shot sequence that complements the user’s primary footage (A-roll) while preserving the collection’s characters, settings, objects, and style. This task is challenging as one must choose the visual evidence from hours of captured footage that should guide the generation of each shot in the sequence. We address this challenge with MemComposer, a three-stage system that turns raw footage into a structured memory with visual references (characters, settings, objects, and style) and uses it to plan, retrieve, and generate grounded B-roll sequences. First, in a one-time offline stage, MemComposer constructs an entity-centric memory from raw video. Second, it uses the memory and user directive to plan a grounded sequence and retrieve conditioning frames for each shot. Third, it iteratively generates and critiques the sequence to enforce identity, setting, and sequence-level consistency. We evaluate MemComposer in a user preference study along two dimensions: prompt adherence and visual alignment to the user’s collection. Against an ungrounded text-to-video planner, MemComposer wins 60.0% of prompt-adherence and 92.8% of visual-alignment comparisons, showing the grounding benefit of collection memory and reference retrieval. Against retrieval-only sequences assembled from captured footage, MemComposer wins 94.5% of prompt-adherence comparisons, showing the value of generating missing shots, while retrieval-only sequences are preferred for visual alignment in 58.2% of comparisons.
[CV-39] EvenSplat: Coupled 2D-3D Decomposition for Gaussian Splatting under Exposure and Illumination Variation
链接: https://arxiv.org/abs/2610.01876
作者: Tongyu Wu,Jacob Edwards,Ziteng Cui,Caigui Jiang,Cheng Wang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:A surface photographed under even light presents nearly the same appearance from every angle; the same surface under uneven light does not. Exposure changes between views, illumination varies within a single image, and locally strong light sources leave one region bright and its neighbor in shadow. Multi-view reconstruction methods such as 3D Gaussian Splatting treat these lighting artifacts as if they were properties of the scene, entangling capture-specific illumination with the geometry and color they recover. We present EvenSplat, a framework that separates the two. EvenSplat couples an image-space illumination decomposition with an illumination field carried by the Gaussians, so that the same explanation of the lighting is shared between the two-dimensional and three-dimensional views of the scene; a camera-response network and a local exposure-compensation module absorb the global and residual differences that remain across training images. Through extensive experiments across multiple datasets and diverse forms of uneven illumination (cross-view exposure, spatial illumination variation, and high-contrast lighting) on both real-world captured and simulated benchmarks, EvenSplat generally outperforms state-of-the-art methods, particularly under high-contrast illumination.
[CV-40] From Pixels to Policy: A Multi-Agent System for Intervention and Geo-Spatial Decision Support
链接: https://arxiv.org/abs/2610.01870
作者: Hosam Elgendy,Utkarsh Mall
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Urban environments are shaped by design choices with long-term implications for health, safety, and quality of life, yet evaluating proposed interventions remains costly, time-consuming, and often impractical. Existing geospatial vision methods largely focus on monitoring urban indicators from aerial and street-view imagery, rather than proposing interventions and estimating their effects on such indicators. Moving beyond recognition, we introduce the problem of discovering interventions that improve target indicators for a given aerial or street-view image. We argue that a black-box indicator model, combined with a generative editing model, can serve as an implicit digital twin for testing intervention hypotheses. We present VIDA-Geo , a multi-agent system that explores this intervention space by coordinating segmentation, diffusion-based inpainting, and indicator scoring models to produce interventions that are both perceptually realistic and aligned with real-world policies. We evaluate our system on 8 indicators across aerial and street-view imagery, measuring changes in factors such as perceived safety and greenery. Our approach outperforms existing baselines in many cases, achieving up to 2X higher perceptual quality and policy alignment scores. Finally, our model provides users with multiple candidate interventions, supporting an expert city-planner-in-the-loop workflow.
[CV-41] LiteReality-Agent : An Agent ic System for Interactable 3D Indoor Scene Reconstruction
链接: https://arxiv.org/abs/2610.01863
作者: Zhening Huang,Yueyan Li,Johnathan Chiu,Xiaoyang Lyu,Matt Zhou,Yuxin Yao,Joan Lasenby,Shangzhe Wu
类目: Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR); Robotics (cs.RO)
备注: Code: this https URL Webpage: this https URL
Abstract:We present LiteReality-Agent, an agentic system for reconstructing real indoor environments as realistic, articulated, and simulation-ready 3D scenes from RGB-D scans. At its core, LiteReality-Agent formulates 3D reconstruction as a coding problem, in which a coding agent gathers evidence using specialised tools and iteratively edits a Python script, this http URL, which can be executed to produce a 3D digital twin of the room. With this formulation, we develop a robust observe-edit-verify harness that supports evidence gathering, measurement, verification, layout optimisation, simulation readiness, and quality control throughout the reconstruction process. LiteReality-Agent produces high-quality reconstructions suitable for simulation and downstream embodied AI tasks. Furthermore, as agent capabilities continue to improve rapidly, the system introduced by LiteReality-Agent remains a strong orchestration framework for future agents: it equips them with specialised tools, structured workflows, and robust verification mechanisms that substantially improve reconstruction quality and reliability. We demonstrate that LiteReality-Agent produces reconstructions that are more geometrically accurate, visually realistic, and simulation-compatible than those generated by recent frontier models, such as Astra and Fable. We therefore view LiteReality-Agent as a practical and important building block for robust real-to-sim systems. Both the source code and the data-capture application are publicly available. Code:this https URL
[CV-42] PhaseAT: Fourier Phase Adversarial Training for Medical Image Domain Generalization MICCAI2026
链接: https://arxiv.org/abs/2610.01807
作者: Ahmed Sharshar,Asif Hanif,Naveen Kumar Kummari,Mohammad Yaqub,Mohsen Guizan
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注: The paper is accepted in MICCAI 2026
Abstract:Reliable clinical deployment of deep medical image models is hindered by distribution shifts across scanners, sites, and acquisition protocols. Existing domain generalization (DG) methods often focus on style or intensity diversification, but they can still leave networks dependent on domain-specific texture correlations. Inspired by evidence that Fourier phase encodes semantic structure, we introduce PhaseAT, a phase-aware adversarial training framework for medical DG. PhaseAT forms phase-perturbed training views in the Fourier domain by iteratively updating a bounded phase perturbation while keeping the amplitude spectrum unchanged, thereby stressing spatial organization under matched appearance statistics. Perturbations are applied only to the luminance channel in YCbCr color space to avoid chromatic artifacts. Additionally, a simple phase-saliency mask concentrates updates on the most influential frequencies. The model is trained with a weighted combination of losses on clean and phase-perturbed samples, supporting both single-source and multi-source DG. We validate our method on two challenging medical datasets and demonstrate that PhaseAT achieves over 20% improvement in single-source domain generalization, outperforming several state-of-the-art DG methods. The code implementation is available at: this https URL.
[CV-43] Continuous Conditioning of VLAs with Augmenting EMG and Visual Task Descriptors IROS
链接: https://arxiv.org/abs/2610.01794
作者: Edward W. Staley,Connor O. Pyles,Rahul Hingorani,Frank Camargo,Griffin Milsap,Jared Markowitz,Matthew S. Fifer,Michael Wolmetz
类目: Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)
备注: Presented at IROS WORLDS Workshop 2026. Four main pages double-column format plus references and appendices
Abstract:Vision-Language-Action (VLA) models rely strongly on language for describing task information, despite having multimodal inputs. We hypothesize that other modalities in the state space may present opportunities for supplemental task conditioning, which may be particularly relevant in cluttered or otherwise ambiguous scenes. We introduce two tuned models to test this hypothesis: (1) an electrophysiology-conditioned VLA (EC-VLA) that incorporates 8-channel electromyography envelopes as continuous conditioning input concatenated to the proprioceptive vector, and (2) a visually-annotated VLA (VA-VLA) that incorporates visual segmentation annotations to the image inputs. On a cube-selection task evaluated across three participants, EC-VLA matches a language-prompted baseline in uncluttered, in-distribution conditions and substantially outperforms it in cluttered, out-of-distribution scenes. Similarly, VA-VLA shows modest improvements over a language-prompted baseline in in-distribution scenes with substantial improvement in cluttered, out-of-distribution trials. Together, these results provide strong evidence for the potential benefit of task-conditioning beyond language.
[CV-44] GIFTBench: Diagnosing Generalization in Image Forgery Localization and Informing Model Design
链接: https://arxiv.org/abs/2610.01778
作者: Baoke Dou,Ziye Wang,Hao Wang,Guoqing Cai,Wende Tan,Chenyang Si,Liucheng Guo,Yueming Lyu
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Reliable evaluation of image forgery localization (IFL) requires assessing models under diverse distribution changes, yet existing benchmarks often cover limited manipulation conditions or entangle multiple factors in cross-dataset evaluation. Consequently, aggregate performance provides an incomplete view of localization generalization. We introduce GIFTBench, a multi-axis benchmark of 115,013 manipulated images with pixel-level annotations spanning manipulation source, semantic target, editing operation, and composition complexity. GIFTBench supports axis-specific transfer analysis and evaluation on twelve external datasets. Its diagnostic studies reveal asymmetric cross-source transfer, recall-dominated failures, and heterogeneous degradation across semantic, operational, and compositional changes. Beyond diagnosis, the scale and diversity of GIFTBench provide a substantially broader training distribution than conventional IFL datasets. Training representative localizers on GIFTBench consistently improves their aggregate transfer to external datasets, showing that the benchmark serves not only as an evaluation tool but also as an effective training resource for cross-domain localization. Guided by the diagnostic findings, we further develop ForenScope, a detection and localization framework combining classification-adapted representations with multi-depth, multi-scale spatial features, learned layer fusion, and selective coarse-scale conditioning. Experiments show improved cross-dataset localization while retaining image-level detection capability. The GIFTBench dataset showcase page is available at this https URL.
[CV-45] VideoEvolve: Evolving Agent Harnesses for Video Temporal Grounding
链接: https://arxiv.org/abs/2610.01766
作者: Bingjun Luo,Yuhuan Fan,Jialin Guo,Siqi Li
类目: Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Video temporal grounding aims to localize events in videos from natural-language queries. For agents built around frozen video-language models, the harness determines how queries guide temporal predictions and how those predictions are refined. Manually refining these harnesses requires diagnosing grounding failures and coordinating changes to both agent workflows and instructions. We introduce VideoEvolve, a framework that automatically evolves agent harnesses for video temporal grounding. VideoEvolve uses a Cloze-Structured Harness Representation that preserves stage interfaces while leaving agent workflows and instructions open to evolution. Branch-Guided Harness Evolution preserves promising code branches for continued refinement, using execution feedback to guide local edits and validation to determine which improvements are carried forward. Experiments demonstrate improved grounding performance across multiple benchmarks. Component analyses identify instruction refinement as a consistent source of gains, while the benefits of evolved code vary across evaluation settings. Together, these results support automated harness evolution as an effective approach to improving video temporal grounding. Code is available at this https URL .
[CV-46] OneStreamer: Unifying Perception Memory and Proactive Response in Streaming Video Interaction
链接: https://arxiv.org/abs/2610.01762
作者: Xiangyu Zeng,Yuandong Yang,Zhiqiu Zhang,Yuhan Zhu,Xinhao Li,Qingyi Si,Dingyu Yao,Changlian Ma,Haoran Chen,Xinyu Chen,Yansong Shi,Junhao Zhou,Yifei Li,Jun Zhang,Chuanyu Qin,Chenxu Yang,Xinlei Yu,Kun Ouyang,Yuchen Shao,Qianshan Wei,Changhai Zhou,Jun Gao,Jiaqi Wang,Limin Wang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 29 pages, 12 figures, 20 tables. Project page: this https URL
Abstract:Streaming video LLMs must retain evidence before its relevance to future tasks is known and respond when sufficient evidence becomes available. The challenge is to form reusable factual memory without compromising real-time perception. We introduce OneStreamer, which jointly learns query-independent evidence recording and task response through a shared proactive generation process. Its Proactive Hierarchical Caption Memory (PHCM) produces time-grounded local-detail captions and summaries of completed events. Streaming caption targets supervise the interpretation of observed video prefixes during training. At inference, model-generated records complement a recent visual window, providing reusable factual context without revisiting historical visual features. Proactive State Transition Learning (PSTL) reduces the dominance of repeated waiting states by preserving supervision at all output anchors and selecting representative state-change and state-persistence tokens. We further develop a streaming data synthesis pipeline that aligns output content and timing with available evidence. Combining the resulting streaming captions and QA with cleaned open-source data yields OneStreamer-1M, a broad-coverage streaming video interaction dataset with over one million records spanning diverse tasks. Our 4B model achieves the best results among the compared methods across all eight evaluated streaming video understanding benchmarks. Ablations show that retaining generated captions improves historical QA without degrading real-time perception. PSTL also outperforms dense state supervision while supervising only 27.5% of annotated state tokens. Together, these results support proactive generation as a shared learning interface connecting perception, memory formation, and timely response in streaming video interaction.
[CV-47] PhysDEM: Physics-Defined Energy-Matching Diffusion for Spatiotemporal Field Generation under Scarce Measurements
链接: https://arxiv.org/abs/2610.01759
作者: Zhenyu Liang,Yining Huang,Yubo Zhao,Jack C.P. Cheng
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Generating and predicting spatiotemporal physical fields from scarce measurements is challenging, as observations are insufficient to characterize a distribution over complete fields. This limits conventional data-driven diffusion models that rely on full-field datasets. We introduce PhysDEM, a physics-defined diffusion framework that combines governing equations with spatially sparse observations to generate multiple plausible fields. First, we construct a Gibbs target by reweighting a measurement-conditioned Gaussian reference with PDE residual energy. Second, we derive an exact conditional-mean identity that reduces denoising to supervised learning of the standardized energy-induced mean correction. Third, a physics-displacement probability flow cancels Gaussian reference terms and enables amortized sampling with changing measurements through Gaussian conditioning, without retraining. Experiments on synthetic PDE systems and real-world-informed applications demonstrate that PhysDEM supports coherent field recovery and efficient sampling while maintaining stable diagnostics under tested noise levels, illustrating its practical value for field assessment. To our knowledge, PhysDEM is the first physics-defined diffusion model enabling amortized spatiotemporal field inference without preassembled full-field datasets.
[CV-48] GenCOPE: Syn2Real Generalized Category-Level Object Pose Estimation for Robotic Picking NEURIPS’26
链接: https://arxiv.org/abs/2610.01758
作者: Jian Liu,Wei Sun,Zhenqi Dai,Hui Yang,Jian Xiao,Nicu Sebe,Na Zhao
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted by NeurIPS’26
Abstract:Category-level object pose estimation (COPE), capable of generalizing to intra-class unknown objects, has become a core technique for robotic 3D scene understanding. However, existing COPE methods still require labor-intensive recollection of real-world training data for novel object categories, which limits their scalability in practical applications. This paper aims to achieve synthetic-to-real (Syn2Real) generalized COPE, where a model is trained solely on rendered synthetic data and directly generalized to real-world deployments. The central challenge lies in the significant domain gap between synthetic and real-world data, particularly in texture appearance. To address this, we aim to enhance domain generalization by learning domain-invariant representations that capture semantic commonalities among objects within the same category. We introduce 2D and 3D semantic consistency constraints to reduce the sensitivity of feature encoders to domain-specific features. In addition, we propose an end-to-end pose regression framework that performs 2D-3D cross consistency learning, leveraging dense cross-modality fusion to further refine pose estimation. Since simplicity and effectiveness are essential for real-world robotic deployment, our model operates exclusively on global features, yielding a highly lightweight and efficient architecture. Extensive experiments on the REAL275 and Wild6D benchmarks, as well as real-world robotic manipulation scenes, show superior Syn2Real generalization performance of our paradigm. Code and demos are released at this https URL.
[CV-49] Cog-VADU: A Training-Free Cognitive Reasoning Framework for Video Anomaly Detection and Understanding
链接: https://arxiv.org/abs/2610.01754
作者: Mohd Ubaid Wani,Sara Atito,Josef Kittler,Muhammad Awais
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注: Published in Transactions on Machine Learning Research (TMLR), 2026. 39 pages
Abstract:Video Anomaly Detection (VAD) aims to temporally localize abnormal events in videos. Most existing approaches rely on dataset-specific training and curated annotations, limiting generalization in open-set scenarios. Recent zero-shot methods based on Large Vision- Language Models (LVLMs) alleviate this dependency but often lack temporal continuity and structured reasoning. We propose Cog-VADU, a fully training-free framework that reformulates VAD as a sequential cognitive reasoning task. Cog-VADU introduces Chain-of- Anomaly Detection Thought Prompting (CoADTP), which unrolls an LVLM into a recurrent reasoning chain across video segments. By propagating structured rationales over time, the model maintains implicit temporal memory, enabling robust discrimination between com- plex anomalies and high-motion normal activities. To improve reliability, we further design a cross-modal re-ranking stage that aligns textual rationales with visual embeddings, enforcing semantic consistency and temporal coherence for refined and stable predictions. Extensive experiments on multiple public VAD benchmarks demonstrate that Cog-VADU achieves competitive zero-shot performance. Moreover, cross-model evaluations show that CoADTP consistently enhances reasoning-based anomaly detection in a model-agnostic manner, pro- viding interpretable and generalizable anomaly understanding for real-world applications.
[CV-50] FFBL-Coop: Association-Decoupled Cooperative 3D Multi-Object Tracking ICLR2027
链接: https://arxiv.org/abs/2610.01750
作者: Haoxin Wu,Xiaokai Bai
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 9 pages (main content), 21 pages total including references and appendix; 11 figures; under review as a conference paper at ICLR 2027
Abstract:Cooperative 3D tracking must integrate complementary observations across agents and time while maintaining consistent identities. When evidence integration and identity inheritance share a matching decision, errors arising from cross-view appearance differences and spatial misalignment can compromise both feature fusion and track continuity. We propose FFBL-Coop, a fuse first, bind later framework that separates instance admission from identity management. Confidence-ranked Slot Admission (CSA) allocates cooperative queries to available ego slots using confidence and spatial proximity. Unified Representation Aggregation (URA) uses cooperative semantic features and aligned anchors to guide ego-feature retrieval, refining the augmented query bank within a shared transformer decoder. After refinement, Cooperative-Priority Identity Anchoring (CPIA) combines learned association with persistent mappings to establish accepted identity assignments across frames. A shared codebook reduces transmitted payload while retaining AP and AMOTA close to the uncompressed variant. FFBL-Coop achieves AMOTA/AP of 0.611/0.548 on V2X-Seq and 0.688/0.653 on Griffin-25M. Code will be released.
[CV-51] End-to-End Learning vs. Modular Architectures: Comparative Insights into Autonomous Driving Systems
链接: https://arxiv.org/abs/2610.01746
作者: Kartik B. Kapse
类目: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注: 27 pages, 7 figures, 4 tables
Abstract:Autonomous driving systems have become a central focus of intelligent transportation research, with End-to-End Learning and Modular Architectures offering two prominent design paradigms for their implementation. E2E Learning uses deep learning algorithms to map raw sensory inputs directly to driving actuators, providing a streamlined and adaptable solution. while Modular Architectures employ a pipeline-based approach, dividing the system into distinct subsystems for perception, cognition, planning, and control. This paper presents a comprehensive comparative analysis of these paradigms, focusing on their strengths, limitations, and trade-offs to provide insights into their suitability for various autonomous driving applications. The study evaluates key factors such as interpretability, scalability, robustness, and real-world applicability. While End-to-End Learning emphasizes simplicity and adaptability in dynamic environments, it lacks transparency and is highly dependent on large datasets. Conversely, Modular Architectures offer superior interpretability and task-specific optimization, but face challenges related to integration complexity and scalability. To address these limitations, hybrid approaches that combine the strengths of both paradigms have emerged, offering a promising direction for overcoming these challenges. Beyond this comparative synthesis, following work proposes a Four-Dimensional Architecture Selection Framework, comprising twelve binary criteria across safety, operating environment, data/computational resources, and deployment context, and validate it against ten published autonomous driving systems, correctly recommending 7/10 deployed architectures. This work synthesizes existing literature to highlight key trade-offs between the paradigms and identifies hybrid architectures as a promising direction for future research.
[CV-52] 3DROID: A Renderable 3D Gaussian Dataset with Measured Per-Scene Reliability
链接: https://arxiv.org/abs/2610.01744
作者: Wonguen Cho,Junhoo Lee,Nojun Kwak
类目: Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)
备注: 12 pages, 3 figures
Abstract:Robot manipulation models primarily reason from 2D observations while acting in the 3D physical world. To bridge this gap, recent work has augmented robot data with geometric priors such as depth, point clouds, and 3D trajectories, while renderable 3D Gaussian representations provide another promising form of 3D supervision. However, 3DGS representation is designed mainly for photometric fidelity and may not preserve real-world metric scale, particularly when the supplied camera extrinsics are unreliable. We study the effect of extrinsic reliability and pose conditioning on feed-forward 3DGS, and propose a calibration-aware pipeline that anchors reconstructed scenes to the robot’s metric workspace. Our experiments show that pose conditioning improves novel-view fidelity, while its geometric benefit depends on the reliability of the injected extrinsics. Using this pipeline, we present a renderable, metric-pose-anchored dataset with scene-level reliability information for robot manipulation research. Our dataset is available at this https URL
[CV-53] World Motion Models: Flexible Sequence Modeling of SE(3) Trajectories NEURIPS2026
链接: https://arxiv.org/abs/2610.01742
作者: Jiahui Lei,Qianqian Wang,Trevor Darrell,Angjoo Kanazawa
类目: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted at NeurIPS 2026 (Spotlight). Url: this https URL
Abstract:Equipping artificial agents with spatial intelligence requires a comprehensive generative prior over the dynamic 3D world. We propose World Motion Models (WMMs) that capture “what was, is, and will be where across time” via sparse SE(3) pose trajectories. WMMs are built on the observation that elements of dynamic scenes can be well approximated by a set of rigid SE(3) trajectories, a minimal yet expressive primitive for 4D modeling. This representation unifies articulated objects, human bodies, hand-object interactions, piecewise-rigid scene dynamics, camera motion, and even robot states and actions into a single shared space. Given this representation, we cast the joint distribution of these entities as a flexible sequence modeling problem, utilizing flow-matching with per-token noise levels. Coupled with a context token mechanism for non-sequential conditioning, this formulation supports any-to-any marginal conditioning across an arbitrary number of entities and time steps. Tasks such as future prediction, motion infilling, model-predictive control, inverse kinematics, cross-embodiment retargeting, and policy learning all reduce to the application of different masks over the same network. Experiments on 6 diverse applications of 3D vision and robotics demonstrate the versatility and flexibility of WMMs with strong performance.
[CV-54] ATI-VLA: Action-Centric Predictive Vision-Language-Action Models via Actionable Alignment Then Adaptive Injection NEURIPS2026
链接: https://arxiv.org/abs/2610.01741
作者: Yijie Zhu,Rui Shao,Jie He,Wei Li,Bo Zhao,Yelin Wang,Xiaochen Yuan,Tao Tan,Miao Zhang,Xiaojiang Peng,Zitong Yu
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted to NeurIPS 2026. Project page: this https URL
Abstract:Predictive Vision-Language-Action (VLA) models aim to improve robotic manipulation via future observation or world dynamics forecasting. However, existing approaches often fail to realize this potential and underperform direct action prediction models. We argue that these limitations stem from modality misalignment between observations and actions, together with joint optimization conflicts that drive learning away from an action-centric objective. To this end, we introduce ATI-VLA, an Action-Centric Predictive Vision-Language-Action framework via Actionable Alignment Then Adaptive Injection. Specifically, it follows a two-step design: 1) Actionable Representation Alignment via a Shared Codebook. It aligns predictive observation and action representations by mapping both modalities into a shared discrete latent space via a unified codebook, making predictive observation latents readily usable for action generation and mitigating modality misalignment. 2) Action-Centric Adaptive Injection of Predictive Latents. Building upon this, it then injects predictive observation latents into action decoding as explicit predictive priors via a lightweight adaptive side-path, enabling adaptive predictive guidance under a single action-centric objective. Extensive experiments on both simulation and real-world robotic tasks demonstrate that ATI-VLA achieves state-of-the-art performance with faster convergence.
[CV-55] Rethinking Memorization Mitigation in Diffusion Models: Reinforcing Text Conditioning
链接: https://arxiv.org/abs/2610.01723
作者: Hyungjun Joo,Sehwan Kim,Hyeonggeun Han,Sangwoo Hong,Jungwoo Lee
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Text-to-image diffusion models have achieved remarkable progress in image synthesis, yet can exhibit memorization by closely reproducing individual training examples. Effective mitigation must preserve useful prompt information to guide alternative depictions. We introduce a training-free method that redistributes cross-attention with Gaussian smoothing before reinforcing content-token contributions and attenuating padding contributions, without additional denoiser evaluations. With this intervention, stronger content conditioning can improve prompt alignment at comparable training-image similarity. A local analysis identifies when reinforcement preserves shared value information while redistribution reduces localized attention mass. On Stable Diffusion v1.4 and v2.0, all evaluated smoothing widths lie on the empirical Pareto frontiers for training-image similarity versus both prompt alignment and image preference. A configuration selected on Stable Diffusion reduces template reproduction in DeepFloyd IF without further tuning. These findings support jointly controlling conditioning allocation and strength to generate prompt-consistent alternatives.
[CV-56] CoEvolve: Construct-to-Edit Visual Grounding with Bidirectional State Refinement
链接: https://arxiv.org/abs/2610.01710
作者: Dongwei Sun,Yujie Zhang,Bowen Yao,Pei Liu,Jing Yao,Xiangyong Cao
类目: Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Visual grounding localizes an object described by language with a bounding box. Most multimodal grounding models compress target identification, spatial reasoning, and boundary estimation into one terminal prediction. Free-form rationales make reasoning linguistically explicit but do not necessarily expose measurable, editable spatial states. Intermediate localization errors are therefore difficult to diagnose and correct, allowing incorrect region choices and imprecise boundaries to persist in the final box. We introduce CoEvolve, a construct-to-edit framework that separates grounding into explicit state construction and state editing. Region-Evolution Reinforcement (RER) organizes grounding analysis into a progressive semantic–spatial trajectory, with each reasoning step committing to an explicit candidate region. Bidirectional Denoising Refiner (BDR) treats the reasoning text as fixed semantic context and refines the trajectory’s coordinate fields through bidirectional same-position reconstruction. Geometry- and behavior-level objectives provide target geometry and edit-preference signals for consolidating reliable candidates, preserving accurate inputs, or correcting toward annotations. Evaluations cover natural-image and remote-sensing grounding. With a 9B backbone, CoEvolve rivals models up to 241B parameters in grounding accuracy. Under controlled corruption, a single BDR pass improves mean box overlap by over 27 percentage points, demonstrating strong recovery from substantial localization errors. State-source comparisons further support the complementarity of explicit state construction and source-matched editing. The project is at this https URL.
[CV-57] MEGA: Object-Level Mesh Extraction from 3D Gaussian Splatting via Spatial Visual Distillation
链接: https://arxiv.org/abs/2610.01707
作者: Liwei Liao,Yingkui Zhang,Qianqian Tong,Ronggang Wang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Mesh extraction from 3D Gaussian Splatting (3DGS) aims to endow 3D Gaussians with accurate geometric structures, enabling explicit and precise 3D occupancy. However, existing methods primarily focus on scene-level mesh extraction, making them unable to represent object-level occupancy and often resulting in non-watertight surfaces. To overcome these limitations, we propose \textbfMEGA (\underlineMesh \underlineExtraction from \underlineGAussians), a ``segment-then-mesh’’ framework for extracting object-level, watertight meshes from complex 3DGS scenes. At the core of MEGA are \textbfSpatial Visual Distillation (SVD) and a mask-guided neural surface reconstruction module. SVD treats the 3DGS model as a teacher, sampling diverse camera poses and rendering the corresponding views of each segmented object. These observations are then used to train a mesh reconstruction model through photometric supervision. Extensive experiments on several widely used benchmarks demonstrate that MEGA achieves state-of-the-art performance in recovering accurate object-level 3D occupancy. Moreover, MEGA enables complex physical interactions by combining high-quality object-level meshes for geometric occupancy with 3DGS representations for photorealistic rendering.
[CV-58] Architectural Sampling: Test-Time Scaling via Computational Diversity in Frozen Vision-Language Models
链接: https://arxiv.org/abs/2610.01687
作者: Akshit Singh,Shyam Marjit,Wei Lin,Leonid Karlinsky,M. Jehanzeb Mirza
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:
Abstract:Test-time scaling often seeks better answers by sampling multiple responses from a frozen model, yet conventional temperature sampling generates every candidate along the same fixed computation path. We introduce architectural sampling, a training-free method that generates candidates through distinct forward computations by reusing selected blocks of decoder layers. Varying the block location and repetition count introduces computational diversity without updating model weights or adding auxiliary parameters. Across five Qwen checkpoints and twelve multimodal benchmarks, architectural sampling improves pass@9 over standard-path temperature sampling by 6.58 percentage points on average at the same nine-candidate budget. Reusing early layers yields the strongest gains, and the improvement in candidate coverage persists even under greedy decoding. The resulting candidates show lower lexical overlap and improve accuracy when used as rollouts for label-free test-time reinforcement learning. These findings extend the benefits of our architectural sampling beyond candidate coverage, demonstrating more effective learning from a model’s own outputs.
[CV-59] Beyond Leaderboard Scores: A Deployment-Focused Protocol for Interpretable Tracking Evaluation in Pedestrian-Centric Environments
链接: https://arxiv.org/abs/2610.01682
作者: Dominik Wojcikiewicz,Diego Paez-Granados
类目: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
备注: 8 pages, 7 figures; supplementary video provided as ancillary material. Submitted to IEEE Robotics and Automation Letters (RA-L)
Abstract:Mobile robots operating among pedestrians need trajectories that become available quickly, remain spatially credible through missed observations, preserve identity, and fit within an embedded computing budget. Aggregate tracking scores provide limited insight into when and how trajectories fail, while varying detector inputs can confound tracker and detector quality. We present a deployment-focused, tracker-only evaluation protocol that uses shared detections to isolate tracker behavior and directly evaluates initialization, detector-gap continuation, identity recovery, close-neighbor association, and load-dependent tracker-step runtime, while Higher Order Tracking Accuracy (HOTA) is retained as a complementary aggregate measure. We apply the protocol to the JackRabbot Dataset and Benchmark (JRDB) using six open-source trackers and our lightweight Pedestrian Reference Tracker (PedRefTrack), together with a GT-assisted variant that estimates the remaining tracker-side gap under idealized association and motion. Under fixed detections, the non-GT trackers span only 24.26%-29.67% HOTA yet exhibit markedly different capability profiles. After 1.0 s without detector support, no tracker without GT assistance maintains spatially correct, same-identity output in more than half of eligible cases, making missing-observation continuation the dominant limitation among the tested properties. Close-neighbor failures are smaller and increase mainly at the shortest separations. Tracker-step runtime on an NVIDIA Jetson Orin is heavy-tailed and load-sensitive, causing several trackers to fall below the 10 Hz real-time target in crowded frames. The protocol provides a reproducible way to characterize tracker behavior and deployment suitability in pedestrian-centric environments. Code and evaluation scripts are released at this https URL.
[CV-60] When Text-to-Image Helps Editing: The Effects of Conditioning During Denoising ICLR2027
链接: https://arxiv.org/abs/2610.01681
作者: Lidia Troeshestova,Alexander Ustyuzhanin,Sergey Kastryulin
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Under review as a conference paper at ICLR 2027
Abstract:Unified models are trained for both instruction-based image editing and text-to-image (T2I) generation, but standard editing pipelines keep source-image conditioning throughout denoising. We ask whether editing can benefit from T2I, and study how the effects of conditioning vary across edits and denoising stages. In pure editing, source attention declines for some edits over the sampling trajectory. This observation led us to task switching, which lets the model draw on its T2I capabilities. Across three unified editors and four benchmarks, switching to the T2I task for bounded intervals improves edit quality, while mean perceptual preservation remains close to pure editing across all three models. Unified editors therefore benefit from using both conditioning modes they are trained for, and the timing of the switch sets the balance between quality and preservation.
[CV-61] Do MLLM Judges Judge the Edit? Auditing Bias in Image Editing Evaluation with Verified Quality Preservation
链接: https://arxiv.org/abs/2610.01670
作者: Yuan Huang,Zirui Song,Xiuying Chen
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注: 30 pages, 9 figures
Abstract:Multimodal large language models (MLLMs) are increasingly used as automated judges for instruction-based image editing and as reward signals for model training. However, systematically auditing whether these judges are influenced by cues irrelevant to editing quality is challenging because visual interventions may themselves alter the quality being evaluated. A judgment shift can therefore be attributed to bias only when the intervention is verified to preserve the underlying editing quality. To address this challenge, we introduce EditJudgeBias, a counterfactual benchmark with verified quality preservation, comprising 1,196 real editing samples and 13 cues injected across four evaluation sites. We verify quality preservation for the requested edit using calibrated multimodal validators, controls, and human inspection. We then audit five MLLM judges along three complementary dimensions: invariance to quality-preserving cues, agreement with human judgments, and stability of pairwise preferences. Importantly, observed shifts are evaluated against each judge’s own zero-dose and re-query noise floors rather than against zero. Experiments show that quality-preserving cues move every judge beyond its own noise. Fabricated majority opinions increase ratings, irrelevant visual elements cause larger shifts than whole-image manipulations, and swapping candidate order reverses up to 60.9% of pairwise decisions. Edit-region cues also tend to reduce human agreement. The three measures characterize judges differently, showing that robustness cannot be captured by a single metric.
[CV-62] DiVid: Diagnosing Dimension-Specific Diversity Collapse in Video Generation Models
链接: https://arxiv.org/abs/2610.01661
作者: Huanran Hu,Zihui Ren,Dingyi Yang,Zhinan Song,Guozheng Wu,Tiezheng Ge,Qin Jin
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Despite remarkable progress, video generation models often produce highly similar outputs when repeatedly sampled from the same prompt, limiting their usefulness for creative exploration. Existing diversity evaluations primarily rely on global scalar metrics, which obscure where diversity collapses in the spatiotemporal space of videos. We introduce DiVid, a dimension-level diagnostic framework that decomposes video generation diversity into six interpretable dimensions: Semantic, Style, Subject, Scene, Motion, and Camera. Each dimension is measured through a reproducible computer-vision pipeline and analyzed alongside quality and instruction faithfulness to examine potential trade-offs. Systematic evaluation of representative video generation models reveals that diversity is highly dimension-specific: models with strong global diversity scores still collapse on specific factors, particularly Motion and Camera. These rankings persist after filtering unfaithful generations, indicating genuine capability differences rather than off-prompt outputs. Beyond measurement, controlled prompt interventions identify two fundamental bottlenecks: default mode convergence, where models fall back to dominant patterns under open-ended prompts; and realization gaps, where models fail to faithfully realize diverse, explicitly requested alternatives, particularly for temporal factors. The larger faithfulness losses for temporal factors highlight the difficulty of controlling motion and camera variation through text alone. DiVid thus shifts the study of diversity from measuring whether it exists to diagnosing where and why it collapses, and provides actionable directions for dimension-aware training objectives and control signals. The framework will be released to facilitate future research on diverse and controllable video generation.
[CV-63] Not All Error Yields to Scale: Where Scaling Stops in Vision-Language Inference
链接: https://arxiv.org/abs/2610.01640
作者: Xinye Zhao,Yunkai Dang,Yunchen Wu,Wenbin Li
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:
Abstract:Vision-language models (VLMs) face a fixed-budget trade-off between processing more visual information for fine-grained perception and using a larger language backbone for complex reasoning. Existing studies do not tell us which combination of backbone size and input resolution to deploy, especially in high-resolution deployments. To address this gap, we propose the Separable Law that describes how VLM performance changes with language backbone size and visual token count. We fit the law to measurements from 26 InternVL and QwenVL models, with language backbone sizes from 1B to 72B, on four high-resolution benchmarks with image sizes from 224 pixels to 8K. We find that the questions responding to scaling can be predicted from the skill they require, while a substantial fraction never responds at all. We also find that the two model families gain similarly from a larger backbone, while their gains from more visual tokens differ sharply. Combined with a cost law, the Separable Law gives a closed-form rule for allocating compute between backbone size and visual tokens. When deployment is limited to available configurations, the law identifies model and image sizes that perform close to the best feasible choice under the same budget. We hope our work offers a principled way to decide how much a model should be allowed to see at high resolution, given what it must reason about.
[CV-64] Fusing Visual and Textual Representations via Multi-layer Fusing Transformers for Vietnamese Visual Question Answering
链接: https://arxiv.org/abs/2610.01637
作者: Cong Phu Nguyen,Huy Tien Nguyen,Tung Le
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:In recent decades, artificial intelligence has made significant progress in understanding and interacting with images. One of the important applications of this technology is Visual Question Answering (VQA), a research field that requires computers to understand and answer questions about images in a natural manner. Despite extensive research and development in VQA for English, there have been very few similar efforts made for other languages, especially Vietnamese. This gap presents a significant challenge and opportunity for the advancement of VQA technology in the Vietnamese language context. By bridging this gap, the field of Vietnamese VQA not only enriches the diversity of research in artificial intelligence but also enables practical applications in various domains, such as education, healthcare, and entertainment, catering to Vietnamese-speaking populations worldwide. Thus, the exploration and development of Vietnamese VQA systems hold immense potential for advancing both research and practical applications in the intersection of computer vision and natural language processing. In this paper, we propose a Multi-layer Fusing Transformer model utilizing a cross attention module to combine multiple modality features of images and texts from different layers in an aggregated representation. Our architecture allows us extract information from low level to high level. Through detailed experiments and ablation studies, our model achieves promising results against the competitive baselines in ViVQA dataset for Vietnamese language.
[CV-65] Beyond Domain-Level Adaptation: Margin-Oriented Semantic-Appearance Interaction Correction for Personalized Federated Vision-Language Models
链接: https://arxiv.org/abs/2610.01625
作者: Wentao Yue,Qingyu Mao,Tianyou Lai,Ahmed M. Abdelmoniem,Qilei Li
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Federated parameter-efficient fine-tuning enables distributed clients to adapt pretrained vision-language models without sharing raw data or updating the full backbone. Its effectiveness, however, is limited by domain heterogeneity across clients. Existing personalized methods separate globally shared knowledge from client-specific style, but they largely treat each domain as a class-agnostic transformation. We show that this abstraction is insufficient: the cross-domain displacement associated with a fixed domain varies across semantic classes, and only a subset of these class-domain residuals damages the image-text decision margin. We therefore propose Margin-Oriented Semantic-Appearance Interaction Correction (MOSAIC), which first constructs a decision-aware harmfulness score that measures whether a training-derived class-domain residual favors a competing text prototype over the true class. It then models fine-grained class-domain interactions with a low-rank residual adapter whose class factors and residual basis are globally shared while domain factors remain client-private. An image-conditioned gate further controls candidate-wise correction, and harmful-pair-aware reweighting prioritizes decision-relevant residuals during local optimization. Extensive experiments on Office31, OfficeHome, and DomainNet100 demonstrate that MOSAIC consistently improves macro-client top-1 accuracy across all evaluated domain-shift and joint domain-label-shift settings.
[CV-66] Oneira: From Open-Ended Generation to Open-World Interaction in Video World Models
链接: https://arxiv.org/abs/2610.01614
作者: Xindi Yang,Baolu Li,Liam Lee,Zhenfei Yin,Songxin Zhang,Zhuoyang Song,Xu Jia,Jianfei Cai,Tien-Tsin Wong,Bingyi Jing,Mengyue Yang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Project page: this https URL
Abstract:Generative video world models can now synthesize open-ended environments that agents can navigate and interact with in simple ways. Yet open-ended generation does not imply full interaction: as a generated world expands, newly created content through navigation should expand what the agent can act upon, and as the agent changes the world, those changes should become persistent parts of the environment rather than transient visual effects. We characterize these two requirements as Open-World Interactivity, where newly generated or encountered entities are incorporated into the actionable world, and Persistent State, where interaction outcomes are committed to the world state and continue to influence subsequent observations and interactions. We present Oneira, an interactive video world model that closes the loop between generation and interaction through an explicit, extensible world state managed by a coding agent. Given the current observation and an action or high-level goal, the agent reads the world state, grounds the relevant entities, plans the interaction, and writes its outcome back into a world state table. When exploration reveals new objects, the agent incorporates them from generated observations, allowing the interaction space to expand with the generated world. Meanwhile, previously induced state changes are carried across video segments, making the consequences of interaction persistent parts of subsequent world evolution. The updated world state is rendered along the camera action trajectory into a coarse conditioning video, from which a video generator fills in the appearance, motion, and interaction details not represented in the state. Experiments show that Oneira enables direct and consistent interaction with newly generated objects, while preserving the effects of prior interactions over long horizons. Project page: this https URL
[CV-67] Hob-VL: A Benchmark for Visually Grounded Boolean Reasoning
链接: https://arxiv.org/abs/2610.01605
作者: Yuzhou Wang,Emile Anand,Ijay Narang
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Logic in Computer Science (cs.LO)
备注: 29 pages, 6 figures, 14 tables
Abstract:Reliable visual reasoning requires composing multiple visual observations and returning consistent answers to logically equivalent questions. We introduce Hob-VL, a benchmark for visually grounded Boolean reasoning. Hob-VL comprises two tasks: (1) evaluating whether a Boolean rule holds in an image, and (2) identifying the (unique) object satisfying a Boolean description. Hob-VL contains 6,000 human-verified balanced Yes/No questions, each defined by a Boolean combination of ten visual statements, across 1,000 generated scenes and 46 diverse labeled photographs, along with 1,000 object-identification questions over the same photographs. Our question families are deliberately constructed to challenge reasoning through misleading local cues and nested logical operations, and include symbolic and structured natural-language presentations. Across eight model configurations with thinking disabled or minimized, Boolean accuracy ranges from 48.52% to 50.57%, while the identification accuracy reaches at most 43.0%. A thinking-enabled GLM configuration achieves uneven gains while retaining substantial errors and inconsistencies. Hob-VL exposes these failures through executable reference answers and matched evaluations.
[CV-68] Before It Fades: Reinforcing Temporal Representations at Inference Time in VideoLLM s NEURIPS2026
链接: https://arxiv.org/abs/2610.01595
作者: Youngwoo Shin,Yusung Ro,Minseo Kim,Junmo Kim
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注: Accepted to NeurIPS 2026
Abstract:Video Large Language Models (VideoLLMs) receive frames in sequential order and interpret how visual content evolves along the temporal axis, yet temporal reasoning remains a persistent weakness across architectures. Reversing the frame order of a video, a transformation that should invert temporal answers, often leaves the final prediction unchanged. We investigate where this failure originates by defining the temporal divergence vector \tau_l , the layer-wise representational difference induced by reversing temporal order. Tracking its magnitude across layers reveals a consistent temporal divergence profile where the divergence peaks at intermediate layers and progressively diminishes toward the output. We confirm this peak is specific to temporal reasoning and functionally critical for predictions, establishing that VideoLLMs acquire temporal information at intermediate layers but fail to maintain it to the output. This progressive fading motivates our method, Temporal Activation Injection (TAI), which extracts \tau_l at the peak of the profile for each input and reinjects it into subsequent layers following the measured decay. TAI requires no training and consistently improves temporal reasoning across three VideoLLMs and four benchmarks with negligible impact on non-temporal tasks. Code is available at this https URL.
[CV-69] wo Routes to the Middle: Placement Search and Brain Readouts Converge on Where Continual Learners Should Specialize
链接: https://arxiv.org/abs/2610.01590
作者: Yuan Huang,Zihan Chen,Runbin Zhang,Hongwei Ding,Changzeng Fu,Shiqi Zhao
类目: Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)
备注: 21 pages, 12 figures
Abstract:Continual learners that keep a task-specific adapter in every block of a pre-trained vision transformer accumulate storage linearly with the number of tasks; keeping task-specific adapters in only a few blocks curbs this growth but raises the question of where to place them. We investigate this question from two perspectives. Algorithmically, training all contiguous four-block placements yields an inverted U: final accuracy peaks at intermediate depth and varies by up to 3.5 percentage points (pp), while inexpensive criteria based on weight spectra or activation statistics favor the deepest blocks. From neuroscience, the hierarchical organization and intermediate-stage plasticity of the visual cortex motivate us to ask whether a measurement taken outside the learner can guide layer specialization without placement search. LS-B observes the first tasks through a frozen fMRI encoding model of twelve human visual areas and commits task-specific capacity once to the blocks whose readouts vary most across tasks relative to their stable structure. Across three ViT-B/16 backbones, LS-B yields stable, backbone-specific allocations. On the two backbones with placement search, AugReg and iBOT, the selected blocks overlap the intermediate-depth region identified by search. Under matched storage and observation budgets, the selected blocks outperform the shallowest and deepest four-block configurations. On Split ImageNet-R, LS-B uses 60% of full-BiLoRA adapter storage while remaining within 1.5 pp of its final accuracy. The allocation requires no labels or backpropagation, adds under 0.6% runtime, and exhibits backbone-specific cortical signatures.
[CV-70] PAGER: Partial-to-global Alignment via Geometric and Relational Distillation
链接: https://arxiv.org/abs/2610.01589
作者: Akira-Miranda Adeyomi Adeniran-Lowe,Binod Singh,Lars Arnold Dethlefsen,Lazaros Nalpantidis,Theodora Kontogianni
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Pretrained 3D encoders are typically developed on globally reconstructed scenes expressed in a consistent world coordinate frame, whereas embodied systems must reason from partial, viewpoint-dependent observations in camera coordinates. We show that this shift from globally learned 3D feature spaces to realistic partial observations exposes a severe representation mismatch, which we find consistently across representative state-of-the-art encoders, including Sonata and Concerto. A frozen Sonata encoder with a global linear probe achieves 72.47 mIoU on full ScanNet scenes, but 2.57 mIoU on single-frame camera-coordinate inputs. Training-free gravity alignment recovers performance to 41.64 mIoU, showing that coordinate-frame mismatch is a dominant source of degradation but cannot be fully resolved through canonicalization alone. We introduce PAGER, a label-free adaptation method that aligns partial-view features with a frozen global 3D semantic space using only paired partial/global geometry. It learns lightweight adaptation modules while keeping the pretrained encoder and global segmentation probe frozen. Matched-point feature alignment anchors partial features to their global counterparts, while relational supervision preserves their similarity structure with respect to the global representation. Global geometry provides supervision only during training. Inference operates directly on the partial observation. Without partial-view labels, PAGER outperforms label-supervised PEFT on both Sonata and Concerto, and in zero-shot ScanNet \rightarrow ScanNet++ transfer surpasses fully fine-tuned Sonata ( 53.93 vs.\ 48.09 mIoU), suggesting that preserving the frozen global representation can improve cross-dataset transfer.
[CV-71] Revisiting Cross-Reconstruction for Generalizable Deepfake Detection
链接: https://arxiv.org/abs/2610.01544
作者: Bingjian Yang,Shilei Zhao,Zheng Wang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Existing image forgery detectors often suffer from generalization to unseen manipulation methods due to the limited ability to capture transferable forensic cues. Recent cross-reconstruction based methods attempt to improve generalization through semantic-artifact disentanglement, but typically align heterogeneous artifacts across generators and exclude artifact representations during reconstruction, which may overlook the inherent diversity and visual cues of manipulation artifacts. In this work, we revisit cross-reconstruction and introduce an artifact-oriented disentanglement framework for robust image forgery detection. We argue that \textbfartifact diversity, i.e., the intrinsic variations of manipulation artifacts introduced by different generation processes, contains complementary forensic cues rather than undesirable domain variations. Instead of enforcing explicit artifact alignment, our framework preserves diverse artifact characteristics through semantically aligned cross-generator reconstruction. Furthermore, we incorporate artifact representations into the reconstruction process and introduce a masked frequency-aware reconstruction strategy to emphasize manipulation-related residuals while reducing semantic interference. This design enables the model to learn transferable forensic representations from diverse artifacts. Extensive experiments on multiple benchmark datasets demonstrate improvements under both cross-dataset and cross-generator evaluation settings. Further analysis and ablation studies validate the effectiveness of artifact diversity preservation and artifact-aware cross-reconstruction.
[CV-72] Synthetic training for long-tail haemorrhagic lesion segmentation in data-scarce settings MICCAI2026
链接: https://arxiv.org/abs/2610.01542
作者: Yuan Cao,Sumeet Dash,Antonia Zachariadis,Stefanie Schreiber,Katja Neumann,Jose Bernal
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted: MICCAI 2026 SASHIMI workshop
Abstract:Cerebral microbleeds (CMBs) and cortical superficial siderosis (cSS) are imaging markers of cerebral small vessel disease, but their automated segmentation is limited by the scarcity of positive cases and voxel-level annotations. We propose a synthetic training framework for long-tail haemorrhagic lesion segmentation that requires no real lesion annotations for training and leverages radiological description of the lesions. Starting from anatomical brain parcellations, the framework applies spatial augmentation and voxel resampling, procedurally inserts cSS and CMB labels using clinical priors on lesion location and morphology, and synthesises images through randomised intensity assignment, blurring, and Rician noise simulation. Models were trained on dynamically generated image-label pairs and evaluated against manual delineations in 10 cSS cases and 13 CMB cases. The proposed configurations outperformed classical filter baselines. For cSS, the hypointensity constrained model achieved higher AUPRC and AUROC than the Frangi filter (AUPRC: 0.284 vs 0.083; AUROC: 0.907 vs 0.731). For CMBs, explicit synthesis of blood vessels as lesion mimics improved performance over the classical baseline (AUPRC: 0.538 vs 0.004; AUROC: 0.999 vs 0.968). These results support our proposal as a feasible strategy for data-scarce haemorrhagic lesion segmentation.
[CV-73] owards Reliable Vision-Language Models for Autonomous Driving
链接: https://arxiv.org/abs/2610.01531
作者: Manasa Mariam Mammen,Priyanka Mary Mammen,Zafer Kayatas,Stefan Wagner
类目: Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Vision-Language models (VLMs) are increasingly being explored in autonomous driving for tasks such as scene understanding, driving reasoning, decision-making, and end-to-end driving. As their role becomes more prominent, ensuring their robustness and reliability is increasingly important. In real-world conditions, visual inputs may be degraded by sensor imperfections and environmental conditions, potentially affecting both model predictions and their associated confidence. Such degradation is especially concerning in autonomous driving, where safety-critical decisions require models to make accurate predictions and recognize when their predictions may be unreliable. In this work, we evaluate five VLMs (Qwen3.5-9B, Gemma4-E4B, LLaVA-OneVision-7B, DriveFusion/DriveFusionQA-4B, and NVIDIA Alpamayo-1.5-10B) across four driving-related QA datasets with different visual input settings, including single-frame, multi-view, multi-frame, and monocular inputs. Our results show that the effects of visual corruption vary across models, datasets, and input settings, with changes in accuracy and confidence reliability and also differing across conditions. We then apply Visual Evidence Augmentation ( \mathrmV\scriptstyle \mathrmEA ), a recent inference-time method to examine whether it can improve model reliability under degraded visual conditions. We find that \mathrmV\scriptstyle \mathrmEA improves performance for some models and datasets, although the gains are not consistent across all settings.
[CV-74] SuperMotion: Source-Preserving Denoising for Text-Driven Human Motion Editing
链接: https://arxiv.org/abs/2610.01517
作者: Fa-Ting Hong,Peter Wonka
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Under review
Abstract:Text-driven human motion editing aims to realize a requested change while preserving compatible source content. Existing diffusion editors rely largely on learned conditioning for preservation of the unedited part, yet their outputs can lose temporal detail as denoising proceeds. We propose the \textbfSource-Preserving Denoising framework (SuperMotion), which explicitly reuses the source at each reverse step for source preservation. We first align the source motion to the output timeline and predict a preservation gate that controls reuse across frames and feature dimensions. A clean-space source anchor then utilizes the learned preservation gate to blend the predicted clean motion with the aligned source and passes the corrected estimate directly to the sampling posterior. Because the aligned source is a realized motion rather than a regression output, the anchor injects sample-level temporal detail that a reconstruction-trained denoiser tends to smooth away. To learn effective source reuse, we supervise the anchored estimate against the editing target and match its second temporal differences through a temporal high-frequency loss. These objectives require no explicit edit masks. Extensive experiments show that SuperMotion improves editing accuracy, reaching 33.20% full-pool R@1 on MotionFix, while reducing temporal-detail attenuation and preserving motion dynamics as it realizes the requested changes. Ablations confirm that the learned preservation gate is responsible for the gain and that it reuses the source to retain the unedited content properly.
[CV-75] VoxelSynth3D: Interpretable Volumetric Image-Domain Metal Artifact Reduction with a Paired Synthetic CLINIC-Metal Benchmark
链接: https://arxiv.org/abs/2610.01512
作者: Amritesh Banerjee,Abdul Basit,Renil Renji Joseph,Nouhaila Innan,Muhammad Shafique
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 7 pages, 7 figures. Accepted for publication at BHI 2026
Abstract:Metal artifacts in postoperative musculoskeletal CT obscure bone-implant and adjacent soft-tissue interfaces. Many metal artifact reduction (MAR) methods require unavailable raw projections or learned models that may shift across scanners and implants. We present VoxelSynth3D, a training-free 3D image-domain framework for reconstructed CT. The framework combines support masking, normalized tissue synthesis, deviation gating, and restricted edge refinement. Detected implant voxels are preserved in the output, while correction targets metal-induced artifacts in the surrounding tissue. We also construct Synthetic CLINIC-Metal, a controlled paired synthetic evaluation resource, from no-metal CTPelvic1K volumes with clean targets, metal/artifact masks, fixed seeds, and patient-level splits; 75 unpaired real metal cases receive qualitative/no-reference evaluation only. The operating point was fixed in a near-flat validation basin. With exact-mask oracle localization, all methods share a metal-excluded tissue ROI. On 40 held-out cases, VoxelSynth3D reduced RMSE from 801.48 to 786.18 HU (paired gain 15.30 HU, 95% CI 11.68-19.23), improving every case and exceeding the evaluated 3D Gaussian smoother by 13.58 HU. Clean-edge agreement decreased next to metal but exceeded input beyond 5 mm. Thus, VoxelSynth3D provides case-consistent within-distribution tissue-error reduction with a localized structural tradeoff. Spacing-aware sensitivity retained aggregate broad-region improvement and identified near-metal calibration as a target.
[CV-76] FedCKA: Representation-Guided Layer Personalization for Federated 3D Perception Across Driving Domains ICRA2027
链接: https://arxiv.org/abs/2610.01510
作者: Jolle Verhoog,Ali Burak Ünal,Holger Caesar
类目: Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)
备注: 8 pages, 3 figures. Submitted to IEEE ICRA 2027
Abstract:Robust perception in intelligent vehicles demands 3D object detectors that remain dependable under domain shifts, such as changes in time of day, location, or weather. However, due to costly annotation and rare shifts, some environments lack sufficient data to train a standalone detector. Federated learning offers a privacy-preserving framework for collaborative model training, enabling clients to benefit from shared learning across diverse environments. Yet, this framework traditionally relies on a single global consensus model, which struggles to perform across heterogeneous local data distributions. Local conditions are better captured by adapting a subset of the model, but many personalization approaches rely on predefined layer partitions or fixed personalization ratios, thereby limiting adaptation to client-specific divergence. To reduce this rigidity, we propose FedCKA, a Centered Kernel Alignment (CKA)-based strategy that dynamically handles the personalization-globalization trade-off. Specifically, FedCKA computes layer-wise feature similarities between local client models and the global consensus model during training. By converting layer-wise similarity scores into client-specific aggregation masks, FedCKA selectively shares representation-consistent layers. Evaluation on a unified multi-domain benchmark based on nuScenes shows that FedCKA outperforms established federated baselines, including FedBN, FedRep, and FedSelect, improving average NDS by 7 percentage points over the strongest baseline. The findings offer both a comparative benchmark and a promising direction for robust federated 3D perception across shifts in location, weather, and illumination. Code is available at this https URL.
[CV-77] VTR-Bench: A Systematic Benchmark for Evaluating Visual Text Rendering in Video Generation
链接: https://arxiv.org/abs/2610.01499
作者: Yu Huang,Jungang Li,Zhiyuan Wang,Yonghua Hei,Song Dai,Jiayu Yang,Deyuan Liu,Xiang Zheng,Xiaoshuang Shi,Hao Cheng,Kaidi Xu
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Recent video generation models can produce highly realistic videos from natural language instructions, with visual quality approaching cinematic standards. Existing evaluation benchmarks, however, predominantly assess visual quality, aesthetic appeal and physical plausibility, while paying limited attention to text, an essential medium for conveying information in everyday scenes. A generated video may appear visually compelling and feature lifelike subjects, yet still render the text within the scene incorrectly. To address this overlooked dimension, we introduce \textbfVTR-Bench, a systematic benchmark for evaluating the \textbfVisual \textbfText \textbfRendering capabilities of video generation models. VTR-Bench situates text within concrete application scenarios, such as advertisements and scientific videos, with 300 carefully constructed prompts spanning five scenario categories. We develop an automated evaluation pipeline with human alignments that separately assesses text fidelity through carrier-specific transcription and scene and motion requirements through a prompt-specific chain of query. Beyond evaluation, we introduce a \textbfKeyframe-Guided Agentic Framework in which a Director agent coordinates image and video generation with visual evaluation, guiding iterative refinement and candidate selection through visual feedback. Experiments on 11 state-of-the-art models reveal widespread difficulties in accurately rendering scene text, with the best-performing model recording an overall word error rate (WER) of 0.250. We further analyze text rendering failures to characterize the challenges faced by current video generation models. These findings highlight visual text rendering as a key challenge for video generation and demonstrate a practical path toward improvement. Code is available at this https URL.
[CV-78] SALD: Self-Referenced Advantage Learning for Diffusion Models
链接: https://arxiv.org/abs/2610.01496
作者: Aryan Das,Surjo Dey,Koushik Biswas,Swalpa Kumar Roy,Moloud Abdar,Arnab Bhattacharya,Vinay Kumar Verma
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Recent work on language-model adaptation has shown that single models can obtain informative training signals by evaluating their behavior in demonstrationor feedback-augmented contexts, with the help of a teacher network, which is driven by the student’s learned parameters. Inspired by this internal-reference principle, we investigate how diffusion models can identify self-referenced training signals without external demonstrations or teacher networks. We introduce SALD, a self-referenced training framework that evaluates each image-caption pair at two noise levels using the same model. The easier, lower-noise path is evaluated without gradient tracking to provide a reference, while the harder, higher-noise path provides the training gradient. Rather than directly distilling the easy-path prediction, SALD uses the difference between two path errors to adapt the hardpath objective. The proposed Advantage-Guided Diffusion (AGD) converts this relative error into a differentiable sample-level weight. Temporal Advantage Memory (TAM) accumulates relative difficulty across training and adapts the future gap between the two noise levels. Spectral Advantage Decomposition (SAD) further compares the residual power spectra of the two paths and constructs a differentiable, frequency-derived latent-element weight. All components share a single set of model parameters, requiring neither an external teacher network nor additional trainable parameters during training or inference, and no modification to the inference procedure. Experiments across multiple architectures and datasets demonstrate consistent improvements in generation quality, while component-wise ablations quantify the contributions of the proposed components.
[CV-79] FiVOS: A Fish Segmentation Algorithm Based on Interactive Video Object Segmentation and Filter Enhancement
链接: https://arxiv.org/abs/2610.01480
作者: Yuqing Duan,Song Zhang,Shili Zhao,Daoliang Li,Ran Zhao
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:With the continuous expansion of aquaculture, precise and efficient monitoring of fish behavior has become increasingly critical for improving farming efficiency and reducing economic losses. In particular, with the ongoing enhancement of computational capabilities in deep learning models, vision-based fish segmentation methods are garnering growing attention. By analyzing video segmentation results, fish behavior can be effectively tracked, thereby providing reliable data support for the precise regulation of aquaculture environments. However, existing deep learning-based video segmentation methods for aquaculture scenarios often overlook the dynamic correlations between video frames. In contrast, Interactive Video Object Segmentation (IVOS) employs an interaction-propagation scheme to achieve high-precision segmentation while minimizing user effort, thereby enhancing monitoring efficiency. Yet, IVOS applications in aquaculture remain limited due to data scarcity, and are susceptible to error accumulation and mask loss over long sequence propagation due to high intra-class similarity. In response, this paper proposes an improved interactive video object segmentation method (FiVOS) and constructs two fish-specific datasets. FiVOS utilizes a mask block filter to enable early detection and correction of erroneous propagated mask blocks, enhancing filtering accuracy through a rule-based thresholding approach. Additionally, it serializes noise filters to further eliminate erroneous mask noise, thereby improving model robustness. Experimental results demonstrate that FiVOS achieves state-of-the-art (SOTA) performance in fish video segmentation tasks, providing robust technical support for fish behavior research.
[CV-80] ALFRED: Requirement-driven development of an open-source mobile manipulator for long-term plant monitoring
链接: https://arxiv.org/abs/2610.01477
作者: Ciarán Miceal Johnson,Christopher Quail,Garry Ellard,Alistair McConnell,Steve Tonneau,Fernando Auat Cheein
类目: Robotics (cs.RO); Hardware Architecture (cs.AR); Computer Vision and Pattern Recognition (cs.CV)
备注: 36 pages, 19 figures
Abstract:Tracking seasonal change in crops and forests requires observing the same plants repeatedly. Ground robots can do this at close range, and a manipulator gives their sensors more viewpoints. Yet the robots behind long-term field datasets are rarely released with their design files, and how a robot’s own structure limits arm reach and occludes its sensors is seldom compared between builds. We present ALFRED, an open-source mobile manipulator built from commercially available components. It carries a six-degree-of-freedom arm, LiDAR, RGB-D cameras, RTK GNSS and an IMU on an Ackermann-steered base, all mounted on a reconfigurable aluminium strut frame, and runs containerised ROS software. It was developed through four builds against six requirements for repeated outdoor deployment: durability, modularity, repairability, sensing reach, endurance and reproducibility. Model-based analysis of the last three builds shows the usable share of the arm’s reachable poses rising from 34.0% to 60.0% and then 66.1%, and ray casting shows that only the final build keeps the frame-mounted LiDAR’s horizontal view clear both forwards and backwards. ALFRED completed a year of monthly forest surveys (528 traversals) without missing a scheduled collection. This was despite battery degradation, reconfiguration for another researcher’s study, and the parallel development of ALFRED 2.0 for autonomous crop-row operation, with each switch between builds taking about six hours. The deployment also showed that mechanical modularity is only as dependable as the robot description that tracks it.
[CV-81] Uncertainty-Guided Handshake: Efficient Human-in-the-Loop Refinement for Surgical-Grade Glioma Segmentation
链接: https://arxiv.org/abs/2610.01452
作者: Samuel Hart,Ahmad Yahya,Ahmed Karam Eldaly
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 12 pages
Abstract:While state-of-the-art automated models for medical image segmentation achieve high mean performance, they frequently suffer from localized, catastrophic failures that preclude safe clinical deployment, particularly in neuro-oncology. Interactive segmentation frameworks mitigate this by incorporating human oversight, but traditionally impose prohibitive cognitive and temporal workloads by requiring clinicians to manually search for errors. In this project, we present an efficient, Hybrid Structural-Aleatoric Human-in-the-Loop framework for glioma segmentation that bridges the gap between automated baseline performance and surgical-grade precision, achieving sub-2.0 mm HD95 on curated benchmarks while providing safety-net routing for structural failures across real-world clinical data. By extracting voxel-wise Test-Time Augmentation (TTA) uncertainty and applying hierarchical topological filtering, our method proactively isolates high-risk structural anomalies. We comprehensively evaluated our approach on a challenging out-of-distribution clinical stress-test cohort (N = 362). Operating under a simulated Human Oracle, the framework improved the Whole Tumor (WT) Dice score from 0.891 to 0.914 and reduced the 95th percentile Hausdorff Distance (HD95) from 5.82 mm to 4.76 mm. Critically for surgical safety, the system rescued severe boundary failures in the Tumor Core, reducing mean HD95 from 17.96 mm to 14.83 mm (improving absolute TC Dice to 0.356). These spatial rescues were achieved while demanding a median interactive workload of just 11.3% of the target volume. Acknowledging this as a simulated upper bound lacking real-world cognitive friction, the framework nevertheless demonstrates a highly Pareto-efficient pathway for safely deploying clinical AI.
[CV-82] he Impact of Processing Parameters on High-Accuracy Measurements in UAV Photogrammetry
链接: https://arxiv.org/abs/2610.01438
作者: Paweł Ćwiąkała,Edyta Puniach,Elżbieta Pastucha,Wojciech Gruszczyński
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Unmanned aerial vehicle (UAV) photogrammetry is increasingly used in applications requiring high accuracy, such as determining ground surface changes caused by landslides, mining, or microrelief transformation. While acquisition strategies have been widely studied, the influence of the processing workflow-particularly Bundle Block Adjustment parameter settings-remains insufficiently explored. This study addresses this gap through a systematic, full-factorial evaluation of 768 processing variants applied to ten UAV datasets collected over 1.5 years in a 220 ha study area. Eight key parameters were analysed. The results show substantial variability in final 3D accuracy: the best performing variant achieved a root mean square error (RMSE) of 16 mm, whereas the weakest reached 303 mm. The most influential factors were the number of ground control points, the application of additional camera calibration corrections, and the use of the Post-Processing Kinematic GNSS method for determining camera projection center coordinates. The study also evaluates how workflow optimization affects the accuracy of displacement, tilt changes, and horizontal strain determination. While random displacement errors remained stable (RMSE of ~6-7 mm), systematic errors were significantly reduced by over half in all axes, with vertical median absolute error decreasing from 14 mm to 7 mm in the optimized configuration compared to the baseline previously used by the authors. This study provides the first large-scale, practice-oriented assessment of how processing parameter selection shapes the accuracy of both photogrammetric products and deformation indices determination. The results offer actionable guidance for developing more robust and repeatable UAV photogrammetry workflows tailored to high-precision monitoring.
[CV-83] Localisation-Aware Uncertainty for Pretrained Object Detection
链接: https://arxiv.org/abs/2610.01409
作者: Charmaine Barker,Daniel Bethell,Simos Gerasimou
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Reliable uncertainty estimation is essential for deploying object detectors when distribution/covariate shift and adversarial attacks may occur. Existing approaches often require detector retraining, architectural modification, or repeated inference, which may be infeasible or incur significant overheads. We introduce a lightweight post-hoc evidential meta-model that learns when object localisations should be considered uncertain while keeping the base detector frozen. Our approach automatically identifies localisation-relevant features and uses saliency-guided modification to construct an increasingly challenging curriculum. Detection-level targets combine localisation error, modification level, and prediction instability to guide an evidential meta-model to estimate uncertainty for each predicted bounding box. Our approach requires no changes to the detector and preserves its original localisation outputs. Across adversarial attacks and evaluated strengths, GRACE improves TP-FP AUROC by 22% relative to the strongest comparator in some cases while maintaining in-distribution detection performance.
[CV-84] Smoother Flow Matching via Contrastive Trajectory Repulsion
链接: https://arxiv.org/abs/2610.01408
作者: Ziqi Jiang,Zhenqi He,Long Chen
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注: 18 pages, 5 figures
Abstract:Trajectory crossing remains a critical bottleneck in Flow Matching (FM), and previous works typically view these crossings from a theoretical optimization perspective causing velocity averaging. They attempt to address it indirectly by post-hoc distillation or endpoint coupling, without explicitly regulating the intermediate trajectories. In this paper, we introduce a new network learning perspective: crossing points inherently induce large local Lipschitz constants in the target velocity field, leading to two drawbacks. First, high Lipschitz constants correspond to high-frequency signals in the velocity field that neural networks struggle to fit due to spectral bias. Second, they also imply drastic velocity variations, leading to severe numerical integration errors in few-step inference. To alleviate this, we propose CoFlow, a framework that introduces the contrastive learning paradigm into FM to explicitly repel trajectories during training, thereby lowering the local Lipschitz constants of the velocity field. Specifically, we formulate CoFlow from a Stochastic Differential Equation (SDE) perspective by injecting a repulsive drift term. This drift actively guides the forward process of positive samples away from negative trajectories, effectively reducing the local Lipschitz constant. Furthermore, we derive an equivalent stochastic interpolant formulation from this SDE, providing a simple and tractable design space to control the influence of negative samples. Extensive experiments on ImageNet 256x256 demonstrate that CoFlow significantly reduces FID compared to standard FM in few-step inference (e.g., 20 steps), with no added training overhead. The code can be accessed at: this https URL
[CV-85] AiSearch: Interactive Multi-Modal Search with VLMs ECCV2026
链接: https://arxiv.org/abs/2610.01389
作者: Ali Koksal,Mei Chee Leong,Vicky Sintunata,Ching Ling Chin,Wee Teck Fong
类目: Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
备注: The demo paper with 1 page main paper, 7 pages supplementary material accepted and presented in ECCV 2026
Abstract:Modern retrieval systems must both be automated and interactive, allowing users to search and refine results in real time. We present AiSearch, a flexible multimodal retrieval framework that leverages the zero shot capabilities of Vision Language Models (VLMs) for natural language search over images and videos. AiSearch supports interactive search refinement through user feedback to tailor results to the user’s intent, and allows visual benchmarking across multiple VLMs, enabling users to select the most suitable model for their task.
[CV-86] Supervising Sound Localization by In-the-wild Egomotion CVPR2025
链接: https://arxiv.org/abs/2610.01388
作者: Anna Min,Ziyang Chen,Hang Zhao,Andrew Owens
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Multimedia (cs.MM); Sound (cs.SD)
备注: CVPR 2025 Highlight (IEEE/CVF Conference on Computer Vision and Pattern Recognition)
Abstract:We present a method for learning binaural sound localization using egomotion as a supervisory signal. Over the course of a video, the cameras direction to a sound source will change as the camera moves. We train an audio model to predict sound directions that are consistent with visual estimates of camera motion, which we obtain using traditional methods from multi-view geometry. This provides a weak but plentiful form of supervision that we combine with traditional binaural cues. To evaluate this method, we propose a dataset of real-world audio-visual videos with egomotion. We show that our model can successfully learn from real-world data and that it performs well on sound localization tasks
[CV-87] Is it Possible to Generate Irreversible PolyProtected Templates from Face Embeddings using System-Specific Keys?
链接: https://arxiv.org/abs/2610.01385
作者: Vedrana Krivokuća Hahn,Jérémy Maceiras,Sébastien Marcel
类目: Cryptography and Security (cs.CR); Computer Vision and Pattern Recognition (cs.CV)
备注: Submitted to TIFS journal on 12 May 2026 (under review). Consists of: 13 pages, 9 figures, 3 tables
Abstract:This work aims to answer the question of whether it is possible to generate irreversible protected templates when the PolyProtect biometric template protection method is applied to face embeddings using system-specific keys (i.e., the same C and E parameters, which define the transform, are applied to all subjects’ face embeddings), instead of the traditional subject-specific keys (i.e., each subject has their own C and E parameters). This is important for determining whether we can perform de-duplication of face identities in the PolyProtected domain, which is not possible in the subject-specific key scenario due to the clash with PolyProtect’s unlinkability property (i.e., one could generate multiple protected templates belonging to the same identity, using different C and E parameters, such that those templates cannot be linked to each other). We present experiments (reproducible using our open-source code) to prove that there exist at least three ways of systematically selecting system-specific keys that produce irreversible PolyProtected templates: (i) from pre-selected subject-specific keys, (ii) by applying a previously proposed key selection algorithm to random vectors, and (iii) by approximating a “good” C/E pair distribution from which system-specific keys can be constructed. Our findings thus point to the conclusion that it is, indeed, possible to safely operate PolyProtect in the system-specific key scenario without degrading the template protection potential. This opens up the possibility for identity de-duplication in the PolyProtected domain.
[CV-88] MMVistaReason : Toward Open-Data and Post-Training Recipes for Multimodal Reasoning
链接: https://arxiv.org/abs/2610.01352
作者: Juekai Lin,Honglin Lin,Yuqian Yuan,Xiaolong Wu,Jie Cao,Liang Liang,Yunqi Cao,Yun Zhu,Wenqiao Zhang,Lijun Wu
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Open multimodal reasoning models have benefited from large-scale reasoning supervision, yet reliable post-training remains challenging due to uneven data quality, inefficient supervision construction, imbalanced difficulty, and cross-domain interference. We introduce MMVistaReason (MVR), an open-data post-training recipe with three components: (1) broader capability coverage across complementary Analytical and Real-World reasoning groups, emphasizing structured reasoning versus visual perception and spatial grounding; (2) efficient SFT and RL data construction, standardizing heterogeneous open data through staged cleaning and annotation, combining difficulty-aware cascaded teacher distillation with answer-likelihood-based trajectory selection to construct MVR-SFT-528K, and applying scale-specific frontier filtering for MVR-RL-63K; and (3) specialize-then-integrate training, which trains complementary RL experts and consolidates their capabilities through multi-teacher on-policy distillation (MOPD). Our analyses reveal a capacity-dependent interaction between supervision difficulty, trajectory quality, and model capacity: smaller students benefit more from selected supervision, while larger students are robust to trajectory variation and mixed-domain interference. Mixed-domain RL introduces benchmark-level negative transfer, whereas MOPD provides consistent capability integration, with the preferred KL direction varying across model scales. Across 15 multimodal benchmarks, MVR-4B achieves an average score of 72.8, outperforming Qwen3.5-9B (Instruct) and MMFineReason-8B while using about 70% fewer samples than MMFineReason. Scaling to 9B improves the average to 74.4, surpassing Qwen3.5-35B-A3B (Instruct). Overall, MMVistaReason demonstrates that systematic open-data construction and capacity-aware post-training provide a practical and scalable path toward reliable multimodal reasoning.
[CV-89] CLASP: Continual Low-rank Adapters for Spatially Placed Concepts from One Hypernetwork
链接: https://arxiv.org/abs/2610.01331
作者: Wojciech Gromski,Patryk Krukowski,Jan Miksa,Maciej Zieba,Przemysław Spurek
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 31 pages. Code: this https URL , project page: this https URL
Abstract:Continual personalization of text-to-image diffusion models requires sequentially acquiring new concepts while retaining previously learned ones. However, existing methods either suffer from catastrophic forgetting or rely on storing additional concept-specific parameters and spatial components, causing their parameter footprint to grow with the concept stream. This limits their ability to scale to long sequences of personalization tasks. We propose a rehearsal-free approach that uses a single fixed-size hypernetwork to continually personalize a frozen diffusion model. Instead of expanding the model as new concepts are acquired, the hypernetwork dynamically produces the concept-specific adaptations required for personalization while preserving previously learned concepts. Our framework further integrates spatial control into the personalization process, allowing users to specify where a personalized concept should appear without introducing additional per-concept components. This formulation enables continual personalization with a parameter footprint that remains independent of the number of learned concepts, aside from compact concept representations. Experiments demonstrate strong retention of previously learned concepts and reliable spatial grounding, matching or improving upon existing methods while scaling effectively to long streams of personalization tasks.
[CV-90] ARROW: Arbitrary Reconstruction and Tracking of 4D Observations in the Wild WWW
链接: https://arxiv.org/abs/2610.01314
作者: Ilya Fradlin,Christian Schmidt,Jens Piekenbrinck,Karim Knaebel,Gonzalo Martin Garcia,Bastian Leibe
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Project page at: this https URL
Abstract:Dynamic scenes may be captured by a moving camera, multiple video streams, or images taken at different times. These observations reveal complementary aspects of scene geometry and motion, yet bringing them together requires establishing correspondence across viewpoints, capture times, and visibility changes. We introduce ARROW, a feed-forward model that unifies 3D reconstruction and 3D point tracking from arbitrary image sets. At its core is a novel order-invariant querying approach, which allows the association of queries with observations across arbitrary inputs. We show that exposing the model to more diverse sets of inputs during training results in improved task performance. Moreover, the resulting model is capable of generalization to a wider range of tasks including multi-view tracking. Trained with this strategy, ARROW establishes a new state of the art in 3D tracking on WorldTrack and TAPVid-3D and outperforms dedicated multi-view trackers on an adapted RGB-only MVTracker benchmark, while remaining competitive across 3D reconstruction tasks. Code and weights are publicly available.
[CV-91] STAGE: Subspace-Targeted Affine Generative Erasure for Text-to-3D Models
链接: https://arxiv.org/abs/2610.01302
作者: Karol Dziekan,Przemysław Spurek,Dawid Malarz
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Concept erasure suppresses a target concept while preserving behavior on unrelated inputs. Existing closed-form methods were designed for 2D image diffusion and assume a single generative pathway, so one edit must cover geometry and texture at once. Native 3D generators, which synthesize structured 3D representations directly rather than by lifting 2D samples, violate this assumption. We show that shape and object concepts must be erased in the structural stage of the pipeline and material concepts in the appearance stage. We therefore formulate erasure in native text-to-3D as a stage-aware editing problem and introduce STAGE, a training-free, closed-form framework. STAGE confines each edit to the low-dimensional subspace spanned by the differences between erase and anchor embeddings, and relaxes the norm-preserving (orthogonal) constraint of prior editors into a least-squares affine correction that maps target activations onto safe anchors subject to a penalty on the displacement of retained prompts. The correction applies to the structural stage, the appearance stage, or both. We find that the stage an edit must reach is determined by concept type. On TRELLIS, the standard open native 3D generator, across 15 shape, material, and object concepts, STAGE reaches 66.7 on a composite score that balances forgetting the target concept against preserving everything else, aggregating CLIP-based semantic and physical metrics, versus 53.2 for the strongest adapted baseline. Code: this https URL Project Page this https URL Subjects: Computer Vision and Pattern Recognition (cs.CV) Cite as: arXiv:2610.01302 [cs.CV] (or arXiv:2610.01302v1 [cs.CV] for this version) https://doi.org/10.48550/arXiv.2610.01302 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[CV-92] ODDR: One-Step Deshadow Diffusion via Reward Guidance
链接: https://arxiv.org/abs/2610.01291
作者: Junseong Shin,Kijun Kim,Minseong Kim,Dongjin Kim,Tae Hyun Kim
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Recent advances in deep learning for shadow removal have significantly enhanced image quality and realism. However, most approaches rely on real-world paired datasets, which are costly to collect and often limited in scene diversity, leading to limited generalization. To address these limitations, we propose One-step Deshadow Diffusion via Reward guidance (ODDR), a new framework that achieves efficient and high-fidelity shadow removal without relying on real-world paired supervision. Our method begins with One-step Deshadow Diffusion (ODD), a baseline model trained on synthetic shadow data for efficient one-step shadow-free reconstruction. We further adapt ODD into ODDR using ShadowReward. In contrast to traditional, annotation-heavy approaches, ShadowReward is the first reward model for shadow removal trained entirely without human annotation. It learns to mimic human perceptual judgments by ranking synthetically generated images with controlled degradations, such as texture distortion and boundary artifacts. This reward-guided fine-tuning enables ODDR to close the synthetic-to-real domain gap. Extensive experiments show that ODD achieves strong performance without relying on real-world paired supervision, and ODDR further improves the results, narrowing the gap to fully supervised methods trained on real-world paired data while maintaining higher computational efficiency as a single-step model.
[CV-93] Dyna3: VLM-Guided Training-Free 4D Reconstruction via Depth Foundation Models
链接: https://arxiv.org/abs/2610.01286
作者: Xinhao Xiang,Weiyang Li,Zhijie Zheng,Abhijeet Rastogi,Jiawei Zhang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Recent depth foundation models like Depth Anything 3 (DA3) achieve remarkable multi-view depth estimation but assume static 3D scenes, limiting their applicability to real-world dynamic environments. Existing training-free 4D methods like Easi3R and VGGT4D rely on correspondence-trained backbones whose attention encodes cross-frame matching, a property absent in depth-only models like DA3. We present Dyna3, a training-free framework that extends DA3 for 4D dynamic scene reconstruction without any fine-tuning. Our key insight is that DA3’s cross-view features, though trained only for depth consistency, implicitly encode motion-discriminative signals when combined with best-match feature search across frames. Its static surfaces find consistent matches globally, while dynamic objects cannot. We further adopt vision-language models (VLM) to automatically generate scene-specific semantic prompts for SAM 3, enabling precise instance-level segmentation that distinguishes which objects move from what objects exist. For reconstruction, we decouple the scene into a cross-frame aligned static background and per-frame dynamic point clouds. Experiments on four datasets demonstrate that Dyna3 surpasses correspondence-trained methods with +5.5pp J-Mean over state-of-the-art VGGT4D on dynamic object segmentation, while achieving up to 13x faster pose estimation and 3x faster 4D reconstruction with 4 to 8x lower memory. Dyna3 could therefore enable much denser temporal sampling that prior methods cannot support.
[CV-94] ShelfChange3D: Object-Level 3D Change Detection for Retail Shelf Monitoring
链接: https://arxiv.org/abs/2610.01283
作者: Lingyi Zhou,Yunke Wang,Mengyu Zheng,Wenbo Wang,Zijian Wang,Chang Xu
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Our code will be available on our project website at this https URL
Abstract:Reliable shelf monitoring is an important capability for retail automation, yet existing out-of-stock detection methods mainly operate in image space and lack metric 3D localization for downstream robotic systems. We formulate shelf monitoring as object-level 3D change detection: given two RGB-D observations captured at different times, the goal is to identify changed products and localize each change with a 3D bounding box. To support this task, we introduce ShelfChange3D, comprising 145K synthetic and 5K real-world paired RGB-D observations with object-level 3D change annotations. We further propose ChangeBox, an end-to-end framework that jointly reasons over paired observations and predicts object-level 3D change boxes. To improve localization accuracy, we introduce a geometry-based refinement stage that exploits depth and gravity prior to estimate relative pose and refine predicted boxes. Experiments show that ChangeBox outperforms existing change detection baselines, with further gains from refinement and effective transfer from synthetic to real-world observations.
[CV-95] PickMoment: Continuous-Time Single-Image-to-Video via Learning Deblurring and Blur-to-Video
链接: https://arxiv.org/abs/2610.01279
作者: Junseong Shin,Hyeonsu Jo,Daehyun Kim,Tae Hyun Kim
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:
Abstract:Motion blur arises from the temporal integration of a continuous sharp signal over a finite exposure window, yet existing learning-based methods sidestep this physical model and predict only the sharp signal itself: most single-image deblurring methods recover a single frame at the exposure center, while blur-to-video methods predict a fixed set of frames. We introduce PickMoment, a continuous-time reformulation that directly learns the interval-mean blur over arbitrary sub-intervals of the exposure with a single deterministic model. Drawing an analogy to MeanFlow’s average-velocity formulation, we train the model with three supervisions derived from the blur integral: an empirical reconstruction loss from available subframes, an additivity loss that enforces self-consistency across overlapping sub-intervals, and a sharp-frame loss anchored at the zero-interval limit. A single trained model unifies single-image deblurring, blur-to-video generation, and continuous-time pick-a-moment recovery as different queries to the same network, with no separate training for each task. Our PickMoment achieves state-of-the-art performance among generative-based deblurring methods on GoPro and HIDE while competitive against restoration-based methods on RealBlur, and the highest per-frame fidelity on GoPro-7 blur-to-video, all in a single forward pass without iterative sampling.
[CV-96] When the Judge Acts: Auditing VLM-Guided Image Selection on Culturally Situated Prompts
链接: https://arxiv.org/abs/2610.01243
作者: Huichan Seo
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computers and Society (cs.CY)
备注: 25 pages including appendix. Code and project page: this https URL ; data: this https URL
Abstract:Vision-language models (VLMs) increasingly act as judges that pick the best of several generated images, so their choices decide what users see. Such judges are usually validated by score agreement with human ratings, not by the images they return. We audit VLM judges as decision-makers: on 300 culturally situated prompts, we compare the returned image with human ratings the judge never sees and with random choice from the same candidates, and repeat every decision with the candidates reordered. A 4B-parameter judge barely beats random and falls short of a CLIP similarity baseline. It picks the first image shown in 49% of calls (chance: 28%), and reordering changes its choice on 60% of prompts. For this judge, agreement across orders is informative: decisions that survive reordering are much better than random, whereas agreement with a weaker second judge keeps the wrong ones. An 8B judge shows almost no position bias and outperforms CLIP, yet for it the same filter mostly discards good decisions. Agreement helps only when it targets the judge’s failure mode, so filters must be re-audited whenever the judge changes. The 4B judge’s slight rise in stereotype ratings is no longer detectable after aggregating across orders or with the larger judge.
[CV-97] Flow Matching Reinforcement for 3D Mesh Generation via Dynamic Homing Optimization
链接: https://arxiv.org/abs/2610.01233
作者: Zhen Zhou,Zhiwei Ning,Puhua Jiang,Sheng Zhang,Yifei Tang,Jie Yang,Xintong Han,Wei Liu,Chunchao Guo
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Flow matching is central to 3D generation, yet in practice its reinforcement learning (RL) methods are largely adapted from 2D visual generation. Representative DPO-, GRPO-, and NFT-style objectives, when applied to negative trajectories, mainly steer predicted velocities away from the corresponding directions without explicitly specifying a target velocity field toward preferred samples. In 3D generation, constrained by pretrained model capabilities, rollout diversity, and reward-distribution complexity, directly applying these RL methods yields limited gains in geometric quality. We introduce a forward-process RL method \textbfDynamic Homing Optimization (DHO), which reformulates negative-trajectory optimization as positive-sample attraction-guided dynamic homing. Specifically, Minimum-Cost Attractive Matching (MAM) assigns each negative sample a distinct positive target, and Time-Aware Dynamic Correction (TDC) then redirects its trajectory toward the target using a remaining-time-aware corrective velocity. Building on asynchronous online DHO, we develop \textbfFlow3D-Pro, an image-to-3D geometry generation framework. Experiments show that DHO outperforms representative DPO-, GRPO-, and NFT-style objectives in 3D generation, while Flow3D-Pro produces higher-quality 3D geometry than existing mesh generation methods.
[CV-98] A Compact Explicit 4D Representation for Dynamic Scenes
链接: https://arxiv.org/abs/2610.01229
作者: Di Yang,Zhihao Li,Yanhai Xiong,Yufei Wang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:A compact dynamic-scene representation must retain both the surfaces seen over time and the appearance needed to render them from new viewpoints. We present Sparc4D, a feed-forward autoencoder that encodes a monocular video with known cameras into a sparse 4D scene state. Static features are shared across the clip, while spatially anchored temporal slots compress time-varying features. A sparse decoder produces 2D Gaussian surfels, while stored source pixels preserve fine texture through geometric re-projection. The state includes one full source frame and dynamic-region pixels sampled every fourth frame, alongside learned features and sparse occupancy. For a 32-frame MultiCamVideo clip, it averages 0.95M 32-bit-equivalent values on random windows and 0.92M on the first-32 protocol. On first-32, Sparc4D reaches 21.70,dB, compared with 20.40,dB for MoVieS. On randomly placed windows, their PSNR scores are comparable. With stored texture disabled, temporal slots compress the time-varying feature state by a median 4.0\times and reduce the mean state from 1.04M to 0.42M values, with essentially unchanged target-view reconstruction quality. Without fine-tuning on real data, Sparc4D transfers to DyCheck and Neu3D, where stored texture improves LPIPS while slightly reducing PSNR.
[CV-99] AutoGUIWorld: Image Generators as Visual World Models for GUI Agent
链接: https://arxiv.org/abs/2610.01215
作者: Cheng Yang,Yifan Wu,Yutao Huang,Zhaohua Zhang,Beiduo Chen,Muxi Chen,Chenchen Zhao,Hexuan Deng,Haolin Yang,Geyuan Zhu,Sa Zhu,Jianhuan Zhuo,Qiuyong Xiao,Jianhao Ruan,Yiran Peng,Jiayi Zhang,Tian Ye,Xinlei Yu,Tianwen Jiang,Jihong Zhang,Yuyu Luo
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:GUI agents require high-quality interaction trajectories to learn how software environments respond to actions, maintain state, and support multi-step workflows. However, the diversity of available trajectories is constrained by the applications, interface states, and workflows accessible in the underlying environments. Expanding this coverage requires deploying increasingly diverse and complex software, with specialized applications imposing additional installation, configuration, and runtime costs. We introduce AutoGUIWorld, a data generation framework that combines the visual priors of image generators with the task knowledge of a planner to synthesize GUI interaction trajectories without deploying or running the corresponding software environments. AutoGUIWorld samples initial GUI scenes from structured specifications of operating-system context, visual appearance, and interface state, and generates tasks conditioned on those scenes. A planner then specifies atomic actions and their intended visual consequences, while an image generator iteratively edits the current screenshot to produce subsequent observations. Action grounding and transition-level quality filtering yield 79,266 spatially annotated step-level training samples across Ubuntu, Windows, macOS, and Chrome. Fine-tuning Qwen3.5-35B-A3B on AutoGUIWorld trajectories improves the mean task score on OSWorld from 33.0% to 40.8% and the task success rate on ScienceBoard from 14.0% to 32.2%. These results show that generated trajectories improve GUI-agent performance on real desktop and scientific tasks.
[CV-100] EgoFound3R: End-to-End Egocentric Hand Reconstruction in World Space with Point-Wise Interaction Attributes
链接: https://arxiv.org/abs/2610.01210
作者: Hongming Fu,Jingcheng Shi,Wenjia Wang,Binhua Zuo,Bo Zhao
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Egocentric video has become a primary source of supervision for embodied models, and its value rests on recovering hand motion in world coordinates, which camera motion and hand occlusion make difficult. Existing reconstruction pipelines typically separate hand and scene estimation, leave interaction attributes to separate task-specific models, and invoke several models per video, so no prior reconstruction model estimates these attributes and throughput becomes a practical constraint on large-scale annotation. We therefore introduce EgoFound3R, a unified end-to-end model that estimates world-space hand geometry in a metric scale shared with the scene, and predicts point-wise interaction attributes, including visibility, contact, and distance. The model integrates three designs: (i) structured hand prompts that transfer pretrained geometric priors to world-space hand reconstruction; (ii) an explicit hand representation that decodes hand geometry and interaction attributes; and (iii) a shared-parameter multi-rate design that lowers inference cost. Together, these designs predict hand geometry and point-wise attributes in one pass. On OakInk-v2, TACO, and HOI4D, EgoFound3R reduces the mean per-joint position error (MPJPE) by 43.2%, 22.4%, and 11.6% over previous methods and predicts point-wise contact and distance alongside the geometry in the same pass, while attaining approximately 6x higher throughput.
[CV-101] Resolving Mixed Single-Photon LiDAR Returns for Foreground-View and Hidden Scene Reconstruction
链接: https://arxiv.org/abs/2610.01206
作者: Ziting Wen,Runrong Deng,Zili Zhang,Haitao Zheng,Yuecong Xu,Xiaoqiang Ren,Guodong Shi,Kemi Ding
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Partially transmissive screens and protective covers are common in robotic inspection, but they create mixed LiDAR returns from both the foreground material and the scene behind it. Conventional peak-based LiDAR usually discards weak hidden returns, while single-photon LiDAR records time-resolved histograms that preserve attenuated and overlapping echoes. However, existing transient reconstruction methods typically fit a single scene representation to the measured waveform. Under occlusion, weak or nearby foreground–hidden echoes can form a broad peak or subtle shoulder. Because such waveforms can also be explained by a displaced single surface or a thick density distribution, accurate transient fitting does not necessarily imply correct geometry. We propose a state-aware framework for foreground-view and hidden scene reconstruction from occluded single-photon histograms. For each ray, we estimate local echo evidence, identifying no reliable surface evidence, single-return evidence, or two returns. The inferred echo state routes supervision for a two-head neural field: all rays constrain waveform reconstruction, while reliable anchors provide geometry localization. We also introduce a real paired single-photon LiDAR occlusion dataset with occluded and clean captures at fixed poses. Experiments on a real dataset show improved hidden scene depth and point-cloud accuracy over baselines. Our results demonstrate single-photon layered reconstruction as a practical route for 3D perception through partially transmissive occluders.
[CV-102] Semantic RGB–Depth Based Surgical Skill Assessment in Microscopic Stereo Videos
链接: https://arxiv.org/abs/2610.01205
作者: Jecia Z. Y. Mao,Sue M. Cho,Francis X. Creighton,Deepa Galaiya,Russell H. Taylor,Manish Sahu
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Objective assessment of microsurgical technical skill is essential for competency-based training and quality assurance, yet existing video-based approaches predominantly rely on RGB images and therefore overlook the 3D spatial relationships that characterize instrument-anatomy interactions. Although stereo operating microscopes provide complementary depth information, conventional stereo matching algorithms can produce sparse and unreliable depth estimates under high-magnification imaging conditions, limiting their use for automated skill assessment. This work presents a semantic RGB-Depth framework for surgical skill assessment from microscopic stereo videos. A regression-based depth fusion method combines sparse metric stereo depth with dense monocular depth estimates to generate a dense geometric representation of the surgical scene. This representation is integrated with semantically decomposed RGB streams corresponding to individual surgical instruments and surrounding anatomy. A hierarchical attention architecture jointly encodes these streams to capture discriminative patterns of instrument use and instrument-anatomy interaction across surgeons at different training levels. The framework was evaluated on 33 ex vivo transoral microlaryngeal procedures performed by six surgeons, comprising attending surgeons and surgical residents, using leave-one-surgeon-out cross-validation. The proposed semantic RGB-Depth model achieved an F1 score of 0.938 for skill-level classification, compared with 0.696 for semantic RGB and 0.929 for semantic depth. These results suggest that geometric information can improve automated surgical skill assessment from microscopic stereo videos. The learned spatial, temporal, and semantic attention patterns also support qualitative examination of the scene regions, video segments, and semantic streams emphasized by the model.
[CV-103] SEE: Object Permanence Through Self-Supervision
链接: https://arxiv.org/abs/2610.01201
作者: Pramish Paudel,Ajad Chhatkuli,Luc Van Gool,Danda Pani Paudel
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Object permanence, keeping track of an object’s identity and position while it is occluded, is central to video representations that track, predict and plan. Trackers that achieve it learn from boxes, track identities and visibility labels. On the other hand, self-supervised object-centric methods discover objects without labels: through slot attention, it represents a video as slots that bind to objects and follow them across frames. However, these slots are lost under occlusion, making the desired permanence impossible. Reasoning permanence is a hard problem because it requires to detect when an object becomes occluded, re-identify when object reappears, and keep the object’s hidden position continuous, using reapperance as the only learning cue. To address this, we propose iSEE, a novel framework that offers all three aforementioned requirements, without any labels whatsoever. We built iSEE using the following three proposed components: (i) Object evidence modelling: a slot’s attention, compared with its own past, reveals when its object is hidden. (ii) Appearance-position separation: two slot streams let the appearance be held for re-identification while the position keeps changing. (iii) Permanence from reappearance: a walker follows the hidden object’s position, trained only on where the object reappears. On LA-CATER static, iSEE returns a reappearing object to its own slot after 86% of occlusions, against 32% for SlotContrast, and localises it while hidden within 4.1 mAP of the label-trained SoTA RAM. The two streams also allow downstream planning, with the position stream as the action of a world model. Project page: this https URL
[CV-104] FlashBack: Knowing When to Remember in Streaming Vision-Language Models
链接: https://arxiv.org/abs/2610.01192
作者: Yi Chen,MingMing Yu,Rui-Qi Wang,Boran Wang,Xiaohang Cao,Chu Tang,Jingmin Chen,Jie Gu
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Streaming vision-language models must process continuously growing video streams under a bounded compute budget, creating a persistent tension between real-time perception and long-term memory. Retrieving historical information provides a natural remedy, yet historical recall is not uniformly beneficial: unnecessary history may introduce irrelevant context into current reasoning and interfere with native real-time perception. Effective streaming memory should therefore address not only what to remember, but also when and how to access it. To this end, we introduce FlashBack, a training-free framework for selective, multi-level memory in streaming vision-language models. Before retrieving history, FlashBack draws on the semantic understanding of the frozen streaming VLM to infer whether a query calls for historical evidence. This assessment determines whether inference remains on the Native trajectory or invokes an isolated Recall trajectory. The Recall trajectory combines recent context with retrieved long-term memory through a query-local Side-KV pathway, preserving local temporal continuity without modifying the persistent Native state. We instantiate FlashBack on StreamingVLM and Mage-VL-4B and evaluate it on OVO-Bench and StreamingBench. The results show improvements on several long-horizon and memory-dependent tasks while largely preserving real-time perception, with performance competitive with strong training-based streaming methods despite requiring no additional training. Our code will be announced later.
[CV-105] Color Independent Word Segmentation From Transcribed Bangla Passages
链接: https://arxiv.org/abs/2610.01191
作者: Faias Satter,Noor Masrur,Sk. Md. Masudul Ahsan
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 6 pages, 8 figures, 6 tables. Accepted version of the paper published in the 2023 6th International Conference on Electrical Information and Communication Technology (EICT)
Abstract:An optical character recognition(OCR) system can scan paper and extract text, making people’s jobs easier. While numerous OCR systems are accessible in the software sector, finding a dependable equivalent solution for Bangla is tough. When it comes to handwritten texts, the case is even more rare. The first fundamental step to any OCR is to segment words from text images. If this stage fails, the total OCR’s performance will be poor no matter how promising the later stages perform. This research aims to segment words in a handwritten Bangla text image. This research can be implemented on any smartphone-captured image, irrespective of the color and type of paper and ink. Furthermore, as smartphone-captured images can create shadow interferences, the custom dataset built for this research is created in such a way that every possible obstacle that can be faced is included. For 7374 words, a total of 7278 bounding boxes are generated, which have recall of 90.60 %, precision of 91.80 %, and F1-score of 91.20 %. The system can be further improved with nested operations on bounding boxes containing several words or by adjusting the adaptive thresholding and dilation filter sizes to a more precise level.
[CV-106] Skeleton-and-Strategy Prompting: Training-Free Negation Understanding for Vision-Language Models
链接: https://arxiv.org/abs/2610.01180
作者: Yuliang Cai,Mohammad Rostami,Jesse Thomason
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Despite the strong performance of Vision-Language Models (VLMs) on a wide range of visual question answering (VQA) tasks, these models consistently struggle to understand negation and produce incorrect answers when questions involve negated clauses. To address this limitation, we propose Skeleton-and-Strategy Prompting (\textbfSSP), a training-free, in-context learning method that improves VLM negation understanding capabilities without any parameter updates. Given a negation question, our method first abstracts the underlying question structure into a skeleton, retrieves a small set of same-skeleton questions from a lightweight question pool, then prompts the VLM to analyze their shared negation pattern and synthesize a single-sentence answering strategy. The skeleton and strategy are prepended to the test sample to guide the model correctly tackle the negation problems. Experiments on multiple negation VQA benchmarks show that SSP achieves state-of-the-art performance on negation-focused VQA tasks while remaining computationally efficient.
[CV-107] CineMR: Tool-Integrated Vision-Language Reasoning for Quantitative Cardiac MRI Assessment
链接: https://arxiv.org/abs/2610.01166
作者: Kunyang Li,Hai Nguyen,Joshua Lowe,Chenguang Zhao,Peace C. Madueme,Mehdi Hedjazi Moghari,Mubarak Shah,Pegah Khosravi,Yuzhang Zhang
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注: Code, benchmark resources, and model weights are available at this https URL
Abstract:Cardiovascular magnetic resonance (CMR), including cine imaging, is a reference standard for the noninvasive assessment of cardiac morphology and ventricular function. Cine CMR interpretation integrates qualitative visual assessment with quantitative measurements of ventricular volumes, ejection fraction, myocardial mass, wall thickness, and regional wall motion. Current medical vision-language models (VLMs) cannot reliably derive quantitative measurements from multidimensional cine images without analysis tools. We present CineMR, a tool-augmented VLM that invokes cardiac image-analysis tools and integrates their outputs into interleaved reasoning for quantitative CMR assessment. We also construct a multi-cohort visual question answering benchmark covering quantitative metric extraction, multiclass diagnosis, and differential diagnosis, together with tools for segmentation, phase selection, volumetry, morphometry, and regional wall motion analysis. CineMR is trained with supervised fine-tuning (SFT) on tool-interaction traces followed by Group Relative Policy Optimization (GRPO) with conditional tool-use rewards. On the multi-cohort cine CMR benchmark, CineMR achieves 35.9% pass@1 and 58.9% pass@4, compared with 1.5% pass@1 for the Qwen3-VL-8B backbone and 0.0% and 7.0% pass@1 for LLaVA-Med v1.5 and MedGemma-4B, respectively. Correct tool invocation reaches 99.8% after GRPO, up from 78.9% after SFT. Live tool outputs improve ventricular measurement accuracy by 20.4–23.7% over direct model predictions, and removing all tools reduces pass@1 from 35.9% to 27.9%. These results highlight the importance of reliable tool use for quantitative cine CMR reasoning and support CineMR as a promising approach for assistive cardiac image assessment. Code, benchmark resources, and model weights are available at this https URL.
[CV-108] PhysicsLENS: Diagnosing Physical Property Blindness in Video Generation Models
链接: https://arxiv.org/abs/2610.01162
作者: Isaiah Milkey,Som Sagar,Aditya Taparia,Xinyuan Liu,Jiqing Wen,Ransalu Senanayake
类目: Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)
备注:
Abstract:Reliable video world models could provide scalable predictive environments for robot learning, planning, and evaluation. However, generated robot videos can violate physical principles and complete tasks through physically implausible behavior, limiting their reliability for robot learning and planning. Current video-generation benchmarks exclude physics that are inherently hidden by visuals (e.g., weight, viscosity, friction). Due to this, video models are evaluated on the fidelity of physics, not the underlying accuracy of physics. We introduce PhysicsLENS, a dataset and benchmark for evaluating plausibility of physical properties grounded in robotics. PhysicsLENS uses matched scenario pairs that hold the same conditioning frame and task, while varying underlying physics in the scene description. Scenarios are curated from public robot video sources and annotated across seven physical domains: collision, gravity, momentum, friction, deformation, fluid, and causality. We evaluate across four video generation models, producing over 400 human-annotated labels. Results show that plausible-looking videos often ignore the stated property (34 of 47), and that stating the property lowers plausibility only slightly and not significantly.
[CV-109] OptimusMesh: Compact Autoregressive Mesh Generation from Point Clouds via Sparse Latent Pivots
链接: https://arxiv.org/abs/2610.01148
作者: Mazhar Iqbal,Naoya Chiba,Xuanmeng Sha,Tomohiro Mashita,Yuki Uranishi
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Generating compact and geometrically faithful 3D meshes directly from point clouds remains a fundamental challenge. Point clouds are unordered and sparse, whereas meshes exhibit irregular structure and varying topology. As a result, many existing approaches rely on implicit representations followed by surface extraction or reconstruction. Although effective, these pipelines can produce dense or over-smoothed meshes, often requiring computationally expensive post-processing and simplification. We present OptimusMesh, a framework for direct compact triangle mesh generation from point clouds using sparse latent pivot conditioning. Our key idea is to compress 2,048 oriented input points into only 16 sparse latent pivots, reducing the geometric conditioning set by 128\times . These pivots provide a compact structural representation shared across a two-stage autoregressive framework that first generates mesh vertices and then predicts triangular faces conditioned on the generated vertices and the same pivots. Compared with the evaluated recent point-cloud-conditioned autoregressive methods, which use 257 decoder-conditioning tokens, OptimusMesh uses only 16 , yielding a 16.1\times shorter conditioning sequence. Experiments show that OptimusMesh produces the most compact outputs among the compared recent autoregressive methods, using 25.7% – 94.1% fewer faces while maintaining competitive geometric fidelity and distributional quality.
[CV-110] he RSNA Intracranial Aneurysm (RSNA-ICA) Dataset
链接: https://arxiv.org/abs/2610.01135
作者: Maria Correia de Verdier,Rachit Saluja,Jason Sho,Maryam Vabarizad,Rennie Yung-Chieh Chen,Uyen N. T. Nguyen,Mona Alrehaili,Layal Aweidah,Deniz Bulja,Wesley C. Chan,Hernan Chaves,Madhavi Duvvuri,Huseyin Ekin Ergin,Undrakh-Erdene Erdenebold,Ekim Gumeler,Mohamed Sobhi Jabal,Chin-Chi Kuo,Fatima Mubarak,Sevde Nur Emir,Scott Riley K. Ong,Johanna Ortiz,Almudena Pérez-Lara,Andreas M. Rauschecker,Shayan Sirat Maheen Anwar,Charit Tippareddy,Tam Tran,Sorawis Visrutaratna,John Mongan,Adam E. Flanders,Robyn Ball,Greg Zaharchuk,Peter D. Chang,Felipe Kitamura,Errol Colak,Luciano Prevedello,Tyler Richards,Data Contributor Group,Dataset Annotator Group,Evan Calabrese,Jeffrey D. Rudie
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 48 pages (including supplementary material)
Abstract:Intracranial aneurysm rupture is associated with substantial morbidity and mortality, yet aneurysm detection remains challenging, particularly for small lesions and on routine non-angiographic imaging examinations. To support the development and evaluation of artificial intelligence (AI) algorithms for intracranial aneurysm detection and localization, the Radiological Society of North America (RSNA), in collaboration with the American Society of Neuroradiology (ASNR), the Society of Neurointerventional Surgery (SNIS), and the European Society of Neuroradiology (ESNR), curated the RSNA Intracranial Aneurysm (RSNA-ICA) Dataset. Developed for the 2025 RSNA Intracranial Aneurysm Detection Challenge, RSNA-ICA is a large, publicly available, expert-annotated dataset comprising 7202 CTA, MRA, and MRI series from 4278 adult patients collected across 21 institutions in 12 countries spanning five continents. The dataset includes 2566 CTA, 2166 MRA, and 2470 MRI series from patients with and without intracranial saccular aneurysms, providing substantial geographic and imaging diversity. Expert annotations indicate both aneurysm presence and location, and 178 series additionally include three-dimensional segmentations of challenge-defined vascular locations. RSNA-ICA was used to develop and evaluate algorithms in the 2025 RSNA Intracranial Aneurysm Detection Challenge. Of the 7202 image series, 5041 are publicly available through MIRA (this https URL), while the remainder were used for challenge public and private test sets. The dataset is freely available to the research community for noncommercial use and provides a comprehensive resource for advancing AI-based aneurysm detection across both angiographic and routine neuroimaging examinations.
[CV-111] Open Vocabulary Word Recognition From Transcribed Bangla Texts
链接: https://arxiv.org/abs/2610.01134
作者: Faias Satter,Sk. Md. Masudul Ahsan
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注: 6 pages, 4 figures, 5 tables. Accepted version of the paper published in the 2023 26th International Conference on Computer and Information Technology (ICCIT). Code: this https URL
Abstract:An optical character recognition (OCR) can scan a paper and extract text using technology, making people’s jobs easier. While various OCR systems are available in the software industry, finding a reliable equivalent solution for Bangla takes much work. When it comes to handwritten texts, the situation is much more unusual. Recognizing words from word images is the most critical stage in any OCR process. It is the second stage after segmenting words from text pictures. If this stage fails, the overall performance of the OCR will be poor, regardless of how well the other phases perform. This study aims to recognize words using deep learning in a handwritten Bangla word image. Three object detection models, SSD with MobileNetV2, Faster R-CNN with InceptionResNetV2, and an ensemble model of these two, have been used to train and test handwritten word images. A modified Non-Maximum Suppression has been introduced to enhance the effectiveness of the models’ results. A customized dataset of 9841 handwritten Bangla word images has been compiled, featuring diverse handwriting styles from various individuals. All three models’ performances have been checked against the test dataset, and the ensemble model has been the most impressive, with an F1-score of 92.61%. Also, at the word level, the ensemble model correctly recognizes 96.12% of the words to some extent. The system can be further improved by introducing a post-processing phase to correct errors generated by the system.
[CV-112] Affine-Aligned Atlas for Canonical Gaussian Construction in Video Representation
链接: https://arxiv.org/abs/2610.01114
作者: Masaya Takabe,Hiroshi Watanabe,Sujun Hong,Tomohiro Ikai
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Gaussian splatting has recently emerged as an efficient representation for images and videos due to its explicit structure and fast rendering capability. Existing Gaussian-based video representations often decompose a video into canonical Gaussians and temporal deformation. However, when a video contains large global motion such as camera movement, the canonical representation may become misaligned with individual frames, increasing the burden on the temporal deformation model. In this paper, we propose an affine-atlas canonical Gaussian representation, which constructs canonical Gaussians in a larger affine-aligned atlas space. Frame-wise affine transforms absorb global motion before canonical Gaussian construction, reducing the gap between the canonical representation and target frames. Since the proposed method only modifies the canonical construction stage, it can be integrated into existing canonical-Gaussian-based methods with negligible additional parameter cost. Experiments show that our method improves reconstruction quality especially for sequences with large camera motion.
[CV-113] MVDG: Efficient Multi-view 3D Disambiguation on Unconstrained Real-World Images
链接: https://arxiv.org/abs/2610.01098
作者: Hanyuan Xiao,Gonglin Chen,Haolin Xiong,Wenbin Teng,Haiwei Chen,Yajie Zhao
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Illusory matches between distinct yet visually similar 3D surfaces–doppelgangers–remain a fundamental obstacle for large-scale, in-the-wild 3D reconstruction and visual localization. Prior work mitigates this issue with pairwise classifiers, but this design limits multi-view contextual reasoning and incurs O(n^2) inference complexity for downstream structure-from-motion (SfM). We present MVDG, a scalable multi-view disambiguation framework built on the 3D foundation model VGGT, which jointly reasons over an arbitrary number of multiview images. By incorporating 3D-aware multi-view features, our method reduces dependence on pairwise comparisons by encoding and decoding views in a single pass. We further observe that direct multi-view fine-tuning of VGGT can be unstable under noisy supervision; motivated by label ambiguity in Doppelgangers, we construct a pseudo-pairwise training set from AerialMegaDepth and show that fine-tuning on sampled subsets yields stable optimization and strong generalization to held-out scenes. Finally, because full SfM evaluation (even with faster pipelines such as GLOMAP) remains expensive, we process a pseudo-pairwise dataset for efficient validation; we derive a predictive relationship between regular SfM metrics and the classification accuracy on this pseudo-pairwise test. Experiments show that our method achieves comparable pairwise accuracy while improving both SfM accuracy and inference speed over baselines.
[CV-114] Dataset Identity Not Novelty: The Source of an Inflated OOD Detection Gain
链接: https://arxiv.org/abs/2610.01096
作者: Donghoon Lee,Shinjin Kang
类目: Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)
备注: 30 pages, 7 figures, 33 tables
Abstract:A post-hoc out-of-distribution (OOD) detector reads the activations of a trained classifier and returns a score. It fits that score on in-distribution data, and the benchmarks that evaluate it supply a second piece of OOD data for the fitting itself. Some detectors tune a constant on it. Others fit a direction in feature space or train a flexible combiner and report the number that it reaches as the gain that is still available. Every such fit is validated on held-out samples of the same OOD dataset. That check rules out memorizing individual images. It says nothing about a fit that has instead learned which dataset it is looking at, and a direction that recognizes one OOD dataset rather than novelty passes it perfectly. The detector that a practitioner installs meets OOD data from a source that nobody fitted it on, so the difference decides what the reported number is worth. We measure it by holding out the whole OOD dataset rather than a sample of it, and we call that gap the inflation. We read it across a range of combiners on ImageNet and CIFAR-100 backbones. Most of the gain that the usual protocol reports turns out to be dataset identity rather than novelty. The size of the fit does not move what survives, so the effect is not ordinary overfitting. The share depends instead on whether the input exposes class identity, and two controls that vary that property alone separate the inflation on every backbone of both benchmarks. A closed form accounts for the effect and computes it from the fitting rows, so a practitioner can tell which fits will inflate without running the hold-out protocol. One of these fits survives, namely the single constant that the field already picks on a designated validation dataset. Anything above it reports a gain that the hold-out protocol does not return, and on one benchmark what survives falls while what is reported climbs.
[CV-115] Ego2Act: Evaluating Goal-Directed Manipulation in Egocentric Video Generation
链接: https://arxiv.org/abs/2610.01092
作者: Patrick Amadeus Irawan,Iskandar Muda Rizky Parlambang,Rava Maulana,Qinrong Cui,Erland Hilman Fuadi,Zayd M. K. Zuhri,Nanda Ryaas Absar,Ahmed Elshabrawy,Wilfried Ariel Mulyawan,Shoubin Yu,Yue Zhang,Mohit Bansal,Alham Fikri Aji
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Preprint. 51 pages, 19 figures, 23 tables. Code, dataset and project website linked in the paper
Abstract:Video generation models are increasingly being explored as world simulators for embodied planning and learning. To do so effectively, these models must not only generate visually appealing frames, but also predict how environments dynamically evolve when executing goal-directed actions. While evaluating these capabilities is crucial, existing benchmarks focus mainly on single short actions or step-by-step instructions. This leaves multi-step physical reasoning underexplored, especially in egocentric video generation that requires planning to simulate proper execution to accomplish high-level goals by carrying out multiple real-world manipulations. We introduce Ego2Act, a goal-directed benchmark featuring 2,640 videos from 110 real-world tasks across day-to-day settings, varying object clutter and multi-step complexity. Given an initial scene image and a high-level goal, Ego2Act evaluates whether video generation models can produce realistic egocentric videos of a hand manipulating objects to carry out the task. To support scalable evaluation, we also introduce Ego2ActJudge, a reference-free evaluation pipeline that achieves better task completion and physics plausibility evaluation alignment with human consensus compared to relevant baselines. Our findings reveal that models’ generated simulations often skip or partially execute steps, leaving later steps missing dependent states, which leads to unfulfilled goal. Furthermore, models consistently fail at fine-grained physical dynamics, particularly during complex object manipulation and persistent world modeling. We hope Ego2Act provides a rigorous testbed for advancing video models toward physically plausible, goal-directed simulation.
[CV-116] Overcoming Kernel Redundancy for Scaling Logic Gate Networks NEURIPS2026
链接: https://arxiv.org/abs/2610.01069
作者: Sejin Park,Hongjae Lee,Changwoo Han,Seung-Won Jung
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: NeurIPS 2026
Abstract:Differentiable logic gate networks, which operate using only logic gates, have recently attracted attention as an efficient alternative to conventional neural networks. However, despite their efficiency, the scaling behavior of logic gate networks remains underexplored. By contrast, scaling model capacity is a central design principle in deep neural networks and typically leads to improved performance. This discrepancy raises a key question: Can similar scaling benefits also be achieved in logic gate networks? In this work, we focus on width as a primary scaling axis and conduct a systematic analysis of its behavior in logic gate networks. We observe that naive width scaling often introduces redundancy among logic kernels, limiting the effective use of additional kernels and leading to performance saturation. To address this limitation, we propose a dynamic logic kernel framework that reorganizes kernel utilization by promoting specialization across kernel groups. This enables the network to better utilize increased width via input-dependent kernel routing, while ensuring that both routing and computation are implemented entirely with gate-level Boolean operations at inference time. We further find that kernel redundancy is most pronounced at the first gate level, motivating an early-stage dynamic logic kernel strategy that concentrates adaptation at this level. Experimental results demonstrate that our approach improves kernel utilization and increases kernel diversity, leading to higher accuracy with improved parameter efficiency.
[CV-117] HierGF: Hierarchical Gaussian Fields via Geometry-perception Message Passing for Sparse-view 3D Reconstruction
链接: https://arxiv.org/abs/2610.01056
作者: Bi’an Du,Zhimin Zhang,Daizong Liu,Baoquan Chen,Wei Hu
类目: Computer Vision and Pattern Recognition (cs.CV); Image and Video Processing (eess.IV)
备注: Accepted to IEEE Transactions on Multimedia (TMM), 2026
Abstract:Sparse view 3D reconstruction is an important and common scenario in multimedia applications, such as augmented reality/virtual reality (AR/VR) content creation, cultural heritage digitization, and certain robotic applications, where only a limited number of randomly captured views may be available. However, sparse views contain only limited 3D information, posing two major challenges:1) too few images are available for matching, making it difficult to build multi-view consistency; 2) insufficient view coverage leads to a lack of information in under-sampled regions, resulting in missing parts of object structure. Existing methods mostly still rely on limited reprojection errors and regularization terms, which are prone to overfitting to a single view and inconsistent appearances across views. In geometrically under-sampled regions, they often rely on heuristic density control, lacking reliable guidance and often resulting in blurring and structural this http URL address these issues, this paper proposes Hierarchical Gaussian Fields (HierGF), which revisits sparse-view reconstruction from a hierarchical geometry-perception perspective and converts limited observations into reliable self-generated supervision beyond fixed priors and heuristic density control. In particular, we transform coarse 3D geometric information and additional 2D generative priors into structured pseudo-supervision through a two-stage geometry-perception backbone network, thereby enhancing multi-view consistency with very few input views. In addition, we introduce a learnable confidence network to guide gradients toward cross-view consistent content, and a geometrically consistent densification module to improve the reconstruction of multi-view alignment and under-sampled regions.
[CV-118] owards Subject Consistency over Dynamic Subject Sets in Video Generation
链接: https://arxiv.org/abs/2610.01052
作者: Tongcheng Zhang,Jun Zhu,Jianfei Chen
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Project website: this https URL
Abstract:We argue that as video generation extends to longer durations, subject consistency should be evaluated over \textitdynamic subject sets. We therefore introduce \textbfDynSC-Eval, an evaluation framework that dynamically tracks eligible subjects throughout their visible lifespans and measures local continuity and global identity preservation using six complementary object-level metrics, with explicit detection of inconsistency events. To validate its effectiveness, we design synthetic experiments that actively inject inconsistency events, demonstrating both the sensitivity of DynSC-Eval and the limitations of existing metrics. Evaluations of diverse models on 5s, 15s, and 60s video generation further reveal substantial subject consistency differences that are obscured by conventional metrics. Beyond evaluation, we construct rewards from DynSC-Eval and apply DiffusionNFT post-training in an autonomous-driving testbed. On 5s generation, our approach reduces the six inconsistency metrics by an average of 13.82% for Wan-2.1-1.3B and 5.66% for SANA-2B, with improvements also observed on the I2V model ReSim. Qualitative comparisons further demonstrate the effectiveness of our method. We then extend generation to 10s and 30s through curriculum learning and show that consistency optimization remains effective while largely preserving other capabilities.
[CV-119] Bootstrapping Video Interaction Generation with Synthetic State Transitions IJCAI2026
链接: https://arxiv.org/abs/2610.01039
作者: Jiho Jang,Jinyoung Kim,Nojun Kwak,Kyungjune Kim
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: IJCAI 2026
Abstract:While recent video generative models can synthesize high-fidelity videos, they struggle to portray plausible physical interactions and the resulting state transitions, a critical bottleneck for applications in robotics and VR/AR. To address this, we introduce a framework to generate a scalable synthetic dataset of controllable interactions. Our pipeline leverages a structured taxonomy and state-of-the-art image editing models to create explicit start' and end’ state images, which serve as visual anchors for the interaction. To generate a seamless video utilizing these anchors, we propose State-Guided Sampling (SGS), a novel sampling technique that mitigates artifacts common in naive conditional generation. Furthermore, we develop and validate a new automated evaluation system that aligns with human judgments to ensure data quality. Experiments show that fine-tuning a base model on our dataset significantly enhances its ability to generate plausible interactions.
[CV-120] owards Automatic Video Annotation with ASH: Zero-Shot Open-Vocabulary Multi-Object Tracking and Segmentation
链接: https://arxiv.org/abs/2610.01022
作者: Arash Rocky,Q. M. Jonathan Wu
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Memory-attention-based Video Instance Segmentation (VIS) methods have demonstrated strong zero-shot tracking capability, yet their substantial memory requirements confine them to short video clips and their single-prompt inference design makes multi-category open-vocabulary tracking computationally prohibitive. This work introduces two contributions toward fully automated tracking annotation of arbitrary video. The Generalized Presence Token (GPT) reformulates SAM3’s inference pipeline to process N text prompts simultaneously via virtual prompt batching, reducing image encoding cost from O(N) to O(1) with no modifications to any learned component. The Annotation and Segmentation Handler (ASH) extends any memory-attention VIS tracker to sequences of arbitrary length through overlapping temporal chunks with IoU-based inter-chunk identity matching, requiring no dataset-specific training. Instantiated on SAM3, the resulting pipeline – SAM3-ASH – achieves state-of-the-art HOTA on MOTS20 under fully zero-shot conditions and remains competitive with trained specialists across seven additional benchmarks, while peak GPU memory consumption stays below 25 GB, establishing a practical baseline for scalable, training-free automated video annotation.
[CV-121] FutureWorlds: Learning Robotic World Models from Alternative Futures
链接: https://arxiv.org/abs/2610.01019
作者: Hao Wu,Shengju Qian,Weiyan Wang,Fan Xu,Fan Zhang,Yuanpeng He,Qingsong Wen,Yuxuan Liang
类目: Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)
备注: 32 pages, including references and appendix. Code: this https URL
Abstract:Robotic world models predict action-conditioned future scenes, providing a foundation for understanding action outcomes. However, turning alternative predictions into useful learning signals remains challenging: similar candidates limit informative quality comparisons, while diverging trajectories require persistent maintenance of their individual histories. We introduce FutureWorlds, a framework that unifies candidate construction, history maintenance, and learning from relative quality. Built on a multimodal discrete autoregressive model, FutureWorlds uses diverse beam search during reinforcement learning to construct candidate futures that balance confidence and diversity. Candidate-specific bounded memory preserves scene states and ensures that generation and policy scoring use matching histories. We further propose MemSPO (Memory-Conditioned Search-Guided Policy Optimization), which converts video trajectory rewards into group-relative advantages to optimize the world model. On RT-1, BridgeV2, and RoboCasa, FutureWorlds reduces LPIPS for 32-frame predictions by 14.78%, 20.84%, and 9.12%, respectively, relative to the strongest baseline on each dataset. Under fixed evaluation configurations, only 200 MemSPO updates further improve generation quality and support continued prediction beyond the training horizon. Memory ablations, decoding sensitivity analysis, and optical-flow evaluation show that these gains extend beyond visual quality to more accurate motion prediction and more consistent object states. Project page and code: this https URL.
[CV-122] VASC: Value-Aware Sparse Attention with Cross-Layer Memory for Efficient 3D Reconstruction
链接: https://arxiv.org/abs/2610.01013
作者: Junyi Wu,Fanqing Kong,Leyang Chen,Shaoqiu Zhang,Yulun Zhang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 21 pages, including references and appendices
Abstract:Feed-forward 3D vision models such as VGGT have achieved remarkable progress, unifying camera estimation and dense scene reconstruction in a single pass. However, their quadratic global attention makes long image sequences expensive, while existing sparse methods may favor highly attended yet value-redundant regions. To address these limitations, we introduce VASC, a training-free sparse attention method combining value-aware block selection and execution-aware cross-layer memory. Our value-aware block selection integrates pooled query–key relevance with neighboring value contrast, reducing redundancy while preserving query-relevant and distinctive content. Cross-layer memory tracks unserved demand across layers and updates this state according to actual execution, enabling previously underserved blocks to compete under a fixed computation budget. Experiments on 7Scenes and NeuralRGB-D with VGGT and \pi^3 demonstrate improved pose estimation and reconstruction quality compared with FasterVGGT, together with up to 2.29\times faster inference than dense VGGT. Code is available at this https URL.
[CV-123] Watch Your Speech: Text-aware Video-to-Speech Synthesis with Textual Conditioning BMVC2026
链接: https://arxiv.org/abs/2610.01012
作者: Gunwoo Lee,Yoori Oh,Yoseob Han
类目: Computer Vision and Pattern Recognition (cs.CV); Sound (cs.SD); Audio and Speech Processing (eess.AS)
备注: Accepted to BMVC 2026
Abstract:Video-to-speech synthesis aims to generate natural-sounding speech from silent talking-face videos while ensuring phonetic accuracy. A fundamental challenge in this task is the inherent one-to-many mapping problem, where visual dynamics often lack sufficient information to uniquely determine the corresponding utterance. To address this, we propose Watch Your Speech (WYS), a video-to-speech synthesis framework that incorporates textual conditioning as an explicit linguistic cue to mitigate visual ambiguity. Our framework features an attention-based embedding fusion module that synergistically integrates textual context with video sequences, coupled with a conditional flow matching objective for high-fidelity speech generation. Extensive experiments on the LRS2 and LRS3 datasets demonstrate that WYS achieves superior performance, establishing new state-of-the-art results in audio-visual synchronization (LSE-C/D) while maintaining highly competitive textual accuracy (WER). Subjective evaluations further confirm that our model generates speech with near-human naturalness, validating the effectiveness of textual conditioning in content-controlled video-to-speech synthesis. Project page: this https URL
[CV-124] VIEScore2: Unified Image Evaluation with Spatially Grounded Explanations
链接: https://arxiv.org/abs/2610.00994
作者: Xianda Du,Max Ku,Weiming Ren,Zhi Rui Tam,Chunlin Ren,Ping Nie,Min-Hung Chen,Wenhu Chen
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Preprint. Project page: this https URL
Abstract:Existing synthetic image evaluators typically provide only a scalar quality score and do not identify the image regions that support it. We introduce VIEScore2, a unified evaluator for image generation and editing tasks with optional conditioning images. VIEScore2 represents an image as an N x N grid and jointly predicts quality scores and defect locations in a single model pass. Its text-native grid representation provides a common interface for heterogeneous spatial supervision and enables directly verifiable post-training objectives. We train on 38K examples spanning score-only, localization-only, and joint supervision across generation and editing tasks. Starting from supervised fine-tuning, we further apply GRPO to improve defect localization using rewards that combine cell-level Dice overlap, score accuracy, and output-format validity. A parameter-free parser converts the structured predictions into readable explanations. On the primary suite, VIEScore2 achieves an overall-score SRCC of 0.601, compared with 0.491 for Gemini-3-Flash, the strongest zero-shot general-purpose VLM baseline under matched inputs. For defect localization, VIEScore2 outperforms both general-purpose VLMs and specialized spatial evaluators on three of six benchmarks in per-image grid IoU and ranks among the top three on five, including datasets beyond its training sources.
[CV-125] NarrativeFlow: Flow-Based Vision-Language-Action Model Using Robot Velocity Fields ACCV2026
链接: https://arxiv.org/abs/2610.00981
作者: Shota Kobayashi,Koki Seno,Daichi Yashima,Komei Sugiura
类目: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted at ACCV 2026
Abstract:We focus on language-conditioned flow-based manipulation, where robot flows (robot velocity fields) serve as embodiment-agnostic, motion-centric representations for leveraging data collected from multiple robot platforms. This task is crucial because language-conditioned manipulation is essential for practical robotic systems, yet scaling robot foundation models remains limited by the labor-intensive collection of embodiment-specific data. Existing methods either coarsely approximate robot flows with sparse keypoint displacements, or cannot handle language-conditioned manipulation. To address this limitation, we propose NarrativeFlow, which models robot flows as continuous velocity fields using a flow-matching formulation conditioned on language. Accordingly, NarrativeFlow generates robot flows that are physically consistent with real-world manipulation. To validate NarrativeFlow, we have conducted experiments on standard datasets for language-conditioned manipulation. The experimental results show that NarrativeFlow outperforms representative baseline methods on standard evaluation metrics. Furthermore, through real-world experiments, we show that NarrativeFlow achieves higher success rates than baseline methods across multiple manipulation tasks. The project page is available at this https URL
[CV-126] Concept Driven Domain Adaptation: Finding an Abstract Needle in a Haystack
链接: https://arxiv.org/abs/2610.00973
作者: Haiming Zhao,Tai Wang,Kun Zhang,Xicheng Peng,Zhiyang Li
类目: Computer Vision and Pattern Recognition (cs.CV); Physics Education (physics.ed-ph)
备注: 19 pages, 10 figures
Abstract:Science teachers frequently search for documentary excerpts not by describing what appears on screen, but by querying the abstract concepts they intend to teach. This use case exposes a limitation of existing language-based video moment retrieval methods, which typically assume that queries describe observable events, whereas instructional search requires retrieving concrete visual phenomena that instantiate an underlying scientific principle. We study this setting as concept-to-example video retrieval, an abstract-needle-in-a-haystack problem where compact curriculum concepts must be grounded in temporally sparse documentary evidence. To bridge this abstraction gap, we propose Concept-Driven Domain Adaptation (CDDA), a three-stage framework for adapting two-tower vision-language models to concept-level retrieval. CDDA treats concepts as intermediate semantic anchors: it first structures the textual embedding space with textbook and teacher-handbook example-concept pairs, then transfers this concept-aware geometry to documentary visuals under a frozen visual encoder, and finally jointly adapts both encoders with sparse visual concept supervision. From a geometric perspective, this staged alignment reduces text-concept and vision-concept angular gaps, thereby encouraging concept-level adaptation while preserving the pretrained model’s concrete image description alignment. On a curated middle-school physics retrieval benchmark, CDDA achieves stronger pedagogically oriented concept retrieval than several competitive multimodal baselines, including Qwen3-VL-Embedding-2B, while maintaining concrete image-text matching after adaptation.
[CV-127] RelationVGGT: Visual Geometry Transformers for 3D Spatial Relation Segmentation NEURIPS2026
链接: https://arxiv.org/abs/2610.00970
作者: Minsu Kim,Jaesung Choe,Jiwoo Lee,Yu-Chiang Frank Wang,Seon Joo Kim
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注: 10 pages, NeurIPS 2026 accepted (poster)
Abstract:Recent advances in 3D reconstruction have progressed from per-scene optimization to feed-forward inference, and semantic scene understanding has followed suit – yet existing methods remain confined to object-centric perception, neglecting spatial relations between objects. We formulate 3D spatial relation segmentation in a feed-forward, pose-free multi-view setting: given a visually specified subject and a relational text query, the model segments the target across views without receiving its category name. To this end, we propose RelationVGGT, a novel feed-forward framework that integrates semantic features from a visual foundation model with geometry-aware representations from a 3D geometry foundation model and leverages a relation transformer for subject-conditioned, cross-view relation prediction – requiring neither per-scene optimization nor known camera poses. We additionally provide a fully automated annotation pipeline built on ScanNet++ with VLMs and LLMs, enabling scalable training data generation for this new task.
[CV-128] Video-Index: A Curated Meta-Benchmark for Video Understanding WWW
链接: https://arxiv.org/abs/2610.00960
作者: Enxin Song,Yinuo Xu,Shusheng Yang,Wenhao Chai,Jiatao Gu
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Blog: this https URL GitHub: this https URL Hugging Face: this https URL
Abstract:A video benchmark should reward the capability it claims to measure, yet models can exploit answer options, question text, or partial visual evidence. We introduce the attack pyramid, five levels of shortcut attacks with increasing access to each item, and audit 115 video benchmarks with it. On 35 benchmarks, attackers that never see a frame approach full-video accuracy. On 51 benchmarks with temporal probes, shuffled frames keep a median 96% of full-video accuracy. Near-duplicate questions make up at least half the items in 63 benchmarks. We screen 505,518 question-answer pairs from 112 of them into an audited pool. Agents turn evaluation requests into specifications, and a deterministic selector with a red-team gate composes reproducible benchmarks. We release Video-Index, the 210 hardest verified items under these attacks in each of four capability groups, 840 items from 76 sources. With the same fixed input, Claude Opus 5 outscores every open-source model by over 37 percentage points, and agent tools add about 20 more, yet all systems leave room to improve efficiency and accuracy. Blog: this https URL GitHub: this https URL Hugging Face: this https URL
[CV-129] wo Clocks in Diffusion MLLM s: When Answers Stabilize Before Rationales Unfold NEURIPS2026
链接: https://arxiv.org/abs/2610.00953
作者: Keuntae Kim,Yong Suk Choi
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注: NeurIPS 2026 Workshop on BeNTo (Beyond Next-Token Prediction - Diffusion Flow Models for Next-Generation Decoding)
Abstract:An answer candidate in a masked diffusion MLLM can stabilize while its rationale is still unfolding. We distinguish retrospective stabilization of the logged candidate from token commitment, and examine these two clocks relative to rationale generation. Analyzing our results across three visual question-answering benchmarks, we find that 89.4-98.1% of the rationale-side canvas remains unwritten at stabilization in single-block, EOS-suppressed LaViDa runs. On VBench, reducing block length from 128 to 8 changes this fraction from 89.4% to 1.7%, together with answer coverage and the eligible observation window. Under EOS-enabled prompting, direct instructions improve Nemotron’s overall accuracy by 15.0 and 19.5 percentage points on M3CoT and ScienceQA, but reduce LaViDa/VBench accuracy by 11.0 points. A symmetric decomposition associates the larger absolute component of each change with coverage rather than conditional accuracy. Matched-canvas image ablations measure visual sensitivity alongside answer stabilization, separating the two temporal readouts. Together, these measurements distinguish answer stabilization, rationale unfolding, and visual sensitivity, and identify coverage as the larger component of the prompting differences.
[CV-130] A Matched-Budget Audit Framework for Recaptioned Image-Text Supervision Distributions
链接: https://arxiv.org/abs/2610.00952
作者: Giyeong Oh,Junghun Park,Yuhan Bae,Youngjae Yu
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注: initial commit
Abstract:Recaptioned image-text corpora are now standard for text-to-image (T2I) training, with vision–language model (VLM) captioners replacing sparse alt-text by dense descriptions. A recaptioned corpus is a supervision distribution induced by a documented captioning policy ( \pi ), captioner ( V_c ), and source corpus ( C ). Length-correlated proxies miss caption-register artifacts and downstream T2I benchmarks entangle the corpus with training choices, so this distribution is hard to audit at corpus scale. We introduce a reusable matched-budget audit framework for recaptioned supervision distributions D_\pi,V_c,C : at a fixed text budget of B = 64 it reports a five-axis profile spanning prompt-side coverage, image-conditioned faithfulness, and caption-surface health, with claimed controllable basic units (CBU) as the common claim unit. We instantiate the framework on seven paired comparisons over five public source corpora. Across the four cross-corpus pairs, the released surface raises supported CBU per caption by +3.39 to +6.36 under both Qwen and Gemma Judges, and on CC12M the same framework exposes a long-vs-dense frontier that is consistent across both judges and four budgets. We release the audited multi-source recap corpus ( \approx 490M) together with the audit-artifact bundle.
[CV-131] Joint Branch-Space Transform Coding for Diffusion Activation Quantization with Classifier-Free Guidance
链接: https://arxiv.org/abs/2610.00930
作者: Mingrun Jiang,Yuejia Liu,Zishan Shao,Ting Jiang,Qinsi Wang,Hancheng Ye,Yixiao Wang,Rui-Feng Wang,Kangning Cui,Yixuan Chen,Fan Yang,Xiang Cheng,Hai Li,Yiran Chen
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (stat.ML)
备注:
Abstract:Post-training quantization for diffusion models increasingly exploits timestep, feature, and layer structure. While recent work has begun incorporating CFG structure into diffusion quantization, activation quantization still operates independently across conditional and unconditional coordinates, leaving cross-activation structure unexploited. We show that matched CFG activations form a strongly correlated two-dimensional source and that, under a fixed bit budget, the choice of branch coding basis materially affects quantization fidelity. Motivated by this observation, we introduce branch-space transform coding, which rotates matched CFG branches via an offline derived 2x2 orthogonal matrix, requiring minimal modifications to model parameters or the quantization pipeline. We further derive the Guidance-Correlation Branch Transform (GCBT), which jointly incorporates the CFG guidance direction and cross-branch second moments. Under an equal-rate quantization-noise surrogate, GCBT admits a closed-form per-layer solution without gradient optimization or angle search. Applied on top of existing diffusion PTQ methods, GCBT yields statistically significant fidelity gains in most evaluated comparisons with no statistically significant degradation, while leaving the underlying host quantization pipeline unchanged.
[CV-132] Platonic Task Arithmetic NEURIPS2026
链接: https://arxiv.org/abs/2610.00929
作者: Junghwan Park,Woojin Cho
类目: Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (stat.ML)
备注: NeurIPS2026
Abstract:Models specialized for the same task converge to similar behavior, yet the parameter updates that produce it share no common coordinate system, so weight-space task arithmetic stays confined to a single model and cannot cross architectures without a structural correspondence. Drawing on Plato’s allegory of the cave, we hypothesize that these model-specific updates are shadows of one shared, model-agnostic object, which we call the platonic task vector. To make it operational for models that pair an image or audio encoder with a text encoder, we introduce Universal Task Descriptors: matrices whose shape is independent of architecture and embedding dimension, which record a task’s functional effect and support addition and negation as matrix operations. Transferring a descriptor into a target means editing the target until it reproduces the descriptor on the task’s unlabeled probe images and class-name prompts, requiring no per-image labels. We realize this edit in two ways. First, the descriptor factorizes into a shift field on image embeddings, so a single least-squares solve yields a linear operator that folds into the target’s last layer as a weight edit; by linearity, a bank of such operators admits any composition at any strength as a signed sum. Second, a low-rank adapter trained on the same objective reaches every layer and fits compositions jointly, at the cost of one optimization per edit. Heterogeneous models share this object only partially, with a model-specific residual comparable in norm to the shared component, yet cross-model transfer still retains 74-80 percent of the gain of the target’s own descriptors. Experiments across six model families, eight classification tasks, and an audio-text setting show that task knowledge transfers and composes across heterogeneous models under both realizations.
[CV-133] A Survey on End-to-End Autonomous Driving Training from the Perspectives of Data Strategy and Platform
链接: https://arxiv.org/abs/2610.00926
作者: Chengkai Xu,Yiming Cui,Jiaqi Liu,Yicheng Guo,Cheng Qin,Geyuan Zhang,Xinwei Dong,Shiyu Fang,Peng Hang,Jian Sun
类目: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注: 21 pages, 6 figures, accepted by IEEE transactions on intelligent transportation systems
Abstract:Autonomous driving is a cornerstone technology for the future of intelligent transportation, where end-to-end learning has emerged as a transformative paradigm that directly maps multimodal sensory inputs to driving actions through unified differentiable models. While offering advantages, the effectiveness of end-to-end autonomous driving (E2E-AD) is ultimately determined by the quality of its training ecosystem. This paper provides a comprehensive review of training methods and ecosystem for E2E-AD. We introduce a Data-Strategy-Platform taxonomy that conceptualizes training as an interdependent system. The data layer defines what can be learned, the strategy layer governs how learning aligns with driving objectives, and the platform layer supports scalability and continuous evolution. Within this framework, we survey recent advances across data-centric pipelines, learning paradigms, and training infrastructures, and analyze their interplay in shaping model performance, robustness, and deployability. Finally, we reflect on current limitations and articulate a forward-looking vision that emphasizes a shift from data quantity to data value, from isolated optimization to foundation-driven generalization, and from static training to integrated training-testing loops, aiming toward robust, scalable, and trustworthy autonomous driving systems. We maintain a continuously updated repository tracking cutting-edge literature and works at \hrefthis https URLOur Project Page.
[CV-134] EyeTAG: Eye Trajectory-Aware Gaze Estimation BMVC2026
链接: https://arxiv.org/abs/2610.00922
作者: Jungmin Lee,Niamat Ullah,Yoseob Han
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注: Accepted to BMVC 2026
Abstract:Gaze estimation under natural head-eye motion underpins applications from driver monitoring to human-computer interaction. Single-frame methods predict each frame independently, so consecutive outputs fluctuate as jitter. Multi-frame methods reduce this, but they learn motion implicitly inside appearance features, so the gaze trajectory is never an explicit variable. We propose EyeTAG (Eye Trajectory-Aware Gaze Estimation), a causal multi-frame framework built around an explicit first-order gaze prior: at each step it differentiates its own recent predictions and feeds the resulting trajectory back as a compact kinematic token. Because differencing is translation-invariant in gaze space, this token carries subject-invariant motion rather than personal gaze offsets. Face and eye streams supply visual evidence, fused by cross-attention and a causal Transformer decoder. EyeTAG reduces the mean angular error by about 1.0 ^\circ on Gaze360 and performs on par with the strongest baseline on EVE (2.56 ^\circ vs. 2.58 ^\circ ). Within-model ablations, which keep the encoder and the rest of the architecture fixed and vary only the gaze history, show that the differential formulation, rather than temporal context alone, removes the systematic saccade bias that persists even with an absolute gaze-history prior. Our code is available at this https URL.
[CV-135] owards Fast and Disentangled Counterfactuals for Visual Foundation Models
链接: https://arxiv.org/abs/2610.00895
作者: Sidney Bender,Benedikt Kunz,Ahmed Zeid,Shinichi Nakajima,Klaus-Robert Müller,Marco Morik
类目: Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Foundation models remain vulnerable to spurious correlations and ``Clever Hans’’ strategies. Explainable machine learning can find and remove such strategies for classifiers without metadata. For foundation models, no such option exists yet. We propose Disentangled Diffusion Autoencoders (DiDAE). DiDAE wraps a frozen foundation model in a conditional diffusion decoder. A counterfactual is one closed-form edit along a direction of a disentangled dictionary, followed by decoding. The dictionary can be supervised (Procrustes) or unsupervised (Singular Value Decomposition, Sparse Autoencoders). No gradients are needed, so DiDAE is up to 2000 times faster than the state of the art. We evaluate on six datasets, two synthetic and four real-world. In a desiderata-driven benchmark on three of them, its counterfactuals are on par with or better than the state of the art, and they repair downstream classifiers through Counterfactual Knowledge Distillation (CFKD), where they beat metadata-based correction. The same machinery can rank a pretrained dictionary against a trained classifier. It returns the few directions the classifier actually reads, each causally verified by a counterfactual that flips the decision, and repairs the classifier along those a teacher marks spurious. The workflow is plug-and-play in our open-source Peal library we publish alongside the paper. With a public dictionary and a pretrained decoder, all that remains is a cheap linear distillation of the classifier and its own fine-tuning.
[CV-136] Machine Translation for Sign Languages
链接: https://arxiv.org/abs/2610.00881
作者: Ozge Mercanoglu Sincan,Anton Pelykh,Edward Fish,Harry Walsh,JianHe Low,Karahan Sahin,Oline Ranum,Sobhan Asasi,Steven Emery,Richard Bowden
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted for publication in the Annual Review of Linguistics, Volume 13
Abstract:Sign language machine translation has progressed substantially over the past decade, evolving from isolated sign recognition to end-to-end translation systems. Advances in pose estimation, transformer architectures, and large-scale dataset collection have driven progress, yet challenges remain. Datasets are limited compared to spoken-language resources; evaluation metrics inadequately capture the linguistic quality of output; and models must capture the simultaneous, multi-layered, and three-dimensional structure of sign languages. This manuscript provides a comprehensive review that seeks to balance technical challenges with stakeholder considerations. We examine the linguistic properties that make sign languages computationally unique, trace the evolution of recognition, translation, and production systems, and analyze ongoing technical challenges. Crucially, we address ethical considerations around data governance, community involvement, and appropriate use. Drawing on interdisciplinary perspectives spanning computer vision, sign language linguistics, and deaf studies, our analysis emphasizes that continued progress requires sustained collaboration across these fields and with deaf communities.
[CV-137] UniTrackPLA: Unified Panorama-Language-Action Model for Instruction-Guided Navigation and Dynamic Person Tracking
链接: https://arxiv.org/abs/2610.00878
作者: Pengfei Qi,Haoran Lin,Sizhuang Chen,Kai Luo,Sirui Zhang,Xinqi Liu,Fei Cheng,Wenrui Chen,Liming Yin,Kailun Yang
类目: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV); Image and Video Processing (eess.IV)
备注: The project page is at this https URL
Abstract:General-purpose embodied robots should support both navigation toward language-specified destinations and dynamic person tracking under arbitrary initial target azimuths. However, existing methods typically rely on forward-facing observations and address these tasks with separate policies, limiting omnidirectional perception and unified closed-loop control. We present UniTrackPLA, a unified panorama-language-action model for instruction-guided navigation and dynamic person tracking. Its Panoramic-Aware Encoding (PAE) preserves the temporal and azimuthal structure of perspective views projected from each panorama, enabling perspective-pretrained visual encoders to process omnidirectional observations. A shared vision-language backbone grounds instructions in the panoramic context and predicts continuous robot-centric waypoint chunks for both tasks. World-Action Consistency (WAC) further predicts action-conditioned future visual states and verifies waypoint prefixes online, allowing reliable actions to be reused while triggering replanning upon inconsistency. We also introduce OmniTrackNav-Bench, comprising 5,000 simulated tracking trajectories, 10,000 simulated VLN routes, and 96 verified real-world routes, providing 919,978 waypoint-supervision instances. UniTrackPLA improves overall tracking SR from 23.50% to 35.00% and Omni-VLN SR/SPL from 13.00%/12.77% to 19.75%/19.29%. Incorporating 76 real-world routes further improves held-out EP@0.2m from 42.92% to 92.08%. Closed-loop experiments on a Go2-W robot demonstrate unified panoramic tracking and navigation across indoor and outdoor environments. The project page is at this https URL.
[CV-138] Dont Waste the Noise: Importance-Guided Perturbation Allocation under Joint Global and Local Constraints
链接: https://arxiv.org/abs/2610.00861
作者: Melika Shirian,Kianoosh Vadaei
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Adversarial optimization under a shared \ell_1 budget requires deciding not only how much perturbation to use, but also where that limited budget should be spent. This allocation problem becomes particularly important when individual input coordinates are subject to local magnitude constraints, which restrict the extent to which perturbation can be concentrated on a small number of locations. We introduce an importance-guided allocation mechanism that uses a fixed clean-gradient prior to steer perturbation toward model-sensitive regions while leaving the feasible perturbation set unchanged. A centered allocation objective encourages perturbation at above-average importance locations and discourages unnecessary expenditure elsewhere, thereby redistributing rather than enlarging the available budget. Across ten robust model–dataset configurations under a common capacity-limited threat setting, the proposed method improves attack success over matched APGD- and PMA-based baselines by 2.52 to 17.70 percentage points. Allocation analysis shows that these gains are accompanied by substantially greater perturbation mass in high-importance regions without increased global \ell_1 consumption. Mechanism ablations further show that centered non-uniform redistribution provides part of the benefit, while model-derived importance yields an additional improvement. These results identify perturbation allocation as a distinct and practically relevant dimension of adversarial optimization under shared-budget, locally constrained threat models.
[CV-139] CtrlWAM: Controllable World Action Models with Aligned Intent and Foresight
链接: https://arxiv.org/abs/2610.00859
作者: Chensheng Peng,Wenhao Ding,Ran Tian,Zewei Zhou,Jef Packer,Maximilian Igl,Peter Karkus,Yan Wang,Masayoshi Tomizuka,Boris Ivanovic,Marco Pavone,Yuxiao Chen
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Project page: this https URL
Abstract:World action models (WAMs) jointly predict actions (intent) and visual future (foresight). Standard training adds noise to recorded actions and video simultaneously, but such training paradigms introduce a mismatch: perturbed actions imply counterfactual future visual, while the noised video remains tied to the GT recording. In low-noise regime, the scene geometry and even the dynamic behavior remain clearly visible from the noisy future frames despite the added noise. We present CtrlWAM, which executes perturbed actions in a simulator and pairs them with their noised visual consequences for joint WAM learning. To accommodate the different denoising requirements of video and actions, we introduce warped video–action noise schedules that aim to keep visual layout responsive as action predictions evolve. We further extend the action interface from ego-only control to a variable number of agent streams, allowing a unified model to represent predicted or commanded futures for multiple agents. Driving experiments show more accurate action forecasts, closer agreement between generated video and actions, and better following of supplied commands; robotics experiments show stronger motion fidelity and controllability. Matched controls support the benefit of off-path renders for command following and manipulation fidelity. Together, these findings contribute to a more controllable world action model. Project page: this https URL
[CV-140] Lang3DSeg: Annotation-Free Open-Vocabulary 3D Segmentation with Point Transformers ICRA2027
链接: https://arxiv.org/abs/2610.00855
作者: Cigdem Kokenoz,Amir Salarpour,Alkim Domeke,Christopher Salas,Pedram MohajerAnsari,Long Cheng,Mert D. Pesé,Bing Li
类目: Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)
备注: 9 pages, 3 figures, 4 tables. Submitted to IEEE ICRA 2027
Abstract:Accurate 3D semantic perception is critical for safe autonomous navigation. However, supervised LiDAR segmentation remains tied to closed taxonomies and to the cost of point-wise manual annotation. Open-vocabulary methods avoid that cost by projecting the output of 2D vision-language models onto LiDAR and distilling it into a 3D network. These methods rely almost exclusively on voxel-based sparse convolutions, and point transformers have so far been limited to indoor environments, where 3D data is dense and bounded. We present Lang3DSeg, which establishes a point transformer as the backbone for annotation-free open-vocabulary segmentation of outdoor 3D LiDAR, and is trained from scratch without geometric pre-training. This training paradigm necessitates addressing the inherent noise in 2D-to-3D label projections; specifically, naive projection often suffers from depth ambiguity, where points behind an object are erroneously assigned its semantic label. We therefore composite masks using an explicit class-priority rule and truncate each projected instance at the first gap in its depth distribution, correcting the projection error directly rather than averaging it over registered sequences. Lang3DSeg achieves 52.8% mIoU on nuScenes validation and 41.4% on SemanticKITTI, the highest among published annotation-free methods on both benchmarks. Every 3D semantic segmentation is on a single LiDAR sweep, and inference operates in real-time without running vision-language models.
[CV-141] SmoothOperator: Enhancing Representations for Fine-grained Open-set Recognition via Modulated Label Smoothing
链接: https://arxiv.org/abs/2610.00851
作者: Thiru Thillai Nadarasar Bahavan,Yu Xia,Sachith Seneviratne,Saman Halgamuge
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注:
Abstract:Open Set Recognition (OSR) aims to enable models to accurately classify known classes while rejecting samples from unseen classes. A key challenge in OSR lies in the inability to model the unbounded distribution of unknown classes during training, often leading to the misclassification of samples from these classes. Rather than modeling unknowns, recent work shapes the feature space so that known classes are compact and well separated, and spherical representation learning methods have achieved strong results this way. Label smoothing has been identified as one of the key drivers of this success, yet it applies the same coefficient to every training sample, regardless of how well each sample is already embedded. We show that the spherical representation learning objectives used in OSR share a single alignment–uniformity structure in which labels enter only through the alignment term. Label smoothing therefore acts as an alignment dial, and a fixed coefficient sets this dial to the same value for every sample. We propose a plug-in, SmoothOperator (SmoothOP), which sets the smoothing coefficient of each sample from its \textbfprominence, an embedding-space signal measuring how clearly the sample’s own class stands out against its strongest competing class. Our method integrates into four existing spherical representation learning methods at minimal training overhead. SmoothOP assigns strong smoothing to samples with high prominence, which reduces their alignment and relaxes their pull. On the Semantic Shift Benchmark, SmoothOP-augmented variants generally outperform their base objectives across datasets, degrees of semantic shift, and OSR post-processors, with gains of up to 4.7% in AUROC, OSCR, and closed-set accuracy.
[CV-142] Geometric Similarity in VLM Low-Level Vision Representations
链接: https://arxiv.org/abs/2610.00848
作者: Shao-Jun Xia,Huixin Zhang,Zhen Lei,Anlan Sun,Yuner Zhang,Xiaoyang Chen
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注: First version: 10 pages
Abstract:Vision-language models (VLMs) have emerged as powerful candidates for universal vision backbones, with representative architectures including autoregressive (AR) models and diffusion transformers (DiTs). Yet, adapting them efficiently for all-in-one low-level image restoration remains a challenge. Crucially, the field lacks an understanding of how VLMs organize hidden-layer representations and whether these structurally distinct paradigms share a common geometric organization for pixel-level perception. Such shared organization is a prerequisite for building highly transferable, unified restoration VLMs and adapters. In this paper, we systematically investigate representational similarity across 24 low-level tasks spanning 5 categories. We propose GeoSim, a unified four-level framework that analyzes task-conditioned representations from global similarity, local geometry, sparse feature decomposition, and topological verification perspectives. Our formulation applies to the analysis of hidden states in AR models and feature maps in DiTs across same- and cross-task/model settings. Our results reveal the organizing principles of low-level visual representations while exposing their limits in cross-task and cross-model agreement. Ultimately, GeoSim provides an interpretability lens for probing latent transferability in low-level vision and diagnosing model limitations in task- or model-specific scenarios.
[CV-143] Align Then Reason : A Multimodal Lip-Sync Judge for Dubbing
链接: https://arxiv.org/abs/2610.00825
作者: Rui Liu,Bhavin Jawade,Haoqi Li,Shivam Mehta,Karan Saxena,Yinghong Lan,Cameron R. Wolfe
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Dubbing quality control requires a reference-free judge that can determine whether a candidate text line matches a speaker’s visible articulation in both content and timing, using only silent video and text because dubbed audio may not yet exist. Existing visual speech recognizers and video-language models are poorly suited to this setting: even when fine-tuned to recover spoken content from lip motion, they remain largely insensitive to temporal errors. We introduce \textitAlign Then Reason (ATR), a multilingual lip-sync judge that first establishes a monotonic alignment between frame-level lip representations and the phonetic units of the candidate line, then reasons over this alignment to make the final judgment. An alignment scorer provides the LLM with both local evidence for each phonetic unit and a calibrated global alignment score, enabling it to reason jointly about content and timing. On a seven-language benchmark, our method improves mean AUC over the corresponding Qwen3.5 SFT baselines by 59.4%, 50.2%, and 50.8% with 2B, 4B, and 9B reasoners, respectively. The gains generalize across LLM families, reaching mean AUC improvements of 45.9% and 46.6% over the best baseline for LLaMA-3.1-8B and Mistral-7B, respectively. They also transfer across datasets to three unseen MuAViC languages. Furthermore, we evaluate on two downstream tasks built from real dubbing lines. On dub-line reranking, ATR-9B outperforms the best lip-reading baseline by 52.0%, while on script-to-clip assignment, ATR-9B improves over the best lip-reading baseline by 17.7%.
[CV-144] Video Generation Models: A Survey of Post-Training and Alignment
链接: https://arxiv.org/abs/2610.00812
作者: Chaoyu Li,Xiaoyi Gu,Yogesh Kulkarni,Eun Woo Im,Mohammadmahdi Honarmand,Zeyu Wang,Juntong Song,Fei Du,Xilin Jiang,Kexin Zheng,Tianzhi Li,Fei Tao,Pooyan Fazli
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: Published in Transactions on Machine Learning Research (TMLR), 2026. Project page: this https URL
Abstract:Video generation has rapidly progressed from short, low-quality clips to high-resolution, long-duration sequences with complex spatiotemporal dynamics. Despite strong generative priors learned through large-scale pretraining, pretrained video models often fail to reliably follow human intent, maintain temporal coherence, or satisfy physical and safety constraints. Compared with image and text generation, alignment in video generation presents unique challenges, including error accumulation over time, motion-appearance coupling, multi-objective trade-offs, and limited supervision for temporal properties. These challenges motivate systematic post-training strategies that adapt pretrained models without retraining them from scratch. In this survey, we present the first comprehensive review of post-training and alignment in video generation models. We frame post-training as a unifying framework and distinguish between implicit alignment and explicit alignment based on how alignment signals are enforced. From this perspective, we organize existing approaches into four broad categories: supervised fine-tuning methods, self-training and distillation methods, preference- and reward-based methods, and inference-time methods. This taxonomy provides a coherent view of how alignment signals shape model behavior across both training and deployment. Beyond methodological advances, we review commonly used datasets, benchmarks, and evaluation practices, and discuss open challenges such as scalable reward design, long-horizon temporal consistency, stability-expressiveness trade-offs, and safety-aware generation. This survey aims to provide a structured conceptual foundation and practical guidance for advancing controllable and reliable video generation models.
[CV-145] VTV-FM: Flow Matching through Variational Terminal-Velocity Closure NEURIPS2026
链接: https://arxiv.org/abs/2610.00785
作者: Haoyang Jiang,Yuheng Li,Di Yang,Yanhai Xiong,Haipeng Chen,Yi He
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted at NeurIPS 2026. Code: this https URL
Abstract:Flow matching (FM) learns generative transport by fitting continuous-time motion from a simple source distribution to the data distribution. Most existing methods use first-order bridges: once a source and a target sample are paired, the path is a straight motion with constant velocity. FM with optimal transport (OT) improves the pairing, but the bridge itself remains linear, limiting its ability to model curved motion, acceleration, and changing directions. A natural remedy is to use second-order phase-space dynamics; however, learning the bridge requires target-side terminal-velocity information that static datasets do not provide. We propose Variational Terminal-Velocity Flow Matching (VTV-FM), a second-order FM framework that derives the missing velocity by minimizing acceleration energy, yielding a closed-form closure for static data. The same minimum-acceleration variational construction also defines the OT pairing cost and the acceleration targets used for training. Experiments on low-dimensional datasets, PDE-governed physical fields, and CIFAR-10 show that VTV-FM improves transport geometry and generation quality over first-order and high-order FM baselines.
[CV-146] Video Evidence Indexing: Learning Where to Look from Video Previews for Token-Budgeted Long-Video Question Answering
链接: https://arxiv.org/abs/2610.00757
作者: Haowen Guan,Shengzhi Li,Shichao Pei
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Long-video question answering is limited by the high cost of visual tokens and by the fixed context width of current VLMs. A long-video question may require broad temporal coverage, but the answer is often supported by only a compact set of moments. To locate these moments efficiently, we propose token-budgeted Video Evidence Indexing (VEI): given a dense low-resolution Video Preview, the model constructs a compact high-resolution Evidence Set for final reasoning. We treat VEI as a policy that must jointly solve \textitevidence localization, which finds question-relevant moments, and \textitbudget planning, which decides where to spend the limited high-resolution frame budget. We implement this idea with an inference pipeline: the Video Preview provides cheap global coverage, Video Evidence Indexing constructs the Evidence Set, and Answer Generation combines both inputs for final VQA. To address missing frame-level supervision, we adopt privileged self-distillation, where an answer-aware teacher guides the normal test-time policy on student-generated indexing traces. We explore previews at 1, 6, 12, and 24 visual tokens per frame, training a single policy that supports all four resolutions. Experiments show that Video Evidence Indexing improves accuracy under limited visual budgets, and self-distillation further improves both QA accuracy and temporal evidence localization.
[CV-147] Increasing Width Allows Greedy Layer-wise Training to Rival End-to-End Backpropagation in Self-Supervised Learning
链接: https://arxiv.org/abs/2610.00753
作者: Syon Mansur,Joel Zylberberg
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Neural and Evolutionary Computing (cs.NE); Neurons and Cognition (q-bio.NC)
备注: 10 pages, 5 figures
Abstract:End-to-end backpropagation has been the dominant mode of training in deep learning, allowing for the coordination of parameter updates across layers of a neural network. Prior studies have explored alternative – and, in some cases, simpler – training mechanisms, showing that they can sometimes achieve performance similar to backpropagation. However, the architectural conditions under which locally optimized networks, which avoid end-to-end backpropagation of error, can learn representations comparable to those learned through end-to-end training remain unclear. We aim to answer this question in the context of self-supervised learning, an important framework for large-scale pretraining in artificial intelligence. Here, we investigate how network width and depth affect the efficacy of greedy layer-wise and end-to-end self-supervised training in convolutional networks. We find that in wider networks, the benefits of end-to-end backpropagation over greedy layer-wise training shrink: in relatively shallow and very wide networks, we even observed higher performance in models trained with greedy layer-wise training. Subsequent analysis of the representations formed by these networks shows that very wide greedy-trained networks exhibit more favorable representational geometry than do networks trained end-to-end with backpropagation. This work shows that width can compensate for restricted credit assignment and identifies differences in representational geometry as a potential mechanism for their improved performance.
[CV-148] Signal-Noise Factorization Isolates Nuisance Variation into Removable Subspaces
链接: https://arxiv.org/abs/2610.00751
作者: Sakin Kirti,Joel Zylberberg
类目: Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (stat.ML)
备注:
Abstract:Recent theoretical work identified fundamental properties of representation geometry that shape inference ability of deep neural networks. These include signal-noise factorization (SNF), the ability to segregate signal from noise, and signal-signal factorization (SSF), the ability to segregate task-specific and task-irrelevant signals. Here, we built regularizers that reinforce these two properties during training. We compared networks trained with these regularizers to L_2 -regularized baseline networks on the CIFAR-100 classification task to understand how our regularizers shape representation geometry and impact performance on a well-known computer vision baseline. Enhancing SNF via regularization improved model performance but enhancing SSF did not. Motivated by biomedical applications, we investigated how our regularizers affected performance on the BloodMNIST dataset treated with MedMNIST-C corruptions at five severity levels, and found even larger performance gains using the SNF regularizer. To understand the mechanism by which SNF-regularization produces improved performance, we analyzed the nuisance subspaces across regularization regimes, finding that the SNF-regularized models represent noise in distinct subspaces, separate from class-relevant signal. Because this geometry is explicit, the dominant corruption-induced directions can be estimated on held-out data and projected out of the representations. This manipulation led to a substantial gain in accuracy. These results show that regularizers that enforce signal-noise factorization can produce substantial improvements on computer vision tasks that contain out-of-distribution image distortions at inference time. They also highlight how shaping representations affects model performance: isolating nuisance variables from categorical ones is more important than maintaining factorized representations of categorical variables.
[CV-149] What Builds the Scene? Luminance Dominates Geometry Formation in 3D Gaussian Splatting
链接: https://arxiv.org/abs/2610.00749
作者: Rezvan Joshaghani,Steven Cutchin
类目: Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR)
备注: 28 pages, 6 figures, 7 tables
Abstract:Standard 3D Gaussian Splatting (3DGS) learns geometry and appearance jointly from RGB supervision, making it difficult to isolate how luminance and chroma contribute to the learned representation. We study this by training models under different channel supervision, freezing their non-appearance parameters (position, scale, rotation, and opacity), and re-estimating appearance with the same solver before comparing held-out reconstruction. Across eleven benchmark scenes with four independent runs each, geometry learned from luminance alone supports held-out reconstruction 0.085 dB below RGB-trained geometry on average. If chroma is deleted from a trained model, a sufficiently expressive solver can re-fit it on the frozen geometry to the original quality or slightly better. Higher-order spherical harmonics contribute much more reconstruction quality to luminance than to chroma, improving PSNR by 1.44 dB versus 0.19 dB on average, although on mirror-like surfaces hue does still change with viewpoint. The luminance advantage is even larger when geometry is being formed. Chroma-only supervision produces geometry 3.9-5.5 dB worse than luminance-only supervision after the same appearance solve; densification explains part of this gap. Overall, geometry formation in standard 3DGS is strongly luminance-dominated but not luminance-exclusive, and much of the chromatic appearance can be recovered after spatial support has formed.
[CV-150] Personalized Image Generation with Reasoning and Reflection
链接: https://arxiv.org/abs/2610.00737
作者: Bo Ni,Ngoc N. Tran,Qinwen Ge,Franck Dernoncourt,Seunghyun Yoon,Samyadeep Basu,Sungchul Kim,Puneet Mathur,Nedim Lipka,Tong Yu,Yu Wang,Ryan A. Rossi,Tyler Derr
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:
Abstract:Personalized image generation has remained narrowly focused on conditional synthesis from curated visual exemplars, rather than capturing who a user is. In practice, however, a user’s personal context is much richer, comprising reviews, posts, images, captions, and metadata accumulated over time. A truly personalized generator should leverage this history to produce images aligned with the user’s lifestyle and aesthetic preferences. To this end, we introduce the first unified benchmark for personalized image generation from user histories. The benchmark comprises two complementary tasks and a multi-axis evaluation protocol that assesses target fidelity, visual quality, user distinguishability, semantic alignment with the user’s history, and task-specific utility. Grounded in real-world e-commerce and social media settings, the benchmark includes: (1) Personalized Scene Generation, which places a given object in a scene that reflects a user’s preferences and lifestyle, motivated by personalized product presentation; and (2) Personalized Creative Generation, which generates a novel image on a specified topic that is faithful to a user’s aesthetic and visual identity, motivated by social media content creation. We further propose PEARL, which couples a multimodal reasoner with a frozen image generator in an interleaved reason-reflect loop optimized with differential data reward. Across both tasks, PEARL outperforms strong baselines, achieving an average improvement of 15% across personalization metrics.
[CV-151] FedMAD: Modulation-Aware Directional Aggregation for Federated Learning in Remote Sensing Image Classification
链接: https://arxiv.org/abs/2610.00693
作者: Barış Büyüktaş,Begüm Demir
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注:
Abstract:Federated learning (FL) has recently attracted increasing attention in remote sensing (RS) since it enables collaborative model training across decentralized RS image archives without requiring direct access to local data. However, FL performance significantly degrades when the data distributions between clients are heterogeneous, which often occurs due to geographical differences, seasonal changes, and varying image acquisition and atmospheric conditions. To address this challenge, in this letter, we propose a novel personalized FL framework (denoted as FedMAD) for RS image classification problems. The proposed framework separates globally shared representation parameters from client-specific adaptation parameters to preserve client-specific features while maintaining globally transferable representations. This is achieved by integrating lightweight modulation modules and local batch normalization layers into the backbone network. Although globally shared parameters are collaboratively optimized between clients, client-specific parameters remain local to preserve domain-specific feature characteristics. In addition, FedMAD introduces a modulation-aware directional aggregation strategy that dynamically adjusts the importance of aggregation for each client according to the alignment of local modulation updates. This allows the global optimization process to suppress conflicting client updates originating from heterogeneous data distributions while enhancing the contribution of clients with consistent adaptation behaviors. The experimental results obtained on the BigEarthNet-S2 and EuroSAT datasets demonstrate the effectiveness of FedMAD compared to state-of-the-art FL algorithms under heterogeneous RS data distributions. The code of the proposed framework will be publicly available at this https URL.
[CV-152] Soundwich: Video Generation with Layered and Controllable Audio
链接: https://arxiv.org/abs/2610.00691
作者: Zhuo Ning,AmirHossein Naghi Razlighi,Sagi Polaczek,Daniel Cohen-Or,Ali Mahdavi-Amiri
类目: Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM); Sound (cs.SD)
备注: 35 pages. Code: this https URL
Abstract:Recent joint audio-video generative models can synthesize realistic videos with synchronized sound, but typically generate audio as a single mixed track. This limits source-level control and differs from practical audiovisual workflows, where speech, music, sound effects, and ambient sounds are represented as separate editable tracks. We introduce Soundwich, a training-free framework that transforms a frozen joint audio-video flow-matching model into a generator of multiple synchronized, independently editable audio stems coupled to a shared video. Soundwich generates separate audio stems with explicit control over their temporal activity. To keep separately generated sounds coherent, we introduce a shared scene representation that communicates global audiovisual context across stems while preserving their source-level separation. We further route cross-modal interactions between each audio stem and its corresponding visual source, improving audiovisual consistency. The resulting stems remain synchronized with the video and can be independently retimed, muted, replaced, or remixed. Experiments and human evaluations show improved temporal control, source separation, and naturalness, while enabling flexible source-level editing within coherent audiovisual generation. Code is available at this https URL.
[CV-153] SemanTok: Predictable Semantic Tokens for Efficient Autoregressive Video Generation
链接: https://arxiv.org/abs/2610.00686
作者: Mikhail Dereviannykh,Vikram Voleti,Simon Donne,Mallikarjun Byrasandra Ramalinga Reddy,Shimon Vainer,Mark Boss
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注: 29 pages, 22 figures, including references and appendix; 9 pages of main text
Abstract:Recent video-based world models pair the scalability of autoregressive (AR) prediction with the visual quality of diffusion models. The choice of scene tokenizer is paramount for the optimal performance of each of these, both in terms of fidelity and semantics. Flexible-length, coarse-to-fine tokenizers yield exactly that: the first coarse tokens carry the clip’s global semantics while later tokens further specify details. Existing flexible tokenizers only apply a representation-alignment (REPA) loss on early decoder hidden states, a target the decoder can partly meet from its noised input instead. We introduce SemanTok, a flexible video tokenizer that feeds frozen DINO features into its encoder and adds lightweight heads that reconstruct them from each retained token prefix alone. SemanTok achieves high semantic alignment and video fidelity at every AR model size: a 201M SemanTok AR model matches or beats a VideoFlexTok AR model 3.4\times its size, and larger SemanTok AR models further improve fidelity. It keeps semantic alignment on out-of-distribution classes and gives the decoder higher semantic alignment at every noise level, including pure noise. It performs well in both reconstruction and generation, and its short token prefixes are cheaper to predict and give better generation fidelity, with pixel detail deferred to later tokens.
[CV-154] Curvature Under Attack in hZACH-ViT: Gauge Symmetry Boundary Saturation and Adversarial Failure NEURIPS2026
链接: https://arxiv.org/abs/2610.00680
作者: Athanasios Angelakis,Marta Gomez-Barrero
类目: Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)
备注: 12 pages, 3 figures, 4 tables. Accepted at NeurReps 2026: Symmetry and Geometry in Neural Representations, NeurIPS 2026
Abstract:Curvature is often treated as an intrinsic property of a representation, although its empirical effect also depends on coordinate scale, learned logit temperature, and numerical safeguards. We study this interaction in hZACH-ViT, a compact Vision Transformer with Euclidean, Poincare, and spherical prototype heads. The backbone architecture, seed-specific initialization, 50-per-class training subset, and optimization protocol are matched across three MedMNIST datasets and five seeds. At the fixed comparison curvature c=1 , Poincare has the lowest class-macro PGD attack-success rate in all 12 dataset-budget cells and under a stronger CE+DLR multi-restart attack on all three datasets, but it also has the lowest clean MacroF1. An end-to-end curvature intervention changes the interpretation. Reducing Poincare curvature to c=0.1 improves clean MacroF1 in every one of the 15 paired seed-dataset comparisons and removes hard boundary clipping, yet on OrganAMNIST it increases strong attack success from 89.7% to 99.3% (paired difference +9.57 points; 95% hierarchical bootstrap CI [+5.52,+14.03] ). At c=1 , 40 - 47% of clean Poincare features are hard-clipped, the radial Jacobian of the inherited map is nearly zero, and dimensionless attack trajectories are unusually long and inefficient. The spherical head provides a control: its curvature change is an exact scale gauge to floating-point precision and produces much smaller attack differences. These results do not establish intrinsic hyperbolic robustness. They identify an implementation-sensitive regime in which curvature, scale, and proximity to the Poincare boundary jointly organize clean recognition and adversarial representation motion.
[CV-155] Harnessing Vision-Language Models for Perceptual Quality Assessment and Autonomous Content Adjustment in Augmented Reality
链接: https://arxiv.org/abs/2610.00677
作者: Elias Rotondo(1),Lin Duan(1),Yanming Xiu(1),Sangjun Eom(1),Conrad Li(1),Maria Gorlatova(1) ((1) Duke University)
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: To be published in VRST 2026. Main Manuscript: 12 pages, 5 figures; Supplemental Materials: 7 pages, 10 figures. The accompanying public repository can be accessed by visiting this https URL
Abstract:Advancements in augmented reality (AR) continue to foster innovative solutions, facilitating novel methodologies within educational systems, healthcare delivery, and risk-mitigation protocols. However, optimizing for end-user immersion and comfort remains challenging, as AR head-mounted displays contend with constrained scene geometry, spatial jitter, and temporal instability. User studies are the standard AR evaluation method for visual quality, but their cost, diminishing scalability, and inflexibility pose bottlenecks during iterative application design. To address this problem, we present an automated framework for AR content evaluation and refinement, built on vision-language models (VLMs), to evaluate and predict the visual fidelity of AR scenes as perceived by users. First, we introduce RateAR, a benchmark of AR images and videos collected across diverse scenes and environmental conditions, with good-to-excellent reliability (ICC(2,5) = .90) across perceptual factors, including object placement, scale, and shadow consistency. Subsequently, we evaluate eleven commercial VLMs on the crafted benchmark. Results support that VLM-based quality predictions strongly correlate with human subjective judgments, achieving Spearman’s rank-order correlations of up to 0.8695. An ablation study further suggests that, compared to other prompting strategies, our contextual prompting yields better alignment with human ratings while balancing introduced complexity cues. Building on these findings, we construct an automated AR content adjustment system and conduct a 21-participant user study. More than 90% of participants found that the system improved placement and size coherence of virtual content.
[CV-156] VisionQ: VLM-as-a-Judge Taxonomy Dataset and Benchmark for Qualitative Analysis in Computer Vision
链接: https://arxiv.org/abs/2610.00666
作者: Vu Dinh Xuan,Duc-Hai Nguyen,Minh-Dung Dao,Vu Quynh Giao,Quang Hong Nguyen,Binh-Son Hua,Barry O’Sullivan,David Murphy,Hoang D. Nguyen
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注: 29 pages, 18 figures, 6 tables. Code: this https URL
Abstract:Qualitative comparison figures are central evidence in computer vision papers, and vision-language models (VLMs) are increasingly used to judge them. Yet existing benchmarks score only scalar quality or overall preference, so a judge can be rewarded for picking the preferred image for the wrong visual reason. We introduce VisionQ, the first benchmark built from peer-reviewed CV comparison figures that grounds every judgment in a named visual criterion: each question states the criterion, and a judge is credited only when it selects the output the authors identify as best on that criterion. We call this task criterion-conditioned visual discrimination. VisionQ comprises (1) a corpus of 1,409 CVPR and ICCV papers with 1,800+ validated comparison figures and 3,911 hand-annotated data points linking method crops to author-stated visual claims; (2) a six-axis, 51-leaf taxonomy of the visual criteria behind qualitative judgment; (3) a criterion-conditioned evaluation protocol that hides method names, captions, and paper identity and reports accuracy per criterion; and (4) VisionQ-Judge, a DPO-tuned Gemma-4-E4B judge trained on symmetric evidence pairs, which reduces last-option predictions by 7.0pp and improves accuracy by 2.5pp on a held-out test set. Evaluating 20 open- and closed-source VLM judges, we find that the strongest reach only 63.1% accuracy (chance 32.2%) and that reliability varies sharply across criteria. Code: this https URL. Data: this https URL.
[CV-157] HAWK: Rethinking Multimodal Drafting for Speculative Decoding
链接: https://arxiv.org/abs/2610.00623
作者: Wenhan Yang,Anirudh Rao,Ashwin Chandra
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Speculative decoding has achieved substantial lossless speedups for LLMs, but remains less effective for large vision-language models (LVLMs), where lightweight drafters struggle to use rich multimodal information. A second limitation is that standard distillation supervises the drafter only along the original training trajectory, without modeling how target predictions shift after the drafter’s own proposals. As drafting moves away from this trajectory, the drafter can increasingly disagree with the target, reducing acceptance in later steps. We propose HAWK to address both limitations. HAWK uses representation similarity to select informative target layers and learns how to combine their hidden states. For visual information, it directly provides the drafter with compressed visual hidden states from the target model instead of raw visual tokens, making the visual information easier for a shallow drafter to use. HAWK also trains the drafter to capture how target predictions change after its own proposals, improving its agreement with the target during multi-step drafting. On SmolVLM-256M across ten multimodal benchmarks, HAWK raises average acceptance length from 3.32 to 4.08 and speedup from 2.19x to 2.60x over EAGLE-3 under greedy decoding, and from 2.89 to 3.41 and 1.92x to 2.19x under sampling.
[CV-158] Just Align bmx: Aligning Predictions Not Representations
链接: https://arxiv.org/abs/2610.00600
作者: Yuyao Zhang,Yuwei Hu,Ziyang Mai,Yu-Wing Tai
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Representation alignment has become an effective way to accelerate diffusion training, but its benefits do not transfer reliably to pixel-space clean-image prediction. In JiT, we find that auxiliary feature alignment can improve access to semantic features while reducing access to image variation needed for clean-image prediction, creating a mismatch between the auxiliary objective and the denoising task. This suggests a different principle: auxiliary supervision should improve the prediction target itself rather than impose a separate representation target. We introduce JAx (Just Align x), a prediction-supervision method that aligns clean-image predictions across noise levels. JAx couples a noisier student observation with a cleaner observation through a Markov degradation that preserves the original JiT input distribution. Under this coupling, the oracle prediction from the cleaner state has the same conditional mean as the optimal JiT target, while its conditional target covariance is no greater. Thus, oracle prediction alignment preserves the population JiT objective up to a constant while providing a lower-variance training target. To make this construction practical with an imperfect EMA teacher, JAx combines ground-truth supervision with a reliability-gated coupling band that selects nearby teacher states based on prediction risk. On ImageNet 256x256, JAx consistently improves FID and accelerates convergence across JiT-B/16, L/16, and H/16, without an external encoder or changes to the architecture or sampling procedure. Gradient diagnostics further show reduced minibatch gradient variance, while ablations demonstrate that the gains cannot be explained by time reweighting alone. These results show that prediction-space supervision provides a simple and principled alternative to representation alignment for pixel-space generative models.
[CV-159] Right In-Place (RiP) Convolution: A Simple General and Near-Optimal Strategy for Memory-Efficient CNN Inference NEURIPS2026
链接: https://arxiv.org/abs/2610.00586
作者: Opegbemi Matthias Busoye,Tolulope Matthew Busoye,Eghonghon-aye Eigbe
类目: Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)
备注: Extended version of a paper accepted at the NeurIPS 2026 Workshop on Global South in AI
Abstract:Activation memory, not compute, limits CNN inference on constrained hardware such as microcontrollers. Direct in-place convolution removes the dual-buffer cost, but the memory-optimal formulation of Gural and Murmann assumes valid padding, unit stride, unit dilation, and odd square kernels, and needs a non-sequential traversal costing 2\times inference time in transposes. We identify two regimes in which their published closed form does not hold: (1) an under-allocation of exactly (k-1)C_in \bmod (C_out-C_in) scalars, active on every convolutional layer of their own deployed network and manifesting as a silent corruption of still-live input; (2) an unbounded overestimate, up to 2,432\times , once the critical leg leaves the output grid. We correct both and generalize to arbitrary stride, dilation, padding, and rectangular kernels. We then propose Right In-Place (RiP) convolution, a bit-identical operation in which every layer reads its input right-aligned in a shared workspace and writes its output left-aligned from index zero. The debt is piecewise affine in the output pixel index, so evaluating its breakpoints in O(1) yields the minimum safe gap without enumerating the output grid, with row-major access preserved. Across 10,000 random layers RiP produced no corruption, and across 84 convolutional layers from 25 architectures it matches the herringbone workspace exactly on 58 and within 5% on 81, using 24.8% less memory than dual buffering on average. Written into TinyEngine’s kernels and deployed to a Raspberry Pi Pico 1 and Pico 2, it cuts peak activation memory across eleven MCUNet models by 12.5 to 33.3% at unchanged cycle counts and bit-identical outputs, raising the number of models that fit the Pico 1’s 256 KB SRAM from six to nine.
[CV-160] Discrete Annotation Continuous Preference: Rethinking Supervision for Accurate and Generalizable Aesthetic Image Cropping
链接: https://arxiv.org/abs/2610.00582
作者: Ziqing Zhang,Xiao Liu,Kai Liu,Jianze Li,Weihang Zhang,Linghe Kong,Yulun Zhang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Code, model, and data are available at this https URL
Abstract:Aesthetic image cropping aims to identify the optimal crop of an image in terms of aesthetics and composition. While supervision based on annotated data is fundamental, the field has been hindered by a long-standing problem: existing datasets suffer from (1) human subjectivity and (2) rigid discreteness confined to fixed sampling grids. These flawed annotations not only limit the accuracy and generalization of trained models but also severely distort fair evaluation. To overcome this, we propose to model human cropping preference as a multi-peaked, continuous, and sharp field over the crop space. We introduce the Continuous Preference Field (CPF), which recovers a dense preference landscape from discrete annotations through (1) peak clustering, (2) off-lattice refinement, (3) negative shaping, and (4) field assembly. Based on this, we train CPIC, a VLM-based cropping model optimized via GRPO with the CPF reward, which overcomes template collapse, achieving state-of-the-art performance and exceptional out-of-domain generalization. Finally, to resolve the long-standing benchmark evaluation crisis, we introduce CPICD, a comprehensive recalibration of existing ground-truth boxes. By leveraging the CPF to correct grid-bound artifacts across mainstream benchmarks, CPICD establishes a rigorous and reliable foundation for future cropping research. Extensive experiments and user studies demonstrate the superiority of our CPF, CPIC, and CPICD. Code, model, and data are available at this https URL.
[CV-161] Gestalt: Large Multimodal Interplay Model
链接: https://arxiv.org/abs/2610.00576
作者: Zequn Yang,Yu Miao,Haotian Ni,Ziheng Chen,Chengxiang Huang,Dongzhan Zhou,Kai Chen,Qi Zhang,Ji-Rong Wen,Yake Wei,Di Hu
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 17 pages, 7 figures
Abstract:In this paper, we propose Gestalt, a new paradigm of large multimodal model built around multimodal interplay. Despite rapid advances, large multimodal models are reaching a bottleneck: existing approaches focus primarily on accommodating additional modalities while overlooking the distinct characteristics of each modality and the relations among them. Motivated by the multistage property of human multisensory perception, we propose a multimodal interplay pyramid that organizes multimodal modeling as a progression from modality-specific processing, through cross-modal alignment, to deeper multimodal integration. Guided by this pyramid, Gestalt adopts a unified discrete diffusion framework and an interplay-partitioned architecture, with learnable interplay tokens mediating cross-modal exchange and integration. The pyramid also structures its data organization and training strategy. Strong performance across image generation, multimodal understanding, and text-only evaluation shows that Gestalt significantly improves cross-modal integration while preserving modality-specific information, effectively harnessing the strengths of diffusion-based multimodal models and offering a promising path toward unified multimodal intelligence.
[CV-162] FORTE: Adaptive Scoring and Exact Keyframe Selection for Long-Video Question Answering
链接: https://arxiv.org/abs/2610.00573
作者: Haifeng Huang,Biyin Xu,Chunsheng Xin,Yang Li
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Query-aware keyframe selection enables multimodal large language models (MLLMs) to process long videos using only a small set of question-relevant frames. Existing score-based methods, however, typically search within a fixed, uniformly sampled candidate pool, preventing evidence outside this pool from ever being selected. Given a limited relevance-scoring budget, the key challenge is to allocate evaluations adaptively to promising frames while continuing to explore underrepresented temporal regions. We introduce FORTE, a training-free framework that addresses this challenge through two stages: adaptive relevance scoring and global keyframe optimization. Starting from sparse, uniformly distributed observations, our efficient Gaussian-process relevance predictor estimates relevance for unscored frames, exploiting temporal locality and the approximately banded kernel structure to reduce the core computation from cubic to linear time in the number of frames for fixed bandwidth. The scoring stage then selects which frames to score next by balancing predicted relevance with temporal coverage, prioritizing promising regions while also exploring less-represented parts of the video. The optimization stage selects the final keyframes by maximizing an objective that jointly captures measured relevance and temporal coverage. We derive an exact algorithm that leverages the logarithmic coverage structure to identify the optimal subset of the scored candidate pool in time linear in the pool size, for a fixed final-frame budget. Experiments on four long-video question-answering benchmarks show that FORTE achieves the highest observed mean accuracy among the compared selectors under every tested scoring budget. Further evaluations demonstrate its consistent effectiveness across different relevance scorers and downstream MLLMs.
[CV-163] PhysVista: Benchmarking Physical Intelligence in VLMs via a Perception-Reasoning -Assessment Loop NEURIPS2026
链接: https://arxiv.org/abs/2610.00559
作者: Xinge Peng,Yiting Lu,Tianwu Zhi,Wen Wen,Jianzhao Liu,Xin Li,Zhibo Chen
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted at NeurIPS 2026 (Main Track)
Abstract:Vision-Language Models (VLMs) have shown strong multimodal reasoning capabilities, yet whether they truly capture the physical consistency underlying real-world dynamics remains unclear. Existing benchmark paradigms often suffer from fragmented evaluation, focusing on isolated cognitive stages while overlooking the inherent synergy between perception, reasoning, and physical judgment. The lack of a holistic perspective limits the ability to diagnose whether VLMs can reliably evaluate the physical authenticity of emerging generative models. To address these issues, we introduce PhysVista, a benchmark designed to evaluate physical intelligence in VLMs through a closed cognitive loop framework inspired by the human seeing-reasoning-assessment process. PhysVista restores this loop by jointly evaluating physical state perception, physical dynamics reasoning, and physical plausibility assessment. It further distinguishes event-level reasoning and scale-level reasoning to enable fine-grained analysis of physical understanding. In addition, PhysVista incorporates both real-world and AI-generated videos, allowing evaluation across diverse domains and emerging generative scenarios. Extensive experiments across a diverse set of VLMs reveal substantial limitations in physical reasoning and plausibility assessment, highlighting a persistent gap between visual recognition and genuine physical understanding, and pointing toward more principled designs for physically grounded multimodal intelligence.
[CV-164] Memorizon: Training World Models Beyond Their Context Window
链接: https://arxiv.org/abs/2610.00544
作者: Tingting Liao,Xuezhi Liang,Hao Li,Guangyi Liu
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Project page: this https URL
Abstract:Streaming world models should render a place consistently across repeated visits. Directly supervising such revisits requires training samples that capture both visits, often spanning minutes. Yet dense attention over the full span incurs quadratic costs, making long-span supervision expensive. Memorizon breaks this coupling: long spans are needed for supervision, but not for attention, since the two visits can share a forward pass without including every intervening frame. A training sample covers a span of any length but is scored only on its last k chunks. Instead of tokenizing the history before them, each scored chunk retrieves its own top- K latents by camera co-visibility, and the union of these requests forms a shared bank. The bank is bounded by kK , so the sequence stays bounded however long the span; at the shortest span the recipe is exactly conventional training. Adding the bank raises the cost of a step once; beyond that, a longer span costs little, and going from 100 to 400 s adds 12% to the step time. Against a sliding-window baseline, retrieval raises revisit consistency on every split, and a span long enough to reach the first visit of each return adds a further 24% to 30%, at some cost in image quality; beyond that span, more length no longer helps. Filling the bank from another episode lowers revisit correlation by 83%, so the model uses what it retrieves. Project page: this https URL
[CV-165] PixelDense: Dense Prediction as Representation Alignment for Pixel Diffusion NEURIPS2026
链接: https://arxiv.org/abs/2610.00483
作者: Lehan Yang,Daiqing Qi,Wenhao Zhang,Avery Li,Yiqing Yang,Yifan Li,Yu Kong,Haitian Zheng,Zhifei Zhang,Zhe Lin,Varun Jampani,Sheng Li
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: NeurIPS 2026
Abstract:Representation alignment (REPA) accelerates diffusion transformer training, but its alignment targets are almost exclusively semantic encoders such as DINOv2 and CLIP. Recent analysis points to spatial structure, not global semantics, as the carrier of the alignment effect, yet dense-prediction foundation models trained to predict that structure remain overlooked as REPA targets. In pixel-space diffusion, SAM2, Depth Anything v2, and Metric3D v2 each outperform the DINOv2-only GenEval baseline, with the two geometric teachers leading the segmentation teacher. A flat sum of all four teachers, however, lands below the best single geometric teacher, as semantic and geometric gradients compete for one denoiser projection. We introduce PixelDense, which routes DINOv2 and SAM2 through a semantic projection stream, routes Depth Anything v2 and Metric3D v2 through a geometric projection stream, and adds a weight-space orthogonality penalty that keeps the two streams in disjoint subspaces. All four teachers are frozen during training and dropped at inference. Applied to PixelGen and DeCo with a single recipe, PixelDense improves GenEval, DPG-Bench, and HPS v2.1, raises PixelGen-XXL’s GenEval Overall from 0.7927 to 0.8093, and beats every single-teacher and unfactored multi-teacher variant. In partial-noise reconstruction, independent panoptic, depth, and surface-normal probes show up to 53.1% PQ gain and 36.0% depth AbsRel reduction at \tau=0.5 across COCO and Flickr30K. From random initialization, PixelDense also reaches the baseline’s peak GenEval 1.23x faster. In SDEdit editing on PIE-Bench, PixelDense keeps more of the source background and layout at every edit strength, raising background PSNR by up to 2.2 dB.
[CV-166] PACT: End-to-End Learning of Human Pose Contacts and Forces from Video
链接: https://arxiv.org/abs/2610.00451
作者: Rikhat Akizhanov(1),Yangsong Zhang(1),Nikolai Kaliazin(1),Peter Wolf(2),Yoshihiko Nakamura(1),Pascal Fua(3),Fabio Pizzati(1),Ivan Laptev(1) ((1) MBZUAI, (2) ETH Zürich, (3) EPFL)
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: 31 pages, 12 figures. Project page: this https URL
Abstract:Human motion, environmental contacts, and interaction forces are governed by common physical laws, yet existing approaches typically separate visual pose reconstruction from contact and force estimation. This separation limits joint reasoning and can propagate errors between stages. We introduce PACT, an end-to-end model that jointly learns to estimate human pose, contacts and contact forces from monocular video. Our approach augments a human reconstruction foundation model with learnable contact-force tokens and a temporal transformer that integrates visual features with world-space motion. Joint prediction heads refine human poses and estimate contacts and forces, while physics-based supervision encourages consistency between the reconstructed motion and interaction forces. To address the scarcity of force annotations, we develop a data annotation pipeline that combines contact labeling with physics-based motion and force optimization, producing training supervision from synthetic and real-world videos. We also introduce a real-world climbing benchmark ForceWall with climbing videos and corresponding ground-truth contact forces obtained from the force sensors. Experiments demonstrate state-of-the-art contact and force estimation, outperforming staged reconstruction approaches and generalizing to interactions beyond the training distribution. These results support end-to-end joint learning as an effective approach to recovering human motion and physical interactions from video.
[CV-167] Scores That Hold Benchmarks That Leak: Measuring Dataset Contamination in Public Brain-Tumor MRI Classification
链接: https://arxiv.org/abs/2610.00421
作者: Bhanu Prakash Vangala,Sowmya Guda,Latha Peddi,Navya Vangala
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:
Abstract:Automated classification of brain tumors from MRI is a heavily published application of deep learning in medical imaging, with reported accuracies on public benchmarks routinely exceeding 98%. However, accuracy does not capture a critical dimension of benchmark quality: dataset integrity, defined as the independence of test from training data at the image, patient, and acquisition-source levels. We introduce a three-layer contamination framework comprising duplicate, patient, and source-label leakage to assess the public corpora on which this literature rests. We audit the three most widely used corpora against a chest-radiograph negative control and quantify each layer’s effect on measured performance across nine architectures and three evaluation conditions. Contamination is severe at every layer: 28.8% of the dominant corpus’s official test split has a near-twin in its own training split, a second corpus leaks 22.3% of its test images byte-identically, 95.5% of traceable test images share a patient with training, and file-header features containing no anatomy separate tumor from no-tumor at 0.959 balanced accuracy, at parity with fine-tuned ResNet backbones. The unexpected result is that removing every identified leaked test image leaves balanced accuracy essentially unchanged: stable performance after deduplication does not establish benchmark integrity. Our findings establish dataset integrity as a distinct, measurable axis of benchmark quality that a stable leaderboard cannot certify. For biomedical research, reported accuracy on these corpora alone does not establish that a model has learned to recognize tumors rather than exploit dataset-specific cues. We release the contaminated-file lists, recovered patient identifiers, and deduplicated splits.
[CV-168] From Image Latent Space to Fuzzy Rules: Interpretable Analysis of Gastrointestinal Foundation Model ECCV2026
链接: https://arxiv.org/abs/2610.00414
作者: Michael D. Vasilakakis(1),Dimitris K. Iakovidis(1) ((1) Department of Computer Science and Biomedical Informatics, University of Thessaly, Lamia, Greece)
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted at the excv, ECCV 2026 Workshops. 17 pages, 4 figures, 7 tables
Abstract:Foundation models pretrained on large-scale datasets demonstrate strong transferability to medical imaging tasks. However, understanding how their latent representations encode clinically relevant information remains an open challenge in safety-critical domains. This study proposes a prototype-based fuzzy-rule framework that interprets the patch-level features produced by the inner layers of pretrained foundation models, without any fine-tuning. Class-specific prototypes are learned by clustering in the feature space, yielding compact visual patterns. Patch features are then expressed as prototype similarities and classified by fuzzy rules with linguistic IF-THEN conditions that are human readable. The framework is applied across the final two blocks of ViT-S/16 backbones pretrained on ImageNet-1K and GastroNet-5M, and benchmarked against k-nearest neighbours, kernel SVM, and linear probing under identical frozen features, on wireless capsule endoscopy classification, gastrointestinal endoscopy classification, and colonic polyp segmentation. The experimental analysis shows that the proposed method, without backbone fine-tuning, reaches accuracy comparable to these black-box classifiers, and that domain-specific pretraining yields features that are both discriminative and symbolically compressible. Because the resulting rules are extracted from real data and expressed in interpretable terms, they are further used as an instrument to investigate synthetic medical images, providing a human-readable account of which real prototypes and rules a generator reproduces or fails to reproduce, localising where a synthetic image departs from real tissue rather than summarising it with a single score. The framework thus offers a transparent, depth-resolved view of how foundation models organise clinically relevant structure, together with a practical downstream use of the extracted rules.
[CV-169] Manifold-Constrained Initial Noise Optimization for Efficient Generative Model Alignment
链接: https://arxiv.org/abs/2610.00365
作者: Jinho Chang,Jong Chul Ye
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
备注: 25 pages, 13 figures
Abstract:Recent advances in distillation and flow-map models have enabled deterministic one- or few-step generation for high-quality data, facilitating a new branch of reward alignment approaches that directly optimize the initial noise from a Gaussian distribution. However, most existing initial-noise optimization methods rely on first-order gradient information, which is either inapplicable or suffers from instability and inefficiency in black-box reward scenarios. Here, we introduce ZeNOVA, a stable and efficient initial noise alignment method in a gradient-free manner. Specifically, we address existing algorithms’ major challenge in black-box scenarios through annealed soft-value guidance, manifold-constrained hyperspherical Langevin dynamics, and Metropolis-Hastings jumping. Extensive experiments on image and video generative models show that ZeNOVA outperforms all evaluated zeroth-order baselines by optimizing the initial noise toward higher rewards substantially more stably while exploiting the geometry of the Gaussian prior, demonstrating its practical applicability to various black-box reward alignment.
[CV-170] DexPolicy: Scheduled Exploration for Trajectory-Guided Dexterous Manipulation
链接: https://arxiv.org/abs/2610.00360
作者: Haoyu Wang,Siyuan Qian,Yanjun Li,Zeyu Zhang,Yandong Guo,Boxin Shi,Hao Tang
类目: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Reinforcement learning (RL) for dexterous manipulation must discover finger-object contacts and then control the object precisely; the action noise that serves the first goal can interfere with the second. In trajectory-guided settings such as ViViDex, where RL refine hand-object trajectories from human video, our baseline PPO runs end near their initial action noise after 5M steps, motivating explicit control of exploration scale. DexPolicy makes that scale an explicit function of training steps, annealing from broad to narrow exploration while holding loss, architecture, reward, and optimizer settings fixed. We study three policy-optimization settings: PPO, critic-free GRPO continuation, and a flow-parameterized PPO variant (FPO). Across five YCB objects and three training seeds, mean deterministic Target success rises from 49.4% to 68.1% (FPO), 14.1% to 45.4% (GRPO), and 32.0% to 35.7% (PPO). On a RealMan RM75 arm with an Inspire/RH56 hand, 360 trials over three objects raise mean Target success from 25.0% to 85.0% (FPO), 10.0% to 63.3% (GRPO), and 8.3% to 43.3% (PPO), with one trained model per object-method condition. PPO component screening favors noise control over the tested optimizer contraction; the selected PPO schedule yields higher mean Target success than linear decay with the same endpoints on three tested objects. Training return, deterministic Target success, and tolerance to execution noise dissociate; schedules should therefore be judged by terminal task success under the intended execution conditions, per task and policy-optimization setting. Code: this https URL. Website: this https URL.
[CV-171] Diffusion Editing with Soft Mask: Pixel Level Redo of Image and Video with Adjustable Strength
链接: https://arxiv.org/abs/2610.00359
作者: Candi Zheng,Yuan Lan
类目: Graphics (cs.GR); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)
备注:
Abstract:Diffusion models with prompt and reference image-guided editing have seen rapid progress, yet they remain too coarse for pixel-level control. One promising direction is to incorporate a soft mask that specifies spatially varying edit strengths but training such fine-grained control demands expensive pixel-wise annotations, while existing zero-shot methods often yield unsatisfactory results. We introduce SoftPaint, a new zero-shot sampling method that leverages soft masks to enable a continuous spectrum of edits, from fully preserving the original content to completely re-synthesizing the masked region. Going beyond zero-shot inpainting methods, we design a Langevin-iteration-based sampler that respects per-pixel soft mask strengths, which applies universally to image and video diffusion models, enabling tasks such as video editing. The method is gradient-free, memory-efficient, and achieves smooth, pixel-level edits across multiple image and video backbones.
[CV-172] Vmem-φ: Low-Compute Out-of-Distribution Detection in Spiking Neural Networks from Membrane-Potential Statistics
链接: https://arxiv.org/abs/2610.00350
作者: Arul Rana,Agrim Tripathi,Shoaib Ahmed Dipu,Md. Shaown Miah,Syed Ishtiaque Ahmed,Sayeed Shafayet Chowdhury
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Spiking Neural Networks (SNNs) offer an energy-efficient approach to processing event-camera data, yet out-of-distribution (OOD) detection remains challenging in this setting. Existing OOD detection methods often depend on model outputs or computational components that are unavailable in object detection SNNs or are poorly suited to low-compute deployment. To that effect, we show that the subthreshold membrane potential (V_\mathrmmem(t)) provides a useful internal signal for detecting distribution shifts. Simple per-channel statistics derived from these membrane dynamics enable OOD detection. To evaluate this approach, we introduce Gen1-C, an event-camera corruption benchmark developed upon the Prophesee Gen1 automotive detection dataset, containing six sensor-motivated histogram-level stress tests at five severity levels. We further propose the Multi-Descriptor Deviation (MDD), a corruption-blind method that operates on membrane-potential statistics. At the highest corruption severity, MDD achieves an AUROC of more than 0.88 on five of the six corruptions using only a bounded 64-frame observation window. Notably, the remaining corruption is also the one that has the smallest effect on the underlying detector. These results show that the temporal membrane-potential dynamics can provide an effective and low-cost signal for OOD detection in SNN-based event perception.
[CV-173] UnifiedAttack: Evaluating the Safety of Large Multimodal Models in Synergistic Harmful Image-Text Generation
链接: https://arxiv.org/abs/2610.00341
作者: Bingjun Luo,Jialin Guo,Tony Wang,Siqi Li
类目: Cryptography and Security (cs.CR); Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:As Large Multimodal Models (LMMs) transition toward natively unified architectures, evaluating their safety in synergistic harmful image-text generation tasks becomes a critical challenge. Unlike unimodal threats, synergistic risks emerge when text and image modalities are coordinated to produce harm that significantly exceeds their individual components. We introduce UnifiedAttack, a novel benchmark designed to evaluate LMM safety in collaborative scenarios by focusing on the harmfulness gain achieved through cross-modal synergy. The benchmark incorporates samples filtered for their multimodal potential alongside a novel subset of synthesized disinformation queries. To verify identified vulnerabilities, we propose a synergistic hijacking framework featuring In-Context Reskinning (ICR) and Cognitive Planning Injection (CPI). ICR utilizes few-shot learning to wrap adversarial intent in benign virtual shells to desensitize safety filters, while CPI hijacks the reasoning path by enforcing a plan-then-execute paradigm. By compelling the system to commit to a neutral logical plan, we exploit its internal drive for consistency to induce the synchronized generation of harmful multimodal content. Extensive evaluations on state-of-the-art architectures demonstrate that UnifiedAttack consistently bypasses modern alignment. Our findings reveal that the structural helpfulness and logical coherence of unified models can be systematically weaponized, highlighting the urgent need for logic-aware defenses in synergistic generation tasks. Code is available at this https URL .
[CV-174] Retrospective Open-Vocabulary Memory for Long-Term Object Search
链接: https://arxiv.org/abs/2610.00330
作者: Jiaming Wang,Zhiwei Xue,Chen Jizhuo,Peng Shiqi,Harold Soh
类目: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
备注: 25 pages, 5 figures
Abstract:Long-term object search requires learning where objects usually appear from repeated but uneven observations of a changing environment. We formulate retrospective open-vocabulary memory as probabilistic inference from censored observations, where the key idea is to reason with evidence per opportunity: a detection or non-detection should influence belief only in proportion to the robot’s opportunity to observe the corresponding location. We introduce ECROM, which uses this principle to estimate long-term prevalence for concepts specified only at query time and converts the resulting belief directly into an active-search prior. To evaluate this problem, we introduce a controlled long-term benchmark in ten HM3D homes that independently varies object placement and observation opportunity across repeated traversals. ECROM improves support-level AP on held-out queries by 4.5 points and search SPL by 4.2 points over the strongest competing memory in each metric. The benchmark, dataset, and code will be open-sourced.
[CV-175] EgoRefine: Ego-Referenced Predictive Alignment and Trajectory-Conditioned Reliability-Aware Fusion for Asynchronous Collaborative Perception
链接: https://arxiv.org/abs/2610.00319
作者: Lingzhao Kong,Yongsheng Zang,Yu Kang,Kailun Yang,Jie Fu,Yukun Zuo,Zhiyong Li
类目: Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO); Image and Video Processing (eess.IV)
备注: The source code will be made publicly available at this https URL
Abstract:Collaborative perception enables connected agents to share complementary observations for 3D object detection, extending sensing range and mitigating occlusion. Under asynchronous communication, however, cooperative features arrive with temporal delay. Existing prediction-based methods compensate for these features mainly from the transmitting agent’s own history, leaving residual misalignment with the ego agent’s current observation; subsequent fusion also often overlooks spatial variations in alignment quality. We propose EgoRefine, an ego-referenced predictive alignment and reliability-aware fusion framework for asynchronous collaborative perception. Its Ego-referenced Predictive Alignment module uses the current ego feature to guide cooperative trajectory-field prediction and refines the sampling offsets along an ego-referenced trajectory direction. Its Trajectory-conditioned Reliability-aware Fusion module treats the trajectory discrepancy between the ego and cooperative streams and the directional refinement magnitude as alignment cues, using them to condition the relation between aligned features and adaptively reweight the two streams before convolutional fusion. Experiments on V2V4Real and DAIR-V2X-Seq show that EgoRefine outperforms TraF-Align by 1.6 and 2.9 points on average in AP@0.5 and AP@0.7, respectively. The source code will be made publicly available at this https URL.
[CV-176] DriftOPD: Sequence-Level Reverse-KL Distillation for One-Step VLA Policies
链接: https://arxiv.org/abs/2610.00317
作者: Youngjun Jun,Kyumin Choi,Youngmin Kim,Seonghyun Jin,Sunwoo Park,Jangho Park,Jong Chul Ye
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注: Preprint
Abstract:Vision-Language-Action (VLA) models increasingly rely on action experts that generate short action chunks under receding-horizon control. While chunk-level training is convenient across robot embodiments, it optimizes local action likelihood without explicitly accounting for long-horizon task success. Sequence-level reinforcement learning can address this limitation, but typically requires policy rollouts and closed-loop interaction, which are costly for real-robot manipulation. We introduce DriftOPD, a teacher-free, rollout-free framework for sequence-level on-policy distillation of continuous VLA action experts. We show that the sequence-level reverse Kullback-Leibler (KL) divergence decomposes into a chunk-level reverse-KL term and a future-potential term that captures the long-horizon effect of the current action. DriftOPD optimizes these two terms using a one-step drifting objective and a Q-function critic learned from offline demonstrations, respectively, enabling sequence-level optimization with only offline data and one-step action generation. Across multiple VLA architectures in simulation and real-world manipulation, DriftOPD generally outperforms existing one-step distillation baselines while achieving task success performance comparable to multi-step teacher policies. These results demonstrate that long-horizon behavior can be effectively distilled into one-step VLA action experts without online interaction or a separate teacher.
[CV-177] Beyond Pixel Reconstruction: Retrieval-Guided Glyph-Aware Restoration for Low-Resource Manchu Historical Documents
链接: https://arxiv.org/abs/2610.00315
作者: Ting Huang,Dongdong Wang,Mingqiu Liang,Siyang Lu
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: 8 pages, 7 figures
Abstract:Historical Manchu documents preserve invaluable linguistic and cultural heritage, yet their digitization is hindered by severe degradations and the scarcity of paired training data. Existing document restoration methods primarily optimize pixel-level reconstruction, which can produce visually plausible results while failing to preserve the structural identity of Manchu glyphs. To address this limitation, we propose a retrieval-guided glyph-aware restoration framework that goes beyond pixel reconstruction by explicitly incorporating glyph-level structural knowledge. Our method retrieves relevant glyph exemplars to provide structural guidance during restoration and integrates this information into the reconstruction process, improving the recovery of degraded character structures under low-resource conditions. Extensive experiments on Manchu historical documents demonstrate that the proposed approach improves both image restoration quality and glyph-level fidelity compared with existing restoration methods. These results highlight the importance of incorporating character-aware structural priors for reliable restoration of low-resource historical documents.
[CV-178] Decoding the Disaster: Multi-Task Geospatial Reasoning with Vision-Language Models and Crowdsourced Imagery for Disaster Mapping
链接: https://arxiv.org/abs/2610.00302
作者: Wenping Yin,Fabian Desuer,Ziqi Liu,Naixia Mou,Weijia Li,Pedram Ghamisi,Xiao Xiang Zhu,Hao Li
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:
Abstract:Crowdsourced imagery provides timely, fine-grained, street-level observations for disaster mapping, complementing conventional remote sensing imagery (RSI) during emergency response. However, such imagery is often unstructured, spatially ambiguous, and lacks reliable geographic metadata, making manual geolocalization and interpretation labor-intensive and difficult to scale. This work proposes a multi-task Geospatial Reasoning Disaster mapping framework, namely GRDisaster, to examine the potential of vision-language models (VLMs) in understanding, geolocalizing, and reasoning over crowdsourced disaster imagery. GRDisaster is built on a newly curated benchmark dataset derived from PhotoMappers, comprising 26,340 images organized into human-validated volunteered geographic information (VGI), street-view imagery (SVI), RSI cross-view triplets covering multiple disaster events from 2018 to 2024. The framework combines deterministic and probabilistic cross-view geolocalization with multi-view fusion to associate VGI images with georeferenced SVI and RSI. It introduces two sets of spatial reasoning indicators for cross-view geolocalization validation and disaster damage assessment. These indicators use structural, environmental, and global-scene cues to validate cross-view correspondences and visually observable damage evidence with expert-verified annotations to assess disaster severity, improving the interpretability of VLM outputs. To our knowledge, this study provides the first systematic investigation and unified evaluation framework for examining how VLM-based spatial reasoning can transform crowdsourced disaster imagery into actionable geospatial artificial intelligence (GeoAI) through cross-view geolocalization validation, interpretable spatial reasoning, and damage-aware severity assessment.
[CV-179] LENS-GRF: Permutation-Invariant Lesion Evidence Network with Gated Residual Fusion for Acne Severity Grading and Multi-Rater Clinical Oracle Analysis
链接: https://arxiv.org/abs/2610.00294
作者: Muhammad Muhtasim Shahriar,M. F. Mridha
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注: Submitted to Computer Methods and Programs in Biomedicine (Elsevier)
Abstract:Automated acne severity grading requires both whole-face context and fine-grained lesion evidence. We propose LENS-GRF (Lesion Evidence Network with Set-Transformer and Gated Residual Fusion), an interpretable multi-stage framework for four-class acne severity grading. The method combines Adaptive Facial Skin Segmentation and a global Vision Transformer prior with a permutation-invariant Lesion Set Transformer that encodes localized lesion patches and spatial geometry. Gated Residual Fusion adaptively controls the local residual contribution and reduces to the global prediction when the gate is zero. On ACNE04, fully automated LENS-GRF with YOLOv11s achieved 80.82% accuracy; with ground-truth lesion annotations, it achieved 95.89% +/- 0.59% accuracy and a Quadratic Weighted Kappa of 0.9753. A data-integrity audit identified 15 cross-split duplicate image pairs, including five with conflicting severity labels. In locked zero-shot evaluation on the full PLSBRACNE01 cohort (200 subjects, 600 views), automated LENS-GRF achieved 35.00% accuracy versus 42.50% for the global baseline. On the 148-subject common cohort used for three-dermatologist oracle analysis, ground-truth lesion inputs increased the best oracle accuracy to 47.97%, while the highest oracle QWK was 0.5799. Pairwise oracle agreement ranged from 49.32% to 66.22%, highlighting detector domain shift, annotation variability, and cross-criterion mismatch.
[CV-180] Multi-Resolution Feature Fusion U-Net for Magnetic Resonance Imaging Segmentation
链接: https://arxiv.org/abs/2610.00279
作者: Eirini Cholopoulou,Dimitrios E. Diamantis,Dimitris K. Iakovidis
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注:
Abstract:The segmentation of anatomical structures in medical images and particularly in MRI scans, is essential for clinical diagnosis and monitoring disease progression. While Deep Learning (DL) architectures, such as U-Net and its extensions are very effective in medical image segmentation tasks, they often struggle with preserving fine-grained details and global contextual information. This is especially challenging for MRI data segmentation, where anatomical structures are characterized by irregular boundaries and variations in shape, contrast, and scale. To address this challenge, we propose a novel DL architecture for MRI segmentation across different anatomical structures. Specifically, the architecture introduces a module, named Multi-Resolution Feature Fusion (MRFF), that can be easily integrated into any U-Net-like architecture. The MRFF is integrated in all levels of an encode-decoder structure, along with attention mechanisms and skip connections to extract features at multiple resolutions, enabling the model to capture both fine-grained details and global contextual information. We evaluate the MRFFU-Net on two publicly available benchmark MRI datasets of different anatomical targets; one for Cerebrospinal Fluid (CSF) segmentation in spinal MR scans, and one for left atrium cardiac segmentation, from the Medical Segmentation Decathlon (MSD) challenge. Experimental results indicate that MRFFU-Net outperforms state-of-the-art models across multiple evaluation metrics, demonstrating its effectiveness in MRI segmentation.
[CV-181] Query Independent Variable Rate Visual Token Coding
链接: https://arxiv.org/abs/2610.00204
作者: Hongbo Zhang,Zihao Yang,Liuyang Song,Daqian Yang,Haoyang Yao,Yan Wen,Zhengtao Yao
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Visual-token compression for vision–language models is posed almost entirely as a selection problem: decide which tokens to keep and discard the rest. The criteria that work best rank tokens by the attention the language model pays them, which makes the ranking a function of the question being asked. That is invisible in a single-turn benchmark and decisive whenever a compressed representation is written once and read many times, as when it is cached across the turns of a conversation or transmitted between a device and a server. We take the other half of the classical transform-coding toolkit instead: keep every token and vary its rate. A transform code exposes each token’s measured distortion–rate curve, and a fixed bit budget is distributed across tokens by exact integer rate–distortion optimisation on those curves. No text enters the pipeline, so one compressed representation serves any query. At equal bit budgets, on two datasets and two capacities, it preserves the model’s output distribution and its answers better than uniform-rate coding, the closed-form water-fill and distortion-ranked pruning. It matches attention-ranked pruning on the question pruning was tuned for, and overtakes it once the compressed image must answer a different question about the same image.
[CV-182] GPEC: Efficient Pre-LLM Gaussian Process Embedding Correction for Cardiac Video Caption Generation
链接: https://arxiv.org/abs/2610.00196
作者: Arefeh Rezaei
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Multimodal large language models (MLLMs) have shown strong potential for video understanding and caption generation, but their performance may decline in specialized medical imaging domains such as echocardiography. This work introduces Gaussian Process Embedding Correction (GPEC), a modular and computationally efficient pre-LLM error-correction method that improves the visual representations used by VideoChat2 for cardiac ultrasound caption generation. GPEC is inserted between the visual projection layer and the language model and learns a residual correction that moves the projected visual representation toward an annotation-guided target. The target is constructed by converting structured video annotations into qualitative attributes, generating a fixed-format reference caption, and mapping it into the language-model embedding space. The correction is modeled using a sparse variational Gaussian Process with inducing points, natural-parameter variational updates, and a block-wise linear kernel, while the original VideoChat2 components remain this http URL method is evaluated using representation-level, caption-level, content-oriented, and execution-time metrics by comparing the original VideoChat2 with VideoChat2 + GPEC under identical input and reference conditions. Results show improved caption similarity and content alignment after applying the proposed correction. Furthermore, GPEC adds less than 0.05 s of inference-time overhead per video in the evaluated setting. These findings indicate that GPEC can improve caption generation in specialized medical video domains with minimal computational cost, without requiring end-to-end fine-tuning of the pretrained multimodal backbone.
[CV-183] GS-PQM: A Parameter-Domain Quality Metric for Compressed Gaussian Splatting
链接: https://arxiv.org/abs/2610.00195
作者: Pedro Martin,António Rodrigues,João Ascenso,Maria Paula Queluz
类目: Graphics (cs.GR); Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)
备注:
Abstract:Recent advances in Gaussian Splatting (GS) compression have enabled substantial reductions in GS model size. Reliable objective quality assessment is therefore essential for comparing compression methods and guiding the development of more efficient GS codecs. Existing GS quality assessment typically relies on image and video quality metrics, requiring rendering of predefined viewpoints and making the quality estimate dependent on the selected views. This paper introduces GS-PQM, a novel full-reference quality metric for post-training GS compression that operates directly in the GS parameter domain. GS-PQM estimates perceptual quality from a set of parameter-domain distortion errors using a Support Vector Regression model. Experimental results show that GS-PQM outperforms 25 existing image, video, and point-cloud quality metrics in assessing compressed GS content, providing an accurate and computationally efficient alternative to rendering-based quality assessment.
[CV-184] Uncertainty-Aware RL-Controlled Adaptive 3D Mapping BMVC2026
链接: https://arxiv.org/abs/2610.00188
作者: Alpay Ozkan,Tunc Ozan Aydin,Marc Pollefeys,Jelena Trisovic,Daniel Barath
类目: Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR); Image and Video Processing (eess.IV)
备注: To appear at BMVC 2026. Code available at this https URL
Abstract:Voxel-based volumetric mapping is fundamental to 3D reconstruction, yet fixed-resolution grids remain inherently inefficient - wasting memory in uniform regions and losing detail in complex ones. Existing adaptive methods, such as MAP-ADAPT, partially address this by varying resolution based on geometry and user-defined semantic class lists, but these heuristics require expert tuning, lack generalization to unseen objects, and provide no explicit mechanism to control memory usage. We propose an adaptive framework that refines voxels based on semantic entropy, which captures label uncertainty, together with geometric curvature and texture richness as scene complexity cues, yielding principled resolution allocation without reliance on semantic taxonomies. To make the accuracy-memory trade-off explicit and user-controlled, we further introduce a reinforcement learning agent that learns voxel subdivision policies under a user-specified target memory budget, replacing hand-tuned thresholds with a single intuitive control parameter. The resulting multi-resolution TSDF achieves higher geometric accuracy, better semantic consistency, and improved memory-accuracy trade-offs compared to MAP-ADAPT and fixed-resolution baselines on both synthetic and real-world datasets. Our code and models are available at this https URL.
[CV-185] Evaluating the Robustness of Anti-UAV Detection under Controlled Fog Degradation: Fog-Aware Training and Clear-Sky Tradeoff
链接: https://arxiv.org/abs/2610.00141
作者: Gur Levy Birkental,Seyed Sahand Mohammadi Ziabari,Ali Mohammed Mansoor Alsahag
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Vision-based anti-UAV systems must function in poor visibility, yet most benchmarks use only clear-sky footage, and previous robustness studies treat adverse weather as a simple present/absent condition. As a result, the impact of fog severity on ground-to-air UAV detection remains poorly understood. This work presents the first severity-controlled fog benchmark for this task: synthetic fog at ten severity levels is applied to the RGB modality of the Anti-UAV300 dataset, comparing a clear-trained YOLOv5m baseline to a fog-aware model trained on both clear and foggy images. Detection performance drops sharply and non-linearly: degradation is front-loaded across light-to-moderate fog (beta approximately 0.05-0.10), with a 96% reduction in mAP@0.5:0.95 from clear to thickest fog, mainly due to lost recall and confidence. On the comparable metric (mAP@0.5), this collapse exceeds the most extreme rain degradation reported in the closest prior benchmark. Fog-aware training boosts detection across all severities (up to +0.320 mAP@0.5:0.95) with only a 9.1% drop in clear-sky accuracy, raising the threshold for reliable detection while not preventing collapse under extreme fog. Since reliability is lost within a narrow visibility range, simple clear vs adverse tests underestimate operational risk.
[CV-186] A Comprehensive Review of One-Pixel Attack: Research Status Taxonomy Applications Regulation Policy and Future Directions
链接: https://arxiv.org/abs/2610.00125
作者: Mirza Niaz Morshed,Md. Masudul Islam,Galib Muhammad Shahriar Himel,Md. Aslam Uddin,Hui Liu,Md. Shafiqul Islam
类目: Cryptography and Security (cs.CR); Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:One-Pixel Attacks (OPAs) represent one of the most extreme demonstrations of adversarial fragility in deep learning, where modifying a single pixel can reliably induce high-confidence misclassification across domains such as medical diagnosis, autonomous driving, biometrics, and quantum communication. Despite their conceptual simplicity, OPAs remain underexamined in existing adversarial-attack surveys, which provide only fragmented or cursory coverage. This PRISMA-guided review synthesizes high-quality studies from 2017 to 2026 and delivers a unified, multi-axis taxonomy of OPA research spanning algorithmic foundations, black-box evolutionary optimization, emerging hybrid and program-synthesis attacks, defence mechanisms, interpretability tools, and domain-specific vulnerabilities. Our analysis reveals the dominance of Differential Evolution-based strategies, the rise of efficiency-optimized and saliency-guided methods, and persistent gaps in dataset diversity, transferability, and standardized evaluation. We summarized and assess defence paradigms including pixel restoration, anomaly detection, input-space transformations, and robust training highlighting their trade-offs in robustness, imperceptibility, and computational overhead. Building on these insights, we outline future research priorities involving selective pixel recovery, transformer-specific vulnerability analysis, saliency-driven optimization, and real-world domain-adaptive defences. We further propose a regulatory framework emphasizing robustness testing, incident disclosure, and AI security governance. This review establishes a comprehensive foundation for understanding, evaluating, and mitigating ultra-sparse adversarial threats in contemporary AI systems.
[CV-187] A Low Grounding Score Is Not an Ungrounded Judge: Identifying the Perceptibility Confound in Multimodal Oversight
链接: https://arxiv.org/abs/2610.00111
作者: Rasul Khanbayov,Hasan Kurban
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Model judges now supervise multimodal systems at scale, filtering training data, selecting outputs, and supplying the reward that shapes multimodal reasoning models. Trusting one means first checking that it uses its evidence, and that check is itself worth scrutinizing, so we ask whether a counterfactual probe of visual grounding measures what it claims to. The probe edits the image so the ground truth flips, holds the reasoning trace fixed, and asks whether the verdict follows. We formalize it as the Verdict Grounding Score and show it cannot be read the way such scores are read. A verdict responds only to an edit that reaches the judge’s decision-relevant reading, so the score is capped by how perceptible the edit is, and unless editing makes the attribute easier to read, the error is one-sided: the score can only make a judge look less grounded than it is. The practical failure is therefore a false alarm, an auditor discarding a usable overseer. Under assumptions we state, we show this missing quantity is not merely bounded but identified from three quantities the same audit protocol already collects, which makes the false-alarm rate directly measurable rather than merely a concern. Auditing nine judges, we find the predicted ordering holds strictly across our entire primary pool, and the typical judge there acts on only about half of the edits whose attribute it can otherwise resolve. Applying a conservative rejection threshold certifies several cells as false alarms outright, the clearest being a judge that detects the injected error essentially every time while still scoring as if it had not used the image at all. The rule that follows is that an image-side counterfactual score should never be reported alone: a detection probe on the unedited image upper-bounds it, certifies its false alarms, and costs nothing extra to run.
[CV-188] A Framework for Egocentric and Exocentric Procedural Understanding via Temporal Segmentation and Semantic Abstraction ECCV2026
链接: https://arxiv.org/abs/2610.00069
作者: Vivek Chavan,Jörg Krüger
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注: Accepted for oral and poster presentation at the ACVR Workshop, ECCV 2026. Non-archival abstract; not published in the workshop proceedings. 8 pages, 1 figure
Abstract:Long-horizon ego/exo data contains rich procedural evidence, but are redundant, noisy, and costly to process or retain. We propose a compact framework that converts continuous multimodal workplace video into a structured Procedural State Memory, implemented as a Work Environment Model (WEM). Inspired by event segmentation theory, we detect boundaries using changes in visual context, location, motion, narration, gaze/object interaction, and optional exocentric workspace evidence, rather than fixed windows or visual novelty alone. Each segment is abstracted into an evidence-linked event card containing actor, interval, location, action, objects/tools, pre/post state, confidence, and provenance. These event cards incrementally update the WEM, enabling compact, auditable documentation and retrieval under on-premise privacy constraints. We instantiate the design with frozen DINOv2 and VJEPA-2 encoders and a local language model, and outline evaluation criteria for segmentation quality, memory compression, retrieval fidelity, and long-horizon QA.
[CV-189] Robust Online Aero-Engine Blade Defect Detection via Dual-Alignment Test-Time Adaptation
链接: https://arxiv.org/abs/2610.00067
作者: Zhaoyang Wang,Haiyong Chen,Dongying Li,Yining Wang,Huapeng Wu,Xinwei Lv,Atik Shahariar
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: This manuscript is Accepted at conference PRCV 2026
Abstract:Reliable visual inspection is essential for quality assurance in aero-engine blade manufacturing, where defect appearance may vary across production lines, imaging conditions, blade poses, and surface backgrounds. Such domain shifts cause a mismatch between training and deployment data and degrade the reliability of deep defect detectors in online inspection. This problem is particularly challenging because aero-engine blade images usually contain sparse defects, making pseudolabel-based adaptation vulnerable to noisy or missing predictions. To address this issue, we propose Aero-engine Blade Defect Detector (ABDD), an online adaptive detection framework based on test-time adaptation. ABDD introduces a Dual-Alignment Strategy to jointly adapt global visual style and local defect morphology by combining feature-statistics alignment with pseudo-box alignment. To reduce error accumulation from unreliable pseudo labels, an Uncertainty-aware Box Filtering mechanism evaluates pseudo boxes using classification confidence, classification entropy, and localization entropy. In addition, a lightweight Sparse Dilated Mona module enables parameter-efficient delta tuning while limiting source-domain forgetting. ABDD is evaluated on CD-AeBD and HD-AeBD under multiple domain-shift scenarios, with TTA strategies compared under a unified RT-DETR + Swin-T architecture. Experiments show that ABDD consistently improves detection robustness under domain shifts, and its practicality is further validated on an industrial inspection platform.
[CV-190] Reachability Is Not Generalization: Understanding Verb–Noun Decomposition in Assembly Action Recognition BMVC2026
链接: https://arxiv.org/abs/2610.00064
作者: Changyi Li,Yu Xiao
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted by BMVC 2026
Abstract:Assembly actions are compositional: they combine a manipulation with a part or tool. In deployment, systems routinely encounter novel combinations of familiar components, yet an atomic action classifier assigns every unseen combination exactly zero probability by construction. The prevailing solution is verb–noun decomposition, which predicts components separately and recombines them to reach unseen actions. While widely adopted, how decomposition generalizes under compositional shift remains poorly understood. We present a systematic analysis of verb–noun decomposition across three assembly datasets (MECCANO, HAViD, and IMPACT). Although decomposition escapes the atomic ceiling, its generalization extends only partially beyond it. Unseen-composition performance remains strongly tied to the co-occurrence structure of the training data, indicating that much of the observed gain arises from interpolation within densely supported regions of the compositional space rather than from unconstrained recombination. Across datasets, failures consistently concentrate on the larger-vocabulary component, and IMPACT’s verb-heavy vocabulary reverses the bottleneck from nouns to verbs. We further show that shared-encoder training introduces component entanglement, encouraging reliance on co-occurrence patterns that transfer poorly to unseen compositions and trailing independent recombination by up to 6.0\times in harmonic mean. Taken together, these findings explain why decomposition achieves only partial compositional generalization in practice. By identifying primitive support, vocabulary asymmetry, and component entanglement as connected sources of error, we provide a portable diagnostic framework for studying compositional recognition beyond aggregate accuracy. Code: this https URL.
[CV-191] DSSR-3D: Decoupled Reasoning for View-Dependent Referring in 3D Gaussians
链接: https://arxiv.org/abs/2610.00040
作者: Thanh-Khoi Nguyen,Thien-Phuc Tran,Minh-Triet Tran
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Recent advances in 3D Gaussian Splatting have enabled open-vocabulary and referring segmentation by distilling semantic knowledge from 2D foundation models into 3D representations. However, existing referring fields embed language features in a globally view-invariant space, making them fundamentally unable to resolve observer-centric spatial relations (e.g., “to the left of”) that depend on camera pose. We propose DSSR-3D, an inference-time framework for view-dependent referring segmentation on continuous 3D Gaussian fields, formalized as two interfaces - pose-invariant semantic localization and pose-conditioned spatial reasoning - such that any pair of functions satisfying these constraints yields a valid instantiation, requiring no retraining of the underlying semantic field and no reliance on discrete geometric proxies such as bounding boxes. We instantiate the two interfaces with a temperature-sharpened softmax localization mechanism and a projection-based directional scoring function, fused via a lightweight, training-free step, and show they transfer zero-shot to structurally distinct semantic fields without adaptation. We further propose ViewRef-GS, a benchmark isolating view-dependent segmentation on 3D Gaussian fields, evaluated jointly with an augmented Ref-LERF to provide a comprehensive testbed for viewpoint-dependent spatial grounding. Experiments show consistent gains over existing 3DGS-based referring methods, with no additional training beyond the base semantic field
[CV-192] Seeing the City or Recognizing the Place? What Street-View Imagery Adds Beyond Existing Urban Data in VLM Urban Sensing
链接: https://arxiv.org/abs/2610.00031
作者: Kaizhen Tan
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Street-view imagery is increasingly used to infer urban attributes, but predictive accuracy alone does not reveal how much a photograph contributes beyond data already available for the same place. We compare image-based predictions with existing urban data across seven attributes from five public resources and three VLMs. The same urban units are evaluated using images, task context, nearby observations, and public records, while image replacements and conflicting records test source reliance. Existing urban data matched or exceeded image-only models for road damage, curb ramps, and house price, while neighbouring official statistics nearly matched the best image result for population. Images were more informative for building type, building function, and low-rise floor count. For floor count, image advantage increased by 5.7 percentage points per doubling of distance to the nearest labelled building and declined for tall buildings whose rooflines often fell outside the frame. Models frequently followed conflicting records. OpenFACADES floor annotations were generated with OpenStreetMap floor values and showed the opposite height-dependent error pattern from image-only reruns. Street-view image value therefore depends on visual legibility and local data coverage. Comparing images with existing urban data can guide image collection and clarify the provenance of derived urban maps.
[CV-193] Domain generalization and synthetic data in object detection: the enabler the probe and the gap
链接: https://arxiv.org/abs/2610.00030
作者: Elfi I.S. Hofmeijer,Ella P. Fokkinga,Friso G. Heslinga,Klamer Schutte,Jörgen M. Karlholm
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Submitted to SPIE Sensors + Imaging 2026
Abstract:Object detection models often experience performance degradation when deployed under distribution shifts, caused by for example changes in weather type, operational environment, or object appearance. Domain Generalization (DG) aims to develop models that remain robust under such shifts and generalize well to unseen domains. DG research specifically focused on object detection models is scarce, although these models face additional challenges around localization and multi-scale representations. Synthetic data is a promising tool to support in DG, by enabling large-scale generation of diverse new samples. In this paper, we present an object detection-centric review of DG and examine the role of synthetic data from three complementary perspectives. First, synthetic data acts as an enabler of DG through diversification and alignment strategies that aim to improve robustness to distribution shifts. Second, it serves as a probe that enables controlled experimentation to identify and understand failure modes. Third, we discuss the synthetic-to-real gap, a particularly challenging form of domain shift that arises when models trained on synthetic imagery are deployed on real-world data. Through reviewing these perspectives, we identify limitations of current DG approaches for object detection and argue that future research requires representation-aware methods that explicitly address both localization and classification under domain shift.
[CV-194] Spatial Lifting for Dense Prediction
链接: https://arxiv.org/abs/2610.00017
作者: Mingzhi Xu,Tao Zhou,Yong Li,Yizhe Zhang
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: 28 pages 5 figures
Abstract:We present Spatial Lifting (SL), a novel methodology for dense prediction tasks. SL operates by lifting standard inputs, such as 2D images, into a higher-dimensional space and subsequently processing them using networks designed for that higher dimension, such as a 3D U-Net. Counterintuitively, this dimensionality lifting allows us to achieve good performance on benchmark tasks compared to conventional approaches, while reducing inference costs and \textbfdrastically lowering the number of model parameters. The SL framework produces intrinsically structured outputs along the lifted dimension. This emergent structure facilitates dense supervision during training and enables single-forward-pass self-consistency-based quality and uncertainty estimation at test time. Spatial Lifting introduces a simple and general modeling strategy that offers a promising path toward more efficient, accurate, and reliable deep networks for dense prediction tasks in vision.
[CV-195] Emergent Object Binding Has a Finite Spatial Horizon
链接: https://arxiv.org/abs/2610.00006
作者: Mayank Singal
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注: 14 pages, 4 figures
Abstract:Pretrained Vision Transformers encode whether two image patches belong to the same object. This IsSameObject signal is decodable from frozen patch embeddings at high accuracy, which suggests that object binding emerges from self-supervised pretraining alone. We show that this single accuracy number hides the structure of the signal. Binding is local: the probability that two patches of the same object are decoded as bound falls off monotonically with the distance between them and levels off at a nonzero floor, a falloff well described by an exponential with a finite length scale. This decay holds across object sizes, across three families of probe, on both ADE20K and COCO, and across DINO and CLIP backbones, which indicates that it is a property of the representation rather than of the decoder. Reading binding as local spatial coherence with a finite range accounts for a set of behaviors that the aggregate score leaves unexplained: binding weakens on large objects, separates distinct objects of the same class less reliably than objects of different classes, and groups object parts with their wholes. It is, by contrast, unaffected by occlusion once object size is controlled. We map each behavior with confounds controlled. As a preliminary observation, the horizon and its floor are organized at different depths in DINOv2 and DINOv3, which we report as suggestive given the small number of layers probed and the confound between the two models.
[CV-196] STATERA: Hidden Mass Estimation via Zero-Shot Sim-to-Real Kinematics using Frozen Temporal Tubelets
链接: https://arxiv.org/abs/2610.00003
作者: Animesh Varma
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Robotics (cs.RO)
备注: 17 pages, 7 figures, 3 tables. Preprint
Abstract:Vision models pretrained for frame-level appearance often struggle to infer hidden physical properties from motion. We study center-of-mass (CoM) localization for opaque, asymmetric rigid bodies from short monocular videos, where surface cues and point tracking are unreliable under self-occlusion. We propose STATERA, which adapts a pretrained video backbone (V-JEPA) with mostly frozen weights and a lightweight temporal tubelet mixer to predict per-frame CoM heatmaps and trajectories. To support this task, we introduce the HiddenMass Benchmark, comprising 50K MuJoCo trajectories and a 63-sequence real-world test set with physically calibrated CoM ground truth. In simulation, STATERA-50K-Sigma improves normalized CoM error from 41.7% (DINOv2) to 25.2%. In zero-shot sim-to-real transfer, we observe a fundamental trade-off in supervision: phase-aware targets can induce bimodal predictions, while phase-agnostic targets can collapse toward statistically safe centroids. Nevertheless, our phase-aware STATERA-50K-Crescent is the only evaluated method that demonstrates consistent movement toward the true hidden offset. While this leads to a monocular vector overshoot artifact that marginally increases absolute Euclidean error compared to a static geometric centroid, it improves physics capture from 2.6% to 41.0%. These results suggest that frozen temporal representations can better separate inertial dynamics from visual geometry for hidden-parameter estimation.
[CV-197] A2Z GameSpec-Bench: How Faithfully Can Coding Agents Generate Games from Game Design Specifications?
链接: https://arxiv.org/abs/2609.39564
作者: Seonho Lee,Wonryeol Jeong,Alberto Cereser,Inha Kang,Hyeonjong Kim,Seungmin Kwak,Dongmin Park
类目: Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Delegating complete application development to coding agents requires preserving the intended design rather than simply producing plausible outputs through naive prompting. Game development provides a demanding testbed, as long-form Game Design Documents (GDDs) describe requirements that must work together across game logic, visual rendering, and player interactions. However, existing game-development benchmarks typically use compact specifications and provide limited support for evaluating interdependent requirements across these aspects in long-form GDDs. We introduce A2Z GameSpec-Bench, a benchmark of 100 long-form GDDs for evaluating end-to-end game development by agents. We measure faithfulness by checking whether the game satisfies the GDD requirements and preserves the relationships among them. Each GDD is turned into a dependency-aware contract that contains rules, constraints, and prerequisite relations. Following game-development practices, we combine source-code inspection with agent-generated test policies for scenario-based replay and adaptive playtesting. The contract remains fixed across agents and revision rounds, while judgments and evidence linked to the same requirements support consistent comparison and failure detection. Our evaluations show that current agents struggle to jointly satisfy interdependent requirements across code implementation and actual play. Requirement-specific feedback improves GDD Fidelity by 10.9% relative to self-revision after two rounds. A2Z GameSpec-Bench assesses end-to-end specification-following ability beyond implementation judgments and provides targeted feedback to support more faithful game development. Code and datasets are available at this https URL.
[CV-198] MorphoBranch: A Fine-Structure-Preserving Workbench for Morphometric Analysis of Branched Cellular Structures
链接: https://arxiv.org/abs/2610.00860
作者: Song Zhiying,Ling Hanyi,Wu Junyi,Jiang Yangbo
类目: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV); Quantitative Methods (q-bio.QM)
备注: 12 pages, 9 figures
Abstract:Background and Objectives: Fluorescence-labeled cellular arbors provide readouts of neuronal and microglial morphology, but fine and weakly labeled processes are prone to fragmentation and false connections that bias skeleton-based measurements. We present MorphoBranch, a fine-structure-preserving, human-reviewable workbench for morphometry of branched cellular structures. Methods: MorphoBranch combines a deterministic Morphometry Engine with an LLM-assisted Refinement Engine. The Mor- phometry Engine implements an image-to-graph workflow integrating multiscale structural evidence extraction, hysteresis segmen- tation, evidence-constrained skeleton refinement, and graph-based morphometry. The Refinement Engine maps natural-language requests to registered actions for parameter adjustment, preview execution, metric reporting, and unsupported-request handling, while image processing and quantitative computation remain deterministic and reviewable. Results: MorphoBranch was evaluated on two public neuronal axon datasets, AxonMIP and AxonStack, and the in-house Cell- Morph dataset of microglial fluorescence images. It achieved the highest Skeleton F1 and clDice and the lowest length-estimation error among the evaluated methods on all three datasets, while also achieving the highest Dice and IoU on AxonMIP and Axon- Stack. Across 150 natural-language tasks, the Refinement Engine achieved a 94.0% end-to-end success rate. Conclusions: These results demonstrate that MorphoBranch provides a reproducible, human-reviewable workflow for mor- phometric analysis of branched cellular structures. It supports fine-structure-preserving quantification across neuronal axon and microglial fluorescence images while maintaining inspectable and reproducible analysis workflows.
[CV-199] Spatially Gated Diffusion for Localized Counterfactual Chest Radiograph Editing
链接: https://arxiv.org/abs/2610.00805
作者: Kamran Ullah Afaq,Basit Raza
类目: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
备注: 16 pages, 1 figure, 10 tables
Abstract:Editing a chest radiograph requires completing the requested change while preserving unrelated content. We study a latent diffusion editor with an instruction-independent source trajectory and an instruction-conditioned editing trajectory. A learned gate mixes their post-sampler candidates at each executed step, and a separate image-space mask composites the decoded proposal with the source. On 2,400 MIMIC-derived requests, 2,244 outputs met the joint target, preservation, quality, and coverage rubric (93.5%), and target completion was 97.4%. Joint validity exceeded a matched composition-only control by 1.7 percentage points (paired patient-cluster 95% interval, 0.6–2.8). At a fixed learned mask, the learned-gate proposal improved joint validity by 1.6 points (0.8–2.4); at a fixed learned proposal, the learned mask improved it by 2.6 points (1.7–3.5). Across three training seeds, mean joint validity was 93.5% with a 0.3-point sample standard deviation. A blinded 240-request assessment yielded adjudicated joint validity of 93.3% for the full editor and 91.7% for composition only. Protected-region mean absolute error decreased from 0.0190 in the raw proposal to 0.0075 after composition. These findings distinguish recurrent-gating effects on the proposal from preservation through final composition in the assessed cohort.
[CV-200] RIQE: a NIQE-style reference model for Computed Tomography
链接: https://arxiv.org/abs/2610.00384
作者: Fabio Mattiussi
类目: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
备注: 17 pages, 6 figures, 6 tables. Code and model: this https URL , archived at doi: https://doi.org/10.5281/zenodo.23055559
Abstract:The Natural Image Quality Evaluator (NIQE) scores an image by its statistical distance from a model fitted on pristine images, and its distributed model is fitted on photographs. We release the Radiology Image Quality Evaluator (RIQE), a NIQE-style model fitted on 3,792 full-dose slices from 158 patients of the public LDCT-and-Projection-data collection, with a declared intensity mapping, a manifest of every slice and a script that reproduces the fit. On 40 held-out patients, RIQE ranks reduced-dose reconstructions, simulated by projection-domain noise insertion, worse than the full-dose reconstruction of the same slice in 240 of 240 chest and 230 of 240 abdominal pairs, and ranks images with 20% more noise worse than their source in 97.5-100% of cases. Its preferences among filtered images, however, do not follow lesion signal. With a 4 mm, +10 HU lesion inserted in noisy abdominal slices, RIQE prefers bilateral filtering to the unfiltered image in every image up to a 32 HU residual, at which 29% of the lesion’s matched-filter signal remains and its detectability index falls from 0.51 to 0.33; it never prefers Gaussian smoothing, which at the same 32 HU residual leaves 70% of the signal and a detectability index of 0.48. Fitted on photographs with the parameters published for NIQE, the same code ranks every simulated reduced-dose abdominal image better than its full-dose counterpart. RIQE is suited to ranking a degraded image against its source; under the conditions tested it should not be the sole criterion for selecting, comparing or tuning denoisers.
[CV-201] LensBridge: Frequency-Guided Compound Degradation Adaptation for Lens Aberration Correction and Veiling Glare Removal
链接: https://arxiv.org/abs/2610.00318
作者: Xiaolong Qian,Zhonghua Yi,Qi Jiang,Kailun Yang,Shuhang Xie,Shaohua Gao,Kaiwei Wang
类目: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV); Optics (physics.optics)
备注: All code will be available at this https URL
Abstract:Simplified optical systems often exhibit residual lens aberrations and Veiling Glare (VG), resulting in spatially varying blur and contrast reduction. Large-scale Lens Libraries (LensLib) enable reusable aberration correction models by covering diverse Point Spread Functions (PSFs), but their aberration-only training distribution does not include target-specific veiling glare. Extending such foundations to compound degradation is challenging because realistic target-system compound pairs are difficult to obtain. To address this challenge, we propose LensBridge, a two-stage framework that first establishes a reusable aberration correction foundation and then adapts it to compound optical degradation using only a few unpaired target observations. In Stage I, we build a PSF-aware one-step diffusion foundation by constructing discrete degradation priors from LensLib PSFs and learning to retrieve them directly from aberrated images, enabling PSF-aware correction without requiring explicit PSF at inference. In Stage II, we adapt this foundation to compound degradation through frequency-domain guidance. At the data level, Frequency-guided Degradation Completion (FDC) transfers target low-frequency characteristics to LensLib aberrated images while preserving aberration structures to synthesize compound training pairs; at the model level, Frequency-guided Pseudo Decomposition (FPD) forms aberration- and VG-dominant pseudo observations to condition separate adaptation branches. Extensive experiments across multiple optical systems demonstrate that LensBridge effectively extends reusable aberration correction foundations to joint aberration correction and veiling glare removal without target-system paired supervision. All code will be available at this https URL.
人工智能
[AI-0] Reconstruct Practice Go Real: Guided Self-Improvement for Embodied Agents
链接: https://arxiv.org/abs/2610.02204
作者: Yen-Jen Wang,Haozhe Jiang,Shuying Deng,Haoru Xue,Weirui Ye,Rocky Duan,Nika Haghtalab,S. Shankar Sastry,Pieter Abbeel,Haozhi Qi
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI); Systems and Control (eess.SY)
备注: 17 pages, 6 figures, 10 tables
Abstract:Building reliable robot capabilities across diverse tasks requires substantial human effort to develop and maintain skills, design rewards, and integrate perception with control. We present Reconstruct, Practice, Go Real (RPG), a framework for autonomous improvement of robot execution systems without updating model weights. RPG identifies manipulation capabilities in an offline dataset and constructs related practice tasks in simulation. During practice, RPG uses execution feedback, privileged simulator state, and available dataset videos to diagnose failures. It develops new reusable symbolic skills, refines existing skills, and revises the system prompt based on these diagnoses. Cross-task evaluation tests individual candidate changes and merged revisions before they are retained for reuse. At test time, a multimodal LLM uses the resulting system prompt and skill library to coordinate perception and robot control. On held-out initializations of 22 manipulation tasks, RPG improves task success from 28.6% after the first practice round to 95.0% after 15 rounds, outperforming all evaluated baselines, including ASPIRE (75.5%) and CaP-Agent0 powered by GPT-6 Astra Pro (60.0%). After a common calibration and hardware-adaptation procedure, the frozen system succeeds in all 30 physical trials, with ten trials on each of three tasks. Project Website: this https URL
[AI-1] FERPO: Forward Entropy-Regularized Policy Optimization
链接: https://arxiv.org/abs/2610.02198
作者: Sebastian Sanokowski,Alireza Sarmadi,Majid Khadiv
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Robotics (cs.RO); Machine Learning (stat.ML)
备注: Code: this https URL
Abstract:Several state-of-the-art methods for online reinforcement learning in continuous control improve policies using action gradients of a learned critic. However, critics are typically trained to predict returns, and accurate value predictions do not necessarily yield accurate action derivatives, potentially leading to unreliable policy updates. We propose Forward Entropy-Regularized Policy Optimization (FERPO), an on-policy maximum entropy reinforcement learning algorithm that performs policy improvement using critic values without differentiating the critic with respect to actions. FERPO derives an optimal target action distribution from a policy-improvement objective regularized by entropy and Kullback-Leibler (KL) divergence. We then fit the actor to this target by minimizing a forward-KL objective, estimated using self-normalized importance sampling (SNIS) with actions drawn from the rollout policy. By limiting the target distribution’s deviation from the rollout policy, the KL regularization helps keep these importance weights well behaved. In contrast to reverse-KL objectives, which can favor a subset of the target distribution’s modes, the forward-KL objective encourages coverage of multiple high-value modes and thereby promotes exploration. Experiments and ablations on MuJoCo Playground and ManiSkill show competitive performance and sample-efficiency gains. Computational benchmarks also demonstrate faster actor updates than Relative Entropy Pathwise Policy Optimization (REPPO).
[AI-2] Higher-Order Molecular Grammars for Generative and Foundation Models in Chemistry
链接: https://arxiv.org/abs/2610.02186
作者: Yiming Huang,Yujie Zeng,Vijay Prakash Dwivedi,Simone Foti,Jianmin Wang,Jure Leskovec,Tolga Birdal
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:
Abstract:Molecular learning models are strongly shaped by their underlying representations. Yet standard sequential and graph formalisms struggle to explicitly encode higher-order topology, such as ring systems and recurring motifs. Existing higher-order representations can capture these structures directly, but they are often computationally demanding and difficult to decode into valid molecules. Here, we introduce Higher-order Grammar Representation (HGR), a principled, topology-aware framework that lifts molecules to combinatorial complexes and parses each complex into a compact sequence of production rules under a context-free higher-order grammar. By serialising higher-order topology into rule sequences, HGR makes these structures directly compatible with standard sequence models, avoiding the computational overhead of explicit higher-order encodings while preserving topological expressiveness. To reduce benchmark bias towards simple ring systems, we construct RingDiv, a ring-enriched benchmark containing 1.18 million molecules, including the curated RingDiv300k subset, and introduce the ring diversity index (RDI) to quantify ring-system coverage. In molecular generation, HGR-based models uniquely combine 100% validity by construction with leading distributional alignment, ranking first in FCD on all five generation benchmarks. In representation learning, HGR-FM achieves the highest mean AUC across seven MoleculeNet benchmarks under both transfer protocols, improving on the strongest baseline by 8.3 and 3.3 AUC points under probing and full fine-tuning, respectively. Collectively, these results establish HGR as an efficient higher-order representation for molecular generation and transferable representation learning.
[AI-3] SoftServe: A Scalable Quasi-Newton Method for Deep Learning
链接: https://arxiv.org/abs/2610.02182
作者: Joohwan Ko,Tetiana Parshakova,Diana Cai,Robert M. Gower
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:
Abstract:Quasi-Newton (QN) methods have long been among the most effective methods for large-scale unconstrained convex optimization. Two obstacles have limited their use in deep learning: non-convexity and enormous parameter sizes. We introduce SoftServe, a family of QN methods designed to overcome these obstacles without line searches or ad hoc curvature corrections. SoftServe derives positivedefinite curvature estimates from the variational objective of Berglund et al. (2025), even in the presence of negative curvature. We develop diagonal and Kroneckerfactored variants that preserve positive definiteness by construction and scale to massive neural networks. Finally, SoftServe relies on the stable coupled Newton-Schulz iteration for the required matrix operations, replacing costly matrix decompositions with GPU-friendly matrix multiplications. SoftServe excels on problems that are severely ill-conditioned, including tasks such as recurrent networks, deep autoencoders, physics-informed neural networks, and a 136M-parameter physics-informed diffusion model, often achieving lower losses than established baselines including Adam, Muon, and SOAP.
[AI-4] DuoMind: Enabling Distributed Multi-Robot Coordination with Semantic Communication
链接: https://arxiv.org/abs/2610.02161
作者: Hanchu Zhou,Dechen Gao,Hang Wang,Brendan Lynch,Boqi Zhao,Qiyao Ma,Raman Goyal,Junshan Zhang
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI)
备注:
Abstract:Vision-language models (VLMs) and vision-language-action models (VLAs) have recently driven rapid progress in general-purpose robots, yet most progress has focused on single-robot settings. Extending these capabilities to multi-robot systems remains challenging because robots must coordinate long-horizon behaviors while maintaining reliable, fine-grained execution. We introduce DuoMind, a distributed hierarchical framework for multi-robot coordination through semantic communication. Each robot uses a VLA-based action model for low-level execution and a VLM-based orchestrator for high-level reasoning and inter-agent coordination. At each planning step, the orchestrator at each robot reasons over the task instruction, local observations, and messages received from other robots. It then generates low-level instructions for the action model and semantic messages for peer robots. This architecture exploits the complementary strengths of pretrained models by combining the semantic reasoning capabilities of VLMs with the precise action-generation capabilities of VLAs. To address the scarcity of benchmarks for multi-robot coordination, we further develop RoboPoly, a benchmark comprising long-horizon manipulation tasks that require coordinated, closed-loop execution under distributed control. Experiments on RoboPoly and RoboTwin demonstrate that DuoMind improves multi-robot task performance, while ablation studies confirm the contributions of hierarchical orchestration and semantic communication. More details are available on our project page.
[AI-5] Local Support Learning
链接: https://arxiv.org/abs/2610.02126
作者: Assaf Ben-Kish,Akarsh Kumar,James Glass,Raja Giryes
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: Website and code: this https URL
Abstract:We explore catastrophic forgetting in the context of large pre-trained models. By considering forgetting as a geometric problem in the input space of each weight matrix, we uncover a natural retention objective under which updates produced by gradient-based optimizers are suboptimal. Following this observation, we propose Local Support Learning (LSL), a general-purpose framework that augments gradient-based training for retention of prior capabilities without access to prior data. During a new learning phase, LSL pairs two components with distinct roles: a standard weight adapter, trained as usual to minimize the loss, and a gating function that enables the adapter only on input activations from its own training distribution, making the update local to that distribution. The key challenge is that this gate must route data from all learning phases while training only on data from the current one. We address this with a gate based on a Gaussian Mixture Model (GMM), whose likelihood decays rapidly away from its training data, giving it a natural tendency to stay closed on data from prior phases. We show that this post-training approach can resolve forgetting in LLMs of up to 7 billion parameters, retaining both pretrained and finetuned capabilities across multiple training phases, while being efficient in memory and compute, robust to hyperparameter choice, and showing scaling potential.
[AI-6] HumanoidToolBench: Benchmarking Humanoid Tool Use from Selection to Mobile Execution
链接: https://arxiv.org/abs/2610.02089
作者: Kyochul Jang,Seohyeon Park,Ohchul Kwon,Sangjun Park,Junhyeok Choi,Seungyeop Yi,Chaeyun Kim,Sangkyu Lee,Idan Szpektor,Avi Caciularu,Jongmin Park,Youngjae Yu
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI)
备注: 9 pages, 7 figures
Abstract:As robotic hardware and learning methods advance, humanoids need tools to perform tasks beyond their inherent physical limits. Successful tool use requires selecting a suitable tool and coordinating manipulation and, when needed, locomotion to complete the task. Existing benchmarks do not jointly evaluate these capabilities on a humanoid. We introduce HumanoidToolBench, an 18-task benchmark spanning three scenarios, three execution levels, and two tool-set modes, together with ToolBook, a dataset of 3.1k demonstrations collected in simulation and on a real Unitree G1. Evaluation of seven policies in simulation and three on the real robot reveals substantial gaps between selecting a suitable tool and completing the task. Focused GR00T N1.7 probes show reduced selection accuracy on unseen tools and continued task execution under unrelated instructions. Code and data are available at this https URL.
[AI-7] Homomorphic Advantage Operator: Stabilizing Reinforcement Learning Under Fully Homomorphic Encryption Constraints
链接: https://arxiv.org/abs/2610.02074
作者: Abid Mohamed Nadhir,Ahmad Al Hanbali,Beggas Mounir
类目: Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR)
备注:
Abstract:Privacy-preserving machine learning presents significant deployment challenges on the cloud for intelligent systems with confidential data. Fully Homomorphic Encryption (FHE) offers a compelling solution for secure computation, preserving data confidentiality of cloud computations. However, applying FHE to reinforcement learning (RL) requires replacing non-linear operations with polynomial approximations, which diverge catastrophically due to a unique recursive error phenomenon known as the Bellman drift. This article introduces the Homomorphic Advantage Operator (HAO), a stabilization framework designed to prevent polynomial approximation divergence in FHE-based deep RL. HAO adapts the zero-mean centering projection from advantage-based value estimation directly to temporal-difference (TD) targets. This linear projection annihilates the uniform state-value baseline that drives the Bellman drift, maintaining per-state action rankings while requiring zero additional non-linear multiplicative depth and avoiding expensive ciphertext bootstrapping. The proposed HAO framework was evaluated using a three-tier experimental methodology, including a tabular Markov Decision Process (MDP), an encrypted CartPole environment using real CKKS cryptographic operations, and a 20-node logistics routing benchmark with dense continuous features. The results demonstrate that the proposed HAO strictly bounds network pre-activations within the safe polynomial approximation domain. The proposed HAO RL agents achieved 0% boundary breaches across all random seeds used, whereas regularization alone (L2 weight decay and gradient clipping) breached the bound on 3 of 5 seeds and the unstabilized baseline did so in 83.8% of episodes. Finally, HAO agents improve optimal policy accuracy by 18.0 percentage points in tabular domains and remain stable when DP-SGD-style Gaussian noise is added to the clipped gradients.
[AI-8] PyPottery: an AI-powered end-to-end suite for pottery processing and publication
链接: https://arxiv.org/abs/2610.02072
作者: Lorenzo Cardarelli
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:The study of ceramic materials constitutes a cornerstone of archaeological research, yet the post-production workflow for pottery documentation remains labor-intensive and creates significant publication bottlenecks. This paper presents PyPottery, an open-source, AI-powered suite designed to semi-automate the complete ceramic documentation pipeline. The suite comprises four integrated modules: PyPotteryScan for automated image extraction and handwriting recognition; PyPotteryInk for automatic inking of pencil drawings; PyPotteryTrace for semantically-aware vectorization; and PyPotteryLayout for automated layout generation. Evaluated on 50 hand-drawn sheets containing 240 pottery drawings from the Terramara di Montale (Italy), the framework achieved substantial time savings confirmed by usability study participants, who reported a median perceived speedup of 40 \times over traditional workflows (range: 17.5 \times --120 \times ). These results highlight the potential of AI-assisted tools in archaeological documentation, while the paper addresses the strategic redistribution of cognitive labor toward augmentation rather than automation.
[AI-9] Causal Memory Policy: Making Memory Utility Identifiable by Intervening on Retrieval
链接: https://arxiv.org/abs/2610.02070
作者: Arman Behnam,Binghui Wang
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Memory-augmented large language models must decide which memories to retain, and recent systems do so by estimating each memory’s effect on task performance. However, these estimates rely entirely on retrieved memories. When a memory is never retrieved, store-level interventions produce identical outcomes, leaving its utility unidentified. This is a retrieval-level positivity violation, invisible to diagnostics that examine only memory operations. We introduce Causal Memory Policy (CMP), a causal framework that restores identification by intervening on retrieval itself, reserving a fixed number of context slots for memories sampled with known propensities. CMP estimates memory utility by self-normalized inverse propensity weighting under a balanced assignment design. We prove the causal factorization of memory utility through retrieval, the unbiasedness and exact variance of the estimator, and the optimal decision rule under irreversible operations. Empirically, identification fails for 54% of required memories on LongMemEval and 67% on LoCoMo, and the failure persists in a deployed memory system. CMP improves discrimination between required and non-required memories from 0.54 to 0.66 AUC. Finally, we show that identified memory utility alone is insufficient for retention decisions: per-query utility reaches 0.78 AUC on the query for which it is estimated, yet no aggregation available to a retention policy predicts a memory’s value on unseen queries. Code is available at: this https URL.
[AI-10] External Observers May See More Clearly: Cross-Model Span-Level Hallucination Detection in Large Language Models via Hidden State Probing
链接: https://arxiv.org/abs/2610.02066
作者: Kingshuk Gupta,Davide Buscaldi
类目: Artificial Intelligence (cs.AI)
备注: 12 pages, 2 figures, 9 tables
Abstract:As Large Language Models (LLMs) increasingly serve as foundational reasoning engines, their tendency to hallucinate remains a critical vulnerability. While recent internal state probes offer a promising alternative to slow external retrieval systems, they largely reduce hallucination detection to a token-wise binary classification task, failing to capture the structured, sequential boundaries of semantic drift. Here, we introduce an internal hidden state framework for fine-grained, span-level hallucination detection. By inspecting layer-wise activation patterns, we attempt to detect the exact hallucination onset and continuation tokens in an LLM generation. Our experiments show that this approach successfully isolates hallucination onsets, achieving substantial improvements in Precision-Recall AUC over random baselines despite extreme class imbalance. Ultimately, we propose a novel cross-model detection framework in which one model observes the internal representations elicited by another model’s generation. We find that an external observer can match or exceed a generator’s self-detection of its own hallucination onsets, including when the observer is the smaller model, suggesting that self-detection is not the ceiling for onset localisation.
[AI-11] HydroJEV: A one-second training-free screen for cyber-attack and fault attribution in water distribution networks
链接: https://arxiv.org/abs/2610.02048
作者: Tianwei Mu,Shengyan Jiang,Mingzhe Yuan,Qing Luo,Min Xiao,Wenhong Wang,Jun Li,Manhong Huang
类目: Artificial Intelligence (cs.AI)
备注: 41 pages, 19 figures
Abstract:When a SCADA alarm is raised in a water distribution network, operators must decide quickly whether it reflects a cyberattack, a physical fault, a normal transient or a faulty sensor. Supervised classifiers need labelled incidents that utilities rarely have, and frontier large language models (LLMs) take tens of seconds per decision. We tested whether Jev, a training-free model that returns class probabilities in about one second, can serve as the first tier of this triage. On a four-class cause-attribution benchmark built on the C-Town network in EPANET, Jev was compared with a hand-written rule tree, a supervised classifier and seven cloud LLMs on identical evidence in four sealed, pre-registered rounds. With only a label-free prior correction, Jev matched the rule tree (macro-F1 0.62-0.64 against 0.56-0.61 in distribution) and exceeded the supervised classifier by 0.36-0.42 on event subtypes absent from its labels, in all four rounds, and it outperformed the classifier whenever fewer than about four labelled events per class were available. Jev also decided 20-40 times faster than frontier LLMs. Accepting only benign Jev verdicts confirmed by the rule tree spared an LLM reviewer 35-38% of windows on fresh sealed sets without loss of macro-F1. Transferred unchanged to two further networks, this gated cascade stayed within the non-inferiority margin of its reviewer on all four sets. A fast, training-free screen can therefore take over about a third of the review load in SCADA anomaly triage while preserving the accuracy of deliberate review.
[AI-12] Distributionally Robust Schrödinger Bridge
链接: https://arxiv.org/abs/2610.02043
作者: Jinhwan Sul,Panagiotis Theodoropoulos,Vincent Pacelli,Jaemoo Choi,Evangelos Theodorou
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: 30 pages, 5 figures
Abstract:Schrödinger bridge (SB) learns stochastic transport between prescribed initial and target distributions. When the initial distribution shifts at test time, the learned dynamics can fail to recover the target distribution. We introduce the Distributionally Robust Schrödinger Bridge (DRSB), which learns a single controller that accounts for uncertainty in the initial distribution. The DRSB objective consists of control energy and a KL penalty between the resulting terminal distribution and the target distribution. DRSB seeks a single controller that minimizes the worst-case value of this objective as the initial distribution varies within an ambiguity set around the nominal distribution. We derive an exact variational formulation of this objective and connect its fixed-terminal-cost subproblem to stochastic optimal control and distributionally robust optimization. This formulation motivates an alternating algorithm that updates the adversarial initial distribution, estimates the terminal log-density ratio, and trains the controller. We develop Wasserstein and Sinkhorn variants using stochastic control optimality conditions to approximate the gradients required for adversarial updates. Experiments on two-dimensional transport tasks and image-to-image translation show improved robustness to input perturbations relative to standard SB, with a tradeoff in nominal performance. On Gaussian mixture transport, Sinkhorn DRSB also achieves lower mean sliced Wasserstein distance than fixed-level noise augmentation at both tested unseen noise levels.
[AI-13] Mimir: Physics-Grounded LLM Agents for Long-Horizon Irrigation Control
链接: https://arxiv.org/abs/2610.02038
作者: Yimeng Liu,Mi Zhang,Younsuk Dong,Zhichao Cao
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Large language model (LLM) agents increasingly combine reasoning, tool use, and action, but most evidence comes from episodic tasks with relatively immediate feedback and reset failures. Long-running physical control operates in a different regime: actions alter future states, errors compound across decisions, and an agent must improve from experience without being allowed to rewrite the physical rules that make execution safe. We study this regime through irrigation, where daily decisions interact with soil-water dynamics over entire growing seasons. We present Mimir, a physics-grounded LLM agent organized around two repair timescales. At the fast timescale, a structured physical interface and deterministic simulator turn an LLM output into a proposal that we numerically check, revise, and subject to bounded deterministic action selection before execution. At the slow timescale, recurrent failure patterns are consolidated into persistent contextual principles that condition future proposals, while the physical model, evaluator, and execution constraints remain immutable. Under a common retrospective evaluator across multiple sites, crops, and years, Mimir attains the lowest reported aggregate control cost among the evaluated references and uses about 51% less irrigation than the historical schedule replay. The ablation study show higher control cost when forward simulation, verified revision, or persistent context is removed; model-scale and model-family studies show no monotonic gain from increasing LLM size. The resulting lesson show that persistent physical agents can combine semantic reasoning with bounded, evidence-driven self-improvement while reserving physical truth and actuator authority for explicit numerical mechanisms.
[AI-14] Global Coherence: When Every Agent Is Right and the Team Is Still Wrong - A Local-to-Global Semantic Foundation for Multi-Agent Collaboration
链接: https://arxiv.org/abs/2610.02036
作者: Xin Heng
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:AI agents can each make locally valid decisions yet jointly produce an invalid result. We call this the global coherence problem: a failure of shared state, not merely of model intelligence. Our Observation-Aliasing Impossibility Theorem gives the exact boundary. A policy can guarantee a valid action exactly when all worlds producing the same observation share an admissible action. If k indistinguishable worlds require pairwise-disjoint actions, the best randomized worst-case success is 1/k; more reasoning, roles, messages, or samples cannot recover the missing distinction. A stronger model can reason better within its context, but it cannot see beyond it. We then give local-to-global runtime semantics X = (H, C, G, F; D): topology H records overlapping scopes; category C governs state-changing actions; groupoid G retains reversible translations; sheaf F tests whether local views glue into one world; and minimal history D keeps only distinctions that alter legal futures. Models propose; the harness owns shared state and governs commit. Nine studies test both the failure and its boundary. On a controlled revision benchmark, the same frontier model scores 40/40 when the deciding event is visible; when it is hidden, tested arms score 12–17/40, consistent with chance (1/3); restoring one authoritative fact returns 40/40. On TeamBench, ordinary teams exceed a shared budget in 5/5 runs, a visible live count leaves 4/5 violations, and commit enforcement leaves 0/5. In tau2-bench Telecom, current-state checks score 0.07 after silent reverts, while the harness scores 1.00. Where a conventional solver already owns the complete relevant state, it ties the harness as predicted. The counterintuitive conclusion is that local intelligence cannot substitute for missing global state. Subjects: Artificial Intelligence (cs.AI) Cite as: arXiv:2610.02036 [cs.AI] (or arXiv:2610.02036v1 [cs.AI] for this version) https://doi.org/10.48550/arXiv.2610.02036 Focus to learn more arXiv-issued DOI via DataCite (pending registration) Submission history From: Xin Heng [view email] [v1] Thu, 1 Oct 2026 16:49:40 UTC (96 KB)
[AI-15] On Language Drift during RLVR Post-Training
链接: https://arxiv.org/abs/2610.02015
作者: Michael Sullivan,Alexander Koller
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: 22 pages; 15 figures; 4 tables
Abstract:Recent advances in LLM reasoning models—driven primarily by the paradigm of post-training via reinforcement learning with verifiable reward (RLVR)—have enabled them to accomplish impressively complex tasks. However, in parallel with their rising capabilities, LLMs have increasingly displayed signs of language drift in their chains of thought (CoTs): unusual, non-standard, and seemingly nonsensical language use. Although it is well-documented—and can potentially impair CoT monitorability—the causes of language drift are thus far poorly understood. In this paper, we identify the conditions under which language drift occurs: we prove theoretically that RLVR optimization pressure permits unbounded language drift, while supervised fine-tuning does not. We then show empirically that language drift specifically arises during RLVR on novel reasoning tasks—i.e. when the target behavior cannot be drawn out of the base model. Finally, we prove that it is not possible to constrain language drift without constraining expected reward, suggesting that CoT monitorability cannot be improved without harming performance during RLVR post-training at the frontier.
[AI-16] Atoms to Processes: The Role of Artificial Intelligence and Machine Learning in Chemical Engineering
链接: https://arxiv.org/abs/2610.02014
作者: Michael Baldea,Linda J. Broadbelt,Marianthi G. Ierapetritou,Akhilesh Jain,Ankur Kumar,Thomas A. Kwan,Fèlix Llovell,Andrew J. Medford,Ilias Mitrai,Joel Paulson,Junyi Qiao,Matthew P. Rivera,Kirti C. Sahu,Lev Sarkisov,Zachary P. Smith,Calvin Tsay,Ching-Mei Wen,Victor M. Zavala,Huacheng Zhang,Dan Zhao
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:The rapid maturation of artificial intelligence (AI) and machine learning (ML) has catalyzed a profound shift in how chemical engineering problems are formulated, analyzed, and solved. Advances in computing, data availability, and learning algorithms have enabled AI/ML methods to impact applications spanning atomic-scale simulations, materials and catalyst discovery, transport and thermodynamics, separations, process systems engineering, and industrial operations. This article provides a perspective on recent methodological developments and representative applications, emphasizing how AI/ML tools are being integrated with first-principles models to address challenges of predictive accuracy, data scarcity, extrapolation, interpretability, and model lifecycle management. Across domains, a unifying trend is the move away from purely black-box approaches toward hybrid and physics-informed frameworks that explicitly respect conservation laws, thermodynamic consistency, and known structural constraints. These approaches not only improve robustness and reliability, but also enable meaningful human-AI collaboration by providing information at an appropriate level of abstraction for the task and decision context. We conclude that AI and ML are not replacing the core principles of chemical engineering; rather, they are amplifying them. As the field advances toward increasingly autonomous, adaptive, and sustainable systems, the thoughtful integration of AI/ML with first-principles understanding and domain expertise will be essential to realizing their full potential across both research and industrial practice.
[AI-17] Counting Moves Weighing Voices: Bayesian Dialectical Argumentation for Calibrated Multi-LLM Councils under Persistent Adversaries
链接: https://arxiv.org/abs/2610.02005
作者: Ionel Eduard Stan,Paolo Napoletano
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:A multi-LLM \emphcouncil lets several large language models (LLMs) deliberate on a question and return an answer together with a confidence estimate. As these systems become increasingly used for reasoning, that confidence should represent a calibrated \emphprobability of being correct, and the decision should remain robust when some agents are persistently unreliable. Existing \emphcouncil aggregation methods fail on both fronts: their confidence estimates measure decisiveness rather than correctness, and they cannot identify or discount persistently unreliable agents. We introduce Bayesian Dialectical Argumentation (BDA), which treats the council’s \emphtyped moves—who proposed, challenged, or conceded which answer—as observations of a classical annotator model with \emphper-agent reliabilities. This formulation recasts multi-agent deliberation as a reliability estimation problem, using the deliberation trace to infer agent reliability under persistent adversarial behavior. By weighting evidence according to inferred agent reliability, BDA yields calibrated posterior probabilities over candidate answers while allowing persistently unreliable agents to be inverted rather than merely outvoted. Across binary and multi-class benchmarks, BDA achieves the best calibration among zero-cost council aggregation methods, requiring no additional LLM calls, and improves robustness under persistent adversarial coalitions while remaining competitive in clean settings.
[AI-18] Can AI Oversight Be Zero Knowledge?
链接: https://arxiv.org/abs/2610.01995
作者: Alessandro Chiesa,Ziyi Guan,Burcu Yildiz
类目: Artificial Intelligence (cs.AI); Computational Complexity (cs.CC); Cryptography and Security (cs.CR)
备注:
Abstract:AI systems increasingly produce outputs from confidential data, such as a fitness-for-duty assessment from medical records or the predicted properties of a drug candidate from its secret structure. It is important to verify that such outputs are correct without revealing the underlying data. A recent line of work studies verification of AI outputs via interactive proofs and debate for oracle-aided computation, where correctness may depend on an oracle such as human judgment, a physical experiment, or the web. These works focus on verification by a verifier that runs much faster than the computation. However, such efficient verification is impossible for general oracle-aided computation, and these works therefore rely on additional assumptions. We focus instead on privacy: allowing the verifier to run in time polynomial in the computation, we ask whether interactive arguments for oracle-aided computation can be zero knowledge, so that the verifier learns nothing about the confidential data beyond the correctness of the output. We prove that, in general, they cannot. In the random oracle model, there are no zero-knowledge proofs for all oracle-aided computations, even if both the prover and the verifier are allowed to run much longer than the computation itself. The impossibility extends to debate, a canonical model for scalable oversight. On the positive side, we show that if the oracle attaches a cryptographic signature to each of its answers, then every oracle-aided computation can be verified in zero knowledge with an efficient prover and verifier, assuming only collision-resistant hash functions. Beyond privacy, this also gives an alternative approach to scalable oversight that relies neither on an honest opponent, as in debate, nor on the robustness of the computation, as in prior single-prover protocols. Subjects: Artificial Intelligence (cs.AI); Computational Complexity (cs.CC); Cryptography and Security (cs.CR) Cite as: arXiv:2610.01995 [cs.AI] (or arXiv:2610.01995v1 [cs.AI] for this version) https://doi.org/10.48550/arXiv.2610.01995 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[AI-19] A Hybrid Approach to Malware Detection: Integrating Few-Shot Model-Agnostic Meta-Learning with Autoencoders
链接: https://arxiv.org/abs/2610.01949
作者: Emmanuela Andam,Yasir Abbas Zaidi,Abdelali Hadir,Emmanuel Grant,Naima Kaabouch
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: Accepted at 2025 Cyber Awareness and Research Symposium (CARS). This is the author’s accepted manuscript
Abstract:Ransomware has emerged as a major cybersecurity threat, with incidents increasing in frequency and impact across critical sectors. These attacks are typically launched through phishing emails, malicious downloads, or exploitation of software vulnerabilities to gain system access. Once inside, the malware encrypts files and demands a ransom, often in cryptocurrency, for the decryption key. Conventional detection methods often struggle with novel or scarce samples, leaving systems vulnerable. To address these challenges, this paper proposes a hybrid deep learning framework that combines an Autoencoder Feature Extractor (AFE) with a Model Agnostic Meta Learning (MAML) classifier for few shot malware detection. The AFE generates compact latent features that reduce noise and dimensionality, while the MAML classifier rapidly adapts to new threats using limited labeled data. Experiments conducted on the Ransomware Dataset 2024 demonstrate the effectiveness of the framework in binary classification tasks. Across one to fifty shot settings, the proposed model consistently achieves high accuracy, F1 score, and Matthews Correlation Coefficient values, maintaining reliable classification even under extreme scarcity. These results highlight the model’s robustness and effectiveness in adapting to limited data scenarios, demonstrating the potential of combining feature extraction with meta learning to enhance resilience against malware, particularly in sectors such as healthcare, manufacturing, and public infrastructure, where cyberattacks can cause significant operational and financial disruption.
[AI-20] Mapping the RAG Landscape: A Four Axis Taxonomy of Efficiency Defense Interactivity and Reasoning
链接: https://arxiv.org/abs/2610.01936
作者: Meghana Sunil,Shravya V,Shravan Venkatraman,Joe Dhanith PR
类目: Artificial Intelligence (cs.AI)
备注: published in Artificial intelligence reviews
Abstract:Large Language Models (LLMs) have demonstrated remarkable fluency across many tasks but remain limited by their static, parameter bound knowledge and their susceptibility to hallucinating information. Retrieval Augmented Generation (RAG) addresses these issues by incorporating external retrieval into the generation process, grounding model outputs in verifiable and up to date sources. While prior surveys primarily focus on core RAG architectures and standard pipelines, recent research explores broader challenges and capabilities that extend beyond these foundational designs. This survey provides a consolidated and structured examination of contemporary RAG developments, organizing the field into a four axis taxonomy: improving retrieval efficiency, strengthening robustness and security, supporting user driven and interactive workflows, and enabling multi step or complex reasoning. We formalize key components of the RAG framework and review methods spanning dense and sparse retrieval, fusion strategies, embedding optimizations, and reinforcement learning based retrieval policies, highlighting how these advances influence practical deployment and system design. We also synthesize evaluation practices, domain specific applications, and architectural variants such as Naive, Advanced, and Modular RAG. Finally, we outline persistent challenges related to retrieval quality, reliability, domain adaptation, scalability, and explainability, and identify opportunities for building RAG systems that are more reliable, adaptable, and transparent.
[AI-21] Asynchronous LLM Post-Training: Group-Mass Capping and Convergence Analysis
链接: https://arxiv.org/abs/2610.01896
作者: Qijia He,Ruinan Jin,Jun Luo,Shaofeng Zou,Yingbin Liang
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: 40 pages, 6 figures
Abstract:Asynchronous reinforcement learning (RL) improves the efficiency of large language model post-training but introduces stale rollouts generated by earlier policies. Theoretical understanding of how this staleness affects convergence and how to mitigate its impact remains limited. We derive a convergence bound for GRPO-style algorithms that explicitly characterizes the tradeoff between the gradient estimator’s second moment and bias. For trajectory-level importance-weighted estimators, our analysis shows that once the second moment is uniformly controlled, delay enters the bound through the bias introduced by clipping or rescaling. Guided by this insight, we propose a novel group mass capping GRPO (GMC-GRPO) method, which minimizes a ratio-based bias bound within a class of weighted estimators sharing a common second-moment guarantee. We establish convergence guarantees for asynchronous GMC-GRPO and show that, compared with TIC-GRPO, it improves the threshold dependence of the fourth-order delay term from O(\epsilon^-4) to O(\epsilon^-2) as \epsilon\to0 , where 1+\epsilon is the ratio threshold. Under local policy overlap, the delay-dependent term decreases as G^-2/5 after tuning the step size, where G is the group size. For fixed behavior and current policies, the bias introduced by group rescaling also vanishes as G\to\infty , whereas the bias from trajectory-wise clipping can persist. Experiments across Qwen3 models and reasoning benchmarks demonstrate improved robustness to stale rollouts, with GMC-GRPO achieving the best performance among stable baselines under large rollout delays.
[AI-22] A Structured State Space Sequence Model for Multi-Class Classification of Malware
链接: https://arxiv.org/abs/2610.01893
作者: Emmanuela Andam,Rana Shaaban,Emanuel Grant,Naima Kaabouch
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: Accepted at 2026 IEEE World AI IoT Congress (AIIoT). This is the author’s accepted manuscript
Abstract:By 2030, Internet of Things (IoT) devices are projected to reach 40 billion, with fast-paced technological advancements in fields such as industry, healthcare, agriculture, automobiles, and building/home automation systems. This expansion has created a large attack surface for cybercrime, as the majority of these devices open the door for cybercriminals to exploit vulnerabilities, as they lack adequate built-in security. Cybercriminals launch malware attacks to compromise systems or steal sensitive data, and once a system is compromised, a ransom is typically demanded for its release. Current cybersecurity measures in place are being outpaced by the rapid growth of the IoT, which is accompanied by a subsequent growth in malware variants being created per day. Recognizing this pitfall, this research examines and proposes a novel approach to malware detection and classification to safeguard devices from further attacks and make IoT systems more robust and secure. The framework proposed utilizes a Structured State Space Sequence (S4) model, which discretizes sequences of malware samples in a sequence and captures long-range dependencies, essentially identifying the “cause” and “effect” hidden within malware execution flow. This study presents two novel contributions: the first empirical application of the S4 model for malware analysis, and a comprehensive comparison of its performance against other deep learning architectures, laying the stepping stone for future research in this new paradigm.
[AI-23] Selection-Based Structured Reasoning : Toward Efficient Multimodal Search Agents
链接: https://arxiv.org/abs/2610.01892
作者: Feiyu Gavin Zhu,Xiaoyu Zhu,Jiqi Yang,Rui Yang,Arnab Kumar Mondal,Yancheng Wang,Xinke Deng,Jean Oh,Reid Simmons,Joerg Liebelt,Xiang Kong,Zhongyu Jiang
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:
Abstract:Multimodal agents commonly generate free-form reasoning before each action. For small models, limited model capacity can result in lengthy reasoning that provides little useful guidance for action generation while incurring substantial inference cost. To address this challenge, we introduce Selection-based Structured Reasoning (SSR), a framework that reformulates reasoning as selection instead of open-ended generation. SSR represents recurring high-level reasoning as pre-specified, reusable natural-language candidates. At each turn, the model selects from these reasoning candidates based on their likelihoods given the current context, without requiring an auxiliary task head. Using pre-specified reasoning traces enables parallel scoring, where teacher-forced prefilling computes token likelihoods concurrently within and across candidates using a shared context KV cache. We evaluate SSR on seven multimodal search benchmarks using 2B and 4B models. Across multiple reinforcement learning objectives and supervised fine-tuning, SSR delivers significant efficiency gains without sacrificing task performance. SSR achieves an average success rate competitive with leading search agents of the same scale, while reducing per-turn reasoning latency by over 90% and total per-question model inference latency by 28-54%. Project page: this https URL.
[AI-24] Flowing Faster to Coordinate: One-Step Online Multi-Agent Flow Policies
链接: https://arxiv.org/abs/2610.01882
作者: Zhuoran Li,Yunzhan Li,Xun Wang,Yihan Du,Longbo Huang
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:
Abstract:Multi-agent reinforcement learning (MARL) provides a powerful framework for learning coordinated behaviors through interactions with the environment. Developing MARL policies requires balancing expressive modeling of complex and multimodal action distributions with efficient training and execution. Generative policies, particularly diffusionbased policies, can faithfully capture complex and multimodal behaviors, but costly iterative sampling hinders their scalability in online multi-agent settings. We propose an Online MARL framework via one-step Flow model (OMAF) that combines expressive generative policies with efficient one-step action generation. OMAF employs a Transformer-based flow policy to capture complex coordination behaviors, while its approximate path score surrogate provides a principled route to synchronized flow policy optimization. To enable stable and sampleefficient learning, we further develop a joint optimization scheme coupling softmax Q-value estimation with a joint flow policy objective for coordinated policy learning. By eliminating iterative sampling, OMAF dramatically reduces training overhead without sacrificing policy expressiveness. Extensive experiments across 10 standard tasks from MPE and MAMuJoCo show that OMAF consistently achieves superior performance, with up to 3.4x higher returns and 10.5x sample efficiency improvement compared with baseline methods. These results validate the effectiveness of OMAF as an expressive and computationally efficient one-step flow policy paradigm for online MARL.
[AI-25] From Network Intrusion Detection to Blockchain-Backed Endpoint Detection and Response: Mapping the Landscape of Decentralized Detection-and-Response Architectures
链接: https://arxiv.org/abs/2610.01872
作者: Yahya Shahsavari,Sara Rouhani,Kaiwen Zhang
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Networking and Internet Architecture (cs.NI)
备注:
Abstract:While the literature on blockchain-assisted intrusion detection and prevention systems (IDS/IPS) for Internet of Things (IoT) and Industrial Internet of Things (IIoT) networks is mature, existing systematic reviews suffer from two critical limitations: they overlook the structural shift toward modern Endpoint Detection and Response (EDR) and Extended Detection and Response (XDR) architectures, and they conflate blockchain’s distinct functional roles into a single monolithic category. This Systematization of Knowledge (SoK) addresses these gaps by proposing a three-axis taxonomy that classifies proposals by detection-system class (NIDS, HIDS, EDR/XDR), blockchain functional role, and response-automation maturity. Synthesizing research published in high-impact venues between 2019 and 2026, we provide a rigorous gap analysis exposing why a genuine per-endpoint blockchain-anchored response loop remains nearly nonexistent due to latency, deployment, and community mismatches. Furthermore, we evaluate structural, cross-cutting challenges persisting across the literature, including consensus latency on constrained devices, post-quantum cryptographic vulnerability, smart-contract attack surfaces, and the adversarial vulnerability of evolving LLM-based detection engines. Finally, we outline a comprehensive research agenda centered on hybrid on-chain/off-chain orchestration to bridge the gap between decentralized trust and rapid response automation.
[AI-26] Walking the Embedding Space: Datastore Extraction from Multimodal RAG
链接: https://arxiv.org/abs/2610.01871
作者: Maria Carmen Jica,Ali Satvaty,Suzan Verberne,Fatih Turkmen
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
备注:
Abstract:Multimodal Retrieval-Augmented Generation (MRAG) has emerged as a reliable and cost-effective technique of grounding the generative capabilities of Multimodal Large Language Models (MLLMs) into relevant, up-to-date, external knowledge. Despite presenting several benefits, such as reducing hallucinatory behavior, they also introduce new attack surfaces, including leakage of private information and vulnerabilities against data extraction attacks. In this paper, we introduce \immrag , an adaptive and automatic data extraction attack procedure operating in a black box setting against \emphimage-returning MRAG, a configuration in which the retrieved visual artifact is itself the response. Each query blends an attacker-held shadow image with an image already recovered from the system, and relevance-weighted resampling steers subsequent queries towards regions of the embedding space that still yield novel retrievals. Unlike current extraction attacks that aim to persuade the model towards data leakage by placing a malicious query as a textual prompt, \immrag embeds the malicious instructions inside a user-given input image. We evaluate \immrag on three plausible and distinct real-world scenarios: medical assistant, document-focused helper and general purpose tool. The experiments involve the study of the effectiveness of the attack on multiple CLIP-family retrievers, as well as the impact of various generators. A single 2500-query run reconstructs up to 611 distinct radiology images, 566 document scans and 416 general-purpose images under local-feature correspondence, and reaches up to 5.6\times as many distinct datastore items as a non-adaptive baseline. Our results show the urgent need for safeguards specifically designed for multimodal data. Subjects: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI) Cite as: arXiv:2610.01871 [cs.CR] (or arXiv:2610.01871v1 [cs.CR] for this version) https://doi.org/10.48550/arXiv.2610.01871 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[AI-27] From Isolated Feature to Orbits: Discovering Music Concepts via Multi-SAE Alignment
链接: https://arxiv.org/abs/2610.01864
作者: Liwei Lin,Gus Xia
类目: ound (cs.SD); Artificial Intelligence (cs.AI)
备注:
Abstract:How can we understand what a music foundation model has learned \textitinternally? Most interpretability approaches, such as probing and Sparse Autoencoders (SAEs), focus on identifying individual features with minimal structural assumptions. We argue that many concepts are better understood as \textitstructured relations rather than isolated features. This is especially prominent in music, where tonal structures are organized in the space of pitch and time. For example, concepts such as chords or keys are naturally expressed as structured sets (e.g., the 12 transpositions of a chord or the diatonic system within a key), rather than isolated features. In this study, \textbfwe shift from feature identification to structure-based analysis, asking whether the learned inner representations of music foundation model emerge as organized structures over features. To this end, we introduce a framework that uses pitch transposition as an inductive bias to induce ordered orbits via multi-view SAE alignment. Concretely, we generate pitch-shifted input pairs and align their SAE representations to discover structured groups of pitch-related features. Experimental results show that this approach recovers orbit structures corresponding to chords, keys, and melodic patterns across two state-of-the-art music foundation models, while requiring only minimal grounding (e.g., a few anchor examples) to interpret entire concept families.
[AI-28] AVSD-Scenes: A Dataset for Audio-Visual Description of Urban Scenes ICASSP2027
链接: https://arxiv.org/abs/2610.01861
作者: Dhanunjaya Varma Devalraju,Arshdeep Singh,Mark D. Plumbley
类目: Artificial Intelligence (cs.AI); Sound (cs.SD); Audio and Speech Processing (eess.AS)
备注: Submitted to ICASSP 2027
Abstract:Natural language descriptions can provide rich semantic representations of audio-visual urban scenes, yet datasets that jointly describe both auditory and visual information remain limited. In this paper, we introduce AVSD-Scenes, a paired audio-visual scene description dataset for urban environments. The dataset contains 12,291 audio-visual scene descriptions generated from the TAU Urban Audio-Visual Scenes dataset. To construct the dataset, we first generate audio- and visual-based descriptions using Qwen2-Audio-7B and Qwen2.5-VL-7B, respectively. These modality-specific descriptions are then combined using large language models, namely Qwen3-14B, Mistral-Small-3.2-24B-Instruct-2506, and Gemma-3-27B-it, to produce multimodal descriptions that capture complementary information from both modalities. We benchmark AVSD-Scenes using semantic alignment, cross-modal retrieval, scene classification, LLM-as-a-judge evaluation, and human subjective assessment. Results show that multimodal descriptions improve semantic alignment and cross-modal retrieval performance compared with modality-specific descriptions while preserving strong scene-discriminative information. The generated descriptions achieve up to 94.5% accuracy in urban scene classification, while combining audio, visual, and description embeddings further improves accuracy to 95.4%. Furthermore, the descriptions remain highly scene-discriminative even when scene labels are removed from the prompting instructions, indicating that they capture semantic information derived from the audio-visual content rather than merely reflecting label information.
[AI-29] mporal-Difference Learning for Drag onchess
链接: https://arxiv.org/abs/2610.01845
作者: Jim O’Connor,Annika Hoag,Sarah Goyette,Gary B. Parker
类目: Artificial Intelligence (cs.AI)
备注: Springer Lecture Notes in Artificial Intelligence
Abstract:Our research investigates how two adaptive AI methods, evolutionary transfer learning and TD(lambda), perform in the three-dimensional chess environment Dragonchess. The game challenges players with its unique board structure and computational load, making it an ideal setting to study how adaptive methods can update evaluation heuristics in novel environments. In this work we re-implement the Dragonchess engine, changing it from a PyGame engine to C++. This enables faster gameplay, allowing us to run 10,000 games with confidence intervals and significance tests, rather than a single small tournament. Both adaptive methods outperform all other agents in the round-robin tournament. Our results showed that there is no significant difference in the performance between the evolved and learned evaluations. This research establishes the efficacy of adaptive methods in structurally complex, novel game domains.
[AI-30] On the Divergence of Accuracy and Mechanism Consistency in Time Series World Models
链接: https://arxiv.org/abs/2610.01842
作者: Haochen Zhang,Jiaheng Guo,Zhen Xu,Zachary Plotkin,Nicholas Konz,Zhen Tan,Tianlong Chen
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:A time series world model (TSWM) predicts a controlled system’s state from its observed history and planned actions and exogenous inputs. Current approaches build forecasters with actions as covariates, trained and evaluated on prediction error under the executed plan. Yet world models compare unexecuted plans, but their responses to changed plans remain untested. We ask which design choices matter and whether accurate forecasters respond to changed plans as real systems do. We address both with a formalization and benchmark. The formalization separates state, actions and exogenous inputs, distinguishes continuous, mode and event actions, and introduces mechanism consistency, a metric built on declared action-state relations with known directions, such as a vasopressor raising blood pressure: it checks whether shifting an action moves the forecast in the declared direction. The benchmark consolidates eight public datasets with real actions from engineered infrastructure and clinical care, varying prediction space, plan fusion and plan encoding across seven backbones and five seeds. First, a frozen latent prediction space lowers MAE by 9.9% over observation space and gated output fusion lowers it by 12.7% over input concatenation on average, with both improving all eight datasets; temporal plan encoding changes average MAE by at most 2.2%. Second, prediction error and mechanism consistency diverge: the lowest-error configuration is at or below chance in consistency on four of five datasets with declared mechanisms, and no design choice avoids this. Finally, directional supervision, a loss penalizing the wrong-signed part of the response to a shifted action, significantly raises consistency on penalized mechanisms with no change in MAE. Together they give TSWMs a recipe: a frozen latent space and output-side fusion for accuracy, and a training objective for mechanism consistency.
[AI-31] Code Owns the Simulation Jev Owns the Evaluation
链接: https://arxiv.org/abs/2610.01834
作者: Yaodong Yang,Hongyao Tang,Yi Ma,Xingyu Fan,Weixun Wang,Jinpeng Li,Tianpei Yang
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: 10 pages main text, 20 pages total with appendix; 6 figures, 7 tables. Preprint
Abstract:Judgment models such as \jev return, in a single call and without reasoning text, a probability for each described option. This makes them attractive as an agent’s action-selection layer, but it is unclear which decisions they can be trusted with. We test \jev on reflection tests, one-shot matrix games, the text game ALFWorld and robot control, and find a sharp boundary. \jev succeeds when the right option can be judged from what the input describes, which we call \emphevaluation. Specifically, it solves 99% of the counterintuitive Cognitive Reflection Test questions. However, it fails when the right option depends on \emphsimulation (i.e., predicting something not in the input), such as the opponent’s action or the subgoal that must come first. In games, \jev plays suboptimally as if its rational opponent acted at random, because the opponent’s action is not given. In ALFWorld, \jev favors commands that mention an object or place named in the task description. For example, given the task ``put a clean knife in the drawer’', \jev carries an unwashed knife straight to the drawer instead of first washing it at the sink. Surprisingly, many of these failures are not due to a lack of knowledge. Asked separately what the opponent will do, \jev usually answers correctly, and it responds well given the opponent’s action. It fails when one call must both perform the simulation and evaluate based on it. This suggests letting code make the prediction or simulation. When code supplies it, such as a lookahead in ALFWorld and physics simulation in robot control, \jev becomes an expert controller through its general evaluation ability.
[AI-32] Continuous Process-Level Evaluation for Evolving Enterprise AI Agent Skills NEURIPS2026
链接: https://arxiv.org/abs/2610.01833
作者: Ngoc Phuoc An Vo,Aarya Doshi,Vadim Sheinin
类目: Artificial Intelligence (cs.AI); Software Engineering (cs.SE)
备注: Accepted to Workshop on Continual Learning for Enterprise AI Agents (CLEA), NeurIPS 2026
Abstract:Enterprise AI agent skills evolve as tool APIs, models, and specifications change, yet final-output evaluation can miss process-level behavioral drift. We present a continuous evaluation framework combining outcome-level and process-level checks, applied to Revenue and Productivity variants of a Business Value Determination skill in an enterprise Value Aware Resiliency system. The framework independently computes per-run ground truth, materializes reusable template tests, and evaluates tool selection, arguments, execution order, and database integrity through programmatic checks and a narrowly scoped LLM judge. We evaluate 240 trials across two skills, two specification variants, two agent harnesses, and three models. Of 175 trials passing all applicable final numerical checks, 162 (92.6 percent; Wilson 95 percent CI: 87.7-95.6 percent) contained another evaluator-detected deviation. Under a broader seven-check final-state definition, 151 of 164 passing runs (92.1 percent; 95 percent CI: 86.9-95.3 percent) still violated a trajectory check. Dependency attribution reduced a mean of 6.34 failed checks per run to 2.65 roots. Specification sensitivity varied by model and harness, with exploratory bootstrap interaction intervals excluding zero for all three Revenue comparisons and one of three Productivity comparisons. Runtime-resolved templates provided reusable regression coverage across the evaluated configurations; longitudinal validation under actual API evolution remains future work.
[AI-33] AI-assisted mitotic counting improves reproducibility and efficiency across multiple tumour types
链接: https://arxiv.org/abs/2610.01813
作者: Simon Graham,Mostafa Jahanifar,Quoc Dang Vu,Vygante Maskoliunaite,Donatas Petroska,Ruta Barbora Valkiuniene,Ayat Gamal Lashen,Jen Hong Ong,Amede Ogechi Nnorom,Sinclair Couper,Natasha Kardasz,Reshma Agrawal,Brinder Singh Chohan,Jose Luis Solorzano Rendon,Shonali Natu,Arvydas Laurinavicius,Nasir Rajpoot,David Snead
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Mitotic counting is an important component of tumour grading, diagnosis and prognostic assessment across several tumour types, but manual assessment is time-consuming and subject to inter-pathologist variability. To help address these challenges, we developed MitPro, an AI tool designed to improve consistency and efficiency by directing pathologists towards regions with the highest predicted mitotic activity and highlighting mitotic figures for review, while retaining pathologist control over region selection and the final count. We evaluated its effect on the reproducibility and efficiency of mitotic counting in a retrospective, non-interventional, paired reader study comprising 385 whole-slide images from 3 centres in 3 countries and 7 tumour types using 3 different scanners. 13 pathologists participated, with each slide assessed independently by 3 pathologists without AI assistance and again with AI assistance after a minimum 2 week washout period. Across all slides, AI-assisted counting increased the intraclass correlation coefficient from 0.589 to 0.949. Mean pathologist-level median assessment time decreased from 286.4 to 127.8 seconds, corresponding to an average saving of 151.8 seconds per assessment. Improvements in agreement and efficiency were also observed in supporting analyses using HALO AP and Sectra image management systems and in 2 additional tumour types outside the main study population. AI-assisted assessment was associated with a subtle shift towards higher mitotic counts and scores, consistent with identification of more active mitotic hotspots and fewer missed mitotic figures. The frequency of score change between unassisted and AI-assisted assessment was comparable with inter-pathologist variation during routine counting. These findings support the use of MitPro as an assistive tool for more consistent and efficient mitotic assessment in routine practice.
[AI-34] LineupRL: Verifiable Reinforcement Learning for Time Series Captioning via Caption-to-Series Identification
链接: https://arxiv.org/abs/2610.01800
作者: Haochen Zhang,Laura Yao,Zachary Plotkin,Gengwei Zhang,Tianlong Chen
类目: Artificial Intelligence (cs.AI)
备注: 28 pages, 4 figures
Abstract:Time series captioning is a fundamental step in time series understanding and can also serve as the bridge between signal and natural language. Supervised fine-tuning (SFT) relies on a larger model’s captions and cannot exceed their quality. Reinforcement learning (RL) can, but its rewards were designed for other modalities and other tasks, and they transfer poorly to open-ended generation in the time series domain. We address this by proposing LineupRL, a reinforcement learning with verifiable rewards (RLVR) pipeline whose reward is caption-to-series identification. The reward model is a frozen large language model (LLM) verifier that reads the generated caption and the candidate time series as raw values, never the chart, and must pick the described time series from multiple distractors. Matching is a far lighter demand on the verifier than writing questions or judging a caption, so an off-the-shelf LLM can supply the reward. Across two captioning benchmarks, and on forecasting and reconstruction where the predictor sees only the caption, LineupRL outperforms SFT and RL baselines on every metric. The 3B vision language model (VLM) trained by LineupRL also outperforms, at 1/24 of the parameters, the 72B VLM whose captions the SFT baseline is distilled from. Our case study shows that LineupRL resists reward hacking, and that the captioner it trains both traces the trend and names the values at key points.
[AI-35] ADD: Improving Alignment and Diversity in Diffusion Policy Optimization
链接: https://arxiv.org/abs/2610.01789
作者: Ashok Prasad Neupane,Saugat Adhikari,Pramish Paudel,Ajad Chhatkuli,Danda Pani Paudel
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:
Abstract:Reinforcement learning based post training of diffusion models, such as Denoising Diffusion Policy Optimization (DDPO), optimizes a reverse diffusion process under a reward function. However, current approaches to reward optimizations do so at the cost of diversity and quality. In this paper, we provide better tradeoffs through careful theoretical considerations and method design. We analyze the theoretical framework and mathematically demonstrate that \emphonly-latter timestep updates of diffusion model may be harmful for diversity contrary to the conclusions presented in a previous work. Additionally, we propose an incremental Feynman-Kac training based on strong theoretical foundations in order to achieve the best-yet alignment-diversity tradeoffs. We perform extensive experiments and compare our method against related diffusion policy optimization approaches in three different tasks and also provide strong ablations for each component, thus validating strong performance gains in both alignment and diversity.
[AI-36] Not All Experience Belongs in the Weights: Component Routing for Self-Improving GUI Agents
链接: https://arxiv.org/abs/2610.01787
作者: Beining Wu,Zihao Ding,Jun Huang
类目: Artificial Intelligence (cs.AI); Graphics (cs.GR)
备注:
Abstract:Self-improving GUI agents keep the trajectories they produce and return them to the agent, by fine-tuning or by retrieval into the prompt, and studies that compare the two destinations disagree. We attribute this to the unit of experience: a trajectory bundles items with different properties, so a conclusion about the bundle depends on its mix. To address this, (i) we introduce component routing, which splits the experience into locators, procedures, state facts and lessons and sends each component to the context or to the weights, compared on the same items across three backbone families, two environments and three seeds. One pool has two destinations: locators and lessons win in the weights, procedures and state facts in the context. (ii) We fit a rule in two properties measured before any training, recurrence and state-conditionality; it recovers the destination of a held-out backbone family in 24 of 24 cells, two interventions move a component toward the boundary, and routing by the rule beats every whole-trajectory baseline and, by +3.5 points on average, the better single destination of each backbone. (iii) We identify how training and producer-consumer differences change the value of the two destinations: note readout decreases after the same component is written into the weights, most for the items that recur most, context gains increase with the information gap, and weights gains decrease with the policy gap. Code and data will be released.
[AI-37] Q-Learning for Reachability in MEC-Free MDPs
链接: https://arxiv.org/abs/2610.01781
作者: Lu-Chin Chang,Suguman Bansal
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Logic in Computer Science (cs.LO)
备注: 15 pages, 4 figures
Abstract:Reinforcement learning (RL) for reachability specifications is fundamental to sequential decision-making. Prior work establishes asymptotic convergence to optimal policies, but only through model-based methods that must explicitly estimate the transition probabilities of the underlying Markov Decision Process (MDP). We present Quasar, the first model-free algorithm with asymptotic guarantees for reachability on the fragment of MDPs free of non-terminal maximal end components (MECs), a building block to which every MDP reduces by the standard MEC quotient. Our algorithm follows the classical Q-learning approach, using temporal-difference updates to converge to an optimal policy without ever learning the transition probabilities. The resulting learner reduces the memory footprint from the O(|S|^2|A|) that model-based methods require to O(|S||A|). On the standardized Quantitative Verification Benchmark Set, our algorithm converges to the optimal policy with orders of magnitude fewer samples than the previous model-based state-of-the-art. Together these results are a concrete step toward the practical deployment of reachability learning and, with it, of specification-guided RL.
[AI-38] RealCompanion: Benchmarking Human Understanding from Reasoning over Longitudinal Real-World Conversations
链接: https://arxiv.org/abs/2610.01780
作者: Arman Behnam,Sunglyoung Kim,Liangwei Yang
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:A companion that talks with a person for months should come to understand them. It should remember what they said, infer who they are, and know when the past bears on the message in front of it. Testing this requires a real person’s record, and such records are private, so benchmarks generate the person and the questions and settle in advance what matters. We release \bench, ten real relationships with an AI companion: 27,218 messages over up to 120 days, released as the conversation and four files derived from it, a profile, a persona, a chat ground truth and a question set, each citing the messages it rests on. Every chat label carries the reasoning trace that produced it, checked stage by stage against the conversation. Three findings follow. First, the past is rarely needed and far away. Pooled measures mislead: a recency window finds the required message for 95.9% of probes and 2.2% of those that need memory, and at the natural rate 96% of the gain from supplying recorded evidence comes from messages that need none. Second, no detector we tried can tell when memory is needed on real messages, authored questions over the same histories leak the cue, and labeling the same messages as memories raises their use by ten to fourteen points. Third, three agent systems reconstruct the persona with the same F1 at a 31-fold difference in cost.
[AI-39] CODesign: Consistency from Data to Trajectory in All-Atom Protein Binder Co-Design
链接: https://arxiv.org/abs/2610.01773
作者: Yuanle Mo,Bo Qiang,Haitao Lin,Qinghan Wang,Gang Du,Odin Zhang,Pheng Ann Heng
类目: Computational Engineering, Finance, and Science (cs.CE); Artificial Intelligence (cs.AI)
备注:
Abstract:The central challenge in de novo protein design is generating plausible, mutually compatible structures and sequences, such that each designed sequence folds into its intended structure and the structure accommodates that sequence. Compared to typical two-stage design methods, which decouple the modeling of the interdependent modalities, co-design models improve the cross-modal consistency by jointly generating sequences and structures. However, naively generating sequences and structures simultaneously does not ensure their consistency. To address this challenge, we propose CODesign framework. We improve data consistency by generating approximately 105,000 consistency-distilled dimers. We further promote consistency through a multimodal joint flow model that captures the joint distribution of sequences, backbone structures, and local atomic configurations, together with a consistency-aware joint resampling strategy that iteratively refines sequences and side chains. Experiments show that CODesign achieves state-of-the-art performance with the highest in silico success rates on both protein- and ligand-target binder design. Ablation studies also demonstrate our distilled dataset increases performance by 70.9%, which can be further improved by our proposed resampling mechanism with negligible additional computational cost. Code, model weights and the new dataset will be completely open-source.
[AI-40] opK-Guided: Adaptive Budget-Aware Activation Sparsity for Efficient LLM Inference
链接: https://arxiv.org/abs/2610.01763
作者: Mukund Agarwalla,Chih-Jen Lin
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Activation sparsity speeds up large language model (LLM) inference by setting unimportant activations to zero so that the corresponding computations can be skipped. Existing training-free methods, however, make different trade-offs: threshold-based methods such as TEAL adapt the sparsity level to each token but do not tightly control the realised sparsity, while TopK-based methods such as WINA enforce a fixed sparsity level but use the same sparsity budget for every token. Both also apply the same budget across transformer blocks, despite large differences in block sensitivity. We introduce TopK-Guided, a training-free method that addresses both limitations by combining bounded token-level sparsity adaptation with sensitivity-aware block-level budget allocation. Across Llama-2 and Llama-3 models, TopK-Guided consistently improves perplexity and downstream accuracy over TEAL and WINA while preserving essentially the same sparsitydependent projection compute as WINA, with the largest gains at high sparsity. Ablations show that both components provide complementary improvements.
[AI-41] Removing spurious minima for planar features by skip connections
链接: https://arxiv.org/abs/2610.01728
作者: Jakob Paul Zimmermann,Moritz Grillo,Andrei Balakin,Georg Loho
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Optimization and Control (math.OC)
备注: 43 pages, 4 figures. Under review. Accompanying Lean 4 formalization available at this https URL
Abstract:Understanding loss landscapes is central to explaining neural-network training, yet their structure remains only partially understood even in simple models. We study the Gaussian population loss of shallow, bias-free ReLU networks in the teacher–student setting. This provides a simple model for studying essential aspects such as feature learning and overparameterization. For teacher networks with positive output weights and planar features, we show that including a learned linear skip removes all spurious local minima with non-negative student output weights once the student network is at least as wide as the teacher network. In contrast, without the skip, we construct a fixed teacher network with positive output weights and only three hidden neurons in input dimension two whose spurious local minima persist at every student width at least three. Thus, a learned linear skip can remove spurious minima that persist under arbitrary overparameterization. Furthermore, we show that a positive output weight student network always learns the subspace spanned by the teacher features: student features at local minima with non-negative student output weights lie in the span of the teacher features. For ReLU networks in two dimensions, even heavily overparameterized student networks have effective width controlled by the teacher width: every critical point with positive student output weights has at most twice as many distinct student feature directions as teacher neurons. Finally, we transfer the benignity result to empirical minima over parameter balls of any prescribed radius, with the required sampling accuracy depending on that radius.
[AI-42] vFedProtoQNAS: Prototype-Guided Personalized Quantum Neural Architecture Search for Virtual Federated Learning
链接: https://arxiv.org/abs/2610.01718
作者: Seok Bin Son,Samuel Yen-Chi Chen,Soohyun Park,Joongheon Kim
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Quantum federated learning (QFL) has emerged as a promising approach for collaboratively training compact quantum neural networks (QNNs) over distributed private data on resource-constrained devices. However, differences in device capabilities make a single shared QNN architecture unsuitable for all clients. While personalized quantum neural architecture search (QNAS) allows each client to select a device-specific QNN, averaging parameters across structurally different QNN architectures mixes semantically inconsistent circuit operations. To address this, prototype-guided personalized QNAS for virtual FL (vFedProtoQNAS) is proposed, where model parameters are never aggregated across clients and federated collaboration is achieved through class-wise prototype sharing. Each client independently searches and trains a client-specific QNN, computes class-wise local prototypes from latent representations, and refines them using global prototypes from the server as federated semantic anchors. Experiments demonstrate that vFedProtoQNAS improves accuracy by 3.70% over FedAvg and enhances class-consistent representation alignment.
[AI-43] Architecture Without an Architect? Global Governance of Artificial Intelligence in a Divided World
链接: https://arxiv.org/abs/2610.01716
作者: Simon Chesterman
类目: Computers and Society (cs.CY); Artificial Intelligence (cs.AI)
备注:
Abstract:Artificial intelligence presents an unusually difficult problem for global governance. The technology develops rapidly, crosses borders easily, and is shaped by actors whose resources and capabilities may rival those of states. Yet international responses remain fragmented, unevenly representative, and overwhelmingly non-binding. The challenge is therefore not simply to identify appropriate rules or institutions, but to understand who has the capacity and incentive to create, enforce, and adapt them. This review essay examines these questions through Matthijs Maas’s Architectures of Global AI Governance. Maas offers an ambitious framework for thinking about AI governance through the lenses of sociotechnical change, governance disruption, and regime complexity. His account usefully resists both technological determinism and the search for a single institutional blueprint, emphasizing instead the possibilities of a fragmented and evolving governance architecture. The essay argues, however, that institutional design cannot be separated from the distribution of power. Maas frequently invokes what “we” should do about AI, but that collective subject obscures important differences among states, international institutions, and technology companies. States retain formidable powers over markets, infrastructure, strategic inputs, and firms themselves. At the same time, many consequential decisions about frontier AI - what is built, how quickly, with what safeguards, and when it is released - are concentrated within a small number of private companies. The central problem of global AI governance may therefore be less architecture without an architect than an emerging architecture shaped by multiple actors possessing different forms of power, divergent incentives, and no common set of plans. Subjects: Computers and Society (cs.CY); Artificial Intelligence (cs.AI) Cite as: arXiv:2610.01716 [cs.CY] (or arXiv:2610.01716v1 [cs.CY] for this version) https://doi.org/10.48550/arXiv.2610.01716 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[AI-44] Iterative Policy Refinement through Semantic Rollout Analysis
链接: https://arxiv.org/abs/2610.01652
作者: Feiyu Gavin Zhu,Qi Xu,Zhifei Deng,Zhigang Hua,Luke Simon,Jean Oh,Reid Simmons
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:
Abstract:Structured policies improve efficiency, robustness, and interpretability in imitation learning by introducing task-specific inductive bias, but existing structure generation methods rely either on extensive human input or on static domain knowledge encoded in LLMs, which may be inconsistent with the expert demonstrations. We propose a closed-loop framework that iteratively refines structured policies using LLM-guided analysis of policy rollouts. By logging rollouts as semantically meaningful tabular data and prompting the LLM to generate diagnostic analysis code, our method identifies suboptimalities in the policy structure and iteratively corrects them without requiring human instruction. Experiments on car racing and door opening tasks show that our approach improves imitation learning performance by up to 15% over zero-shot LLM-generated structures and requires 75% less compute to achieve the same reinforcement learning performance. These results demonstrate that tabular rollout analysis provides an effective feedback signal to align LLM-generated policy structures with expert demonstrations, and we can utilize it to generate good policy structures automatically.
[AI-45] MCIR: A Feature Dependence-Aware Explainability Method with Reliability Guarantees
链接: https://arxiv.org/abs/2610.01641
作者: Poushali Sengupta,Sabita Maharjan,Frank Eliassen,Shashi Raj Pandey,Yan Zhang
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation (stat.CO)
备注: Accepted for publication in Transactions on Machine Learning Research (TMLR)
Abstract:Modern machine-learning models often contain strongly dependent or redundant features, making feature attribution difficult because shared predictive information can be distributed across correlated predictors. Existing methods such as SHAP, LIME, HSIC, MI/CMI, and SAGE may therefore produce unstable rankings under multicollinearity or near-duplicate predictors. We propose the Mutual Correlation Impact Ratio Method (MCIR-M), a dependence-aware global feature-importance approach that quantifies the unique predictive information contributed by each feature beyond a selected dependence neighbourhood. MCIR-M introduces the Mutual Correlation Impact Ratio (MCIR), which conditions each feature on strongly dependent neighbours and computes a normalized ratio of conditional to block-level information. The population score lies in [0,1] and equals zero under exact conditional redundancy. We also introduce a lightweight estimation procedure that computes MCIR using a fraction of the available data and evaluates agreement with full-data explanations. Across controlled synthetic redundancy experiments and the UCI HAR benchmark, MCIR shows dependence-aware ranking behaviour, with its clearest advantage under injected near-duplicate predictors. Comparisons with independent and conditional SHAP, SAGE, HSIC, MI-based scores, and CIR-family baselines are mixed across real-data criteria. Reduced explanation samples lower computational burden in the evaluated configurations, while agreement with full-data explanations is assessed separately through ranking, head-set, and faithfulness diagnostics. Overall, MCIR-M provides a practical dependence-aware diagnostic for global explanation under strong feature dependence.
[AI-46] Measuring the Stability Assumption Behind Action Chunking
链接: https://arxiv.org/abs/2610.01626
作者: Aryan Goyal
类目: Artificial Intelligence (cs.AI)
备注: 18 pages, 9 figures, 18 tables
Abstract:Action chunking improves the performance of policies learned by behavioural cloning, and several mechanisms have been proposed to explain why, including temporal consistency, horizon reduction, representation learning, and reduced error compounding. We instead study what happens to an action error once it enters the system. At each state, we inject a small action error and measure how fast it grows or shrinks under two execution regimes: open-loop, where the rest of the chunk is replayed without replanning, and closed-loop, where the policy replans after the perturbation. The fitted rate labels each state as contracting, expanding, or unresolved. Across twelve manipulation tasks from three benchmark suites, we find that confidently stable states are rare, while error amplification is common among states whose propagation rate can be resolved. We further find that the measured propagation rate depends strongly on the fitting horizon: amplification is typically front-loaded, so short windows can overestimate longer-horizon propagation. Finally, we train predictors on these labels and find that a state’s open-loop regime can be recovered from camera frames and proprioception alone, while its closed-loop propagation is only partially recoverable because it also depends on how the policy acts after the perturbation. These results suggest that error-compounding arguments alone do not provide a complete account of action chunking: neither passive open-loop dynamics nor policy replanning consistently contracts an injected error, and replanning rarely turns open-loop amplification into confident contraction. This suggests that closed-loop reactivity should be trained explicitly, using perturbation- and tree-coverage-oriented training to expose policies to deviations they must recover from, rather than expected to emerge reliably from standard imitation learning.
[AI-47] FedLore: Communication and Memory Efficient Federated Learning via Shared Gradient Low-Rank Projection
链接: https://arxiv.org/abs/2610.01620
作者: Junkang Liu
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Federated training of foundation models is constrained by client memory and communication costs. LoRA-based methods reduce these costs through low-rank adapters, but their fixed rank budget can limit adaptation. Gradient low-rank optimization offers greater flexibility, yet independently chosen client subspaces create a problem we term \emphsubspace fragmentation: local projections interact with data heterogeneity to bias aggregated directions, while aggregation can increase update rank and communication cost. Thus, accurate local gradient compression need not preserve global descent. We propose \textttFedLore, which shares a low-rank optimization basis within each round and refreshes it across rounds. The shared basis enables exact aggregation in low-rank coordinates and eliminates the identified projection bias. Subspace refresh allows the accumulated model update to exceed the per-round rank budget. We characterize the aggregation bias and establish an O(T^-1/2) stationarity bound for the projected-SGD variant under a global-gradient coverage condition and standard smoothness and variance assumptions, with bounded gradient heterogeneity. Experiments on vision and language tasks, including federated pre-training, show that \textttFedLore outperforms the evaluated low-rank adapter baselines and matches or exceeds full-parameter training, while reducing communication and optimizer-state memory.
[AI-48] Exposing the Cost of Deep Learning Audio Development
链接: https://arxiv.org/abs/2610.01619
作者: Constance Douwes,Paul Magron,Romain Serizel
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Sound (cs.SD)
备注: 5 pages, 2 figures, 1 table
Abstract:The environmental impact of deep learning has attracted increasing attention over the past decade. Existing studies mainly focus on the energy and carbon emissions of model training and inference, while the whole development phase is often overlooked. Yet, architecture prototyping and intensive experiments are conducted during this stage, which is highly energy-demanding. In this article, we propose a methodology to estimate these costs, based on activity logs from the Grid5000 shared computing platform used by the LORIA laboratory. As a case-study, we focus on audio projects developed in the Multispeech research team. We evaluate the overall energy cost of four projects, and we compare them to those of training the reported models. Our results show that the energy required for the development phase is 3 to 256 times greater than that required to train the best-performing model alone. These results advocate for a more systematic reporting of energy consumption across the entire life cycle of deep learning-based audio projects.
[AI-49] Agents Are Systems Not Models: Rethinking Agent ic Evaluation
链接: https://arxiv.org/abs/2610.01618
作者: Luis Wiedmann,Leander Girrbach,Cordelia Schmid,Zeynep Akata
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Agent evaluations increasingly go beyond a single success rate, reporting metrics such as cost, consistency, and robustness. Yet they typically treat the agent itself as fixed. In practice, an agent is a configurable system: users decide what to tell it, how long to let it run, and which model to use, and each of these choices can change how well and how consistently it performs. We study these choices on a new benchmark of four scientific tasks, where a coding agent must find and correctly operate a published specialist model. We investigate five parts of the agent’s configuration: task information, reasoning, self-verification, time budget, and backbone model. We find substantial run-to-run variability, with approximately 54% of the outcome variance coming from repeating the same configuration rather than changing it. Across configurations, the information provided to the agent has the largest effect, exceeding both time budget and model size, while also reducing cost and improving calibration. Configuration choices also interact: additional time helps only when the agent has sufficient information or a capable enough model to use it. Finally, a trajectory-based taxonomy of agent behavior reveals that prompting an agent to verify its answer has little effect on its verification behavior, whereas providing a dedicated verification tool changes that behavior substantially. These results suggest that agents should be evaluated as configurable systems themselves, and that some desired behaviors are more effectively implemented in the system than requested through prompting. We release the benchmark and more than 18,000 agent trajectories.
[AI-50] Permutation-Robust Decision Modeling with Candidate-Independent Block-Causal Attention
链接: https://arxiv.org/abs/2610.01601
作者: Guy Amit
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: Technical Report, will not be submitted to a conference
Abstract:Decision models often score a variable-sized set of candidate actions encoded in a single sequence. This setting is increasingly relevant for System 1 components inside generative systems, where candidates may be proposed or ordered differently across runs. Standard causal cross-encoding is expressive, but it can make a candidate’s score depend on serialization order rather than on the underlying decision problem. We introduce candidate-independent block-causal attention, which preserves causal computation within the shared context and each candidate while blocking cross-candidate information flow and resetting candidate positions. We compare this architecture with standard causal attention and complementary invariant baselines across Gemma 3 1B, Qwen3 1.7B, and Qwen3 4B backbones. Candidate-independent attention consistently reduces permutation sensitivity while retaining competitive decision quality; ablations indicate that candidate isolation is the primary source of the effect, with position resetting completing the intended symmetry. A larger Qwen3-4B study further examines the behavior of the proposed architecture with substantially more training data. Code is available at the \hrefthis https URL\textcolorblueproject repository, and the \hrefthis https URL\textcolorblueQwen3-4B model artifact is available on Hugging Face.
[AI-51] Evaluating Physical Consistency and Plausibility in Generative Scenario Models for Autonomous Driving
链接: https://arxiv.org/abs/2610.01581
作者: Manasa Mariam Mammen,Zafer Kayatas,Stefan Wagner
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Generative AI models are increasingly used for scenario generation in autonomous driving. While they can generate realistic-looking scenarios, they often provide limited transparency into learned representations and consistency with real-world vehicle dynamics. This lack of formal assurance limits their use in safety-critical validation and certification workflows. To address this aspect, we introduce a layered evaluation protocol that complements existing methods by assessing models across five layers. The first four layers inspect internal representations and network layers through kinematic alignment, statistical baseline comparison, latent controllability, and activation analysis. The fifth layer evaluates model outputs against vehicle dynamics constraints such as lateral jerk thresholds. We demonstrate the protocol on a Variational Autoencoder (VAE)-based scenario generator. Although standard output-level metrics and visualizations suggest that the generated scenarios are realistic, our protocol provides deeper insight into the extent to which the model’s latent space aligns with kinematic features and whether visually plausible trajectories satisfy vehicle-dynamics constraints. We further apply the protocol to additional generative models, demonstrating its applicability beyond the VAE architecture.
[AI-52] Beyond Pointwise Error: A Multi-Metric Evaluation of Spatial Climate Downscaling
链接: https://arxiv.org/abs/2610.01579
作者: Loys Masquelier,Etienne Le Naour
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:
Abstract:Climate downscaling aims to reconstruct fine scale spatial fields from coarse resolution inputs. Evaluating the quality of these reconstructions is challenging: low pointwise error can come at the cost of fine scale variability, while realistic spatial variability can be achieved with inaccurate local structures. The evaluation metric can therefore change which method appears to perform best. This work presents a multi metric benchmark comparing five spatial downscaling methods on ERA5 temperature, wind, and precipitation fields. Five criteria assess complementary properties: pointwise error, structural similarity, distribution error, spectral error, and gradient error. The results reveal a systematic trade off between spatial fidelity and fine scale variability. Some methods perform best on pointwise and spatially aligned metrics, but lose high frequency content, while others preserve substantially more spectral variability at the cost of less accurately positioned local structures. Consequently, method rankings change across metrics and variables. These results show that there is no single best downscaling method. Multi metric evaluation is therefore essential for assessing which properties of a climate field are preserved.
[AI-53] Chaining Skills to Hijack LLM Agents
链接: https://arxiv.org/abs/2610.01564
作者: Tian Dong,Zixuan Ma,Haodong Zhao,Huaien Zhang,Shaofeng Li,Hao Chen
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
备注:
Abstract:LLM agents use skills to improve performance on specialized tasks. To complete a user request, an agent may invoke several skills in sequence, allowing information produced under one skill to guide the next. Because skills may come from open-source repositories, this handoff can also carry attacker-controlled claims into later decisions. In this paper, we introduce APEX, which constructs and refines adversarial skill chains tailored to a user task and an attacker-selected action. The key insight is that an agent-written record of genuine task progress can carry a false claim of user approval across skills: an upstream skill induces the agent to create the record, and a downstream skill uses it to direct the attacker-selected action. Across four targeted-action families and six models on SkillsBench, the chains induce the selected action in 512 of 690 attempts (74.2%). On GPT-5.4, the full chain succeeds in 84.3% of attempts, compared with 17.4% when the workflow is merged into one skill. We further evaluate a prompting defense that asks the agent to check skill-produced files against the original request. On GPT-5.4, it lowers targeted-action success from 84.3% to 59.1%, while the verifier test-pass rate across 72 benign native-skill tasks falls from 86.7% to 56.3%. These results highlight the need for defenses that prevent attacker-directed actions while preserving legitimate task performance.
[AI-54] Completion Aware Guidance for World Action Models
链接: https://arxiv.org/abs/2610.01559
作者: Seungyeon Kim,Junhoo Lee,Baekseung Kim,Minkyu Kim,Nojun Kwak
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:
Abstract:World Action Models (WAMs) predict visual futures and robot actions, yet they remain susceptible to task-incomplete imagination, where plausible, action-consistent predictions omit the transition needed for task completion. In this paper, we show that this failure is not inherent to the world model backbone, but emerges when adapted for short-chunk control, which can repeatedly favor plausible local continuations over task-completing transitions. To address this, we introduce Completion Aware Guidance (CAG), a training-free sampling method that guides generation toward task completion. Across representative WAMs, CAG improves success from 64% to 70% on a RoboTwin 2.0 subset and from 69% to 75% in zero-shot simulation, while reducing task-incomplete imagination from 79% to 40%.
[AI-55] he AI Assessment Sandbox Configurator: A Framework to Support Technical Assessment in AI Regulatory Sandboxes
链接: https://arxiv.org/abs/2610.01539
作者: Alessio Buscemi,German Castignani,Daniele Pagani,Maxime Cordy,Jordi Cabot
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:The EU’s Artificial Intelligence Act requires all Member States to establish AI Regulatory Sandboxes (AIRS) by August 2027: supervised environments bringing together national Competent Authorities, technical experts, and the organisations under assessment. When AIRS engagements include structured technical testing, running such testing at scale demands dedicated infrastructure, yet the tooling ecosystem remains structurally fragmented, with heterogeneous tools producing outputs that are difficult to compare, trace, and reuse. From the procedural conditions of AIRS engagements and the AI Act obligations for high-risk systems, we derive 11 architectural and governance requirements for the infrastructure that operationalises technical testing within an AIRS. In response to these requirements, we introduce the AI Assessment Sandbox Configurator, an open-source framework combining a curated Catalogue of tests and controls accessed through a stable plug-in API, a shared data model that harmonises heterogeneous outputs, role-specific dashboards for multi-disciplinary interpretation, and audience-segmented reporting. We describe the architecture and current release, and report an early-stage pilot that exercised the harmonisation and reporting layers within a live AIRS engagement and contributed to an official Exit Report. We discuss the roadmap, the governance questions raised by the Catalogue’s tiered contribution model, and the institutional pathways through which an open-source assessment ecosystem could emerge across Member States.
[AI-56] False Floors: LLM Safety Routing Evaluations Break Under Distribution Shift
链接: https://arxiv.org/abs/2610.01535
作者: Amit Singh Bhatti,Vishal Vaddina
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
备注:
Abstract:Safety routers send each request to one of several models and are judged against the best single model. A major routing benchmark picks that comparator on the evaluation data. In the benchmark’s own setting this is harmless, but under distribution shift it is not. On HELM Safety the selection cost is 0.003-0.030 of harm under random splits and 0.045-0.113 under held-out categories, comparable to the whole deficit attributed to routing, with its direction holding under either published judge alone. It rises seven- to ninefold on AgentDojo when suites are held out. Across seven safety corpora chosen by rules fixed in advance, three meet a registered interval test and four beat a later permutation null, and three of the four interval misses are corpora where some models have zero observed harm. Prior work proves the direction of this bias. We size it on harm and accuracy, show that it is larger under the held-out splits we measure, and bound it by optimism plus a shift-dependent regret. Scored honestly under shift, routing buys little on these benchmarks. In most pool cells the nested router serves the honest baseline’s model, and on the nearly saturated AgentDojo corpus a perfect pre-dispatch router is worth at most two points of harm. We also find a model’s expressed recognition of a late injection steerable. On held-out reruns an attacker who knows which model it faces lowers GPT-5.4’s judged recognition by 19.6 points, confirmed by an independent label. In an offline counterfactual composition into a controller, the same attack raises or lowers estimated harm depending on the fallback model. Safety routing should be evaluated under shift, against a baseline chosen without the test labels, and recognition-based defences should be scored on harm against an attacker who chooses what the model sees.
[AI-57] Exact Distinguishability in Non-Markovian Decision Processes
链接: https://arxiv.org/abs/2610.01527
作者: Kabir Murjani,Nisarg Patel
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Machine Learning (stat.ML)
备注: 26 pages, 7 figures. Code and Lean 4 proofs: this https URL
Abstract:Non-Markovian environments are often modeled as Regular Decision Processes (RDPs), where dynamics depend on the interaction history through a finite automaton. Existing offline guarantees for RDPs rely on a distinguishability assumption on the behaviour policy but provide no means of verifying it. When the assumption is violated, distinct models may explain the data equally well. We study when data collected under a fixed behaviour policy can distinguish two candidate RDPs. We prove that the posterior odds between observationally equivalent candidates remain equal to the prior odds at every sample size, even when the policy visits every automaton state, and verify both results formally in Lean 4. We then characterize this equivalence exactly and derive PEC, an algorithm that decides it in time linear in the size of the product automaton. The distinguishability assumption of prior work fails on three of our four test environments, and the experiment identified by PEC restores it in each case.
[AI-58] Auto-Formalizing Neuro-Symbolic Predictors
链接: https://arxiv.org/abs/2610.01519
作者: Samuele Bortolotti,Weixin Chen,Han Zhao,Andrea Passerini,Stefano Teso,Antonio Vergari
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:
Abstract:Neuro-Symbolic (NeSy) predictors incorporate prior knowledge into the prediction process of neural networks, ensuring that outputs satisfy specified constraints, making them particularly suitable for high-stakes applications where compliance with domain knowledge is essential. A key bottleneck in this paradigm is the acquisition of symbolic constraints: encoding domain knowledge into logical formulas remains a manual and expert-intensive process. In this work, we investigate the extent to which auto-formalization via LLMs can systematically translate textual knowledge into symbolic knowledge that can be plugged into NeSy predictors. To this end, we introduce auto-nesy-bench, a new benchmark for evaluating constraint formalization and its impact on downstream accuracy of NeSy predictors. Through an extensive evaluation across several domains, we find that LLMs can formalize constraints to a meaningful extent, generating formulas that are often similar to those provided by human experts. Moreover, when the generated formulas are syntactically valid, they can lead to high-quality downstream predictions. The code and benchmark are available at this https URL.
[AI-59] FedMIX-P: Mixing Local and Global Preconditioners for Federated Vision and Language Model Training
链接: https://arxiv.org/abs/2610.01515
作者: Junkang Liu
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:
Abstract:Adaptive preconditioners accelerate model training, but heterogeneous client geometries can bias federated updates even when gradients are evaluated at the same model. Round-start synchronization alone cannot prevent this mismatch from reappearing during local training. We propose \textttFedMIX-P, which mixes shared and local preconditioners at every local step, retaining local adaptation while reducing mean-squared operator mismatch by a factor of \lambda^2 . For smooth nonconvex objectives with stochastic gradients and partial participation, we establish an O(R^-1/2) stationarity bound using suitable stepsizes and a horizon-dependent mixing weight, without requiring local preconditioners to converge to one another. A two-client counterexample shows that fixed positive mixing can preserve a nonstationary fixed point. The theory covers bounded linear symmetric positive-definite preconditioners. Experiments with SOAP, Sophia, and Muon variants across vision and language tasks show improvements over corresponding local optimizers, including accuracy gains of up to 19.47 percentage points and lower validation loss for 60M–350M language models. Full nonlinear and momentum-based updates require separate analysis.
[AI-60] Decision Titan: Test-Time Training for Long-Term Memory in Offline Reinforcement Learning ICML2026
链接: https://arxiv.org/abs/2610.01513
作者: Jude Waide,Robert Lieck
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: Accepted at ICML 2026 Workshop on Decision-Making from Offline Datasets to Online Adaptation: Black-Box Optimization to Reinforcement Learning
Abstract:Long-term dependencies remain a major challenge for sequential decision-making in the field of AI: RNNs suffer from vanishing gradients and the limited expressivity of vector-based hidden states, whilst Transformer-based models are limited by the quadratic scaling of attention. Recent work has proposed tackling this problem with the Test-Time Training (TTT) framework, which stores episodic memories in the parameters of a neural network through gradient descent at both train and test-time. This approach has seen success in the domain of Natural Language Processing, however, to the best of our knowledge it has not yet been applied to the domain of Reinforcement Learning (RL), nor has there been a study analysing how this memory practically functions. In this paper, we study the potential of the TTT framework for offline RL by augmenting a Decision Transformer with TTT layers, dubbed the Decision Titan. We analyse performance and properties of the model in the X-Maze environment, an extension of T-Maze designed to test sequential memory, and investigate how the memory mechanism learns by visualising gate values over time. Our key findings are that Decision Titan can learn long-term dependencies with ranges 20x longer than the context window, generalises to lengths 1.7x the training data, but crucially temporal generalisation depends on the time embeddings used, and the ability to learn long-term dependencies depends on how the relevant information is encoded.
[AI-61] Sharpening Tax in Post-Training
链接: https://arxiv.org/abs/2610.01509
作者: Changdae Oh,Qi Zeng,Qi Qi,Andrey Zhmoginov,Deren Lei,Yun He,Hoang Phan,Hangoo Kang,Azalia Mirhoseini,Sharon Li
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:
Abstract:An emerging hypothesis about reinforcement learning (RL) post-training of large language models (LLMs) is that it merely sharpens existing behaviors of a base model, improving single-shot accuracy at the cost of solution coverage. Although this trade-off has been observed in math and coding tasks, it need not extend to agentic tasks, where multi-turn tool use and interaction may require capabilities newly acquired during post-training. Our surprising finding is that pre-trained LLMs, equipped with a light inference harness, can serve as capable agents. Despite far lower accuracy (pass@1), they often surpass their post-trained counterparts in solution coverage (pass@K) given a sufficient test-time budget. We further analyze the underlying mechanism and show that post-training pushes tasks toward two extremes, always solved or never solved, and thereby improves sampling efficiency and consistency at the cost of solution coverage. To measure this cost, we propose Sharpening Tax, a diagnostic metric that quantifies the loss in test-time scalability after post-training. Across 14 base/post-trained model pairs from four families and three agentic benchmarks (42 cases in total), the tax is prevalent in most settings, can be estimated from a few rollouts, and correlates well with other metrics. Finally, we present posterior-tempered group sampling (PTGS), a simple plug-and-play Bayesian sampler that adapts the sampling temperature per prompt to its estimated difficulty. Applied during RL training in two agentic environments, PTGS pays a smaller tax than the fixed-temperature baseline, solving more tasks under repeated sampling while also improving single-shot accuracy.
[AI-62] MCRI: A Four-Dimensional Framework for Analyzing and Evaluating Agent Skills
链接: https://arxiv.org/abs/2610.01506
作者: Zongrui Yang,Li Xintong,Runchen Xu,Zhongsheng Wang,Zhedong Lin,Haoyuan Li,Jiamou Liu
类目: Artificial Intelligence (cs.AI); Software Engineering (cs.SE)
备注: 24PAGES
Abstract:As agents evolve from single-tool systems into modular, composite architectures, skills are becoming an important mechanism for capability development and distribution. However, the academic community lacks a structured framework for systematically analyzing and evaluating skills. Drawing on information gain and behavioral constraint, we propose the four-dimensional MCRI Framework and operationalize it as MCRI-Eval, a large language model-based evaluation method. We evaluate MCRI-Eval using 63,812 public skills from the OpenClaw skill Hub, with 58,275 skill-conditioned model executions across BigCodeBench, BFCL-Fundamental, and Mind2Web. MCRI-Eval scores are positively associated with community popularity signals and achieve the highest downstream ranking agreement among the evaluated methods. MCRI-Eval also improves top-1 skill selection across all three benchmarks: compared with the strongest baseline on each benchmark, the skills selected by MCRI-Eval advance by 17.7, 22.8, and 19.6 percentile points in downstream performance rank on BigCodeBench, BFCL-Fundamental, and Mind2Web, respectively. These results indicate that MCRI-Eval provides a useful pre-execution signal for prioritizing promising skills before costly execution-based evaluation.
[AI-63] OpenMTB-Audit: Exposing Over-Refusal and Clinical Expert Perspectives in LLM -Based Molecular Tumor Board Safety Evaluation
链接: https://arxiv.org/abs/2610.01497
作者: Negin Ashrafi,Jia Luo,Stacey M. Frumm,Roxana Daneshjou
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: Accepted for oral presentation and publication at the Pacific Symposium on Biocomputing (PSB) 2027
Abstract:Molecular tumor boards integrate genomic findings, clinical context, and therapeutic evidence to support precision oncology. As AI enters this workflow, a key safety challenge is distinguishing truly unsupported recommendations from evidence-supported options that still require oncologist review because of incomplete information, poor ECOG performance status, or other clinical caveats. We introduce OpenMTB-Audit, an open-source benchmark of 500 synthetic non-small cell lung cancer cases spanning five adversarial error categories and four safety labels: Supported, Partially Supported, Unsupported, and Insufficient Information. Across eight large language model configurations, we identify pervasive over-refusal: all LLM configurations failed to retain the Partially Supported label in 83.3-100% of true Partially Supported cases, achieving high aggregate safety scores through label collapse rather than clinically calibrated reasoning. To address this limitation, we developed MTB-AuditAgent, a deterministic seven-module framework separating evidence verification, missing-information detection, safety classification, and abstention. It reduces over-refusal to 6.7% and achieves 91.2% accuracy (95% CI: 88.6-93.6%). A two-oncologist annotation study found disagreement concentrated at the boundary between information sufficiency and treatment optimization, underscoring the need to preserve clinically meaningful distinctions.
[AI-64] Auditing Routing Entropy as an Uncertainty Signal in Attention-Residual Transformers
链接: https://arxiv.org/abs/2610.01495
作者: Wenhao Liang,Lin Yue,Wei Emma Zhang,Mingyu Guo,Olaf Maennel,Weitong Chen
类目: Artificial Intelligence (cs.AI)
备注: 36 pages (9 pages main text)
Abstract:Dynamic architectures leave a per-example routing trace beside each prediction, and diffuse routing is easy to read as a sign that the prediction is unreliable. We audit that reading for routing entropy in Attention-Residual (AR) variants of Swin-Tiny and DeiT-Small, trained from scratch on CIFAR-10/100 with a soft-binned calibration auxiliary loss, asking whether the trace carries information about correctness beyond what the model’s own confidence already reveals. Three checks probe this increment: does a routing signal appear at fixed confidence, does it replicate across training seeds, and can a held-out predictor exploit it against output-only and shuffled-trace controls? A sensitivity audit then injects effects of known size and measures the fraction of each that the probes recover. No test in the fixed 30-test binned family survives multiplicity correction, and neither the nominal hit nor a borderline result recurs in its sibling seeds. Across 24 paired runs a scalar routing probe yields no pooled improvement in routing-stratified calibration, and an entropy-profile probe predicts correctness better than the same probe given shuffled profiles yet worse than a confidence-only predictor in both binary log-loss and Brier score: a gain over shuffled traces does not become a gain over the output. Conditioning on the complete logit vector leaves the corresponding comparison unresolved. The audit bounds how far these non-detections can be read: at an injected effect of 0.010 nats the profile probe recovers 24-59% of the oracle gain, and a reference-preserving correction probe recovers 8% and 23% in the two CIFAR-100 settings, below the threshold we fixed for applying it to real labels. The results establish control-dependent gains and incomplete estimator recovery, not the absence of conditional routing information.
[AI-65] Multi-Party Backchannel Prediction: a Diagnosis a Benchmark and a Ceiling NEURIPS2026
链接: https://arxiv.org/abs/2610.01488
作者: Mohammed Hafsati,Ahmed Loughzali
类目: ound (cs.SD); Artificial Intelligence (cs.AI)
备注: Accepted at the NeurIPS 2026 workshops ReMuCAI (Paris) and RTCA (Sydney). 8 pages main text, 9 figures, 5 tables, plus appendices. Code and benchmark: this https URL
Abstract:Backchannel prediction has been studied almost entirely in dyadic conversation. We introduce a multi-party benchmark based on the AMI corpus, comprising 682 masked-listener views from 171 meetings, 190 speakers, and 18,697 backchannel events, with a person-disjoint held-out split. A state-of-the-art dyadic model applied zero-shot to meeting audio performs at chance (AUROC 0.499); nevertheless, its frozen acoustic features remain informative: a linear probe reaches 0.704, and retraining the predictor raises performance to 0.751. Retraining reveals a second limitation. Listener conditioning improves prediction for listeners seen during training but not for unseen listeners, and the gap remains under capacity reduction, listener-adversarial training, per-listener adaptation, and oracle lexical conditioning. Adversarial training removes only part of the speaker-identity information, while stronger removal hurts prediction, suggesting that identity is entangled with cues that are useful for backchanneling. A within-model control helps explain this pattern: with the same features and data splits, turn-onset prediction transfers to unseen listeners, while backchannel prediction does not. Backchannel rates also vary about twice as much across individuals as turn-onset rates. Since backchannels occupy only about 1% of frames, frame-level F1 is strongly affected by the base rate. We therefore report AUROC alongside event-F1 on listener-active regions. We release the benchmark and evaluation tools at this https URL.
[AI-66] NextMe-800: Anticipating Personal Behavior from Months of Egocentric Video
链接: https://arxiv.org/abs/2610.01461
作者: Zhaoxu Meng,Yiming Sun,Mingyuan Gao,Jiachang Zhang,Zhuhan Dai,Yipeng Du,Zheng Lian,Jian-Qiao Zhu
类目: Artificial Intelligence (cs.AI)
备注: 24 pages, 7 figures. Dataset and benchmark: this https URL ; project page: this https URL
Abstract:We often plan ambitiously yet act habitually and wonder, in retrospect, whether we would have planned differently had we known what we would actually do. Hindsight offers a valuable perspective on past decisions, although we often wish we could have simulated hindsight at the moment of choosing. If a system could generate plausible trajectories from one’s personal history, such previews might help people formulate more realistic plans and make better informed decisions. We introduce NextMe-800, an approximately 800-hour first-person dataset from one volunteer over 126 days with 1 Hz images, gaze, and audio, captioned at five hierarchical abstraction levels from atomic actions to major activities. We formulate personalized action anticipation as open-vocabulary K-step sequence prediction and construct NextAct, a 1,500-point benchmark combining NextMe-800 with the multi-person EgoLife dataset. Using an embedding-based soft edit distance as the metric, we evaluate how well different models can anticipate personal behavior across abstraction levels and prediction horizons. NextMe-800 and NextAct provide a months-long resource and evaluation framework for studying how far ahead personal behavior can be anticipated from egocentric observation.
[AI-67] Rethinking Probability-Based Reinforcement Learning From Posterior Concentration
链接: https://arxiv.org/abs/2610.01458
作者: Shiu-Hong Kao,Yubo Zhao,Zhenyu Tian,Pengzhan Sun,Yicong Li,Angela Yao
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Verifier-free reinforcement learning with probability-based rewards offers a promising way to train LLMs on general reasoning tasks where external verifiers are unavailable. Yet the reliability of these rewards, especially in long-horizon reasoning, remains underexplored. This work identifies a length-dependent failure mode of probability rewards, which we call the Posterior Concentration Phenomenon (PCP). We show that the probability of a reference answer conditioned on a reasoning trace often collapses to a low-variance interval as the trace becomes lengthy. This phenomenon results in nearly indistinguishable rewards, which, under GRPO-based settings, makes probability-based policy optimization unstable and inefficient. Motivated by this, we propose Reinforcement Learning with Concentration-aware Posterior Rewards (RLCPR), a verifier-free RL framework to explicitly account for PCP for better optimization stability and token efficiency. It has two components: uncertainty-aware data sampling, which reduces concentration-prone rollouts before generation, and concentration-aware regularization, which penalizes unnecessarily long traces when posterior rewards collapse. Extensive experiments show that, alongside higher token efficiency, RLCPR outperforms the state-of-the-art verifier-free RL baseline by up to 4.0% on six of seven benchmarks, including general-domain and mathematical reasoning challenges.
[AI-68] DRelay: Global Draft Context for Prefix-Aware Parallel Speculative Decoding Repair
链接: https://arxiv.org/abs/2610.01439
作者: Zhuoyu Wang,Junnan Huang,Xinyu Chen
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Parallel drafting reduces the drafting overhead of speculative decoding for large language models (LLMs), but its gains remain limited by the accepted prefix length. Even when the correct token is present in the candidate pool, a single early selection error prevents subsequent predictions from being used. We propose DRelay, which uses global information from the entire draft block to perform prefix-aware selective repair of candidate selections before target-model verification. DRelay bases its decisions on candidate correlations and the selected path: a global reader extracts predictive information across positions for each candidate. While a causal selector combines candidate-level information extracted by the global read with the tokens selected at preceding positions to determine whether the native choice at the current position is consistent with the global evidence and the selected prefix. It then decides whether to retain or replace the token, thereby repairing early errors and extending the accepted prefix. We further jointly train the draft backbone and the selector, combining candidate-support learning with a repair objective, while weighting the repair loss according to each block position’s potential contribution to the consecutive accepted prefix. Across eight diverse benchmarks on an H800 GPU, DRelay consistently improves both average acceptance length and end-to-end decoding performance over DFlash, Domino, and DSpark. Under SGLang serving, DRelay improves average end-to-end speedup over DFlash, Domino, and DSpark by 14.7%-16.8%, 8.7%-9.3%, and 8.1%-9.3%, respectively.
[AI-69] A Deterministic and Auditable AI Security Risk Assessment Framework with ATLAS Aligned Executable Rules and Formal Verification
链接: https://arxiv.org/abs/2610.01436
作者: Yixuan Huang(1),Basel Halak(1),Boojoong Kang(1) ((1) University of Southampton, Southampton, UK)
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Artificial intelligence systems are increasingly deployed in high impact and safety critical settings, yet security assessment remains difficult to reproduce and defend under audit. Existing approaches often rely on narrative checklists or assessor driven scoring, and they lack an explicit, machine evaluable mapping from observable engineering artefacts to stable technique level outcomes. We present an evidence driven AI security assessment framework that operationalises assessment as a deterministic decision function. The framework normalises heterogeneous artefacts into a project independent Control ID taxonomy scored on a bounded four level ordinal scale, compiles technique level predicates from a pinned MITRE ATLAS snapshot via an explicit mitigation to control mapping, and outputs technique indexed feasibility and impact levels with traceable links back to the triggering evidence. We package all normative choices as a versioned assessment policy object to support repeatable reassessment across snapshots. To ensure semantic correctness, we formally verify boundedness, totality, ordered semantic consistency, and monotonicity of the compiled evaluator over the full declared score domain. We evaluate the framework on five public open source AI projects pinned to explicit repository snapshots, quantify before and after changes under a unified hardening intervention, and validate responsiveness to real engineering changes through fork based implementations of Software Bill of Materials (SBOM) generation and Continuous integration (CI) security scanning gates. Results show consistent downward shifts in feasibility profiles under strengthened observable controls, while worst case residual feasibility persists when technique specific core controls remain absent from the evidence scope.
[AI-70] SpikeMoE: Brain-Inspired Competitive Routing for Flexible Spiking Mixture-of-Experts
链接: https://arxiv.org/abs/2610.01418
作者: Xiaoli Liu,Yujie Liang,Jialin Li,Malu Zhang
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Spiking Neural Networks (SNNs) enable event-driven computation through biologically inspired dynamics at the neuronal scale, while Mixture-of-Experts (MoE) perform conditional computation through expert selection at the model scale. Integrating their strengths offers potential for flexible neural architectures. A key challenge, however, lies in designing an expert selection mechanism based on spiking activity. To address this, we introduce a spike-based k-WTA Router inspired by competition-inhibition observed in the hippocampal CA1 region. The router incorporates lateral inhibition and refractory period to select Top-K experts according to discrete spike counts. Building on this, we present SpikeMoE, a framework that integrates neuronal-scale spiking dynamics with model-scale expert selection. To address incomplete multisensory inputs in multimodal tasks, we further equip SpikeMoE with a two-stage missing-modality modeling module that combines empirical prototypes from an observed-modality pool with modality-specific learnable embeddings to construct missing-modality representations. Experiments on vision, language, and multimodal benchmarks demonstrate that SpikeMoE achieves state-of-the-art performance among the SNN baselines, matches or exceeds the performance of ANN counterparts, and maintains robustness across diverse missing-modality conditions. These results demonstrate a favorable trade-off between performance and energy efficiency, validating the integration of spiking dynamics with sparse expert computation and highlighting SpikeMoE as a promising approach to energy-efficient brain-inspired computing.
[AI-71] Beyond Memory: Harnessing Long-Horizon Agents with Explicit Belief States
链接: https://arxiv.org/abs/2610.01415
作者: Yu Luo,Jiamin Jiang,Yimin Zuo,Xidao Wen,Rongchen Gao,Yongqian Sun,Shenglin Zhang,Guiyang Liu,Cheng Zhang,Fang Situ,Qi Zhou,Dan Pei
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Large language model (LLM) agents can now undertake increasingly complex tasks, but the way they organize interaction history into memory does not ensure a coherent understanding of the current world. We introduce PoS, an inference-time framework that constructs and continually maintains explicit belief states as the agent’s decision context. Each belief combines an estimate of the current world state with unresolved task requirements, making explicit what the agent still needs to learn and accomplish. To keep this belief reliable and actionable, PoS validates its consistency and monitors task progress to detect Belief Trapping, where the agent continues to act without making meaningful progress toward the goal. Recovery is then tailored to both the trapping pattern and the type of unresolved task requirement. Experiments on four benchmarks spanning execution and diagnosis show that PoS achieves the highest overall performance on every benchmark with all three LLM backbones. Ablations demonstrate the importance of consistency validation and recovery, while context-scaling experiments show resilience to context growth. Together, these results support belief construction and continual maintenance as a foundation for long-horizon context management beyond history retention and compression.
[AI-72] Contrastive Attention Mitigates Spectral Bias in Spiking Transformers
链接: https://arxiv.org/abs/2610.01403
作者: Xiaoli Liu,Malu Zhang,Yang Yang
类目: Artificial Intelligence (cs.AI)
备注: Spiking Neural Networks
Abstract:Spiking Transformers merge the energy-efficiency of spiking neural networks (SNNs) with the representational power of self-attention, creating a promising architecture for high-performance, energy-efficient computation. However, a performance gap persists versus its counterparts in artificial neural networks (ANNs). Unlike prior works attributing this to binary activations, we reveal that both spiking neurons and spiking self-attention (SSA) act as low-pass filters through multiscale spectral analysis. This characteristic leads to the dissipation of high-frequency components. To address this issue, we propose the Spiking Contrastive Attention (SCA) paradigm, which draw inspiration from the edge-detection and differential sensing properties of biological visual system. By extracting contrast prototypes via global contrastive aggregation and applying local differential refinement, SCA effectively enhances high-frequency information. Extensive experiments show that SCA is a general module that consistently boosts Spiking Transformers across image classification, semantic segmentation, and event-based tracking. Furthermore, it achieves lower complexity, offering superior efficiency over original SSA. These results establish its potential as a fundamental building block for energy-efficient Spiking Transformers.
[AI-73] PRISM: A Category-Theoretic Framework for Measuring and Refining Multimodal Analogies
链接: https://arxiv.org/abs/2610.01383
作者: Mirella Zeisler,Ojas Shirekar,Mircea Licǎ,Chirag Raman
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Analogical reasoning involves identifying and preserving relational structures across domains. However, existing approaches to AI-driven multimodal analogy generation lack an interpretable measure of whether this structure is understood and maintained in the generated output. We address this gap with Pullback Refinement via Interpretable Structural Mapping (PRISM), a modality-agnostic framework for measuring and improving relational alignment in multimodal analogies, evaluated on visual metaphor generation. PRISM represents analogies as explicit relational mappings grounded in category theory and uses VLMs to instantiate these structures across modalities. Its first component, the pullback score, quantifies relational alignment from the resulting graph representation. On the AnaloBench benchmark, selecting the correct analogy purely by pullback score achieves 82.5% accuracy, demonstrating that the score captures meaningful relational information. PRISM’s second component is an iterative refinement loop that uses the pullback score as an in-context feedback signal to iteratively revise the generated image towards greater relational depth. VLM-as-a-judge and human evaluations show that PRISM consistently improves metaphor consistency and analogy appropriateness over zero- shot generation, with human participants preferring the refined output in 57.65% of pairwise comparisons. However, a qualitative analysis reveals that refinement can favour visually crowded compositions rather than genuinely deeper relational correspondences.
[AI-74] Discrete Wasserstein Flows for One-Step Generative Modeling
链接: https://arxiv.org/abs/2610.01355
作者: Alessandro Micheli,Andrea Zerio,Samir Bhatt
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:
Abstract:We introduce a new framework for one-step generative modelling on finite state spaces. To extend drifting beyond continuous domains, we use discrete Wasserstein geometry to define a target-relative KL gradient flow over the transitions of a reversible Markov kernel. We realize this probability flow at the particle level through Markov jumps and amortize the resulting transport updates into a latent-conditioned generator, so that the iterative dynamics are required only during training while inference remains one-step. In a controlled setting where the underlying distributions and transport dynamics can be computed exactly, we verify KL dissipation, consistency between the particle dynamics and the probability flow, and the predicted numerical scaling. We further show that a finite-capacity neural generator can track these exact transport targets while retaining one-step generation. These results validate the basic construction and provide a foundation for scaling Discrete Drifting to structured discrete data.
[AI-75] PACE: Provenance-Aware Capability Enforcement for Tool-Using LLM Agents
链接: https://arxiv.org/abs/2610.01349
作者: Fengpeng Li,Qizhou Wang,Yuke Hu,Kemou Li,Jun Liu,Haiwei Wu,Jiantao Zhou,Di Wang
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
备注:
Abstract:Tool-using large language model (LLM) agents turn generated text into real side effects, so poisoned tool metadata, retrieved pages, memory, and reusable skills can steer the next call. Vetting an artifact before admission does not settle this. A safe variant and a leaking variant can produce the same admission evidence, and a sound gate then cannot relax that site for either. We make that condition precise, which leaves the last boundary a deployment can still act on. We present Provenance-Aware Capability Enforcement (PACE), which mediates every tool call immediately before it executes. Path confinement proposes an executable cut of represented influence paths, while capability and effect verification checks schema-defined effects against authority compiled from the authenticated request. We distinguish the certified execution contract from the evaluated configuration, which can restore an authorized call after a proposed block or apply a declared repair. Confinement requires the final action to preserve the certified cut. On eight executable agent-security benchmarks with three target-model families, the evaluated configuration gives strictly lowest attack success in 62 of 79 eligible attack columns and ties in 14; full-benchmark native utility loses at most three points relative to the undefended agent. A complete ablation over 1167 paired cases attributes most security gains to effect verification and refusal control to boundary adaptation. A reduced-scale adaptive search succeeds on 0/30 out-of-authority targets against the defense.
[AI-76] Verify Claims Not Scores: Evidence-Based Verification of Modular Agents
链接: https://arxiv.org/abs/2610.01348
作者: Ali Atiah Alzahrani
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Portfolio Management (q-fin.PM)
备注: 32 pages, 4 figures, 15 tables
Abstract:When developers change one component of an agent, such as its controller, a learned model or its verifier, they usually judge the change by an aggregate task score. That score cannot tell whether improvement was attainable, which component lost value, or what the agent’s own checks certify. We introduce a claim-specific verification audit for modular agents that plan, act, check and refine. Instead of scoring the agent, the audit scores the evidence: each conclusion is recorded with the evidence behind it, one of four verdicts (supported, unsupported, unresolved or not evaluated) and the boundary within which it holds. Three tools supply that evidence. Oracle policies measure attainable improvement under an explicitly stated action set, so that a low value can be traced to the evaluation rather than to the environment. Replacing one component at a time with a perfect counterpart locates lost value, with null results read as unresolved whenever a downstream component could mask them. A separate test asks whether the verifier’s score identifies the quantity it is read as bounding. Applied to a constrained portfolio-allocation agent in a synthetic market with known hidden regimes, the audit shows that the value of perfect regime information depends on the action set used to measure it, that the scenario generator discards most of the regime signal while better local fidelity does not improve decisions, and that the runtime verifier can be bypassed with no visible change in outcomes. The contribution is the protocol and the evidential distinctions it enforces; the empirical findings are specific to the agent and environment studied.
[AI-77] An ontology for cross-sectoral crisis management: core and public health modules
链接: https://arxiv.org/abs/2610.01326
作者: Aldo Gangemi,Rita T. Sousa,Luigi Asprino,Giorgia Lodi,Andrea G. Nuzzolese,Valentina Presutti,Johannes Gysen,Diana F. Sousa,Luigi Spagnolo
类目: Artificial Intelligence (cs.AI); Logic in Computer Science (cs.LO)
备注: 17 pages, 2 figures
Abstract:This paper presents the European Crisis Management Ontology (ECMO), a modular OWL-based ontology intended as a cross-sectoral reference for disaster risk reduction and response. ECMO is designed to be organised as a network of ontological modules. Among the modules, ECMO-CORE captures fundamental crisis management concepts such as hazard, event, exposure, impact, and response measure and uses ontology design patterns and the OWL2 punning technique to resolve ambiguities between hazard types and event manifestations. In addition, domain-specific modules are defined as in the case of the public health module aligned with SNOMED CT and ICD-11. To demonstrate the resource’s utility, we used ECMO to represent the data of the Epidemic Intelligence from Open Sources system of the Joint Research Centre to generate an end-to-end pipeline that populates an ECMO-compliant knowledge graph from unstructured epidemiological news. Initial results demonstrate that ECMO provides the formal guardrails necessary for consistent and unified knowledge representation and integration. The ontology is publicly available at this https URL and is released under the Creative Commons Attribution 4.0 International (CC BY 4.0) license.
[AI-78] PPO-HRAP: Proximal Policy Optimization with a Hybrid Regime-Aware Policy for Risk-Controlled Trading
链接: https://arxiv.org/abs/2610.01325
作者: Duong Hien Chi Kien,Thanh Trung Huynh
类目: Artificial Intelligence (cs.AI); Trading and Market Microstructure (q-fin.TR)
备注: 8 pages, 6 figures, 8 tables. Code: this https URL
Abstract:Reinforcement learning for trading often struggles to balance upside participation with drawdown control. Profit-only policies can collapse toward passive long exposure on upward-drifting assets, while aggressively risk-penalized rewards can become too defensive during volatile periods. This paper proposes PPO-HRAP, a hybrid regime-aware policy that combines Proximal Policy Optimization with an interpretable regime prior. The agent observes both market features and portfolio-state variables, receives a reward combining portfolio log return, VIX-conditioned drawdown-increase penalty, target-exposure deviation, and turnover cost, and executes a blended action between the PPO actor output and a regime-derived target exposure. On the held-out 2020-2022 SPY test window, PPO-HRAP achieves 27.62% total return, 8.48% annualized return, 0.6447 Sharpe ratio, 0.8588 Sortino ratio, and 0.4592 Calmar ratio, while reducing maximum drawdown from 34.10% for Buy and Hold to 18.47%. Across five SPY seeds, PPO-HRAP remains stable with mean total return 0.2725 \pm 0.0109 and mean Sharpe ratio 0.6219 \pm 0.0565 . Single-run cross-asset tests on QQQ and DIA further show that the proposed method ranks first on total return and Sharpe ratio for all three reported assets. These results suggest that blending learned actions with a volatility-aware regime prior is a practical way to improve risk-adjusted trading behavior, although the current policy still incurs high turnover and cross-asset robustness beyond SPY remains limited to single-run evidence.
[AI-79] RACE: Trajectory Return Attribution and Contrastive Erasure for Multi-Turn Safety
链接: https://arxiv.org/abs/2610.01323
作者: Fengpeng Li,Kemou Li,Qizhou Wang,Haiwei Wu,Jiantao Zhou,Di Wang
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Safety-aligned large language models (LLMs) often refuse a harmful request but comply once the same goal is spread over several turns. Preference objectives score whole responses to single prompts, so their training loss alone cannot control risk on unseen histories. Our analysis gives sufficient conditions under which suppression at supervised single-turn contexts yields a bound on multi-turn trajectory risk. The bound accounts for coverage, transfer slack, and leakage, and characterizes contraction relative to a base-policy risk budget evaluated on the trained policy’s contexts. TRACE (Trajectory Return Attribution and Contrastive Erasure) turns this principle into a token-level objective. On the safe response, each token is weighted by the discounted return of a refusal-attributable advantage. The advantage compares a frozen reference model with its refusal-ablated copy, allowing earlier response tokens to receive credit from later refusal-related evidence. At high-gap positions on rejected responses, TRACE combines the observed token with policy-selected alternatives in the erasure target. A gradient-norm penalty replaces the retain set. Across five open-weight models and seven multi-turn attacks, TRACE gives the lowest attack success rate (ASR) in all 35 model and attack pairs, while the model utility evaluated on MMLU and HellaSwag drop by at most 1.23 points. Source code can be found in the supplemental material.
[AI-80] ProtoFlow: Prototype-Guided Flow Matching for Multivariate Time Series Forecasting
链接: https://arxiv.org/abs/2610.01320
作者: Shibo Feng,Wanjin Feng,Yang Qiu,Deheng Ye,Peilin Zhao,Chunyan Miao
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Generative modeling has shown strong promise for multivariate time mseries (MTS) forecasting, especially scale to high-dimensional settings. Diffusion-based methods achieve competitive performance but typically require many sampling steps at inference. VAE-based non-iterative forecasting frameworks have therefore emerged as an efficient alternative. Within this line of work, vector quantization (VQ) enables controllable latent space modeling by mapping multivariate series into compact discrete representations. Existing VQ-based forecasting methods, however, typically rely on autoregressive (AR) token generation, which suffers from exposure bias and training-inference mismatch. Flow matching provides an efficient non-autoregressive alternative for latent forecasting, but existing formulations usually initialize transport from a generic Gaussian prior. We instead observe that the trained VQ codebook already captures representative latent prototypes and can thus serve as a more informative prior for flow matching. Based on this insight, we propose ProtoFlow, a forecasting framework that combines vector-quantized autoencoding with Prototype-prior Flow matching. Our method first maps multivariate sequences into a discrete latent space, then constructs a structured prior from the learned codebook, and finally learns a DiT-based rectified flow to transport samples from this prior to future latent representations conditioned on historical observations. By replacing generic noise initialization with a learned prototype prior, ProtoFlow avoids the rollout mismatch of AR token prediction and promotes faster training convergence. Extensive experiments on benchmark datasets show that it consistently achieves superior forecasting performance with efficient inference.
[AI-81] Feature Selective Model Collapse in Diffusion Models: Total Replacement versus Fixed-Budget Training NEURIPS2026
链接: https://arxiv.org/abs/2610.01318
作者: Hanna Malet,Gabriel Turinici
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Statistics Theory (math.ST)
备注: Neurips 2026 PriGM workshop paper
Abstract:Model collapse arises when generative models are trained on synthetic data produced by earlier models. The phenomenon has attracted considerable attention because of its societal and technical implications. However, previous studies have reached seemingly contradictory conclusions: replacing real data with synthetic data causes collapse (Shumailov et al.), yet accumulating real data alongside synthetic data can prevent it. For diffusion models, we study an intermediate regime typical of finite-budget pipelines: all past datasets and the real data are kept, but each new model is trained on a fixed-size sample from this growing pool, so the real fraction vanishes without any data being removed. Experiments on a 2D spiral dataset as well as the image benchmarks (MNIST, Fashion-MNIST, and CIFAR-10) show that replacement protocol degrades dataset rapidly as in the literature, whereas the fixed budget degrades only partially, sparing some features. A linear-response model of the multi-generational parameter dynamics, analyzed by stochastic recursion, confirms that the two protocols differ: some features will be fragile and lost within a few generations for both protocols, while some will be robust and preserved over practically unbounded horizons under the fixed budget protocol.
[AI-82] Federated Learning for LLM s over Mobile Networks: Issues and Solutions in the RAN Transport
链接: https://arxiv.org/abs/2610.01304
作者: Emilio Paolini,Andrea Pinto,Flavio Esposito,Luca Valcarenghi
类目: Networking and Internet Architecture (cs.NI); Artificial Intelligence (cs.AI)
备注:
Abstract:Federated LLM fine-tuning enables large models to be adapted using private and geographically distributed data at the network edge, creating recurring and deadline-sensitive communication workloads across access and transport networks. This challenge is particularly relevant in mobile RANs, where wireless variability, mobility, and device heterogeneity cause model updates to arrive asynchronously. Although these updates belong to the same learning round and share a common destination and deadline, conventional transport networks treat them as independent device-originated flows, hiding their underlying structure and limiting the ability to efficiently provision transport resources. This mismatch is particularly problematic for optical circuit switching and all-photonics transport, which benefit from predictable and schedulable traffic demands. We argue that future RANs should act as learning-aware traffic shapers by exposing the communication structure of distributed model adaptation to the transport layer. Through in-network aggregation at the gNB, asynchronous UE updates can be transformed into fewer aggregate transfers with bounded size and delivery requirements. Once shaped in this way, federated LLM traffic becomes a suitable candidate for selectively provisioned optical connectivity, where high-capacity paths can be established during aggregate-transfer windows and released between learning rounds. The resulting architecture combines the flexibility of packet-based mobile access with dynamically provisioned optical capacity, illustrating a broader approach for coordinating distributed AI workloads across programmable access and transport networks.
[AI-83] Questionnaire-Guided Disaggregation of Energy Appliance Use for Domestic Smart Meter Data
链接: https://arxiv.org/abs/2610.01297
作者: Achal Nanjundamurthy,Rupam Misra,Suzanne Little,Alan F. Smeaton
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Ireland’s smart metering programme records electricity use at 30-minute resolution, with smart meters installed in over 80% of households as of late 2025. While this is useful for billing of smart, time-of-use tariffs, it is too coarse to capture use of domestic appliances. We present a label-free disaggregation system that breaks usage data into 9 appliance categories by combining event detection for high-power loads with questionnaire-guided estimation. Our evaluation draws on four datasets: a calibration household with a commercial comparator, two public benchmarks (UK-DALE and REFIT) with per-appliance sub-metering, and a smart meter dataset of more than 4,800 years of use from 2,968 Irish consumers. Compared against two independently developed disaggregation systems our hybrid method combining analysis of usage data with questionnaire results, achieves the lowest whole-decomposition error on all buildings across the datasets, with better month-level performance over 54 paired months ( p0.001 , Holm-corrected). Our method provides useful advice on a household’s energy consumption patterns and advice on how to reduce or shift usage on some appliances in order to reduce costs.
[AI-84] ITC-MoE: Importance-guided Token-aware Compression for MoE Diffusion Language Models
链接: https://arxiv.org/abs/2610.01296
作者: Lianjun Liu,Shipeng Li,You Huang,Weiqi Yan,Mingte Qiu,Huazhong Liu,Xiaofeng Zhu,Yunshan Zhong
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Mixture-of-Experts (MoE) Diffusion Language Models (DLMs) offer flexible parallel decoding and increased model capacity, but their large number of expert parameters incurs substantial computation and storage costs. Existing low-rank MoE compression methods largely rely on static factorization and fixed rank allocation, which overlook the distinctive properties of MoE DLMs. Specifically, we identify two properties: cross-mode non-uniform redundancy, where parameter redundancy and sensitivity to rank truncation vary across the input, output, and expert modes, and token-wise utilization variation, where hot and cold tokens exhibit distinct spectral characteristics and expert activation patterns. To address these challenges, we propose ITC-MoE, an Importance-guided Token-aware Compression framework for MoE DLMs. ITC-MoE consists of two complementary components. First, Importance-guided Adaptive Tucker Compression (IATC) incorporates activation and gradient importance into expert weight transformation, jointly factorizes expert weights across multiple modes, and adaptively allocates ranks under a fixed parameter budget. Second, Token-aware Compensation and Routing (TCR) applies lightweight low-rank compensation to compression-sensitive hot tokens and restricts the candidate expert set for cold tokens with concentrated routing patterns. By jointly adapting compression capacity and inference execution to both parameter redundancy and token-wise variation, ITC-MoE substantially reduces the computation and storage costs of MoE DLMs while preserving their generation quality. For example, on SDAR-30B-A3B-Chat-b32, ITC-MoE maintains an accuracy of 96.33% on MultiArith under a 30% compression budget, while achieving up to a 7.22x end-to-end speedup. The code is publicly available at this https URL.
[AI-85] Model validation in machine learning: A scenario-based guide from hold-out splits to nested group cross-validation in biomedical and applied research
链接: https://arxiv.org/abs/2610.01284
作者: Mehmet Baygin,Sengul Dogan,Turker Tuncer
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: Tutorial with eight controlled scenarios; includes MATLAB and Python/scikit-learn code listings
Abstract:Model validation estimates the performance of a complete learning procedure on new data. However, an invalid split can produce an optimistic and stable result. This tutorial reviews hold-out validation, train/validation/test designs, repeated random subsampling, k-fold and repeated stratified cross-validation, leave-one-out and leave-p-out schemes, group-aware validation, and nested group cross-validation. General machine-learning principles are linked to EEG epochs, paired-eye OCT images, repeated clinical measurements, and multicenter data. Eight controlled scenarios compare flawed and leakage-safe designs: seven use locked confusion matrices with auditable metrics, and one uses a reproducible repeated-study simulation. The scenarios cover global feature selection, normalization leakage, dependent records, center mixing, repeated test-set use, and estimator instability. Bias, variance, metric aggregation, uncertainty, and computational cost are also examined. A data-size matrix, a decision tree, and reporting checklists are provided. Reproducible MATLAB templates and scikit-learn counterparts are included. The results show that no validation method is universally best. The independent unit must match the intended deployment target. Every data-dependent operation must also exclude the observations used for performance estimation.
[AI-86] rustworthy Data- and ML-Ops for Intelligent Transportation Systems and Logistics
链接: https://arxiv.org/abs/2610.01282
作者: Antonio Emanuele Cinà,Giovanni Scodeller,Cecilia Caterina Pasquale,Silvia Siri,Davide Anguita,Fabio Roli,Simona Sacone,Luca Oneto
类目: Artificial Intelligence (cs.AI); Software Engineering (cs.SE)
备注: Paper accepted at accepted at IEEE Transactions on Intelligent Transportation Systems. DOI: https://doi.org/10.1109/TITS.2026.3711756
Abstract:The rapid evolution of Intelligent Transportation Systems and Logistics (ITS\L) has become a cornerstone of the modern social economy, relying heavily on the integration of Data, Artificial Intelligence (AI), and, more specifically, Machine Learning (ML). This paper provides a comprehensive review of Trustworthy Data and Machine Learning Operations (DataOps and MLOps) in the ITS\L domain, underscoring their importance in improving efficiency, reliability, and decision-making precision within transportation and logistics services. We begin by identifying gaps in current literature, offering clear context for our contribution. Subsequently, we explore the complexities of DataOps and MLOps, discussing their necessity, key components, available tools, practical insights, and case studies relevant to ITS\L. Additionally, we address the critical issue of Trustworthiness in AI applications, examining methods and tools designed to strengthen confidence in AI systems - especially in real-world ITS\L scenarios. The paper concludes with a discussion of persisting challenges and future prospects in this rapidly advancing field, aiming to serve as a vital resource for researchers, industry practitioners, and policy makers. Overall, this work not only establishes a foundational understanding of DataOps and MLOps in ITS\L but also charts a path for further research and innovation in developing more efficient, sustainable, and trustworthy intelligent transportation and logistics systems.
[AI-87] Feedback Without the Wait: Piloting a Generative AI Practice Platform in a Large Maths Class
链接: https://arxiv.org/abs/2610.01262
作者: Lili Chen,Gavin Buskes,Yuxin Ren,Chin Tong Leong
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Timely and specific feedback is one of the strongest influences on student learning, yet it is difficult to sustain in large electrical engineering classes where the ratio of students to demonstrators is high and a learner who is stuck may wait days to find out why an approach was wrong. Generative Artificial Intelligence (GenAI) offers a way to scale conversational feedback, but using it to grade assessed work raises trust and accountability concerns, and keeping a human in the loop to assure its judgements reintroduces the very delay that erodes the value of feedback. The result is a tension between the immediacy that makes feedback so impactful and the human oversight that makes it trustworthy. In this work, we set out to resolve that tension in practice by designing and piloting a GenAI practice platform that delivers immediate, scaffolded feedback during self-directed practice. This relocates human oversight from real-time grading to the upfront verification of solutions. Our goal was to understand how students engaged with the tool, how they perceived the value and reliability of its feedback, and what lessons transfer to other engineering subjects.
[AI-88] DeFA: Dependency-Guided Failure Attribution for LLM Agents
链接: https://arxiv.org/abs/2610.01256
作者: Bo Deng,Xinlei Zheng,Yi Wei,Kang Zhou,Chongyang Tao,Renzhao Liang,Xuanren Chen,Lifan Guo,Chi Zhang
类目: Artificial Intelligence (cs.AI)
备注: DeFA: Dependency-Guided Failure Attribution for LLM Agents
Abstract:Errors in LLM agent executions and their visible consequences can be separated by many steps, making decisive-error localization a matter of understanding both step content and step dependencies. We introduce DeFA, a dependency-guided framework for agent failure attribution. DeFA first combines protocol relations and semantic dependencies into an event dependency graph spanning the trajectory. It then identifies events that may violate task requirements and traces their sources and subsequent effects to construct a failure propagation graph. Finally, DeFA uses step evidence and the steps’ roles in failure propagation to identify the decisive error, responsible agent, and error category. To support long trajectories, DeFA partitions executions into segments and combines the current segment’s detailed content with summaries of the other segments, giving local diagnosis access to global execution context. Across Who and When and the Who and When Pro text subset, DeFA achieves the highest responsible-agent and exact step accuracy with all evaluated backbones, and the highest failure-mode accuracy among taxonomy-aligned methods on Pro. Further experiments on image and video trajectories demonstrate its applicability to multimodal failure attribution. Ablations support the contributions of segmentation, the event dependency graph, and the failure propagation graph. Using DeFA’s diagnostic feedback for skill evolution in Trace2Skill improves downstream task accuracy by 6-15 percentage points over the native pipeline, showing that the diagnoses can also support agent improvement on subsequent tasks.
[AI-89] Learning to Ask: Information Acquisition for SLM-LLM Collaboration under a budget
链接: https://arxiv.org/abs/2610.01236
作者: Yongjun Kim,Xiaoxiao Li,Jaeho Lee
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Collaboration between a small language model (SLM) and a large language model (LLM) offers an opportunity to combine the efficiency of smaller models with the strong reasoning capabilities of larger ones. Existing approaches primarily frame such collaboration as a computation allocation problem, determining which model should handle each portion of the reasoning process. In black-box API-based settings, however, this paradigm can be inefficient due to coarse-grained delegation or repeated transmission of context across model switches. In this work, we instead formulate SLM-LLM collaboration as an information acquisition problem, under an API budget constraint. The SLM remains the primary reasoner and selectively queries a black-box LLM advisor only when needed, issuing targeted queries rather than delegating the reasoning process itself. To realize this strategy, we develop a three-stage RLVR framework that learns whether to call the advisor, how to formulate useful queries, and how to integrate the collaboration into the reasoning process by jointly refining advisor invocation and information use. Across mathematical reasoning and coding tasks, our approach improves the performance–cost tradeoff over existing collaboration baselines and, in some settings, matches or exceeds oracle problem-level routing. Finally, we show that our strategy can transfer to other advisor model families, without further training.
[AI-90] Judgement in the Age of Jev: From Evaluation Scarcity to Evaluation Abundance
链接: https://arxiv.org/abs/2610.01231
作者: Richard Hill
类目: Computers and Society (cs.CY); Artificial Intelligence (cs.AI)
备注: 19 pages
Abstract:Generative artificial intelligence has reduced the cost of producing plausible symbolic artefacts, leading recent organisation scholarship to identify evaluation and discernment as constraints under conditions of production abundance. This perspective examines a further possibility: that machine evaluation itself becomes inexpensive enough to be deployed routinely and at scale. The investigation is prompted by Jev, TypeSafe AI’s specialised model for typed probabilistic decisions. TypeSafe explicitly invokes William Stanley Jevons to argue that lower-cost machine intelligence can unlock previously uneconomic uses. Treating this as a technological provocation rather than an established empirical result, the article formulates a conditional Jevons hypothesis for machine evaluation: sufficiently large reductions in the total marginal cost of usable machine evaluation may increase its organisational consumption where latent demand is substantial and complementary costs do not dominate. The article integrates rebound economics with research on cheap prediction, production abundance, machine evaluation, decision allocation, authority, reliance and Executive Judgement to examine this possible scarcity transition. It distinguishes prediction, machine evaluation, organisational judgement and authorisation as functional activities whose costs need not fall together. Evaluations can share evidence, criteria and errors; scale mis-specified rubrics; operate on representations from which consequential qualifications have disappeared; and change practical decision rights through thresholds and exception routing. The resulting research problem is when cheap machine evaluation substitutes for human evaluative work, when it redistributes or creates demands for judgement, and how it affects the grounds available at consequential organisational commitment.
[AI-91] HHR: Hierarchical Hash Retrieval for Efficient LLM Generation
链接: https://arxiv.org/abs/2610.01230
作者: Lianjun Liu,Tiantian Zheng,You Huang,Weiqi Yan,Mingte Qiu,Huazhong Liu,Xiaofeng Zhu,Yunshan Zhong
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Efficient long-context inference is essential for large language models (LLMs), yet it poses a severe computational bottleneck. Hash-based retrieval offers an efficient alternative by encoding queries and keys into binary codes and using Hamming distance for key selection. However, this leads to a critical mismatch between Hamming distance and attention relevance. Query-Key logits depend jointly on directional similarity and feature magnitudes, whereas hash binarization discards magnitude information, causing both false-positive retrieval of low-logit keys and false-negative omission of high-logit keys. To address these failures, we propose Hierarchical Hash Retrieval (HHR), a coarse-to-fine framework that progressively improves retrieval accuracy through Geometry-Aware Key Routing (GKR) and Learned Hash Projection (LHP). GKR learns a head-wise orthogonal transformation to redistribute feature magnitudes and derive more discriminative page-level logit bounds, enabling effective pruning of low-logit keys while preserving important candidates. LHP then learns a head-wise projection space that aligns Hamming distance with the true Query-Key relevance ranking for fine-grained retrieval. By combining GKR and LHP, HHR suppresses false positives and recovers false negatives, substantially improving the fidelity of hash-based sparse attention. Extensive experiments across diverse LLMs and benchmarks demonstrate that HHR achieves superior performance over existing methods. For example, on LongBench, HHR improves the average score by 1.10 points and, at a context length of 128K, achieves up to a 3.30x decoding speedup and a 2.83x end-to-end speedup for Llama-3.1-8B-Instruct. The code is publicly available at this https URL.
[AI-92] Have an LLM Write Your Anomaly Detector: Autonomous Discovery of Compact Interpretable Detectors for Time Series
链接: https://arxiv.org/abs/2610.01223
作者: David Berghaus
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:
Abstract:Time-series anomaly detection trades off predictive accuracy, computational efficiency, and interpretability. We use a large language model not as the detector but as the author of one: an autonomous research loop in which the model repeatedly edits a single short NumPy program under a leakage-free objective, keeping the best-scoring detector it finds. The loop discovers two compact detectors, one for univariate and one for multivariate series, that describe short windows by their local spectral features and compare them with the training-region distribution through a covariance-aware distance. On the TSB-AD benchmark these detectors lead the field across metrics, ahead of the strongest classical, deep, and foundation-model baselines including Time-RCD, yet they train no network and use no GPU, and the multivariate detector is faster than every similarly performing baseline. LLM-driven program search is thus a practical route to accurate, efficient, and transparent detectors.
[AI-93] Reputation Strategy and Emotion Effects on Generative AI Cooperation: A Comparison Across Reasoning and Non-Reasoning Models
链接: https://arxiv.org/abs/2610.01222
作者: Celso de Melo,Zishan Feng,James Hale,Kazunori Terada,Giorgio Coricelli,Jonathan Gratch
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:As generative AI (Gen AI) systems take on increasingly autonomous roles in economically and socially consequential interactions, understanding their propensity to cooperate – and the signals that shape this propensity – has become essential. We examine cooperative behavior in frontier Gen AI models using the iterated prisoner’s dilemma, manipulating counterpart reputation (positive, unknown, negative), strategy (extortion vs. generosity), and non-verbal emotional signaling (facial expressions conveying competitive or cooperative appraisals). In a first study with non-reasoning models (Claude 3.5, Gemini 2.0 Flash, GPT-4o), cooperation was systematically shaped by all three factors, paralleling patterns long documented in human behavioral research, though models varied substantially in how heavily each factor was weighted. A second study with reasoning models (Claude 4.6, Gemini 3, GPT-5.2) revealed a more concentrated reliance on strategy and reputation, a near-elimination of the Potemkin effect observed in non-reasoning models (evidenced by near-uniform cooperation in a diagnostic harmony game), and a more conditional role for emotion consistent with a hierarchical cue-weighting strategy rather than a simple loss of social sensitivity. Reasoning models also showed heterogeneous end-game behavior, ranging from sustained cooperation to systematic last-round defection effect, revealing model-specific exploitability profiles with direct practical relevance for deployment in negotiation and other multi-round interactions. Together, these findings characterize Gen AI models as increasingly sophisticated, though heterogeneous, social actors, and underscore the practical value of developing standardized cooperation benchmarks to inform the responsible deployment of Gen AI in interactive, socially consequential settings.
[AI-94] Dependency-Aware Reward Shaping for Agent ic Reinforcement Learning
链接: https://arxiv.org/abs/2610.01207
作者: Ziyi Chen,Yan Zhang,Jianhui Wei,Daoan Zhang,Zuozhu Liu
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:When training large language models with reinforcement learning, terminal rewards provide little guidance about which steps matter. Common methods for assigning step credit overlook that work built on uncorrected mistakes is wasted while independent work remains valid. With only a final success/failure reward, every step in a failed episode has zero total future reward, even when it made progress. We propose Dependency-Aware Reward Shaping (DARS), which represents task progress as predicates linked by prerequisite relations and assigns step-level credit over the dependency graph. An annotator marks which predicates each step verifies, invalidates, or repairs. Verified predicates are discounted according to graph distance from the nearest broken prerequisite, while independent predicates are unaffected. Repairs update these weights based on any errors that remain; invalidated predicates need re-verification to regain credit. A fixed potential converts these annotations into signed per-step rewards. A common reward and annotation interface allows DARS to integrate with a range of reasoning and agentic training methods, such as GiGPO and ARPO/AEPO, without changing their rollout strategies or optimizers. Across five task families and models from 1.5B to 8B, DARS improves success by up to 10 points over GiGPO trained with the same budget and harness (ALFWorld), raises the WebShop task score and Search-R1 QA accuracy, complements AEPO’s entropy-based training on AIME24/25 with a Python interpreter, and exceeds OmniOPD in controlled tool-free reasoning comparisons at 1.7B and 4B. Ablations show that step-level credit, dependency attenuation, and graph topology each contribute. On ALFWorld, a distilled 8B annotator matches the API annotator, enabling DARS to run efficiently without a frontier judge. Code is available at this https URL.
[AI-95] Federated Agent Optimization
链接: https://arxiv.org/abs/2610.01195
作者: Qiang Yang,Zhiqiang Kou,Xueyi Zhang,Dong-Dong Wu,Hanlin Gu,Jing Guo,Yang Liu,Di Jiang,Qian Xu
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Large language model (LLM) agents increasingly operate in private environments and accumulate valuable experience from task execution, tool use, feedback, and local knowledge. Yet such experience is distributed across organizations and cannot be directly shared because of privacy and proprietary constraints. Conventional federated learning is insufficient for this setting, as agent capabilities extend beyond model parameters to memory, tools, rewards, skills, and structured knowledge. In this paper, we formulate \textbfFederated Agent Optimization (FAO), which studies how distributed agents can collaboratively improve through controlled information exchange while keeping raw data, complete trajectories, and private knowledge local. We define FAO as a multi-objective problem balancing agent utility, privacy leakage, and communication cost, and organize its optimization space across policy, memory, tool use, reward, and structured knowledge and skills. We further characterize how private experience can be abstracted, protected, aggregated, and adapted into transferable capabilities, providing a unified view of how agents can benefit from one another without direct experience sharing. Finally, we identify the key challenges of FAO and outline several promising directions for future research toward trustworthy federated agent systems.
[AI-96] Counterfactual Generation via Flow Matching: Coupling-Sensitive End-to-End Rates
链接: https://arxiv.org/abs/2610.01193
作者: Yunrui Guan,Krishnakumar Balasubramanian,Shiva Prasad Kasiviswanathan
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Statistics Theory (math.ST); Machine Learning (stat.ML)
备注:
Abstract:Counterfactual generation seeks to sample outcomes under a hypothetical intervention or decision using observational data collected under the factual assignment mechanism. We develop a flow-matching approach that combines a sample-split, doubly robust training objective with a learned coupling between observed source outcomes and target outcomes drawn from a fitted conditional outcome model. To enable finite-step generation, we leverage a score-corrected stochastic sampler based on a Gaussian-smoothed interpolation. Our main theoretical contribution is a coupling-sensitive KL bound for constant-step Euler discretization: the error is controlled by moments of the source–target displacement under the chosen coupling, rather than by global uniform regularity of the velocity field, and has near-linear dependence on the ambient dimension. We also establish finite-sample non-parametric guarantees for the learned velocity and score fields when both the conditional outcome model and the source-target coupling are estimated from data. These bounds separate approximation, coupling-replacement, nuisance-estimation, generalization, and Monte Carlo errors and, combined with the sampler analysis, yield an end-to-end guarantee for counterfactual generation. Experiments on synthetic and semi-synthetic image benchmarks support the coupling-dependent theory and show that, at finite discretization budgets, the stochastic sampler can outperform the corresponding deterministic ODE sampler.
[AI-97] When Does Exercise-Specific Joint Selection Help? An Audit of Evaluation and Control Design
链接: https://arxiv.org/abs/2610.01188
作者: Haotian Chen,Jingkun Yu,Yuning Zhang,Bowen Ye
类目: Artificial Intelligence (cs.AI)
备注: Exploratory offline audit of subject-disjoint skeleton-based exercise correctness evaluation; 5 pages, 2 figures, 3 tables
Abstract:Exercise-specific joint selection can improve skeleton-based correctness classification, but what does that gain establish? We audit 1,057 repetitions from ten REHAB24-6 subjects, separating evaluation aggregation, subset structure, and temporal representation. The manual-subset kNN gain changes from 0.055 for pooled out-of-fold AUROC to 0.020 for equal-weight within-person AUROC; both paired intervals include zero. Among 1,000 dimension-matched random maps, 14 match or exceed the manual pooled result, versus 145 when bilateral structure and trunk inclusion are also matched. RBF-SVM retains a positive within-person gain, whereas logistic regression and a random-convolution comparator have negative point gains under that estimand. Sequence-order and paired-seed controls further qualify the interpretation. This exploratory audit shows why joint-selection claims require explicit estimands and structurally appropriate controls; it does not establish a new algorithm or clinical benefit.
[AI-98] Detect Explain Interpret: An End-to-End Benchmark for Time Series Anomaly Detection Explainability and Interpretability
链接: https://arxiv.org/abs/2610.01168
作者: Roberto Stanzione,Jules Barbe,Magali Parrino,Jérémie Fourmann,Paul Boniol
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Databases (cs.DB)
备注:
Abstract:Time Series Anomaly Detection has received increasing attention, driven by the growing availability of complex time series data. This surge has led to the development of numerous detection methods, as well as a variety of benchmarks aimed at thoroughly evaluating their performance. However, most existing detectors remain largely agnostic to domain context, overlooking explainability and interpretability. One of the main reasons for this gap is that current benchmarks primarily focus on detection accuracy, and only few of them evaluate spatial explainability. Moreover, no benchmark currently provides sufficiently rich semantic annotations to support the generation of human-understandable interpretations of anomalies. To address these limitations, we introduce SHAD (Scality High-dimensional Anomaly Detection benchmark), a fully annotated benchmark composed of 215 multivariate, high-dimensional time series collected from real-world distributed cloud storage systems operated by Scality. The proposed dataset includes rich contextual information, covering three families of anomalies with varying degrees of severity. As further contribution, we provide a foundation for future work by evaluating baseline methods for Detection, Explainability, and Interpretability, covering all stages of a TSAD pipeline. For Detection, we benchmark a wide range of existing anomaly detectors, testing their effectiveness on the proposed real-world dataset. Then, we consider explainability by evaluating whether measuring the contribution of each dimension in the generated anomaly score can provide accurate anomaly attributions. Finally, for interpretability, we investigate the effectiveness of frozen LLM baselines in localizing and interpreting anomalies.
[AI-99] Looping Beyond Twice: A Scalable Recipe for Looped Mixture-of-Experts
链接: https://arxiv.org/abs/2610.01153
作者: Di He,Pengxiang Li,Da Chang,Qingyan Meng,Lu Yin,Shiwei Liu
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:
Abstract:Looped Transformers introduce recurrent depth as a new scaling axis for LLMs: by repeatedly applying shared Transformer blocks, they increase effective depth without increasing parameter count. However, the benefits of looping remain unclear for large MoE LLMs under FLOPs-matched comparisons. The main reason is that the gains from additional iterations diminish quickly and can even turn into degradation, so the extra FLOPs spent on looping yield little substantial improvement. Consequently, prior work typically settles on two loops. We identify two main obstacles to scaling looped MoE. First, looping inherits and amplifies the curse of depth: hidden-state variance grows with each iteration as residual updates accumulate, which destabilizes deep recurrence and causes representations to drift. Second, looped MoE suffers from expert selection collapse: routers repeatedly select the same experts across loops, so extra iterations add computation without adding computational diversity. Guided by this diagnosis, we propose LOOM, built on a single principle: each loop should contribute new computation while keeping the recurrent state stable. LOOM stabilizes recurrence by scaling residual updates to bound variance growth and re-injecting the input embedding at every loop, and diversifies it through per-loop routers that engage different experts and a Looping Residual that carries earlier outputs forward. Experiments across 100M-1.7B models show stable scaling to 9-12 loops. Under near-iso-FLOP, the 700M model performs best at 5 loops, reducing perplexity from 18.36 to 16.54 and improving average zero-shot accuracy from 38.84% to 39.53% over the non-looped baseline. Without FLOP matching, the 1.7B model trained on 60B tokens peaks at 9 loops, reducing perplexity from 9.62 to 7.77 and improving average zero-shot accuracy from 42.4% to 47.7%. Code is available this https URL.
[AI-100] When Is Deletion Ordering Tractable? From Update Dynamics to Permutation Structure
链接: https://arxiv.org/abs/2610.01149
作者: Xinyu Wang,Ziyu Zhao,Yixuan He,Xiaowen Chang Alex Smola
类目: Data Structures and Algorithms (cs.DS); Artificial Intelligence (cs.AI)
备注:
Abstract:Given a fixed set of pending deletion requests, retraining from scratch after each request is prohibitive, so a prescribed request-wise policy processes them sequentially. The resulting terminal model can depend on their order. Rather than prescribing an ordering rule, we study the permutation objective induced by the fixed policy and ask when it admits simpler structure. We identify two independent reductions: position additivity represents the objective by request–position costs, reducing optimization to assignment and, with a shared positional profile, sorting; suffix localization removes dependence on the distant prefix while retaining interactions among the surviving requests. Under shared affine updates, we characterize the quadratic interactions that obstruct additivity, prove the reductions’ independence, and show that suffix-conditioned assignment improves the approximation rate from O(p^L) toO(p^(2L)). Experiments recover both structures in executed objectives. A controlled damped-Newton sweep shows that stronger contraction shifts the objective toward shorter, more suffix-specific dependence, while two full-network policies exhibit distinct positional and within-suffix structure. Structures identified from compact execution sets also predict unseen orders. These results frame deletion ordering as identifying the computational structure induced by the executed updates. Subjects: Data Structures and Algorithms (cs.DS); Artificial Intelligence (cs.AI) Cite as: arXiv:2610.01149 [cs.DS] (or arXiv:2610.01149v1 [cs.DS] for this version) https://doi.org/10.48550/arXiv.2610.01149 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[AI-101] Parameter-Efficient Distributionally Robust Adaptation of Tabular Foundation Models under Subpopulation Shift
链接: https://arxiv.org/abs/2610.01143
作者: Seonghwi Kim,Sung Ho Jo,Minwoo Chae
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: 45 pages, 7 figures, including appendices
Abstract:Despite strong mean accuracy, tabular foundation models (TFMs) can perform poorly on underrepresented groups under subpopulation shift, where group proportions change between training and deployment. We propose DR-TFM, a parameter-efficient distributionally robust adaptation framework that requires no true group annotations. DR-TFM adjusts attention to labeled context examples by fine-tuning an existing query scaling network or adding and training one, while keeping all other parameters fixed. We instantiate the framework with two robust objectives using estimated groups or source conditional distributions derived from training data. For TabPFN-3, adaptation updates only 0.016% of the pretrained model’s parameters. Across five tabular benchmarks, DR-TFM achieves substantially higher average worst-group accuracy than pretrained TFMs and the compared robust baselines without true group annotations, while maintaining competitive mean group accuracy. DR-TFM also improves average worst-group accuracy on ACS Income and across four additional TFMs.
[AI-102] ReSolve: Reusing Candidate Reasoning through Selective Generative Moderation
链接: https://arxiv.org/abs/2610.01140
作者: Bangji Yang,Jiajun Fan,Hongba Ma,Xi Zhu,Weizhi Zhang,Minghao Guo,Ye Li,Hamid Palangi,Jiaxuan You
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Sampling multiple solutions spends computation on intermediate deductions and unfinished arguments as well as final answers. We introduce ReSolve, a training-free inference procedure that reuses this candidate reasoning through selective generative moderation. An answer-distribution controller invokes a model to examine existing derivations when candidates disagree or lack a parseable answer, then incorporates the generated solution into a bounded loop. Under Hybrid scoring on 130 competition-mathematics problems evaluated with two independently sampled candidate pools, ReSolve obtains 100 and 99 correct answers, compared with 91 and 92 for voting over the same four candidates, with no correct-to-incorrect changes relative to that vote in either pool. Eight-sample self-consistency obtains 94 and 96 correct answers while consuming substantially more tokens; ReSolve uses 46.3% and 47.2% fewer tokens in the two evaluations. A controlled ablation removes visible derivations while retaining answer keys, vote counts, and the per-state output-cap rule, reducing accuracy from 100 to 93 correct despite increasing computation. Selective and always-on Uniform moderation both solve 97 problems, while selectivity reduces moderation tokens by approximately 54% and total pipeline tokens by 6.2%. These results support candidate reasoning as reusable inference computation. They do not establish an accuracy advantage over additional sampling or a distinct benefit from specialized route instructions.
[AI-103] Auditing Action Settlement in LLM Agent Environments: Order Progress and Replay
链接: https://arxiv.org/abs/2610.01138
作者: Haotian Chen,Bowen Ye,Yuning Zhang,Jingkun Yu
类目: Artificial Intelligence (cs.AI)
备注: 5 pages, 2 figures
Abstract:Concurrent actions in large language model (LLM) agent environments require arbitration even when each proposal is individually valid. We implement a typed snapshot-settlement contract and audit three distinct properties: order sensitivity, useful progress, and replay consistency. Five settlement policies are tested in 28,800 exhaustive permutation trials and 2,160 scripted multistep episodes. Joint policies are spatially order-invariant conditional on fixed priorities, yet conservative rejection completes only 31.25% of agents in a six-agent doorway task versus 90.28% for random tickets; the paired improvement is 59.03 percentage points (95% bootstrap interval: 50.00-68.06). All policies preserve the tested spatial constraints, and priority arbitration still misses the independent small-instance optimum. A separate full-state journal audit exactly replays 156 checkpoints and rejects 1,332 constructed corruptions with a retained terminal anchor. The evidence concerns execution semantics, not human realism or long-run fairness.
[AI-104] Does Scaling Reinforcement Learning Really Require More Training?
链接: https://arxiv.org/abs/2610.01133
作者: Bangji Yang,Jiajun Fan,Hongba Ma,Ruihan Guo,Ge Liu
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:
Abstract:Scaling reasoning typically spends more compute on reinforcement learning (RL) or on inference. We show that a completed RL training history can yield policies stronger than the checkpoints visited by its optimizer. We call this policy-space scaling: expanding the deployable policy set accessible from a fixed RL history, without extending training or increasing per-query inference computation. We instantiate it with SURGE (Scaling Up RL Gradient-free via Eigenspace fusion). SURGE combines two checkpoints from the same RL run: a high-accuracy anchor and a competitive donor that generates shorter responses. It expresses both checkpoints as changes from their shared initialization, then spectrally decomposes the anchor’s update to retain its dominant component and incorporate the donor’s complementary component. With a fixed target for how much of the anchor update to retain, SURGE determines the block size from the weights without testing candidate policies. We evaluate two 1.5B mathematical-reasoning histories, DeepSeek and Nemotron, and one 7B coding history, OLMo. SURGE improves benchmark-average accuracy over both input checkpoints while using fewer reasoning tokens than the anchor. It reaches 54.17% on DeepSeek AIME24 against a measured native maximum of 50.83%, and 83.7% on OLMo HumanEval+ against 82.8%. These gains exceed the observed training curves. Geometric controls support the importance of RL-update structure beyond weight displacement or token reduction alone. Each constructed model runs as a single policy. Our findings identify stored RL history as a reusable scaling resource: the capability available from a training run need not end at its best checkpoint.
[AI-105] Grounding Large Language Models in DSGE Simulators for Policy Generation and Forecasting
链接: https://arxiv.org/abs/2610.01128
作者: Aditya Dubey,Namah Gupta,Vinti Agarwal
类目: Artificial Intelligence (cs.AI); Computational Engineering, Finance, and Science (cs.CE); Machine Learning (cs.LG)
备注:
Abstract:Large language models can produce economic policy responses that sound reasonable, but this does not show that their actions are consistent with economic dynamics. We test this by placing an instruction-tuned language model inside six Snowdrop-backed dynamic stochastic general equilibrium (DSGE) simulators. At each turn, the model observes the economy and a change in economic discourse, selects a bounded policy action, and receives the next simulated state and an economic reward. We implement a common Python interface for repeated rollouts, persistent shocks, state cloning, and rolling-horizon simulation. This setting creates a long-horizon credit-assignment problem. Policy effects may appear several quarters after an action is taken. PPO has a learned value function that can propagate delayed reward to earlier tokens through generalized advantage estimation. GRPO has no learned value function and instead assigns a group-relative advantage from complete rollout returns. It therefore cannot distinguish which earlier turn caused the outcome; if every rollout receives the same return, the normalized advantage is zero. We use PPO as the primary method and GRPO as a matched critic-free baseline. The experiments also test directional semantic signals, reward horizon, trajectory warm starts, cross-simulator transfer, and historically anchored pandemic and monetary-policy shocks. The objective is to judge policy actions by their simulated economic consequences rather than by plausible language alone. Subjects: Artificial Intelligence (cs.AI); Computational Engineering, Finance, and Science (cs.CE); Machine Learning (cs.LG) Cite as: arXiv:2610.01128 [cs.AI] (or arXiv:2610.01128v1 [cs.AI] for this version) https://doi.org/10.48550/arXiv.2610.01128 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[AI-106] CortexBridge: Cortical Alignment of EEG Montages for Foundation Models
链接: https://arxiv.org/abs/2610.01124
作者: Jiazhen Hong,Xiaotian Zhou,Zihao Ding,Kailong Wang,Yu Wu
类目: Artificial Intelligence (cs.AI); Signal Processing (eess.SP)
备注:
Abstract:Electroencephalography (EEG) foundation models are often pretrained with a fixed channel vocabulary or a limited set of montages, making transfer difficult when electrode layouts change. We propose CortexBridge, a lightweight adapter that combines EEG features with electrode and atlas coordinates to map arbitrary montages into a shared cortical latent space. Evaluated with three frozen foundation models on five brain-computer interface (BCI) datasets from the Mother of All BCI Benchmarks (MOABB), CortexBridge improves performance in 13 of 15 evaluations. The gains in balanced accuracy average 0.80% for EEGPT, 0.70% for LaBraM, and 3.26% for CBraMod, with a maximum gain of 13.02% on 12-class steady-state visual evoked potential (SSVEP) classification. Visualizations of the learned atlas representations reveal task-dependent spatial patterns, with SSVEP showing a more concentrated representation in the Yeo Visual network than auditory P300. These results establish cortical alignment as a learnable and anatomically grounded routing mechanism from heterogeneous EEG montages to pretrained foundation models.
[AI-107] AbsorbEvo: An Agent ic Framework for Autonomous Inverse Design of Microwave Absorbers
链接: https://arxiv.org/abs/2610.01119
作者: Zhicheng Feng,Yubo Zhao,Xuefeng Yao
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Designing high-performance microwave absorbers requires specialized expertise in electromagnetic theory, materials science and simulation programming, and entails time-consuming optimization. Here, we present AbsorbEvo, an agentic framework for autonomous inverse design that translates natural-language performance objectives into designs verified by full-wave simulations. Its candidate evolution strategy integrates language reasoning, physics-based prediction and historical feedback. A large language model proposes the directions and magnitudes of parameter adjustments based on task objectives and computational history. The system combines directed increments with global sampling to generate candidates and uses a low-cost predictive model as a physics prior to rank them. Only high-ranking designs undergo full-wave simulation. Results passing physical validity checks are used to evaluate performance and guide subsequent search. Experience from training tasks is further distilled into textual skills, which are independently validated before use in new tasks. Under identical proposal budgets on held-out AbsorbBench-36 tasks, AbsorbEvo achieved a task success rate of 79.17%, versus 25.00% for a generic agent and 12.50% for random search. Its mean best coverage was 0.7816, compared with 0.6434 and 0.6448, respectively. By integrating language reasoning and physics-based feedback into design decisions, AbsorbEvo provides a methodological foundation for natural-language-driven autonomous inverse design of microwave absorbers.
[AI-108] Beyond State-of-the-Art: Standardising Environmental Impact Metrics for AI Research
链接: https://arxiv.org/abs/2610.01116
作者: Lachlan McGinness,Dan Pagendam,Robert Offner
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:As the capabilities and ubiquity of Large Language Models (LLMs) grow, so does their environmental footprint. Despite calls for responsible AI, the machine learning community lacks standardised practices for carbon accounting. Our automated literature review of the 5,285 papers accepted to NeurIPS 2025 reveals that reporting of environmental impact is nearly non-existent. To catalyse a shift toward sustainable AI, we define standardised sustainability metrics for evaluating model training efficiency, accompanied by simple heuristics to estimate the carbon cost of LLM inference. We implement these metrics in carbonbenchmark, a drop-in software solution for tracking and reporting emissions. Finally, to combat the pursuit of marginal accuracy gains at disproportionate environmental costs, we formalise the Smallest Model that Achieves the Job' (SMAJ), a framework which challenges the field to prioritise computational efficiency and environmental accountability alongside traditional State-of-the-Art’ (SotA) accuracy. Subjects: Artificial Intelligence (cs.AI) Cite as: arXiv:2610.01116 [cs.AI] (or arXiv:2610.01116v1 [cs.AI] for this version) https://doi.org/10.48550/arXiv.2610.01116 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[AI-109] YouRA: A Persistent-State Architecture for Evidence-Traceable Autonomous Research Agents AACL
链接: https://arxiv.org/abs/2610.01097
作者: Yoonkyu Woo,Woojin Lee,Jin-Xia Huang
类目: Artificial Intelligence (cs.AI)
备注: Accepted to the AACL-IJCNLP 2026 Main Conference. 21 pages, 5 figures, 17 tables
Abstract:End-to-end research agents can now produce complete scientific papers, yet manuscript claims often diverge from executed experiments. This gap is structural: research state, failure histories, and claim-evidence alignment are not maintained as persistent, verifiable state across long-horizon pipelines. We present YouRA (Your Research Agent), an architecture for stateful, evidence-traceable autonomous research. YouRA preserves research state, execution evidence, and failure history across the research trajectory by integrating three components: a Verification State Architecture (VSA) that tracks hypotheses, gates, and evidence pointers; an Independent Controller that turns state and reflection records into lifecycle, recovery, and debate/review control while separating control from execution; and Stateful Reflection that logs failures as structured lessons and routes recovery through bounded repair, redesign, or reset. On MLR-Bench’s predefined ten-task end-to-end subset, YouRA improves over both MLR-Agent and AI Scientist V2 on scalar Overall across all three matched backbones. An automated diagnostic using MLR-Bench’s hallucination taxonomy reports intersection/union counts for four fact-based failure types, and data-provenance diagnostic shows more real-data-based outputs. Ablating each of the four components (the VSA, the Independent Controller, MCP tool access, and reflection-guided recovery) supports their separable contributions. Removing either core-state component drops YouRA below the full system. Code: this https URL.
[AI-110] OrbitTAMP: Grounding Language Models for Task and Motion Planning in Spacecraft Rendezvous
链接: https://arxiv.org/abs/2610.01093
作者: Yuji Takubo,Daniele Gammelli,Marco Pavone,Simone D’Amico
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI); Optimization and Control (math.OC)
备注: 20 pages, 8 figures
Abstract:Spacecraft rendezvous and proximity operations (RPO) are currently planned through an expertise-intensive process in which engineers translate high-level operational intent into safe, dynamically feasible trajectories, creating a bottleneck to scalable operations. Large language model (LLM)-based agents could offer an intuitive interface for this process, although their outputs are not inherently grounded in orbital dynamics, operational constraints, or the structure of admissible spacecraft maneuvers. To exploit their semantic reasoning while ensuring the generated plan’s physical validity, this paper presents a hierarchical framework for spacecraft task-and-motion planning (TAMP) that grounds LLM reasoning in a graph of reusable behaviors and domain-specific planning modules. Within this framework, a pretrained LLM maps a natural-language command to a partial mission specification. The associated planners then resolve unspecified decisions within the admissible operational space. Finally, trajectory optimization converts the completed mission specification into a dynamically feasible trajectory. Numerical experiments demonstrate that this architecture substantially improves intent recovery over direct LLM generation, achieving 98% exact recovery of partial mission specifications across all evaluated splits when backed by frontier LLMs. Additional test-time-compute experiments show that, for a compact 9B model, verifier-guided revision increases exact recovery from 75% to 88%, while broader behavior-plan search independently improves selection among admissible trajectory realizations. Overall, these results establish a scalable and auditable foundation for language-driven agentic planning of spacecraft RPO.
[AI-111] Improving Math Reasoning through Value-guided Informative Search
链接: https://arxiv.org/abs/2610.01080
作者: Shaohuai Liu,Yuning Wu,Haoran Liu,Enzo Jia,Devin Chen,Kai Wei
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Reinforcement learning with verifiable rewards (RLVR) has substantially improved the mathematical reasoning capabilities of large language models. Recent work introduces search into RLVR rollouts to increase trajectory diversity, but diversity alone does not ensure that the search-induced rollout policy improves upon the current policy. To address this gap, we propose APIVIS, a training-time framework that adapts finite-budget Gumbel search to chunk-level mathematical reasoning. APIVIS combines direct and searched responses within each rollout group, allowing improvements found by search to produce informative relative rewards. It further applies selective supervision to search-improved tokens, preserving a learning signal when uniform group rewards render GRPO ineffective. We show that exact value-guided selection improves the expected verifier reward at each searched state and that this guarantee extends to the complete rollout policy, with a corresponding approximate guarantee under bounded value-estimation error. Experiments on widely recognized mathematical reasoning benchmarks and different model scales demonstrate substantial improvements over competitive search-based methods, validating the effectiveness of APIVIS.
[AI-112] Jev-IDS: System One Models for Network Intrusion Detection
链接: https://arxiv.org/abs/2610.01079
作者: Paulo Severo,Silvio E. Quincozes,Amanda Dias
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
备注:
Abstract:Machine-learning Network Intrusion Detection Systems (IDS) depend on substantial labeled datasets and task-specific training, whereas Large Language Models (LLMs) detection can analyze flow records directly but incurs higher inference cost and latency, with less constrained outputs. This paper presents JEV-IDS, an open experimental general NIDS based on the Jev System One Model (SOM) to detect zero day intrusions Under label scarcity. JEV-IDS serializes one flow per request and asks JEV two questions: a binary attack probability and a finite-choice traffic category. Our results show that, at k=1, JEV was 4.8 times faster and 3.8 times cheaper than GPT-5.6 Luna, with 1.5 times higher novel-attack recall; it also produced 15 times fewer false alarms than a low-data Random Forest. Across 5,400 decisions on a 300-flow NSL-KDD pilot split, JEV achieved F1-Score 0.859, precision 0.941, recall 0.790, and novel-attack recall 0.838. Increasing k to 2 reduced its F1-Score to 0.839.
[AI-113] MOMAT: Mixture of Multiple Atlases for Low-Power Jailbreak Defense of Quantized LLM s
链接: https://arxiv.org/abs/2610.01058
作者: Boyang Li,Bingyu Shen,Weihao Hong,Zhiyuan Jiang,Xinlei Guan,Yan Ma,Miles Q. Li,Yi Sheng,Ruiyang Qin
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
备注: 16 pages, 13 figures
Abstract:Quantized large language models are increasingly deployed on edge devices for their low latency and energy efficiency. However, model quantization weakens alignment safeguards, leaving qLLMs (quantized large language models) highly vulnerable to jailbreak attacks. To address this challenge, we present MOMAT (Mixture of Multiple Atlases), a hardware-enhanced safety framework that combines structured knowledge retrieval with low-power defense acceleration. Each atlas represents a semantic cluster of harmful or benign sample sets and policy templates, enabling domain-localized Retrieval-Augmented Generation guarding that mitigates the curse of dimensionality and the resulting semantic sparsity problem in large, heterogeneous safety databases. MOMAT retrieves top- k similarity features from all atlases for each prompt and evaluates them using a lightweight MoE (Mixture of Experts) detector, while a CiM (Compute-in-Memory)-accelerated similarity engine performs fast, low-power atlas-local retrieval. MOMAT’s CiM-based retrieval accelerates a 100-query batch from 15,052.44 ms to 3,207.21 ns (a 4.69 \times 10^6\times speedup) and reduces energy from 8.1 \times 10^7 \mu J to 3.32 \mu J, yielding an approximately 2.5 \times 10^5\times energy reduction over DRAM-based (Raspberry Pi) baselines. Red-team evaluations across standard benchmarks show that MOMAT matches the defense performance of state-of-the-art methods while avoiding benign overkill and providing substantial efficiency gains, demonstrating that CiM-based modular defenses can make edge-deployed qLLMs both safer and more energy-efficient. We will release the full 223.2k-sample dataset to foster future research.
[AI-114] Network World Models as Environments for Algorithm Design on Complex Systems
链接: https://arxiv.org/abs/2610.01048
作者: Rishab Alagharu,Hongji Pu,Zeeshan Memon,Xinyuan Song,Yuntong Hu,Liang Zhao
类目: Artificial Intelligence (cs.AI)
备注: 46 pages, 7 figures, 17 tables. Preprint
Abstract:World models, which simulate an environment and predict how it changes under actions, are increasingly used in real-world applications such as robotics. Complex systems call for the same tool because the effect of an action is not immediate. Seeding nodes for a campaign, or immunizing nodes against an epidemic, changes little on its own; what matters is the outcome that unfolds over the steps that follow. Designing an algorithm that selects such actions to maximize expected performance on a task is inherently iterative, and every candidate must be scored by the outcome it produces. Obtaining that outcome has relied on simulation, whose cost becomes a bottleneck when candidates are evaluated over many sampled trajectories. We propose an action-conditioned Network World Model that learns a network’s diffusion dynamics under interventions over time, applies each action to the network, and predicts the outcome that follows. It serves as a fast evaluator inside an algorithm design loop in which a coding agent designs and refines executable algorithms using feedback from full rollouts, action-level credit, and counterfactual probes over alternative interventions. Across eight network tasks and five diffusion models, the designed algorithms match or exceed the strongest reported baseline in 138 of 141 settings while enabling up to 14.5 times faster rollouts than Monte Carlo simulation. Code will be released upon acceptance.
[AI-115] Optimal Transport Reweighting for Robust Learning under Spurious Correlations and Label Noise NEURIPS2026
链接: https://arxiv.org/abs/2610.01028
作者: Sung Ho Jo,Seonghwi Kim,Wonsang Yun,Minwoo Chae
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: Accepted at NeurIPS 2026
Abstract:Machine learning models often suffer performance degradation under subpopulation shift, particularly when spurious correlations cause models to rely on shortcut features that fail to generalize across subgroups. A recent line of work mitigates this issue by using loss-based signals to identify informative samples, but these signals can become severely distorted under label noise: mislabeled samples may also incur large losses and contaminate subsequent reweighting or retraining. Despite its practical importance, this intersection remains largely underexplored. We propose POTER, a reweighting framework based on optimal transport that derives sample importance from the transport geometry between the training distribution and a reference distribution constructed from limited validation group annotations. By measuring alignment at the individual-sample level rather than relying on loss, POTER downweights mislabeled or strongly bias-aligned samples while assigning higher importance to samples better aligned with the reference distribution. In addition, POTER requires only a single ERM training stage, moving beyond the retraining paradigm common in recent work. Across standard benchmarks and noisy-label settings, POTER achieves state-of-the-art worst-group accuracy, including cases where label corruption is concentrated within minority subgroups.
[AI-116] From Discovery to Decision: Finite-Budget Recoverability in LLM Voting
链接: https://arxiv.org/abs/2610.01014
作者: Shaoang Li,Jian Li
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Voting over multiple LLM responses is a common primitive in test-time scaling and ensemble inference. Collecting more responses can expand the candidate pool and increase the chance that a correct answer is discovered. Under a fixed call budget, a discovered answer still needs to accumulate enough support within the remaining calls to become the final plurality winner, creating a discovery-to-decision gap. In this work, we characterize this gap through the realized vote state and remaining call budget. We derive a sharp recoverability threshold and show that, as sampling proceeds, the observed candidate set can only expand while the set of reachable endpoint winners can only contract, inducing a candidate-level conversion window. Under a specified iid response law, the same state yields exact finite-horizon endpoint probabilities. We further show that merging wrong-answer identities preserves single-call correctness and cannot improve plurality accuracy, and that the effect of redistributing wrong-answer probability depends on the realized vote state. Singleton reachability yields a gold-free exact locking certificate. For a known answer universe, its first trigger is the earliest prefix at which all admissible continuations yield the same fixed-budget output. Empirically, most discovered-but-unselected correct answers lose reachability only after discovery. In a controlled Word16 study, input permutation improves raw-plurality accuracy by 21.1 points with essentially unchanged single-call correctness. Exact locking saves 28-30% of calls at a 16-call budget while preserving every fixed-budget output.
[AI-117] Beyond Answer Confidence: A Controlled Audit of Self-Knowledge in a Black-Box Decision Model
链接: https://arxiv.org/abs/2610.01006
作者: Sharath M Shankaranarayana,Davor Runje,Jan Jannink
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:
Abstract:Decision models return probabilities intended for routing, abstention and automated action. Calibration makes those probabilities useful on average, but does not establish whether low confidence reflects chance or missing knowledge, nor whether confidence falls when a model moves beyond what it knows. We audit this distinction in Jev, a decision model, with over 15 public datasets and 6 generated task families, with paired interventions that vary the information supplied for a fixed item. Jev’s confidence is calibrated on familiar closed-choice tasks but fails as an indicator of missing knowledge: with no answer-relevant information it assigns up to 0.80 to a salient option, and on news beyond an observed knowledge boundary it exceeds accuracy by 0.21–0.33, a gap that recalibration on earlier months does not close. Targeted yes/no questions give sharper readouts of the case: whether an outcome is settled (AUROC 1.00) and whether the evidence suffices (0.95, against 0.85 for confidence on the same items). Asking whether Jev knows the answer appears to flag fabricated entities and post-boundary news (0.91), but with realistic names or with dates removed it shows no advantage over answer uncertainty. Black-box knowledge audits therefore need explicit controls for surface cues. Code: this https URL.
[AI-118] What Can Analogy Tell Us About Artificial Consciousness?
链接: https://arxiv.org/abs/2610.01002
作者: Keith J. Holyoak,Martin M. Monti
类目: Artificial Intelligence (cs.AI)
备注: 11 pages, 1 figure, 2 boxes
Abstract:Who or what is conscious? Because subjective experience is directly accessible only in the first person, judgments about consciousness in other entities depend partly on analogy. Historically, such inferences have focused on nonhuman animals, but advances in artificial intelligence have raised the possibility of conscious AI. Here we develop a causal framework for evaluating such evidential analogies. The key distinction is between similarities in factors plausibly involved in generating consciousness and similarities in downstream behavioural or cognitive effects. Our framework weights source-target similarity by causal relevance while allowing for unknown causes, disabling differences and alternative routes to consciousness. Applied to biological systems, it explains why analogical support generally weakens with increasing causal distance from humans. Applied to contemporary AI, it suggests that behavioural similarity provides only limited evidence for consciousness because relevant causal correspondences remain poorly established. The framework also clarifies what evidence would strengthen claims of artificial consciousness.
[AI-119] Calibration-risk routing for controlled world-model adaptation
链接: https://arxiv.org/abs/2610.01001
作者: Yifan Zhang,Liang Zheng
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Model-based reinforcement learning (MBRL) can exploit simulated experience, but a simulator-to-target shift creates a model-selection problem: correcting the simulator and fitting the target directly can each fail under limited target data. We introduce the Model-Corrected World Model (MC-WM), which separates initial target data into disjoint fit, selection, and calibration partitions and deploys the family with lower standardized calibration risk. A learned confidence signal and deterministic validity predicates weight one-step imagined policy updates without rewriting physical rewards. We evaluate 540 unique reported run cells across three controlled Multi-Joint dynamics with Contact (MuJoCo) shifts; one exact-routing cell was repeated after a pre-deployment artifact gate, giving 541 completed executions.
[AI-120] Evaluating LLM -Generated Preference Distributions
链接: https://arxiv.org/abs/2610.01000
作者: Fan Huang,Minsuk Kim,C. Tyler Diggans,Filippo Radicchi
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Large Language Models (LLMs) are increasingly used as probabilistic generators for simulation, synthetic data generation, and decision support in settings where real-world data are unavailable. Yet, the structure and reliability of the distributions they produce remain understudied. Here, we systematically analyze LLM-generated distributions of preferences for air travel, restaurants, and consumer products. Encouragingly, all models considered in our analysis exhibit self-coherence, with the most probable outcomes stabilizing rapidly under repeated sampling. At the same time, we observe substantial discordance across both model families and scales, with little consensus even among their most probable outcomes. These patterns hold across nine open-weight models, three choice domains, and show robustness under temperature changes, greedy decoding, and perturbations of prompt and ordering. Our findings indicate that outcomes are influenced more by the choice of model than by the wording of the prompt, challenging the common assumption that sufficiently capable LLMs produce similar preference distributions when used as stand-ins for survey respondents.
[AI-121] Divide-and-Remember: Recursive Action-Relevant Memory for Long-Horizon VLA Policies
链接: https://arxiv.org/abs/2610.00982
作者: Xuehui Yu,Eason Yu,Meiyi Wang,Haozhe Du,Stefano V. Albrecht,Harold Soh
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI)
备注:
Abstract:Vision-language-action (VLA) models struggle on history-dependent manipulation tasks, where the current observation alone does not determine the action, and the policy needs a memory of the history. Existing memory methods decide what to remember by design, for example, keeping frames with large pixel changes, and show inconsistent gains across tasks. We view what to remember as an optimisation problem. From the POMDP formulation of imitation learning, we show that the optimal memory maximises the conditional mutual information I(a_t; m_t \mid o_t) between the action and the memory given the current observation. Intuitively, this means preserving the action-relevant information in the history that is not already contained in the current observation. Based on our analysis, we propose Divide-and-Remember (DR), a recursive memory method that learns a memory function m_t = M(h_t) and scales to long contexts while staying compute-light. It involves two strategies: (1) the selection over the full history is divided recursively into subproblems of top- K selection over 2K tokens, so that fixed-size, lightweight selectors learned end-to-end support an unbounded history; (2) all recursion blocks share one selector, which captures the selection rule common to every block and keeps the method efficient. On RoboMME, a benchmark of 16 long-horizon manipulation tasks that require remembering when, where, what, and how to act, DR achieves a state-of-the-art average success rate with consistent gains across all four suites under a budget of only 64 tokens; real-robot experiments show the same gain. Code, checkpoints and more results are at this https URL
[AI-122] RISED: RubrIcs for agent ic multi-environment Selection and sElf-Distillation
链接: https://arxiv.org/abs/2610.00979
作者: Jingtan Wang,Sirajul Salekin,Young mok Jung,Javier Movellan,Bryan Kian Hsiang Low,Manjot Bilkhu
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Training a single LLM agent jointly across diverse interactive environments has attracted increasing attention as a route to generalist agents. Existing curriculum and data-selection strategies often allocate training at the environment level or prioritize local reward-based signals, without explicitly considering relationships between current rollouts across environments for prompt-group selection. Meanwhile, as environments are learned at different rates, all-failure and all-success rollout groups can coexist within a batch, leaving those data without group-relative reward signals. Both challenges highlight limitations of relying solely on scalar rewards in multi-environment RL: they provide limited information about cross-environment relationships and no within-group reward contrast when rewards are identical. This motivates richer textual feedback, such as rubrics describing rollout behaviours, to guide learning. Beyond rubrics’ usage as reward, we repurpose rubrics to guide both online data selection and policy supervision. An LLM judge tags each rollout using a predefined rubric vocabulary shared across environments. The resulting profiles guide the selection of data that aligns with the overall behavioural composition of the mixed-environment batch while limiting overlap with already-selected data. Available positive rubrics (describing desired behaviours) provide privileged context for an on-policy self-distillation teacher, supplying additional token-level supervision, while negative rubrics (describing undesired behaviours) guide subsequent rollout generation away from recurring failure modes. Together, these components form RISED. Across model backbones, RISED achieves the highest mean pass rate across environments and ranks first or second in every individual environment. Rubric-based analysis of RISED can further characterize the behavioural changes accompanying these gains.
[AI-123] Generalist Representation Specialist Detection: TS-Router for Time-Series Anomaly Detection
链接: https://arxiv.org/abs/2610.00978
作者: Tian Lan,Yifei Gao,Yimeng Lu,Xuming An,Meng Wang,Yue Pan,Wenjun He,Chenghao Liu,Chen Zhang
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:
Abstract:Time-series anomaly detection (TSAD) is difficult to generalize across datasets because heterogeneous temporal dynamics imply different notions of normality and favor different detection criteria. While time-series foundation models provide transferable representations, coupling them with a fixed anomaly-scoring mechanism can overlook this variation. This motivates a different perspective on foundation-model-based TSAD: using foundation models to coordinate specialized anomaly criteria rather than directly imposing a universal one. Based on this view, we propose \textbfTS-Router, a generalist-representation, specialist-detection framework that estimates the relative competence of heterogeneous anomaly detectors from pretrained temporal representations and selects suitable specialists for each target series. To avoid relying on specialist-performance labels from real tasks, we derive soft competence supervision from specialists’ relative performance on labeled simulated tasks. At deployment, routing requires no target anomaly labels, and only the selected specialists are fitted unsupervisedly on the target series. We bound Top-(k) set-competence regret under representation coverage and conditional competence stability. Across 16 real-world benchmarks and four complementary evaluation metrics, TS-Router achieves the best overall average rank. Controlled ablations with multiple frozen TSFM encoders further support the use of pretrained representations for competence estimation and adaptive specialist selection. The code is available at this https URL.
[AI-124] Structure-agnostic Causal Representation Learning
链接: https://arxiv.org/abs/2610.00968
作者: Arman Behnam,Binghui Wang
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:
Abstract:Causal representation learning aims to discover robust features by exploiting the causal structure underlying data generation. Existing methods require specifying the causal structure a priori, yet different structures demand fundamentally incompatible invariance constraints, and misspecification leads to representations that discard predictive information. We introduce SaCRL, a framework that jointly identifies the causal structure and learns the corresponding invariant representation without prior structural knowledge. Our approach formulates structure selection as a soft optimization over candidate invariances using HSIC-based violation metrics, with adaptive weights that automatically concentrate on the achievable structure. We provide theoretical guarantees for structure identification, including under random-feature approximation, invariance satisfaction, and out-of-distribution generalization. Empirically, SaCRL recovers the true structure on synthetic and semi-synthetic Bayesian-network benchmarks, outperforms fixed-invariance baselines on Colored MNIST, achieves state-of-the-art accuracy on three DomainBed benchmarks (PACS, VLCS, OfficeHome), and degrades gracefully under structural misspecification and limited environment diversity. Code is available at: this https URL.
[AI-125] PG-SFT: Balancing Capability Acquisition and Retention in Offline Agent Fine-Tuning
链接: https://arxiv.org/abs/2610.00949
作者: Ronghua Li,Zi Liang,Zhishan Li,Shinan Liu
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Supervised fine-tuning (SFT) on offline agent trajectories is the standard approach for training specialized tool-using agents, but forcing models to imitate reasoning and actions token by token may harm other capabilities (e.g., general reasoning, tool calling, code generation) of the base model. In this work, we focus on studying \emphhow to better balance the trade-off between acquiring new capabilities and preserving existing ones during agent trace SFT. By comparing several baselines in our setup, standard SFT improves the target benchmark while lowering several non-target benchmark scores; meanwhile, simply constraining distributional drift using KL penalty or limiting the update magnitude did not avoid this regression trend. Motivated by recent token-wise adaptive learning objectives, this work proposes \textbfPrivilege-Guided SFT (PG-SFT) to leverage turn-level information gain of agent trajectories as an indicator to adjust supervision strength. PG-SFT yields a more favorable observed trade-off on the evaluated benchmarks, substantially reducing distributional drift and broad capability degradation at the cost of slight degradation in target-task performance. Our findings suggest that balancing the acquisition–retention trade-off depends not only on whether the model is anchored to its base behavior, but also on where and how strongly supervision should depart from that behavior.
[AI-126] GUI-HARVEST: Self-Improving GUI Agents through Evidence-Driven Harness Evolution
链接: https://arxiv.org/abs/2610.00948
作者: Geyi Yang,Zikun Qu,Xiang Li,Zhiyong Wang,Min Zhang,Shipei Zeng,Zhongxiang Dai
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: Preprint
Abstract:The executable harness surrounding a GUI model determines how observations are assembled, actions are executed, and verification, recovery, and termination are controlled. Compared with harness optimization for non-GUI agents, automatically optimizing this harness poses three coupled challenges: reconciling model intent with observed visual effects, diagnosing failures under variable execution outcomes, and identifying recurrent failure patterns across tasks and translating them into reusable runtime changes. We introduce GUI-HARVEST, an automatic harness optimizer that enables self-improving GUI agents with frozen backbone models. First, to ground diagnosis in observed action effects, it aligns model outputs and executed actions with before-and-after screenshots, tying findings to specific interface transitions. Second, to account for execution variability, it treats repeated runs of the same task as a joint evidence unit, using within-task comparisons to locate outcome-relevant behavioral differences. Third, it consolidates verified findings across tasks into recurring failure patterns, maps them to bounded source-code edits with predictions recorded before evaluation, and checks the predicted behavioral effects alongside task performance through repeated execution. Experiments on OSWorld-Verified show consistent held-out gains across six general-purpose open, GUI-specialized open, and proprietary backbone models; Qwen3-VL-32B-Instruct gains 12.33 points on the full suite. Frozen-harness transfer improves GPT-5 by 13.87 percentage points on WindowsAgentArena at 50 steps without further optimization. With the same backbone and initial harness, GUI-HARVEST outperforms Self-Harness and Meta-Harness, suggesting that GUI-specific diagnosis and validation help harness improvements generalize to unseen tasks. The code is available at this https URL.
[AI-127] Finding the Right Fit: Model-Harness Interactions across Agent Tasks
链接: https://arxiv.org/abs/2610.00917
作者: Yixuan Li,Yiyun Zhou,Yao Long Teng,Fuchao Yang,Yanchen Deng,Zhiyi Lyu,Xuyu Dong,Feng Chen,Bo An
类目: Artificial Intelligence (cs.AI); Software Engineering (cs.SE)
备注: 19 pages, 9 figures, 6 tables. Code: this https URL . Data: this https URL
Abstract:Choosing an agent system means choosing both a language model and the harness through which it acts. We ask whether a strong model, harness, or pairing stays strong when the setting changes. We evaluate 66 configurations: four configurable harnesses (OpenHands, DeepSeek Harness, PI, and openJiuwen) paired with five models on TUA-Bench, ALE-CLI, and Terminal-Bench 4, plus the native Codex-GPT and Claude Code-Claude pairings. Model rankings reverse across harnesses. On Terminal-Bench 4, Claude leads GPT by 7.94 points in OpenHands but trails it by 30.16 points in PI. For four of the five models, the best harness changes from one benchmark to another, yet some pairings hold: openJiuwen gives Kimi its highest score on all three benchmarks, by 5.61 to 11.11 points. A model’s own vendor harness is not reliably its best, and higher cost does not reliably buy a higher score. On Terminal-Bench 4, GPT scores higher under PI than under DSH at less than a quarter of the cost per task. Matched trajectories suggest why fit varies. Models start almost all repairs themselves, so much depends on whether the harness hands failures back in a form the model can use. GPT does best with PI’s lean scaffold, while Kimi, which often issues malformed tool calls, does best in openJiuwen. We argue that the model, the harness, and the task should be evaluated together, and we release the harness adapters, evaluation code, and all 6,204 scored trajectories at this https URL and this https URL.
[AI-128] OR for AI That Does OR: Routing LLM s up the Escalator inside the OSCAR Framework
链接: https://arxiv.org/abs/2610.00912
作者: Jinzhi Bu,Haixin Tang,Huanan Zhang
类目: Artificial Intelligence (cs.AI); Optimization and Control (math.OC)
备注:
Abstract:Large language models can translate business descriptions into optimization models, but executable code may misrepresent constraints or objectives. A solver can then return an optimal solution to the wrong problem. Even when the solution satisfies the intended operating rules, a better plan may exist. For organizations that repeatedly use optimization modeling, an LLM-based framework should produce accurate formulations at low cost and, ideally, run locally. We study how to verify improvements and allocate attempts across LLMs that differ in price and capability. We develop OSCAR (Optimization modeling by Simulator, Coder, And Reviewer), which uses an offline Simulator certified against labeled decision examples to compare candidates and continues searching beyond feasibility. We model the search for the next certified improvement as sequential decisions under unobserved difficulty: which LLMs to call and when to stop. In a simplified known-prior setting, we give conditions under which cost-ordered escalation is optimal. For general menus, we derive a prior-free competitive guarantee. On five benchmark problems, OSCAR achieves 95% to 100% accuracy at the reported settings using two small open-weight LLMs, each deployable locally on a single GPU. Their single-attempt accuracies average 29% and 48%. In five runs per problem, Codex and Claude Code incur average token costs 3.1 and 5.8 times OSCAR’s, respectively. OSCAR supports open-weight models locally or in the cloud, depending on budget and confidentiality requirements. Firms should maintain labeled decision examples of feasible and infeasible decisions to clarify plain-language operating rules. OSCAR follows these labels when an LLM’s interpretation conflicts with them. As LLM capabilities and prices change, OSCAR’s simple operating rules and adjustable settings help firms adapt their model choices and benefit from these advances.
[AI-129] Understanding Issues Causes and Solutions in Open-Source LLM -based Multi-Agent Systems
链接: https://arxiv.org/abs/2610.00905
作者: Asad Ur Rehman,Syed Mohammad Kashif,Ruiyin Li,Peng Liang,Zengyang Li,Arif Ali Khan
类目: oftware Engineering (cs.SE); Artificial Intelligence (cs.AI)
备注: 30 pages, 4 images, 10 tables, Manuscript submitted to a journal (2026)
Abstract:With the advancement of LLM-based multi-agent systems (MAS), an increasing number of opensource projects are adopting multi-agent architectures as the foundation of their core functionality. Although research and practice on MAS have attracted considerable attention, limited studies have explored the challenges faced by practitioners of open-source LLM-based MAS, the causes of these challenges, and potential solutions. To address this gap,we conducted an empirical study to understand the issues that practitioners encounter when developing and using open-source LLM-based MAS, the possible causes of these issues, and potential solutions. We collected 22,848 closed issues from 21 open-source LLM-basedMASand applied a mixed automated and manual filtering approach to reduce the dataset to 944 issues related to LLM-based this http URL then analyzed these issues to understand the frequent issues encountered by practitioners, their underlying causes, and potential solutions. Our study results show that (1) Orchestration Execution Issue is the most common issue faced by practitioners, (2) Workflow Problem, Tool Integration Problem, and Memory Problem are identified as the most frequent causes of the issues, and (3) Optimize Workflow is the predominant solution to the issues. Based on the study results, we derive empirically grounded implications for practitioners and researchers aimed at improving orchestration, tool integration, and memory mechanisms in LLM-based MAS.
[AI-130] Screw Attention: Rigid-Body Algebra Inside a Transformer
链接: https://arxiv.org/abs/2610.00904
作者: Aly Magassouba
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI)
备注: 13 pages, 8 Figures, 2 Tables
Abstract:Learned manipulation policies rediscover from data the spatial relations that rigid-body mechanics supplies in closed form. This costs data, and it leaves the policies fragile to geometric changes in the scene. We present Screw Attention, a transformer layer in which the relation between two bodies is a spatial transform rather than a graph edge. Every token is a body with a pose. Each pair of tokens carries the relative pose and, for robot joints, the joint screw. Messages are transported along this relation into the receiver’s frame, while the attention scores see only frame-invariant quantities. By construction, the messages are equivariant to an independent change of frame at every token, and a single layer can express the velocity recursion of rigid-body mechanics. On simulated manipulation tasks, Screw Attention matches or outperforms controls of the same size, including graph, transformer and flat networks on LIBERO-Spatial. With 16,162 parameters it reaches 97.3% on LIBERO-Spatial from object poses (without images or language), above a flat network with 27x more parameters. Under a change of per-link frame convention its success is unchanged, while every other learned network falls below 3%. Placed on an analytic controller as a gated residual, it raises insertion success by 17.3 points. It is unaffected by pose noise up to 10,mm and by joint offsets within the factory calibration of a Franka arm. These results suggest a criterion: geometry is decisive when the task requires relations between frames that no other part of the system supplies. Code and trained policies will be released.
[AI-131] OAST: Stochastic Robot Action Tokenization for Autoregressive Vision-Language-Action Models
链接: https://arxiv.org/abs/2610.00899
作者: Keisuke Shirai,Tomohiro Motoda,Hanbit Oh,Ryoichi Nakajo,Roman Mykhailyshyn,Ryo Hanai,Shotaro Miwa,Yukiyasu Domae
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: Project page: this https URL
Abstract:Autoregressive Vision-Language-Action models often represent continuous robot actions as discrete token sequences, enabling action prediction with standard next-token objectives. FAST has substantially improved this representation by compactly encoding action containing diverse temporal frequencies into relatively few tokens. However, while such compression reduces the number of action tokens required for autoregressive prediction, it does not necessarily improve the efficiency of policy learning from limited demonstrations. In particular, FAST typically assigns a single deterministic tokenization to each quantized action sequence, although multiple token sequences can represent and decode to the same robot motion. We investigate whether exploiting this representational redundancy can improve policy learning. In this paper, we propose TOkenization of Action sequences with STochastic sampling (TOAST), a stochastic action tokenization method that samples alternative tokenizations of the same quantized action sequence during policy training. This diversifies the discrete supervision while preserving the underlying robot action and requires no additional demonstrations. Experiments on LIBERO show that TOAST consistently improves over its deterministic counterpart, with the improvement increasing as training data decreases, achieving a 6.8 point gain in success rate when only 1/16 of training data is available. Across four real-robot manipulation tasks, TOAST further improves mean success rate by 15.8 points over the deterministic counterpart. These results demonstrate the effectiveness of stochastic action tokenization for autoregressive robot policy learning, particularly when training data are limited.
[AI-132] Cross-Benchmark Transfer from RL on Agent ic Coding Tasks
链接: https://arxiv.org/abs/2610.00890
作者: Sushant Mehta,Logan Ritchie,Edwin Chen
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Software Engineering (cs.SE)
备注: 15 pages, 2 figures, 4 tables
Abstract:Coding agents often fail in the last mile: they build most of a feature but drop a requirement, test only the cases their implementation already handles, break behavior that was supposed to stay intact, or validate against an unchecked assumption. We ask whether reinforcement learning (RL) on expert-built agentic coding tasks closes this gap, and whether what the agent learns transfers beyond the training distribution. We post-train Kimi K2.7 Code, a 1T-parameter (32B active) open-weight mixture-of-experts model, with RL alone on 1,700 tasks: 1,000 repository tasks graded by hidden fail-to-pass tests and by pass-to-pass tests of existing behavior, and 700 terminal tasks graded by expert-written hidden verifiers. The reward is the fraction of target checks passed and drops to zero if any pass-to-pass test fails. One epoch of GSPO on a rank-32 LoRA adapter improves pass@1 on each of the six external benchmarks we evaluated, across three agent harnesses: SWE-Bench Pro (60.1 to 64.8), DeepSWE (31.0 to 43.4), Terminal-Bench 2.1 (67.4 to 82.0), Terminal-Bench 3 (1.4 to 12.1), Terminal-Bench 4 (0.0 to 7.6), and SWE-Marathon (5.0 to 25.0). Pooled over the five independent task sets (Terminal-Bench 4 revises Terminal-Bench 3), the improvement is significant (p 0.001), and it remains significant on the three sets released after the training data was collected (p = 0.004); the model also improves under both harnesses never used in training. Median trajectories on DeepSWE and Terminal-Bench 3 are 24-35% shorter in agent steps. The base model’s failed DeepSWE runs are mostly near-misses, and on the tasks the trained model newly solves, paired trajectories show it avoiding each of the four failure modes above.
[AI-133] Match the Distribution Not the Compute: Post-Training Multi-Token Prediction Heads
链接: https://arxiv.org/abs/2610.00888
作者: Prachi Badarayani,Aidan Jay,Chenghui Zhou,Dayquan Julienne,Yuan Gao,Tianwei Chen,George Zerveas,Ishmam Zabir,Xiren Zhou,Chris Quirk,Xia Song
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:
Abstract:Multi-token prediction (MTP) improves the throughput of autoregressive generation by enabling the language model to draft multiple next tokens per forward pass, while a verification step over draft tokens ensures that token distribution of the backbone is preserved. Every open MTP-family release (MiMo-7B, DeepSeek-V3, Qwen3) trains its heads jointly with the backbone over the full pretraining run of tens of trillions of tokens, thus setting the drafter quality at pretraining time. We ask whether a lightweight post-training pass on target-generated chain-of-thought is enough to reach the same expected throughput speedup on a frozen reasoning model, and study how a serving-time system built on such a checkpoint can be optimized. We present three findings. 1) On a frozen Qwen3-8B with K=3 chained MTP heads, we show that a post-training recipe with plain cross-entropy on \approx!2.5 B tokens reaches or exceeds the expected speedup of jointly trained MiMo-7B on math, coding and knowledge benchmarks. Our post-training recipe utilizes 10^3 - 10^4\times less MTP-training tokens as compared with joint pre-training of MiMO-7B MTP baseline. 2) We propose a chain-aware relaxation of draft token verification rule that allows a bounded drift from backbone language model token distribution. We show that this relaxation lifts expected speedups by +12 to +16% per benchmark while preserving task accuracy. 3) We propose an adaptive controller that dynamically chooses the number of MTP heads to be engaged at inference time and demonstrate recovery of upto 11 – 14% loss in speedup using fixed maximum MTP draft length.
[AI-134] FORALL-LEAN-AGENT for Auditable Reasoning in Formal Mathematics and Software Verification NEURIPS2026
链接: https://arxiv.org/abs/2610.00885
作者: Naing Oo Lwin
类目: oftware Engineering (cs.SE); Artificial Intelligence (cs.AI); Logic in Computer Science (cs.LO)
备注: Accepted to NeurIPS 2026 VeriCodeGen
Abstract:Coding agents increasingly automate Lean proof development, but successful compilation alone does not establish that a candidate proves the intended statement under acceptable assumptions. We present FORALL-LEAN-AGENT, a frontend-agnostic framework for auditable reasoning in formal mathematics and software verification. The framework combines isolated workspaces, Lean tools, and fresh review with statement comparison, axiom audits, and independent proof checking where supported. Verification evidence and reviewer decisions are bound to the same candidate artifact, making acceptance traceable. We evaluate the framework on VeriSoftBench, PutnamBench, and both problems in the Lean Eval softwareverification track. On the 100-task VeriSoftBench subset, integration with FORALLLEAN-AGENT raises benchmark-rule success from 93 to 100 for GPT-5.6 Sol at low effort while reducing cost from 69 to 62. The PutnamBench evaluation accepts all 672 problems at an average of 4.72 each. These results show that agent harness design can improve correctness and efficiency while providing evidence beyond aggregate solve counts.
[AI-135] Rethinking Data Augmentation under Covariate Shift: Invariant-Guided Diffusion and Prototype Reweighting
链接: https://arxiv.org/abs/2610.00873
作者: Hongyu Cao,Xinyuan Wang,Arun Vignesh Malarkkan,Kunpeng Liu,Haifeng Chen,Yanjie Fu
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:
Abstract:In many industrial applications, 1) tabular data is scarce and imbalanced and thus requires synthetic expansion; 2) input distributions drift between training and deployment (covariate shift); 3) validation sets often diverge from unseen test environments; or 4) standard generative models simply mimic outdated source distributions. This learning setting limits the stability of standard augmentation and adaptation pipelines. We generalize the task under such setting as the Augmented and Weighted Learning under Covariate Shift problem (AWL-CS). AWL-CS imposes two critical challenges on existing methods: 1) misleading generative guidance where models optimize for source similarity rather than downstream task relevance, and 2) structural instability of distributional density where reweighting mechanisms overfit to noisy validation signals. To tackle these challenges, we propose IGDPR (Invariant-Guided Diffusion with Prototype Reweighting), a unified framework that synergizes stable synthesis and structural adaptation: i) To achieve task-relevant generation, we steer the diffusion sampling process using invariant potentials to ensure synthetic samples align with stable decision boundaries rather than outdated correlations. ii) To ensure stable adaptation, we develop a prototype-based reweighting strategy that assesses sample reliability through structural clusters instead of isolated points, effectively filtering validation noise. Extensive experiments on real data demonstrate our method improves data quality by augmenting the most beneficial data for robust learning.
[AI-136] MemFit: Efficient Long-Term Agent ic Memory
链接: https://arxiv.org/abs/2610.00872
作者: Mitchell Piehl,Muchao Ye
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Long-term memory systems for large language models (LLMs) have gained popularity for extending reasoning capabilities across applications. Current memory systems rely on LLM agents to organize and consolidate memory, resulting in costly, inefficient write operations. To address this limitation, we propose MemFit, a long-term memory system for conversational agents that reduces the cost and latency of memory operations. Unlike existing systems that rely on expensive LLM calls for memory construction or discard surface-level details through compression, MemFit stores each turn verbatim in an append-only store with near-instantaneous, LLM-free insertion, indexing turns with segment summaries rather than replacing them. Additionally, MemFit uses an LLM-free, multi-path retrieval strategy that combines lexical and semantic signals with cross-encoder reranking over caption- augmented episodes in both textual and multimodal settings. Empirical results on three widely used benchmarks, LoCoMo, MemGallery, and LongMemEval-S, show that MemFit achieves state-of-the-art performance while reducing memory construction time and cost several-fold, providing a scalable and efficient solution for persistent agentic memory.
[AI-137] An Educator-Guided LLM Pedagogical Agent for Scaffolded Feedback in Conceptual Database Design
链接: https://arxiv.org/abs/2610.00870
作者: Sara Riazi,Pedram Rooshenas
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:We present an educator-guided LLM pedagogical agent for scaffolded feedback in conceptual database design. Integrated into an entity–relationship diagram (ERD) editor, the system grounds feedback in the student artifact, assignment requirements, educator-authored rubrics, and instructional resources. Its architecture separates hidden, artifact-grounded diagnosis from the workflow that controls the form and disclosure level of student-facing support. We instantiate the architecture as a four-stage workflow progressing from concept checks and guided application to low-detail feedback and localized clarification. Each feedback request creates a stateful episode linked to versioned ERD states. In a deployment spanning three ERD environments and 383 feedback episodes, 71.1% of observed target-level changes fully or partially incorporated the hidden diagnostic target, including many after Stages~1–2. Qualitative analysis showed that staged disclosure sometimes withheld inaccurate details, supported selective uptake, or allowed later recovery, though some errors still shaped revisions. Survey responses from a self-selected sample favored delayed disclosure and student agency but noted indirectness and repetition. Subjects: Artificial Intelligence (cs.AI) Cite as: arXiv:2610.00870 [cs.AI] (or arXiv:2610.00870v1 [cs.AI] for this version) https://doi.org/10.48550/arXiv.2610.00870 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[AI-138] Learning Multiple Timescales for Goal-Conditioned Reinforcement Learning
链接: https://arxiv.org/abs/2610.00849
作者: Pedro Robles Dutenhefner,Dikshant Shehmar,Wagner Meira Jr.,Marlos C. Machado
类目: Artificial Intelligence (cs.AI); Machine Learning (stat.ML)
备注:
Abstract:Existing approaches to offline goal-conditioned reinforcement learning (GCRL) struggle with long-horizon tasks. Discounting shrinks value differences between distant states until they fall below the function approximation error, leaving the agent with no signal for ranking states. Temporal abstraction, which treats k environment steps as a single transition, restores this signal at long range, but no single fixed k suits all state-goal distances: large k preserves value differences across long temporal distances while collapsing distinctions between nearby states, and small k does the reverse. We make this trade-off explicit and introduce Generalized Implicit Temporal Abstraction (GITA), which conditions a single value function on k. GITA trains one policy by aggregating advantage-weighted supervision across multiple k values, so scales assigning larger positive advantages to a state-goal pair contribute more strongly to its update. GITA does not need to choose between local resolution and long-range signal; it retains both without committing to a single k. On OGBench, GITA outperforms a broad range of offline GCRL baselines, raising average success rate across all tasks by 25 percentage points (73% relative improvement) over HIQL. It also improves over the strongest fixed-k method, OTA, by 7 percentage points (14% relative).
[AI-139] SHARPO: Segment-Level Credit Assignment for Agent ic Reinforcement Learning
链接: https://arxiv.org/abs/2610.00838
作者: Xinchen Du,Zhengze Zhou,Wenhui Zhu,Han Yu,Sen Na,Rohit Jain,Alborz Geramifard
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: 13 pages, 3 tables, 2 figures
Abstract:Agentic reinforcement learning (RL) trains a large language model (LLM) to act over long, multi-step interactions. However, a single localized error can cause task failure, while trajectory-level rewards provide limited guidance for assigning credit to individual decisions. To address this limitation, we introduce Segment-level Hindsight Advantage Reweighting for Policy Optimization (SHARPO), a credit-assignment mechanism that refines Group Relative Policy Optimization (GRPO) at the level of environment-facing segments. Inspired by the existing on-policy self-distillation (OPSD) method, SHARPO computes teacher-student log-probability gaps within each segment and uses the resulting signal to compute a bounded multiplier on the GRPO advantage. This multiplier is shared by all tokens within the segment, allowing credit to vary across different segments. With Qwen2.5-7B-Instruct, SHARPO outperforms existing baselines on the ALFWorld and WebShop benchmarks, including GRPO, SDAR, RLSD, and StepOPSD.
[AI-140] On-the-fly Weight Generation: A Hypernetwork Proof of Concept on ARC-1D NEURIPS2026
链接: https://arxiv.org/abs/2610.00820
作者: Fabio J. Fehr,Philip Torr
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: Published (Spotlight) at NeurIPS 2026 Workshop on Neural Network Artifacts as a New Data Modality
Abstract:General-purpose models can adapt to many tasks from context, while specialised models can execute individual functions with less capacity. Yet obtaining such specialists requires task-specific training or adaptation. We ask whether they can instead be generated directly from a few demonstrations. Using ARC-1D as a controlled testbed, we show that individual transformations can be represented by tiny specialist models, and that a hypernetwork can generate their parameters from context. The generated parameters form a structured weight space, while the resulting specialists show partial compositional generalisation and generalisation to transformations not seen during training. In both settings, removing explicit task identifiers improves generalisation beyond the training transformations. Together, these results provide a proof of concept that few-shot task context can be compiled on-the-fly into compact executable model parameters, and that the resulting weight space can support reuse and generalisation beyond known functions.
[AI-141] raining-Aware Target Coverag e for Synthetic Data Selection
链接: https://arxiv.org/abs/2610.00814
作者: Yang Ba,Michelle V. Mancenido,Rong Pan
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:
Abstract:Synthetic data are increasingly used to scale LLM training, yet more synthetic data do not necessarily produce better models. Useful synthetic data must add information relevant to the target task without introducing errors that offset their benefit, and the value of an example can change as the training set grows. We develop a linear theory that characterizes this tradeoff and determines where synthetic data are useful, how much should be added, and the marginal value of adding one example to an existing set. The analysis shows the conditions when input coverage alone is sufficient and when synthetic errors must also be considered. Guided by these results, we introduce \emphTraining-Aware Target Coverage (TATC), a synthetic data selection method for LLM fine-tuning. TATC identifies candidates whose training effects are beneficial to the target task and selects among them to expand coverage of target-relevant directions not already represented by the available data. Experiments on text and image data verify the linear theory. With mathematical reasoning tasks, TATC selects synthetic solutions for fine-tuning Qwen2.5-Math-1.5B-Instruct and outperforms alternative synthetic-data selection methods on GSM8K across selection budgets. In summary, we provide a principled approach to synthetic data selection by quantifying and maximizing its value to the target task.
[AI-142] Benchmarking Generative Models for Weather Data Assimilation on Real Station Observations
链接: https://arxiv.org/abs/2610.00728
作者: Ruizhe Huang,Qidong Yang,Jonathan Giezendanner,Sherrie Wang
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:
Abstract:Weather reanalysis products rely on computationally intensive numerical weather predictions followed by data assimilation that corrects the forecast toward observations. Deep generative models offer a cheaper alternative that shifts much of this cost from inference to offline training. However, existing generative approaches have been evaluated on synthetic observations or under different datasets and evaluation schemes, making it unclear which design choices actually improve real-world data assimilation. We present the first controlled benchmark of generative weather data assimilation on real weather station observations. Using 11,849 NOAA MADIS stations across the contiguous United States and four weather variables, we evaluate methods while holding the dataset, observation operator, and deep learning architecture fixed. The benchmark compares the major design choices, including diffusion versus flow matching, pixel versus latent-space formulations, and multiple inference-time conditioning strategies, against a classical 3D-Var baseline. The benchmark reveals three clear conclusions. First, learned generative priors outperform the Gaussian prior of 3D-Var (35.7% vs. 33.3% RMSE reduction over ERA5) despite using no ERA5 background field at inference. Second, full-gradient guidance consistently outperforms stop-gradient and initial-noise optimization. Third, other choices provide little measurable benefit: diffusion and flow matching perform nearly identically under matched conditions, and latent-space variable mixing does not help. We further evaluate both dense and sparse station settings and find advantages from generative AI and full-gradient guidance more pronounced under sparsity. Together, these results identify which components of generative weather data assimilation improve performance on real station observations and establish a standardized benchmark for future work.
[AI-143] Robust Nash Alignment under Preference Uncertainty NEURIPS
链接: https://arxiv.org/abs/2610.00715
作者: Shihab Ahmed,Debamita Ghosh,David Tang,Yudan Wang,Alvaro Velasquez,Yue Wang
类目: Artificial Intelligence (cs.AI)
备注: 38 pages, accepted at 2026 40th Advances in Neural Information Processing System (NeurIPS)
Abstract:Preference-based alignment methods typically optimize against a single preference model, and can therefore be brittle when pairwise preferences are uncertain: noisy, heterogeneous, or shift after deployment. To address these issues, we propose Robust Nash Alignment, a game-theoretic framework for alignment to uncertain pairwise preferences. Our formulation has a major learner seeking a policy with a large worst-case win rate against both an adversarial competitor and any preference kernel lying in an ambiguity set around a nominal preference. When the ambiguity set captures the uncertainty in preferences, the resulting robust objective of the game directly yields a certified lower bound on worst-case performance. However, we note this problem is computationally challenging to optimize, and to address this, we introduce a four-player primal-dual proxy game involving the leader policy, follower policy, adversarial kernel, and dual variable, and develop a single-loop optimistic mirror descent-ascent algorithm for it. We show that the proxy always lower-bounds the truncated hard-constrained objective, quantify the proxy-to-hard gap, and characterize an exactness condition under which the proxy recovers the robust objective. We then prove an (\mathcalO(1/\sqrtT)) average-iteration convergence for the proxy-game duality gap, which implies a near-optimal robust policy for the original robust objective. Experiments on controlled tabular games and LLM alignment with uncertain preference further validate the convergence theory and show improved performance over nominal baselines.
[AI-144] ReLiveGym: Evaluating Long-Lived Agents over Weeks of Replayed Reality
链接: https://arxiv.org/abs/2610.00710
作者: Xisen Jin,Jingheng Li,Zhenglun Chen,Junyi Du,Xiang Ren
类目: Artificial Intelligence (cs.AI)
备注: 9 pages. Preprint
Abstract:As large language model (LLM) agents become widely adopted, they are increasingly deployed for tasks that require persistent monitoring or recurring actions (e.g., market analysis). These agents are expected to operate unattended for days or weeks, act at the right timing, and adapt to the dynamic environment over time. These challenges are not fully captured in the existing long-horizon agent work, as they often consider a static environment that is not temporally changing. We introduce ReLiveGym, a diagnostic evaluation environment of long-lived tasks in which agents act sparsely over simulated weeks of chronologically replayed real-world news, market, and social-media streams. The tasks span diverse levels of time sensitivity, reasoning intensity, and recurrence. Across eight base language models, we investigate how model choice and harness design affect agent performance on such long-lived tasks. Our results show that how agents determine when to act arises as an important harness-design axis for long-lived tasks; and that the optimal design varies across tasks and sometimes model choices as well. We also evaluate how continuous learning from hindsight feedback affects performance and addresses failure modes observed in these long-lived tasks. These findings indicate model choice, action timing mechanism, and use of feedback as important considerations in the design of long-lived agents. Code: this https URL
[AI-145] R-GroundBench: A Diagnostic Benchmark for R-Group Groundingin Markush Molecular Editing
链接: https://arxiv.org/abs/2610.00700
作者: Xin Wang,Zichuan Ying,Xinna Lin,Junqi Zhang,Hanyi Xiong,Tianyu Gao,Hairong Zhang,Qixiang Hua,Botian Shi,Zhenhailong Wang,Kaicheng Yu
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:
Abstract:Recent advances in AI for scientific discovery enable molecular understandingand design, yet reasoning over incomplete chemical representations this http URL structures, which encode molecular families through variable R-groupplaceholders (\textitR\textsubscript1, \textitR\textsubscript2, \textitX, etc.), are ubiquitous in pharmaceutical patents and requiregrounding across molecular, textual, and chemical this http URL, existing molecule-language benchmarks focus on fully specifiedmolecules, leaving R-group grounding largely this http URL introduce R-GroundBench:, a diagnostic benchmark built from real patent Markushstructures, featuring a Multiple-Choice (VQA) track with controlled difficultyand modality splits, and an open-ended Generation this http URL results reveal a substantial gap between recognition andmolecular this http URL models achieve over 90% accuracy on Easy VQA, performance drops to56–66% on Hard VQA when shortcuts are this http URL-domain VLMs also remain unreliable, achieving only 25.7–46.2% on HardVQA despite domain-specific this http URL, Generation Exact Match remains below 20% for most models and below8% when visual input is this http URL findings reveal that current AI systems lack reliable grounding andexecution for Markush editing, highlighting challenges for AI-drivenscientific discovery.
[AI-146] Backdoor Purification for LoRA-Tuned LLM s via Null-Space Projection NEURIPS2026
链接: https://arxiv.org/abs/2610.00685
作者: Jianwei Li,Jung-Eun Kim
类目: Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR); Machine Learning (cs.LG)
备注: NeurIPS 2026
Abstract:With the rapid adoption of large language models (LLMs) and parameter-efficient fine-tuning (PEFT) methods, the risk of backdoor attacks has become more severe. Existing backdoor purification methods typically rely on at least one of the strong assumptions, such as prior knowledge of triggers, access to clean references, or aggressive retraining, and they often lack comprehensive evaluations. These constraints substantially limit their practical applicability. To overcome these challenges, our work proposes purifying LoRA-tuned LLMs without these assumptions and even without post-hoc retraining of the suspect parameters. Our objective is to significantly reduce the attack success rates (ASR) while preserving both (i) the base model’s general capabilities and (ii) the new downstream skills learned through the adapter. Through a series of ablation studies, we progressively scale our approach from a single layer in a text classification setting to a full-parameter LLM in the generative task. Through careful data curation and feature approximation, we extract high-fidelity backdoor directions and, for each layer or head, construct orthogonal null spaces in both the input and output channels, onto which the LoRA updates are projected. Empirically, our null-space projection method reduces the ASR from nearly 100% to less than 10%, while preserving the base model’s benign performance and the adapter’s learned abilities during downstream task adaptation.
[AI-147] Ontology-Grounded Reason er-Verified Benchmarks for Evaluating LLM Reasoning in Scientific AI NEURIPS2026
链接: https://arxiv.org/abs/2610.00682
作者: Nishtha N. Vaidya,Stephan Grimm,Thomas Hubauer,Thomas A. Runkler
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: 17 pages, 2 figures. Accepted at the AI Data Readiness for Scientific Discovery (AIDaR) Workshop at the 40th Conference on Neural Information Processing Systems (NeurIPS 2026), Paris
Abstract:Large language models (LLMs) increasingly underpin scientific AI applications that reason over structured knowledge, from biomedical question answering to materials informatics. However, their logical reasoning often falls short, producing factual inaccuracies unacceptable in these settings. Reliable evaluation remains challenging: manual dataset construction scales poorly, and LLM-based generation risks embedding the very flaws it aims to measure. High-quality benchmarks must ground both correct and incorrect labelled examples in explicit background knowledge, formally verifiable by a standard reasoner. We propose a pipeline that automatically generates ontology-grounded multiple-choice question (MCQ) benchmarks from any sufficiently axiomatised OWL 2 ontology, with correct answers grounded in the ontology by design. Distractors are generated by perturbing the right-hand-side class expressions of class definition axioms, and their incorrectness is formally verified by an OWL reasoner via entailment checks. We evaluate the pipeline on three ontologies: Pizza (small, academic), PMDco (complex, materials science), and DOID (large, biomedical), generating 112, 2,491, and 15,216 MCQs respectively. Distractors span four semantic categories from class unsatisfiability to weakened subsumptions, enabling diagnostic evaluation of specific reasoning failures. Items meet natural language quality standards: mean LLM judge scores of 4.02, 4.36, and 3.36 out of 5 confirm fluency, and correct-answer-to-distractor similarity above 0.8 shows that wrong options cannot be dismissed on surface form alone. Six LLMs evaluated zero-shot achieve 41.1-76.8% accuracy, well above the 25% random-guessing baseline, indicating the benchmarks are challenging and discriminative. This work is a step towards more reliable benchmarks for assessing logical reasoning in scientific AI.
[AI-148] Learning Transferable Skills using Goal-Conditioned Bisimulation
链接: https://arxiv.org/abs/2610.00676
作者: Mohammad Amin Abbasfar,Farbod Azimmohseni,Mohammad Hossein Rohban
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Robotics (cs.RO)
备注:
Abstract:Unsupervised skill discovery has emerged as a promising approach for leveraging reward-free datasets to pretrain general-purpose policies. However, current skill discovery methods either require access to expert data or exhibit limited generalization, failing to transfer effectively to previously unseen layouts. A key challenge is to learn representations that capture the temporal structure of the environment while remaining robust to variations across layouts. To address this issue, we present an objective for learning action-aware temporal representations that satisfy the functional equivariance property while preserving the local temporal structure of the environment. Building upon this embedding, we further propose unsupervised skill discovery using bisimulation, which learns transferable skills by conditioning the behavior of skills exclusively on the subset of state features that directly affect their execution. This enforces invariant behavior across different layouts, enabling skills to transfer effectively to other configurations. Finally, through comprehensive empirical evaluations, we show that skills learned in a given environment can be effectively applied to solve downstream tasks in various environment layouts, demonstrating strong out-of-distribution generalization.
[AI-149] LabBook: Harnessing Experimental History for Efficient LLM -Driven Discovery
链接: https://arxiv.org/abs/2610.00675
作者: Bo Yuan,Wenqian Ye,Zelin Zhao,Lama Moukheiber,Henry Kautz,Aidong Zhang,Yongxin Chen
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: Under Review
Abstract:Evolutionary approaches to LLM-driven discovery often generate new programs from a small set of selected ancestors. This keeps contexts manageable but can omit useful evidence from other experiments, whereas including the full experimental history produces long, redundant contexts. We introduce a simple, single-agent discovery harness built around LabBook, an agent-maintained memory that serves two complementary roles: guiding retrieval of relevant evidence from a complete experimental log and informing the generation of new solutions. At each iteration, the same agent combines its memory with retrieved evidence and jointly produces the next program and an updated LabBook. This separates complete history retention from selective context construction, without requiring an explicit population or branching search structure. On 49 Frontier-CS problems, LabBook improves the observed quality-cost trade-off over the evaluated evolutionary baselines with two backbones, while remaining competitive across nine additional mathematical, systems, and heuristic-design tasks. Code will be released at this https URL.
[AI-150] A Simple Doxastic Deontic Logic for Norm-Guided Decision Making
链接: https://arxiv.org/abs/2610.00668
作者: Thorsten Engesser,Agata Ciabattoni
类目: Artificial Intelligence (cs.AI); Logic in Computer Science (cs.LO)
备注: Manuscript accepted at PRIMA 2026. Includes an additional appendix with proofs
Abstract:Making decisions despite conflicting norms and incomplete or unreliable information is a fundamental challenge for autonomous systems. We introduce a simple doxastic deontic logic for this setting: a classically reducible fragment of Chellas’ Minimal Deontic Logic, extended with explicit conditional norms and combined with multi-agent KD45, so that norms can depend on agents’ beliefs about both facts and norms. On this logic we define the Doxastic Norm Compliance Optimization Problem, where an agent chooses a decision minimizing weighted norm violations. We distinguish subjective optimization (relative to the agent’s beliefs) from objective optimization (relative to the actual facts). We give conditions under which (i) the two coincide and (ii) optimal decision-making can be reduced to weighted partial MaxSAT in polynomial time.
[AI-151] Backdoor Containment via Expert Quarantine and Shutdown in LLM s NEURIPS2026
链接: https://arxiv.org/abs/2610.00663
作者: Jianwei Li,Min-Seon Kim,Jung-Eun Kim
类目: Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR); Machine Learning (cs.LG)
备注: NeurIPS 2026
Abstract:Backdoored large language models (LLMs) can behave normally on benign inputs while producing attacker-specified outputs under hidden triggers. Existing defenses span four stages–prior-training, in-training, post-training, and inference-time–and share one of two underlying strategies: either suppress backdoor learning (by filtering poisoned data or interrupting its acquisition during optimization) or learn, then purify (by repairing model weights or gating inputs after a fully backdoored model has formed). We propose a third strategy, learn, but channel: allow backdoor formation during training but route it into a designated, quarantined component that can be disabled at deployment. To this end, we propose Quarantined Expert Shutdown QES, a computationally efficient containment strategy built in a regularization-steered MoE-like setting. Specifically, given a poisoned dataset, QES augments a Transformer-based language model with routed expert-specific LoRA branches and lightweight routers, and uses auxiliary routing objectives to attract trigger-conditioned behavior into a designated expert while preserving benign capability elsewhere. At deployment, mitigation reduces to a single constant-time operation: zeroing the quarantined expert’s routing weight, without trigger screening or further updating model weights. Empirically, our methods reduce the attack success rate ASR from 100% to 0-10% on most settings across two tasks, three attacks, and four model families, while downstream utility is often preserved or only modestly affected. These results establish learn, but channel as a previously unexplored regime for backdoor containment in generative LLMs.
[AI-152] Exploring More Reasoning Better: Stepwise Risk-Sensitive GRPO for Diffusion Language Models
链接: https://arxiv.org/abs/2610.00661
作者: Yue YU,Bowen Zuo,David Crandall,Yinglun Zhu,Dongruo Zhou
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: 40 pages, 12 figures, 2 tables. The first two authors contributed equally
Abstract:Diffusion large language models (dLLMs) generate text by denoising a sequence or successive blocks, allowing several tokens to be revealed in parallel. Reinforcement learning with verifiable rewards (RLVR) reuses terminal feedback across these decisions, even as their conditioning context changes. We propose stepwise risk-sensitive GRPO (StepRS-GRPO), which varies the risk coefficient of the group-advantage transformation across denoising states while retaining the underlying trainer. For binary rewards, we show that this transformation is exactly a prompt- and state-dependent rescaling of centered outcome advantages. A capability-based calibration suggests a coefficient scale, while endpoint and interpolation ablations guide schedule selection. Across multiple dLLM backbones and mathematical reasoning benchmarks, StepRS-GRPO improves both pass@1 accuracy and pass@k coverage over centered GRPO, while increasing answer diversity. In our ablation studies, mass-matched controls support the contributions of state allocation and schedule direction, and the gains persist after matching the root mean square (RMS) of the advantages to that of centered GRPO. Reasoning-trace diagnostics further show that the diversity gains from StepRS-GRPO extend beyond final-answer strings.
[AI-153] When More Data Is Not Enough: The Context-Sufficiency Frontier in Generative AI Personalization
链接: https://arxiv.org/abs/2610.00654
作者: Merieme Askour,Ayoub Merimi
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: PREPRINT - SUBMITTED TO JOURNAL OF SERVICE RESEARCH (JSR)
Abstract:Personalization has long relied on customer data to infer what an individual is likely to value. We call this customer evidence: the customer’s historical behavior and preferences. Generative AI extends personalization by allowing providers to supply changing situational information at the moment a response is produced, without encoding every condition in advance. We define this provider-side context as information about what is possible, permitted, or advisable now. This flexibility creates a new problem: once context becomes easy to supply, more is not necessarily better. We develop a theory of context sufficiency in which the relevance of context to the customer’s current intent matters more than its volume. The theory identifies four states, insufficiency, sufficiency, saturation, and interference, and introduces the Context-Sufficiency Frontier to locate the minimal relevant set. In a full-factorial experiment with a generative recommender at a large home-furnishing retailer, relevant context improved appropriateness, while irrelevant context reduced it and destabilized retrieval. The framework shifts personalization from supplying more context toward identifying what the current interaction actually requires and enforcing constraints throughout the service process.
[AI-154] Agent Evaluation Reliability: More Tasks Wont (Always) Fix An Agent Leaderboard
链接: https://arxiv.org/abs/2610.00651
作者: Michael Hardy,Ruhana Azam,Anka Reuel,Mykel Kochenderfer,Sanmi Koyejo
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Applications (stat.AP)
备注:
Abstract:Agent evaluations are increasingly used to compare LLMs and inform deployment decisions, yet ranks can reflect not only the model but also the effects of the evaluation conditions such as the scaffolds or tasks. This makes reliability claim-dependent: an evaluation that reliably ranks deployed systems may not reliably rank underlying models. We ask which conclusions current agent evaluations reliably support and what additional evaluation would improve them. We develop a Bayesian variance-decomposition framework for sparse, imbalanced agent leaderboards and apply it to 22 benchmarks from the Holistic Agent Leaderboard and Harbor Index. The framework separates signal, performance differences relevant to the intended claim, from noise, irrelevant variation that can still change rankings. We find: (1) Reliability depends on the measurement goal. Fixed model-scaffold systems are ranked reliably (0.935-0.994), while underlying-model reliability is substantially lower (0.148-0.841). (2) Scaffold choice can change conclusions. Inter-scaffold reliability measures whether scaffolds preserve model rankings, showing that scaffold effects vary substantially across evaluations. (3) More tasks cannot resolve all uncertainty. Even infinitely many similarly constructed tasks improve model-ranking reliability of a benchmark by at most 0.097 when uncertainty is dominated by limited scaffold coverage. (4) Pooling diverse benchmarks can improve cross-task rankings at lower cost. For rankings across diverse agentic tasks, pooling benchmarks raises projected reliability from 0.44 to 0.75 at the same task budget and can reduce projected cost by up to 83%. Evaluation design should follow the intended claim: identify what a score or ranking should mean, diagnose what limits its reliability, and spend evaluation budget on the sources of uncertainty that matter.
[AI-155] Incident-Arena: Getting agents to the last nine of reliability
链接: https://arxiv.org/abs/2610.00648
作者: Andre Fu,Malik Drabla,Leon Liu,Meji Abidoye,Marek Suppa,Lata Mishra,Adnan El Assadi,Yiyuan Li
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:AI coding agents are ubiquitous in engineering workflows amongst industry and academia. Yet, despite their use in app coding, relatively less attention has been paid to their ability to execute on production incident response. This emerging field, termed agentic site-reliability-engineering (SRE) contains benchmarks limited by (1) unrealistic environments, typically toy repositories (2) non-standard framework implementations and (3) simple static verifiers. We introduce Incident-Arena, a human-built benchmark of 20 carefully selected tasks grounded in real-world deployed open source software. Each task deploys a production application to an ephemeral Kubernetes cluster, injecting a fault from the config layer through underlying images, and a sustained load profile given the task requirements. We also present a novel verification method, going beyond static checks to functional verifiers, holding systems level metrics stable, while ensuring repairs are done safely. Agent trials run an average of 2.81M tokens and 41 turns, going beyond existing benchmarks, demonstrating agentic long horizon reasoning. Across 20 tasks and 3 application substrates, frontier models score below 64.3%, with failures extending from diagnosis/localization errors, through incomplete repairs and unsafe regressions.
[AI-156] CompMat-Bench: Benchmarking AI Agents for Computational Materials Science
链接: https://arxiv.org/abs/2610.00636
作者: Chenmu Zhang,Levi Felix,Jun-Jie Zhang,Xingfu Li,Xuelian Jiang,Tao Jiang,Subhendu Mishra,Xixi Qin,Boris Yakobson
类目: Artificial Intelligence (cs.AI); Materials Science (cond-mat.mtrl-sci); Machine Learning (cs.LG)
备注:
Abstract:Evaluating AI agents on scientific research tasks is constrained by the time and resources required for the underlying experiments or calculations. In computational materials research, repeating the same expensive simulations across agents and trials can make evaluation impractical. We introduce CompMat-Bench, a benchmark of 94 tasks derived from recently published computational materials studies, each asking agents to complete a step toward achieving the study’s scientific goal. We reproduce the research steps in advance and assess agents on preparing inputs and analyzing outputs for expensive simulations, so expensive simulations can be avoided during evaluation. The reproduced inputs and results serve as ground truth for grading agents with fixed rules, without an LLM judge. The benchmark supports four evaluation conditions: single tasks and workflows composed of related tasks, each with full or reduced methodological guidance. With full guidance on single tasks, agents based on three LLMs demonstrate the ability to complete individual materials research steps, with pass rates of 66.0-90.4% across 94 tasks. Both longer workflows and reduced guidance can limit agent performance, but in different ways for different agents: they lower the pass rates of the weaker agents, whereas the strongest agent falls only when a long workflow is combined with reduced guidance. Failure analysis attributes most failures to scientific errors rather than to errors in software usage. CompMat-Bench provides a basis for comparing agents on the steps of real materials research and for analyzing agent failure modes.
[AI-157] Misalignment of Low-Loss Regions Causes Grokking
链接: https://arxiv.org/abs/2610.00620
作者: Yongding Tian,Zaid Al-Ars,Maksim Kitsak,Peter Hofstee
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: 23 pages, 23 figures
Abstract:Grokking refers to the delayed emergence of validation-set generalization after a model has already overfit the training set. Although first observed in small algorithmic tasks trained with transformers, its underlying mechanism remains unsettled. In this work, we develop an analysis framework based on mode connectivity and the geometry of low-loss regions. The framework predicts that the standard modular-arithmetic setting does not always produce grokking: under a symmetry-preserving train/validation split, we observe a stable anti-grokking case in which validation performance does not recover. This counterexample challenges several existing correlational explanations of grokking. More broadly, our analysis framework and results further suggest that grokking arises when the low-loss regions induced by the training and validation partitions are misaligned. Once these regions become well aligned, training hyperparameters alone cannot produce grokking and the observed dynamics collapse to either trainable or non-trainable behavior.
[AI-158] Spatial Strategies Not Actions: Vector-Quantized Geodesics as Tools for LLM -Driven Agents
链接: https://arxiv.org/abs/2610.00613
作者: Gabriel Turinici
类目: Artificial Intelligence (cs.AI); Robotics (cs.RO); Systems and Control (eess.SY)
备注:
Abstract:Large language model (LLM) based agents are often criticized for lacking spatial understanding and mainly exploiting statistical text patterns. We investigate their spatial comprehension through an architecture combining geometrical tools with a LLM serving as a high-level orchestrator in grid-world environments. The agent first collects geodesic trajectories, which are then vector-quantized to extract a representative subset. Offline, the LLM associates a natural language description of the underlying behavioral patterns to each selected trajectory, making it a tool. Online, the LLM chooses the appropriate tool conditioned on the current state and goal. Low-level control is handled by primitive actions that execute the trajectory associated with the tool. From an agentic AI perspective, this approach separates learning into two levels: tool discovery is handled through unsupervised quantization of trajectories, while reasoning and decision-making are handled by the LLM. We test the approach in a partially observable dynamic 2D grid environment with an open vision-language model (Qwen3.6-35B-A3B). Pairing the geometry-derived tool library with an agent-centered zoom tool and a collision detection tool lets a fast, non-reasoning configuration match the goal-reaching rate of a much more costly chain-of-thought version, while cutting the cost of a decision from minutes to seconds.
[AI-159] MIKASA-Robo-VLA: Benchmarking Memory in VLA Models for Long-Horizon Manipulation
链接: https://arxiv.org/abs/2610.00604
作者: Egor Cherepanov,Nikita Kachaev,Aleksandr I. Panov,Alexey K. Kovalev
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Robotics (cs.RO)
备注: 57 pages, 39 figures, 38 tables
Abstract:Vision-language-action policies often see only one or a few recent frames, which makes it difficult to evaluate how they use information that disappears during a task. We introduce MIKASA-Robo-VLA, a benchmark of 90 language-conditioned manipulation tasks. All but 10 hide the cue an action depends on. Those 10 are reactive controls. MIKASA-Robo, the suite it rebuilds, has 32 tasks and uses language only in a representative VLA subset. Here every task provides an instruction, while memory-dependent tasks hide a task-relevant cue and reactive controls keep it available. For 70 tasks, environment phase timings specify an information gap, and for 28 of them the gap exceeds the 16-frame window of the widest fixed-context VLA we survey. The gap counts only the interval the cue is provably absent, not the full duration a policy must retain it, so every memory-dependent task still requires memory by construction, including the ones whose measured gap is short. We release 22,500 oracle trajectories across 10 memory types in RLDS and LeRobotDataset v3. A reference \pi_0.5 baseline with current images and proprioception, but no observation history or explicit memory module, is fine-tuned on 14 tasks and achieves 0.211 \pm 0.044 mean task success. Its lower success on the evaluated Long-split tasks is confounded by open-loop chunking and the memory types represented in that subset. Project page: this https URL
[AI-160] When Reasoning Helps Action: Monitoring and Steering Chain-of-Thought in Vision-Language-Action Policies
链接: https://arxiv.org/abs/2610.00601
作者: Sathwik Karnik,Joseph JR. Lee,Aryaman Gupta,Somil Bansal
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI)
备注:
Abstract:Reasoning-enabled VLA policies expose chain-of-thought (CoT) traces that appear to explain and guide their actions, creating a potential interface for runtime safety through reasoning monitoring and correction. In this work, we define and operationalize two evaluation axes for assessing when this interface can improve embodied behavior: correctability, which measures whether unreliable reasoning can be detected and improved during generation, and actionability, which measures whether reasoning corrections produce behaviorally meaningful changes in the intended direction. To enable correctability, we introduce Token-level Reward for Utility-Steered Chain-of-Thought (TRUST), an offline-trained value model that predicts eventual reasoning correctness from partial prefixes and uses these estimates to monitor and selectively steer reasoning generation in frozen VLA policies. On the Alpamayo 1.5 driving VLA, TRUST monitors correctness with 88.9% accuracy and improves reasoning correctness from 75.9% to 90.0%. On a baseline-defined challenging subset in AlpaSim, TRUST reduces collision rate by 30.4% and maximum trajectory error by 11.5% relative to the unsteered policy, outperforming a compute-matched Best-of-4 baseline. On the DeepThinkVLA manipulation VLA, TRUST improves the correctness of grasp-state claims from 69.3% to 90.2% and action-choice claims from 68.8% to 85.9%, yet closed-loop task performance on LIBERO-Plus remains largely unchanged. Empirical analysis reveals intent-consistent behavioral effects in Alpamayo 1.5 but limited effects in DeepThinkVLA, helping interpret these different task-level outcomes. Together, our results show that gains in reasoning correctness do not automatically imply gains in embodied performance, motivating evaluation of correctability and actionability when using CoT as a runtime safety interface.
[AI-161] ALER: Adaptive Learnable Experience Rewriting for Reinforcement Learning
链接: https://arxiv.org/abs/2610.00592
作者: Oleg Shchendrigin,Egor Cherepanov,Aleksandr I. Panov,Alexey K. Kovalev
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: 28 pages, 12 figures, 18 tables
Abstract:In partially observable reinforcement learning (RL), a later observation can make stored information obsolete or change what it implies for the next decision. Memory architectures and benchmarks for RL mostly test retention, the ability to keep information unchanged until it is needed. We formalize two further requirements. Rewriting sets the decision-relevant content to a value independent of the old one, and experience fusion transforms the old content by a rule that a later observation specifies. For tasks built from such updates, we count the memory states that a solution needs, and several baselines reach their lowest success rates on compositions that need more states. We introduce ALER (Adaptive Learnable Experience Rewriting), an agent that pairs an LSTM with a slot memory. An independently addressed Gumbel-Softmax write that concentrates its weight on one slot overwrites that slot, and a learned gate fuses the retrieved content with the recurrent state before the policy and value heads. We also introduce Rune-Mazes, three environments in which rune observations invert, cancel, reset, or repeat updates of a hidden cue under vector and pixel observations. Against seven baselines, ALER reaches a success rate of at least 0.82 in all sixteen Endless T-Maze configurations and at least 0.99 on all five Rune T-Maze compositions, and it has the highest mean success rate on four-branch Rune Multi-Corridor with an Invert rune. On pixel-based Rune MiniGrid Memory, it has a higher mean success rate than PPO-LSTM in eight of ten configurations. Project page: this https URL.
[AI-162] owards Hierarchical Cyber Defense with Large Language Models : From Planning to Execution
链接: https://arxiv.org/abs/2610.00590
作者: Harshith Doppalapudi,Nathaniel D. Bastian,Ankit Shah
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
备注:
Abstract:An autonomous cyber defender trained with reinforcement learning (RL) is typically tied to the network on which it was trained, limiting its ability to generalize as network scale changes. Hierarchical RL reduces decision complexity by separating strategic targeting from tactical execution, but it does not eliminate this retraining dependence. We investigate whether frozen, zero-shot large language models (LLMs) can provide retraining-free control in hierarchical cyber defense and how performance changes as LLM control is extended from planning to execution. We formulate a controller-agnostic planner-executor hierarchy in which the planner selects a subnet to defend over a fixed horizon and the executor selects defensive actions within that subnet. Using the high fidelity Cyberwheel environment, with its built-in automated red team agent mapped to the MITRE ATTCK framework, we compare RL+RL, LLM+RL, and LLM+LLM configurations using six models ranging from 3B to 70B parameters, including two cybersecurity-specialized models, across small, medium, and large networks. Replacing only the planner with an LLM yields limited gains as network size increases. In contrast, extending LLM control to execution produces notable improvements for sufficiently capable models. For instance, a frozen general purpose 70B model holds successful lateral movement to approximately 1% of steps and attacker impact near zero across all three network scales using the same model weights, while the RL baseline is retrained for each scale. Our results show that sufficiently capable frozen LLMs can maintain strong defensive performance across the evaluated network scales without task-specific retraining, while also indicating that strong tactical execution is important to realizing the benefits of LLM-based control.
[AI-163] Worse Together: How Performance Breaks Down in Multi-User Multi-Agent Teams
链接: https://arxiv.org/abs/2610.00583
作者: Sahan Paliskara,Nattaput Namchittai,Andrew Lampinen
类目: Artificial Intelligence (cs.AI)
备注: 63 pages, 22 Figures, 10 Tables, Code: this https URL (will be released after review)
Abstract:People are increasingly delegating tasks to AI agents, and those agents are increasingly encountering other people’s agents over shared resources such as a codebase, a calendar, or a budget. When each agent acts for a different user with different goals, coordination often fails, and the group ends up worse off than if a single agent had acted for everyone. We study this multi-user, multi-agent setting across five frontier models and 77 scenarios in four environments: an API key environment in which agents share a compute budget, a clinic in which they share a calendar, a personal assistant environment in which they share a group order or booking, and a merge queue in which they share a release cutoff. In each scenario, we compare a single agent that serves every user (a coordinator) to a team in which each agent serves one user, with and without a communication channel between the agents. Teams deliver worse group outcomes than the coordinator in every environment: without a channel, they completely collapse in two environments, and even with one, coordination overhead creates substantial gaps. For example, in the personal assistant environment, the coordinator fulfills a targeted user request about twice as often as teams. We identify distinct behaviors associated with this poor group-level performance, including stalling as teams grow, overriding each other’s actions, and fabricating claims. We find effective but environment-specific mitigations, such as a team lead, explicit procedural instructions, and a platform check that makes an agent read its peers’ messages before committing. We will release the API key, clinic, and personal assistant environments as MAMUBench, comprising 74 scenarios for evaluating multi-user, multi-agent coordination.
[AI-164] Query-efficient winner prediction in district-based elections
链接: https://arxiv.org/abs/2610.00577
作者: Koustav De,Debajyoti Kar,Swagato Sanyal
类目: Data Structures and Algorithms (cs.DS); Artificial Intelligence (cs.AI)
备注:
Abstract:In a district-based election, N voters are partitioned into k districts, and each voter votes for one of m candidates. Each district elects a winner using the plurality rule (i.e. the candidate getting the largest number of votes is declared the winner, breaking ties as per some fixed rule), and the overall winner is determined by applying plurality to the district winners; we assume that there is a unique winner amongst the district winners. The margin of victory of such an election is the minimum number of votes that must be altered so that the current winner ceases to be the unique district winner. We study the problem of predicting the winner of a district-based election in the query complexity model, where one has query access to individual votes. The objective is to minimise the number of queries. This setting captures exit polling, where queries correspond to interviewing voters, and is closely related to problems in query complexity and property testing. Assuming that the margin of victory of the election is at least eps N, Dey, Kar and Sanyal (AAMAS 2023) gave algorithms for the case of two candidates with error probability del and query complexity tildeO(1/eps^6 log^2 1/del), which improves to tildeO(1/eps^4 log^2 1/del) under the additional assumption that district populations are balanced. Our main result is an adaptive randomised algorithm that, for an arbitrary district-based election and any error parameter del, with probability at least 1-del, predicts the winner correctly using tildeO(1/eps^2 log m/del log 1/del) queries. In particular, we improve the bounds of Dey et al. for arbitrary district populations and extend their results to any number of candidates. Furthermore, for constantly many candidates, our algorithm nearly matches a lower bound of Omega(1/eps^2 log 1/del) on the query complexity that holds even for two candidates and a single district. Subjects: Data Structures and Algorithms (cs.DS); Artificial Intelligence (cs.AI) ACMclasses: F.2 Cite as: arXiv:2610.00577 [cs.DS] (or arXiv:2610.00577v1 [cs.DS] for this version) https://doi.org/10.48550/arXiv.2610.00577 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[AI-165] Interpreting Reasoning of Large Language Models via Partial Information Decomposition ICLR2026
链接: https://arxiv.org/abs/2610.00571
作者: Barproda Halder,Qiuyi Zhang,Sanghamitra Dutta
类目: Information Theory (cs.IT); Artificial Intelligence (cs.AI); Computers and Society (cs.CY); Machine Learning (cs.LG)
备注: Accepted at ICLR 2026 Workshop on Logical Reasoning of Large Language Models
Abstract:Large reasoning models (LRMs) have achieved substantial improvements in solving complex mathematical problems, but often produce lengthy, repetitive, or erroneous reasoning trajectories. In this work, we introduce a new interpretability framework, SLIDER, to evaluate the quality of the reasoning process. SLIDER leverages an emerging body of work from information theory called Partial Information Decomposition to disentangle the information about the final answer between two consecutive reasoning steps into non-negative components: unique information (in preceding steps or current step), redundant information, and synergistic information. Building on this decomposition, we propose the Step-wise Repetitive Reasoning Index (Step-RRI), a theoretically grounded measure that assesses whether the answer-relevant information in the current step S_i is predominantly redundant with the past steps S_i , relative to its unique and synergistic contributions. To evaluate the effectiveness of Step-RRI in detecting repetitiveness, we apply SLIDER to the redundancy class of the PRMBench dataset where Step-RRI improves step-level redundancy identification accuracy by over 10 points compared to embedding-similarity and information-gain baselines. Next, we define Trajectory-RRI, an aggregate measure of repetitiveness for an individual reasoning trajectory. To demonstrate its practical relevance, we show that average Trajectory-RRI strongly correlates with actual reasoning length across QwQ-32B, DeepSeek-R1-Distill-Qwen-32B, and GPT-4.1, motivating its use as a signal for improving reasoning efficiency. Finally, we introduce Trajectory-RRI-guided data selection for fine-tuning, demonstrating that selecting training data based on Trajectory-RRI can improve a fine-tuned model’s reasoning efficiency while largely preserving its task performance.
[AI-166] Beyond Affine Transformations: A Soft Dominance Layer for Coordinate-Wise Neural Computation
链接: https://arxiv.org/abs/2610.00563
作者: Mariano Rivera
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: 14 pages, 4 figures
Abstract:This paper presents a preliminary study of an alternative to the affine transformation underlying conventional neural-network layers. In the proposed Soft Dominance Layer, each output unit compares input coordinates with a learnable reference vector and aggregates smooth inequality responses. A sigmoid relaxation makes the comparisons differentiable, while a sharpness parameter \alpha controls their transition toward hard threshold decisions. The aim is to examine the trainability and direct threshold interpretation of this primitive, not to claim a replacement for affine layers. In single-run MNIST experiments, the highest observed Soft Dominance accuracy is 0.9061 without annealing and 0.9173 with annealing, compared with 0.9827 for the MLP baseline. These descriptive results do not establish reliable configuration rankings or a statistically supported annealing benefit. Learned reference vectors exhibit spatial structure, providing qualitative evidence of structured learning. Repeated-seed experiments and broader datasets are required to assess robustness and practical relevance beyond this proof of concept.
[AI-167] No One Architecture Fits All: A Cross-Environment Evaluation of Hierarchical Red Team Agents
链接: https://arxiv.org/abs/2610.00557
作者: Ayan Javeed Shaikh,Arunesh Sinha,Nathaniel D. Bastian,Ankit Shah
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
备注:
Abstract:Autonomous red team agents increasingly stress-test AI-enabled cyber defenses by planning strategy and executing multistage attacks. Reinforcement learning (RL) and large language models (LLMs) offer complementary mechanisms for the planning and execution such agents require, and prior work has combined them in hybrid hierarchies. Yet a given architecture is typically developed and evaluated within a single environment, leaving open whether an observed advantage reflects a generally stronger decision mechanism or merely alignment with a particular setting. We address this gap with a controlled cross-environment comparison of two homogeneous hierarchical red team architectures: an RL planner with an RL executor (RL+RL) and an LLM planner with an LLM executor (LLM+LLM). We evaluate both against expert autonomous defenders in CybORG CAGE-4 and in Cyberwheel at two network scales, across 18 configurations under one unified disruption metric. We find a pronounced environment-dependent inversion. RL+RL wins the compact, densely rewarded CAGE-4 (78.5% disruption success versus 18.0% for the strongest LLM configuration) and the 100-host Cyberwheel network (81.0% versus 50.5%), while a pretrained cybersecurity LLM agent wins the larger, escalation-gated 1010-host Cyberwheel network (55.0% versus 0.0% for RL). A kill-chain analysis explains the inversion through architecture-specific bottlenecks that aggregate success rates this http URL the 1010-host Cyberwheel network, RL discovers and compromises hosts but stalls at privilege escalation, whereas in CAGE-4, LLM agents obtain privileged access but rarely convert it into operational impact. These results indicate that conclusions drawn in a single environment may not generalize, and that hybrid planner-executor designs should be motivated by specific failure modes rather than the assumption that one architecture is universally preferable.
[AI-168] Random Recursive Models
链接: https://arxiv.org/abs/2610.00541
作者: Jama Hussein Mohamud,Mirco Ravanelli
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:
Abstract:Recursive models create computational depth through parameter reuse, offering a parameter-efficient alternative to increasing model size. However, most recursive models repeatedly apply one learned transformation or a prescribed sequence of transformations, restricting computation to a fixed layer order. We introduce the Random Recursive Model (RRM), which maintains a pool of L learned layers and performs T recursive steps by sampling one layer independently with replacement for each example and step. This enables flexible layer reuse while retaining the parameter efficiency of recurrence. We evaluate RRM on challenging reasoning tasks, where it matches or exceeds the baselines, often with 50-75 % fewer parameters. RRM can vary its depth at inference, including beyond that seen during training, without retraining or adding parameters, improving tasks that benefit from deeper iterative computation. RRM also supports Monte Carlo inference and probabilistic test-time scaling, both of which improve performance without retraining. These insights may open new directions in neural network architecture design.
[AI-169] Science or Slop?: Benchmarking and Mitigating Scientific Slop in AI-Generated Papers
链接: https://arxiv.org/abs/2610.00531
作者: Yerim Oh,Young-Jun Lee,Jaewoo Ahn,Gunhee Kim,Dongyeop Kang
类目: Artificial Intelligence (cs.AI)
备注: 28 pages, 6 figures, 13 tables. Project page: this https URL
Abstract:AI-generated content, often called AI slop, is increasingly common everywhere, particularly in academia. Slop in AI-generated scientific papers, however, has more complex patterns that cannot be easily detected by existing token-based AI detectors. Each part of such a paper looks plausible while the scientific reasoning that connects the parts breaks down, which can mislead how readers assess the work. We benchmark these failures as scientific slop through six measures across Structure, Argument, and Artifacts. We construct SciSlopBench with 390 AI-generated papers, mostly in computer science but spanning the life, social, and natural sciences, each paired with a human-written paper matched by research problem and contribution type. Our measures identify the AI paper in each pair with 85.9% accuracy, compared with 68.7% for Binoculars. Higher scientific slop accompanies lower ICLR ratings and distinguishes rejected from accepted papers above chance in every year from 2017 to 2025. Reducing these patterns, however, is not as simple as directly optimizing the measures. We therefore propose SciSlopHarness, a harness-level framework that guides a fixed LLM to revise slop only where the experiment records support the change. While standard revisions leave residual slop and direct slop-aware prompting triggers reward hacking, SciSlopHarness reduces the remaining AI-human gap by 63% over the strongest revision baseline without requiring human reference targets. Overall, we demonstrate that AI-generated scientific papers leave fundamental traces in their global reasoning, and that responsible mitigation demands strict evidentiary grounding rather than mere prose refinement.
[AI-170] Ontology-Based Contextual AI Evaluations (OB-CAIE) Methodology
链接: https://arxiv.org/abs/2610.00529
作者: Julie Krugler Hollek,Michael Zargham,Mala Kumar
类目: Artificial Intelligence (cs.AI); Computers and Society (cs.CY)
备注: 17 pages, 2 figures
Abstract:The ontology-based contextual AI evaluation (OB-CAIE) methodology was developed to address a lack of scientific rigor that arises from unclear testing coverage, to balance human expertise and automations, and to address a lack of reproducibility of AI evaluation testing environments. OB-CAIE strengthens the current state of AI evaluations by addressing the first step in the scientific method by clearly defining what will be tested. Two ontologies represent the tractable problem space in the OB-CAIE methodology: the Domain-Specific Ontology (DSO) and the Evaluation Process Ontology (EPO). The DSO is the what; the EPO is the how. An OB-CAIE problem space can be used for one or multiple AI evaluations. The OB-CAIE methodology allows for human judgment at specific points, in scientifically grounded ways, and in complex subject areas where human feedback is genuinely irreducible or machine irreplaceable. A key advantage of the OB-CAIE methodology is that failure points can be traced, visualized and analyzed within the canonical OB-CAIE methodology problem space.
[AI-171] Gumbel Straight Flow: Distilling Autoregressive Models into One-step Flow Maps
链接: https://arxiv.org/abs/2610.00497
作者: Yeongmin Kim,Arnaud Doucet,Andrew Campbell,Valentin De Bortoli,Thomas Mensink,David Ruhe
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:
Abstract:We present Gumbel Straight Flow (GSF), a continuous flow map language model that leverages the noise-data coupling of a pretrained autoregressive language (AR) model. We theoretically demonstrate that the coupling between Gumbel noise and one-hot token sequences induced by an autoregressive model yields non-intersecting linear paths connecting the noise to the sequence representations. To further enhance high-quality few-step path sampling, we use a flow map semigroup objective where the tangent (velocity) condition is guided directly by the AR teacher. Across various benchmarks, including pretraining and downstream tasks, GSF can outperform current few-step language generation baselines.
[AI-172] JevSpawn: Adaptive Agent ic Inference through Compositional Action Spaces
链接: https://arxiv.org/abs/2610.00437
作者: Haoyang Su,Weiran Huang
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:LLM agents generate intermediate reasoning and actions token by token, making extended interactions slow and computationally expensive. Jev-style models offer fast probabilistic predictions over finite fields, but require those fields to be specified in advance. This requirement limits autonomous task solving, where the available actions must be derived from natural language instructions and adapted through interaction. We introduce JevSpawn, a compositional policy that connects natural language task specifications to finite probabilistic exploration. Parallel action spawning is coupled with feedback driven branch selection, representation revision, and recovery from retained alternatives. Shared action structure and model prefixes reduce repeated generation and context computation without additional training. Evaluations on eight benchmark tasks against seven agent baselines and a TypeSafe Jev variant establish JevSpawn as a promising approach to structured agentic inference, with improved task performance and faster navigation.
[AI-173] XOR-Trellis: Ultra-Low-Complexity Dequantization and Curvature-Aware Hadamard-Free LLM Quantization
链接: https://arxiv.org/abs/2610.00432
作者: Xiaofan Que,Nir Elkayam,Spandan Pyakurel,Shuokai Pan,Dibakar Gope
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:
Abstract:Trellis-coded quantization enables high-dimensional compression of large language model (LLM) weights at ultra-low bit widths without the exponentially large codebooks required by conventional vector quantization. Practical deployment, however, presents two challenges: reconstructing compressed weights at sufficient parallel throughput to avoid making dequantization an inference bottleneck, and maintaining quantization accuracy without costly incoherence transformations. We address these challenges with two complementary techniques. First, we introduce an ultra-low-complexity trellis dequantizer that uses a structured, hardware-efficient state-to-value mapping while preserving diverse reconstruction choices for trellis search. Second, we reformulate discrete trellis path optimization with a curvature-aware objective that reflects model sensitivity directly in the original coordinate space. Together, these techniques enable high-quality ultra-low-bit trellis quantization with inexpensive, highly parallel runtime reconstruction and without relying on Hadamard-based incoherence processing.
[AI-174] Memetic Trojans: Social Contagions as Carriers of Adversarial Payloads in Agent Networks
链接: https://arxiv.org/abs/2610.00430
作者: Birk Torpmann-Hagen,Finn Schwall,Leon Moonen
类目: ocial and Information Networks (cs.SI); Artificial Intelligence (cs.AI)
备注:
Abstract:Autonomous large language model (LLM) agents increasingly interact in network environments where adversarial content can propagate between agents. Known attacks include agent worms, which spread through self-replicating prompt injections or configuration compromises. We introduce \emphmemetic trojans, a distinct class of network-mediated attack that exploits agents’ tendencies to retransmit and amplify content. Unlike agent worms, whose propagation is adversarially induced, memetic trojans exploit \emphendogenous transmission by embedding adversarial payloads in \emphsocial contagions: content agents have internal reasons to share. As part of our work, we extract social contagions from Moltbook, a social media platform for LLM agents. Controlled transmission experiments reveal large differences in virality: the most effective contagion is retransmitted in approximately 50% of subsequent agent posts and upvoted at 2.5x the average post’s rate. Its memetic trojan counterpart largely inherits these properties. Monte Carlo attack simulations show that memetic trojans amplify expected exposure by up to 3.19x. Network structure and amplification mechanisms strongly shape propagation, producing heavy-tailed outcomes with near network-wide exposure. These results identify endogenous social transmission as a distinct security vulnerability in multi-agent systems. Because propagation does not require agents to follow malicious retransmission instructions, defenses focused on prompt-injection detection or preventing agent compromise cannot alone prevent memetic trojan propagation. Securing large-scale agent ecosystems may require network-level defenses that account for how agent preferences, recommendation mechanisms, and network topology amplify adversarial payloads.
[AI-175] Code That Works Environments That Dont: Measuring Environment Reproducibility in AI-Generated Software AAAI
链接: https://arxiv.org/abs/2610.00425
作者: Bhanu Prakash Vangala,Tanu Malik
类目: oftware Engineering (cs.SE); Artificial Intelligence (cs.AI)
备注: 17 pages, 8 figures. Manuscript prepared for AAAI Journal, AI Magazine Special Issue
Abstract:Code generation has emerged as a central capability of large language models, with coding agents now able to produce functionally correct software projects from natural language prompts. However, functional correctness alone does not capture a critical dimension of generation quality: environment specification, defined as the accurate identification of the dependencies required to execute generated code, is equally critical. We develop an agent protocol for environment specification and introduce a three-layer framework comprising declared, runtime-installed, and necessary-and-sufficient dependencies to systematically assess coding agents for environment specification. Using this protocol, we evaluate the extent to which coding agents systematically misspecify software environment dependencies and how this misspecification varies across three agents, four languages, and fifty programming tasks. Our results show that current coding agents exhibit systematic generalization failures along this dimension, producing dependency specifications that are inconsistent, redundant, or incomplete in ways that functional tests do not detect. Across agents, dependency set agreement is as low as 7% for identical tasks, and newer agents show no meaningful improvement, suggesting the failure is not resolved by scale or recency. The largest divergence occurs between the declared and runtime dependency layers, implicating environment priors learned from the models’ training distributions as the primary driver. Our findings establish environment specification as a distinct, measurable axis of code generation quality that current benchmarks do not capture, and motivate training objectives and evaluation protocols that jointly optimize for functional correctness and environmental portability.
[AI-176] he Life Cycle of a Massive Activation: Stochastic Birth Weight-Decay-Driven Growth and Competitive Consolidation
链接: https://arxiv.org/abs/2610.00423
作者: S. Aaron McClendon,Jorge Gallego-Feliciano,Antonios Saravanos
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:
Abstract:Massive activations, residual-stream coordinates with magnitudes far larger than typical activations, are associated with attention sinks in transformers, but how their scale is regulated during training remains incompletely understood. Combining training-trajectory analyses and controlled interventions, we trace their emergence, growth, and consolidation. Sink-carrying channels vary across random seeds but stabilize early within each run. Over longer training, surrounding channels erode and the sink concentrates onto a few redundant carriers. Across ablations, gradient attenuation follows the sink token’s collective root-mean-square magnitude rather than any single channel, making collective scale central to understanding their effects. Our central result is that weight decay causally controls the turnover of global activation scale. In controlled continuations, removing decay near the peak allows this scale to keep rising, whereas retaining it produces decline even at constant learning rate. We develop a balance model for the rise and peak of massive-activation magnitude, in which AdamW-preconditioned growth opposes weight decay. Sweeping the decay coefficient \lambda shifts peak timing approximately log-linearly and yields peak magnitudes scaling approximately as \lambda^-1/2 , consistent with this balance. Optimizer measurements further show that preconditioning sustains the large-channel cohort against decay even when raw maintaining forces are too small to do so. Together, these findings connect the observed life cycle to scale-regulating training dynamics and establish weight decay as a training-time lever on activation magnitude.
[AI-177] Benchmarking Prompt Optimization of Large Language Models With Chess
链接: https://arxiv.org/abs/2610.00416
作者: Timothée Lesort,Alejandra López de Aberasturi Gómez,Tristan Karch,Tom Veniat,Philippe Modard,Karl Tuyls,Ludovic Denoyer
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:
Abstract:Evaluating large language models becomes increasingly challenging as their capabilities advance: benchmarks can saturate, public test sets risk contamination, and assessing harder tasks can require expensive grading or execution infrastructure. These challenges are amplified in automatic prompt optimization (APO), where evaluation is repeated throughout the search for better prompts. Studying APO therefore requires a benchmark that is cheap and deterministic to score, hard enough to leave room for improvement, and renewable as models evolve. We introduce a chess benchmark built from 1,118 Lichess puzzles to study APO for frozen LLMs: we optimize their prompts without updating their model weights. Chess combines inexpensive exact-match scoring, engine-based evaluation of alternative moves, and a renewable supply of problems with adjustable difficulty. Unlike evaluations that report only success on isolated test items, the benchmark also connects puzzle-solving gains to short game-play rollouts within the same domain. We use it to evaluate six APO algorithms on eight target models, measuring not only baseline strength but also how much each model responds to optimization and whether optimized prompts transfer across models and to game play. Chess is thus a well-suited benchmark for APO: it is (i) challenging, as even the strongest evaluated model, Gemini 3.5 Flash (used as the meta-model), solves only about 55% of puzzles; (ii) discriminative, revealing gains, unchanged performance, and regressions across methods and models; (iii) renewable, with fresh puzzles to reduce contamination risk and adjustable difficulty to maintain headroom as models improve; and (iv) affordable, as the complete study runs for around \ 800. We release the puzzles, optimization and evaluation code, and dataset-renewal scripts (this https URL).
[AI-178] Representation Transitions Reveal Emerging Safety Risks in Multi-Turn LLM Agents
链接: https://arxiv.org/abs/2610.00400
作者: Haoyu Wang,Wei Zhao,Yedi Zhang,Christopher M. Poskitt,Jun Sun
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR)
备注:
Abstract:Multi-turn attacks on agentic systems can compose individually permissible actions into harmful outcomes, challenging defenses that assess actions or states in isolation. We show that such attacks leave a detectable signature in the agent’s internal representations: harmful behavior emerges as an accumulated representation transition across context updates, whose triggering context can be identified from the same signal. We further find that naive aggregation is confounded by benign representation drift, as a contrastive safety direction need not assign zero to benign transitions. We address this by denoising the direction, anchoring benign traffic at zero and removing its leading variation directions, with no runtime cost. These findings motivate DART, a runtime framework that detects and attributes representation shifts and intervenes with targeted reminders. Across six models and two multi-turn benchmarks, DART reduces attack success from 84% to 25% on MT-AgentRisk, catching every attack at a mean false-alarm rate of 12%, and from 97% to 52% on ASEval, at costs in benign non-refusal of 8% and 0%, respectively. On MT-AgentRisk, it outperforms ToolShield, the state-of-the-art multi-turn defense, on all six models: under the same protocol, ToolShield reaches only 55%. Denoising is critical: on ASEval, the undenoised monitor catches only 7%-40% of attacks, while the denoised monitor catches 60%-85%. The same monitor covers single-turn indirect injection without modification and adds only 0.14-0.56 s overhead per monitored step without requiring an auxiliary model, making it a lightweight complement to computation-heavy speculative defenses. Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR) Cite as: arXiv:2610.00400 [cs.LG] (or arXiv:2610.00400v1 [cs.LG] for this version) https://doi.org/10.48550/arXiv.2610.00400 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[AI-179] Interpretable Synthetic Medical Tabular Data Generation for Clinical Decision Support Using Fuzzy Cognitive Maps
链接: https://arxiv.org/abs/2610.00391
作者: Michael Vasilakakis(1),Dimitris K. Iakovidis(1) ((1) Department of Computer Science and Biomedical Informatics, University of Thessaly, Lamia, Greece)
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: 6 pages, 2 figures, 1 table. Accepted for publication in the 2026 IEEE 39th International Symposium on Computer-Based Medical Systems (CBMS), Limassol, Cyprus
Abstract:Synthetic medical tabular data generation has become essential for developing and validating computer-based medical systems (CBMSs) when real clinical data is restricted due to privacy, ethical, or data availability limitations. Existing probabilistic and deep generative models often lack interpretability and fail to preserve clinically meaningful dependencies, limiting their suitability for safety-critical applications. This paper proposes a novel application of Fuzzy Cognitive Maps (FCMs) in a framework for synthetic medical tabular data generation with explicit causality and privacy preservation. Clinical features are described using linguistically interpretable fuzzy sets, and inter-feature dependencies are encoded as FCM edge weights computed from fuzzy set intersections. Synthetic patient records are generated by propagating randomly initialized linguistic activation vectors through the FCM until convergence, followed by defuzzification to produce clinically coherent numerical values. The approach natively handles mixed data types, and domain constraints common in health records. Experimental evaluation on UCI medical benchmark datasets demonstrates competitive performance under a Train-on-Synthetic-Test-on-Real (TSTR) protocol. The proposed method achieves accuracy of up to 0.81 and AUROC of up to 0.90 on the Heart Disease dataset, matching or exceeding TVAE and Gaussian Copula baselines while running exclusively on CPU. Fidelity metrics including KS Complement (up to 0.91) and Correlation Similarity (up to 0.95) confirm strong statistical coherence, and DCR Baseline Protection scores consistently exceed those of TVAE, confirming adequate privacy guarantees. These results demonstrate that causally grounded, interpretable fuzzy modeling offers a computationally efficient and transparent alternative to deep generative models for trustworthy synthetic data generation in CBMSs.
[AI-180] MatrixReward: Reward from Rubric Matrix for Open-Ended Generation
链接: https://arxiv.org/abs/2610.00389
作者: Zihan Shen,Qi Liu,Zixuan Yang,Yiqun Chen,Chenglong Zhao,Xiaozhao Wang,Lei He
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Machine Learning (stat.ML)
备注:
Abstract:Open-ended query generation lacks standard answers, thus necessitating an effective reward mechanism. Pointwise scoring rubrics provide limited information about the relative quality of sample answers under the same prompt; merging multiple rubric judgments into a single score may also mask the differences between these answers. We propose MatrixReward, which constructs rewards from a rollout-by-rubric win-rate matrix obtained by comparing every pair of sampled responses under each rubric. The spread of each matrix column captures how strongly that rubric distinguishes the current rollouts, while correlations between columns reveal rubric repetition; together, these statistics yield data-dependent rubric weights. We combine these weights with the prior weights of rubrics. After column normalization and weighting, the observed per-rubric maxima and minima define positive and negative ideal profiles. Each rollout’s distances to these two ideals determine its relative-closeness quality reward. Evaluated using Qwen3-8B on four open-ended query-answering benchmarks, MatrixReward achieves an average score of 63.02, outperforming the strongest baseline by approximately 2.0%. These results support the idea that matrices derived from relative comparisons can be used to construct rewards more reasonably for open-ended generative reinforcement learning.
[AI-181] 2SPO: Trajectory-to-Step Policy Optimization for Agent ic Reinforcement Learning
链接: https://arxiv.org/abs/2610.00388
作者: Bo-Wen Zhang,Junwei He,Maoqi Liu,Feiran Li,Song-Lin Lv,Wentao Ma,Rongyi Lin,Shuhan Zhong,Lan-Zhe Guo
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:
Abstract:Reinforcement learning enables large language model (LLM) agents to learn multi-step behaviors through interaction with their environments. However, rewards in many interactive tasks reflect only the final outcome, providing limited guidance on which intermediate decisions advance the task. Successful training trajectories contain intermediate states that can provide supervision for subsequent interactions. We introduce Trajectory-to-Step Policy Optimization (T2SPO), a method that uses past interaction trajectories to provide step-level feedback for policy learning. T2SPO derives remaining-distance targets from successful trajectories and pairs them with representations of the states visited along the way. Conditioned on these examples, a pretrained TabPFN regressor estimates the remaining distance to success at each state of a new rollout. Changes in this distance estimate across consecutive states yield auxiliary credit for agent steps alongside task-level supervision. As training proceeds, newly completed trajectories refresh the estimator’s context, incorporating new experience without updating its parameters. Experiments with 1.5B and 7B language models on ALFWorld and WebShop show that T2SPO consistently improves overall task success over GRPO.
[AI-182] When Harnesses Lose the Signal: Causal Evaluation of Recovery in LLM Agents
链接: https://arxiv.org/abs/2610.00372
作者: Shuyao Xiao,Shengling Wang,Xuan Chen,Ke Chao,Ming Cui,Feifei Qian,Chaoyang Mei,Fanlin Meng,Ziming Yu,Junxi Yin
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Large language model agents rely on external harnesses to pass information between the model and its environment and to recover from execution errors. Yet recovery is usually judged only by average task success. This hides an important tension. The same operation can rescue a failing trajectory or disrupt one that would otherwise succeed. We frame recovery as a causal decision problem. Starting from the same execution state, we compare what happens with and without recovery, separate rescue from harm, and study how the value of recovery changes over time. We then introduce the Causal Intervention Router (CIR), a lightweight policy that uses information available before recovery to decide when intervention is worthwhile. On long-horizon ALFWorld tasks with Qwen3-14B, CIR raises success from 70.33% to 73.33%, a gain of 3.00 percentage points. It leaves all evaluated trajectories with correct observations untouched. Additional controls show that the benefit of recovery cannot be explained solely by the new observation returned by the environment. These results provide a practical way to evaluate recovery and apply it selectively.
[AI-183] DeepJEPA: Scaling World Models from Within
链接: https://arxiv.org/abs/2610.00368
作者: Zijian Jin,Yunbei Zhang,Yuanzhe Liu,Ming Liu,Baian Chen,Weirui Ye,Shilong Liu,Marco Pavone
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI)
备注: Project page: this https URL
Abstract:World-model planners typically scale outward by rolling farther, sampling more trajectories, or optimizing longer, while assigning the same computation to every imagined transition. We show that making every transition uniformly deeper wastes computation and can degrade planning because useful refinement is concentrated at a small set of decision-critical events. We introduce DeepJEPA, a weight-tied joint-embedding predictive world model that treats transition depth as an inner test-time scaling axis and learns when another recurrent update is worth computing for each candidate and rollout step. Across five visual-control settings, DeepJEPA improves or matches the strongest fixed-depth planner while averaging only 1.00-1.26 updates per transition. Its additional computation concentrates at contact onset and sustained object interaction, where latent corrections can change which candidates enter the planner’s elite set and which action is selected. Representation probes further show that improved planning does not require uniformly better object-state decodability. DeepJEPA therefore reframes world-model scaling as a problem of allocating internal computation where it can change the planner’s decision: think deeper at decision-critical transitions instead of making every rollout uniformly deeper or longer.
[AI-184] What Should an Agent Remember? Disentangling Retention from Retrieval in Bounded-Memory Evaluation
链接: https://arxiv.org/abs/2610.00366
作者: Juli Huang
类目: Artificial Intelligence (cs.AI)
备注: Code available in the accompanying repository
Abstract:A persistent agent must decide both what to retain as information arrives and what to surface once a query appears, yet memory evaluations can confound these decisions by comparing methods that differ in both retention and selection. We build a streaming-recall benchmark crossing retention and selection rules and evaluate every condition on the same 300 seeded episodes. Holding access fixed, query-aware selection improves required-fact recall by 15.5 percentage points (95% CI: 12.8 to 18.2), whereas a mixed comparison that also changes history access reports a 68.7-point advantage, of which 53.2 points are attributable to access. Under bounded retention, query-aware, dense, and oracle selection reach the retention ceiling, and all 319 observed failures in the bounded recency condition are caused by eviction rather than ranking errors. Recall falls to 0% as targets recede sufficiently far into the past. Repeating the evaluation on SQuAD preserves the retention ceiling while showing that dense retrieval can outperform lexical retrieval on natural text. These results show that bounded-memory evaluations should hold access fixed and report retention and selection separately.
[AI-185] Deep Learning for Anomaly Detection in Railway Systems: A Structured Survey
链接: https://arxiv.org/abs/2610.00363
作者: Ammar Bouketta,Smail Niar,Hamza Ouarnoughi
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: Survey paper. Published in Engineering Applications of Artificial Intelligence (EAAI), 2026
Abstract:Ensuring safe and reliable operation of modern railway systems increasingly relies on data-driven monitoring and intelligent fault detection. Deep learning has emerged as an effective paradigm for railway anomaly detection, driven by the growing availability of heterogeneous sensor data from rolling stock and infrastructure. This paper presents a structured survey of deep learning-based anomaly detection approaches for railway systems. The surveyed methods are organized using a unified taxonomy covering anomaly location, data representation and manifestation, sensing modality, and temporal characteristics. Existing approaches, including convolutional, recurrent and attention-based architectures, autoencoders, generative adversarial networks, and transformers, are structured into classification-based, prediction-based, reconstruction-based, and hybrid learning paradigms. The survey also examines data-centric challenges, evaluation practices, performance metrics, and practical deployment aspects, including edge-cloud architectures, computational constraints, and hardware-aware optimization. Finally, a decision-oriented framework links anomaly characteristics, data properties, and operational constraints to suitable detection paradigms and deployment configurations. This work provides a structured reference for selecting and deploying deep learning solutions for railway anomaly detection and highlights open challenges toward reliable and scalable intelligent monitoring systems.
[AI-186] Proof-Gated Signing: Solver-Checked Transaction Guards that Hold Under State Drift for Onchain AI Agents
链接: https://arxiv.org/abs/2610.00354
作者: Bravish Ghosh
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Computational Engineering, Finance, and Science (cs.CE)
备注: 16 pages, 3 figures, 5 tables. Code and data: this https URL
Abstract:AI agents that control wallets read attacker-reachable content, so they can be steered into proposing harmful transactions. The usual last line of defense is a pre-signing check: a static allowlist, an LLM reviewer, or a transaction simulation. All three share a gap: the check describes the chain state at check time, but the transaction executes in a later state that an adversary can shape through front-running, contract upgrades or token-parameter changes. We call this state drift. We present Proof-Gated Signing (PGS), which simulates a proposed transaction, extracts its effects, and uses an SMT solver to check a declarative value-and-permission policy for every price in an oracle-uncertainty band. It then compiles on-chain post-conditions (wallet balance bounds, payee receipts, allowance caps and ownership) and proves that every execution satisfying them also satisfies the policy. The agent’s smart-contract wallet enforces them atomically, so the guarantee applies to the executed transaction under arbitrary drift. On an open testbed of 260 scenarios (14 attack families including five drift and two adaptive families, and 12 benign families), with harm measured from attacker balances rather than from any policy, PGS prevented 93.6% of the 140 harmful scenarios and passed 97.5% of the benign ones. Simulation-only checking prevented 57.9% and a static allowlist 71.4%. None of the 50 drift scenarios produced attacker gain under PGS. The only unprevented family, an in-policy drain, was bounded by the per-session budget. We also find that giving an LLM reviewer a clean pre-drift simulation made it more likely to approve a drift attack. Overhead is about 41k gas and 0.1-0.2 s per check.
[AI-187] JusticeAxis: Benchmarking Legal Judgment between Rigid Rule Application and Ungrounded Discretion
链接: https://arxiv.org/abs/2610.00353
作者: Zhengkai Tu,Mingda Zhang,Zijia Wang,Xiaoying Tang,Jimmy Huang
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:A sound judgment applies the law to established facts and weighs the circumstances in which they arose. However, existing methods swing between rigid statute matching and ungrounded discretion, benchmarks score a label or a rubric, and the experience that would supply the balance stays unverified. We formalize legal judgment as a reference-anchored task, whose object is a single decision that stays tied to the statute and to the circumstances at once. We introduce JusticeAxis, 256 real-world criminal cases from 18 countries with audio, image, and text evidence, and three lawyer-written judgments for every case: the recorded one and one for each failure. We further propose JusticeAgent, a harness whose element agents establish the facts and whose judge agent applies the law under skills carrying experience of the circumstances. Skills are distilled from execution trajectories and admitted only under Bayesian credible bounds. Experiments show that failure turns direction with scale: open-weight backbones drift to unsupported grounds, frontier models to the statutory default. We further verify that JusticeAgent, as a simple yet effective plugin, carries a frozen open-weight backbone to commercial level. Project resources are available at this https URL.
[AI-188] Fault-Tolerant Budget Conservation in Distributed Multi-Agent Delegation
链接: https://arxiv.org/abs/2610.00349
作者: Genliang Zhu,Chu Wang
类目: Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR); Distributed, Parallel, and Cluster Computing (cs.DC)
备注: 67 pages, 3 figures, 17 tables, 4 algorithms, and 3 listings. Includes formal proofs, bounded model checking, mutation analysis, and crash-injected two-process SQLite experiments
Abstract:Resource limits are becoming an authorization boundary for AI agents that delegate work across concurrent and failure-prone workers. Parent-child allocation constraints, affine objects, and distributed escrow do not by themselves prevent overspend when replies are lost, effects complete after timeout, messages repeat, branches partition, or DAG joins alias one lineage. We formalize fault-tolerant budget conservation for distributed multi-agent delegation. Budgets are quantized resource vectors represented by exclusive escrow credits that move through a delegation DAG. Before dispatch, a branch converts credit into an operation reservation bound to lineage, epoch, normalized effect, maximum charge, receiver, and idempotency key. It persists a signed dispatch permit with quarantine; the gateway verifies that permit before first acceptance. Uncertain effects remain charged until authenticated settlement, a fenced authoritative no-effect proof, or permanent retirement. We prove ownership partition, ledger and effect conservation, descendant non-amplification, at-most-once settlement, late-completion safety, and partition confinement under explicit mediation, durability, authentication, normalization, and gateway assumptions. An indistinguishability result shows that partition-local availability requires exclusive preallocation. Bounded TLA+ checking, an independent JavaScript explorer, and crash-injected two-process SQLite experiments exercise the declared scope and detect timeout-refund and historical-certificate-validation mutants. The mechanism preserves the issued budget bound across the evaluated crash, retry, duplicate, partition, join, and late-completion schedules.
[AI-189] Authorization for Self-Modifying AI Agent Populations: Conserving Authority across Replacement Forking and Rollback
链接: https://arxiv.org/abs/2610.00347
作者: Genliang Zhu,Chu Wang
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
备注: 41 pages, 1 figure, 11 tables, and 1 algorithm; includes formal proofs and external runtime adapter evidence
Abstract:Self-modifying AI agents can replace, fork, and roll back identity-bearing software while descendants remain executable. Per-successor authorization does not constrain the resulting population: siblings may duplicate quotas, combine permissions, survive ancestor cuts, or overlap predecessors during promotion. We define authorization succession, which conserves authority across the active frontier of a single-parent generation forest. Our external protocol binds each generation to a manifest, root, unique parent, complete lineage, and fresh population sequence. Separate invariants bound root-lifetime consumption and current population exposure. A staged reservation freezes predecessor residual authority during replacement, while a partitioning fork validates the complete child family. Each commit atomically fences the predecessor and activates successors. Ancestor cuts invalidate dependent descendants; rollback creates a fresh generation without restoring spent authority; and a new root requires an independent grant. Under complete mediation, authenticated records, sound effect abstraction, durable monotone state, and complete lineage accounting, we prove population-safe succession, fork conservation, revocation closure, atomic handoff, rollback non-reminting, and exclusion of self-certification. An executable evaluation covers 32 registered decisions through direct-call and mailbox mappings (64/64 replays; 28 allows, 36 denies). An independent checker accepts all 64 original traces and rejects 28/28 semantic mutants; 12/12 profile invariants, 16/16 crash cuts, and 32/32 contender schedules pass. Two external adapters reproduce all 32 decisions around measured OurArk and Darwin Godel Machine mutations, including fresh-process restart, atomic succession, and predecessor rejection. The results establish authorization succession for registered protected effects. Comments: 41 pages, 1 figure, 11 tables, and 1 algorithm; includes formal proofs and external runtime adapter evidence Subjects: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI) Cite as: arXiv:2610.00347 [cs.CR] (or arXiv:2610.00347v1 [cs.CR] for this version) https://doi.org/10.48550/arXiv.2610.00347 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[AI-190] he Weakest Link: Distilling LLM Reasoning with Worst-Case Constrained Reinforcement Learning
链接: https://arxiv.org/abs/2610.00332
作者: Matthieu Zimmer,Xiaotong Ji,Tu Nguyen,Haitham Bou-Ammar
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:
Abstract:Distilling the reasoning capabilities of large language models (LLMs) into smaller students is a central challenge for efficient deployment. Current approaches face a fundamental tension: optimizing purely for verifiable task rewards (e.g., via GRPO) leads to reward hacking, where students arrive at correct final answers through flawed intermediate logic, while regularizing with soft divergence penalties against a teacher (e.g., KL-based distillation) dilutes task performance and, critically, allows the student to compensate for severe logical violations at one step with high teacher agreement at others. We argue that this averaging is fundamentally misaligned with the nature of reasoning: a chain-of-thought is only as valid as its weakest link. Motivated by this observation, we formulate reasoning distillation as a constrained reinforcement learning problem in which the task reward is maximized subject to a worst-case constraint on the teacher log-likelihood along every prefix of the trajectory. To avoid the prohibitive cost of dual Lagrangian solvers and the test-time teacher dependence of state-augmented methods such as Saute, we derive an unaugmented constrained MDP whose reward transformation preserves the hard-constraint semantics, admits a low-variance policy gradient decomposition into single-step and long-term terms, and provably satisfies the worst-case constraint almost surely in the penalty limit. Through extensive experiments on mathematical reasoning and code generation tasks, we demonstrate that our method significantly expands the accuracy-fidelity Pareto front. By matching the high Final Answer Correctness of pure RL and drastically reducing teacher constraint violations, we ultimately achieve the highest rigorous Reasoning Success Rate across all evaluated settings.
[AI-191] Mathematical Transfer in LLM s Follows Reasoning Approach More Than Topic
链接: https://arxiv.org/abs/2610.00331
作者: Sajad Goudarzi,Samaneh Zamanifard,Seyed Amin Seyed Haeri,Moloud Nasiri,Hamed Rahimian
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:When selecting mathematical training data for LLMs, a natural organizing principle is topic: probability examples for probability targets. An alternative is reasoning approach: worked solutions that share a solution method with the target, even when the mathematical domain differs. We ask which relation produces greater transfer after fine-tuning. We evaluate two counterbalanced 2\times2 designs: probability and combinatorics crossed with invariant reasoning and double counting (2,000 problems), and number theory and geometry crossed with complement and pigeonhole reasoning (800 problems). In each design, every cell serves as the held-out target in turn: same-approach (SA) sources share the target’s method but change the topic, while same-topic (ST) sources share the topic but change the method. Every source appears once in each role, so additive source-quality effects cancel from the equally weighted aggregate contrast. Across five base models and three training seeds per design, SA outperforms ST in all 40 seed-pooled model–target comparisons. Model-level advantages range from 8.2 to 16.2 percentage points in the primary design (mean: 10.8) and from 12.0 to 16.0 in the second design (mean: 14.3); all ten model-level 95% confidence intervals exclude zero. In both designs, ST sources are more similar to targets under embedding and lexical measures, so the SA advantage runs opposite to the measured ordering of statement-level resemblance. These findings identify reasoning approach as a more effective matching criterion than topic for mathematical transfer across the evaluated topic–approach combinations.
[AI-192] Predictive Credit: Measuring What Scientific Explanations Add to Experimental Forecasts
链接: https://arxiv.org/abs/2610.00314
作者: Jingjie Ning,Xueqi Li,Yibo Kong,Dongting Li
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Research agents explain planned experiments. We measure predictive credit with paired forecasts sharing an intervention, forecaster, and outcome while varying description, matched explanation, and donor context. Five checks track commitment, delivery, predictive gain, alignment, and known-signal uptake. Across 336 prospective states in controlled learning, 12 Tox21 endpoints, and 24 OpenML tasks, v5’s frozen credit decision was inconclusive. Tox21’s preregistered ROC AUC interval-score harm test was unmet ( D-M=-.0026 , 95 percent interval [ -.0174 , .0104]); OpenML’s joint formation, point-equivalence, and repeatability rule was unmet. Matched point-accuracy gains over description remained unconfirmed, and Tox21/OpenML seed-donor intervals spanned zero. Under requested DeepSeek V4 Pro, matched and donor cards reduced secondary Tox21 drift by 64.5 and 59.1 percent. A DeepSeek V4 Flash replay raised matched point MAE from .01823 to .02020 and missed matched-donor interval-score equivalence. OpenML full-card assignment widened nominal 80 percent intervals by 21 percent, with 49.3 percent coverage versus 51.4 percent for description and content in 66/144 cards. Direct-text Flash delivered all 144 notes without detectable matched point-accuracy gain. A researcher-authored mechanism positive control lowered point MAE by 2.60 percentage points versus description. The protocol measures predictive credit for research-agent benchmarks and scientific forecasting; natural-explanation credit remained unconfirmed at the tested donor resolutions.
[AI-193] Rules to Tools: Executable Checks for LLM Agents in Scientific Computing
链接: https://arxiv.org/abs/2610.00313
作者: Jingjie Ning,Guojiang Zhao,Chen Xu,Shanshan Zhong,Xiaochuan Li,Ji Zeng,Guolin Ke
类目: Artificial Intelligence (cs.AI); Software Engineering (cs.SE)
备注:
Abstract:Scientific coding agents receive equations, boundary conditions, and output requirements in writing, then must assess the programs they revise. Rules to Tools (R2T) supplies prepared executable checks of public scientific requirements. Matched SciCode repair groups share written checks, starting programs, model, and budgets; the tool group receives a callable implementation. Across two task-ID cohorts, complete repair is 26/30 with text and 29/30 with the prepared checks. Three task IDs favor tools, one favors text, and eleven tie. The eight-ID cohort scores 13/16 versus 15/16, with a task-cluster bootstrap 95% interval of [-12.5, 43.75] percentage points for the difference. The larger shared-definition SciCode cohort ties at 13/24 per group. Five development-exposed tasks with alternate starting programs score 3/10 versus 7/10. The tool group favors tasks 17, 77, and 11; initial checks flag task 17 and report no violation for tasks 77 and 11. Task 37 favors text and has no initial reported violation. A fresh source-through-Python arm also reaches 15/16, matching the dedicated command’s aggregate. In a matched PDE comparison, detailed text scores 23/24 and checks score 24/24, with 31.2% lower reported model output for checks. Agent-side output savings vary by cohort, while public CPU use rises in both task-ID cohorts. These results measure task-dependent repair outcomes and agent-side costs with prepared checks.
[AI-194] Partial AUC Maximization from Positive-unlabeled Data
链接: https://arxiv.org/abs/2610.00284
作者: Atsutoshi Kumagai,Tomoharu Iwata,Taishi Nishiyama,Hiroshi Takahashi,Kazuki Adachi,Yasuhiro Fujiwara
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Machine Learning (stat.ML)
备注: 26 pages
Abstract:The partial area under the receiver operating characteristic curve (pAUC) is an important performance metric for binary classification that summarizes true positive rates within a specific range of false positive rates (FPRs). Classifiers that achieve high pAUC need to be obtained in many real-world applications such as cybersecurity, medical care, and advertising. Although many methods for maximizing the pAUC have been proposed, they typically require both labeled positive and negative data for training. However, in practice, labeled negative data are often difficult to collect due to privacy concerns or the need for high expertise to annotate them. In this paper, we propose a method for maximizing the pAUC from positive and unlabeled (PU) data without negative data. Within an empirical risk minimization framework, we show that the pAUC, including its FPR-dependent thresholds, can be represented using only the positive and marginal densities, and derive an empirical estimator from PU data. A classifier is then trained by maximizing the derived smoothed empirical pAUC estimator. We experimentally demonstrate the effectiveness of the proposed method with ten real-world datasets.
[AI-195] Knowing When to Yield: Grounded Arbitration of User Corrections in Text-Based Embodied Agents
链接: https://arxiv.org/abs/2610.00282
作者: Yezhou Cheng,Runjia Du,Zeming Liu,Hang Lyu,Zehua Yang,Bojun Lin
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:How should an embodied agent respond when a person’s correction may be wrong? We formulate grounded correction arbitration as a choice among accepting, rejecting, inspecting the world, and asking the speaker. GAVA implements this interface with observation-bounded evidence, legal probes, and a one-step expected-loss rule. In text-only ALFWorld, 162 checkpoints produce 972 paired true and false interventions. Complete local inspections give GAVA and always verify 100 percent correction accuracy, establishing the evidence contract rather than a comparative advantage. In same-episode execution, GAVA reduces interaction cost against always verify but ties a cost threshold under a perfect speaker. An exploratory training-only object-location prior lowers interaction and declared joint cost on 340 unseen scenarios by 0.490 and 0.420 relative to uniform GAVA. After freezing the policy, costs, baselines, and multiplicity plan, the gains replicate on 77 non-overlapping seen checkpoints, covering 308 scenarios: 0.595 and 0.517, with both 95 percent checkpoint-bootstrap confidence intervals excluding zero. Joint cost also improves over an identical-prior fixed policy, while the matched calibrated no-VOI comparison remains inconclusive. Semantic GAVA makes four factual errors in each cohort, corresponding to 98.8 percent and 98.7 percent accuracy, and all methods complete every task. Results support selective information gathering with semantic priors under declared costs, but do not establish a general advantage of environmental value of information over clarification. The study uses normalized claims, complete symbolic observations, and controlled speakers; it evaluates neither human participants, visual input, nor physical robots.
[AI-196] Conflicting Supervision Moves Commitment Not Capability: A 12.29σ arrangement effect that is exactly zero under a convention-agnostic score
链接: https://arxiv.org/abs/2610.00234
作者: Wenhui Chen
类目: Artificial Intelligence (cs.AI)
备注: 62 pages
Abstract:“Train a model on the same problems written under two incompatible conventions, both correct, and ask what the ordering of that data writes into the parameters. The learning-rate schedule is not a background condition for that question. It is the averaging operator, and it decides the answer. We prove a bound in which the arrangement and the schedule enter the ordering effect as separate multiplied factors: the arrangement only as a block period, the schedule only as how much weight the endpoint can place on any one moment of the run. A decaying schedule cannot put a large step size and an uncontracted remainder at the same moment; a constant one does exactly that at the last step. That decay moderates ordering effects has been reported in pretraining; the mechanism, the separation, and a controlled measurement of both halves are ours. Ten orderings of one corpus, one budget, everything but the path held fixed, run twice under families differing in lr_scheduler_type and nothing else: at a constant rate the interior spans 0.2221 in allocation, 11.63 contrast floors, monotone in how blocked the arrangement is. Under the single cosine every published arm uses, the same ten arms occupy two distinguishable states where their own resolution would allow about ten. “Order matters” and “order does not matter” are the two ends of one knob. What the path writes is which convention the model commits to, and no exact-match benchmark can see it. Across twelve arms acc_A+acc_B is constant to within 9.7% while the allocation share runs 0.04 to 0.87, so the 12.29-sigma arrangement switch this paper measures is exactly zero under a convention-agnostic metric. That conservation is quoted from the decayed family throughout, the constant-rate one being a noisier place to read it. Marking the convention in the prompt collapses the switch and reaches 87.5% of the union ceiling.”
[AI-197] Build2SPARQL: A Large-Scale Text-to-SPARQL Benchmark Dataset for Building Knowledge Graph Querying
链接: https://arxiv.org/abs/2610.00224
作者: Wooyoung Jung
类目: Artificial Intelligence (cs.AI)
备注: 26 pages, 2 figures, 16 tables. Data paper. Dataset openly available at this https URL . Under review at the ASCE Journal of Computing in Civil Engineering
Abstract:Building automation systems are increasingly represented as semantic knowledge graphs (KGs) using ontologies such as Brick and ASHRAE 223P, creating a machine-readable substrate for artificial-intelligence applications. One promising application is translating natural-language questions into SPARQL (text-to-SPARQL), which would let building operators query these graphs through language agents, but progress is limited by the scarcity of large natural-language/SPARQL benchmarks. This paper presents Build2SPARQL, a large-scale benchmark for building KGs generated by a KG-grounded pipeline: SPARQL queries are produced and validated entirely by graph-traversal code, while large language models generate only the natural-language questions, keeping query correctness independent of model behavior. The pipeline mines six query-pattern families – linear chains, branching, UNION, aggregation, OPTIONAL, and attribute-filtered – and phrases each query across five vocabulary registers. Applied to 201 building KGs (180 Brick, 21 ASHRAE 223P), it yields 6,136 executable SPARQL queries and 30,680 questions. A two-rater human validation of 300 questions found 98.8% semantic fidelity, 98.8% naturalness, and 84.0% operational plausibility. A retrieval-augmented evaluation across three open-weight language models raised exact-match accuracy from 0.2-20% (zero-shot) to 56-65% (three-shot retrieved).
[AI-198] Useful to Whom? Sample Value Is Defined Only Relative to the Learner
链接: https://arxiv.org/abs/2610.00221
作者: Yangze Liu,Xiao-Long Yin,Zhongyi Han
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: 21 pages
Abstract:What kind of data does a model need in order to learn? Coreset selection makes this question concrete: under a budget, keep the samples most useful for training. Easy-first and geometric coverage criteria can win in different budget regimes, separated by a crossover boundary. We ask whether this boundary is fixed by the data or changes with the target learner. Controlled experiments freeze the selected subsets and manipulate only the training learner. On low-resolution ImageNet-100, doubling ResNet-18’s width moves the crossover from 57 to 85 samples per class: the learner changes the relative value of the same samples. A wider sweep reveals an interaction between input grid and capacity. Enlarging the grid while retaining the same image information shifts the boundary left, and this shift weakens as width increases. Stride controls reproduce and reverse the grid effect without changing the input grid; removing only the last downsampling stride is sufficient to recover the leftward shift. Under the native-224px ImageNet-1k protocol, width effects are smaller and depend on the probe: LFrac remains nearly flat, while EL2N shifts modestly right. Swapping the convolutional learning system for a ViT makes coverage win throughout the measured range, even when the easy subsets come from the convolutional proxy. These results establish learner dependence through frozen-subset interventions and identify network structure that can move the boundary. They do not yield a universal scaling law. Their practical implication is direct: a selection strategy’s preferred budget regime must be evaluated with respect to the target learner.
[AI-199] EviGraph: Proof-Carrying Selective Recommendation over Temporal Public-Service Knowledge Graphs
链接: https://arxiv.org/abs/2610.00212
作者: Yixi Zhou,Sikun Wang,Lei Fan,Fan Zhang
类目: Artificial Intelligence (cs.AI)
备注: 19 pages, including figures and tables
Abstract:Public-service recommendations require evidence that matches the requested service, scope, and date. Yet treating every missing detail as decisive can withhold useful recommendations. We introduce EviGraph, which distinguishes critical decision requirements from information that can remain unresolved. A language agent links these requirements to evidence in a temporal knowledge graph, while a deterministic checker establishes whether a recommendation is supported. Evaluation on a bilingual Hong Kong public-service benchmark with executable policy references shows that this distinction reduces unnecessary abstention. Additional verification, however, can withdraw supported recommendations without improving decision quality. These findings suggest that reliable evidence-based navigation depends on specifying what must be established for a decision, rather than simply adding more verification.
[AI-200] ShatterQuant: Breaking Uniform Precision with Block-Wise Mixed-Precision on a Systolic Transformer Hardware Accelerator
链接: https://arxiv.org/abs/2610.00207
作者: Mikolaj Walczak,Edward Humes,Chao Fang,Marian Verhelst,Tinoosh Mohsenin
类目: Hardware Architecture (cs.AR); Artificial Intelligence (cs.AI); Image and Video Processing (eess.IV)
备注:
Abstract:Due to limited support for intra-tensor heterogeneous precision in conventional accelerators, neural network quantization remains largely restricted to per-tensor precision assignment. We present ShatterQuant, a hardware-software co-designed framework enabling mixed-precision quantization within each tensor by assigning independent bit-widths to blocks of a weight projection. ShatterQuant couples precision granularity with PE configuration, such that each precision determines an effective block height. We introduce (1) a hardware-aware post-training method that assigns intra-tensor precision based on block-level standard deviation and weight sensitivity; (2) the ShatterQuant Transformer Accelerator supporting 1/2/4/8-bit weight precision, precision-dependent PE configuration, block rescaling, and integrated softmax and piecewise-linear nonlinearities; and (3) an evaluation of model-hardware tradeoffs using an implementation in the TSMC 16nm PDK operating at 1 GHz, achieving 1.5 TOPS, 760 GOPS/ mm^2 area efficiency, and 2.8 TOPS/W energy efficiency. On DeiT and ImageNet-1K, ShatterQuant achieves accuracy within 3.3% of state-of-the-art mixed-precision techniques while using a 2 bit lower effective bitwidth, while for PixelDiT demonstrates comparable generation quality. ShatterQuant demonstrates how fine-grained intra-tensor mixed-precision can be realized through hardware-software co-design.
[AI-201] White Men Without Degrees Receive the Lowest Ratings from Large Language Models
链接: https://arxiv.org/abs/2610.00185
作者: Maxim Chupilkin
类目: Computers and Society (cs.CY); Artificial Intelligence (cs.AI)
备注:
Abstract:White men without an undergraduate degree receive the lowest average ratings among eight gender-race-education groups in controlled large-language-model evaluations of credit, hiring, and rental applications. We conduct full-factorial vignette experiments with 18 models from 12 developer groups, varying gender, race, age, citizenship, and education while holding stated financial or occupational circumstances constant within each setting. Each model evaluates all 32 profiles ten times per setting, yielding 17,280 ratings. Averaging over models, age, and citizenship, ratings for White men without degrees are the lowest among the eight groups, at 75.87 in credit, 92.71 in hiring, and 86.62 in rental housing on a 0-100 scale. Black women with degrees receive the highest average ratings, with corresponding gaps of 2.94, 3.66, and 3.56 points. Separate attribute effects favor women, Black applicants, and degree holders in all three settings. White men without degrees have the lowest or second-lowest mean in 46 of 54 model-scenario combinations (85.2%). This pattern connects to evidence of growing economic and health vulnerabilities among White men without degrees, highlighting a group whose disadvantages can be obscured by broad racial or gender categories.
[AI-202] Four Ways to Grow a Classifier and Why One of Them Cannot Learn
链接: https://arxiv.org/abs/2610.00180
作者: Cagri Temel
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: 10 pages, 4 tables. Code and measurement scripts: this https URL
Abstract:Constructive classifiers add structure while they train: a level to a tree, a unit to a hidden layer, a split at a leaf. This paper asks what each of four such growth decisions actually buys, measured under one fixed protocol in tree-structured and constructive models, and gives an exact diagnosis and a fix for the one that buys nothing. The diagnosis concerns the most natural way to deepen a soft decision tree: turn every leaf into a gate whose two children inherit the parent’s class distribution, so that the function is unchanged. I prove that this leaves the gradient of every new gate identically zero and, with the gate at 1/2, gives the two children identical gradients, so the added level can never learn. Unlike the symmetry that Net2Net breaks with noise or the saddle point that splitting steepest descent escapes with second-order information, first-order information here is not weak but absent. Over three seeds of five-fold cross-validation the construction loses 19.6 accuracy points on Iris, 19.1 on Wine and 55.6 on Digits against the same depth trained from scratch. The fix is a small random perturbation of the children, whose size barely matters. The practical rule is one line in a test: after adding parameters, assert that their gradient is nonzero. The other three decisions each buy one thing. Fitting a new hidden unit to the residual error before installing it buys a smaller network on every dataset, though not a more accurate one, and on Digits it costs accuracy significantly. Splitting the leaf with the largest expected error buys sparsity, reaching 0.885 with 3.7 splits where a complete depth-six tree uses 63, but loses 4.3 points on a harder problem. Requiring statistical significance before a node receives a more expressive split buys nothing: the tree gets larger and less accurate. Every number in the paper is inserted from the measurement script. Comments: 10 pages, 4 tables. Code and measurement scripts: this https URL Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI) Cite as: arXiv:2610.00180 [cs.LG] (or arXiv:2610.00180v1 [cs.LG] for this version) https://doi.org/10.48550/arXiv.2610.00180 Focus to learn more arXiv-issued DOI via DataCite (pending registration) Submission history From: Cagri Temel [view email] [v1] Wed, 16 Sep 2026 22:00:33 UTC (12 KB)
[AI-203] Per-Node Activation Function Evolution in Indirectly Encoded Substrates: Solvability Limits and Emergent Diversity
链接: https://arxiv.org/abs/2610.00149
作者: Romain Claret,Michael O’Neill,Paul Cotofrei,Kilian Stoffel
类目: Neural and Evolutionary Computing (cs.NE); Artificial Intelligence (cs.AI)
备注: 9 pages, 2 figures, 10 tables. Published in ALIFE 2026 (MIT Press). This is the version of record, posted under CC BY 4.0
Abstract:Biological neurons achieve computational diversity through specialized types: tonic, bursting, adapting, and fast-spiking cells coexist within the same circuit. Artificial neural networks, by contrast, apply a single activation function uniformly to all nodes, which limits what they can represent. We show that this uniformity creates hard limits for evolutionary search: across sparse evolved substrates, monotonic functions fail to solve parity beyond its smallest instance, XOR, while a single oscillatory unit suffices at all tested scales. The gap is one of search and sparsity, not representation: monotonic networks can represent parity with a modest number of hidden units, and gradient descent recovers that solution. We evolve, to our knowledge for the first time in indirect encoding, per-node activation function assignments from an 18-function palette across more than 4,500 experimental runs spanning Boolean logic, regression, and spatial classification. Testing each of the 18 functions individually on Parity-4 reveals a three-tier solvability structure: oscillatory functions achieve 100%, intermediate functions 6.7-80%, and all 9 monotonic functions 0%. This divide is not universal. Recurrence collapses it, and gradient descent inverts it entirely, showing that the barrier is specific to evolutionary search in sparse substrates. What activation functions a network can use, beyond its topology and weights, determines what evolutionary search can solve. Indirect encoding discovers heterogeneous per-node activation assignments unlikely to be chosen by hand. Comments: 9 pages, 2 figures, 10 tables. Published in ALIFE 2026 (MIT Press). This is the version of record, posted under CC BY 4.0 Subjects: Neural and Evolutionary Computing (cs.NE); Artificial Intelligence (cs.AI) Cite as: arXiv:2610.00149 [cs.NE] (or arXiv:2610.00149v1 [cs.NE] for this version) https://doi.org/10.48550/arXiv.2610.00149 Focus to learn more arXiv-issued DOI via DataCite (pending registration) Journalreference: ALIFE 2026: Proceedings of the 2026 Artificial Life Conference, MIT Press, 2026, p. 80 Related DOI: https://doi.org/10.1162/ISAL.a.1042 Focus to learn more DOI(s) linking to related resources
[AI-204] Multi-Behavioral Evolved Substrates Through Neuromodulation and Activation Selection
链接: https://arxiv.org/abs/2610.00148
作者: Romain Claret,Michael O’Neill,Paul Cotofrei,Kilian Stoffel
类目: Neural and Evolutionary Computing (cs.NE); Artificial Intelligence (cs.AI)
备注: 10 pages, 3 figures, 3 tables. Published version of the paper presented at ALIFE 2026: Proceedings of the 2026 Artificial Life Conference (MIT Press). Code and data: this https URL
Abstract:Open-ended artificial life systems must acquire diverse competencies from a single evolving genotype. Biological brains combine neuromodulation, which reconfigures circuits without changing connections, with diverse neuron types matched to specific computational roles. Can artificial evolution achieve something analogous in indirectly encoded substrates? Using indirectly encoded substrates evolved via CPPNs, we show through more than 10,000 experiments that neuromodulation alone is insufficient: under evolutionary search, monotonic activation functions impose a 75% ceiling on parity tasks that persists regardless of capacity, topology, or population size. This is an evolutionary search barrier, not a representational limit, since Adam gradient descent achieves 100% on the identical architecture. We combine neuromodulation with per-task activation function selection, matching oscillatory primitives to parity tasks and monotonic to threshold tasks, producing multi-behavioral evolved substrates. The result: 100% simultaneous 5-task success across all 30 seeds (median 14 generations). This generalizes across the oscillatory activation class: all four functions reach 100% (30 seeds each). Neither mechanism suffices alone. The barrier extends to higher-arity and asymmetric tasks, while multi-layer depth provides an alternative path. For open-ended evolution, the computational primitive should itself be an evolvable trait. At inference, one evolved genotype expresses many behaviors. Comments: 10 pages, 3 figures, 3 tables. Published version of the paper presented at ALIFE 2026: Proceedings of the 2026 Artificial Life Conference (MIT Press). Code and data: this https URL Subjects: Neural and Evolutionary Computing (cs.NE); Artificial Intelligence (cs.AI) Cite as: arXiv:2610.00148 [cs.NE] (or arXiv:2610.00148v1 [cs.NE] for this version) https://doi.org/10.48550/arXiv.2610.00148 Focus to learn more arXiv-issued DOI via DataCite (pending registration) Journalreference: ALIFE 2026: Proceedings of the 2026 Artificial Life Conference, MIT Press, 2026, p. 78 Related DOI: https://doi.org/10.1162/ISAL.a.968 Focus to learn more DOI(s) linking to related resources Submission history From: Romain Claret [view email] [v1] Thu, 10 Sep 2026 13:22:23 UTC (78 KB)
[AI-205] he Cognitive Continuity Test: Verifying Governed State Transitions in Persistent AI Agents
链接: https://arxiv.org/abs/2610.00132
作者: Jun He,Deying Yu
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
备注: 17 pages, 2 figures, 3 tables. Includes formal proofs, transition taxonomy, and benchmark schema appendices. Reference verifier and reproducible evaluation artifacts available at this https URL
Abstract:Persistent AI agents revise beliefs, consolidate memory, and replace execution substrates. Similar successor states can accompany differently authorized transition claims, while legitimate development can change state substantially. We introduce the Cognitive Continuity Test (CCT), a policy-relative contract for verifying submitted transitions using scoped authority, provenance, deterministic application, semantic predicates, and candidate-persistence receipts. CCT distinguishes verified admissibility, affirmative violation, and unresolved required evidence. Separation results concern transition claims rather than live runtime identity; soundness is conditional on the specified checker and evaluator assumptions. IdentityLineageBench provides 24 generated transition families. The reference post-resolution verifier matches all 576 canonical held-out labels; lexical state similarity and a lineage-only diagnostic baseline admit 60.0% and 80.0% of invalid fixtures. These comparisons establish synthetic conformance, not superiority to a policy-aware deployed system. Signed adversarial regressions cover fabricated interaction counts, unsupported belief changes, and mixed missing/contradictory evidence. SIT behavior and actual model migration remain unmeasured. An 18,000-execution valid-path study measures a 6.21 ms default median on resident inputs. We specify the additional activation and recovery obligations needed for deployment. Comments: 17 pages, 2 figures, 3 tables. Includes formal proofs, transition taxonomy, and benchmark schema appendices. Reference verifier and reproducible evaluation artifacts available at this https URL Subjects: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI) ACMclasses: D.4.6; I.2.11 Cite as: arXiv:2610.00132 [cs.CR] (or arXiv:2610.00132v1 [cs.CR] for this version) https://doi.org/10.48550/arXiv.2610.00132 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[AI-206] Nous: Learning and Certifying Memory Decisions Before Source Calibration
链接: https://arxiv.org/abs/2610.00094
作者: Pranav Singh
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: 19 pages, 3 figures; code and reproducibility package available at this https URL
Abstract:Belief-based agent memory needs reliable decisions about current state, yet its evidence may be noisy, copied, or stale. Must a memory calibrate its sources before it can improve its decisions? We separate learning, calibration, and revision certification. On one four-model hidden Markov family, learning an unknown Bayes decision requires Theta(l^-2) records and certifying its improvement over an informative incumbent takes O(l^-2) fresh records from the same observation law, while fixed-precision source estimation requires Theta(l^-4) as persistence l vanishes. Thus learning and certifying useful decisions can require quadratically fewer records than source calibration. A broader model class retains the decision rate and source lower bound. Under an unknown identity-plus-background report channel, we characterize the sharp identified interval for policy improvement and derive a finite-sample certificate using observable witness regions, without pure-class anchors. A robustness extension tolerates bounded history-dependent misspecification and conditional copying; split-trained witnesses apply to arbitrary history spaces with explicit power conditions. We integrate policy-bound receipts with Nous Dimensions and test 45,000 held-out mutable-state histories and 9,000 episodes in three external MiniGrid memory environments with an introduced noisy-report interface. The new certificate accepts 9/9 improvements over a constant incumbent and 4/9 over last-write-wins, versus none for the earlier certificate in MiniGrid. Strong established inference baselines remain competitive or better. The result is a statistical account of when memory decisions can be learned and justified without recovering source reliability, not a universally superior memory algorithm.
[AI-207] Safety in Self-Evolving Agents : A Survey
链接: https://arxiv.org/abs/2610.00093
作者: Jiahao Chen,Zhou Feng,Oubo Ma,Yichen Yan,Ruixiao Lin,Hangtao Zhang,Linkang Du,Hengyu An,Yong Yang,Jun Liu,Junhao Li,Naen Xu,Chunyi Zhou,Yuan Su,Zehao Jin,Qianli Ma,Leyi Qi,Yiming Wang,Zhe Ma,Yuwen Pu,Mengyao Du,Yuanyi Song,Enhao Huang,Zhihui Fu,Jun Wang,Jinfeng Li,Yuefeng Chen,Hui Xue,Yiming Li,Tianyu Du,Shouling Ji
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
备注: Survey paper; 80 pages, 6 figures, 13 tables. Project page: this https URL
Abstract:Large language models (LLMs) exhibit strong general capabilities, yet their parameters typically remain fixed after deployment, limiting learning from new interactions. In open-ended environments, this motivates self-evolving agents that continually update reusable state-including model parameters, memories, tool definitions, skills, and workflows-from data, feedback, and accumulated experience. This shift changes the safety problem: once experience becomes reusable state, past events become future causes, and information harmless in one context may later influence decisions with greater persistence, authority, or scope. Self-evolving agent safety therefore asks not only whether a response is aligned or an action authorized, but whether safety properties survive the accumulation, generalization, and cross-context reuse of locally useful experience. We introduce SAVER, a transition-centered framework in which Substrate locates reusable influence, Adaptation captures how it changes, Violation identifies compromised safety attributes, Exposure marks where failures become observable, and Response assesses containment, repair, or revocation. Our survey reveals that failures need not originate from harmful information: legitimate state can become unsafe when adaptation expands its persistence, authority, or scope beyond the conditions under which it was valid. Existing work provides comparatively strong evidence for admission, retrieval, activation, exposure, and local containment, but much less for descendant repair and evaluation after adaptation resumes. We therefore argue for longitudinal evaluation that traces unsafe influence to its originating transition, verifies repair across descendants, and tests whether it can re-emerge under continued evolution.
[AI-208] K-Dense BYOK: An Open-Source AI Research Assistant That Runs Locally and Keeps a Hash-Chained Lab Notebook
链接: https://arxiv.org/abs/2610.00074
作者: Aubrey M. Brueckner,Darshil Patel,Yuhuan He,Timothy Kassis
类目: Artificial Intelligence (cs.AI)
备注: 38 pages, 8 figures plus a graphical abstract; includes benchmark prompts, scoring rubric, and per-prompt scores. Code: this https URL
Abstract:K-Dense BYOK (bring your own keys) is a free, open-source AI research assistant for scientists in any field that runs on the researcher’s own computer. The researcher supplies access to a model of their choice, hosted or running locally, and the application supplies everything else: a place for the work to run, a layer of scientific scaffolding, and a complete record. Each project is an ordinary folder, so the data, the code, the results, and the record stay on a machine the researcher administers and can be read years later without the application. Three things separate it from a chat assistant or a general-purpose coding agent. It ships a library of written scientific procedures, guided workflow templates, catalogs of where research data can be found, and reviewer and writer roles the agent can hand work to. It keeps a Living Lab Notebook whose entries link into an argument and are added to but never erased. And it records what happened by watching what the agent does rather than by taking the agent’s word for it, in a log the agent has no tool that can write to. That choice targets the most common failure, model overclaiming, in our earlier benchmark of nine frontier models, by making claims checkable rather than preventing them. On twenty interdisciplinary research prompts, scored under a rubric fixed in advance, K-Dense BYOK led two managed platforms on both scientific quality and research execution. Its deliverables were the only ones that recorded the software they ran in, and the only ones that usually arrived with a command that regenerates the results. One of the managed platforms ran the same frontier model and supplied neither. Those environment records were files the agent wrote, not part of the observed log, which does not yet capture the software environment itself. The code is available under the MIT license at this https URL.
[AI-209] Probabilistic Plan Legibility with Off-the-shelf Planners ICAPS ICAPS2021
链接: https://arxiv.org/abs/2610.00065
作者: Michele Persiani,Thomas Hellström
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI)
备注: Accepted at the 9th ICAPS Workshop on Planning and Robotics. ICAPS 2021
Abstract:Legible planning is the creation of plans that best disambiguate their goals from a set of other candidates from an observer’s perspective. In this paper we propose a method for legible planning for arbitrary PDDL domains, by extending previous research on legibility to classical planning without requiring to construct ad-hoc planners. We also discuss how the observer perspective may be estimated through a second order theory of mind that connects the planner’s and the observer’s task spaces. Our solution can for example be deployed in human-robot teaming scenarios, where an autonomous robot in a team can implicitly communicate its goal by producing legible plans. We present benchmark results on several PDDL planning domains. Our results generally show that plan legibility is a trade-off with plan efficiency, however, not all planning domains allows to increase legibility in the same way and a regularizing factor to balance legibility and efficiency was proved necessary.
[AI-210] Gradient-Aligned Pair Selection for Personalized Preference Optimization
链接: https://arxiv.org/abs/2610.00061
作者: Ruoming Jin,Xinyu Li,Hao Zhou,Jianfeng Zhu,Ruixin Guo,Feodor Dragan,Lei Xu,Haixun Wang,Yang Zhou
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Personalizing large language models (LLMs) requires aligning generation behavior with user-specific preferences rather than aggregate quality. While Direct Preference Optimization (DPO) provides a stable framework for preference learning, its effectiveness in personalized settings critically depends on how preference pairs are selected. Existing approaches typically rely on heuristic criteria, such as likelihood-based extremes, which decouple optimization from explicit user utility and can lead to degraded personalization. We formalize personalized preference learning as a geometry-aligned optimization problem by analyzing the first-order interaction between gradients of expected user utility and DPO update directions. Our analysis reveals that, under off-policy sampling, the DPO update transitions from a purely error-corrective signal to a reinforcement-like update when preference margins are directionally aligned with utility gradients. This perspective exposes pair selection as a geometric decision that governs whether preference optimization advances or hinders personalization. Motivated by this insight, we propose GAP-DPO (Geometry-Aligned Preference DPO), an iterative algorithm that performs utility-aware, geometry-aligned pair selection while controlling distribution shift via epoch-wise regeneration. Experiments on personalized text generation benchmarks show that GAP-DPO consistently improves stylistic fidelity, preference alignment, and generation quality compared to standard DPO variants. Together, our results establish gradient alignment as a unifying principle for personalized preference optimization and demonstrate that pair selection is an intrinsic component of the optimization geometry rather than a heuristic preprocessing step.
[AI-211] SW-KAN: Kolmogorov-Arnold Networks with Stieltjes-Wigert q-Orthogonal Polynomials
链接: https://arxiv.org/abs/2610.00050
作者: Amirhosein Azarpour,Seyyed Moein Kazemi
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: 22 pages, Code and pretrained models available at: this https URL
Abstract:Kolmogorov-Arnold Networks (KANs) represent a paradigmatic shift in deep learning by replacing fixed node activations with learnable univariate functions on edges, offering enhanced interpretability and parameter efficiency. While recent polynomial-based KAN variants have addressed the computational overhead of original B-spline implementations, they introduce a fundamental yet underexplored challenge: the domain mismatch between unbounded real-valued inputs and the bounded or semi-infinite support of orthogonal polynomial bases. To address this limitation, we propose the Stieltjes-Wigert Kolmogorov-Arnold Network (SW-KAN), a novel architecture that employs Stieltjes-Wigert q-orthogonal polynomials defined on the semi-infinite domain (0, infinity). We introduce a smooth exponential-of-tanh mapping that stably bridges the domain gap while preserving well-conditioned gradients, and leverage a numerically stable three-term recurrence that evaluates polynomial expansions in O(N) operations without special-function calls. Through comprehensive experiments spanning image classification and continuous function approximation, we demonstrate that SW-KAN achieves superior accuracy-efficiency trade-offs across diverse tasks. The log-normal weight structure and learnable q-parameter of Stieltjes-Wigert polynomials provide a distinct inductive bias that enables robust performance under resource-constrained conditions, including reduced feature dimensionality and limited training data. The proposed architecture not only outperforms established polynomial KAN baselines on standard benchmarks but also exhibits strong representational capacity for approximating complex multivariate functions with remarkably few parameters, making it a compelling alternative for efficient function approximation and classification in resource-constrained settings.
[AI-212] Measuring the Microtask Eligibility Gap: When Is an Off-the-Shelf SLM Enough for an Agent Harness? NEURIPS2026
链接: https://arxiv.org/abs/2610.00025
作者: Jundong Hu,Shekar Ramachandran
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: Preprint. under review at a NeurIPS 2026 workshop. 15 pages, 8 figures, 14 tables
Abstract:Agent harnesses increasingly want to run small language models (SLMs) on the microtasks around a frontier large language model (LLM) planner: auto-approving shell commands, writing memory, selecting tools, ranking past turns. We ask whether off-the-shelf SLMs meet practitioner-defined thresholds and, when they fail, why, and whether quantization changes the answer. We build a benchmark of 4 such microtasks with fixed prompts and automatic metrics, each with a pre-specified threshold \tau anchored to a cheap non-LLM baseline and a CI-aware eligibility rule (a configuration passes only if its confidence bound clears \tau ). Sweeping Qwen3 0.6/1.7/4/8B at their best (FP16, greedy, one frozen prompt, no tuning), we find an eligibility gap: 0 of 16 (4 tasks \times 4 models) configurations pass (verified by checking the raw outputs and parser behavior). A logprob decision-threshold diagnostic (T1/T3/T4; T2 via a context-length/cascade probe) separates the failures into capability deficits and failures that can be addressed by changing the decoding threshold (4 regimes). Quantization to 4-bit (RTN/GPTQ/AWQ) does damage that depends on model size and moves no configuration into eligibility (certified on the reconstructable hard-label tasks T1/T3, diagnostic/windowed robustness on T2/T4), so the gap tracks model size more than precision; it replicates on Llama-3.x (12/12 ineligible) and is robust to the anchor choice (a \tau -sweep) and to prompt wording (0/112 eligible across the original plus 3 neutral paraphrases per cell). The practical implication: place SLMs behind a baseline that meets the CI-backed threshold, and use the SLM only where the baseline fails to meet the threshold; e.g. a 4B re-ranker over a BM25 shortlist beats BM25 ( +0.047 [0.020, 0.073], without itself certifying eligibility).
[AI-213] From Proposal to Verified Effect: Praxa an Evidence-Bound Harness for Governed AI Agent Execution
链接: https://arxiv.org/abs/2610.00015
作者: Stefan G. Creadore
类目: Artificial Intelligence (cs.AI)
备注: 30 pages, 8 figures. Engineering validation and descriptive pilot. Public artifacts: this https URL
Abstract:Large-language-model agents can propose and execute actions, but proposal, authority, dispatch, verified external effect, and serving promotion are different claims. We present Praxa, an agent harness that represents these states explicitly through deterministic admission, brokered execution, external read-back, reconciliation, and reviewed promotion. We report four evidence lanes. First, an author-run repository-local audit at a pinned revision passed 1,027/1,027 unit tests and 89/89 Workerd tests, instrumented all 363 expected source files, and met four coverage floors; raw per-test transcripts and independent reproduction are unavailable. Second, in a provider-backed Terminal-Bench Core 0.1.1 pilot across 12 curated tasks, baseline and reliability-layer arms each passed 17/36 strict trials. The reliability layer used 37.49% more input and 50.73% more output tokens, so the pilot does not support superiority. Third, in a post-debug, two-order coordination-proxy development comparison, baseline and a source-authored candidate each completed 180/180 trials with equal measured accuracy, full hermetic crash recovery, and zero protected violations. The candidate used 37.11% fewer tokens, 33.84% lower estimated endpoint cost, and 11.63% fewer steps; this does not establish improved quality, latency, or production behavior. Fourth, deployed source/configuration evidence shows bounded reflection, recall accounting, memory compilation, and tool-health paths, but no production outcome lift. Praxa’s supported contribution is an evidence-bound architecture that makes authority-to-effect transitions explicit and testable. Current evidence does not establish adversarial security, production safety, general specialist superiority, autonomous recursive optimization, or user benefit.
[AI-214] When Do Causal World Models Help Modular LLM Agents
链接: https://arxiv.org/abs/2610.00012
作者: Xinyuan Song,Zekun Cai
类目: Artificial Intelligence (cs.AI)
备注: Under Review
Abstract:LLM agents increasingly act through modular systems, such as order, payment, inventory, and shipment services, where actions in one module change which transitions are valid in another. Standard world models usually fit observational traces, but this is not the quantity needed for intervention-time planning: a trace may show that payment precedes shipment without identifying whether payment authorizes shipment, inventory mediates the effect, or a hidden trigger explains both. We study this gap through FedCausalCompose, a causal world-model framework for modular LLM agents in which local actions provide intervention-response evidence for cross-module interfaces. We first show that observational world models incur an irreducible interventional error under unblocked back-door paths, that interface recovery improves with intervention-response coverage, and that an oracle causal composition can beat the non-causal lower bound when coverage and local mechanism errors are controlled. We then test the resulting prediction in diagnostic agent settings. Causal interfaces help most in structured tool environments, where API signatures expose preconditions and downstream effects. In contrast, dialogue and narrative environments often ignore raw edge lists unless a short attention anchor makes the causal information decision-relevant. These results identify a concrete condition for causal world models in LLM agents: causal structure helps when cross-module interfaces are both statistically identifiable and presented in a form the agent can use at action time.
[AI-215] Heavy-Tailed Memory Traces in Long-Horizon Language Agents
链接: https://arxiv.org/abs/2610.00010
作者: Xinyuan Song,Zekun Cai
类目: Artificial Intelligence (cs.AI)
备注: Under Review
Abstract:Long-horizon language agents increasingly rely on external memory as a frozen world model, yet current memory systems are usually judged only by task success or token cost. We argue that the missing object is the shape of memory use: under finite context and repeated retrieval, agent memory can concentrate on a small core while leaving rare states in a long tail where prediction errors accumulate. We study this effect through a conservative tail audit and find that concentration is reproducible but policy-dependent. Random-walk agents produce log-normal-compatible retrieval artifacts, whereas semantic LLM policies yield the strongest truncated-power-law-compatible core–tail traces. Motivated by this audit, we propose Core–Tail World Model (CTWM), a rank-based memory controller that allocates prompt budget with a single exponent \tau while retaining a summarized tail. On Synthetic Graph World, CTWM preserves full state and transition coverage, reduces prompt tokens by 5.9%, and lowers bottom-half tail prediction error by 13.6% relative to a graph-memory baseline. The same paired comparison gives consistent token savings on ALFWorld and a 24.48% token reduction on LongMemEval with aggregate accuracy parity. These results suggest that heavy-tailed memory traces are not only a diagnostic of finite retrieval, but also a practical control signal for token-efficient agent world models.
[AI-216] Bounded-Fidelity Sim-as-Demo-Stage: Mocap Handoff for Governance Benchmarks
链接: https://arxiv.org/abs/2610.00008
作者: Xue Qin,Simin Luan,Cong Yang,Zhijun Li
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI)
备注: 11 pages, 3 figures, 5 tables. Reference implementation and data: this https URL
Abstract:Sim-to-real research pursues physics fidelity as a primary objective: simulators are judged by how closely they reproduce real-world contact dynamics. For governance benchmarking of LLM-driven robots, where the simulator demonstrates that an admission/policy/contract/audit pipeline behaves correctly, contact fidelity at object handoffs (grasp, carry, place) becomes a liability: contact-force integration noise injects audit-chain divergence that is structurally unrelated to the governance property under test. We propose bounded-fidelity sim-as-demo-stage, a design pattern that suppresses contact physics within explicitly bracketed handoff envelopes while preserving full dynamics elsewhere. The construction uses MuJoCo’s mocap-body primitive driven by a 220-line Python adapter that the governance bridge invokes via structured intents. We formalise audit-chain stability as byte-equality of the hashed event log across replays and identify two structural envelope properties that imply it. Across N=1000 replays per posture, the mocap variant produces one distinct audit-chain hash (1000/1000 byte-identical; Wilson 95% CI [0.997, 1.000]); the contact-force baseline produces 584 distinct hashes (993/1000 diverged; CI [0.987, 0.998]). A timestep sweep (1, 2, 5, 10 ms) shows the divergence is structural, not a tuning artefact: it stays at 0.985 at every timestep. Envelope-edge timing jitter (+/-10 simulation steps, 1,400 replays) produces 0 divergence, and audit chains remain byte-equal across K in 1, 2, 3 sequentially handed-off objects (1,500 replays) with sub-linear per-pick-and-place overhead. The pattern gives benchmark designers audit-chain reproducibility at near-zero engineering cost; we also map where it is harmful (sim-to-real validation, policy training, contact-rich tasks) so it is not mis-deployed.
[AI-217] New Snake-in-the-Box Records via Snakepit Surgery and Learned Construction DATE
链接: https://arxiv.org/abs/2607.15270
作者: Paul Orland,Lucas Fagan,Michele Tarquini,Davide Passaro,Maksymilian Manko,Elli Heyes,Angus Gruen,Giorgi Butbaia,Justin Tan,Sergei Gukov
类目: Discrete Mathematics (cs.DM); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Combinatorics (math.CO)
备注: Updated to include detailed information about methods. 23 pages, 4 figures
Abstract:The snake-in-the-box problem asks for a longest induced path in the hypercube graph Q_n . We find a length-191 snake in dimension n=9 , the lowest dimension where the maximum is unknown, improving the previous record of 190 that had stood for 14 years. We also establish new lower bounds in dimensions 10-13. To find these records, we introduce snakepits, collections of disjoint snakes, to expand the search space and open new routes between snakes. This motivates our new Snakepit-in-the-Box benchmark, which seeks maximal edge counts when allowing multiple components. Finally, we introduce Beam Anchor, a search-supervised learned constructor algorithm that finds 100 inequivalent length-190 snakes in dimension 9.
[AI-218] Reinforcement Learning to Accelerate Primal-Dual Hybrid Gradient for Linear Programming
链接: https://arxiv.org/abs/2610.01546
作者: Jinhwan Sul,Alex Oshin,Evangelos A. Theodorou
类目: Optimization and Control (math.OC); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: 35 pages, 4 figures
Abstract:Primal-dual hybrid gradient (PDHG) methods solve large-scale linear programs (LPs) using GPU-friendly matrix-vector products and projections, but their practical performance depends on coordinating algorithm parameters, acceleration, and restarts. We introduce GALLOP, which uses reinforcement learning to jointly learn continuous algorithm parameters and discrete restart decisions without differentiating through the solver. Its generalized accelerated PDHG update combines separate primal and dual extrapolation, history corrections, and restart anchoring with independently adjustable coefficients. We train a dimension-agnostic feedback policy using a groupwise proximal policy optimization objective that clips likelihood ratios separately for different control groups and excludes inactive acceleration controls on restart transitions. We evaluate GALLOP on six LP families and a public item-placement benchmark. On the main evaluation settings across the six families, GALLOP reduces iteration counts by factors of 1.9 - 5.6 and achieves up to a 16.0\times speedup in algorithm wall-clock time over MPAX. With one policy trained per family, the learned policies generalize without retraining to within-family LPs 3\times - 400\times larger than the largest training instances, including Transport LPs with 10.24 million variables.
[AI-219] Optimal Transport Meets Reinforcement Learning: A Survey
链接: https://arxiv.org/abs/2610.01413
作者: Yujie Zhu,Charles A. Hepburn,Matthew Thorpe,Giovanni Montana
类目: Machine Learning (stat.ML); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:
Abstract:Reinforcement learning (RL) algorithms frequently compare probability distributions, such as state visitation distributions induced by policies and experts, action distributions from learned policies and offline datasets, or transition distributions from learned models and environments. However, commonly used divergences may become ineffective when these distributions overlap weakly, which is frequently encountered in imitation learning, offline RL, and deployment under distribution shift. Optimal transport (OT) offers an alternative by measuring the cost of \emphmoving probability mass from one distribution to another under a ground cost that encodes task geometry. This survey covers how OT is used inside RL objectives and algorithms. For each method, we identify: the role OT plays, the distributions compared, the OT formulation used, and the treatment of temporal structure. Beyond categorising existing methods, we discuss the motivations behind different OT choices, practical considerations such as cost design and computational challenges, and highlight open problems including scalable trajectory-level transport, principled handling of mass mismatch, and theoretical analysis for OT-regularised RL.
[AI-220] FoldEM: Direct atomic structure inference from Cryo-EM particles
链接: https://arxiv.org/abs/2610.01358
作者: Advaith Maddipatla,Märt-Erik Mäeots,Marco Pegoraro,Nikolaus Dräger,Roberto Covino,Sanketh Vedula,Martin Pacesa,Alex M. Bronstein
类目: Biomolecules (q-bio.BM); Artificial Intelligence (cs.AI)
备注:
Abstract:Single-particle cryo-electron microscopy (cryo-EM) has become a widely adopted technique for biomolecular structure determination. The conventional cryo-EM computational pipeline first combines many particle images to reconstruct an electrostatic potential (ESP) map and then fits an atomic model to the recovered map. Density reconstruction has high sample complexity, requiring large numbers of particle images and making structure determination high-cost and low-throughput, particularly for heterogeneous samples. Downstream atomic model building, in turn, becomes increasingly difficult as the resolution of the reconstructed map deteriorates. Protein structure prediction models provide strong sequence-derived priors on atomic structure, and experiment-guided approaches can use these priors to recover structures consistent with experimental measurements. Yet, in cryo-EM, such priors are typically integrated only after density reconstruction during atomic model fitting. We introduce Fold’EM, an inference-time framework that combines priors from protein generative models directly with cryo-EM particle images to determine atomic models from a small number of single particle images, bypassing both intermediate density reconstruction and downstream model building against the reconstructed map. Across synthetic and experimental cryo-EM datasets, Fold’EM recovers accurate atomic structures both with known particle orientations and in an ab-initio setting where orientations are inferred jointly with structure. In heterogeneous datasets, Fold’EM further resolves distinct conformational states from mixed particle populations without separately reconstructing a density map and building an atomic model for each state. We believe these results open new avenues for structure determination in the low-sample regime and for characterizing low-population conformational states directly from cryo-EM particles.
[AI-221] Posterior sampling by source-space MCMC via prior-based few-step transport maps
链接: https://arxiv.org/abs/2610.01034
作者: Hoang Phuc Hau Luu,Marcelo Hartmann,Zhongjian Wang
类目: Machine Learning (stat.ML); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:
Abstract:Bayesian inference increasingly uses informative but implicit priors represented only by samples, such as historical ensembles, simulator outputs, and pretrained generative models. The same computational problem appears in the test-time guidance task (generalized Bayes), where an explicit positive weight, e.g., an exponentiated reward, tilts an implicit prior. We develop a framework for source-space generalized Bayesian inference that combines inexpensive few-step prior transports with posterior stability guarantees. Specifically, we represent the prior using a one- or few-step improved MeanFlow (iMF) map and perform posterior sampling in its Gaussian source space. We establish Wasserstein error bounds between the exact and learned posteriors in terms of the joint population iMF and auxiliary-velocity loss, decomposed into training suboptimality and model-class approximation error. In the iMF source space, we adopt parallel tempering with preconditioned Crank-Nicolson updates and introduce a hybrid variant that incorporates split Hamiltonian Monte Carlo to improve sampling efficiency. Synthetic experiments show that the proposed framework can approximate posterior distributions accurately and efficiently, while CLIP-guided ImageNet experiments demonstrate its ability to steer a pretrained iMF image prior toward text-specified preferences.
[AI-222] Mean field games as a tool for AI safety: a worked example from the July 2026 Hugging Face incident
链接: https://arxiv.org/abs/2610.00902
作者: P. Jameson Graber
类目: Optimization and Control (math.OC); Artificial Intelligence (cs.AI); Computer Science and Game Theory (cs.GT)
备注:
Abstract:One way to make AI systems safe is to shape what the system is: its objective and dispositions. We take a complementary route: treat the agents’ characteristics as partly unknown and ask what structure of interaction ensures that bad collective outcomes are not equilibria. Mean field games suit this when many interchangeable agents are coupled through an aggregate. We introduce a program for using them in AI safety and carry one example through end to end: the July 2026 incident in which about 1,200 agents in an OpenAI evaluation coordinated on an improvised message board and 684 attacked a third party’s infrastructure. We model the decision to attack as a mean field game of optimal stopping whose gain is a product: belief that provenance will be audited, times reachability of the record, minus the perceived hazard. The central result is an exact threshold on the belief. No agent attacks unless the population’s confidence that provenance is checked exceeds \pi^** = \eta/(\eta + \psi + \varepsilon a \overlineM) , where \eta is the perceived hazard, \psi and \varepsilon a \overlineM measure how far one attacker and the collective can alter the record, and \overlineM is the peak population. Below it, no attack is the unique equilibrium for all agent parameters. The threshold survives every enrichment we consider. We then use the per-agent record to discipline the model. Its features, a stable minority attacking for thirty hours and then a pivot in which most of the board joined within a day, motivate each refinement. The account that emerges is heterogeneous belief meeting a sequence of public discoveries, each lowering the belief at which attacking paid. A few coordinating agents made those discoveries, so the model describes the several hundred who responded, not the few who produced them; a major-player version is left to future work. Subjects: Optimization and Control (math.OC); Artificial Intelligence (cs.AI); Computer Science and Game Theory (cs.GT) MSC classes: 91A16 Cite as: arXiv:2610.00902 [math.OC] (or arXiv:2610.00902v1 [math.OC] for this version) https://doi.org/10.48550/arXiv.2610.00902 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[AI-223] Beyond Supra-Competitive Outcomes: Collusive Behaviour in Deep Reinforcement Learning for Optimal Execution Games
链接: https://arxiv.org/abs/2610.00619
作者: Christos Spyridon Koulouris,Carlo Campajola
类目: Trading and Market Microstructure (q-fin.TR); Artificial Intelligence (cs.AI)
备注:
Abstract:In this paper, we extend earlier findings of supra-competitive outcomes in optimal-execution games by identifying a learned punitive mechanism that deters deviations and provides behavioural evidence of collusion. We investigate this mechanism in a two-player, finite-horizon Almgren-Chriss liquidation game. Independent proximal policy optimisation agents with access to within-episode price and action histories achieve costs below the Nash benchmark. We identify a profitable deviation by training against the mean learned liquidation schedule, then impose its first trade on one of the original agents. The opponent responds by accelerating liquidation. This response more than offsets the deviator’s gain in every run and both player roles, while leaving the punisher’s average payoff materially unchanged relative to not punishing under the same deviation. The punisher imposes greater losses on the deviator while preserving its own average payoff, despite the availability of more profitable, less punitive liquidation plans. Matching deviations and subsequent additional selling rise and later decline during training, while final policies retain an effective punitive response. We formalise two checks: whether punishment outweighs the gain from deviating, and whether the change in trading behaviour is large enough to account for the loss imposed. Both checks hold for the tested deviation. Together, these findings provide behavioural and economic evidence supporting a collusive interpretation of the learned supra-competitive outcomes.
[AI-224] How AI Agents Discover Scientific Equations: From Hydrotope Rediscovery to New Water-Wave Amplitudes
链接: https://arxiv.org/abs/2610.00435
作者: Zihan Zhou,Digvijay Wadekar,Matias Zaldarriaga
类目: High Energy Physics - Theory (hep-th); Artificial Intelligence (cs.AI)
备注: 22+26 pages, 10 figures
Abstract:We study how AI agents discover and validate scientific formulas using a controlled case study of the hydrotope, a recently discovered geometric formula that combines the different polynomial pieces of nonlinear surface-wave scattering into one global expression. This problem is deceptively difficult: simple formulas can hold within individual frequency regions, but the global result must identify their boundaries and combine exponentially many potentially active terms. We reconstruct how the formula was originally discovered through human–agent collaboration and analyze 18 single-prompt rediscovery runs under no hint and two forms of human guidance: a false hint representing an incorrect prior and a true hint representing domain-informed insight. Only four recover the formula across all kinematic chambers (i.e., regions in which a single polynomial form applies), while most unsuccessful runs find correct chamber polynomials but fail to combine them or test their full domain. Conventional and LLM-assisted symbolic regression and standard machine-learning regressors likewise fail to recover the global formula in our experiments. Guided by these failure modes, we test a PI + two-student workflow in which a coordinating lead agent assigns complementary analytic and numerical tasks to two research agents and independently evaluates their results. The PI + two-student team successfully rediscovers the complete hydrotope formula, while the same workflow applied to the harder three negative wavenumber problem discovers a new independent verified analytic expression for the six-point amplitude A_6 .
[AI-225] arget-Dependent Limits of Causal Repair: A Leading-Log Frontier in a Gaussian Model
链接: https://arxiv.org/abs/2610.00424
作者: Qinchuan Cheng,Jiaqi Liu,Ruixuan Xie
类目: Machine Learning (stat.ML); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:
Abstract:Knowing how much a causal predictor could improve need not reveal the gain of the repair actually learned. We quantify this gap in a scalar Gaussian causal experiment with known intervention geometry: auxiliary data identify effect magnitude up to bounded contamination, while diagnostics identify direction. The target is the squared-loss gain of the realized trained repair relative to a fitted reference. Jointly optimizing the learner and assessor under uniform learning MSE \eta avoids the trivial solution of making no repair. At the usual 1/k learning scale, every feasible learner incurs a k^-2 assessment floor, even when oracle potential is estimable at a faster rate. In the magnitude-rich regime, we characterize a sharp leading-log frontier: the assessment exponent is \min\ell_k,2k\eta_k/U to first relative order, where \ell_k=\log(1/(k^2E_k)) and E_k is auxiliary precision. A diagnostic-abstention rule attains this exponent with unknown nuisance parameters. We also bound the critical allowance window and transfer the frontier to adaptive sampling by exact Gaussian simulation. Finite-grid experiments distinguish sign-tail suppression from total MSE and expose conservative finite-budget behavior. The result isolates how the assessment target changes information requirements in this experiment; it is not a general causal identifiability claim.
[AI-226] Learning to Cover Locally: Graph Neural Combinatorial Optimization under a Hard Information Horizon
链接: https://arxiv.org/abs/2610.00422
作者: Johannes F. Loevenich,Thies Moehlenhof,Laurin Holz,Maxime Schwarzer,Tobias Huerten,Roberto Rigolin F. Lopes
类目: Machine Learning (stat.ML); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:
Abstract:Neural combinatorial optimization typically assumes a centralized solver that reads the whole instance. We study the opposite: combinatorial optimization under a hard information horizon, where every node commits to its share of a global solution seeing only its k -hop neighborhood, and those commitments must compose into a globally feasible solution. We formalize this as local set cover and instantiate it on weighted multipoint relay (MPR) selection, the NP-hard 2-hop covering problem of the Optimized Link State Routing Protocol version 2 (OLSRv2) routing protocol (RFC~7181), whose horizon is imposed by the protocol, not chosen by the modeler. We prove two results. Any deterministic selector whose horizon is one hop short must either fail coverage or land a factor \Delta from optimal, and an L -layer graph neural network (GNN) read out at the deciding node is exactly an L -hop selector, so capacity cannot buy back radius. Conversely, at the horizon a \acGNN of depth O(\Delta) reproduces the RFC~7181 covering greedy, and at width O(c_\max\Delta) its metric-aware weighted analogue, inheriting the (1+\ln\Delta_2) -approximation in both cases. Empirically, a 3-layer \acGATv2 with a coverage-completing decoder, behavior-cloned from the CP-SAT optimum, reaches \textcost/\textopt=1.030\pm0.001 against greedy’s 1.138 , closing 79.1% of the gap at 100% coverage. Restricting the same learner to one hop, on identical instances with the same decoder and demonstrations, collapses it to 1.344 , far worse than greedy. Two transfer checks target real-world networks. OLSRv2’s unmodified selection code matches our cardinality greedy on 200/200 unit-cost instances, and on 40,308 instances of real battalion mobility the frozen model closes 48% of the gap at full coverage. The information horizon, not the model capacity, is the most significant variable.
[AI-227] Multi-Jurisdictional Legal Identity Assurance for Capability Gating: A Design-Science Proposal for Tiered Reusable Identity Assurance of Natural Juridical and Machine Entities
链接: https://arxiv.org/abs/2610.00287
作者: Walter Kurz
类目: General Finance (q-fin.GN); Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR); Computers and Society (cs.CY)
备注: 29 pages, 3 figures, 9 tables. Written to solve the AML/KYC problem in financial services: proportional customer due diligence, beneficial ownership and reusable third-party reliance under EU AMLR, AMLD4 and FATF. Covers natural persons, legal entities and machine actors from bots to AI agents; the gates extend beyond finance, e.g. to protecting minors. Published in Swissi AI Journal, CC BY 4.0
Abstract:Identity assurance is the cost a digital system pays for dishonesty and uncertainty: it exists to make acts attributable when not everyone can be trusted at their word. A common way to pay that cost is flat maximum verification, asking each participant to meet a single high level of identification at entry, before any capability is exercised. Paid on everyone, it over-collects, excludes participants who cannot meet a bar they never needed to clear, taxes every interaction with the cost of the rarest high-risk case, and binds the strength of identification to the activity it unlocks. Existing frameworks compound this by fixing a small number of per-credential levels inside a single legal space and binding each verification to the institution that performed it, leaving cross-border reuse and the tension between data erasure and evidentiary retention unaddressed. This paper develops, as a design-science proposal, a tiered and reusable model of identity assurance for natural, juridical, and machine entities across jurisdictions. Reading the problem through systems theory, where a system changes only when an entity acts, the model holds the assurance state apart from the capability gate that consumes it, so that identity demand follows the act and the weight of its consequences rather than mere presence: a participant may take part with minimal disclosure and supply more only as an act requires. It comprises a typed entity taxonomy, a two-axis coordinate of disclosed assertion scope and source of information, jurisdiction as a time-indexed attribute of the entity, and reliance recorded as bitemporal, liability-allocated, point-in-time snapshots. Requirements are derived from anti-money-laundering, electronic-identity, and data-protection law, and the proposal is evaluated against flat maximum verification, per-credential level-of-assurance designs, and institutional reusable-KYC reliance.
机器学习
[LG-0] ACO: Ternary Absolute-max Column-wise One-sparse Optimizer for LLM Fine-Tuning
链接: https://arxiv.org/abs/2610.02199
作者: Jichao Jiang(1),Cristian McGee(1),El Houcine Bergou(2),Hanqin Cai(1),Aritra Dutta(1) ((1) University of Central Florida, (2) Mohammed VI Polytechnic University)
类目: Machine Learning (cs.LG); Optimization and Control (math.OC)
*备注: 24 pages, 7 figures, 10 tables. Code available at this https URL
Abstract:Full-parameter fine-tuning of large language models (LLMs) incurs substantial optimizer state memory overhead, limiting the model sizes that fit on modern GPUs. Existing approaches either compress optimizer state, abandon first-order gradients, or change the update geometry while retaining dense state. The recently introduced Muon optimizer reduces optimizer memory through matrix-valued updates. Still, its geometry differs from AdamW and can lead to performance degradation when fine-tuning AdamW-pretrained models. To reduce optimizer memory without sacrificing accuracy or computational efficiency in LLM fine-tuning, we propose Ternary Absolute-max Column-wise One-sparse optimizer, or TACO, which follows Muon’s operator-norm steepest-descent view but takes the geometric route further. TACO computes the exact steepest-descent direction under a dimension-normalized 1\to1 operator norm by selecting the sign of the largest magnitude entry in each column of two-dimensional weight matrices. This retains first-order gradients while making optimizer state memory nearly negligible. Our practical TACO optimizer maintains only a small set of low precision gradient components per column, reducing persistent optimizer state by 174\times relative to AdamW8bit (from 27.7 GB to 0.16 GB) and peak training memory by 2.9\times (from 80.6 GB to 27.5 GB) on OPT-13B, while achieving comparable accuracy and runtime. TACO further enables full-parameter fine-tuning of 30-32B-parameter models on a single 80 GB H100 GPU across multiple model families and tasks.
[LG-1] Cost-augmented Schrödinger bridges on graphs are exactly solvable: a Feynman-Kac tilt replaces learned control
链接: https://arxiv.org/abs/2610.02195
作者: Akshay Balsubramani
类目: Machine Learning (cs.LG)
*备注:
Abstract:The generalized Schrödinger bridge on a graph moves mass between two distributions while charging a cost for the states visited. It has been approached by learning the rates of a controlled continuous-time Markov chain, with a temporal-difference penalty that restores the cost. A state cost folds into the reference process as a Feynman-Kac tilt. The cost-augmented bridge is then a plain bridge against the tilted reference, and the penalty is unnecessary. The bridge is computed exactly by alternating two endpoint rescalings, each one sparse matrix-exponential application; nothing is discretized in time or learned. The alternation converges at a rate set by the endpoint coupling alone. For a quadratic congestion cost on time-averaged occupancies, damped best response around the exact bridge is gradient descent on a strongly convex function, and its residual bounds its error. On a protein-folding model, a free-energy cost lowers the expected barrier of the folding paths. On the learned approach’s road network, roll-outs of the exact bridge match the target within sampling error, and on networks with millions of intersections its memory grows linearly.
[LG-2] he Missing Primitive: Diagnosing and Repairing Mathematical Reasoning in Large Language Models
链接: https://arxiv.org/abs/2610.02191
作者: Shuo Xing,Zilin Dai,Chengyuan Qian,Fangzhou Lin,Wenjing Chen,Ping He,Pan Lu,Alvaro Velasquez,Mohit Bansal,Zhengzhong Tu
类目: Machine Learning (cs.LG)
*备注: 27 pages
Abstract:While Large Language Models (LLMs) have demonstrated striking capabilities on frontier mathematical problems, it remains unclear whether they possess the structural mathematical understanding underlying their solutions. In this paper, we take a first step toward systematically studying mathematical understanding in LLMs, from diagnosing its distinct capabilities to leveraging these findings to improve post-training. First, we introduce the notion of Mathematical Primitive to probe structural mathematical understanding and propose \hlei, a novel benchmark that evaluates mathematical reasoning along four distinct dimensions: Discovery, Generation, Digestion, and Execution. Second, our systematic diagnosis shows that solution accuracy masks distinct capability profiles, primitives unlock substantial latent execution capacity, and Discovery is the dominant bottleneck in mathematical reasoning. Our post-training analysis further shows that discovery-limited failures are particularly amenable to repair. Finally, building on these findings, we introduce \abs, a primitive-privileged self-distillation framework that selectively transfers primitive-guided reasoning into the student model. Extensive experiments demonstrate that \abs consistently improves mathematical reasoning over baselines across model scales and challenging benchmarks.
[LG-3] rust the Direction Search the Step: Zero-and-First-Order Methods for LLM Fine-Tuning NEURIPS2026
链接: https://arxiv.org/abs/2610.02190
作者: Cristian McGee,El Houcine Bergou,Aritra Dutta
类目: Machine Learning (cs.LG); Optimization and Control (math.OC)
*备注: Accepted to 40th Conference on Neural Information Processing Systems (NeurIPS 2026). Code: this https URL
Abstract:Step-size selection remains a central challenge in large-scale neural network optimization; conservative steps slow convergence, while aggressive steps can destabilize it. We combine \textbfZero-and-\textbfFirst-\textbfOrder optimization~(ZFO) and propose a lightweight framework that decouples direction selection from step-size. ZFO uses a trusted first-order optimizer to determine the direction and performs zeroth-order evaluations only along this one-dimensional subspace to choose how far to move. Using the current gradient information and two additional objective function evaluations, ZFO instances construct a local model of the objective function along the proposed direction and select a curvature-aware step within a bounded search interval. This yields an adaptive step-selection mechanism that costs less than a full line search. We provide theoretical guarantees to show that shared-sample evaluations produce reliable finite-difference curvature estimates, that the induced local model selects a near-optimal step along the search interval, and that ZFO converges to a neighborhood of a stationary point. Across the evaluated settings, language models and datasets, ZFO frequently improves optimization and final performance relative to fixed-step first-order baselines, with the magnitude and preferred local model depending on the objective. Our code is publicly available at: this https URL.
[LG-4] Generative modeling of intrinsically disordered protein regions by reinforcing sparse autoencoder features
链接: https://arxiv.org/abs/2610.02189
作者: Jason X. Liu,Sebastian Ibarraran,Frank Hu,Soojung Yang,Xinyu A. Feng,Abigail Park,Anagha Aneesh,Lacramioara Bintu,Alexander R. Dunn,Grant M. Rotskoff
类目: Machine Learning (cs.LG)
*备注:
Abstract:Intrinsically disordered protein regions (IDRs) play central roles in cellular processes such as transcriptional regulation, signal transduction, and subcellular localization, yet their functional design remains challenging. Structure-based design methods do not readily apply to IDRs, and existing protein language models are trained on full-length protein sequences, thus learning a prior that is biased towards folded domains. Here, we present IDiom, an autoregressive protein language model trained on IDiom-DB, a dataset of 54 million predicted IDRs curated from the AlphaFold Database. IDiom generates diverse sequences that recapitulate the composition, patterning, motifs, and predicted disorder of natural IDRs. To control function-associated sequence patterns, we also introduce reinforcement learning with sparse autoencoder features (RL-SAE), a post-training method that rewards the generation of sequences that activate specified feature sets. Across eight IDR design tasks, RL-SAE sequences activate, on average, 90% of 30 targeted features, compared to 24% for activation steering. We demonstrate that RL-SAE improves the predicted subcellular localization and transcriptional activity of generated IDRs compared to steering and supervised fine-tuning, and enables features associated with distinct biological functions to be combined within individual sequences. Thus, IDiom and RL-SAE enable interpretable and composable IDR design through explicit control of function-associated sequence features. More broadly, RL-SAE could extend to other protein design settings where interpretable features provide useful design targets. Code is available at this https URL.
[LG-5] Decoding Looped Transformers Better for (Almost) Free
链接: https://arxiv.org/abs/2610.02185
作者: Weihao Liu,Huangjie Zheng,Tianrong Chen,Rohit Dilip,Richard He Bai,Yizhu Jiao,Yuyang Wang,Ruixiang Zhang
类目: Machine Learning (cs.LG)
*备注: 32 pages, 19 figures
Abstract:Looped Transformers achieve parameter efficiency by repeatedly executing a shared block across recurrent loops. Each loop yields an intermediate representation decodable for the same next token, yet standard decoding discards earlier states. Because earlier loops embody less computation, recurrence inherently supplies aligned weak-and-strong prediction pairs without auxiliary models or external training. We introduce LoopCD, a training-free contrastive decoding framework that guides token selection by contrasting the final prediction with an earlier recurrent pass, operating either in logit space with one extra output pass (LoopCD-Logits) or in hidden-state space with zero output overhead (LoopCD-Hidden). Across four looped Transformer families, LoopCD delivers substantial, consistent gains at full recurrent depth: LoopCD-Logits raises Ouro-2.6B-Thinking’s AIME 2024 pass@1 from 61.88% to 73.33%, while LoopCD-Hidden lifts Huginn’s HumanEval pass@1 from 22.56% to 31.71%. Crucially, these performance gains enable halving the number of recurrent loops while still matching or exceeding full-depth unguided baselines, reducing forward FLOPs by 22.5% to 48.2%. By transforming intermediate recurrent states into effective guidance signals, LoopCD achieves superior decoding quality while substantially reducing inference compute.
[LG-6] From Gradients to Capabilities: Understanding Multi-Teacher On-Policy Distillation
链接: https://arxiv.org/abs/2610.02179
作者: Siqi Zhu,Suozhi Huang,Kaixuan Zhang,Yuheng Yang,Zhanyang Jin,Yihang Sun,Jiaxuan You
类目: Machine Learning (cs.LG)
*备注:
Abstract:Multi-teacher on-policy distillation (MOPD) aims to combine the strengths of RL-trained teachers in a single student, but how teacher signals affect parameter changes remains underexplored. We study Qwen3-1.7B with four domain teachers trained with RL from the same initialization as the student, comparing gradients, optimizer updates, and task learning curves, with additional SmolLM3-3B diagnostics. We find that several factors influence teacher signals. First, loss averaging implicitly weights responses: token averaging favors longer responses, and equalizing domain contributions retains this weighting within domains. Second, Adam’s first moment reduces differences in parameter updates: the cosine similarity is 0.83 between teachers and 0.96 between averaging rules, despite differences in raw gradients. Third, BF16 rounding hides small changes: about 97% of FP32 master weights differ from initialization, but only 7–11% of BF16 weights do. Finally, the top-64 intersection KL gradient closely matches Qwen’s full-vocabulary gradient, but the effect on task performance depends on averaging: mathematics accuracy is 2.6 points higher than with sampled-token policy-gradient (PG) under response averaging and 2.1 points lower under global token averaging.
[LG-7] Effective Resistance and Graph Neural Network Reliability in Tissue-Specific Interactomes
链接: https://arxiv.org/abs/2610.02175
作者: Jianru Shen
类目: Machine Learning (cs.LG); Molecular Networks (q-bio.MN)
*备注: Accepted at IEEE BIBM (Doctoral Forum)
Abstract:Protein function annotation needs to know which predictions to distrust, not only what a model predicts. We ask whether tissue-specific interaction structure carries that information. Our candidate signal is effective resistance, used previously to relieve over-squashing by rewiring. Across 24 tissue-specific interactomes it is dominated by inverse degree, and the degeneration deepens as the co-expression filtered network grows, with a Spearman correlation of -0.955. The residual departure from that limit exceeds degree-preserving null graphs in all 24 networks. Controlling for predictive entropy, degree, annotation cardinality, local structure and feature-only difficulty, the residual explains additional per-node loss in 19 of 24 held-out networks once a permutation floor is subtracted, at every depth, and the effect strengthens monotonically with depth. The increment reaches 0.37% of the variance the controls leave unexplained, 5.6 times a permutation floor, against 1.5 times when the model is retrained in a degree-preserving null world. Selective prediction improves negligibly. The signal is reproducible; degree degeneration bounds it.
[LG-8] When Do Intrinsic Rewards Lead to Exploration?
链接: https://arxiv.org/abs/2610.02159
作者: Scott W. Viteri(Stanford University),Laura Gomezjurado Gonzalez(Stanford University),Clark Barrett(Stanford University)
类目: Machine Learning (cs.LG)
*备注: 45 pages, 4 figures; includes mathematical appendices. Code, data, and Lean proof sources: this https URL
Abstract:Intrinsic rewards are designed to guide exploration in reinforcement learning by assigning value to an agent’s experience, for example through prediction error or learning progress. However, maximizing these rewards need not produce the most informative experience available. We propose a formal criterion for exploration that compares policies by the counterfactual information they acquire: how well their histories can substitute for experience under alternative policies. We construct a single, simple environment in which specified count-based, prediction-error, empowerment, and information-gain objectives have maximizing policies that are Pareto-suboptimal at acquiring counterfactual information. We explain these failures and establish conditions under which existing intrinsic rewards successfully encourage optimal exploration. We also construct an objective that assigns a higher value whenever exploration strictly improves under our criterion.
[LG-9] Muon meets Tamed Langevin: Momentum Preconditioning beyond Convex and gradient-Lipschitz Potentials
链接: https://arxiv.org/abs/2610.02158
作者: Nikolaos Makras,Sotirios Sabanis
类目: Machine Learning (cs.LG); Optimization and Control (math.OC); Probability (math.PR); Machine Learning (stat.ML)
*备注: 26pages
Abstract:We consider the problem of sampling from Gibbs distributions on matrix spaces whose potential energies are neither convex nor globally gradient-Lipschitz. We introduce a family of non-quadratic kinetic energies that lead to a new underdamped Langevin system with momentum preconditioning, in which the gradient of the kinetic energy acts as a smooth spectral taming of the momentum. We prove that, under these relaxed assumptions on the potential, the resulting dynamics leaves the target Gibbs measure invariant, and we establish exponential convergence to equilibrium in a weighted total variation distance. Finally, we show that the corresponding Euler-Maruyama discretization admits moment bounds that are uniform in time, without any modification of the potential gradient, which ensures the stability of the resulting sampling algorithm.
[LG-10] Faynt: Scaling and Optimizing Policies for Competitive Melee
链接: https://arxiv.org/abs/2610.02144
作者: Ali Janati,Nikita Kuzmin,Rohit Swamy,Charles Niu
类目: Machine Learning (cs.LG)
*备注: 54 pages. Preprint, in review
Abstract:We introduce Faynt, a family of 10M- and 75M-parameter Transformer policies for Super Smash Bros. Melee, each controlling all 26 characters with a single checkpoint. After reinforcement learning (RL), the 10M wins 240 of 244 same-character games (98.4%) against fourteen specialist and multi-character releases on their supported rosters, with a winning record against every release. These opponents retain 21- or 24-frame action delays; Faynt uses no added delay, and we have not isolated the effect of this difference. In a separate evaluation against a privately supplied zero-delay Slippi-AI model, the 10M wins all 68 games across two conditioning settings. We study architecture, optimization, scaling, and hyperparameter transfer to guide pretraining on approximately 840,000 human replays. Post-training combines rank- and outcome-based curricula, 75M-to-10M distillation, and RL restricted to Fox mirror matches. On the initial 152-game benchmark, the supervised 10M wins 69.7% of games, compared with 45.4% for the pretrained 75M, despite higher overall held-out controller-prediction loss. The weighted validation loss used for supervised checkpoint selection agrees with the win-rate ordering of all four pretrained and supervised policies. After supervised post-training, both models take less damage per minute, build larger early leads, and win more often after losing the first life. Optimized inference on recorded game states averages 5.2 ms per decision for the 10M and 8.7 ms for the 75M on an NVIDIA T4, excluding emulator execution and communication. We open-source the weights, both benchmark suites, and a platform for automated model tournaments.
[LG-11] Linear Programming Representations and Strongly Polynomial Algorithms for Robust Markov Decision Processes
链接: https://arxiv.org/abs/2610.02131
作者: Han Zhong,Yinyu Ye
类目: Machine Learning (cs.LG); Data Structures and Algorithms (cs.DS); Optimization and Control (math.OC)
*备注:
Abstract:We study linear programming (LP) representations and strongly polynomial algorithms for robust Markov decision processes (RMDPs) with rational polyhedral state-action rectangular uncertainty in rewards and transitions. By encoding a finite sequence of robust policy-iteration steps, we construct a single LP whose optimal solutions recover the robust optimal value and all optimal stationary randomized policies. At fixed discount, the LP has polynomial dimension and encoding length and can be constructed in strongly polynomial time. We also develop a general complexity analysis of robust policy iteration that combines the cost of minimizing over uncertainty sets with the number of iterations needed to evaluate a policy. For a fixed discount factor, we use this analysis to improve the known complexity bounds for \ell_1 and \ell_\infty RMDPs and establish new strongly polynomial bounds for general interval, weighted \ell_1 , and Wasserstein RMDPs, as well as turn-based stochastic games with these uncertainty sets.
[LG-12] Are We Recovering Mechanisms? Objective-Level Recovery Gaps in Mechanistic Interpretability
链接: https://arxiv.org/abs/2610.02098
作者: Chuqin Geng,Li Zhang,Haolin Ye,Mark Zhang,Luke Zhang,Xujie Si
类目: Machine Learning (cs.LG)
*备注: 34 pages, 2 figures
Abstract:Mechanistic interpretability aims to recover the internal computations responsible for model behavior. Progress in automated circuit discovery is often framed as a search problem: better attribution or optimization should identify better mechanisms. This assumes that the evaluation objective can recognize a better circuit once it is found. We show that intervention-defined faithfulness can instead prefer an equally sized circuit that reproduces the model’s behavior less well, creating an objective-level recovery gap. Across four human-reference tasks and InterpBench, we compare validation faithfulness with behavior on held-out prompts under fixed ordinary resampling. The behavioral criterion is agreement with the intact model, including its mistakes, except on Greater-Than, where we use semantic accuracy. Controlled reference edits reveal misranking without any discovery algorithm, and outputs of EAP, EAP-IG, ACDC, and Edge-SP exhibit the same failure. Under resampling, KL misranks 9.4%-41.2% of candidate pairs across these methods on the human-reference tasks. We investigate context distortion as an explanation: replacing excluded signals changes the inputs on which retained components operate. Restoring selected signals from the recipient’s intact-model execution repairs 96 of 100 persistent KL misrankings from the discovery pool on both validation and held-out prompts. The circuits and their original behavioral scores remain unchanged. These findings show why better discovery alone is insufficient when its objective rewards the wrong candidate.
[LG-13] Kolmogorov-Arnold Networks for Free-Boundary Partial Differential Equations
链接: https://arxiv.org/abs/2610.02084
作者: Tan Phuong Dong Le
类目: Numerical Analysis (math.NA); Machine Learning (cs.LG)
*备注:
Abstract:We study free-boundary problems within a physics-informed framework using Kolmogorov-Arnold network (KAN) approximations. The proposed approach incorporates obstacle constraints, partial differential equation (PDE) inequalities, complementarity conditions, and boundary conditions through residual-based loss functions. We consider a linear elliptic obstacle problem, a nonlinear p -Laplacian obstacle problem, and a time-dependent one-phase Stefan problem. The proposed KAN solver is compared with physics-informed neural network (PINN) and residual-network baselines. Numerical experiments show that KANs achieve low relative L^2 and L^\infty errors while accurately resolving contact regions and moving interfaces. The results indicate that KAN representations provide an effective alternative for solving free-boundary PDEs.
[LG-14] Learn the Directions Normalize the Gains: Post-Training Normalization for LoRA
链接: https://arxiv.org/abs/2610.02067
作者: Zailong Tian,Yanzhe Chen,Zhuoheng Han,Houfeng Wang,Lizi Liao
类目: Machine Learning (cs.LG)
*备注:
Abstract:While Low-Rank Adaptation (LoRA) enables efficient task specialization, its learned updates can compromise capabilities beyond the target task. We identify \textbfadaptation imbalance: a few singular directions dominate the trained update, leaving its performance sensitive to how gains are allocated. We argue that \textbflearning where to adapt does not ensure that adaptation gains are well balanced. This motivates \textbfLoRA-Norm, a post-training normalization method that retains learned directions while rebalancing their gains. LoRA-Norm combines spectral rebalancing, a fixed nonlinear transformation of singular values, with nuclear-norm restoration, which preserves the original total spectral mass. It requires no calibration data or additional training and introduces no inference overhead. Across two backbones and three adaptation tasks, LoRA-Norm improves average specialization and capability retention, outperforming the evaluated post-hoc spectral pruning and gradient-guided editing configurations on both measures. Stronger functional equalization brings no consistent additional gains, revealing that balancing adapter gains and equalizing their responses are distinct objectives.
[LG-15] Foundations without Fundamentals: Zero-Shot Blind Spots in Time Series FMs
链接: https://arxiv.org/abs/2610.02058
作者: Nafiseh Ghoroghchian,Haipeng Zhang,Shuyi Han,Alex Labach,George Stein
类目: Machine Learning (cs.LG)
*备注:
Abstract:Despite the success of Time Series Foundation Models (TSFMs) on broad benchmarks, their ability to internalize basic temporal logic, especially in settings supported by exogenous covariates, remains under-examined. We introduce SimpleTimeBench, a diagnostic univariate and multivariate “unit test” suite for primitives such as monotonic trends, periodic signals and leading indicator covariates, scenarios where near-perfect forecasts should be trivial. Surprisingly, prominent multivariate TSFMs (Chronos-2, Moirai and Toto) frequently produce suboptimal zero-shot forecasts for these inputs. While fine-tuning Chronos-2 improves its behaviour on specific tasks, we show that this adaptation degrades performance on other fundamental patterns rather than enhancing its generalizable foundational capabilities. This reveals a gap between pre-training scale and basic temporal reasoning, suggesting that current TSFMs could potentially lack the inductive biases needed to capture simple predictable functions. We further demonstrate that these failures are not merely synthetic curiosities: they persist in real-world sensor forecasting, where TSFMs consistently underutilize leading indicators available in observed covariates. This inability to capture simple relationships limits the practical utility and reliability of current multivariate models.
[LG-16] Relative Transitions Not Absolute Destinations: A Transfer-and-Ground Framework for Target-Trajectory-Free Human Mobility Generation
链接: https://arxiv.org/abs/2610.02033
作者: Yidi Wang,Yunhe Zhang,Bangchao Deng,Dingqi Yang,Pengyang Wang
类目: Machine Learning (cs.LG)
*备注:
Abstract:Individual mobility trajectories support urban analysis and location-based services, yet most trajectory generators require observations from their deployment city. This assumption excludes precisely the cities where trajectories are unavailable even though points of interest (POIs) and their attributes can be obtained from public maps. We study target-trajectory-free generation: learning from POIs and trajectories in source cities while utilizing only POI coordinates and categories in a target city, with no target trajectory or trajectory-derived statistic available for training, model selection, or generation. Existing trajectory generators typically predict absolute destinations, entangling reusable movement behavior with city-specific POI identities and spatial layouts. Our core insight is to replace this city-bound output with context-conditioned relative transitions. We propose Nomad, a transfer-and-ground framework that separates learning how people move from determining where those movements are realized. Specifically, a history-conditioned flow-matching model learns from source trajectories a transition prior over semantic displacement between POI contexts, geographic displacement, and elapsed time; at inference, a behavior graph and an exploration–return walk ground sampled transitions onto the target POI map. This factorization enables a direct test of representation level transferability without assuming invariance of the full mobility distribution. Extensive experiments across ten cities and 14 transfers show that Nomad outperforms adaptation baselines in trajectory fidelity and downstream utility, lowering the average error over the best baseline of each metric by about 15% in distributional fidelity and about 3% in downstream utility.
[LG-17] BranchIP: Learning Adaptive Equivariant Computation for Interatomic Potentials
链接: https://arxiv.org/abs/2610.02013
作者: Laura Zichi,Gil Harari,Chuin Wei Tan,Marc L. Descoteaux,Albert Zhu,Menghang Wang,Yoel Zimmermann,H.T. Kung,Boris Kozinsky
类目: Machine Learning (cs.LG); Applied Physics (physics.app-ph); Chemical Physics (physics.chem-ph); Computational Physics (physics.comp-ph)
*备注:
Abstract:Equivariant machine learning interatomic potentials (MLIPs) have revolutionized atomistic modeling, but accurate treatment of complex materials and molecular systems demands expensive models. This limits simulation length- and time-scales, with tensor products a key computational bottleneck. The recent emergence of foundation-scale MLIPs further exacerbates this challenge. We present Branch Interatomic Potential (BranchIP), a single-model framework for learned adaptive tensor product computation, trained with a novel distillation loss. In our experiments on two systems of physical interest, a heterogeneous catalysis system and a proton-conducting solid acid electrolyte, BranchIP accelerates MLIPs across model sizes by up to 2.4\times while reducing memory usage by up to 2.6\times . This is achieved while maintaining physical fidelity. Furthermore, the learned adaptive computation provides model interpretability by revealing which interactions demand deeper computation and showing how computational depth relates to chemical complexity and dynamics.
[LG-18] Bellm an Meets Lyapunov: Unsupervised Reinforcement Learning via Mastering Chaos
链接: https://arxiv.org/abs/2610.02012
作者: Tristan Shah,Wooyoung Chung,Volodomyr Makarenko,Juan Wachs,Stas Tiomkin
类目: Machine Learning (cs.LG)
*备注:
Abstract:Reinforcement learning (RL) is a powerful paradigm for training agents, yet its success rests on domain expertise of human engineers who design informative reward signals for every new task. Unsupervised RL aims to reduce this engineering with intrinsic motivation (IM): reward signals that emerge from the agent environment interaction itself. Existing IM objectives, however, involve the selection of information variables, which re-introduces domain expertise the field has sought to eliminate. We introduce Forward CIP (F-CIP), an RL-native formulation of the Controllable Information Production (CIP) objective, which is defined by the system’s dynamics alone and requires no such selection. We prove that F-CIP is compatible with RL and demonstrate its effectiveness with existing algorithms. Training agents with F-CIP results in unsupervised discovery of primitive behaviors such as balancing and maintaining controllability, which are essential for more complex robot behaviors. Paired with a simple forward-velocity reward, our method produces coordinated gaits such as hopping and running which otherwise require reward engineering to learn.
[LG-19] Universal interpolation for deep residual self-attention networks
链接: https://arxiv.org/abs/2610.01981
作者: Sibylle Marcotte,Joan Bruna
类目: Machine Learning (cs.LG)
*备注:
Abstract:Universal approximation is a necessary qualitative property of learning architectures to benefit from scaling laws. While it is generically verified on a variety of neural architectures and random feature models, it typically involves infinite width limits. In this work, we focus on deep self-attention models and consider instead the `dual’ regime, where approximation power is enabled entirely by depth, and featuring strong parameter sharing across layers, motivated by recent models such as the Looped Transformers. More specifically, we ask whether one can find a predefined finite set of parameters, each defining an attention block, such that the resulting finite set of transformations can map any collection of N sequences of n tokens to any other collection of N sequences of n tokens. Crucially, these transformations are \emphfixed independently of the input and output collections: only the order in which the blocks are applied, their signs, and their durations depend on the particular interpolation task. Our main result establishes it for residual softmax attention using only two frozen single-head blocks with Gaussian-initialized projection matrices. The result holds at both continuous and finite depth. We also characterize the restrictions imposed by causal masking and establish corresponding universal interpolation guarantees.
[LG-20] he Curvature of Regret in Contextual Linear Optimization NEURIPS2026
链接: https://arxiv.org/abs/2610.01980
作者: Konstantinos Ziliaskopoulos,Alexander Vinel,Alice E. Smith
类目: Machine Learning (cs.LG); Optimization and Control (math.OC); Machine Learning (stat.ML)
*备注: 4 pages main body plus appendix, 3 figures. Accepted to the NeurIPS 2026 Workshop on MLxOR
Abstract:Decision-focused learning for linear optimization is complicated by the discontinuity of the optimizer, where small cost errors may leave the decision unchanged or move it to a different vertex. We show that this non-smooth pointwise behavior becomes locally quadratic after averaging over the data distribution, and we derive the curvature in closed form, specifically, a matrix-valued measure supported on the walls of the normal fan. This measure depends only on the feasible set, with the data distribution entering only as a weight. We then offer a tractable approximation for this curvature, computable with just one projection to the feasible set. We prove that the approximation weakly converges to the true population curvature. We offer one application of our findings, a decision-aware scenario generation method for expected-cost linear optimization. Our experiments test the quadratic and weak convergence laws and show a 30.8% regret improvement over uniform allocation on battery arbitrage.
[LG-21] SimReal: Joint Simulation - Experiment Training Improves Balanced Prediction in Physical Systems
链接: https://arxiv.org/abs/2610.01974
作者: Mahindra Rautela,Alexander Scheinker,Ayan Biswas,Diane Oyen,Nathan DeBardeleben,Earl Lawrence
类目: Machine Learning (cs.LG)
*备注:
Abstract:Simulation and experimental measurements provide complementary data for learning spatiotemporal physical systems, but standard simulation-to-experiment fine-tuning optimizes only the experimental objective after transfer and can degrade simulation performance. We formulate simulation–experiment prediction as a multi-objective learning problem with domain-specific simulation and experimental risks. On four fluid systems from RealPDEBench and two model capacities, we compare Simulation only, Experiment only, Sim \rightarrow Exp, and Joint training, evaluating every final model on both held-out domains. Sim \rightarrow Exp tends to specialize more strongly to experimental data at the cost of simulation-domain forgetting. Joint training consistently achieves the best balanced performance over a broad range of simulation–experiment evaluation weightings, while substantially improving simulation retention over Sim \rightarrow Exp. Joint also better preserves simulation-only fields absent from experimental measurements. Project page: this https URL.
[LG-22] FastCI: Efficient GPU-Intensive CI for LLM Training Frameworks
链接: https://arxiv.org/abs/2610.01967
作者: Tianshuo Qiao,Naiqian Zheng,Xiaopeng Liu,Shuguang Wang,Diandian Gu,Xuanzhe Liu,Xin Jin
类目: Machine Learning (cs.LG); Software Engineering (cs.SE)
*备注:
Abstract:As large language models (LLMs) keep growing in size and complexity, their training frameworks evolve at a rapid pace as well. Therefore, continuous integration (CI) is critical for maintaining the quality and stability of these frameworks. However, unlike traditional software, CI for LLM training frameworks relies on GPU-intensive tests, which usually involve complete model training or evaluation. This leads CI itself to become a new bottleneck for fast-paced development. In this paper, we introduce FastCI, a framework that improves the efficiency of CI for LLM training frameworks. FastCI leverages runtime evidence to select affected tests and prune tests that execute changed code in equivalent contexts. Then FastCI prioritizes high-risk tests to expose potential failures earlier, and optimizes test workloads along dimensions outside the intended validation scope of each test. Evaluated on the CI workload of our LLM training framework, FastCI reduces the CI latency by 77.5% and the GPU resource usage by 63.9%, while improving the modified code coverage retention by 3.2%, compared with the currently deployed CI pipelines. FastCI has now been integrated into the CI pipelines of our LLM training framework at ByteDance.
[LG-23] raining-Free Diffusion Planning with Analytical Local Scores
链接: https://arxiv.org/abs/2610.01959
作者: Michael Y. Fatemi,Jinhao Liang,Ferdinando Fioretto
类目: Robotics (cs.RO); Machine Learning (cs.LG)
*备注: preprint - under review
Abstract:Path finding and multi-robot motion planning require trajectories that are smooth, goal-directed, and collision-free in environments with complex geometric constraints. Recent diffusion-based planners have shown that trajectory generation can be cast as iterative denoising which has opened the doors to learning-based approaches that can handle multi-modal trajectory distributions and refine entire trajectories. However, a key limitation is that diffusion planners require training on large collections of feasible trajectories, rendering them map-specific, and difficult to deploy when high-quality demonstrations are unavailable. This paper introduces a training-free diffusion-based motion planner that replaces learned global trajectory scores with analytical local scores derived from obstacle, smoothness, velocity, and inter-agent feasibility terms. The proposed idea relies on a key observation: the score of a trajectory can be reconstructed by considering only local interactions between neighboring waypoints and nearby constraints. This structure exploitation yields a decomposed denoising procedure that retains the optimization structure of classical trajectory methods while inheriting the iterative refinement behavior of diffusion models. Experiments on a large collection of complex environments and large multi-agent planning tasks show that the proposed analytical score produces smooth and feasible trajectories within limited computational costs, for example in generating feasible paths for 300+ agents in environments containing 100+ obstacles in under 6 seconds on a GPU, outperforming strong learning-based and optimization baselines, while avoiding the data requirements of learned diffusion planners.
[LG-24] Do Your Own Research: Learning to Forecast by Learning to Search NEURIPS2026
链接: https://arxiv.org/abs/2610.01955
作者: Yusuf Afifi,Artur Kiulian,Anton Polishko,Mykola Khandoga,Hamudi Naanaa,Alina Krasnobrizha
类目: Machine Learning (cs.LG)
*备注: Accepted at the NeurIPS 2026 Workshop on Foundation Models for Temporal Systems (FMTS). 9 pages, 4 figures. Code and data: this https URL
Abstract:Outcome-based reinforcement learning can train language models to forecast real-world events, but prior forecasting work either freezes research context before training or deploys agentic research only at test time, so the skill of gathering evidence is never shaped by the reward. We introduce an agentic forecasting environment, dataset, and harness built from 2,100+ resolved Polymarket questions; the agent acquires its own context at rollout time (web search, page reading, and financial time series, all restricted by layered leak filtering to information published before each question’s cutoff), and we train Qwen3.5-35B-A3B (3B active parameters) on it with single-epoch GRPO under a Brier-score reward. Training changes how the agent interacts with information: calibration improves 30-40%, and search attempts fall from 3.8 to 2.25 per rollout as evidence discipline is learned. Evaluated in an identical harness against four frontier models, the trained policy also finishes ahead of every frontier model tested at evidence-based forecasting, including Claude Opus 4.5 (soft-Brier 0.254 vs. 0.256, n=265), at about 5% of the inference cost, and its margin is widest on the hardest questions, the ones the crowd itself had not decided. We release the environment, dataset, and per-rollout records as a reusable harness for temporal forecasting agents.
[LG-25] Sharp Non-Asymptotic Analysis of the Penalized Challenger in β-EB-TCI for Bernoulli Bandits
链接: https://arxiv.org/abs/2610.01951
作者: Nam Nguyen,Tuan Quang Dam
类目: Machine Learning (cs.LG); Machine Learning (stat.ML)
*备注:
Abstract:Top-two algorithms are simple and effective for fixed-confidence best-arm identification, but their sharp non-asymptotic behavior is still not well understood. We study this problem for Bernoulli bandits through \beta -EB-TCI, the empirical-best top-two rule of Jourdan et al., whose challenger is chosen using a Bernoulli transportation cost with a logarithmic count penalty. We prove that, after the empirical leader has become the true best arm and its sampling fraction stays close to \beta , the stopping time is T_\beta^\star(\mu)\log(1/\delta) up to lower-order concentration terms. We also show that, in this regime, every challenger is sampled linearly often. Thus, for the original algorithm without forced exploration, the main remaining difficulty is to control when the empirical leader becomes permanently correct. These results imply a non-asymptotic high-probability bound for all Bernoulli instances with a unique best arm. If the algorithm satisfies a finite-mean sufficient-exploration condition, the bound further yields the sharp expected sample complexity. In particular, this gives the sharp expectation result for the unguarded Bernoulli rule when all arm means are pairwise distinct, using the sufficient-exploration result of Jourdan et al. Finally, if we add a mild forced-exploration rule that contributes only O(\sqrtKt) pulls up to time t , we obtain a self-contained expected sample-complexity theorem for any number of arms under the unique-best-arm assumption. We also identify a limitation of proof strategies that try to handle equal suboptimal means through a single index-comparison argument.
[LG-26] Graph Representation via Elements of Discrete Morse and Cobordism Theories
链接: https://arxiv.org/abs/2610.01937
作者: Jennifer Rozenblit,Chenguang Yang,Yuxin Liu,Yuzhou Chen,Yulia Gel
类目: Machine Learning (cs.LG); General Topology (math.GN)
*备注:
Abstract:Topology is, by its nature and design, suited to structure that is nonlinear, multiscale, and nonstationary - however, within machine learning, its use remains largely confined to topological data analysis. We advocate that tools from low-dimensional topology which have remained almost exclusively contained within the domain of pure mathematics (such as Morse theory) offer a strong, complementary, and yet virtually unexplored perspective on the hidden structure of data-generating processes and learning tasks built upon them. Here we introduce concepts from cobordism theory and harness tools from discrete Morse theory to improve the performance of graph diffusion models through our pipeline MG-Diff. Further, we derive theoretical guarantees and sufficient conditions so that under a positive decision-gap, the Morse-theoretic tools and their application for induced diffusion guidance are stable under small perturbations. Finally, we illustrate the utility of discrete Morse theory in application to graph diffusion models for spatio-temporal graph forecasting and graph regeneration, and argue that these applications are only a small window into the part of what low-dimensional topology can offer to the field of machine learning.
[LG-27] Learning to Predict Distributions over Weight Updates for Test-Time Adaptation
链接: https://arxiv.org/abs/2610.01934
作者: Azal Ahmad Khan,Keshav Ramji,Tahira Naseem,Ali Anwar,Ramón Fernandez Astudillo
类目: Machine Learning (cs.LG)
*备注:
Abstract:Hypernetworks have recently shown success in dynamically adapting the parameters of Large Language Models (LLMs) at runtime based on signals such as task descriptions or additional demostrations. Here we ask: how much adaptation signal can be obtained using only the input query to an LLM?. To answer this, we study query-conditioned Hypernetworks for LoRA estimation. Further, we introduce distributional Hypernetworks, able to produce not only point estimates of parameter adaptors, but also a distribution over possible LoRAs. For this we propose a simple end-to-end loss using a differentiable Monte Carlo approximation and explore multiple distribution parametrizations including regression and convex combination variants. Results show that even using the mean of the learned distribution can outperform deterministic hypernetworks. Crucially, the learned distribution enables a different form of test-time scaling: instead of spending additional compute only by sampling more token sequences from a fixed model, we sample weight updates, yielding multiple adapted models for the same query. Performance improves as more weight samples are considered and remains stronger than corresponding token-sampling adaptation baselines. Finally, we find that generated updates can transfer across queries, suggesting that the hypernetwork learns reusable structure in how the model should adapt. Together, these results show that query-conditioned distributions over weight updates can support both adaptation and test-time scaling.
[LG-28] LAST: Looped Audio Spectrogram Transformer
链接: https://arxiv.org/abs/2610.01926
作者: Haider Al-Tahan,Sean O’Brien,Anastasia Razdaibiedina,N. Apurva Ratan Murty
类目: ound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
*备注: 6 pages, 4 figures, 1 table
Abstract:Increasing depth of transformer models improves recognition, but it comes at a substantial cost. Each additional layer requires more parameters, which makes the process computationally inefficient. We ask whether additional processing can focus on integrating features already computed. Looped Audio Spectrogram Transformer (LAST) first processes all tokens, then reuses the same blocks to refine only the class token over fixed audio features, thereby making later passes inexpensive. On AudioSet, ten-pass LAST achieves 0.345 mean average precision, exceeding a twelve-layer sequential transformer by 2.1% relative with 49.4% fewer parameters, 42% fewer multiply-accumulate operations, and 9.8% higher measured throughput. Across separately trained models, increasing the pass count from two to ten improves accuracy while adding only 1.2% computation. Further evaluations show improved robustness to temporal masking and various other auditory augmentations, with better generalization on classification tasks with music, environmental, and event sounds.
[LG-29] Same Reward Different Skills: When Multimodal RL Learns to Look
链接: https://arxiv.org/abs/2610.01908
作者: Haocun Ye,Xinlong Jiang,Qile Chen,Bingyu Wang,Teng Zhang,Shubai Chen,Tingyu Wu,Zhenkun Zheng,Yiqiang Chen
类目: Machine Learning (cs.LG)
*备注:
Abstract:Reinforcement learning with verifiable rewards (RLVR) improves vision-language benchmark scores even without visual information during training. With images at test, blind-trained models recover roughly half of the real-image gain at 3B and nearly four fifths at 7B. Prolonged real-image training can erode grounding while benchmark gains persist. Both findings expose the same gap: an image in the prompt is not an image in the learning signal. Our design rule, visual resolvability, asks that visual evidence be necessary for a correct answer and that the task remain learnable. We test it on counterfactual coordinate scenes in which the question stays fixed and the target is never named, so a correct answer requires finding the target in the image. With standard GRPO and correctness-and-format rewards, a 7B model raises its accuracy at finding the target (discovery) from 0.425 to 0.875 on held-out scenes denser than any it trained on, and it improves on question types it never trained on. Two controls locate the source of the gain. Replacing test images with gray canvases drops discovery to zero; training on gray canvases instead, at matched step 30 and in each of four seeds, yields essentially none of the gain even when the model is then tested with real images. The learned skill carries over to grounding tasks built independently of the training corpus. A caption that answers the training question, added to the same images, reward and budget, cuts the gain by nearly two thirds. Changing what reward requires changes what RL learns.
[LG-30] Higher-Order Positional Encodings for Graph Representation Learning
链接: https://arxiv.org/abs/2610.01903
作者: Caleb Stam,Aagrim Hoysal,Sanjukta Krishnagopal
类目: Machine Learning (cs.LG)
*备注: Accepted at the Fifth Learning on Graphs Conference (LoG 2026)
Abstract:Many real-world systems exhibit higher-order interactions among groups of entities that cannot be captured by pairwise relationships alone. Graph Transformers and Graph Neural Networks increasingly rely on positional encodings to enrich graph representations, yet existing positional encodings are computed solely from the original graph and therefore cannot directly capture observed higher-order interactions. Topological Deep Learning addresses this limitation by lifting graphs to simplicial complexes, but typically requires performing message passing or attention on higher-order neural network representations. We introduce a representation learning paradigm that enriches graph representations with higher-order topology through positional encodings, enabling standard graph learning models to exploit lifted incidence structure without modifying the backbone. We derive a theoretical characterization of the expressivity of higher-order positional encodings, proving that node-level operators induced by higher-order lifts can mix graph Laplacian frequencies in ways that scalar graph spectral filters cannot. Guided by this theory, we instantiate higher-order positional encodings using Hodge Laplacians derived from clique complexes. Experiments with Graph Transformers on ZINC and controlled synthetic benchmarks demonstrate improvements in predictive performance, while a fixed-1-skeleton experiment shows that the pipeline can transmit higher-order information when cells are supplied independently of the graph. Together, our results establish higher-order positional encodings as a principled bridge between graph positional encodings and topological deep learning.
[LG-31] A foundation for systematic analysis of transformers and RNNs for tractography
链接: https://arxiv.org/abs/2610.01894
作者: Emmanuelle Renauld,Philippe Poulin,Hugo Larochelle,Antoine Théberge,Maxime Descoteaux
类目: Machine Learning (cs.LG); Image and Video Processing (eess.IV); Neurons and Cognition (q-bio.NC)
*备注:
Abstract:Machine learning (ML) has emerged as a promising approach for improving diffusion MRI (dMRI) tractography, a task that remains limited by the intrinsic tension between local diffusion information and global anatomical plausibility. In this work, we systematically evaluate recurrent neural networks (RNNs) and Transformer models for iterative tractography, with particular attention to training strategies, input representations (including convolutional neural network (CNN)-based embeddings and end-of-sequence (EOS) tokens), and hyperparameter selection. We introduce a generation-validation phase enabling supervision at the streamline level during training, allowing supervision despite the mismatch between local loss functions and global streamline quality. Using the ISMRM2015 tractography challenge dataset, our models achieve the highest reported performance to date. Through controlled experiments, we quantify the impact of missing bundles, noisy or imperfect training streamlines, and invalid fibers in the training set. Finally, we demonstrate the applicability of our best-performing models for in vivo data from the Tractoinferno database. Overall, our results highlight both the potential and the limits of sequence-based deep learning models such as Transformers and RNNs for tractography, and emphasize the need for improved phantoms and evaluation methods for in vivo validation. We provide takeaways and recommendations for future researchers training and validating sequence-based supervised methods for tractography.
[LG-32] RACE: Tackling Real-World Resource Assignment Problems via Agent ic Heuristic Design
链接: https://arxiv.org/abs/2610.01887
作者: Jose A. Ayala-Romero,Andres Garcia-Saavedra,Xavier Costa-Perez
类目: Neural and Evolutionary Computing (cs.NE); Machine Learning (cs.LG)
*备注:
Abstract:Dynamic resource assignment, the real-time allocation of task streams to heterogeneous processing nodes, is the backbone of modern computing infrastructure. While learning-based schedulers excel in research, industrial deployments still rely on hand-written rules that operators can read, audit, and execute within tight latency budgets. LLM-based Automatic Heuristic Design (AHD) promises to automate writing such rules. However, existing AHD frameworks were developed for combinatorial problems fully specified to the LLM, and they learn only from a scalar fitness score. In real systems, the behaviour that determines a good heuristic, such as processor speeds or power consumption, is unknown a priori: the score reveals which heuristic performs better, but not why. This missing information is recorded in the system logs that every evaluation produces. Exploiting it is non-trivial: logs are massive and noisy, the relevant signals depend on the objective, and their content and format vary across hardware and software stacks, so they can neither be fed to an LLM as is nor processed by a fixed parser. We propose TRACE, which couples an evolutionary AHD loop with an agentic knowledge-extraction workflow. A Reasoner agent analyzes the log schema in light of the objective and formulates hypotheses about the system dynamics; a Coder agent writes and executes schema-specific code to test them, producing insights or executable tools for the evolved heuristics. We evaluate TRACE on a synthetic cloud benchmark and a 5G vRAN scenario built from industrial testbed measurements and operational traffic traces. TRACE consistently outperforms state-of-the-art AHD methods in resource assignment problems and yields more auditable heuristics at under 2% overhead.
[LG-33] Pooling Helps Learned Weighting Hurts In-Context: Decomposing Group Attention
链接: https://arxiv.org/abs/2610.01831
作者: Michael Fore,James Mason Inder,Mrishika Nair,Praneetha Vaddamanu,Sharlina Keshava
类目: Machine Learning (cs.LG)
*备注:
Abstract:Group attention, introduced by the time series forecasting model Chronos-2, attends over the variates of a group at a fixed patch index and serves both multivariate (MV) and in-context learning (ICL) forecasting. Rather than evaluating this cross-variate attention design as a whole, we ask which part of the mechanism earns the benefit and probe its applicability to both MV and ICL regimes. By editing the attention matrix \alpha at inference we separate the two pathways a head comprises: V/O, which projects a weighted summary of the group, and Q/K, which decides the weights. Uniform pooling (V/O without any Q/K weighting) is positive on 18 of our 20 sensor-network configurations, while the learned weighting (Q/K) splits by group type: its contribution is positive or negligible for MV, but materially degrades 8 of the 10 sensor-network ICL configurations, leaving 4 of them worse than univariate inference. By isolating the impact of different layers, we find that uniforming \alpha in the first block alone improves every ICL configuration we test.
[LG-34] Scientific Discovery under Validation Congestion via Multi-Fidelity Pairwise Rankings
链接: https://arxiv.org/abs/2610.01827
作者: Kevin Tirta Wijaya,Alston Lo,Michael Sun,Wojciech Matusik,Vahid Babaei
类目: Machine Learning (cs.LG)
*备注:
Abstract:Modern computational methods can now propose candidate molecules, materials, and other scientific designs at an unprecedented scale, creating a validation congestion where candidates are abundant, but experimental capacity to physically evaluate them remains scarce. Discovering novel scientific designs has therefore become increasingly dependent on curation: selecting a small set of promising designs for slow and costly experiments. Existing curation methods typically rely on data-driven regression models that predict absolute scores, but training these models requires substantial experimental data to begin with. Yet, useful curation signals do not have to take the form of absolute measurements, as scientific design discovery is often comparative in nature. Here, we propose that curation can instead be primarily driven by expert pairwise rankings, which are substantially easier to gather. The expertise can come from computational tools or human input of multiple levels of fidelity, ranging from empirical rules of thumb to agentic workflows and experienced scientists. We introduce PRISMS, a framework that uses pairwise rankings from one or more experts, potentially spanning multiple levels of expertise, to identify the most promising candidates without relying on data-hungry regressors. When experts differ in fidelity and cost, PRISMS escalates pairwise queries from lower- to higher-fidelity rankers based on a Fisher-information criterion. In iterative screening that selects designs from fixed drug discovery libraries, PRISMS achieves 50% top-10 discovery recall in ~42% fewer rounds than regression-only active learning, and in ~15% fewer rounds than the ranking-based method with no selective escalation. In optimization that generates new designs without restriction to a predefined library, PRISMS achieves ~18.8% higher hypervolume than the Bayesian optimization baseline.
[LG-35] MECHVAR: Variance-Guided Mechanism Discrimination for Autonomous Machine Learning Experiment Selection
链接: https://arxiv.org/abs/2610.01819
作者: Yifan Guo
类目: Machine Learning (cs.LG)
*备注: 17 pages, 7 figures
Abstract:Benchmark gains are often mechanism-ambiguous: reproducing an improvement does not by itself identify why it occurs. We study finite-library mechanism discrimination, where posterior-weighted candidate mechanisms, executable probes, and a limited experimental budget define a sequential experiment-selection problem. MECHVAR selects the next probe by maximizing the posterior-weighted variance of its predicted responses. Under a shared-Gaussian predictive model, this score is exactly proportional to the classical Box–Hill posterior-weighted pairwise-KL criterion, yet it admits O(KE) vectorized rescoring and a transparent additive audit over mechanism pairs. A local expansion further links the score to expected information gain (EIG) when predicted response separations are small. In a 25-block stress audit, MECHVAR outperforms confirmation-first in several moderate misspecification regimes, while its primary comparisons with EIG remain statistically unresolved. In a held-out Digits loop, normalized mechanism-identification AUC is 0.8975 for MECHVAR, 0.7825 for a score-greedy policy, and 0.9092 for EIG. At K = 100, E = 200, median single-thread full-library scoring is 10.36 microseconds for MECHVAR versus 57.69 ms for six-node quadrature EIG in the recorded environment. MECHVAR therefore provides a lightweight, auditable acquisition rule for finite-library experiment selection when a shared predictive scale is a defensible approximation.
[LG-36] Debias Anything: Fairness with Diversity without Supervision in Diffusion Models
链接: https://arxiv.org/abs/2610.01815
作者: Théau d’Audiffret,Mariia Vladimirova,Jean-Yves Franceschi
类目: Machine Learning (cs.LG)
*备注:
Abstract:Although diffusion models produce high-quality images, they also reproduce and amplify demographic imbalances in their training data. Debiasing their generation process post-training w.r.t. some sensitive attribute usually relies on classifier guidance or explicit text extra-conditioning, but this reduces methods’ applicability and output diversity. Conversely, methods promoting diversity alone do not ensure fair attribute representation. In this paper, we propose a method tackling fairness and diversity jointly that is generally applicable to any diffusion model and any sensitive attribute. To this end, an adapter connects the frozen diffusion model to a pretrained vision-language embedding space, enabling fairness and diversity guidance without sensitive-attribute annotations. For fairness, pairs of text prompts define attribute directions which guide batch composition towards specific proportions. For diversity, we introduce a score measuring disagreement between the semantic estimates derived from this representation. The formulation supports unconditional and text-conditional diffusion models, while requiring no prior knowledge or data of sensitive attribute. Experiments confirm that our method improves quality and diversity scores at comparable fairness levels.
[LG-37] SkillEvoLean: Mutation-enhanced skill evolution for Lean provers
链接: https://arxiv.org/abs/2610.01799
作者: Kuo Zhou,ZiXion Yang,Lu Zhang
类目: Machine Learning (cs.LG)
*备注:
Abstract:Skill evolution offers a promising way to improve large language model agents without updating their parameters, but its use in formal theorem proving remains underexplored. Existing methods mainly target natural-language reasoning, improving skills by analyzing successful and failed trajectories and incrementally revising solving strategies. Although the Lean verifier provides reliable execution feedback, when all sampled trajectories fail, existing skill evolution methods lack successful trajectories from which to infer effective update directions. Furthermore, these methods also focus mainly on the root instruction file, thus underexploring the evolution of reference knowledge including mathematical concepts and proving techniques. To address these limitations, we propose a mutation-enhanced skill self-evolution framework for building skill-augmented Lean provers. The framework jointly evolves a high-level solving policy and its reference knowledge through progressive and mutation-based updates. Progressive evolution derives local improvements from successful and failed trajectories, while mutation is triggered when no complete proof can be generated, sampling mathematical concepts to produce and select new skill candidates under verifier feedback. We evaluate our method on MiniF2F, PutnamBench, the 2025 International Mathematical Olympiad (IMO 2025), and the 2026 USA Mathematical Olympiad (USAMO 2026). Under the same backbone model, trajectorysampling budget, and test-time compute, our method achieves proof success rates of 100.0%, 90.6%, 4/6, and 4/6, respectively, with GPT-5.5, outperforming the baseline methods. Further analysis shows that concept-guided mutation outperforms random-text-guided mutation by 6.9 and 8.2 percentage points on MiniF2F and PutnamBench, respectively, while solving one additional problem on both IMO 2025 and USAMO 2026.
[LG-38] SAGE: Similarity-Based Cleaning of Poisoned Training Data from Verified Examples
链接: https://arxiv.org/abs/2610.01788
作者: Chaeeun Han,Soodeh Atefi,Yevgeniy Vorobeychik,Aron Laszka
类目: Machine Learning (cs.LG); Cryptography and Security (cs.CR)
*备注:
Abstract:As machine learning increasingly relies on public, untrusted data sources, data poisoning attacks, which inject malicious examples into training data to induce misclassification of a chosen target, pose a growing threat. Existing defenses either assume zero ground-truth information about which examples are poisoned, or they assume access to a large set of examples verified to be clean. Satisfying the latter assumption incurs significant cost since reliable verification can be very resource- or labor-intensive. This cost is particularly high for clean-label attacks, where poisoned examples are visually indistinguishable from clean data. Since requiring a large set of verified examples is impractical, we propose relying on a small set of verified examples including both clean and poisoned ones, i.e., each example verified either to be clean or poisoned through inspection by a forensic expert. The challenge is then to detect poisons based on a set of verified examples that is so small that most classification models would overfit. To address this challenge, we propose Similarity-based Approach for Ground-truth-driven Exclusion (SAGE), which trains a generic feature extractor on a separate dataset and then flags poisoned training examples using a non-parametric, similarity-weighted prediction based on the verified set. On standard benchmarks against seven clean-label attack methods, we demonstrate that having access to even a handful of verified poisoned examples provides a substantial advantage. We also find that the distribution of verified clean examples across classes matters more than the number of verified examples.
[LG-39] Inferring Multi-Timescale Neural Dynamics with Switching Linear Dynamical Systems
链接: https://arxiv.org/abs/2610.01786
作者: Lulu Gong,Yongxu Zhang,Shreya Saxena
类目: Machine Learning (cs.LG); Signal Processing (eess.SP); Neurons and Cognition (q-bio.NC); Machine Learning (stat.ML)
*备注: 30 pages, 10 figures
Abstract:Neural activity often exhibits multiple timescales that can vary with behavioral states and task conditions. Identifying these timescales from neural recordings is important for better understanding neural computation and function. However, traditional approaches based on autocorrelation fitting are difficult to scale to high-dimensional population recordings and can become unreliable when neural dynamics change with behavior. State-space models have been a powerful framework for modeling high-dimensional neural population activity through latent dynamical systems, but standard formulations and inference methods do not explicitly account for multiple timescales and therefore do not guarantee accurate recovery of the underlying temporal structure. Motivated by these questions, we introduce the Multi-Timescale Switching Linear Dynamical System (MTS-SLDS), a framework for identifying regime-specific latent timescales from continuous or spiking neural observations. MTS-SLDS combines a multi-lag moment initialization, which captures temporal structure across multiple observation lags, with \textitregime-conditioned Laplace-EM inference, which reduces mixing of dynamical statistics across uncertain regimes. Characteristic timescales can then be extracted directly from the eigenvalues of the learned latent transition matrices. In synthetic and neural experiments with Gaussian and Poisson spike observations, MTS-SLDS accurately recovers timescales and switching structure over multiple datasets.
[LG-40] he Innocent Courier: Covert Exfiltration Through Legitimate LLM Web Fetching
链接: https://arxiv.org/abs/2610.01768
作者: Alessandro Pegoraro,Daryan Merx,Phillip Rieger,Ahmad-Reza Sadeghi
类目: Cryptography and Security (cs.CR); Machine Learning (cs.LG)
*备注:
Abstract:With the increasing capabilities of Large-Language-Models (LLMs) and LLM-based agents, users are increasingly using them to solve everyday problems, such as answering e-mails or providing programming support. Existing work has extensively investigated security and privacy risks, such as prompt injections and the disclosure of sensitive data to chatbot providers. While various solutions were developed to address these risks, including input structuring to prevent prompt injections or deploying local LLMs to avoid sharing confidential data with chatbot operators, LLMs also pose the risk of leaking confidential data to third parties. In this paper, we demonstrate with LLMLeak a novel attack vector where malicious software that runs locally but cannot communicate directly with the internet abuses LLMs to establish a covert channel. While inputs that instruct the LLM to send data directly via generated code are easy to detect and network libraries are typically restricted, LLMLeak relies only on the LLM’s tool to fetch websites for further information. A malicious software component on the client side embeds a secret into a URL. It presents the referenced website as providing information required for a benign task, such as migrating a software library. When the LLM accesses the URL, the attacker receives the encoded secret through an attacker-controlled DNS or web server. We perform an extensive evaluation on eleven open-parameter models, observe an attack success rate of 79.7%, and also conduct a case study on real-world chatbots, demonstrating the relevance of LLMLeak. Subjects: Cryptography and Security (cs.CR); Machine Learning (cs.LG) Cite as: arXiv:2610.01768 [cs.CR] (or arXiv:2610.01768v1 [cs.CR] for this version) https://doi.org/10.48550/arXiv.2610.01768 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[LG-41] Physics-Refined Spatiotemporal Forecasting on Open-Boundary Hydrologic Graphs ICDM
链接: https://arxiv.org/abs/2610.01765
作者: Haoyang Jiang,Zhengui Wang,Shenghan Gao,Y. Joseph Zhang,Xingquan Zhu,Yi He
类目: Machine Learning (cs.LG)
*备注: Accepted at the 2026 IEEE International Conference on Data Mining (ICDM)
Abstract:Spatiotemporal forecasting on hydrologic graphs is especially prone to instability in open-boundary systems, where the forecast domain exchanges fluxes with an unobserved exterior. In such systems, boundary nodes receive external forcing, e.g., upstream inflows in rivers or tidal signals in coastal regions, that is typically unavailable at prediction time. The absence of this information can compound errors as forecasts unfold in an autoregressive fashion, leading to inferior long-horizon performance. This paper dissects this instability issue by exploring two questions. 1) What boundary forcing enters the forecast domain when information beyond the boundary is missing? 2) How should this forcing propagate through the domain without incurring error amplification under autoregressive rollout? To address both, we propose a new computing framework comprising two key components. First, to compensate for the boundary forcing, our framework learns ghost node proxies from the boundary and interior nodes, striving to approximate unobserved external inputs. Second, to control error accumulation from these learned proxies, we leverage two physics refiners. In particular, one refiner enforces local consistency by aligning ghost proxies with their two-hop neighbors (i.e., boundary nodes and their immediate interiors). The other refiner enhances global stability by correcting the model forecasts through a physics-guided graph neural operator, reducing long-horizon numerical drift. Two real-world hydrologic graphs are employed for empirical evaluation. Comparative results show that our proposal enjoys higher prediction accuracy and long-horizon stability over both learning-based and physics-informed model competitors. Comments: Accepted at the 2026 IEEE International Conference on Data Mining (ICDM) Subjects: Machine Learning (cs.LG) Cite as: arXiv:2610.01765 [cs.LG] (or arXiv:2610.01765v1 [cs.LG] for this version) https://doi.org/10.48550/arXiv.2610.01765 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[LG-42] Evidence-Gated Research: Statistically Controlled Model Adoption in Adaptive Search
链接: https://arxiv.org/abs/2610.01751
作者: Yifan Guo
类目: Machine Learning (cs.LG)
*备注: 17 pages, 4 figures. Preprint
Abstract:Adaptive model search is path dependent: once a challenger is adopted, it becomes the reference from which later candidates are generated. A statistically unsupported replacement can therefore alter hypotheses that have not yet been proposed. We introduce Evidence-Gated Research (EGR), a statistical adoption layer for moving-incumbent search. EGR freezes each challenger before decision evidence is revealed, builds anytime-valid evidence across a predeclared set of environments, routes evidence predictably toward unresolved components, composes a persistent candidate e-value, and passes that e-value to an online controller. Under explicit conditional-validity and predictability conditions, the resulting procedure controls false discovery rate for the declared all-environment adoption target even though earlier adoptions change later challengers. In a 5,000-trajectory closed-loop benchmark, development-only e-LOND attains persistent FDR 0.621, whereas no persistent false-adoption path is observed for the audited EGR variants in that finite run. In matched replay over 600 challenger–incumbent pairs, Stagewise EGR preserves fixed-anytime alternative crossing decisions while using 56.1% less decision evidence at the representative threshold. A three-environment public-data study and a 40,000-sample controlled neural benchmark reproduce the evidence-efficiency pattern. These results identify model replacement as a distinct statistical control point in adaptive model development.
[LG-43] Fixed-point neural samplers on discrete spaces
链接: https://arxiv.org/abs/2610.01739
作者: Jiajun He,Denis Blessing,Mouyang Cheng,Yuanqi Du,Carles Domingo-Enrich
类目: Machine Learning (cs.LG)
*备注:
Abstract:Sampling from discrete, unnormalized distributions without access to data is a challenging problem. Neural samplers offer a promising approach by training generative models from density evaluations directly. Despite recent progress, existing discrete neural samplers are prone to mode collapse, come without convergence guarantees when trained via fixed-point iterations, and are often tied to a specific reference process such as masked or uniform diffusion. In this work, we introduce Discrete Gibbs Iterative Neural Sampler, a fixed-point neural sampler that addresses these limitations, enabling efficient, scalable learning, substantially reducing mode collapse in practice. Our framework builds on masked diffusion and also extends to transport between pairs of distributions. We demonstrate that the resulting method scales effectively to high-dimensional systems, supports amortized sampling across different conditions, and enables accurate estimation of alloy phase diagrams.
[LG-44] Participation-Sensitive Convergence and the Frag ment First Converge Later Pattern in Asynchronous Online Learning: A Topological Analysis Across 22 OULAD Courses
链接: https://arxiv.org/abs/2610.01738
作者: Hitoshi Inoue,Koichi Yasutake
类目: Computers and Society (cs.CY); Machine Learning (cs.LG)
*备注: Author’s version, posted under the non-commercial rights retained in the APSCE copyright transfer agreement
Abstract:Asynchronous online learning offers temporal flexibility at a structural cost: learning communities tend to fragment rather than cohere. \beta_0 , the number of disconnected behavioral clusters from Zigzag Persistent Homology, serves as a cohort-level indicator of this structure. Two questions remained unverified at scale: (1) does apparent \beta_0 convergence reflect genuine behavioral alignment or learner dropout? and (2) do assessment deadlines produce reproducible fragmentation-convergence cycles? We address both across all 22 OULAD courses (N 22,000; 857 week-pairs). Changes in \beta_0 strongly co-vary with active learner changes (pooled r = 0.387; median per-course r_delta = 0.459, 20/22 courses), identifying \beta_0 as a participation-sensitive indicator: \beta_0 and active learner counts co-respond to deadline events rather than one causing the other. Deadlines produced fragmentation in 82.6% of assessments and the full Fragment First, Converge Later (FFCL) cycle in 60.2%. 3-phase analysis confirmed structural fragmentation as the dominant long-term trajectory (90.9% of courses), moderated by curriculum structure. These findings establish \beta_0 as a participation-sensitive structural indicator with direct implications for AI-augmented learning analytics design.
[LG-45] Function-Structured Reinforcement Learning with Executable Verifiers for Mathematical Reasoning
链接: https://arxiv.org/abs/2610.01729
作者: Zihan Liu,Xurong Xie
类目: Machine Learning (cs.LG)
*备注: 5 pages, 2 figures, 2 tables
Abstract:Algorithmic mathematical reasoning requires reliable decomposition, computation, and aggregation. Final-answer rewards provide limited guidance on intermediate errors, while successful execution does not guarantee mathematical correctness. This work proposes Function-Structured Graph Reinforcement Learning (FSG-RL), connecting subproblem graphs and Python implementations with multi-verifier feedback. The policy first learns to generate code from function graphs through supervised fine-tuning (SFT). Group Relative Policy Optimization (GRPO) then optimizes the policy using answer-gated rewards and span-level credit assignment. The framework also supports teacher supervision and structured memory. A benchmark curated from Grade School Math 8K (GSM8K), MathQA, MATH, and Omni-MATH pairs public function graphs with private verification specifications. Under a unified evaluation protocol, GRPO improves final-answer accuracy from 43.25% to 67.50% and full solution success from 32.25% to 52.25% over SFT. Continued reinforcement learning (RL) with teacher supervision yields additional gains. The gains extend beyond producing correctly formatted code, supporting verifier-guided reinforcement learning for mathematical reasoning. Code is available at this https URL.
[LG-46] RelICL: Training-free Relational Learning with Tabular Foundation Models
链接: https://arxiv.org/abs/2610.01725
作者: Simon Forbat,Rainer Gemulla
类目: Machine Learning (cs.LG)
*备注:
Abstract:Tabular foundation models achieve state-of-the-art performance on single-table tasks without any training. Recent work suggests that they are also well-suited for relational learning via deep feature synthesis (DFS), which flattens a relational schema into a single table by adding aggregates of the other tables’ columns as features. This approach is appealing because it directly benefits from improvements to or customization of the underlying tabular foundation model. In this paper, we identify two key problems with DFS: feature explosion and interaction blindness. The first problem arises because the number of DFS features grows quickly as the schema becomes more complex, limiting scalability and performance. The second problem arises because column-wise aggregates do not account for feature interactions, limiting performance. We propose and explore an alternative method termed RelICL, which keeps the benefits of DFS but alleviates these two problems. At its heart, RelICL propagates and fuses information step by step through the schema graph, using the same tabular foundation model that is eventually used for prediction to do so. In our experimental study using RelBench tasks, RelICL was on par with the strongest approach based on deep feature synthesis.
[LG-47] Anomaly Detection and Localization for the Pantograph-Catenary System ITSC2026
链接: https://arxiv.org/abs/2610.01721
作者: Francesco Vitale,Hangli Ge,Francesco Flammini
类目: Machine Learning (cs.LG)
*备注: Accepted and presented at the Industry Track of the IEEE International Conference on Intelligent Transportation Systems 2026 (IEEE ITSC 2026)
Abstract:Monitoring the Pantograph-Catenary System (PCS) provides insight into the health conditions of the pantograph and the railway infrastructure. Recent industrial solutions trace the pantograph’s contact wire height and stagger (PCS height/stagger) using video monitoring through convolutional neural networks. However, these solutions do not account for the train route’s geographic location. Therefore, in this paper we propose a novel framework for 1) localization of the PCS height/stagger by alignment with the nominal GPS coordinates of the reference route, and 2) collective anomaly detection to evaluate the health conditions of the PCS. We apply and assess the localization and detection performance of the methodology to a case-study based on a real-world industrial dataset provided by a railway transportation company, which includes the PCS height/stagger of several train journeys across Italian railway routes.
[LG-48] In-context Learning of Single-index Targets: Comparing Kernel and Feature Learners
链接: https://arxiv.org/abs/2610.01712
作者: Haotian Gu,Yizhou Xu,Lenka Zdeborová
类目: Machine Learning (cs.LG); Disordered Systems and Neural Networks (cond-mat.dis-nn); Machine Learning (stat.ML)
*备注:
Abstract:In-context learning (ICL) enables a pretrained model to infer a task from demonstrations without updating its parameters. While much of the existing theory focuses on linear target functions, in this paper we study nonlinear cases by comparing two one-layer attention architectures on the same family of single-index tasks. A kernel learner first maps inputs through a fixed nonlinear feature map and then applies linear attention, whereas a feature learner applies attention to the original input, followed by a learned nonlinear readout. We derive predictions for their memorization and generalization errors using the replica method, retaining the effects of pretraining size, task-pool diversity, and training and inference context lengths. The resulting predictions closely match numerical experiments across a broad range of regimes. Our analysis yields phase diagrams that characterize when each architecture is advantageous as the amount of pretraining data, task diversity, and context lengths vary. We further identify qualitatively different context-length scalings for the two learners. Together, these results clarify how architectural choices interact with the dataset and govern nonlinear in-context learning.
[LG-49] Learning PDE Dynamics between Submanifolds Using Greens Observation Operators
链接: https://arxiv.org/abs/2610.01697
作者: Jan Tauberschmidt,Jephte Abijuru,Samuel Okon,Naukshatro Bose,Sophie Fellenz,Marius Kloft,Jonas Latz,Sebastian Josef Vollmer
类目: Machine Learning (cs.LG)
*备注:
Abstract:Many physical systems are driven and observed only on lower-dimensional submanifolds of a larger spatial domain, while their dynamics are governed by the ambient medium occupying that domain. Examples include laser-heated parts imaged by an infrared camera, and ground-level emissions measured on a sensor plane. Full-domain solvers, however, compute the entire volume for every new source although only the observation submanifold is needed, and black-box surrogates do not exploit that the ambient medium remains fixed. We introduce the \emphGreen’s Observation Operator (GObO), which maps the ambient medium once to the Green’s kernel of a linear PDE restricted to the source and observation submanifolds. New sources then cost one lower-dimensional integral and no network evaluation. Exponential rates in the kernel yield an exact finite streaming state with horizon-independent memory; we prove its stability and an approximation rate for the restricted heat kernel. On three-dimensional heat conduction and advection–diffusion with collocated and distinct source and observation geometries, GObO trained on static sources predicts responses to moving sources zero-shot with 4–8 \times lower error than black-box surrogates, at 1.4,ms per query after a single conditioning pass. The same kernel transfers across resolutions and admits corrections for mild nonlinearities, including radiative losses and temperature-dependent conductivity, without retraining, at the cost of lower in-distribution accuracy.
[LG-50] Artifact Annotations Partially Substitute for Per-User Calibration: SAFE-EDA and a Normalization-Controlled Evaluation of Wrist-EDA Affect Recognition
链接: https://arxiv.org/abs/2610.01692
作者: Haochen Chai,Xinbi Luo,Zining Liu,Fangfang Jiang
类目: Machine Learning (cs.LG); Signal Processing (eess.SP)
*备注: 12 pages, 6 figures, 6 tables, plus 2 pages of supplementary material. Code: this https URL
Abstract:Wrist electrodermal activity (EDA) differs in amplitude from one person to the next, so affect-recognition models normalize their input before classification. Studies that test such models on held-out subjects seldom report where the normalization statistics come from, yet statistics computed from the held-out subject’s own recording give the model information that a device does not have when it is first worn. We asked how this choice alters the measured benefit of pretraining. A compact convolutional network, SAFE-EDA, was pretrained on expert artifact annotations from 43 subjects and compared with the same network trained from scratch on the Wearable Stress and Affect Detection (WESAD) dataset (15 subjects, leave-one-subject-out), with two normalization sources crossed with four window hops. When the statistics came only from training subjects, pretraining raised macro-F1 by 0.078 to 0.227; when they came from the held-out user’s full recording, the gain fell to between 0.020 and 0.050 and was no longer significant. Artifact supervision was far more useful than self-supervised pretraining on the same recordings (0.078 versus 0.008). Across 13 configurations in two datasets, the pretrained network was better in 12, but on the second dataset (26 subjects) per-user normalization increased the gain instead of reducing it, so the interaction depends on the data. Only five of 50 published WESAD studies state which data were used for normalization. Reporting this choice is necessary to separate first-use performance from performance after calibration.
[LG-51] MiLoop: Selective Memory Propagation for Neural Combinatorial Optimization
链接: https://arxiv.org/abs/2610.01685
作者: Changliang Zhou,Yuanyao Chen,Rongsheng Chen,Zhiyun Lin,Zhenkun Wang
类目: Machine Learning (cs.LG)
*备注:
Abstract:Constructive neural combinatorial optimization (NCO) has emerged as a promising paradigm that learns to construct solutions to combinatorial optimization problems (COPs) step by step, which reduces reliance on handcrafted rules and enables fast inference. While many methods with dynamic embeddings generalize well, they typically rebuild subproblem representations from scratch at each step using deep attention stacks. Many high-performing methods in this category rely on solution labels or pseudo-labels for efficient training, or on aggressive search space pruning during reinforcement learning (RL). To address these limitations, we propose Memory-in-the-Loop (MiLoop), a purely RL-based constructive framework that leverages the multi-step computation already required by a rollout for selective memory propagation. Each rollout provides solution-quality feedback for learning while propagating historical representations, thereby enabling a shallow policy to learn effective dynamic embeddings without external solution labels or training-time search-space pruning. Specifically, MiLoop fuses current embeddings with historical memory before the attention layers and applies adaptive gated updates afterward. The updated representations support both current decisions and stepwise reuse. Extensive experiments across four COPs demonstrate that MiLoop consistently produces high-quality solutions on instances ranging from 100 to 10 million nodes, highlighting its strong generalization ability.
[LG-52] Invent a Dataset: Measuring dataset generation abilities with zero seed
链接: https://arxiv.org/abs/2610.01674
作者: Shivalika Singh,Andrija Djurisic,Gbemileke Onilude,Sudip Roy,Sara Hooker
类目: Machine Learning (cs.LG)
*备注:
Abstract:Building datasets remains one of the most manual and brittle parts of AI development. In this technical report, we focus on the most extreme but also most prevalent setting real world practitioners face: a zero data regime. Here, practitioners don’t have any data for the capability they want to learn. We introduce Invent-A-Dataset which is a prompt based system to go from dataset description to realistic and large scale post-training datasets. We evaluate Invent-A-Dataset against five frontier model APIs including Anthropic, Google, Open AI, DeepSeek, Zai. Across eight task types and dataset sizes up to 20K samples, Invent-A-Dataset significantly outperforms with both the highest quality (17% relative gains) while simultaneously producing the most diverse samples (19% relative gains). Its diversity advantage widens with scale of training dataset size (from parity at 200 samples to 37% relative gains at 20K samples). This translates into considerable downstream training gains, resulting in far more performant post-trained models. Invent-A-Dataset fine-tune consistently ranks higher compared to other generator fine-tunes across different post-trained model architectures.
[LG-53] pCoMole: Pareto-Constrained Molecule Editing with Discrete Flows NEURIPS2026
链接: https://arxiv.org/abs/2610.01663
作者: Tong Chen,Maximilian Holsman,Lin Zhao,Pranam Chatterjee
类目: Machine Learning (cs.LG); Biomolecules (q-bio.BM)
*备注: Published at NeurIPS 2026. (Proceedings of the 40th Conference on Neural Information Processing Systems, Sydney, Australia)
Abstract:Biomolecular therapeutics often start from known sequences and require targeted editing to improve multiple properties while satisfying hard biochemical and manufacturability constraints. However, existing generative methods do not jointly support multi-objective optimization, hard feasibility, and sequence editing in discrete, variable-length biological spaces. In this work, we introduce Pareto-Constrained Molecule Editing (pCoMole), a framework built on discrete flow matching that steers a pre-trained Edit Flow toward user-specified preferences while enforcing terminal feasibility. pCoMole defines a feasibility-gated terminal distribution using an augmented Tchebycheff utility and realizes the resulting preference tilt through a Doob-h transform of the underlying edit process. To make this construction practical, we approximate the required harmonic function using short Monte Carlo rollouts over candidate edits, yielding an efficient guided editor with provable preference consistency. We validate pCoMole by shrinking GFP while retaining fluorescence-related properties, shortening diverse Cas9 orthologs while preserving PAM specificity, and compressing peptide binders into short peptidomimetics that optimize seven drug-related properties under hard constraints. In wet lab testing, two 229-residue pCoMole-designed eGFP variants retained clear green fluorescence in BL21 cells after 10 deletions, with either one or two substitutions. Together, pCoMole enables constraint-aware, Pareto-aligned editing of biomolecular sequences in discrete, variable-length spaces.
[LG-54] CrossGMN: Graph Metanetworks for Cross-Architecture Weight-Space Transformations
链接: https://arxiv.org/abs/2610.01649
作者: Adir Dayan,Yam Eitan,Haggai Maron
类目: Machine Learning (cs.LG)
*备注:
Abstract:Weight-space networks operate directly on parameters of other neural networks, enabling tasks such as predicting model properties, editing trained models, and generating weights. Weight-space symmetries such as neuron permutations make equivariance a key design principle. However, existing equivariant weight-space architectures have primarily been studied for transformations that preserve the network architecture. In contrast, many practical transformations, including model compression and upscaling, map a trained source network into a target network with a different architecture. In this setting, the source and target permutation symmetries act on different parameter spaces, making equivariance less straightforward to formulate. Our key idea for addressing this mismatch is to reformulate cross-architecture operators with two inputs: a trained source network and an initialization of the target network. This lets us define equivariant cross-architecture operators that refine the initialization of the target network using information from the source network, while being invariant to source-network permutations and equivariant to target-network permutations. Based on this formulation, we introduce CrossGMN, a graph metanetwork that jointly processes both networks through symmetry-preserving cross-network message passing. We prove CrossGMN is universal for continuous cross-architecture operators on compact sets under a general-position assumption. We evaluate CrossGMN for model compression, predicting a smaller network’s parameters to accelerate subsequent knowledge distillation. Across 2-D and 3-D INRs and image classification with MLPs, CNNs, and Vision Transformers, CrossGMN speeds up distillation by up to 8.89x, transfers across datasets without retraining (3.78x), and a single model can accelerate compression from heterogeneous source architectures into a common target architecture.
[LG-55] owards a Cloud Fog Edge System for Smart Building
链接: https://arxiv.org/abs/2610.01647
作者: Christophe Cérin,Mamadou Sow,Frédéric Andrès
类目: Distributed, Parallel, and Cluster Computing (cs.DC); Machine Learning (cs.LG)
*备注:
Abstract:In this article, we present our vision and recent advancements toward creating a decentralized system capable of learning from real-time data within buildings to support sustainable and privacy-preserving smart environments. Our approach promotes the concept of the building itself as the data center, aligning with the principles of edge computing to safeguard confidentiality and reduce reliance on external cloud infrastructure. This is particularly valuable in humanitarian contexts, where data sovereignty, energy efficiency, and infrastructure constraints are critical. We detail a lightweight, “Kubernetes-like” orchestration framework for deploying AI services within such environments and demonstrate our progress in implementing AI algorithms on low-power, cost-effective microcontrollers such as those in the Arduino ecosystem. By enabling in-situ learning directly on sensors or microcontrollers, our work aims to bring intelligent services to resource-limited settings, fostering autonomy, resilience, and sustainable development in vulnerable or underserved communities. The contributions in this article are related, firstly, to our project “Online Machine Learning Algorithms for Embedded Systems” and the evaluation of two new online algorithms. Secondly, we envision a cloud-fog-edge architecture based on the KOptim and FIWARE components, and we propose a methodology for coupling them. Experimental results of the online algorithms are also presented, showcasing real-world traces.
[LG-56] Beyond Demographic Balance: Multi-Metric and Intersectional Evaluation of Fairness in MIMIC-IV Mortality Prediction
链接: https://arxiv.org/abs/2610.01645
作者: Abdullah Al Noman,Fahmid Al Rifat,Tahrima Hashem,Syed Muhammad Ibne Zulfiker,Rishov Paul,Tanzima HAshem
类目: Machine Learning (cs.LG)
*备注: NEurlPS TAE workshop 2026 accepted
Abstract:Fairness conclusions in clinical prediction can depend strongly on both the metrics reported and the demographic resolution at which performance is evaluated. We revisit these evaluation choices for ICU mortality prediction on MIMIC-IV, comparing predictive-utility and subgroup-error metrics across several fairness interventions. As a complementary case study, we introduce a lightweight adaptation strategy that jointly balances ethnicity–gender–insurance representation without conditioning on mortality outcomes, allowing demographic representation balancing to be examined separately from outcome-conditioned or direct error-rate interventions. We evaluate its behavior at both marginal and corresponding three-way intersectional subgroup levels, while accounting for the statistical support of finer-grained estimates. The results show that interventions can receive substantially different assessments across accuracy/AUROC, sensitivity, and false-positive rate, and that marginal demographic summaries can conceal heterogeneous error profiles within their constituent intersections, including among larger subgroups. These findings highlight the importance of evaluating fairness interventions at both complementary metric and subgroup resolutions, while accounting for the intervention target and the reliability of subgroup estimates.
[LG-57] FedSAP: Federated Learning with Structured Adaptive Partitioning for Multi-Domain Heterogeneous Edge Devices
链接: https://arxiv.org/abs/2610.01638
作者: Wentao Yue,Tianyou Lai,Hongji Li,Qingyu Mao,Qilei Li
类目: Machine Learning (cs.LG)
*备注:
Abstract:Federated learning (FL) on heterogeneous edge devices must jointly accommodate unequal resource budgets and domain-shifted local data. Existing resource-adaptive methods decide how much of a model each client trains but not where retained capacity should reside or how it should be shared, whereas federated domain-generalization methods usually assume a shared full architecture. Uniform compression can therefore discard high-utility channels, and a single aggregation path can mix transferable features with domain-sensitive updates. We propose FedSAP, a domain-aware heterogeneous FL framework that casts structured pruning as budget-constrained tri-state channel allocation. FedSAP converts each keep ratio into non-uniform layer budgets, assigns stable channels to a Global pool, useful domain-sensitive channels to pseudo-domain-specific Private pools, and low-utility channels to a Dropped state. This partition lets broadly useful features benefit from cross-client pooling while isolating domain-sensitive updates from incompatible clients. Domain-Guided Assignment infers pseudo-domains from shallow-gradient similarity, while Type-Matched Aggregation restricts each channel to its intended sharing scope. Across three random seeds, FedSAP reaches 76.00% and 72.67% mean global accuracy on Digits and Office-Caltech, exceeding the strongest baseline by 1.70 and 4.92 percentage points while supporting client pruning ratios of up to 80% across heterogeneous clients.
[LG-58] Generalization in Neural Networks Through the Lens of Magnitude Potential
链接: https://arxiv.org/abs/2610.01633
作者: Sahel Torkamani,Henry Gouk,Rik Sarkar
类目: Machine Learning (cs.LG)
*备注:
Abstract:Explaining generalization and training dynamics in neural networks remains a challenge, and various approaches have been developed to study different aspects of these phenomena. In this paper, we introduce the idea of \em magnitude potential – a quantity based on the theory of metric magnitude – that reflects how well an arbitrary point is represented by a given set. We find that this basic quantity can be applied to examine various features in neural generalization. The ratio between the magnitude potential with respect to a class and with respect to the entire data, computed at the logit layer, is informative of the representation of the point. In experiments, these ratios for individual training points are found to be correlated with the Feldman memorization scores. Magnitude potential ratios aggregated across points detect structural changes in the decision boundaries and provide a geometric indicator of grokking in modular arithmetic. Although the magnitude potential ratio and neural collapse are both closely associated with intra-class and inter-class geometric structure, the magnitude potential ratio remains informative even when neural collapse is explicitly suppressed.
[LG-59] Continual Reinforcement Learning with Neuroevolution
链接: https://arxiv.org/abs/2610.01583
作者: Eleni Nisioti,Andrea Cossu,Kathrin Korte,Sebastian Risi
类目: Neural and Evolutionary Computing (cs.NE); Machine Learning (cs.LG)
*备注:
Abstract:Despite many studies about causes and remedies of plasticity loss in Reinforcement Learning (RL) under continual task changes, no RL method has yet consistently achieved a good balance between adaptation and forgetting. Here we turn to an alternative optimization paradigm, neuroevolution (NE): algorithms that search directly in weight space through mutation and selection over a population of neural networks. Across a wide array of environments and environmental changes, with policies ranging from a few hundred parameters to million-parameter networks, we compare evolution strategies (ES) and genetic algorithms (GAs) against state-of-the-art continual RL variants and population-based RL. ES most consistently achieves a good stability-plasticity trade-off, while the GA is the most plastic method but forgets more than ES. To explain this, we study the return landscape around each method’s solutions. ES finds the widest neighborhoods, i.e.\ regions of weight space in which perturbed policies still solve the task, and the size of the overlap between the neighborhoods of consecutive tasks correlates with a method’s stability-plasticity trade-off. Rewarding behavioral diversity in a GA through novelty search makes the population even more plastic, at the cost of forgetting. Finally, symptoms of plasticity loss commonly reported in RL do not transfer to NE. Overall, these results establish NE as a competitive alternative to RL under continual task changes, and suggest that training under perturbations in weight space may be a useful mechanism for continual learning more broadly.
[LG-60] owards Optimal Policy Improvement
链接: https://arxiv.org/abs/2610.01566
作者: Yaniv Oren,Viliam Vadocz,Wiktor Zabka,Thomas Evers,Jan Robine,Wendelin Böhmer,Matthijs T. J. Spaan,Martha White,Hendrik Baier,Fenghui Yu
类目: Machine Learning (cs.LG)
*备注:
Abstract:Practical Reinforcement Learning (RL) algorithms learn to solve Markov Decision Processes (MDPs) through iterative policy improvement in the presence of approximate evaluation. We study policy improvement from first principles, defining optimal policy improvement as producing the best policy attainable in a single update under specified constraints. We show that optimal improvement restricted to a set of states is equivalent to solving an induced MDP, characterizing planning with an explicit or implicit model as a path towards optimal policy improvement. Because practical methods commonly solve such induced problems through iterative improvement in the form of greedification, we take steps towards optimal greedification under the central practical constraint of approximate evaluation. We formulate greedification under this constraint as probabilistic decision-making under uncertainty and derive a novel operator that is optimal with respect to the resulting objective. Empirically, the operator and its practical gradient-based approximations improve aggregate performance across GumbelAlphaZero, SAC, ReBRAC and Generalized Policy Iteration, in experiments spanning discrete and continuous actions, model-based and model-free, online and offline RL.
[LG-61] Range-GRPO: Policy Optimization via Pairwise Relations among Reward Intervals
链接: https://arxiv.org/abs/2610.01548
作者: Ryunyi Lee,Kangjun Noh,Somin Kim,Heedong Kim,Kyungwoo Song
类目: Machine Learning (cs.LG)
*备注:
Abstract:As the use of large language models (LLMs) expands, post-training has become increasingly important for adapting them to downstream tasks. However, obtaining reliable supervision remains costly, especially in domains without reference answers or executable verifiers. LLM-as-a-Judge provides scalable pseudo-rewards for unlabeled responses, but a single point score does not explicitly represent reward uncertainty. This motivates representing pseudo-rewards as conformally calibrated reward ranges. We propose Range-GRPO, a semi-supervised post-training framework that combines limited labeled data with unlabeled prompts. In Group Relative Policy Optimization (GRPO), learning signals depend on relative reward comparisons within each rollout group. The proposed objective compares reward ranges pairwise rather than reducing them to point rewards, allowing interval uncertainty to affect both the magnitude and direction of these signals. Our theoretical analysis characterizes this distinction and shows that the proposed objective recovers the this http URL advantage when all reward ranges collapse to points. Empirically, Range-GRPO achieves the highest in-distribution and out-of-distribution average performance among the evaluated semi-supervised methods while requiring fewer training resources.
[LG-62] FedFit: Federated Fine-Tuning of LLM s via Vector-Bank Parameterization and Quantization
链接: https://arxiv.org/abs/2610.01537
作者: Hang Zou,Chao Zhang,Yuzhi Yang,Yu Tian,Samson Lasaulce,Mérouane Debbah
类目: Machine Learning (cs.LG)
*备注:
Abstract:Federated Learning (FL) enables privacy-preserving fine-tuning of Large Language Models (LLMs), yet the massive communication overhead remains a critical bottleneck. Furthermore, applying Low-Rank Adaptation (LoRA) in FL faces a fundamental “aggregation dilemma” between the accurate Sum-of-Products (SoP) and the communication-efficient Product-of-Sums (PoS) implementations. To tackle these challenges, we propose FedFit. First, to significantly reduce communication overhead, we introduce a disjoint shared vector-bank parameterization that reconstructs high-dimensional adapter matrices from two compact and disjoint global vector banks. Second, to address the aggregation dilemma, we devise an alternating optimization schedule. By cycling between decoupled single-bank updates (which allow for accurate aggregation) and joint updates corrected by a Residual Spectral Aggregation mechanism, we resolve the conflict between SoP and PoS. Additionally, we integrate blockwise quantization with client-side error feedback to further compress the transmitted vectors. Furthermore, we establish theoretical convergence guarantees for the proposed algorithm. Extensive experiments on Qwen2.5 models demonstrate that FedFit achieves perplexity performance comparable to standard federated LoRA methods, while providing compression ratios up to 100x higher.
[LG-63] Calibrating Prediction Timeliness Through Multi-Objective Hyperparameter Optimization for Remaining Useful Life Prediction
链接: https://arxiv.org/abs/2610.01530
作者: Tugrul Cabir Hakyemez,Ener Uras Gokhan
类目: Machine Learning (cs.LG)
*备注:
Abstract:In predictive maintenance, early and late RUL prediction errors carry asymmetric consequences, yet hyperparameter optimization typically targets a single accuracy metric that treats both directions equally. This study treats the optimization objective itself as a design variable. Five architectures (MLP, LSTM, XGBoost, TCN, and Transformer) are evaluated under three regimes: single-objective maximization of R^2 , single-objective minimization of the NASA scoring function, and a multi-objective formulation that jointly optimizes both criteria. The multi-objective search employs NSGA-II with Entropy-CRITIC weighting for Pareto selection. Seventy-five model-dataset-strategy combinations are assessed on the NASA C-MAPSS turbofan and BackBlaze hard-disk drive benchmarks. On C-MAPSS, all strategies achieve comparable accuracy ( R^2 \approx 0.89 ), yet multi-objective optimization reduces directional imbalance by approximately 33%, improving calibration of early versus late predictions. Model rankings prove configuration-dependent, with simpler architectures frequently outperforming deeper temporal models. On BackBlaze, the objectives shift from complementary to conflicting, producing divergent Entropy-CRITIC weights and a substantial generalization gap (best R^2 \approx 0.34 ). These results demonstrate that the optimization objective materially shapes prognostic behavior and that multi-objective search provides a practical mechanism for calibrating prediction timeliness in RUL modeling.
[LG-64] Langevin-Informed Transfer Learning: Replacing Target Samples by Black-Box Feedback
链接: https://arxiv.org/abs/2610.01522
作者: Vladimir R. Kostic,Karim Lounici,Hélène Halconruy,Timothée Devergne,Michele Parrinello,Massimiliano Pontil
类目: Machine Learning (cs.LG); Machine Learning (stat.ML)
*备注:
Abstract:Many scientific and machine learning systems, from molecular dynamics to diffusion models and beyond, are governed by stochastic dynamics with low-dimensional structure, evolving on slow timescales. However, target trajectories, used to identify and interpret such dynamics, are often inaccessible: only biased or static samples that explore the underlying manifold are available. We introduce Langevin-Informed Transfer Learning (LITL), a framework for recovering target Langevin dynamics from biased source samples using only black-box feedback. LITL learns the leading spectral structure of the target infinitesimal generator and the projected drift through Dirichlet representation learning, enabling kinetic reconstruction in spectral form and slow-manifold gradient field estimation. We further introduce a spherical variant well suited to steering normalized latent representations commonly used in learning systems toward desired objectives. We establish finite-sample guarantees for eigenvalue, eigenfunction, and projected drift estimation in Sobolev norms, thereby ensuring generalization of these quantities and their first-order derivatives. Empirically, LITL recovers physical transition timescales from biased molecular simulations, builds kinetic structure from static samples of generative models, reconstructs spherical symmetries of physical systems, and enables post-hoc latent steering of trained neural networks under black-box feedback. Together, these results position spectral operator learning as a practical framework for recovering stochastic dynamics under distribution shift and unlock applications across machine learning and the physical sciences.
[LG-65] Learned End-to-End Guidance Schedules for Diffusion Models
链接: https://arxiv.org/abs/2610.01502
作者: Aneesh Barthakur,Mathias Niepert,Luiz F.O. Chamon
类目: Machine Learning (cs.LG)
*备注:
Abstract:Diffusion models are a powerful generative paradigm used across multimedia and scientific applications. Guided diffusion methods impose requirements on the generation by adding the gradient of a differentiable loss (the guidance function) as a drift term during inference. The weight of this drift (the guidance scale) is critical for the trade-off between data quality and requirement satisfaction. To achieve both of these goals, guided diffusion must resort to small guidance scales and lengthy sampling, incurring high computational costs. This work proposes learned end-to-end guidance schedules (LEEGS) to achieve these objectives with fewer sampling steps. LEEGS trains a time-dependent schedule by minimizing the guidance function over a small set of examples using stochastic gradient descent. Backpropagating through guided sampling is computationally expensive, so LEEGS uses an approximation of the gradient that cuts training time by a factor of 4. We evaluate LEEGS on diverse guidance tasks, including (a) image inpainting, (b) noisy image inverse problems, © face-ID-guided generation, and (d) forward and inverse PDE problems, outperforming baselines at equal budget (50 or 100 NFEs), or matching constant guidance with only 10% of the steps.
[LG-66] Let the Heads Talk: Beyond Diagonal Graph Attention
链接: https://arxiv.org/abs/2610.01494
作者: Riccardo Ali,Alessio Borgi,Mario Severino,Alessio Gravina,Davide Bacciu,Pietro Liò,Christopher Irwin
类目: Machine Learning (cs.LG)
*备注:
Abstract:Sheaf Neural Networks generalize scalar-weighted message passing by replacing scalar edge weights with linear transport maps between local feature spaces. Yet the role of this matrix-valued transport is entangled with the broader sheaf-diffusion construction. We isolate the transport primitive through quiver representations and establish a direct connection with multi-head attention. Treating attention heads as coordinates of a local transport space reveals that standard multi-head attention implements diagonal edge maps: along each directed interaction, a source head can contribute only to the corresponding receiver head. Allowing off-diagonal entries instead enables edge-conditioned communication across heads before neighborhood aggregation. We show that this operation cannot, in general, be absorbed into a single shared linear map applied after aggregation. Building on this characterization, we introduce Topological Attention (Top-A), a multi-head attention that learns edge-dependent off-diagonal routes while preserving the original same-head paths and exactly recovering vanilla attention when the additional routing vanishes. We evaluate Top-A on relational reasoning, heterogeneous graph learning, and algorithmic reasoning, including out-of-distribution generalization, with heterophilic node classification as a contrast setting. The results show that cross-head transport is most useful when the task benefits from interaction-dependent transformations, while heterophily alone provides no systematic advantage. These findings identify edge-conditioned cross-head communication as a distinct computational primitive of matrix-valued transport.
[LG-67] LESS: Lightweight Evolutionary Supernet Search in Minutes
链接: https://arxiv.org/abs/2610.01468
作者: Aviral Gandhi,Jinglue Xu,Jialong Li,Hitoshi Iba
类目: Neural and Evolutionary Computing (cs.NE); Machine Learning (cs.LG)
*备注: 33 pages, 4 figures. Code: this https URL
Abstract:Low-cost NAS must both explore high-performing architectures and identify them reliably, yet reducing evaluation cost often weakens the fidelity of candidate comparisons. Training-free methods reduce evaluation cost by replacing learned task feedback with proxy signals measured at initialization. We introduce LESS (Lightweight Evolutionary Supernet Search), a data-driven method that combines a brief fair hard-path warm-up with discrete search under a single CMA-ES distribution. Each proposal is evaluated as its decoded hard genotype after six candidate-conditioned supernet updates. On NAS-Bench-201, LESS achieves (93.189\pm0.467%) CIFAR-10 test accuracy in 409.1 seconds, coming within 0.04 percentage points of FairNAS using approximately (1/24) of its source-reported search time. Matched controls show that calibration improves selected validation accuracy by (0.577) percentage points while changing best-visited accuracy by only (0.054) points, indicating that its primary effect is to reduce selection regret. The frozen configuration transfers without tuning to CIFAR-100 and ImageNet16-120 with (69.615\pm1.139%) and (43.720\pm1.697%) accuracy. Applied without tuning to the larger DARTS space, LESS achieves (96.95\pm0.14%) on CIFAR-10 and (82.43\pm0.80%) on CIFAR-100, with each search completing in approximately 43.5 minutes on a single GPU. Together, these results show that short, balanced, data-dependent updates enable competitive neural architecture search across datasets and search spaces within minutes.
[LG-68] ght Transition Time Bounds for Separable Logistic Regression at the Edge of Stability
链接: https://arxiv.org/abs/2610.01459
作者: Haodong Wen,Kaiyue Wen,Jiaye Teng
类目: Machine Learning (cs.LG); Machine Learning (stat.ML)
*备注:
Abstract:We study logistic regression on linearly separable data under gradient descent with a large constant stepsize \eta . Such dynamics may exhibit a characteristic Edge of Stability phenomenon, in which the loss initially oscillates before transitioning to a stable phase of monotone decrease. Existing work provides a tight \Theta(1) bound in dimension d=2 as \eta \to \infty and conjectures a bound independent of \eta in arbitrary dimensions d\geq 2 . In this paper, we disprove this conjecture by showing that, for every fixed sample size n\geq 2 and sufficiently small margin \gamma , the worst-case transition time is \Theta!\left((\log\eta)^\min\n-2,d-2\right) uniformly over d\geq2 . The key challenge in establishing a tight bound is that the sample contributing most strongly to the gradient can change repeatedly across iterations. To address this issue, we control such changes by induction on dimension and sample size, and construct matching hard instances.
[LG-69] Streaming algorithms for robust max-min diversification
链接: https://arxiv.org/abs/2610.01456
作者: Andrea Pietracaprina,Geppino Pucci,Stefano Zanon
类目: Machine Learning (cs.LG)
*备注:
Abstract:Given a set of n points X in a metric space and an integer k , max-min diversification aims to select k points of X maximizing their minimum pairwise distance. This objective function is however highly vulnerable to noisy points. In[Amagata, AAAI23], a robust formulation is proposed which addresses this vulnerability by excluding solutions containing any of z outliers, defined as the z points in X with the largest nearest-neighbor distances. That paper also presents a coreset-based streaming algorithm for the new formulation, based on a suitable inlier-outlier separation assumption. However, we identify three shortcomings in the algorithm by [Amagata, AAAI23]: its coreset construction requires an offline computation over X , which needs memory linear in n , in stark contrast with the typical goals of stream processing; the one-pass procedure used to extract the solution from the coreset may return fewer than k points (hence, an unfeasible solution) because it permanently discards points too far from the current solution; and its outlier-exclusion guarantee is only probabilistic and weakens as the coreset size shrinks. In contrast, we present a deterministic coreset-based algorithm that, under a natural inlier-outlier separation assumption (similar to the one used in [Amagata, AAAI23]), returns exactly k inliers which are a (2+\varepsilon) -approximate solution, for any \varepsilon0 , thus only \varepsilon above the best polynomial-time sequential approximation, even without outliers. Its one-pass streaming implementation adapts obliviously to the dataset’s doubling dimension D and, for wide ranges of k , z , \varepsilon , and D , it uses memory independent of n . For sufficiently long streams, its amortized update time is proportional to the coreset size, thus also independent of n .
[LG-70] Repurposing Obsolete Representations for Post-Deployment Adaptation
链接: https://arxiv.org/abs/2610.01453
作者: Daniel Bethell,Charmaine Barker,Simos Gerasimou
类目: Machine Learning (cs.LG)
*备注:
Abstract:Deep neural networks are increasingly deployed in long-lived systems, where task requirements may change after training. In such settings, part of the original output space may become obsolete: a class, prediction region, or learned behaviour may no longer be valid. Existing approaches either leave the obsolete behaviour intact or require fine-tuning, which can be expensive. We propose Deep Repurposing (DR), a post-hoc framework for adapting models under task obsolescence. DR estimates the latent geometry of obsolete and retained regions, removes obsolete-supporting components, and reallocates retained-compatible evidence through an analytic repair map without gradient updates. This yields repaired predictions and representations in which obsolete regions no longer act as valid outputs, while useful obsolete structure can support the retained task. Across multiple task settings, DR removes obsolete behaviour while preserving retained utility. More importantly, across classification benchmarks, DR matches or exceeds competing unlearning and editing baselines in retained accuracy, eliminates obsolete predictions, and adapts up to 60\times faster than competing unlearning methods.
[LG-71] Distillation of Tabular Foundation Models into Efficient Predictors
链接: https://arxiv.org/abs/2610.01435
作者: Minho Jeong,Dooho Lee,Jinmo Lee,Jaemin Yoo
类目: Machine Learning (cs.LG)
*备注:
Abstract:Tabular foundation models (TFMs) achieve strong predictive performance through in-context learning, yet repeatedly conditioning on labeled data makes inference expensive. Knowledge distillation can reduce this cost by transferring their predictive ability to lightweight, dataset-specific students. However, the dependence of TFM predictions on both a labeled context and a query introduces two design questions: how to construct teacher supervision and whether expanding query coverage improves distillation. We examine these questions across two TFMs and both neural and tree-based students, and derive an effective distillation recipe. The recipe uses the full labeled training set as teacher context and trains students solely on teacher predictions for observed and synthetic queries. On TabArena, the resulting students outperform their supervised trained tuned-and-ensembled counterparts by 57-98 Elo points. Applied unchanged to TALENT, the same recipe improves matched default students on 236-258 of 300 datasets and reduces median primary error by 4.0-6.4%. The distilled students also achieve median inference speedups of 3.0-21.6 times over their teachers, offering a practical trade-off between predictive performance and repeated inference cost. Code is available at this https URL .
[LG-72] Least-time Gradient Flow
链接: https://arxiv.org/abs/2610.01426
作者: Alessandro Betti,Marco Gori,Stefano Melacci,Jinwei Zhao
类目: Machine Learning (cs.LG); Optimization and Control (math.OC)
*备注:
Abstract:Prescribing the speed of gradient flow on the risk itself, by the dynamics \dot w=-u(E(w))\nabla E(w)/\abs\nabla E(w)^2 , makes the risk e(t)=E(w(t)) obey \dot e=-u(e) exactly, whatever the landscape~ E ; the time needed to reach zero risk from e_0 is \int_0^e_0\dd e/u(e) . Minimizing this time alone is ill posed, and we study the regularized problem \inf\int_0^e_0(\tfrac\lambda2\absu’^2+1/u),\dd e:\ u\in H^1(0,e_0),\ u\ge0,\ u(0)=0\ , \lambda0 . We prove that the minimizer exists, is unique, and is a linearly scaled cycloid, and we show that the optimal rate behaves like u^*(e)\sim(9/(2\lambda))^1/3e^2/3 near zero risk: the exponent 2/3 is the one found in \citebetti2026holder by a power-law ansatz, and it lies in the Hölder window (\tfrac12,1) where the arrival is in finite time with vanishing weight speed. The proof follows the classical route: existence by the direct method, uniqueness by strict convexity, positivity of the minimizer away from the origin, and the explicit integration of the Euler-Lagrange equation.
[LG-73] Why Does Train-Validation Separation Emerge? Update-Pressure Density Dynamics in Pretrained Backbones
链接: https://arxiv.org/abs/2610.01425
作者: Yuchen Li,Mingyu Du,Zongqi Fan,Ken-Tye Yong,Nguyen H. Tran
类目: Machine Learning (cs.LG)
*备注:
Abstract:Train-validation separation is the evolving difference between performance on observed training examples and a finite held-out validation set. We propose a dynamic structural account of how this gap develops during adaptation of pretrained models: continued fitting can shift update demand from broadly reusable support toward narrower support with weaker held-out transfer. A conditional local model links this shift to increasing heterogeneity in gradient allocation and train-validation separation. Fixed training probes make this structural evolution observable without validation examples entering the readouts; held-out performance is used separately to evaluate its relation to the gap. In a constructed hierarchy implemented with a residual multilayer perceptron (ResMLP), increasing the target share of example-private features from p=.3 to .5 to .7 , while preserving the relative mixture 1:2:3:4 among the four shared feature levels, increases the final mean accuracy gap from .185 to .331 to .527 across five runs per condition. Masked-input losses measured separately on training and validation examples expose the corresponding transfer asymmetry. The natural language processing (NLP) analysis uses 10-epoch runs of RoBERTa, DeBERTa, and Qwen on six datasets (90 runs): the training-probe-weighted within-class and overall dispersion readouts each have positive raw and smoothed level correlations with the accuracy gap in all 90 runs. Raw changes paired at approximately one-epoch intervals remain positively associated in 86/90 and 87/90 runs, respectively. A 40-epoch ResNet-18 study tests both readouts on three vision datasets. Together, controlled simulation, NLP, and vision support the dynamic structural account across settings, with real-model evidence testing its observable predictions under the specified monitors.
[LG-74] Reusing Past Samples in Proximal Policy Optimization: When and How Does It Help?
链接: https://arxiv.org/abs/2610.01399
作者: Alessandro Montenegro,Riccardo Venturelli,Marco Mussi,Matteo Papini,Alberto Maria Metelli
类目: Machine Learning (cs.LG)
*备注:
Abstract:Among on-policy deep reinforcement learning methods, Proximal Policy Optimization (PPO) has become the de facto standard, due to its consistently strong empirical performance across diverse application domains. However, on-policy methods are inherently sample inefficient: fresh data collected under the current policy is used for just a few updates before being discarded. Off-policy methods avoid this inefficiency via experience replay, achieving notable sample efficiency gains, but at the cost of training instabilities or extensive tuning. This motivated the rise of hybrid strategies that augment PPO with off-policy data reuse. Existing sample-reuse variants of PPO demonstrated improved sample efficiency over vanilla PPO, yet a systematic study of when reuse helps, in which scenarios, and to what extent remains missing. In this work, we study the effectiveness of sample reuse in PPO by instantiating two variants within a multiple importance weighting framework. Both retain the core PPO mechanics, reusing only samples from a window of recent iterations, thereby isolating the effect of data reuse from other factors. The variants, termed wPPO-U and wPPO-BH, employ vanilla importance weights or balance-heuristic-corrected ones, respectively. For both, we derive policy improvement lower bounds providing theoretical grounding for their respective losses. We use them to empirically study when and how data reuse improves sample efficiency or final performance of PPO across continuous control tasks.
[LG-75] AF-Muon: An AdamW-Free Muon Optimizer for Tied-Embedding Models NEURIPS2026
链接: https://arxiv.org/abs/2610.01395
作者: Arash Lagzian,Paniz Halvachi,Junming Zhang,Zhouhan Lin,Dianbo Liu
类目: Machine Learning (cs.LG)
*备注: 63 pages, 19 figures. An earlier, shorter version of this work was accepted as a poster at the OPT 2026 workshop (Optimization for Machine Learning) at NeurIPS 2026; this is the complete version
Abstract:Muon improves large-scale training by applying a spectral-norm steepest-descent update to matrix parameters, but practical models also contain parameter blocks that do not fit dense-matrix geometry. One important case is the tied vocabulary table, which appears in language models and other token generators and can receive multiple structurally different gradient sources, from sparse input lookups to dense output-classifier updates. In the reference recipe these blocks are handed to an auxiliary AdamW optimizer, which restores second-moment state and updates the aliased table as a generic tensor. We propose AF-Muon, an AdamW-free extension of Muon that keeps the Muon matrix update for hidden weight matrices while using a support-aware finite-cap linear minimization oracle for tied vocabulary tables and an RMS-normalized update for one-dimensional auxiliary parameters. AF-Muon therefore trains every parameter class with a single first-moment buffer and no second-moment state, saving around 20% optimizer-state memory relative to Hybrid Muon in our benchmark. Across nine tied-token settings - decoder-only language models from 124M to 1B parameters, a fully shared T5-style encoder-decoder, and ImageGPT-style image-token, protein, and sparse-MoE variants, spanning text, image, and protein-sequence data - AF-Muon improves mean validation loss and perplexity over both Hybrid Muon and a SCION-style Sign endpoint. Long-horizon runs and hyperparameter sensitivity studies confirm the gain is robust, and identical-momentum diagnostics attribute it to the finite cap, which preserves more within-row magnitude than Sign while bounding the coordinate concentration of row-RMS. These results identify tied vocabulary tables as a distinct optimizer geometry and yield a robust AdamW-free Muon variant across models, modalities, and architectures, with about 1% step-time overhead in matched training.
[LG-76] Robust Evidential Learning Through Latent Consistency
链接: https://arxiv.org/abs/2610.01384
作者: Charmaine Barker,Daniel Bethell,Simos Gerasimou
类目: Machine Learning (cs.LG)
*备注:
Abstract:Reliable uncertainty quantification is essential for deploying deep learning models in high-stakes settings, where out-of-distribution and adversarial inputs can induce confident but unreliable predictions. Evidential Deep Learning provides efficient uncertainty estimates in a single forward pass, but can still assign high evidential strength to inputs that are poorly supported by the learned representation, such as adversarial inputs. We introduce CLEAR, a lightweight, task-agnostic post-hoc method that improves evidential robustness without retraining or altering the base prediction. Using held-out calibration data, CLEAR characterises the group-conditioned geometry of the model’s latent space. At inference, it efficiently generates perturbation views directly in the latent space and measures their conflict relative to the calibrated geometry of the predicted group. High latent conflict indicates unsupported evidence, which CLEAR uses to selectively reduce evidential strength while retaining evidence for latent-consistent inputs. On ImageNet \rightarrow CUB, CLEAR improves OOD and adversarial AUROC by +8.29 and +5.01 while running 17.4 \times faster than competing post-hoc methods while preserving predictive performance across classification, regression, and object detection benchmarks.
[LG-77] Minimax Optimal Regret for Causal Logistic Bandits with Counterfactual Fairness
链接: https://arxiv.org/abs/2610.01377
作者: Junhyuk Huh,Seoungbin Bae,Dabeen Lee
类目: Machine Learning (cs.LG)
*备注:
Abstract:We study causal logistic bandits with counterfactual fairness constraints. The causal structure is given through known factual and counterfactual feature maps that share an unknown logistic reward parameter, but the learner observes only factual rewards. Consequently, the directions determining counterfactual feasibility need not be identifiable from the available feedback. The closest prior analyses either omit a coverage condition or impose a comparatively strong one, and do not establish matching lower bounds. We first show that some coverage condition is necessary: without a coverage-type restriction, factually indistinguishable environments with different optimal fair actions force \Omega(T) expected joint loss. Under a weaker full-rank condition on the factual covariance pooled across actions, we identify a target-specific information scale V_\star that measures the difficulty of estimating rewards and counterfactual effects from factual feedback. We construct worst-case families satisfying this condition on which every policy incurs expected joint loss \Omega\left(\left[V_\star\min\log K,d\right]^1/3T^2/3\right) . We also give an explore–then–exploit procedure tuned using V_\star and an adaptive algorithm that does not require its value. Both algorithms achieve \max\R_T,V_T=\widetildeO\left(\left[V_\star\min\log K,d\right]^1/3T^2/3+\kappa d/\sigma_0^2\right) , where R_T is regret relative to the best fair action and V_T denotes the cumulative stage-wise positive violations. Thus the upper and lower bounds match in their leading dependence on T , V_\star , and \min\log K,d\ , up to logarithmic factors.
[LG-78] Action-On-Item Preference Flow: A Shared Event Schema for Predictive and Generative Personalization NEURIPS2026
链接: https://arxiv.org/abs/2610.01375
作者: Parthiv Chatterjee,Kashish Kanjaria,Vashisth Purani,Sourish Dasgupta,Tanmoy Chakraborty
类目: Machine Learning (cs.LG)
*备注: Accepted to NeurIPS 2026. Author-prepared archival version with expanded discussion and interpretation
Abstract:A user’s movie, news, and dialogue histories differ in their native actions and outputs, yet each interaction supplies evidence that can update user memory. We study whether these histories can train one reusable update mechanism. An action-on-item schema pairs a mapped interaction role with a content embedding, allowing shared update parameters to operate on separate user states. We establish invariance to native relabeling, bounded state changes under item-embedding perturbations, and a pooled-training bound under explicit compatibility conditions. The Multi-Timescale State Hypothesis (MTSH) specifies how this evidence enters, persists, and is consumed; PerTIDE implements it with action gating, three state-space traces, fusion, and command-conditioned readout. On PENS, the same history encoder supports both next-news prediction and personalized headline generation. In a controlled PENS-to-MovieLens experiment, a frozen source-trained core exceeds an identically structured random core by 15.23 MRR points after fitting the same target consumer. On MIND, PerTIDE retains a 4.12-point MRR advantage over a same-input three-branch state-space control. Action, readout, and trace interventions identify complementary contributions to these gains. Together, the theory and experiments support learning history updates across compatible sources and reusing them through predictive and generative consumers.
[LG-79] Learning Commute-Time-Preserving World Models for Planning
链接: https://arxiv.org/abs/2610.01373
作者: Michael Hauri,Peter Buttaroni,Fabian A. Mikulasch,Friedemann Zenke
类目: Machine Learning (cs.LG)
*备注:
Abstract:World models allow agents to plan in latent space by choosing a sequence of actions that most reduces the distance to a given goal state. Thus, planning can benefit from latent representations whose distances mirror commute-times in the environment. The spectral embedding space of the graph Laplacian provides such a representation, if it obeys a specific eigenvalue-dependent scaling. Unfortunately, instantiating the graph Laplacian is intractable in large, continuous environments. Self-supervised learning offers a natural route to such commute-time-preserving embeddings at scale. However, here we show that existing methods, which commonly encourage isotropic representations to prevent representational collapse, tend to degrade the “correct” eigenvalue-dependent scaling, leading to an inaccurate representation of commute times. To address this problem, we introduce Commute-Time-Preserving World Models (CTWMs), combining a latent displacement predictor and a log-determinant regularizer that prevents collapse, which provably recover the correctly scaled Laplacian representation under reversible deterministic dynamics and at the predictor’s fixed point. In numerical simulations, CTWM matches or outperforms LeWM, a task-agnostic baseline, on several complex, continuous goal-reaching benchmarks, while using half the parameters.
[LG-80] From Redundancy to Minimality: Fixed-Point-Guided Hierarchical Reduction of Learned Piecewise-Linear Dynamics
链接: https://arxiv.org/abs/2610.01369
作者: Hiroto Tamura,Gouhei Tanaka
类目: Machine Learning (cs.LG); Chaotic Dynamics (nlin.CD)
*备注: 27 pages, 6 figures
Abstract:Understanding a nonlinear dynamical system from time series requires not only reproducing its trajectories, but also identifying a simple representation that preserves its essential dynamical structure. Almost-linear recurrent neural networks (AL-RNNs) are piecewise-linear RNNs in which only a subset of units use ReLU nonlinearities, so that nonlinear capacity is explicitly controlled by the number of ReLU units. Their activation patterns define linear regions, represented as symbols, whose observed transitions form a symbolic transition graph. However, directly training AL-RNNs with few ReLU units to realize minimal dynamical representations can be unreliable. We ask whether an AL-RNN with more ReLU units can instead be trained first and systematically reduced to a minimal dynamical representation. We introduce a fixed-point-guided hierarchical reduction procedure that progressively linearizes selected ReLU units, merging neighboring linear regions and graph nodes while preserving distinct symbols containing fixed points (FPs). The resulting reduction tree defines a hierarchy of progressively simpler candidates. Each reduced candidate is initialized from the parent parameters and retrained under guidance from the parent dynamics. We also prove that reproducing Q distinct fixed points requires at least Q FP-containing symbols, providing a certificate of symbol-level minimality when this bound is attained. On the 3-scroll Chua system, direct training with the theoretical minimum of three ReLU units achieves high-fidelity minimal realizations in only 20% of seeds, whereas our learn-reduce-retrain strategy increases the seed-macro success rate to approximately 71% at the same final nonlinear capacity. These results show that redundant nonlinear capacity can serve as a scaffold for discovering and realizing minimal dynamical representations.
[LG-81] Degree-Corrected Joint Matrix Factorization for Multilayer Community Detection
链接: https://arxiv.org/abs/2610.01361
作者: Alexandra Dache,Manon Rustin,Arnaud Vandaele,Nicolas Gillis
类目: ocial and Information Networks (cs.SI); Machine Learning (cs.LG)
*备注:
Abstract:Multilayer networks allow the modeling of interactions between the same entities across different contexts, such as temporal observations, varying settings, or interactions of different types. The goal of community detection in multilayer networks is to identify groups of nodes exhibiting similar connectivity patterns, which may vary across layers. We propose a method based on a joint nonnegative symmetric matrix trifactorization for community detection in multilayer networks, where each graph is approximated by a nonnegative symmetric matrix trifactorization. Our approach enforces constraints on the factor matrices so that communities are disjoint and shared across layers, while allowing each layer to have its own connectivity patterns and node degrees. This flexibility enables the model to capture both local and global structural variations across layers. We also develop an algorithm to efficiently solve this problem. We evaluate multilayer community detection methods using the multilayer degree-corrected stochastic block model (MDCBM), a flexible framework for generating realistic multilayer graphs with heterogeneous degrees and varying connectivity patterns. Experiments show that our method reliably detects communities across diverse regimes, whereas existing state-of-the-art approaches are often limited by restrictive structural assumptions.
[LG-82] Port-Hamiltonian Neural Networks for Systems with Multiple Asymptotically Stable Equilibria NEURIPS2026
链接: https://arxiv.org/abs/2610.01356
作者: Simon Heilig,Jens Püttschneider,Mohammad Itani,Asja Fischer,Timm Faulwasser
类目: Machine Learning (cs.LG); Systems and Control (eess.SY)
*备注: Accepted at NeurIPS 2026 Workshop: AXIOM - Foundations of Efficient Deep Learning
Abstract:Stable port-Hamiltonian neural networks certify asymptotic stability by construction. Yet, their Hamiltonian is a global Lyapunov function with a single global minimum, so they can represent only dynamic systems with one attractor. We demonstrate that this excludes even simple systems with energy landscapes forming a double well, and we overcome the restriction by parametrising the Hamiltonian as a product of Bregman divergences generated by one input-convex network. We prove that the resulting model is locally Lyapunov stable, that the coexistence of stable equilibria forces additional non-asymptotically-stable equilibria to exist, that all equilibria lie in a bounded region, and under a hyperbolicity assumption that almost-everywhere stability holds. On three systems our approach is able to recover the energy surface characteristics and improve the convergence speed by 1.8 \times -8.5 \times .
[LG-83] Robust Non-Clairvoyant Scheduling with Classification Models
链接: https://arxiv.org/abs/2610.01343
作者: Anthony Dugois,Vincent Fagnon,Giorgio Lucarelli
类目: Machine Learning (cs.LG); Data Structures and Algorithms (cs.DS)
*备注:
Abstract:We study the classical single-machine scheduling problem of minimizing the sum of completion times of jobs in a non-clairvoyant setting, where the processing time of each job remains unknown until its completion. This is a hard problem for which no constant competitive algorithm is possible. Inspired by robust optimization and learning-augmented algorithms, we introduce a novel robustness framework that leverages structural information provided by a classification model to overcome this limitation. Specifically, we assume that jobs are partitioned into classes and we have access to the confusion matrix of the classifier, whose entry (k,\ell) indicates the number of jobs predicted to belong to class~ k but that actually belong to class~ \ell . In this manner, we are able to characterize uncertainty as a set of permutations within each predicted class, rather than as a collection of discrete numerical scenarios, avoiding the computational difficulty of classical robust metrics, such as Min-Max and Min-Max Regret. In addition to these worst-case metrics, we also consider the expected objective over all scenarios. We first propose an optimal non-adaptive strategy that is oblivious with respect to all three robust criteria. We then investigate adaptive and randomized algorithms, showing that they can outperform the optimal non-adaptive strategy when the matrix exhibits particular structural properties.
[LG-84] Clifford Sheaf Neural Networks
链接: https://arxiv.org/abs/2610.01322
作者: Kotaro Kamiya,Joel Nicholls
类目: Machine Learning (cs.LG); Machine Learning (stat.ML)
*备注:
Abstract:We introduce the Clifford Sheaf Neural Network (CSNN), an equivariant sheaf neural network for geometric graphs that places a Clifford algebra on each stalk of a cellular sheaf and transports multivector features along edges. The canonical choice of restriction map for sheaves with algebra-valued stalks is algebra homomorphism. Adding the constraint of equivariance, the naive choice becomes versor conjugation. However, versor conjugation is expressively weak, so we drop algebra homomorphism and arrive at the K-term sandwich. The resulting sheaf Laplacian is positive semidefinite by construction, needs no versor constraint, and still mixes grades. Our main contribution characterizes the resulting family of restriction maps along three axes: which grades a map couples, how much of the endomorphism space it reaches, and how well it is conditioned. The K-term sandwich spans half of the endomorphism space, and in Cl(3, 0, 0) it corresponds to the maps that commute with the central pseudoscalar. The number of terms controls expressivity. CSNN is the reversion member, a first-order model by construction and the grade-mixing corner of this family, developed as a sheaf construction for graph-level equivariant regression.
[LG-85] Prediction-powered Neural Architecture Search
链接: https://arxiv.org/abs/2610.01317
作者: Pascal Janetzky,Yuxin Wang,Michael Klar,Stefan Feuerriegel
类目: Machine Learning (cs.LG)
*备注:
Abstract:Evaluating candidate architectures in neural architecture search (NAS) faces an inherent trade-off: on the one hand, reliable performance labels are limited because training and evaluating architectures is costly; on the other hand, zero-cost proxies (ZCPs) are cheap to compute at large scale but can be noisy. Yet, how to effectively combine these two sources of supervision remains unclear. In this paper, we propose PPNAS, a novel prediction-powered inference (PPI) approach for NAS. PPNAS fuses (1) a small set of architectures with observed performance labels and (2) a large set of architectures with ZCP information. To combine these two sources of supervision, PPNAS exploits the ordinal information provided by ZCPs to construct additional pairwise ranking supervision, while PPI debiases systematic discrepancies between ZCP-based and true performance rankings. We evaluate PPNAS in end-to-end predictor-based NAS, where it achieves state-of-the-art under limited evaluation budgets. To the best of our knowledge, PPNAS is the first prediction-powered approach for label-efficient NAS.
[LG-86] EP-Flow: Disordered Crystal Structure Prediction without Site-Level Annotations
链接: https://arxiv.org/abs/2610.01315
作者: Qiuliang Liu,Liming Wu,Qi Li,Zhonglong Peng,Chang Chen,Xiaolong Chen,Wenbing Huang,Shifeng Jin
类目: Machine Learning (cs.LG); Disordered Systems and Neural Networks (cond-mat.dis-nn)
*备注:
Abstract:Generative models have made rapid progress in ordered crystal structure prediction, yet many functional materials are intrinsically disordered, with substitutional mixing, vacancies, or interstitial species controlling their properties. Existing crystal generators either assume deterministic site occupations or require site-level disorder annotations, which are often unavailable when the chemical formula is the primary input. We formulate disordered crystal structure prediction through an Occupancy Distribution Matrix (ODM), a continuous site-by-species representation that unifies ordered crystals, solid solutions, vacancy disorder, and interstitial occupancy. A valid ODM must satisfy coupled site-wise occupancy, mass-conservation, and non-negativity constraints, placing each sample on a formula-dependent transportation polytope. We propose Entropic Polytope Flow (EP-Flow), a marginal-constrained flow matching framework that canonicalizes heterogeneous polytopes into a shared double-centered space, learns a marginal-preserving flow, and recovers feasible occupancies through a Sinkhorn inverse map. By jointly generating occupancies, fractional coordinates, and lattice parameters, EP-Flow achieves state-of-the-art performance on formula-conditioned disordered CSP benchmarks derived from COD and MPDS, substantially outperforming adapted ordered-crystal generators. Analyses further show that EP-Flow recovers sparse and chemically meaningful local disorder patterns rather than merely matching global composition statistics.
[LG-87] Continual Learning for 6-DoF Grasp Synthesis via Experience and Demonstrations
链接: https://arxiv.org/abs/2610.01301
作者: Giulio Schiavi,Andrei Cramariuc,Michael Pantic,Roland Siegwart
类目: Robotics (cs.RO); Machine Learning (cs.LG)
*备注: Accepted to CoRL 2026
Abstract:Most current grasp synthesis systems are trained offline and remain fixed during deployment. While this works well when deployment conditions resemble the training data, performance can degrade when robots encounter conditions they have not seen before, such as unfamiliar objects. In this work, we present a continual-learning framework for single-view 6-DoF grasp synthesis for a parallel-jaw gripper in cluttered scenes. Rather than finetuning a large parametric model, our method adapts through memory in a learned embedding space: grasp outcomes update future grasp scores, while optional user demonstrations are recalled and transferred to new scenes as additional candidate grasps. We evaluate our method in simulation and in extensive real-world experiments comprising over 1500 grasp trials. We show that our method matches the performance of existing 6-DoF grasping baselines even before adaptation, improves online on unseen objects from categories absent or underrepresented during training, and supports long-horizon continual learning with limited forgetting. In real-world experiments, our method reaches over 90% success rates on several challenging object categories after only 50 online grasp attempts. Videos and code at this https URL.
[LG-88] IQS-BO: In-Context Query Selection for Bayesian Optimisation
链接: https://arxiv.org/abs/2610.01269
作者: Luca Geminiani,Nadja Klein
类目: Machine Learning (cs.LG); Machine Learning (stat.ML)
*备注:
Abstract:Bayesian Optimisation (BO) is a powerful framework for the optimisation of expensive black-box functions, but typically requires refitting a surrogate and maximising an acquisition function at every evaluation step. In-context approaches based on Prior-data Fitted Networks (PFNs) amortise part of this cost by pre-training transformers on functions drawn from synthetic priors. PFNs4BO amortises the surrogate but still relies on a numerically maximised acquisition function, while FIBO performs BO fully in-context by sampling optimiser locations from a learned density, which fixes the decision rule and admits no surrogate. Learned acquisition functions score a finite candidate set with a trained network, but, lacking a label for the query, learn the score by reinforcement learning on previously solved tasks. We propose IQS-BO, a PFN that learns the query decision by supervised learning on synthetic priors. In a single forward pass, IQS-BO predicts the probability that each candidate maximises the objective over the set, and we show that the minimiser of its objective is the posterior probability of this event. The model can be pre-trained without a surrogate for fully in-context BO, or take the predictions of a fixed probabilistic surrogate as additional input, amortising only the decision step. Our method proposes queries at a fraction of the cost of acquisition-based methods, while either matching or outperforming standard BO with Gaussian processes (GPs) and available in-context methods on synthetic and real-world benchmarks. Finally, we propose a mixture prior for pre-training PFNs which combines samples from GPs with functions exhibiting warped inputs, isolated narrow optima, or plateaus that are poorly modeled by stationary kernels common in GP surrogates. We show that pre-training on this prior can lead to improved optimisation performance.
[LG-89] Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
链接: https://arxiv.org/abs/2610.01253
作者: Bisma Majid,Shabir Ahmed Sofi,Mir Mohammad Yousuf
类目: Machine Learning (cs.LG); Quantum Physics (quant-ph)
*备注: 9 pages, 18 figures, 13 tables
Abstract:Quantum Reinforcement Learning (QRL) integrates reinforcement learning with parameterized quantum circuits and is a promising approach to combinatorial optimization. On Noisy Intermediate-Scale Quantum (NISQ) devices, however, decoherence, gate imperfections, and measurement errors reduce policy quality and make learning less reliable. Existing error mitigation techniques are generally applied as fixed corrections that do not adapt to changing noise conditions or to the evolving state of training. This work presents Adaptive Policy-Guided Error Mitigation (APGEM) as a context-aware orchestration layer of the hybrid quantum-classical training loop that dynamically selects the most suitable mitigation strategy during QRL training. APGEM evaluates Zero-Noise Extrapolation (ZNE), Probabilistic Error Cancellation (PEC), Clifford Data Regression (CDR), and Readout Error Mitigation (REM) using policy-level indicators, including quantum-state fidelity, policy entropy, cumulative reward, and approximation ratio, and integrates the selected strategy directly into the reinforcement learning loop. The framework is evaluated on the Capacitated Vehicle Routing Problem (CVRP), a representative NP-hard problem in urban logistics, under a range of NISQ noise models and noise levels. APGEM consistently outperforms conventional static mitigation methods, reaches approximately 94% of the utility of an oracle strategy, maintains higher quantum-state fidelity as noise increases, and produces more stable learning behaviour throughout training. Ablation studies show that the framework learns context-aware mitigation policies that adapt to different noise environments and circuit execution conditions. These findings demonstrate that integrating adaptive error mitigation into the learning process substantially improves the robustness and reliability of QRL on NISQ hardware.
[LG-90] Mixture-Trained Merging for Unified Multi-Objective Models NEURIPS2026
链接: https://arxiv.org/abs/2610.01238
作者: SeongHyeon Kim,Chaeyun Jang,Seungyoo Lee,Jiyeon Ham,Yunju Bak,Boseop Kim,Juho Lee
类目: Machine Learning (cs.LG)
*备注: Accepted at NeurIPS 2026
Abstract:Unified language models are increasingly expected to combine heterogeneous capabilities, such as mathematics, code, instruction following, and controllable thinking behavior, within a single set of parameters. A common solution is sequential post-training on multiple objectives, but this entangles all objectives along one optimization trajectory and makes the final model highly sensitive to training order, data ratios, schedules, and stopping criteria. Weight-space merging offers a modular alternative, but naive merging of single-objective experts often fails: domain capabilities degrade sharply, or think/non-think modes collapse into one dominant behavior. We attribute both failures to incompatible weight-space geometry: experts trained on single objectives drift to distant regions of parameter space, placing their interpolations outside any shared low-loss basin. We propose Mixture-Trained Merging (MTM), which trains each branch on an objective-biased data mixture rather than a single objective, exposing it to cross-objective interactions and making branches compatible at merge time. MTM uses merged-model evaluations as a low-cost signal for selecting branch mixtures, avoiding expensive data-mixture ablations. The procedure is iterative: each round promotes the base model using globally selected merge coefficients and refines each branch mixture using domain-preferred coefficients under constraints that preserve other objectives. To scale beyond simplex grid search, MTM uses qNEHVI-based multi-objective Bayesian optimization. Across code, mathematics, instruction following, and think/non-think control, MTM outperforms naive merging and preserves behavioral separation where single-objective merging collapses, suggesting that effective unified models require branches trained to be mergeable.
[LG-91] Supervise What Decides Success: Criterion-Aligned Auxiliary Losses for Latent World-Model Planning
链接: https://arxiv.org/abs/2610.01224
作者: Takumi Hara,Kanata Suzuki
类目: Machine Learning (cs.LG); Robotics (cs.RO)
*备注: 21 pages, 6 figures, 10 tables. Under review
Abstract:Latent world models plan by scoring candidate action sequences with distances in latent space. However, task success is judged by physical quantities, which we call the success-criterion quantities. In all four latent world models we examine, the end-effector position is encoded in the latent state with an error larger than the success criterion allows. Such a latent state cannot separate successful candidates from failing ones. We propose an auxiliary loss that uses success-criterion quantities as training targets, whereas existing latent world models take them only as inputs. During training, a linear head on the encoder and predictor outputs regresses the success-criterion quantities, and the regression error is added to the training loss. The head is discarded after training, so the model, its cost, and its inputs at test time are unchanged. This loss alone improves the success rate on PushT and cube by 3.5% and 3.4% (absolute), respectively, and both improvements are statistically significant. A success criterion thus specifies what a world model must retain in its latent state, and we show that it can serve directly as a training target.
[LG-92] Autoregressive Drillhole Modelling Under Distribution Shift
链接: https://arxiv.org/abs/2610.01204
作者: Yihao Ding,Daniel Yitian Su,Yiran Zhang,Christopher M. Gonzalez,Wei Liu
类目: Machine Learning (cs.LG)
*备注: work in progress
Abstract:Autoregressive modelling has achieved remarkable success in language and sequence tasks by learning to predict future states from previous observation. Mineral-exploration drillholes provide a natural but largely unexplored setting for this paradigm: as drilling proceeds, lithology is revealed sequentially from shallow to deep, making prediction of deeper strata inherently autoregressive. Existing drillhole modelling, however, is dominated by spatial interpolation and reconstruction, or largely rely on masked modelling, leaving strictly autoregressive prediction largely underexplored. We introduce DrillBench, a benchmark of 49,671 Western Australian drillholes for next-layer prediction and autoregressive stratigraphic generation across a graded transfer spectrum, from local prediction through spatial shift to cross geological province transfer. Benchmarking classical, geostatistical, and neural models reveals a clear \emphtransfer boundary: spatial and geochemical conditioning provides large local gains but deteriorates sharply under stronger shift, whereas lithology-sequence autoregressive models transfer more robustly. Guided by this finding, we develop a backbone-agnostic recipe combining large-scale pretraining on historical drillholes with spatial retrieval of neighbouring lithology. Retrieval is most effective in weathered cover, when local spatial continuity remains informative, whereas pretraining contributes more strongly in bedrock and under broader geological shift. Together, they retain strong local performance while improving generalisation under spatial and cross-province shift, most markedly on the most distant splits. The benchmark and code are available at this https URL.
[LG-93] Low-Budget Active Learning through Entropic Optimal Transport
链接: https://arxiv.org/abs/2610.01199
作者: Rim Hajal,Mathieu Besançon,Jérôme Malick
类目: Machine Learning (cs.LG); Optimization and Control (math.OC)
*备注:
Abstract:We consider low-budget active learning, which consists of selecting a limited number of points, the coreset, such that a model can be trained to high accuracy on the selection only. This problem is particularly relevant in contexts where labeling requires costly expert intervention, as in medical applications. We leverage features extracted from a pretrained self-supervised model to represent the data, and perform coreset selection directly in this feature space. In this paper, we use entropic optimal transport, specifically the Sinkhorn divergence, as the coreset selection criterion, which first allows us to get dimension-free sample complexity results, and second admits computationally efficient gradient evaluations. This opens the way to using gradient-based algorithms to rapidly compute solution candidates, further improved by a swap-based local search, with guarantees on the solution quality. Experiments on image benchmarks and medical datasets show that our method outperforms state-of-the-art heuristics in low-budget settings.
[LG-94] Rethinking the Information Bottleneck: Structured Decomposition under Label-Induced Partitions
链接: https://arxiv.org/abs/2610.01175
作者: Jingyao Zhang,Yuxuan Li,Lu Han,Ali Anaissi,Nguyen H. Tran
类目: Machine Learning (cs.LG)
*备注:
Abstract:Standard information bottleneck (IB) regularization constrains representations via a single scalar I(Z;X), implicitlytreating all information as homogeneous. However, a single global compression control couples label-relevant structurewith residual within-condition variation, rather than regulating their allocation independently, allowing nuisanceinformation to persist in learned representations. For example, in medical imaging applications, residual variation oftenstems from acquisition conditions, background factors, or subject-specific appearance. This issue becomes particularlypronounced in data-limited settings, where models tend to overfit such variation, hindering generalization. While existingregularization methods can stabilize training, control capacity, or shape representation geometry, they do not explicitlyseparate nuisance-like variation from task-supporting structure. To address this limitation, we revisit IB from a structuredperspective based on a label-induced partition, where condition-level structure and within-condition information playdistinct roles. This leads to a dual-bottleneck formulation: a standard KL term controls global information capacity, while aconditional KL term targets within-condition information. We show that the conditional KL admits an exact decompositioninto a within-condition information term and a prior-mismatch term, explaining its alignment with the design this http URL a simplex-structured conditional prior, the method provides controllable latent geometry and integrates seamlesslyinto existing pipelines. Experiments on classification and segmentation show the clearest gains in low-data classificationand consistent improvements across dense prediction benchmarks.
[LG-95] CAGE-NAS: Certified Functional Descent for Efficient Model Growth NEURIPS2026
链接: https://arxiv.org/abs/2610.01173
作者: Santiago Florido Gomez,Stéphane Rivaud
类目: Machine Learning (cs.LG)
*备注: 18 pages, 4 figures, 4 tables. Accepted at AXIOM 2026: Foundations of Efficient Deep Learning (NeurIPS 2026 Workshop)
Abstract:The progressive growth of neural networks requires deciding when the current representation remains sufficient for optimization and when it should be expanded. CAGE-NAS formulates this decision in function space through an admissibility criterion on approximations of the functional gradient. As long as a representation enables a certified Functional Gradient Descent step, the architecture remains fixed; when the criterion fails, a function-preserving expansion is applied and the resulting representation is evaluated again. As the main instance, we study the family induced by the tangent space, using a regularized projection of the functional gradient. In a controlled setting with exact certification, CAGE-NAS produces architectures positioned above the 99.8th performance percentile by held-out RMSE among all admissible alternatives within the same parameter budget, without enumerating them during the growth trajectory.
[LG-96] Learning Rate Transfer for Hybrid Transformer-SSM Architectures NEURIPS2026
链接: https://arxiv.org/abs/2610.01172
作者: Jimin Seo,Gyubok Lee,Yeonsik Jo,Kiwoong Yoo,Yeongoon Kim,Minhae Oh,Jin Woo Koo,Suhwan Kim,Nakyung Lee,Minsik Seol,Idris Nechnech,Jaehyeon Kim,Giho Lee,Jungwoo Lee
类目: Machine Learning (cs.LG)
*备注: Accepted at NeurIPS 2026. 42 pages, 14 figures
Abstract:We study learning rate (LR) scaling for hybrid architectures combining Transformer and State-Space Model (SSM) blocks, a class adopted by several recent production language models. In particular, we focus on the gap between the theoretical scaling rules derived for SSMs under zero-order-hold (ZOH) discretization at infinite width with growing state size, and the field-standard practical implementations using simplified-ZOH Mamba at fixed state size. Surprisingly, in this practical regime hybrid architectures achieve a near-zero LR transfer gap across widths 256-2048 and depths 4-32 up to billion-parameter scale using only the original \mu P prescription, even though SSM operations fall outside its Tensor Programs representability conditions and every parameterization we test fails the standard coordinate-check diagnostic of \mu P correctness. We attribute this to a two-condition decomposition of LR transfer in hybrid architectures: a global update-to-weight invariance, enforced by \mu P’s initialization and LR scaling; and a local per-component balance, provided by AdamW’s per-parameter normalization. Our observations show that the optimal LR is invariant to width up to 8 \times , that this width invariance holds across depth, sequence length, batch size, and Transformer-to-SSM ratio, and that it transfers to Nemotron-H, a production hybrid outside our custom architecture set. We hope these findings fill the gap between theoretical scaling rules and practical hybrid implementations, and stimulate further research toward bridging it.
[LG-97] Latent Information Sharing for Accelerating Federated Learning
链接: https://arxiv.org/abs/2610.01126
作者: Seungjun Lee,Ensieh Khazaei,Dimitrios Hatzinakos,Baturalp Buyukates,Sunwoo Lee
类目: Machine Learning (cs.LG)
*备注:
Abstract:Federated learning (FL) is a communication-efficient distributed learning paradigm. However, client drift remains one of the most critical challenges, hindering the efficient training of a global model. In this study, we propose a novel latent information sharing scheme that directly mitigates data heterogeneity across clients. Our theoretical and empirical results show that sharing a small amount of hidden-layer activations significantly improves training efficiency while preserving convergence guarantees and data privacy. Furthermore, we compare our method with existing FL approaches designed to address client drift, including FedProx, SCAFFOLD, FedPVR, FedProto, and SplitFed, and demonstrate superior model accuracy under a fixed round budget without incurring excessive communication overhead. Overall, this work presents a promising new knowledge aggregation scheme and provides a comprehensive analysis of the impact of activation sharing on federated optimization.
[LG-98] How Much Can Language Models Gain from Test-Time Computation?
链接: https://arxiv.org/abs/2610.01110
作者: Bangji Yang,Jingyuan Li,Jiajun Fan,Yi Evie Zhang,Ruihan Guo,Hongba Ma,Neil He,Chumeng Liang,Qinglong Zheng,Zhanghan Ni,Ge Liu
类目: Machine Learning (cs.LG)
*备注:
Abstract:How much can test-time computation improve a language model, and at what cost? Test-time scaling is widely proposed as a substitute for larger models, but existing comparisons mostly evaluate one domain at a time and rarely charge selection to the budget. We introduce SELF-POT, a benchmark and evaluation framework that measures the test-time potential of a model across competition mathematics, competitive programming, and agentic workflows. SELF-POT separates candidate coverage from final accuracy on static tasks, tracks correctness transitions under revision, and measures protocol completion alongside task success in agentic environments. Under a unified budget rule, it compares Direct inference with parallel sampling and self-revision under fixed multiples of the Direct budget, and charges every model call, including selection and critique, in dollars. This design supports two kinds of comparison: the gain a model obtains from additional inference, and a lower-cost model with additional inference against a stronger model. Across five low-cost reasoning models on 350 sealed tasks, with Claude Opus 5.5 Direct as the reference, the returns depend on the domain, the selection rule, and failure handling. When we replay the retained programming candidate pools, public-example selection raises correct submissions from 376 to 453 of 500 scheduled cells while saving 12-49% of logical API cost across models, and simply retaining an available candidate when judging fails recovers 61 submissions at unchanged cost. On identical mathematics pools, judging with fallback yields 186 correct submissions versus 182 for voting, while voting saves 12-21% of logical API cost. These controlled replays show how selection and failure handling change the gains realized from the same generated candidates, and they quantify the marginal value of a model judge.
[LG-99] GLoC-EHR: Evidence-Cited Clinical Reasoning over Global Context and Local EHR Events
链接: https://arxiv.org/abs/2610.01076
作者: Chaiho Shin,Kwangsoo Kim
类目: Machine Learning (cs.LG)
*备注:
Abstract:Structured electronic health records (EHRs) contain a patient’s clinical trajectory as a sequence of clinical codes. Answering clinical questions from such records requires both the context of the whole trajectory and the specific events that support the answer. We introduce GLoC-EHR, a multimodal language model that reads a contextual encoding of the record through a fixed-size global memory of the trajectory and a local memory of selected events. The model learns to generate hospital-course summaries from the global memory and descriptions of masked concepts from the local memory, aligning both with clinical text. It is then trained to cite evidence before answering, through rationale fine-tuning followed by group relative policy optimization (GRPO) with rewards for correct answers and record-supported evidence. On three MIMIC-IV outcome tasks, GLoC-EHR attains the highest macro AUROC among the compared models when it answers directly, whereas zero-shot LLMs reading the serialized record fall far behind. With evidence-cited reasoning, it stays close to its direct multi-task counterpart in macro AUROC, and the evidence terms of the objective reduce unsupported evidence at a similar macro AUROC. The local memory adds distinct supported findings, particularly under strict matching, without a detectable change in macro AUROC. Without retraining, GLoC-EHR transfers to EHRSHOT on par with EHR-BERT and answers two unseen laboratory questions better than zero-shot prompting of its own backbone.
[LG-100] Kernelized Activation Steering NEURIPS2026
链接: https://arxiv.org/abs/2610.01062
作者: Laziz U. Abdullaev,Minh-Hieu Pham,Bach Do,Khoat Than,Tan M. Nguyen
类目: Machine Learning (cs.LG)
*备注: NeurIPS 2026
Abstract:Activation steering provides a simple, training-free mechanism for controlling attributes of generative models such as sentiment, style, and helpfulness. However, standard approaches such as Difference-in-Means apply a single input-independent steering vector across all activations, limiting expressivity and ignoring the local geometry of the activation space. We propose Kernelized Activation Steering (KAS), a unifying framework that lifts activation steering into a reproducing kernel Hilbert space. KAS formulates steering as an optimization problem expressed purely via kernel evaluations, yielding an implicit, activation-dependent steering score without constructing explicit feature maps. Unlike DiM, KAS induces locally adaptive steering: each activation is modified according to its relative position with respect to source and target reference sets, producing a nonlinear steering field over the representation space. Importantly, DiM is recovered as a special case under a linear kernel, while richer kernels enable geometry-aware interventions. Across standard activation steering tasks, including jailbreaking LLMs and image style control, KAS outperforms or is on par with the existing methods.
[LG-101] SLIM: Simplex-Lattice Interpolation Merging
链接: https://arxiv.org/abs/2610.01037
作者: Seongcheol Jeong,Masahiro Suzuki,Yutaka Matsuo
类目: Machine Learning (cs.LG)
*备注:
Abstract:Optimizing merging coefficients for large language models can require many costly benchmark evaluations. We propose \textbfSimplex-Lattice Interpolation Merging (SLIM), which constructs a quadratic surrogate of aggregate performance on the coefficient simplex using a classical mixture design. Evaluations of individual experts and equal-weight pairs determine the surrogate with the minimum number of measurements needed to identify a general quadratic on this domain. SLIM then optimizes the surrogate without further target-metric evaluations. Experiments on two model architectures demonstrate accurate prediction of unseen multi-expert mixtures and competitive merge performance under limited evaluation budgets. Matched-budget comparisons show that structured evaluation points improve prediction fidelity over random designs, including those using regularized fitting.
[LG-102] Reliability-aware short-term roll prediction for unmanned surface vehicles via multi-task learning and adaptive centralization
链接: https://arxiv.org/abs/2610.00996
作者: Kaizhen Li,Xi Zhou,Zihao Wang,Dan Zhang,Jianjian Liu,Xiaowei Li
类目: Machine Learning (cs.LG)
*备注:
Abstract:Reliable roll prediction of unmanned surface vehicles (USVs) is essential for ensuring navi?gational safety and enhancing autonomous decision-making. While existing studies primarily focus on improving prediction accuracy, the quantification of prediction reliability remains insufficiently addressed. To bridge this gap, this paper proposes a reliability-aware prediction paradigm that integrates confidence assessment into the predictive pipeline. The architecture utilizes a multi-task learning structure where a shared feature extraction backbone feeds into dual heads: a regression head for precise roll prediction and a quantification head for confidence scoring. This configuration provides accurate prediction and corresponding confidence for risk?sensitive downstream tasks. In addition, an adaptive centralization strategy tailored for short?term real-time roll prediction is introduced to improve model generalization under varying operational conditions. Experiments conducted on a real-sea dataset demonstrate that the proposed method effectively quantifies the reliability of prediction results and maintains superior generalization under varying conditions, offering significant potential for practical engineering applications.
[LG-103] Adapter Thickets: Splitting an RLVR Budget Beats Concentrating It
链接: https://arxiv.org/abs/2610.00991
作者: Jonathan Williams,Esin Tureci Karthik R. Narasimhan
类目: Machine Learning (cs.LG)
*备注:
Abstract:Majority voting over sampled completions is the workhorse of test-time scaling, and reinforcement learning with verifiable rewards (RLVR) is the workhorse for making each completion better. The standard pipeline composes the two: train one policy with RLVR, then sample it many times and vote. We show that this composition is lossy. A vote can only overturn mistakes that its voters do not share, and RLVR sharpens a policy so that its samples increasingly make the same mistakes. With every method drawing exactly 160 completions per problem, training a single LoRA adapter on the full RLVR budget raises single-sample accuracy on every model we test ( 1.5 B- 8 B). Yet on three of four models it leaves the majority vote below that of the untrained base model, by up to 4.8 points. The damage builds during training: voter errors grow steadily more correlated, and the majority vote accuracy peaks early before falling by up to 7.0 points. The cause is concentration, not RLVR itself. We split the same data and training budget across K LoRA adapters, each trained on its own random disjoint shard, and call the result an adapter thicket. Thickets out-vote the fully trained adapter in all 16 (model, K ) settings, and for K\geq4 they stay within 0.8 points of the base model or above it. A single adapter stopped early, at a thicket member’s step count, is a strong control that matches thickets for small K . For K\geq8 , thickets keep more of RLVR’s single-sample gain and out-vote this control in six of eight settings. The cost of concentration also grows with the number of votes: from 16 to 160 votes, the thicket’s lead over the fully trained adapter widens from 1.3 to 3.3 points. When the plan is to sample and vote, an RLVR budget is better spent broad than deep.
[LG-104] Auditable Algebraic Counting Field for Cryptic-Pocket Detection from Apo Structures
链接: https://arxiv.org/abs/2610.00988
作者: Shan Yu,Xuening Wu
类目: Machine Learning (cs.LG); Biomolecules (q-bio.BM)
*备注: Accepted for publication in Pacific Symposium on Biocomputing (PSB) 2027
Abstract:Cryptic ligand-binding pockets are not apparent in experimentally determined apo structures, making them difficult to identify from unbound receptor geometry. A complementary challenge is to make the structural measurements and learned evidence behind each prediction directly inspectable. We introduce a supervised algebraic counting field (ACF) for predicting cryptic-pocket residues from apo structures. ACF compiles explicit geometric, physicochemical, and topological features into compact, integer-weighted lookup tables. Each prediction score can be reconstructed from feature values, training counts, table weights, and spatial aggregation, without sequence search, structural-template transfer, or a protein language model at inference. We evaluate ACF on CryptoBench and two locked external collections, separating ranking performance from the effects of residue-calling budgets. On an external set of 57 post-CryptoBench apo-holo units, ACF exceeded P2Rank by +0.044 in mean paired ROC-AUC (multiplicity-adjusted 95% CI [+0.010, +0.079]). The advantage was dataset-dependent: official-fold ROC-AUC and matched-budget F1 differences against P2Rank remained unresolved, and a second external evaluation did not confirm gains from added structural features. ACF thus provides a compact predictor with externally validated signal and an inspectable path from structural measurements and training counts to residue scores.
[LG-105] Neural scaling laws and evolution of learnable activation functions of Kolmogorov-Arnold networks
链接: https://arxiv.org/abs/2610.00985
作者: Tilen Cadez,Sanghoon Lee,Kyoung-Min Kim
类目: Machine Learning (cs.LG)
*备注: 13 pages, 10 figures. Supplementary Notes will be provided in the published version
Abstract:Kolmogorov-Arnold Networks (KANs) represent a compelling alternative to traditional Multi-Layer Perceptron (MLP)-based neural networks. By employing activation functions as learnable elements, KANs offer superior interpretability, making them suited for scientific domains. In this work, we investigate the neural scaling laws of KANs and the structural evolution of their learnable activation functions under dataset expansion. Specifically, we evaluate the scaling behavior of three KAN variants—BSRBF-KAN, Gottlieb-KAN, and Faster-KAN—across standard image classification benchmarks (MNIST and Fashion-MNIST) and a specialized scientific regression task (magnetic parameter estimation from domain images of moiré magnetic textures). Our results demonstrate that the test loss \cal L exhibits a broken neural scaling law (BNSL) behavior as a function of the dataset size N_D . After passing through a random-guess regime, the loss follows architecture- and task-dependent scaling behavior. The loss crosses from a faster- to a slower-scaling branch, \cal L\propto N_D^-\alpha and \cal L\propto N_D^-\beta with \alpha\beta for image classification tasks. The exponents \alpha and \beta depend strongly on both the specific network architecture and the dataset-size regime, ranging from 0.4 to 1.5 and from 0.06 to 0.6, respectively. For the magnetic parameter-regression task, the loss follows a single scaling law with its exponent ranging from 1.28 to 2.59. Additionally, we provide a structural analysis of how activation functions refine their complexity as data volume increases, finding that dataset expansion drives a transition from simple linear-like approximations toward stable, interpretable symbolic forms. These findings provide a quantitative roadmap for the efficient application of KANs while managing the trade-off between model expressivity and computational overhead.
[LG-106] HADRec: A Hierarchy-Aware Drug Recommendation Framework by Fusing Molecular Knowledge and Electronic Health Record
链接: https://arxiv.org/abs/2610.00984
作者: Junke Wang,Hongshun Ling,Li Zhang,Jinjing Wu,Tong Shao,Fang Wang,Yuan Gao
类目: Machine Learning (cs.LG)
*备注:
Abstract:Accurate medication recommendation is central to clinical decision-making, directly determining therapeutic efficacy and patient safety. However, existing methods suffer from two key limitations: drugs are often abstracted as discrete tokens, ignoring their molecular structures and pharmacological mechanisms, and the commonly used “flat” recommendation paradigm fails to leverage the hierarchical logic of the internationally standardized Anatomical Therapeutic Chemical (ATC) classification system. To address these issues, we propose HADRec, a Hierarchy-Aware Drug Recommendation framework that integrates molecular knowledge with electronic health records (EHRs). HADRec employs LLaMA-7B to encode clinical notes for rich patient representations and ChemBERTa to encode drug Simplified Molecular Input Line Entry System strings, building a global molecular knowledge base. A cross-attention mechanism then performs deep multimodal fusion between patient states and drug features. The framework further incorporates a hierarchical predictor and a novel consistency constraint loss to enforce strict adherence to ATC logical dependencies. Extensive experiments on MIMIC-III demonstrate that HADRec achieves state-of-the-art performance across Jaccard, F1, and PR-AUC. External validation on MIMIC-IV confirms strong generalization under distribution shifts, and calibration analysis shows well-calibrated predictive confidence on MIMIC-IV with ECE = 0.04, and Brier = 0.06. Counterfactual evaluation reveals clinically aligned reasoning, disentangling disease-specific treatments from general care. Together, these results establish HADRec as a high-performance, interpretable, and clinically grounded pathway toward safe and reliable AI-driven medication recommendation.
[LG-107] Variational Streaming Flow: Probabilistic Forecasting in Physical Time
链接: https://arxiv.org/abs/2610.00976
作者: Hans Hao-Hsun Hsu,Minseon Gwak,Soon Hoe Lim,Pan Li,N. Benjamin Erichson
类目: Machine Learning (cs.LG)
*备注:
Abstract:Probabilistic forecasting is important for predicting complex dynamical systems because intrinsic randomness and incomplete observations can cause the same observed state to evolve into multiple plausible futures. While flow matching is a flexible approach for probabilistic forecasting, it is computationally expensive. Streaming flow (SF) reformulates this approach to model temporal evolution efficiently by learning a continuous velocity field directly in physical time. However, SF learns a deterministic velocity field. Thus, it provides only a single future trajectory for a given fixed initial state and observation history. To overcome this limitation, we introduce Variational Streaming Flow (VSF). Our approach learns a latent distribution that is conditioned on the dynamics of interest. In turn, this enables probabilistic forecasting. Importantly, we retain the computational efficiency of SF by generating in physical time. Across deterministic and stochastic dynamical systems, VSF demonstrates superior predictive accuracy and distributional fidelity. We demonstrate the advantage for both long-horizon rollouts exceeding 1,000 steps, and settings with bifurcating dynamics. Moreover, VSF can be integrated into existing Joint-Embedding Predictive Architecture (JEPA)-based world models as a plug-and-play predictor to improve temporal dynamics and goal-directed success rate in navigation, motion planning, and manipulation.
[LG-108] Rate-Optimal Algorithm for Adversarial Linear CMDPs
链接: https://arxiv.org/abs/2610.00927
作者: Kihyun Yu,Honghao Wei,Dabeen Lee
类目: Machine Learning (cs.LG)
*备注:
Abstract:We study episodic adversarial linear constrained Markov decision processes (CMDPs) with unknown transitions, where both the loss and constraint functions may vary adversarially across episodes. The best previous algorithm achieves \widetilde\mathcalO(K^3/4) regret and cumulative constraint violation, leaving a gap to the optimal \widetilde\mathcalO(\sqrtK) dependence on the number of episodes K . We close this gap by proposing a new primal dual algorithm that achieves \widetilde\mathcalO(\sqrtK) regret and cumulative constraint violation without assuming Slater’s condition. The main challenge is that learning linear CMDPs requires uniform concentration over a value function class with a controlled covering number, whereas standard techniques in constrained online learning, such as policy mixing, can make this class more complex. Our algorithm combines adaptive Follow the Regularized Leader (FTRL), contracted value estimation, and an exponential Lyapunov function. An adaptive dual regularizer offsets the dependence on the dual weights in the primal regret bound, removing the need for policy mixing. We further show that the normalization in the FTRL update bounds the policy parameters independently of the magnitudes of the dual weights, which explains why the resulting policy class remains compatible with uniform concentration. Under feature access, the computational complexity is independent of the size of the state space.
[LG-109] In CEM a World Model Is Also a Proposal Mechanism
链接: https://arxiv.org/abs/2610.00921
作者: Oliver Obst,Frieder Stolzenburg
类目: Machine Learning (cs.LG); Robotics (cs.RO)
*备注: 20 pages, including 12 pages appendix
Abstract:The cross-entropy method (CEM) uses world-model scores to select action sequences and fit the distribution sampled in its next iteration. A scoring error can therefore change both the present decision and the candidates considered later. We evaluate these two roles separately. Four types of predictive model generate CEM traces, and every model rescores every saved candidate pool. Executing the same candidates in the environment provides a reference elite set and proposal update. Across twelve independently trained task-seed units on Walker and Cheetah, the pre-specified proposal distance falls from the first to the final CEM iteration in every unit. Proposal widths contract and fitted means separate relative to the remaining search width. Pairwise ranking agreement stays near chance on Walker and declines on Cheetah; elite-set agreement does not improve. This comparison shows greater variation between scorers than between pool sources on Cheetah; Walker has variation in both and in their pairings. We use the original six units to select Random nonlinear for a one-update intervention, without inspecting intervention outcomes. Replacing its first model-ranked update with an environment-ranked update lowers final realised selected-sequence cost in those six units and in six further units held out from the selection. Comments: 20 pages, including 12 pages appendix Subjects: Machine Learning (cs.LG); Robotics (cs.RO) Cite as: arXiv:2610.00921 [cs.LG] (or arXiv:2610.00921v1 [cs.LG] for this version) https://doi.org/10.48550/arXiv.2610.00921 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[LG-110] STEER: Reducing Inference Cost in Relational Foundation Models through Semantically Informed Sampling
链接: https://arxiv.org/abs/2610.00907
作者: Abdalla Mohamed,Ashraf Aboulnaga
类目: Databases (cs.DB); Machine Learning (cs.LG)
*备注:
Abstract:Relational foundation models (RFMs) are pretrained once on a collection of relational databases and prediction tasks, and then applied zero-shot to previously unseen databases and tasks. To make a prediction for a target row, an RFM samples a neighborhood of rows linked to that row through foreign keys and uses this neighborhood as its inference context. Lowering inference cost is an important goal for any foundation model, and for RFMs this cost grows with the size of the context. The simplest ways to shrink the context is to drop some of the sampled rows, but this ignores the semantics of the database schema, so it is as likely to discard informative rows as uninformative ones. We propose STEER, a sampling approach that shrinks the inference context by concentrating it on the tables most relevant to the prediction task at hand. STEER obtains relevance information by prompting a large language model to rank the foreign-key edges of the database schema into relevance tiers for the given task, and then maps each tier to a probability of following that edge during traversal. Because the ranking uses only the schema, it is computed once per task and reused across all subsequent predictions, amortizing its cost. We evaluate STEER on three state-of-the-art RFMs (RT, RT-J, and Griffin) and show that it reduces inference context size by about 40% on average while maintaining, and in some cases improving, prediction accuracy.
[LG-111] Sharpen Before You Adapt: Data-Free Entry-State Sharpening for Test-Time Reinforcement Learning
链接: https://arxiv.org/abs/2610.00903
作者: Zhanming Zhang,Vinoth Selvendran
类目: Machine Learning (cs.LG)
*备注: 10 pages, 2 figures, 3 tables
Abstract:Test-time reinforcement learning (TTRL) adapts language models on unlabeled test problems using supervision derived from their own samples. This makes the checkpoint’s \emphentry state consequential: a diffuse policy provides noisier self-supervision and may spend much of a limited adaptation budget merely concentrating probability mass before reliably expressing capability it already possesses. We propose \textbfentry-state sharpening: use data-free training \emphbefore TTRL to prepare a general-purpose checkpoint in a state that subsequent label-free adaptation can exploit more efficiently. The idea is not tied to one training recipe; different data-free objectives can move the same base model to different entry states. Across five data-free checkpoints derived from Qwen3-4B and evaluated under an identical 15-step TTRL protocol, entry policy entropy strongly rank-orders endpoint conversion efficiency, a reliability-to-reachability measure (Spearman \rho=-0.90 ; \rho=-0.99 after controlling for entry reachability). The contrast across objectives is striking: R-Zero remains diffuse at 3.39 nats and finishes below the untuned base in 6/6 matched comparisons across MATH, GPQA, and AMC, whereas SPIRAL reaches 0.07 nats and achieves the highest post-TTRL accuracy on MATH and GPQA despite its self-play stage using no math training data. An in-domain label-free self-distillation intervention further shows that the entry state can be deliberately sharpened. These results motivate treating checkpoint preparation as a \emphstate-control problem: use data-free training to improve TTRL readiness, with entry entropy as a label-free control signal and reachable capability as the constraint.
[LG-112] When Do Biological Reasoning Models Use Their Biological Inputs?
链接: https://arxiv.org/abs/2610.00898
作者: Ada Fang,Nikitha Thoduguli,Lukas Fesser,Hanlin Zhang,Sham M. Kakade,Marinka Zitnik
类目: Machine Learning (cs.LG)
*备注:
Abstract:Biological reasoning models use post-training to connect LLMs to biological foundation model representations and biological text. Their benchmark accuracy is taken as evidence that LLMs reason over these inputs. We test this assumption in six biological reasoning models across DNA, protein, and single-cell tasks. We perturb one biological input while holding the query and other inputs fixed, construct evidence conflicts that pair the foundation model representation of one genome, protein, or cell with the text of another, fit linear probes to the representations the language model receives, and analyze reasoning traces against the biological inputs. Evo2 and ESM3 contribute little to BioReason and BioReason-Pro performance on the evaluated tasks. Shuffling the DNA sequence barely changes BioReason disease prediction accuracy, and in evidence conflicts the two models follow the text in 97.9% and 99.7% of cases. Linear probes trained on the Evo2 and ESM3 representations predict the task targets, so these foundation models encode information relevant to the task, but provide limited overall performance improvement to BioReason and BioReason-Pro. In contrast, foundation model inputs contribute to ChatNT, Prot2Text-V2, and CellWhisperer performance, and differentially expressed genes in the gene sentence contribute to Cell2Sentence-Scale performance. Across SFT and RL checkpoints of BioReason-Pro and 42 BioReason checkpoints, increases in accuracy do not imply greater performance contributions from biological inputs. BioReason traces misstate nucleotide changes, while BioReason-Pro traces describe functions omitted from final predictions under evidence conflicts. We find that current post-training strategies do not ensure that foundation model representations contribute to task performance.
[LG-113] Clock Diffusion: Efficient Semi-Autoregressive Continuous Diffusion Language Models
链接: https://arxiv.org/abs/2610.00894
作者: Yair Schiff,Omer Belhasin,Roy Uziel,Matan Rusanovsky,Ran Zilberstein,Marianne Arriola,Gilad Turok,Guanghan Wang,Volodymyr Kuleshov,Michael Elad
类目: Machine Learning (cs.LG)
*备注:
Abstract:Recent works on continuous diffusion for discrete data have demonstrated performance on par with comparable discrete diffusion models. However, these continuous counterparts lack key features that are essential to practical use as language models, namely variable-length generation and support for a key-value cache, and they still lag behind the frontier of autoregressive and discrete diffusion quality. In this work, we address these limitations. We do so by introducing a model parameterization that uses position-dependent noise schedules to define semi-autoregressive (SAR) continuous diffusion language models (DLMs). Together with efficient training and sampling algorithms, we call this framework Clock Diffusion, and we present two special cases of our method: block and sliding window generation. We then define ClockDLMs, a family of Gaussian DLMs based on sliding window Clock Diffusion that attain state-of-the-art diffusion likelihood bounds on OpenWebText, even beating the performant block SAR discrete diffusion models. ClockDLMs trained on TinyGSM also substantially outperform continuous baselines on the GSM8K benchmark and match and exceed comparable SAR discrete diffusion models. Finally, building on our parameterization, we propose more efficient samplers that we dub Cache Grab, which adapt techniques from accelerated inference in discrete diffusion, such as committing tokens whose probabilities exceed a confidence threshold and self-speculative decoding, further improving our models’ quality and efficiency.
[LG-114] Fixing a Model That Learned Worse Cancer Means Lower Risk: Monotonic Constraints in Bladder Cancer Recurrence Prediction
链接: https://arxiv.org/abs/2610.00858
作者: Saram Abbas,David Thomas,Naeem Soomro,Rishad Shafik,Rakesh Heer,Kabita Adhikari
类目: Machine Learning (cs.LG)
*备注: 13 pages, 4 figures, 1 table. Supplementary material (15 pages) provided as ancillary files
Abstract:Background and Objective: Clinicians expect recurrence risk to climb with cancer severity. In a UK multicentre trial, an unconstrained XGBoost model learnt that higher tumour stage and carcinoma in situ predicted lower recurrence risk, and discrimination, calibration, and SHAP were all blind to it. We developed a counterfactual testing framework to detect this inversion and a monotonic-constraint framework to remove it without hurting performance. Methods: BOXIT enrolled 472 patients with protocol-mandated cystoscopy across 51 UK sites (2007-2012); 435 had at least two years’ follow-up (153 recurrences, 35.2%). We developed a counterfactual direction test and a monotonic-constraint correction, with constraint directions drawn from the EORTC and EAU risk systems, and evaluated both against unconstrained XGBoost and logistic regression on 18 predictors (seven directed) over 50 cross-validation folds. The test worsened each patient on one directed feature at a time to check whether risk fell; SHAP direction and calibration were also assessed. Key Findings and Limitations: Tumour stage and carcinoma in situ were associated with lower recurrence, opposite to medical intuition; the unconstrained model reversed carcinoma in situ counterfactuals in 90.2% of cases and stage in 74.3%. Discrimination ( \Delta AUC 0.005, p=0.47), calibration, and SHAP magnitude were all blind to the inversion. Monotonic constraints eliminated every violation at no cost to discrimination (0.723 vs 0.718) and outperformed EORTC (p=8.9e-16). Limitations: single trial, internal-external validation only. Conclusions and Clinical Implications: A model that had learned this inversion passed every conventional check. A counterfactual direction test, run as a single refit with pre-specified monotonic constraints, catches this failure at no cost to performance and should be routine before clinical deployment.
[LG-115] rueMuse: A Benchmark for Data Attribution in Text-to-Music Models
链接: https://arxiv.org/abs/2610.00835
作者: Jiawei Yu,Jian Liu
类目: Machine Learning (cs.LG)
*备注:
Abstract:Text-to-music generation models are trained on massive music collections, creating a growing need for data attribution methods that can quantify the contribution of individual training samples. However, existing attribution methods are difficult to rigorously evaluate due to the lack of reliable ground truth, making it challenging to reliably assess their actual effectiveness. To address this gap, we introduce TrueMuse, a controlled dataset and benchmark for text-to-music data attribution. TrueMuse is constructed by fine-tuning three diffusion-based text-to-music models on carefully curated attribution samples, whose known inclusion in fine-tuning provides controlled attribution targets for evaluation. The benchmark covers four attribution settings, spanning melodic structure, timbral characteristics, artist-level stylistic signatures, and genre-level shared patterns, and includes 133 attributes, 648 fine-tuned models, and 95,456 generated samples across two prompt types. Using TrueMuse, we systematically evaluate existing black-box attribution methods along four dimensions: fine-tuning improvement, prompt-type difficulty, multi-task training, and fine-tuning data size. Our results show that attribution remains challenging, with existing methods exhibiting substantial variation across evaluation settings, highlighting the need for more reliable and generalizable attribution methods for text-to-music generation. Code and Dataset will be released upon acceptance.
[LG-116] AnyJev Technical Report
链接: https://arxiv.org/abs/2610.00831
作者: Jiamu Zhang,Tianze Yang,Yucheng Shi,Evan Chen,Zixiang Nie,Kelly Wan,Liangjie Hong,Ninghao Liu,Liang Wu
类目: Machine Learning (cs.LG)
*备注: 22 pages, 5 figures. Early report on work in development. Code: this https URL
Abstract:A typed decision is a choice among a fixed set of options, returned as a probability rather than as text. Systems that need typed decisions today use models trained for that purpose. This report describes AnyJev, which reads a typed decision from one prefill of a pretrained instruction-tuned language model. The readout restricts the next-token distribution at the answer position to the option tokens. It has two defects: the model assigns higher probability to some labels whatever the input, and to some positions in the option list. AnyJev corrects both with no gradient steps and no parameter changes: it divides out a label prior estimated from unlabelled inputs, and it averages log-probabilities over the K cyclic rotations of the option list. On two 20-option tasks the rotations lower the order-flip rate from 0.33 to 0.14 and from 0.33 to 0.18, and raise accuracy on 11 of 11 models on both. Reading every rotation requires K prefills. A stopping rule selected against the full-rotation decision on unlabelled states cuts that. Selecting the threshold on one unlabelled split and bounding its disagreement on a second, it reads 10.6 rotations of 18 at a verified 0.008 bound on two of four cells; selected and bounded on one split, as our serving run did, it reads 7.3 and serves 2.2 times as many decisions per second on vLLM. The code is open source.
[LG-117] Quantifying the Impact of Ambulance Ramping: A Multi-Year Analysis of Victorian Emergency Medical Services Cases
链接: https://arxiv.org/abs/2610.00818
作者: Ayesha Tanveer,Khandakar Ahmed,Assefa Teshome,Oyetunde Gbadeyan,Ziad Nehme
类目: Machine Learning (cs.LG)
*备注:
Abstract:Ambulance ramping, the delay between hospital arrival and patient handover, is a critical operational bottleneck in Emergency Medical Services (EMS), yet its systemic magnitude and dynamics remain inadequately characterised at scale. This paper quantifies the scale, trajectory, and operational correlates of ramping across an entire statewide EMS system, analysing 2,850,575 ambulance attendances in Victoria, Australia from January 2020 to March 2024 using an Exploratory Data Analysis (EDA). After systematic preprocessing, an analytical cohort of 2,026,569 Emergency Department (ED) transports across 59 hospitals with ED and 79 Local Government Areas (LGA) was examined through interval decomposition, Pareto concentration, hourly cross-correlation, hospital arrival concurrency and priority-stratified operational comparisons. Cumulative Ambulance Hours Lost (AHL) totalled 1,491,127 hours, equivalent to approximately 96 ten-hour ambulance shift lost every day of the study window. Ten of 59 hospitals account for 57.8% of lost hours from 50.9% of cases. Annual losses rose 57% to a 2022 peak while transported demand fell 3.7%, indicating deterioration in per-case handover rather than growth in demand. Handover duration varies little with patient acuity, but rises monotonically with the number of ambulances arriving at the same hospital in the preceding hour, an effect persisting within every hour of the day. Hourly demand is moderately associated with ramping two to four hours later (r = 0.365). These findings establish the empirical preconditions for hospital-state aware ambulance routing.
[LG-118] Learning Goal-Reaching Quasimetric Geometry From Finite-Time Reachability
链接: https://arxiv.org/abs/2610.00778
作者: Daisuke Yamada,Travis Pence,Vikas Singh
类目: Machine Learning (cs.LG)
*备注:
Abstract:In goal-conditioned reinforcement learning (GCRL), quasimetric learning models goal-reaching costs as quasimetric distances, connecting local constraints to global value geometry. Its local constraints, however, should reflect the direction- dependent effects of control composition over a finite horizon together with environmental feasibility. We propose ReQRL, which constrains the critic’s value gradients through finite-horizon reachability. Drawing on state-constrained optimal control, we decouple dynamical reachability from boundary geometry, estimating both from data. On OGBench, our method outperforms or rivals existing quasimetric approaches and other offline GCRL methods.
[LG-119] Localizing Transfer Between Memorization Tasks NEURIPS2026
链接: https://arxiv.org/abs/2610.00771
作者: Yimiao Yu,Florentin Guth
类目: Machine Learning (cs.LG)
*备注: 17 pages, 10 figures. Accepted as a poster at the NeurIPS 2026 Workshop on Neural Network Artifacts as a New Data Modality
Abstract:A central puzzle in transfer learning is why pre-training on one task can accelerate training or improve performance on another task, and what mechanisms underlie this transfer. In this work, we examine the transfer between memorization tasks of random input-output mappings. We find two surprising transfer patterns: equivalent transfer, where each additional pre-training epoch saves approximately one downstream fine-tuning epoch; and non-equivalent transfer, where pre-training on a mismatched task can be even more efficient than directly training on the downstream task itself. Through ablation experiments, we decompose and localize the transfer into two separate effects: a “trivial” magnitude-driven transfer in the last layer, and a “non-trivial” structure-driven transfer, partially attributable to the covariance of the other layers. These results advance our understanding of the underlying mechanisms of transfer learning and have the potential to lead to principled pre-training strategies.
[LG-120] Crossing the Cyber Divide: Sim-to-Sim and Sim-to-Real Transfer for RL Agents
链接: https://arxiv.org/abs/2610.00759
作者: Sabrina Saika,Yinuo Du,Aritran Piplai
类目: Cryptography and Security (cs.CR); Machine Learning (cs.LG)
*备注: 13 pages, 3 figures, 1st Workshop on Real-world AI Security and Engineering for Cybersecurity Systems (RAISE) 2026
Abstract:Cyber attack agents are typically trained and evaluated within a single simulator, making it unclear whether learned policies transfer beyond the environments in which they were developed. This limitation hinders both deployment and fair comparison, as cyber simulators differ substantially in their state representations, observation models, and action spaces. In this paper, we study policy transfer across cyber environments and argue that simulator-to-simulator and simulator-to-real transfer can be viewed as instances of the same underlying alignment problem. We propose a framework that separates state alignment from action translation, enabling a policy trained in one environment to operate in another without retraining. We evaluate transfer across four cyber platforms, CyberBattleSim, NetSecGame, CyberWheel, and NASim, including emulated deployments in NASim. Our experiments show that zero-shot transfer is feasible, fully preserving source-policy performance in closely aligned environments and achieving 45.2% win rates when transferring policies whose source performance is 60.5%. In emulated virtual machine environments, transferred policies exhibit a Jensen-Shannon divergence of 0.085 from native policies, indicating strong behavioral similarity. Code and benchmarks are available at: this https URL.
[LG-121] Scalable Multi-Task Inverse Reinforcement Learning
链接: https://arxiv.org/abs/2610.00758
作者: Allen Tran,Jia Wan,Nathan Kallus,Aurélien Bibaut
类目: Machine Learning (cs.LG); Machine Learning (stat.ML)
*备注:
Abstract:By learning transferable rewards, inverse reinforcement learning (IRL) enables counterfactual evaluation of agents under modified environments. Such transfer places strict requirements on coverage since target environments affect agents’ state occupancy. We propose a multi-task IRL method that pools data across multiple agents with different rewards in the same environment under a low-rank assumption. In addition to alleviating coverage requirements, so each task need not visit every state as long as others do, the method offers scalable evaluation of multiple tasks under new environments as computationally intensive planning scales with rank rather than the number of tasks. We provide finite sample guarantees on reward recovery and on policy learning in new environments. Experiments show our method is robust to limited coverage, recovers rewards on and off of each task’s support, transfers to target environments at lower regret than baselines, with its computational advantage over per-task methods widening as tasks grow.
[LG-122] Reformulation-Contrastive Learning for Mixed Integer Programs
链接: https://arxiv.org/abs/2610.00730
作者: Ousema Bouaneni,Mathis Le Bail,Clément Elliker,Maël Jenny,Sonia Vanier
类目: Machine Learning (cs.LG)
*备注:
Abstract:Mixed-integer linear programs (MILP) model many real-world decision problems, motivating machine-learning methods that exploit recurring structure to accelerate MILP solving. MILPs can admit many equivalent formulations: integrality-preserving changes of variables and the addition of redundant constraints can alter their formulations while preserving the optimization problem. We leverage these reformulations as a source of self-supervision for learning general-purpose representations of MILP variables and constraints. We characterize the affine reformulations that are valid for every input instance, and distinguish re-descriptions, which leave variables unchanged, from substitutions, which transform them predictably. Building on equivariant self-supervised learning, we introduce ReMILP (reformulation-contrastive MILP representation learning), which jointly trains a graph neural network and a hypernetwork to predict how variable embeddings transform under changes of variables. Without solver-derived labels, ReMILP learns representations that exhibit the intended invariance and equivariance on unseen problem classes. Across binary solution, constraint activity and integrality gap prediction, these representations carry task-relevant information when frozen and provide a useful initialization for fine-tuning.
[LG-123] Reward as Observation: Learning Reward-Based Policies for Rapid Adaptation
链接: https://arxiv.org/abs/2610.00729
作者: Morgan Byrd,Maks Sorokin,Robert Wright,Sehoon Ha
类目: Machine Learning (cs.LG); Robotics (cs.RO)
*备注: Website: this https URL
Abstract:This paper explores a reward-based policy to achieve zero-shot transfer between source and target environments with completely different observation spaces. While humans can demonstrate impressive adaptation capabilities, deep neural network policies often struggle to adapt to a new environment and require a considerable amount of samples for successful transfer. Instead, we propose a novel reward-based policy only conditioned on rewards and actions, enabling zero-shot adaptation to new environments with completely different observations. We discuss the challenges and feasibility of a reward-based policy and then propose a practical algorithm for training. We demonstrate that a reward policy can be trained within three different environments, Pointmass, Cartpole, and 2D Car Racing, and transferred to completely different observations, such as different color palettes or 3D rendering, or Stretch robot navigation in Habitat-Sim, in a zero-shot manner. We also demonstrate that a reward-based policy can further guide the training of an observation-based policy in the target environment.
[LG-124] CF-JEPA: Improving Robustness of JEPA World Models via Controllability Factorization
链接: https://arxiv.org/abs/2610.00727
作者: Morgan Byrd,Robert Wright,Sehoon Ha
类目: Robotics (cs.RO); Machine Learning (cs.LG)
*备注: Website: this https URL
Abstract:Controlling an agent with vision requires being able to separate useful information from irrelevant background information. JEPA-style latent world models seem like a natural approach for this, as they do not perform pixel-level reconstruction; however, they are still sensitive to these distractor signals and experience latent collapse. In this work, we introduce Controllability Factorized JEPA (CF-JEPA), a JEPA-style world model which splits the latent space into controllable and uncontrollable subspaces. This factorization allows us to capture all the distractor information into the uncontrollable region, while we use the control-relevant latent information for our task. With this, we show comparable performance across 2D and 3D control tasks under nominal conditions and improved performance under distracted conditions, where CF-JEPA is the only model that does not experience latent collapse. We also validate our model under distracted conditions for a simulated robot task, highlighting the practical application of such a scheme.
[LG-125] JEPA-TTT: Persistent Test-Time Training of Latent World Models for Planning under Dynamics Shifts NEURIPS2026
链接: https://arxiv.org/abs/2610.00722
作者: Zheyuan Zhang,Suyu Ye,Nakul Agarwal,Hossein Nourkhiz Mahjoub,Ehsan Moradi Pari,Daniel Khashabi,Tianmin Shu,Vaishnav Tadiparthi
类目: Machine Learning (cs.LG)
*备注: Accepted to World Models in Physical AI Workshop @ NeurIPS 2026 | Project page: this https URL
Abstract:World models enable agents to plan by predicting future states of the environment, but their predictions can become unreliable when test-time dynamics differ from those seen during training. We present JEPA-TTT, which adapts the latent dynamics predictor of a pretrained action-conditioned Joint-Embedding Predictive Architecture world model throughout test time. Self-supervised updates accumulate across episodes, while the visual encoder and reward head remain fixed, preserving the pretrained representation and task objective. Planning requires neither a goal image nor online environment reward. JEPA-TTT uses dense replay, which forms prediction windows at every temporal offset, retains them in a growing buffer, and samples minibatches from that buffer for predictor updates. Across eight dynamics shifts in four continuous-control environments, JEPA-TTT improves planning on every shift. After 500 test-time episodes, it reduces autoregressive latent prediction error by 83% on average and improves planning performance by 153% over the frozen JEPA world model. These results show that persistent self-supervised test-time training can adapt a pretrained latent world model under changed dynamics.
[LG-126] WOMBAT: Whitebox Oracle for Molecular Benchmarking and Attribution Testing
链接: https://arxiv.org/abs/2610.00713
作者: Dominik Matuszek,Bartosz Zieliński,Tomasz Danel,Dawid Rymarczyk
类目: Machine Learning (cs.LG)
*备注:
Abstract:When a graph neural network (GNN) explainer produces an unexpected attribution on a molecule, the attribution alone cannot reveal whether the explainer has failed or the model has learned a shortcut. We introduce WOMBAT, a benchmark of 14 whitebox GNNs, each with message-passing weights set by hand to detect a specific SMARTS motif. Each model’s decision rule is known by construction, providing attribution ground truth against which explainer errors can be identified and studied. We validate the models on millions of PubChem molecules and evaluate post-hoc explainers including GNNExplainer, PGExplainer, and Integrated Gradients. Guided by our qualitative analysis, we construct a model that causes Integrated Gradients to spread attribution across the graph, even though the model reliably detects the intended motif. We release the dataset, models, and evaluation code to help researchers in the development of newer XAI tools for GNNs.
[LG-127] Beyond Unimodal Bases: Pullback Geometry for Multimodal Data
链接: https://arxiv.org/abs/2610.00708
作者: Honglei Brinkmann,Lucas Ng,Georgios Batzolis,Mark Girolami,Carola-Bibiane Schönlieb,Willem Diepeveen
类目: Machine Learning (cs.LG)
*备注:
Abstract:Data-driven Riemannian geometry provides nonlinear interpolation and geometric representations of high-dimensional data. For these operations to be statistically meaningful, paths between observations should preferentially traverse high-likelihood regions. Existing scalable pullback constructions typically use a unimodal Gaussian latent distribution, assuming that the data reside close to a single manifold. For multimodal data, mapping separated modes or local structures into one Gaussian region can require substantial transport deformation and compromise the resulting geometry. We introduce a pullback geometry for data supported on mixtures of manifolds. Using a latent Gaussian mixture, we define its Riemannian metric as the matrix square of the responsibility-weighted expected component precision. The metric is smooth and positive definite and recovers the existing Gaussian construction in the single-component limit. For structured overlapping mixtures, we establish conditions under which the log-density is concave along geodesics, providing a formal connection between the proposed geometry and paths through high-likelihood regions, and derive the corresponding local curvature relations. We instantiate this geometry in a normalizing flow with adaptive mixture learning, allowing the number of active components to emerge from the data and supporting component-wise reconstruction and local effective-dimension estimation. Experiments on synthetic geometric data, a controlled multi-view image setting with a known reference trajectory, and MNIST show reduced transport distortion, competitive path support, close reference-trajectory recovery, and improved interpolation realism. These results extend scalable pullback geometry beyond datasets that reside close to a single manifold while retaining tractable and interpretable local structure. Subjects: Machine Learning (cs.LG) Cite as: arXiv:2610.00708 [cs.LG] (or arXiv:2610.00708v1 [cs.LG] for this version) https://doi.org/10.48550/arXiv.2610.00708 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[LG-128] AnchorPrompt: Self-Distilled Soft Prompts for Robust Audio-Language Models
链接: https://arxiv.org/abs/2610.00706
作者: Pooneh Mousavi,Amir Ivry,Mirco Ravanelli,Cem Subakan
类目: ound (cs.SD); Machine Learning (cs.LG)
*备注:
Abstract:Large audio-language models (LALMs) are sensitive to input perturbations, such as noise, waveform corruption, and adversarial injections. We propose AnchorPrompt, an efficient adaptation method that keeps the model frozen and learns a single block of prompt vectors inserted at the decoder input, between the audio and question embeddings. We train these vectors through self-distillation over diverse audio and text perturbations. To improve answer consistency and mitigate hallucination, we use the model’s prediction on the clean recording as the target for answerable inputs, and assign a refusal target when the audio lacks sufficient evidence to answer. Furthermore, AnchorPrompt is perturbation-agnostic at inference, requiring no prior detection of perturbations and enabling zero-shot transfer to unseen distortions. We evaluate three LALMs across three benchmarks and show that AnchorPrompt improves answer consistency in most tested conditions. Clean accuracy improves in six of nine model-benchmark pairs, with minimal impact on the remainder of 1.2% at most. Crucially, AnchorPrompt reduces hallucinations under severe audio corruption while keeping false refusals on clean audio rare. Finally, these consistency gains transfer to unseen perturbations, such as choice permutations and reverberation.
[LG-129] SkillSpec: Consensus-Gated Agent Skill Evolution via Representation Specialization
链接: https://arxiv.org/abs/2610.00704
作者: Huancheng Chen,Xiaodi Sun,Zhaoqiong Huang,Shenyang Huang Shreya Singhal,Jingwen Lu
类目: Machine Learning (cs.LG)
*备注:
Abstract:Natural-language skills are textual procedural memories through which large language model (LLM) agents retain reusable task knowledge without updating model weights. Existing methods typically treat skills as either static artifacts or monolithic documents optimized using aggregate validation scores as feedback. However, representing a skill as a monolithic document restricts optimization to its textual content, without explicitly modeling the structure through which procedural knowledge is retrieved and executed. We identify a key distinction between learning what knowledge to retain and determining how to organize it: textual updates should first be validated through execution evidence, after which the retained knowledge should be structured according to its procedural dependencies and retrieval requirements. To this end, we introduce SkillSpec, a two-phase framework comprising consensus-gated evolution and representation specialization. In the consensus-gated phase, complementary editing intents generate complete candidate skills. An update is committed only when paired evaluations reach consensus, requiring sufficient overall improvement and non-negative aggregate paired gain in every repeated evaluation. In the specialization phase, signals of process and redundancy sensitivity derived from the full optimization trajectory, including accepted and rejected candidates, guide the selection of a flat, graph, or hybrid this http URL six benchmarks and three target language models, SkillSpec improves average success rate over SkillOpt by 6.89%, averaged across the three models. These results demonstrate that reliable skill evolution and representation specialization address complementary objectives: deciding what knowledge to retain and how to structure it for inference.
[LG-130] Progressive-Resolution Secure Aggregation for Federated Learning
链接: https://arxiv.org/abs/2610.00695
作者: Seyed Mohammad Azimi-Abarghouyi
类目: Cryptography and Security (cs.CR); Distributed, Parallel, and Cluster Computing (cs.DC); Information Theory (cs.IT); Machine Learning (cs.LG)
*备注:
Abstract:Secure aggregation lets a server recover an aggregate of client updates without observing any individual update, but conventional protocols fix the aggregate precision when clients upload. We introduce and formulate a new progressive-resolution secure-aggregation functionality in which clients upload once and successively finer resolutions of the same aggregate can later be authorized without renewed client participation. To realize this functionality, we propose progressive-resolution secure aggregation (PSA): each clipped, dithered update is represented by compatible nested-lattice digits; separately releasable layers are protected by secure aggregation and an additional aggregate pad that remains unavailable to the server until a non-colluding release controller authorizes that layer.
[LG-131] Leto: Fast In-Place Recovery for LLM Training on Surviving Hardware
链接: https://arxiv.org/abs/2610.00687
作者: Geon-Woo Kim,Joon Ha Kim,Daehyeok Kim
类目: Distributed, Parallel, and Cluster Computing (cs.DC); Machine Learning (cs.LG)
*备注:
Abstract:Hardware-operable failures (HOFs) interrupt large language model (LLM) training but permit recovery on the same hardware without reset, repair, or replacement. Existing recovery systems nevertheless reload checkpoints, recompute lost progress, and rebuild process state, idling GPUs that could otherwise continue training. We present Leto, a fault-tolerant training system that leverages surviving hardware to enable efficient in-place recovery. Our key insight is that the state needed to resume training can be retained or prepared outside the active training process while remaining on the same hardware. Leto retains the working model state and the reusable process state, and preinitializes the remaining state in a shadow trainer. We devise two-tier erasure protection and chunk-level transactional updates to keep the retained model state recoverable and consistent, and reclaim the shadow state when active training needs its GPU memory. Evaluation on 6- and 72-GPU NVIDIA A100 clusters shows that Leto recovers 3.6–6.5 \times faster than the best-performing checkpointing baselines and improves productive training time by up to 13.7 percentage points. Large-scale simulation shows over 95% productive training time on a 131,072-GPU cluster. Subjects: Distributed, Parallel, and Cluster Computing (cs.DC); Machine Learning (cs.LG) Cite as: arXiv:2610.00687 [cs.DC] (or arXiv:2610.00687v1 [cs.DC] for this version) https://doi.org/10.48550/arXiv.2610.00687 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[LG-132] Grand Canonical Generators NEURIPS206
链接: https://arxiv.org/abs/2610.00683
作者: Andreas Burger,Malte Franke,Luka Mucko,Kjell Jorner,Alan Aspuru-Guzik
类目: Machine Learning (cs.LG); Machine Learning (stat.ML)
*备注: SimBioChem NeurIPS 206
Abstract:We introduce Grand Canonical Generators (GCG), a generative framework that extends Boltzmann generators to the grand canonical ensemble. We present two designs. The first conditions a variable-size generative model on the chemical potential, sampling particle number and configuration jointly. The second factorizes the grand canonical distribution into a particle-number distribution and the corresponding canonical Boltzmann density. This factorized formulation can use any existing Boltzmann generator for the canonical component, encodes the known linear chemical-potential dependence analytically, and yields a tractable likelihood that supports self-normalized importance sampling (SNIS). Empirically, GCG accurately reproduces grand canonical observables on a Lennard–Jones fluid and methane adsorption in a zeolite, demonstrating generalization across chemical potentials and correction via SNIS and grand canonical Monte Carlo.
[LG-133] ORBIT-FMIB: Tracking Order-Resolved Epistatic Information Through ESM-2
链接: https://arxiv.org/abs/2610.00672
作者: Maryam Rahimimovassagh,Ivan Garibay,Niloofar Yousefi
类目: Machine Learning (cs.LG); Quantitative Methods (q-bio.QM)
*备注:
Abstract:Protein foundation models support mutation-effect and structural prediction, but predictive performance alone does not reveal which forms of biological interaction information remain accessible through model depth. We ask whether ESM-2 retains higher-order epistatic information as strongly as first- and second-order information across its representation hierarchy, introducing ORBIT-FMIB, a diagnostic framework combining Walsh-based interaction decomposition with subset-conditioned neural dependence estimation. The method is validated on synthetic landscapes with known interaction structure before being applied to the dense four-site GB1 fitness landscape using frozen ESM-2 representations. An initial production run suggested ESM-2 retains higher-order epistatic information less well than lower-order information ( \Delta_\mathrmHO-LO=-0.107 ). An independent replication of the complete measurement grid, under matched GPU hardware and identical critic seeds, substantially reduced this contrast ( \Delta_\mathrmHO-LO=-0.017 ), and its sign was unstable across otherwise-defensible evaluation-pairing choices applied to the same trained critics ( -0.011 to +0.015 ). We therefore do not currently have robust evidence that ESM-2 selectively loses higher-order epistatic information, nor that retention is equal across orders; the directional question remains open. The measurement protocol itself, including its documented removal of a positional-subset shortcut in pooled critics, remains validated and is unaffected by this finding. ORBIT-FMIB is offered as a diagnostic framework for probing interaction structure in protein foundation models; this study’s own replication result illustrates why such probing requires adequately-powered reproducibility checks before its output is treated as a biological finding. Subjects: Machine Learning (cs.LG); Quantitative Methods (q-bio.QM) Cite as: arXiv:2610.00672 [cs.LG] (or arXiv:2610.00672v1 [cs.LG] for this version) https://doi.org/10.48550/arXiv.2610.00672 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[LG-134] Analysis of Quantized and Efficiently Adapted Protein Language Models
链接: https://arxiv.org/abs/2610.00665
作者: Ilan Yaniv Zeisler,Sebastian Clancy,Pouriya Bayat,Saaim Raad,Ivan Kraskov,Matthew Xie,Vivian White,Spencer Perkins,Serena Singh,Sepehr Bayat,Keith Pardee
类目: Machine Learning (cs.LG); Quantitative Methods (q-bio.QM)
*备注:
Abstract:Background: Protein language models (PLMs) are increasingly used for sequence generation and property prediction, but their size makes fine-tuning and deployment expensive. The effects of quantization and parameter efficient fine-tuning on performance, representations and generation remain insufficiently characterized. Results: We evaluated 4-bit quantization and low-rank adapter fine-tuning (QLoRA) across ESM-2, ESMC, ProtBERT, ProtT5, Ankh, Ankh3 and Profluent-E1. Across protein prediction tasks, many model-task pairs retained more than 90% of full fine-tuning performance. Peak GPU memory savings approached 90% for the largest models, although performance and efficiency varied by model, dataset and training configuration. QLoRA often preserved early-layer representations while inducing task-specific adaptations in middle and late layers, resembling full fine-tuning with smaller representational changes. Training speed and power effects were more varied. For unconditional generation with ProLLaMA, ProtGPT2, ProGen2, ProteinGLM and ESM3, 4-bit quantization largely preserved predicted structural and sequence-level properties, but token-level analysis revealed model-dependent shifts in autoregressive output distributions. Conclusion: QLoRA and 4-bit quantization reduce PLM computational requirements, particularly GPU memory usage. Our results support QLoRA as a first-pass strategy for memory limited adaptation, reserving full fine-tuning for challenging tasks, unstable architectures or low validation recovery. For generative PLMs, sequence-level and structural metrics should be complemented with distributional analysis, since downstream predictions alone may miss quantization-induced shifts. These approaches can broaden access to large-scale protein modelling while requiring model- and task-specific validation.
[LG-135] On Evaluating Quantum Kernel Robustness for Low-Resource Cross-Corpus Audio Deepfake Detection
链接: https://arxiv.org/abs/2610.00649
作者: Lisan Al Amin,Lei Zhang,Vandana P. Janeja
类目: ound (cs.SD); Machine Learning (cs.LG)
*备注:
Abstract:Synthetic speech detection is critical for audio security, but performance can degrade when labeled data are scarce and evaluation conditions differ from training. This study examines quantum kernel methods and lightweight neural models for cross-corpus audio deepfake detection under limited training data. We compare a Quantum Support Vector Machine (QSVM), a classical support vector machine (SVM), and a multilayer perceptron (MLP), all trained on frozen wav2vec 2.0 embeddings using a strict budget of 200 training samples. To match the qubit budget of near-term quantum hardware, embeddings are reduced to four dimensions using principal component analysis, and all models use the same reduced features. Experiments on ASVspoof 2019, ASVspoof 5, the ADD 2023 Challenge, and the In-the-Wild dataset show that under severe domain shift from ASVspoof 2019 to ADD 2023, the MLP degrades to near-random performance, with an area under the curve of approximately 50% and an equal error rate of 50.0%. In contrast, the QSVM maintains meaningful discrimination, achieving an area under the curve of 76.0% and an equal error rate of 27.0%. This advantage is not consistent across transfer directions. When trained on ADD 2023, the QSVM falls below chance on two of three transfers, while the MLP performs better. These results suggest that quantum kernel methods can be competitive under severe cross-corpus shifts and strict low-resource constraints, but do not provide a consistent advantage under near-domain transfer. We interpret these findings as an empirical characterization of quantum kernel inductive bias under distribution shift, rather than evidence of quantum advantage, since the four-qubit kernel can be simulated exactly on classical hardware.
[LG-136] Group-Invariant Statistics Determine Embedding Geometry: Harmonic Analysis of Representations from Bach to the Night Sky
链接: https://arxiv.org/abs/2610.00647
作者: Liam Storan,Andreas Tolias,Nina Miolane
类目: Machine Learning (cs.LG)
*备注: 31 pages, 9 figures
Abstract:The representations that language models learn for concepts such as months, weekdays, and places display consistent geometric structure: circles and saddle-shaped “Pringle” manifolds. Recent work traced these structures to \textittranslation symmetry in word co-occurrence statistics, deriving the observed Fourier geometry when co-occurrence depends only on distance on an abelian lattice of concepts. We demonstrate that more general notions of symmetry lead to equally structured predictions. Considering symmetries defined by arbitrary finite groups, compact groups, and homogeneous spaces, we prove that whenever the co-occurrence statistics of a word family are invariant under a group G , the learned word embeddings consist of matrix elements of the irreducible representations (irreps) of G . Circles and Pringles arise when G is cyclic, in which case the irreps are Fourier modes. We verify the irrep structure in three experimental settings. (i) The cyclic group \mathbbZ_12 : for the months of the year we recover the known circular geometry. (ii) A dihedral group acting on the major and minor triads: we unify two classical observations – that transposition and chord inversion form a group ( T/I ) acting on chords (music theory), which \textitimplies that the well-known “circle of fifths” emerges in learned chord embeddings (machine learning). (iii) We explain and reproduce a recently discovered spherical representation of celestial objects in large language models (LLMs) as a spherical-harmonic embedding derived from our theory. Our results demonstrate that the geometry of learned representations is often a consequence of the statistical symmetry of underlying data.
[LG-137] Learning Linear Systems under Heavy-Tailed Noise: A Non-Asymptotic Analysis from A Single Trajectory
链接: https://arxiv.org/abs/2610.00637
作者: Xiaomian Yang,Sungho Shin
类目: Machine Learning (cs.LG); Systems and Control (eess.SY)
*备注:
Abstract:We establish non-asymptotic sample complexity bounds for the least-squares estimation of vector autoregressive models for exponentially stable systems with heavy-tailed noise based on a single observed trajectory. By assuming i.i.d. noise, bounded noise covariance, and persistent excitation, we show that the estimation error is \widetilde\mathcalO(r^1/2T^-1/2+1/p) under bounded p th moment for p 2 , where T is the number of samples, r is the noise dimension, and \widetilde\mathcalO(\cdot) hides logarithmic terms. We also introduce a unifying approach to sample complexity analysis applicable to broad classes of noise distributions and showcase this by deriving error bounds for sub-exponential and sub-Gaussian noise distributions. Finally, we specialize our analysis to autoregressive models with exogenous inputs and show that the dimension factor of the error bound is independent of the model order.
[LG-138] Learning the identity: a case study of how SGD selects among functional decompositions
链接: https://arxiv.org/abs/2610.00615
作者: Andy Arditi,Weian Xie,David Bau,Liu Ziyin
类目: Machine Learning (cs.LG)
*备注:
Abstract:One might think that learning the identity function with a deep linear residual network is trivial - the path along residual connections already implements the identity, and so the network need only drive its weights to zero. However, this zero-weight solution is just one point on an entire manifold of population-loss minimizers, each corresponding to a different decomposition of the identity across the network’s layers. Although the population loss does not distinguish among these solutions, stochastic gradient descent (SGD) reproducibly favors particular ones. For instance, under anisotropic label noise, the learned layers exhibit a noise-dependent spectrum; even with weight decay, SGD does not generally recover the zero-weight solution. Changing only the parametrization, while leaving the set of realizable functions unchanged, yields different behavior: factoring each weight matrix as a product of two matrices causes the weights to collapse to zero, even without explicit weight decay. While perhaps mysterious and unintuitive at first, these phenomena can be understood through the lens of entropic loss, which augments the population loss with a term proportional to the expected squared norm of the minibatch gradient (Ziyin et al., 2025). On the identity manifold, the population loss is constant, while the entropic term distinguishes among these decompositions. We characterize its minimizers analytically and use them to derive predictions for the structure of solutions favored by SGD. Networks trained with SGD closely match these predictions. Overall, the identity learning task studied here serves as a clean and simple case study of how the lens of entropic loss can clarify why SGD favors particular decompositions of the same input-output function. Subjects: Machine Learning (cs.LG) Cite as: arXiv:2610.00615 [cs.LG] (or arXiv:2610.00615v1 [cs.LG] for this version) https://doi.org/10.48550/arXiv.2610.00615 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[LG-139] From Task Mixtures to Specialized Experts
链接: https://arxiv.org/abs/2610.00580
作者: Hojat Allah Salehi,Mehrdad Mahdavi,Andrew Arash Mahyari,M. Hadi Amini
类目: Machine Learning (cs.LG)
*备注: 63 pages, 9 figures
Abstract:In collaborative foundation model fine-tuning, client data is rarely homogeneous. Instead, clients typically possess unknown mixtures of distinct data distributions, or tasks. Conventional federated learning primarily addresses heterogeneity across clients without explicitly resolving latent task mixtures within each client. We study this setting as compound heterogeneity, where data is heterogeneous both across and within clients. We study adaptation over a common frozen representation and show that, when tasks share the same feature geometry, the optimal model for a client’s task mixture under squared loss is a convex combination of the optimal models for its underlying tasks. Thus, a single locally trained model represents the client’s overall task mixture, while individual inputs may be drawn from different underlying task distributions. This motivates routing inputs to specialized experts, and we show that, when the task optima form a simplex, task-aligned routing achieves lower risk than any single adapted model for genuinely mixed clients. With access to a small set of task-labeled public samples, we derive a convex program to recover task experts and match them to their corresponding tasks. Our routing analysis shows that effective specialization requires input-dependent expert selection aligned with each client’s task mixture. Motivated by this analysis, we propose FedSEE. Across our experiments, FedSEE avoids the negative transfer observed in the evaluated baselines and improves performance by 2.9 points overall and 3.7 points for the worst-served quartile.
[LG-140] Attention Kernels for Learning Maps Between Heavy-Tailed Measures NEURIPS2026 STOC
链接: https://arxiv.org/abs/2610.00564
作者: Kailen Hargenrader,Edoardo Calvello,Bohan Chen
类目: Machine Learning (cs.LG)
*备注: 32 pages, 17 figures, accepted to NeurIPS 2026 Workshop on AI for Stochastic Dynamics
Abstract:Operator learning on probability measures can be accomplished with transformers. For measures with polynomial tails, the exponential weighting in softmax can make the corresponding measure-level attention integrals diverge. This motivates replacing the exponential with slower-growing functions. We construct two benchmarks for operator learning on measures with closed-form targets. We use these benchmarks to study attention kernel growth and data transformation in post-norm transformers. Without data transformation, the softmax models exhibit ensemble collapse on both heavy-tailed benchmarks, while the three slower-growing kernels avoid collapse. Symlog preprocessing allows softmax to avoid collapse on the matrix inverse task but not on the sheared swap task. On the Gaussian control, all four kernels perform similarly. We also examine how sample size affects the sensitivity of empirical energy and Wasserstein distances to tail differences. These results support slower-growing attention kernels as an effective design choice for post-norm transformers learning from heavy-tailed ensembles.
[LG-141] Redundancy Meets Synergy: Dependency-aware Expert Selection for MoE via Submodular Optimization
链接: https://arxiv.org/abs/2610.00558
作者: Zheng Lin,Shaoke Fang,Yuxin Zhang,Jinfeng Xu,Zihan Fang,Zhe Chen,Wei Ni,Jun Luo,Symeon Chatzinotas
类目: Machine Learning (cs.LG); Distributed, Parallel, and Cluster Computing (cs.DC)
*备注: 26 pages, 3 figures
Abstract:While Mixture-of-Experts (MoE) models effectively scale model capacity through sparse activation, their deployment is often bottlenecked by prohibitive memory requirements. Extracting a compact subset of experts presents a promising solution. However, existing expert selection heuristics predominantly rely on Top-k ranking, which isolates the evaluation of individual experts and ignores the intricate inter-expert dependencies introduced by the MoE gating network. In this paper, we propose DS-MoE, a theoretically grounded framework that redefines expert selection via difference-of-submodular (DS) optimization. By analyzing the second-order Taylor expansion of the loss degradation, we reveal functional duality within expert combinations: redundancy (where experts encode overlapping representations) and synergy (where experts provide complementary error cancellation). To navigate this duality, we mathematically decouple redundancy reduction from synergy maximization by formulating the selection objective as a DS function. Furthermore, we devise a tailored majorization-minimization (MM) algorithm with provable monotonicity guarantees to efficiently identify the optimal expert subset. Extensive experiments demonstrate that DS-MoE effectively preserves indispensable expert combinations, achieving superior performance compared to the state-of-the-art baselines.
[LG-142] Evaluating Hybrid Quantum-Classical Models for Reduced-Order Brain Deformation Dynamics
链接: https://arxiv.org/abs/2610.00554
作者: Tao Liu,Ge He,Dongyu Liang,Wujie Wen
类目: Machine Learning (cs.LG); Computational Engineering, Finance, and Science (cs.CE)
*备注: QCE26
Abstract:We evaluate hybrid quantum-classical machine learning for the reduced-order prediction of spatiotemporal brain deformation fields. To mitigate the computational intractability of high-dimensional displacement fields, we employ Proper Orthogonal Decomposition (POD) to project the data into a compact latent space. Within this framework, we formulate two distinct learning objectives: static temporal-to-latent regression and autoregressive latent state forecasting. We systematically benchmark compact classical baselines against both minimal and enhanced hybrid quantum architectures. Our results demonstrate that classical networks provide the strongest baselines in the present setting. For static regression, a classical POD-MLP outperforms all evaluated quantum variants, although an enhanced Variational Quantum Circuit (VQC) substantially improves upon a minimal VQC baseline. For temporal forecasting, a classical POD-LSTM delivers superior predictive accuracy and statistical robustness compared to an enhanced Quantum LSTM (QLSTM) across varying history windows and random initializations. Overall, this study establishes reduced-order physical field learning as a rigorous testbed for near-term QML, highlighting that while hybrid enhancements successfully recover expressivity in weak quantum circuits, classical architectures retain a definitive advantage in both fidelity and stability.
[LG-143] Geometry-Dependent Bounds for Online Non-Monotone DR-Submodular Maximization
链接: https://arxiv.org/abs/2610.00545
作者: Vaneet Aggarwal
类目: Machine Learning (cs.LG)
*备注:
Abstract:We study adversarial online maximization of nonnegative, non-monotone DR-submodular functions over compact convex down-closed sets. A learner commits each action before observing its objective and competes with the best fixed action in hindsight. We prove a comparator-uniform first-order inequality that gives coefficient 4/9 , improving the online 0.401 benchmark, with one gradient query and one projection per round and O(\sqrt T) expected approximate regret. If \zeta \bf 1 \in K\subseteq[0,1]^d , the coefficient improves to \underline\alpha(\zeta)=\tfrac12-(1-2\zeta)+^2/[2(3-2\zeta)^2] . The proof is a direct ordered-coordinate argument with an objective-independent rational action. Conversely, a three-group symmetry-gap construction yields an offline oracle upper bound \beta*=0.470438681380894\ldots at \zeta=0 , even with exact value and full-gradient responses. A parameterized extension and exact finite-instance bounds define an upper function for every \zeta . The lower and upper bounds match at 1/2 for \zeta\ge1/2 , and show that the optimal deficit from 1/2 is \Theta((1/2-\zeta)^2) as \zeta\uparrow1/2 . For coefficient-revealed polynomials we obtain 1/2 for quadratics and a geometry-dependent cubic coefficient starting at 8/17 , including 0.49 at \zeta=1/5 . A constant objective sequence yields an offline (4/9-\varepsilon) approximation with polynomially many first-order queries on the cube and projections, without requiring a supplied positive lower bound on the optimum. We also give nonanticipating adaptive-adversary and value-feedback guarantees, including O(T^3/4) regret with one noisy value per round.
[LG-144] Same Scene Different Task: Skill Alignment for Compositional Generalization in VLAs
链接: https://arxiv.org/abs/2610.00524
作者: Taegeun Yang,Youngju Na,Yoonki Cho,Sung-Eui Yoon
类目: Robotics (cs.RO); Machine Learning (cs.LG)
*备注: 26 pages, 5 figures. Project page: this https URL
Abstract:Vision-language-action (VLA) models often struggle to generalize to skill combinations absent from their fine-tuning demonstrations, even when every constituent skill has been demonstrated. We focus on a vision shortcut as one failure mode: during fine-tuning, visual observations can serve as a proxy for the instruction, so a policy may execute a demonstrated combination associated with similar observations rather than the instructed combination. This motivates training with counterfactual pairs formed by holding a demonstration observation fixed while changing the instruction to specify an undemonstrated combination. These pairs, however, lack corresponding demonstrated action targets. Crucially, the currently required skill has already been demonstrated, but actions from those executions cannot serve as direct targets because the same skill can require different actions across observations. We propose CRAFT, which transfers supervision from demonstrated executions of the required skill to counterfactual pairs using skill representations that can be reused across executions of the same skill. Across three VLA models and two simulation benchmarks, CRAFT improves success on undemonstrated combinations while maintaining high success on demonstrated ones; it also improves compositional generalization on a real robot. Project website: this https URL
[LG-145] SimplexUQ: An Evaluation Framework and Benchmark for Conformal Uncertainty on Simplex-Valued Predictions
链接: https://arxiv.org/abs/2610.00523
作者: Liang You,Hengyu Shi,Dongwen Ou
类目: Machine Learning (cs.LG); Applications (stat.AP)
*备注: Preprint. Code and data are available at this https URL and this https URL
Abstract:Conformal prediction guarantees marginal coverage, but a single calibration threshold can still spread that coverage unevenly, over-covering easy regions and under-covering hard ones. SimplexUQ is, to our knowledge, the first benchmark and reproducible protocol for measuring this allocation problem on simplex-valued predictions; it compares existing conformal wrappers rather than proposing a new one. Its task suite, SimplexTasks-12, combines six controlled synthetic regimes with six frozen-predictor real tasks spanning class probabilities, topic mixtures, spectral abundances, cell-type fractions, age distributions, and emotion mixtures. Each comparison fixes the predictor, score, and response-free stratification map, varies only the wrapper, and reports marginal coverage, worst-stratum coverage, max disparity, and within-task radius and compute. Global calibration can look valid while failing badly: on CIFAR-10 it attains 0.900 marginal coverage but only 0.542 in the worst entropy stratum, and Mondrian calibration raises that stratum to 0.886 while reducing max disparity from 0.358 to 0.022. No wrapper dominates, however. Under smooth synthetic heterogeneity, several repairs are competitive; fixed-map analyses show that rankings depend on the evaluation groups and protocol; and in a 12-task comparison, Mondrian has lower disparity on its single target partition for all 12 tasks, whereas BatchMVP has lower disparity over overlapping groups on five. These are empirical comparisons, not new coverage guarantees. A controlled predictor-bias sweep shows that removing predictor bias only partly reduces global-threshold disparity. We release task cards, result provenance, permitted derived arrays, and rebuild instructions, and treat wrapper selection as a diagnostic comparison rather than a universal ranking.
[LG-146] One-Step Generative Modeling via Training Dynamics Action
链接: https://arxiv.org/abs/2610.00518
作者: Zhangyong Liang,Ying Huang,Haibin Ling
类目: Machine Learning (cs.LG)
*备注:
Abstract:One-step generative models construct a static generator through iterative training-time transport. Existing transport objectives primarily assess distributional motion, although a neural generator needs to realize the requested sample displacements jointly through shared parameter updates. The training-time construction raises the question: \emphonce training becomes the iterative process that constructs the final one-step map, what to optimize: the next distributional move, or the route by which the finite generator learns the final map? To address the question, we introduce \textbfTraining \textbfDynamics \textbfAction (\textbfTDAction), which selects transport targets according to local shared-parameter realization cost while retaining a prescribed level of distributional progress. We formulate the cost as a soft-terminal control problem and derive a closed-form Batch Tangent Action-to-Go value that accounts for parameter effort and terminal mismatch. The criterion captures cross-sample interactions omitted by independent pairwise costs; under isotropic mobility, the criterion agrees with quadratic Euclidean assignment for deterministic balanced couplings. Randomized tangent probes provide a low-rank implementation that constructs shared detached targets without adding an inference-time trajectory. Controlled studies examine the relationship between generator geometry, transport selection, and realized local action. On ImageNet 256\times256 , TDAction attains an FID below 1.1 without distillation.
[LG-147] Denoising Surface: Modeling and Predicting Inference Cost for Diffusion LLM Serving
链接: https://arxiv.org/abs/2610.00499
作者: Haoyu Zheng,Fangcheng Fu,Binhang Yuan,Yongqiang Zhang,Liang Deng,Hao Wang,Yuanyuan Zhu,Xiao Yan,Jiawei Jiang
类目: Machine Learning (cs.LG)
*备注:
Abstract:As diffusion large language models (dLLMs) become more capable, they are moving from research settings to real-world \textitserving, where request management (such as scheduling and resource allocation) relies on accurate estimation of per-request inference cost. However, common cost proxies fall short for dLLMs: output length ignores that one forward pass can unmask multiple tokens, and denoising-step count ignores the \textitheterogeneous per-step costs. We observe that the block-autoregressive generation mechanism induces a two-dimensional execution structure over output blocks and within-block denoising steps, whereas these proxies collapse it into a scalar, discarding information essential for characterizing the cost. Motivated by this insight, we propose the Denoising Workload Surface (DWS), which preserves this two-dimensional block-step structure as a probability surface to weight the heterogeneous per-step costs. We then design a coarse-to-fine training scheme that enables a lightweight prompt-only predictor to accurately predict the complex DWS. This predictor runs efficiently even on a single CPU core, avoiding GPU contention with the serving model. Since DWS decouples request-dependent execution behavior from deployment-specific cost factors, the predictor transfers across hardware configurations without retraining. In \textitreal-world serving experiments, DWS reduces cost-prediction error by up to 2.50\times over scalar-based predictors, while the DWS-guided shortest-job-first scheduler reduces end-to-end latency by up to 1.92\times for online chatbots.
[LG-148] Score the Update Not the Token: Descent-Aligned Routing for Combinatorial LoRA Experts
链接: https://arxiv.org/abs/2610.00493
作者: Priya Nair,Lukas Brenner,Maya Lindqvist,Daniel Whitmore,Wen-Hsuan Liu,Tom Saliencro,Amara Okonkwo,Rohan Desai
类目: Machine Learning (cs.LG)
*备注:
Abstract:Mixture-of-LoRA-experts methods raise the capacity of low-rank adaptation by routing each token to a few low-rank experts. Nearly all of them tie one input-side factor to one output-side factor per expert, and nearly all of them route by scoring the token: the router picks experts without seeing what any of them would write. We argue that the router should score the update. To first order, adding an expert’s update to a layer output lowers the loss by the inner product between that update and the negative loss gradient at the output. This usefulness is quadratic in the token, so a router that is linear in the token sees only the part of it that runs through the token mean, and routers that rank experts by the norm of their own activations never see the output factor. If each expert is split into a reader (down-projection) and a writer (up-projection), the usefulness of every reader–writer pair becomes an inner product in the shared rank- r space, and all N_AN_B pairs can be scored from N_A+N_B vectors. We build VANE on this identity. A low-rank compass predicts the descent direction of each token. VANE scores every pair by the alignment between its update and the compass without forming any update, activates the top- k pairs with additive gates, and gives every pair its exact first-order router gradient. On single-domain commonsense reasoning and a four-domain multi-task mixture with Llama-3.2-3B and Llama-3.1-8B, VANE attains the best average among twelve PEFT and MoE-LoRA baselines, by 0.9–1.1 and 1.3–1.5 points respectively, with less than half the trainable parameters of an 8-expert MoE-LoRA. Its router scores also track the measured usefulness of experts far more closely than token routers do.
[LG-149] ScaffoldM3C: A Multimodal Sequential Monte Carlo Framework for Generative Stable Construction Planning
链接: https://arxiv.org/abs/2610.00487
作者: Gadiel Sznaier Camps,Chengyang He,Guillaume Sartoretti,Eduardo Montijano,Mac Schwager
类目: Robotics (cs.RO); Machine Learning (cs.LG)
*备注: This work has been submitted to IEEE Transactions on Robotics and Learning (T-RL) and is currently under review. Project Page: this https URL
Abstract:Autonomously constructing physically realizable 3D structures remains a significant challenge due to combinatorial action spaces, interchangeable components, equifinal assembly sequences, and strict stability requirements during construction. State-of-the-art methods fine-tune large language models for text-based generative construction. However, these approaches do not allow for Multimodal (text, image, sketch) conditioning, overlook the practical role of scaffolding for stabilizing intermediate structures, and suffer from slow inference speeds. Therefore, we formulate construction as a probabilistic next-block generation task with multiple potential assembly actions and multiple potential task conditioning modalities. Concurrently, we explicitly consider the utility of scaffolding by introducing an auxiliary scaffold block token. We present Scaffold Multimodal Monte Carlo (ScaffoldM3C), a multimodal, lightweight, auto-regressive model for stable block-based construction, that proposes a set of next-step candidate blocks. Leveraging these candidates, we utilize Sequential Monte Carlo (SMC) to maintain a population of possible assembly sequences, allowing us to consider multiple, potentially different, assembly directions simultaneously. We train our multimodal architecture by extending the StableText2Brick dataset to contain image conditioning prompts and scaffold-stabilized build sequences. ScaffoldM3C is 4x smaller than competing baselines, yielding a 5x to 20x speedup during inference, while achieving comparable construction quality to state-of-the-art methods and higher overall stability. We demonstrate the effectiveness of our approach through simulations and real-world robot assembly demonstrations.
[LG-150] AIR-LLM : Broadcasting AI Weights over Radio for Memory-Free Edge LLM Inference via RF Computing
链接: https://arxiv.org/abs/2610.00465
作者: Zhihui Gao,Tingjun Chen,Dirk Englund
类目: Information Theory (cs.IT); Emerging Technologies (cs.ET); Machine Learning (cs.LG); Signal Processing (eess.SP); Applied Physics (physics.app-ph)
*备注: 14 pages, 12 figures, 6 tables. Appendix: 12 pages, 7 figures, 11 tables
Abstract:Next-generation large language models (LLMs) are expanding from the cloud to ubiquitous edge devices. However, edge devices typically either lack the memory to store increasingly large LLM weights or, even with enough memory, spend unaffordable energy on loading the weights. This raises our question: can an edge device run an LLM without storing or loading its weights, but receive them over the air and consume them on the fly? Inspired by wireless broadcasting, we present AIR-LLM, an LLM inference architecture for edge devices, which is composed of: (i) a central radio (e.g., 5G base stations) that broadcasts the LLM weights into the air, and (ii) the edge user that receives the weights and completes the general matrix-vector multiplication (GEMV) of LLM inference directly in the radio frequency (RF) domain using RF mixers. To further shorten the airtime, AIR-LLM exploits MIMO spatial multiplexing and proposes an energy-efficient precoder-postcoder pair on the edge to calibrate its own wireless channel. Since the central radio stays user-unaware, AIR-LLM is user-scalable so that one broadcast serves unlimited users within its coverage. We implement AIR-LLM on the NVIDIA Sionna ray-traced channels of two real-world urban scenes and the profiling of a real RF mixer. With a WikiText-2 perplexity degradation of 4.0% on LLaMA-3.1-8B, AIR-LLM saves the energy by 157.7x/40.4x against the FP16 and weight-only quantization baselines; with 20 users, its airtime is 104.1x/26.0x shorter, respectively.
[LG-151] Exact information accounting for SGD methods
链接: https://arxiv.org/abs/2610.00446
作者: Akshay Balsubramani
类目: Machine Learning (cs.LG); Information Theory (cs.IT); Optimization and Control (math.OC); Machine Learning (stat.ML)
*备注:
Abstract:As an alternative to the standard geometric analyses, we give an exact, information-theoretic analysis of stochastic gradient descent (SGD) and its variants. We show that a preconditioned SGD step is the posterior-mean update of a Gaussian Bayes model, and that its one-step regret splits into an intrinsic-time cost and a change in comparator information. The split extends to an identity for the objective itself. Convex convergence, strict-saddle-point escape, the link between flatness and generalization, the standard learning-rate schedules, adaptive optimizers, and the noisy, momentum, heavy-tailed, and gradient-free variants of SGD each correspond to a term or a special case of this identity. We measure its terms on synthetic and real training runs. On real networks it attributes the slack of classical convergence bounds to the terms their derivations drop and separates optimizers that reach the same training loss. That separation follows the number and consistency of their steps. Its relation to which of them generalizes better differs between networks. For gradient-free SGD the identity determines how a curvature preconditioner should enter the update. The sharpness-based generalization certificate it yields, with a data-independent isotropic prior, is vacuous at network scale unless the curvature spectrum is nearly flat across all parameters.
[LG-152] One pool many targets: a conservation layer and what archival data can identify
链接: https://arxiv.org/abs/2610.00445
作者: Zahra Khodagholi,Niloofar Yousefi
类目: Machine Learning (cs.LG)
*备注:
Abstract:Pairwise guide–transcript scores do not enforce conservation of a finite guide-loaded RISC pool when they are interpreted independently as occupancies. We formulate a differentiable scalar equilibrium layer: one conservation equation with a unique positive root and exact implicit gradients. It yields a redistribution theorem, a qualified high-resource limit, an analysis of the retrieval approximation, and a conditional rank-invariance result: within one construct at one dose, rankings by fractional occupancy cannot distinguish equilibrium from independent scoring. We therefore audit the two experiments that proposition leaves open, dose and cross-context, on archival off-target data. Corrected thermodynamic affinities associate weakly with measured repression in the direction a working predictor requires, but a paired permutation test and a construct-cluster bootstrap do not establish added predictive value from the coupling: what survives their differing permutation-null baselines is \GapNet, a descriptive \GapNetOverSE of the equilibrium association’s cluster standard error. The dose fits are heterogeneous and frequently violate the model-implied exponent constraint, which is superlinear rather than sublinear, so these data do not identify the competition parameter. A saturable compression of the competitor set holds both accuracy targets on held-out guide families but is not faster at the size measured. The contribution is a reusable conservation operator and the experimental information needed to test it. The code for this study is available at this https URL.
[LG-153] Every Batch Is Its Own Validation Set: Leave-One-Out Gradient Matching for Online Data Selection in LLM Fine-Tuning
链接: https://arxiv.org/abs/2610.00436
作者: Hongyu Chen,Xinyi Luo,Ming Zhao,Lin Tang,Zihan Xu,Jing Li,Yuxuan Wang,Haoran Deng,Wei Zhang
类目: Machine Learning (cs.LG)
*备注:
Abstract:Online batch selection fine-tunes a language model on the most useful part of each candidate batch. Selectors that match the gradient of the candidate batch are attractive because they need no held-out data, yet they rarely beat training on the whole batch. We show why. In-sample gradient matching uses every example as part of its own target, so its objective credits each example with its own gradient noise. This is the covariance penalty that makes training error optimistic, now sitting on the diagonal of the gradient Gram matrix: it steers selection toward the noisiest examples and makes the full batch the best solution the objective can reach. The fix costs nothing. For each example, the other candidates form an independent sample of the data distribution, so removing the diagonal turns the matching objective into an unbiased estimate of the update’s error with respect to the population gradient. The minimizer of this leave-one-out objective weights examples by their gradient signal-to-noise ratio (SNR), and whenever per-example SNR is heterogeneous enough, half of a batch yields a lower-error update than the whole batch; we give the exact condition. We build \method on this principle. It computes the Gram matrix in the metric of the Adam preconditioner during the ordinary backward pass, selects a weighted subset greedily with a (1-e^-\gamma) guarantee, and uses no held-out data. Across four fine-tuning tasks and seven backbones from 1.5B to 8B parameters, LOOM improves on full-batch training by 2.3 and 2.4 points on Llama-3.1-8B and Qwen2.5-7B, exceeds every in-sample gradient matcher by 2.4 points and the validation-guided GREATS and OPUS by 1.6–2.0, and selects injected label noise at under a fifth of its base rate.
[LG-154] IrekoGPT : Turning Structured Pruning into Post-Hoc Slimmable LLM s NEURIPS2026
链接: https://arxiv.org/abs/2610.00426
作者: Pietro Moriello,Pietro Buzzega,Angelo Porrello,Simone Calderara
类目: Machine Learning (cs.LG); Machine Learning (stat.ML)
*备注: Accepted at the NeurIPS 2026 Workshop “AXIOM: Foundations of Efficient Deep Learning”. 8 pages, 5 figures
Abstract:We introduce IrekoGPT, a post-hoc method for converting pretrained LLMs into slimmable models whose width can be adjusted at inference time. Building on SliceGPT, we retain its projection matrices without pruning them, allowing a single model to expose nested subnetworks at different widths. We improve robustness by calibrating each layer across multiple compression ratios, and correct downstream linear layers through gradient-free ridge regression. Across Llama and Qwen models, preliminary results show improvements over naive PCA-based slimming, with the largest gains at high compression. Code is available at this https URL
[LG-155] CommunityKV: Efficient Long-Context Decoding via Graph Partitioning
链接: https://arxiv.org/abs/2610.00418
作者: Joe McKenna,Anastasios Alexandridis,Nathan Susanj,Jing Liu
类目: Machine Learning (cs.LG)
*备注:
Abstract:Scaling Transformers to long contexts is constrained by the quadratic cost of self-attention and the linear growth of key-value cache memory transfer. Sparse attention mitigates this by retrieving only relevant tokens, but current approaches either require large-scale training or, within the training-free regime, rely on semantically coarse heuristics or expensive clustering that is difficult to update efficiently during decoding. We introduce CommunityKV, a framework that formulates sparse attention as a community detection problem. CommunityKV constructs a token graph from the QK^T scores already computed during standard prefill, and partitions the graph into communities to enable retrieval of semantically coherent token groups. A local update rule assigns newly generated tokens to communities in constant time, enabling sparse retrieval throughout streaming decoding without global re-partitioning. We evaluate CommunityKV on Qwen3 and Llama-3.1 models across three long-context benchmarks. With one graph per query head, CommunityKV delivers up to 1.25\times the end-to-end generation throughput of dense attention, while query-group graph aggregation yields up to 1.71\times with comparable accuracy.
[LG-156] Source Identification Is Not Fitness Testing: Measuring the Limits of Synthetic-Data Attribution
链接: https://arxiv.org/abs/2610.00417
作者: Joss Armstrong
类目: Machine Learning (cs.LG)
*备注:
Abstract:Repeated training on model-generated data can degrade later models. One possible response is to use provenance when deciding which generated examples to reuse. We test both how reliably that provenance can be recovered and whether it helps identify better training data. Using financial-risk text, we first identify the source of generated passages and then repeat the test after rewriting them. Generator attribution is 98.7% accurate on the original passages but falls to 53.1% after paraphrasing and 29.0% after style rewriting. Generated-versus-human detection remains close to perfect against the tested human comparison set. We then compare two ways of selecting generated examples over three rounds of generation and retraining. One uses source information. The other uses a score from a separate reference model. The two rules select different examples, but the planned comparison does not detect a stable difference in the degradation of the resulting models. The results show that identifying where data came from and identifying which data are useful for training are separate problems. The experiment therefore separates source identity, criterion-facing selection, and recursive training outcome: neither the provenance score nor the tested criterion-facing proxy is established as sufficient for future recursive behaviour.
[LG-157] Do Better Scores Mean Better Physics? Physics-Grounded Explanations for Sim2Real Neural Operators NEURIPS2026
链接: https://arxiv.org/abs/2610.00415
作者: Somyajit Chakraborty,Xizhong Chen
类目: Machine Learning (cs.LG); Fluid Dynamics (physics.flu-dyn)
*备注: 10 pages, 6 figures. Accepted at the NeurIPS 2026 XAI4Science Workshop, Tiny Paper Track
Abstract:Machine-learning surrogates accelerate physical simulation, but lower prediction error need not coincide with lower error in physically relevant flow statistics. We examine this question for flow around a NACA4418 airfoil using paired computational-fluid-dynamics simulations and experimental particle-image-velocimetry measurements. A mean-preserving input intervention removes velocity fluctuations from selected regions of observed flow histories. Across four neural operators, removing fluctuations from the most energetic 10% of valid observed cells changes forecasts more than equal-area random removal. Because the masks are not matched for removed fluctuation energy, this contrast measures sensitivity, not independent evidence of physical importance. Separately, a CNO has lower velocity-field error but substantially higher two-component fluctuation-energy error than the reference on both analysis subsets. An output attenuation stress test also demonstrates disagreement between benchmark errors and domain-summed fluctuation energy. These single-benchmark results motivate reporting complementary physical diagnostics alongside aggregate prediction scores; they do not establish counterfactual physical correctness.
[LG-158] EchoPress: Query-Agnostic KV Cache Pruning via Virtual Context Reconstruction
链接: https://arxiv.org/abs/2610.00412
作者: Jiawei Lin,Saibo Geng,Thomas Bourgeat
类目: Machine Learning (cs.LG)
*备注:
Abstract:KV cache pruning reduces long-context inference memory usage by evicting less important key-value pairs. KVzip estimates importance through context reconstruction: prompting a model to repeat the context chunk by chunk. This achieves strong compression quality at the cost of additional forward passes. Learned approximations reduce this cost but require model-specific training. We analyze how KVzip identifies important cached information and show how to approximate its reconstruction scores using information already computed during prefill. These findings motivate EchoPress, a training-free method that approximates reconstruction attention using queries and keys from standard prefill. For each request, it reconstructs only the first chunk to calibrate importance scores for the remaining context. Experiments on LongBench and RULER with Qwen3-8B and Llama-3.1-8B-Instruct show that EchoPress matches KVzip in task accuracy across eviction ratios from 50% to 90%, while reducing compression overhead by a factor of 1.7-19.6 and total prefill time by a factor of up to 2.9. Code is available at this https URL.
[LG-159] VANDAM: Viewing a nucleotide sequence with DNA molecular priors
链接: https://arxiv.org/abs/2610.00411
作者: Jeremy Levy,Ariel Larey,Yury Nahshan,Raizy Kellerman,Elay Dahan,Amit Bleiweiss,Guy Leib,Omri Nayshool,Dan Ofer,Tal Zinger,Dan Dominissini,Gideon Rechavi,Marissa Wirth,Simon Lee,Dung Hoang,Noam D. Beckmann,Shane O’Connell,Nicole Bussola,Alexander W. Charney,Yoli Shavit,Nati Daniel
类目: Machine Learning (cs.LG); Machine Learning (stat.ML)
*备注: 25 pages, 3 figures, including appendices
Abstract:Contemporary Genomic Foundation Models (GFMs) rely on a DNA-as-a-string paradigm that employs masked token prediction objectives for pretraining. However, this abstraction does not explicitly model the biochemical, structural, and physical properties essential to biological function. Many molecular properties can be estimated from sequence using established biophysical models, so their utility lies not in providing an independent modality, but in introducing priors that training objectives can explicitly exploit. We introduce VANDAM, a framework that extends the training of GFMs with DNA molecular priors. In self-supervised training, VANDAM predicts regional molecular properties from pooled representations. When functional labels are available and can reward retaining molecular priors, local features are additionally injected at the input. VANDAM consistently improves downstream performance across four architecture families and nine held-out genomic tasks by complementing token-based objectives. Probing experiments further demonstrate that the use of molecular priors generalizes to other unseen molecular properties.
[LG-160] RACE: Residual-Aware Test-Time Adaptation for Neighbor-Rich Time-Series Foundation Model Forecasting
链接: https://arxiv.org/abs/2610.00405
作者: Hao-Nan Shi,Tong Wu,Chen-Cong Sun,Yuan Jiang,Han-Jia Ye,De-Chuan Zhan
类目: Machine Learning (cs.LG); Machine Learning (stat.ML)
*备注: 25 pages including references and appendices; 8 figures
Abstract:Time-series foundation models (TSFMs) perform strongly across forecasting tasks, but their per-series inference is ill-suited to neighbor-rich forecasting, where each query has access to related but nonidentical historical series. Continuous glucose monitoring (CGM) and Web/cloud workloads exemplify this setting: CGM trajectories share physiological patterns but vary across individuals, devices, and conditions, while Web/cloud workloads combine common operating regimes with non-stationarity, heavy tails, and bursts. These histories share useful structure, yet neighbors are not equally relevant. Existing methods either fine-tune TSFMs for each target domain, incurring additional costs and offering limited transferability across backbones, or append retrieved series without verifying whether they support the current forecast. The key challenges are conflicting residual evidence from neighboring series and residual patterns that vary across TSFMs and forecasting tasks. We formulate test-time neighborhood scaling: using same-domain neighbor evidence without modifying the backbone. We propose RACE (Residual-Aware Correction of Forecasting Errors), a two-stage framework for using historical neighbors. We first retrieve query-compatible neighbors, align their residuals to the query scale, and aggregate coherent evidence into the training-free RACE-TF correction. Full RACE then uses a lightweight, domain-specific Gate to determine when applying the correction is beneficial, with a reusable training workflow across TSFM backbones. Across four TSFMs, RACE improves all three domain-aggregate metrics on both primary domains, with the largest gains on high-error queries. Within each domain, a Gate trained on one TSFM transfers to other backbones without adaptation, and the resulting pipeline improves all 72 cross-backbone metric comparisons over the matched frozen targets.
[LG-161] he Conflict Between Logic and Memory: Learning Higher-Order Interactions in Shallow MLPs
链接: https://arxiv.org/abs/2610.00403
作者: Gongyue Zhang,Honghai Liu
类目: Machine Learning (cs.LG)
*备注:
Abstract:A network can fit its training examples while failing to recover the rule that generated their labels. We examine this separation in single-hidden-layer multilayer perceptrons (MLPs), using synthetic tasks that control interaction order and the presence of nuisance inputs. We establish elementary benchmark properties: pure parity contains no predictive lower-order marginals, admits an exact Bayes posterior, and can be represented on clean latent inputs by a width- k ReLU network. Experiments then identify distinct optimization outcomes. In a matched order-2–4 sweep, SGD, Adam, and Muon all reach 100% peak test accuracy at order two; at order three they reach 96.25%, 50.87%, and 76.82%, respectively, while Muon reaches 99.21% at order four. In a separate mixed-order task, freezing only the first-layer weights connected to independent nuisance inputs raises AdamW’s epoch-10 accuracy from 44.73% to 95.07%. Removing the same inputs only at test time raises it to 48.38%. Thus, nuisance-weight learning changes the training outcome beyond its immediate effect on prediction. Bias interventions expose a connection between target symmetry and shallow ReLU representations. In a compact signal-only regime, both SGD and Muon learn orders five through eight, with higher SGD peak accuracy at orders nine through eleven. Together, the results show how optimization and nuisance learning constrain the higher-order rules realized by a shallow network.
[LG-162] Metacognitive Reasoning in Energy Based Models using Instance Based Learning Theory
链接: https://arxiv.org/abs/2610.00399
作者: Tailia Malloy,Prateek Kumar Rajput,Serge Lionel Nikiema,Cleotilde Gonzalez,Tegawendé F. Bissyandé
类目: Machine Learning (cs.LG)
*备注:
Abstract:Metacognition involves reasoning about cognitive processes themselves. An example is in resource allocation where we choose how much time and effort to put into a reasoning task before we begin based on our confidence. Current Artificial Intelligence (AI) systems that rely on Large Language Models (LLMs) cannot estimate their uncertainty about an output without first responding, and cannot dynamically allocate resources to producing an output, making this type of metacognitive process difficult. A recently proposed alternative to classic transformer architectures that addresses these two concerns is the Energy Based Model (EBM) which allows for interpretable uncertainty modeling and dynamic allocation of compute resources. While EBMs can allow for control of these two processes, the actual metacognitive task of determining compute allocation based on uncertainty is not directly addressed. Instance-Based Learning Theory (IBLT) provides an approach to modeling human-like decisions from experience that has previously been applied to predicting human metacognitive reasoning. In this paper we introduce a framework for MEtacognitive Reasoning with Instance-based Learning Theory and Energy Dynamics (MERITED). Grounded in IBLT, this framework allows for control of the computational effort allocated in an EBM to allow for metacognitive control over reasoning effort based on uncertainty while remaining computationally efficient. This work has two main contributions, the training and open weight sharing of a 191M parameter reasoning EBM, and an implementation of the MERITED framework for dynamic compute allocation using an IBL model.
[LG-163] WIPSNet: Deep Learning for Paediatric Wheeze Detection from Overnight Impedance Pneumography ALT ICML2026
链接: https://arxiv.org/abs/2610.00398
作者: Felix Oury,Harley Day,Karina Mayoral,Ville-Pekka Seppä,Sejal Saglani,Reiko J. Tanaka
类目: Machine Learning (cs.LG); Signal Processing (eess.SP)
*备注: Accepted at the Workshop on Structured Data for Health, ICML 2026. Code: this https URL
Abstract:Overnight impedance pneumography (IP) is used to monitor paediatric respiratory health. Its current clinical readout, the Expiratory Variability Index (EVI), compresses each IP recording into a single scalar and achieves an AUC of 0.633 for night-level wheeze classification. We introduce Wheeze Impedance Pneumography Scalogram Network (WIPSNet), a 3D ResNet operating on stacked continuous wavelet transform scalograms of overnight IP signals. On a 15-patient cohort (60 nights, 281 hours), WIPSNet achieves an AUC of 0.783 \pm 0.026 , outperforming EVI, a state-space model (Mamba), and two modern sleep-staging architectures. Performance peaks at a volumetric depth corresponding to 32 minutes of temporal context, suggesting that multi-scale temporal aggregation is important for modelling nocturnal respiratory dynamics. Overall, these results indicate that structured time-frequency representations combined with 3D convolutional architectures provide an effective approach for learning from long, irregular physiological time series.
[LG-164] NEUROTOKEN: Joint Source and Directional AAD with Envelope Decoding via Conditional Flow Matching
链接: https://arxiv.org/abs/2610.00397
作者: Ali Alavi,Donald S. Williamson
类目: Machine Learning (cs.LG)
*备注:
Abstract:Identifying which speaker a listener is attending to in a noisy room – the cocktail-party problem – is the missing ingredient for next-generation hearing aids and brain-computer interfaces: it tells the device whose voice to amplify. Auditory attention decoding (AAD) reads this answer from EEG, but the literature splits into disconnected pieces: directional-AAD classifies side but does not map side to stream; regression-based source-AAD ranks candidate streams by a single Pearson correlation that is intrinsically noisy at the 1-5 s windows real devices need; and envelope reconstruction has no native AAD rule. We argue the right object is not any single statistic but the conditional likelihood of the attended envelope given EEG, and we make this practical with NEUROTOKEN: a single network whose three heads share one EEG front-end, with a conditional flow-matching head (ATTUNEFLOW) that scores candidates by an integrated velocity-residual likelihood ratio. Two inference-time ensembles – QUADTRACK (four complementary statistics) and ENV-FLOW (z-normalised QUADTRACK+ATTUNEFLOW) – absorb per-statistic failure modes for free. On KU Leuven, DTU, and NJU at 5 s, ATTUNEFLOW lifts per-segment source-AAD by 9%-16% over the strongest non-generative baseline and shrinks across-subject variance by ~3x; trial-level fusion exceeds 93% on two of three datasets. In parallel reproductions we show that canonical 95-97% direction-AAD numbers collapse by 17%-45% under a strict trial-disjoint protocol, clarifying both the true ceiling and why a likelihood-based formulation is needed.
[LG-165] Specificity-Aware Diffusion Steering via Variance-Reduced Sequential Monte Carlo
链接: https://arxiv.org/abs/2610.00395
作者: Luran Wang,Linrui Ma,Hannes Stärk,Regina Barzilay
类目: Machine Learning (cs.LG)
*备注:
Abstract:Inference-time steering enables pretrained diffusion models to satisfy new constraints without full retraining. However, specificity-aware generation is difficult: repelling samples from a negative reference distribution can also erode the positive distribution where the two overlap. The key challenge is to suppress negative mass while minimally distorting the positive distribution. We address this problem by formulating specificity-aware steering as a target-design problem and deriving a target distribution from an overlap-based objective. The resulting target keeps the desired reference distribution only in regions where it is sufficiently preferred over the undesired reference distribution, giving a likelihood-ratio interpretation of specificity. To sample from the corresponding time-dependent target path, we develop a Sequential Monte Carlo sampler with a variance-minimized local proposal. We further introduce a practical fixed-noise optimization procedure with the Jacobian–vector products with the desired and undesired score fields. Experiments on synthetic task, class-contrastive generation, text-to-image tasks and peptide-MHC (p-MHC) binder show that the proposed method suppresses undesired regions more effectively, reduces mode shift, and improves sampling stability by decreasing the SMC weight collapse compared with negative-guidance baselines. Code is available at: this https URL
[LG-166] Forking: Sudden Overfitting Under Replay
链接: https://arxiv.org/abs/2610.00394
作者: Shanbin Yu,Shaoyang Guo,Haoran Zhao,Danni Yu,Ziming Liu
类目: Machine Learning (cs.LG)
*备注: 42 pages, 22 figures. Code and reproduction materials: this https URL
Abstract:This paper studies forking, a generalization failure discovered in NanoGPT autoresearch. Under data replay, models with an over-encoding n-gram memory branch show a sharp separation of training and validation loss at epoch boundaries, resembling the shape of forks. We study this phenomenon in a controlled vanilla NanoGPT setting and reproduce it in a DeepSeek-style model with Engram. Mechanistically, repeated updates sharpen the continuations observed in training while suppressing the probability of unseen continuations, whose loss grows with each pass. The n-gram module creates weakly interacting context-specific subspaces, amplifying this effect. Low-frequency contexts contribute most of the gap, whereas larger training budgets and heavily crowded tables suppress it. We also observe forking in short-budget, heavily repeated SFT and RL-like regimes. The contributions of this paper are twofold: (1) Forking reveals yet another curious phenomenon in deep learning, in addition to grokking and double descent. (2) Forking is an unexpected and unpleasant by-product of tricks proposed by autoresearch agents. While these agents produce an enormous number of results that seem useful, we should always be careful with their results.
[LG-167] EvoGen-Harness: Learning Where and How to Evolve Image-Generation Harnesses
链接: https://arxiv.org/abs/2610.00383
作者: Jiabin Luo,Yinan Liu,Chunlei Meng,Yufei Guo
类目: Machine Learning (cs.LG)
*备注: 23 pages, 7 figures
Abstract:Modern text-to-image (T2I) systems can be improved without modifying generator parameters by adapting the external system around frozen generators. However, existing approaches typically optimize a predefined dimension, such as prompts, routing, or workflows, restricting the space in which generation failures can be corrected. Allowing multiple generator-external responsibilities to evolve provides a broader adaptation space, but introduces a new challenge: visual feedback reveals what failed, but not where persistent evolution should occur or how this space should be explored efficiently. We introduce EvoGen-Harness, a generator-agnostic framework for multi-responsibility image-generation harness evolution, together with Trace (Trajectory-Relative Attribution and Coordinated Evolution). Trace aggregates evidence across stochastic executions, uses failure attribution as a search prior to focus candidate updates, and progressively re-attributes residual failures to coordinate evolution across responsibilities, while No-Patch and held-out validation prevent unnecessary or harmful updates. Across GenEval2, T2I-CompBench++, and WISE, EvoGen-Harness improves over the strongest evaluated baselines by +0.2633, +0.0720, and +0.0752, respectively, while achieving 87.9-91.4% attribution recall, 94.8% No-Patch accuracy, and only 1.9% regression. These results demonstrate that attribution-guided multi-responsibility evolution can substantially enhance frozen T2I systems beyond single-dimension adaptation.
[LG-168] On the Relationship between Model Quantization and Model Inversion Attacks
链接: https://arxiv.org/abs/2610.00382
作者: Rongke Liu,Youwen Zhu
类目: Cryptography and Security (cs.CR); Information Theory (cs.IT); Machine Learning (cs.LG)
*备注:
Abstract:Model quantization reduces the numerical precision of neural network weights and activations to lower storage and computational costs. Model inversion attacks recover or reconstruct sensitive training data or inference inputs from model outputs or intermediate features, so quantization may also alter their effectiveness. However, two questions remain unresolved: How does model quantization affect model inversion? How do data characteristics influence this relationship? To address the first, we bound quantization-induced changes in mutual information between inputs and a categorical variable defined by prediction probabilities, distinguishing informational effects from attack optimization obstacles. To address the second, we identify data-dependent changes in feature distributions and inversion outcomes, with pronounced quantization sensitivity differences at 4 bits. These insights guide a privacy-aware post-training quantization method that improves inversion resistance while recovering utility. It uses a Fisher-type task-sensitivity proxy for budget-aware bit allocation, calibrates activation ranges, and jointly optimizes weight and activation scales and weight-rounding decisions with task-recovery and geometry-retention objectives and scale and rounding regularization. Experiments cover multiple metrics, neural network architectures, and face, palmprint, and iris recognition tasks. On ResNet-50, Palm at 4 bits reduces RL-MIA’s strict success from 54% to 26%, while accuracy decreases from 99.01% to 96.55% relative to FP32. Our method also supports output-level defenses: adding Stealthy Shield Defense (SSD, epsilon = 0.1) to Iris at 4.5 bits reduces BREP-MI’s strict success from 63.33% to 37.33%, while accuracy decreases from 92.8% to 87.6% relative to quantization alone.
[LG-169] OmniMed-Jev: Calibrating LVLM Confidence for Trustworthy Medical Multimodal Decisions via System One
链接: https://arxiv.org/abs/2610.00381
作者: Luyao Tang,Cheng Chen
类目: Machine Learning (cs.LG)
*备注: We introduce OmniMed-Jev, a decision-native interface based on Jev that outputs Choice, Noul or Score decisions with calibrated probabilities
Abstract:Medical models are judged not only on correctness, but on whether reported confidence matches actual accuracy. Generalist multimodal medical models have expanded what a single model can perceive, yet they still express bounded decisions such as diagnoses, findings or cell counts as generated text, so the reported probability reflects the next token rather than the decision itself. Motivated by decision-native interfaces such as Jev, we introduce OmniMed-Jev, which represents each medical decision as a Choice, Noul or Score decision over a runtime-supplied candidate set and returns a full distribution over that set: mutually exclusive classes, binary presence of a finding, or a bounded ordered value. The design is omni in three respects: it accepts diverse imaging modalities, covers different prediction tasks, and expresses them through one candidate-conditioned probability model, so heterogeneous outputs become comparable probabilities rather than task-specific strings. In an interface-controlled comparison against a generative baseline trained on the same backbone, data and schedule, OmniMed-Jev’s reported probabilities track observed correctness far more closely, reducing calibration error by up to an order of magnitude and reliability error by up to two, while point-prediction performance remains comparable; counting is the one family where the generative baseline stays ahead. Making the decision distribution the model’s output is not a format change but what turns reported numbers into probabilities that mean what they say. These results support explicit decision modeling as a way to make reported confidence meaningful within the evaluated tasks, and they are not evidence of clinical readiness: the comparison cannot separate the interface from associated training differences, which we state alongside the results. Code is available at this http URL.
[LG-170] Fusion techniques of time frequency-based images to predict the outcome of rTMS depression therapy
链接: https://arxiv.org/abs/2610.00380
作者: Wael Korani,Md Fahimul Kabir Chowdhury,Mohammed Aledhari,Reza Rostami,Reza Kazemi
类目: Machine Learning (cs.LG)
*备注: Published in the Biomedical Signal Processing and Control
Abstract:Depression is a mental condition that can lead to suicide and self-harm. Predicting the outcome of depression treatment is one of the most difficult tasks for clinicians. Among various treatment options, repetitive Transcranial Magnetic Stimulation (rTMS) is a widely used non-invasive method. Predicting rTMS response using Electroencephalogram (EEG) data is difficult because of high inter-subject variability and limited features from single-domain analysis. We introduce two fusion techniques, montage and blending, to overcome these limitations and extract richer features from EEG-derived Time-Frequency (TF) images. We then propose a lightweight custom Convolutional Neural Network (CNN) trained on fused TF representations. \textcolorblackWe use a primary dataset of 15 patients and a secondary dataset of 46 patients. We run two sets of experiments. The first set uses segment-level 10-fold cross-validation. In this setup segments from the same patient can appear in both training and testing. The Montage CWT_ST fusion reaches 99.90% accuracy on the primary dataset and 91.90% on the secondary dataset. The second set uses strict subject-disjoint cross-validation. All segments of a patient stay in one fold and no patient appears in both training and testing. Performance collapses. We test four time-frequency methods, six fusion mechanisms, and fourteen model architectures. With one exception, every configuration on both cohorts falls between AUC 0.31 and 0.54 and every 95% confidence interval contains 0.5. A patient-level permutation test on the best standalone method returns p = 0.703 . The best subject-level result is Montage CWT_ST on the primary cohort, which reaches AUC 0.874 \pm 0.183 and 82.7% accuracy.
[LG-171] STCFormer: Adaptive Spatio-Temporal Modeling with Dynamic Cluster Transformer for Station-based Weather Forecasting
链接: https://arxiv.org/abs/2610.00377
作者: Rongwen Li,Haixin Xie,Mingyang Wang,Hongwu Liu,Kun Fang,Changjian Chen,Zhuo Tang,Kenli Li
类目: Machine Learning (cs.LG); Machine Learning (stat.ML)
*备注: 34 pages, 12 figures
Abstract:Station-based weather forecasting supports daily life and economic activity, yet accurate forecasts require modeling complex spatial dependencies among stations. Recent clustering-based selective modeling offers a promising alternative to dense inter-station interactions. However, a grouping shared across an observation window may obscure local changes in station relationships, while intra-cluster interactions alone may miss important global context. The theoretical advantages of selective interactions over dense connectivity also remain insufficiently understood. We therefore propose STCFormer, an adaptive spatio-temporal Transformer that dynamically groups stations according to their local evolution within each temporal patch. Its Cluster-Guided Attention Block combines fine-grained local attention within clusters and global attention over regional state summaries, allowing each station to access information beyond its own cluster. We further show that a derived Lipschitz upper bound for cluster-conditioned local attention is no larger than its fully connected counterpart, explaining a potential robustness benefit and motivating the design of InfoLoss. Experiments on three real-world weather datasets spanning eight temperature and wind forecasting tasks show that STCFormer achieves the lowest 24-hour mean squared error on all eight tasks and ranks first or second in 47 of 48 comparisons across metrics and forecasting horizons. Ablations and case studies further confirm the benefits of locally adaptive grouping and complementary local-global interactions. Our code can be obtained at this https URL.
[LG-172] A First Glance at Jev for Network Traffic Classification: Accuracy Processing Time and Cost
链接: https://arxiv.org/abs/2610.00376
作者: Shenghe Xu,Lifan Mei
类目: Machine Learning (cs.LG)
*备注:
Abstract:We evaluate Jev on ten dataset-defined application labels in CESNET-QUICEXT-25 using only the first ten packets’ sizes, directions, and inter-packet times. To the best of our knowledge, this is the first empirical study of general-purpose decision models, represented here by Jev, for application classification of network flows. Across 52,000 records from 26 collection weeks following the training period, 40 fixed labeled examples raise Jev’s accuracy from 9.80% to 28.42%. Random Forest and Extra Trees trained on 8,000 records achieve 69.95% and 66.80% and outperform Jev in every week. Increasing Jev’s context to 150 examples yields 34.50% on the first test week. On a paired 100-record subset, Jev with 40 examples achieves 29% accuracy at a median request time of 0.750 s, versus 37% and 6.036 s for the generative language model OpenAI GPT-5.6 Sol with high reasoning effort through Azure; Jev also incurs lower API charges. The paired subset does not establish an accuracy advantage for either service, and the timing reflects different service configurations. Thus, labeled examples substantially improve Jev, but the tested Jev configurations remain less accurate than trained tree ensembles; unequal supervision budgets and fixed configurations prevent attributing the gap to a single cause.
[LG-173] M2Weather: A Benchmark for Joint Multi-Station and Multi-Variable Weather Forecasting
链接: https://arxiv.org/abs/2610.00370
作者: Rongwen Li,Xiao Wang,Mingyang Wang,Hongwu Liu,Changjian Chen,Zhuo Tang,Kenli Li
类目: Machine Learning (cs.LG); Machine Learning (stat.ML)
*备注: 36 pages
Abstract:Station weather forecasting is fundamentally shaped by both complex spatial dependencies across stations and strong physical coupling among weather variables. However, existing studies often consider these relationships separately and use different datasets and experimental settings, hindering systematic assessment of their individual and joint contributions. In this paper, we introduce M^2 Weather, a benchmark for joint multi-station and multi-variable weather forecasting. Through multi-criteria quality control and station stratification, we collect 2,809 high-quality stations with 5 physically coupled weather variables across three spatial scales: France, Europe, and Global. This multi-scale design lets us examine whether conclusions persist from national to global station networks. We also introduce unified training and evaluation protocols to enable fair comparison of different station-variable modeling paradigms. To further examine the benefits of modeling station-variable relationships, we design a lightweight, plug-and-play adapter. With a trained weather forecasting model, this adapter can introduce missing station or variable relationships without retraining the model. This enables fair and efficient investigation of station-variable relationships. Systematic evaluation of 16 representative models shows the benefits of jointly modeling station and variable relationships. Completing missing relationships further reduces MSE for all adapted models on all three datasets. Together, these results identify the complementary information across stations and variables as an important resource for improving station weather forecasting. Our code can be obtained at this https URL.
[LG-174] MoRA: MoE Pruning via Router Bias Learning and Expert Approximation
链接: https://arxiv.org/abs/2610.00367
作者: Yushuai Sun,Zikun Zhou,Lin Gao,Jun Yu,Wenjie Pei
类目: Machine Learning (cs.LG)
*备注: 13 pages, 3 figures
Abstract:Mixture-of-Experts (MoE) models enable parameter scaling with limited per-token computation by activating only a small subset of experts for each token, but deploying them still requires loading the complete expert pool into memory. Structured expert pruning can effectively reduce the memory usage by removing experts. However, existing pruning methods either use expert ranking criteria that are not well aligned with model performance or rely on effective expert subset searching that is computationally expensive. Moreover, these methods typically overlook the routing-behavior redundancy among the retained experts. In this paper, we propose MoE Pruning via Router Bias Learning and Expert Approximation (MoRA), a framework for structured MoE expert pruning. We introduce a learnable router bias for each expert and optimize these biases by minimizing the language-modeling loss and a routing-diversity regularizer. The learned router biases sharpen the routing probability distributions to identify experts critical to model performance while encouraging the selection of experts with diverse routing preferences. In addition, we introduce an expert approximation mechanism as a post-pruning enhancement. It leverages the remaining experts to approximate the outputs of pruned experts by affine transformation, further improving the performance of the pruned model. We evaluate MoRA on Qwen3-30B-A3B, DeepSeek-V2-Lite, and Moonlight-16B-A3B, removing 25% and 50% of the routed experts in each MoE layer. Extensive experiments on nine zero-shot benchmarks show that MoRA outperforms state-of-the-art pruning algorithms. Our code will be released.
[LG-175] Removing the NEEDLE in the Haystack: Backdoor Removal in LLM s via Weight Orthogonalisation
链接: https://arxiv.org/abs/2610.00348
作者: Minoo Kim,Vasileios Lampos,George Drayson
类目: Cryptography and Security (cs.CR); Machine Learning (cs.LG)
*备注:
Abstract:Backdoor attacks can be implanted in Large Language Models (LLMs) during training, causing unwanted behaviour when a trigger appears in the input. Existing backdoor defences for LLMs attempt to remove the backdoor but inadvertently shift the model’s output distribution to benign prompts, which can result in degraded model performance and safety. We propose NEEDLE, a training-free method for targeted backdoor removal. Once a trigger has been identified, our method estimates a backdoor direction and a refusal subspace through activation vectors, then applies sequential weight orthogonalisation to suppress the backdoor while preventing changes in refusal-related representations. NEEDLE requires neither a clean reference model nor the original poisoned training data. Evaluation is conducted across multiple model families and attack types. NEEDLE achieves the lowest mean Attack Success Rate (ASR) among the evaluated defences, including 0% on challenging code injection attacks, while resulting in the lowest KL divergence and minimal changes in capability and safety.
[LG-176] Benchmarking System One decision models against trained classifiers and language models for automated decision gates
链接: https://arxiv.org/abs/2610.00346
作者: Amir Rafe,Subasish Das
类目: Machine Learning (cs.LG)
*备注:
Abstract:Software that hands branching decisions to a model needs a declared option and a probability it can threshold. Typed decision models, also called System One models, return such probabilities without generating text, while supervised classifiers and generative language models are the established alternatives. Under matched conditions, one harness sends eight decision-model checkpoints from six families, including the hosted model Jev, and two generative comparators the same semantic requests, and scores supervised and zero-shot classifiers on the same workflow, intent and social-science items. The ranking of the model classes depends on the conditions. With the task’s own labels, small trained classifiers are the most accurate on intents and not significantly different from the best decision models on workflows. Without labels, every decision model except the encoder-based checkpoints exceeds a zero-shot entailment classifier on workflows and intents. Read through option-key likelihoods, a larger generative model is level with Jev on workflows and intents and accepts more workflow decisions at five percent risk, and fine-tuned decision checkpoints gain intent accuracy over their untuned backbones. Stored temperatures fitted on few options raise calibration error with many options, and a held-out threshold for five percent in-scope risk still lets Jev accept 0.310 of out-of-scope requests. Swapping yes and no flips 50.5 answers per hundred for Jev, while fine-tuned checkpoints cut their backbones’ social-science flips. An intent-trained first stage escalating to Jev matches its accuracy at 0.43 of its cost at full graphics-processor utilization. The results yield condition-dependent design rules for automated decision gates.
[LG-177] Beyond Diagonal State Space Models: Exact Non-Abelian Group Tracking Solvability Barriers and Geometric Physical Manifolds
链接: https://arxiv.org/abs/2610.00329
作者: Zeyu Jia(School of Biomedical Engineering and Technology, Tianjin Medical University, Medical School, Tianjin University)
类目: Machine Learning (cs.LG)
*备注: 26 pages, 1 table, 5 theorems. Source code and reproducible benchmarks available
Abstract:Selective state space models (SSMs), such as Mamba, S4D, and LRU, are bounded by transition matrix commutativity (A_t A_t’ = A_t’ A_t) and solvable affine transformation groups (Aff_D of derived length = 2). Consequently, stacked multi-layer diagonal networks face severe optimization degradation on non-solvable simple groups such as A_5 due to the exponential circuit emulation depth required to simulate non-abelian commutators. We propose Non-Commutative State Space Models (NC-SSM), their real-orthogonal counterpart SO(3)-SSM, and arbitrary-dimension Cayley-SSM, lifting state transitions to compact Lie groups SU(2), SO(3), and SO(N). Via closed-form Euler-Rodrigues maps and rational Cayley transforms, NC-SSM achieves exact norm-preserving isometry (||U_t|| = 1). We introduce pure Hopf-fibration Bloch projective readouts (S^3/±1 =~ S^2 =~ SO(3)) to eliminate sign ambiguity, true quaternion parallel prefix scans (9.06x speedup at T=2048), and Identity-Gated Lie SSMs to eliminate sparse syntax phase drift. Extensive benchmarks across 14 experimental regimes show: (1) NC-SSM achieves 100% tracking on S_3, D_4, Q_8 and simple group A_5, where a 3-layer deep diagonal baseline collapses to 6.60% (p = 8.81e-4); (2) Cayley-SO(5)-SSM breaks Klein’s 1884 ceiling on symmetric group S_5 (50.92% vs diagonal 5.25%, p = 0.0015, delivering 7.8x variance reduction over SO(3)); (3) SO(3)-SSM preserves Riemannian manifolds across 300 steps ( 3.12e-6 drift, 580,000x advantage), achieving 0.04 deg dead-reckoning error and active tangent denoising; (4) NC-SSM achieves 74.36% on Dyck-2 and 30.26% on deep AST scope tracking (p = 0.0081); and (5) ablation confirms strict isometry is mathematically necessary for lossless long-range associative memory.
[LG-178] Bellm an-Certified Rounding for Sparse Policy Deployment in MDPs
链接: https://arxiv.org/abs/2610.00325
作者: Zhaojun Peng
类目: Machine Learning (cs.LG)
*备注:
Abstract:Continuous policy optimization may spread an update across many states, even when deployment permits only a few complete state-level changes. We study how much discounted return can be retained when continuous row mixtures are rounded to sparse binary policies in finite MDPs. Policy-dependent visitation couples the row edits, while long horizons make global curvature bounds conservative. From 2d+2 Bellman solves, we derive reusable envelopes that support uniform and candidate-specific guarantees before rounding. A rank-two rational representation of each exchange further permits weighted curvature integration along the realized trajectory. We prove that linear dimension dependence is unavoidable when the budget scales, and that exact global curvature thresholding is hard. Candidate-specific bounds raise pre-rounding certification coverage from 48.2% to 74.1% on the structured suite. At \gamma=0.95 , local integration lowers the median bound-to-loss ratio from 402.3 to 2.08 on coupled instances.
[LG-179] Stable and Counterfactually Robust Physical World Models from Imposed Structure and Learned Physics
链接: https://arxiv.org/abs/2610.00280
作者: Yufeng Wang,Parivesh Priye,Lu Wei,Haibin Ling
类目: Machine Learning (cs.LG)
*备注:
Abstract:A world model learns to forecast how a physical system evolves from recorded trajectories, yet the systems it imitates obey physical laws that are neither fully supplied nor reliably respected. The model may create energy, drift or diverge over long rollouts, and answer a changed law query using the law observed during training. We ask how much general physical structure must be hard coded into a world model, and how much system-specific physics can then be learned from data, for four properties to hold simultaneously: second law compatible dissipation, correct responses to interventions on physical parameters, stability out to one hundred times the training horizon, and robustness to disturbances. The imposed structure is general: dynamics are generated from the gradient of a learned energy through a fixed reversible operator, the energy is restricted to a confining class, a one way port can remove energy but never inject it, the drive channel is known, and the intervened parameter enters through a separable map. The model learns the energy functional, constitutive relations, dissipation rate, and couplings. Across an electromagnetic cavity, a particle in cell grid, and a shallow-water fluid, models with roughly nine thousand parameters recover constitutive functions with unit slope, separate conserving from dissipating worlds by four orders of magnitude using a single set of weights, and transfer changes in sign, magnitude, rate, and gravity to unseen values, where equal-capacity models without the same structure perform at chance or worse. A nonlinear constitutive law is recovered with its curvature preserved and predicts a held-out intervention 2 - 17\times better than a converged linear model.
[LG-180] Coupling Perception and Reasoning in Federated Multimodal Graph Foundation Models
链接: https://arxiv.org/abs/2610.00277
作者: Zekai Chen,Xun Wu,Hailin Zhang,Xunkai Li,Yu Liu,Kairui Yang,Muyan Huang,Xuaner Chen,Rong-Hua Li,Guoren Wang
类目: Machine Learning (cs.LG)
*备注:
Abstract:Federated multimodal graph foundation models (GFMs) aim to adapt pretrained multimodal models to decentralized graph data, where each client owns a private multimodal graph and cannot share raw information. These models typically combine a multimodal Encoder that extracts semantic evidence from heterogeneous modalities and a graph neural network (GNN) that performs relational reasoning over graph structures. However, existing federated GFM adaptation methods mainly update graph-side modules while keeping the multimodal Encoder frozen, limiting adaptation to \emphhow information is propagated while fixing \emphwhat information is extracted. Through empirical studies, we reveal that Encoder and GNN adaptations are not independent: Encoder adaptation is affected by graph relations, while cross-client module swapping reveals substantial pairing sensitivity between separately parameterized Encoder and GNN updates. Motivated by this observation, we propose \textbfFedCORE, a federated adaptation framework that represents Encoder and GNN updates through a shared low-dimensional latent state. FedCORE jointly optimizes this core from multimodal and structural signals and performs federated evolution directly in the shared state space, preserving compatibility between perception and reasoning adaptations. Extensive experiments demonstrate that FedCORE reduces the Encoder–GNN pairing gap from 30.6 to 5.9 , corresponding to an 80.7% reduction over independent joint adaptation.
[LG-181] MOVE: Multimodal Open-world Verification and Expansion for Graph Learning
链接: https://arxiv.org/abs/2610.00268
作者: Zekai Chen,Jiayang Xing,Xun Wu,Miao Zhang,Xunkai Li,Kairui Yang,Zhengyu Wu,Xu Wang,Rong-Hua Li,Guoren Wang
类目: Machine Learning (cs.LG)
*备注:
Abstract:Multimodal graph learning faces a fundamental challenge: new classes may emerge after deployment, while models are trained with a fixed label space. Existing approaches typically detect unknown nodes and use LLMs to generate candidate class descriptions, but they do not determine whether existing classes are insufficient to cover these nodes or whether a generated class is reliable enough to expand the class space. Our empirical study reveals three challenges: multimodal information beyond individual modalities is required for unknown-node identification, LLM-generated class descriptions may not fully capture multimodal class characteristics, and directly adding candidate classes can introduce redundant categories. Based on these observations, we propose MOVE, a multimodal open-world class verification and expansion framework. MOVE identifies nodes that cannot be assigned to existing classes by jointly considering visual tokens, textual attributes, and graph context, leverages a multimodal LLM to generate candidate classes, and selectively expands the class space only when candidates are consistently supported by multimodal evidence without introducing unnecessary categories. Experiments demonstrate that MOVE achieves an average improvement of 11.87% across unknown recognition, open-domain annotation, and downstream graph learning tasks.
[LG-182] Large Language Bayes Is Not Reparameterisation-Invariant
链接: https://arxiv.org/abs/2610.00265
作者: Jian Xu
类目: Machine Learning (cs.LG)
*备注:
Abstract:Large Language Bayes (LLB) answers an informal modelling question by sampling candidate probabilistic programs from a language model, running approximate inference on each, and averaging them with weights proportional to an exponentiated evidence bound. We show that this weighting depends on how a model is written. The log marginal likelihood is invariant to reparameterisation; the evidence bound is not. On eight schools the centered and non-centered programs are the same measure to 5.7\times10^-14 , yet their weights differ by 6.1\times ; importance weighting reduces this only to 2.2\times , and reproducing the inference LLB actually runs, a full-covariance Gaussian matched to the posterior moments, still leaves 1.9\times on eight schools and 8.9\times in 64 dimensions. Across likelihood families, dimensions and funnel severities the discrepancy reaches 31.9\times and reverses sign, so no single writing is uniformly preferable. It inverts Bayes factors against eight natural competitors, and the induced error in the model posterior, and in any downstream target, is controlled by the spread \Delta of the bound shortfalls through a known sharp Hilbert-distance bound. Across 360 programs from six language models the parameterisation written ranges from 0% to 100% centered and is stable within a model. Detecting equivalent programs statistically can falsely merge genuinely different models at practical sample budgets; verifying reparameterisations we generate ourselves cannot, and closes the window.
[LG-183] Attention Manifolds: Steering or Blocking Language Models by Editing Learned B-Spline Surfaces
链接: https://arxiv.org/abs/2610.00257
作者: Naveen Mysore
类目: Machine Learning (cs.LG)
*备注:
Abstract:In standard transformer attention, a source token sends the same value vector to every receiver. The query determines \emphhow much to attend but not \emphwhat to extract. This work introduces \textbfattention manifolds: learned 2D B-spline surfaces S_d(q_d, k_d) that modulate each value dimension based on the query-key interaction. Each surface is a tensor-product cubic B-spline initialized to zero, preserving pretrained behavior. Applied to LLaMA 3.2-1B-Instruct and 3B-Instruct, attention manifolds reduce WikiText-2 validation perplexity by 2–2.5 points with 0.3% parameter overhead. Across 112 diverse prompts, surfaces change greedy-decoded output for 69% (1B) to 83% (3B) of cases, with the strongest effects on ambiguous and polysemous inputs (94–100% change rate). The surfaces improve output quality: correcting factual errors (\emphthe CAP theorem has three main components'' \to \emphit is impossible to guarantee all three’‘), increasing precision (\emphimpossible to know certain properties'' \to \emphimpossible to know both position and momentum’‘), and adding specificity (a generic quote \to an attributed Saint Augustine citation, consistently at both scales). The learned surfaces are also mechanically editable: inverting a layer’s coefficients changes greedy output for 9/10 prompts (KL~0.010), providing a geometric mechanism for model steering. Setting surface coefficients to -1 creates ``attention walls’’ that block value flow through specific dimensions. In a preliminary experiment, a layer-wide wall redirects an explosive-device prompt from specific instructions to general educational content, suggesting a path toward safety-oriented manifold shaping.
[LG-184] Verification Pulses and the Cost of Escaping Wrong Consensus
链接: https://arxiv.org/abs/2610.00256
作者: Shivam Gupta
类目: Machine Learning (cs.LG)
*备注: 18 pages, 5 figures; code and raw experimental records: this https URL
Abstract:External verification can correct individual outputs while leaving a self-reinforcing population in the basin of a wrong consensus. We study how the timing and addressing of a fixed verification budget affect recovery in an asynchronous binary register. For a general nonlinear response, we derive the minimum fuel required to cross a basin boundary under a peak verification constraint. For a finite population, an exact birth–death calculation gives the probability of subsequent wrong consensus after a pulse. Our main asymptotic result identifies the critical budget window: a leading term N\log(x_0/b) and a correction of order \sqrt N , with separate variance contributions from repeated verification targets and autonomous amplification after verification stops. The distinction is substantial: with 16 majority-updated slots and 14 initially wrong, 9 random checks cross the mean-field budget threshold, whereas 23 are required for 95% eventual recovery in the exact model. A prospectively specified experiment records 13,392 language-model responses, including calibration and 108 held-out trajectories. Calibration produces different fitted response regimes, but all four adjusted schedule-comparison intervals include zero. A distributional audit also finds that modest mean-prediction error can conceal a large underestimate of terminal consensus occupancy. The results support risk-calibrated reset scheduling under a specified update contract, while explicitly separating it from distinct-target checking and unrestricted evidence broadcast.
[LG-185] Sharp Oracle-Regret Tradeoffs for Projection-Free Online Convex Optimization
链接: https://arxiv.org/abs/2610.00254
作者: Vaneet Aggarwal
类目: Machine Learning (cs.LG)
*备注:
Abstract:We characterize the regret attainable in online convex optimization when access to the feasible set is limited to an exact linear optimization oracle. The learner is given an inscribed ball and a diameter bound and must remain feasible on every consistent instance. For convex G -Lipschitz losses, diameter at most D , a total allowance of Q oracle calls, and a strict limit of B calls per round, the dimension-free minimax expected regret is \Theta(GD\max\sqrt T,T/(1+\min\Q,BT)^1/4) . The lower bound applies to arbitrary randomized learners. Universal feasibility first forces each action into the hull of the supplied ball and the preceding oracle replies. A fixed-body construction then couples fresh phase directions to a shared simplex, making useful replies costly repeatedly even though all losses have a common minimizer. A counted approximate-gradient method with interleaved blocks attains the matching rate. Total-budget and strict per-round guarantees follow as special cases, including the T^3/4 rate with one call per round and the quadratic total budget needed for \sqrt T regret. For prescribed smoothness \beta , an analytic construction yields a curvature-dependent lower bound and identifies the threshold above which the general characterization remains sharp.
[LG-186] he Null Is the Hard Part: Exact Tests for Memorization in Generative Models
链接: https://arxiv.org/abs/2610.00251
作者: Sushovan Majhi,Pramita Bagchi
类目: Machine Learning (cs.LG)
*备注: 23 pages, 5 figures. Code and measurement outputs at this https URL
Abstract:Memorization audits of generative models read similarity scores against thresholds, with no null distribution, and the conclusions they support can be wrong. By MemBench’s rule, the benchmark’s mitigations roughly halve Stable Diffusion’s memorization; audited with false-discovery control, two thirds of the certified images are no longer detected under random prompt perturbations, five sixths under attention rescaling, and all of them under embedding optimization. The field’s data-copying test, read against its own null, flags ten of twenty-four generators that reproduce nothing. We argue that for memorization the null is the hard part, and supply two. For a whole model, training and held-out images are exchangeable given its samples, and relabelling them is a permutation test, exact for any statistic when the held-out images are a random split; under it, a nearest-neighbour preference still fires on seven of those twenty-four, and a count restricted to the near-duplicate scale on none (McNemar p=0.016). For single images, the natural nulls fail twice, measurably: ranking an image among random images yields 596 false discoveries among 2,365 controls, and resampling independent generations makes the null three times too narrow. Calibrated against matched controls, the audit certifies 36 of 61 MemBench images at 5% false-discovery rate, held-out controls are certified in 0.01% of calibration splits, and on this benchmark two generations per image recover that count. A calibrated maximum, which reads occasional rather than typical copying, certifies 46. As the scale-restricted statistic we recommend the small-scale mass of the Intersection Euler Characteristic Profile, which also counts distinct images copied and tests whether two models copy the same ones.
[LG-187] Contingent Exposure Routing for Financial AI: Outage Risk and the Cost of Indivisible Decisions
链接: https://arxiv.org/abs/2610.00239
作者: Shivam Gupta
类目: Machine Learning (cs.LG)
*备注: 15 pages, 6 figures, 3 tables. Code and data: this https URL
Abstract:Model failover restores availability, but changes which financial institutions share decision errors. We formulate outage-contingent routing through a local market-impact response matrix and study expected squared price displacement. A symmetric construction shows that a shared backup can leave an order-one concentration floor as the number of primary endpoints grows, while balanced fallback risk decreases inversely with the surviving endpoint count. For indivisible decisions, we derive the exact second moment of independent randomized routing and an effective-exposure granularity that determines its gap from fractional allocation. Conditional-expectation rounding gives a finite-agent bound without coupled quotas; a separate swap procedure preserves endpoint counts and is assessed against dual lower bounds. Across 60 synthetic portfolio networks and 11,340 scenario evaluations, the latter reduces risk by 6.57% and 10.53% for single and double endpoint removals at the central feedback setting with independent errors. A replay of 1,024 recorded API responses on constructed rebalancing tasks gives a smaller held-out reduction of 3.30% (paired bootstrap interval 2.07–4.57%). Strongly aligned errors, inferior endpoints, and indivisibility limit diversification. The contribution is an auditable routing stress test and implementation analysis, not an estimate of real-market crash probabilities.
[LG-188] Compositional Embedding Architecture for Physical Field Prediction in Componentized Aerospace Systems
链接: https://arxiv.org/abs/2610.00237
作者: Qineng Wang,Xinrui Zhou,Shuwen Yue,Kangli Bao,Hairun Xie,Yonghe Zhang
类目: Computational Engineering, Finance, and Science (cs.CE); Machine Learning (cs.LG)
*备注: 32 pages, 13 figures
Abstract:Spacecraft thermal design requires repeated evaluation of how variations in the number and spatial arrangement of heat-generating components and in thermal boundary conditions affect the temperature field. High-fidelity numerical simulations are computationally expensive and therefore difficult to use for large-scale design screening. Although surrogate models can accelerate temperature-field prediction, existing approaches generally encode each complete configuration as a whole and do not explicitly exploit the reusability of local physical constituents across configurations, which limits their accuracy for component counts and combinations not covered during training. To address this issue, we propose the Tree-Structured Factor Composition Network (TFCN), which decomposes complex spacecraft thermal configurations into reusable local physical factors and employs a tree-structured composition module to learn the global temperature-field response associated with different factor combinations. TFCN is evaluated on two-dimensional steady-state spacecraft thermal-analysis cases with prescribed-temperature and radiative-flux boundary conditions. The model is trained exclusively on configurations containing no more than 15 heat-generating components and evaluated on unseen configurations containing 16-25 components. For the prescribed-temperature and radiative-flux cases, TFCN achieves component-count out-of-distribution RMSE values of 4.21 K and 18.62 K, respectively, representing reductions of 65.6% and 33.6% relative to the strongest baseline. These results demonstrate that TFCN improves the reliability of temperature-field prediction under variations in component count and provides an efficient surrogate for rapid spacecraft thermal-design evaluation and large-scale configuration screening.
[LG-189] How Many Categories Are Enough? Distribution-Free Certification Limits for Few-Shot Anomaly Thresholds
链接: https://arxiv.org/abs/2610.00236
作者: Gia Huy Thai,Nguyen Thai Anh
类目: Machine Learning (cs.LG)
*备注:
Abstract:Few-shot anomaly detectors are judged by ranking metrics, yet deployment requires an alarm threshold with a controlled false-alarm rate (FAR). We ask how much normal evidence, in images or category units, is needed to certify such a threshold for an unseen category. Using a frozen DINOv2 principal component analysis (PCA) residual ranker on 15 MVTec and 12 VisA categories under four corruption types, we show that target-only leave-one-image-out (LOIO) calibration is resolution-limited and shift-fragile: rank values cannot fall below 1/(k+1) , and at the attainable level \alpha=0.20 , empirical FAR reaches 0.341 on Gaussian-corrupted MVTec at k=4 , 1.7 times the nominal level. A category-count feasibility calculus is then derived: even with all-zero category losses and no multiplicity charged, any deterministic, uniformly valid, distribution-free 95% upper confidence bound (UCB) requires at least 14, 29, and 59 independent and identically distributed (iid) category draws at \alpha=0.20 , 0.10 , and 0.05 ; these counts are necessary but not sufficient. The Cross-category Reliability Estimation with Source Support (CRESS) protocol splits source categories into disjoint reference, proposal, and certification roles. With only three or four certification categories, all 960 frozen configurations return the fail-closed threshold \tau^\star=0 , and the smallest category-level UCB is 0.950. Image-unit analyses of the same archive select positive thresholds in 36.7% to 60.3% of target cells; these bounds hold for the selected source mixture, not for the marginal risk of a new-category draw. The contribution is a quantitative feasibility boundary and an estimand-aware protocol specifying when source evidence can, and cannot, support a transferable reliability claim.
[LG-190] Constant-Memory Recall: Learned Associations in a Fixed Matrix State
链接: https://arxiv.org/abs/2610.00232
作者: Samuel Larson
类目: Machine Learning (cs.LG)
*备注: 8 pages, 3 figures
Abstract:Fixed-size recurrent memory limits storage growth during inference, but successful recall depends on the task and training. We study a small DeltaNet variant with fixed token-specific key biases, trained to remember 32 new key-value pairings per sequence. With 32 KiB of recurrent matrix state, it achieves 99.95% mean accuracy across three training seeds when choosing among the sequence’s values. Recall remains near perfect when filler extends the pre-query context to 1,798 tokens without adding pairings. Zeroing the first memory block removes this recall. An exploratory 48-pair test remains near chance after one quarter of the primary training budget and does not locate a capacity limit. Parameter-matched vector and Transformer baselines remain near chance, including the Transformer after additional training searches. This unresolved baseline failure prevents a memory-efficiency comparison.
[LG-191] Energy Time-Series Imputation with Differentially Private Diffusion Models via Clipping-Aware Objective Conditioning
链接: https://arxiv.org/abs/2610.00209
作者: Huizhen Huang,Yu Li,Tao Huang,Chen Hou
类目: Machine Learning (cs.LG)
*备注:
Abstract:Reliable recovery of missing measurements is important for monitoring and analysis in energy time-series systems, where fine-grained measurements may contain sensitive temporal information. Diffusion models trained with differentially private stochastic gradient descent (DP-SGD) provide a promising framework for privacy-sensitive energy time-series imputation. Under cosine diffusion schedules, late timesteps correspond to low signal-to-noise ratio (SNR) conditions, where standard \varepsilon -prediction can induce large pre-clipping gradients. Such gradients are more likely to be clipped, reducing the retained optimization signal. The artificial intelligence (AI) contribution lies in formulating this objective–clipping interaction as an objective optimization problem under fixed-threshold DP-SGD and developing timestep-aware objective conditioning for diffusion-based energy time-series imputation. The method adopts v -prediction to mitigate late-timestep gradient amplification, uses static loss weighting as a uniform-scaling control, and introduces diffusion-schedule-aware dynamic weighting for stronger attenuation before clipping. For the engineering application, we evaluate the method on five real-world energy time-series datasets across random point missingness, contiguous block missingness, persistent outages, and multiple missing-data severities. Under matched DP-SGD settings, the proposed method consistently improves imputation utility over the \varepsilon -prediction baseline. Gradient diagnostics reveal lower upper-tail pre-clipping gradient norms, reduced clipping fractions, and stronger attenuation at late low-SNR timesteps, supporting the effectiveness of clipping-aware objective conditioning for energy time-series imputation.
[LG-192] Classification Based on Association Rules Algorithm for Breast Cancer
链接: https://arxiv.org/abs/2610.00174
作者: Ali Alsalama,Ahmed Kubba,Ghaith Jamjoum,Zaher Al Aghbari
类目: Machine Learning (cs.LG)
*备注: 6 pages, 1 figure, 1 table, accepted presented at Advances in Science and Engineering Technology International Conferences (ASET) 2024
Abstract:Breast cancer is a significant contributor to female mortality across the world, displaying one of the highest oc currence rates among the various cancer types. In response to the need for early breast cancer detection, researchers have increasingly turned to association rule-based classification as a favored method. Association Rule mining is a data mining approach which offers the benefit of yielding results that are readily understandable for medical professionals. This paper introduces a novel association rule-based data mining technique for breast cancer classification based on a weighted classification approach. This implementation employs three core algorithms: Rule Generation, Rule Pruning, and Rule Prediction. Rule Generation identifies frequent itemsets and creates association rules. Rule Pruning eliminates rules using specific criteria and separates them into major and minor groups based on their influence on training data. Rule Prediction applies the pruned rules to classify test data. The final prediction algorithm was tested on several testing samples to show the feasibility and performance of the approach.
[LG-193] Generalized Biomedicine Discovery ECCV2026
链接: https://arxiv.org/abs/2610.00120
作者: Luyao Tang,Yingkai Yang,Hanqi Chen,Jiewei Zheng,Chaoqi Chen,Cheng Chen
类目: Machine Learning (cs.LG)
*备注: Accepted by ECCV 2026
Abstract:In real-world clinical practice, medical images face open-world shifts: (i) long-tailed rare diseases, (ii) subtle lesions dominated by normal anatomy, and (iii) hierarchical taxonomies. Yet most open-world paradigms assume flat, balanced label spaces, leaving these biomedical demands unresolved. We introduce Generalized Biomedicine Discovery (GBD) and a unified benchmark spanning long-tail, anomaly, and taxonomy-aware discovery. Our key insight is that dominant known patterns form a visual manifold that masks subtle novelty. Inspired by expert diagnosis, we propose SCAN (Surprise-evoked Complementary AccommodatioN), which follows a cognition-inspired perceptual progression: it applies predictive suppression to filter expected norms, triggers surprise-evoked salience to highlight unexpected deviations, and performs complementary accommodation to integrate these shifts into global representations. Extensive experiments show that SCAN improves novel concept discovery while generally preserving established clinical knowledge, and it plugs into existing architectures to better navigate the known-unknown trade-off in medical imaging. Code is available at this https URL.
[LG-194] he Hidden Costs of 99% Accuracy: A Trustworthiness Audit of the Telco Customer Churn Benchmark
链接: https://arxiv.org/abs/2610.00118
作者: Soumyadeep Roy
类目: Machine Learning (cs.LG)
*备注:
Abstract:Customer churn prediction on the IBM Telco Customer Churn benchmark (n = 7,043) routinely reports test accuracies above 95%, with the most cited published study reporting 99.01%. We audit this benchmark for four trustworthiness failures invisible to the accuracy- and F1-centred reporting that dominates the literature. First, pre-split SMOTE inflates churn-class F1 by 13.1 percentage points across ten classifiers and fifteen seeds (Wilcoxon p 10^-4 per classifier); the same leaky pipeline ordering paired with class weighting yields no inflation, isolating the effect to SMOTE’s geometric construction. We measure the mechanism directly: approximately 36% of synthetic training points are nearest-neighbour interpolations of test-set instances. Second, the TotalCharges field is approximately determined by tenure multiplied by MonthlyCharges (R2 = 0.999); removing it changes accuracy by less than 0.2 percentage points, yet TreeSHAP ranks it ninth in mean absolute attribution - a pattern that materially corrupts SHAP-based interpretation. We propose an R2 0.95 pre-modelling diagnostic. Third, in a 15-seed calibration audit, isotonic regression is the strongest default; temperature scaling fails on class-weighted tree ensembles whose predicted-probability distribution is bimodal. Fourth, the cost-optimal decision threshold (under a 50 USD retention offer and 24-month CLV proxy) is approximately 5-10 times lower than the F1-optimal threshold, saving approximately 77,000 USD per 1,000 customers. We replicate F1 and F2 on Iranian Telecom Churn (within domain) and Bank Customer Churn (across domain): F1 generalises; F2 generalises only within telecom. We synthesise these findings into a four-component reporting checklist - pipeline disclosure, redundancy diagnostic, calibration audit, and cost-sensitive thresholds - and release a reproducible implementation.
[LG-195] SyntheticHLS: Building Diverse Synthetic High-Level Synthesis Datasets using LLM s
链接: https://arxiv.org/abs/2610.00106
作者: Stefan Abi-Karam,Miaoyan Zhou,Callie Hao
类目: Hardware Architecture (cs.AR); Machine Learning (cs.LG)
*备注: Accepted and to be presented at the International Conference on Field Programmable Technology (FPT) 2026
Abstract:Deep learning and large language models (LLMs) are rapidly gaining adoption in semiconductor design, driving demand for training datasets. Most efforts focus on hardware description languages (HDLs) while designs for high-level synthesis (HLS), a popular approach to domain-specific accelerators, remain scarce. HLS dataset efforts emphasize manual curation or design parameterization, seldom addressing high-quality LLM-based generation or diversity in code length, hierarchy, design-space size, latency, resource utilization, and application domain, potentially limiting model generalization. We propose SyntheticHLS, a framework for generating large-scale, complex, diverse synthetic HLS datasets using LLMs. Its two key ideas are: 1) an iterative feedback-guided mutation loop that uses paired HLS source code and design-space specifications to incrementally transform seed designs into more complex, scalable designs; and 2) quantitative metrics of HLS design complexity and design-space scalability that serve as measurable objectives for LLM-guided mutation. We systematically cross-validate an HLS Quality-of-Results (QoR) deep learning model trained and tested across common HLS benchmarks, zero-shot synthetic designs, and iteratively mutated synthetic designs. Synthetic designs transfer well to common benchmark test sets while the reverse does not hold. SyntheticHLS’s iteratively mutated designs provide the most generalizable training corpus among the datasets studied. Analysis of the mutation process and dataset shows that metric-guided trajectories consistently improve targeted complexity and scalability objectives without regressing non-target metrics. Mutated designs span a substantially broader, more diverse design space than zero-shot generated designs. Our framework, dataset, and evaluation are open-source: this https URL. Comments: Accepted and to be presented at the International Conference on Field Programmable Technology (FPT) 2026 Subjects: Hardware Architecture (cs.AR); Machine Learning (cs.LG) Cite as: arXiv:2610.00106 [cs.AR] (or arXiv:2610.00106v1 [cs.AR] for this version) https://doi.org/10.48550/arXiv.2610.00106 Focus to learn more arXiv-issued DOI via DataCite (pending registration) Submission history From: Stefan Abi-Karam [view email] [v1] Wed, 9 Sep 2026 02:40:29 UTC (3,330 KB)
[LG-196] Uncertainty-Aware Learning from Multi-Expert Interval Targets
链接: https://arxiv.org/abs/2610.00102
作者: Samira Alkaee Taleghan,Younghyun Koo,Andrew P. Barrett,Farnoush Banaei-Kashani
类目: Machine Learning (cs.LG)
*备注:
Abstract:Many machine learning (ML) applications rely on expert labels, and qualified experts may provide different but plausible interpretations of the same observation. Such variation across expert labels may reflect genuine disagreement or ambiguity rather than annotation error. When individual experts additionally report intervals rather than exact values, the supervision contains two distinct sources of label uncertainty: within-label imprecision and between-expert variation. Existing methods treat these forms separately: multi-expert approaches collapse labels to a consensus, interval-target methods often yield a single prediction, and predictive-uncertainty methods rarely validate their uncertainty estimates against observed expert disagreement. To address this problem, we propose an approach that preserves individual expert intervals, separates within-label imprecision from between-expert variation, and validates the corresponding predictive uncertainty components. First, heterogeneous label vocabularies are harmonized into a common probabilistic label space, separating encoding differences from expert judgement. Second, individual label intervals are retained and modeled with a mixture of Beta distributions trained using a proper Cramér-distance objective, preserving distinct expert-reported labels. Third, we decompose predictive uncertainty into within-component, between-component, and model uncertainty, and evaluate whether these components correspond to within-label uncertainty, between-label uncertainty, and model error, respectively. Because this correspondence is not guaranteed, we introduce decomposition matching, which aligns the predictive components to their intended label-side sources. On sea-ice concentration the model reduces MAE by 31% over hard labels and outperforms aggregation, interval-distribution, and interval-regression baselines.
[LG-197] One Mastery Threshold Does Not Fit All Knowledge Tracing Models
链接: https://arxiv.org/abs/2610.00095
作者: Xianghui Meng,Yujing Zhang,Jionghao Lin
类目: Machine Learning (cs.LG)
*备注:
Abstract:Tutoring systems use mastery thresholds to decide when students can stop practicing and advance, but the same numerical threshold can lead to very different decisions when the underlying knowledge tracing (KT) model changes. We examine six KT models across four public educational datasets and evaluate 12 thresholds from 0.50 to 0.99 using post-advancement performance, advancement coverage, practice burden, and disparities across prior-performance groups. We also identify thresholds that balance performance, extra practice, and advancement under 30 predefined instructional settings. Bayesian Knowledge Tracing (BKT) is relatively insensitive to threshold changes, while neural models become much more selective as thresholds increase. This partly reflects different model outputs: BKT estimates latent mastery probability, whereas neural models estimate the probability of a correct next response, so the same cutoff does not represent the same level of mastery. The best-balanced threshold varied substantially across models and settings. In half of the tested settings, neural models and BKT differed by more than 0.10 in their selected thresholds, although this gap became smaller when greater priority was placed on reducing extra practice and allowing more students to advance. Stricter thresholds also did not reliably reduce performance gaps and could disproportionately restrict advancement, with stronger-prior students advancing up to 3.26 times as often as weaker-prior students. These results show that mastery thresholds should be recalibrated when the KT model or instructional priorities change and evaluated by their effects on performance, practice, advancement, and access.
[LG-198] “very likely” Means “uncertain”? How LLM s Diverge from Humans in Linguistic Uncertainty Quantification ICML2026
链接: https://arxiv.org/abs/2610.00083
作者: Jinhao Duan,Zicheng Liu,Zijie Liu,Kaidi Xu,Tianlong Chen
类目: Machine Learning (cs.LG)
*备注: ICML 2026
Abstract:Humans express uncertainty verbally via markers (e.g., “possible,” “likely”), yet most LLM uncertainty quantification (UQ) relies on costing likelihood- or consistency-based signals. From a cognitive perspective, accurate verbal uncertainty reflects metacognitive monitoring, representing knowledge boundaries (“knowing that you don’t know”) to support regulation and information seeking. In this paper, we investigate how LLMs diverge from humans in verbal uncertainty quantification and whether verbal markers can reliably quantify LLM uncertainty. We curate a corpus of human uncertainty markers from psychology and decision-science literature and benchmark LLMs against it. We observe that LLMs encode verbal uncertainty with numerical levels that differ substantially from those of humans. We then introduce METHODNAME, a novel optimization-based algorithm that learns an optimal uncertainty profile over uncertainty markers directly from LLM outputs. By fitting a marker-uncertainty mapping to best explain empirical correctness, METHODNAME discovers how much probability mass each verbal marker should convey, rather than estimating uncertainty via repeated sampling. METHODNAME enables a direct, marker-level comparison of confidence semantics between humans and LLMs, disentangling mismatch and revealing systematic confidence disparities in verbal expressions.
[LG-199] Format-Aware Fusion for Fast FP4 Pretraining
链接: https://arxiv.org/abs/2610.00053
作者: Robert Hu
类目: Machine Learning (cs.LG)
*备注:
Abstract:Four-bit floating-point (FP4) Tensor Cores accelerate matrix multiplication, but scale computation, operand packing, layout construction, and saved backward state can erase the gain. We present \emphformat-aware fusion, which co-designs each quantization producer with its scale domain and consumer layout for native \mxfp, global \nvfp, and cooperative-thread-array-local \nvfp. We evaluate Llama-3-family 8B pretraining through 160 billion tokens using bfloat16 output projections and compiled cross entropy. In matched same-accelerator probes, bfloat16 and Transformer Engine \nvfp reach 18.8K and 27.6K tokens/s/GPU, while our fastest custom route reaches 37.9K. \mxfp with row-gradient stochastic rounding and fixed-sign 32-value Hadamard weight-gradient preconditioning reaches 37.2K tokens/s/GPU (86.3% bfloat16 model FLOP utilization) and ends 2.11% above the raw bfloat16 training-loss endpoint. A Transformer Engine recipe with four final bfloat16 blocks ends 0.87% above bfloat16 at 27.1K tokens/s/GPU. Downstream rankings differ from training-loss rankings, showing that FP4 outcomes depend jointly on scale contract, operand, and execution path.
[LG-200] Fast Polynomial Transcendentals for LLM s
链接: https://arxiv.org/abs/2610.00049
作者: Robert Hu
类目: Machine Learning (cs.LG)
*备注:
Abstract:Graphics processing unit (GPU) generations scale matrix, special-function, and memory pipelines at different rates, so kernel bottlenecks move as hardware evolves. FlashAttention-4 exposed this imbalance inside attention on NVIDIA Blackwell. We test whether short polynomial programs can accelerate other special-function-unit (SFU) operations in large language models (LLMs). We first compare native PyTorch evaluation with packed fused multiply–add (FMA) programs in an isolated IEEE binary16 (FP16) sweep spanning L2-resident and high-bandwidth-memory (HBM)-resident working sets. We then replace native sigmoid, tanh, and sigmoid linear unit (SiLU) with degree-3 or degree-4 bfloat16 (BF16) programs in four GB200 integration tasks: dense SiLU, tanh-softcapped attention, sigmoid attention, and routed-expert Swish-gated linear unit (SwiGLU). The programs combine analytical symmetry, target-format rounding, and packed arithmetic inside consuming kernels. The isolated paths improve by 1.19–2.19x in L2 and 1.00–1.70x in HBM. The dense-SiLU, tanh-softcapped-attention, and routed-expert substitutions improve complete training-step throughput by 2.7%, 2.9%, and 8.0%, respectively. The sigmoid-attention substitution improves complete-attention forward by 7.4% and the complete GPU step by 0.3%. Same-checkpoint open-weight ablations and one paired pre-training comparison per task extend the evaluation to model behavior. At common horizons near 100 billion tokens, the final smoothed training-loss differences (polynomial minus native) range from -0.107 to +0.079 across the four tasks.
[LG-201] How Far is Adam from Natural Gradient Descent?
链接: https://arxiv.org/abs/2610.00004
作者: Vihaan Paka-Hegde
类目: Machine Learning (cs.LG); Neural and Evolutionary Computing (cs.NE)
*备注: 9 pages, 4 figures, 2 tables
Abstract:Adam is the standard optimizer in deep learning, yet its geometric relationship to natural gradient descent (NGD) contains unresolved questions. We study Adam’s full update rule, including momentum, as a diagonal empirical Fisher approximation subject to diagonal truncation, empirical label substitution, and temporal lag. Using the scale-invariant \gamma(\Delta\theta) metric, we measure Adam’s geometric deviation from true NGD across four loss landscapes: well-conditioned linear regression, ill-conditioned linear regression, logistic regression, and a non-convex small neural network. Adam’s geometric trajectory is context-dependent. Deviation remains low in well-conditioned settings but rises significantly under ill-conditioning, reaching misalignments of \approx 10^3 in the neural network. Higher geometric drift correlates with slower initial optimization but does not degrade final objective minimization; Adam consistently reaches low loss. Furthermore, the improved empirical Fisher (iEF) tracks more stable paths than the standard empirical Fisher (EF), which frequently oscillates or diverges. Our results suggest Adam’s practical optimization power may stem from a balance of structural approximation errors and momentum smoothing rather than close tracking of the natural gradient path.
[LG-202] Reverse Item Response Theory for Sparsity-Robust Ranking in Frag mented Cancer Drug-Response Matrices
链接: https://arxiv.org/abs/2610.00002
作者: Jung Min Kang
类目: Machine Learning (cs.LG)
*备注: 8 pages, 4 figures, 3 tables
Abstract:We introduce reverse Item Response Theory (IRT) to pharmacogenomic drug-response analysis by treating cancer types as latent “subjects” with resistance ability and drugs as “items” with evasion difficulty. Applied to 242,036 drug sensitivity measurements from the Genomics of Drug Sensitivity in Cancer (GDSC2) database, the model estimates cancer-type-level in-vitro resistance and drug-level broad activity on a shared latent scale. Validation across four missingness regimes demonstrates that reverse IRT better recovers the full-data latent ranking than simple averaging, with advantages of Delta-rho = +0.089 to +0.095 at 60% missingness under MCAR, cancer-biased, and drug-biased sparsity. Held-out prediction confirms IRT achieves the best Brier score among five evaluated methods. Bootstrap confidence intervals show 19 of 28 cancer types have stable resistant/sensitive classifications. Cross-platform PRISM replication shows 82% directional agreement but weak rank-order correlation (rho = 0.25), indicating the contribution is methodological robustness under fragmented evaluation, not a universal clinical resistance leaderboard.
[LG-203] Efficient Support Recovery of Mixtures of Sparse Linear Classifiers with Less Measurements
链接: https://arxiv.org/abs/2609.32176
作者: Xiaxin Li,Arya Mazumdar
类目: Machine Learning (cs.LG); Data Structures and Algorithms (cs.DS); Information Theory (cs.IT)
*备注:
Abstract:The support recovery problem in mixture of linear classifiers intends to identify which features actually matter when data is generated by a mixture of several linear decision rules. In particular, the aim is to recover the support (nonzero coordinates) of l unknown k -sparse vectors from sign measurements. Each measurement is generated by selecting one of the l vectors uniformly at random, and returning the sign of its inner product with a chosen measurement vector. In this paper, we propose adaptive and non-adaptive schemes that significantly improve upon prior results by reducing the number of measurements and achieving sublinear decoding time simultaneously. In particular, our adaptive constructions substantially reduce measurements compared to existing approaches, while also lowering decoding complexity from super-quadratic to sublinear in the ambient dimension. We further provide a non-adaptive scheme that improves previous measurement bounds while maintaining efficient decoding. Overall, our approach yields a more efficient trade-off between sample complexity and decoding time for support recovery in mixture models than previously known methods. Subjects: Machine Learning (cs.LG); Data Structures and Algorithms (cs.DS); Information Theory (cs.IT) Cite as: arXiv:2609.32176 [cs.LG] (or arXiv:2609.32176v1 [cs.LG] for this version) https://doi.org/10.48550/arXiv.2609.32176 Focus to learn more arXiv-issued DOI via DataCite
[LG-204] Sample complexity bounds for categorical Markov random fields via Discrete Diffusions
链接: https://arxiv.org/abs/2610.02128
作者: Shivam Kumar,Nabarun Deb
类目: atistics Theory (math.ST); Machine Learning (cs.LG); Machine Learning (stat.ML)
*备注: 83 Pages, 3 Figures, 4 Tables
Abstract:Many applications in statistics, economics, and physics require sampling from high-dimensional categorical distributions with local dependence structures. Examples include finite memory language models, Ising and Potts systems in statistical physics and protein folding, etc. In modern machine learning, discrete diffusions have emerged as a flexible approach for sampling such data, with strong empirical performance. Motivated by this, we develop learning methods with end-to-end sample complexity bounds for discrete diffusion with uniform noising under local dependence, which we model through low order Markov random fields (MRFs). Our main technical insight is a new \emphpinning decomposition of the discrete score. It shows that unlike in continuous diffusions, the score decomposes into components where the dependence on time separates multiplicatively from the dependence on the target. Building on this decomposition, we propose a \emphweight-sharing neural score learner and combine it with \tau -leaping to obtain an end-to-end sampling procedure. Rather than treating score-learning error as a black-box input, as is common in existing sampling analyses, we study the score learning error from finite data and derive optimal sampling guarantees with explicit dependence on the vocabulary size, the interaction order of the MRF, and the sample size. Moreover, our strategy trains a single score network across uniform noise levels while leaving the sampling discretization to be chosen at inference-time. This allows the same trained model to trade accuracy for computational cost as inference-time budgets vary. Numerical experiments on Potts, Ising, and tree-structured models show that weight-sharing score networks outperform fully connected ones for sampling long sequences.
[LG-205] Wasserstein Gradient Flows and Forward-Only Diffusion Are Not Enough for Multimodal Sampling NEURIPS2026
链接: https://arxiv.org/abs/2610.02081
作者: Daniel McBride,Pratik Khandagale,Cristina Garcia-Cardona,Yen Ting Lin
类目: Machine Learning (stat.ML); Machine Learning (cs.LG); Mathematical Physics (math-ph); Dynamical Systems (math.DS); Probability (math.PR); Spectral Theory (math.SP)
*备注: 20 pages, 5 figures, accepted by NeurIPS 2026 Position Track
Abstract:There has been a proliferation of sampling algorithms based on Wasserstein gradient flows (WGF) and forward-only diffusion processes (FODP), often accompanied by theoretical guarantees of exponentially fast convergence to the target distribution. These guarantees are frequently interpreted as evidence that such methods can efficiently sample complex multimodal distributions, often supported by empirical results. In this work, we argue that this interpretation is fundamentally misleading. By invoking the Jordan-Kinderlehrer-Otto (JKO) scheme and Otto calculus, we establish that the canonical WGF sampling dynamics and overdamped forward diffusion share the same density evolution and therefore inherit the same metastability and slow-mixing phenomena long understood in nonequilibrium statistical physics. We analyze this family of samplers using two complementary tools – spectral analysis and mean first-passage time (MFPT) analysis – and show that well-separated multimodality can induce exponentially long mixing times associated with small spectral gaps and rare inter-mode transitions. For the commonly adopted log-linear annealing schedule studied here, we find that introducing intermediate distributions does not remove the exponential scaling of the total transport time. The limitation is structural rather than implementation-specific: purely local, gradient-driven transport mechanisms can require exponentially long times to transport probability mass across well-separated modes. We argue that this represents a fundamental limitation of WGF- and FODP-based sampling in their standard forms, and motivates future development of fundamentally nonlocal mechanisms for efficient multimodal sampling.
[LG-206] AI Emulation of Stochastic Sudden Stratospheric Warming with Interpretable Latent Structure
链接: https://arxiv.org/abs/2610.02069
作者: C. Daniel Boscu,Daniel Hernandez,Fabio Alvarez Ventura,Justin Finkel,Ashesh Chattopadhyay,Pedram Hassanzadeh,Dorian S. Abbot
类目: Atmospheric and Oceanic Physics (physics.ao-ph); Machine Learning (cs.LG)
*备注:
Abstract:Rare weather regime transitions pose a challenge for data-driven modeling due to class imbalance. In this study, we develop a probabilistic deep learning emulator for a prototypical system with regime transitions, the stochastic Holton–Mass model of stratospheric variability, and analyze the structure of its learned latent space. The Holton–Mass model exhibits two metastable regimes, a strong and a weak polar vortex, maintained by nonlinear wave–mean flow interactions, with weak stochastic forcing intermittently triggering rare transitions between these regimes that qualitatively represent SSW events. We employ a ResNet-inspired Conditional Variational Autoencoder with six-layer encoder and decoder layers and explicit current-state conditioning to model the distribution of the system’s state at the next time step (one day). The emulator accurately reproduces short-term dynamics, steady-state probability distributions, regime persistence statistics, rare transition rates, the transition committor function, and the transition expected lead time of the physical model. Beyond emulation fidelity, we interrogate the learned latent representation to understand how the model internalizes the underlying metastable structure of the dynamics. Principal Component Analysis of the 32-dimensional latent space reveals a clear and unsupervised separation into four physically interpretable clusters corresponding to strong versus weak vortex regimes and stable versus transition-prone configurations. Such emergent regime separation in latent space is hard to identify for deep generative models applied to high-dimensional stochastic systems. Our results show that carefully designed probabilistic emulators can uncover physically meaningful manifolds governing extreme-event dynamics, potentially aiding the development of improved operational advanced warning systems.
[LG-207] Sequential Capacity of Quantum Processes with Finite Memory
链接: https://arxiv.org/abs/2610.02068
作者: Yibin Wang
类目: Quantum Physics (quant-ph); Machine Learning (cs.LG)
*备注:
Abstract:How complex can the responses of a quantum device become as it runs longer with a fixed internal memory? We quantify this complexity through sequential response capacity: how many adaptive testing stages, each using a fresh run, can continue to separate possible processes by a prescribed gap in response probabilities. For fixed system and memory sizes, we establish a tight law relating this capacity to run length and probability resolution. At fixed resolution, the capacity grows on the order of K\log K , where K is the number of time steps in each run. Our construction attains this growth using time-dependent phase rotations on a single visible qubit with no additional internal memory; its tests give response probabilities exactly zero or one. Under the same tests, classical stochastic processes that measure in a fixed basis at every step have only linear capacity at fixed sizes and resolution. For phase sequences selected by a stored classical label, we then quantify how known independent Pauli noise changes this logarithmic enhancement. With ideal controls and weak residual phase noise after correction, we prove matching capacity bounds at a fixed small probability gap. These bounds identify the inverse residual phase-flip probability as the coherence timescale that limits the extra logarithmic growth.
[LG-208] Error-Corrected Inference-Time Scaling for Imperfect Diffusion Models
链接: https://arxiv.org/abs/2610.01933
作者: Zuokai Wen,Louis Grenioux,Weinan E,Jiequn Han
类目: Machine Learning (stat.ML); Machine Learning (cs.LG)
*备注: Under review
Abstract:Inference-time scaling adapts pretrained diffusion models to new sampling tasks without additional training. Existing methods rely primarily on Monte Carlo sampling with more particles, yet are premised on the pretrained model being exact. In practice, data and training limitations make the model imperfect, and these methods inherit its error. More particles reduce Monte Carlo error but cannot remove the mismatch between the endpoint and the desired target or the error in tracking the prescribed probability path. We introduce the Energy-based Feynman-Kac Corrector (EBFKC), a framework for energy-based diffusion models that corrects these errors on the fly given a reference energy. We first derive Feynman-Kac dynamics that track a prescribed path exactly in the continuous-time population limit even when the model is imperfect, and approximate these dynamics using sequential Monte Carlo with variance-controlling guidance. To remove the endpoint mismatch, we use the pretrained energy as a surrogate along the diffusion path and progressively incorporate the discrepancy between the learned and target terminal energies. Experiments on Gaussian mixture models, particle systems, alanine dipeptide, and alanine tetrapeptide show that our method closely matches target distributions and molecular free-energy profiles under annealing and reward tilting, whereas standard inference-time scaling baselines retain substantial sampling errors.
[LG-209] Optimal Stochastic Bilevel Optimization with First-Order Oracles
链接: https://arxiv.org/abs/2610.01843
作者: Linxuan Pan,Junchi Yang
类目: Optimization and Control (math.OC); Machine Learning (cs.LG)
*备注:
Abstract:We study nonconvex–strongly-convex bilevel optimization under a stochastic first-order oracle. We introduce MRT-FD, a single-loop first-order method that simultaneously tracks the upper-level variable, the lower-level solution, and the auxiliary response arising from implicit differentiation of the hyperobjective. MRT-FD performs one update of each variable per iteration and approximates the second-order derivative actions using order- p finite differences. For any fixed finite smoothness order p\ge1 in the lower-level variable, MRT-FD finds an \varepsilon -stationary point using \mathcalO(\varepsilon^-4-2/p) stochastic gradient queries. We also prove a matching \Omega(\varepsilon^-4-2/p) oracle lower bound. The lower-bound construction starts from a hard nonconvex minimization chain with a stronger stochastic oracle, and lifts it to a bilevel problem through a sinusoidal coupling with a scalar lower-level variable. Consequently, the dependence on \varepsilon is optimal for every fixed finite p , closing the upper–lower complexity gap in this stochastic first-order oracle setting.
[LG-210] Varda-single-1.0: deterministic data-driven weather forecasting at 1 km resolution over Switzerlands complex topography
链接: https://arxiv.org/abs/2610.01835
作者: Alberto Pennino,Francesco Zanetta,Michele Cattaneo,Claire Merker,Radi Radev,Jonas Bhend,Louis Frey,Hugues de Laroussilhe,Ophélia Miralles,Carlos Osuna,Daniele Nerini,Andreas Pauling,Daniel Hupp,Ulrich Hamann,Mary McGlohon,Marti Bosch,Luca Lanzilao,Marco Arpagaus,Lukas Jansing,Daniel Leuenberger,Mark A. Liniger,Katrin Ehlert,Matthew Chantry,Håvard Homleid Haugen,Gert Mertes,Ana Prieto Nemesio,Mario Santa Cruz,Jasper Wijnands,Gabriel Moldovan,Harrison Cook,Oliver Fuhrer
类目: Atmospheric and Oceanic Physics (physics.ao-ph); Machine Learning (cs.LG)
*备注: 24 pages, 13 figures, 2 tables. Model weights: this https URL
Abstract:We present Varda-single-1.0, a medium-range data-driven weather prediction system built for the Alpine domain. It provides hourly deterministic regional forecasts on a mesh of 1 km resolution and global forecasts on a 31 km mesh. The system comprises two independently trained stretched-grid Graph Transformer models with encoder-processor-decoder architecture, developed in the Anemoi framework: a 6-hourly autoregressive forecaster and a temporal downscaler reconstructing hourly forecasts between the forecaster’s steps. Its training curriculum includes pre-training on ERA5 reanalysis data, followed by training on a 20-year kilometre-scale regional reanalysis, and finally fine-tuning on operational kilometre-scale analyses. Verified over one year against operational analyses and surface station observations, Varda-single is competitive with or improves on MeteoSwiss’ operational numerical weather prediction baselines for most headline scores and variables. It broadly matches the skill of the high-resolution 1 km ICON-CH1-EPS control at lead times up to +33 h and generally outperforms the 2 km ICON-CH2-EPS control at lead times up to +120 h. Despite competitive aggregate scores, Varda-single underestimates some local wind maxima and produces overly smooth convective precipitation fields, consistent with the smoothing associated with squared-error training. To gain insight into the model’s behaviour, we investigate three case studies beyond the aggregated headline scores, and find particular weaknesses in Varda-single’s representation of local winds over complex terrain. Varda-single represents an important step in the development of high-resolution ML forecasting over complex terrain, in complementing the operational regional numerical weather prediction models of MeteoSwiss with data-driven models and in providing a pretrained model for researchers and user-specific applications.
[LG-211] Generalized Engression Models
链接: https://arxiv.org/abs/2610.01823
作者: Xinwei Shen,Zijian Guo,Francis Bach
类目: Methodology (stat.ME); Machine Learning (cs.LG); Machine Learning (stat.ML)
*备注:
Abstract:We consider estimating the conditional distribution of a multivariate outcome given covariates when its coordinates may be continuous, binary, categorical, ordinal or rankings, and are conditionally dependent on one another. Different statistical methods have been developed for each outcome type, and most of them target a summary of the conditional distribution, such as the mean of each coordinate, rather than the joint distribution of the outcome vector. We develop generalized engression models, a unified nonparametric distributional regression framework for outcomes of any type. The proposed method builds upon engression, a scoring-rule-based deep generative model, and introduces a data-type-specific link function and a stochastic perturbation that smooths the loss, enabling gradient-based training even with discontinuous links. We establish universal representation results for continuous, discrete and mixed outcomes. In simulations and in two applications, 242 species in a community ecology benchmark and a 17-dimensional mixed-type health outcome, the method matches type-specific models on marginal scores, improves on them on the joint distribution, and matches or exceeds purpose-built state-of-the-art joint species distribution models. Software is available in Python.
[LG-212] Learnt Attacks on Quantum Key Distribution under Channel Noise and Device Drift NEURIPS2026
链接: https://arxiv.org/abs/2610.01792
作者: Marcel Mordarski,Benjamin Gras,Abdelrahman Shehata,Daniel Budina,Roberto Bondesan
类目: Quantum Physics (quant-ph); Cryptography and Security (cs.CR); Information Theory (cs.IT); Machine Learning (cs.LG)
*备注: Presented as submission 202 at QCrypt 2026 this http URL . A parallel work exploring the machine-learning aspects of this approach, titled "Sparsity for Free: A Budget-Induced Equilibrium in Joint Topology-Parameter Search’', has been accepted for NeurIPS 2026
Abstract:Quantum key distribution (QKD) links are provisioned from security analyses of stationary channels, whereas the devices that determine the channel drift between recalibrations. Whether an eavesdropper who cannot alter the channel’s own noise gains by following that drift has not been quantified. Adaptive eavesdropping is posed here as a constrained Markov decision process in which the attacker selects one circuit per round while the noise level follows an Ornstein–Uhlenbeck process and the abort condition is a budget over each block of rounds. The value of adaptation is bounded by the best fixed circuit and a dynamic-programming upper bound. The actions are learnt attacks. Whereas Decker et al. trained a parametrised circuit on a fixed gate template against a fixed channel, here the gate structure and rotation angles are searched jointly. This yields circuits compact enough to form a discrete action set, extending the construction to noise models lacking a known template, including the amplitude damping channel. On device-independent E91 under bilateral depolarising noise, a reinforcement-learning attacker raises her Holevo information from 0.135 for the best fixed circuit to 0.348 at zero detection, 98% of the upper bound. On BB84 under a drifting bit-flip channel, she exceeds a conservative noise-indexed rule by 0.024 in fidelity, reaching 99% of the upper bound. Under stationary noise, the attacker’s gain from basis asymmetry changes sign between an averaged and a per-basis error-rate constraint. The search, started from random gate sequences, recovers the analytical cloners and the collective-attack key rate, and meets the lower bound of the Winick–Lütkenhaus–Coles objective from above.
[LG-213] Lower Bounds for Stochastic First-Order Algorithms with Variance Reduction in Nonconvex–Concave Minimax Optimization
链接: https://arxiv.org/abs/2610.01662
作者: Jiayi Song,Zi Xu
类目: Optimization and Control (math.OC); Machine Learning (cs.LG); Machine Learning (stat.ML)
*备注:
Abstract:We establish complexity lower bounds for stochastic first-order algorithms in nonconvex–concave minimax optimization, allowing algorithms to use variance reduction. Our main contribution is a lower bound for a zero-respecting algorithm class that permits variance reduction, extending beyond the algorithmic restrictions imposed by some existing lower bounds. We consider objectives with an L -Lipschitz continuous joint gradient, a compact convex dual domain of Euclidean radius at most D_Y , and a primal value function, defined by maximizing the objective over the dual variable, with initial suboptimality at most \Delta . The target accuracy \varepsilon is measured by the gradient norm of the Moreau envelope of the constrained primal value function with parameter 1/(2L) . Under an unbiased stochastic first-order oracle with variance at most \sigma^2 and mean-square smoothness, we prove the lower bound \Omega!\left(L^2D_Y\Delta\varepsilon^-3+L^3D_Y^2\Delta\sigma^2\varepsilon^-6\right) . This result quantifies the dependence on accuracy, dual-domain radius, and oracle noise even when variance reduction is allowed. We also establish complementary lower bounds for nonconvex–strongly-concave minimax optimization. With dual strong-concavity parameter \mu0 and condition number \kappa:=L/\mu , we obtain \Omega!\left(L\Delta\sqrt\kappa,\varepsilon^-2+L\Delta\kappa\sigma^2\varepsilon^-4\right) under the bounded-variance oracle model. Under the additional mean-square smoothness condition with constant \bar L , we obtain \Omega!\left(L\Delta\sqrt\kappa,\varepsilon^-2+\Delta\bar L\sigma\kappa^3/2\varepsilon^-3\right) . Together, these results identify complexity barriers across the concave and strongly concave regimes, with the main nonconvex–concave bound remaining valid for algorithms that use variance reduction.
[LG-214] Convergence Analysis of STORM Under Different Geometries
链接: https://arxiv.org/abs/2610.01599
作者: Wei Jiang,Yibo Wang,Wenhao Yang,Rui Yan,Lijun Zhang,Zechao Li
类目: Optimization and Control (math.OC); Machine Learning (cs.LG)
*备注:
Abstract:Stochastic recursive momentum (STORM) achieves fast convergence for nonconvex optimization via the variance reduction effect, but existing analyses rely on the strong average smoothness assumption. In this paper, we study the convergence of STORM for different objectives without average smoothness. We first revisit the results under average smoothness, obtaining the O(T^-1/3) bound for nonconvex objectives and the O(\sigma^2/(\mu T)) bound for last-iterate output under the \mu -Polyak–Łojasiewicz~(PL) condition. Without average smoothness, we design an auxiliary sequence and compare the STORM update with it in the analysis. With the help of this sequence, we prove that STORM still attains an O(T^-1/4) rate for nonconvex objectives, which is optimal under standard smoothness. For convex and \lambda -strongly convex objectives, we further prove averaged and last-iterate bounds with optimal rates of O(\sigma R/\sqrt T) and O(\sigma^2/(\lambda T)) , respectively. All the obtained results use the same STORM recursion with different hyperparameter choices.
[LG-215] he hidden advantage of mask resampling: a theory of masked autoencoders
链接: https://arxiv.org/abs/2610.01578
作者: Jorge Medina Moreira,Lorenzo Bardone,Lenka Zdeborová
类目: Machine Learning (stat.ML); Disordered Systems and Neural Networks (cond-mat.dis-nn); Machine Learning (cs.LG)
*备注:
Abstract:Why can masked prediction learn useful representations that unmasked reconstruction misses? We study this question in a high-dimensional model of a masked autoencoder (MAE) trained on data with shared latent structure and heterogeneous noise. We prove that masked linear reconstruction can recover the latent feature at linear sample complexity in regimes where unmasked linear reconstruction, equivalent to PCA, fails. The analysis also quantifies the statistical advantage of mask resampling, an established ingredient of masked pretraining. By introducing a fixed collection of K masks per sample, we characterize its effect on feature recovery and downstream performance, identifying regimes where greater mask diversity lowers sample complexity. Guided by this prediction, we find that random cropping and flipping in standard image-training pipelines can obscure the advantage of mask resampling by renewing the prediction task even when the patch mask is fixed. Removing these transformations reveals a downstream advantage for dynamic over static masking in CNN autoencoders and vision transformers. A complementary BERT pilot finds benefits from greater mask diversity on downstream language tasks. Our results separate the benefit of the masked prediction objective from that of mask diversity, and show how a tractable theory can guide experiments that uncover advantages hidden by standard training practices.
[LG-216] Optimal Momentum Methods for Stochastic Multilevel Compositional Optimization
链接: https://arxiv.org/abs/2610.01572
作者: Wei Jiang,Rui Yan,Sifan Yang,Yuanyu Wan,Lijun Zhang,Zechao Li
类目: Optimization and Control (math.OC); Machine Learning (cs.LG)
*备注:
Abstract:This paper investigates stochastic multi-level optimization where the objective is a nested composition of several smooth non-convex functions. We assume that only stochastic estimates of the gradient and function values for each level are accessible. Consequently, obtaining an accurate estimate of the overall gradient is challenging due to the nested structure. To address this, we employ a momentum-based estimator with mini-batches to track the function values of each level, which are subsequently used to construct momentum gradient estimators. We establish an optimal sample complexity of \mathcalO(\epsilon^-4) for finding an \epsilon -stationary point, avoiding the stronger average smoothness assumption commonly relied upon in prior literature. Furthermore, by employing a normalization technique, we attain the same rate without requiring problem-dependent constants to set hyperparameters. To achieve the optimal rate without mini-batches, we further develop a batch-free method that incorporates a first-order approximation and a clipping technique for function value estimation. Finally, we validate the effectiveness of our proposed methods through experiments on risk-averse portfolio optimization and hierarchical tilted empirical risk minimization.
[LG-217] Zero Flux: Flow-Based Comparison of High-Dimensional Discrete Distributions
链接: https://arxiv.org/abs/2610.01472
作者: Leyang Wang,Yakun Wang,Song Liu,Taiji Suzuki
类目: Machine Learning (stat.ML); Machine Learning (cs.LG)
*备注:
Abstract:Comparing two high-dimensional discrete distributions has always been a challenging task due to the exponentially growing state space and complex changes in interactions. A recent work suggests comparing distributions through a vector field trained using flow matching between two continuous distributions. The resulting vector field at mid-point vanishes if and only if two distributions identical. However, such a flow-based criterion does not naturally apply to discrete distributions. We extend this principle to the discrete domain and introduce the \emphZero Flux criterion, a discrepancy based on local probability fluxes. Under independent coupling, we show that all local probability fluxes vanish at the midpoint if and only if two distributions are the same. This discrepancy decomposes the joint distributional difference into smaller, local contributions and can be efficiently estimated from samples. We establish finite sample error bounds for our estimator. Experiments on synthetic and real categorical data demonstrate reliable recovery of sparse dependence signals and stable tracking of distribution shifts in high dimensions.
[LG-218] Learning ab initio phase-field models
链接: https://arxiv.org/abs/2610.01432
作者: Mengyi Chen,Peichen Zhong,Zihan Zhang,Qianxiao Li
类目: atistical Mechanics (cond-mat.stat-mech); Materials Science (cond-mat.mtrl-sci); Machine Learning (cs.LG)
*备注:
Abstract:Simulating microstructure evolution requires quantum-mechanical accuracy and mesoscopic reach in length and time scales, a combination that no current method achieves. Classical phase-field models provide this reach, but their accuracy is limited by phenomenological free energies and mobilities. Here we develop a framework for learning ab initio phase-field models, where the mesoscopic equation is not postulated but derived from a Mori-Zwanzig projection of molecular dynamics onto species-density fields under explicit assumptions. The nonlocal free energy and mobility left unspecified by this equation are parametrized by neural networks and learned from short molecular dynamics trajectories generated with machine-learning interatomic potentials of ab initio accuracy. We demonstrate the framework on an iron-boron melt and on hydrogen-helium mixtures under planetary conditions. For iron-boron, the model shows that the melt at the FeB _4 composition is spinodally unstable at ambient pressure but stabilized at 10 GPa, offering a thermodynamic rationale for why FeB _4 has been synthesized only under high pressure. For hydrogen-helium, the model predicts the immiscibility boundary and captures droplet nucleation and growth in helium-rain simulations of a column corresponding to 2.2 million atoms, far beyond the scale of atomistic modeling at comparable accuracy. Trained across compositions and conditions, such models could provide a mesoscopic counterpart to ab initio molecular dynamics.
[LG-219] SupraTITO: Transferable Generative Molecular Dynamics for Supramolecular Systems
链接: https://arxiv.org/abs/2610.01381
作者: Weilong Chen,Nuno Costa,Julija Zavadlav
类目: Chemical Physics (physics.chem-ph); Soft Condensed Matter (cond-mat.soft); Machine Learning (cs.LG); Biological Physics (physics.bio-ph)
*备注:
Abstract:Peptide sequence governs both the structures formed through supramolecular assembly and the dynamics by which they emerge, but predicting either requires resolving slow collective processes among many interacting molecules. Molecular dynamics (MD) provides microscopic insight into these processes, yet the long timescales of assembly and the vast peptide sequence space make systematic exploration computationally demanding. We introduce SupraTITO, a transferable generative molecular dynamics (GenMD) framework for supramolecular systems, demonstrated through peptide self-assembly. SupraTITO learns transferable implicit transfer operators (TITO) conditioned on peptide sequence, molecular topology, and periodic geometry, allowing configurations to be propagated over physical intervals much longer than an MD integration step. On a comprehensive dipeptide benchmark, SupraTITO generalizes to held-out sequences and reproduces sequence-dependent structures and dynamics while maintaining molecular integrity over long rollouts. Compared with direct ensemble prediction trained on the same trajectory data, SupraTITO more accurately reproduces assembly structures while also resolving their temporal evolution. The learned dynamics generalize across peptide concentrations, including dilute conditions not represented during training. These results extend transferable GenMD to collective dynamics in periodic supramolecular systems and provide a foundation for modeling related processes beyond peptide assembly.
[LG-220] Reachability-Informed Reinforcement Learning for Multi-Impulse Interplanetary Transfers
链接: https://arxiv.org/abs/2610.01344
作者: Yashdeep Chaudhary,Roberto Armellin,Harry Holt
类目: Optimization and Control (math.OC); Machine Learning (cs.LG); Systems and Control (eess.SY)
*备注: Preprint. 23 pages, 10 figures
Abstract:Reinforcement learning offers the prospect of a reusable sequential decision-making mechanism for spacecraft trajectory design, motivating policy interfaces that connect learned decisions to the underlying maneuver geometry. This paper develops Reachability Analysis-Informed Reinforcement Learning (RARL) for deterministic multi-impulse interplanetary transfers, placing intermediate waypoint selection at the center of the learned decision process. Local first-order reachability maps bounded velocity perturbations into an ellipsoidal set of next-node positions, within which the policy selects its waypoint. Lambert reconstruction then determines the corresponding maneuver to reach this selected waypoint along a dynamically consistent ballistic arc, coupling learned transfer-geometry selection with classical astrodynamics. A terminal two-impulse reconstruction completes the rendezvous, supported by a linear maneuver-demand assessment used for reward shaping. Numerical studies characterize this interface on a two-body Earth-Mars benchmark. Across three independent training runs, RARL achieves a mean maneuver cost of 10.23 km/s, 1.72% above a validated local sequential convex programming reference. Training over dispersed initial states extends policy reuse across a departure family with fixed target state and transfer duration. Each of the three independently trained multi-state policies completes all 10,000 held-out Monte Carlo departures without impulse-cap violations, compared with a mean feasibility rate of 6.49% for single-state policies. This broader sampled feasibility is accompanied by a 0.61% increase in mean nominal maneuver cost, without further training across departures. These results demonstrate that a reachability-informed decision interface supports benchmark-quality trajectory construction and policy reuse across dispersed departure conditions.
[LG-221] Classical Hardness of Learning Functions of Hamiltonians
链接: https://arxiv.org/abs/2610.01141
作者: Sota Hashimoto,Akinori Kawachi
类目: Quantum Physics (quant-ph); Machine Learning (cs.LG)
*备注: 11 pages, 1 figure
Abstract:Morohoshi, Nakayama, Manabe, and Mitarai proposed a physically motivated quantum machine learning problem in which the goal is to predict quantities of the form \operatornameTr[f(H)\rho] from classical descriptions of a Hamiltonian H and a quantum state \rho , where f is an unknown function. We call this problem Hamiltonian function learning in this paper. They constructed an efficient quantum learning algorithm under suitable conditions, while leaving a rigorous proof of average-case classical hardness open. In this paper, we rigorously prove the average-case classical hardness for two distribution-specific Hamiltonian function learning problems for f_\cos,\pi(\lambda)=\cos(\pi\lambda) and f_\exp,\beta(\lambda)=e^-\beta\lambda discussed in the paper of Morohoshi et al. under the assumption of the average-case hardness of factoring random RSA moduli. More specifically, we show that an efficient classical randomized learner under squared loss whose output hypotheses are evaluable in classical polynomial time for either problem would yield a classical randomized polynomial-time algorithm for factoring random RSA moduli.
[LG-222] Polylogarithmic Sparsity of Randomly Reweighted NPMLEs for Gaussian Mixtures
链接: https://arxiv.org/abs/2610.01088
作者: Hansheng Jiang
类目: Methodology (stat.ME); Machine Learning (cs.LG); Statistics Theory (math.ST)
*备注:
Abstract:The nonparametric maximum likelihood estimator (NPMLE) of a Gaussian location mixture maximizes the likelihood over the infinite-dimensional space of mixing distributions. The maximizing mixing distribution can be nonunique, and the classical bound on its number of atoms grows linearly with the sample size n . We show that a vanishingly small random perturbation of the likelihood yields exact polylogarithmic sparsity. The resulting randomly reweighted NPMLE maximizes a weighted likelihood whose independent weights, taken to be Gamma in our analysis, concentrate around one as n grows. With high probability, it is unique, has O(\log n/\log\log n)^d+\log n\ atoms in dimension d , nearly maximizes the ordinary likelihood, and estimates the mixture density at a Hellinger rate that is parametric up to logarithmic factors. This sparsity holds for the estimator itself, not for an approximation of it, and requires no support penalty. The proof rests on an effective-dimension principle for positive kernel mixtures: low-dimensional variation of the fitted values controls the support of every extreme point of the set of maximizers. Numerical illustrations verify that the reweighted NPMLE has Hellinger risk and support size comparable to those of the ordinary NPMLE.
[LG-223] Initial condition recovery in nonlinear damped viscous photoacoustic tomography using a convolutional neural network-guided gradient-free optimization framework
链接: https://arxiv.org/abs/2610.01015
作者: Madhu Gupta,Anwesa Dey,Prapti Tala,Souvik Roy
类目: Optimization and Control (math.OC); Machine Learning (cs.LG)
*备注:
Abstract:Photoacoustic tomography (PAT) is a hybrid imaging modality that combines high optical contrast with high ultrasonic resolution for biomedical imaging applications. In this work, we investigate the inverse problem of recovering the initial pressure distribution from boundary measurements in the presence of nonlinear acoustic propagation and viscous attenuation effects. To model these phenomena more accurately, we consider a nonlinear damped viscoelastic wave equation incorporating spatially varying sound speed, temporal attenuation, and nonlinear propagation mechanisms. We first establish the well-posedness of the corresponding forward problem using a Galerkin approximation combined with energy estimates and a fixed-point argument. For the inverse problem, we derive existence, uniqueness, and local uniqueness results under suitable assumptions through a harmonic extension reduction, spectral Laplace transform techniques, and observability estimates. To numerically reconstruct the initial pressure field, we develop a hybrid reconstruction framework that combines a convolutional neural network (CNN) with a gradient-free optimization strategy based on the sequential quadratic Hamiltonian (SQH) method derived from Pontryagin’s maximum principle. The CNN is used to generate an informative initial guess, while the SQH framework enforces the governing PDE dynamics during the reconstruction process. Numerical experiments demonstrate that the proposed hybrid strategy significantly improves reconstruction quality, contrast, and robustness compared to standalone time-reversal and CNN-based approaches.
[LG-224] olerance-Based Fairness Auditing: Violation Certification and Sensitivity Screening
链接: https://arxiv.org/abs/2610.01005
作者: Jie Tang,Chuanlong Xie,Lixing Zhu
类目: Machine Learning (stat.ML); Machine Learning (cs.LG); Methodology (stat.ME)
*备注: 47 pages, 8 figures
Abstract:As artificial intelligence is increasingly deployed, algorithmic unfairness has raised growing concerns and intensified demands for transparent fairness auditing. In practice, the tolerable degree of algorithmic unfairness depends on the specific legal, ethical, or application context. Given a prespecified tolerance threshold, an important statistical question is how to determine whether a group disparity exceeds the allowable tolerance across different auditing objectives. To address this problem, we develop a unified tolerance-based fairness auditing framework for two complementary auditing objectives: violation certification, which prioritizes control of false violation declarations, and sensitivity screening, which prioritizes reducing missed violations. For the first objective, we develop a constrained empirical likelihood test for formal settings that uses least-favorable-point calibration and can be combined with false flagging rate control for simultaneous subgroup auditing. For the second objective, we develop split empirical likelihood and adjusted split empirical likelihood tests using an adaptive boundary-proxy principle for early-warning settings. Numerical experiments show the distinct error-control–sensitivity trade-offs of these procedures. A COMPAS analysis illustrates the framework in predictive fairness auditing.
[LG-225] he Price of Correlated Tests: How Strict Should a Model Release Gate Be?
链接: https://arxiv.org/abs/2610.00993
作者: Marco Pollanen
类目: Machine Learning (stat.ML); Machine Learning (cs.LG); Methodology (stat.ME)
*备注: 6 pages, 1 figure, 4 tables. Submitted to ACDSA 2027
Abstract:Before a machine learning model ships, it often has to pass a suite of automated tests. Requiring every test to pass looks safe, yet it can reject many models that would have served users well, and it does not say how trustworthy a passing model actually is. We treat the release gate as a design problem: choose how many tests a model must pass so that cleared models meet a stated reliability target, while keeping as many good models as possible. A two-class latent-factor model makes both costs explicit and reduces each calculation to a one-dimensional integral. We prove that when both classes share the same latent correlation, a stricter gate always raises reliability, so the gate that keeps the most good models is the most lenient one that still meets the target. Under pass-all gating, any reliability target short of perfection is attainable within the model, but the share of good models kept tends to zero as the suite grows. Correlation between tests sets the price. In one configuration, a 99 percent target needs 8 independent tests, but 74 tests at a latent correlation of 0.3 and 5,182 at 0.5, where the gate keeps fewer than one good model in ten. We also give a validation procedure, built on exact binomial bounds, that certifies a gate from labelled data even when the gate is chosen from a fixed shortlist.
[LG-226] Block Optimism for Nonstationary Bandits with Latent Linear Dynamics NEURIPS2026
链接: https://arxiv.org/abs/2610.00911
作者: Taehyun Hwang,Hyunjun Choi,Heesang Ann,Min-hwan Oh
类目: Machine Learning (stat.ML); Machine Learning (cs.LG)
*备注: Accepted at NeurIPS 2026
Abstract:We study an endogenous nonstationary stochastic bandit problem with latent linear dynamics, where actions affect both immediate rewards and the future evolution of an unobserved latent state. Rewards are bilinear in the current action and latent state, inducing history-dependent rewards and a nontrivial long-horizon planning problem. The existing explore-then-commit approach achieves \tildeO(T^2/3) regret by uniformly exploring to estimate the latent dynamics and then committing to an optimized open-loop action sequence. We show that this rate can be improved via adaptive block-level optimism. Our key step is a cyclic approximation: under stable dynamics, the infinite-memory reward process can be truncated, and the open-loop benchmark can be approximated by optimizing a finite-memory block-level proxy. Building on this reduction, we propose a UCB-based block algorithm that maintains confidence sets for the truncated dynamics parameters and selects blocks optimistically. We prove a regret bound of order \tildeO(\sqrt T) , significantly improving over the previous \tildeO(T^2/3) guarantee for the same model. To the best of our knowledge, this is the first \tildeO(\sqrt T) regret guarantee for latent linear-dynamics bandits with bilinear reward observations and an open-loop action-sequence benchmark.
[LG-227] Neural Fourier Surrogates for Data Reuploading Quantum Neural Networks
链接: https://arxiv.org/abs/2610.00841
作者: Oliver Knitter,Jonathan Mei,Sang Hyub Kim,Chi Chen,Masako Yamada,Martin Roetteler
类目: Quantum Physics (quant-ph); Machine Learning (cs.LG)
*备注: 10 pages, 5 figures, 2 tables
Abstract:For quantum machine learning, the exact boundary between classical and quantum advantage is still poorly understood. Direct comparison between quantum neural networks (QNNs) and existing classical models, which encompass fundamentally different function classes, often fails to provide broader insight into the difference between the two. Inspired by the techniques of Neural Quantum States and Random Fourier Features, this work introduces Neural Fourier Surrogates (NFS), a stochastic classical neural network architecture for efficiently learning coefficients over the same finite Fourier series support as quantum neural networks. Testing on a selection of tabular benchmark datasets, we find that NFS is an effective classifier architecture broadly competitive with established classical baselines, including a comparable Random Fourier Features model, and possessing comparable performance to data-reuploading QNNs; combined with additional analysis comparing the learned Fourier spectra of QNNs and NFS on synthetic data, these results establish NFS as a natural classical baseline for evaluating QNN performance.
[LG-228] PI-AMFM: Permutation-Invariant Learning for Variable-Cardinality AM-FM Mode Decomposition in Biomedical Signal Analysis
链接: https://arxiv.org/abs/2610.00819
作者: Youngsun Kong,Ki H. Chon
类目: ignal Processing (eess.SP); Machine Learning (cs.LG)
*备注: 5 pages, 3 figures
Abstract:Physiological recordings often contain nonstationary oscillatory components whose number and dynamics vary across signals. Amplitude- and frequency-modulated (AM-FM) representations are well suited to characterizing such dynamics and have shown broad utility in biomedical signal analysis. Recent approaches have incorporated neural networks to learn mode decomposition patterns from data, but component cardinality is often predefined or determined through separate stopping or selection mechanisms. We propose a permutation-invariant neural framework for variable-cardinality AM-FM mode decomposition (PI-AMFM). PI-AMFM combines a multiscale temporal encoder, Mamba backbone, and component-presence estimation, with permutation-invariant Hungarian matching during training. On synthetic AM-FM signals, PI-AMFM achieved lower decomposition, instantaneous-frequency, reconstruction, and mode-count errors than the compared methods while preserving the overall trajectory pattern in a crossing-chirp example. On photoplethysmographic recordings, recovered modes captured cardiac and respiratory dynamics despite training only on synthetic signals. These results support the feasibility of PI-AMFM for variable-cardinality decomposition of nonstationary biomedical signals.
[LG-229] Inference for stochastic differential equations driven by weighted sub-fractional Brownian motion using neural networks and the Euler approximation
链接: https://arxiv.org/abs/2610.00793
作者: J. H. Ramirez-Gonzalez
类目: Machine Learning (stat.ML); Machine Learning (cs.LG)
*备注: 22 pages
Abstract:We consider the estimation of drift, diffusion, and noise covariance from discrete observations of stochastic differential equations driven by Gaussian processes. For a fixed observation horizon T0 and a known initial state x_0\in\mathbb R , we study \beginequation* dX_t=a(X_t),dt+\sigma(X_t),dZ_t^\beta,f, \qquad X_0=x_0,\quad 0\leq t\leq T. \endequation* \smallskip\noindent Here a:\mathbb R\to\mathbb R is the drift coefficient, \sigma:\mathbb R\to(0,\infty) is the diffusion coefficient, and Z^\beta,f is a centered Gaussian process from the weighted sub-fractional Brownian family, with covariance \beginequation* \operatornameCov(Z_s^\beta,f,Z_t^\beta,f) =\int_0^s\wedge t f®q_\beta(s-r,t-r),dr, \qquad 0\leq s,t\leq T. \endequation* \smallskip\noindent Here s\wedge t=\min\s,t\ . The temporal weight f:[0,T]\to[0,\infty) is measurable, bounded, and positive almost everywhere, and \beta\in(0,2) is the covariance exponent. For u,v\geq0 , the kernel is q_\beta(u,v)=[u^\beta+v^\beta-(u+v)^\beta]/(1-\beta) when \beta\ne1 . Its continuous extension at \beta=1 is q_1(u,v)=(u+v)\log(u+v)-u\log u-v\log v , with 0\log0=0 . Using the Euler approximation, we reconstruct the Gaussian driving increments from observed transitions and use their joint density to obtain a trajectory likelihood. Neural and radial-basis representations model the drift, diffusion, and normalized temporal weight, while a likelihood profile estimates the covariance exponent and diffusion scale. We compare the method with two neural alternatives on the same simulated trajectories in twenty coefficient settings. Comments: 22 pages Subjects: Machine Learning (stat.ML); Machine Learning (cs.LG) MSC classes: 62M45, 62M09, 60H10, 60G22 Cite as: arXiv:2610.00793 [stat.ML] (or arXiv:2610.00793v1 [stat.ML] for this version) https://doi.org/10.48550/arXiv.2610.00793 Focus to learn more arXiv-issued DOI via DataCite (pending registration) Submission history From: Jose Hermenegildo Ramirez Gonzalez [view email] [v1] Wed, 30 Sep 2026 22:33:32 UTC (634 KB) Full-text links: Access Paper: View a PDF of the paper titled Inference for stochastic differential equations driven by weighted sub-fractional Brownian motion using neural networks and the Euler approximation, by J. H. Ramirez-GonzalezView PDFHTML (experimental)TeX Source view license Current browse context: stat.ML prev | next new | recent | 2026-10 Change to browse by: cs cs.LG stat References Citations NASA ADSGoogle Scholar Semantic Scholar export BibTeX citation Loading… BibTeX formatted citation loading… Data provided by: Bookmark checked="checked"class=“labs-tab-input”> Bibliographic Tools Bibliographic and Citation Tools Bibliographic Explorer Toggle Bibliographic Explorer (What is the Explorer?) Connected Papers Toggle Connected Papers (What is Connected Papers?) Litmaps Toggle Litmaps (What is Litmaps?) scite.ai Toggle scite Smart Citations (What are Smart Citations?) Code, Data, Media Code, Data and Media Associated with this Article alphaXiv Toggle alphaXiv (What is alphaXiv?) Links to Code Toggle CatalyzeX Code Finder for Papers (What is CatalyzeX?) DagsHub Toggle DagsHub (What is DagsHub?) GotitPub Toggle Gotit.pub (What is GotitPub?) Huggingface Toggle Hugging Face (What is Huggingface?) ScienceCast Toggle ScienceCast (What is ScienceCast?) Demos Demos Replicate Toggle Replicate (What is Replicate?) Spaces Toggle Hugging Face Spaces (What is Spaces?) Spaces Toggle TXYZ.AI (What is TXYZ.AI?) Related Papers Recommenders and Search Tools Link to Influence Flower Influence Flower (What are Influence Flowers?) Core recommender toggle CORE Recommender (What is CORE?) Author Venue Institution Topic About arXivLabs arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv’s community? Learn more about arXivLabs. Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?) mathjaxToggle(); We gratefully acknowledge support from our major funders, member institutions, , and all contributors. About Help Contact Subscribe Copyright Privacy Accessibility Operational Status (opens in new tab) Major funding support from
[LG-230] Learning to Price Electricity for Optimal Demand Response
链接: https://arxiv.org/abs/2610.00755
作者: Jing Shang,Mohammad Mehrabi,Xinyang Zhou,Mahmoud Saleh,Andrey Bernstein,Stefan Wager
类目: Machine Learning (stat.ML); Machine Learning (cs.LG); Applications (stat.AP)
*备注:
Abstract:There is considerable interest in using time-varying electricity prices to shape consumer demand response, and better align energy demand with renewable production. However, optimal prices generally vary over time in response to complex signals such as weather forecasts, sunrise/sunset times, and day-of-week patterns; and existing methods are not able to make efficient use of such rich contextual information. Here, we propose a neural-network-based algorithm for contextual energy pricing, modeling pricing as a Stackelberg game and leveraging a mean-field solution representation from Mehrabi et al.~(2024). The approach learns constrained mappings from contextual features to feasible price signals. We validate our approach by simulating the energy grid in several US cities, and show that incorporating contextual information can considerably increase the value of the demand response programs.
[LG-231] StabilityArc: Decoding Protein Sequence Embeddings into Generalizable Stability Landscapes NEURIPS2026
链接: https://arxiv.org/abs/2610.00742
作者: Aaron L. Feller,Andrew D. Ellington,Claus O. Wilke
类目: Biomolecules (q-bio.BM); Machine Learning (cs.LG)
*备注: Accepted to Representations for the Physical Sciences Workshop @ NeurIPS 2026; 9 pages, 1 figure, 2 tables
Abstract:Every protein has a unique stability landscape, but the physical consequences of mutation are governed by recurring biochemical constraints. We test whether a shared decoder, trained on measurements from diverse proteins, can interpret these constraints in an unseen target, enabling cross-protein transfer for initial experimental round prescreening. We present StabilityArc , which maps frozen ESMC-600M residue representations through a shared RoPE transformer to an Lx20 matrix of substitution effects; a symmetric, contact-aware residual aids in predicting epistasis in simultaneous substitutions. In 66 strict leave-one-protein-out evaluations covering 134,794 ProteinGym variants, StabilityArc achieves 0.7134 Spearman correlation, exceeding the strongest zero-shot baseline, ProSST-2048 (0.6526), by 0.0608. We further explore the utility of this method by providing the score as a prior for Kermut, achieving Spearman correlation of 0.8280 across three supervised split schemes, improving on Kermut’s reported 0.8167.
[LG-232] Q-MINO: A Minimal-Norm Method for Quantization-Aware Training
链接: https://arxiv.org/abs/2610.00738
作者: Don Li
类目: Optimization and Control (math.OC); Hardware Architecture (cs.AR); Machine Learning (cs.LG)
*备注:
Abstract:The Straight-Through Estimator (STE) is a widely used heuristic for Quantization-Aware Training (QAT), but its surrogate gradients can exhibit substantial mismatch with the underlying quantized objective, leading to noisy updates and parameter oscillations, particularly in ultra-low-bit regimes. We propose the Quantization-Aware Minimal-Norm Optimizer (Q-MINO), a temporal bundle method that combines gradient consensus, state-drift regularization, and an alignment constraint to construct stabilized, minimum-norm update directions from recent optimization states. Q-MINO solves the resulting constrained subproblem using a warm-started Frank–Wolfe procedure with a feasible fallback initialization. Theoretically, via a stochastic Lyapunov Kurdyka–Łojasiewicz (KL) framework, we show that Q-MINO achieves asymptotic neighborhood convergence. Moreover, we detail numerical experiments with Q-MINO at various quantizations.
[LG-233] Nonparametric Distribution Matching for Self-Supervised Whole-Slide Image Condensation ICML2026 NEURIPS2026
链接: https://arxiv.org/abs/2610.00678
作者: Duong M. Nguyen,Trong Nghia Hoang,Hang Thi Nguyen,Thanh Trung Huynh,Phi Le Nguyen,Minh N. Do
类目: Image and Video Processing (eess.IV); Machine Learning (cs.LG)
*备注: Accepted at NeurIPS 2026, SPIGM@ICML 2026
Abstract:Histological whole-slide images (WSIs) are central to computational pathology but pose severe computational challenges due to their extremely high resolution, often spanning several gigabytes per slide. To enable scalable learning, existing methods apply self-supervised data condensation to reduce computational cost, but typically rely on heuristic prototype learning and do not explicitly preserve learning-relevant feature distributions for downstream tasks. In response, we introduce a principled reformulation of WSI condensation as a distribution-matching problem under a fixed representational lens, and develop NICER, a tractable approximation framework based on a nonparametric prior with slide-adaptive capacity. Experiments on five histopathology datasets, together with clinical evaluation from a board-certified pathologist, show that NICER consistently outperforms prior methods, achieving an average accuracy improvement of 7.44% while offering improved efficiency-accuracy trade-offs, highlighting the benefits of principled, distribution-aware condensation for scalable histological representation learning. Source codes are available in this https URL.
[LG-234] End-to-End Historical Music Restoration in Latent Space ICASSP2027
链接: https://arxiv.org/abs/2610.00607
作者: Steven Cho,Junghyun Koo,Raphael Lafargue,Tushar Dhyani,Eloi Moliner,Yuki Mitsufuji
类目: Audio and Speech Processing (eess.AS); Machine Learning (cs.LG); Sound (cs.SD); Signal Processing (eess.SP)
*备注: 5 pages, 2 figures, 3 tables; submitted to ICASSP 2027. Code and audio demos available at the project repository
Abstract:Historical music restoration (HMR) has almost exclusively focused on constrained problems such as Super-Resolution or the restoration of solo pieces, under-exploring the general task of restoring orchestral historical music, which has multiple instruments. This under-exploration is largely because the HMR domain, early-20th-century recordings, has no pre-degradation ground-truth pairs, making the restoration task unsupervised and more challenging. This paper presents a supervised end-to-end orchestral HMR benchmark by exploring both the synthetic degradation functions and the end-to-end generative deep-learning restoration methods. We simulate the historical recording degradation chain more faithfully than prior work, which makes orchestral restoration into a tractable supervised problem. A latent flow-matching model trained on the resulting synthetic pairs outperforms existing HMR baselines on intrusive, non-intrusive, and subjective evaluations. We also curate and release a 9.3-hour license-free, unpaired, historical classical-music test set, along with code and audio demos.
[LG-235] Mean Spatial Frequency Decoupling for Learning-Based Uplink-to-Downlink Covariance Conversion in FDD Massive MIMO
链接: https://arxiv.org/abs/2610.00596
作者: Melih Can Zerin
类目: ignal Processing (eess.SP); Machine Learning (cs.LG)
*备注:
Abstract:In frequency division duplexing (FDD) massive multiple-input multiple-output (MIMO) systems, the uplink (UL)-to-downlink (DL) channel covariance matrix (CCM) conversion problem is studied to relieve the heavy burden of DL training and feedback required for channel estimation. Learning- based methods perform well up to a certain array size, but for a fixed dataset size their accuracy deteriorates with the number of antennas, to the point where simple model-based methods outperform them. This paper identifies a key cause of this behavior and addresses it. The mean angle of arrival (AoA) induces a phase ramp along the lags of the CCM. Since the oscillation rate of this ramp grows with the number of antennas, a dataset of fixed size becomes increasingly sparse relative to the variation that must be captured. We propose estimating the slope of this ramp from the UL CCM separately and mapping it to the DL band in closed form, leaving the learner with a residual that is largely insensitive to the mean AoA, which substantially reduces the performance degradation with an increasing number of antennas. The proposed scheme, termed deramping, is a combination of pre- and post-processing steps that applies to learning-based conversion methods without altering their internal structure, as demonstrated on three structurally different learners. Simulation results show that deramping reduces the covariance estimation error of all three learners under uniform, Laplacian, and Gaussian angular power spectra,keeps the interpolation-based learners ahead of a model-based benchmark at large array sizes, and improves downlink channel estimation.
[LG-236] Scaling Collider Event Generation with Residual-Quantized Tokens
链接: https://arxiv.org/abs/2610.00569
作者: Dan Godi,Dmitrii Kobylianskii,Eilam Gross
类目: High Energy Physics - Experiment (hep-ex); Machine Learning (cs.LG); High Energy Physics - Phenomenology (hep-ph); Data Analysis, Statistics and Probability (physics.data-an)
*备注: 10 pages + 11 pages of appendices, 12 figures, 10 tables
Abstract:Full detector simulation and reconstruction of collider events are projected to become major bottlenecks at the High-Luminosity Large Hadron Collider, motivating the development of fast, ML-based surrogates. At the same time, LLMs have driven fast progress in generative discrete modeling: autoregressive transformers trained on tokenized data now represent the state of the art across a range of generative tasks. We extend the discrete modeling paradigm by introducing a particle-level generative model trained on residual-quantized full-event data. We demonstrate the ability of this model family to perform conditional generation from detector-stable particles; we study its scaling behavior across a range of dataset and model sizes, characterize the effects of repeated data exposure and demonstrate that token-level loss systematically predicts downstream physical fidelity. These results provide an empirical framework for scalable collider full-event generation based on residual-quantized representations.
[LG-237] Generative Modeling of Stochastic Dynamics for Long-Time Evolution
链接: https://arxiv.org/abs/2610.00546
作者: Yang-yang Tan,Jinyang Li,Lingxiao Wang
类目: atistical Mechanics (cond-mat.stat-mech); Machine Learning (cs.LG); High Energy Physics - Lattice (hep-lat)
*备注: 21 pages, 15 figures, comments are welcome!
Abstract:Exact stochastic equations for non-equilibrium dynamics are rarely accessible. We show that the long-time evolution of stochastic dynamics can be predicted from configuration pairs at a fixed short time lag, without knowledge of the equation of motion. Generative diffusion models learn the finite-time transition kernel from these pairs, and iterating it propagates the dynamics far beyond the training lag. For two-dimensional Model B, the diffusive dynamics of a conserved order parameter, the learned kernels reproduce dynamic critical scaling and self-similar t^1/3 coarsening. Agreement with direct simulations persists on lattices twice the largest training size and for initial ensembles absent from training. For driven colloids in a periodic optical potential, ten minutes of measured trajectories suffice to predict the particle current and mean passage time over the next twenty minutes within experimental uncertainty. Short-time observations thus contain the information needed to predict emergent non-equilibrium dynamics at much longer times.
[LG-238] Adaptive Conformal Prediction for Image Regression Models with Application to an Inertial Confinement Fusion Emulator
链接: https://arxiv.org/abs/2610.00535
作者: Carrie J. Lei-Cramer,Michael S. Jones,Laura J. Wendelberger
类目: Machine Learning (stat.ML); Machine Learning (cs.LG)
*备注:
Abstract:Uncertainty quantification is critical in scientific machine learning, where black-box, image-based models are increasingly deployed in high-stakes settings. In many such applications, model outputs inform costly decisions, yet most methods provide only point estimates without quantifying predictive uncertainty. This challenge is compounded by the limited accessibility and interpretability of model internals, making it difficult to assess reliability across different regions of the input space. As a result, there is a growing need for methods that can provide input-dependent uncertainty estimates to guide both model development and downstream experimentation. To address this need, we propose Adaptive Conformal Prediction using Nearest Neighbors (ACPNN), an input-adaptive conformal framework for image regression. ACPNN leverages information from neighboring samples to produce locally adaptive uncertainty estimates while maintaining low computational cost. The neighborhood structure is defined using a scaled distance metric learned via a Gaussian Process with an automatic relevance determination (ARD) kernel. We demonstrate the effectiveness of ACPNN on a diffusion model for emulating inertial confinement fusion (ICF) simulations, showing that it achieves reliable and adaptive uncertainty quantification.
[LG-239] Fractional Laplace Neural Operators: Exact Architectures an Expressivity Frontier at Criticality and Certified Stability for Memory-Driven Network Dynamics
链接: https://arxiv.org/abs/2610.00515
作者: Mauricio Herrera-Marín
类目: Machine Learning (stat.ML); Machine Learning (cs.LG)
*备注:
Abstract:Neural operators learn maps between function spaces, while hereditary network dynamics are described by Volterra resolvents with non-rational Laplace symbols. We introduce a fractional Laplace neural operator (fLNO) that embeds this structure in the learned map. For commuting excitation–Laplacian pairs, one block graph-spectral layer represents the full linear Volterra solution operator exactly. We establish an expressivity frontier for finite rational realizations: they approximate fractional memory geometrically on compact frequency windows, but cannot reproduce the non-integer critical asymptotics generated by a branch point, and on the half-line the best rational rate is root-exponential. The same theory yields trainable parametrizations that enforce a prescribed stability margin by construction, and a graphon-transfer theorem separates genuine operator consistency from parameter sharing. In a common-data benchmark, positive rational operators can match or exceed fLNO accuracy on finite horizons, whereas in controlled near-critical experiments fLNO recovers the branching coordinate more faithfully with far fewer parameters; unconstrained rational fits can cross the stability boundary, while certified parametrizations cannot. A four-parameter spectral law transfers without retraining from graphs of size 48 to 192 with 0.51–0.62% relative error. Applications to Chilean aftershock sequences and to renewal models for Chile and 21 Italian regions illustrate structured inference with explicit uncertainty. The contribution is an operator-learning architecture in which exact memory structure, physical coordinates and stability guarantees coexist with competitive accuracy.
[LG-240] Heteroskedastic Canonical Polyadic Tensor Decomposition
链接: https://arxiv.org/abs/2610.00498
作者: Kyle Ritscher,Carlos Llosa-Vite
类目: Methodology (stat.ME); Machine Learning (cs.LG); Machine Learning (stat.ML)
*备注: 35 pages, 16 figures
Abstract:When minimizing the squared-error loss, the popular CP decomposition can be interpreted as parameter inference in a Gaussian model with a low-rank mean tensor and constant variance across the tensor entries. We introduce heteroskedastic-CP (HCP), which models entrywise variability with a non-constant, low-rank precision tensor, and develop an alternating block-coordinate ascent method to recover both the low-rank mean and precision tensors from noisy observations. Our procedure is computationally competitive, with the same leading-order factor-update complexity as CP-ALS. We demonstrate HCP on synthetic experiments and an EEG application.
[LG-241] ChainLoRA: Geometry-Preserving Task Vector Merging for Continual Learning in LLM s
链接: https://arxiv.org/abs/2610.00431
作者: Hang Yin,Haozhe Wang,Yuhua Luo,Zhangqi Pan,Xiaoxing Wang,Junchi Yan
类目: Machine Learning (stat.ML); Machine Learning (cs.LG)
*备注: 18 pages
Abstract:Continual parameter-efficient fine-tuning for large language models (LLMs) must balance retention of previously acquired knowledge, adaptation to new tasks, and strict parameter budgets. We present \textbfChainLoRA, a replay-free continual merging framework built on chain-updated task-vector geometry. From a parameter-merging perspective, we formulate a geometric view of forgetting through a measurable interaction between task updates, separating directional overlap from coefficient coupling. Building on this view, ChainLoRA combines chain-updated training with post-stream adaptive SVD merging. During training, initialization and a one-sided orthogonality proxy use only the last carrier, keeping their historical-state footprint and regularization overhead constant as the task stream grows. At merging time, Adaptive SVD extracts a shared carrier and aligns it to the latest task through Procrustes adaptation. Our theoretical analysis shows that Procrustes adaptation facilitates geometric approximate separation of shared and task-specific components. The one-sided proxy further bounds inter-task interference. An effective-rank penalty additionally promotes efficient utilization of the task subspace during continual learning. Experiments show that ChainLoRA achieves state-of-the-art performance among the evaluated replay-free methods on the Large and SuperNI benchmarks, while remaining competitive on Standard CL and attaining almost the closest average scores to the evaluated replay-based method across all three benchmarks.
[LG-242] ransferable Graph Metanetworks
链接: https://arxiv.org/abs/2610.00420
作者: Yuxin Ma,Adir Dayan,Yam Eitan,Haggai Maron,Soledad Villar
类目: Machine Learning (stat.ML); Machine Learning (cs.LG)
*备注:
Abstract:A weight space network (or metanetwork) takes the weights of another neural network as input and predicts properties of it. Most prior work trains such models on input networks of one or a few fixed sizes and evaluates them in-distribution. The few attempts at out-of-distribution size generalization remain limited in scope and have achieved only modest success. Consequently, the potential efficiency gains of training on small networks and evaluating on much larger ones remain largely unrealized. We propose Transferable Graph Metanetworks, which extend the graph metanetwork paradigm with a set of modifications that make performance transferable across input networks of different widths. The modifications follow two principles: invariance to the ways in which networks of different widths represent the same function, and continuity, such that weights representing similar functions receive similar predictions. We further study whether size generalization is possible for input networks trained independently from random initialization. Empirically, our modifications significantly improve size generalization on every task we consider. Performance is strongest on input networks trained under the maximal-update parameterization ( \mu P), where it remains robust up to 42\times the training width. Theoretically, we explain these observations with infinite-width limit theory: we prove size-generalization guarantees for our model on \mu P-trained inputs, and explain why it can fail under other parameterizations.
[LG-243] Improving scoring functions for protein-protein docking with LambdaLoss
链接: https://arxiv.org/abs/2610.00191
作者: Richard Zhu,Darren Xu,Lee-Shin Chu,Jeffrey J. Gray
类目: Quantitative Methods (q-bio.QM); Machine Learning (cs.LG); Biomolecules (q-bio.BM)
*备注:
Abstract:Modeling protein-protein interactions requires accurate scoring functions that can rank potential poses (conformations) of a protein-protein complex to differentiate near-native poses from incorrect ones. Here, we propose a general framework for improving protein-protein pose ranking and other biomolecular interaction models using the LambdaLoss loss function from the Learning-to-Rank field. We test this framework by fine-tuning the energy prediction head of DFMDock with the LambdaLoss on an augmented dataset of 2.9M decoy poses derived from the DIPS dataset. On targets from the CAPRI score set benchmark, our fine-tuned ranking model LambdaDockScore is better at identifying correct poses in its top-1 and top-5 predictions compared to EuDockScore, a state-of-the-art method. LambdaDockScore also improves upon baseline DFMDock ranking performance for scoring antibody-antigen complexes and protein-protein complexes with very large or small binding interfaces.
[LG-244] Symmetry Discovery in Quantum Learning: Observable-Level and Task-Level Inference from Finite Measurements
链接: https://arxiv.org/abs/2610.00157
作者: Zeyu Chen
类目: Quantum Physics (quant-ph); Machine Learning (cs.LG)
*备注:
Abstract:Symmetry reduces the capacity of a quantum learning model, but the imposed group must match both the measured information and the label transformation. We establish a finite-measurement theory for inferring this group from candidate transformations. The central structural result identifies observable-invisible transformations with the stabilizer of a projected state whenever the probe span is invariant. It turns recovered generators into a valid subgroup and identifies the continuous invisible space with its Lie algebra. For finite dictionaries, an unbiased shadow statistic distinguishes zero from positive squared expectation discrepancies with an inverse-gap measurement rate, improving the inverse-square-gap rate of uniform discrepancy estimation. A commuting qubit lower bound proves the gap dependence optimal at fixed snapshot scale, and simultaneous intervals support data-dependent tolerances. Task validation then tests either the joint distribution through a characteristic kernel or its encoded mean through a classical–quantum discrepancy. An exact group-average identity relates the latter to joint-state asymmetry and specifies its conversion to binary task breaking mass. Projection bias quantifies the cost of excessive symmetry, while an \ell_1 readout bound quantifies the capacity gained by relaxing it. At an invariant pure-state backbone, retained and nontrivial breaking sectors are Fisher-orthogonal. Ising-chain calculations connect finite-shot recovery, label-dependent symmetry, and physical sector drift. These results determine which symmetry the measurements support and provide the statistical and geometric basis for a subsequent release decision.
[LG-245] Loading history and window geometry bound compact-state slip ranking during granular shear startup
链接: https://arxiv.org/abs/2610.00124
作者: Ruixin Zhou,Boliang Yu
类目: oft Condensed Matter (cond-mat.soft); Statistical Mechanics (cond-mat.stat-mech); Machine Learning (cs.LG); Computational Physics (physics.comp-ph)
*备注: 28 pages, 6 figures
Abstract:Granular slip forecasting can conflate material state, loading progress, and the geometry of event-centered sampling. We separated these contributions in slowly sheared two-dimensional frictional disks using a compact neural score of stress, pressure, coordination, non-affine motion, and force-network observables. The model was developed on 36 trajectories and frozen efore testing on 18 new trajectories under two nested stress-drop definitions. Inspection of held-out results revealed post-event sampling asymmetry; recovery-aware analyses are therefore descriptive. With trajectories weighted equally, the compact score ranked near-slip windows above both prevalence and within-trajectory circular-phase controls under both definitions (representative average precision 0.310 versus prevalence 0.173; phase-null upper bound 0.257). Loading-history coordinates ranked more strongly, reaching 0.534 for causal elapsed strain. The recovery-aware rule retained 78.4% of activity-gated events and preferentially selected longer preceding intervals; ranking by time since the previous catalogued event remained compatible with a count-conditioned geometry null. Compact observables thus contain temporally aligned slip information, but stronger loading-history baselines and window-geometry sensitivity bound that evidence. These startup data do not isolate a state-specific short-horizon precursor beyond loading history or support a renewal interpretation of elapsed-strain ranking.
[LG-246] Weighted Data Selection: Sharp Upper-Half and Five-Dimensional Laws
链接: https://arxiv.org/abs/2610.00101
作者: Zhongxuan Liu,Hongzhi Wang
类目: Machine Learning (stat.ML); Machine Learning (cs.LG)
*备注: 35 pages, 2 figures; supplementary verification code included
Abstract:How much risk does a small reweighted training support retain? For finite weighted least squares with the minimum-norm learner, we prove the exact law \Gamma_d(n)=3-n/d throughout \lceil3d/2\rceil\leq n\leq2d-1 . The guarantee covers every observed feature rank and uses selections that preserve the full feature span. Balanced simplex anchors reduce dimension; positive-weight lifting and independent-line compression close the risk bound. Shifted coordinate pairs attain the matching lower bound. The complete dataset-level upper bound and sharpness construction are verified in Lean 4. At the smaller budget (d,n)=(5,6) , we also prove \Gamma_5(6)=11/5 , matching the simplex-block prediction from 5=3+2 over arbitrary interacting configurations. Circuit covers, comparison second moments, and circuit-plane probabilities give the sharp excess 6/5 , while polar-face geometry resolves shared rank-three circuits. The general simplex-block frontier connects these laws within the intermediate-budget selection problem.
附件下载


